EDBT 2026 Demo / reviewers in the wild / expert
Amit P. Sheth
dblp:s/AmitPSheth
· DBLP profile ↗
209ranked-venue papers
29as first author
30since 2021 · last 2026
0000-0002-0021-5293ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 126 · 23 first-author · 5 since 2021Artificial intelligence and machine learning · 61 · 2 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 43 · 6 since 2021Human-computer interaction and ubiquitous computing · 20 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 11 since 2021Software engineering, systems software and programming languages · 15 · 1 first-authorSecurity and privacy · 6Systems, architecture and hardware · 3 · 2 first-authorComputer networks · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PAL: Personal Adaptive LearnerabstractAI-driven education platforms have made some progress in personalisation, yet most remain constrained to static adaptation—predefined quizzes, uniform pacing, or generic feedback—limiting their ability to respond to learners’ evolving understanding. This shortfall highlights the need for systems that are both context-aware and adaptive in real time. We introduce PAL (Personal Adaptive Learner), an AI-powered platform that transforms lecture videos into interactive learning experiences. PAL continuously analyzes multimodal lecture content and dynamically engages learners through questions of varying difficulty, adjusting to their responses as the lesson unfolds. At the end of a session, PAL generates a personalized summary that reinforces key concepts while tailoring examples to the learner’s interests. By uniting multimodal content analysis with adaptive decision-making, PAL contributes a novel framework for responsive digital learning. Our work demonstrates how AI can move beyond static personalization toward real-time, individualized support, addressing a core challenge in AI-enabled education. Megha Chakraborty, Darssan Eswaramoorthi, Madhur Thareja, Het Riteshkumar Shah, Finlay Palmer, Aryaman Bahl, Michelle A. Ihetu, Amit P. Sheth |
AAAI | 8 |
| 2026 | In-Situ Eval: A Modular Framework for Custom and Real-Time RAG BenchmarkingabstractRetrieval-Augmented Generation (RAG) has become the standard approach for integrating domain knowledge into Large Language Models (LLMs). However, fair comparison of RAG pipelines remains difficult: data preparation is often ad hoc, subsampling methods are opaque, parameters vary across implementations, and evaluation is fragmented. We present In-Situ Eval, a unified and reproducible framework that operationalizes the full RAG pipeline with configurable subsampling strategies and both RAG-specific and generic evaluation metrics. The platform supports two execution modes: an offline Dataset mode for evaluating precomputed outputs, and a live Retrieval mode for benchmarking RAG variants with state-of-the-art LLMs. Users can flexibly select datasets, retrieval techniques, models, and metrics, enabling side-by-side comparisons, ablations, and targeted analyses. This holistic approach reduces computational costs, clarifies the impact of subsampling techniques, and provides actionable insights for real-world deployments. By facilitating transparent, customizable, and interactive benchmarking, In-Situ Eval empowers both researchers and practitioners to make informed decisions in adapting RAG pipelines to domain-specific needs. Ritvik Garimella, Kaushik Roy 0009, Chathurangi Shyalika, Amit P. Sheth |
AAAI | 4 |
| 2026 | DETONATE - A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference OptimizationabstractAlignment is crucial for text-to-image (T2I) models to ensure that the generated images faithfully capture user intent while maintaining safety and fairness. Direct Preference Optimization (DPO) has emerged as a key alignment technique for large language models (LLMs), and its influence is now extending to T2I systems. This paper introduces DPO-Kernels for T2I models, a novel extension of DPO that enhances alignment across three key dimensions: (i) Hybrid Loss, which integrates embedding-based objectives with the traditional probability-based loss to improve optimization; (ii) Kernelized Representations, leveraging Radial Basis Function (RBF), Polynomial, and Wavelet kernels to enable richer feature transformations, ensuring better separation between safe and unsafe inputs; and (iii) Divergence Selection, expanding beyond DPO’s default Kullback–Leibler (KL) regularizer by incorporating alternative divergence measures such as Wasserstein and Rényi divergences to enhance stability and robustness in alignment training. We introduce DETONATE, the first large-scale benchmark of its kind, comprising approximately 100K curated image pairs, categorized as chosen and rejected. This benchmark encapsulates three critical axes of social bias and discrimination: Race, Gender, and Disability. The prompts are sourced from the hate speech datasets, while the images are generated using state-of-the-art T2I models, including Stable Diffusion 3.5 Large (SD-3.5), Stable Diffusion XL (SD-XL), and Midjourney. Furthermore, to evaluate alignment beyond surface metrics, we introduce the Alignment Quality Index (AQI) for T2I systems: a novel geometric measure that quantifies latent space separability of safe/unsafe image activations, revealing hidden model vulnerabilities. While alignment techniques often risk overfitting, we empirically demonstrate that DPO-Kernels preserve strong generalization bounds using the theory of Heavy-Tailed Self-Regularization (HT-SR). Renjith Prasad Kaippilly Mana, Abhilekh Borah, Hasnat Md Abdullah, Chathurangi Shyalika, Ritvik Garimella, Rajarshi Roy 0007, Harshul Raj Surana, Nasrin Imanpour, Suranjana Trivedy, Amit P. Sheth, Amitava Das 0001 |
AAAI | 11 |
| 2026 | Chatsparent: An Interactive System for Detecting and Mitigating Cognitive Fatigue in LLMsabstractLLMs are increasingly being deployed as chatbots, but today’s interfaces offer little to no friction: users interact through seamless conversations that conceal when the model is drifting, hallucinating or failing. This lack of transparency fosters blind trust, even as models produce unstable or repetitive outputs. We introduce an interactive demo that surfaces and mitigates cognitive fatigue, a failure mode where LLMs gradually lose coherence during auto-regressive generation. Our system, Chatsparent, instruments real-time, token-level signals of fatigue, including attention-to-prompt decay, embedding drift, and entropy collapse, and visualizes them as a unified fatigue index. When fatigue thresholds are crossed, the interface allows users to activate lightweight interventions such as attention resets, entropy-regularized decoding, and self-reflection checkpoints. The demo streams live text and fatigue signals, allowing users to observe when fatigue arises, how it affects output quality, and how interventions restore stability. By turning passive chatbot interaction into an interactive diagnostic experience, our system empowers users to better understand LLM behavior while improving reliability at inference time. Riju Marwah, Vishal Pallagani, Ritvik Garimella, Amit P. Sheth |
AAAI | 4 |
| 2026 | CausalPulse: Agentic Copilot for Root Cause Analysis in Smart ManufacturingabstractModern manufacturing systems demand real-time, trustworthy, and interpretable insights into anomalies and their underlying causes. However, conventional pipelines treat anomaly detection, causal inference, and decision-making as siloed tasks, lacking integration, explainability, and adaptability. We present CausalPulse, an intelligent, multi-agent copilot for automated Root Cause Analysis (RCA) in industrial settings. Built on a modular and extensible architecture, the system leverages standard agentic protocols, including Model Context Protocol (MCP), Agent2Agent (A2A), and LangGraph for dynamic tool and agent discovery and seamless orchestration of tasks. Agents dynamically interact to perform data preprocessing, anomaly detection, causal discovery, and root cause analysis through a neurosymbolic workflow that combines symbolic reasoning with neural methods. Intelligent postprocessing pipelines enable automatic chaining of agent tasks, enhancing contextual awareness and adaptability. CausalPulse is evaluated using both an academic public dataset (i.e., Future Factories) and an industrial proprietary dataset (i.e., Planar Oxygen Sensor Element) and shows that the system outperforms traditional baselines in interpretability, trustworthiness, and operational utility. Chathurangi Shyalika, Utkarshani Jaimini, Cory A. Henson, Amit P. Sheth |
AAAI | 4 |
| 2026 | CausalTrace: A Neurosymbolic Causal Analysis Agent for Smart ManufacturingabstractModern manufacturing environments demand not only accurate predictions but also interpretable insights to process anomalies, root causes, and potential interventions. Existing AI systems often function as isolated black boxes, lacking the seamless integration of prediction, explanation, and causal reasoning required for a unified decision-support solution. This fragmentation limits their trustworthiness and practical utility in high-stakes industrial environments. In this work, we present CausalTrace, a neurosymbolic causal analysis module integrated into the SmartPilot industrial CoPilot. CausalTrace performs data-driven causal analysis enriched by industrial ontologies and knowledge graphs, including advanced functions such as causal discovery, counterfactual reasoning, and root cause analysis (RCA). It supports real-time operator interaction and is designed to complement existing agents by offering transparent, explainable decision support. We conducted a comprehensive evaluation of CausalTrace using multiple causal assessment methods and the C3AN framework (i.e. Custom, Compact, Composite AI with Neurosymbolic Integration), which spans principles of robustness, intelligence, and trustworthiness. In an academic rocket assembly testbed, CausalTrace achieved substantial agreement with domain experts (ROUGE-1: 0.91 in ontology QA) and strong RCA performance (MAP@3: 94%, PR@2: 97%, MRR: 0.92, Jaccard: 0.92). It also attained 4.59/5 in the C3AN evaluation, demonstrating precision and reliability for live deployment. Chathurangi Shyalika, Aryaman Sharma, Fadi El Kalach, Utkarshani Jaimini, Cory A. Henson, Ramy F. Harik, Amit P. Sheth |
AAAI | 7 |
| 2026 | Exploring the Potential of Large Language Models for Assisting with Mental Health Diagnostic AssessmentsabstractLarge language models (LLMs) are increasingly attracting the attention of healthcare professionals for their potential to assist in diagnostic assessments, which could alleviate the strain on the healthcare system caused by a high patient load and a shortage of providers. For LLMs to be effective in supporting diagnostic assessments, it is essential that they closely replicate the standard diagnostic procedures used by clinicians. In this paper, we specifically examine the diagnostic assessment processes described in the Patient Health Questionnaire-9 (PHQ-9) for major depressive disorder (MDD) and the Generalized Anxiety Disorder-7 (GAD-7) questionnaire for generalized anxiety disorder (GAD). We investigate various prompting and fine-tuning techniques to guide both proprietary and open source LLMs in adhering to these processes, and we evaluate the agreement between LLM-generated diagnostic outcomes and expert-validated ground truth. For fine-tuning, we utilize the MentaLLaMa and Llama models, while for prompting, we experiment with proprietary models like GPT-3.5 and GPT-4o, as well as open source models such as llama-3.1-8b and mixtral-8x7b. Software Availability . We make all software artifacts available at this GitHub link ( https://github.com/kauroy1994/Large-Language-Models-for-Assisting-with-Mental-Health-Diagnostic-Assessments ). Institutional Review Board (IRB) . This study does not require approval from the IRB. It involves using clinician-annotated social media posts, authorized for research purposes. The primary objective is to evaluate the effectiveness of LLMs that incorporate diagnostic criteria for major depressive disorder and general anxiety disorder for assisting with mental health assessments. Kaushik Roy 0009, Harshul Surana, Darssan Eswaramoorthi, Yuxin Zi, Vedant Palit, Ritvik Garimella, Amit P. Sheth |
ACM Trans. Comput. Heal. | 7 |
| 2025 | Pic2Prep: A Multimodal Conversational Agent for Cooking AssistanceabstractAs the demand for healthier, personalized culinary experiences grows, so does the need for advanced food computation models that offer more than basic nutritional insights. However, current food computation models lack the depth to provide actionable insights like ingredient substitution or alternative cooking actions to suit users’ dietary goals. To address this, we introduce and demonstrate Pic2Prep, a multimodal conversational system that generates detailed cooking instructions, actions and ingredient lists from both images and text provided by users. The system is developed using a novel dataset generated through Stable Diffusion, where the input consists of recipe titles and ingredient lists from the Recipe1M dataset to create synthesized food images with variations. This dataset is used to fine-tune the Bootstrapping Language-Image Pre-training (BLIP) model to extract cooking instructions and ingredients from food images. Pic2Prep also employs the CookGen model, a small-scale custom generative model to derive specific cooking actions from cooking instructions. A custom mapper, trained on the Mistral model, links these actions to the corresponding ingredients, creating a comprehensive understanding of the cooking process. The system features an interactive user interface that allows users to input images and ask targeted questions, receiving real-time responses. Renjith Prasad Kaippilly Mana, Chathurangi Shyalika, Revathy Venkataramanan, Darssan Eswaramoorthi, Amit P. Sheth |
AAAI | 5 |
| 2025 | KnowledgePrompts: Exploring the Abilities of Large Language Models to Solve Proportional Analogies via Knowledge-Enhanced PromptingabstractMaking analogies is fundamental to cognition. Proportional analogies, which consist of four terms, are often used to assess linguistic and cognitive abilities. For instance, completing analogies like “Oxygen is to Gas as < blank > is to < blank >" requires identifying the semantic relationship (e.g., “type of”) between the first pair of terms (“Oxygen” and “Gas”) and finding a second pair that shares the same relationship (e.g., “Aluminum” and “Metal”). In this work, we introduce a 15K Multiple-Choice Question Answering (MCQA) dataset for proportional analogy completion and evaluate the performance of contemporary Large Language Models (LLMs) in various knowledge-enhanced prompt settings. Specifically, we augment prompts with three types of knowledge: exemplar, structured, and targeted. Our results show that despite extensive training data, solving proportional analogies remains challenging for current LLMs, with the best model achieving an accuracy of 55%. Notably, we find that providing targeted knowledge can better assist models in completing proportional analogies compared to providing exemplars or collections of structured knowledge. Our code and data are available at: https://github.com/Thiliniiw/KnowledgePrompts/ Thilini Wijesiriwardene, Ruwan Wickramarachchi, Sreeram Vennam, Vinija Jain, Aman Chadha, Amitava Das 0001, Ponnurangam Kumaraguru, Amit P. Sheth |
COLING | 8 |
| 2025 | SmartPilot: Agent-Based CoPilot for Intelligent Manufacturing
Chathurangi Shyalika, Renjith Prasad, Alaa T. Al Ghazo, Darssan Eswaramoorthi, Sara Shree Muthuselvam, Amit P. Sheth |
AAMAS | 6 |
| 2025 | NSF-MAP: Neurosymbolic Multimodal Fusion for Robust and Interpretable Anomaly Prediction in Assembly PipelinesabstractIn modern assembly pipelines, identifying anomalies is crucial in ensuring product quality and operational efficiency. Conventional single-modality methods fail to capture the intricate relationships required for precise anomaly prediction in complex predictive environments with abundant data and multiple modalities. This paper proposes a neurosymbolic AI and fusion-based approach for multimodal anomaly prediction in assembly pipelines. We introduce a time series and image-based fusion model that leverages decision-level fusion techniques. Our research builds upon three primary novel approaches in multimodal learning: time series and image-based decision-level fusion modeling, transfer learning for fusion, and knowledge-infused learning. We evaluate the novel method using our derived and publicly available multimodal dataset and conduct comprehensive ablation studies to assess the impact of our preprocessing techniques and fusion model compared to traditional baselines. The results demonstrate that a neurosymbolic AI-based fusion approach that uses transfer learning can effectively harness the complementary strengths of time series and image data, offering a robust and interpretable approach for anomaly prediction in assembly pipelines with enhanced performance. \noindent The datasets, codes to reproduce the results, supplementary materials, and demo are available at https://github.com/ChathurangiShyalika/NSF-MAP. Chathurangi Shyalika, Renjith Prasad, Fadi El Kalach, Revathy Venkataramanan, Ramtin Zand, Ramy F. Harik, Amit P. Sheth |
IJCAI | 7 |
| 2025 | A Cross Attention Approach to Diagnostic Explainability Using Clinical Practice Guidelines for DepressionabstractThe lack of explainability in using relevant clinical knowledge hinders the adoption of artificial intelligence-powered analysis of unstructured clinical dialogue. A wealth of relevant, untapped Mental Health (MH) data is available in online communities, providing the opportunity to address the explainability problem with substantial potential impact as a screening tool for both online and offline applications. Inspired by how clinicians rely on their expertise when interacting with patients, we leverage relevant clinical knowledge to classify and explain depression-related data, reducing manual review time and engendering trust. We developed a method to enhance attention in contemporary transformer models and generate explanations for classifications that are understandable by mental health practitioners (MHPs) by incorporating external clinical knowledge. We propose a domain-general architecture called ProcesS knowledgeinfused cross ATtention (PSAT) that incorporates clinical practice guidelines (CPG) when computing attention. We transform a CPG resource focused on depression, such as the Patient Health Questionnaire (e.g. PHQ-9) and related questions, into a machine-readable ontology using SNOMED-CT. With this resource, PSAT enhances the ability of models like GPT-3.5 to generate application-relevant explanations. Evaluation of four expert-curated datasets related to depression demonstrates PSAT's applicationrelevant explanations. PSAT surpasses the performance of twelve baseline models and can provide explanations where other baselines fall short. Sumit Dalal, Deepa Tilwani, Manas Gaur, Sarika Jain 0001, Valerie L. Shalin, Amit P. Sheth |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | GEAR-Up: Generative AI and External Knowledge-Based Retrieval: Upgrading Scholarly Article Searches for Systematic ReviewsabstractThis paper addresses the time-intensive nature of systematic reviews (SRs) and proposes a solution leveraging advancements in Generative AI (e.g., ChatGPT) and external knowledge augmentation (e.g., Retrieval-Augmented Generation). The proposed system, GEAR-Up, automates query development and translation in SRs, enhancing efficiency by enriching user queries with context from language models and knowledge graphs. Collaborating with librarians, qualitative evaluations demonstrate improved reproducibility and search strategy quality. Access the demo at https://youtu.be/zMdP56GJ9mU. Kaushik Roy 0009, Vedant Khandelwal, Valerie Vera, Harshul Surana, Heather Heckman, Amit P. Sheth |
AAAI | 6 |
| 2024 | A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19abstractMonitoring public sentiment via social media is potentially helpful during health crises such as the COVID-19 pandemic. However, traditional frequency-based and data-driven neural network-based approaches can miss newly relevant content due to the evolving nature of language in a dynamic environment. Human-curated symbolic knowledge sources, such as lexicons for standard language and slang terms, can potentially elevate social media signals in evolving language. We introduce a neurosymbolic method that integrates neural networks with symbolic knowledge sources, improving the detection and interpretation of mental health-related tweets relevant to COVID-19. Our method was evaluated using a corpus of large datasets (~12 billion tweets, 2.5 million subreddit data, and 700k news articles) and multiple knowledge graphs. This method dynamically adapts to evolving language, outperforming purely data-driven models with an F1 score exceeding 92%. This approach also showed faster adaptation to new data and lower computational demands than fine-tuning pre-trained large language models (LLMs). This study demonstrates the benefit of neurosymbolic methods in interpreting text in a dynamic environment for tasks such as health surveillance. Vedant Khandelwal, Manas Gaur, Ugur Kursuncu, Valerie L. Shalin, Amit P. Sheth |
IEEE Big Data | 5 |
| 2024 | On the Prospects of Incorporating Large Language Models (LLMs) in Automated Planning and Scheduling (APS)abstractAutomated Planning and Scheduling is among the growing areas in Artificial Intelligence (AI) where mention of LLMs has gained popularity. Based on a comprehensive review of 126 papers, this paper investigates eight categories based on the unique applications of LLMs in addressing various aspects of planning problems: language translation, plan generation, model construction, multi-agent planning, interactive planning, heuristics optimization, tool integration, and brain-inspired planning. For each category, we articulate the issues considered and existing gaps. A critical insight resulting from our review is that the true potential of LLMs unfolds when they are integrated with traditional symbolic planners, pointing towards a promising neuro-symbolic approach. This approach effectively combines the generative aspects of LLMs with the precision of classical planning methods. By synthesizing insights from existing literature, we underline the potential of this integration to address complex planning challenges. Our goal is to encourage the ICAPS community to recognize the complementary strengths of LLMs and symbolic planners, advocating for a direction in automated planning that leverages these synergistic capabilities to develop more advanced and intelligent planning systems. We aim to keep the categorization of papers updated on https://ai4society.github.io/LLM-Planning-Viz/, a collaborative resource that allows researchers to contribute and add new literature to the categorization. Vishal Pallagani, Bharath Muppasani, Kaushik Roy 0009, Francesco Fabiano, Andrea Loreggia, Keerthiram Murugesan, Biplav Srivastava, Francesca Rossi 0001, Lior Horesh, Amit P. Sheth |
ICAPS | 10 |
| 2024 | AssemAI: Interpretable Image-Based Anomaly Detection for Manufacturing PipelinesabstractAnomaly detection in manufacturing pipelines remains a critical challenge, intensified by the complexity and variability of industrial environments. This paper introduces AssemAI, an interpretable image-based anomaly detection system tailored for smart manufacturing pipelines. Utilizing a curated image dataset from an industry-focused rocket assembly pipeline, we address the challenge of imbalanced image data and demonstrate the importance of image-based methods in anomaly detection. Our primary contributions include deriving an image dataset, fine-tuning an object detection model YOLO-FF, and implementing a custom anomaly detection model for assembly pipelines. The proposed approach leverages domain knowledge in data preparation, model development and reasoning. We implement several anomaly detection models on the derived image dataset, including a Convolutional Neural Network, Vision Transformer (ViT), and pretrained versions of these models. Additionally, we incorporate explainability techniques at both user and model levels, utilizing ontology for user-level explanations and SCORE-CAM for indepth feature and model analysis. Finally, the best-performing anomaly detection model and YOLO-FF are deployed in a real-time setting. Our results include ablation studies on the baselines and a comprehensive evaluation of the proposed system. This work highlights the broader impact of advanced image-based anomaly detection in enhancing the reliability and efficiency of smart manufacturing processes. The image dataset, codes to reproduce the results and additional experiments are available at https:/github.com/renjithk4/AssemAI. Renjith Prasad, Chathurangi Shyalika, Fadi El Kalach, Revathy Venkataramanan, Ramtin Zand, Ramy F. Harik, Amit P. Sheth |
ICMLA | 7 |
| 2023 | Demo Alleviate: Demonstrating Artificial Intelligence Enabled Virtual Assistance for Telehealth: The Mental Health CaseabstractAfter the pandemic, artificial intelligence (AI) powered support for mental health care has become increasingly important. The breadth and complexity of significant challenges required to provide adequate care involve: (a) Personalized patient understanding, (b) Safety-constrained and medically validated chatbot patient interactions, and (c) Support for continued feedback-based refinements in design using chatbot-patient interactions. We propose Alleviate, a chatbot designed to assist patients suffering from mental health challenges with personalized care and assist clinicians with understanding their patients better. Alleviate draws from an array of publicly available clinically valid mental-health texts and databases, allowing Alleviate to make medically sound and informed decisions. In addition, Alleviate's modular design and explainable decision-making lends itself to robust and continued feedback-based refinements to its design. In this paper, we explain the different modules of Alleviate and submit a short video demonstrating Alleviate's capabilities to help patients and clinicians understand each other better to facilitate optimal care strategies. Kaushik Roy 0009, Vedant Khandelwal, Raxit Goswami, Nathan Dolbir, Jinendra Malekar, Amit P. Sheth |
AAAI | 6 |
| 2023 | CLUE-AD: A Context-Based Method for Labeling Unobserved Entities in Autonomous Driving DataabstractGenerating high-quality annotations for object detection and recognition is a challenging and important task, especially in relation to safety-critical applications such as autonomous driving (AD). Due to the difficulty of perception in challenging situations such as occlusion, degraded weather, and sensor failure, objects can go unobserved and unlabeled. In this paper, we present CLUE-AD, a general-purpose method for detecting and labeling unobserved entities by leveraging the object continuity assumption within the context of a scene. This method is dataset-agnostic, supporting any existing and future AD datasets. Using a real-world dataset representing complex urban driving scenes, we demonstrate the applicability of CLUE-AD for detecting unobserved entities and augmenting the scene data with new labels. Ruwan Wickramarachchi, Cory A. Henson, Amit P. Sheth |
AAAI | 3 |
| 2023 | FACTIFY-5WQA: 5W Aspect-based Fact Verification through Question AnsweringabstractAnku Rani, S.M Towhidul Islam Tonmoy, Dwip Dalal, Shreya Gautam, Megha Chakraborty, Aman Chadha, Amit Sheth, Amitava Das. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Anku Rani, S. M. Towhidul Islam Tonmoy, Dwip Dalal, Shreya Gautam, Megha Chakraborty, Aman Chadha, Amit P. Sheth, Amitava Das 0001 |
ACL (1) | 7 |
| 2023 | FACTIFY3M: A benchmark for multimodal fact verification with explainability through 5W Question-AnsweringabstractMegha Chakraborty, Khushbu Pahwa, Anku Rani, Shreyas Chatterjee, Dwip Dalal, Harshit Dave, Ritvik G, Preethi Gurumurthy, Adarsh Mahor, Samahriti Mukherjee, Aditya Pakala, Ishan Paul, Janvita Reddy, Arghya Sarkar, Kinjal Sensharma, Aman Chadha, Amit Sheth, Amitava Das. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Megha Chakraborty, Khushbu Pahwa, Anku Rani, Shreyas Chatterjee, Dwip Dalal, Harshit Dave, Ritvik Garimella, Preethi Gurumurthy, Adarsh Mahor, Samahriti Mukherjee, Aditya Pakala, Ishan Paul, Janvita Reddy, Arghya Sarkar, Kinjal Sensharma, Aman Chadha, Amit P. Sheth, Amitava Das 0001 |
EMNLP | 17 |
| 2023 | Counter Turing Test (CT2): AI-Generated Text Detection is Not as Easy as You May Think - Introducing AI Detectability Index (ADI)abstractMegha Chakraborty, S.M Towhidul Islam Tonmoy, S M Mehedi Zaman, Shreya Gautam, Tanay Kumar, Krish Sharma, Niyar Barman, Chandan Gupta, Vinija Jain, Aman Chadha, Amit Sheth, Amitava Das. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Megha Chakraborty, S. M. Towhidul Islam Tonmoy, S. M. Mehedi Zaman, Shreya Gautam, Tanay Kumar, Krish Sharma, Niyar R. Barman, Chandan Gupta, Vinija Jain, Aman Chadha, Amit P. Sheth, Amitava Das 0001 |
EMNLP | 11 |
| 2023 | The Troubling Emergence of Hallucination in Large Language Models - An Extensive Definition, Quantification, and Prescriptive RemediationsabstractVipula Rawte, Swagata Chakraborty, Agnibh Pathak, Anubhav Sarkar, S.M Towhidul Islam Tonmoy, Aman Chadha, Amit Sheth, Amitava Das. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Vipula Rawte, Swagata Chakraborty, Agnibh Pathak, Anubhav Sarkar, S. M. Towhidul Islam Tonmoy, Aman Chadha, Amit P. Sheth, Amitava Das 0001 |
EMNLP | 7 |
| 2023 | Cook-Gen: Robust Generative Modeling of Cooking Actions from RecipesabstractAs people become more aware of their food choices, food computation models have become increasingly popular in assisting people in maintaining healthy eating habits. For example, food recommendation systems analyze recipe instructions to assess nutritional contents and provide recipe recommendations. The recent and remarkable successes of generative AI methods, such as auto-regressive Large Language Models, can enable robust methods for a more comprehensive understanding of recipes for healthy food recommendations beyond surface-level nutrition content assessments. In this study, we investigate the use of generative AI methods to extend current food computation models, primarily involving the analysis of nutrition and ingredients, to also incorporate cooking actions (e.g., add salt, fry the meat, boil the vegetables, etc.), Cooking actions are notoriously hard to model using statistical learning methods due to irregular data patterns - significantly varying natural language descriptions for the same action (e.g., marinate the meat vs. marinate the meat and leave overnight) and infrequently occurring patterns (e.g., add salt occurs far more frequently than marinating the meat). The prototypical approach to handling irregular data patterns is to increase the volume of data that the model ingests by orders of magnitude. Unfortunately, in the cooking domain, these problems are further compounded with larger data volumes presenting a unique challenge that is not easily handled by simply scaling up. In this work, we propose novel aggregation-based generative AI methods, Cook-Gen, that reliably generate cooking actions from recipes, despite difficulties with irregular data patterns, while also outperforming Large Language Models and other strong baselines. Revathy Venkataramanan, Kaushik Roy 0009, Kanak Raj, Renjith Prasad, Yuxin Zi, Vignesh Narayanan, Amit P. Sheth |
SMC | 7 |
| 2022 | A Computational Approach to Understand Mental Health from Reddit: Knowledge-Aware Multitask Learning Framework
Usha Lokala, Aseem Srivastava, Triyasha Ghosh Dastidar, Tanmoy Chakraborty 0002, Md. Shad Akhtar, Maryam Panahiazar, Amit P. Sheth |
ICWSM | 7 |
| 2022 | International Workshop on Knowledge Graphs: Open Knowledge NetworkabstractKnowledge networks/graphs provide a powerful approach for data discovery, integration, and reuse. The NSF's new Convergence Accelerator program, which focuses on transitioning research to practice and translational research, announced Track A on the Open Knowledge Network (OKN). The program calls for multidisciplinary and multi-sector teams to work together to build a cooperative and shared open knowledge network infrastructure to drive innovation across science, engineering, and humanities. This workshop aims to invite researchers, practitioners, and the general public to brainstorm the ideas related to OKN, collaboratively build KGs for different domains or applications, develop AI algorithms to provide intelligent services based on OKN, and discuss the social and economic implications related to OKN. Ying Ding 0001, Amit P. Sheth, Krzysztof Janowicz, Sergio Baranzini, Sharat Israni, Ilkay Altintas, Lilit Yeghiazarian, Ellie Young, Sam Klein |
KDD | 2 |
| 2022 | Context-Enriched Learning Models for Aligning Biomedical Vocabularies at Scale in the UMLS MetathesaurusabstractThe Unified Medical Language System (UMLS) Metathesaurus construction process mainly relies on lexical algorithms and manual expert curation for integrating over 200 biomedical vocabularies. A lexical-based learning model (LexLM) was developed to predict synonymy among Metathesaurus terms and largely outperforms a rule-based approach (RBA) that approximates the current construction process. However, the LexLM has the potential for being improved further because it only uses lexical information from the source vocabularies, while the RBA also takes advantage of contextual information. We investigate the role of multiple types of contextual information available to the UMLS editors, namely source synonymy (SS), source semantic group (SG), and source hierarchical relations (HR), for the UMLS vocabulary alignment (UVA) problem. In this paper, we develop multiple variants of context-enriched learning models (ConLMs) by adding to the LexLM the types of contextual information listed above. We represent these context types in context-enriched knowledge graphs (ConKGs) with four variants ConSS, ConSG, ConHR, and ConAll. We train these ConKG embeddings using seven KG embedding techniques. We create the ConLMs by concatenating the ConKG embedding vectors with the word embedding vectors from the LexLM. We evaluate the performance of the ConLMs using the UVA generalization test datasets with hundreds of millions of pairs. Our extensive experiments show a significant performance improvement from the ConLMs over the LexLM, namely +5.0% in precision (93.75%), +0.69% in recall (93.23%), +2.88% in F1 (93.49%) for the best ConLM. Our experiments also show that the ConAll variant including the three context types takes more time, but does not always perform better than other variants with a single context type. Finally, our experiments show that the pairs of terms with high lexical similarity benefit most from adding contextual information, namely +6.56% in precision (94.97%), +2.13% in recall (93.23%), +4.35% in F1 (94.09%) for the best ConLM. The pairs with lower degrees of lexical similarity also show performance improvement with +0.85% in F1 (96%) for low similarity and +1.31% in F1 (96.34%) for no similarity. These results demonstrate the importance of using contextual information in the UVA problem. Vinh Nguyen 0002, Hong Yung Yip, Goonmeet Bajaj, Thilini Wijesiriwardene, Vishesh Javangula, Srinivasan Parthasarathy 0001, Amit P. Sheth, Olivier Bodenreider |
WWW | 7 |
| 2022 | Defining and detecting toxicity on social media: context and knowledge are key
Amit P. Sheth, Valerie L. Shalin, Ugur Kursuncu |
Neurocomputing | 1 |
| 2021 | Designing Children's New Learning Partner: Collaborative Artificial Intelligence for Learning to Solve the Rubik's CubeabstractDeveloping the problem solving skills of children is a challenging problem that is crucial for the future of our society. Given that artificial intelligence (AI) has been used to solve problems across a wide variety of domains, AI offers unique opportunities to develop problem solving skills using a multitude of tasks that pique the curiosity of children. To make this a reality, it is necessary to address the uninterpretable “black-box” that AI often appears to be. Towards this goal, we design a collaborative artificial intelligence algorithm that uses a human-in-the-loop approach to allow students to discover their own personalized solutions to problems. This collaborative algorithm builds on state-of-the-art AI algorithms and leverages additional interpretable structures, namely knowledge graphs and decision trees, to create a fully interpretable process that is able to explain solutions in their entirety. We describe this algorithm when applied to solving the Rubik’s cube as well as our planned user-interface and assessment methods. Forest Agostinelli, Mihir Mavalankar, Vedant Khandelwal, Hengtao Tang, Dezhi Wu, Barnett Berry, Biplav Srivastava, Amit P. Sheth, Matthew Irvin |
IDC | 8 |
| 2021 | Knowledge Infused Policy Gradients with Upper Confidence Bound for Relational Bandits
Kaushik Roy 0009, Qi Zhang 0038, Manas Gaur, Amit P. Sheth |
ECML/PKDD (1) | 4 |
| 2021 | Don't Handicap AI without Explicit Knowledge : Keynote 3abstractSummary form only given, as follows. The complete presentation was not made available for publication as part of the conference proceedings. Knowledge representation as expert system rules or using frames and variety of logics, played a key role in capturing explicit knowledge during the hay days of AI in the past century. Such knowledge, aligned with planning and reasoning are part of what we refer to as Symbolic AI. The resurgent AI of this century in the form of Statistical AI has benefitted from massive data and computing. On some tasks, deep learning methods have even exceeded human performance levels. This gave the false sense that data alone is enough, and explicit knowledge is not needed. But as we start chasing machine intelligence that is comparable with human intelligence, there is an increasing realization that we cannot do without explicit knowledge. Neuroscience (role of long-term memory, strong interactions between different specialized regions of data on tasks such as multimodal sensing), cognitive science (bottom brain versus top brain, perception versus cognition), brain-inspired computing, behavioral economics (system 1 versus system 2), and other disciplines point to need for furthering AI to neuro-symbolic AI (i.e., hybrid of Statistical AI and Symbolic AI, also referred to as the third wave of AI). As we make this progress, the role of explicit knowledge becomes more evident. I will specifically look at our endeavor to support human-like intelligence, our desire for AI systems to interact with humans naturally, and our need to explain the path and reasons for AI systems’ workings. Nevertheless, the variety of knowledge needed to support understanding and intelligence is varied and complex. Using the example of progressing from NLP to NLU, I will demonstrate the dimensions of explicit knowledge, which may include, linguistic, language syntax, common sense, general (world model), specialized (e.g., geographic), and domain-specific (e.g., mental health) knowledge. I will also argue that despite this complexity, such knowledge can be scalability created and maintained (even dynamically or continually). Finally, I will describe our work on knowledge-infused learning as an example strategy for fusing statistical and symbolic AI in a variety of ways. Amit P. Sheth |
SERVICES | 1 |
| 2020 | Identifying Depressive Symptoms from Tweets: Figurative Language Enabled Multitask Learning FrameworkabstractExisting studies on using social media for deriving mental health status of users focus on the depression detection task.However, for case management and referral to psychiatrists, healthcare workers require practical and scalable depressive disorder screening and triage system.This study aims to design and evaluate a decision support system (DSS) to reliably determine the depressive triage level by capturing fine-grained depressive symptoms expressed in user tweets through the emulation of Patient Health Questionnaire-9 (PHQ-9) that is routinely used in clinical practice.The reliable detection of depressive symptoms from tweets is challenging because the 280-character limit on tweets incentivizes the use of creative artifacts in the utterances and figurative usage contributes to effective expression.We propose a novel BERT based robust multi-task learning framework to accurately identify the depressive symptoms using the auxiliary task of figurative usage detection.Specifically, our proposed novel task sharing mechanism, co-task aware attention, enables automatic selection of optimal information across the BERT layers and tasks by soft-sharing of parameters.Our results show that modeling figurative usage can demonstrably improve the model's robustness and reliability for distinguishing the depression symptoms. Shweta Yadav 0001, Jainish Chauhan, Joy Prakash Sain, Krishnaprasad Thirunarayan, Amit P. Sheth, Jeremiah Schumm |
COLING | 5 |
| 2020 | Medical Knowledge-enriched Textual Entailment FrameworkabstractOne of the cardinal tasks in achieving robust medical question answering systems is textual entailment.The existing approaches make use of an ensemble of pre-trained language models or data augmentation, often to clock higher numbers on the validation metrics.However, two major shortcomings impede higher success in identifying entailment: (1) understanding the focus/intent of the question and (2) ability to utilize the real-world background knowledge to capture the context beyond the sentence.In this paper, we present a novel Medical Knowledge-Enriched Textual Entailment framework that allows the model to acquire a semantic and global representation of the input medical text with the help of a relevant domain-specific knowledge graph.We evaluate our framework on the benchmark MEDIQA-RQE dataset and manifest that the use of knowledgeenriched dual-encoding mechanism help in achieving an absolute improvement of 8.27% over SOTA language models.We have made the source code available here. 1 Shweta Yadav 0001, Vishal Pallagani, Amit P. Sheth |
COLING | 3 |
| 2020 | Assessing the Severity of Health States based on Social Media PostsabstractThe unprecedented growth of Internet users has resulted in an abundance of unstructured information on social media including health forums, where patients request health-related information or opinions from other users. Previous studies have shown that online peer support has limited effectiveness without expert intervention. Therefore, a system capable of assessing the severity of health state from the patients' social media posts can help health professionals (HP) in prioritizing the user's post. In this study, we inspect the efficacy of different aspects of Natural Language Understanding (NLU) to identify the severity of the user's health state in relation to two perspectives(tasks) (a) Medical Condition (i.e., Recover, Exist, Deteriorate, Other) and (b) Medication (i.e., Effective, Ineffective, Serious Adverse Effect, Other) in online health communities. We propose a multiview learning framework that models both the textual content as well as contextual-information to assess the severity of the user's health state. Specifically, our model utilizes the NLU views such as sentiment, emotions, personality, and use of figurative language to extract the contextual information. The diverse NLU views demonstrate its effectiveness on both the tasks and as well as on the individual disease to assess a user's health. Shweta Yadav 0001, Joy Prakash Sain, Amit P. Sheth, Asif Ekbal, Sriparna Saha 0001, Pushpak Bhattacharyya |
ICPR | 3 |
| 2020 | eDarkFind: Unsupervised Multi-view Learning for Sybil Account DetectionabstractDarknet crypto markets are online marketplaces using crypto currencies (e.g., Bitcoin, Monero) and advanced encryption techniques to offer anonymity to vendors and consumers trading for illegal goods or services. The exact volume of substances advertised and sold through these crypto markets is difficult to assess, at least partially, because vendors tend to maintain multiple accounts (or Sybil accounts) within and across different crypto markets. Linking these different accounts will allow us to accurately evaluate the volume of substances advertised across the different crypto markets by each vendor. In this paper, we present a multi-view unsupervised framework (eDarkFind) that helps modeling vendor characteristics and facilitates Sybil account detection. We employ a multi-view learning paradigm to generalize and improve the performance by exploiting the diverse views from multiple rich sources such as BERT, stylometric, and location representation. Our model is further tailored to take advantage of domain-specific knowledge such as the Drug Abuse Ontology to take into consideration the substance information. We performed extensive experiments and demonstrated that the multiple views obtained from diverse sources can be effective in linking Sybil accounts. Our proposed eDarkFind model achieves an accuracy of 98% on three real-world datasets which shows the generality of the approach. Ramnath Kumar, Shweta Yadav 0001, Raminta Daniulaityte, Francois R. Lamy, Krishnaprasad Thirunarayan, Usha Lokala, Amit P. Sheth |
WWW | 7 |
| 2019 | Predicting public opinion on drug legalization: social media analysis and consumption trendsabstractIn this paper, we focus on the collection and analysis of relevant Twitter data on a state-by-state basis for (i) measuring public opinion on marijuana legalization by mining sentiment in Twitter data and (ii) determining the usage trends for six distinct types of marijuana. We overcome the challenges posed by the informal and ungrammatical nature of tweets to analyze a corpus of 306,835 relevant tweets collected over the four-month period, preceding the November 2015 Ohio Marijuana Legalization ballot and the four months after the election for all states in the US. Our analysis revealed two key insights: (i) the people in states that have legalized recreational marijuana express greater positive sentiments about marijuana than the people in states that have either legalized medicinal marijuana or have not legalized marijuana at all; (ii) the states that have a high percentage of positive sentiment about marijuana is more inclined to authorize (e.g., by allowing medical marijuana) or broaden its legal usage (e.g., by allowing recreational marijuana in addition to medical marijuana). Our analysis shows that social media can provide reliable information and can serve as an alternative to traditional polling of public opinion on drug use and epidemiology research. Farahnaz Golrooy Motlagh, Saeedeh Shekarpour, Amit P. Sheth, Krishnaprasad Thirunarayan, Michael L. Raymer |
ASONAM | 3 |
| 2019 | Who Should Be the Captain This Week?Leveraging Inferred Diversity-Enhanced Crowd Wisdom for a Fantasy Premier League Captain Prediction
Shreyansh P. Bhatt, Keke Chen, Valerie L. Shalin, Amit P. Sheth, Brandon S. Minnery |
ICWSM | 4 |
| 2019 | A Pipeline for Disaster Response and Relief CoordinationabstractNatural disasters such as floods, forest fires, and hurricanes can cause catastrophic damage to human life and infrastructure. We focus on response to hurricanes caused by both river water flooding and storm surge. Using models for storm surge simulation and flood extent prediction, we generate forecasts about areas likely to be highly affected by the disaster. Further, we overlay the simulation results with information about traffic incidents to correlate traffic incidents with other data modality. We present these results in a modularized, interactive map-based visualization, which can help emergency responders to better plan and coordinate disaster response. Pranav Maneriker, Nikhita Vedula, Hussein Al-Olimat, Jiayong Liang, Omar El-Khoury, Ethan J. Kubatko, Krishnaprasad Thirunarayan, Valerie L. Shalin, Amit P. Sheth, Srinivasan Parthasarathy 0001 |
SIGIR | 10 |
| 2019 | kBot: Knowledge-Enabled Personalized Chatbot for Asthma Self-ManagementabstractThere is a well-recognized need for a shift to proactive asthma care given the impact asthma has on overall healthcare costs. The demand for continuous monitoring of patient's adherence to the medication care plan, assessment of environmental triggers, and management of asthma can be challenging in traditional clinical settings and taxing on clinical professionals. Recent years have seen a robust growth of general purpose conversational systems. However, they lack the capabilities to support applications such an individual's health, which requires the ability to contextualize, learn interactively, and provide the proper hyper-personalization needed to hold meaningful conversations. In this paper, we present kBot, a knowledge-enabled personalized chatbot system designed for health applications and adapted to help pediatric asthmatic patients (age 8 to 15) to better control their asthma. Its core functionalities include continuous monitoring of the patient's medication adherence and tracking of relevant health signals and environment data. kBot takes the form of an Android application with a frontend chat interface capable of conversing in both text and voice, and a backend cloud-based server application that handles data collection, processing, and dialogue management. It achieves contextualization by piecing together domain knowledge from online sources and inputs from our clinical partners. The personalization aspect is derived from patient answering questionnaires and day-to-day conversations. kBOT's preliminary evaluation focused on chatbot quality, technology acceptance, and system usability involved eight asthma clinicians and eight researchers. For both groups, kBot achieved an overall technology acceptance value of greater than 8 on the 11-point Likert scale and a mean System Usability Score (SUS) greater than 80. Dipesh Kadariya, Revathy Venkataramanan, Hong Yung Yip, Maninder Kalra, Krishnaprasad Thirunarayan, Amit P. Sheth |
SMARTCOMP | 6 |
| 2019 | Knowledge Graph Enhanced Community Detection and CharacterizationabstractRecent studies show that by combining network topology and node attributes, we can better understand community structures in complex networks. However, existing algorithms do not explore "contextually" similar node attribute values, and therefore may miss communities defined with abstract concepts. We propose a community detection and characterization algorithm that incorporates the contextual information of node attributes described by multiple domain-specific hierarchical concept graphs. The core problem is to find the context that can best summarize the nodes in communities, while also discovering communities aligned with the context summarizing communities. We formulate the two intertwined problems, optimal community-context computation, and community discovery, with a coordinate-ascent based algorithm that iteratively updates the nodes' community label assignment with a community-context and computes the best context summarizing nodes of each community. Our unique contributions include (1) a composite metric on Informativeness and Purity criteria in searching for the best context summarizing nodes of a community; (2) a node similarity measure that incorporates the context-level similarity on multiple node attributes; and (3) an integrated algorithm that drives community structure discovery by appropriately weighing edges. Experimental results on public datasets show nearly 20 percent improvement on F-measure and Jaccard for discovering underlying community structure over the current state-of-the-art of community detection methods. Community structure characterization was also accurate to find appropriate community types for four datasets. Shreyansh P. Bhatt, Swati Padhee, Amit P. Sheth, Keke Chen, Valerie L. Shalin, Derek Doran, Brandon S. Minnery |
WSDM | 3 |
| 2019 | Knowledge-aware Assessment of Severity of Suicide Risk for Early InterventionabstractMental health illness such as depression is a significant risk factor for suicide ideation, behaviors, and attempts. A report by Substance Abuse and Mental Health Services Administration (SAMHSA) shows that 80% of the patients suffering from Borderline Personality Disorder (BPD) have suicidal behavior, 5-10% of whom commit suicide. While multiple initiatives have been developed and implemented for suicide prevention, a key challenge has been the social stigma associated with mental disorders, which deters patients from seeking help or sharing their experiences directly with others including clinicians. This is particularly true for teenagers and younger adults where suicide is the second highest cause of death in the US. Prior research involving surveys and questionnaires (e.g. PHQ-9) for suicide risk prediction failed to provide a quantitative assessment of risk that informed timely clinical decision-making for intervention. Our interdisciplinary study concerns the use of Reddit as an unobtrusive data source for gleaning information about suicidal tendencies and other related mental health conditions afflicting depressed users. We provide details of our learning framework that incorporates domain-specific knowledge to predict the severity of suicide risk for an individual. Our approach involves developing a suicide risk severity lexicon using medical knowledge bases and suicide ontology to detect cues relevant to suicidal thoughts and actions. We also use language modeling, medical entity recognition and normalization and negation detection to create a dataset of 2181 redditors that have discussed or implied suicidal ideation, behavior, or attempt. Given the importance of clinical knowledge, our gold standard dataset of 500 redditors (out of 2181) was developed by four practicing psychiatrists following the guidelines outlined in Columbia Suicide Severity Rating Scale (C-SSRS), with the pairwise annotator agreement of 0.79 and group-wise agreement of 0.73. Compared to the existing four-label classification scheme (no risk, low risk, moderate risk, and high risk), our proposed C-SSRS-based 5-label classification scheme distinguishes people who are supportive, from those who show different severity of suicidal tendency. Our 5-label classification scheme outperforms the state-of-the-art schemes by improving the graded recall by 4.2% and reducing the perceived risk measure by 12.5%. Convolutional neural network (CNN) provided the best performance in our scheme due to the discriminative features and use of domain-specific knowledge resources, in comparison to SVM-L that has been used in the state-of-the-art tools over similar dataset. Manas Gaur, Amanuel Alambo, Joy Prakash Sain, Ugur Kursuncu, Krishnaprasad Thirunarayan, Ramakanth Kavuluru, Amit P. Sheth, Randy S. Welton, Jyotishman Pathak |
WWW | 7 |
| 2019 | Processing social media in real-time
Damiano Spina, Arkaitz Zubiaga, Amit P. Sheth, Markus Strohmaier |
Inf. Process. Manag. | 3 |
| 2019 | Social determinants of health in mental health care and research: a case for greater inclusionabstractSocial determinants of health (SDOH) are known to influence mental health outcomes, which are independent risk factors for poor health status and physical illness. Currently, however, existing SDOH data collection methods are ad hoc and inadequate, and SDOH data are not systematically included in clinical research or used to inform patient care. Social contextual data are rarely captured prospectively in a structured and comprehensive manner, leaving large knowledge gaps. Extraction methods are now being developed to facilitate the collection, standardization, and integration of SDOH data into electronic health records. If successful, these efforts may have implications for health equity, such as reducing disparities in access and outcomes. Broader use of surveys, natural language processing, and machine learning methods to harness SDOH may help researchers and clinical teams reduce barriers to mental health care. Joseph DeFerio, Scott Breitinger, Dhruv Khullar, Amit P. Sheth, Jyotishman Pathak |
J. Am. Medical Informatics Assoc. | 4 |
| 2019 | Modeling Islamist Extremist Communications on Social Media using Contextual Dimensions: Religion, Ideology, and HateabstractTerror attacks have been linked in part to online extremist content. Online conversations are cloaked in religious ambiguity, with deceptive intentions, often twisted from mainstream meaning to serve a malevolent ideology. Although tens of thousands of Islamist extremism supporters consume such content, they are a small fraction relative to peaceful Muslims. The efforts to contain the ever-evolving extremism on social media platforms have remained inadequate and mostly ineffective. Divergent extremist and mainstream contexts challenge machine interpretation, with a particular threat to the precision of classification algorithms. Radicalization is a subtle long-running persuasive process that occurs over time. Our context-aware computational approach to the analysis of extremist content on Twitter breaks down this persuasion process into building blocks that acknowledge inherent ambiguity and sparsity that likely challenge both manual and automated classification. Based on prior empirical and qualitative research in social sciences, particularly political science, we model this process using a combination of three contextual dimensions -- religion, ideology, and hate -- each elucidating a degree of radicalization and highlighting independent features to render them computationally accessible. We utilize domain-specific knowledge resources for each of these contextual dimensions such as Qur'an for religion, the books of extremist ideologues and preachers for political ideology and a social media hate speech corpus for hate. The significant sensitivity of the Islamist extremist ideology and its local and global security implications require reliable algorithms for modelling such communications on Twitter. Our study makes three contributions to reliable analysis: (i) Development of a computational approach rooted in the contextual dimensions of religion, ideology, and hate, which reflects strategies employed by online Islamist extremist groups, (ii) An in-depth analysis of relevant tweet datasets with respect to these dimensions to exclude likely mislabeled users, and (iii) A framework for understanding online radicalization as a process to assist counter-programming. Given the potentially significant social impact, we evaluate the performance of our algorithms to minimize mislabeling, where our context-aware approach outperforms a competitive baseline by 10.2% in precision, thereby enhancing the potential of such tools for use in human review. Ugur Kursuncu, Manas Gaur, Carlos Castillo 0001, Amanuel Alambo, Krishnaprasad Thirunarayan, Valerie L. Shalin, Dilshod Achilov, Ismailcem Budak Arpinar, Amit P. Sheth |
Proc. ACM Hum. Comput. Interact. | 9 |
| 2018 | "Let Me Tell You About Your Mental Health!": Contextualized Classification of Reddit Posts to DSM-5 for Web-based InterventionabstractSocial media platforms are increasingly being used to share and seek advice on mental health issues. In particular, Reddit users freely discuss such issues on various subreddits, whose structure and content can be leveraged to formally interpret and relate subreddits and their posts in terms of mental health diagnostic categories. There is prior research on the extraction of mental health-related information, including symptoms, diagnosis, and treatments from social media; however, our approach can additionally provide actionable information to clinicians about the mental health of a patient in diagnostic terms for web-based intervention. Specifically, we provide a detailed analysis of the nature of subreddit content from domain expert's perspective and introduce a novel approach to map each subreddit to the best matching DSM-5 (Diagnostic and Statistical Manual of Mental Disorders - 5th Edition) category using multi-class classifier. Our classification algorithm analyzes all the posts of a subreddit by adapting topic modeling and word-embedding techniques, and utilizing curated medical knowledge bases to quantify relationship to DSM-5 categories. Our semantic encoding-decoding optimization approach reduces the false-alarm-rate from 30% to 2.5% over a comparable heuristic baseline, and our mapping results have been verified by domain experts achieving a kappa score of 0.84. Manas Gaur, Ugur Kursuncu, Amanuel Alambo, Amit P. Sheth, Raminta Daniulaityte, Krishnaprasad Thirunarayan, Jyotishman Pathak |
CIKM | 4 |
| 2018 | A Practical Incremental Learning Framework For Sparse Entity ExtractionabstractThis work addresses challenges arising from extracting entities from textual data, including the high cost of data annotation, model accuracy, selecting appropriate evaluation criteria, and the overall quality of annotation. We present a framework that integrates Entity Set Expansion (ESE) and Active Learning (AL) to reduce the annotation cost of sparse data and provide an online evaluation method as feedback. This incremental and interactive learning framework allows for rapid annotation and subsequent extraction of sparse data while maintaining high accuracy. We evaluate our framework on three publicly available datasets and show that it drastically reduces the cost of sparse entity annotation by an average of 85% and 45% to reach 0.9 and 1.0 F-Scores respectively. Moreover, the method exhibited robust performance across all datasets. Hussein Al-Olimat, Steven Gustafson, Jason Mackay, Krishnaprasad Thirunarayan, Amit P. Sheth |
COLING | 5 |
| 2018 | Location Name Extraction from Targeted Text Streams using Gazetteer-based Statistical Language ModelsabstractExtracting location names from informal and unstructured social media data requires the identification of referent boundaries and partitioning compound names. Variability, particularly systematic variability in location names (Carroll, 1983), challenges the identification task. Some of this variability can be anticipated as operations within a statistical language model, in this case drawn from gazetteers such as OpenStreetMap (OSM), Geonames, and DBpedia. This permits evaluation of an observed n-gram in Twitter targeted text as a legitimate location name variant from the same location-context. Using n-gram statistics and location-related dictionaries, our Location Name Extraction tool (LNEx) handles abbreviations and automatically filters and augments the location names in gazetteers (handling name contractions and auxiliary contents) to help detect the boundaries of multi-word location names and thereby delimit them in texts. We evaluated our approach on 4,500 event-specific tweets from three targeted streams to compare the performance of LNEx against that of ten state-of-the-art taggers that rely on standard semantic, syntactic and/or orthographic features. LNEx improved the average F-Score by 33-179%, outperforming all taggers. Further, LNEx is capable of stream processing. Hussein Al-Olimat, Krishnaprasad Thirunarayan, Valerie L. Shalin, Amit P. Sheth |
COLING | 4 |
| 2018 | Enhancing Crowd Wisdom Using Explainable Diversity Inferred from Social MediaabstractA crowd sampled from a set of individuals can provide a more accurate prediction in aggregate than most individuals.This effect, referred to as wisdom of crowd, exists when crowd members bring diverse perspectives to decision making. Such diversity leads to uncorrelated prediction errors that cancel out in aggregate. As crowd members' judgments are often the result of solution strategies, diversity in solution strategies can enhance crowd wisdom. One of the most challenging tasks in sampling such a crowd is to determine the individual's solution strategy for a prediction problem. As participating individuals often share their perspectives through social media, we can use such data to identify an individual's solution strategy. In this paper, we propose a crowd selection approach using social media posts (tweets) indicating diverse solution strategies. We use tweet classification to identify participants' prediction strategies and categorize participants based on the binomial test to identify sets of participants that apply a similar strategy. We then form a diverse crowd by sampling participants from different sets. Using the domain of Fantasy Sports, we show that such a diverse crowd can outperform crowd selected at random and 90% of individual participants, and participant categorization schemes using word2vec. Further, we use a knowledge graph to investigate the factors forming such a diverse crowd and how these factors can lead to a better decision. Relative to bottom-up (data-driven) processes the approach presented here provides an explanation of diverse crowd behavior. Shreyansh P. Bhatt, Manas Gaur, Beth Bullemer, Valerie L. Shalin, Amit P. Sheth, Brandon S. Minnery |
WI | 5 |
| 2018 | What's ur Type? Contextualized Classification of User Types in Marijuana-Related Communications Using Compositional Multiview EmbeddingabstractWith 93% of pro-marijuana population in US favoring legalization of medical marijuana, high expectations of a greater return for Marijuana stocks, and public actively sharing information about medical, recreational and business aspects related to marijuana, it is no surprise that marijuana culture is thriving on Twitter. After the legalization of marijuana for recreational and medical purposes in 29 states, there has been a dramatic increase in the volume of drug-related communications on Twitter. Specifically, Twitter accounts have been established for promotional and informational purposes, some prominent among them being American Ganja, Medical Marijuana Exchange, and Cannabis Now. Identification and characterization of different user types can allow us to conduct more fine-grained spatiotemporal analysis to identify dominant or emerging topics in the echo chambers of marijuana-related communities on Twitter. In this research, we mainly focus on classifying Twitter accounts created and run by ordinary users, retailers, and informed agencies. Classifying user accounts by type can enable better capturing and highlighting of aspects such as trending topics, business profiling of marijuana companies, and state-specific marijuana policymaking. Furthermore, type-based analysis can provide more profound understanding and reliable assessment of the implications of marijuana-related communications. We developed a comprehensive approach to classifying users by their types on Twitter through contextualization of their marijuana-related conversations. We accomplished this using compositional multiview embedding synthesized from People, Content, and Network views achieving 8% improvement over the empirical baseline. Ugur Kursuncu, Manas Gaur, Usha Lokala, Anurag Illendula, Krishnaprasad Thirunarayan, Raminta Daniulaityte, Amit P. Sheth, Ismailcem Budak Arpinar |
WI | 7 |
| 2018 | Building IoT-Based Applications for Smart Cities: How Can Ontology Catalogs Help?abstractThe Internet of Things (IoT) plays an ever-increasing role in enabling smart city applications. An ontology-based semantic approach can help improve interoperability between a variety of IoT-generated as well as complementary data needed to drive these applications. While multiple ontology catalogs exist, using them for IoT and smart city applications require significant amount of work. In this paper, we demonstrate how can ontology catalogs be more effectively used to design and develop smart city applications? We consider four ontology catalogs that are relevant for IoT and smart cities: 1) READY4SmartCities; 2) linked open vocabulary (LOV); 3) OpenSensingCity (OSC); and 4) LOVs for IoT (LOV4IoT). To support semantic interoperability with the reuse of ontology-based smart city applications, we present a methodology to enrich ontology catalogs with those ontologies. Our methodology is generic enough to be applied to any other domains as is demonstrated by its adoption by OSC and LOV4IoT ontology catalogs. Researchers and developers have completed a survey-based evaluation of the LOV4IoT catalog. The usefulness of ontology catalogs ascertained through this evaluation has encouraged their ongoing growth and maintenance. The quality of IoT and smart city ontologies have been evaluated to improve the ontology catalog quality. We also share the lessons learned regarding ontology best practices and provide suggestions for ontology improvements with a set of software tools. Amelie Gyrard, Antoine Zimmermann, Amit P. Sheth |
IEEE Internet Things J. | 3 |
| 2017 | RQUERY: Rewriting Natural Language Queries on Knowledge Graphs to Alleviate the Vocabulary Mismatch ProblemabstractFor non-expert users, a textual query is the most popular and simple means for communicating with a retrieval or question answering system.However, there is a risk of receiving queries which do not match with the background knowledge.Query expansion and query rewriting are solutions for this problem but they are in danger of potentially yielding a large number of irrelevant words, which in turn negatively influences runtime as well as accuracy.In this paper, we propose a new method for automatic rewriting input queries on graph-structured RDF knowledge bases.We employ a Hidden Markov Model to determine the most suitable derived words from linguistic resources.We introduce the concept of triple-based co-occurrence for recognizing co-occurred words in RDF data.This model was bootstrapped with three statistical distributions.Our experimental study demonstrates the superiority of the proposed approach to the traditional n-gram model. Saeedeh Shekarpour, Edgard Marx, Sören Auer, Amit P. Sheth |
AAAI | 4 |
| 2017 | Semi-Supervised Approach to Monitoring Clinical Depressive Symptoms in Social MediaabstractWith the rise of social media, millions of people are routinely expressing their moods, feelings, and daily struggles with mental health issues on social media platforms like Twitter. Unlike traditional observational cohort studies conducted through questionnaires and self-reported surveys, we explore the reliable detection of clinical depression from tweets obtained unobtrusively. Based on the analysis of tweets crawled from users with self-reported depressive symptoms in their Twitter profiles, we demonstrate the potential for detecting clinical depression symptoms which emulate the PHQ-9 questionnaire clinicians use today. Our study uses a semi-supervised statistical model to evaluate how the duration of these symptoms and their expression on Twitter (in terms of word usage patterns and topical preferences) align with the medical findings reported via the PHQ-9. Our proactive and automatic screening tool is able to identify clinical depressive symptoms with an accuracy of 68% and precision of 72%. Amir Hossein Yazdavar, Hussein Al-Olimat, Monireh Ebrahimi, Goonmeet Bajaj, Tanvi Banerjee, Krishnaprasad Thirunarayan, Jyotishman Pathak, Amit P. Sheth |
ASONAM | 8 |
| 2017 | Domain-specific hierarchical subgraph extraction: A recommendation use caseabstractHierarchical relationships play a key role in knowledge graphs. Particularly, large and well-known knowledge graphs such as DBpedia contain significant number of facts expressed with hierarchical relationships in comparison to the other types of relationships. These hierarchical relationships are extensively harnessed by applications such as personalization, question answering, and recommendation systems. However, the presence of large number of facts with hierarchical relationships makes the applications computationally intensive. Additionally, the applications can be domain-specific and may not require all the hierarchical facts available, but only require those that are specific to the domain. In this paper, we present an approach to extract domain-specific hierarchical subgraph from large knowledge graphs by identifying the domain-specificity of the categories in the hierarchy. Given a domain, the domain-specificity of categories are determined by combining different types of evidence using a probabilistic framework. We show the effectiveness of our approach with a recommendation use case for movie and book domains. Our evaluation demonstrates that the domain-specific hierarchical subgraphs extracted by our approach can reduce the baseline subgraph by 40% to 50% without compromising the accuracy of the recommendations. Furthermore, the presented approach outperforms the recommendation results obtained with a state-of-the-art domain-specific subgraph extraction technique which uses supervised learning. Sarasi Lalithsena, Sujan Perera, Pavan Kapanipathi, Amit P. Sheth |
IEEE BigData | 4 |
| 2017 | EmojiNet: An Open Service and API for Emoji Sense Discovery
Sanjaya Wijeratne, Lakshika Balasuriya, Amit P. Sheth, Derek Doran |
ICWSM | 3 |
| 2017 | Relatedness-based Multi-Entity SummarizationabstractRepresenting world knowledge in a machine processable format is important as entities and their descriptions have fueled tremendous growth in knowledge-rich information processing platforms, services, and systems. Prominent applications of knowledge graphs include search engines (e.g., Google Search and Microsoft Bing), email clients (e.g., Gmail), and intelligent personal assistants (e.g., Google Now, Amazon Echo, and Apple's Siri). In this paper, we present an approach that can summarize facts about a collection of entities by analyzing their relatedness in preference to summarizing each entity in isolation. Specifically, we generate informative entity summaries by selecting: (i) inter-entity facts that are similar and (ii) intra-entity facts that are important and diverse. We employ a constrained knapsack problem solving approach to efficiently compute entity summaries. We perform both qualitative and quantitative experiments and demonstrate that our approach yields promising results compared to two other stand-alone state-of-the-art entity summarization approaches. Kalpa Gunaratna, Amir Hossein Yazdavar, Krishnaprasad Thirunarayan, Amit P. Sheth, Gong Cheng 0001 |
IJCAI | 4 |
| 2017 | Enhancing crowd wisdom using measures of diversity computed from social media dataabstract"Wisdom of Crowds" (WoC) refers to a form of collective intelligence in which the aggregate judgment of a group of individuals is, in most instances, superior to that of any one group member. For a crowd to be wise, its members must possess diverse knowledge and viewpoints. Such diversity leads to uncorrelated judgment errors that cancel out in aggregate. Yet despite the fact that diversity is known to be an essential ingredient in WoC, little research aims to measure and exploit diversity in human social systems for the purpose of maximizing crowd intelligence. Here we quantify the diversity of a group of individuals through semantic analysis of their social media (Twitter) communications. Focusing on the domain of fantasy sports, we show that virtual crowds of fantasy team owners selected based on the diversity of their tweet content can outperform both non-diverse and randomly sampled crowds. Our results suggest a new approach for intelligent crowd assembly in which measures of diversity extracted from online social media communications can guide the selection of crowd members. These results have implications for numerous domains that utilize aggregated judgments - from consumer reviews, to econometrics, to geopolitical forecasting and intelligence analysis. Shreyansh P. Bhatt, Brandon S. Minnery, Srikanth Nadella, Beth Bullemer, Valerie L. Shalin, Amit P. Sheth |
WI | 6 |
| 2017 | Knowledge will propel machine understanding of content: extrapolating from current examplesabstractMachine Learning has been a big success story during the AI resurgence. One particular stand out success relates to learning from a massive amount of data. In spite of early assertions of the unreasonable effectiveness of data, there is increasing recognition for utilizing knowledge whenever it is available or can be created purposefully. In this paper, we discuss the indispensable role of knowledge for deeper understanding of content where (i) large amounts of training data are unavailable, (ii) the objects to be recognized are complex, (e.g., implicit entities and highly subjective content), and (iii) applications need to use complementary or related data in multiple modalities/media. What brings us to the cusp of rapid progress is our ability to (a) create relevant and reliable knowledge and (b) carefully exploit knowledge to enhance ML/NLP techniques. Using diverse examples, we seek to foretell unprecedented progress in our ability for deeper understanding and exploitation of multimodal data and continued incorporation of knowledge in learning techniques. Amit P. Sheth, Sujan Perera, Sanjaya Wijeratne, Krishnaprasad Thirunarayan |
WI | 1 |
| 2017 | Adaptive training instance selection for cross-domain emotion identificationabstractThis paper exploits a large number of self-labeled emotion tweets as the training data from the source domain to improve emotion identification in target domains (i.e., blogs and fairy tales), where there is a short supply of labeled data. Due to the noisy and ambiguous nature of self-labeled emotion training data, the existing domain adaptation methods that typically depend on high-quality labeled source-domain data do not work satisfactorily. This paper describes an adaptive source-domain training instance selection method to address the problem of noisy source-domain training data. The proposed approach can effectively identify the most informative training examples based on three carefully designed measures: consistency, diversity, and similarity. It uses an iterative method that consists of the following steps in each iteration: selecting informative samples from the source domain with the informativeness measures, merging with the target-domain training data, evaluating the performance of learned classifier for the target domain, and updating the informativeness measures for the next iteration. It stops until no new training instance is selected or in a designated number of iterations. Experiments show that our approach performs effectively for cross-domain emotion identification and consistently outperforms baseline approaches across four domains. Wenbo Wang 0002, Keke Chen, Krishnaprasad Thirunarayan, Amit P. Sheth |
WI | 5 |
| 2017 | A semantics-based measure of emoji similarityabstractEmoji have grown to become one of the most important forms of communication on the web. With its widespread use, measuring the similarity of emoji has become an important problem for contemporary text processing since it lies at the heart of sentiment analysis, search, and interface design tasks. This paper presents a comprehensive analysis of the semantic similarity of emoji through embedding models that are learned over machine-readable emoji meanings in the EmojiNet knowledge base. Using emoji descriptions, emoji sense labels and emoji sense definitions, and with different training corpora obtained from Twitter and Google News, we develop and test multiple embedding models to measure emoji similarity. To evaluate our work, we create a new dataset called EmoSim508, which assigns human-annotated semantic similarity scores to a set of 508 carefully selected emoji pairs. After validation with EmoSim508, we present a real-world use-case of our emoji embedding models using a sentiment analysis task and show that our models outperform the previous best-performing emoji embedding model on this task. The EmoSim508 dataset and our emoji embedding models are publicly released with this paper and can be downloaded from http://emojinet.knoesis.org/. Sanjaya Wijeratne, Lakshika Balasuriya, Amit P. Sheth, Derek Doran |
WI | 3 |
| 2016 | Understanding City Traffic Dynamics Utilizing Sensor and Textual ObservationsabstractUnderstanding speed and travel-time dynamics in response to various city related events is an important and challenging problem. Sensor data (numerical) containing average speed of vehicles passing through a road link can be interpreted in terms of traffic related incident reports from city authorities and social media data (textual), providing a complementary understanding of traffic dynamics. State-of-the-art research is focused on either analyzing sensor observations or citizen observations; we seek to exploit both in a synergistic manner. We demonstrate the role of domain knowledge in capturing the non-linearity of speed and travel-time dynamics by segmenting speed and travel-time observations into simpler components amenable to description using linear models such as Linear Dynamical System (LDS). Specifically, we propose Restricted Switching Linear Dynamical System (RSLDS) to model normal speed and travel time dynamics and thereby characterize anomalous dynamics. We utilize the city traffic events extracted from text to explain anomalous dynamics. We present a large scale evaluation of the proposed approach on a real-world traffic and twitter dataset collected over a year with promising results. Pramod Anantharam, Krishnaprasad Thirunarayan, Surendra Marupudi, Amit P. Sheth, Tanvi Banerjee |
AAAI | 4 |
| 2016 | Finding street gang members on Twitterabstractscore with a low false positive rate. Lakshika Balasuriya, Sanjaya Wijeratne, Derek Doran, Amit P. Sheth |
ASONAM | 4 |
| 2016 | Harnessing relationships for domain-specific subgraph extraction: A recommendation use caseabstractApplications on the Web such as search engines and recommendation systems are increasingly adapting semantic approaches by leveraging knowledge graphs. While some applications require processing of the whole knowledge graph, most are domain-specific and require only a relevant subset of it. For example, a movie or a book recommendation system would require a subgraph that comprises knowledge relevant to the specific domain. In such scenarios, processing the whole knowledge graph, particularly the commonly used, large, and openly available knowledge graphs on the Web, is computationally intensive and the irrelevant portion may negatively impact the performance of the application. This necessitates the identification and extraction of relevant subgraphs that adequately captures entities and their relationships for a given application domain and/or task. In this work, we present an approach to identify a minimal domain-specific subgraph by utilizing statistic and semantic-based metrics. Our approach highlights the importance of relationships as first-class elements to capture the domain specificity of a subgraph. We demonstrate the applicability of this approach for a recommendation use case on two domains, i.e. movie and book. Our evaluation demonstrates a reduction of 80% to 90% of the knowledge graph with orders of magnitude decrease in time for computation without compromising accuracy. Sarasi Lalithsena, Pavan Kapanipathi, Amit P. Sheth |
IEEE BigData | 3 |
| 2016 | Gleaning Types for Literals in RDF Triples with Application to Entity Summarization
Kalpa Gunaratna, Krishnaprasad Thirunarayan, Amit P. Sheth, Gong Cheng 0001 |
ESWC | 3 |
| 2016 | Implicit Entity Linking in Tweets
Sujan Perera, Pablo N. Mendes, Adarsh Alex, Amit P. Sheth, Krishnaprasad Thirunarayan |
ESWC | 4 |
| 2016 | Clustering for Simultaneous Extraction of Aspects and Features from ReviewsabstractLu Chen, Justin Martineau, Doreen Cheng, Amit Sheth. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Justin Martineau, Doreen Cheng, Amit P. Sheth |
HLT-NAACL | 4 |
| 2015 | FACES: Diversity-Aware Entity Summarization Using Incremental Hierarchical Conceptual ClusteringabstractSemantic Web documents that encode facts about entities on the Web have been growing rapidly in size and evolving over time. Creating summaries on lengthy Semantic Web documents for quick identification of the corresponding entity has been of great contemporary interest. In this paper, we explore automatic summarization techniques that characterize and enable identification of an entity and create summaries that are human friendly. Specifically, we highlight the importance of diversified (faceted) summaries by combining three dimensions: diversity, uniqueness, and popularity. Our novel diversity-aware entity summarization approach mimics human conceptual clustering techniques to group facts and picks representative facts from each group to form concise (i.e., short) and comprehensive (i.e., improved coverage through diversity) summaries. We evaluate our approach against the state-of-the-art techniques and show that our work improves both the quality and the efficiency of entity summarization. Kalpa Gunaratna, Krishnaprasad Thirunarayan, Amit P. Sheth |
AAAI | 3 |
| 2015 | Knowledge Enabled Approach to Predict the Location of Twitter Users
Revathy Krishnamurthy, Pavan Kapanipathi, Amit P. Sheth, Krishnaprasad Thirunarayan |
ESWC | 3 |
| 2015 | Analyzing the social media footprint of street gangsabstractGangs utilize social media as a way to maintain threatening virtual presences, to communicate about their activities, and to intimidate others. Such usage has gained the attention of many justice service agencies that wish to create better crime prevention and judicial services. However, these agencies use analysis methods that are labor intensive and only lead to basic, qualitative data interpretations. This paper presents the architecture of a modern platform to discover the structure, function, and operation of gangs through the lens of social media. Preliminary analysis of social media posts shared in the greater Chicago, IL region demonstrate the platform's capability to understand gang members' social media usage patterns. Sanjaya Wijeratne, Derek Doran, Amit P. Sheth, Jack L. Dustin |
ISI | 3 |
| 2015 | Context-driven automatic subgraph creation for literature-based discovery
Delroy Cameron, Ramakanth Kavuluru, Thomas C. Rindflesch, Amit P. Sheth, Krishnaprasad Thirunarayan, Olivier Bodenreider |
J. Biomed. Informatics | 4 |
| 2015 | Extracting City Traffic Events from Social StreamsabstractCities are composed of complex systems with physical, cyber, and social components. Current works on extracting and understanding city events mainly rely on technology-enabled infrastructure to observe and record events. In this work, we propose an approach to leverage citizen observations of various city systems and services, such as traffic, public transport, water supply, weather, sewage, and public safety, as a source of city events. We investigate the feasibility of using such textual streams for extracting city events from annotated text. We formalize the problem of annotating social streams such as microblogs as a sequence labeling problem. We present a novel training data creation process for training sequence labeling models. Our automatic training data creation process utilizes instance-level domain knowledge (e.g., locations in a city, possible event terms). We compare this automated annotation process to a state-of-the-art tool that needs manually created training data and show that it has comparable performance in annotation tasks. An aggregation algorithm is then presented for event extraction from annotated text. We carry out a comprehensive evaluation of the event annotation and event extraction on a real-world dataset consisting of event reports and tweets collected over 4 months from the San Francisco Bay Area. The evaluation results are promising and provide insights into the utility of social stream for extracting city events. Pramod Anantharam, Payam M. Barnaghi, Krishnaprasad Thirunarayan, Amit P. Sheth |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2015 | Application Portability in Cloud Computing: An Abstraction-Driven PerspectiveabstractCloud computing has changed the way organizations create, manage, and evolve their applications. While the abundance of computing resources at low cost opens up many possibilities for migrating applications to the cloud, this migration also comes at a price. Cloud applications, in many cases, depend on certain provider specific features or services. In moving applications to the cloud, application developers face the challenge of balancing these dependencies to avoid vendor lock-in. We present an abstraction-driven approach to address the application portability issues and focus on the application development process. We also present our theoretical basis and experience in two practical projects where we have applied the abstraction-driven approach. Ajith Ranabahu, E. Michael Maximilien, Amit P. Sheth, Krishnaprasad Thirunarayan |
IEEE Trans. Serv. Comput. | 3 |
| 2014 | Active Learning with Efficient Feature Weighting Methods for Improving Data Quality and Classification AccuracyabstractMany machine learning datasets are noisy with a substantial number of mislabeled instances. This noise yields sub-optimal classification performance. In this paper we study a large, low quality annotated dataset, created quickly and cheaply using Amazon Mechanical Turk to crowdsource annotations. We describe computationally cheap feature weighting techniques and a novel non-linear distribution spreading algorithm that can be used to iteratively and interactively correcting mislabeled instances to significantly improve annotation quality at low cost. Eight different emotion extraction experiments on Twitter data demonstrate that our approach is just as effective as more computationally expensive techniques. Our techniques save a considerable amount of time. Justin Martineau, Doreen Cheng, Amit P. Sheth |
ACL (1) | 4 |
| 2014 | Analysis of Online Information Searching for Cardiovascular Diseases on a Consumer Health Information Portal
Ashutosh Jadhav, Amit P. Sheth, Jyotishman Pathak |
AMIA | 2 |
| 2014 | An Analysis of Mayo Clinic Search Query Logs for Cardiovascular Diseases
Ashutosh Jadhav, Amit P. Sheth, Jyotishman Pathak |
AMIA | 2 |
| 2014 | Applications of multimodal physical (IoT), cyber and social data for reliable and actionable insightsabstractPhysical objects with embedded sensors are increasingly be- ing networked together using wireless and internet technologies to form Internet of Things (IoT). However, early applications that rely on IoT data fail to provide comprehensive situational awareness. This often requires combining physical Amit P. Sheth, Pramod Anantharam, Krishnaprasad Thirunarayan |
CollaborateCom | 1 |
| 2014 | Cursing in English on twitterabstractCursing is not uncommon during conversations in the physical world: 0.5% to 0.7% of all the words we speak are curse words, given that 1% of all the words are first-person plural pronouns (e.g., we, us, our). On social media, people can instantly chat with friends without face-to-face interaction, usually in a more public fashion and broadly disseminated through highly connected social network. Will these distinctive features of social media lead to a change in people's cursing behavior? In this paper, we examine the characteristics of cursing activity on a popular social media platform - Twitter, involving the analysis of about 51 million tweets and about 14 million users. In particular, we explore a set of questions that have been recognized as crucial for understanding cursing in offline communications by prior studies, including the ubiquity, utility, and contextual dependencies of cursing. Wenbo Wang 0002, Krishnaprasad Thirunarayan, Amit P. Sheth |
CSCW | 4 |
| 2014 | User Interests Identification on Twitter Using a Hierarchical Knowledge Base
Pavan Kapanipathi, Prateek Jain 0001, Chitra Venkatramani, Amit P. Sheth |
ESWC | 4 |
| 2014 | Transforming Big Data into Smart Data: Deriving value via harnessing Volume, Variety, and Velocity using semantic techniques and technologiesabstractBig Data has captured a lot of interest in industry, with anticipation of better decisions, efficient organizations, and many new jobs. Much of the emphasis is on the challenges of the four V's of Big Data: Volume, Variety, Velocity, and Veracity, and technologies that handle volume, including storage and computational techniques to support analysis (Hadoop, NoSQL, MapReduce, etc). However, the most important feature of Big Data, the raison d'etre, is none of these 4 V's — but value. In this talk, I will forward the concept of Smart Data that is realized by extracting value from a variety of data, and how Smart Data for growing variety (e.g., social, sensor/IoT, health care) of Big Data enable a much larger class of applications that can benefit not just large companies but each individual. This requires organized ways to harness and overcome the four V-challenges. In particular, we will need to utilize metadata, employ semantics and intelligent processing, and go beyond traditional reliance on ML and NLP. Amit P. Sheth |
ICDE | 1 |
| 2014 | YouRank: Let User Engagement Rank Microblog Search Results
Wenbo Wang 0002, Lei Duan, Anirudh Koul, Amit P. Sheth |
ICWSM | 4 |
| 2014 | On Understanding the Divergence of Online Social Group Discussion
Hemant Purohit, Yiye Ruan, David Fuhry, Srinivasan Parthasarathy 0001, Amit P. Sheth |
ICWSM | 5 |
| 2014 | Don't like RDF reification?: making statements about statements using singleton propertyabstractStatements about RDF statements, or meta triples, provide additional information about individual triples, such as the source, the occurring time or place, or the certainty. Integrating such meta triples into semantic knowledge bases would enable the querying and reasoning mechanisms to be aware of provenance, time, location, or certainty of triples. However, an efficient RDF representation for such meta knowledge of triples remains challenging. The existing standard reification approach allows such meta knowledge of RDF triples to be expressed using RDF by two steps. The first step is representing the triple by a Statement instance which has subject, predicate, and object indicated separately in three different triples. The second step is creating assertions about that instance as if it is a statement. While reification is simple and intuitive, this approach does not have formal semantics and is not commonly used in practice as described in the RDF Primer. In this paper, we propose a novel approach called Singleton Property for representing statements about statements and provide a formal semantics for it. We explain how this singleton property approach fits well with the existing syntax and formal semantics of RDF, and the syntax of SPARQL query language. We also demonstrate the use of singleton property in the representation and querying of meta knowledge in two examples of Semantic Web knowledge bases: YAGO2 and BKR. Our experiments on the BKR show that the singleton property approach gives a decent performance in terms of number of triples, query length and query execution time compared to existing approaches. This approach, which is also simple and intuitive, can be easily adopted for representing and querying statements about statements in other knowledge bases. Vinh Nguyen 0002, Olivier Bodenreider, Amit P. Sheth |
WWW | 3 |
| 2014 | Identifying Seekers and Suppliers in Social Media Communities to Support Crisis Coordination
Hemant Purohit, Andrew J. Hampton, Shreyansh P. Bhatt, Valerie L. Shalin, Amit P. Sheth, John M. Flach |
Comput. Support. Cooperative Work. | 5 |
| 2014 | Comparative trust management with applications: Bayesian approaches emphasis
Krishnaprasad Thirunarayan, Pramod Anantharam, Cory A. Henson, Amit P. Sheth |
Future Gener. Comput. Syst. | 4 |
| 2014 | Semantics Driven Approach for Knowledge Acquisition From EMRsabstractSemantic computing technologies have matured to be applicable to many critical domains such as national security, life sciences, and health care. However, the key to their success is the availability of a rich domain knowledge base. The creation and refinement of domain knowledge bases pose difficult challenges. The existing knowledge bases in the health care domain are rich in taxonomic relationships, but they lack nontaxonomic (domain) relationships. In this paper, we describe a semiautomatic technique for enriching existing domain knowledge bases with causal relationships gleaned from Electronic Medical Records (EMR) data. We determine missing causal relationships between domain concepts by validating domain knowledge against EMR data sources and leveraging semantic-based techniques to derive plausible relationships that can rectify knowledge gaps. Our evaluation demonstrates that semantic techniques can be employed to improve the efficiency of knowledge acquisition. Sujan Perera, Cory A. Henson, Krishnaprasad Thirunarayan, Amit P. Sheth, Suhas Nair |
IEEE J. Biomed. Health Informatics | 4 |
| 2014 | A hybrid approach to finding relevant social media content for complex domain specific information needs
Delroy Cameron, Amit P. Sheth, Nishita Jaykumar, Krishnaprasad Thirunarayan, Gaurish Anand, Gary A. Smith |
J. Web Semant. | 2 |
| 2013 | Twitris v3: From Citizen Sensing to Analysis, Coordination and Action
Hemant Purohit, Amit P. Sheth |
ICWSM | 2 |
| 2013 | Automatic Domain Identification for Linked Open DataabstractLinked Open Data (LOD) has emerged as one of the largest collections of interlinked structured datasets on the Web. Although the adoption of such datasets for applications is increasing, identifying relevant datasets for a specific task or topic is still challenging. As an initial step to make such identification easier, we provide an approach to automatically identify the topic domains of given datasets. Our method utilizes existing knowledge sources, more specifically Freebase, and we present an evaluation which validates the topic domains we can identify with our system. Furthermore, we evaluate the effectiveness of identified topic domains for the purpose of finding relevant datasets, thus showing that our approach improves reusability of LOD datasets. Sarasi Lalithsena, Pascal Hitzler, Amit P. Sheth, Prateek Jain 0001 |
Web Intelligence | 3 |
| 2013 | Characterising Concepts of Interest Leveraging Linked Data and the Social WebabstractExtracting and representing user interests on the Social Web is becoming an essential part of the Web for personalisation and recommendations. Such personalisation is required in order to provide an adaptive Web to users, where content fits their preferences, background and current interests, making the Web more social and relevant. Current techniques analyse user activities on social media systems and collect structured or unstructured sets of entities representing users' interests. These sets of entities, or user profiles of interest, are often missing the semantics of the entities in terms of: (i) popularity and temporal dynamics of the interests on the Social Web and (ii) abstractness of the entities in the real world. State of the art techniques to compute these values are using specific knowledge bases or taxonomies and need to analyse the dynamics of the entities over a period of time. Hence, we propose a real-time, computationally inexpensive, domain independent model for concepts of interest composed of: popularity, temporal dynamics and specificity. We describe and evaluate a novel algorithm for computing specificity leveraging the semantics of Linked Data and evaluate the impact of our model on user profiles of interests. Fabrizio Orlandi, Pavan Kapanipathi, Amit P. Sheth, Alexandre Passant |
Web Intelligence | 3 |
| 2013 | A graph-based recovery and decomposition of Swanson's hypothesis using semantic predications
Delroy Cameron, Olivier Bodenreider, Hima Yalamanchili, Tu Danh, Sreeram Vallabhaneni, Krishnaprasad Thirunarayan, Amit P. Sheth, Thomas C. Rindflesch |
J. Biomed. Informatics | 7 |
| 2013 | PREDOSE: A semantic web platform for drug abuse epidemiology using social media
Delroy Cameron, Gary A. Smith, Raminta Daniulaityte, Amit P. Sheth, Drashti Dave, Gaurish Anand, Robert Carlson, Kera Z. B. Watkins, Russel Falck |
J. Biomed. Informatics | 4 |
| 2012 | PhylOnt: A domain-specific ontology for phylogeny analysisabstractPhylogenetic analyses can resolve historical relationships among genes, organisms or higher taxa. Understanding such relationships can elucidate a wide range of biological phenomena including the role of adaptation as a driver of diversification, the importance of gene and genome duplications in the evolution gene function, or the evolutionary consequences of biogeographic shifts. The variety of methods of analysis and data types typically employed in phylogenetic analyses can pose challenges for semantic reasoning due to significant representational and computational complexity. These challenges could be ameliorated with the development of an ontology designed to capture and organize the variety of concepts used to describe phylogenetic data, methods of analysis and the results of phylogenetic analyses. In this paper, we discuss the development of PhylOnt - an ontology for phylogenetic analyses, which establishes a foundation for semantics-based workflows including meta-analyses of phylogentic data and trees. PhylOnt is an extensible ontology, which describes the methods employed to estimate trees given a data matrix, models and programs used for phylogenetic analysis and descriptions of phylogenetic trees including branch-length information and support values. The relational vocabulary included in PhylOnt will facilitate the integration of heterogeneous data types derived from both structured and unstructured sources. To illustrate the utility of PhylOnt, we annotated scientific literature to support semantic search. The semantic annotations can subsequently support workflows that requiring the exchange and integration of heterogeneous phylogenetic information. Maryam Panahiazar, Ajith Ranabahu, Vahid Taslimitehrani, Hima Yalamanchili, Arlin Stoltzfus, Jim Leebens-Mack, Amit P. Sheth |
BIBM | 7 |
| 2012 | Data driven knowledge acquisition method for domain knowledge enrichment in the healthcareabstractSemantic computing technologies have matured to be applicable to many critical domains, such as life sciences and health care. However, the key to their success is the rich domain knowledge which consists of domain concepts and relationships, whose creation and refinement remains a challenge. In this paper, we develop a technique for enriching domain knowledge, focusing on populating the domain relationships. We determine missing relationships between the domain concepts by validating domain knowledge against real world data sources. We evaluate our approach in the healthcare domain using Electronic Medical Record(EMR) data, and demonstrate that semantic techniques can be used to semi-automate labour intensive tasks without sacrificing fidelity of domain knowledge. Sujan Perera, Cory A. Henson, Krishnaprasad Thirunarayan, Amit P. Sheth, Suhas Nair |
BIBM | 4 |
| 2012 | Extracting Diverse Sentiment Expressions with Target-Dependent Polarity from Twitter
Wenbo Wang 0002, Meena Nagarajan, Amit P. Sheth |
ICWSM | 5 |
| 2012 | Finding Influential Authors in Brand-Page Communities
Hemant Purohit, Jitendra Ajmera, Sachindra Joshi, Ashish Verma 0001, Amit P. Sheth |
ICWSM | 5 |
| 2012 | An Efficient Bit Vector Approach to Semantics-Based Machine Perception in Resource-Constrained Devices
Cory A. Henson, Krishnaprasad Thirunarayan, Amit P. Sheth |
ISWC (1) | 3 |
| 2012 | A new landscape for distributed and parallel data management
Amit P. Sheth |
Distributed Parallel Databases | 1 |
| 2012 | The SSN ontology of the W3C semantic sensor network incubator groupabstractThe W3C Semantic Sensor Network Incubator group (the SSN-XG) produced an OWL 2 ontology to describe sensors and observations — the SSN ontology, available at http://purl.oclc.org/NET/ssnx/ssn. The SSN ontology can describe sensors in terms of capabilities, measurement processes, observations and deployments. This article describes the SSN ontology. It further gives an example and describes the use of the ontology in recent research projects. Michael Compton, Payam M. Barnaghi, Luis Bermudez, Raúl García-Castro, Óscar Corcho, Simon J. D. Cox, John B. Graybeal, Manfred Hauswirth, Cory A. Henson, Arthur Herzog, Vincent Huang 0002, Krzysztof Janowicz, W. David Kelsey, Danh Le Phuoc, Laurent Lefort, Myriam Leggieri, Holger Neuhaus, Andriy Nikolov, Kevin R. Page, Alexandre Passant, Amit P. Sheth, Kerry L. Taylor |
J. Web Semant. | 21 |
| 2011 | Semantic Predications for Complex Information Needs in Biomedical LiteratureabstractMany complex information needs that arise in biomedical disciplines require exploring multiple documents in order to obtain information. While traditional information retrieval techniques that return a single ranked list of documents are quite common for such tasks, they may not always be adequate. The main issue is that ranked lists typically impose a significant burden on users to filter out irrelevant documents. Additionally, users must intuitively reformulate their search query when relevant documents have not been not highly ranked. Furthermore, even after interesting documents have been selected, very few mechanisms exist that enable document-to-document transitions. In this paper, we demonstrate the utility of assertions extracted from biomedical text (called semantic predications) to facilitate retrieving relevant documents for complex information needs. Our approach offers an alternative to query reformulation by establishing a framework for transitioning from one document to another. We evaluate this novel knowledge-driven approach using precision and recall metrics on the 2006 TREC Genomics Track. Delroy Cameron, Ramakanth Kavuluru, Olivier Bodenreider, Pablo N. Mendes, Amit P. Sheth, Krishnaprasad Thirunarayan |
BIBM | 5 |
| 2011 | Contextual Ontology Alignment of LOD with an Upper Ontology: A Case Study with Proton
Prateek Jain 0001, Peter Z. Yeh, Kunal Verma, Reymonrod G. Vasquez, Mariana Damova, Pascal Hitzler, Amit P. Sheth |
ESWC (1) | 7 |
| 2011 | Privacy-Aware and Scalable Content Dissemination in Distributed Social Networks
Pavan Kapanipathi, Julia Anaya, Amit P. Sheth, Brett Slatkin, Alexandre Passant |
ISWC (2) | 3 |
| 2011 | A unified framework for managing provenance information in translational researchabstractBACKGROUND: A critical aspect of the NIH Translational Research roadmap, which seeks to accelerate the delivery of "bench-side" discoveries to patient's "bedside," is the management of the provenance metadata that keeps track of the origin and history of data resources as they traverse the path from the bench to the bedside and back. A comprehensive provenance framework is essential for researchers to verify the quality of data, reproduce scientific results published in peer-reviewed literature, validate scientific process, and associate trust value with data and results. Traditional approaches to provenance management have focused on only partial sections of the translational research life cycle and they do not incorporate "domain semantics", which is essential to support domain-specific querying and analysis by scientists. RESULTS: We identify a common set of challenges in managing provenance information across the pre-publication and post-publication phases of data in the translational research lifecycle. We define the semantic provenance framework (SPF), underpinned by the Provenir upper-level provenance ontology, to address these challenges in the four stages of provenance metadata:(a) Provenance collection - during data generation(b) Provenance representation - to support interoperability, reasoning, and incorporate domain semantics(c) Provenance storage and propagation - to allow efficient storage and seamless propagation of provenance as the data is transferred across applications(d) Provenance query - to support queries with increasing complexity over large data size and also support knowledge discovery applicationsWe apply the SPF to two exemplar translational research projects, namely the Semantic Problem Solving Environment for Trypanosoma cruzi (T.cruzi SPSE) and the Biomedical Knowledge Repository (BKR) project, to demonstrate its effectiveness. CONCLUSIONS: The SPF provides a unified framework to effectively manage provenance of translational research data during pre and post-publication phases. This framework is underpinned by an upper-level provenance ontology called Provenir that is extended to create domain-specific provenance ontologies to facilitate provenance interoperability, seamless propagation of provenance, automated querying, and analysis. Satya Sanket Sahoo, Vinh Nguyen 0002, Olivier Bodenreider, Priti Parikh, Todd Minning, Amit P. Sheth |
BMC Bioinform. | 6 |
| 2010 | A Study in Hadoop Streaming with Matlab for NMR Data ProcessingabstractApplying Cloud computing techniques for analyzing large data sets has shown promise in many data-driven scientific applications. Our approach presented here is to use Cloud computing for Nuclear Magnetic Resonance (NMR)data analysis which normally consists of large amounts of data. Biologists often use third party or commercial software for ease of use. Enabling the capability to use this kind of software in a Cloud will be highly advantageous in many ways. Scripting languages especially designed for clouds may not have the flexibility biologists need for their purposes. Although this is true, they are familiar with special software packages that allow them to write complex calculations with minimum effort, but are often not compatible with a Cloud environment. Therefore, biologists who are trying to perform analysis on NMR data, acquire many advantages due to our proposed solution. Our solution gives them the flexibility to Cloud-enable their familiar software and it also enables them to perform calculations on a significant amount of data that was not previously possible. Our study is also applicable to any other environment in need of similar flexibility. We are currently in the initial stage of developing a framework for NMR data analysis. Kalpa Gunaratna, Ajith Ranabahu, Amit P. Sheth |
CloudCom | 4 |
| 2010 | Power of Clouds in Your Pocket: An Efficient Approach for Cloud Mobile Hybrid Application DevelopmentabstractThe advancements in computing have resulted in a boom of cheap, ubiquitous, connected mobile devices as well as seemingly unlimited, utility style, pay as you go computing resources, commonly referred to as Cloud computing. However, taking full advantage of this mobile and cloud computing landscape, especially for the data intensive domains has been hampered by the many heterogeneities that exist in the mobile space as well as the Cloud space. Our research focuses on exploiting the capabilities of the mobile and cloud landscape by defining a new class of applications called cloud mobile hybrid (CMH) applications and a Domain Specific Language (DSL) based methodology to develop these applications. We define Cloud-mobile hybrid as a collective application that has a Cloud based back-end and a mobile device front-end. Using a single DSL script, our toolkit is capable of generating a variety of CMH applications. These applications are composed of multiple combinations of native Cloud and mobile applications. Our approach not only reduces the learning curve but also shields developers from the complexities of the target platforms. We provide a detailed description of our language and present the results obtained using our prototype generator implementation. We also present a list of extensions that will enhance the various aspects of this platform. Ashwin Manjunatha, Ajith Ranabahu, Amit P. Sheth, Krishnaprasad Thirunarayan |
CloudCom | 3 |
| 2010 | Semantics Centric Solutions for Application and Data Portability in Cloud ComputingabstractCloud computing has become one of the key considerations both in academia and industry. Cheap, seemingly unlimited computing resources that can be allocated almost instantaneously and pay-as-you-go pricing schemes are some of the reasons for the success of Cloud computing. The Cloud computing landscape, however, is plagued by many issues hindering adoption. One such issue is vendor lock-in, forcing the Cloud users to adhere to one service provider in terms of data and application logic. Semantic Web has been an important research area that has seen significant attention from both academic and industrial researchers. One key property of Semantic Web is the notion of interoperability and portability through high level models. Significant work has been done in the areas of data modeling, matching, and transformations. The issues the Cloud computing community is facing now with respect to portability of data and application logic are exactly the same issue the Semantic Web community has been trying to address for some time. In this paper we present an outline of the use of well established semantic technologies to overcome the vendor lock-in issues in Cloud computing. We present a semantics-centric programming paradigm to create portable Cloud applications and discuss MobiCloud, our early attempt to implement the proposed approach. Ajith Ranabahu, Amit P. Sheth |
CloudCom | 2 |
| 2010 | A Qualitative Examination of Topical Tweet and Retweet Practices
Meena Nagarajan, Hemant Purohit, Amit P. Sheth |
ICWSM | 3 |
| 2010 | Ontology Alignment for Linked Open Data
Prateek Jain 0001, Pascal Hitzler, Amit P. Sheth, Kunal Verma, Peter Z. Yeh |
ISWC (1) | 3 |
| 2010 | Provenance Context Entity (PaCE): Scalable Provenance Tracking for Scientific RDF Data
Satya Sanket Sahoo, Olivier Bodenreider, Pascal Hitzler, Amit P. Sheth, Krishnaprasad Thirunarayan |
SSDBM | 4 |
| 2010 | Linked Open Social SignalsabstractIn this paper we discuss the collection, semantic annotation and analysis of real-time social signals from micro blogging data. We focus on users interested in analyzing social signals collectively for sense making. Our proposal enables flexibility in selecting subsets for analysis, alleviating information overload. We define an architecture that is based on state-of-the-art Semantic Web technologies and a distributed publish-subscribe protocol for real time communication. In addition, we discuss our method and application in a scenario related to the health care reform in the United States. Pablo N. Mendes, Alexandre Passant, Pavan Kapanipathi, Amit P. Sheth |
Web Intelligence | 4 |
| 2010 | Multimodal social intelligence in a real-time dashboard system
Daniel Gruhl, Meena Nagarajan, Jan Pieper, Christine Robson, Amit P. Sheth |
VLDB J. | 5 |
| 2009 | Context and Domain Knowledge Enhanced Entity Spotting in Informal Text
Daniel Gruhl, Meena Nagarajan, Jan Pieper, Christine Robson, Amit P. Sheth |
ISWC | 5 |
| 2009 | Monetizing User Activity on Social Networks - Challenges and ExperiencesabstractThis work summarizes challenges and experiences in monetizing user activity on public forums on social network sites. We present a approach that identifies the monetization potential of user posts and eliminates off-topic content to identify the most relevant and monetizable keywords for advertising. Preliminary studies using data from MySpace and Facebook show that 52% of ad impressions generated using keywords from our system were more targeted compared to the 30% relevant impressions generated without using our system. Meena Nagarajan, Kamal Baid, Amit P. Sheth |
Web Intelligence | 3 |
| 2009 | Spatio-Temporal-Thematic Analysis of Citizen Sensor Data: Challenges and Experiences
Meena Nagarajan, Karthik Gomadam, Amit P. Sheth, Ajith Ranabahu, Raghava Mutharaju, Ashutosh Jadhav |
WISE | 3 |
| 2008 | Unsupervised Discovery of Compound Entities for Relationship Extraction
Cartic Ramakrishnan, Pablo N. Mendes, Amit P. Sheth |
EKAW | 4 |
| 2008 | Relationship Web: Spinning the Web from Trailblazing to Semantic Analytics
Amit P. Sheth |
ER | 1 |
| 2008 | Graph Summaries for Subgraph Frequency Estimation
Angela Maduko, Kemafor Anyanwu, Amit P. Sheth, Paul Schliekelman |
ESWC | 3 |
| 2008 | A Faceted Classification Based Approach to Search and Rank Web APIsabstractWeb application hybrids, popularly known as mashups, are created by integrating services on the Web using their APIs. Support for finding an API is currently provided by generic search engines or domain specific solutions such as Google and ProgrammableWeb. Shortcomings of both these solutions in terms of and reliance on user tags make the task of identifying an API challenging. Since these APIs are described in HTML documents, it is essential to look beyond the boundaries of current approaches to Web service discovery that rely on formal descriptions. In this work, we present a faceted approach to searching and ranking Web APIs that takes into consideration attributes or facets of the APIs as found in their HTML descriptions. Our method adopts current research in document classification and faceted search and introduces the serviut score to rank APIs based on their utilization and popularity. We evaluate classification, search accuracy and ranking effectiveness using available APIs while contrasting our solution with existing ones. Karthik Gomadam, Ajith Ranabahu, Meena Nagarajan, Amit P. Sheth, Kunal Verma |
ICWS | 4 |
| 2008 | Joint Extraction of Compound Entities and Relationships from Biomedical LiteratureabstractIn this paper we identify some limitations of contemporary information extraction mechanisms in the context of biomedical literature. We present an extraction mechanism that generates structured representations of textual content. Our extraction mechanism achieves this by extracting compound entities, and relationships between them, occuring in text. A detailed evaluation of the relationship and compound entities extracted is presented. Our results show over 62% average precision across 8 relationship types tested with over 82% average precision for compound entity identification. Cartic Ramakrishnan, Pablo N. Mendes, Rodrigo A. T. S. da Gama, Guilherme C. N. Ferreira, Amit P. Sheth |
Web Intelligence | 5 |
| 2008 | Growing Fields of Interest - Using an Expand and Reduce Strategy for Domain Model ExtractionabstractDomain hierarchies are widely used as models underlying information retrieval tasks. Formal ontologies and taxonomies enrich such hierarchies further with properties and relationships but require manual effort; therefore they are costly to maintain, and often stale. Folksonomies and vocabularies lack rich category structure. Classification and extraction require the coverage of vocabularies and the alterability of folksonomies and can largely benefit from category relationships and other properties. With Doozer, a program for building conceptual models of information domains, we want to bridge the gap between the vocabularies and Folksonomies on the one side and the rich, expert-designed ontologies and taxonomies on the other. Doozer mines Wikipedia to produce tight domain hierarchies, starting with simple domain descriptions. It also adds relevancy scores for use in automated classification of information. The output model is described as a hierarchy of domain terms that can be used immediately for classifiers and IR systems or as a basis for manual or semi-automatic creation of formal ontologies. Christopher Thomas 0001, Pankaj Mehra, Roger Brooks, Amit P. Sheth |
Web Intelligence | 4 |
| 2008 | WS3: international workshop on context-enabled source and service selection, integration and adaptation (CSSSIA 2008)abstractThis write-up provides a summary of the International Workshop on Context enabled Source and Service Selection, Integration and Adaptation (CSSSIA 2008), organized in conjunction with WWW 2008, at Beijing, China on April 22nd 2008. We outline the motivation for organizing the workshop, briefly describe the organizational details and program of the workshop, and summarize each of the papers accepted by the workshop. More information about the workshop can be found at http://www.cs.adelaide.edu.au/~csssia08/. Quan Z. Sheng, Ullas Nambiar, Amit P. Sheth, Biplav Srivastava, Zakaria Maamar, Said Elnaffar |
WWW | 3 |
| 2008 | Business process management
Schahram Dustdar, José Luiz Fiadeiro, Amit P. Sheth |
Data Knowl. Eng. | 3 |
| 2008 | An ontology-driven semantic mashup of gene and biological pathway information: Application to the domain of nicotine dependence
Satya Sanket Sahoo, Olivier Bodenreider, Joni L. Rutter, Karen J. Skinner, Amit P. Sheth |
J. Biomed. Informatics | 5 |
| 2008 | Scalable semantic analytics on social networks for addressing the problem of conflict of interest detectionabstractIn this article, we demonstrate the applicability of semantic techniques for detection of Conflict of Interest (COI). We explain the common challenges involved in building scalable Semantic Web applications, in particular those addressing connecting-the-dots problems. We describe in detail the challenges involved in two important aspects on building Semantic Web applications, namely, data acquisition and entity disambiguation (or reference reconciliation). We extend upon our previous work where we integrated the collaborative network of a subset of DBLP researchers with persons in a Friend-of-a-Friend social network (FOAF). Our method finds the connections between people, measures collaboration strength, and includes heuristics that use friendship/affiliation information to provide an estimate of potential COI in a peer-review scenario. Evaluations are presented by measuring what could have been the COI between accepted papers in various conference tracks and their respective program committee members. The experimental results demonstrate that scalability can be achieved by using a dataset of over 3 million entities (all bibliographic data from DBLP and a large collection of FOAF documents). Boanerges Aleman-Meza, Meena Nagarajan, Li Ding 0001, Amit P. Sheth, Ismailcem Budak Arpinar, Anupam Joshi, Tim Finin |
ACM Trans. Web | 4 |
| 2007 | A Semantic Framework for Identifying Events in a Service Oriented ArchitectureabstractWe propose a semantic framework for automatically identifying events as a step towards developing an adaptive middleware for Service Oriented Architecture (SOA). Current related research focuses on adapting to events that violate certain non-functional objectives of the service requestor. Given the large of number of events that can happen during the execution of a service, identifying events that can impact the non-functional objectives of a service request is a key challenge. To address this problem we propose an approach that allows service requestors to create semantically rich service requirement descriptions, called semantic templates. We propose a formal model for expressing semantic templates and for measuring the relevance of an event to both the action being performed and the nonfunctional objectives. This model is extended to adjust the relevance of the events based on feedback from the underlying adaptation framework. We present an algorithm that utilizes multiple ontologies for identifying relevant events and present our evaluations that measure the efficiency of both the event identification and the subsequent adaptation scheme. Karthik Gomadam, Ajith Ranabahu, Lakshmish Ramaswamy, Amit P. Sheth, Kunal Verma |
ICWS | 4 |
| 2007 | Visualization of Events in a Spatially and Multimedia Enriched Virtual EnvironmentabstractSemantic Event Tracker (SET) is a highly interactive visualization tool for tracking and associating activities (events) in a spatially and Multimedia Enriched Virtual Environment. SET provides integrated views of information spaces while providing overview and detail to improve perception and evaluation of complex scenarios. We model an event as an object that describes an action and its location, time, and relations to other objects. Real world event information is extracted from Internet sources, then stored and processed using Semantic Web technologies that enable us to discover semantic associations between events. We use RDF graphs to represent semantic metadata and ontologies. SET is capable of visualizing as well as navigating through the event data in all three aspects of space, time and theme. Leonidas Deligiannidis, Farshad Hakimpour, Amit P. Sheth |
ISI | 3 |
| 2007 | Semantic Convergence of Wikipedia ArticlesabstractSocial networking, distributed problem solving and human computation have gained high visibility. Wikipedia is a well established service that incorporates aspects of these three fields of research. For this reason it is a good object of study for determining quality of solutions in a social setting that is open, completely distributed, bottom up and not peer reviewed by certified experts. In particular, this paper aims at identifying semantic convergence of Wikipedia articles; the notion that the content of an article stays stable regardless of continuing edits. This could lead to an automatic recommendation of good article tags but also add to the usability of Wikipedia as a Web Service and to its reliability for information extraction. The methods used and the results obtained in this research can be generalized to other communities that iteratively produce textual content. Christopher Thomas 0001, Amit P. Sheth |
Web Intelligence | 2 |
| 2007 | SPARQ2L: towards support for subgraph extraction queries in rdf databasesabstractMany applications in analytical domains often have the need to "connect the dots" i.e., query about the structure of data. In bioinformatics for example, it is typical to want to query about interactions between proteins. The aim of such queries is to "extract" relationships between entities i.e. paths from a data graph. Often, such queries will specify certain constraints that qualifying results must satisfy e.g. paths involving a set of mandatory nodes. Unfortunately, most present day Semantic Web query languages including the current draft of the anticipated recommendation SPARQL, lack the ability to express queries about arbitrary path structures in data. In addition, many systems that support some limited form of path queries rely on main memory graph algorithms limiting their applicability to very large scale graphs. In this paper, we present an approach for supporting Path Extraction queries. Our proposal comprises (i) a query language SPARQ2L which extends SPARQL with path variables and path variable constraint expressions, and (ii) a novel query evaluation framework based on efficient algebraic techniques for solving path problems which allows for path queries to be efficiently evaluated on disk resident RDF graphs. The effectiveness of our proposal is demonstrated by a performance evaluation of our approach on both real world based and synthetic dataset. Kemafor Anyanwu, Angela Maduko, Amit P. Sheth |
WWW | 3 |
| 2007 | Estimating the cardinality of RDF graph patternsabstractMost RDF query languages allow for graph structure search through a conjunction of triples which is typically processed using join operations. A key factor in optimizing joins is determining the join order which depends on the expected cardinality of intermediate results. This work proposes a pattern-based summarization framework for estimating the cardinality of RDF graph patterns. We present experiments on real world and synthetic datasets which confirm the feasibility of our approach. Angela Maduko, Kemafor Anyanwu, Amit P. Sheth, Paul Schliekelman |
WWW | 3 |
| 2007 | Altering document term vectors for classification: ontologies as expectations of co-occurrenceabstractIn this paper we extend the state-of-the-art in utilizing background knowledge for supervised classification by exploiting the semantic relationships between terms explicated in Ontologies. Preliminary evaluations indicate that the new approach generally improves precision and recall, more so for hard to classify cases and reveals patterns indicating the usefulness of such background knowledge. Meena Nagarajan, Amit P. Sheth, Marcos K. Aguilera, Kimberly Keeton, Arif Merchant, Mustafa Uysal |
WWW | 2 |
| 2007 | Welcome to Prof. Amit Sheth
Ahmed K. Elmagarmid, Amit P. Sheth |
Distributed Parallel Databases | 2 |
| 2007 | SwetoDblp ontology of Computer Science publications
Boanerges Aleman-Meza, Farshad Hakimpour, Ismailcem Budak Arpinar, Amit P. Sheth |
J. Web Semant. | 4 |
| 2006 | Modular Ontology Design Using Canonical Building Blocks in the Biochemistry Domain
Christopher J. Thomas, Amit P. Sheth, William S. York |
FOIS | 2 |
| 2006 | Analyzing theme, space, and time: an ontology-based approachabstractThe W3C's Semantic Web Activity is illustrating the use of semantics for information integration, search, and analysis. However, the majority of the work in this community has focused more on the thematic aspects of information and has paid less attention to its spatial and temporal dimensions. In this paper, we present an integrative ontology-based framework incorporating the thematic, spatial, and temporal dimensions of information. This framework is built around the RDF metadata model. Our ultimate goal is to provide an information system which allows searching and analysis of relationships in any or all of the three dimensions of space, time, and theme. Toward this end, we present an upper-level ontology combining concepts and relationships from both the thematic and spatial dimensions and show how to incorporate temporal semantics into this ontology. We also introduce the notion of a thematic context linking entities of differing dimensions and define a set of query operators built upon these contexts. Matthew Perry, Farshad Hakimpour, Amit P. Sheth |
GIS | 3 |
| 2006 | Semantic Interoperability of Web Services - Challenges and ExperiencesabstractWith the rising popularity of Web services, both academia and industry have invested considerably in Web service description standards, discovery, and composition techniques. The standards based approach utilized by Web services has supported interoperability at the syntax level. However, issues of structural and semantic heterogeneity between messages exchanged by Web services are far more complex and crucial to interoperability. It is for these reasons that we recognize the value that schema/data mappings bring to Web service descriptions. In this paper, we examine challenges to interoperability; classify the types of heterogeneities that can occur between interacting services and present a possible solution for data mediation using the mapping support provided by WSDL-S, the extensibility features of WSDL and the popular SOAP engine, Axis 2 Meena Nagarajan, Kunal Verma, Amit P. Sheth, John A. Miller 0001, Jon Lathem |
ICWS | 3 |
| 2006 | Optimal Adaptation in Web Processes with Coordination ConstraintsabstractWe present methods for optimally adapting Web processes to exogenous events while preserving inter-service constraints that necessitate coordination. For example, in a supply chain process, orders placed by a manufacturer may get delayed in arriving. In response to this event, the manufacturer has the choice of either waiting out the delay or changing the supplier. Additionally, there may be compatibility constraints between the different orders, thereby introducing the problem of coordination between them if the manufacturer chooses to change the suppliers. We focus on formulating the decision making models of the managers, who must adapt to external events while satisfying the coordination constraints, using Markov decision processes. Our methods range from being centralized and globally optimal in their adaptation but not scalable, to decentralized that is suboptimal but scalable to multiple managers. We also develop a hybrid approach that improves on the performance of the decentralized approach with a minimal loss of optimality Kunal Verma, Prashant Doshi, Karthik Gomadam, John A. Miller 0001, Amit P. Sheth |
ICWS | 5 |
| 2006 | Semantic Analytics Visualization
Leonidas Deligiannidis, Amit P. Sheth, Boanerges Aleman-Meza |
ISI | 2 |
| 2006 | SemanticSpy: Suspect Tracking Using Semantic Data in a Multimedia Environment
Amit Mathew, Amit P. Sheth, Leonidas Deligiannidis |
ISI | 2 |
| 2006 | A Framework for Schema-Driven Relationship Discovery from Unstructured Text
Cartic Ramakrishnan, Krys J. Kochut, Amit P. Sheth |
ISWC | 3 |
| 2006 | Active Semantic Electronic Medical Record
Amit P. Sheth, Subodh Agrawal, Jon Lathem, Nicole Oldham, Harry Wingate, Prem Yadav, Kelly Gallagher |
ISWC | 1 |
| 2006 | Semantic analytics on social networks: experiences in addressing the problem of conflict of interest detectionabstractIn this paper, we describe a Semantic Web application that detects Conflict of Interest (COI) relationships among potential reviewers and authors of scientific papers. This application discovers various 'semantic associations' between the reviewers and authors in a populated ontology to determine a degree of Conflict of Interest. This ontology was created by integrating entities and relationships from two social networks, namely "knows," from a FOAF (Friend-of-a-Friend) social network and "co-author," from the underlying co-authorship network of the DBLP bibliography. We describe our experiences developing this application in the context of a class of Semantic Web applications, which have important research and engineering challenges in common. In addition, we present an evaluation of our approach for real-life COI detection. Boanerges Aleman-Meza, Meena Nagarajan, Cartic Ramakrishnan, Li Ding 0001, Pranam Kolari, Amit P. Sheth, Ismailcem Budak Arpinar, Anupam Joshi, Tim Finin |
WWW | 6 |
| 2006 | Semantic WS-agreement partner selectionabstract(Under the direction of Amit P. Sheth) In a dynamic service oriented environment it is desirable for service consumers and providers to offer and obtain guarantees regarding their capabilities and requirements. WS-Agreement defines a language and protocol for establishing agreements between two parties. The agreements are complex and expressive to the extent that the manual matching of these agreements would be expensive both in time and resources. It is essential to develop a method for matching agreements automatically. This work presents the framework and implementation of an innovative tool for the matching providers and consumers based on WS-Agreements. The approach utilizes Semantic Web technologies to achieve rich and accurate matches. A key feature is the novel and flexible approach for achieving user personalized matches. Nicole Oldham, Kunal Verma, Amit P. Sheth, Farshad Hakimpour |
WWW | 3 |
| 2006 | Knowledge modeling and its application in life sciences: a tale of two ontologiesabstractHigh throughput glycoproteomics, similar to genomics and proteomics, involves extremely large volumes of distributed, heterogeneous data as a basis for identification and quantification of a structurally diverse collection of biomolecules. The ability to share, compare, query for and most critically correlate datasets using the native biological relationships are some of the challenges being faced by glycobiology researchers. As a solution for these challenges, we are building a semantic structure, using a suite of ontologies, which supports management of data and information at each step of the experimental lifecycle. This framework will enable researchers to leverage the large scale of glycoproteomics data to their benefit.In this paper, we focus on the design of these biological ontology schemas with an emphasis on relationships between biological concepts, on the use of novel approaches to populate these complex ontologies including integrating extremely large datasets ( 500MB) as part of the instance base and on the evaluation of ontologies using OntoQA [38] metrics. The application of these ontologies in providing informatics solutions, for high throughput glycoproteomics experimental domain, is also discussed. We present our experience as a use case of developing two ontologies in one domain, to be part of a set of use cases, which are used in the development of an emergent framework for building and deploying biological ontologies. Satya Sanket Sahoo, Christopher Thomas 0001, Amit P. Sheth, William S. York, Samir Tartir |
WWW | 3 |
| 2005 | Demonstrating Dynamic Configuration and Execution of Web Processes
Karthik Gomadam, Kunal Verma, Amit P. Sheth, John A. Miller 0001 |
ICSOC | 3 |
| 2005 | OpenWS-Transaction: Enabling Reliable Web Service Transactions
Ivan Vasquez, John A. Miller 0001, Kunal Verma, Amit P. Sheth |
ICSOC | 4 |
| 2005 | Autonomic Web Processes
Kunal Verma, Amit P. Sheth |
ICSOC | 2 |
| 2005 | A Semantic Template Based Designer for Web ProcessesabstractThe growing popularity of service oriented computing based on Web services standards is creating a need for paradigms to represent and design business processes. Significant work has been done in the representation aspects with regards to WSBPEL. However, design and modeling of business processes is still an open issue. In this paper, we present a novel designer for business processes, which allows for intuitive modeling of Web processes, as well as using a template based approach for semi-automatically integrating partners either at design time or at deployment time. This work has been done as part of the METEOR-S project, which concentrates on adding semantics to the entire Web process lifecycle. Ranjit Mulye, John A. Miller 0001, Kunal Verma, Karthik Gomadam, Amit P. Sheth |
ICWS | 5 |
| 2005 | An Ontological Approach to the Document Access Problem of Insider Threat
Boanerges Aleman-Meza, Phillip Burns, Matthew Eavenson, Devanand Palaniswami, Amit P. Sheth |
ISI | 5 |
| 2005 | Template Based Semantic Similarity for Security Applications
Boanerges Aleman-Meza, Christian Halaschek-Wiener, Satya Sanket Sahoo, Amit P. Sheth, Ismailcem Budak Arpinar |
ISI | 4 |
| 2005 | SemRank: ranking complex relationship search results on the semantic webabstractWhile the idea that querying mechanisms for complex relationships (otherwise known as Semantic Associations) should be integral to Semantic Web search technologies has recently gained some ground, the issue of how search results will be ranked remains largely unaddressed. Since it is expected that the number of relationships between entities in a knowledge base will be much larger than the number of entities themselves, the likelihood that Semantic Association searches would result in an overwhelming number of results for users is increased, therefore elevating the need for appropriate ranking schemes. Furthermore, it is unlikely that ranking schemes for ranking entities (documents, resources, etc.) may be applied to complex structures such as Semantic Associations.In this paper, we present an approach that ranks results based on how predictable a result might be for users. It is based on a relevance model SemRank, which is a rich blend of semantic and information-theoretic techniques with heuristics that supports the novel idea of modulative searches, where users may vary their search modes to effect changes in the ordering of results depending on their need. We also present the infrastructure used in the SSARK system to support the computation of SemRank values for resulting Semantic Associations and their ordering. Kemafor Anyanwu, Angela Maduko, Amit P. Sheth |
WWW | 3 |
| 2005 | Semantics for the Semantic Web: The Implicit, the Formal and the PowerfulabstractEnabling applications that exploit heterogeneous data in the Semantic Web will require us to harness a broad variety of semantics. Considering the role of semantics in a number of research areas in computer science, we organize semantics in three forms — implicit, formal, and powerful — and explore their roles in enabling some of the key capabilities related to the Semantic Web. The central message of this article is that building the Semantic Web purely on description logics will artificially limit its potential, and that we will need to both exploit well-known techniques that support implicit semantics, and develop more powerful semantic techniques. Amit P. Sheth, Cartic Ramakrishnan, Christopher Thomas 0001 |
Int. J. Semantic Web Inf. Syst. | 1 |
| 2005 | Semantic Association Identification and Knowledge Discovery for National Security ApplicationsabstractPublic and private organizations have access to a vast amount of internal, deep Web and open Web information. Transforming this heterogeneous and distributed information into actionable and insightful information is the key to the emerging new classes of business intelligence and national security applications. Although the role of semantics in search and integration has been often talked about, in this paper we discuss semantic approaches to support analytics on vast amounts of heterogeneous data. In particular, we bring together novel academic research and commercialized Semantic Web technology. The academic research related to semantic association identification is built upon commercial Semantic Web technology for semantic metadata extraction. A prototypical demonstration of this research and technology is presented in the context of an aviation security application of significance to national security. Amit P. Sheth, Boanerges Aleman-Meza, Ismailcem Budak Arpinar, Clemens Bertram, Yashodhan S. Warke, Cartic Ramakrishnan, Christian Halaschek-Wiener, Kemafor Anyanwu, David Avant, Fatma Sena Arpinar, Krys J. Kochut |
J. Database Manag. | 1 |
| 2004 | Services Oriented Architecture and Semantic Web Processes
Francisco Curbera, Amit P. Sheth, Kunal Verma |
ICWS | 2 |
| 2004 | Discovery of Web Services in a Federated Registry EnvironmentabstractThe potential of a large scale growth of private and semi-private registries is creating the need for an infrastructure which can support discovery and publication over a group of autonomous registries. Recent versions of UDDI have made changes to accommodate interactions between distributed registries. In this paper, we discuss METEOR-S Web service Discovery Infrastructure, which provides an ontology-based infrastructure to access a group of registries that are divided based on business domains and grouped into federations. We also discuss how Web service discovery is carried out within a federation. Kaarthik Sivashanmugam, Kunal Verma, Amit P. Sheth |
ICWS | 3 |
| 2004 | Discovering and Ranking Semantic Associations over a Large RDF Metabase
Christian Halaschek-Wiener, Boanerges Aleman-Meza, Ismailcem Budak Arpinar, Amit P. Sheth |
VLDB | 4 |
| 2004 | Meteor-s web service annotation frameworkabstractThe World Wide Web is emerging not only as an infrastructure for data, but also for a broader variety of resources that are increasingly being made available as Web services. Relevant current standards like UDDI, WSDL, and SOAP are in their fledgling years and form the basis of making Web services a workable and broadly adopted technology. However, realizing the fuller scope of the promise of Web services and associated service oriented architecture will requite further technological advances in the areas of service interoperation, service discovery, service composition, and process orchestration. Semantics, especially as supported by the use of ontologies, and related Semantic Web technologies, are likely to provide better qualitative and scalable solutions to these requirements. Just as semantic annotation of data in the Semantic Web is the first critical step to better search, integration and analytics over heterogeneous data, semantic annotation of Web services is an equally critical first step to achieving the above promise. Our approach is to work with existing Web services technologies and combine them with ideas from the Semantic Web to create a better framework for Web service discovery and composition. In this paper we present MWSAF (METEOR-S Web Service Annotation Framework), a framework for semi-automatically marking up Web service descriptions with ontologies. We have developed algorithms to match and annotate WSDL files with relevant ontologies. We use domain ontologies to categorize Web services into domains. An empirical study of our approach is presented to help evaluate its performance. Abhijit A. Patil, Swapna A. Oundhakar, Amit P. Sheth, Kunal Verma |
WWW | 3 |
| 2004 | Quality of service for workflows and web service processes
Jorge Cardoso 0001, Amit P. Sheth, John A. Miller 0001, Jonathan P. Arnold, Krys J. Kochut |
J. Web Semant. | 2 |
| 2003 | Adding Semantics to Web Services Standards
Kaarthik Sivashanmugam, Kunal Verma, Amit P. Sheth, John A. Miller 0001 |
ICWS | 3 |
| 2003 | Semantic Web Processes: Semantics Enabled Annotation, Discovery, Composition and Orchestration of Web Scale ProcessesabstractThis paper deals with the evolution of inter-enterprise and Web scale process to support e-commerce and e-services. It taps into the promises of two of the hottest R&D and technology areas: Web services and the semantic Web. It presents how applying semantics to each of the steps in the semantic Web process lifecycle can help address critical issues in reuse, integration and scalability. Jorge Cardoso 0001, Amit P. Sheth |
WISE | 2 |
| 2003 | ?-Queries: enabling querying for semantic associations on the semantic webabstractThis paper presents the notion of Semantic Associations as complex relationships between resource entities. These relationships capture both a connectivity of entities as well as similarity of entities based on a specific notion of similarity called r-isomorphism. It formalizes these notions for the RDF data model, by introducing a notion of a Property Sequence as a type. In the context of a graph model such as that for RDF, Semantic Associations amount to specific certain graph signatures. Specifically, they refer to sequences (i.e. directed paths) here called Property Sequences, between entities, networks of Property Sequences (i.e. undirected paths), or subgraphs of r-isomorphic Property Sequences.The ability to query about the existence of such relationships is fundamental to tasks in analytical domains such as national security and business intelligence, where tasks often focus on finding complex yet meaningful and obscured relationships between entities. However, support for such queries is lacking in contemporary query systems, including those for RDF. Kemafor Anyanwu, Amit P. Sheth |
WWW | 2 |
| 2003 | IntelliGEN: A Distributed Workflow System for Discovering Protein-Protein Interactions
Krys J. Kochut, Jonathan P. Arnold, Amit P. Sheth, John A. Miller 0001, Eileen T. Kraemer, Ismailcem Budak Arpinar, Jorge Cardoso 0001 |
Distributed Parallel Databases | 3 |
| 2003 | Exception Handling for Conflict Resolution in Cross-Organizational Workflows
Zongwei Luo, Amit P. Sheth, Krys J. Kochut, Ismailcem Budak Arpinar |
Distributed Parallel Databases | 2 |
| 2003 | Semantic E-Workflow Composition
Jorge Cardoso 0001, Amit P. Sheth |
J. Intell. Inf. Syst. | 2 |
| 2003 | Complex relationships and knowledge discovery support in the InfoQuilt system
Amit P. Sheth, Sanjeev Thacker, Shuchi Patel |
VLDB J. | 1 |
| 2002 | Semantic technology applications for homeland securityabstractSemantic Content Organization and Retrieval Engine (SCORE) is among the earliest commercialized Semantic Web technologies. Based on supporting and exploiting domain specific ontologies, it offers advanced capability in heterogeneous content processing analysis, and integration at a higher semantic level-- rather than merely syntactical and structural level approaches based on XML and RDF. These capabilities are now being demonstrated in addressing requirements of very demanding Homeland Security and National Security applications. This paper briefly describes two of them. David Avant, M. Baum, Clemens Bertram, M. Fisher, Amit P. Sheth, Yashodhan S. Warke |
CIKM | 5 |
| 2002 | Authorization and Access Control of Application Data in Workflow Systems
Shengli Wu 0001, Amit P. Sheth, John A. Miller 0001, Zongwei Luo |
J. Intell. Inf. Syst. | 2 |
| 2001 | Planning and Optimizing Semantic Information Requests Using Domain Modeling and Resource Characteristics
Shuchi Patel, Amit P. Sheth |
CoopIS | 2 |
| 2000 | Exception Handling in Workflow Systems
Zongwei Luo, Amit P. Sheth, Krys J. Kochut, John A. Miller 0001 |
Appl. Intell. | 2 |
| 2000 | OBSERVER: An Approach for Query Processing in Global Information Systems Based on Interoperation Across Pre-Existing Ontologies
Eduardo Mena, Arantza Illarramendi, Vipul Kashyap, Amit P. Sheth |
Distributed Parallel Databases | 4 |
| 2000 | Imprecise Answers in Distributed Environments: Estimation of Information Loss for Multi-Ontology Based Query ProcessingabstractThe World Wide Web is fast becoming a ubiquitous computing environment. Prevalent keyword-based search techniques are scalable, but are incapable of accessing information based on concepts. We investigate the use of concepts from multiple, real-world pre-existing, domain ontologies to describe the underlying data content and support information access at a higher level of abstraction. It is not practical to have a single domain ontology to describe the vast amounts of data on the Web. In fact, we expect multiple ontologies to be used as different world views and present an approach to "browse" ontologies as a paradigm for information access. A critical challenge in this approach is the vocabulary heterogeneity problem. Queries are rewritten using interontology relationships to obtain translations across ontologies. However, some translations may not be semantics preserving, leading to uncertainty or loss in the information retrieved. We present a novel approach for estimating loss of information based on the navigation of ontological terms. We define measures for loss of information based on intensional information as well as on well established metrics like precision and recall based on extensional information. These measures are used to select results having the desired quality of information. Eduardo Mena, Vipul Kashyap, Arantza Illarramendi, Amit P. Sheth |
Int. J. Cooperative Inf. Syst. | 4 |
| 1999 | A Multilevel Secure Workflow Management System
Myong H. Kang, Judith N. Froscher, Amit P. Sheth, Krys J. Kochut, John A. Miller 0001 |
CAiSE | 3 |
| 1998 | ZEBRA Image Access SystemabstractThe ZEBRA system, which is part of the VisualHarness platform for managing heterogeneous data, supports three types of access to distributed image repositories: keyword based, attribute based, and image content based. A user can assign different weights (relative importance) to each of the three types, and within the last type of access, to each of the image properties. The image based access component (IBAC) supports access based on computable image properties such as those based on spatial domain, frequency domain or statistical and structural analysis. However, it uses a novel black box approach of utilizing a Visual Information Retrieval (VIR) engine to compute corresponding metadata that is then independently managed in a relational database to provide query processing involving image features and information correlation. That is, one overcomes the difficulties in using the feature vectors that are proprietary to a VTR engine, as one does not require any knowledge of the internal representation or format of the image feature used by a VIR engine. Srilekha Mudumbai, Kshitij Shah, Amit P. Sheth, Krishnan Parasuraman, Clemens Bertram |
ICDE | 3 |
| 1998 | WebWork: METEOR2's Web-Based Workflow Management System
John A. Miller 0001, Devanand Palaniswami, Amit P. Sheth, Krys J. Kochut |
J. Intell. Inf. Syst. | 3 |
| 1997 | Perspectives in Modeling: Simulation, Database, and Workflow
John A. Miller 0001, Amit P. Sheth, Krys J. Kochut |
Conceptual Modeling | 2 |
| 1997 | The Carnot Heterogeneous Database Project: Implemented Applications
Munindar P. Singh, Philip Cannata, Michael N. Huhns, Nigel Jacobs, Tomasz Ksiezyk, KayLiang Ong, Amit P. Sheth, Christine Tomlinson, Darrell Woelk |
Distributed Parallel Databases | 7 |
| 1996 | OBSERVER: An Approach for Query Processing in Global Information Systems based on Interoperation across Pre-existing OntologiesabstractThe huge number of autonomous and heterogeneous data repositories accessible on the "global information infrastructure" makes it impossible for users to be aware of the locations structure/organization, query languages and semantics of the data in various repositories. There is a critical need to complement current browsing, navigational and information retrieval techniques with a strategy that focuses on information content and semantics. In any strategy that focuses on information content, the most critical problem is that of different vocabularies used to describe similar information across domains. We discuss a scalable approach for vocabulary sharing. The objects in the repositories are represented as intensional descriptions by preexisting ontologies expressed in Description Logics characterizing information in different domains. User queries are rewritten by using interontology relationships to obtain semantics preserving translations across the ontologies. Eduardo Mena, Vipul Kashyap, Amit P. Sheth, Arantza Illarramendi |
CoopIS | 3 |
| 1996 | What's in a WWW Link? - Panel
Amit P. Sheth, Robert Meersman, Erich J. Neuhold, Calton Pu, V. S. Subrahmanian |
ICDE | 1 |
| 1996 | Bellcore's ADAPT/X Harness System for Managing Information on Internet and Intranets
Amit P. Sheth |
VLDB | 1 |
| 1996 | Supporting State-Wide Immunisation Tracking Using Multi-Paradigm Workflow Technology
Amit P. Sheth, Krys J. Kochut, John A. Miller 0001, Devashish Worah, Chenye Lin, Devanand Palaniswami, John Lynch, Ivan Shevchenko |
VLDB | 1 |
| 1996 | Bounding the Effects of Compensation under Relaxed Multi-level Serializability
Piotr Krychniak, Marek Rusinkiewicz, Andrzej Cichocki, Amit P. Sheth, Gomer Thomas |
Distributed Parallel Databases | 4 |
| 1996 | Semantic and Schematic Similarities Between Database Objects: A Context-Based Approach
Vipul Kashyap, Amit P. Sheth |
VLDB J. | 2 |
| 1995 | InfoHarness: Use of Automatically Generated Metadata for Search and Retrieval of Heterogeneous Information
Leon A. Shklar, Amit P. Sheth, Vipul Kashyap, Kshitij Shah |
CAiSE | 2 |
| 1995 | Workflow Automation: Applications, Technology, and Research (Tutorial)abstractWith increasing global exposure, today's enterprises must react quickly to changes, rapidly develop new services and products, and at the same time improve productivity and quality and reduce cost. Business process re-engineering and workflow automation to coordinate activities throughout the enterprise are recognized as important emerging technologies to support these requirements. Rosy estimates of a multi-billion dollar marketplace for workflow software has resulted in significant commercial activities in the area, with nearly hundred products now claiming to support workflow automation. While many help to automate document- and image-driven office applications, therefore helping to improve the productivity of small groups, most current products fail to support:• mission critical and enterprise-wide applications with requirements such as failure handling and recovery, and• interoperability with existing heterogeneous information systems.In this tutorial, we will discuss requirements for applications involving workflow automation, present an overview of the current state-of-the-art in products, and present some of the research efforts that are attempting to respond to unmet challenges. Amit P. Sheth |
SIGMOD Conference | 1 |
| 1995 | InfoHarness: A System for Search and Retrieval of Heterogeneous InformationabstractEnormous amounts of heterogeneous information have been accumulated within corporations, government organizations and universities. It is becoming increasingly easier to create new information, but the knowledge about the existence, location, and means of retrieval of information, have become so confusing as to give rise to the phenomenon of write-only databases. Leon A. Shklar, Amit P. Sheth, Vipul Kashyap, Satish Thatte |
SIGMOD Conference | 2 |
| 1995 | An Overview of Workflow Management: From Process Modeling to Workflow Automation Infrastructure
Dimitrios Georgakopoulos 0001, Mark F. Hornick, Amit P. Sheth |
Distributed Parallel Databases | 3 |
| 1995 | Managing Hetergeneous Multi-system Tasks to Support Enterprise-Wide Operations
Narayanan Krishnakumar, Amit P. Sheth |
Distributed Parallel Databases | 2 |
| 1994 | Semantics-Based Information BrokeringabstractThe rapid advances in computer and communication technologies, and their merger, is leading to a global information market place. It will consist of federations of very large number of information systems that will cooperate to varying extents to support the users' information needs. We discuss an approach to information brokering in the above environment. We discuss two of its tasks: information resource discovery, which identifies relevant information sources for a given query, and query processing, which involves the generation of appropriate mapping from relevant but structurally heterogeneous objects. Query processing consists of information focusing and information correlation. Vipul Kashyap, Amit P. Sheth |
CIKM | 2 |
| 1994 | Transactional Workflows: Research, Enabling Technologies, and Applications (Abstract)abstractAbstract only given, as follows. Need to increase productivity and reduce cost have lead to reengineering and automation of operations across corporations. Some of the applications involve imaging, document processing and routing. These tasks can be effectively automated using current genre of workflow automation products. Some other applications involve tasks that can be modeled in client-server style using traditional transactions. These can be supported by distributed transaction processing/monitoring systems. Finally, there is an important class of more complicated applications that involve heterogeneous but automated tasks, with varying levels of transaction properties, and performed at heterogeneous systems. We look at three aspects of the emerging technology of transactional workflow management that aims to support such applications: Identify properties of a class of multi-system applications and the environments that can be supported by transactional workflow systems: Discuss how a transactional workflow management system is different from, but "builds upon" the current transaction processing and workflow automation technologies: Discuss some of the relevant database research as well as software system and application prototyping experiences, especially those related to the extended/relaxed transaction models. Much of the discussion is based on our study of some real (mostly telecommunications) applications, and research and prototyping done at Bellcore in collaboration with U. of Houston and MCC's Carnot project.> Amit P. Sheth |
ICDE | 1 |
| 1994 | Using Tickets to Enforce the Serializability of Multidatabase TransactionsabstractTo enforce global serializability in a multidatabase environment the multidatabase transaction manager must take into account the indirect (transitive) conflicts between multidatabase transactions caused by local transactions. Such conflicts are difficult to resolve because the behavior or even the existence of local transactions is not known to the multidatabase system. To overcome these difficulties, we propose to incorporate additional data manipulation operations in the subtransactions of each multidatabase transaction. We show that if these operations create direct conflicts between subtransactions at each participating local database system, indirect conflicts can be resolved even if the multidatabase system is not aware of their existence. Based on this approach, we introduce optimistic and conservative multidatabase transaction management methods that require the local database systems to ensure only local serializability. The proposed methods do not violate the autonomy of the local database systems and guarantee global serializability by preventing multidatabase transactions from being serialized in different ways at the participating database systems. Refinements of these methods are also proposed for multidatabase environments where the participating database systems allow schedules that are cascadeless or transactions have analogous execution and serialization orders. In particular, we show that forced local conflicts can be eliminated in rigorous local systems, local cascadelessness simplifies the design of a global scheduler, and that local strictness offers no significant advantages over cascadelessness.> Dimitrios Georgakopoulos 0001, Marek Rusinkiewicz, Amit P. Sheth |
IEEE Trans. Knowl. Data Eng. | 3 |
| 1993 | Concurrency Control and Recovery of Multidatabase Work Flows in Telecommunication ApplicationsabstractIn a research and technology application project at Bellcore, we used multidatabase transactions to model multisystem work flows of telecommunication applications. During the project a prototype scheduler for executing multi-database transactions was developed. Two of the issues addressed in this project were concurrent execution of multi-database transactions and their failure recovery. This paper discusses our use of properties of the application and the telecommunication systems to develop simple and efficient solutions to the concurrency control and recovery problems. W. Woody Jin, Marek Rusinkiewicz, Linda A. Ness, Amit P. Sheth |
SIGMOD Conference | 4 |
| 1993 | Multidatabase Interdependencies in IndustryabstractIn this paper we address the problem of data consistency between interrelated data. In industrial environments,lack of consistent data creates difficulties in interoperation between systems and often requires manual interventions to restart operations that fail due to inconsistent data. We report the results of a study to understand applicability, adequacy and advantages of a framework we had proposed earlier to specify interdatabase dependencies in multidatabase environments. We studied several existing Bellcore systems and identified examples of interdependent data. The examples demonstrate that the framework allows precise and detailed specification of complex interdependencies that lead to efficient strategies to enforce the consistency requirements among the corporate data managed in multiple databases. We believe that our specification framework can help in the maintenance of data that meet a business’s consistency needs, reduce time consuming and costly manual operations, and provide data of better quality to end users. 1 Amit P. Sheth, George Karabatis |
SIGMOD Conference | 1 |
| 1993 | Task Scheduling Using Intertask Dependencies in CarotabstractThe Carnot Project at MCC is addressing the problem of logically unifying physically-distributed, enterprise-wide, heterogeneous information. Carnot will provide a user with the means to navigate information efficiently and transparently, to update that information consistently, and to write applications easily for large, heterogeneous, distributed information systems. A prototype has been implemented which provides services for (a) enterprise modeling and model integration to create an enterprise-wide view, (b) semantic expansion of queries on the view to queries on individual resources, and (c) inter-resource consistency management. This paper describes the Carnot approach to transaction processing in environments where heterogeneous, distributed, and autonomous systems are required to coordinate the update of the local information under their control. In this approach, subtransactions are represented as a set of tasks and a set of intertask dependencies that capture the semantics of a particular relaxed transaction model. A scheduler has been implemented which schedules the execution of these tasks in the Carnot environment so that all intertask dependencies are satisfied. Darrell Woelk, Paul C. Attie, Philip Cannata, Greg Meredith, Amit P. Sheth, Munindar P. Singh, Christine Tomlinson |
SIGMOD Conference | 5 |
| 1993 | Specifying and Enforcing Intertask Dependencies
Paul C. Attie, Munindar P. Singh, Amit P. Sheth, Marek Rusinkiewicz |
VLDB | 3 |
| 1993 | On Automatic Reasoning for Schema IntegrationabstractSuccess in database schema integration depends on the ability to capture real world semantics of the schema objects, and to reason about the semantics. Earlier schema integration approaches mainly rely on heuristics and human reasoning. In this paper, we discuss an approach to automate a significant part of the schema integration process. Our approach consists of three phases. An attribute hierarchy is generated in the first phase. This involves identifying relationships (equality, disjointness and inclusion) among attributes. We discuss a strategy based on user-specified semantic clustering. In the second phase, a classification algorithm based on the semantics of class subsumption is applied to the class definitions and the attribute hierarchy to automatically generate a class taxonomy. This class taxonomy represents a partially integrated schema. In the third phase, the user may employ a set of well-defined comparison operators in conjunction with a set of restructuring operators, to further modify the schema. These operators as well as the automatic reasoning during the second phase are based on subsumption. The formal semantics and automatic reasoning utilized in the second phase is based on a terminological logic as adapted in the CANDIDE data model. Classes are completely defined in terms of attributes and constraints. Our observation is that the inability to completely define attributes and thus completely capture their real world semantics imposes a fundamental limitation on the possibility of automatically reasoning about attribute definitions. This necessitates human reasoning during the first phase of the integration approach. Amit P. Sheth, Sunit K. Gala, Shamkant B. Navathe |
Int. J. Cooperative Inf. Syst. | 1 |
| 1992 | Using Flexible Transactions to Support Multi-System Telecommunication Applications
Mansoor Ansari, Linda A. Ness, Marek Rusinkiewicz, Amit P. Sheth |
VLDB | 4 |
| 1992 | Multidatabase Applications: Semantic and System Issues
Marek Rusinkiewicz, Amit P. Sheth |
VLDB | 2 |
| 1992 | Structural schema integration with full and partial correspondence using the Dual Model
James Geller, Yehoshua Perl, Erich J. Neuhold, Amit P. Sheth |
Inf. Syst. | 4 |
| 1991 | On Serializability of Multidatabase Transactions Through Forced Local ConflictsabstractA multidatabase transaction management mechanism called the optimistic ticket method (OTM) is introduced for enforcing global serializability. It permits the commitment of multidatabase transactions only if their relative serialization order is the same in all participating local database systems (LDBSs). OTM requires the LDBSs to guarantee only local serializability. The basic idea in OTM is to create direct conflicts between multidatabase transactions at each LDBS in order to determine the relative serialization order of their subtransactions. A refinement of OTM, called the implicit ticket method (ITM), is also introduced that uses implicit tickets and eliminates ticket conflicts but works only when the participating LDBSs use rigorous transaction scheduling mechanisms. ITM uses the local commitment order of each subtransaction to determine its implicit ticket value. It achieves global serializability by controlling the commitment (execution order) and thus the serialization order of multidatabase transactions. Both OTM and ITM do not violate the autonomy of the LDBSs and can be combined in a single comprehensive mechanism.> Dimitrios Georgakopoulos 0001, Marek Rusinkiewicz, Amit P. Sheth |
ICDE | 3 |
| 1991 | The Architecture of BrAID: A System for Bridging AI/DB SystemsabstractThe design of BrAID (a bridge between artificial intelligence and database management systems), an experimental system for the efficient integration of logic-based artificial intelligence (AI) and databases (DB) technologies, is described. Features provided by BrAID include (a) access to conventional DBMSs, (b) support for multiple inferencing strategies, (c) a powerful caching subsystem that manages views and uses subsumption to facilitate the reuse of previously cached data, (d) lazy or eager evaluation of queries submitted by the AI system, and (e) the generation of advice by the AI system to aid in cache management and query execution planning. Some of the key aspects of the BrAID architecture are discussed, focusing on the generation of advice by the AI system and its use by a cache management system to increase efficiency in accessing remote DBMSs through the selective application of such techniques as prefetching, query generalization, result caching, attribute indexing, and lazy evaluation.> Amit P. Sheth, Anthony B. O'Hare |
ICDE | 1 |
| 1991 | Federated Database Systems for Managing Distributed, Heterogeneous, and Autonomous Databases
Amit P. Sheth |
VLDB | 1 |
| 1991 | Cooperative Database Design (Panel)
Stefano Spaccapietra, Shamkant B. Navathe, Erich J. Neuhold, Amit P. Sheth |
VLDB | 4 |
| 1991 | Updating relational views using knowledge at view definition and view update time
James A. Larson, Amit P. Sheth |
Inf. Syst. | 2 |
| 1989 | Fault tolerance in a very large database system: a strawman analysisabstractA simple model is used to study the effect of fault-tolerance techniques and system design on system availability. A generic multiprocessor architecture is used that can be configured in different ways to study the effect of system architectures. Important parameters studied are different system architectures and hardware fault-tolerance techniques, mean time to failure of basic components, database size and distribution, interconnect capacity, etc. Quantitative analysis compares the relative effect of different parameter values. Results show that the effect of different parameter values on system availability can be very significant. System architecture, use of hardware fault tolerance (particularly mirroring), and data storage methods emerge as very important parameters under the control of a system designer.> Amit P. Sheth |
ICDCS | 1 |
| 1989 | Does Loose AI-DBMS Coupling Stand a Chance?abstractThe concept of loose coupling of artificial intelligence systems and database systems is defined. The issues that arise in determining the scope of loose integration are identified. A cache-based architecture being implemented in a knowledge management system project is described. The objective is to develop an interface that is useful for generalized database access from an expert system.> Amit P. Sheth |
ICDE | 1 |
| 1988 | TAILOR, A Tool for Updating Views
Amit P. Sheth, James A. Larson, Evan Watkins |
EDBT | 1 |
| 1988 | Managing and Integrating Unstructured and Structured Data: Problems of Representation, Features, and Abstraction (position paper)
Amit P. Sheth |
ICDE | 1 |
| 1988 | A Tool for Integrating Conceptual Schemas and User ViewsabstractAn interactive tool has been developed to assist database designers and administrators (DDA) in integrating schemas. It collects the information required for integration from a DDA, performs essential bookkeeping, and integrates schemas according to the semantics provided. The authors present the capabilities of this tool by discussing the integration methodology and the user interface of the tool.> Amit P. Sheth, James A. Larson, Aloysius Cornelio, Shamkant B. Navathe |
ICDE | 1 |
| 1987 | Performance Analysis of Resiliency Mechanisms in Distributed Datbase SystemsabstractDegradation in system performance due to component failures is an important factor that prevents a distributed database management system (DDBMS) from achieving its full potential of better availability, response time, and system throughput. Resiliency mechanisms that help to continue system operation in spite of failures introduce overhead due to the need to maintain redundant information. Earlier performance studies of DDBMS have either altogether ignored failure or studied only limited aspects of the effect of failures and performance of resiliency mechanisms. In this paper an analytical model is used to comprehensively characterize the effect of failures and resiliency mechanisms on the performance of DDBMS. Two new performance measures are introduced. The methodology is illustrated by comparatively evaluating the performance of three algorithms. Amit P. Sheth, Anoop Singhal, Ming T. Liu |
ICDE | 1 |
| 1986 | Integrating Locking and Optimistic Concurrency Control in Distributed Database Systems
Amit P. Sheth, Ming T. Liu |
ICDCS | 1 |
| 1986 | Deadlock Detection Algorithms in Distributed Database SystemsabstractIn this paper, a centralized deadlock detection algorithm with multiple outstanding requests (CDDMOR) is proposed for use in distributed database systems and transaction-processing systems. This algorithm allows a process to request many resources simultaneously. While a centralized scheme is superior to a completely distributed scheme in terms of performance, a major problem of such a scheme is congestion. Therefore, an important extension to the basic CDDMOR, a partially distributed scheme, is proposed to alleviate the problem of congestion, as well as to take advantage of the result presented by several researchers that global (multisite) deadlocks are infrequent. It takes care of the local (single site) deadlocks without involving other sites and uses centralized deadlock detection only when there is a possibility of global deadlock. Ahmed K. Elmagarmid, Amit P. Sheth, Ming T. Liu |
ICDE | 2 |
| 1985 | An Analysis of the Effect of Network Parameters on the Performance of Distributed Database SystemsabstractPerformance analysis studies of distributed database systems in the past have assumed that the message transmission time between any two nodes of a network is constant. They disregard the effect of communication network parameters such as network traffic, network topology, and capacity of transmission channels. In this paper, an analytical model is used to estimate the delays in transmission channels of the long haul network supporting the distributed database system. The analysis shows that the constant transmission time assumption cannot be justified in many cases, and that the response time is sensitive to the parameters mentioned above. Extensions and performance analysis in the context of interconnection networks are also discussed. Amit P. Sheth, Anoop Singhal, Ming T. Liu |
IEEE Trans. Software Eng. | 1 |
| 1984 | An Adaptive Concurrency Control Strategy for Distributed Database SystemsabstractPerformance of a Concurrency Control Algorithm (CCA) managing a distributed database system will deteriorate considerably when the configuration of the network supporting it will change due to either communication link failures or the communication delays introduced by varying load patterns. To get a good performance in spite of the changing configurations, we propose a scheme that involves breaking down the network into ‘weakly connected’ clusters. The problem to identify the clusters of a network is NP-hard. However, we present a heuristic strategy to identify the clusters of a network that works in polynomial time. Any of the present CCAs can be modified to work on a network that is partitioned into clusters by our scheme that uses (what we term as) multiple controllers. As an example, we present a Centralized Locking Algorithm with Acknowledgment using Multiple Controllers (CLAA/MC). Performance gain achieved using multiple controllers is also discussed. Amit P. Sheth, Anoop Singhal, Ming T. Liu |
ICDE | 1 |