EDBT 2026 Demo / reviewers in the wild / expert
Yolanda Gil
dblp:88/2686
· DBLP profile ↗
96ranked-venue papers
42as first author
3since 2021 · last 2022
0000-0001-8465-8341ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 20 first-author · 2 since 2021Databases, data management, data science and information retrieval · 29 · 19 first-authorApplied, interdisciplinary, general and emerging computing · 22 · 7 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 14 · 4 first-authorSoftware engineering, systems software and programming languages · 12 · 4 first-author · 1 since 2021Systems, architecture and hardware · 11Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-authorSecurity and privacy · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
16 papers |
Knowledge representation and reasoning · 84% Planning, search and constraint satisfaction · 6% Vision and language · 3% | |
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Computational science and engineering · 63% Bioinformatics and computational biology · 20% Computing education · 17% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
High-performance computing · 56% Performance modeling and evaluation · 25% Distributed systems · 12% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% | |
| Databases, data mining, and information retrieval
5 papers |
Data mining · 52% Information retrieval · 42% Data integration and cleaning · 5% |
Topics — the 28 heaviest of 39, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computational science and engineering › AI for science
automated scientific discovery |
0.4 | 1 | 2020 | Embedding the Scientific Record on the Web: Towards Automating Scientific Discoveries · WWW 2020 |
Computational science and engineering
scientific workflow |
0.4 | 1 | 2020 | Embedding the Scientific Record on the Web: Towards Automating Scientific Discoveries · WWW 2020 |
Computing education › STEM education
data science education |
0.2 | 1 | 2016 | Teaching Big Data Analytics Skills with Intelligent Workflow Systems · AAAI 2016 |
Multimedia analysis and retrieval
multimodal fusion |
0.2 | 1 | 2013 | Large-scale multimedia content analysis using scientific workflows · ACM Multimedia 2013 |
Information retrieval › web search
scholarly search |
0.1 | 1 | 2020 | Embedding the Scientific Record on the Web: Towards Automating Scientific Discoveries · WWW 2020 |
Performance modeling and evaluation
parameter space exploration |
0.1 | 1 | 2009 | An integrated framework for performance-based optimization of scientific workflows · HPDC 2009 |
High-performance computing
scientific workflow |
0.1 | 1 | 2009 | An integrated framework for performance-based optimization of scientific workflows · HPDC 2009 |
High-performance computing › scientific workflow
workflow optimization |
0.1 | 1 | 2009 | An integrated framework for performance-based optimization of scientific workflows · HPDC 2009 |
Data mining
big data analytics |
0.1 | 1 | 2016 | Teaching Big Data Analytics Skills with Intelligent Workflow Systems · AAAI 2016 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge acquisition |
0.1 | 1 | 2007 | Incorporating tutoring principles into interactive knowledge acquisition · Int. J. Hum. Comput. Stud. 2007 |
Computer vision › Vision and language
multimodal understanding |
0.0 | 1 | 2013 | Large-scale multimedia content analysis using scientific workflows · ACM Multimedia 2013 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
ontology |
0.0 | 1 | 2004 | Incremental formalization of document annotations through ontology-based paraphrasing · WWW 2004 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology
ontology-based annotation |
0.0 | 1 | 2004 | Incremental formalization of document annotations through ontology-based paraphrasing · WWW 2004 |
Natural language and speech › Information extraction and text analysis › data annotation
semantic annotation |
0.0 | 1 | 2004 | Incremental formalization of document annotations through ontology-based paraphrasing · WWW 2004 |
Distributed systems
grid computing |
0.0 | 1 | 2004 | Artemis: Integrating Scientific Data on the Grid · AAAI 2004 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
interactive knowledge capture |
0.0 | 1 | 2003 | Proactive Dialogue for Interactive Knowledge Capture · IJCAI 2003 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
knowledge extraction |
0.0 | 1 | 2003 | Proactive Dialogue for Interactive Knowledge Capture · IJCAI 2003 |
Natural language and speech › Question answering and dialogue systems
proactive dialogue |
0.0 | 1 | 2003 | Proactive Dialogue for Interactive Knowledge Capture · IJCAI 2003 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge level analysis |
0.0 | 1 | 2001 | Knowledge Analysis on Process Models · IJCAI 2001 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
process models |
0.0 | 1 | 2001 | Knowledge Analysis on Process Models · IJCAI 2001 |
Computational science and engineering
spatial data analysis |
0.0 | 1 | 2009 | An integrated framework for performance-based optimization of scientific workflows · HPDC 2009 |
Cloud and datacenter computing
resource allocation |
0.0 | 1 | 2008 | Self-Configuring Applications for Heterogeneous Systems: Program Composition and Optimization Using Cognitive Techniques · Proc. IEEE 2008 |
High-performance computing
scientific computing systems |
0.0 | 1 | 2007 | Wings for Pegasus: Creating Large-Scale Scientific Applications Using Semantic Representations of Computational Workflows · AAAI 2007 |
Data integration and cleaning
scientific data integration |
0.0 | 1 | 2004 | Artemis: Integrating Scientific Data on the Grid · AAAI 2004 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge base
knowledge base refinement |
0.0 | 1 | 1994 | Knowledge Refinement in a Reflective Architecture · AAAI 1994 |
Machine learning › Trustworthy machine learning › interpretability
explanation-based learning |
0.0 | 1 | 1989 | Explanation-Based Learning: A Problem Solving Perspective · Artif. Intell. 1989 |
Empirical software engineering
experimental methodology |
0.0 | 2 | 1993 | Efficient Domain-Independent Experimentation · ICML 1993 A Domain-Independent Framework for Effective Experimentation in Planning · ML 1991 |
Requirements engineering and software design › software architecture › architectural style
reflective architecture |
0.0 | 1 | 1994 | Knowledge Refinement in a Reflective Architecture · AAAI 1994 |
Methods — techniques the papers use, named apart from their topics
semantic web · 1.3workflow · 0.9provenance · 0.9workflow systems · 0.8ontology · 0.8semantic representation · 0.5scientific workflows · 0.3machine learning · 0.3linked data · 0.3computer vision · 0.3semantic representations · 0.2performance-based optimization · 0.2natural language interpretation · 0.1cognitive techniques · 0.1workflow generation · 0.1interactive knowledge acquisition · 0.1learning by experimentation · 0.0incremental refinement · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Towards Capturing Scientific Reasoning to Automate Data Analysis
Yolanda Gil, Deborah Khider, Maximiliano Osorio, Varun Ratnakar, Hernán Vargas, Daniel Garijo, Suzanne A. Pierce |
CogSci | 1 |
| 2021 | Towards Democratizing Modeling at ScaleabstractWe use AI techniques to create a modeling environment that makes sophisticated models accessible to non-experts. Our AI framework for Model INTegration (MINT) assists users to explore scenarios, which MINT can run in a local environment or at scale in a supercomputing facility. We are using MINT with hydrology, agriculture, and drought models for food security. Yolanda Gil, Maximiliano Osorio, Varun Ratnakar, Suzanne A. Pierce, Je'aime H. Powell, Nicolas Thorne, Peter Lubbs |
e-Science | 1 |
| 2021 | Artificial Intelligence for Modeling Complex Systems: Taming the Complexity of Expert Models to Improve Decision MakingabstractMajor societal and environmental challenges involve complex systems that have diverse multi-scale interacting processes. Consider, for example, how droughts and water reserves affect crop production and how agriculture and industrial needs affect water quality and availability. Preventive measures, such as delaying planting dates and adopting new agricultural practices in response to changing weather patterns, can reduce the damage caused by natural processes. Understanding how these natural and human processes affect one another allows forecasting the effects of undesirable situations and study interventions to take preventive measures. For many of these processes, there are expert models that incorporate state-of-the-art theories and knowledge to quantify a system's response to a diversity of conditions. A major challenge for efficient modeling is the diversity of modeling approaches across disciplines and the wide variety of data sources available only in formats that require complex conversions. Using expert models for particular problems requires integration of models with third-party data as well as integration of models across disciplines. Modelers face significant heterogeneity that requires resolving semantic, spatiotemporal, and execution mismatches, which are largely done by hand today and may take more than 2 years of effort. We are developing a modeling framework that uses artificial intelligence (AI) techniques to reduce modeling effort while ensuring utility for decision making. Our work to date makes several innovative contributions: (1) an intelligent user interface that guides analysts to frame their modeling problem and assists them by suggesting relevant choices and automating steps along the way; (2) semantic metadata for models, including their modeling variables and constraints, that ensures model relevance and proper use for a given decision-making problem; and (3) semantic representations of datasets in terms of modeling variables that enable automated data selection and data transformations. This framework is implemented in the MINT (Model INTegration) framework, and currently includes data and models to analyze the interactions between natural and human systems involving climate, water availability, agricultural production, and markets. Our work to date demonstrates the utility of AI techniques to accelerate modeling to support decision-making and uncovers several challenging directions for future work. Yolanda Gil, Daniel Garijo, Deborah Khider, Craig A. Knoblock, Varun Ratnakar, Maximiliano Osorio, Hernán Vargas, Minh Pham 0004, Jay Pujara, Basel Shbita, Yao-Yi Chiang, Dan Feldman, Yijun Lin 0001, Hayley Song, Vipin Kumar 0001, Ankush Khandelwal, Michael S. Steinbach, Kshitij Tayal, Shaoming Xu, Suzanne A. Pierce, Lissa Pearson, Daniel Hardesty-Lewis, Ewa Deelman, Rafael Ferreira da Silva, Rajiv Mayani, Armen R. Kemanian, Lorne Leonard, Scott D. Peckham, Maria Stoica 0001, Kelly M. Cobourn, Zeya Zhang, Christopher J. Duffy, Lele Shu |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2020 | Keynote Speaker: Yolanda GilabstractDr. Yolanda Gil Dr. Yolanda Gil is Director of Knowledge Technologies and Associate Division Director at the Information Sciences Institute of the University of Southern California, and Research Professor in Computer Science and in Spatial Sciences. She is also Associate Director of Interdisciplinary Programs in Informatics. She received her M.S. and Ph. D. degrees in Computer Science from Carnegie Mellon University, with a focus on artificial intelligence. Her research is on intelligent interfaces for knowledge capture and discovery, which she investigates in a variety of projects concerning knowledge-based planning and problem solving, information analysis and assessment of trust, semantic annotation and metadata, and community-wide development of knowledge bases. Dr. Gil collaborates with scientists in different domains on semantic workflows and metadata capture, social knowledge collection, computer-mediated collaboration, and automated discovery. Dr. Gil has served in the Advisory Committee of the Computer Science and Engineering Directorate of the National Science Foundation. She initiated and chaired the W3C Provenance Group that led to a community standard in this area. Dr. Gil is a Fellow of the Association for Computing Machinery (ACM), and Past Chair of its Special Interest Group in Artificial Intelligence. She is also Fellow of the Association for the Advancement of Artificial Intelligence (AAAI), and was elected as its 24th President in 2016. Yolanda Gil |
KDD | 1 |
| 2020 | Embedding the Scientific Record on the Web: Towards Automating Scientific DiscoveriesabstractFuture AI systems will be key contributors to science, but this is unlikely to happen unless we reinvent our current publications and embed our scientific records in the Web as structured Web objects. This implies that our scientific papers of the future will be complemented with explicit, structured descriptions of the experiments, software, data, and workflows used to reach new findings. These scientific papers of the future will not only culminate the promise of open science and reproducible research, but also enable the creation of AI systems that can ingest and organize scientific methods and processes, re-run experiments and re-analyze results, and explore their own hypothesis in systematic and unbiased ways. In this talk, I will describe guidelines for writing scientific papers of the future that embed the scientific record on the Web, and our progress on AI systems capable of using them to systematically explore experiments. I will also outline a research agenda with seven key characteristics for creating AI scientists that will exploit the Web to independently make new discoveries [1]. AI scientists have the potential to transform science and the processes of scientific discovery [2, 3]. Yolanda Gil |
WWW | 1 |
| 2019 | OKG-Soft: An Open Knowledge Graph with Machine Readable Scientific Software MetadataabstractScientific software is crucial for understanding, reusing and reproducing results in computational sciences. Software is often stored in code repositories, which may contain human readable instructions necessary to use it and set it up. However, a significant amount of time is usually required to understand how to invoke a software component, prepare data in the format it requires, and use it in combination with other software. In this paper we introduce OKG-Soft, an open knowledge graph that describes scientific software in a machine readable manner. OKG-Soft includes: 1) an ontology designed to describe software and the specific data formats it uses; 2) an approach to publish software metadata as an open knowledge graph, linked to other Web of Data objects; and 3) a framework to annotate, query, explore and curate scientific software metadata. OKG-Soft supports the FAIR principles of findability, accessibility, interoperability, and reuse for software. We demonstrate the benefits of OKG-Soft with two applications: a browser for understanding scientific models in the environmental and social sciences, and a portal to combine climate, hydrology, agriculture, and economic software models. Daniel Garijo, Maximiliano Osorio, Deborah Khider, Varun Ratnakar, Yolanda Gil |
eScience | 5 |
| 2019 | Towards human-guided machine learningabstractAutomated Machine Learning (AutoML) systems are emerging that automatically search for possible solutions from a large space of possible kinds of models. Although fully automated machine learning is appropriate for many applications, users often have knowledge that supplements and constraints the available data and solutions. This paper proposes human-guided machine learning (HGML) as a hybrid approach where a user interacts with an AutoML system and tasks it to explore different problem settings that reflect the user's knowledge about the data available. We present: 1) a task analysis of HGML that shows the tasks that a user would want to carry out, 2) a characterization of two scientific publications, one in neuroscience and one in political science, in terms of how the authors would search for solutions using an AutoML system, 3) requirements for HGML based on those characterizations, and 4) an assessment of existing AutoML systems in terms of those requirements. Yolanda Gil, James Honaker, Shikhar Gupta, Yibo Ma, Vito D'Orazio, Daniel Garijo, Shruti Gadewar, Neda Jahanshad |
IUI | 1 |
| 2018 | Semantic Software Metadata for Workflow Exploration and EvolutionabstractScientific workflow management systems play a major role in the design, execution and documentation of computational experiments. However, they have limited support for managing workflow evolution and exploration because they lack rich metadata for the software that implements workflow components. Such metadata could be used to support scientists in exploring local adjustments to a workflow, replacing components with similar software, or upgrading components upon release of newer software versions. To address this challenge, we propose OntoSoft-VFF (Ontology for Software Version, Function and Functionality), a software metadata repository designed to capture information about software and workflow components that is important for managing workflow exploration and evolution. Our approach uses a novel ontology to describe the functionality and evolution through time of any software used to create workflow components. OntoSoft-VFF is implemented as an online catalog that stores semantic metadata for software to enable workflow exploration through understanding of software functionality and evolution. The catalog also supports comparison and semantic search of software metadata. We showcase OntoSoft-VFF using machine learning workflow examples. We validate our approach by testing that a workflow system could compare differences in software metadata, explain software updates and describe the general functionality of workflow steps. Lucas Augusto Montalvão Costa Carvalho, Daniel Garijo, Claudia Bauzer Medeiros, Yolanda Gil |
eScience | 4 |
| 2018 | PSM-Flow: Probabilistic Subgraph Mining for Discovering Reusable Fragments in WorkflowsabstractScientific workflows define computational processes needed for carrying out scientific experiments. Existing workflow repositories contain hundreds of scientific workflows, where scientists can find materials and knowledge to facilitate workflow design for running related experiments. Identifying reusable fragments in growing workflow repositories has become increasingly important. In this paper, we present PSM-Flow, a probabilistic subgraph mining algorithm designed to discover commonly occurring fragments in a workflow corpus using a modified version of the Latent Dirichlet Allocation algorithm. The proposed model encodes the geodesic distance between workflow steps into the model for implicitly modeling fragments. PSM-Flow captures variations of frequent fragments while maintaining its space complexity bounded polynomially, as it requires no candidate generation. We applied PSM-Flow to three real-world scientific workflow datasets containing more than 750 workflows for neuroimaging analysis. Our results show that PSM-Flow outperforms three state of the art frequent subgraph mining techniques. We also discuss other potential future improvements of the proposed method. Chin Wang Cheong, Daniel Garijo, William Kwok-Wai Cheung, Yolanda Gil |
WI | 4 |
| 2017 | Towards Continuous Scientific Data Analysis and Hypothesis EvolutionabstractScientific data is continuously generated throughout the world. However, analyses of these data are typically performed exactly once and on a small fragment of recently generated data. Ideally, data analysis would be a continuous process that uses all the data available at the time, and would be automatically re-run and updated when new data appears. We present a framework for automated discovery from data repositories that tests user-provided hypotheses using expert-grade data analysis strategies, and reassesses hypotheses when more data becomes available. Novel contributions of this approach include a framework to trigger new analyses appropriate for the available data through lines of inquiry that support progressive hypothesis evolution, and a representation of hypothesis revisions with provenance records that can be used to inspect the results. We implemented our approach in the DISK framework, and evaluated it using two scenarios from cancer multi-omics: 1) data for new patients becomes available over time, 2) new types of data for the same patients are released. We show that in all scenarios DISK updates the confidence on the original hypotheses as it automatically analyzes new data. Yolanda Gil, Daniel Garijo, Varun Ratnakar, Rajiv Mayani, Ravali Adusumilli, Hunter Boyce, Arunima Srivastava, Parag Mallick |
AAAI | 1 |
| 2017 | Towards Automating Data NarrativesabstractWe propose a new area of research on automating data narratives. Data narratives are containers of information about computationally generated research findings. They have three major components: 1) A record of events, that describe a new result through a workflow and/or provenance of all the computations executed; 2) Persistent entries for key entities involved for data, software versions, and workflows; 3) A set of narrative accounts that are automatically generated human-consumable renderings of the record and entities and can be included in a paper. Different narrative accounts can be used for different audiences with different content and details, based on the level of interest or expertise of the reader. Data narratives can make science more transparent and reproducible, because they ensure that the text description of the computational experiment reflects with high fidelity what was actually done. Data narratives can be incorporated in papers, either in the methods section or as supplementary materials. We introduce DANA, a prototype that illustrates how to generate data narratives automatically, and describe the information it uses from the computational records. We also present a formative evaluation of our approach and discuss potential uses of automated data narratives. Yolanda Gil, Daniel Garijo |
IUI | 1 |
| 2017 | A Controlled Crowdsourcing Approach for Practical Ontology Extensions and Metadata Annotations
Yolanda Gil, Daniel Garijo, Varun Ratnakar, Deborah Khider, Julien Emile-Geay, Nicholas McKay |
ISWC (2) | 1 |
| 2017 | Abstract, link, publish, exploit: An end to end framework for workflow sharing
Daniel Garijo, Yolanda Gil, Óscar Corcho |
Future Gener. Comput. Syst. | 2 |
| 2017 | Scientific workflows in data analysis: Bridging expertise across multiple domains
Ricky J. Sethi, Yolanda Gil |
Future Gener. Comput. Syst. | 2 |
| 2016 | Teaching Big Data Analytics Skills with Intelligent Workflow SystemsabstractWe have designed an open and modular course for data science and big data analytics using a workflow paradigm that allows students to easily experience big data through a sophisticated yet easy to use instrument that is an intelligent workflow system. A key aspect of this work is the use of semantic workflows to capture and reuse end-to-end analytic methods that experts would use to analyze big data, and the use of an intelligent workflow system to elaborate the workflow and manage its execution and resulting datasets. Through the exposure of big data analytics in a workflow framework, students will be able to get first-hand experiences with a breadth of big data topics, including multi-step data analytic and statistical methods, software reuse and composition, parallel distributed programming, high-end computing. In addition, students learn about a range of topics in AI, including semantic representations and ontologies, machine learning, natural language processing, and image analysis. Yolanda Gil |
AAAI | 1 |
| 2016 | OntoSoft: A distributed semantic registry for scientific softwareabstractOntoSoft is a distributed semantic registry for scientific software. This paper describes three major novel contributions of OntoSoft: 1) a software metadata registry designed for scientists, 2) a distributed approach to software registries that targets communities of interest, and 3) metadata crowdsourcing through access control. Software metadata is organized using the OntoSoft ontology along six dimensions that matter to scientists: identify software, understand and assess software, execute software, get support for the software, do research with the software, and update the software. OntoSoft is a distributed registry where each site is owned and maintained by a community of interest, with a distributed semantic query capability that allows users to search across all sites. The registry has metadata crowdsourcing capabilities, supported through access control so that software authors can allow others to expand on specific metadata properties. Yolanda Gil, Daniel Garijo, Varun Ratnakar |
eScience | 1 |
| 2016 | Reproducibility in computer vision: Towards open publication of image analysis experiments as semantic workflowsabstractReproducibility of research is an area of growing concern in computer vision. Scientific workflows provide a structured methodology for standardized replication and testing of state-of-the-art models, open publication of datasets and software together, and ease of analysis by re-using pre-existing components. In this paper, we present initial work in developing a framework that will allow reuse and extension of many computer vision methods, as well as allowing easy reproducibility of analytical results, by publishing dadasets and workflows packaged together as linked data. Our approach uses the WINGS semantic workflow system which validates semantic constraints of the computer vision algorithms, making it easy for non-experts to correctly apply state-of-the-art image processing methods to their data. We show the ease of use of semantic workflows for reproducibility in computer vision by both utilizing pre-developed workflow fragments and developing novel computer vision workflow fragments for a video activity recognition task, analysis of multimedia web content, and the analysis of artistic style in paintings using convolutional neural networks. Ricky J. Sethi, Yolanda Gil |
eScience | 2 |
| 2015 | A Task-Centered Framework for Computationally-Grounded Science CollaborationsabstractCollaboration is ubiquitous in today's science, yet there is limited support for coordinating scientific work. The general-purpose tools that are typically used (e.g., email, shared document editing, social coding sites), have still not replaced in-person meetings, phone calls, and extensive emails needed to coordinate and track collaborative activities. Scientists with diverse knowledge and skills around the globe could collaborate by opening scientific processes that expose all tasks and activities publicly to achieve a shared scientific question. This paper describes the Organic Data Science framework to support scientific collaborations that revolve around complex science questions that require significant coordination, entice contributors to remain engaged for extended periods of time, and enable continuous growth to accommodate new contributors as the work evolves over time. We discuss how the design of this framework incorporates principles followed by successful on-line communities. We present initial results to date of several communities that are collaborating using this framework. Yolanda Gil, Felix Michel, Varun Ratnakar, Matheus Hauder, Christopher J. Duffy, Hilary Dugan, Paul C. Hanson |
e-Science | 1 |
| 2015 | Supporting Open Collaboration in Science Through Explicit and Linked Semantic Description of Processes
Yolanda Gil, Felix Michel, Varun Ratnakar, Jordan S. Read, Matheus Hauder, Christopher J. Duffy, Paul C. Hanson, Hilary Dugan |
ESWC | 1 |
| 2015 | The Provenance Bee Wiki: Tracking the Growth of Semantic Wiki CommunitiesabstractContributors in hundreds of semantic wiki sites are creating structured information in RDF every day, thus growing the semantic content of the Web in spades. Although wikis have been analyzed extensively, there has been little analysis of the use of semantic wikis. The Provenance Bee Wiki was created to gather and aggregate data from these sites, show how this content is growing over time, and to make all this detailed data readily available to the research community. We also present a high-level analysis of the almost 600 wikis indexed in Provenance Bee Wiki that have less than 5,000 pages. Yolanda Gil, Dipsy Kapoor, Reed Markham, Varun Ratnakar |
K-CAP | 1 |
| 2015 | OntoSoft: Capturing Scientific Software MetadataabstractThis paper presents OntoSoft, an ontology to describe metadata for scientific software. The ontology is designed considering how scientists would approach the reuse and sharing of software. This includes supporting a scientist to: 1) identify software, 2) understand and assess software, 3) execute software, 4) get support for the software, 5) do research with the software, and 6) update the software. The ontology is available in OWL and contains more than fifty terms. We are using OntoSoft to structure a software registry for geosciences, and to develop user interfaces to capture its metadata. Yolanda Gil, Varun Ratnakar, Daniel Garijo |
K-CAP | 1 |
| 2015 | Human Tutorial Instruction in the RawabstractHumans learn procedures from one another through a variety of methods, such as observing someone do the task, practicing by themselves, reading manuals or textbooks, or getting instruction from a teacher. Some of these methods generate examples that require the learner to generalize appropriately. When procedures are complex, however, it becomes unmanageable to induce the procedures from examples alone. An alternative and very common method for teaching procedures is tutorial instruction, where a teacher describes in general terms what actions to perform and possibly includes explanations of the rationale for the actions. This article provides an overview of the challenges in using human tutorial instruction for teaching procedures to computers. First, procedures can be very complex and can involve many different types of interrelated information, including (1) situating the instruction in the context of relevant objects and their properties, (2) describing the steps involved, (3) specifying the organization of the procedure in terms of relationships among steps and substeps, and (4) conveying control structures. Second, human tutorial instruction is naturally plagued with omissions, oversights, unintentional inconsistencies, errors, and simply poor design. The article presents a survey of work from the literature that highlights the nature of these challenges and illustrates them with numerous examples of instruction in many domains. Major research challenges in this area are highlighted, including the difficulty of the learning task when procedures are complex, the need to overcome omissions and errors in the instruction, the design of a natural user interface to specify procedures, the management of the interaction of a human with a learning system, and the combination of tutorial instruction with other teaching modalities. Yolanda Gil |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2015 | Special Issue of the Journal of Web Semantics on Geospatial Semantics
Yolanda Gil, Raphaël Troncy |
J. Web Semant. | 1 |
| 2014 | Workflow Reuse in Practice: A Study of Neuroimaging Pipeline UsersabstractWorkflow reuse is a major benefit of workflow systems and shared workflow repositories, but there are barely any studies that quantify the degree of reuse of workflows or the practical barriers that may stand in the way of successful reuse. In our own work, we hypothesize that defining workflow fragments improves reuse, since end-to-end workflows may be very specific and only partially reusable by others. This paper reports on a study of the current use of workflows and workflow fragments in labs that use the LONI Pipeline, a popular workflow system used mainly for neuroimaging research that enables users to define and reuse workflow fragments. We present an overview of the benefits of workflows and workflow fragments reported by users in informal discussions. We also report on a survey of researchers in a lab that has the LONI Pipeline installed, asking them about their experiences with reuse of workflow fragments and the actual benefits they perceive. This leads to quantifiable indicators of the reuse of workflows and workflow fragments in practice. Finally, we discuss barriers to further adoption of workflow fragments and workflow reuse that motivate further work. Daniel Garijo, Óscar Corcho, Yolanda Gil, Meredith N. Braskie, Derrek P. Hibar, Xue Hua, Neda Jahanshad, Paul M. Thompson, Arthur W. Toga |
eScience | 3 |
| 2014 | FragFlow Automated Fragment Detection in Scientific WorkflowsabstractScientific workflows provide the means to define, execute and reproduce computational experiments. However, reusing existing workflows still poses challenges for workflow designers. Workflows are often too large and too specific to reuse in their entirety, so reuse is more likely to happen for fragments of workflows. These fragments may be identified manually by users as sub-workflows, or detected automatically. In this paper we present the FragFlow approach, which detects workflow fragments automatically by analyzing existing workflow corpora with graph mining algorithms. FragFlow detects the most common workflow fragments, links them to the original workflows and visualizes them. We evaluate our approach by comparing FragFlow results against user-defined sub-workflows from three different corpora of the LONI Pipeline system. Based on this evaluation, we discuss how automated workflow fragment detection could facilitate workflow reuse. Daniel Garijo, Óscar Corcho, Yolanda Gil, Boris Gutman, Ivo D. Dinov, Paul M. Thompson, Arthur W. Toga |
eScience | 3 |
| 2014 | Common motifs in scientific workflows: An empirical analysis
Daniel Garijo, Pinar Alper, Khalid Belhajjame, Óscar Corcho, Yolanda Gil, Carole A. Goble |
Future Gener. Comput. Syst. | 5 |
| 2014 | Similarity assessment and efficient retrieval of semantic workflows
Ralph Bergmann, Yolanda Gil |
Inf. Syst. | 2 |
| 2014 | Ten Simple Rules for the Care and Feeding of Scientific DataabstractAuthor(s): Goodman, Alyssa; Pepe, Alberto; Blocker, Alexander W; Borgman, Christine L; Cranmer, Kyle; Crosas, Merce; Di Stefano, Rosanne; Gil, Yolanda; Groth, Paul; Hedstrom, Margaret; Hogg, David W; Kashyap, Vinay; Mahabal, Ashish; Siemiginowska, Aneta; Slavkovic, Aleksandra | Editor(s): Bourne, Philip E Alyssa Goodman, Alberto Pepe, Alexander W. Blocker, Christine L. Borgman, Kyle Cranmer, Mercè Crosas, Rosanne Di Stefano, Yolanda Gil, Paul Groth, Margaret L. Hedstrom, David W. Hogg, Vinay L. Kashyap, Ashish Mahabal, Aneta Siemiginowska, Aleksandra B. Slavkovic |
PLoS Comput. Biol. | 8 |
| 2013 | Towards task-centered network models through semantic workflowsabstractVirtual organizations conduct network operations by executing complex tasks over resources that are distributed and vulnerable to attack. Current network models do not have a task-oriented representation of the mission, which is crucial to manage the accomplishment of mission goals while ongoing attacks and deception are occurring in the network. We describe a new approach to model high-level tasks to be accomplished in the network through semantic workflows. The system can then map tasks dynamically to logical and physical resources in the network, and reassign those mappings if any resources are compromised. Yolanda Gil |
ISI | 1 |
| 2013 | Detecting common scientific workflow fragments using templates and execution provenanceabstractProvenance plays a major role when understanding and reusing the methods applied in a scientific experiment, as it provides a record of inputs, the processes carried out and the use and generation of intermediate and final results. In the specific case of in-silico scientific experiments, a large variety of scientific workflow systems (e.g., Wings, Taverna, Galaxy, Vistrails) have been created to support scientists. All of these systems produce some sort of provenance about the executions of the workflows that encode scientific experiments. However, provenance is normally recorded at a very low level of detail, which complicates the understanding of what happened during execution. In this paper we propose an approach to automatically obtain abstractions from low-level provenance data by finding common workflow fragments on workflow execution provenance and relating them to templates. We have tested our approach with a dataset of workflows published by the Wings workflow system. Our results show that by using these kinds of abstractions we can highlight the most common abstract methods used in the executions of a repository, relating different runs and workflow templates with each other. Daniel Garijo, Óscar Corcho, Yolanda Gil |
K-CAP | 3 |
| 2013 | Knowledge capture in the wild: a perspective from semantic wiki communitiesabstractSemantic wikis augment wikis with semantic properties that can be used to structure content that can therefore be aggregated and queried through reasoning. Semantic wikis have been adopted by many communities for very diverse purposes, such as organizing genomic knowledge, coding software, learn about hobbies, and tracking environmental data. Although wikis have been analyzed extensively, there has been little analysis of the use of semantic wikis. In this paper, we analyze the formalization of knowledge in 230 semantic wiki communities. We report our findings in terms of the edits of semantic concepts and properties, as well as the communities of editors for these semantic features of the wikis. Yolanda Gil, Varun Ratnakar |
K-CAP | 1 |
| 2013 | Large-scale multimedia content analysis using scientific workflowsabstractAnalyzing web content, particularly multimedia content, for security applications is of great interest. However, it often requires deep expertise in data analytics that is not always accessible to non-experts. Our approach is to use scientific workflows that capture expert-level methods to examine web content. We use workflows to analyze the image and text components of multimedia web posts separately, as well as by a multimodal fusion of both image and text data. In particular, we re-purpose workflow fragments to do the multimedia analysis and create additional components for the fusion of the image and text modalities. In this paper, we present preliminary work which focuses on a Human Trafficking Detection task to help deter human trafficking of minors by thus fusing image and text content from the web. We also examine how workflow fragments save time and effort in multimedia content analysis while bringing together multiple areas of machine learning and computer vision. We further export these workflow fragments using linked data as web objects. Ricky J. Sethi, Yolanda Gil, Hyunjoon Jo, Andrew Philpot |
ACM Multimedia | 2 |
| 2013 | Structured analysis of the ISI Atomic Pair Actions dataset using workflows
Ricky J. Sethi, Hyunjoon Jo, Yolanda Gil |
Pattern Recognit. Lett. | 3 |
| 2012 | Common motifs in scientific workflows: An empirical analysisabstractWhile workflow technology has gained momentum in the last decade as a means for specifying and enacting computational experiments in modern science, reusing and repurposing existing workflows to build new scientific experiments is still a daunting task. This is partly due to the difficulty that scientists experience when attempting to understand existing workflows, which contain several data preparation and adaptation steps in addition to the scientifically significant analysis steps. One way to tackle the understandability problem is through providing abstractions that give a high-level view of activities undertaken within workflows. As a first step towards abstractions, we report in this paper on the results of a manual analysis performed over a set of real-world scientific workflows from Taverna and Wings systems. Our analysis has resulted in a set of scientific workflow motifs that outline i) the kinds of data intensive activities that are observed in workflows (data oriented motifs), and ii) the different manners in which activities are implemented within workflows (workflow oriented motifs). These motifs can be useful to inform workflow designers on the good and bad practices for workflow development, to inform the design of automated tools for the generation of workflow abstractions, etc. Daniel Garijo, Pinar Alper, Khalid Belhajjame, Óscar Corcho, Yolanda Gil, Carole A. Goble |
eScience | 5 |
| 2012 | Reproducibility and Efficiency of Scientific Data Analysis: Scientific Workflows and Case-Based Reasoning
Yolanda Gil |
ICCBR | 1 |
| 2012 | Capturing Common Knowledge about Tasks: Intelligent Assistance for To-Do ListsabstractAlthough to-do lists are a ubiquitous form of personal task management, there has been no work on intelligent assistance to automate, elaborate, or coordinate a user’s to-dos. Our research focuses on three aspects of intelligent assistance for to-dos. We investigated the use of intelligent agents to automate to-dos in an office setting. We collected a large corpus from users and developed a paraphrase-based approach to matching agent capabilities with to-dos. We also investigated to-dos for personal tasks and the kinds of assistance that can be offered to users by elaborating on them on the basis of substep knowledge extracted from the Web. Finally, we explored coordination of user tasks with other users through a to-do management application deployed in a popular social networking site. We discuss the emergence of Social Task Networks, which link users‘ tasks to their social network as well as to relevant resources on the Web. We show the benefits of using common sense knowledge to interpret and elaborate to-dos. Conversely, we also show that to-do lists are a valuable way to create repositories of common sense knowledge about tasks. Yolanda Gil, Varun Ratnakar, Timothy Chklovski, Paul Groth, Denny Vrandecic |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2011 | A Framework for Efficient Data Analytics through Automatic Configuration and Customization of Scientific WorkflowsabstractData analytics involves choosing between many different algorithms and experimenting with possible combinations of those algorithms. Existing approaches however do not support scientists with the laborious tasks of exploring the design space of computational experiments. We have developed a framework to assist scientists with data analysis tasks in particular machine learning and data mining. It takes advantage of the unique capabilities of the Wings workflow system to reason about semantic constraints. We show how the framework can rule out invalid workflows and help scientists to explore the design space. We demonstrate our system in the domain of text analytics, and outline the benefits of our approach. Matheus Hauder, Yolanda Gil, Yan Liu 0002 |
eScience | 2 |
| 2011 | Retrieval of Semantic Workflows with Knowledge Intensive Similarity Measures
Ralph Bergmann, Yolanda Gil |
ICCBR | 2 |
| 2011 | A formal framework for combining natural instruction and demonstration for end-user programmingabstractWe contribute to the difficult problem of programming via natural language instruction. We introduce a formal framework that allows for the use of program demonstrations to resolve several types of ambiguities and omissions that are common in such instructions. The framework effectively combines some of the benefits of programming by demonstration and programming by natural instruction. The key idea of our approach is to use non-deterministic programs to compactly represent the (possibly infinite) set of candidate programs for given instructions, and to filter from this set by means of simulating the execution of these programs following the steps of a given demonstration. Due to the rigorous semantics of our framework we can prove that this leads to a sound algorithm for identifying the intended program, making assumptions only about the types of ambiguities and omissions occurring in the instruction. We have implemented our approach and demonstrate its ability to resolve ambiguities and omissions by considering a list of classes of such issues and how our approach resolves them in a concrete example domain. Our empirical results show that our approach can effectively and efficiently identify programs that are consistent with both the natural instruction and the given demonstrations. Christian Fritz 0001, Yolanda Gil |
IUI | 2 |
| 2011 | TellMe: learning procedures from tutorial instructionabstractThis paper describes an approach to allow end users to define new procedures through tutorial instruction. Our approach allows users to specify procedures in natural language in the same way that they would instruct another person, while the system handles incompleteness and ambiguity inherent in natural human instruction and formulates follow up questions. We describe the key features of our approach, which include exposing prior knowledge, deductive and heuristic reasoning, shared learning state, and selectively asking questions to the user. We also describe how those key features are realized in our implemented TellMe system, and present preliminary user studies where non-programmers were able to easily specify complex multi-step procedures. Yolanda Gil, Varun Ratnakar, Christian Fritz 0001 |
IUI | 1 |
| 2011 | Want world domination? win at risk!: matching to-do items with how-tos from the webabstractTo-Do lists are widely used for personal task management. We propose a novel approach to assist users in managing their To-Dos by matching them to How-To knowledge from the Web. We have implemented a system that, given a To-Do item, provides a number of possibly matching How-Tos, broken down into steps that can be used as new To-Do entries. Our implementation is in the form of a web service that can be easily integrated into existing To-Do applications. This can help users by providing them with an approach to tackle the To-Do by listing smaller, more actionable To-Dos. In this paper we present our implementation, an evaluation of the matching component over two sets of To-Do corpora with very different characteristics, and a discussion of the results. Denny Vrandecic, Yolanda Gil, Varun Ratnakar |
IUI | 2 |
| 2011 | LinkedDataLens: linked data as a network of networksabstractWith billions of assertions and counting, the Web of Data represents the largest multi-contributor interlinked knowledge base that ever existed. We present a novel framework for analyzing and using the Web of Data based on extracting and analyzing thematic subsets of it. We view the Web of Data as a "network of networks" from which to extract meaningful subsets that can be converted them into self-contained networks to be further analyzed and reused. These extracted networks can then be analyzed through network analysis and discovery algorithms, and the results of these analyses can be published back on the Web of Data. We describe LinkedDataLens, an implementation of this framework that uses the Wings workflow system to represent multi-step network extraction and analysis processes. Yolanda Gil, Paul Groth |
K-CAP | 1 |
| 2011 | Mind Your Metadata: Exploiting Semantics for Configuration, Adaptation, and Provenance in Scientific Workflows
Yolanda Gil, Pedro A. Szekely, Sandra R. Villamizar, Thomas C. Harmon, Varun Ratnakar, Maria Muslea, Fabio Silva, Craig A. Knoblock |
ISWC (2) | 1 |
| 2011 | The Open Provenance Model core specification (v1.1)
Luc Moreau 0001, Ben Clifford, Juliana Freire, Joe Futrelle, Yolanda Gil, Paul Groth, Natalia Kwasnikowska, Simon Miles, Paolo Missier, James D. Myers, Beth Plale, Yogesh L. Simmhan, Eric G. Stephan, Jan Van den Bussche |
Future Gener. Comput. Syst. | 5 |
| 2011 | A semantic framework for automatic generation of computational workflows using distributed data and component cataloguesabstractComputational workflows are a powerful paradigm to represent and manage complex applications, particularly in large-scale distributed scientific data analysis. Workflows represent application components that result in individual computations as well as their interdependences in terms of dataflow. Workflow systems use these representations to manage various aspects of workflow creation and execution for users, such as the automatic assignment of execution resources. This article describes an approach to automating a new aspect of the process: the selection of application components and data sources. We present a novel approach that enables users to specify varying degrees of detail and amount of constraints in a workflow request, including the specification of constraints on input, intermediate or output data in the workflow, abstract workflow component classes rather than specific component implementations, and generic reusable workflow templates that express a pre-defined combination of components. The algorithm elaborates the user request into a set of fully ground workflows with specific choices of data sources and codes to be used so that they can be submitted for mapping and execution. The algorithm searches through the space of possible candidate workflows by creating increasingly more specialized versions of the original template and eliminating candidates that violate constraints cumulated in the candidate workflow as components and data sources are selected. A novel feature of our approach is that it assumes a distributed architecture where data and component catalogues are separate from the workflow system. The algorithm explicitly poses queries to external catalogues, and therefore any reasoning regarding data or component properties is not assumed to occur within the workflow system. We describe our implementation of this approach in the Wings workflow system. This implementation uses the W3C Web Ontology Language and associated reasoners to implement the workflow system as well as the data and component catalogues. This research demonstrates the use of artificial intelligence techniques to support the kinds of automation envisioned by the scientific community for large-scale distributed scientific data analysis. Yolanda Gil, Pedro A. González-Calero, Jihie Kim, Joshua Moody, Varun Ratnakar |
J. Exp. Theor. Artif. Intell. | 1 |
| 2011 | Using provenance in the Semantic Web
Yolanda Gil, Paul Groth |
J. Web Semant. | 1 |
| 2011 | Shortipedia aggregating and curating Semantic Web data
Denny Vrandecic, Varun Ratnakar, Markus Krötzsch, Yolanda Gil |
J. Web Semant. | 4 |
| 2010 | Principles for interactive acquisition and validation of workflowsabstractWorkflows, also known as process models, are essential in many science and engineering fields. Workflows express compositions of individual steps or tasks that assembled together account for various aspects of an overall process. When workflows include dozens of components and many links among them, the creation of valid workflows becomes challenging since users have to track many interdependencies and constraints. This article describes principles for assisting users to create valid workflows that are based on two knowledge acquisition systems that we have developed. A shared goal in these projects was to enable end users who do not have computer science backgrounds, such as biologists, military officers or engineers, to create valid end-to-end process models or workflows. Our approach exploits knowledge-rich descriptions of the individual components and their constraints in order to validate the composition, and uses artificial intelligence planning techniques in order to systematically verify formal properties of valid workflows. Both systems analyse partial workflows created by the user, determine whether they are consistent with the background knowledge that the system has, notifies the user of issues to be resolved in the current workflow, and suggests to the user what actions could be taken to correct those issues. Jihie Kim, Yolanda Gil, Marc Spraragen |
J. Exp. Theor. Artif. Intell. | 2 |
| 2009 | Expressive Reusable Workflow TemplatesabstractWorkflow systems can manage complex scientific applications with distributed data processing. Although some workflow systems can represent collections of data with very compact abstractions and manage their execution efficiently, there are no approaches to date to manage collections of application components required to express some scientific applications. We present an approach to handle collections of components and data alike in expressive workflow templates whose basic structure is reusable. We also present an algorithm that can elaborate abstract compact workflow templates into execution-ready workflows that enumerate all computations to be carried out. We implemented the proposed approach in the Wings workflow system. Our work is motivated by real-world complex scientific applications that require handling of nested collections of both components and data. Yolanda Gil, Paul Groth, Varun Ratnakar, Christian Fritz 0001 |
eScience | 1 |
| 2009 | An integrated framework for performance-based optimization of scientific workflowsabstractData analysis processes in scientific applications can be expressed as coarse-grain workflows of complex data processing operations with data flow dependencies between them. Performance optimization of these workflows can be viewed as a search for a set of optimal values in a multi-dimensional parameter space. While some performance parameters such as grouping of workflow components and their mapping to machines do not a ect the accuracy of the output, others may dictate trading the output quality of individual components (and of the whole workflow) for performance. This paper describes an integrated framework which is capable of supporting performance optimizations along multiple dimensions of the parameter space. Using two real-world applications in the spatial data analysis domain, we present an experimental evaluation of the proposed framework. Vijay S. Kumar, P. Sadayappan, Gaurang Mehta, Karan Vahi, Ewa Deelman, Varun Ratnakar, Jihie Kim, Yolanda Gil, Mary W. Hall, Tahsin M. Kurç, Joel H. Saltz |
HPDC | 8 |
| 2009 | A scientific workflow construction command lineabstractWorkflows have emerged as a common tool for scientists to express their computational analyses. While there are a multitude of visual data flow editors for workflow construction, to date there are none that support the input of workflows using natural language. This work presents the design of a hybrid system that combines natural language input through a command line with a visual editor. Paul Groth, Yolanda Gil |
IUI | 2 |
| 2009 | Workflow matching using semantic metadataabstractWorkflows are becoming an increasingly more common paradigm to manage scientific analyses. As workflow repositories start to emerge, workflow retrieval and discovery becomes a challenge. Studies have shown that scientists wish to discover workflows given properties of workflow data inputs, intermediate data products, and data results. However, workflows typically lack this information when contributed to a repository. Our work addresses this issue by augmenting workflow descriptions with constraints derived from properties about the workflow components used to process data as well as the data itself. An important feature of our approach is that it assumes that component and data properties are obtained from catalogs that are external to the workflow system, consistent with current architectures for computational science. Yolanda Gil, Jihie Kim, Gonzalo Flórez Puga, Varun Ratnakar, Pedro A. González-Calero |
K-CAP | 1 |
| 2008 | Automating To-Do Lists for Users: Interpretation of To-Dos for Selecting and Tasking Agents
Yolanda Gil, Varun Ratnakar |
AAAI | 1 |
| 2008 | Designing and parameterizing a workflow for optimization: A case study in biomedical imagingabstractThis paper describes our experience to date employing the systematic mapping and optimization of large- scale scientific application workflows to current and future parallel platforms. The overall goal of the project is to integrate a set of system layers - application program, compiler, run-time environment, knowledge representation, optimization framework, and workflow manager - and through a systematic strategy for workflow mapping, our approach will exploit the vast machine resources available in such parallel platforms to dramatically increase the productivity of application programmers. In this paper, we describe the representation of a biomedical imaging application as a workflow, our early experiences in integrating the set of tools brought together for this project, and implications for future applications. Vijay S. Kumar, Mary W. Hall, Jihie Kim, Yolanda Gil, Tahsin M. Kurç, Ewa Deelman, Varun Ratnakar, Joel H. Saltz |
IPDPS | 4 |
| 2008 | Towards intelligent assistance for to-do listsabstractAssisting users with to-do lists presents new challenges for intelligent user interfaces. This paper presents a detailed analysis of to-do list entries jotted by users of a system that automates tasks for users that we would like to extend to assist users with their to-do entries. We also present four distinct stages of interpretation of to-do entries that can be accomplished and evaluated separately. A system that has good performance in any of these four stages can provide intelligent assistance that is useful to users. Author Keywords User interfaces, to-do lists, automated assistance, natural language interpretation, knowledge acquisition, knowledge collection from web volunteers, office assistants. ACM Classification Keywords H5.m. Information interfaces and presentation (e.g., HCI): Yolanda Gil, Varun Ratnakar |
IUI | 1 |
| 2008 | Provenance trails in the Wings/Pegasus systemabstractAbstract Our research focuses on creating and executing large‐scale scientific workflows that often involve thousands of computations over distributed, shared resources. We describe an approach to workflow creation and refinement that uses semantic representations to (1) describe complex scientific applications in a data‐independent manner, (2) automatically generate workflows of computations for given data sets, and (3) map the workflows to available computing resources for efficient execution. Our approach is implemented in the Wings/Pegasus workflow system and has been demonstrated in a variety of scientific application domains. This paper illustrates the application‐level provenance information generated Wings during workflow creation and the refinement provenance by the Pegasus mapping system for execution over grid computing environments. We show how this information is used in answering the queries of the First Provenance Challenge. Copyright © 2007 John Wiley & Sons, Ltd. Jihie Kim, Ewa Deelman, Yolanda Gil, Gaurang Mehta, Varun Ratnakar |
Concurr. Comput. Pract. Exp. | 3 |
| 2008 | Special Issue: The First Provenance ChallengeabstractAbstract The first Provenance Challenge was set up in order to provide a forum for the community to understand the capabilities of different provenance systems and the expressiveness of their provenance representations. To this end, a functional magnetic resonance imaging workflow was defined, which participants had to either simulate or run in order to produce some provenance representation, from which a set of identified queries had to be implemented and executed. Sixteen teams responded to the challenge, and submitted their inputs. In this paper, we present the challenge workflow and queries, and summarize the participants' contributions. Copyright © 2007 John Wiley & Sons, Ltd. Luc Moreau 0001, Bertram Ludäscher, Ilkay Altintas, Roger S. Barga, Shawn Bowers, Steven P. Callahan, George Chin, Ben Clifford, Shirley Cohen, Sarah Cohen Boulakia, Susan B. Davidson, Ewa Deelman, Luciano A. Digiampietri, Ian T. Foster, Juliana Freire, James Frew, Joe Futrelle, Tara Gibson, Yolanda Gil, Carole A. Goble, Jennifer Golbeck, Paul Groth, David A. Holland, Jihie Kim, David Koop, Ales Krenek, Timothy M. McPhillips, Gaurang Mehta, Simon Miles, Dominic Metzger, Steve Munroe, James D. Myers, Beth Plale, Norbert Podhorszki, Varun Ratnakar, Emanuele Santos, Carlos Scheidegger, Karen Schuchardt, Margo I. Seltzer, Yogesh L. Simmhan, Cláudio T. Silva, Peter Slaughter, Eric G. Stephan, Robert Stevens 0001, Daniele Turi, Huy T. Vo, Michael Wilde, Jun Zhao 0003, Yong Zhao 0009 |
Concurr. Comput. Pract. Exp. | 19 |
| 2008 | Self-Configuring Applications for Heterogeneous Systems: Program Composition and Optimization Using Cognitive TechniquesabstractThis paper describes several challenges facing programmers of future edge computing systems, the diverse many-core devices that will soon exemplify commodity mainstream systems. To call attention to programming challenges ahead, this paper focuses on the most complex of such architectures: integrated, power-conserving systems, inherently parallel and heterogeneous, with distributed address spaces. When programming such complex systems, new concerns arise: computation partitioning across functional units, data movement and synchronization, managing a diversity of programming models for different devices, and reusing existing legacy and library software. We observe that many of these challenges are also faced in programming applications for large-scale heterogeneous distributed computing environments, and current solutions as well as future research directions in distributed computing can be adapted to commodity computing environments. Optimization decisions are inherently complex due to large search spaces of possible solutions and the difficulty of predicting performance on increasingly complex architectures. Cognitive techniques are well suited for managing systems of such complexity, citing recent trends of using cognitive techniques for code mapping and optimization support. Combining these, we describe a fundamentally new programming paradigm for complex heterogeneous systems, where programmers design self-configuring applications and the system automates optimization decisions and manages the allocation of heterogeneous resources. Mary W. Hall, Yolanda Gil, Robert F. Lucas |
Proc. IEEE | 2 |
| 2007 | Wings for Pegasus: Creating Large-Scale Scientific Applications Using Semantic Representations of Computational Workflows
Yolanda Gil, Varun Ratnakar, Ewa Deelman, Gaurang Mehta, Jihie Kim |
AAAI | 1 |
| 2007 | Intelligent Optimization of Parallel and Distributed ApplicationsabstractThis paper describes a new project that systematically addresses the enormous complexity of mapping applications to current and future parallel platforms. By integrating the system layers - domain-specific environment, application program, compiler, run-time environment, performance models and simulation, and workflow manager - and through a systematic strategy for application mapping, our approach exploit the vast machine resources available in such parallel platforms to dramatically increase the productivity of application programmers. This project brings together computer scientists in the areas represented by the system layers (i.e., language extensions, compilers, run-time systems, workflows) together with expertise in knowledge representation and machine learning. With expert domain scientists in molecular dynamics (MD) simulation, we are developing our approach in the context of a specific application class which already targets environments consisting of several hundreds of processors. In this way, we gain valuable insight into a generalizable strategy, while simultaneously producing performance benefits for existing and important applications. Bhupesh Bansal, Ümit V. Çatalyürek, Jacqueline Chame, Chun Chen 0002, Ewa Deelman, Yolanda Gil, Mary W. Hall, Vijay S. Kumar, Tahsin M. Kurç, Kristina Lerman, Aiichiro Nakano, Yoon-Ju Lee Nelson, Joel H. Saltz, Ashish Sharma 0001, Priya Vashishta |
IPDPS | 6 |
| 2007 | Incorporating tutoring principles into interactive knowledge acquisition
Jihie Kim, Yolanda Gil |
Int. J. Hum. Comput. Stud. | 2 |
| 2007 | A survey of trust in computer science and the Semantic Web
Donovan Artz, Yolanda Gil |
J. Web Semant. | 2 |
| 2007 | Towards content trust of web resources
Yolanda Gil, Donovan Artz |
J. Web Semant. | 1 |
| 2007 | Introduction to the special issue of JWS with selected papers from ISWC 2005
Yolanda Gil, Enrico Motta |
J. Web Semant. | 1 |
| 2006 | Managing Large-Scale Scientific Workflows in Distributed Environments: Experiences and ChallengesabstractIn this paper we discuss several challenges associated scientific workflow design and management in distributed, heterogeneous environments. Based on our prior work with a number of scientific applications, we describe the workflow lifecycle and examine our experiences and the challenges ahead as they pertain to the user experience, planning the workflow execution and managing the execution itself. Ewa Deelman, Yolanda Gil |
e-Science | 2 |
| 2006 | Semantic Metadata Generation for Large Scientific Workflows
Jihie Kim, Yolanda Gil, Varun Ratnakar |
ISWC | 2 |
| 2006 | Towards content trust of web resourcesabstractTrust is an integral part of the Semantic Web architecture. While most prior work focuses on entity-centered issues such as authentication and reputation, it does not model the content, i.e. the nature and use of the information being exchanged. This paper discusses content trust as an aggregate of other trust measures that have been previously studied. The paper introduces several factors that users consider in deciding whether to trust the content provided by a Web resource. Many of these factors are hard to capture in practice, since they would require a large amount of user input. Our goal is to discern which of these factors could be captured in practice with minimal user interaction in order to maximize the system's trust estimates. The paper also describes a simulation environment that we have designed to study alternative models of content trust. Yolanda Gil, Donovan Artz |
WWW | 1 |
| 2006 | On agents and grids: Creating the fabric for a new generation of distributed intelligent systems
Yolanda Gil |
J. Web Semant. | 1 |
| 2005 | An Analysis of Knowledge Collected from Volunteer Contributors
Timothy Chklovski, Yolanda Gil |
AAAI | 2 |
| 2005 | Task scheduling strategies for workflow-based applications in gridsabstractGrid applications require allocating a large number of heterogeneous tasks to distributed resources. A good allocation is critical for efficient execution. However, many existing grid toolkits use matchmaking strategies that do not consider overall efficiency for the set of tasks to be run. We identify two families of resource allocation algorithms: task-based algorithms, that greedily allocate tasks to resources, and workflow-based algorithms, that search for an efficient allocation for the entire workflow. We compare the behavior of workflow-based algorithms and task-based algorithms, using simulations of workflows drawn from a real application and with varying ratios of computation cost to data transfer cost. We observe that workflow-based approaches have a potential to work better for data-intensive applications even when estimates about future tasks are inaccurate. Jim Blythe, Ewa Deelman, Yolanda Gil, Karan Vahi, Anirban Mandal, Ken Kennedy |
CCGRID | 4 |
| 2005 | User interfaces with semi-formal representations: a study of designing argumentation structuresabstractWhen designing mixed-initiative systems, full formalization of all potentially relevant knowledge may not be cost-effective or practical. This paper motivates the need for semi-formal representations that combine machine-processable structures with free text statements, and discusses the need to design them in a way that makes the free text more amenable to automated structuring and processing. Our work is done in the context of argumentation systems, and has explored a range of tradeoffs in combining informal free-text statements with formal connectors. The paper compares alternative argument representations which combine structured argument connectors with free text. We discuss merits of the systems based on a variety of analysis structures that we have collected from Web users to date. Timothy Chklovski, Varun Ratnakar, Yolanda Gil |
IUI | 3 |
| 2005 | Improving the design of intelligent acquisition interfaces for collecting world knowledge from web contributorsabstractAn emerging approach to knowledge acquisition is to collect statements from volunteer contributors over the Web. In this approach, the design of the acquisition interface is key to focusing on statements of interest, avoiding spurious entries, retaining the contributors, etc. Several such volunteer-contribution-based systems have been deployed to date, each with its own idiosyncratic interface. This paper discusses some key challenges faced by volunteer collection interfaces, and outlines the design features that we have found effective in addressing some aspects of those challenges. The paper discusses how these features have been implemented in deployed collection systems, and reflects on the data collected to extract lessons for future work in this research area. Timothy Chklovski, Yolanda Gil |
K-CAP | 2 |
| 2004 | Artemis: Integrating Scientific Data on the Grid
Rattapoom Tuchinda, Snehal Thakkar, Yolanda Gil, Ewa Deelman |
AAAI | 3 |
| 2004 | An intelligent assistant for interactive workflow compositionabstractComplex applications in many areas, including scientific computations and business-related web services, are created from collections of components to form computational workflows. In many cases end users have requirements and preferences that depend on how the workflow unfolds, and that cannot be specified beforehand. Workflow editors enable users to formulate workflows, but the editors need to be augmented with intelligent assistance in order to help users in several key aspects of the task, namely: 1) keeping track of detailed constraints across selected components and their connections; 2) specifying the workflow flexibly, e.g., top-down, bottom-up, from requirements, or from available data; and 3) taking partial or incomplete descriptions of workflows and understanding the steps needed for their completion. We present an approach that combines knowledge bases (that have rich representations of components) together with planning techniques (that can track the relations and constraints among individual steps). We illustrate the approach with an implemented system called CAT (Composition Analysis Tool) that analyzes workflows and generates error messages and suggestions in order to help users compose complete and consistent workflows. Jihie Kim, Marc Spraragen, Yolanda Gil |
IUI | 3 |
| 2004 | Incremental formalization of document annotations through ontology-based paraphrasingabstractFor the manual semantic markup of documents to become wide-spread, usersmust be able to express annotations that conform to ontologies (orschemas) that have shared meaning. However, a typical user is unlikelyto be familiar with the details of the terms as defined by the ontology authors. In addition, the idea to be expressed may not fit perfectly within a pre-defined ontology. The ideal tool should help users find apartial formalization that closely follows the ontology where possiblebut deviates from the formal representation where needed. We describe animplemented approach to help users create semi-structured semantic annotations for a document according to an extensible OWL ontology. In our approach, users enter a short sentence in free text to describe allor part of a document, and the system presents a set of potential paraphrases of the sentence that are generated from valid expressions inthe ontology, from which the user chooses the closest match. We use a combination of off-the-shelf parsing tools and breadth-first search of expressions in the ontology to help users create valid annotations starting from free text. The user can also define new terms to augmentthe ontology, so the potential matches can improve over time. Jim Blythe, Yolanda Gil |
WWW | 2 |
| 2004 | A short study on the success of the Gene Ontology
Michael Bada, Robert Stevens 0001, Carole A. Goble, Yolanda Gil, Michael Ashburner, Judith A. Blake, J. Michael Cherry, Midori A. Harris, Suzanna Lewis |
J. Web Semant. | 4 |
| 2003 | A Knowledge Acquisition Tool for Course of Action Analysis
Kim Barker, Jim Blythe, Gary C. Borchardt, Vinay K. Chaudhri, Peter Clark, Paul R. Cohen, Julie Fitzgerald, Kenneth D. Forbus, Yolanda Gil, Boris Katz, Jihie Kim, Gary W. King, Sunil Mishra, Clayton T. Morrison, Kenneth S. Murray, Charley Otstott, Bruce W. Porter, Robert Schrag, Tomás E. Uribe, Jeffrey M. Usher, Peter Z. Yeh |
IAAI | 9 |
| 2003 | Transparent Grid Computing: A Knowledge-Based Approach
Jim Blythe, Ewa Deelman, Yolanda Gil, Carl Kesselman |
IAAI | 3 |
| 2003 | Proactive Dialogue for Interactive Knowledge Capture
Jihie Kim, Yolanda Gil |
IJCAI | 2 |
| 2003 | Mapping Abstract Complex Workflows onto Grid Environments
Ewa Deelman, Jim Blythe, Yolanda Gil, Carl Kesselman, Gaurang Mehta, Karan Vahi, Kent Blackburn, Albert Lazzarini, Adam Arbree, Richard Cavanaugh, Scott Koranda |
J. Grid Comput. | 3 |
| 2002 | IKRAFT: Interactive Knowledge Representation and Acquisition from Text
Yolanda Gil, Varun Ratnakar |
EKAW | 1 |
| 2002 | TRELLIS: An Interactive Tool for Capturing Information Analysis and Decision Making
Yolanda Gil, Varun Ratnakar |
EKAW | 1 |
| 2002 | Deriving Acquisition Principles from Tutoring Principles
Jihie Kim, Yolanda Gil |
Intelligent Tutoring Systems | 2 |
| 2002 | Trusting Information Sources One Citizen at a Time
Yolanda Gil, Varun Ratnakar |
ISWC | 1 |
| 2001 | Electric Elves: Applying Agent Technology to Support Human Organizations
Hans Chalupsky, Yolanda Gil, Craig A. Knoblock, Kristina Lerman, Jean Oh, David V. Pynadath, Thomas A. Russ, Milind Tambe |
IAAI | 2 |
| 2001 | Knowledge Analysis on Process Models
Jihie Kim, Yolanda Gil |
IJCAI | 2 |
| 2001 | An integrated environment for knowledge acquisitionabstractThis paper describes an integrated acquisition interface that includes several techniques previously developed to support users in various ways as they add new knowledge to an intelligent system. As a result of this integration, the individual techniques can take better advantage of the context in which they are invoked and provide stronger guidance to users. We describe the current implementation using examples from a travel planning domain, and demonstrate how users can add complex knowledge to the system. Jim Blythe, Jihie Kim, Surya Ramachandran, Yolanda Gil |
IUI | 4 |
| 2001 | Knowledge entry as the graphical assembly of componentsabstractDespite some successes, the lack of tools to allow subject matter experts to directly enter, query, and debug formal domain knowledge in a knowledge-base still remains a major obstacle to their deployment. Our goal is to create such tools, so that a trained knowledge engineer is no longer required to mediate the interaction. This paper presents our work on the knowledge entry part of this overall knowledge capture task, which is based on several claims: that users can construct representations by connecting pre-fabricated, representational components, rather than writing low-level axioms; that these components can be presented to users as graphs; and the user can then perform composition through graph manipulation operations. To operationalize this, we have developed a novel technique of graphical dialog using examples of the component concepts, followed by an automated process for generalizing the user's graphically-entered assertions into axioms. We present these claims, our approach, the system (called SHAKEN) that we are developing, and an evaluation of our progress based on having users encode knowledge using the system. Keywords Graphical knowledge entry, knowledge acquisition, components, composition, knowledge-based systems. Peter Clark, John A. Thompson, Ken Barker 0002, Bruce W. Porter, Vinay K. Chaudhri, Andres C. Rodriguez, Jérôme Thoméré, Sunil Mishra, Yolanda Gil, Patrick J. Hayes, Thomas Reichherzer |
K-CAP | 9 |
| 2001 | User studies of knowledge acquisition tools: methodology and lessons learnedabstractKnowledge acquisition research concerned with the development of knowledge acquisition tools is in need of a methodological approach to evaluation. This paper describes experimental methodology to conduct studies and experiments of users modifying knowledge bases with knowledge acquisition tools. The paper also reports on the lessons learned from several experiments that have been performed using this methodology. The hope is that it will help others design user evaluations of knowledge acquisition tools. Ideas are discussed for improving the current methodology and some open issues that remain. Marcelo Tallis, Jihie Kim, Yolanda Gil |
J. Exp. Theor. Artif. Intell. | 3 |
| 2000 | User studies of an interdependency-based interface for acquiring problem-solving knowledgeabstractThis paper describes a series of experiments with a range of users to evaluate an intelligent interface for acquiring problem-solving knowledge to describe how to accomplish a task. The tool derives the interdependencies between different pieces of knowledge in the system and uses them to guide the user in completing the acquisition task. The paper describes results obtained when the tool was tested with a wide range of users, including end users. The studies show that our acquisition interface saves users an average of 32% of the time it takes to add new knowledge, and highlight some interesting differences across user groups. The paper also describes what are the areas that need to be addressed in future research in order to make these tools usable by end users. Jihie Kim, Yolanda Gil |
IUI | 2 |
| 1999 | IUI and Agents for the New Millennium (Panel)abstractNo abstract available. Henry Lieberman, Jeffrey M. Bradshaw, Yolanda Gil, Ted Selker |
IUI | 3 |
| 1994 | Knowledge Refinement in a Reflective Architecture
Yolanda Gil |
AAAI | 1 |
| 1994 | Learning by Experimentation: Incremental Refinement of Incomplete Planning Domains
Yolanda Gil |
ICML | 1 |
| 1993 | Efficient Domain-Independent Experimentation
Yolanda Gil |
ICML | 1 |
| 1991 | A Domain-Independent Framework for Effective Experimentation in Planning
Yolanda Gil |
ML | 1 |
| 1989 | Explanation-Based Learning: A Problem Solving Perspective
Steven Minton, Jaime G. Carbonell, Craig A. Knoblock, Daniel Kuokka, Oren Etzioni, Yolanda Gil |
Artif. Intell. | 6 |