VLDB 2026 Research / reviewers in the wild / expert
Natasha F. Noy
dblp:n/NatalyaFridmanNoy · also Natalya Fridman Noy, Natasha Fridman Noy
· DBLP profile ↗
67ranked-venue papers
13as first author
3since 2021 · last 2024
0000-0002-7437-0624ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 38 · 10 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 2 first-authorHuman-computer interaction and ubiquitous computing · 6 · 1 first-authorArtificial intelligence and machine learning · 4 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Relationships Are Complicated! An Analysis of Relationships Between Datasets on the Web
Kate Lin, Tarfah Alrashed, Natasha F. Noy |
ISWC (1) | 3 |
| 2023 | Will LLMs reshape, supercharge, or kill data science?abstractLarge language models (LLMs) have recently taken the world by storm, promising potentially game changing opportunities in multiple fields. Naturally, there is significant promise in applying LLMs to the management of structured data, or more generally, to the processes involved in data science. At the very least, LLMs have the potential to provide substantial advancements in long-standing challenges that our community has been tackling for decades. On the other hand, they may introduce completely new capabilities that we have only dreamed of thus far. This panel will bring together a few leading experts who have been thinking about these opportunities from various perspectives and fielding them in research prototypes and even in commercial applications. Alon Y. Halevy, Yejin Choi 0001, Avrilia Floratou, Michael J. Franklin, Natasha F. Noy, Haixun Wang |
Proc. VLDB Endow. | 5 |
| 2021 | Dataset or Not? A Study on the Veracity of Semantic Markup for Dataset PagesabstractAbstract Semantic markup, such as , allows providers on the Web to describe content using a shared controlled vocabulary. This markup is invaluable in enabling a broad range of applications, from vertical search engines, to rich snippets in search results, to actions on emails, to many others. In this paper, we focus on semantic markup for datasets, specifically in the context of developing a vertical search engine for datasets on the Web, Google’s Dataset Search. Dataset Search relies on to identify pages that describe datasets. While was the core enabling technology for this vertical search, we also discovered that we need to address the following problem: pages from 61% of internet hosts that provide markup do not actually describe datasets. We analyze the veracity of dataset markup for Dataset Search’s Web-scale corpus and categorize pages where this markup is not reliable. We then propose a way to drastically increase the quality of the dataset metadata corpus by developing a deep neural-network classifier that identifies whether or not a page with markup is a dataset page. Our classifier achieves 96.7% recall at the 95% precision point. This level of precision enables Dataset Search to circumvent the noise in semantic markup and to use the metadata to provide high quality results to users. Tarfah Alrashed, Dimitris Paparas, Omar Benjelloun, Ying Sheng 0002, Natasha F. Noy |
ISWC | 5 |
| 2020 | Google Dataset Search by the Numbers
Omar Benjelloun, Natasha F. Noy |
ISWC (2) | 3 |
| 2020 | When the Web is your Data Lake: Creating a Search Engine for Datasets on the WebabstractThere are thousands of data repositories on the Web, providing access to millions of datasets. National and regional governments, scientific publishers and consortia, commercial data providers, and others publish data for fields ranging from social science to life science to high-energy physics to climate science and more. Access to this data is critical to facilitating reproducibility of research results, enabling scientists to build on others' work, and providing data journalists easier access to information and its provenance. In this talk, I will discuss our work on Dataset Search, which provides search capabilities over potentially all dataset repositories on the Web. I will talk about the open ecosystem for describing and citing datasets that we hope to encourage and the technical details on how we went about building Dataset Search. Finally, I will highlight research challenges in building a vibrant, heterogeneous, and open ecosystem where data becomes a first-class citizen. Natasha F. Noy |
SIGMOD Conference | 1 |
| 2019 | Google Dataset Search: Building a search engine for datasets in an open Web ecosystemabstractThere are thousands of data repositories on the Web, providing access to millions of datasets. National and regional governments, scientific publishers and consortia, commercial data providers, and others publish data for fields ranging from social science to life science to high-energy physics to climate science and more. Access to this data is critical to facilitating reproducibility of research results, enabling scientists to build on others' work, and providing data journalists easier access to information and its provenance. In this paper, we discuss Google Dataset Search, a dataset-discovery tool that provides search capabilities over potentially all datasets published on the Web. The approach relies on an open ecosystem, where dataset owners and providers publish semantically enhanced metadata on their own sites. We then aggregate, normalize, and reconcile this metadata, providing a search engine that lets users find datasets in the “long tail” of the Web. In this paper, we discuss both social and technical challenges in building this type of tool, and the lessons that we learned from this experience. Dan Brickley, Matthew Burgess, Natasha F. Noy |
WWW | 3 |
| 2016 | Goods: Organizing Google's DatasetsabstractEnterprises increasingly rely on structured datasets to run their businesses. These datasets take a variety of forms, such as structured files, databases, spreadsheets, or even services that provide access to the data. The datasets often reside in different storage systems, may vary in their formats, may change every day. In this paper, we present GOODS, a project to rethink how we organize structured datasets at scale, in a setting where teams use diverse and often idiosyncratic ways to produce the datasets and where there is no centralized system for storing and querying them. GOODS extracts metadata ranging from salient information about each dataset (owners, timestamps, schema) to relationships among datasets, such as similarity and provenance. It then exposes this metadata through services that allow engineers to find datasets within the company, to monitor datasets, to annotate them in order to enable others to use their datasets, and to analyze relationships between them. We discuss the technical challenges that we had to overcome in order to crawl and infer the metadata for billions of datasets, to maintain the consistency of our metadata catalog at scale, and to expose the metadata to users. We believe that many of the lessons that we learned are applicable to building large-scale enterprise-level data-management systems in general. Alon Y. Halevy, Flip Korn, Natasha F. Noy, Christopher Olston, Neoklis Polyzotis, Sudip Roy 0002, Steven Euijong Whang |
SIGMOD Conference | 3 |
| 2016 | Discovering Structure in the Universe of Attribute NamesabstractRecently, search engines have invested significant effort to answering entity--attribute queries from structured data, but have focused mostly on queries for frequent attributes. In parallel, several research efforts have demonstrated that there is a long tail of attributes, often thousands per class of entities, that are of interest to users. Researchers are beginning to leverage these new collections of attributes to expand the ontologies that power search engines and to recognize entity--attribute queries. Because of the sheer number of potential attributes, such tasks require us to impose some structure on this long and heavy tail of attributes. This paper introduces the problem of organizing the attributes by expressing the compositional structure of their names as a rule-based grammar. These rules offer a compact and rich semantic interpretation of multi-word attributes, while generalizing from the observed attributes to new unseen ones. The paper describes an unsupervised learning method to generate such a grammar automatically from a large set of attribute names. Experiments show that our method can discover a precise grammar over 100,000 attributes of {\sc Countries} while providing a 40-fold compaction over the attribute names. Furthermore, our grammar enables us to increase the precision of attributes from 47\% to more than 90\% with only a minimal curation effort. Thus, our approach provides an efficient and scalable way to expand ontologies with attributes of user interest. Alon Y. Halevy, Natasha F. Noy, Sunita Sarawagi, Steven Euijong Whang |
WWW | 2 |
| 2015 | Discovering Subsumption Relationships for Web-Based OntologiesabstractAs search engines are becoming smarter at interpreting user queries and providing meaningful responses, they rely on ontologies to understand the meaning of entities. Creating ontologies manually is a laborious process, and resulting ontologies may not reflect the way users think about the world, as many concepts used in queries are noisy, and not easily amenable to formal modeling. There has been considerable effort in generating ontologies from Web text and query streams, which may be more reflective of how users query and write content. In this paper, we describe the LATTE system that automatically generates a subconcept--superconcept hierarchy, which is critical for using ontologies to answer queries. LATTE combines signals based on word-vector representations of concepts and dependency parse trees; however, LATTE derives most of its power from an ontology of attributes extracted from the Web that indicates the aspects of concepts that users find important. LATTE achieves an F1 score of 74%, which is comparable to expert agreement on a similar task. We additionally demonstrate the usefulness of LATTE in detecting high quality concepts from an existing resource of IsA links. Dana Movshovitz-Attias, Steven Euijong Whang, Natasha F. Noy, Alon Y. Halevy |
WebDB | 3 |
| 2015 | How to apply Markov chains for modeling sequential edit patterns in collaborative ontology-engineering projects
Simon Walk, Philipp Singer, Markus Strohmaier, Denis Helic, Natasha F. Noy, Mark A. Musen |
Int. J. Hum. Comput. Stud. | 5 |
| 2015 | Using the wisdom of the crowds to find critical errors in biomedical ontologies: a study of SNOMED CTabstractOBJECTIVES: The verification of biomedical ontologies is an arduous process that typically involves peer review by subject-matter experts. This work evaluated the ability of crowdsourcing methods to detect errors in SNOMED CT (Systematized Nomenclature of Medicine Clinical Terms) and to address the challenges of scalable ontology verification. METHODS: We developed a methodology to crowdsource ontology verification that uses micro-tasking combined with a Bayesian classifier. We then conducted a prospective study in which both the crowd and domain experts verified a subset of SNOMED CT comprising 200 taxonomic relationships. RESULTS: The crowd identified errors as well as any single expert at about one-quarter of the cost. The inter-rater agreement (κ) between the crowd and the experts was 0.58; the inter-rater agreement between experts themselves was 0.59, suggesting that the crowd is nearly indistinguishable from any one expert. Furthermore, the crowd identified 39 previously undiscovered, critical errors in SNOMED CT (eg, 'septic shock is a soft-tissue infection'). DISCUSSION: The results show that the crowd can indeed identify errors in SNOMED CT that experts also find, and the results suggest that our method will likely perform well on similar ontologies. The crowd may be particularly useful in situations where an expert is unavailable, budget is limited, or an ontology is too large for manual error checking. Finally, our results suggest that the online anonymous crowd could successfully complete other domain-specific tasks. CONCLUSIONS: We have demonstrated that the crowd can address the challenges of scalable ontology verification, completing not only intuitive, common-sense tasks, but also expert-level, knowledge-intensive tasks. Jonathan Mortensen, Evan P. Minty, Michael Januszyk, Timothy E. Sweeney, Alan L. Rector, Natasha F. Noy, Mark A. Musen |
J. Am. Medical Informatics Assoc. | 6 |
| 2014 | Reasoning Based Quality Assurance of Medical Ontologies: A Case Study
Matthew Horridge, Bijan Parsia, Natasha F. Noy, Mark A. Musen |
AMIA | 3 |
| 2014 | An empirically derived taxonomy of errors in SNOMED CT
Jonathan Mortensen, Mark A. Musen, Natasha F. Noy |
AMIA | 3 |
| 2014 | WebProtégé: a collaborative Web-based platform for editing biomedical ontologiesabstractUNLABELLED: WebProtégé is an open-source Web application for editing OWL 2 ontologies. It contains several features to aid collaboration, including support for the discussion of issues, change notification and revision-based change tracking. WebProtégé also features a simple user interface, which is geared towards editing the kinds of class descriptions and annotations that are prevalent throughout biomedical ontologies. Moreover, it is possible to configure the user interface using views that are optimized for editing Open Biomedical Ontology (OBO) class descriptions and metadata. Some of these views are shown in the Supplementary Material and can be seen in WebProtégé itself by configuring the project as an OBO project. AVAILABILITY AND IMPLEMENTATION: WebProtégé is freely available for use on the Web at http://webprotege.stanford.edu. It is implemented in Java and JavaScript using the OWL API and the Google Web Toolkit. All major browsers are supported. For users who do not wish to host their ontologies on the Stanford servers, WebProtégé is available as a Web app that can be run locally using a Servlet container such as Tomcat. Binaries, source code and documentation are available under an open-source license at http://protegewiki.stanford.edu/wiki/WebProtege. Matthew Horridge, Tania Tudorache, Csongor Nyulas, Jennifer Vendetti, Natasha F. Noy, Mark A. Musen |
Bioinform. | 5 |
| 2014 | Discovering Beaten Paths in Collaborative Ontology-Engineering Projects using Markov Chains
Simon Walk, Philipp Singer, Markus Strohmaier, Tania Tudorache, Mark A. Musen, Natasha F. Noy |
J. Biomed. Informatics | 6 |
| 2013 | A Family-Based Framework for Supporting Quality Assurance of Biomedical Ontologies in BioPortal
Zhe He 0001, Christopher Ochs, Ankur Agrawal, Yehoshua Perl, Dimitris Zeginis, Konstantinos A. Tarabanis, Gai Elhanan, Michael Halper, Natasha F. Noy, James Geller |
AMIA | 9 |
| 2013 | Crowdsourcing the Verification of Relationships in Biomedical Ontologies
Jonathan Mortensen, Mark A. Musen, Natasha F. Noy |
AMIA | 3 |
| 2013 | Indented Tree or Graph? A Usability Study of Ontology Visualization Techniques in the Context of Class Mapping Evaluation
Bo Fu 0005, Natasha F. Noy, Margaret-Anne D. Storey |
ISWC (1) | 2 |
| 2013 | Simplified OWL Ontology Editing for the Web: Is WebProtégé Enough?
Matthew Horridge, Tania Tudorache, Jennifer Vendetti, Csongor Nyulas, Mark A. Musen, Natasha F. Noy |
ISWC (1) | 6 |
| 2013 | Getting Lucky in Ontology Search: A Data-Driven Evaluation Framework for Ontology Ranking
Natasha F. Noy, Paul R. Alexander, Rave Harpaz, Patricia L. Whetzel, Ray W. Fergerson, Mark A. Musen |
ISWC (1) | 1 |
| 2013 | Using Semantic Web in ICD-11: Three Years Down the Road
Tania Tudorache, Csongor Nyulas, Natasha F. Noy, Mark A. Musen |
ISWC (2) | 3 |
| 2013 | PragmatiX: An Interactive Tool for Visualizing the Creation Process Behind Collaboratively Engineered OntologiesabstractWith the emergence of tools for collaborative ontology engineering, more and more data about the creation process behind collaborative construction of ontologies is becoming available. Today, collaborative ontology engineering tools such as Collaborative Protégé offer rich and structured logs of changes, thereby opening up new challenges and opportunities to study and analyze the creation of collaboratively constructed ontologies. While there exists a plethora of visualization tools for ontologies, they have primarily been built to visualize aspects of the final product (the ontology) and not the collaborative processes behind construction (e.g. the changes made by contributors over time). To the best of the authors’ knowledge, there exists no ontology visualization tool today that focuses primarily on visualizing the history behind collaboratively constructed ontologies. Since the ontology engineering processes can influence the quality of the final ontology, they believe that visualizing process data represents an important stepping-stone towards better understanding of managing the collaborative construction of ontologies in the future. In this application paper, the authors present a tool – PragmatiX – which taps into structured change logs provided by tools such as Collaborative Protégé to visualize various pragmatic aspects of collaborative ontology engineering. The tool is aimed at managers and leaders of collaborative ontology engineering projects to help them in monitoring progress, in exploring issues and problems, and in tracking quality-related issues such as overrides and coordination among contributors. The paper makes the following contributions: (i) They present PragmatiX, a tool for visualizing the creation process behind collaboratively constructed ontologies (ii) the authors illustrate the functionality and generality of the tool by applying it to structured logs of changes of two large collaborative ontology-engineering projects and (iii) they conduct a heuristic evaluation of the tool with domain experts to uncover early design challenges and opportunities for improvement. Finally, the authors hope that this work sparks a new line of research on visualization tools for collaborative ontology engineering projects. Simon Walk, Jan Pöschko, Markus Strohmaier, Keith Andrews, Tania Tudorache, Natasha F. Noy, Csongor Nyulas, Mark A. Musen |
Int. J. Semantic Web Inf. Syst. | 6 |
| 2013 | How ontologies are made: Studying the hidden social dynamics behind collaborative ontology engineering projects
Markus Strohmaier, Simon Walk, Jan Pöschko, Daniel Lamprecht, Tania Tudorache, Csongor Nyulas, Mark A. Musen, Natasha F. Noy |
J. Web Semant. | 8 |
| 2012 | Applications of Ontology Design Patterns in Biomedical Ontologies
Jonathan Mortensen, Matthew Horridge, Mark A. Musen, Natasha F. Noy |
AMIA | 4 |
| 2012 | Deriving an Abstraction Network to Support Quality Assurance in OCRe
Christopher Ochs, Ankur Agrawal, Yehoshua Perl, Michael Halper, Samson W. Tu, Simona Carini, Ida Sim, Natasha F. Noy, Mark A. Musen, James Geller |
AMIA | 8 |
| 2012 | Using SPARQL to Query BioPortal Ontologies and Metadata
Manuel Salvadores, Matthew Horridge, Paul R. Alexander, Ray W. Fergerson, Mark A. Musen, Natasha F. Noy |
ISWC (2) | 6 |
| 2012 | CrowdMap: Crowdsourcing Ontology Alignment with Microtasks
Cristina Sarasua, Elena Simperl, Natasha F. Noy |
ISWC (1) | 3 |
| 2012 | The National Center for Biomedical OntologyabstractThe National Center for Biomedical Ontology is now in its seventh year. The goals of this National Center for Biomedical Computing are to: create and maintain a repository of biomedical ontologies and terminologies; build tools and web services to enable the use of ontologies and terminologies in clinical and translational research; educate their trainees and the scientific community broadly about biomedical ontology and ontology-based technology and best practices; and collaborate with a variety of groups who develop and use ontologies and terminologies in biomedicine. The centerpiece of the National Center for Biomedical Ontology is a web-based resource known as BioPortal. BioPortal makes available for research in computationally useful forms more than 270 of the world's biomedical ontologies and terminologies, and supports a wide range of web services that enable investigators to use the ontologies to annotate and retrieve data, to generate value sets and special-purpose lexicons, and to perform advanced analytics on a wide range of biomedical data. Mark A. Musen, Natasha F. Noy, Nigam H. Shah, Patricia L. Whetzel, Christopher G. Chute, Margaret-Anne D. Storey, Barry Smith 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | Where to publish and find ontologies? A survey of ontology libraries
Mathieu d'Aquin, Natasha F. Noy |
J. Web Semant. | 2 |
| 2011 | A knowledge base driven user interface for collaborative ontology developmentabstractScientists and researchers often use ontologies to describe their data, to share and integrate this data from heterogeneous sources. Ontologies are formal computer models that describe the main concepts and their relationships in a particular domain. Ontologies are usually authored by a community of users with different roles and levels of expertise. To support collaboration among distributed teams and to provision for distinct authoring requirements of each of the user roles and of individual users, we designed a configurable Web-based ontology editor, WebProtege. WebProtege extends Protege, a widely popular ontology editor with more than 150,000 registered users. The user interface layout and configuration for WebProtege is model-based and declarative: we represent it in a knowledge base, with an ontology defining its structure, and linking the interface configuration to the users, their roles, and access policies. We will discuss how the knowledge base driven configuration of the user interface supports the reuse and modularization of layout configurations. Such configuration is also highly flexible and extensible, and is easier to manage than many traditional approaches. Tania Tudorache, Natasha F. Noy, Sean M. Falconer, Mark A. Musen |
IUI | 2 |
| 2011 | An analysis of collaborative patterns in large-scale ontology development projectsabstractToday, distributed teams collaboratively create and maintain more and more ontologies. To support this type of ontology development, software engineers are introducing a new generation of tools. However, we know relatively little about how existing large-scale collaborative ontology development works and what user workflows the tools must support. In this paper, we analyze our experience in supporting several such projects. We describe a visual and interactive project-management tool that we have developed, which helps ontology developers explore historical ontology change and discussion data. We present the results of qualitative and quantitative studies of the collaborative activity associated with three large-scale ontology-development projects. Based on the analysis, we conclude that domain and ontology experts have different patterns of ontology editing behavior, which has important implications for ontology-development tools. Sean M. Falconer, Tania Tudorache, Natasha F. Noy |
K-CAP | 3 |
| 2011 | From mappings to modules: using mappings to identify domain-specific modules in large ontologiesabstractThe problem of ontology modularization is an active area of research in the Semantic Web community. With the emergence and wider use of very large ontologies, in particular in fields such as biomedicine, more and more application developers need to extract meaningful modules of these ontologies to use in their applications. Researchers have also noted that many ontology-maintenance tasks would be simplified if we could extract modules from ontologies. These tasks include ontology matching: If we can separate ontologies into modules based on the topics that these modules cover, we can simplify and improve ontology matching. In this paper, we study a complementary problem: Can we use existing mappings between ontologies to facilitate modularization? We present a novel approach to modularization based on mappings between ontologies. We validate and analyze our approach by applying our methods to identify modules for National Cancer Institutes Thesaurus (NCI Thesaurus) and Systematized Nomenclature of Medicine--Clinical Terms (SNOMED-CT). Amir Ghazvinian, Natasha F. Noy, Mark A. Musen |
K-CAP | 2 |
| 2011 | vSPARQL: A view definition language for the semantic web
Marianne Shaw, Landon Fridman Detwiler, Natasha F. Noy, James F. Brinkley, Dan Suciu |
J. Biomed. Informatics | 3 |
| 2011 | The Biomedical Resource Ontology (BRO) to enable resource discovery in clinical and translational researchabstractThe biomedical research community relies on a diverse set of resources, both within their own institutions and at other research centers. In addition, an increasing number of shared electronic resources have been developed. Without effective means to locate and query these resources, it is challenging, if not impossible, for investigators to be aware of the myriad resources available, or to effectively perform resource discovery when the need arises. In this paper, we describe the development and use of the Biomedical Resource Ontology (BRO) to enable semantic annotation and discovery of biomedical resources. We also describe the Resource Discovery System (RDS) which is a federated, inter-institutional pilot project that uses the BRO to facilitate resource discovery on the Internet. Through the RDS framework and its associated Biositemaps infrastructure, the BRO facilitates semantic search and discovery of biomedical resources, breaking down barriers and streamlining scientific research that will improve human health. Jessica D. Tenenbaum, Patricia L. Whetzel, Kent Anderson, Charles D. Borromeo, Ivo D. Dinov, Davera Gabriel, Beth A. Kirschner, Barbara Mirel, Timothy D. Morris, Natasha F. Noy, Csongor Nyulas, David Rubenson, Paul R. Saxman, Nancy Whelan, Zachary C. Wright, Brian D. Athey, Michael J. Becich, Geoffrey S. Ginsburg, Mark A. Musen, Kevin A. Smith 0001, Alice F. Tarantal, Daniel L. Rubin, Peter Lyster |
J. Biomed. Informatics | 10 |
| 2011 | NCBO Resource Index: Ontology-based search and mining of biomedical resources
Clément Jonquet, Paea LePendu, Sean M. Falconer, Adrien Coulet, Natasha F. Noy, Mark A. Musen, Nigam H. Shah |
J. Web Semant. | 5 |
| 2010 | Ontology Development for the Masses: Creating ICD-11 in WebProtégé
Tania Tudorache, Sean M. Falconer, Natasha F. Noy, Csongor Nyulas, Tevfik Bedirhan Üstün, Margaret-Anne D. Storey, Mark A. Musen |
EKAW | 3 |
| 2010 | Optimize First, Buy Later: Analyzing Metrics to Ramp-Up Very Large Knowledge Bases
Paea LePendu, Natasha F. Noy, Clément Jonquet, Paul R. Alexander, Nigam H. Shah, Mark A. Musen |
ISWC (1) | 2 |
| 2010 | Will Semantic Web Technologies Work for the Development of ICD-11?
Tania Tudorache, Sean M. Falconer, Csongor Nyulas, Natasha F. Noy, Mark A. Musen |
ISWC (2) | 4 |
| 2009 | Creating Mappings For Ontologies in Biomedicine: Simple Methods Work
Amir Ghazvinian, Natasha F. Noy, Mark A. Musen |
AMIA | 2 |
| 2009 | What Four Million Mappings Can Tell You about Two Hundred Ontologies
Amir Ghazvinian, Natasha F. Noy, Clément Jonquet, Nigam H. Shah, Mark A. Musen |
ISWC | 2 |
| 2008 | Developing Biomedical Ontologies Collaboratively
Natasha F. Noy, Tania Tudorache, Sherri de Coronado, Mark A. Musen |
AMIA | 1 |
| 2008 | A Generic Ontology for Collaborative Ontology-Development Workflows
Abraham Sebastian, Natasha F. Noy, Tania Tudorache, Mark A. Musen |
EKAW | 2 |
| 2008 | Collecting Community-Based Mappings in an Ontology Repository
Natasha F. Noy, Nicholas Griffith, Mark A. Musen |
ISWC | 1 |
| 2008 | Supporting Collaborative Ontology Development in Protégé
Tania Tudorache, Natasha F. Noy, Samson W. Tu, Mark A. Musen |
ISWC | 2 |
| 2008 | Biomedical ontologies: a functional perspectiveabstractThe information explosion in biology makes it difficult for researchers to stay abreast of current biomedical knowledge and to make sense of the massive amounts of online information. Ontologies--specifications of the entities, their attributes and relationships among the entities in a domain of discourse--are increasingly enabling biomedical researchers to accomplish these tasks. In fact, bio-ontologies are beginning to proliferate in step with accruing biological data. The myriad of ontologies being created enables researchers not only to solve some of the problems in handling the data explosion but also introduces new challenges. One of the key difficulties in realizing the full potential of ontologies in biomedical research is the isolation of various communities involved: some workers spend their career developing ontologies and ontology-related tools, while few researchers (biologists and physicians) know how ontologies can accelerate their research. The objective of this review is to give an overview of biomedical ontology in practical terms by providing a functional perspective--describing how bio-ontologies can and are being used. As biomedical scientists begin to recognize the many different ways ontologies enable biomedical research, they will drive the emergence of new computer applications that will help them exploit the wealth of research data now at their fingertips. Daniel L. Rubin, Nigam H. Shah, Natasha F. Noy |
Briefings Bioinform. | 3 |
| 2008 | Translating the Foundational Model of Anatomy into OWL
Natasha F. Noy, Daniel L. Rubin |
J. Web Semant. | 1 |
| 2007 | Searching ontologies based on content: experiments in the biomedical domainabstractAs more ontologies become publicly available, finding the "right" ontologies becomes much harder. In this paper, we address the problem of ontology search: finding a collection of ontologies from an ontology repository that are relevant to the user's query. In particular, we look at the case when users search for ontologies relevant to a particular topic (e.g., an ontology about anatomy). Ontologies that are most relevant to such query often do not have the query term in the names of their concepts (e.g., the Foundational Model of Anatomy ontology does not have the term "anatomy" in any of its concepts' names). Thus, we present a new ontology-search technique that helps users in these types of searches. When looking for ontologies on a particular topic (e.g., anatomy), we retrieve from the Web a collection of terms that represent the given domain (e.g., terms such as body, brain, skin, etc. for anatomy). We then use these terms to expand the user query. We evaluate our algorithm on queries for topics in the biomedical domain against a repository of biomedical ontologies. We use the results obtained from experts in the biomedical-ontology domain as the gold standard. Our experiments demonstrate that using our method for query expansion improves retrieval results by a 113%, compared to the tools that search only for the user query terms and consider only class and property names (like Swoogle). We show 43% improvement for the case where not only class and property names but also property values are taken into account. Harith Alani, Natasha F. Noy, Nigam H. Shah, Nigel Shadbolt, Mark A. Musen |
K-CAP | 2 |
| 2006 | A Framework for Ontology Evolution in Collaborative Environments
Natasha F. Noy, Abhita Chugh, William Liu, Mark A. Musen |
ISWC | 1 |
| 2005 | OMEN: A Probabilistic Ontology Mapping Tool
Prasenjit Mitra 0001, Natasha F. Noy, Anuj R. Jaiswal |
ISWC | 2 |
| 2005 | EZPAL: Environment for composing constraint axioms by instantiating templates
Chih-Sheng Johnson Hou, Mark A. Musen, Natasha F. Noy |
Int. J. Hum. Comput. Stud. | 3 |
| 2004 | The Protégé OWL Plugin: An Open Development Environment for Semantic Web Applications
Holger Knublauch, Ray W. Fergerson, Natasha F. Noy, Mark A. Musen |
ISWC | 3 |
| 2004 | Tracking Changes During Ontology Evolution
Natasha F. Noy, Sandhya Kunnatur, Michel C. A. Klein, Mark A. Musen |
ISWC | 1 |
| 2004 | Specifying Ontology Views by Traversal
Natasha F. Noy, Mark A. Musen |
ISWC | 1 |
| 2004 | Pushing the envelope: challenges in a frame-based representation of human anatomy
Natasha F. Noy, Mark A. Musen, José L. V. Mejino Jr., Cornelius Rosse |
Data Knowl. Eng. | 1 |
| 2004 | Ontology Evolution: Not the Same as Schema Evolution
Natasha F. Noy, Michel C. A. Klein |
Knowl. Inf. Syst. | 1 |
| 2003 | Protégé-2000: An Open-Source Ontology-Development and Knowledge-Acquisition Environment: AMIA 2003 Open Source Expo
Natasha F. Noy, Monica Crubézy, Ray W. Fergerson, Holger Knublauch, Samson W. Tu, Jennifer Vendetti, Mark A. Musen |
AMIA | 1 |
| 2003 | Knowledge acquisition, consistency checking and concurrency control for Gene Ontology (GO)abstractMOTIVATION: A critical element of the computational infrastructure required for functional genomics is a shared language for communicating biological data and knowledge. The Gene Ontology (GO; http://www.geneontology.org) provides a taxonomy of concepts and their attributes for annotating gene products. As GO increases in size, its ongoing construction and maintenance becomes more challenging. In this paper, we assess the applicability of a Knowledge Base Management System (KBMS), Protégé-2000, to the maintenance and development of GO. RESULTS: We transferred GO to Protégé-2000 in order to evaluate its suitability for GO. The graphical user interface supported browsing and editing of GO. Tools for consistency checking identified minor inconsistencies in GO and opportunities to reduce redundancy in its representation. The Protégé Axiom Language proved useful for checking ontological consistency. The PROMPT tool allowed us to track changes to GO. Using Protégé-2000, we tested our ability to make changes and extensions to GO to refine the semantics of attributes and classify more concepts. AVAILABILITY: Gene Ontology in Protégé-2000 and the associated code are located at http://smi.stanford.edu/projects/helix/gokbms/. Protégé-2000 is available from http://protege.stanford.edu. Iwei Yeh, Peter D. Karp, Natasha F. Noy, Russ B. Altman |
Bioinform. | 3 |
| 2003 | The evolution of Protégé: an environment for knowledge-based systems development
John H. Gennari, Mark A. Musen, Ray W. Fergerson, William E. Grosso, Monica Crubézy, Henrik Eriksson, Natasha F. Noy, Samson W. Tu |
Int. J. Hum. Comput. Stud. | 7 |
| 2003 | The PROMPT suite: interactive tools for ontology merging and mapping
Natasha F. Noy, Mark A. Musen |
Int. J. Hum. Comput. Stud. | 1 |
| 2002 | Jambalaya: an interactive environment for exploring ontologiesabstractNo abstract available. Margaret-Anne D. Storey, Natasha F. Noy, Mark A. Musen, Casey Best, Ray W. Fergerson, Neil A. Ernst |
IUI | 2 |
| 2001 | Representation of Structural Relationships in the Foundational Model of Anatomy
José L. V. Mejino Jr., Natasha F. Noy, Mark A. Musen, James F. Brinkley, Cornelius Rosse |
AMIA | 2 |
| 2001 | Protege-2000: A Plug-in Architecture to Support Knowledge Acquisition, Knowledge Visualization, and the Semantic Web
Mark A. Musen, Ray W. Fergerson, Natasha F. Noy, Monica Crubézy |
AMIA | 3 |
| 2000 | Ontology acquisition from on-line knowledge sources
Philip Shilane, Natasha F. Noy, Mark A. Musen |
AMIA | 3 |
| 2000 | The Knowledge Model of Protégé-2000: Combining Interoperability and Flexibility
Natasha F. Noy, Ray W. Fergerson, Mark A. Musen |
EKAW | 1 |
| 1996 | Ontological Foundations for Biology Knowledge Models
Carole D. Hafner, Natasha F. Noy |
ISMB | 2 |
| 1994 | Creating a Knowledge Base of Biological Research Papers
Carole D. Hafner, Kenneth Baclawski, Robert P. Futrelle, Natasha F. Noy, Shobana Sampath |
ISMB | 4 |
| 1993 | Database Techniques for Biological Materials and Methods
Kenneth Baclawski, Robert P. Futrelle, Natasha F. Noy, Maurice J. Pescitelli |
ISMB | 3 |