Natasha F. Noy

dblp:n/NatalyaFridmanNoy · also Natalya Fridman Noy, Natasha Fridman Noy · DBLP profile ↗
← Back
38ranked-venue papers in the field
10as first author
3since 2021 · last 2024
0000-0002-7437-0624ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 30 (7 first)Database Systems & Data Management · 5 (2 first)Information Retrieval & Web Search · 2Data Mining & Knowledge Discovery · 1 (1 first)
YearPublicationVenuePosition
2024 Relationships Are Complicated! An Analysis of Relationships Between Datasets on the Web
Kate Lin, Tarfah Alrashed, Natasha F. Noy
ISWC (1)3
2023 Will LLMs reshape, supercharge, or kill data science?
abstract
Large language models (LLMs) have recently taken the world by storm, promising potentially game changing opportunities in multiple fields. Naturally, there is significant promise in applying LLMs to the management of structured data, or more generally, to the processes involved in data science. At the very least, LLMs have the potential to provide substantial advancements in long-standing challenges that our community has been tackling for decades. On the other hand, they may introduce completely new capabilities that we have only dreamed of thus far. This panel will bring together a few leading experts who have been thinking about these opportunities from various perspectives and fielding them in research prototypes and even in commercial applications.
Alon Y. Halevy, Yejin Choi 0001, Avrilia Floratou, Michael J. Franklin, Natasha F. Noy, Haixun Wang
Proc. VLDB Endow.5
2021 Dataset or Not? A Study on the Veracity of Semantic Markup for Dataset Pages
abstract
Abstract Semantic markup, such as , allows providers on the Web to describe content using a shared controlled vocabulary. This markup is invaluable in enabling a broad range of applications, from vertical search engines, to rich snippets in search results, to actions on emails, to many others. In this paper, we focus on semantic markup for datasets, specifically in the context of developing a vertical search engine for datasets on the Web, Google’s Dataset Search. Dataset Search relies on to identify pages that describe datasets. While was the core enabling technology for this vertical search, we also discovered that we need to address the following problem: pages from 61% of internet hosts that provide markup do not actually describe datasets. We analyze the veracity of dataset markup for Dataset Search’s Web-scale corpus and categorize pages where this markup is not reliable. We then propose a way to drastically increase the quality of the dataset metadata corpus by developing a deep neural-network classifier that identifies whether or not a page with markup is a dataset page. Our classifier achieves 96.7% recall at the 95% precision point. This level of precision enables Dataset Search to circumvent the noise in semantic markup and to use the metadata to provide high quality results to users.
Tarfah Alrashed, Dimitris Paparas, Omar Benjelloun, Ying Sheng 0002, Natasha F. Noy
ISWC5
2020 Google Dataset Search by the Numbers
Omar Benjelloun, Natasha F. Noy
ISWC (2)3
2020 When the Web is your Data Lake: Creating a Search Engine for Datasets on the Web
abstract
There are thousands of data repositories on the Web, providing access to millions of datasets. National and regional governments, scientific publishers and consortia, commercial data providers, and others publish data for fields ranging from social science to life science to high-energy physics to climate science and more. Access to this data is critical to facilitating reproducibility of research results, enabling scientists to build on others' work, and providing data journalists easier access to information and its provenance. In this talk, I will discuss our work on Dataset Search, which provides search capabilities over potentially all dataset repositories on the Web. I will talk about the open ecosystem for describing and citing datasets that we hope to encourage and the technical details on how we went about building Dataset Search. Finally, I will highlight research challenges in building a vibrant, heterogeneous, and open ecosystem where data becomes a first-class citizen.
Natasha F. Noy
SIGMOD Conference1
2019 Google Dataset Search: Building a search engine for datasets in an open Web ecosystem
abstract
There are thousands of data repositories on the Web, providing access to millions of datasets. National and regional governments, scientific publishers and consortia, commercial data providers, and others publish data for fields ranging from social science to life science to high-energy physics to climate science and more. Access to this data is critical to facilitating reproducibility of research results, enabling scientists to build on others' work, and providing data journalists easier access to information and its provenance. In this paper, we discuss Google Dataset Search, a dataset-discovery tool that provides search capabilities over potentially all datasets published on the Web. The approach relies on an open ecosystem, where dataset owners and providers publish semantically enhanced metadata on their own sites. We then aggregate, normalize, and reconcile this metadata, providing a search engine that lets users find datasets in the “long tail” of the Web. In this paper, we discuss both social and technical challenges in building this type of tool, and the lessons that we learned from this experience.
Dan Brickley, Matthew Burgess, Natasha F. Noy
WWW3
2016 Goods: Organizing Google's Datasets
abstract
Enterprises increasingly rely on structured datasets to run their businesses. These datasets take a variety of forms, such as structured files, databases, spreadsheets, or even services that provide access to the data. The datasets often reside in different storage systems, may vary in their formats, may change every day. In this paper, we present GOODS, a project to rethink how we organize structured datasets at scale, in a setting where teams use diverse and often idiosyncratic ways to produce the datasets and where there is no centralized system for storing and querying them. GOODS extracts metadata ranging from salient information about each dataset (owners, timestamps, schema) to relationships among datasets, such as similarity and provenance. It then exposes this metadata through services that allow engineers to find datasets within the company, to monitor datasets, to annotate them in order to enable others to use their datasets, and to analyze relationships between them. We discuss the technical challenges that we had to overcome in order to crawl and infer the metadata for billions of datasets, to maintain the consistency of our metadata catalog at scale, and to expose the metadata to users. We believe that many of the lessons that we learned are applicable to building large-scale enterprise-level data-management systems in general.
Alon Y. Halevy, Flip Korn, Natasha F. Noy, Christopher Olston, Neoklis Polyzotis, Sudip Roy 0002, Steven Euijong Whang
SIGMOD Conference3
2016 Discovering Structure in the Universe of Attribute Names
abstract
Recently, search engines have invested significant effort to answering entity--attribute queries from structured data, but have focused mostly on queries for frequent attributes. In parallel, several research efforts have demonstrated that there is a long tail of attributes, often thousands per class of entities, that are of interest to users. Researchers are beginning to leverage these new collections of attributes to expand the ontologies that power search engines and to recognize entity--attribute queries. Because of the sheer number of potential attributes, such tasks require us to impose some structure on this long and heavy tail of attributes. This paper introduces the problem of organizing the attributes by expressing the compositional structure of their names as a rule-based grammar. These rules offer a compact and rich semantic interpretation of multi-word attributes, while generalizing from the observed attributes to new unseen ones. The paper describes an unsupervised learning method to generate such a grammar automatically from a large set of attribute names. Experiments show that our method can discover a precise grammar over 100,000 attributes of {\sc Countries} while providing a 40-fold compaction over the attribute names. Furthermore, our grammar enables us to increase the precision of attributes from 47\% to more than 90\% with only a minimal curation effort. Thus, our approach provides an efficient and scalable way to expand ontologies with attributes of user interest.
Alon Y. Halevy, Natasha F. Noy, Sunita Sarawagi, Steven Euijong Whang
WWW2
2015 Discovering Subsumption Relationships for Web-Based Ontologies
abstract
As search engines are becoming smarter at interpreting user queries and providing meaningful responses, they rely on ontologies to understand the meaning of entities. Creating ontologies manually is a laborious process, and resulting ontologies may not reflect the way users think about the world, as many concepts used in queries are noisy, and not easily amenable to formal modeling. There has been considerable effort in generating ontologies from Web text and query streams, which may be more reflective of how users query and write content. In this paper, we describe the LATTE system that automatically generates a subconcept--superconcept hierarchy, which is critical for using ontologies to answer queries. LATTE combines signals based on word-vector representations of concepts and dependency parse trees; however, LATTE derives most of its power from an ontology of attributes extracted from the Web that indicates the aspects of concepts that users find important. LATTE achieves an F1 score of 74%, which is comparable to expert agreement on a similar task. We additionally demonstrate the usefulness of LATTE in detecting high quality concepts from an existing resource of IsA links.
Dana Movshovitz-Attias, Steven Euijong Whang, Natasha F. Noy, Alon Y. Halevy
WebDB3
2013 Indented Tree or Graph? A Usability Study of Ontology Visualization Techniques in the Context of Class Mapping Evaluation
Bo Fu 0005, Natasha F. Noy, Margaret-Anne D. Storey
ISWC (1)2
2013 Simplified OWL Ontology Editing for the Web: Is WebProtégé Enough?
Matthew Horridge, Tania Tudorache, Jennifer Vendetti, Csongor Nyulas, Mark A. Musen, Natasha F. Noy
ISWC (1)6
2013 Getting Lucky in Ontology Search: A Data-Driven Evaluation Framework for Ontology Ranking
Natasha F. Noy, Paul R. Alexander, Rave Harpaz, Patricia L. Whetzel, Ray W. Fergerson, Mark A. Musen
ISWC (1)1
2013 Using Semantic Web in ICD-11: Three Years Down the Road
Tania Tudorache, Csongor Nyulas, Natasha F. Noy, Mark A. Musen
ISWC (2)3
2013 PragmatiX: An Interactive Tool for Visualizing the Creation Process Behind Collaboratively Engineered Ontologies
abstract
With the emergence of tools for collaborative ontology engineering, more and more data about the creation process behind collaborative construction of ontologies is becoming available. Today, collaborative ontology engineering tools such as Collaborative Protégé offer rich and structured logs of changes, thereby opening up new challenges and opportunities to study and analyze the creation of collaboratively constructed ontologies. While there exists a plethora of visualization tools for ontologies, they have primarily been built to visualize aspects of the final product (the ontology) and not the collaborative processes behind construction (e.g. the changes made by contributors over time). To the best of the authors’ knowledge, there exists no ontology visualization tool today that focuses primarily on visualizing the history behind collaboratively constructed ontologies. Since the ontology engineering processes can influence the quality of the final ontology, they believe that visualizing process data represents an important stepping-stone towards better understanding of managing the collaborative construction of ontologies in the future. In this application paper, the authors present a tool – PragmatiX – which taps into structured change logs provided by tools such as Collaborative Protégé to visualize various pragmatic aspects of collaborative ontology engineering. The tool is aimed at managers and leaders of collaborative ontology engineering projects to help them in monitoring progress, in exploring issues and problems, and in tracking quality-related issues such as overrides and coordination among contributors. The paper makes the following contributions: (i) They present PragmatiX, a tool for visualizing the creation process behind collaboratively constructed ontologies (ii) the authors illustrate the functionality and generality of the tool by applying it to structured logs of changes of two large collaborative ontology-engineering projects and (iii) they conduct a heuristic evaluation of the tool with domain experts to uncover early design challenges and opportunities for improvement. Finally, the authors hope that this work sparks a new line of research on visualization tools for collaborative ontology engineering projects.
Simon Walk, Jan Pöschko, Markus Strohmaier, Keith Andrews, Tania Tudorache, Natasha F. Noy, Csongor Nyulas, Mark A. Musen
Int. J. Semantic Web Inf. Syst.6
2013 How ontologies are made: Studying the hidden social dynamics behind collaborative ontology engineering projects
Markus Strohmaier, Simon Walk, Jan Pöschko, Daniel Lamprecht, Tania Tudorache, Csongor Nyulas, Mark A. Musen, Natasha F. Noy
J. Web Semant.8
2012 Using SPARQL to Query BioPortal Ontologies and Metadata
Manuel Salvadores, Matthew Horridge, Paul R. Alexander, Ray W. Fergerson, Mark A. Musen, Natasha F. Noy
ISWC (2)6
2012 CrowdMap: Crowdsourcing Ontology Alignment with Microtasks
Cristina Sarasua, Elena Simperl, Natasha F. Noy
ISWC (1)3
2012 Where to publish and find ontologies? A survey of ontology libraries
Mathieu d'Aquin, Natasha F. Noy
J. Web Semant.2
2011 An analysis of collaborative patterns in large-scale ontology development projects
abstract
Today, distributed teams collaboratively create and maintain more and more ontologies. To support this type of ontology development, software engineers are introducing a new generation of tools. However, we know relatively little about how existing large-scale collaborative ontology development works and what user workflows the tools must support. In this paper, we analyze our experience in supporting several such projects. We describe a visual and interactive project-management tool that we have developed, which helps ontology developers explore historical ontology change and discussion data. We present the results of qualitative and quantitative studies of the collaborative activity associated with three large-scale ontology-development projects. Based on the analysis, we conclude that domain and ontology experts have different patterns of ontology editing behavior, which has important implications for ontology-development tools.
Sean M. Falconer, Tania Tudorache, Natasha F. Noy
K-CAP3
2011 From mappings to modules: using mappings to identify domain-specific modules in large ontologies
abstract
The problem of ontology modularization is an active area of research in the Semantic Web community. With the emergence and wider use of very large ontologies, in particular in fields such as biomedicine, more and more application developers need to extract meaningful modules of these ontologies to use in their applications. Researchers have also noted that many ontology-maintenance tasks would be simplified if we could extract modules from ontologies. These tasks include ontology matching: If we can separate ontologies into modules based on the topics that these modules cover, we can simplify and improve ontology matching. In this paper, we study a complementary problem: Can we use existing mappings between ontologies to facilitate modularization? We present a novel approach to modularization based on mappings between ontologies. We validate and analyze our approach by applying our methods to identify modules for National Cancer Institutes Thesaurus (NCI Thesaurus) and Systematized Nomenclature of Medicine--Clinical Terms (SNOMED-CT).
Amir Ghazvinian, Natasha F. Noy, Mark A. Musen
K-CAP2
2011 NCBO Resource Index: Ontology-based search and mining of biomedical resources
Clément Jonquet, Paea LePendu, Sean M. Falconer, Adrien Coulet, Natasha F. Noy, Mark A. Musen, Nigam H. Shah
J. Web Semant.5
2010 Ontology Development for the Masses: Creating ICD-11 in WebProtégé
Tania Tudorache, Sean M. Falconer, Natasha F. Noy, Csongor Nyulas, Tevfik Bedirhan Üstün, Margaret-Anne D. Storey, Mark A. Musen
EKAW3
2010 Optimize First, Buy Later: Analyzing Metrics to Ramp-Up Very Large Knowledge Bases
Paea LePendu, Natasha F. Noy, Clément Jonquet, Paul R. Alexander, Nigam H. Shah, Mark A. Musen
ISWC (1)2
2010 Will Semantic Web Technologies Work for the Development of ICD-11?
Tania Tudorache, Sean M. Falconer, Csongor Nyulas, Natasha F. Noy, Mark A. Musen
ISWC (2)4
2009 What Four Million Mappings Can Tell You about Two Hundred Ontologies
Amir Ghazvinian, Natasha F. Noy, Clément Jonquet, Nigam H. Shah, Mark A. Musen
ISWC2
2008 A Generic Ontology for Collaborative Ontology-Development Workflows
Abraham Sebastian, Natasha F. Noy, Tania Tudorache, Mark A. Musen
EKAW2
2008 Collecting Community-Based Mappings in an Ontology Repository
Natasha F. Noy, Nicholas Griffith, Mark A. Musen
ISWC1
2008 Supporting Collaborative Ontology Development in Protégé
Tania Tudorache, Natasha F. Noy, Samson W. Tu, Mark A. Musen
ISWC2
2008 Translating the Foundational Model of Anatomy into OWL
Natasha F. Noy, Daniel L. Rubin
J. Web Semant.1
2007 Searching ontologies based on content: experiments in the biomedical domain
abstract
As more ontologies become publicly available, finding the "right" ontologies becomes much harder. In this paper, we address the problem of ontology search: finding a collection of ontologies from an ontology repository that are relevant to the user's query. In particular, we look at the case when users search for ontologies relevant to a particular topic (e.g., an ontology about anatomy). Ontologies that are most relevant to such query often do not have the query term in the names of their concepts (e.g., the Foundational Model of Anatomy ontology does not have the term "anatomy" in any of its concepts' names). Thus, we present a new ontology-search technique that helps users in these types of searches. When looking for ontologies on a particular topic (e.g., anatomy), we retrieve from the Web a collection of terms that represent the given domain (e.g., terms such as body, brain, skin, etc. for anatomy). We then use these terms to expand the user query. We evaluate our algorithm on queries for topics in the biomedical domain against a repository of biomedical ontologies. We use the results obtained from experts in the biomedical-ontology domain as the gold standard. Our experiments demonstrate that using our method for query expansion improves retrieval results by a 113%, compared to the tools that search only for the user query terms and consider only class and property names (like Swoogle). We show 43% improvement for the case where not only class and property names but also property values are taken into account.
Harith Alani, Natasha F. Noy, Nigam H. Shah, Nigel Shadbolt, Mark A. Musen
K-CAP2
2006 A Framework for Ontology Evolution in Collaborative Environments
Natasha F. Noy, Abhita Chugh, William Liu, Mark A. Musen
ISWC1
2005 OMEN: A Probabilistic Ontology Mapping Tool
Prasenjit Mitra 0001, Natasha F. Noy, Anuj R. Jaiswal
ISWC2
2004 The Protégé OWL Plugin: An Open Development Environment for Semantic Web Applications
Holger Knublauch, Ray W. Fergerson, Natasha F. Noy, Mark A. Musen
ISWC3
2004 Tracking Changes During Ontology Evolution
Natasha F. Noy, Sandhya Kunnatur, Michel C. A. Klein, Mark A. Musen
ISWC1
2004 Specifying Ontology Views by Traversal
Natasha F. Noy, Mark A. Musen
ISWC1
2004 Pushing the envelope: challenges in a frame-based representation of human anatomy
Natasha F. Noy, Mark A. Musen, José L. V. Mejino Jr., Cornelius Rosse
Data Knowl. Eng.1
2004 Ontology Evolution: Not the Same as Schema Evolution
Natasha F. Noy, Michel C. A. Klein
Knowl. Inf. Syst.1
2000 The Knowledge Model of Protégé-2000: Combining Interoperability and Flexibility
Natasha F. Noy, Ray W. Fergerson, Mark A. Musen
EKAW1