Christopher A. Welty

dblp:w/CAWelty · also Chris Welty · DBLP profile ↗
← Back
36ranked-venue papers
19as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 15 · 7 first-author · 2 since 2021Software engineering, systems software and programming languages · 9 · 9 first-authorTheory of computation · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 50% Learning and educational technologies · 50%
Artificial intelligence
5 papers
Trustworthy machine learning · 88% Knowledge representation and reasoning · 12%
Software engineering, system software, and programming languages
4 papers
Empirical software engineering · 95% Software maintenance and evolution · 2% Requirements engineering and software design · 2%
Databases, data mining, and information retrieval
3 papers
Machine learning and data management · 54% Information retrieval · 31% Knowledge graphs · 11%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
annotator disagreement
1.012026
Forest vs Tree: The (N, K) Trade-off in Reproducible ML Evaluation · AAAI 2026
Empirical software engineering
reproducibility
1.012026
Forest vs Tree: The (N, K) Trade-off in Reproducible ML Evaluation · AAAI 2026
Learning and educational technologies
active learning
0.612022
Vexation-Aware Active Learning for On-Menu Restaurant Dish Availability · KDD 2022
Collaborative and social computing
crowdsourcing
0.612022
Vexation-Aware Active Learning for On-Menu Restaurant Dish Availability · KDD 2022
Information retrieval
question answering
0.212022
Vexation-Aware Active Learning for On-Menu Restaurant Dish Availability · KDD 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
ontology
0.142004
Evaluating Ontology Cleaning · AAAI 2004
Panel: Are Upper-Level Ontologies Worth the Effort? · KR 2002
A Formal Ontology for Re-Use of Software Architecture Documents · ASE 1999
Knowledge graphs › semantic web
semantic web application
0.112006
Supporting online problem-solving communities with the semantic web · WWW 2006
Empirical software engineering › open source software
open source communities
0.112006
Supporting online problem-solving communities with the semantic web · WWW 2006
Requirements engineering and software design
software architecture
0.011999
A Formal Ontology for Re-Use of Software Architecture Documents · ASE 1999
Software maintenance and evolution
program comprehension
0.011997
Augmenting Abstract Syntax Trees for Program Understanding · ASE 1997
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology
ontology engineering
0.012002
Panel: Are Upper-Level Ontologies Worth the Effort? · KR 2002
Software maintenance and evolution
software reuse
0.011999
A Formal Ontology for Re-Use of Software Architecture Documents · ASE 1999
Program analysis › program representation
abstract syntax tree analysis
0.011997
Augmenting Abstract Syntax Trees for Program Understanding · ASE 1997

Methods — techniques the papers use, named apart from their topics

statistical analysis · 3.0simulation · 3.0active learning · 1.1semantic annotation · 0.1ontology-based modeling · 0.1knowledge-based tool · 0.0formal ontology · 0.0ontology · 0.0automated reasoning · 0.0
YearPublicationVenuePosition
2026 Forest vs Tree: The (N, K) Trade-off in Reproducible ML Evaluation
abstract
Reproducibility is a cornerstone of scientific validation and of the authority it confers on its results. Reproducibility in machine learning evaluations leads to greater trust, confidence, and value. However, the ground truth responses used in machine learning often necessarily come from humans, among whom disagreement is prevalent, and surprisingly little research has studied the impact of effectively ignoring disagreement in these responses, as is typically the case. One reason for the lack of research is that budgets for collecting human-annotated evaluation data are limited, and obtaining more samples from multiple raters for each example greatly increases the per-item annotation costs. We investigate the trade-off between the number of items (N) and the number of responses per item (K) needed for reliable machine learning evaluation. We analyze a diverse collection of categorical datasets for which multiple annotations per item exist, and simulated distributions fit to these datasets, to determine the optimal (N, K) configuration, given a fixed budget (N x K), for collecting evaluation data and reliably comparing the performance of machine learning models. Our findings show, first, that accounting for human disagreement may come with N x K at no more than 1000 (and often much lower) for every dataset tested on at least one metric. Moreover, this minimal N x K almost always occurred for K > 10. Furthermore, the nature of the tradeoff between K and N, or if one even existed, depends on the evaluation metric, with metrics that are more sensitive to the full distribution of responses performing better at higher levels of K. Our methods can be used to help ML practitioners get more effective test data by finding the optimal metrics and number of items and annotations per item to collect to get the most reliability for their budget.
Deepak Pandita, Flip Korn, Christopher A. Welty, Christopher Homan
AAAI3
2022 Vexation-Aware Active Learning for On-Menu Restaurant Dish Availability
abstract
Here we leverage the power of the crowd: online users who are willing to answer questions about dish availability at restaurants visited. While motivated users are happy to contribute knowledge, they are much less likely to respond to "silly'' or embarrassing questions (e.g., "DoesPizza Hut serve pizza?'' or "DoesMike's Vegan Restaurant serve steak?'')
Jean-François Kagy, Flip Korn, Afshin Rostamizadeh, Christopher A. Welty
KDD4
2021 Rapid Instance-Level Knowledge Acquisition for Google Maps from Class-Level Common Sense
abstract
Successful knowledge graphs (KGs) solved the historical knowledge acquisition bottleneck by supplanting an expert focus with a simple, crowd-friendly one: KG nodes represent popular people, places, organizations, etc., and the graph arcs represent common sense relations like affiliations, locations, etc. Techniques for more general, categorical, KG curation do not seem to have made the same transition: the KG research community is still largely focused on methods that belie the common-sense characteristics of successful KGs. In this paper, we propose a simple approach to acquiring and reasoning with class-level attributes from the crowd that represent broad common sense associations between categories. We pick a very real industrial-scale data set and problem: how to augment an existing knowledge graph of places and products with associations between them indicating the availability of the products at those places, which would enable a KG to provide answers to questions like, "Where can I buy milk nearby?" This problem has several practical challenges, not least of which is that only 30% of physical stores (i.e. brick & mortar stores) have a website, and fewer list their product inventory, leaving a large acquisition gap to be filled by methods other than information extraction (IE). Based on a KG-inspired intuition that a lot of the class-level pairs are part of people's general common sense, e.g. everyone knows grocery stores sell milk and don't sell asphalt, we acquired a mixture of instance- and class- level pairs (e.g. , , resp.) from a novel 3-tier crowdsourcing method, and demonstrate the scalability advantages of the class-level approach. Our results show that crowdsourced class-level knowledge can provide rapid scaling of knowledge acquisition in this and similar domains, as well as long-term value in the KG.
Christopher A. Welty, Lora Aroyo, Flip Korn, Sara M. McCarthy, Shubin Zhao
HCOMP1
2020 Embedding Semantic Taxonomies
abstract
A common step in developing an understanding of a vertical domain, e.g.shopping, dining, movies, medicine, etc., is curating a taxonomy of categories specific to the domain.These human created artifacts have been the subject of research in embeddings that attempt to encode aspects of the partial ordering property of taxonomies.We compare Box Embeddings, a natural containment-based representation of taxonomies, to partial-order embeddings and a baseline Bayes Net, in the context of representing the Medical Subject Headings (MeSH) taxonomy given a set of 300K PubMed articles with subject labels from MeSH.We deeply explore the experimental properties of training box embeddings, including preparation of the training data, sampling ratios and class balance, initialization strategies, and propose a fix to the original box objective.We then present first results in using these techniques for representing a bipartite learning problem (i.e.collaborative filtering) in the presence of taxonomic relations within each partition, inferring disease (anatomical) locations from their use as subject labels in journal articles.Our box model substantially outperforms all baselines for taxonomic reconstruction and bipartite relationship experiments.This performance improvement is observed both in overall accuracy and the weighted spread by true taxonomic depth.
Alyssa Lees, Christopher A. Welty, Shubin Zhao, Jacek Korycki, Sara Mc Carthy
COLING2
2018 Capturing Ambiguity in Crowdsourcing Frame Disambiguation
abstract
FrameNet is a computational linguistics resource composed of semantic frames, high-level concepts that represent the meanings of words. In this paper, we present an approach to gather frame disambiguation annotations in sentences using a crowdsourcing approach with multiple workers per sentence to capture inter-annotator disagreement. We perform an experiment over a set of 433 sentences annotated with frames from the FrameNet corpus, and show that the aggregated crowd annotations achieve an F1 score greater than 0.67 as compared to expert linguists. We highlight cases where the crowd annotation was correct even though the expert is in disagreement, arguing for the need to have multiple annotators per sentence. Most importantly, we examine cases in which crowd workers could not agree, and demonstrate that these cases exhibit ambiguity, either in the sentence, frame, or the task itself, and argue that collapsing such cases to a single, discrete truth value (i.e. correct or incorrect) is inappropriate, creating arbitrary targets for machine learning.
Anca Dumitrache, Lora Aroyo, Christopher A. Welty
HCOMP3
2018 Crowdsourcing Ground Truth for Medical Relation Extraction
abstract
Cognitive computing systems require human labeled data for evaluation and often for training. The standard practice used in gathering this data minimizes disagreement between annotators, and we have found this results in data that fails to account for the ambiguity inherent in language. We have proposed the CrowdTruth method for collecting ground truth through crowdsourcing, which reconsiders the role of people in machine learning based on the observation that disagreement between annotators provides a useful signal for phenomena such as ambiguity in the text. We report on using this method to build an annotated data set for medical relation extraction for the cause and treat relations, and how this data performed in a supervised training experiment. We demonstrate that by modeling ambiguity, labeled data gathered from crowd workers can (1) reach the level of quality of domain experts for this task while reducing the cost, and (2) provide better training data at scale than distant supervision. We further propose and validate new weighted measures for precision, recall, and F-measure, which account for ambiguity in both human and machine performance on this task.
Anca Dumitrache, Lora Aroyo, Christopher A. Welty
ACM Trans. Interact. Intell. Syst.3
2013 Long-Distance Time-Event Relation Extraction
Alessandro Moschitti, Siddharth Patwardhan, Christopher A. Welty
IJCNLP3
2012 When Did that Happen? - Linking Events and Relations to Timestamps
Dirk Hovy, James Fan, Alfio Massimiliano Gliozzo, Siddharth Patwardhan, Christopher A. Welty
EACL5
2012 Query Driven Hypothesis Generation for Answering Queries over NLP Graphs
Christopher A. Welty, Ken Barker 0002, Lora Aroyo, Shilpa Arora
ISWC (2)1
2012 A Comparison of Hard Filters and Soft Evidence for Answer Typing in Watson
Christopher A. Welty, J. William Murdock, Aditya Kalyanpur, James Fan
ISWC (2)1
2011 Leveraging Community-Built Knowledge for Type Coercion in Question Answering
Aditya Kalyanpur, J. William Murdock, James Fan, Christopher A. Welty
ISWC (2)4
2010 Learning to Predict Readability using Diverse Linguistic Features
Rohit J. Kate, Xiaoqiang Luo, Siddharth Patwardhan, Martin Franz, Radu Florian, Raymond J. Mooney, Salim Roukos, Christopher A. Welty
COLING8
2006 OntOWLClean: Cleaning OWL ontologies with OWL
Christopher A. Welty
FOIS1
2006 A Reusable Ontology for Fluents in OWL
Christopher A. Welty, Richard Fikes
FOIS1
2006 Semantic Web: The Story of the RIFt so Far
Christopher A. Welty
ICLP1
2006 A Model Driven Approach for Building OWL DL and OWL Full Ontologies
Saartje Brockmans, Robert M. Colomb, Peter Haase 0001, Elisa F. Kendall, Evan K. Wallace, Christopher A. Welty, Guo Tong Xie
ISWC6
2006 Explaining Conclusions from Diverse Knowledge Sources
J. William Murdock, Deborah L. McGuinness, Paulo Pinheiro 0001, Christopher A. Welty, David A. Ferrucci
ISWC4
2006 Towards Knowledge Acquisition from Information Extraction
Christopher A. Welty, J. William Murdock
ISWC1
2006 Supporting online problem-solving communities with the semantic web
abstract
The Web plays a critical role in hosting Web communities, their content and interactions. A prime example is the open source software (OSS) community, whose members, including software developers and users, interact almost exclusively over the Web, constantly generating, sharing and refining content in the form of software code through active interaction over the Web on code design and bug resolution processes. The Semantic Web is an envisaged extension of the current Web, in which content is given a well defined meaning, through the specification of metadata and ontologies, increasing the utility of the content and enabling information from heterogeneous sources to be integrated. We developed a prototype Semantic Web system for OSS communities, Dhruv. Dhruv provides an enhanced semantic interface to bug resolution messages and recommends related software objects and artifacts. Dhruv uses an integrated model of the OpenACS community, the software, and the Web interactions, which is semi-automatically populated from the existing artifacts of the community.
Anupriya Ankolekar, Katia P. Sycara, James D. Herbsleb, Robert E. Kraut, Christopher A. Welty
WWW5
2004 Evaluating Ontology Cleaning
Christopher A. Welty, Ruchi Mahindru, Jennifer Chu-Carroll
AAAI1
2002 Ontology-Driven Conceptual Modeling
Christopher A. Welty
CAiSE1
2002 Panel: Are Upper-Level Ontologies Worth the Effort?
Christopher A. Welty
KR1
2001 FOIS introduction: Ontology - towards a new synthesis
abstract
This introduction to the Second International Conference on Formal Ontology and Information Systems presents a brief history of ontology as a discipline spanning the boundaries of philosophy and information science. We sketch some of the reasons for the growth of ontology in the information science field, and offer a preliminary stocktaking of how the term 'ontology' is currently used. We conclude by suggesting some grounds for optimism as concerns the future collaboration between philosophical ontologists and information scientistsPhilosophical ontology is the science of what is, of the kinds and structures of objects, properties, events, processes and relations in every area of reality. Philosophical ontology takes many forms, from the metaphysics of Aristotle to the object-theory of Alexius Meinong. The term 'ontology' (or ontologia) was itself coined in 1613, independently, by two philosophers, Rudolf Göckel (Goclenius), in his Lexicon philosophicum and Jacob Lorhard (Lorhardus), in his Theatrum philosophicum. Its first occurrence in English as recorded by the OED appears in Bailey's dictionary of 1721, which defines ontology as 'an Account of being in the Abstract'Regardless of its name, what we now refer to as philosophical ontology has sought the definitive and exhaustive classification of entities in all spheres of being. It can thus be conceived as a kind of generalized chemistry. The taxonomies which result from philosophical ontology have been intended to be definitive in the sense that they could serve as answers to such questions as: What classes of entities are needed for a complete description and explanation of all the goings-on in the universe? Or: What classes of entities are needed to give an account of what makes true all truths? They have been designed to be exhaustive in the sense that all types of entities should be included, including also the types of relations by which entities are tied together.
Barry Smith 0001, Christopher A. Welty
FOIS2
2001 Supporting ontological analysis of taxonomic relationships
Christopher A. Welty, Nicola Guarino
Data Knowl. Eng.1
2000 Identity, Unity, and Individuality: Towards a Formal Toolkit for Ontological Analysis
Nicola Guarino, Christopher A. Welty
ECAI2
2000 A Formal Ontology of Properties
Nicola Guarino, Christopher A. Welty
EKAW2
2000 Ontological Analysis of Taxonomic Relationships
Nicola Guarino, Christopher A. Welty
ER2
1999 A Formal Ontology for Re-Use of Software Architecture Documents
abstract
Software architecture has been established as a viable level of representation for reuse in practical software engineering efforts. The main reason for this is that an architectural view of software is sufficiently abstract to have many instantiations. Even with technologies such as CORBA and JavaBeans, which emphasize reuse of components, the realization of widespread reuse has been severely limited. While architectural reuse has been successful, it has thus far suffered from an ad-hoc semantics, and even savvy architecture practitioners are unsure precisely what is being reused. We have been engaged in research into reuse of software documents, such as design documents, statements of work, contracts, etc., that capture and reuse architectural level knowledge of software solutions. We have found that, given a sufficiently robust knowledge based tool for maintaining documents, a formal ontology or meta-model for software architectures is required to achieve reuse of these architecture-level documents. We present such an ontology.
Christopher A. Welty, David A. Ferrucci
ASE1
1999 Guest Editorial
Christopher A. Welty, Michael R. Lowry, Yves Ledru
Autom. Softw. Eng.1
1999 Formal Ontology for Subject
Christopher A. Welty, Jessica Jenkins
Data Knowl. Eng.1
1999 Report on the 1998 International Workshop on Description Logics (DL'98)
abstract
E Franconi, G De Giacomo, IR Horrocks, DL McGuinness, W Nutt, PF Patel-Schneider, CA Welty; Conferences. Report on the 1998 International Workshop on Descriptio
Enrico Franconi, Giuseppe De Giacomo, Ian Horrocks 0001, Deborah L. McGuinness, Werner Nutt, Peter F. Patel-Schneider, Christopher A. Welty
J. Log. Comput.7
1998 Editorial
Christopher A. Welty
Autom. Softw. Eng.1
1998 Desert Island Column
Christopher A. Welty
Autom. Softw. Eng.1
1997 Augmenting Abstract Syntax Trees for Program Understanding
abstract
Program understanding efforts by individual maintainers are dominated by a process known as discovery, which is characterized by low-level searches through the source code and documentation to obtain information that is important to the maintenance task. Discovery is complicated by the delocalization of information in the source code, and can consume from 40-60% of a maintainer's time. This paper presents an ontology for representing code-level knowledge based on abstract syntax trees, that was developed in the context of studying maintenance problems in a small software company. The ontology enables the utilization of automated reasoning to counter delocalization, and thus to speed up discovery.
Christopher A. Welty
ASE1
1997 Artificial Intelligence and Software Engineering: Breaking the Toy Mold
Christopher A. Welty, Peter G. Selfridge
Autom. Softw. Eng.1
1994 The Eighth Annual Knowledge-Based Software Engineering Conference
Christopher A. Welty
Autom. Softw. Eng.1