Steven Minton

dblp:m/SMinton · DBLP profile ↗
← Back
42ranked-venue papers
20as first author
0since 2021 · last 2018
0000-0001-8640-2905ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 20 first-authorDatabases, data management, data science and information retrieval · 13 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 13 · 10 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
10 papers
Data integration and cleaning · 77% Knowledge graphs · 11% Information retrieval · 4%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%
Artificial intelligence
15 papers
Information extraction and text analysis · 33% Planning, search and constraint satisfaction · 17% Learning paradigms · 13%

Topics — the 27 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational social science and digital humanities › social network analysis
criminal network analysis
0.312018
The Impact of Environmental Stressors on Human Trafficking · ICDM 2018
Data integration and cleaning › data extraction
web data extraction
0.222011
Materializing multi-relational databases from the web using taxonomic queries · WSDM 2011
Mixed-initiative, multi-source information assistants · WWW 2001
Data integration and cleaning › data generation
multi-relational data synthesis
0.112011
Materializing multi-relational databases from the web using taxonomic queries · WSDM 2011
Data integration and cleaning › data extraction
table extraction
0.112011
Materializing multi-relational databases from the web using taxonomic queries · WSDM 2011
Data integration and cleaning › data extraction › web data extraction
RSS feed generation
0.112006
Overview of AutoFeed: An Unsupervised Learning System for Generating Webfeeds · AAAI 2006
Knowledge graphs
entity linking
0.112005
A Heterogeneous Field Matching Method for Record Linkage · ICDM 2005
Natural language and speech › Information extraction and text analysis
web information extraction
0.012004
Using the Structure of Web Sites for Automatic Segmentation of Tables · SIGMOD Conference 2004
Natural language and speech › Information extraction and text analysis › web information extraction
wrapper induction
0.012003
Active Learning with Strong and Weak Views: A Case Study on Wrapper Induction · IJCAI 2003
Knowledge graphs › relation learning
relation discovery
0.012011
Materializing multi-relational databases from the web using taxonomic queries · WSDM 2011
Machine learning › Efficient and distributed learning
active learning
0.012002
Active + Semi-supervised Learning = Robust Multi-View Learning · ICML 2002
Machine learning › Representation and self-supervised learning
multi-view learning
0.012002
Active + Semi-supervised Learning = Robust Multi-View Learning · ICML 2002
Machine learning › Learning paradigms
semi-supervised learning
0.012002
Active + Semi-supervised Learning = Robust Multi-View Learning · ICML 2002
Data integration and cleaning
object identification
0.012002
Learning domain-independent string transformation weights for high accuracy object identification · KDD 2002
Data integration and cleaning › data transformation
string transformation
0.012002
Learning domain-independent string transformation weights for high accuracy object identification · KDD 2002
Data mining › text mining
information extraction
0.012001
Mixed-initiative, multi-source information assistants · WWW 2001
Information retrieval
search engines
0.012001
Mixed-initiative, multi-source information assistants · WWW 2001
Data integration and cleaning
web data integration
0.022001
ARIADNE: A System for Constructing Mediators for Internet Sources · SIGMOD Conference 1998
Mixed-initiative, multi-source information assistants · WWW 2001
Machine learning › Trustworthy machine learning › interpretability
explanation-based learning
0.031991
Integrating Abstraction and Explanation-Based Learning in PRODIGY · AAAI 1991
Quantitative Results Concerning the Utility of Explanation-based Learning · Artif. Intell. 1990
Explanation-Based Learning: A Problem Solving Perspective · Artif. Intell. 1989
Web and social media mining
web mining
0.012004
Using the Structure of Web Sites for Automatic Segmentation of Tables · SIGMOD Conference 2004
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
heuristic selection
0.011993
Integrating Heuristics for Constraint Satisfaction Problems: A Case Study · AAAI 1993
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
scheduling
0.021992
Solving Large-Scale Constraint-Satisfaction and Scheduling Problems Using a Heuristic Repair Method · AAAI 1990
Minimizing Conflicts: A Heuristic Repair Method for Constraint Satisfaction and Scheduling Problems · Artif. Intell. 1992
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › classical planning
partial-order planning
0.011992
Total Order vs. Partial Order Planning: Factors Influencing Performance · KR 1992
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
hierarchical planning
0.011991
Integrating Abstraction and Explanation-Based Learning in PRODIGY · AAAI 1991
Query processing and optimization
query rewriting
0.011998
ARIADNE: A System for Constructing Mediators for Internet Sources · SIGMOD Conference 1998
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › generalized planning
plan generalization
0.011985
Selectively Generalizing Plans for Problem-Solving · IJCAI 1985
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.011985
Controlling Search in Flexible Parsing · IJCAI 1985
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › machine learning for planning
plan learning
0.011984
Constraint-Based Generalization: Learning Game-Playing Plans From Single Examples · AAAI 1984

Methods — techniques the papers use, named apart from their topics

probabilistic soft logic · 0.3probabilistic graphical model · 0.3collective inference · 0.3taxonomic queries · 0.1join discovery · 0.1active learning · 0.1probabilistic inference · 0.1constraint satisfaction · 0.1unsupervised learning · 0.1machine learning · 0.1expert-like rule learning · 0.1weak supervision · 0.0supervised learning · 0.0semi-supervised learning · 0.0constraint propagation · 0.0heuristic repair · 0.0explanation-based learning · 0.0heuristic integration · 0.0
YearPublicationVenuePosition
2018 The Impact of Environmental Stressors on Human Trafficking
abstract
Severe environmental events have extreme effects on all segments of society, including criminal activity. Extreme weather events, such as tropical storms, fires, and floods create instability in communities, and can be exploited by criminal organizations. Here we investigate the potential impact of catastrophic storms on the criminal activity of human trafficking. We propose three theories of how these catastrophic storms might impact trafficking and provide evidence for each. Researching human trafficking is made difficult by its illicit nature and the obscurity of high-quality data. Here, we analyze online advertisements for services which can be collected at scale and provide insights into traffickers' behavior. To successfully combine relevant heterogenous sources of information, as well as spatial and temporal structure, we propose a collective, probabilistic approach. We implement this approach with Probabilistic Soft Logic, a probabilistic programming framework which can flexibly model relational structure and for which inference of future locations is highly efficient. Furthermore, this framework can be used to model hidden structure, such as latent links between locations. Our proposed approach can model and predict how traffickers move. In addition, we propose a model which learns connections between locations. This model is then adapted to have knowledge of environmental events, and we demonstrate that incorporating knowledge of environmental events can improve prediction of future locations. While we have validated our models on the impact of severe weather on human trafficking, we believe our models can be generalized to a variety of other settings in which environmental events impact human behavior.
Sabina Tomkins, Golnoosh Farnadi, Brian Amanatullah, Lise Getoor, Steven Minton
ICDM5
2015 Building and Using a Knowledge Graph to Combat Human Trafficking
Pedro A. Szekely, Craig A. Knoblock, Jason Slepicka, Andrew Philpot, Chengye Yin, Dipsy Kapoor, Premkumar Natarajan, Daniel Marcu, Kevin Knight, David Stallard, Subessware S. Karunamoorthy, Rajagopal Bojanapalli, Steven Minton, Brian Amanatullah, Todd Hughes, Mike Tamayo, David Flynt, Rachel Artiss, Shih-Fu Chang, Tao Chen 0015, Gerald Hiebel, Lidia Silva Ferreira
ISWC (2)14
2011 Monitoring Entities in an Uncertain World: Entity Resolution and Referential Integrity
Steven Minton, Sofus A. Macskassy, Peter M. LaMonica, Kane See, Craig A. Knoblock, Greg Barish, Matthew Michelson, Raymond A. Liuzzi
IAAI1
2011 Materializing multi-relational databases from the web using taxonomic queries
abstract
Recently, much attention has been given to extracting tables from Web data. In this problem, the column definitions and tuples (such as what "company" is headquartered in what "city,") are extracted from Web text, structured Web data such as lists, or results of querying the deep Web, creating the table of interest. In this paper, we examine the problem of extracting and discovering multiple tables in a given domain, generating a truly multi-relational database as output. Beyond discovering the relations that define single tables, our approach discovers and leverages "within column" set membership relations, and discovers relations across the extracted tables (e.g., joins). By leveraging within-column relations our method can extract table instances that are ambiguous or rare, and by discovering joins, our method generates truly multi-relational output. Further, our approach uses taxonomic queries to bootstrap the extraction, rather than the more traditional "seed instances." Creating seeds often requires more domain knowledge than taxonomic queries, and previous work has shown that extraction methods may be sensitive to which input seeds they are given. We test our approach on two real world domains: NBA basketball and cancer information. Our results demonstrate that our approach generates databases of relevant tables from disparate Web information, and discovers the relations between them. Further, we show that by leveraging the "within column" relation our approach can identify a significant number of relevant tuples that would be difficult to do so otherwise.
Matthew Michelson, Sofus A. Macskassy, Steven Minton, Lise Getoor
WSDM3
2009 Interactive Data Integration through Smart Copy & Paste
Zachary G. Ives, Craig A. Knoblock, Steven Minton, Marie Jacob, Partha P. Talukdar, Rattapoom Tuchinda, José Luis Ambite, Maria Muslea, Cenk Gazen
CIDR3
2006 Overview of AutoFeed: An Unsupervised Learning System for Generating Webfeeds
Bora Gazen, Steven Minton
AAAI2
2006 Active Learning with Multiple Views
abstract
Active learners alleviate the burden of labeling large amounts of data by detecting and asking the user to label only the most informative examples in the domain. We focus here on active learning for multi-view domains, in which there are several disjoint subsets of features (views), each of which is sufficient to learn the target concept. In this paper we make several contributions. First, we introduce Co-Testing, which is the first approach to multi-view active learning. Second, we extend the multi-view learning framework by also exploiting weak views, which are adequate only for learning a concept that is more general/specific than the target concept. Finally, we empirically show that Co-Testing outperforms existing active learners on a variety of real world domains such as wrapper induction, Web page classification, advertisement removal, and discourse tree parsing.
Ion Muslea, Steven Minton, Craig A. Knoblock
J. Artif. Intell. Res.2
2005 A Heterogeneous Field Matching Method for Record Linkage
abstract
Record linkage is the process of determining that two records refer to the same entity. A key subprocess is evaluating how well the individual fields, or attributes, of the records match each other. One approach to matching fields is to use hand-written domain-specific rules. This "expert systems" approach may result in good performance for specific applications, but it is not scalable. This paper describes a new machine learning approach that creates expert-like rules for field matching. In our approach, the relationship between two field values is described by a set of heterogeneous transformations. Previous machine learning methods used simple models to evaluate the distance between two fields. However, our approach enables more sophisticated relationships to be modeled, which better capture the complex domain specific, common-sense phenomena that humans use to judge similarity. We compare our approach to methods that rely on simpler homogeneous models in several domains. By modeling more complex relationships we produce more accurate results.
Steven Minton, Claude J. Nanjo, Craig A. Knoblock, Martin Michalowski, Matthew Michelson
ICDM1
2005 AutoFeed: an unsupervised learning system for generating webfeeds
abstract
Our goal is to automatically extract data from semi-structured webn sites. Previously, researchers have developed two types of supervised learning approaches for extracting web data: methods that create precise, site-specific extraction rules and methods that learn less-precise site-independent extraction rules. In either case, significant training is required. In this paper, we describe a third, more ambitious approach, where we use unsupervised learning to analyze sites and discover their structure. Our method relies on a set of heterogeneous "experts", each of which is capable of identifying certain types of generic structure. Each expert represents its discoveries as "hints". Based on these hints, our system clusters the pages and identifies semi-structured data that can be extracted. To identify a good clustering, we use a probabilistic model of the hint-generation process. The paper describes our formulation of the fully-automatic web-extraction problem, our clustering approach, and our results on a set of experiments.
Bora Gazen, Steven Minton
K-CAP2
2004 Using the Structure of Web Sites for Automatic Segmentation of Tables
abstract
Many Web sites, especially those that dynamically generate HTML pages to display the results of a user's query, present information in the form of list or tables. Current tools that allow applications to programmatically extract this information rely heavily on user input, often in the form of labeled extracted records. The sheer size and rate of growth of the Web make any solution that relies primarily on user input is infeasible in the long term. Fortunately, many Web sites contain much explicit and implicit structure, both in layout and content, that we can exploit for the purpose of information extraction. This paper describes an approach to automatic extraction and segmentation of records from Web tables. Automatic methods do not require any user input, but rely solely on the layout and content of the Web source. Our approach relies on the common structure of many Web sites, which present information as a list or a table, with a link in each entry leading to a detail page containing additional information about that item. We describe two algorithms that use redundancies in the content of table and detail pages to aid in information extraction. The first algorithm encodes additional information provided by detail pages as constraints and finds the segmentation by solving a constraint satisfaction problem. The second algorithm uses probabilistic inference to find the record segmentation. We show how each approach can exploit the web site structure in a general, domain-independent manner, and we demonstrate the effectiveness of each algorithm on a set of twelve Web sites.
Kristina Lerman, Lise Getoor, Steven Minton, Craig A. Knoblock
SIGMOD Conference3
2003 Active Learning with Strong and Weak Views: A Case Study on Wrapper Induction
Ion Muslea, Steven Minton, Craig A. Knoblock
IJCAI2
2003 Wrapper Maintenance: A Machine Learning Approach
abstract
The proliferation of online information sources has led to an increased use of wrappers for extracting data from Web sources. While most of the previous research has focused on quick and efficient generation of wrappers, the development of tools for wrapper maintenance has received less attention. This is an important research problem because Web sources often change in ways that prevent the wrappers from extracting data correctly. We present an efficient algorithm that learns structural information about data from positive examples alone. We describe how this information can be used for two wrapper maintenance applications: wrapper verification and reinduction. The wrapper verification system detects when a wrapper is not extracting correct data, usually because the Web source has changed its format. The reinduction algorithm automatically recovers from changes in the Web source by identifying data on Web pages so that a new wrapper may be generated for this source. To validate our approach, we monitored 27 wrappers over a period of a year. The verification algorithm correctly discovered 35 of the 37 wrapper changes, and made 16 mistakes, resulting in precision of 0.73 and recall of 0.95. We validated the reinduction algorithm on ten Web sources. We were able to successfully reinduce the wrappers, obtaining precision and recall values of 0.90 and 0.80 on the data extraction task.
Kristina Lerman, Steven Minton, Craig A. Knoblock
J. Artif. Intell. Res.2
2002 Active + Semi-supervised Learning = Robust Multi-View Learning
Ion Muslea, Steven Minton, Craig A. Knoblock
ICML2
2002 Adaptive View Validation: A First Step Towards Automatic View Detection
Ion Muslea, Steven Minton, Craig A. Knoblock
ICML2
2002 Learning domain-independent string transformation weights for high accuracy object identification
abstract
The task of object identification occurs when integrating information from multiple websites. The same data objects can exist in inconsistent text formats across sites, making it difficult to identify matching objects using exact text match. Previous methods of object identification have required manual construction of domain-specific string transformations or manual setting of general transformation parameter weights for recognizing format inconsistencies. This manual process can be time consuming and error-prone. We have developed an object identification system called Active Atlas [18], which applies a set of domain-independent string transformations to compare the objects' shared attributes in order to identify matching objects. In this paper, we discuss extensions to the Active Atlas system, which allow it to learn to tailor the weights of a set of general transformations to a specific application domain through limited user input. The experimental results demonstrate that this approach achieves higher accuracy and requires less user involvement than previous methods across various application domains.
Sheila Tejada, Craig A. Knoblock, Steven Minton
KDD3
2001 AN intelligent user interface for mixed-initiative multi-source travel planning
abstract
Amixed-initiative plannerin our context is one in which either the human or the computer can spontaneously provide the content of the same input fields. A multi-source planneris one that accesses multiple external information sources in parallel, using separate threads. This type of highly dynamic user interface is desirable but presents a challenge in "keeping the user in control" because it can be confusing to understand which fields of the form currently "belong" to the user, which ones "belong" to the system, how these two interact, and when and how their ownership changes.
Maria Muslea, Jean Oh, Steven Minton, Craig A. Knoblock
IUI4
2001 Mixed-initiative, multi-source information assistants
abstract
While the information resources on the Web are vast, the sources are often hard to nd, painful to use, and dicult to integrate. Wehavedeveloped the Heracles framework for building Web-based information assistants. This framework provides the infrastructure to rapidly construct new applications that extract information from multiple Web sources and interactively integrate the data using a dynamic, hierarchical constraint network. This paper describes the core technologies that comprise the framework, including information extraction, hierarchical template representation, and constraint propagation. In addition, we present an application of this framework, the ###### #########, which is an interactivetravel planning system. We also briey describe our experience using the same framework to build a second application, the ######### #########, which extracts and integrates geographic-related data about countries thorughout the world. We believe these types of information assistants provide a signicant step forward in fully exploiting the information available on the Internet. 1.
Craig A. Knoblock, Steven Minton, José Luis Ambite, Maria Muslea, Jean Oh
WWW2
2001 Hierarchical Wrapper Induction for Semistructured Information Sources
Ion Muslea, Steven Minton, Craig A. Knoblock
Auton. Agents Multi Agent Syst.2
2001 The Ariadne Approach to Web-Based Information Integration
abstract
The Web is based on a browsing paradigm that makes it difficult to retrieve and integrate data from multiple sites. Today, the only way to do this is to build specialized applications, which are time-consuming to develop and difficult to maintain. We have addressed this problem by creating the technology and tools for rapidly constructing information agents that extract, query, and integrate data from web sources. Our approach is based on a uniform representation that makes it simple and efficient to integrate multiple sources. Instead of building specialized algorithms for handling web sources, we have developed methods for mapping web sources into this uniform representation. This approach builds on work from knowledge representation, databases, machine learning and automated planning. The resulting system, called Ariadne, makes it fast and easy to build new information agents that access existing web sources. Ariadne also makes it easy to maintain these agents and incorporate new sources as they become available.
Craig A. Knoblock, Steven Minton, José Luis Ambite, Naveen Ashish, Ion Muslea, Andrew Philpot, Sheila Tejada
Int. J. Cooperative Inf. Syst.2
2001 Learning object identification rules for information integration
Sheila Tejada, Craig A. Knoblock, Steven Minton
Inf. Syst.3
2000 TheaterLoc: Using Information Integration Technology to Rapidly Build Virtual Applications
abstract
Although much has been written about various information integration technologies, little has been said regarding how to combine these technologies together to build an entire application. We demonstrate TheaterLoc, an information integration application that allows users to retrieve information about theatres and restaurants for various U.S. cities, including an interactive map depicting their relative locations. The data retrieved by TheaterLoc comes from five distinct heterogeneous and distributed sources. The enabling technology used to achieve the integration includes: the Ariadne information mediator, a Web site wrapper learning tool, the Theseus execution system, and a mechanism for distributed spatial query planning. Our system is novel because it demonstrates how "virtual applications" can be rapidly built from a set of integration tools and existing online data sources.
Greg Barish, Yi-Shin Chen, Dan DiPasquo, Craig A. Knoblock, Steven Minton, Ion Muslea, Cyrus Shahabi
ICDE5
1998 ARIADNE: A System for Constructing Mediators for Internet Sources
abstract
The Web is based on a browsing paradigm that makes it difficult to retrieve and integrate data from multiple sites. Today, the only way to achieve this integration is by building specialized applications, which are time-consuming to develop and difficult to maintain. We are addressing this problem by creating the technology and tools for rapidly constructing information mediators that extract, query, and integrate data from web sources. The resulting system, called Ariadne, makes it feasible to rapidly build information mediators that access existing web sources.
José Luis Ambite, Naveen Ashish, Greg Barish, Craig A. Knoblock, Steven Minton, Pragnesh Jay Modi, Ion Muslea, Andrew Philpot, Sheila Tejada
SIGMOD Conference5
1997 Configurable Solvers: Tailoring General Methods to Specific Applications
Steven Minton
CP1
1994 Small is Beautiful: A Brute-Force Approach to Learning First-Order Formulas
Steven Minton, Ian Underwood
AAAI1
1994 Total-Order and Partial-Order Planning: A Comparative Analysis
abstract
For many years, the intuitions underlying partial-order planning were largely taken for granted. Only in the past few years has there been renewed interest in the fundamental principles underlying this paradigm. In this paper, we present a rigorous comparative analysis of partial-order and total-order planning by focusing on two specific planners that can be directly compared. We show that there are some subtle assumptions that underly the wide-spread intuitions regarding the supposed efficiency of partial-order planning. For instance, the superiority ofpartial-order planning can depend critically upon the search strategy and the structure of the search space. Understanding the underlying assumptions is crucial for constructing efficient planners.
Steven Minton, John L. Bresina, Mark Drummond
J. Artif. Intell. Res.1
1993 Integrating Heuristics for Constraint Satisfaction Problems: A Case Study
Steven Minton
AAAI1
1993 An Analytic Learning System for Specializing Heuristics
Steven Minton
IJCAI1
1992 Why EBL Produces Overly-Specific Knowledge: A Critique of the PRODIGY Approaches
Oren Etzioni, Steven Minton
ML2
1992 Total Order vs. Partial Order Planning: Factors Influencing Performance
Steven Minton, Mark Drummond, John L. Bresina, Andrew B. Philips
KR1
1992 Minimizing Conflicts: A Heuristic Repair Method for Constraint Satisfaction and Scheduling Problems
Steven Minton, Mark D. Johnston, Andrew B. Philips, Philip Laird
Artif. Intell.1
1991 Integrating Abstraction and Explanation-Based Learning in PRODIGY
Craig A. Knoblock, Steven Minton, Oren Etzioni
AAAI2
1991 Commitment Strategies in Planning: A Comparative Analysis
Steven Minton, John L. Bresina, Mark Drummond
IJCAI1
1991 A Reply to Zito-Wolf's Book Review of Learning Search Control Knowledge: An Explanation-Based Approach
Steven Minton
Mach. Learn.1
1990 Solving Large-Scale Constraint-Satisfaction and Scheduling Problems Using a Heuristic Repair Method
Steven Minton, Mark D. Johnston, Andrew B. Philips, Philip Laird
AAAI1
1990 Issues in the Design of Operator Composition Systems
Steven Minton
ML1
1990 Quantitative Results Concerning the Utility of Explanation-based Learning
Steven Minton
Artif. Intell.1
1989 Explanation-Based Learning: A Problem Solving Perspective
Steven Minton, Jaime G. Carbonell, Craig A. Knoblock, Daniel Kuokka, Oren Etzioni, Yolanda Gil
Artif. Intell.1
1988 Quantitative Results Concerning the Utility of Explanation-Based Learning
Steven Minton
AAAI1
1987 Strategies for Learning Search Control Rules: An Explanation-based Approach
Steven Minton, Jaime G. Carbonell
IJCAI1
1985 Selectively Generalizing Plans for Problem-Solving
Steven Minton
IJCAI1
1985 Controlling Search in Flexible Parsing
Steven Minton, Philip J. Hayes, Jill Fain
IJCAI1
1984 Constraint-Based Generalization: Learning Game-Playing Plans From Single Examples
Steven Minton
AAAI1