David M. Hart

dblp:29/6821 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3Artificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 50% Information retrieval · 33% Spatial and temporal data management · 17%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
query understanding
0.012004
Automatic recognition of reading levels from user queries · SIGIR 2004
Spatial and temporal data management › time series compression
piecewise linear approximation
0.012001
An Online Algorithm for Segmenting Time Series · ICDM 2001
Data mining › temporal data mining
time series mining
0.012001
An Online Algorithm for Segmenting Time Series · ICDM 2001
Data mining › temporal data mining › time series mining
time series representation
0.012001
An Online Algorithm for Segmenting Time Series · ICDM 2001
Data mining › time series analysis
time series segmentation
0.012001
An Online Algorithm for Segmenting Time Series · ICDM 2001

Methods — techniques the papers use, named apart from their topics

automatic reading level classification · 0.0online segmentation · 0.0
YearPublicationVenuePosition
2026 A framework for evaluating predicted sperm trajectories in crowded microscopy videos
abstract
Since the 1980s, semi-automated sperm motility analysis of phase contrast microscopy videos has been used to measure and categorize sperm motility patterns. Motility categories are determined from various kinematic parameters such as Curvilinear Velocity (VCL) and Beat Cross Frequency (BCF). These measures ultimately rely on the quality of the tracking for each individual sperm in the microscopy video. However, common approaches to sperm tracking require sample dilution and shortening the time window of observation (less than 1 to 2 seconds) to avoid tracking errors that occur when sperm cross paths. The post-ejaculatory lifespan of sperm can exceed several hours to days in some species, and long-term adaptive changes in motility pattern may be an important distinguishing factor for predictive modeling of sperm fertilizing competence. Improving the predictive value of computer assisted semen analysis will require accurate tracking of sperm trajectories over physiologically-relevant time scales and at the high cell densities typically found in semen. In this work, we identify a framework for accurately assessing the quality of sperm trajectory tracking that is independent of standard motility measures. We utilize cell tracking metrics adapted from the more common task of tracking adherent somatic cells and propose modifications based on the unique challenges of sperm video-microscopy. We also provide a small dataset of microscopy videos that includes 340 labeled sperm trajectories to allow for future comparisons and developments. Finally, we demonstrate that variations in configuration can lead to as much as a 30% improvement on metrics, showcasing their effectiveness at analyzing tracking quality.
David M. Hart, Kylie D. Cashwell, Anita Bhandari, Jayath Premasinghe, Cameron A. Schmidt
PLoS Comput. Biol.1
2025 Modeling diffusive search by non-adaptive sperm: Empirical and computational insights
abstract
During fertilization, mammalian sperm undergo a winnowing selection process that reduces the candidate pool of potential fertilizers from ~106-1011 cells to 101-102 cells (depending on the species). Classical sperm competition theory addresses the positive or 'stabilizing' selection acting on sperm phenotypes within populations of organisms but does not strictly address the developmental consequences of sperm traits among individual organisms that are under purifying selection during fertilization. It is the latter that is of utmost concern for improving assisted reproductive technologies (ART) because low-fitness sperm may be inadvertently used for fertilization during interventions that rely heavily on artificial sperm selection, such as intracytoplasmic sperm injection (ICSI). Importantly, some form of sperm selection is used in nearly all forms of ART (e.g., differential centrifugation, swim-up, or hyaluronan binding assays, etc.). To date, there is no unifying quantitative framework (i.e., theory of sperm selection) that synthesizes causal mechanisms of selection with observed natural variation in individual sperm traits. In this report, we reframe the physiological function of sperm as a collective diffusive search process and develop multi-scale computational models to explore the causal dynamics that constrain sperm fitness during fertilization. Several experimentally useful concepts are developed, including a probabilistic measure of sperm fitness as well as an information theoretic measure of the magnitude of sperm selection, each of which are assessed under systematic increases in microenvironmental selective pressure acting on sperm motility patterns.
Benjamin M. Brisard, Kylie D. Cashwell, Stephanie M. Stewart, Logan M. Harrison, Aidan C. Charles, Chelsea V. Dennis, Ivie R. Henslee, Ethan L. Carrow, Heather A. Belcher, Debajit Bhowmick, Paul W. Vos, Maciej Majka, Martin Bier, David M. Hart, Cameron A. Schmidt
PLoS Comput. Biol.14
2004 Automatic recognition of reading levels from user queries
abstract
No abstract available.
W. Bruce Croft, David M. Hart
SIGIR4
2002 Iterative Deepening Dynamic Time Warping for Time Series
abstract
1 Introduction Time series are a ubiquitous form of data occurring in virtually every scientific discipline and business application. There has been much recent work on adapting data mining algorithms to time series databases. For example, Das et al. attempt to show how association rules can be learned from time series [7]. Debregeas and Hebrail [8] demonstrate a technique for scaling up time series clustering algorithms to massive datasets. Keogh and Pazzani introduced a new, scalable time series classification algorithm [16]. Almost all algorithms that operate on time series data need to compute the similarity between them. Euclidean distance, or some extension or modification thereof, is typically used. However as we will demonstrate in Section 2.1, Euclidean distance can be an extremely brittle distance measure.
Selina Chu, Eamonn J. Keogh, David M. Hart, Michael J. Pazzani
SDM3
2001 An Online Algorithm for Segmenting Time Series
abstract
In recent years, there has been an explosion of interest in mining time-series databases. As with most computer science problems, representation of the data is the key to efficient and effective solutions. One of the most commonly used representations is piecewise linear approximation. This representation has been used by various researchers to support clustering, classification, indexing and association rule mining of time-series data. A variety of algorithms have been proposed to obtain this representation, with several algorithms having been independently rediscovered several times. In this paper, we undertake the first extensive review and empirical comparison of all proposed techniques. We show that all these algorithms have fatal flaws from a data-mining perspective. We introduce a novel algorithm that we empirically show to be superior to all others in the literature.
Eamonn J. Keogh, Selina Chu, David M. Hart, Michael J. Pazzani
ICDM3
1994 Tols for Experiments in Planning
abstract
The paper describes two separate but synergistic tools for running experiments on large Lisp systems such as artificial intelligence planning systems, by which we mean systems that produce plans and execute them in some kind of simulator. The first tool, called CLIP (Common Lisp Instrumentation Package), allows the researcher to define and run experiments, including experimental conditions (parameter values of the planner or simulator) and data to be collected. The data are written out to data files that can be analyzed by statistics software. The second tool, called CLASP (Common Lisp Analytical Statistics Package), allows the researcher to analyze data from experiments by using graphics, statistical tests, and various kinds of data manipulation. CLASP has a graphical user interface (using CLIM, the Common Lisp Manager) and also allows data to be directly processed by Lisp functions.>
Scott D. Anderson, Adam Carlson, David L. Westbrook, David M. Hart, Paul R. Cohen
ICTAI4
1990 Addressing Real-Time Constraints in the Design of Autonomous Agents
Adele E. Howe, David M. Hart, Paul R. Cohen
Real Time Syst.2
1989 A declarative representation of control knowledge
abstract
An explicit representation of control called strategy frames is described. The control of several well-known expert systems can be described in terms of strategy frames, although their control is actually encoded in an interpreter. One advantage of strategy frames is that complex control strategies emerge from their interaction, so complex interpreters are not necessary. This idea is illustrated in the context of a process control problem.>
Paul R. Cohen, Jefferson DeLisio, David M. Hart
IEEE Trans. Syst. Man Cybern.3