VLDB 2026 Research / reviewers in the wild / expert
Jennifer L. Leopold
dblp:68/6646
· DBLP profile ↗
23ranked-venue papers
1as first author
2since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 1 since 2021Computer networks · 2Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data mining · 46% Spatial and temporal data management · 46% Data stream processing · 4% | |
| Artificial intelligence
1 paper |
Graph learning · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning › graph neural network
graph convolutional network |
0.4 | 1 | 2019 | Beyond Geo-First Law: Learning Spatial Representations via Integrated Autocorrelations and Complementarity · ICDM 2019 |
Machine learning › Graph learning
graph neural network |
0.4 | 1 | 2019 | Beyond Geo-First Law: Learning Spatial Representations via Integrated Autocorrelations and Complementarity · ICDM 2019 |
Data mining › structured data mining
spatial data mining |
0.4 | 1 | 2019 | Beyond Geo-First Law: Learning Spatial Representations via Integrated Autocorrelations and Complementarity · ICDM 2019 |
Spatial and temporal data management › spatial representation
spatial representation learning |
0.4 | 1 | 2019 | Beyond Geo-First Law: Learning Spatial Representations via Integrated Autocorrelations and Complementarity · ICDM 2019 |
Data stream processing
continuous query processing |
0.0 | 1 | 2002 | Webformulate: a web-based visual continual query system · WWW 2002 |
Data integration and cleaning
heterogeneous data sources |
0.0 | 1 | 2002 | Webformulate: a web-based visual continual query system · WWW 2002 |
User interface design and tools › visual programming
visual query formulation |
0.0 | 1 | 2002 | Webformulate: a web-based visual continual query system · WWW 2002 |
Methods — techniques the papers use, named apart from their topics
multi-view graph construction · 0.8adversarial autoencoder · 0.8spreadsheet-based query formulation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | NodeSense2Vec: Spatiotemporal Context-Aware Network Embedding for Heterogeneous Urban Mobility DataabstractThe problem of learning latent representations of heterogeneous networks with spatial and temporal attributes has been gaining traction in recent years, given its myriad of real-world applications. Most systems with applications in the field of transportation, urban economics, medical information, online e-commerce, etc., handle big data that can be structured into Spatiotemporal Heterogeneous Networks (SHNs), thereby making efficient analysis of these networks extremely vital.In this paper, we propose a spatiotemporal context-aware network embedding framework that jointly captures the spatial regularities between objects and the sequential transition patterns of human mobility. First, we model the heterogeneous urban mobility data collected from multiple sources as an SHN using a probabilistic weighted degree centrality measure. To learn the sequential transition patterns of human mobility in urban regions, we perform meta-path constrained random walks (MPCRWs) on the constructed SHN, which captures the proximities between multi-typed objects via their rich spatiotemporal links. By treating the generated meta-path instances as sentences, we capture multiple contrastive context senses associated with nodes in an SHN produced due to multiplex of spatial and temporal dependencies between objects in urban mobility data by performing spectral graph clustering. We then map the learned contrastive contextual node senses with respective meta-path instances. Finally, we learn latent embeddings of the mapped meta-path instances by using the word2vec model Skip-gram. We evaluate the performance of our proposed model on real-world application problems. Experimental results demonstrate the effectiveness of our model over state-of-the-art alternatives. Dakshak Keerthi Chandra, Jennifer L. Leopold, Yanjie Fu |
IEEE BigData | 2 |
| 2021 | Discriminative Pattern Mining for Runtime Security Enforcement of Cyber-Physical Point-of-Care Medical TechnologyabstractPoint-of-care diagnostics are a key technology for various safety-critical applications from providing diagnostics in developing countries lacking adequate medical infrastructure to fight infectious diseases to screening procedures for border protection. Digital microfluidics biochips are an emerging technology that are increasingly being evaluated as a viable platform for rapid diagnosis and point-of-care field deployment. In such a technology, processing errors are inherent. Cyber-physical digital biochips offer higher reliability through the inclusion of automated error recovery mechanisms that can reconfigure operations performed on the electrode array. Recent research has begun to explore security vulnerabilities of digital microfluidic systems. This paper expands previous work that exploits vulnerabilities due to implicit trust in the error recovery mechanism. In this work, a discriminative data mining approach is introduced to identify frequent bioassay operations that can be cyber-physically attested for runtime security protection. Fred Love, Jennifer L. Leopold, Bruce M. McMillin |
COMPSAC | 2 |
| 2020 | Collective Embedding with Feature Importance: A Unified Approach for Spatiotemporal Network EmbeddingabstractIn the last decade, there has been great progress in the field of machine learning and deep learning. These models have been instrumental in addressing a great number of problems. However, they have struggled when it comes to dealing with high dimensional data. In recent years, representation learning models have proven to be quite efficient in addressing this problem as they are capable of capturing effective lower-dimensional representations of the data. However, most of the existing models are quite ineffective when it comes to dealing with high dimensional spatiotemporal data as they encapsulate complex spatial and temporal relationships that exist among real-world objects. High-dimensional spatiotemporal data of cities represent urban communities. By learning their social structure we can better quantitatively depict them and understand factors influencing rapid growth, expansion, and changes. Dakshak Keerthi Chandra, Pengyang Wang, Jennifer L. Leopold, Yanjie Fu |
CIKM | 3 |
| 2019 | Collective Representation Learning on Spatiotemporal Heterogeneous Information NetworksabstractRepresentation learning is a technique that is used to capture the underlying latent features of complex data. Representation learning on networks has been widely implemented for learning network structure and embedding it in a low dimensional vector space. In recent years, network embedding using representation learning has attracted increasing attention, and many deep architectures have been widely proposed. However, existing network embedding techniques ignore the multi-class spatial and temporal relationships that crucially reflect the complex nature among vertices and links in spatiotemporal heterogeneous information networks(SHINs). Dakshak Keerthi Chandra, Pengyang Wang, Jennifer L. Leopold, Yanjie Fu |
SIGSPATIAL/GIS | 3 |
| 2019 | Beyond Geo-First Law: Learning Spatial Representations via Integrated Autocorrelations and ComplementarityabstractSpatial representation learning (SRL) is to automatically learn feature representations that characterize spatial entities. In this paper, we study the problem of improving spatial representation learning using spatial structure knowledge. We consider two types of structure knowledge: (1) spatial autocorrelations refer to the pattern that similar spatial entities are more likely to share similar roles and configurations. (2) spatial complementarity refers to the effect that the role of a spatial entity can be complemented and augmented by other different yet compatible spatial entities. Along this line, we develop a step-by-step SRL framework to integrate spatial autocorrelations and complementarity. This framework includes four testable steps. First, we construct multi-view POI-POI(Point of Interest) graphs to characterize the static and dynamic patterns of each spatial region. We then use the graphs as inputs to train an adversarial autoencoder (AAE) that can preserve the spatial autocorrelation property and learn representations of spatial entities. Later, with the learned representations extracted from AAE inputs, a Graph Convolutional Network (GCN) is trained in an unsupervised fashion in order to overcome label sparsity and capture the spatial complementarity effect. In this way, we significantly improve the quality of spatial representations. In addition, we apply the proposed method to characterize residential communities for predicting real estate prices. Finally, we present intensive experimental results with real-world real estate data to demonstrate the proposed method effectiveness. Jiadi Du, Yunchao Zhang, Pengyang Wang, Jennifer L. Leopold, Yanjie Fu |
ICDM | 4 |
| 2017 | Graph compaction in analyzing large scale online social networksabstractThe real-world large scale networks motivate the need for parallel and distributed evaluation of network analysis and computational tasks for computational efficiency and application effectiveness. One of the essential tasks for parallel and distributed evaluation, is to have partitions over the underlying network graph. Over these partitions the computational or network analysis tasks are in turn processed in a distributed or parallel manner. It is interesting to use intrinsic communities of social networks as partitions, to be used as basic components in parallel and distributed computation. We propose two novel graph compaction algorithms that generate the desired compact graph of communities as a preprocessing stage to the parallel and distributed evaluation of computational tasks. To comply with heterogeneity in community structure and size, we use a flexible limit on them. We evaluate the structure and quality of our algorithms and hence its resulting communities over two distinct application networks. We show that the generated community structure, reasonably complies with the modular structure of the network. We evaluate the quality of the partitions, relative to the partitions generated using existing state-of-the-art approach, and compare the approaches to show better quality of our partitions in terms of number of graph cuts. Sima Das, Jennifer L. Leopold, Susmita Ghosh, Sajal K. Das 0001 |
ICC | 2 |
| 2016 | Graph Partitioning in Parallelization of Large Scale NetworksabstractReal world large scale networks exhibit intrinsic community structure, with dense intra-community connectivity and sparse inter-community connectivity. Leveraging their community structure for parallelization of computational tasks and applications, is a significant step towards computational efficiency and application effectiveness. We propose a weighted depth-first-search graph partitioning algorithm for community formation that preserves the needed community dependency without any cycles. To comply with heterogeneity in community structure and size of the real world networks, we use a flexible limiting value for them. Further, our algorithm is a diversion from the existing modularity based algorithms. We evaluate our algorithm as the quality of the generated partitions, measured in terms of number of graph cuts. Sima Das, Jennifer L. Leopold, Susmita Ghosh, Sajal K. Das 0001 |
LCN | 2 |
| 2013 | Efficient determination of spatial relations using composition tables and decision treesabstractIn order for Qualitative Spatial Reasoning applications to be both useful and usable, the information feedback loop between the computational engine and the user must be as seamless as possible. Inherently, computational geometry can be quite expensive, and every effort must be made to avoid inefficient or unnecessary calculations. Within the field of Region Connection Calculi, the 9-Intersection model often is used to determine the spatial relation between two regions. Consequently, optimization efforts typically focus on calculations involving the intersections between the interiors, boundaries, and exteriors of the regions, or the use of composition tables to narrow down the possibilities for the relations that can hold between two regions. The few implementations of spatial reasoners that have been attempted have been simply proofs-of-concept and/or have been limited to two dimensions. Herein we present a novel approach that combines the use of composition tables and decision trees to efficiently determine the spatial relation between two objects in 3D considering both connectivity and obscuration. This approach has been fully implemented for the VRCC-3D+ spatial reasoning system, and benchmarks are included to corroborate our claims of efficiency. Nathan W. Eloe, Jennifer L. Leopold, Chaman L. Sabharwal, Douglas McGeehan |
CIMSIVP | 2 |
| 2013 | Smooth transition neighborhood graphs for 3D spatial relationsabstractDistance between two relations can be defined by using some metric based on the qualitative or quantitative representation of the relations [1]. However, qualitative distances cannot be expressed by conventional measures. Most differentiating measures are derived from observation and experience in an ad hoc manner. The outcomes are cognitively acceptable only if they match the user's concept of distance. We have designed an algorithm based on heuristics to derive a conceptual neighborhood supporting smooth transitions between the relations. Herein we present the results of applying the algorithm to the well-known region connection calculus, RCC-8, and to an additional model, VRCC-3D+, that considers both 3D connectivity and obscuration. Chaman L. Sabharwal, Jennifer L. Leopold |
CIMSIVP | 2 |
| 2012 | Protein secondary structure prediction using BLAST and exhaustive RT-RICO, the search for optimal segment length and thresholdabstractProtein secondary structure prediction from its amino acid sequence is a well studied computational problem in bioinformatics and data mining. It can be viewed as an intermediate research objective to solving the more challenging protein three-dimensional structure prediction problem, which is one of the most important research goals of bioinformatics. Although the secondary structure prediction problem was first defined in the 1960s, the prediction accuracy of the most modern methods still hovers around 80%. In [1] this research team presented a protein secondary structure prediction method, BLAST-RT-RICO (Relaxed Threshold Rule Induction from Coverings), that employs a modified association rule learning approach, utilizing multiple sequence alignment information, to predict secondary structures. Despite producing higher prediction accuracy than many other contemporary methods, that preliminary research study identified some crucial areas in need of improvements, such as determining the optimal segment length, finding the optimal threshold value, and improving the time complexity for the rule generation algorithm. In this paper, we present a modified method, BLAST-ERT-RICO (Exhaustive Relaxed Threshold Rule Induction from Coverings), which has an improved time complexity, as well as more optimal choices of segment length and threshold value. Preliminary test results showed that with a segment length of 9 amino acid residues, and a threshold value of 0.8, BLAST-ERT-RICO achieved a Q3score of 92.19% on the standard test dataset RS126, which suggests that this approach may be even more useful as a secondary structure prediction method in the future. Leong Lee, Jennifer L. Leopold, Ronald L. Frank |
CIBCB | 2 |
| 2012 | Exhaustive RT-RICO algorithm for mining association rules in protein secondary structureabstractPrediction of a protein's secondary structure from its amino acid sequence is a well studied computational problem in bioinformatics, and has significant practical research value. Although the secondary structure prediction problem was first defined almost fifty years ago, the accuracy of most modern methods still hovers around 80%. In [1] this research team presented a promising protein secondary structure prediction method, BLAST-RT-RICO (Relaxed Threshold Rule Induction from Coverings), that employs a modified association rule learning approach, utilizing multiple sequence alignment information. BLAST-RT-RICO achieved Q3scores of 89.93% and 87.71% on the standard test datasets RS126 and CB396, respectively. However, there were some areas of the algorithm that were in need of improvement; most importantly, the time complexity for the rule generation step needed to be reduced. Recently, we developed a modified rule generation algorithm, ERT-RICO (Exhaustive Relaxed Threshold Rule Induction from Coverings), that addresses this issue. The research team now is able to run much larger test datasets with different choices of segment length and threshold value; preliminary test results achieved a Q3score of 92.19% on the standard test dataset RS126. The modified algorithm, its mathematical definitions, and the improved time/space complexity are discussed in this paper. Leong Lee, Jennifer L. Leopold, Ronald L. Frank |
CIBCB | 2 |
| 2012 | Guest Editors' Introduction
Gennaro Costagliola, Jennifer L. Leopold, Paolo Nesi, Kia Ng |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2011 | ROAR: A Reference Ontology for Anatomical RelationsabstractThe ontology has become a useful model for organizing knowledge. This is particularly true in the field of biomedicine, where individual ontologies have been created for specific data domains ranging from genomics through species morphologies to human anatomical reference ontologies. Although specific sets of relationships have been proposed to improve the accuracy and consistency of such ontologies, there has been little to nothing proposed concerning the organization of those relationships. To help address this deficiency, herein we present a Reference Ontology of Anatomical Relations (ROAR). ROAR extends the concepts used in existing biomedical ontologies by defining and hierarchically organizing temporal, spatial, functional, and taxonomic relations based on generalization/specialization and semantic relatedness. Also provided in this paper are examples of how the use of such a reference ontology would significantly increase the ease with which data from multiple ontologies could be developed and integrated and would improve the information base for other computational intelligence activities. Alton B. Coalter, Jennifer L. Leopold |
CIBCB | 2 |
| 2011 | Protein secondary structure prediction using BLAST and Relaxed Threshold Rule Induction from CoveringsabstractProtein structure prediction has been a very important and challenging research problem in bioinformatics for years. Yet the determination of protein structures by time-consuming and relatively expensive experimental methods continues to lag far behind the explosive discovery of protein sequences. With the recent breakthrough of combining multiple sequence alignment information and artificial intelligence algorithms to predict protein secondary structure, the Q3accuracy of the best computational prediction methods has finally exceeded 80%. Herein we present a rule-based data-mining approach called BLAST-RT-RICO (Relaxed Threshold Rule Induction from Coverings) that utilizes multiple sequence alignment information to predict protein secondary structure. This method uses the PSI-BLAST algorithm to identify suitable proteins, and then generates rules from these proteins that can be used to predict secondary structure. By also utilizing known homologous template secondary structures in the Protein Data Bank (PDB) database, BLAST-RT-RICO achieved a Q3score of 89.93% on the standard test dataset RS126 and a Q3score of 87.71% on the standard test dataset CB396. These successful preliminary results suggest that this rule-based method may be the foundation for even more accurate prediction of protein secondary structure in the future. Leong Lee, Jennifer L. Leopold, Ronald L. Frank |
CIBCB | 2 |
| 2010 | RCC-3D: Qualitative Spatial Reasoning in 3D
Julia Albath, Jennifer L. Leopold, Chaman L. Sabharwal, Anne M. Maglia |
CAINE | 2 |
| 2010 | Representation and validation of domain and range restrictions in a relational database-driven ontology maintenance systemabstractAn ontology can be used to represent and organize the objects, properties, events, processes, and relations that embody an area of reality [1]. These knowledge bases may be created manually (by individuals or groups), and/or automatically using software tools, such as those developed for information retrieval and data mining. Recently the National Science Foundation funded a large collaborative development project for the semi-automated construction of an ontology of amphibian anatomy (AmphibAnat [2]). To satisfy the extensive community curation requirements of that project, a generic, Web-based, multi-user, relational database ontology management system (RDBOM [3]) was constructed, based upon a novel theoretical ontology model called an Ontology Abstract Machine (OAM [4]). The need to support concurrent data entry by multiple users with different levels of access privileges (as determined and assigned by the administrators) made it critical to ensure that the entered data were semantically correct. In particular, the ability to define and enforce restrictions on relations would help to identify inconsistencies in the ontology, maintain a higher level of overall integrity, and avoid erroneous conclusions that could be made by automated reasoners. In this paper we present a modified OAM model that accommodates one type of data restriction, domain and range, and facilitates associated validation. As proof of concept, we also describe how this modified abstract model has been implemented in RDBOM. Patrick G. Edgett, Leong Lee, Jennifer L. Leopold, Alton B. Coalter |
IDEAS | 3 |
| 2010 | Efficient Reasoning with RCC-3D
Julia Albath, Jennifer L. Leopold, Chaman L. Sabharwal, Kenneth Perry |
KSEM | 2 |
| 2010 | Automated Ontology Generation Using Spatial Reasoning
Alton B. Coalter, Jennifer L. Leopold |
KSEM | 2 |
| 2009 | Protein secondary structure prediction using rule induction from coveringsabstractWith the increase of data from genome sequencing projects comes the need for reliable and efficient methods for the analysis and classification of protein motifs and domains. Experimental methods currently used to determine protein structure are accurate, yet expensive both in terms of time and equipment. Therefore, various computational approaches to solving the problem have been attempted, although their accuracy has rarely exceeded 75%. In this paper, a rule-based method to predict protein secondary structure is presented. This method uses a newly developed data-mining algorithm called RT-RICO (Relaxed Threshold Rule Induction from Coverings), which identifies dependencies between amino acids in a protein sequence, and generates rules that can be used to predict secondary structures. The average prediction accuracy on sample data sets, or Q3score, using RT-RICO was 80.3%, an improvement over comparable computational methods Leong Lee, Jennifer L. Leopold, Ronald L. Frank, Anne M. Maglia |
CIBCB | 2 |
| 2007 | Determining Domain Similarity and Domain-Protein Similarity Using Functional Similarity Measurements of Gene Ontology TermsabstractProtein domains typically correspond to major functional sites of a protein. Therefore, determining similarity between domains can aid in the comparison of protein functions, and can provide a basis for grouping domains based on function. One strategy for comparing domain similarity and domain-protein similarity is to use similarity measurements of annotation terms from the Gene Ontology (GO). In this paper five methods are analyzed in terms of their usefulness for comparing domains, and comparing domains to proteins based on GO terms. Lisa M. Guntly, Jennifer L. Leopold, Anne M. Maglia |
BIBE | 2 |
| 2007 | Alternative Splicing: Associating Frequency with IsoformsabstractIn the simplest model of protein production, a gene gives rise to a single protein; DNA is transcribed to form pre-mRNA, which is converted to mRNA by splicing or removing introns. The result is a chain of exons that is translated to form a protein. Alternative splicing of exons may result in the formation of multiple proteins from the same gene sequence. However, not all of these proteins may be functional. Thus, we ask whether we can predict and rank (in order of frequency of occurrence and functional importance) the set of possible proteins for a gene. Herein we describe a tool that predicts the relative frequencies of isoforms that can be produced from a given gene. Anuradha Roy, Jennifer L. Leopold, Anne M. Maglia |
BIBE | 2 |
| 2004 | Identifying Character Non-Independence in Phylogenetic Data Using Data Mining Techniques
Anne M. Maglia, Jennifer L. Leopold, Venkat Ram Ghatti |
APBC | 2 |
| 2002 | Webformulate: a web-based visual continual query systemabstractToday there is a plethora of data accessible via the Internet. The Web has greatly simplified the process of searching for, accessing, and sharing information. However, a considerable amount of Internet-distributed data still goes unnoticed and unutilized, particularly in the case of frequently-updated, Internet-distributed databases. In this paper we give an overview of WebFormulate, a Web-based visual continual query system that addresses the problems associated with formulating temporal ad hoc analyses over networks of heterogeneous, frequently-updated data sources. The main distinction between this system and existing Internet facilities to retrieve information and assimilate it into computations is that WebFormulate provides the necessary facilities to perform continual queries, developing and maintaining dynamic links such that Web-based computations and reports automatically maintain themselves. A further distinction is that this system is specifically designed for users of spreadsheet-level ability, rather than professional programmers. Jennifer L. Leopold, Meg Heimovics, Tyler Palmer |
WWW | 1 |