VLDB 2026 Research / reviewers in the wild / expert
Cliff A. Joslyn
dblp:53/5019
· DBLP profile ↗
17ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-5923-5547ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Theory of computation · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Talking to GDELT Through Knowledge GraphsabstractIn this work we study various Retrieval Augmented Regeneration (RAG) approaches to gain an understanding of the strengths and weaknesses of each approach in a question-answering analysis. To gain this understanding we use a case-study subset of the Global Database of Events, Language, and Tone (GDELT) dataset as well as a corpus of raw text scraped from the online news articles. To retrieve information from the text corpus we implement a traditional vector store RAG as well as state-of-the-art large language model (LLM) based approaches for automatically constructing KGs and retrieving the relevant subgraphs. In addition to these corpus approaches, we develop a novel ontology-based framework for constructing knowledge graphs (KGs) from GDELT directly which leverages the underlying schema of GDELT to create structured representations of global events. For retrieving relevant information from the ontology-based KGs we implement both direct graph queries and state-of-the-art graph retrieval approaches. We compare the performance of each method in a question-answering task. We find that while our ontology-based KGs are valuable for question-answering, automated extraction of the relevant subgraphs is challenging. Conversely, LLM-generated KGs, while capturing event summaries, often lack consistency and interpretability. Our findings suggest benefits of a synergistic approach between ontology and LLM-based KG construction, with proposed avenues toward that end. Audun Myers, Max Vargas, Sinan G. Aksoy, Cliff A. Joslyn, Lee Burke, Tom Grimes |
NeSy | 4 |
| 2023 | Experimental Observations of the Topology of Convolutional Neural Network ActivationsabstractTopological data analysis (TDA) is a branch of computational mathematics, bridging algebraic topology and data science, that provides compact, noise-robust representations of complex structures. Deep neural networks (DNNs) learn millions of parameters associated with a series of transformations defined by the model architecture resulting in high-dimensional, difficult to interpret internal representations of input data. As DNNs become more ubiquitous across multiple sectors of our society, there is increasing recognition that mathematical methods are needed to aid analysts, researchers, and practitioners in understanding and interpreting how these models' internal representations relate to the final classification. In this paper we apply cutting edge techniques from TDA with the goal of gaining insight towards interpretability of convolutional neural networks used for image classification. We use two common TDA approaches to explore several methods for modeling hidden layer activations as high-dimensional point clouds, and provide experimental evidence that these point clouds capture valuable structural information about the model's process. First, we demonstrate that a distance metric based on persistent homology can be used to quantify meaningful differences between layers and discuss these distances in the broader context of existing representational similarity metrics for neural network interpretability. Second, we show that a mapper graph can provide semantic insight as to how these models organize hierarchical class knowledge at each layer. These observations demonstrate that TDA is a useful tool to help deep learning practitioners unlock the hidden structures of their models. Emilie Purvine, Davis Brown, Brett A. Jefferson, Cliff A. Joslyn, Brenda Praggastis, Archit Rathore, Madelyn Shapiro, Bei Wang 0001, Youjia Zhou |
AAAI | 4 |
| 2023 | Topological Analysis of Temporal Hypergraphs
Audun Myers, Cliff A. Joslyn, Bill Kay, Emilie Purvine, Gregory Roek, Madelyn Shapiro |
WAW | 2 |
| 2022 | High-order Line Graphs of Non-uniform Hypergraphs: Algorithms, Applications, and Experimental AnalysisabstractHypergraphs offer flexible and robust data representations for many applications, but methods that work directly on hypergraphs are not readily available and tend to be prohibitively expensive. Much of the current analysis of hypergraphs relies on first performing a graph expansion – either based on the nodes (clique expansion), or on the hyperedges (line graph) − and then running standard graph analytics on the resulting representative graph. However, this approach suffers from massive space complexity and high computational cost with increasing hypergraph size. Here, we present efficient, parallel algorithms to accelerate and reduce the memory footprint of higher-order graph expansions of hypergraphs. Our results focus on the hyperedge-based s-line graph expansion, but the methods we develop work for higher-order clique expansions as well. To the best of our knowledge, ours is the first framework to enable hypergraph spectral analysis of a large dataset on a single shared-memory machine. Our methods enable the analysis of datasets from many domains that previous graph-expansion-based models are unable to provide. The proposed s-line graph computation algorithms are orders of magnitude faster than state-of-the-art sparse general matrix-matrix multiplication methods, and obtain approximately 2–31× speedup over a prior state-of-the-art heuristic-based algorithm for$s$-line graph computation. Xu T. Liu, Jesun Sahariar Firoz, Sinan G. Aksoy, Ilya Amburg, Andrew Lumsdaine, Cliff A. Joslyn, Brenda Praggastis, Assefaw Hadish Gebremedhin |
IPDPS | 6 |
| 2021 | Parallel Algorithms for Efficient Computation of High-Order Line Graphs of HypergraphsabstractThis paper considers structures of systems beyond dyadic (pairwise) interactions and investigates mathematical modeling of multi-way interactions and connections as hyper-graphs, where captured relationships among system entities are set-valued. To date, in most situations, entities in a hypergraph are considered connected if there is at least one common “neighbor”. However, minimal commonality sometimes discards the “strength” of connections and interactions among groups. To this end, considering the “width” of a connection, referred to as the s-overlap of neighbors, provides more meaningful insights into how closely the communities or entities interact with each other. In addition, s-overlap computation is the fundamental kernel to construct the line graph of a hypergraph, a low-order approximation of the hypergraph which can carry significant information about the original hypergraph. Subsequent stages of a data analytics pipeline then can apply highly tuned graph algorithms on the line graph to reveal important features. Given a hypergraph, computing the s-overlaps by exhaustively considering all pairwise entities can be computationally prohibitive. To tackle this challenge, we develop efficient algorithms to compute s-overlaps and the corresponding line graph of a hypergraph. We propose several heuristics to avoid execution of redundant work and improve performance of the s-overlap computation. Our parallel algorithm, combined with these heuristics, is orders of magnitude (more than 10x) faster than the naive algorithm in all cases and the SpGEMM algorithm with filtration in most cases (especially with large$s$value). Xu T. Liu, Jesun Sahariar Firoz, Andrew Lumsdaine, Cliff A. Joslyn, Sinan G. Aksoy, Brenda Praggastis, Assefaw Hadish Gebremedhin |
HiPC | 4 |
| 2021 | Hypergraph models of biological networks to identify genes critical to pathogenic viral responseabstractBACKGROUND: Representing biological networks as graphs is a powerful approach to reveal underlying patterns, signatures, and critical components from high-throughput biomolecular data. However, graphs do not natively capture the multi-way relationships present among genes and proteins in biological systems. Hypergraphs are generalizations of graphs that naturally model multi-way relationships and have shown promise in modeling systems such as protein complexes and metabolic reactions. In this paper we seek to understand how hypergraphs can more faithfully identify, and potentially predict, important genes based on complex relationships inferred from genomic expression data sets. RESULTS: We compiled a novel data set of transcriptional host response to pathogenic viral infections and formulated relationships between genes as a hypergraph where hyperedges represent significantly perturbed genes, and vertices represent individual biological samples with specific experimental conditions. We find that hypergraph betweenness centrality is a superior method for identification of genes important to viral response when compared with graph centrality. CONCLUSIONS: Our results demonstrate the utility of using hypergraphs to represent complex biological systems and highlight central important responses in common to a variety of highly pathogenic viruses. Emily Heath, Brett A. Jefferson, Cliff A. Joslyn, Henry Kvinge, Hugh D. Mitchell, Brenda Praggastis, Amie J. Eisfeld, Amy C. Sims, Larissa B. Thackray, Shufang Fan, Kevin B. Walters, Peter J. Halfmann, Danielle Westhoff-Smith, Qing Tan, Vineet D. Menachery, Timothy P. Sheahan, Adam S. Cockrell, Jacob F. Kocher, Kelly G. Stratton, Natalie C. Heller, Lisa M. Bramer, Michael S. Diamond, Ralph S. Baric, Katrina M. Waters, Yoshihiro Kawaoka, Jason E. McDermott, Emilie Purvine |
BMC Bioinform. | 4 |
| 2020 | Hypergraph Analytics of Domain Name System Relationships
Cliff A. Joslyn, Sinan G. Aksoy, Dustin Arendt, Jesun Sahariar Firoz, Louis Jenkins, Brenda Praggastis, Emilie Purvine, Marcin Zalewski |
WAW | 1 |
| 2017 | When Labels Fall Short: Property Graph Simulation via Blending of Network Structure and Vertex AttributesabstractProperty graphs can be used to represent heterogeneous networks with labeled (attributed) vertices and edges. Given a property graph, simulating another graph with same or greater size with the same statistical properties with respect to the labels and connectivity is critical for privacy preservation and benchmarking purposes. In this work we tackle the problem of capturing the statistical dependence of the edge connectivity on the vertex labels and using the same distribution to regenerate property graphs of the same or expanded size in a scalable manner. However, accurate simulation becomes a challenge when the attributes do not completely explain the network structure. We propose the Property Graph Model (PGM) approach that uses a label augmentation strategy to mitigate the problem and preserve the vertex label and the edge connectivity distributions as well as their correlation, while also replicating the degree distribution. Our proposed algorithm is scalable with a linear complexity in the number of edges in the target graph. We illustrate the efficacy of the PGM approach in regenerating and expanding the datasets by leveraging two distinct illustrations. Our open-source implementation is available on GitHub. Arun V. Sathanur, Sutanay Choudhury, Cliff A. Joslyn, Sumit Purohit |
CIKM | 3 |
| 2014 | Optimizing graph queries with graph joins and Sprinkle SPARQLabstractBig data problems are often more akin to sparse graphs rather than relational tables. As such we argue that graph-based physical representations provide advantages in terms of both size and speed for executing queries. Drawing from research in sparse matrices, we use a compressed sparse row (CSR) format to model graph-oriented data. We also present two novel mechanisms for exploiting the CSR format that both find optimal join strategies and also prune variable bindings before expensive join operations occur. The first tactic we call Sprinkle SPARQL, which takes triple patterns of SPARQL queries and performs low-cost, linear-time set intersections to produce a constrained list of variable bindings for each variable in a query. Besides constrained lists of variable bindings, Sprinkle SPARQL also produces metrics that are consumed by the join algorithm to select an optimal execution path. The second tactic, graph joins, utilizes the CSR data structure as an index to efficiently join two variables expressed in a triple pattern together. We evaluate our approach on two data sets with over a billion edges: LUBM(8000) and an R-MAT graph generated with Graph5001parameters and extended to have edge labels. Eric L. Goodman, Edward Jimenez, Cliff A. Joslyn, David J. Haglin, Sinan Al-Saffar, Dirk Grunwald |
IEEE BigData | 3 |
| 2013 | Research towards a systematic signature discovery processabstractIn its most general form, a signature is a unique or distinguishing measurement, pattern, or collection of data that identifies a phenomenon (object, action, or behavior) of interest. The discovery of signatures is an important aspect of a wide range of disciplines from basic science to national security for the rapid and efficient detection and/or prediction of phenomena. Current practice in signature discovery is typically accomplished by asking domain experts to characterize and/or model individual phenomena to identify what might compose a useful signature. What is lacking is an approach that can be applied across a broad spectrum of domains and information sources to efficiently and robustly construct candidate signatures, validate their reliability, measure their quality, and overcome the challenge of detection - all in the face of dynamic conditions, measurement obfuscation, and noisy data environments. Our research has focused on the identification of common elements of signature discovery across application domains and the synthesis of those elements into a systematic process for more robust and efficient signature development. In this way, a systematic signature discovery process lays the groundwork for leveraging knowledge obtained from signatures to a particular domain or problem area, and, more generally, to problems outside that domain. This paper presents the initial results of this research by discussing a mathematical framework for representing signatures and placing that framework in the context of a systematic signature discovery process. Additionally, the basic steps of this process are described with details about the methods available to support the different stages of signature discovery, development, and deployment. Nathan A. Baker, Jonathan L. Barr, George T. Bonheyo, Cliff A. Joslyn, Kannan Krishnaswami, Mark E. Oxley, Rich Quadrel, Landon H. Sego, Mark F. Tardiff, Adam S. Wynne |
ISI | 4 |
| 2013 | Orders on intervals over partially ordered sets: extending Allen's algebra and interval graph results
Francisco Zapata, Vladik Kreinovich, Cliff A. Joslyn, Emilie Hogan |
Soft Comput. | 3 |
| 2012 | An Analysis of Multi-type Relational Interactions in FMA Using Graph Motifs with Disjointness Constraints
Guo-Qiang Zhang 0001, Lingyun Luo, Chimezie Ogbuji, Cliff A. Joslyn, José L. V. Mejino Jr., Satya Sanket Sahoo |
AMIA | 4 |
| 2011 | Structure Discovery in Large Semantic Graphs Using Extant Ontological Scaling and Descriptive SemanticsabstractAs semantic datasets grow to be very large and divergent, there is a need to identify and exploit their inherent semantic structure for discovery and optimization. Towards that end, we present here a novel methodology to identify the semantic structures inherent in an arbitrary semantic graph dataset. We first present the concept of an extant ontology as a statistical description of the semantic relations present amongst the typed entities modeled in the graph. This serves as a model of the underlying semantic structure to aid in discovery and visualization. We then describe a method of ontological scaling in which the ontology is employed as a hierarchical scaling filter to infer different resolution levels at which the graph structures are to be viewed or analyzed. We illustrate these methods on three large and publicly available semantic datasets containing more than one billion edges each. Sinan Al-Saffar, Cliff A. Joslyn, Alan R. Chappell |
Web Intelligence | 2 |
| 2009 | View Discovery in OLAP Databases through Statistical Combinatorial Optimization
Cliff A. Joslyn, John Burke, Terence Critchlow, Nicolas W. Hengartner, Emilie Hogan |
SSDBM | 1 |
| 2005 | Protein annotation as term categorization in the gene ontology using word proximity networksabstractBACKGROUND: We participated in the BioCreAtIvE Task 2, which addressed the annotation of proteins into the Gene Ontology (GO) based on the text of a given document and the selection of evidence text from the document justifying that annotation. We approached the task utilizing several combinations of two distinct methods: an unsupervised algorithm for expanding words associated with GO nodes, and an annotation methodology which treats annotation as categorization of terms from a protein's document neighborhood into the GO. RESULTS: The evaluation results indicate that the method for expanding words associated with GO nodes is quite powerful; we were able to successfully select appropriate evidence text for a given annotation in 38% of Task 2.1 queries by building on this method. The term categorization methodology achieved a precision of 16% for annotation within the correct extended family in Task 2.2, though we show through subsequent analysis that this can be improved with a different parameter setting. Our architecture proved not to be very successful on the evidence text component of the task, in the configuration used to generate the submitted results. CONCLUSION: The initial results show promise for both of the methods we explored, and we are planning to integrate the methods more closely to achieve better results overall. Karin Verspoor, Judith D. Cohn, Cliff A. Joslyn, Susan M. Mniszewski, Andreas Rechtsteiner, Luis M. Rocha, Tiago Simas |
BMC Bioinform. | 3 |
| 1998 | Towards a Formal Taxonomy of Hybrid Uncertainty Representations
Cliff A. Joslyn, Luis M. Rocha |
Inf. Sci. | 1 |
| 1996 | Aggregation and Completion of Random Sets with Distributional fuzzy MeasuresabstractThe two known information theories, probability and possibility theory, are based on t-conorm decomposable fuzzy measures, so that bijective mappings exist between their set-valued measures and their point-valued distributions. Further, their random set (Dempster-Shafer evidence theoretical) interpretations have simple topological structures, with bijective mappings between the subset focal elements and the point singletons. We introduce the concepts of distributional and aggregable random sets and random set completion, and first use them as a model in which to cast probability and possibility measures and distributions. Then, towards the goal of deriving new forms of information theory, general Sugeno conorm decomposable fuzzy measures and ring-like aggregable random sets with set-intersection structural aggregation are examined, but it is shown that in these two cases no new information theories are forthcoming. Cliff A. Joslyn |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |