VLDB 2026 Research / reviewers in the wild / expert
Dennis P. Groth
dblp:77/751
· DBLP profile ↗
19ranked-venue papers
12as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 16 · 11 first-authorGraphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-authorApplied, interdisciplinary, general and emerging computing · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Visualization and visual analytics › information visualization › metadata visualization
provenance visualization |
0.1 | 1 | 2006 | Provenance and Annotation for Visual Exploration Systems · IEEE Trans. Vis. Comput. Graph. 2006 |
Data mining › pattern mining
apriori algorithm |
0.0 | 1 | 2004 | Average-Case Performance of the Apriori Algorithm · SIAM J. Comput. 2004 |
Data mining › pattern mining › itemset mining
frequent itemset mining |
0.0 | 1 | 2004 | Average-Case Performance of the Apriori Algorithm · SIAM J. Comput. 2004 |
Data mining
pattern mining |
0.0 | 1 | 2004 | Average-Case Performance of the Apriori Algorithm · SIAM J. Comput. 2004 |
Algorithms and data structures › analysis of algorithms
average-case analysis |
0.0 | 1 | 2004 | Average-Case Performance of the Apriori Algorithm · SIAM J. Comput. 2004 |
Methods — techniques the papers use, named apart from their topics
random shopper model · 0.1probabilistic analysis · 0.1spatio-temporal visualization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | An application of participatory action research in advising-focused learning analyticsabstractAdvisors assist students in developing successful course pathways through the curriculum. The purpose of this project is to augment advisor institutional and tacit knowledge with knowledge from predictive algorithms (i.e., Matrix Factorization and Classifiers) specifically developed to identify risk. We use a participatory action research approach that directly involves key members from both advising and research communities in the assessment and provisioning of information from the predictive analytics. The knowledge gained from predictive algorithms is evaluated using a mixed method approach. We first compare the predictive evaluations with advisors evaluations of student performance in courses and actual outcomes in those courses We next expose and classify advisor knowledge of student risk and identify ways to enhance the value of the prediction model. The results highlight the contribution that this collaborative approach can give to the constructive integration of Learning Analytics in higher education settings. Stefano Fiorini, Adrienne Sewell, Mathew Bumbalough, Pallavi Chauhan, Linda Shepard, George Rehrey, Dennis P. Groth |
LAK | 7 |
| 2017 | Developing institutional learning analytics 'communities of transformation' to support student successabstractInstitutional implementation of learning analytics calls for thoughtful management of cultural change. This interactive halfday workshop responds to the LA literature describing the benefits and challenges of institutional LA implementation by offering participants an opportunity to learn about and begin planning for a program to actively engage faculty as leaders of data exploration around the theme of 'student success'. This session will share experiences from five institutions actively engaged in fostering Learning Analytics Communities (LAC) by identifying key issues, sharing lessons learned, and considering structural frameworks that are transferable to other institutional contexts. Structured discussion and activities will engage participants in developing an action plan for establishing an LAC on their own campus. Leah Macfadyen, Dennis P. Groth, George Rehrey, Linda Shepard, Jim E. Greer, Douglas Ward, Caroline Bennett, Jake Kaupp, Marco Molinaro 0002, Matthew Steinwachs |
LAK | 2 |
| 2014 | Students in Transit: Understanding the Migration of Students between Disciplines and Disciplinary GroupingsabstractIn this paper we investigate methods for visualizing the flow of university-level students through the undergraduate academic experience. We use the lens of transit between major declarations, both in the context of specific major to major mappings, as well as mappings between disciplinary clusters. Our motivation is twofold. First, we wish to understand the implications for students, in terms of time to degree, fidelity to initial major interest, and intellectual distance travelled by students as they cross boundaries of major clusters. Second, we wish to understand the implications for degree programs, as they seek to attract and retain top students, or manage the consequences of declining enrollments. Our approach utilizes multiple visualization techniques, as we first try and develop high-level understanding of the migratory patterns of students, and then drill down to more fine-grained approaches yielding insights that are aimed at providing specific, actionable information to a diverse set of stakeholders, including students, faculty, and administrators. Dennis P. Groth, Spencer Hayden, Michael J. Sauer |
IV | 1 |
| 2008 | Improving computer science diversity through summer campsabstractSummer camps offer a ripe opportunity for increasing computer science diversity. This panel provides several examples of summer camps that specifically recruit from traditionally underrepresented demographics. The panelists run camps at a community college, a private liberal-arts college, and public universities. The camps are residential and day camps, coed and all-female camps, ranging from three-days to two-weeks long, with campers from 10-years-olds to high school seniors. In addition to describing their camps, the panelists will provide information on securing funding, recruiting campers from underrepresented populations, measuring impact, and lessons learned along the way. Demonstrations of what campers accomplished will also be shown. Dennis P. Groth, Helen H. Hu, Betty Lauer, Hwajung Lee |
SIGCSE | 1 |
| 2007 | How Students Perceive Risk: A Study of Senior Capstone Project TeamsabstractOne view of software engineering project management is that it is the portion of a project that is involved with managing tasks, resources, assignments and deadlines. Certainly, many of the daily activities for a software project manager might be categorized in such a way; however, these activities merely serve as surrogates for the principal responsibility of the project manager: risk management. Risk management includes identification, analysis, and mitigation of risks that, should the risks materialize, have the potential for negatively impacting the completion of a project. In this paper we approach the problem of risk management for IT projects from the perspective of the developers, and pursue an understanding of how student developers perceive project risks using Tiwana and Keil's "one-minute risk assessment tool". We present the results of a study of perceived project risk for different IT projects completed by more than 50 senior undergraduate students during the 2005/2006 academic year. Dennis P. Groth, Matthew P. Hottell |
CSEE&T | 1 |
| 2007 | Tracking and Organizing Visual Exploration Activities across Systems and ToolsabstractModern knowledge discovery activities occur in highly dynamic environments. Specific activities may involve multiple tools, techniques, systems, individuals, and locations. In addition to these complexities, the span of time involved with discovery may vary from short to long, as well as being contiguous or disjoint. This paper presents a framework for tracking the history, or provenance, of the discovery process across applications, systems, and users. The resulting capabilities provide fine-grained provenance information relative to the discovered information. Along with the provenance framework, a prototype system is used to demonstrate the main concepts of the proposed approach. Dennis P. Groth |
IV | 1 |
| 2006 | Designing and Developing an Informatics Capstone Project CourseabstractInformatics is the study of the application of information technology. The focus of studies within an Informatics program is on particular problem domain areas, including scientific domains like biology or chemistry, but also non-scientific domains like music, fine arts, or business. In this paper we describe our experience with the development of a capstone course for undergraduate Informatics students. In particular, we focus on process-related activities we have undertaken to deal with the rapid increase in enrollment for our course, which has grown from 15 students during the 2001/2002 academic year to more than 100 students currently enrolled. During the first 4 years that we have offered the course, over 300 students have completed more than 80 capstone group projects, with an expected 30 more projects this year. We present profiles of our past projects to illustrate the diversity and challenges to be expected of an Informatics program, which is significantly different than a traditional Computer Science program Dennis P. Groth, Matthew P. Hottell |
CSEE&T | 1 |
| 2006 | Visualizing Distributions and Classification AccuracyabstractData mining is the search for novel, actionable information within data. It is important to note that the number of records in the data being analyzed is only one (and perhaps a small) factor in determining the complexity of a given data mining technique. Most complexity in data mining arises from the distribution of values contained in the data - not the number of records. In this paper, we utilize straightforward histogram-based visualizations to gain insight into how the performance of a well-studied data mining technique, the naive-Bayes classifier, performs under various discretization schemes for both continuous and discrete values. The resulting visualization system provides users with a tool that describes the underlying model of the data used by the classifier. Exploratory visualizations of the distributions of training data can be selected based on expert domain knowledge and then combined to apply to the test data Dennis P. Groth |
IV | 1 |
| 2006 | USEable security: interface design strategies for improving securityabstractAs people start depending more on technology and the internet they are opening themselves up to new risks. In this project, we specifically investigated wireless router interfaces to understand the needs of users when they configure security. Two studies were conducted: a baseline study comparing the interfaces of two routers on the market and a study comparing a prototype and the Linksys interface. The baseline study showed that there was no difference between the current interfaces. We then conducted a controlled experiment with a prototype that gave visual feedback. The prototype showed significant improvement in level of security achieved. Amanda L. Stephano, Dennis P. Groth |
VizSEC | 2 |
| 2006 | Provenance and Annotation for Visual Exploration SystemsabstractExploring data using visualization systems has been shown to be an extremely powerful technique. However, one of the challenges with such systems is an inability to completely support the knowledge discovery process. More than simply looking at data, users will make a semipermanent record of their visualizations by printing out a hard copy. Subsequently, users will mark and annotate these static representations, either for dissemination purposes or to augment their personal memory of what was witnessed. In this paper, we present a model for recording the history of user explorations in visualization environments, augmented with the capability for users to annotate their explorations. A prototype system is used to demonstrate how this provenance information can be recalled and shared. The prototype system generates interactive visualizations of the provenance data using a spatio-temporal technique. Beyond the technical details of our model and prototype, results from a controlled experiment that explores how different history mechanisms impact problem solving in visualization environments are presented. Dennis P. Groth, Kristy Streefkerk |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2005 | Enhancing 3D File Search with Landscapes and Personal Histories: Exploring the Possibilities of TerraSearch+abstractMost users today have difficulty locating files due to memory, i.e., an ability to recreate the personal history of where they placed the document in the first place. Microsoft XP search provides the standard file search functions according to human semantic memory. However, depending exclusively on this human system for file acquisition can be problematic and very limiting for users who perform well in 3D space. The prototype system utilizes three-dimensional environments that are augmented with second-order information that depicts the relationships between the first-order properties by a visualization system. This system maintains user trace data in a way that enhances their ability to locate files according to place, time, and content. Early results of cognitive walkthroughs and questionnaires indicate users could more easily find files through search tools, but that unfamiliarity with the interface could be a barrier to acceptance. Anthony Faiola, Dennis P. Groth |
IV | 2 |
| 2004 | A collaborative annotation system for data visualizationabstractWe present Collaborative Annotations on Visualizations (CAV), a system for annotating visual data in remote and collocated environments. Our system consists of a network framework, and a client application built for tablet PC's. CAV is designed to support the collection and sharing of annotations, through the use of mobile devices connected to visualization servers. We have developed a working system prototype based on tablet PC's that supports digital ink, voice and text annotation, and illustrates our approach in a variety of application domains, including biology, chemistry, and telemedicine. We have created an XML based open standard that supports access to a variety of client devices by publishing visualizations (data and annotations) as streams of images. CAV's primary goal is to enhance scientific discovery by supporting collaboration in the context of data visualizations. Sean E. Ellis, Dennis P. Groth |
AVI | 2 |
| 2004 | Information Provenance and the Knowledge Rediscovery ProblemabstractVisualizations leverage innate human capabilities for recognizing interesting aspects of data. Even if users might agree on what is interesting about a visualization, the steps that they use in the knowledge discovery process may be significantly different. This results in an inability to effectively recreate the exact conditions of the discovery process, which we call the knowledge rediscovery problem. Because we cannot expect a user to fully document each of their interactions, there is a need for visualization systems to maintain user trace data in a way that enhances a user's ability to communicate what they found to be interesting, as well as how they found it. We present a model for representing user interactions that articulates with a corresponding set of annotations, or observations that are made during the exploration. Such ability is critical to addressing the knowledge rediscovery problem, and is a fundamental component for systems that must provide information provenance. Dennis P. Groth |
IV | 1 |
| 2004 | Average-Case Performance of the Apriori AlgorithmabstractThe failure rate of the Apriori Algorithm is studied analytically for the case of random shoppers. The time needed by the Apriori Algorithm is determined by the number of item sets that are output (successes: item sets that occur in at least k baskets) and the number of item sets that are counted but not output (failures: item sets where all subsets of the item set occur in at least k baskets but the full set occurs in less than k baskets). The number of successes is a property of the data; no algorithm that is required to output each success can avoid doing work associated with the successes. The number of failures is a property of both the algorithm and the data. We find that under a wide range of conditions the performance of the Apriori Algorithm is almost as bad as is permitted under sophisticated worst-case analyses. In particular, there is usually a bad level with two properties: (1) it is the level where nearly all of the work is done, and (2) nearly all item sets counted are failures. Let l be the level with the most successes, and let the number of successes on level l be approximately ${m\choose l}$ for some m. Then, typically, the Apriori Algorithm has total output proportional to approximately ${m\choose l}$ and total work proportional to approximately ${m\choose l+1}$. In addition m is usually much larger than l, so the ratio of work to output is proportional to approximately $m/(l+1)$. The analytical results for random shoppers are compared against measurements for three data sets. These data sets are more like the usual applications of the algorithm. In particular, the buying patterns of the various shoppers are highly correlated. For most thresholds, these data sets also have a bad level. Thus, under most conditions nearly all of the work done by the Apriori Algorithm consists in counting item sets that fail. Paul W. Purdom, Dirk Van Gucht, Dennis P. Groth |
SIAM J. Comput. | 3 |
| 2003 | Visual Representation of Database Queries using Structural SimilarityabstractIt is often useful to get high-level views of datasets in order to identify areas of interest worthy of further exploration. In relational databases, the high-level view can be described using entity-relationship diagrams, which identify relationships between entities in the data model. Such high-level views are useful for database design activities, and can be used to generate user interfaces for constructing queries. This research introduces techniques for visualizing structural similarity of database queries. We demonstrate that individual queries can be visualized using graph visualization techniques. A distance measure based on query structure is proposed that provides database designers and administrators with a high-level perspective of relationships in the underlying data. Dennis P. Groth |
IV | 1 |
| 2003 | Development of an Informatics Tool for Crystallography Laboratory AdministratorsabstractWith increased demand for storage of scientific data comes a corresponding demand for efficient retrieval mechanisms necessary for analytical and reporting purposes. As is often the case, the design of databases to support scientific data is (rightly so) driven by the science. Unfortunately, this "science-centric" view of data management does not make the subsequent reporting of information easy. The natural solution to such a problem is, of course, a data warehouse approach to scientific data management. The focus of this project is the design of a data warehousing architecture for scientific data. Our goal is to separate the science-specific aspects of the information from the reporting and analysis requirements. The problem domain we are currently investigating is crystallographic structure data, coupled with metadata concerning the structure. The paper describes ITCLA - an Informatics Tool for Crystallography Laboratory Administrators. We report on the current state of ITCLA, as well as our future plans. Leah Sandvoss, Dennis P. Groth |
SSDBM | 2 |
| 2002 | An integrated approach to database visualizationabstractWe present an architecture that enables information visualization activities within a database environment. Our approach presents an abstraction of this transformation process, which we call mapping. The implementation of the mapping process is controlled by the end-user through a Map, which can be used to add order and scale to data. Dennis P. Groth, Edward L. Robertson |
AVI | 1 |
| 2002 | An Integrated System for Database VisualizationabstractThis paper present details of an integrated database visualization system. The system supports the visualization process from an end-to-end perspective. Included in the system is a mechanism for performing transformations to the data being visualized through the use of database relations. This mapping process provides an abstract mechanism for supporting data to geometry transformations under the control of a user-defined, declarative language. The system supports a wide variety of visualization techniques, including scatterplots, bar charts and surface plots. Dennis P. Groth, Edward L. Robertson |
IV | 1 |
| 2001 | It's All about Process: Project Oriented Teaching of Software EngineeringabstractProcess considerations are a central part of the material for a software engineering course; they are also central to accomplishing full-lifecycle, team-based systems development projects in such a course. This paper discusses the ways in which we have achieved an effective process structure within an academic context of full-year project courses. The key features are a kernel project plan and a process management mechanism. The project plan is a schedule including eight milestones with fixed due dates and quite explicit deliverables. The management is accomplished through an advanced full-year course, whose participants guide the project teams through the process. Dennis P. Groth, Edward L. Robertson |
CSEE&T | 1 |