VLDB 2026 Research / reviewers in the wild / expert
Suzanne J. Matthews
dblp:159/0150
· DBLP profile ↗
29ranked-venue papers
12as first author
8since 2021 · last 2025
0000-0001-9170-2240ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 20 · 6 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 first-authorSystems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | U.S. Government-Funded Opportunities for CS Educators
Joel Adams 0001, Cynthia Bailey, Suzanne J. Matthews, Paul T. Tymann |
SIGCSE (2) | 3 |
| 2025 | Fostering Creativity: Student-Generative AI Teaming in an Open-Ended CS0 AssignmentabstractThe increasing ubiquity of web-based generative artificial intelligence technologies necessitates that all students experience teaming with such technologies -- exploring their strengths and limitations and learning how to create synergy with them. To aid in this effort, we designed an open-ended generative AI project for the freshmen taking our general-education introduction to computing course. Students were required to team with generative AI to create something beyond what they alone (or the AI alone) could accomplish. Upon completion, students submitted a short written critical analysis documenting their experiences and presented a three-minute demonstration of their project in class. Despite limited course coverage of AI and generative AI prior to this project, we were impressed by the creativity and sophistication of the submitted final products as well as the breadth of generative AI tools explored. Student reflections on the experience illustrated numerous insights into the strengths and limitations of the tools they employed. Our results underscore that students can learn about the benefits and limitations of generative AI in as little as a single assignment and that covering such topics need not require extensive amounts of course time and resources. Daniel Filcik, Edward Sobiesk, Suzanne J. Matthews |
SIGCSE (1) | 3 |
| 2025 | ASM Visualizer: A Learning Tool for Assembly ProgrammingabstractWe present ASM Visualizer, a tool that is designed to help students learn assembly programming, aiding in their understanding of how assembly instructions are executed and the relationship between assembly and equivalent high-level language code.Our tool allows a user to step both forward and backward through the execution of an assembly program, one instruction at a time, seeing how instructions use and modify values in stack memory and CPU registers.ASM Visualizer presents three user-interface modes, supporting different stages of learning assembly programming.Beginners can step through basic arithmetic instructions, whereas more advanced learners can trace through function call/return sequences, stack frame manipulation, or entire assembly programs.We present our experiences using ASM Visualizer in introductorylevel courses at our two institutions, and we discuss other ways in which our tool could be used by educators in both introductory and advanced CS courses.Results from a preliminary assessment of students using our tool show that students gain confidence in their understanding of different aspects of assembly programming.We feel that the visual interface to assembly code execution that ASM Visualizer provides is key to helping students understand assembly. Tia Newhall, Kevin C. Webb 0001, Isabel Romea, Emma Stavis, Suzanne J. Matthews |
SIGCSE (1) | 5 |
| 2025 | Hands-on parallel & distributed computing with Raspberry Pi devices and clustersabstractParallel and distributed computing (PDC) concepts are now required topics for accredited undergraduate computer science programs. However, introducing PDC into the CS curriculum is challenging for several reasons, including an instructor's lack of PDC knowledge and difficulties in accessing PDC hardware. This paper addresses both of these challenges by presenting free, interactive, web-based PDC teaching modules using inexpensive Raspberry Pi single board computers (SBCs). Our materials include a free disk image that makes it possible for instructors to build Raspberry Pi clusters in minutes and use our software in a variety of curricular contexts. Our multi-year assessment of these materials on students and faculty members indicates that: (i) our materials increased students' confidence regarding important PDC concepts and motivated them to study PDC further; and (ii) our materials increased faculty members' confidence and preparedness in teaching key PDC concepts at their own institutions. • Free online interactive modules for learning PDC with Raspberry Pis and Pi clusters. • Self-organizing cluster: connects disparate Pis into a working cluster in minutes. • Free disk image pre-loaded with all activities for painless classroom adoption. • Our materials increase student confidence about PDC and motivation to learn more PDC. • Our materials increase faculty confidence and preparedness to teach PDC. Elizabeth Shoop, Suzanne J. Matthews, Richard A. Brown, Joel Adams 0001 |
J. Parallel Distributed Comput. | 2 |
| 2022 | Teaching Distributed Computing Fundamentals using Raspberry Pi ClustersabstractThe 2019 ABET computer science criteria requires that all computing students learn parallel and distributed computing (PDC) as undergraduates, and CS2013 recommends at least fifteen hours of PDC in the undergraduate curriculum. Consequently, many educators are looking for easy ways to integrate PDC into courses at their institutions. This hands-on workshop introduces Message Passing Interface (MPI) basics in Python or C/C++ using clusters of Raspberry Pi single-board computers. MPI is a multi-language, platform independent, industry-standard library for PDC. Raspberry Pis are an inexpensive and engaging hardware platform for studying PDC as early as CS1. Participants will experience how to teach distributed computing essentials with MPI by means of reusable, effective "parallel patterns," including single program multiple data (SPMD) execution, send-receive message passing, and parallel loop patterns. No prior experience with MPI, PDC, or the Raspberry Pi is expected; participants will explore short programs designed to help students understand MPI basics, plus longer "exemplar" programs that use MPI to solve significant applied problems. The workshop includes: (i) personal experience with the Raspberry Pi (clusters provided); (ii) instructions on how to deploy Raspberry Pi clusters quickly in the classroom; (iii) self-paced hands-on experimentation with MPI programs; and (iv) a discussion of how to use Raspberry Pi clusters to align courses with CS2013 and ABET. All materials from this workshop are available from CSinParallel.org; participants should bring a laptop to access materials and connect to the Raspberry Pi clusters. Elizabeth Shoop, Richard A. Brown, Joel Adams 0001, Suzanne J. Matthews |
SIGCSE (2) | 4 |
| 2021 | Teaching Parallel and Distributed Computing in the Time of COVIDabstractWith both CS2013 and the most recent ABET computing criteria requiring coverage of parallel and distributed computing (PDC), many CS faculty are looking for ways to incorporate PDC concepts into their curricula. However, the COVID-19 pandemic and the switch to remote teaching introduces new difficulties to teaching a subject that many already find challenging. This BOF provides a forum for computing educators to discuss strategies they have used to teach PDC remotely. The organizers will share techniques that they have found effective (including the different tradeoffs those techniques involve) and foster a discussion in which others can share novel ways of teaching PDC remotely. The end-goals are: (i) to provide a venue in which those who have taught PDC remotely can share their experiences, in the hopes of identifying best practices; and (ii) to enable participants who are new to PDC to learn from the experiences of faculty who have already taught such courses in this new environment. No laptop required; any materials resulting from this session will be distributed via CSinParallel.org. Joel Adams 0001, Richard A. Brown, Suzanne J. Matthews, Elizabeth Shoop |
SIGCSE | 3 |
| 2021 | TextbooksForAll: Free Textbooks and Their Place in Computer Science EducationabstractThe expense of textbooks is a common source of frustration among students. Furthermore, the lack of affordable textbooks can inadvertently limit who studies computer science. One way educators and faculty can mitigate these issues is through adopting (and writing) free online textbooks. This panel discusses the benefits, challenges, and development practices of free online textbooks. The panelists, who have authored three widely used textbooks, characterize the role of free textbooks in CS education, describe their experiences writing a free textbook, and offer advice to faculty interested in incorporating them into their courses. Suzanne J. Matthews, Chris Mayfield, Remzi H. Arpaci-Dusseau, Kevin C. Webb 0001 |
SIGCSE | 1 |
| 2021 | Dive into Systems: A Free, Online Textbook for Introducing Computer SystemsabstractThis paper presents our experiences, motivations, and goals for developing Dive into Systems [17], a new, free, online textbook that introduces computer systems, computer organization, and parallel computing. Our book's topic coverage is designed to give readers a gentle and broad introduction to these important topics. It teaches the fundamentals of computer systems and architecture, introduces skills for writing efficient programs, and provides necessary background to prepare students for advanced study in computer systems topics. Our book assumes only a CS1 background of the reader and is designed to be useful to a range of courses as a primary textbook for courses that introduce computer systems topics or as an auxiliary textbook to provide systems background in other courses. Results of an evaluation from students and faculty at 18 institutions who used a beta release of our book show overwhelmingly strong support for its coverage of computer systems topics, its readability, and its availability. Chapters are reviewed and edited by external volunteers from the CS education community. Their feedback, as well as that of student and faculty users, is continuously incorporated into its online content. We anticipate releasing version 1.0 of the book in spring of 2021, and a release candidate is currently available at https://diveintosystems.org. Suzanne J. Matthews, Tia Newhall, Kevin C. Webb 0001 |
SIGCSE | 1 |
| 2020 | Incorporating Parallel Computing in the Undergraduate Computer Science CurriculumabstractTeaching parallel and distributed computing (PDC) concepts is an ongoing and pressing concern for many undergraduate educators. The ACM/IEEE CS Joint Task Force on Computing Curricula (CS2013) recommends 15 hours of PDC education in the undergraduate curriculum. Most recently, the 2019 ABET Criteria for Accrediting Computer Science requires coverage of PDC topics. For faculty who are unfamiliar with PDC, the prospect of incorporating parallel computing into their courses can seem very daunting. For example, should PDC concepts be covered in a single required course (perhaps computer systems) or be scattered throughout different courses in the undergraduate curriculum? What languages are the best/easiest for students to learn PDC? How much revision is truly needed? This Birds of a Feather session provides a platform for computing educators to discuss the common challenges they face when attempting to incorporate PDC into their curricula and share potential solutions. Chiefly, the organizers are interested in identifying "gap areas" that hinder a faculty member's ability to integrate PDC into their undergraduate courses. The multiple viewpoints and expertise provided by the BOF leaders should lead to lively discourse and enable experienced faculty to share their strategies with those beginning to add PDC across their curricula. We anticipate that this session will be of interest to all CS faculty looking to integrate PDC into their courses and curricula. Suzanne J. Matthews, Joel Adams 0001, Richard A. Brown, Elizabeth Shoop |
SIGCSE | 1 |
| 2020 | Introducing Beginners to Distributed Computing using Raspberry Pi ClustersabstractThe 2019 ABET computer science criteria requires that all computing students learn parallel and distributed computing (PDC) as undergraduates, and CS2013 recommends at least fifteen hours of PDC in the undergraduate curriculum. Consequently, many educators look for easy ways to integrate PDC into courses at their institutions. This hands-on workshop introduces Message Passing Interface (MPI) basics in C/C++ and Python using clusters of Raspberry Pis. The Message Passing Interface (MPI) is a multi-language, platform independent, industry-standard library for parallel and distributed computing. Raspberry Pis are an inexpensive and engaging hardware platform for studying PDC as early as the first course. Participants will experience how to teach distributed computing essentials with MPI by means of reusable, effective "parallel patterns", including single program multiple data (SPMD) execution, send-receive message passing, the master-worker pattern, parallel loop patterns, and other common patterns, plus longer "exemplar" programs that use MPI to solve significant applied problems. The workshop includes: (i) personal experience with the Raspberry Pi (clusters provided for workshop use); (ii) assembly of Beowulf clusters of Raspberry Pis quickly in the classroom; (iii) self-paced hands-on experimentation with the working MPI programs; and (iv) a discussion of how these may be used to achieve the goals of CS2013 and ABET. No prior experience with MPI, PDC, or the Raspberry Pi is expected. All materials from this workshop will be freely available from CSinParallel.org; participants should bring a laptop to access these materials. Elizabeth Shoop, Joel Adams 0001, Richard A. Brown, Suzanne J. Matthews |
SIGCSE | 4 |
| 2019 | The Adventures of ScriptKitty: Teaching Middle School Students Cyber Awareness with Comics on the Raspberry PiabstractCyber security and on-line safety practices are not commonly taught in schools. However, there is an increasing need for education in these topics as children are joining the Internet community at a much earlier age than previous generations. It is crucial that young people understand the risks they may face on-line and how to mitigate them, ideally as soon as they begin using the Internet unsupervised. The Adventures of ScriptKitty (AOSK) introduces students to basic cyber security concepts using the Raspberry Pi, a single board computer that retails for $35.00. We created AOSK to help facilitate a culture of good cyber security practices and raise interest in STEM. The material is presented in the form of comics paired with instructional sections, including sections of more detailed technical information for readers who wish to learn more about key concepts. We piloted a portion of AOSK to a group of local middle school students. Our time with the students was limited, so we administered a short quiz, then discussed the Raspberry Pi. Next, students completed the packet sniffing exercise from Chapter 2, with the authors available to answer questions and help troubleshoot. Students were asked to re-take the quiz afterward. Our preliminary results show that students achieved a greater understanding of the material, with improved scores of 14%. A custom Pi image preloaded with Kali Linux and all needed software is included with the material. All the materials are published and available for free through GitBook at: https://suzannejmatthews.gitbooks.io/aosk/content Ovidiu-Gabriel Baciu-Ureche, Carlie Sleeman, Karlee Scott, William C. Moody, Suzanne J. Matthews |
SIGCSE | 5 |
| 2019 | Exploring Parallel Computing with OpenMP on the Raspberry PiabstractThe ACM/IEEE CS 2013 report recommends fifteen hours of parallel & distributed computing (PDC) education for every undergraduate. This workshop illustrates the use of the Raspberry Pi as an inexpensive, multicore platform for teaching shared-memory parallel programming. The inexpensive and tactile nature of the Raspberry Pi enables each student to experience her own parallel multiprocessor through sight and touch. In this hands-on workshop, we will teach attendees how they can leverage the Raspberry Pi and the OpenMP library to teach shared-memory parallel concepts in their own classrooms. All CS educators who are interested in learning about the Raspberry Pi, shared memory parallelism, and OpenMP are encouraged to attend. In Part I of the workshop, each participant will connect to and learn about the Raspberry Pi's multicore capabilities. In Part II, each participant will engage in self-paced, hands-on exploration of basic parallel computing concepts using the OpenMP "patternlets" from CSinParallel.org. In Part III, participants will investigate more complex applications, such as numeric integration and drug design and study how these applications can be parallelized using OpenMP. We will conclude the workshop with a series of lightning talks discussing how the Raspberry Pi has been used to teach parallel computing concepts at different institutions. We will also present a summary of student perceptions of the Raspberry Pi. All materials from this workshop will be freely available from CSinParallel.org. Space is limited to 20 participants. A laptop is required. Suzanne J. Matthews, Joel Adams 0001, Richard A. Brown, Elizabeth Shoop |
SIGCSE | 1 |
| 2018 | Leveraging the Raspberry Pi for CS EducationabstractThe Raspberry Pi (R-Pi) is a single board computer priced at 35 USD -- less than the cost of many textbooks. The current model (3B) includes a quad-core ARM 64-bit CPU, 1GB of RAM, a GPU, and numerous communication ports, including USB, HDMI, Ethernet, WiFi and Bluetooth. This combination of low cost and high functionality creates many new pedagogical possibilities for CS educators, ranging from using the R-Pi to teach assembly language to using it as a multiprocessor. Relatedly, mathematics educators have produced an extensive literature on the use of pedagogical tools known as "manipulatives" that have been shown to be effective at starting students through a "concrete, representational, abstract" progression of understanding of an abstract topic. We believe that by using the R-Pi as a manipulative, this same "concrete, representational, abstract" progression can be used to help CS students master many topics that are often taught as abstractions. By providing a "concrete" foundation on which to build, a single board computer like the R-Pi can provide the first step in helping students build mental models of such abstractions, and thus enhance student learning. Experience also indicates that many students find the R-Pi to be a fun and enjoyable way to learn about these abstractions. In this panel session, four CS educators will share their experiences using the R-Pi in their courses, followed by a Q&A conversation between the audience and the panelists. Joel Adams 0001, Richard A. Brown, Jalal Kawash, Suzanne J. Matthews, Elizabeth Shoop |
SIGCSE | 4 |
| 2018 | Teaching Parallel and Distributed Computing with MPI on Raspberry Pi Clusters: (Abstract Only)abstractCS2013 brings parallel and distributed computing (PDC) into the CS curricular mainstream. The Message Passing Interface (MPI) is a platform independent, industry-standard library for parallel and distributed computing. The MPI standard includes support for C, C++, and Fortran; third parties have created implementations for Python and Java. This hands-on workshop introduces MPI basics and applications in C/C++ using Raspberry Pi single-board computers, as an inexpensive and engaging hardware platform for studying PDC. The workshop includes: (i) personal experience with the Raspberry Pi (units provided) accessed via participant laptops (Windows, Mac, or Linux); (ii) assembly of Beowulf clusters of Raspberry Pis quickly in the classroom; (iii) self-paced hands-on experimentation with the working MPI programs; and (iv) a discussion of how such clusters can be used to engage students in and out of the classroom. Participants will experience how to teach distributed computing essentials with MPI by means of reusable, effective "parallel programming patterns," including single program multiple data (SPMD) execution, send-receive message passing, the master-worker, parallel loop, and other common patterns. Participants will then explore more in-depth "exemplar" applications, such as drug design and epidemiology. All materials including the Raspberry Pi software system setup from this workshop will be freely available from CSinParallel.org. No prior experience with MPI, PDC, or the Raspberry Pi is required. Windows, Mac, or Linux laptop required. Richard A. Brown, Joel Adams 0001, Suzanne J. Matthews, Elizabeth Shoop |
SIGCSE | 3 |
| 2018 | Portable Parallel Computing with the Raspberry PiabstractWith the requirement that parallel & distributed computing (PDC) topics be covered in the core computer science curriculum, educators are exploring new ways to engage students in this area of computing. In this paper, we discuss the use of the Raspberry Pi single-board computer (SBC) to provide students with hands-on multicore learning experiences. We discuss how the authors use the Raspberry Pi to teach parallel computing, and present assessment results that indicate such devices are effective at achieving CS2013 PDC learning outcomes, as well as motivating further study of parallelism. We believe our results are of significant interest to CS educators looking to integrate parallelism in their classrooms, and support the use of other SBCs for teaching parallel computing. Suzanne J. Matthews, Joel Adams 0001, Richard A. Brown, Elizabeth Shoop |
SIGCSE | 1 |
| 2017 | Teaching Parallel Computing with OpenMP on the Raspberry Pi (Abstract Only)abstractParallel computing is one of the new knowledge units in the ACM/IEEE CS 2013 curriculum recommendations. This workshop will present the Raspberry Pi as an inexpensive hardware platform for providing each student with her own parallel processor. The tactile and visceral benefits of each student having her own machine and being able to take full advantage of its multicore capabilities are significant. In this hands-on workshop, we show how parallelism can be used to spread the workload of compute-intensive applications across the multiple cores of a Raspberry Pi, and explore its use as an inexpensive hardware platform for teaching parallel computing. CS educators who are interested in learning about parallel computing, OpenMP, and how to teach these concepts on a Raspberry Pi are encouraged to attend. Attendees will enjoy a hands-on hardware/software experience, exploring how parallel computations operate and work in practice. In Part I of the workshop, attendees will set up and explore a Raspberry Pi multi-core computer in small teams. In Part II, each team will use the parallel capabilities of the Raspberry Pi to explore parallel computation through the use of OpenMP "patternlets" published on CSinParallel.org. Part III explores applications of the Raspberry Pi to parallel applications such as image processing and population dynamics, using OpenMP. All materials from this workshop will be freely available from CSinParallel.org. Suzanne J. Matthews, Joel Adams 0001, Richard A. Brown, Elizabeth Shoop |
SIGCSE | 1 |
| 2017 | Can we really do it?: Conducting Significant Computer Science Research in Primarily Undergraduate Institutions (PUIs) (Abstract Only)abstractUndergraduate research is a critical component of high-quality education in any discipline, including Computer Science (CS). Over the past few years, there has been a dramatic increase in CS undergraduate research activities at colleges and universities, and predominantly undergraduate institutions (PUIs) have an important role to play. Not every university has abundant resources to devote to research, and teaching-focused institutions may face the greatest challenges in this respect. Faculty at PUIs, for example, may face funding and infrastructure challenges and may find themselves stretched thin due to especially high teaching and service expectations. A frequently asked question by new faculty at these institutions is: Is it really possible to conduct meaningful research in such a fast-paced discipline as CS, while juggling a very high teaching and service load? Not only is the answer to this question "Yes!" but there are advantages to conducting research at a non-research institution. Faculty here has access to some of the brightest young minds who will potentially be future graduate students in research-intensive universities. They may have the freedom to do research that is too risky for graduate students. They can work on projects they are interested in, rather than those they know must work. With good time management techniques and careful selection of collaborators and student researchers, faculty here really can conduct important CS research. Thus, the focus of this BOF is to share methods that are helpful in conducting significant and meaningful CS research in a primarily undergraduate or teaching institution. Farzana Rahman, Suzanne J. Matthews, Kelly A. Shaw 0001, Andrea Pohoreckyj Danyluk |
SIGCSE | 2 |
| 2016 | The Micro-Cluster Showcase: 7 Inexpensive Beowulf Clusters for Teaching PDCabstractJust as a micro-computer is a personal, portable computer, a micro-cluster is a personal, portable, Beowulf cluster. In this special session, six cluster designers will bring and demonstrate micro-clusters they have built using inexpensive single-board computers (SBCs). The educators will describe how they have used their clusters to provide their students with hands-on experience using the shared-memory, distributed-memory, and heterogeneous computing paradigms, and thus achieve the parallel and distributed computing (PDC) objectives of CS 2013 [1]. Joel Adams 0001, Jacob Caswell, Suzanne J. Matthews, Charles Peck 0002, Elizabeth Shoop, David Toth, James Wolfer |
SIGCSE | 3 |
| 2015 | Accurate simulation of large collections of phylogenetic treesabstractPhylogenetic analyses are growing at a rapid rate, producing increasingly large collections of trees. Scientists rarely share their tree collections, making it difficult for researchers to develop methods that anticipate and respond to this growth of data. While common methods for simulating phylogenetic trees focus on random topologies, the tree collections returned from phylogenetic search are rarely random and contain a high degree of topological similarity. In this paper, we introduce TreeSim, a software package that simulates large tree collections from published consensus trees. TreeSim implements our new simulation algorithm, the combined consensus. Our experimental results indicate that simulating trees based on the combined consensus produces collections whose topological diversity most closely resemble the trees returned from phylogenetic search. We expect that TreeSim will play a critical role in guiding the algorithmic development of new approaches that support the growth of phylogenetic data. Suzanne J. Matthews |
BIBM | 1 |
| 2015 | Budget Beowulfs: A Showcase of Inexpensive Clusters for Teaching PDCabstractIn response to the shift to multicore processors, the ACM-IEEE CS2013 curriculum recommendations [1] include parallel and distributed computing (PDC) as a new core knowledge area. Some of the key concepts in PDC are the distinctions between shared-memory, distributed-memory, and heterogeneous system architectures. Joel Adams 0001, Jacob Caswell, Suzanne J. Matthews, Charles Peck 0002, Elizabeth Shoop, David Toth |
SIGCSE | 3 |
| 2015 | Parallel Author Verification of E-mail (Abstract Only)abstractCyber-crime is becoming alarmingly common through the use of anonymous e-mails. Author attribution helps digital forensics investigators filter through a large set of possible authors and focus traditional investigative techniques on the most probable culprits. A recent promising technique is to construct a write-print for each known author, and compare it to the write-print extracted from the anonymous message(s). A write-print is a unique digital fingerprint created by mining frequent patterns from a particular author's writing style. However, the process for generating a write-print is very slow, making it a poor choice for author attribution situations of a time-sensitive nature such as anonymous threats of attack, exposure, or ongoing harassment. Andreas Kellas, Alexander Molnar, Leo St. Amour, Frederick Ulrich, Suzanne J. Matthews |
SIGCSE | 5 |
| 2015 | Nifty AssignmentsabstractA great CS assignment is a delight to all, but the path to one can be most roundabout. Many CS students have had their characters built up on assignments that worked better as an idea than as an actual assignment. Assignments are hard to come up with, yet they are the key to student learning. The Nifty Assignments special session is all about promoting and sharing the ideas and ready-to-use materials of successful assignments. Nick Parlante, Julie Zelenski, Peter-Michael Osera, Marty Stepp, Mark Sherriff, Luther A. Tychonievich, Ryan Layer, Suzanne J. Matthews, Allison Obourn, David R. Raymond, Josh Hug, Stuart Reges |
SIGCSE | 8 |
| 2015 | Heterogeneous Compression of Large Collections of Evolutionary TreesabstractCompressing heterogeneous collections of trees is an open problem in computational phylogenetics. In a heterogeneous tree collection, each tree can contain a unique set of taxa. An ideal compression method would allow for the efficient archival of large tree collections and enable scientists to identify common evolutionary relationships over disparate analyses. In this paper, we extend TreeZip to compress heterogeneous collections of trees. TreeZip is the most efficient algorithm for compressing homogeneous tree collections. To the best of our knowledge, no other domain-based compression algorithm exists for large heterogeneous tree collections or enable their rapid analysis. Our experimental results indicate that TreeZip averages 89.03 percent (72.69 percent) space savings on unweighted (weighted) collections of trees when the level of heterogeneity in a collection is moderate. The organization of the TRZ file allows for efficient computations over heterogeneous data. For example, consensus trees can be computed in mere seconds. Lastly, combining the TreeZip compressed (TRZ) file with general-purpose compression yields average space savings of 97.34 percent (81.43 percent) on unweighted (weighted) collections of trees. Our results lead us to believe that TreeZip will prove invaluable in the efficient archival of tree collections, and enables scientists to develop novel methods for relating heterogeneous collections of trees. Suzanne J. Matthews |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2011 | An efficient and extensible approach for compressing phylogenetic treesabstractBACKGROUND: Biologists require new algorithms to efficiently compress and store their large collections of phylogenetic trees. Our previous work showed that TreeZip is a promising approach for compressing phylogenetic trees. In this paper, we extend our TreeZip algorithm by handling trees with weighted branches. Furthermore, by using the compressed TreeZip file as input, we have designed an extensible decompressor that can extract subcollections of trees, compute majority and strict consensus trees, and merge tree collections using set operations such as union, intersection, and set difference. RESULTS: On unweighted phylogenetic trees, TreeZip is able to compress Newick files in excess of 98%. On weighted phylogenetic trees, TreeZip is able to compress a Newick file by at least 73%. TreeZip can be combined with 7zip with little overhead, allowing space savings in excess of 99% (unweighted) and 92%(weighted). Unlike TreeZip, 7zip is not immune to branch rotations, and performs worse as the level of variability in the Newick string representation increases. Finally, since the TreeZip compressed text (TRZ) file contains all the semantic information in a collection of trees, we can easily filter and decompress a subset of trees of interest (such as the set of unique trees), or build the resulting consensus tree in a matter of seconds. We also show the ease of which set operations can be performed on TRZ files, at speeds quicker than those performed on Newick or 7zip compressed Newick files, and without loss of space savings. CONCLUSIONS: TreeZip is an efficient approach for compressing large collections of phylogenetic trees. The semantic and compact nature of the TRZ file allow it to be operated upon directly and quickly, without a need to decompress the original Newick file. We believe that TreeZip will be vital for compressing and archiving trees in the biological community. Suzanne J. Matthews, Tiffani L. Williams |
BMC Bioinform. | 1 |
| 2010 | TreeZip: A New Algorithm for Compressing Large Collections of Evolutionary TreesabstractThe primary advantage of TreeZip is its use of semantic compression, which allows us to uniquely store tree relationship information. Phylogenetic trees are stored in a format known as a Newick representation, which uses nested parentheses to represent the evolutionary relationships (or subtrees) within a phylogenetic tree. TreeZip uses two universal hashing functions in order to represent compactly all of the shared evolutionary relationships in the tree collection. In our previous work, we have used successively universal hash functions in our HashCS and HashRF algorithms that build consensus trees and topological distance matrices, respectively. Once the hash table is constructed, the TreeZip then writes the compressed file depicting the information contained in the collection of phylogenetic trees. Suzanne J. Matthews, Seung-Jin Sul, Tiffani L. Williams |
DCC | 1 |
| 2010 | A Novel Approach for Compressing Phylogenetic Trees
Suzanne J. Matthews, Seung-Jin Sul, Tiffani L. Williams |
ISBRA | 1 |
| 2010 | MrsRF: an efficient MapReduce algorithm for analyzing large collections of evolutionary treesabstractBACKGROUND: MapReduce is a parallel framework that has been used effectively to design large-scale parallel applications for large computing clusters. In this paper, we evaluate the viability of the MapReduce framework for designing phylogenetic applications. The problem of interest is generating the all-to-all Robinson-Foulds distance matrix, which has many applications for visualizing and clustering large collections of evolutionary trees. We introduce MrsRF (MapReduce Speeds up RF), a multi-core algorithm to generate a t x t Robinson-Foulds distance matrix between t trees using the MapReduce paradigm. RESULTS: We studied the performance of our MrsRF algorithm on two large biological trees sets consisting of 20,000 trees of 150 taxa each and 33,306 trees of 567 taxa each. Our experiments show that MrsRF is a scalable approach reaching a speedup of over 18 on 32 total cores. Our results also show that achieving top speedup on a multi-core cluster requires different cluster configurations. Finally, we show how to use an RF matrix to summarize collections of phylogenetic trees visually. CONCLUSION: Our results show that MapReduce is a promising paradigm for developing multi-core phylogenetic applications. The results also demonstrate that different multi-core configurations must be tested in order to obtain optimum performance. We conclude that RF matrices play a critical role in developing techniques to summarize large collections of trees. Suzanne J. Matthews, Tiffani L. Williams |
BMC Bioinform. | 1 |
| 2009 | Using tree diversity to compare phylogenetic heuristicsabstractBACKGROUND: Evolutionary trees are family trees that represent the relationships between a group of organisms. Phylogenetic heuristics are used to search stochastically for the best-scoring trees in tree space. Given that better tree scores are believed to be better approximations of the true phylogeny, traditional evaluation techniques have used tree scores to determine the heuristics that find the best scores in the fastest time. We develop new techniques to evaluate phylogenetic heuristics based on both tree scores and topologies to compare Pauprat and Rec-I-DCM3, two popular Maximum Parsimony search algorithms. RESULTS: Our results show that although Pauprat and Rec-I-DCM3 find the trees with the same best scores, topologically these trees are quite different. Furthermore, the Rec-I-DCM3 trees cluster distinctly from the Pauprat trees. In addition to our heatmap visualizations of using parsimony scores and the Robinson-Foulds distance to compare best-scoring trees found by the two heuristics, we also develop entropy-based methods to show the diversity of the trees found. Overall, Pauprat identifies more diverse trees than Rec-I-DCM3. CONCLUSION: Overall, our work shows that there is value to comparing heuristics beyond the parsimony scores that they find. Pauprat is a slower heuristic than Rec-I-DCM3. However, our work shows that there is tremendous value in using Pauprat to reconstruct trees-especially since it finds identical scoring but topologically distinct trees. Hence, instead of discounting Pauprat, effort should go in improving its implementation. Ultimately, improved performance measures lead to better phylogenetic heuristics and will result in better approximations of the true evolutionary history of the organisms of interest. Seung-Jin Sul, Suzanne J. Matthews, Tiffani L. Williams |
BMC Bioinform. | 2 |
| 2008 | New Approaches to Compare Phylogenetic Search HeuristicsabstractWe present new and novel insights into the behavior of two maximum parsimony heuristics for building evolutionary trees of different sizes. First, our results show that the heuristics find different classes of good-scoring trees, where the different classes of trees may have significant evolutionary implications. Secondly, we develop a new entropy-based measure to quantify the diversity among the evolutionary trees found by the heuristics. Overall, topological distance measures such as the Robinson-Foulds distance identify more diversity among a collection of trees than parsimony scores, which implies more powerful heuristics could be designed that use a combination of parsimony scores and topological distances. Thus, by understanding phylogenetic heuristic behavior, better heuristics could be designed, which ultimately leads to more accurate evolutionary trees. Seung-Jin Sul, Suzanne J. Matthews, Tiffani L. Williams |
BIBM | 2 |