VLDB 2026 Research / reviewers in the wild / expert
David R. Musicant
dblp:56/2195
· DBLP profile ↗
25ranked-venue papers
7as first author
2since 2021 · last 2026
0000-0002-6580-382XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 12 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 9 · 3 first-authorDatabases, data management, data science and information retrieval · 4 · 1 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Computing education · 69% Computational social science and digital humanities · 16% Bioinformatics and computational biology · 11% | |
| Artificial intelligence
4 papers |
Legged, aerial and field robots · 38% Kernel, tree and ensemble methods · 38% Probabilistic and Bayesian machine learning · 12% | |
| Human-computer interaction and pervasive computing
1 paper |
Collaborative and social computing · 100% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 100% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 100% |
Topics — the 13 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Collaborative and social computing › crowdsourcing
volunteered geographic information |
0.2 | 1 | 2015 | Barriers to the Localness of Volunteered Geographic Information · CHI 2015 |
Data mining › predictive modeling
classification |
0.1 | 1 | 2007 | Supervised Learning by Training on Aggregate Outputs · ICDM 2007 |
Data mining › predictive modeling
supervised learning |
0.1 | 1 | 2007 | Supervised Learning by Training on Aggregate Outputs · ICDM 2007 |
Machine learning › Kernel, tree and ensemble methods
support vector machine |
0.1 | 2 | 2001 | Lagrangian Support Vector Machines · J. Mach. Learn. Res. 2001 Active Support Vector Machine Classification · NIPS 2000 |
Mathematical optimization › continuous optimization
convex optimization |
0.1 | 2 | 2000 | Robust Linear and Support Vector Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2000 Active Support Vector Machine Classification · NIPS 2000 |
Mathematical optimization › continuous optimization › nonlinear optimization
quadratic programming |
0.1 | 2 | 2000 | Robust Linear and Support Vector Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2000 Active Support Vector Machine Classification · NIPS 2000 |
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis |
0.0 | 1 | 2004 | Mass Spectrum Labeling: Theory and Practice · ICDM 2004 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
regression |
0.0 | 1 | 2000 | Robust Linear and Support Vector Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2000 |
Machine learning › Learning theory › statistical estimation › robust statistics
robust regression |
0.0 | 1 | 2000 | Robust Linear and Support Vector Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2000 |
Machine learning › Kernel, tree and ensemble methods › support vector machine
support vector regression |
0.0 | 1 | 2000 | Robust Linear and Support Vector Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2000 |
Mathematical optimization › continuous optimization › nonlinear optimization › quadratic programming
active set method |
0.0 | 1 | 2000 | Active Support Vector Machine Classification · NIPS 2000 |
Privacy and data protection
privacy-preserving data analysis |
0.0 | 1 | 2007 | Supervised Learning by Training on Aggregate Outputs · ICDM 2007 |
Environmental and earth informatics
environmental monitoring |
0.0 | 1 | 2004 | Mass Spectrum Labeling: Theory and Practice · ICDM 2004 |
Methods — techniques the papers use, named apart from their topics
language analysis · 0.4geographic analysis · 0.4support vector machine · 0.1neural network · 0.1k-nearest neighbors · 0.1k-nearest neighbor · 0.1statistical learning theory · 0.1sherman-morrison-woodbury formula · 0.1huber m-estimator · 0.1convex quadratic programming · 0.1active set strategy · 0.1lagrangian optimization · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Kotlin in the Data Structures Classroom
David R. Musicant |
SIGCSE (2) | 1 |
| 2024 | Playing with Matches: Adopting Gale-Shapley for Managing Student Enrollments Beyond CS2abstractEnrollment in computer science has increased dramatically in recent years, straining capacities and leading to various strategies for managing enrollment. But some strategies increase student competition and may have disproportionate negative impacts on students from underrepresented groups. We believe success in computing education necessitates a more equitable approach to course enrollment. In this experience report, we describe our new enrollment mechanism, "the Match." Building on the Gale--Shapley stable matching algorithm, the Match was designed to encourage a liberal arts approach to course selection and attempt to broaden participation in computing. Drawing on data from three years of use, we find high student participation, with the vast majority of students having their enrollment preferences met. With Match registration, our courses have tended to be a bit more inclusive of younger students. The Match appears not to have disparate negative impacts like those of competitive enrollment, but has increased workload in the Registrar's Office. Overall, we believe the Match has decreased student and faculty angst around registration, and we argue that systems like the Match can help manage enrollment pressures in ways that are consistent with educational values. Anna N. Rafferty, David Liben-Nowell, David R. Musicant, Emy Farley, Allie Lyman, Ann May |
SIGCSE (1) | 3 |
| 2020 | NoSQL in Undergrad Courses is NoProblemabstractRelational databases have dominated both industry usage and academic database courses for decades. More recently, there has been a dramatic increase in the use of NoSQL database systems, especially in data science. Bringing NoSQL databases to the classroom is important not only to prepare our students for the technology that they will face, but also because the underlying paradigm introduces new ideas that are not typically emphasized in a relational-only database course alone. However, NoSQL systems seem to receive dramatically less coverage in undergraduate database curricula than relational systems do. This panel discusses examples of how NoSQL can be introduced into undergraduate education, and the possible challenges faced in doing so. Questions from the audience will be invited. Margaret S. Menzin, Sriram Mohan, David R. Musicant, Raja Sooriamurthi |
SIGCSE | 3 |
| 2019 | CS Principles Higher Education PathwaysabstractWith approximately 37,000 students entering college with an Advanced Placement CS Principles credit, students, parents, and teachers are wondering how that AP credit "counts" in college. In this panel, we will share perspectives from the College Board and higher education institutions of various pathways from CS Principles to a major in CS. While over 500 institutions have indicated they have credit and placement policies for CSP, they may vary quite a bit. (Credit gives students units of credit on their college record while placement means the AP credit can meet a particular course requirement.) For example, some institutions offer credit only, while others allow the course to count as part of a general education program, as a major elective, or even require CSP in the major. Panelists will share institutional context and reasoning behind their policies. Audience members will have a better understanding of the breadth of credit and placement options and considerations for their own institution. Crystal Furman, Owen L. Astrachan, Dan Garcia 0001, David R. Musicant, Jennifer Rosato |
SIGCSE | 4 |
| 2018 | Team-Teaching with Colleagues in the Arts and HumanitiesabstractThis panel will include experience reports from five computer science faculty members who have team-taught courses with professors from outside the sciences. Specifically, we will discuss lessons learned and best practices with collaborating with faculty from the arts and humanities. Courses that look outward have the potential to broaden participation and promote computing's role in the broader world beyond software engineering concerns. The panelists will highlight how to: find a topic, find a collaborator(s), design the course, maintain rigor in both disciplines, target the right audience, assess how well it worked, and do it more than once. Keith J. O'Hara, Sven Anderson, David R. Musicant, Amber Stubbs, Thomas P. Way |
SIGCSE | 3 |
| 2017 | Open-Ended Robotics Exploration Projects for Budding ResearchersabstractThere are many benefits to introducing students to the idea of doing projects where the outcome is unknown or unsure. Some have proposed that engaging students in research can help with retention of underrepresented groups. In this paper, we report on a particular approach we have used to introduce high school students to open-ended robotics projects in a three-week summer program. We describe the structure of our summer program, how we ramp the students up to speed, and we summarize the five open-ended "research" projects that the students work on. These projects can be adopted for open-ended work elsewhere by high school students or undergraduates. David R. Musicant, Abha Laddha, Tom Choi |
AAAI | 1 |
| 2017 | Elegit: Git Learning Tool for Students (Abstract Only)abstractVersion control systems are crucial tools for computer scientists, and the need for students to be fluent in them is well-recognized. However, Git and other version control systems (VCSs) are difficult to learn and use. Elegit is a new Git client that we created to help students learn how Git works while using it. Our approach is different from other GUI Git clients in that our key goals are not only to help students successfully use Git, but equally importantly to help students learn about how Git works in its own native way. We preserve standard Git terminology wherever possible, and place a high priority on not modifying the standard Git model. Simultaneously, we strive to make Elegit easy for beginners to use. This demo provides a brief tutorial on using Elegit, discussion on the process of designing the tool to do this, evaluation of the effectiveness of the tool, and improvements made based on this evaluation and our own learning of Git while developing the application. Information about Elegit can be found at http://elegit.org. This work is supported by a SIGCSE Special Projects Grant, and by Carleton College. Eric Walker, Julia Connelly, David R. Musicant |
SIGCSE | 3 |
| 2015 | Barriers to the Localness of Volunteered Geographic InformationabstractLocalness is an oft-cited benefit of volunteered geographic information (VGI). This study examines whether localness is a constant, universally shared benefit of VGI, or one that varies depending on the context in which it is produced. Focusing on articles about geographic entities (e.g. cities, points of interest) in 79 language editions of Wikipedia, we examine the localness of both the editors working on articles and the sources of the information they cite. We find extensive geographic inequalities in localness, with the degree of localness varying with the socioeconomic status of the local population and the health of the local media. We also point out the key role of language, showing that information in languages not native to a place tends to be produced and sourced by non-locals. We discuss the implications of this work for our understanding of the nature of VGI and highlight a generalizable technical contribution: an algorithm that determines the home country of the original publisher of online content. Shilad Sen, Heather Ford, David R. Musicant, Oliver S. B. Keyes, Brent J. Hecht |
CHI | 3 |
| 2015 | Engaging High School Students in Modeling and Simulation through Educational MediaabstractThe new AP CS Principles curriculum has a significant component regarding modeling and simulation, and many teachers will need to figure out how to accomplish this in their classroom. We have developed a new publicly available media-enhanced approach for teaching modeling and simulation, designated as TrafficJam. The approach consists of two core activities, both of which involve students optimizing traffic signals (also known as traffic lights, or stoplights). TrafficJam is likely best distinguished from other classroom simulation exercises in that these activities are introduced, motivated, and demonstrated in a video created by Twin Cities Public Television (TPT). The video uses an inquiry-driven format to feature four high school students who take on the task of improving signal timing in their own neighborhood. A pilot study of TrafficJam in four schools indicates that students find the video engaging, the activities relevant and interesting, and that they gain understanding of modeling and simulation from the experience. David R. Musicant, S. Selcen Guzey |
SIGCSE | 1 |
| 2013 | Getting to the source: where does Wikipedia get its information from?abstractWe ask what kinds of sources Wikipedians value most and compare Wikipedia's stated policy on sources to what we observe in practice. We find that primary data sources developed by alternative publishers are both popular and persistent, despite policies that present such sources as inferior to scholarly secondary sources. We also find that Wikipedians make almost equal use of information produced by associations such as nonprofits as from scholarly publishers, with a significant portion coming from government information sources. Our findings suggest the rise of new influential sources of information on the Web but also reinforce the traditional geographic patterns of scholarly publication. This has a significant effect on the goal of Wikipedians to represent "the sum of all human knowledge." Heather Ford, Shilad Sen, David R. Musicant, Nathaniel Miller |
OpenSym | 3 |
| 2010 | It seemed like a good idea at the timeabstractNo abstract available. Jonas Boustedt, Robert McCartney, Josh Tenenberg, Edward F. Gehringer, Raymond Lister, David R. Musicant |
SIGCSE | 6 |
| 2007 | Predicting User-Perceived Quality Ratings from Streaming Media DataabstractMedia stream quality is highly dependent on underlying network conditions, but identifying scalable, unambiguous metrics to discern the user-perceived quality of a media stream in the face of network congestion is a challenging problem. User-perceived quality can be approximated through the use of carefully chosen application layer metrics, precluding the need to poll users directly. We discuss the use of data mining prediction techniques to analyze application layer metrics to determine user-perceived quality ratings on media streams. We show that several such prediction techniques are able to assign correct (within a small tolerance) quality ratings to streams with a high degree of accuracy. The time it takes to train and tune the predictors and perform the actual prediction are short enough to make such a strategy feasible to be executed in real time and on real computer networks. Amy Csizmar Dalal, David R. Musicant, Jamie F. Olson, Brandy McMenamy, Sami Benzaid, Ben Kazez, Erica Bolan |
ICC | 2 |
| 2007 | Supervised Learning by Training on Aggregate OutputsabstractSupervised learning is a classic data mining problem where one wishes to be be able to predict an output value associated with a particular input vector. We present a new twist on this classic problem where, instead of having the training set contain an individual output value for each input vector, the output values in the training set are only given in aggregate over a number of input vectors. This new problem arose from a particular need in learning on mass spectrometry data, but could easily apply to situations when data has been aggregated in order to maintain privacy. We provide a formal description of this new problem for both classification and regression. We then examine how k-nearest neighbor, neural networks, and support vector machines can be adapted for this problem. David R. Musicant, Janara M. Christensen, Jamie F. Olson |
ICDM | 1 |
| 2007 | It seemed like a good idea at the timeabstractNo abstract available. Jonas Boustedt, Robert McCartney, Josh Tenenberg, Titus Winters, Stephen H. Edwards, Briana B. Morrison, David R. Musicant, Ian Utting, Carol Zander |
SIGCSE | 7 |
| 2007 | Mechanics of undergraduate research at liberal arts colleges: lessons learned
David R. Musicant, Amruth N. Kumar, Douglas Baldwin, Ellen Lowenfeld Walker |
SIGCSE | 1 |
| 2006 | Learning from Aggregate ViewsabstractIn this paper, we introduce a new class of data mining problems called learning from aggregate views. In contrast to the traditional problem of learning from a single table of training examples, the new goal is to learn from multiple aggregate views of the underlying data, without access to the un-aggregated data. We motivate this new problem, present a general problem framework, develop learning methods for RFA (Restriction-Free Aggregate) views defined using COUNT, SUM, AVG and STDEV, and offer theoretical and experimental results that characterize the proposed methods. Bee-Chung Chen, Lei Chen 0003, Raghu Ramakrishnan 0001, David R. Musicant |
ICDE | 4 |
| 2006 | Adapting K-Medians to Generate Normalized Cluster CentersabstractMany applications of clustering require the use of normalized data, such as text or mass spectra mining. The spherical K-means algorithm [6], an adaptation of the traditional K-means algorithm, is highly useful for data of this kind because it produces normalized cluster centers. The K-medians clustering algorithm is also an important clustering tool because of its wellknown resistance to outliers. K-medians, however, is not trivially adapted to produce normalized cluster centers. We introduce a new algorithm (called MN), inspired by spherical K-means, that integrates with Kmedians clustering to produce locally optimal normalized cluster centers. We then show theoretically and experimentally that MN produces clusters of significantly higher quality than one would obtain via a simple scaling of the cluster centers produced from traditional K-medians. Benjamin J. Anderson, Deborah S. Gross, David R. Musicant, Anna M. Ritz, Thomas G. Smith, Leah E. Steinberg |
SDM | 3 |
| 2006 | A data mining course for computer science: primary sources and implementationsabstractAn undergraduate elective course in data mining provides a strong opportunity for students to learn research skills, practice data structures, and enhance their understanding of algorithms. I have developed a data mining course built around the idea of using research-level papers as the primary reading material for the course, and implementing data mining algorithms for the assignments. Such a course is accessible to students with no prerequisites beyond the traditional data structures course, and allows students to experience both applied and theoretical work in a discipline that straddles multiple areas of computer science. This paper provides detailed descriptions of the readings and assignments that one could use to build a similar course. David R. Musicant |
SIGCSE | 1 |
| 2004 | Mass Spectrum Labeling: Theory and PracticeabstractWe introduce the problem of labeling a particle's mass spectrum with the substances it contains, and develop several formal representations of the problem, taking into account practical complications such as unknown compounds and noise. This task is currently a bottle-neck in analyzing data from a new generation of instruments for real-time environmental monitoring. Lei Chen 0003, Jin-Yi Cai, Deborah S. Gross, David R. Musicant, Raghu Ramakrishnan 0001, James J. Schauer, Stephen J. Wright 0001 |
ICDM | 5 |
| 2004 | Active set support vector regressionabstractThis paper presents active set support vector regression (ASVR), a new active set strategy to solve a straightforward reformulation of the standard support vector regression problem. This new algorithm is based on the successful ASVM algorithm for classification problems, and consists of solving a finite number of linear equations with a typically large dimensionality equal to the number of points to be approximated. However, by making use of the Sherman-Morrison-Woodbury formula, a much smaller matrix of the order of the original input space is inverted at each step. The algorithm requires no specialized quadratic or linear programming code, but merely a linear equation solver which is publicly available. ASVR is extremely fast, produces comparable generalization error to other popular algorithms, and is available on the web for download. David R. Musicant, Alexander Feinberg |
IEEE Trans. Neural Networks | 1 |
| 2002 | Large Scale Kernel Regression via Linear Programming
Olvi L. Mangasarian, David R. Musicant |
Mach. Learn. | 2 |
| 2001 | Lagrangian Support Vector Machines
Olvi L. Mangasarian, David R. Musicant |
J. Mach. Learn. Res. | 2 |
| 2000 | Active Support Vector Machine ClassificationabstractAn active set strategy is applied to the dual of a simple reformula(cid:173) tion of the standard quadratic program of a linear support vector machine. This application generates a fast new dual algorithm that consists of solving a finite number of linear equations, with a typically large dimensionality equal to the number of points to be classified. However, by making novel use of the Sherman-Morrison(cid:173) Woodbury formula, a much smaller matrix of the order of the orig(cid:173) inal input space is inverted at each step. Thus, a problem with a 32-dimensional input space and 7 million points required inverting positive definite symmetric matrices of size 33 x 33 with a total run(cid:173) ning time of 96 minutes on a 400 MHz Pentium II. The algorithm requires no specialized quadratic or linear programming code, but merely a linear equation solver which is publicly available. Olvi L. Mangasarian, David R. Musicant |
NIPS | 2 |
| 2000 | Robust Linear and Support Vector RegressionabstractThe robust Huber M-estimator, a differentiable cost function that is quadratic for small errors and linear otherwise, is modeled exactly, in the original primal space of the problem, by an easily solvable simple convex quadratic program for both linear and nonlinear support vector estimators. Previous models were significantly more complex or formulated in the dual space and most involved specialized numerical algorithms for solving the robust Huber linear estimator. Numerical test comparisons with these algorithms indicate the computational effectiveness of the new quadratic programming model for both linear and nonlinear support vector problems. Results are shown on problems with as many as 20000 data points, with considerably faster running times on larger problems. Olvi L. Mangasarian, David R. Musicant |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1999 | Successive overrelaxation for support vector machinesabstractSuccessive overrelaxation (SOR) for symmetric linear complementarity problems and quadratic programs is used to train a support vector machine (SVM) for discriminating between the elements of two massive datasets, each with millions of points. Because SOR handles one point at a time, similar to Platt's sequential minimal optimization (SMO) algorithm which handles two constraints at a time and Joachims' SVMlight which handles a small number of points at a time, SOR can process very large datasets that need not reside in memory. The algorithm converges linearly to a solution. Encouraging numerical results are presented on datasets with up to 10,000,000 points. Such massive discrimination problems cannot be processed by conventional linear or quadratic programming methods, and to our knowledge have not been solved by other methods. On smaller problems, SOR was faster than SVMlight and comparable or faster than SMO. Olvi L. Mangasarian, David R. Musicant |
IEEE Trans. Neural Networks | 2 |