David R. Musicant

dblp:56/2195 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
2since 2021 · last 2026
0000-0002-6580-382XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 12 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 9 · 3 first-authorDatabases, data management, data science and information retrieval · 4 · 1 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Computing education · 69% Computational social science and digital humanities · 16% Bioinformatics and computational biology · 11%
Artificial intelligence
4 papers
Legged, aerial and field robots · 38% Kernel, tree and ensemble methods · 38% Probabilistic and Bayesian machine learning · 12%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 100%
Databases, data mining, and information retrieval
2 papers
Data mining · 100%
Theoretical computer science
2 papers
Mathematical optimization · 100%

Topics — the 13 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Collaborative and social computing › crowdsourcing
volunteered geographic information
0.212015
Barriers to the Localness of Volunteered Geographic Information · CHI 2015
Data mining › predictive modeling
classification
0.112007
Supervised Learning by Training on Aggregate Outputs · ICDM 2007
Data mining › predictive modeling
supervised learning
0.112007
Supervised Learning by Training on Aggregate Outputs · ICDM 2007
Machine learning › Kernel, tree and ensemble methods
support vector machine
0.122001
Lagrangian Support Vector Machines · J. Mach. Learn. Res. 2001
Active Support Vector Machine Classification · NIPS 2000
Mathematical optimization › continuous optimization
convex optimization
0.122000
Robust Linear and Support Vector Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2000
Active Support Vector Machine Classification · NIPS 2000
Mathematical optimization › continuous optimization › nonlinear optimization
quadratic programming
0.122000
Robust Linear and Support Vector Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2000
Active Support Vector Machine Classification · NIPS 2000
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis
0.012004
Mass Spectrum Labeling: Theory and Practice · ICDM 2004
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
regression
0.012000
Robust Linear and Support Vector Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2000
Machine learning › Learning theory › statistical estimation › robust statistics
robust regression
0.012000
Robust Linear and Support Vector Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2000
Machine learning › Kernel, tree and ensemble methods › support vector machine
support vector regression
0.012000
Robust Linear and Support Vector Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2000
Mathematical optimization › continuous optimization › nonlinear optimization › quadratic programming
active set method
0.012000
Active Support Vector Machine Classification · NIPS 2000
Privacy and data protection
privacy-preserving data analysis
0.012007
Supervised Learning by Training on Aggregate Outputs · ICDM 2007
Environmental and earth informatics
environmental monitoring
0.012004
Mass Spectrum Labeling: Theory and Practice · ICDM 2004

Methods — techniques the papers use, named apart from their topics

language analysis · 0.4geographic analysis · 0.4support vector machine · 0.1neural network · 0.1k-nearest neighbors · 0.1k-nearest neighbor · 0.1statistical learning theory · 0.1sherman-morrison-woodbury formula · 0.1huber m-estimator · 0.1convex quadratic programming · 0.1active set strategy · 0.1lagrangian optimization · 0.0
YearPublicationVenuePosition
2026 Kotlin in the Data Structures Classroom
David R. Musicant
SIGCSE (2)1
2024 Playing with Matches: Adopting Gale-Shapley for Managing Student Enrollments Beyond CS2
abstract
Enrollment in computer science has increased dramatically in recent years, straining capacities and leading to various strategies for managing enrollment. But some strategies increase student competition and may have disproportionate negative impacts on students from underrepresented groups. We believe success in computing education necessitates a more equitable approach to course enrollment. In this experience report, we describe our new enrollment mechanism, "the Match." Building on the Gale--Shapley stable matching algorithm, the Match was designed to encourage a liberal arts approach to course selection and attempt to broaden participation in computing. Drawing on data from three years of use, we find high student participation, with the vast majority of students having their enrollment preferences met. With Match registration, our courses have tended to be a bit more inclusive of younger students. The Match appears not to have disparate negative impacts like those of competitive enrollment, but has increased workload in the Registrar's Office. Overall, we believe the Match has decreased student and faculty angst around registration, and we argue that systems like the Match can help manage enrollment pressures in ways that are consistent with educational values.
Anna N. Rafferty, David Liben-Nowell, David R. Musicant, Emy Farley, Allie Lyman, Ann May
SIGCSE (1)3
2020 NoSQL in Undergrad Courses is NoProblem
abstract
Relational databases have dominated both industry usage and academic database courses for decades. More recently, there has been a dramatic increase in the use of NoSQL database systems, especially in data science. Bringing NoSQL databases to the classroom is important not only to prepare our students for the technology that they will face, but also because the underlying paradigm introduces new ideas that are not typically emphasized in a relational-only database course alone. However, NoSQL systems seem to receive dramatically less coverage in undergraduate database curricula than relational systems do. This panel discusses examples of how NoSQL can be introduced into undergraduate education, and the possible challenges faced in doing so. Questions from the audience will be invited.
Margaret S. Menzin, Sriram Mohan, David R. Musicant, Raja Sooriamurthi
SIGCSE3
2019 CS Principles Higher Education Pathways
abstract
With approximately 37,000 students entering college with an Advanced Placement CS Principles credit, students, parents, and teachers are wondering how that AP credit "counts" in college. In this panel, we will share perspectives from the College Board and higher education institutions of various pathways from CS Principles to a major in CS. While over 500 institutions have indicated they have credit and placement policies for CSP, they may vary quite a bit. (Credit gives students units of credit on their college record while placement means the AP credit can meet a particular course requirement.) For example, some institutions offer credit only, while others allow the course to count as part of a general education program, as a major elective, or even require CSP in the major. Panelists will share institutional context and reasoning behind their policies. Audience members will have a better understanding of the breadth of credit and placement options and considerations for their own institution.
Crystal Furman, Owen L. Astrachan, Dan Garcia 0001, David R. Musicant, Jennifer Rosato
SIGCSE4
2018 Team-Teaching with Colleagues in the Arts and Humanities
abstract
This panel will include experience reports from five computer science faculty members who have team-taught courses with professors from outside the sciences. Specifically, we will discuss lessons learned and best practices with collaborating with faculty from the arts and humanities. Courses that look outward have the potential to broaden participation and promote computing's role in the broader world beyond software engineering concerns. The panelists will highlight how to: find a topic, find a collaborator(s), design the course, maintain rigor in both disciplines, target the right audience, assess how well it worked, and do it more than once.
Keith J. O'Hara, Sven Anderson, David R. Musicant, Amber Stubbs, Thomas P. Way
SIGCSE3
2017 Open-Ended Robotics Exploration Projects for Budding Researchers
abstract
There are many benefits to introducing students to the idea of doing projects where the outcome is unknown or unsure. Some have proposed that engaging students in research can help with retention of underrepresented groups. In this paper, we report on a particular approach we have used to introduce high school students to open-ended robotics projects in a three-week summer program. We describe the structure of our summer program, how we ramp the students up to speed, and we summarize the five open-ended "research" projects that the students work on. These projects can be adopted for open-ended work elsewhere by high school students or undergraduates.
David R. Musicant, Abha Laddha, Tom Choi
AAAI1
2017 Elegit: Git Learning Tool for Students (Abstract Only)
abstract
Version control systems are crucial tools for computer scientists, and the need for students to be fluent in them is well-recognized. However, Git and other version control systems (VCSs) are difficult to learn and use. Elegit is a new Git client that we created to help students learn how Git works while using it. Our approach is different from other GUI Git clients in that our key goals are not only to help students successfully use Git, but equally importantly to help students learn about how Git works in its own native way. We preserve standard Git terminology wherever possible, and place a high priority on not modifying the standard Git model. Simultaneously, we strive to make Elegit easy for beginners to use. This demo provides a brief tutorial on using Elegit, discussion on the process of designing the tool to do this, evaluation of the effectiveness of the tool, and improvements made based on this evaluation and our own learning of Git while developing the application. Information about Elegit can be found at http://elegit.org. This work is supported by a SIGCSE Special Projects Grant, and by Carleton College.
Eric Walker, Julia Connelly, David R. Musicant
SIGCSE3
2015 Barriers to the Localness of Volunteered Geographic Information
abstract
Localness is an oft-cited benefit of volunteered geographic information (VGI). This study examines whether localness is a constant, universally shared benefit of VGI, or one that varies depending on the context in which it is produced. Focusing on articles about geographic entities (e.g. cities, points of interest) in 79 language editions of Wikipedia, we examine the localness of both the editors working on articles and the sources of the information they cite. We find extensive geographic inequalities in localness, with the degree of localness varying with the socioeconomic status of the local population and the health of the local media. We also point out the key role of language, showing that information in languages not native to a place tends to be produced and sourced by non-locals. We discuss the implications of this work for our understanding of the nature of VGI and highlight a generalizable technical contribution: an algorithm that determines the home country of the original publisher of online content.
Shilad Sen, Heather Ford, David R. Musicant, Oliver S. B. Keyes, Brent J. Hecht
CHI3
2015 Engaging High School Students in Modeling and Simulation through Educational Media
abstract
The new AP CS Principles curriculum has a significant component regarding modeling and simulation, and many teachers will need to figure out how to accomplish this in their classroom. We have developed a new publicly available media-enhanced approach for teaching modeling and simulation, designated as TrafficJam. The approach consists of two core activities, both of which involve students optimizing traffic signals (also known as traffic lights, or stoplights). TrafficJam is likely best distinguished from other classroom simulation exercises in that these activities are introduced, motivated, and demonstrated in a video created by Twin Cities Public Television (TPT). The video uses an inquiry-driven format to feature four high school students who take on the task of improving signal timing in their own neighborhood. A pilot study of TrafficJam in four schools indicates that students find the video engaging, the activities relevant and interesting, and that they gain understanding of modeling and simulation from the experience.
David R. Musicant, S. Selcen Guzey
SIGCSE1
2013 Getting to the source: where does Wikipedia get its information from?
abstract
We ask what kinds of sources Wikipedians value most and compare Wikipedia's stated policy on sources to what we observe in practice. We find that primary data sources developed by alternative publishers are both popular and persistent, despite policies that present such sources as inferior to scholarly secondary sources. We also find that Wikipedians make almost equal use of information produced by associations such as nonprofits as from scholarly publishers, with a significant portion coming from government information sources. Our findings suggest the rise of new influential sources of information on the Web but also reinforce the traditional geographic patterns of scholarly publication. This has a significant effect on the goal of Wikipedians to represent "the sum of all human knowledge."
Heather Ford, Shilad Sen, David R. Musicant, Nathaniel Miller
OpenSym3
2010 It seemed like a good idea at the time
abstract
No abstract available.
Jonas Boustedt, Robert McCartney, Josh Tenenberg, Edward F. Gehringer, Raymond Lister, David R. Musicant
SIGCSE6
2007 Predicting User-Perceived Quality Ratings from Streaming Media Data
abstract
Media stream quality is highly dependent on underlying network conditions, but identifying scalable, unambiguous metrics to discern the user-perceived quality of a media stream in the face of network congestion is a challenging problem. User-perceived quality can be approximated through the use of carefully chosen application layer metrics, precluding the need to poll users directly. We discuss the use of data mining prediction techniques to analyze application layer metrics to determine user-perceived quality ratings on media streams. We show that several such prediction techniques are able to assign correct (within a small tolerance) quality ratings to streams with a high degree of accuracy. The time it takes to train and tune the predictors and perform the actual prediction are short enough to make such a strategy feasible to be executed in real time and on real computer networks.
Amy Csizmar Dalal, David R. Musicant, Jamie F. Olson, Brandy McMenamy, Sami Benzaid, Ben Kazez, Erica Bolan
ICC2
2007 Supervised Learning by Training on Aggregate Outputs
abstract
Supervised learning is a classic data mining problem where one wishes to be be able to predict an output value associated with a particular input vector. We present a new twist on this classic problem where, instead of having the training set contain an individual output value for each input vector, the output values in the training set are only given in aggregate over a number of input vectors. This new problem arose from a particular need in learning on mass spectrometry data, but could easily apply to situations when data has been aggregated in order to maintain privacy. We provide a formal description of this new problem for both classification and regression. We then examine how k-nearest neighbor, neural networks, and support vector machines can be adapted for this problem.
David R. Musicant, Janara M. Christensen, Jamie F. Olson
ICDM1
2007 It seemed like a good idea at the time
abstract
No abstract available.
Jonas Boustedt, Robert McCartney, Josh Tenenberg, Titus Winters, Stephen H. Edwards, Briana B. Morrison, David R. Musicant, Ian Utting, Carol Zander
SIGCSE7
2007 Mechanics of undergraduate research at liberal arts colleges: lessons learned
David R. Musicant, Amruth N. Kumar, Douglas Baldwin, Ellen Lowenfeld Walker
SIGCSE1
2006 Learning from Aggregate Views
abstract
In this paper, we introduce a new class of data mining problems called learning from aggregate views. In contrast to the traditional problem of learning from a single table of training examples, the new goal is to learn from multiple aggregate views of the underlying data, without access to the un-aggregated data. We motivate this new problem, present a general problem framework, develop learning methods for RFA (Restriction-Free Aggregate) views defined using COUNT, SUM, AVG and STDEV, and offer theoretical and experimental results that characterize the proposed methods.
Bee-Chung Chen, Lei Chen 0003, Raghu Ramakrishnan 0001, David R. Musicant
ICDE4
2006 Adapting K-Medians to Generate Normalized Cluster Centers
abstract
Many applications of clustering require the use of normalized data, such as text or mass spectra mining. The spherical K-means algorithm [6], an adaptation of the traditional K-means algorithm, is highly useful for data of this kind because it produces normalized cluster centers. The K-medians clustering algorithm is also an important clustering tool because of its wellknown resistance to outliers. K-medians, however, is not trivially adapted to produce normalized cluster centers. We introduce a new algorithm (called MN), inspired by spherical K-means, that integrates with Kmedians clustering to produce locally optimal normalized cluster centers. We then show theoretically and experimentally that MN produces clusters of significantly higher quality than one would obtain via a simple scaling of the cluster centers produced from traditional K-medians.
Benjamin J. Anderson, Deborah S. Gross, David R. Musicant, Anna M. Ritz, Thomas G. Smith, Leah E. Steinberg
SDM3
2006 A data mining course for computer science: primary sources and implementations
abstract
An undergraduate elective course in data mining provides a strong opportunity for students to learn research skills, practice data structures, and enhance their understanding of algorithms. I have developed a data mining course built around the idea of using research-level papers as the primary reading material for the course, and implementing data mining algorithms for the assignments. Such a course is accessible to students with no prerequisites beyond the traditional data structures course, and allows students to experience both applied and theoretical work in a discipline that straddles multiple areas of computer science. This paper provides detailed descriptions of the readings and assignments that one could use to build a similar course.
David R. Musicant
SIGCSE1
2004 Mass Spectrum Labeling: Theory and Practice
abstract
We introduce the problem of labeling a particle's mass spectrum with the substances it contains, and develop several formal representations of the problem, taking into account practical complications such as unknown compounds and noise. This task is currently a bottle-neck in analyzing data from a new generation of instruments for real-time environmental monitoring.
Lei Chen 0003, Jin-Yi Cai, Deborah S. Gross, David R. Musicant, Raghu Ramakrishnan 0001, James J. Schauer, Stephen J. Wright 0001
ICDM5
2004 Active set support vector regression
abstract
This paper presents active set support vector regression (ASVR), a new active set strategy to solve a straightforward reformulation of the standard support vector regression problem. This new algorithm is based on the successful ASVM algorithm for classification problems, and consists of solving a finite number of linear equations with a typically large dimensionality equal to the number of points to be approximated. However, by making use of the Sherman-Morrison-Woodbury formula, a much smaller matrix of the order of the original input space is inverted at each step. The algorithm requires no specialized quadratic or linear programming code, but merely a linear equation solver which is publicly available. ASVR is extremely fast, produces comparable generalization error to other popular algorithms, and is available on the web for download.
David R. Musicant, Alexander Feinberg
IEEE Trans. Neural Networks1
2002 Large Scale Kernel Regression via Linear Programming
Olvi L. Mangasarian, David R. Musicant
Mach. Learn.2
2001 Lagrangian Support Vector Machines
Olvi L. Mangasarian, David R. Musicant
J. Mach. Learn. Res.2
2000 Active Support Vector Machine Classification
abstract
An active set strategy is applied to the dual of a simple reformula(cid:173) tion of the standard quadratic program of a linear support vector machine. This application generates a fast new dual algorithm that consists of solving a finite number of linear equations, with a typically large dimensionality equal to the number of points to be classified. However, by making novel use of the Sherman-Morrison(cid:173) Woodbury formula, a much smaller matrix of the order of the orig(cid:173) inal input space is inverted at each step. Thus, a problem with a 32-dimensional input space and 7 million points required inverting positive definite symmetric matrices of size 33 x 33 with a total run(cid:173) ning time of 96 minutes on a 400 MHz Pentium II. The algorithm requires no specialized quadratic or linear programming code, but merely a linear equation solver which is publicly available.
Olvi L. Mangasarian, David R. Musicant
NIPS2
2000 Robust Linear and Support Vector Regression
abstract
The robust Huber M-estimator, a differentiable cost function that is quadratic for small errors and linear otherwise, is modeled exactly, in the original primal space of the problem, by an easily solvable simple convex quadratic program for both linear and nonlinear support vector estimators. Previous models were significantly more complex or formulated in the dual space and most involved specialized numerical algorithms for solving the robust Huber linear estimator. Numerical test comparisons with these algorithms indicate the computational effectiveness of the new quadratic programming model for both linear and nonlinear support vector problems. Results are shown on problems with as many as 20000 data points, with considerably faster running times on larger problems.
Olvi L. Mangasarian, David R. Musicant
IEEE Trans. Pattern Anal. Mach. Intell.2
1999 Successive overrelaxation for support vector machines
abstract
Successive overrelaxation (SOR) for symmetric linear complementarity problems and quadratic programs is used to train a support vector machine (SVM) for discriminating between the elements of two massive datasets, each with millions of points. Because SOR handles one point at a time, similar to Platt's sequential minimal optimization (SMO) algorithm which handles two constraints at a time and Joachims' SVMlight which handles a small number of points at a time, SOR can process very large datasets that need not reside in memory. The algorithm converges linearly to a solution. Encouraging numerical results are presented on datasets with up to 10,000,000 points. Such massive discrimination problems cannot be processed by conventional linear or quadratic programming methods, and to our knowledge have not been solved by other methods. On smaller problems, SOR was faster than SVMlight and comparable or faster than SMO.
Olvi L. Mangasarian, David R. Musicant
IEEE Trans. Neural Networks2