VLDB 2026 Research / reviewers in the wild / expert
Mehran Sahami
dblp:96/481
· DBLP profile ↗
52ranked-venue papers
20as first author
6since 2021 · last 2025
0000-0001-9549-6974ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 28 · 12 first-author · 4 since 2021Artificial intelligence and machine learning · 16 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 11 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement But May Increase Adopters' Exam PerformancesabstractLarge language models (LLMs) are quickly being adopted in a wide range of learning experiences, especially via ubiquitous and broadly accessible chat interfaces like ChatGPT. This type of interface is readily available to students and teachers around the world. Coding education is an interesting test case, both because LLMs have strong performance on coding tasks, and because LLM-powered support tools are rapidly becoming part of the workflow of professional software engineers. To help understand the impact of generic LLM use on coding education, we conducted a large-scale randomized control trial with 5,831 students from 146 countries in an online coding class in which we provided some students with access to a chat interface with GPT-4. Under some assumptions, we estimate positive benefits on exam performance for adopters, the students who used the tool, but over all students, the advertisement of GPT-4 led to a significant average decrease in exam participation. We observe similar decreases in other forms of course engagement. However, this decrease is modulated by the student's country of origin. Offering access to LLMs to students from low human development index countries increased their exam participation rate on average. Our results suggest there may be promising benefits to using LLMs in an introductory coding class, but also potential harms for engagement, which makes their longer term impact on student success unclear. Our work highlights the need for additional investigations to help understand the potential impact of future adoption and integration of LLMs into classrooms. Allen Nie, Yash Chandak, Miroslav Suzara, Ali Malik, Juliette Woodrow, Matt Peng, Mehran Sahami, Emma Brunskill, Chris Piech |
L@S | 7 |
| 2025 | Infinite Story
Chris Piech, Mehran Sahami, Yasmine Alonso, Katie Liu, Javokhir Arifov, Anjali Sreenivas, Dan Webber, Tina Zheng, Ngoc Nguyen, Iddah Mlauzi, Juliette Woodrow |
SIGCSE (2) | 2 |
| 2023 | Teaching Responsible Computing in Context: Models, Practices, and ToolsabstractRecent news and national reports have significantly increased interest in new approaches for teaching responsible computing to help students understand, evaluate, and address the social impact of existing and emerging computing technologies. This 3-hour workshop will be offered in two workshop format sessions: in-person and online. The first part of each session will introduce responsible computing and its connections to RESPECT and Cultural Competence in Computing (3C). Next, we will provide a short overview of our own work in teaching responsible computing along with frameworks, tools, and best practices. We will showcase four different approaches to teaching responsible computing across institutional settings (high school, college, university), interdisciplinary partnerships (computing, philosophy, STS, digital humanities), and instructional formats (dedicated courses, embedded lessons, design challenges, bootcamps). The workshop presentations will focus on practical advice about how to get started, available resources, securing support from administration and colleagues, and other considerations for this work. In the second half of the workshop, participants will work in small groups to co-design potential lessons based on shared topical interests, institutional settings, and/or learning objectives. Facilitators will provide guidance, recommendations, and classroom examples to help the small groups to complete draft lessons that will be disseminated among workshop participants and on the workshop website. A laptop or internet connected device is needed to participate in the small group activity. Handouts/materials will be provided on the workshop website. Stacy A. Doore, Atri Rudra, Omowumi Ogunyemi, Trystan S. Goetze, Mehran Sahami, Thomas J. Cortina, Kiran Bhardwaj, Crystal Lee |
SIGCSE (2) | 5 |
| 2022 | Should the AP Computer Science A Exam Switch to Using Python?abstractChanging the language used to teach the AP Computer Science A course is an expensive, time-consuming, and ultimately controversial endeavor. However, it is worth raising the question from time to time to determine if such a change may now be appropriate. This panel offers a set of speakers specifically chosen to provide a diverse set of viewpoints and experiences (high school teacher, curriculum developer, exam reader, exam development committee member, university credit-granting administrator) on the question of whether the language in the AP CS A exam should be switched to Python in the near future. Mehran Sahami, Owen L. Astrachan, Sandy Czajka, Adrienne Decker, Jennifer Rosato |
SIGCSE (2) | 1 |
| 2021 | SimGrade: Using Code Similarity Measures for More Accurate Human Grading
Sonja Johnson-Yu, Nicholas Bowman, Mehran Sahami, Chris Piech |
EDM | 3 |
| 2021 | Code in Place: Online Section Leading for Scalable Human-Centered LearningabstractCould it be the case that the number of people who want to teach computer science, and have the potential, is roughly proportional to the number of people who want to learn? During the time of COVID-19 we offered a free CS1 class to people around the world. Well-aware of the high drop-out rates reported in many massive open-access online courses (MOOCs), we augmented our course with a scalable, human-centered solution: section leading. Section leaders teach small, weekly interactive learning sessions. We hypothesize that the personalized attention adds a sense of responsibility for both student and teacher which drives learning. We recruited over 900 volunteer section leaders and more than 10,000 students in the class. To our knowledge this is the largest group of section leaders in a single CS1 course offering and the most small group interactions. The completion rate in our class was more than 10 times that usually reported for similar MOOCs. Additionally, 99% of the volunteer section leaders taught through the entire span of the course, showing the potential for large scale volunteer-driven education, and the benefit that teachers themselves derive. We also discovered the potential for replication of this model, as 34% of students in a representative-sample survey indicated they would serve as section leaders for a future offering of the course. This level of participation would be more than sufficient to field additional offerings of the course sustainably. We believe this is an intriguing case study of a model for significantly scaling human-centric CS education for all. Chris Piech, Ali Malik, Kylie Jue, Mehran Sahami |
SIGCSE | 4 |
| 2020 | Assignments that Blend Ethics and TechnologyabstractWith the 2018 revision of the ACM Code of Ethics and Professional Conduct, there is a growing interest in how computer science faculty can integrate these principles into the education of future practitioners. This special session illustrates one approach by highlighting assignments that blend ethics and technology. These assignments can be used in a variety of courses, including CS1, CS2, and later courses. Presenters will provide an overview of each assignment and gather feedback from the audience. All materials, including descriptions, starter files, and guidelines for instructors, will be published at https://ethics.acm.org/SIGCSE2020. Stacy A. Doore, Casey Fiesler, Michael S. Kirkpatrick, Evan M. Peck, Mehran Sahami |
SIGCSE | 5 |
| 2020 | Teaching Computer Ethics: A Deeply Multidisciplinary ApproachabstractWe report on a curricular experiment at Stanford University focused on teaching computer ethics. After nearly a year of preparation, we launched a new course at the intersection of ethics, public policy, and technology that deeply marries the humanities, social sciences, and computer science. While the teaching of computer ethics courses dates back decades, such courses are often taught by a (single) CS faculty member without significant training in ethics, do not include a policy component, and are meant for CS students. By contrast, we take a deeply multidisciplinary approach, where three faculty instructors, from philosophy, political science, and CS, each bring their respective lens to four related course modules: algorithmic decision-making, data privacy and civil liberties, AI and autonomous systems, and the power of platform companies. Panels of guest speakers drawn from academia, industry, civil society, and government provide a practitioner's view of the topics addressed. Additionally, custom case studies were developed under the direction of the course staff. These materials (videos of the speaker panels and the case studies) are freely available for use by the broader community. We report on the details of the course structure, including how multiple disciplines are integrated throughout the course, including lectures, discussions, and assignments. We discuss aspects of the course that worked well as well as challenges in making the course broadly accessible (beyond just CS majors). Importantly, we also include a discussion of students' response to the course, showing that a deeply multidisciplinary approach resonates strongly with them. Rob Reich, Mehran Sahami, Jeremy M. Weinstein, Hilary Cohen |
SIGCSE | 2 |
| 2019 | Wrestling with Retention in the CS Major: Report from the ACM Retention CommitteeabstractThis panel focuses on the challenges of collecting and analyzing data relating to retention of students in undergraduate computer science education programs. The panelists will share learnings and recommendations from the final report of the ACM Retention Committee and share their individual perspectives on data collection challenges, promising interventions, and recommendations for actively addressing retention for all students. Alison Derbenwick Miller, Christine Alvarado, Mehran Sahami, Elsa Q. Villa, Stuart H. Zweben |
SIGCSE | 3 |
| 2018 | Five Slides About: Abstraction, Arrays, Uncomputability, Networks, Digital Portfolios, and the CS Principles Explore Performance TaskabstractSIGCSE is packed with teaching insights and inspiration. However, we get these insights and inspiration from hearing our colleagues talk about their teaching. Why not just watch them teach? This session does exactly that. Each of six exceptional educators will be given ten minutes to teach the audience something. After this, the moderator will draw the attention of the audience to particular pedagogical moves that the instruction included. Attendees can see a new approach to introducing a topic or a new pedagogical move. No matter what, we expect attendees will be taking ideas from this session directly back to their teaching! The format is based upon a practice in chemistry of sharing "Five Slides About," which introduce a topic in a novel or concise way (https://www.ionicviper.org/types/five_slides_about). Resources from each of the presenters will be shared on the website CSTeachingTips.org. Colleen M. Lewis, Leslie Aaronson, Eric Allatta, Zachary Dodds, Jeffrey Forbes 0001, Kyla A. McMullen, Mehran Sahami |
SIGCSE | 7 |
| 2018 | Challenges and Approaches for Data Collection to Understand Student Retention: (Abstract Only)abstractFor many years, computing faculty have devoted substantial time and energy to the retention of diverse populations. But how are we doing really? The ACM Retention Committee has identified at least 5 populations of interest in tracking student retention: * Students who start college expecting to major in computing. * Students who enter college with some interest in computing, but also with other interests. * Students who enter college with interests outside computing, but who take computing early as part of a broad education. * Students who enter college with little or no interest in computing, but need a computing course to satisfy a general education requirement or a prerequisite in another discipline. * Students who transfer into a four-year university from a two-year college, partway into a computer science program. In practice, each group has different characteristics, and retention rates may vary dramatically. On some campuses, gathering data for the first group may be manageable--particularly if students declare majors as they enter college. Data collection and tracking for others is difficult, since these populations may not be known in early years. This BoF will identify approaches for tracking students and for exploring retention rates. Further, this BoF will encourage sharing and brainstorming for further mechanisms to help data collection. As we better identify retention rates among various populations, the ACM Retention Committee hopes we can better understand obstacles and opportunities related to retention. Session Agenda: Context/Introduction, Data most relevant locally, What data are currently tracked, Thoughts about a common data gathering instrument Henry MacKay Walker, Mehran Sahami, Christine Alvarado |
SIGCSE | 2 |
| 2018 | TMOSS: Using Intermediate Assignment Work to Understand Excessive Collaboration in Large ClassesabstractAs computer science classes grow, instructor workload also increases: teachers must simultaneously teach material, provide assignment feedback, and monitor student progress. At scale, it is hard to know which students need extra help, and as a result some students can resort to excessive collaboration--using online resources or peer code--to complete their work. In this paper, we present TMOSS, a tool that analyzes the intermediate steps a student takes to complete a programming assignment. We find that for three separate course offerings, TMOSS is almost twice as effective as traditional software similarity detectors in identifying the number of students who exhibit excessive collaboration. We also find that such students spend significantly less time on their assignment, use fewer class tutoring resources, and perform worse on exams than their peers. Finally, we provide a theory of the parametric distribution of typical student assignment similarity, which allows for probabilistic interpretation. Lisa Yan, Nick McKeown, Mehran Sahami, Chris Piech |
SIGCSE | 3 |
| 2017 | Scaling Introductory Courses Using Undergraduate Teaching AssistantsabstractUndergraduates are widely used in support of Computer Science (CS) departments' teaching missions as teaching assistants, peer mentors, section leaders, course assistants, and tutors. Those undergraduates engaged in teaching have the opportunity to deeply engage with CS concepts and develop key communication and social competencies. As enrollments surge, undergraduate teaching assistants (UTAs) play a larger role in student experience and outcomes. While faculty and graduate student instructional support does not necessarily increase with the number of students in our courses, the number of qualified undergraduate teaching assistants for introductory CS courses should scale with the number of students in our courses. With large courses, the significance of the UTAs' role in students' learning likely also increases. Students have relatively little interaction with the instructor, and faculty may have more challenges monitoring and supporting individual UTAs. UTAs have a major role in affecting climate in computer science courses. The climate in large courses has substantial implications for students from groups traditionally underrepresented in computing. This panel will discuss how undergraduate teaching assistants can serve as a scalable effective teaching resource that benefits both the students in the course and the UTAs themselves. Jeffrey Forbes 0001, David J. Malan, Heather Pon-Barry, Stuart Reges, Mehran Sahami |
SIGCSE | 5 |
| 2016 | Statistical Modeling to Better Understand CS StudentsabstractWhile educational data mining has often focused on modeling behavior at the level of individual students, we consider developing statistical models to give us insight into the dynamics of student populations. In this talk, we consider two case studies in this vein. The first involves analyzing the evolution of gender balance in a college computer science program, showing that focusing on percentages of underrepresented groups in the overall population may not always provide an accurate portrayal of the impact of various program changes. We propose a new statistical model based on Fisher's Noncentral Hypergeometric Distribution that better captures how program changes are impacting the dynamics of gender balance in a population, especially in the case where the overall population is rapidly increasing (as has been the case in CS in recent years). Our second study looks at the performance of student populations in an introductory college programming course during the past eight years to better understand the evolving mix of students' abilities given the rapid growth in the number of students taking CS courses. Often accompanying such growth is a concern from faculty that the additional students choosing to pursue computing may not have the same aptitude for the subject as was seen in prior student populations. To directly address this question, we present a statistical analysis of students' performance using mixture modeling. Importantly, in this setting many variables that would normally confound such a study are directly controlled for. We find that the distribution of student performance during this period, as reflected in their programming assignment scores, remains remarkably stable despite the large growth in course enrollments. The results of this analysis also show how conflicting perceptions of students' abilities among faculty can be consistently explained. Mehran Sahami |
ITiCSE | 1 |
| 2016 | As CS Enrollments Grow, Are We Attracting Weaker Students?abstractIn recent years, enrollments in undergraduate computer science programs have seen tremendous growth nationally. Often accompanying such growth is a concern from faculty that the additional students choosing to pursue computing may not have the same aptitude for the subject as was seen in prior student populations. Thus such students may exhibit weaker performance in computing courses. To help address this question, we present a statistical analysis using mixture modeling of students' performance in an introductory programming class at Stanford University over an eight year period, during which enrollments in the course more than doubled. Importantly, in this setting many variables that would normally confound such a study are directly controlled for. We find that the distribution of student performance during this period, as reflected in their programming assignment scores, remains remarkably stable despite the large growth in enrollment. We then explain how the notion of having "more weak students" and the fact that the distribution of student ability is unchanged can readily co-exist and lead to misperceptions about the quality of incoming students during an enrollment boom. Mehran Sahami, Chris Piech |
SIGCSE | 1 |
| 2015 | Learning Program Embeddings to Propagate Feedback on Student CodeabstractProviding feedback, both assessing final work and giving hints to stuck students, is difficult for open-ended assignments in massive online classes which can range from thousands to millions of students. We introduce a neural network method to encode programs as a linear mapping from an embedded precondition space to an embedded postcondition space and propose an algorithm for feedback at scale using these linear maps as features. We apply our algorithm to assessments from the Code.org Hour of Code and Stanford University’s CS1 course, where we propagate human comments on student assignments to orders of magnitude more submissions. Chris Piech, Jonathan Huang, Andy Nguyen, Mike Phulsuksombati, Mehran Sahami, Leonidas J. Guibas |
ICML | 5 |
| 2015 | Autonomously Generating Hints by Inferring Problem Solving PoliciesabstractExploring the whole sequence of steps a student takes to produce work, and the patterns that emerge from thousands of such sequences is fertile ground for a richer understanding of learning. In this paper we autonomously generate hints for the Code.org `Hour of Code,' (which is to the best of our knowledge the largest online course to date) using historical student data. We first develop a family of algorithms that can predict the way an expert teacher would encourage a student to make forward progress. Such predictions can form the basis for effective hint generation systems. The algorithms are more accurate than current state-of-the-art methods at recreating expert suggestions, are easy to implement and scale well. We then show that the same framework which motivated the hint generating algorithms suggests a sequence-based statistic that can be measured for each learner. We discover that this statistic is highly predictive of a student's future success. Chris Piech, Mehran Sahami, Jonathan Huang, Leonidas J. Guibas |
L@S | 2 |
| 2015 | Deep Knowledge TracingabstractKnowledge tracing, where a machine models the knowledge of a student as they interact with coursework, is an established and significantly unsolved problem in computer supported education.In this paper we explore the benefit of using recurrent neural networks to model student learning.This family of models have important advantages over current state of the art methods in that they do not require the explicit encoding of human domain knowledge,and have a far more flexible functional form which can capture substantially more complex student interactions.We show that these neural networks outperform the current state of the art in prediction on real student data,while allowing straightforward interpretation and discovery of structure in the curriculum.These results suggest a promising new line of research for knowledge tracing. Chris Piech, Jonathan Bassen, Jonathan Huang, Surya Ganguli, Mehran Sahami, Leonidas J. Guibas, Jascha Sohl-Dickstein |
NIPS | 5 |
| 2014 | Panel: online learning platforms and data scienceabstractThe software platforms that mediate online learning experiences are the common ground where learning science and computer science intersect. This panel will discuss the affordances of current online learning platforms and lessons learned in using them with students. The goal of the panel is to help learning scientists and computer scientists understand each others' needs and how they might be effectively addressed in these platforms. The panelists, who have experience creating/using these platforms and interacting with learning scientists, will discuss how current platforms for learning at scale might evolve to better serve the community. Mehran Sahami, Jace Kohlmeier, Peter Norvig, Andreas Paepcke, Amin Saberi |
L@S | 1 |
| 2014 | Experiences mapping and revising curricula with CS2013abstractNo abstract available. David W. Reed, Andrea Pohoreckyj Danyluk, Elizabeth K. Hawthorne, Mehran Sahami, Henry MacKay Walker |
SIGCSE | 4 |
| 2014 | ACM/IEEE-CS computer science curricula 2013: implementing the final reportabstractFor over 40 years, the ACM and IEEE-Computer Society have sponsored international curricular guidelines for undergraduate programs in computing. The rapid evolution and expansion of the computing field and the growing number of topics in computer science have made regular revision of curricular recommendations necessary. Thus, the Computing Curricula volumes are updated on an approximately 10-year cycle, with the aim of keeping curricula modern and relevant. The latest volume in the series, Computer Science Curricula 2013 (CS2013), is due for release in the Fall of 2013. This panel seeks to inform the SIGCSE community about the final version of the report, provide insight on interpreting the CS2013 guidelines, and give guidance regarding how the guidelines may be implemented at different institutions. Mehran Sahami, Steve Roach, Ernesto Cuadros-Vargas, Elizabeth K. Hawthorne, Amruth N. Kumar, Richard LeBlanc, David W. Reed, Remzi Seker |
SIGCSE | 1 |
| 2013 | Special session: The CS2013 Computer Science curriculum guidelines projectabstractThe ACM/IEEE-Computer Society CS2013 Computer Science Curricula task force is working to update the previous curricular guidelines published in 2008 and 2001. The CS2013 guidelines are scheduled to be published in the latter half of 2013. This special session is devoted to exploring the guidelines with an emphasis on migrating current curricula to curricula aligned with the new guidelines. A number of significant changes from the 2008 and 2001 guidelines have been made, including the addition of new knowledge areas (including Parallel and Distributed Computing and Security and Information Assurance) as well as the reorganization and refactoring of previous areas to create a Systems Fundamentals area and a Software Development Fundamentals area. These changes are intended to identify significant changes in the computing field over the past decade, look forward to future changes, provide greater flexibility in the design and implementation of Computer Science curricula, provide stronger guidance with respect to student outcomes, and provide diverse examples of fielded curricula. Steve Roach, Mehran Sahami, Richard LeBlanc, Remzi Seker |
FIE | 2 |
| 2013 | A large-scale quantitative study of women in computer science at Stanford UniversityabstractIn this paper, we analyze gender dynamics in the undergraduate Computer Science program at Stanford University through a quantitative analysis of 7209 academic transcripts and 536 survey responses. We examine previously studied effects as well as present new findings. We also introduce Fisher's Noncentral Hypergeometric Distribution as a model for estimating the impact of program changes on underrepresented populations and explain why it is a more robust measure than changes in the percentage of minority participants. Katie Redmond, Sarah Evans, Mehran Sahami |
SIGCSE | 3 |
| 2013 | The revolution will be televised: perspectives on massive open online educationabstractNo abstract available. Mehran Sahami, Mark Guzdial, Fred G. Martin, Nick Parlante |
SIGCSE | 1 |
| 2013 | ACM/IEEE-CS computer science curriculum 2013: reviewing the ironman reportabstractFor over 40 years, the ACM and IEEE-Computer Society have sponsored the creation of international curricular guidelines for undergraduate programs in computing. These Computing Curricula volumes are updated approximately every 10-year cycle, with the aim of keeping curricula modern and relevant. The next volume in the series, Computer Science 2013 (CS2013), is currently in progress. This panel seeks to update and engage the SIGCSE community in providing feedback on a complete draft of the CS2013 report (called the Ironman report), which will be released shortly before SIGCSE. Since the Ironman report is the penultimate draft of the CS2013 report, this panel is an especially important venue for starting the last round of feedback that will impact the final CS2013 curricular guidelines. Mehran Sahami, Steve Roach, Ernesto Cuadros-Vargas, Richard LeBlanc |
SIGCSE | 1 |
| 2012 | Special session: The CS2013 Computer Science curriculum guidelines projectabstractWork began in late 2010 on a project to revise the ACM/IEEE-Computer Society Computer Science volume of Computing Curricula 2001 and the interim review CS 2008. The new guidelines for computer science are scheduled for release in 2013. This interactive session will give the computing education community an opportunity to review current working documents created by the CS2013 Steering Committee and thus influence the CS2013 volume scheduled for release in 2013. The Strawman draft of CS2013 was released in early 2012, containing new knowledge areas as well as significant updates to previous knowledge areas. The Steering Committee will present the Strawman draft and discuss changes to that draft that are being considered as a result of the first round of public comments. The majority of the session will be devoted to engaging attendees in guided discussions of sections for which additional community input is particularly relevant. Steve Roach, Mehran Sahami, Richard LeBlanc |
FIE | 2 |
| 2012 | Modeling how students learn to programabstractDespite the potential wealth of educational indicators expressed in a student's approach to homework assignments, how students arrive at their final solution is largely overlooked in university courses. In this paper we present a methodology which uses machine learning techniques to autonomously create a graphical model of how students in an introductory programming course progress through a homework assignment. We subsequently show that this model is predictive of which students will struggle with material presented later in the class. Chris Piech, Mehran Sahami, Daphne Koller, Steve Cooper, Paulo Blikstein |
SIGCSE | 2 |
| 2012 | Computer science curriculum 2013: reviewing the strawman report from the ACM/IEEE-CS task forceabstractBeginning over 40 years ago with the publication of Curriculum 68, the major professional societies in computing--ACM and IEEE-Computer Society--have sponsored various efforts to establish international curricular guidelines for undergraduate programs in computing. As the field has grown and diversified, so too have the recommendations for curricula. There are now guidelines for Computer Engineering, Information Systems, Information Technology, and Software Engineering in addition to Computer Science. These volumes are updated regularly with the aim of keeping computing curricula modern and relevant. In the Fall of 2010, work on the next volume in the series, Computer Science 2013 (CS2013), began. Considerable work on the new volume has already been completed and a first draft of the CS2013 report (known as the Strawman report) will be complete by the beginning of 2012. This panel seeks to update and engage the SIGCSE community in providing feedback on the Strawman report, which will be available shortly prior to the SIGCSE conference. Mehran Sahami, Steve Roach, Ernesto Cuadros-Vargas, David W. Reed |
SIGCSE | 1 |
| 2011 | Special session - The CS2013 computer science curriculum guidelines projectabstractWork began in late 2010 on a project to revise the ACM/IEEE-Computer Society Computer Science volume of Computing Curricula 2001 and the interim review CS 2008. The new guidelines for computer science are scheduled for release in 2013. This interactive session will give the computing education community an opportunity to review current working documents created by the CS2013 Steering Committee and thus influence a draft of the CS 2013 volume scheduled for release in December 2011. In early 2011, the Steering Committee began an extensive review of the Body of Knowledge and Characteristics of Graduates sections that are key parts of this series of curriculum volumes. New Knowledge Areas are being defined (Parallel and Distributed Computing, Security and Information Assurance, and Systems Fundamentals) and existing areas are being reorganized and updated. The Steering Committee will present a draft of the new Body of Knowledge as part of this session and will engage attendees in guided discussions of sections for which community input is particularly relevant. Similarly, a draft of the new Characteristics of Graduates description will be presented and discussed as part of the session. Mehran Sahami, Steve Roach, Richard LeBlanc |
FIE | 1 |
| 2011 | A course on probability theory for computer scientistsabstractDuring the past 20 years, probability theory has become a critical element in the development of many areas in computer science. Commensurately, in this paper, we argue for expanding the coverage of probability in the computing curriculum. Specifically, we present details of a new course we have developed on Probability Theory for Computer Scientists. An analysis of course evaluation data shows that students find the contextualized content of this class more relevant and valuable than general presentations of probability theory. We also discuss different models for expanding the role of probability in different curricular programs that may not have the capacity to teach a full course on the subject. Mehran Sahami |
SIGCSE | 1 |
| 2011 | Setting the stage for computing curricula 2013: computer science - report from the ACM/IEEE-CS joint task forceabstractFollowing a roughly 10 year cycle, the Computing Curricula volumes have helped to set international curricular guidelines for undergraduate programs in computing. In the summer of 2010, planning for the next volume in the series, Computer Science 2013, began. This panel seeks to update and engage the SIGCSE community on the Computer Science 2013 effort. The development of curricular guidelines in Computer Science is particularly challenging given the rapid evolution and expansion of the field. Moreover, the growing diversity of topics in Computer Science and the integration of computing with other disciplines create additional challenges and opportunities in defining computing curricula. As a result, it is particularly important to engage the broader computer science education community in a dialog to better understand new opportunities, local needs, and novel successful models of computing curriculum. The last complete Computer Science curricular volume was released in 2001 [3] and followed by a review effort that concluded in 2008 [2]. While the review helped to update some of the knowledge units in the 2001 volume, it was not aimed at producing an entirely new curricular volume and deferred some of the more significant questions that arose at the time. The Computer Science 2013 effort seeks to provide a new volume reflecting the current state of the field and highlighting promising future directions through revisiting and redefining the knowledge units in CS, rethinking the essentials necessary for a CS curriculum, and identifying working exemplars of courses and curricula along these lines. Mehran Sahami, Mark Guzdial, Andrew D. McGettrick, Steve Roach |
SIGCSE | 1 |
| 2011 | Educational advances in artificial intelligenceabstractIn 2010 a new annual symposium on Educational Advances in Artificial Intelligence (EAAI) was launched as part of the AAAI annual meeting. The event was held in cooperation with ACM SIGCSE and has many similar goals related to broadening and disseminating work in computer science education. EAAI has a particular focus, however, as the event is specific to educational work in Artificial Intelligence and collocated with a major research conference (AAAI) to promote more interaction between researchers and educators in that domain. This panel seeks to introduce participants to EAAI as a way of fostering more interaction between educational communities in computing. Specifically, the panel will discuss the goals of EAAI, provide an overview of the kinds of work presented at the symposium, and identify potential synergies between that EAAI and SIGCSE as a way of better linking the two communities going forward. Mehran Sahami, Marie desJardins, Zachary Dodds, Todd W. Neller |
SIGCSE | 1 |
| 2010 | Expanding the frontiers of computer science: designing a curriculum to reflect a diverse fieldabstractWhile the discipline of computing has evolved significantly in the past 30 years, Computer Science curricula have not as readily adapted to these changes. In response, we have recently completely redesigned the undergraduate CS curriculum at Stanford University, both modernizing the program as well as highlighting new directions in the field and its multi-disciplinary nature. As we explain in this paper, our restructured major features a streamlined core of foundation courses followed by a depth concentration in a track area as well as additional elective courses. Since its deployment this past year, the new program has proven to be very attractive to students, contributing to an increase of over 40% in the number of CS major declarations. We analyze feedback we received on the program from students, as well as commentary from industrial affiliates and other universities, providing further evidence of the promise this new curriculum holds. Mehran Sahami, Alex Aiken, Julie Zelenski |
SIGCSE | 1 |
| 2009 | Nifty assignmentsabstractAssignments determine much of what students actually take away from a course. Sadly, creating successful assignments is difficult and error prone. With that in mind, the Nifty Assignments session is about promoting and sharing successful assignment ideas, and more importantly, making the assignment materials available for others to adopt. Nick Parlante, Thomas P. Murtagh, Mehran Sahami, Owen L. Astrachan, David W. Reed, Christopher A. Stone, Brent Heeringa, Karen L. Reid |
SIGCSE | 3 |
| 2006 | Combinatorial Markov Random Fields
Ron Bekkerman, Mehran Sahami, Erik G. Learned-Miller |
ECML | 2 |
| 2006 | A web-based kernel function for measuring the similarity of short text snippetsabstractDetermining the similarity of short text snippets, such as search queries, works poorly with traditional document similarity measures (e.g., cosine), since there are often few, if any, terms in common between two short text snippets. We address this problem by introducing a novel method for measuring the similarity between short text snippets (even those without any overlapping terms) by leveraging web search results to provide greater context for the short texts. In this paper, we define such a similarity kernel function, mathematically analyze some of its properties, and provide examples of its efficacy. We also show the use of this kernel function in a large-scale system for suggesting related queries to search engine users. Mehran Sahami, Timothy D. Heilman |
WWW | 1 |
| 2005 | Adaptive Product Normalization: Using Online Learning for Record Linkage in Comparison ShoppingabstractThe problem of record linkage focuses on determining whether two object descriptions refer to the same underlying entity. Addressing this problem effectively has many practical applications, e.g., elimination of duplicate records in databases and citation matching for scholarly articles. In this paper, we consider a new domain where the record linkage problem is manifested: Internet comparison shopping. We address the resulting linkage setting that requires learning a similarity function between record pairs from streaming data. The learned similarity function is subsequently used in clustering to determine which records are co-referent and should be linked. We present an online machine learning method for addressing this problem, where a composite similarity function based on a linear combination of basis functions is learned incrementally. We illustrate the efficacy of this approach on several real-world datasets from an Internet comparison shopping site, and show that our method is able to effectively learn various distance functions for product data with differing characteristics. We also provide experimental results that show the importance of considering multiple performance measures in record linkage evaluation. Mikhail Bilenko, Sugato Basu, Mehran Sahami |
ICDM | 3 |
| 2005 | Evaluating similarity measures: a large-scale study in the orkut social networkabstractOnline information services have grown too large for users to navigate without the help of automated tools such as collaborative filtering, which makes recommendations to users based on their collective past behavior. While many similarity measures have been proposed and individually evaluated, they have not been evaluated relative to each other in a large real-world environment. We present an extensive empirical comparison of six distinct measures of similarity for recommending online communities to members of the Orkut social network. We determine the usefulness of the different recommendations by actually measuring users' propensity to visit and join recommended communities. We also examine how the ordering of recommendations influenced user selection, as well as interesting social issues that arise in recommending communities within a real social network. Ellen Spertus, Mehran Sahami, Orkut Buyukkokten |
KDD | 2 |
| 2004 | Efficient face orientation discriminationabstractThe paper presents efficient methods to address the problem of discriminating between live facial orientations. We present the most efficient methods for this task to date, which can accurately discriminate between five facial orientations with approximately 92% accuracy using fewer than 30 pixel comparisons and greater than 99% accuracy using 150 pixel comparisons. We achieve these rates by using a boosting method to select from a large set of extremely simple features. Comparisons to other methods are given. Shumeet Baluja, Mehran Sahami, Henry A. Rowley |
ICIP | 2 |
| 2004 | The Happy Searcher: Challenges in Web Information Retrieval
Mehran Sahami, Vibhu O. Mittal, Shumeet Baluja, Henry A. Rowley |
PRICAI | 1 |
| 2003 | QProber: A system for automatic classification of hidden-Web databasesabstractThe contents of many valuable Web-accessible databases are only available through search interfaces and are hence invisible to traditional Web "crawlers." Recently, commercial Web sites have started to manually organize Web-accessible databases into Yahoo!-like hierarchical classification schemes. Here we introduce QProber, a modular system that automates this classification process by using a small number of query probes, generated by document classifiers. QProber can use a variety of types of classifiers to generate the probes. To classify a database, QProber does not retrieve or inspect any documents or pages from the database, but rather just exploits the number of matches that each query probe generates at the database in question. We have conducted an extensive experimental evaluation of QProber over collections of real documents, experimenting with different types of document classifiers and retrieval models. We have also tested our system with over one hundred Web-accessible databases. Our experiments show that our system has low overhead and achieves high classification accuracy across a variety of databases. Luis Gravano, Panagiotis G. Ipeirotis, Mehran Sahami |
ACM Trans. Inf. Syst. | 3 |
| 2001 | Probe, Count, and Classify: Categorizing Hidden Web DatabasesabstractThe contents of many valuable web-accessible databases are only accessible through search interfaces and are hence invisible to traditional web “crawlers.” Recent studies have estimated the size of this “hidden web” to be 500 billion pages, while the size of the “crawlable” web is only an estimated two billion pages. Recently, commercial web sites have started to manually organize web-accessible databases into Yahoo!-like hierarchical classification schemes. In this paper, we introduce a method for automating this classification process by using a small number of query probes. To classify a database, our algorithm does not retrieve or inspect any documents or pages from the database, but rather just exploits the number of matches that each query probe generates at the database in question. We have conducted an extensive experimental evaluation of our technique over collections of real documents, including over one hundred web-accessible databases. Our experiments show that our system has low overhead and achieves high classification accuracy across a variety of databases. Panagiotis G. Ipeirotis, Luis Gravano, Mehran Sahami |
SIGMOD Conference | 3 |
| 1998 | Inductive Learning Algorithms and Representations for Text CategorizationabstractText categorization – the assignment of natural language texts to one or more predefined categories based on their content – is an important component in many information organization and management tasks. We compare the effectiveness of five different automatic learning algorithms for text categorization in terms of learning speed, realtime classification speed, and classification accuracy. We also examine training set size, and alternative document representations. Very accurate text classifiers can be learned automatically from training examples. Linear Support Vector Machines (SVMs) are particularly promising because they are very accurate, quick to train, and quick to evaluate. 1.1 Keywords Text categorization, classification, support vector machines, machine learning, information management. Susan T. Dumais, John C. Platt, David Hecherman, Mehran Sahami |
CIKM | 4 |
| 1997 | Hierarchically Classifying Documents Using Very Few Words
Daphne Koller, Mehran Sahami |
ICML | 2 |
| 1996 | Toward Optimal Feature Selection
Daphne Koller, Mehran Sahami |
ICML | 2 |
| 1996 | Applying the Multiple Cause Mixture Model to Text Categorization
Mehran Sahami, Marti A. Hearst, Eric Saund |
ICML | 1 |
| 1996 | Error-Based and Entropy-Based Discretization of Continuous Features
Ron Kohavi, Mehran Sahami |
KDD | 2 |
| 1996 | Learning Limited Dependence Bayesian Classifiers
Mehran Sahami |
KDD | 1 |
| 1995 | Generating Neural Networks Through the Induction of Threshold Logic Unit Trees (Extended Abstract)
Mehran Sahami |
ECML | 1 |
| 1995 | Learning Classification Rules Using Lattices (Extended Abstract)
Mehran Sahami |
ECML | 1 |
| 1995 | Supervised and Unsupervised Discretization of Continuous Features
James Dougherty, Ron Kohavi, Mehran Sahami |
ICML | 3 |
| 1993 | Learning Non-Linearly Separable Boolean Functions With Linear Threshold Unit Trees and Madaline-Style Networks
Mehran Sahami |
AAAI | 1 |