Jérémie O. Lumbroso

dblp:06/7357 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-5563-687XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Theory of computation · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Bit-array-based alternatives to HyperLogLog
abstract
We present a family of algorithms for the problem of estimating the number of distinct items in an input stream that are simple to implement and are appropriate for practical applications. Our algorithms are a logical extension of the series of algorithms developed by Flajolet and his coauthors starting in 1983 that culminated in the widely used HyperLogLog algorithm. These algorithms divide the input stream into M substreams and lead to a time-accuracy tradeoff where a small number of bits per substream are saved to achieve a relative accuracy proportional to 1 / M . Our algorithms use just one or two bits per substream. Their effectiveness is demonstrated by a proof of approximate normality, with explicit expressions for standard errors that inform parameter settings and allow proper quantitative comparisons with other methods. Performance hypotheses are validated through experiments using a realistic input stream, with the general conclusion that our algorithms are significantly more accurate than HyperLogLog when using the same amount of memory, and they use significantly less memory than HyperLogLog to achieve a given accuracy. • Efficient algorithms for estimating the number of distinct items in a data stream. • Explicit characterization of the distribution of reported values. • Detailed fair comparisons with other algorithms in the literature. • Sufficient detail to enable development of real-world implementations.
Svante Janson, Jérémie O. Lumbroso, Robert Sedgewick
Theor. Comput. Sci.2
2024 Bit-Array-Based Alternatives to HyperLogLog
abstract
We present a family of algorithms for the problem of estimating the number of distinct items in an input stream that are simple to implement and are appropriate for practical applications. Our algorithms are a logical extension of the series of algorithms developed by Flajolet and his coauthors starting in 1983 that culminated in the widely used HyperLogLog algorithm. These algorithms divide the input stream into M substreams and lead to a time-accuracy tradeoff where a constant number of bits per substream are saved to achieve a relative accuracy proportional to 1/√M. Our algorithms use just one or two bits per substream. Their effectiveness is demonstrated by a proof of approximate normality, with explicit expressions for standard errors that inform parameter settings and allow proper quantitative comparisons with other methods. Hypotheses about performance are validated through experiments using a realistic input stream, with the conclusion that our algorithms are more accurate than HyperLogLog when using the same amount of memory, and they use two-thirds as much memory as HyperLogLog to achieve a given accuracy.
Svante Janson, Jérémie O. Lumbroso, Robert Sedgewick
AofA2
2022 Affirmative Sampling: Theory and Applications
Jérémie O. Lumbroso, Conrado Martínez
AofA1
2022 Tools and Techniques for Increasing and Measuring Student Engagement With Pre-Recorded Videos
Ananda Gunawardena, Jérémie O. Lumbroso
SIGCSE (2)2
2020 Reviewing Computing Education Papers
abstract
Peer review is a mainstay of academic publication - indeed, it is the peer-review process that provides much of the publications' credibility. This working group is examining the ways peer review is used in various computing education venues and will use this examination to articulate community standards for peer review in this discipline.
Marian Petre, Kate Sanders 0001, Robert McCartney, Marzieh Ahmadzadeh, Cornelia Connolly, Sally Hamouda, Brian Harrington 0001, Jérémie O. Lumbroso, Joseph Maguire 0001, Lauri Malmi, Monica McGill, Jan Vahrenhold
ITiCSE8
2020 Making Manual Code Review Scale
abstract
Although personalized feedback is believed to improve student learning more than correctness tests alone, resource constraints make it difficult for many CS programs to provide personalized feedback on student programming work. This workshop discusses strategies for reducing the cost of code review, as well as pedagogical benefits that can be reaped once a course implements a robust code review process. We showcase the code review process implemented by Princeton's CS1 and CS2 courses which leverages codePost (https://codepost.io), and will teach participants to implement a similar process.
Jérémie O. Lumbroso
SIGCSE1
2017 Nailing the TA Interview: Using a Rubric to Hire Teaching Assistants
abstract
Where would we be without them? Teaching assistants (TAs) make it possible for us to deliver high-quality large-scale computer science courses with relatively few faculty. Though their responsibilities vary by institution, TAs often play a crucial role in student learning. The use of teaching assistants in computer science courses is a common and longstanding practice and, yet, little has been published about how to choose the best TAs among those interested in the job. This paper describes the development of an interview rubric in use by faculty teaching a large introductory computer science course to score applicant responses in a formal in-person 30-minute interview. We describe the motivation behind developing such a rubric, the initial development process, its refinement based on feedback provided by students about their TAs, and the preliminary results of implementing this hiring system.
Dan Leyzberg, Jérémie O. Lumbroso, Christopher Moretti
ITiCSE2