VLDB 2026 Research / reviewers in the wild / expert
Michael Roberts
dblp:31/1197
· DBLP profile ↗
10ranked-venue papers
3as first author
2since 2021 · last 2022
0000-0002-1441-7363ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorComputer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Programming languages and type systems · 100% | |
| Human-computer interaction and pervasive computing
2 papers |
Interaction techniques and input · 49% User interface design and tools · 49% Ubiquitous computing and smart environments · 2% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 100% | |
| Theoretical computer science
1 paper |
Quantum computing and quantum information · 100% |
Topics — the 10 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Programming languages and type systems › language semantics › formal semantics
denotational semantics |
0.5 | 1 | 2021 | Universal Semantics for the Stochastic λ-Calculus · LICS 2021 |
Programming languages and type systems › language semantics › formal semantics
operational semantics |
0.5 | 1 | 2021 | Universal Semantics for the Stochastic λ-Calculus · LICS 2021 |
Programming languages and type systems › probabilistic programming
stochastic lambda calculus |
0.5 | 1 | 2021 | Universal Semantics for the Stochastic λ-Calculus · LICS 2021 |
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome assembly |
0.2 | 2 | 2013 | The MaSuRCA genome assembler · Bioinform. 2013 Figaro: a novel statistical method for vector sequence removal · Bioinform. 2008 |
Bioinformatics and computational biology › sequence analysis › sequence assembly › genome assembly › de novo assembly
de bruijn graph assembly |
0.2 | 1 | 2013 | The MaSuRCA genome assembler · Bioinform. 2013 |
Bioinformatics and computational biology › sequence analysis › sequence assembly › genome assembly
de novo assembly |
0.2 | 1 | 2013 | The MaSuRCA genome assembler · Bioinform. 2013 |
Bioinformatics and computational biology › sequence analysis › sequence assembly › genome assembly
hybrid assembly |
0.2 | 1 | 2013 | The MaSuRCA genome assembler · Bioinform. 2013 |
Bioinformatics and computational biology
sequence analysis |
0.1 | 1 | 2008 | Figaro: a novel statistical method for vector sequence removal · Bioinform. 2008 |
Recommender systems
context-aware recommendation |
0.1 | 1 | 2008 | Activity-based serendipitous recommendations with the Magitti mobile leisure guide · CHI 2008 |
Bioinformatics and computational biology › sequence analysis
sequence comparison |
0.0 | 1 | 2004 | Reducing storage requirements for biological sequence comparison · Bioinform. 2004 |
Methods — techniques the papers use, named apart from their topics
pen-based interaction · 1.1handwriting recognition · 1.1operational semantics · 0.5denotational semantics · 0.5overlap-based assembly · 0.2de bruijn graph · 0.2prototype evaluation · 0.2field study · 0.2poisson statistics · 0.1oligonucleotide frequency analysis · 0.1minimizer · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Notational Programming for Notebook Environments: A Case Study with Quantum CircuitsabstractWe articulate a vision for computer programming that includes pen-based computing, a paradigm we term notational programming. Notational programming blurs contexts: certain typewritten variables can be referenced in handwritten notation and vice-versa. To illustrate this paradigm, we developed an extension, Notate, to computational notebooks which allows users to open drawing canvases within lines of code. As a case study, we explore quantum programming and designed a notation, Qaw, that extends quantum circuit notation with abstraction features, such as variable-sized wire bundles and recursion. Results from a usability study with novices suggest that users find our core interaction of implicit cross-context references intuitive, but suggests further improvements to debugging infrastructure, interface design, and recognition rates. Throughout, we discuss questions raised by the notational paradigm, including a shift from ‘recognition’ of notations to ‘reconfiguration’ of practices and values around programming, and from ‘sketching’ to writing and drawing, or what we call ‘notating.’ Ian Arawjo, Anthony DeArmas, Michael Roberts, Shrutarshi Basu, Tapan S. Parikh |
UIST | 3 |
| 2021 | Universal Semantics for the Stochastic λ-CalculusabstractWe define sound and adequate denotational and operational semantics for the stochastic lambda calculus. These two semantic approaches build on previous work that used an explicit source of randomness to reason about higher-order probabilistic programs. Pedro H. Azevedo de Amorim, Dexter Kozen, Radu Mardare, Prakash Panangaden, Michael Roberts |
LICS | 5 |
| 2015 | Unsupervised Modeling of Users' Interests from their Facebook Profiles and ActivitiesabstractUser interest profiles have become essential for personalizing information streams and services, and user interfaces and experiences. In today's world, social networks such as Facebook or Twitter provide users with a powerful platform for interest expression and can, thus, act as a rich content source for automated user interest modeling. This, however, poses significant challenges because the user generated content on them consists of free unstructured text. In addition, users may not explicitly post or tweet about everything that interests them. Moreover, their interests evolve over time. In this paper, we propose a novel unsupervised algorithm and system that addresses these challenges. It models a broad range of an individual user's explicit and implicit interests from her social network profile and activities without any user input. We perform extensive evaluation of our system, and algorithm, with a dataset consisting of 488 active Facebook users' profiles and demonstrate that it can accurately estimate a user's interests in practice. Preeti Bhargava, Oliver Brdiczka, Michael Roberts |
IUI | 3 |
| 2013 | The MaSuRCA genome assemblerabstractMOTIVATION: Second-generation sequencing technologies produce high coverage of the genome by short reads at a low cost, which has prompted development of new assembly methods. In particular, multiple algorithms based on de Bruijn graphs have been shown to be effective for the assembly problem. In this article, we describe a new hybrid approach that has the computational efficiency of de Bruijn graph methods and the flexibility of overlap-based assembly strategies, and which allows variable read lengths while tolerating a significant level of sequencing error. Our method transforms large numbers of paired-end reads into a much smaller number of longer 'super-reads'. The use of super-reads allows us to assemble combinations of Illumina reads of differing lengths together with longer reads from 454 and Sanger sequencing technologies, making it one of the few assemblers capable of handling such mixtures. We call our system the Maryland Super-Read Celera Assembler (abbreviated MaSuRCA and pronounced 'mazurka'). RESULTS: We evaluate the performance of MaSuRCA against two of the most widely used assemblers for Illumina data, Allpaths-LG and SOAPdenovo2, on two datasets from organisms for which high-quality assemblies are available: the bacterium Rhodobacter sphaeroides and chromosome 16 of the mouse genome. We show that MaSuRCA performs on par or better than Allpaths-LG and significantly better than SOAPdenovo on these data, when evaluated against the finished sequence. We then show that MaSuRCA can significantly improve its assemblies when the original data are augmented with long reads. AVAILABILITY: MaSuRCA is available as open-source code at ftp://ftp.genome.umd.edu/pub/MaSuRCA/. Previous (pre-publication) releases have been publicly available for over a year. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Aleksey V. Zimin, Guillaume Marçais, Daniela Puiu, Michael Roberts, Steven Salzberg, James A. Yorke |
Bioinform. | 4 |
| 2010 | The "3D Wiki": Blending virtual worlds and Web architecture for remote collaborationabstractWhile a lot of technical data is available on the Web, conveying information about detailed procedures for the assembly and repair of complex machinery has so far been limited mostly to 2D drawings and textual content. In this paper we describe the technology behind our 3D Wiki, a system meant to address this problem. By blending the functionality of virtual worlds (3D visualization and navigation) with the benefits of social software (editability, traceability, and hyperlinking), the 3D Wiki can help users tasked with the remote maintenance and repair of technical equipment. Our environment provides a kind of “living technical manual”, seamlessly blending spatial and textual content. Michael Roberts, Nicolas Ducheneaut, Trevor F. Smith |
ICME | 1 |
| 2009 | Collaborative Filtering Is Not Enough? Experiments with a Mixed-Model Recommender for Leisure Activities
Nicolas Ducheneaut, Kurt Partridge, Qingfeng Huang, Bob Price, Michael Roberts, Ed H. Chi, Victoria Bellotti, James Begole |
UMAP | 5 |
| 2008 | Activity-based serendipitous recommendations with the Magitti mobile leisure guideabstractThis paper presents a context-aware mobile recommender system, codenamed Magitti. Magitti is unique in that it infers user activity from context and patterns of user behavior and, without its user having to issue a query, automatically generates recommendations for content matching. Extensive field studies of leisure time practices in an urban setting (Tokyo) motivated the idea, shaped the details of its design and provided data describing typical behavior patterns. The paper describes the fieldwork, user interface, system components and functionality, and an evaluation of the Magitti prototype. Victoria Bellotti, James Begole, Ed H. Chi, Nicolas Ducheneaut, Ji Fang, Ellen Isaacs, Tracy Holloway King, Mark W. Newman, Kurt Partridge, Bob Price, Paul Rasmussen, Michael Roberts, Diane J. Schiano, Alan Walendowski |
CHI | 12 |
| 2008 | Scalable architecture for context-aware activity-detecting mobile recommendation systemsabstractOne of the main challenges in building multi-user mobile information systems for real-world deployment lies in the development of scalable systems. Recent work on scaling infrastructure for conventional web services using distributed approaches can be applied to the mobile space, but limitations inherent to mobile devices (computational power, battery life) and their communication infrastructure (availability and quality of network connectivity) challenge system designers to carefully design and optimize their software architectures. Additionally, notions of mobility and position in space, unique to mobile systems, provide interesting directions for the segmentation and scalability of mobile information systems. In this paper we describe the implementation of a mobile recommender system for leisure activities, codenamed Magitti, which was built for commercial deployment under stringent scalability requirements. We present concrete solutions addressing these scalability challenges, with the goal of informing the design of future mobile multi-user systems. Michael Roberts, Nicolas Ducheneaut, James Begole, Kurt Partridge, Bob Price, Victoria Bellotti, Alan Walendowski, Paul Rasmussen |
WOWMOM | 1 |
| 2008 | Figaro: a novel statistical method for vector sequence removalabstractMOTIVATION: Sequences produced by automated Sanger sequencing machines frequently contain fragments of the cloning vector on their ends. Software tools currently available for identifying and removing the vector sequence require knowledge of the vector sequence, specific splice sites and any adapter sequences used in the experiment-information often omitted from public databases. Furthermore, the clipping coordinates themselves are missing or incorrectly reported. As an example, within the approximately 1.24 billion shotgun sequences deposited in the NCBI Trace Archive, as many as approximately 735 million (approximately 60%) lack vector clipping information. Correct clipping information is essential to scientists attempting to validate, improve and even finish the increasingly large number of genomes released at a 'draft' quality level. RESULTS: We present here Figaro, a novel software tool for identifying and removing the vector from raw sequence data without prior knowledge of the vector sequence. The vector sequence is automatically inferred by analyzing the frequency of occurrence of short oligo-nucleotides using Poisson statistics. We show that Figaro achieves 99.98% sensitivity when tested on approximately 1.5 million shotgun reads from Drosophila pseudoobscura. We further explore the impact of accurate vector trimming on the quality of whole-genome assemblies by re-assembling two bacterial genomes from shotgun sequences deposited in the Trace Archive. Designed as a module in large computational pipelines, Figaro is fast, lightweight and flexible. AVAILABILITY: Figaro is released under an open-source license through the AMOS package (http://amos.sourceforge.net/Figaro). James Robert White, Michael Roberts, James A. Yorke, Mihai Pop |
Bioinform. | 2 |
| 2004 | Reducing storage requirements for biological sequence comparisonabstractMOTIVATION: Comparison of nucleic acid and protein sequences is a fundamental tool of modern bioinformatics. A dominant method of such string matching is the 'seed-and-extend' approach, in which occurrences of short subsequences called 'seeds' are used to search for potentially longer matches in a large database of sequences. Each such potential match is then checked to see if it extends beyond the seed. To be effective, the seed-and-extend approach needs to catalogue seeds from virtually every substring in the database of search strings. Projects such as mammalian genome assemblies and large-scale protein matching, however, have such large sequence databases that the resulting list of seeds cannot be stored in RAM on a single computer. This significantly slows the matching process. RESULTS: We present a simple and elegant method in which only a small fraction of seeds, called 'minimizers', needs to be stored. Using minimizers can speed up string-matching computations by a large factor while missing only a small fraction of the matches found using all seeds. Michael Roberts, Wayne B. Hayes, Brian R. Hunt, Stephen M. Mount, James A. Yorke |
Bioinform. | 1 |