VLDB 2026 Research / reviewers in the wild / expert
Michael Shilman
dblp:68/2885
· DBLP profile ↗
14ranked-venue papers
5as first author
0since 2021 · last 2008
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 3 first-authorArtificial intelligence and machine learning · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 5 · 1 first-authorSystems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Human-computer interaction and pervasive computing
3 papers |
Interaction techniques and input · 75% Design research and methods · 16% User interface design and tools · 5% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% | |
| Artificial intelligence
1 paper |
Language models and text generation · 50% Information extraction and text analysis · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Embedded and real-time systems · 100% |
Topics — the 11 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Interaction techniques and input
pen input |
0.2 | 3 | 2007 | SketchWizard: Wizard of Oz prototyping of pen-based user interfaces · UIST 2007 InkSeine: In Situ search for active note taking · CHI 2007 CueTIP: a mixed-initiative interface for correcting handwriting errors · UIST 2006 |
Design research and methods › research methodology
wizard of oz prototyping |
0.1 | 1 | 2007 | SketchWizard: Wizard of Oz prototyping of pen-based user interfaces · UIST 2007 |
Natural language and speech › Language models and text generation › decoding
sequence decoding |
0.1 | 1 | 2006 | Online decoding of Markov models under latency constraints · ICML 2006 |
Interaction techniques and input › text entry
error correction |
0.1 | 1 | 2006 | CueTIP: a mixed-initiative interface for correcting handwriting errors · UIST 2006 |
Interaction techniques and input › pen input
handwriting recognition |
0.1 | 1 | 2006 | CueTIP: a mixed-initiative interface for correcting handwriting errors · UIST 2006 |
Image and video processing
document image analysis |
0.1 | 1 | 2005 | Learning Non-Generative Grammatical Models for Document Analysis · ICCV 2005 |
Image and video processing › document image analysis
document layout analysis |
0.1 | 1 | 2005 | Learning Non-Generative Grammatical Models for Document Analysis · ICCV 2005 |
Image and video processing › document image analysis › graphics recognition
mathematical expression recognition |
0.1 | 1 | 2005 | Learning Non-Generative Grammatical Models for Document Analysis · ICCV 2005 |
Programming languages and type systems
language design |
0.0 | 1 | 1998 | Design and Specification of Embedded Systems in Java Using Successive, Formal Refinement · DAC 1998 |
Embedded and real-time systems
embedded system design |
0.0 | 1 | 1998 | Design and Specification of Embedded Systems in Java Using Successive, Formal Refinement · DAC 1998 |
Human-AI interaction
mixed-initiative interaction |
0.0 | 1 | 2006 | CueTIP: a mixed-initiative interface for correcting handwriting errors · UIST 2006 |
Methods — techniques the papers use, named apart from their topics
wizard-of-oz · 0.1user study · 0.1ink-based query formulation · 0.1window-based decoding · 0.1viterbi algorithm · 0.1grammatical parsing · 0.1discriminative learning · 0.1program transformation · 0.0formal refinement · 0.0class-library extension · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2008 | Aggregate documents: making sense of a patchwork of topical documentsabstractWith the dramatic increase in quantity and diversity of online content, particularly in the form of user generated content, we now have access to unprecedented amounts of information. Whether you are researching the purchase of a new cell phone, planning a vacation, or trying to assess a political candidate, there are now countless resources at your fingertips. However, finding and making sense of all this information is laborious and it is difficult to assess high-level trends in what is said. Web sites like Wikipedia and Digg democratize the process of organizing the information from countless document into a single source where it is somewhat easier to understand what is important and interesting. In this talk, I describe a complementary set of automated alternatives to these approaches, demonstrate these approaches with a working example, the commercial web site Wize.com, and derive some basic principles for aggregating a diverse set of documents into a coherent and useful summary. Michael Shilman |
ACM Symposium on Document Engineering | 1 |
| 2007 | InkSeine: In Situ search for active note takingabstractUsing a notebook to sketch designs, reflect on a topic, or capture and extend creative ideas are examples of active note taking tasks. Optimal experience for such tasks demands concentration without interruption. Yet active note taking may also require reference documents or emails from team members. InkSeine is a Tablet PC application that supports active note taking by coupling a pen-and-ink interface with an in situ search facility that flows directly from a user's ink notes (Fig. 1). InkSeine integrates four key concepts: it leverages preexisting ink to initiate a search; it provides tight coupling of search queries with application content; it persists search queries as first class objects that can be commingled with ink notes; and it enables a quick and flexible workflow where the user may freely interleave inking, searching, and gathering content. InkSeine offers these capabilities in an interface that is tailored to the unique demands of pen input, and that maintains the primacy of inking above all other tasks. Ken Hinckley, Shengdong Zhao 0001, Raman Sarin, Patrick Baudisch, Edward Cutrell, Michael Shilman, Desney S. Tan |
CHI | 6 |
| 2007 | SketchWizard: Wizard of Oz prototyping of pen-based user interfacesabstractSketchWizard allows designers to create Wizard of Oz prototypes of pen-based user interfaces in the early stages of design. In the past, designers have been inhibited from participating in the design of pen-based interfaces because of the inadequacy of paper prototypes and the difficulty of developing functional prototypes. In SketchWizard, designers and end users share a drawing canvas between two computers, allowing the designer to simulate the behavior of recognition or other technologies. Special editing features are provided to help designers respond quickly to end-user input. This paper describes the SketchWizard system and presents two evaluations of our approach. The first is an early feasibility study in which Wizard of Oz was used to prototype a pen-based user interface. The second is a laboratory study in which designers used SketchWizard to simulate existing pen-based interfaces. Both showed that end users gave valuable feedback in spite of delays between end-user actions and wizard updates. Richard C. Davis, T. Scott Saponas, Michael Shilman, James A. Landay |
UIST | 3 |
| 2006 | Combining Multiple Classifiers for Faster Optical Character Recognition
Kumar Chellapilla, Michael Shilman, Patrice Y. Simard |
Document Analysis Systems | 2 |
| 2006 | Online decoding of Markov models under latency constraintsabstractThe Viterbi algorithm is an efficient and optimal method for decoding linear-chain Markov Models. However, the entire input sequence must be observed before the labels for any time step can be generated, and therefore Viterbi cannot be directly applied to online/interactive/streaming scenarios without incurring significant (possibly unbounded) latency. A widely used approach is to break the input stream into fixed-size windows, and apply Viterbi to each window. Larger windows lead to higher accuracy, but result in higher latency.We propose several alternative algorithms to the fixed-sized window decoding approach. These approaches compute a certainty measure on predicted labels that allows us to trade off latency for expected accuracy dynamically, without having to choose a fixed window size up front. Not surprisingly, this more principled approach gives us a substantial improvement over choosing a fixed window. We show the effectiveness of the approach for the task of spotting semi-structured information in large documents. When compared to full Viterbi, the approach suffers a 0.1 percent error degradation with a average latency of 2.6 time steps (versus the potentially infinite latency of Viterbi). When compared to fixed windows Viterbi, we achieve a 40x reduction in error and 6x reduction in latency. Mukund Narasimhan, Paul A. Viola, Michael Shilman |
ICML | 3 |
| 2006 | CueTIP: a mixed-initiative interface for correcting handwriting errorsabstractWith advances in pen-based computing devices, handwriting has become an increasingly popular input modality. Researchers have put considerable effort into building intelligent recognition systems that can translate handwriting to text with increasing accuracy. However, handwritten input is inherently ambiguous, and these systems will always make errors. Unfortunately, work on error recovery mechanisms has mainly focused on interface innovations that allow users to manually transform the erroneous recognition result into the intended one. In our work, we propose a mixed-initiative approach to error correction. We describe CueTIP, a novel correction interface that takes advantage of the recognizer to continually evolve its results using the additional information from user corrections. This significantly reduces the number of actions required to reach the intended result. We present a user study showing that CueTIP is more efficient and better preferred for correcting handwriting recognition errors. Grounded in the discussion of CueTIP, we also present design principles that may be applied to mixed-initiative correction interfaces in other domains. Michael Shilman, Desney S. Tan, Patrice Y. Simard |
UIST | 1 |
| 2005 | Learning Non-Generative Grammatical Models for Document AnalysisabstractWe present a general approach for the hierarchical segmentation and labeling of document layout structures. This approach models document layout as a grammar and performs a global search for the optimal parse based on a grammatical cost function. Our contribution is to utilize machine learning to discriminatively select features and set all parameters in the parsing process. Therefore, and unlike many other approaches for layout analysis, ours can easily adapt itself to a variety of document analysis problems. One need only specify the page grammar and provide a set of correctly labeled pages. We apply this technique to two document image analysis tasks: page layout structure extraction and mathematical expression interpretation. Experiments demonstrate that the learned grammars can be used to extract the document structure in 57 files from the UWIII document image database. We also show that the same framework can be used to automatically interpret printed mathematical expressions so as to recreate the original LaTeX Michael Shilman, Percy Liang, Paul A. Viola |
ICCV | 1 |
| 2005 | Efficient Geometric Algorithms for Parsing in Two DimensionsabstractGrammars are a powerful technique for modeling and extracting the structure of documents. One large challenge, however, is computational complexity. The computational cost of grammatical parsing is related to both the complexity of the input and the ambiguity of the grammar. For programming languages, where the terminals appear in a linear sequence and the grammar is unambiguous, parsing is O(N). For natural languages, which are linear yet have an ambiguous grammar, parsing is O(N/sup 3/). For documents, where the terminals are arranged in two dimensions and the grammar is ambiguous, parsing time can be exponential in the number of terminals. In this paper we introduce (and unify) several types of geometrical data structures which can be used to significantly accelerate parsing time. Each data structure embodies a different geometrical constraint on the set of possible valid parses. These data structures are very general, in that they can be used by any type of grammatical model, and a wide variety of document understanding tasks, to limit the set of hypotheses examined and tested. Assuming a clean design for the parsing software, the same parsing framework can be tested with various geometric constraints to determine the most effective combination. Percy Liang, Mukund Narasimhan, Michael Shilman, Paul A. Viola |
ICDAR | 3 |
| 2005 | Grouping Text Lines in Freeform Handwritten NotesabstractHandwritten text lines are prominent structures in freeform digital ink notes and their reliable detection is the foundation to a natural and intelligent interface for note editing and repurposing. This paper presents an optimization method for text line grouping. The global'cost function is designed to find the simplest stroke partitioning to maximize the likelihood of the resulting lines and the consistency of their configuration. A dynamic programming algorithm provides an initial segmentation of the time-ordered stroke sequence. Then a local gradient-descent algorithm iteratively evaluates splitting and merging hypotheses to minimize the global cost function. On average, the proposed technique processes each note page in less than a second at 90% accuracy. Herry Sutanto, Sashi Raghupathy, Michael Shilman |
ICDAR | 5 |
| 2005 | DIZI: A Digital Ink Zooming Interface for Document Annotation
Maneesh Agrawala, Michael Shilman |
INTERACT | 2 |
| 2004 | Recognizing Freeform Digital Ink Annotations
Michael Shilman, Zile Wei |
Document Analysis Systems | 1 |
| 2004 | Robust sketched symbol fragmentation using templatesabstractAnalysis of sketched digital ink is often aided by the division of stroke points into perceptually-salient fragments based on geometric features. Fragmentation has many applications in intelligent interfaces for digital ink capture and manipulation, as well as higher-level symbolic and structural analyses. It is our intuitive belief that the most robust fragmentations closely match a user's natural perception of the ink, thus leading to more effective recognition and useful user feedback. We present two optimal fragmentation algorithms that fragment common geometries into a basis set of line segments and elliptical arcs. The first algorithm uses an explicit template in which the order and types of bases are specified. The other only requires the number of fragments of each basis type. For the set of symbols under test, both algorithms achieved 100% fragmentation accuracy rate for symbols with line bases, ›99% accuracy for symbols with elliptical bases, and ›90% accuracy for symbols with mixed line and elliptical bases. Heloise Hwawen Hse, Michael Shilman, A. Richard Newton |
IUI | 2 |
| 2003 | Discerning Structure from Freeform Handwritten NotesabstractThis paper presents an integrated approach to parsing textual structure in freeform handwritten notes. Text-graphics classification and text layout analysis are classical problems in printed document analysis, but the irregularity in handwriting and content in freeform notes reveals limitations in existing approaches. We advocate an integrated technique that solves the layout analysis and classification problems simultaneously: the problems are so tightly coupled that it is not possible to solve one without the other for real user notes. We tune and evaluate our approach on a large corpus of unscripted user files and reflect on the difficult recognition scenarios that we have encountered in practice. Michael Shilman, Zile Wei, Sashi Raghupathy, Patrice Y. Simard |
ICDAR | 1 |
| 1998 | Design and Specification of Embedded Systems in Java Using Successive, Formal RefinementabstractSuccessive, formal refinement is a new approach for specificationof embedded systems using a general-purpose programming language.Systems are formally modeled as Abstractable SynchronousReactive systems, and Java is used as the design inputlanguage. A policy of use is applied to Java, in the form of languageusage restrictions and class-library extensions, to ensureconsistency with the formal model. A process of incremental,user-guided program transformation is used to refine a Java programuntil it is consistent with the policy of use. The final productis a system specification possessing the properties of the formalmodel, including deterministic behavior, bounded memory usage,and bounded execution time. This approach allows systems designto begin with the flexibility of a general-purpose language, followedby gradual refinement into a more restricted form necessaryfor specification. James Shin Young, Josh MacDonald, Michael Shilman, Abdallah Tabbara, Paul N. Hilfinger, A. Richard Newton |
DAC | 3 |