VLDB 2026 Research / reviewers in the wild / expert
Jerod J. Weinman
dblp:50/2682
· DBLP profile ↗
25ranked-venue papers
15as first author
6since 2021 · last 2026
0000-0002-2247-8174ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 11 first-author · 4 since 2021Databases, data management, data science and information retrieval · 9 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TICLS: Tightly Coupled Language Text SpotterabstractScene text spotting aims to detect and recognize text in real-world images, where instances are often short, fragmented, or visually ambiguous. Existing methods primarily rely on visual cues and implicitly capture local character dependencies, but they overlook the benefits of external linguistic knowledge. Prior attempts to integrate language models either adapt language modeling objectives without external knowledge or apply pretrained models that are misaligned with the word-level granularity of scene text. We propose TICLS, an end-to-end text spotter that explicitly incorporates external linguistic knowledge from a character-level pretrained language model. TICLS contains a pretrained linguistic decoder that fuses visual and linguistic features, enabling robust recognition of ambiguous or fragmented text. Experiments on ICDAR 2015, Total-Text, and CTW1500 demonstrate that TICLS achieves state-of-the-art performance, validating the effectiveness of PLM-guided linguistic integration for scene text spotting. The code is available at https://github.com/knowledge-computing/TiCLS. Leeje Jang, Yijun Lin 0001, Yao-Yi Chiang, Jerod J. Weinman |
WACV | 4 |
| 2025 | LIGHT: Multi-modal Text Linking on Historical Maps
Yijun Lin 0001, Rhett M. Olson, Junhan Wu, Yao-Yi Chiang, Jerod J. Weinman |
ICDAR (2) | 5 |
| 2025 | ICDAR 2025 Competition on Historical Map Text Detection, Recognition, and Linking
Yijun Lin 0001, Solenn Tual, Zekun Li 0007, Leeje Jang, Yao-Yi Chiang, Jerod J. Weinman, Joseph Chazalon, Edwin Carlinet, Julien Perret, Nathalie Abadie, Bertrand Dumenieu, Ta-Chien Chan, Hsiung-Ming Liao, Wen-Rong Su, Mengjie Zou, Tianhao Dai, Rémi Petitpierre, Beatrice Vaienti, Frédéric Kaplan, Isabella diLenardo, Youngmin Baek, Michael Hentschel, Yu Nakagome, Ichimura Shuta, Jeongtae Lee, Chankyu Choi |
ICDAR (5) | 6 |
| 2024 | ICDAR 2024 Competition on Historical Map Text Detection, Recognition, and Linking
Zekun Li 0007, Yijun Lin 0001, Yao-Yi Chiang, Jerod J. Weinman, Solenn Tual, Joseph Chazalon, Julien Perret, Bertrand Dumenieu, Nathalie Abadie |
ICDAR (6) | 4 |
| 2024 | Counting the Corner Cases: Revisiting Robust Reading Challenge Data Sets, Evaluation Protocols, and Metrics
Jerod J. Weinman, Amelia Gómez Grabowska, Dimosthenis Karatzas |
ICDAR (4) | 1 |
| 2023 | Descriptive Attributes for Language-Based Object Keypoint Detection
Jerod J. Weinman, Serge J. Belongie, Stella Frank |
ICVS | 1 |
| 2019 | Deformable Part Models for Automatically Georeferencing Historical Map ImagesabstractLibraries are digitizing their collections of maps from all eras, generating increasingly large online collections of historical cartographic resources. Aligning such maps to a modern geographic coordinate system greatly increases their utility. This work presents a method for such automatic georeferencing, matching raster image content to GIS vector coordinate data. Given an approximate initial alignment that has already been projected from a spherical geographic coordinate system to a Cartesian map coordinate system, a probabilistic shape-matching scheme determines an optimized match between the GIS contours and ink in the binarized map image. Using an evaluation set of 20 historical maps from states and regions of the U.S., the method reduces average alignment RMSE by 12%. Nicholas R. Howe, Jerod J. Weinman, John Gouwar, Aabid Shamji |
SIGSPATIAL/GIS | 2 |
| 2019 | Deep Neural Networks for Text Detection and Recognition in Historical MapsabstractWe introduce deep convolutional and recurrent neural networks for end-to-end, open-vocabulary text reading on historical maps. A text detection network predicts word bounding boxes at arbitrary orientations and scales. The detected word images are then normalized for a robust recognition network. Because accurate recognition requires large volumes of training data but manually labeled data is relatively scarce, we introduce a dynamic map text synthesizer providing a practically infinite stream of training data. Results are evaluated on a labeled data set of 30 maps featuring over 30,000 text labels. Jerod J. Weinman, Benjamin Gafford, Nathan Gifford, Abyaya Lamsal, Liam Niehus-Staab |
ICDAR | 1 |
| 2017 | Geographic and Style Models for Historical Map Alignment and Toponym RecognitionabstractRecognizing the place names within textual labels on historical maps is complicated by many factors, such as curvilinear baselines and dense overlap with other textual or graphical elements. However, maps' alignment with known geography and inter-label typographic style consistencies provide strong cues for resolving uncertainty and reducing text recognition errors. We present a unified probabilistic model to leverage the mutual information between text labels and styles and their geographical locations and categories. This work also introduces likelihood functions to model label placement for polyline and polygon geographical features, such as rivers or provinces. We evaluate the methods on 30 maps from 1866-1927. By interleaving automated map georeferencing with text recognition, we reduce word recognition error by 36% over OCR alone. Incorporating category-style links reduces toponym matching error by 32%. Jerod J. Weinman |
ICDAR | 1 |
| 2016 | Reading and Writing Like Computer Scientists: How to Promote Critical Thinking and Student Engagement (Abstract Only)abstractThis workshop introduces participants to an informal writing process that promotes student engagement and critical thinking with easily-assessed, low-stakes assignments. Unlike the formal writing typically used in software development or capstone courses to demonstrate knowledge, informal writing supports student learning (i.e., writing as thinking). Participants will use the Prioritize, Translate, and Analogize (PTA) Process in a model assignment; discuss how it works; and use it to develop a writing assignment. Participants will receive materials for the workshop assignment, samples of prompts employing the PTA process in a variety of courses, and other support materials. Participants are encouraged to bring an assignment idea to develop at the workshop. The workshop is intended for computer science instructors who want to learn about strategies for integrating writing in their courses to engage students and improve their critical thinking while limiting time for instruction and evaluation. No laptop is required. Mark E. Hoffman, Jerod J. Weinman |
SIGCSE | 2 |
| 2015 | Teaching Computing as Science in a Research ExperienceabstractMany instructors and institutions offer research experiences and training in computing research methods. However, in a national survey, we find that undergraduate students rate their computing research experiences lower than students in other STEM fields. To address this learning gap, we have offered summer undergraduate research experiences in computing that include not only instruction in the important mechanics of research but also grounding in a philosophy of computing science that emphasizes generalized explanation of behavior as a means for control and prediction. After five years, survey results indicate the experience helps close the gap between CS and other STEM fields in benefits gained. Jerod J. Weinman, David D. Jensen, David Lopatto |
SIGCSE | 1 |
| 2014 | Toward Integrated Scene Text ReadingabstractThe growth in digital camera usage combined with a worldly abundance of text has translated to a rich new era for a classic problem of pattern recognition, reading. While traditional document processing often faces challenges such as unusual fonts, noise, and unconstrained lexicons, scene text reading amplifies these challenges and introduces new ones such as motion blur, curved layouts, perspective projection, and occlusion among others. Reading scene text is a complex problem involving many details that must be handled effectively for robust, accurate results. In this work, we describe and evaluate a reading system that combines several pieces, using probabilistic methods for coarsely binarizing a given text region, identifying baselines, and jointly performing word and character segmentation during the recognition process. By using scene context to recognize several words together in a line of text, our system gives state-of-the-art performance on three difficult benchmark data sets. Jerod J. Weinman, Zachary Butler, Dugan Knoll, Jacqueline L. Feild |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Toponym Recognition in Historical Maps by Gazetteer AlignmentabstractHistorical map documents are increasingly digitized for widespread access, but most are only coarsely indexed with meta-data while the contents are largely unsearchable. We propose to increase search ability by automatically recognizing the place names in these digitized artifacts. Using a word recognition system that produces a noisy ranked list of initial hypotheses from a lexicon of viable toponyms, we form a joint probabilistic model for inferring the most likely latent alignment between image toponyms and a gazetteer of known place locations. After a robust generalized RANSAC algorithm identifies the global alignment, we rerank the toponym hypotheses by their posterior probability. Experiments demonstrate a significant boost in word recognition accuracy on a manually annotated set of 19th century U.S. state and regional maps. Jerod J. Weinman |
ICDAR | 1 |
| 2013 | Building knowledge and confidence with mediascripting: a successful interdisciplinary approach to CS1abstractAs the Media Computing approach has shown, writing programs that make images excites a wide variety of students. In this paper, we report on five years of experience with a new approach to media computation, which we call "media scripting". In our introductory class, students build images by interactively scripting an application, so they can experiment easily and mix work "by hand" and "by code"; we collaborate with studio art faculty, so students build works meeting design criteria; and we emphasize multiple paradigms, so students make images using functional, declarative, imperative, and object-oriented techniques. Our approach has proven quite successful--enrollments are up (at least 33% in CS1, 50% in CS2) and we attract more women (currently 40% of the students in the first course, 25% of those in the second course). Other outcomes are equally positive. For example, comparative data show that our students gain significantly more confidence in their abilities than students in other introductory science courses. Samuel A. Rebelsky, Janet Davis, Jerod J. Weinman |
SIGCSE | 3 |
| 2012 | MediaScripting: teaching introductory CS by through interactive graphics scripting (abstract only)abstractupswing, Computer science teachers continue to strive for new examples and problems to interest millenials. The Media Computation approach (Guzdial 2003) has proven successful in attracting students in contexts from community colleges to R1 universities - students are clearly excited by writing programs that make images. In this project, take Media Computing in new directions: we have students build images by interactively scripting an application, which means that they can more easily experiment and mix work that they create "by hand" and work that they create "by programming"; we work collaboratively with studio art faculty, so students build works that must meet underlying design criteria; we teach using the workshop approach, so most classes involve students working in small teams on a set of problems; and we use a multi-paradigm approach - students make images using functional, declarative, imperative, and object-oriented techniques. Janet Davis, Samuel A. Rebelsky, Jerod J. Weinman |
SIGCSE | 3 |
| 2012 | Imaging college educators (abstract only)abstractWithin computing, the imaging field includes computer vision, image understanding, and image processing. While much research and teaching is done at the graduate level, the typical imaging educator at an undergraduate institution is the only specialist in his or her department. This BOF brings together educators who currently teach imaging courses or may be interested in expanding curricular offerings. We will emphasize sharing best practices, ideas, and resources as well as building a network for continued cooperation. Discussion topics may include course organization, assignments and projects, and lecture aids or other materials. Our network will include a mailing list for participants to ask questions and share ideas about imaging pedagogy and other means of sharing course materials. Jerod J. Weinman, Ellen Lowenfeld Walker |
SIGCSE | 1 |
| 2012 | On Learning Conditional Random Fields for Stereo - Exploring Model Structures and Approximate Inference
Christopher Joseph Pal, Jerod J. Weinman, Lam C. Tran, Daniel Scharstein |
Int. J. Comput. Vis. | 2 |
| 2010 | Typographical Features for Scene Text RecognitionabstractScene text images feature an abundance of font style variety but a dearth of data in any given query. Recognition methods must be robust to this variety or adapt to the query data's characteristics. To achieve this, we augment a semi-Markov model-integrating character segmentation and recognition-with a bigram model of character widths. Softly promoting segmentations that exhibit font metrics consistent with those learned from examples, we use the limited information available while avoiding error-prone direct estimates and hard constraints. Incorporating character width bigrams in this fashion improves recognition on low-resolution images of signs containing text in many fonts. Jerod J. Weinman |
ICPR | 1 |
| 2009 | Scene Text Recognition Using Similarity and a Lexicon with Sparse Belief PropagationabstractScene text recognition (STR) is the recognition of text anywhere in the environment, such as signs and storefronts. Relative to document recognition, it is challenging because of font variability, minimal language context, and uncontrolled conditions. Much information available to solve this problem is frequently ignored or used sequentially. Similarity between character images is often overlooked as useful information. Because of language priors, a recognizer may assign different labels to identical characters. Directly comparing characters to each other, rather than only a model, helps ensure that similar instances receive the same label. Lexicons improve recognition accuracy but are used post hoc. We introduce a probabilistic model for STR that integrates similarity, language properties, and lexical decision. Inference is accelerated with sparse belief propagation, a bottom-up method for shortening messages by reducing the dependency between weakly supported hypotheses. By fusing information sources in one model, we eliminate unrecoverable errors that result from sequential processing, improving accuracy. In experimental results recognizing text from images of signs in outdoor scenes, incorporating similarity reduces character recognition error by 19 percent, the lexicon reduces word recognition error by 35 percent, and sparse belief propagation reduces the lexicon words considered by 99.9 percent with a 12X speedup and no loss in accuracy. Jerod J. Weinman, Erik G. Learned-Miller, Allen R. Hanson |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Efficiently Learning Random Fields for Stereo Vision with Sparse Message Passing
Jerod J. Weinman, Lam C. Tran, Christopher Joseph Pal |
ECCV (1) | 1 |
| 2008 | A discriminative semi-Markov model for robust scene text recognitionabstractWe present a semi-Markov model for recognizing scene text that integrates character and word segmentation with recognition. Using wavelet features, it requires only approximate location of the text baseline and font size; no binarization or prior word segmentation is necessary. Our system is aided by a lexicon, yet it also allows non-lexicon words. To facilitate inference with a large lexicon, we use an approximate Viterbi beam search. Our system performs robustly on low-resolution images of signs containing text in fonts atypical of documents. Jerod J. Weinman, Erik G. Learned-Miller, Allen R. Hanson |
ICPR | 1 |
| 2007 | Fast Lexicon-Based Scene Text Recognition with Sparse Belief PropagationabstractUsing a lexicon can often improve character recognition under challenging conditions, such as poor image quality or unusual fonts. We propose a flexible probabilistic model for character recognition that integrates local language properties, such as bigrams, with lexical decision, having open and closed vocabulary modes that operate simultaneously. Lexical processing is accelerated by performing inference with sparse belief propagation, a bottom-up method for hypothesis pruning. We give experimental results on recognizing text from images of signs in outdoor scenes. Incorporating the lexicon reduces word recognition error by 42% and sparse belief propagation reduces the number of lexicon words considered by 97%. Jerod J. Weinman, Erik G. Learned-Miller, Allen R. Hanson |
ICDAR | 1 |
| 2007 | Techniques and Applications for Persistent Backgrounding in a Humanoid Torso RobotabstractOne of the most basic capabilities for an agent with a vision system is to recognize its own surroundings. Yet surprisingly, despite the ease of doing so, many robots store little or no record of their own visual surroundings. This paper explores the utility of keeping the simplest possible persistent record of the environment of a stationary torso robot, in the form of a collection of images captured from various pan-tilt angles around the robot. We demonstrate that this particularly simple process of storing background images can be useful for a variety of tasks, and can relieve the system designer of certain requirements as well. We explore three uses for such a record: auto-calibration, novel object detection with a moving camera, and developing attentional saliency maps. David Walker Duhon, Jerod J. Weinman, Erik G. Learned-Miller |
ICRA | 2 |
| 2006 | Improving Recognition of Novel Input with SimilarityabstractMany sources of information relevant to computer vision and machine learning tasks are often underused. One example is the similarity between the elements from a novel source, such as a speaker, writer, or printed font. By comparing instances emitted by a source, we help ensure that similar instances are given the same label. Previous approaches have clustered instances prior to recognition. We propose a probabilistic framework that unifies similarity with prior identity and contextual information. By fusing information sources in a single model, we eliminate unrecoverable errors that result from processing the information in separate stages and improve overall accuracy. The framework also naturally integrates dissimilarity information, which has previously been ignored. We demonstrate with an application in printed character recognition from images of signs in natural scenes. Jerod J. Weinman, Erik G. Learned-Miller |
CVPR (1) | 1 |
| 2003 | Nonlinear Diffusion Scale-Space and Fast Marching Level Sets for Segmentation of MR Imagery and Volume Estimation of Stroke Lesions
Jerod J. Weinman, George Dean Bissias, Joseph Horowitz, Edward M. Riseman, Allen R. Hanson |
MICCAI (2) | 1 |