Jerod J. Weinman

dblp:50/2682 · DBLP profile ↗
← Back
25ranked-venue papers
15as first author
6since 2021 · last 2026
0000-0002-2247-8174ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 11 first-author · 4 since 2021Databases, data management, data science and information retrieval · 9 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2026 TICLS: Tightly Coupled Language Text Spotter
abstract
Scene text spotting aims to detect and recognize text in real-world images, where instances are often short, fragmented, or visually ambiguous. Existing methods primarily rely on visual cues and implicitly capture local character dependencies, but they overlook the benefits of external linguistic knowledge. Prior attempts to integrate language models either adapt language modeling objectives without external knowledge or apply pretrained models that are misaligned with the word-level granularity of scene text. We propose TICLS, an end-to-end text spotter that explicitly incorporates external linguistic knowledge from a character-level pretrained language model. TICLS contains a pretrained linguistic decoder that fuses visual and linguistic features, enabling robust recognition of ambiguous or fragmented text. Experiments on ICDAR 2015, Total-Text, and CTW1500 demonstrate that TICLS achieves state-of-the-art performance, validating the effectiveness of PLM-guided linguistic integration for scene text spotting. The code is available at https://github.com/knowledge-computing/TiCLS.
Leeje Jang, Yijun Lin 0001, Yao-Yi Chiang, Jerod J. Weinman
WACV4
2025 LIGHT: Multi-modal Text Linking on Historical Maps
Yijun Lin 0001, Rhett M. Olson, Junhan Wu, Yao-Yi Chiang, Jerod J. Weinman
ICDAR (2)5
2025 ICDAR 2025 Competition on Historical Map Text Detection, Recognition, and Linking
Yijun Lin 0001, Solenn Tual, Zekun Li 0007, Leeje Jang, Yao-Yi Chiang, Jerod J. Weinman, Joseph Chazalon, Edwin Carlinet, Julien Perret, Nathalie Abadie, Bertrand Dumenieu, Ta-Chien Chan, Hsiung-Ming Liao, Wen-Rong Su, Mengjie Zou, Tianhao Dai, Rémi Petitpierre, Beatrice Vaienti, Frédéric Kaplan, Isabella diLenardo, Youngmin Baek, Michael Hentschel, Yu Nakagome, Ichimura Shuta, Jeongtae Lee, Chankyu Choi
ICDAR (5)6
2024 ICDAR 2024 Competition on Historical Map Text Detection, Recognition, and Linking
Zekun Li 0007, Yijun Lin 0001, Yao-Yi Chiang, Jerod J. Weinman, Solenn Tual, Joseph Chazalon, Julien Perret, Bertrand Dumenieu, Nathalie Abadie
ICDAR (6)4
2024 Counting the Corner Cases: Revisiting Robust Reading Challenge Data Sets, Evaluation Protocols, and Metrics
Jerod J. Weinman, Amelia Gómez Grabowska, Dimosthenis Karatzas
ICDAR (4)1
2023 Descriptive Attributes for Language-Based Object Keypoint Detection
Jerod J. Weinman, Serge J. Belongie, Stella Frank
ICVS1
2019 Deformable Part Models for Automatically Georeferencing Historical Map Images
abstract
Libraries are digitizing their collections of maps from all eras, generating increasingly large online collections of historical cartographic resources. Aligning such maps to a modern geographic coordinate system greatly increases their utility. This work presents a method for such automatic georeferencing, matching raster image content to GIS vector coordinate data. Given an approximate initial alignment that has already been projected from a spherical geographic coordinate system to a Cartesian map coordinate system, a probabilistic shape-matching scheme determines an optimized match between the GIS contours and ink in the binarized map image. Using an evaluation set of 20 historical maps from states and regions of the U.S., the method reduces average alignment RMSE by 12%.
Nicholas R. Howe, Jerod J. Weinman, John Gouwar, Aabid Shamji
SIGSPATIAL/GIS2
2019 Deep Neural Networks for Text Detection and Recognition in Historical Maps
abstract
We introduce deep convolutional and recurrent neural networks for end-to-end, open-vocabulary text reading on historical maps. A text detection network predicts word bounding boxes at arbitrary orientations and scales. The detected word images are then normalized for a robust recognition network. Because accurate recognition requires large volumes of training data but manually labeled data is relatively scarce, we introduce a dynamic map text synthesizer providing a practically infinite stream of training data. Results are evaluated on a labeled data set of 30 maps featuring over 30,000 text labels.
Jerod J. Weinman, Benjamin Gafford, Nathan Gifford, Abyaya Lamsal, Liam Niehus-Staab
ICDAR1
2017 Geographic and Style Models for Historical Map Alignment and Toponym Recognition
abstract
Recognizing the place names within textual labels on historical maps is complicated by many factors, such as curvilinear baselines and dense overlap with other textual or graphical elements. However, maps' alignment with known geography and inter-label typographic style consistencies provide strong cues for resolving uncertainty and reducing text recognition errors. We present a unified probabilistic model to leverage the mutual information between text labels and styles and their geographical locations and categories. This work also introduces likelihood functions to model label placement for polyline and polygon geographical features, such as rivers or provinces. We evaluate the methods on 30 maps from 1866-1927. By interleaving automated map georeferencing with text recognition, we reduce word recognition error by 36% over OCR alone. Incorporating category-style links reduces toponym matching error by 32%.
Jerod J. Weinman
ICDAR1
2016 Reading and Writing Like Computer Scientists: How to Promote Critical Thinking and Student Engagement (Abstract Only)
abstract
This workshop introduces participants to an informal writing process that promotes student engagement and critical thinking with easily-assessed, low-stakes assignments. Unlike the formal writing typically used in software development or capstone courses to demonstrate knowledge, informal writing supports student learning (i.e., writing as thinking). Participants will use the Prioritize, Translate, and Analogize (PTA) Process in a model assignment; discuss how it works; and use it to develop a writing assignment. Participants will receive materials for the workshop assignment, samples of prompts employing the PTA process in a variety of courses, and other support materials. Participants are encouraged to bring an assignment idea to develop at the workshop. The workshop is intended for computer science instructors who want to learn about strategies for integrating writing in their courses to engage students and improve their critical thinking while limiting time for instruction and evaluation. No laptop is required.
Mark E. Hoffman, Jerod J. Weinman
SIGCSE2
2015 Teaching Computing as Science in a Research Experience
abstract
Many instructors and institutions offer research experiences and training in computing research methods. However, in a national survey, we find that undergraduate students rate their computing research experiences lower than students in other STEM fields. To address this learning gap, we have offered summer undergraduate research experiences in computing that include not only instruction in the important mechanics of research but also grounding in a philosophy of computing science that emphasizes generalized explanation of behavior as a means for control and prediction. After five years, survey results indicate the experience helps close the gap between CS and other STEM fields in benefits gained.
Jerod J. Weinman, David D. Jensen, David Lopatto
SIGCSE1
2014 Toward Integrated Scene Text Reading
abstract
The growth in digital camera usage combined with a worldly abundance of text has translated to a rich new era for a classic problem of pattern recognition, reading. While traditional document processing often faces challenges such as unusual fonts, noise, and unconstrained lexicons, scene text reading amplifies these challenges and introduces new ones such as motion blur, curved layouts, perspective projection, and occlusion among others. Reading scene text is a complex problem involving many details that must be handled effectively for robust, accurate results. In this work, we describe and evaluate a reading system that combines several pieces, using probabilistic methods for coarsely binarizing a given text region, identifying baselines, and jointly performing word and character segmentation during the recognition process. By using scene context to recognize several words together in a line of text, our system gives state-of-the-art performance on three difficult benchmark data sets.
Jerod J. Weinman, Zachary Butler, Dugan Knoll, Jacqueline L. Feild
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Toponym Recognition in Historical Maps by Gazetteer Alignment
abstract
Historical map documents are increasingly digitized for widespread access, but most are only coarsely indexed with meta-data while the contents are largely unsearchable. We propose to increase search ability by automatically recognizing the place names in these digitized artifacts. Using a word recognition system that produces a noisy ranked list of initial hypotheses from a lexicon of viable toponyms, we form a joint probabilistic model for inferring the most likely latent alignment between image toponyms and a gazetteer of known place locations. After a robust generalized RANSAC algorithm identifies the global alignment, we rerank the toponym hypotheses by their posterior probability. Experiments demonstrate a significant boost in word recognition accuracy on a manually annotated set of 19th century U.S. state and regional maps.
Jerod J. Weinman
ICDAR1
2013 Building knowledge and confidence with mediascripting: a successful interdisciplinary approach to CS1
abstract
As the Media Computing approach has shown, writing programs that make images excites a wide variety of students. In this paper, we report on five years of experience with a new approach to media computation, which we call "media scripting". In our introductory class, students build images by interactively scripting an application, so they can experiment easily and mix work "by hand" and "by code"; we collaborate with studio art faculty, so students build works meeting design criteria; and we emphasize multiple paradigms, so students make images using functional, declarative, imperative, and object-oriented techniques. Our approach has proven quite successful--enrollments are up (at least 33% in CS1, 50% in CS2) and we attract more women (currently 40% of the students in the first course, 25% of those in the second course). Other outcomes are equally positive. For example, comparative data show that our students gain significantly more confidence in their abilities than students in other introductory science courses.
Samuel A. Rebelsky, Janet Davis, Jerod J. Weinman
SIGCSE3
2012 MediaScripting: teaching introductory CS by through interactive graphics scripting (abstract only)
abstract
upswing, Computer science teachers continue to strive for new examples and problems to interest millenials. The Media Computation approach (Guzdial 2003) has proven successful in attracting students in contexts from community colleges to R1 universities - students are clearly excited by writing programs that make images. In this project, take Media Computing in new directions: we have students build images by interactively scripting an application, which means that they can more easily experiment and mix work that they create "by hand" and work that they create "by programming"; we work collaboratively with studio art faculty, so students build works that must meet underlying design criteria; we teach using the workshop approach, so most classes involve students working in small teams on a set of problems; and we use a multi-paradigm approach - students make images using functional, declarative, imperative, and object-oriented techniques.
Janet Davis, Samuel A. Rebelsky, Jerod J. Weinman
SIGCSE3
2012 Imaging college educators (abstract only)
abstract
Within computing, the imaging field includes computer vision, image understanding, and image processing. While much research and teaching is done at the graduate level, the typical imaging educator at an undergraduate institution is the only specialist in his or her department. This BOF brings together educators who currently teach imaging courses or may be interested in expanding curricular offerings. We will emphasize sharing best practices, ideas, and resources as well as building a network for continued cooperation. Discussion topics may include course organization, assignments and projects, and lecture aids or other materials. Our network will include a mailing list for participants to ask questions and share ideas about imaging pedagogy and other means of sharing course materials.
Jerod J. Weinman, Ellen Lowenfeld Walker
SIGCSE1
2012 On Learning Conditional Random Fields for Stereo - Exploring Model Structures and Approximate Inference
Christopher Joseph Pal, Jerod J. Weinman, Lam C. Tran, Daniel Scharstein
Int. J. Comput. Vis.2
2010 Typographical Features for Scene Text Recognition
abstract
Scene text images feature an abundance of font style variety but a dearth of data in any given query. Recognition methods must be robust to this variety or adapt to the query data's characteristics. To achieve this, we augment a semi-Markov model-integrating character segmentation and recognition-with a bigram model of character widths. Softly promoting segmentations that exhibit font metrics consistent with those learned from examples, we use the limited information available while avoiding error-prone direct estimates and hard constraints. Incorporating character width bigrams in this fashion improves recognition on low-resolution images of signs containing text in many fonts.
Jerod J. Weinman
ICPR1
2009 Scene Text Recognition Using Similarity and a Lexicon with Sparse Belief Propagation
abstract
Scene text recognition (STR) is the recognition of text anywhere in the environment, such as signs and storefronts. Relative to document recognition, it is challenging because of font variability, minimal language context, and uncontrolled conditions. Much information available to solve this problem is frequently ignored or used sequentially. Similarity between character images is often overlooked as useful information. Because of language priors, a recognizer may assign different labels to identical characters. Directly comparing characters to each other, rather than only a model, helps ensure that similar instances receive the same label. Lexicons improve recognition accuracy but are used post hoc. We introduce a probabilistic model for STR that integrates similarity, language properties, and lexical decision. Inference is accelerated with sparse belief propagation, a bottom-up method for shortening messages by reducing the dependency between weakly supported hypotheses. By fusing information sources in one model, we eliminate unrecoverable errors that result from sequential processing, improving accuracy. In experimental results recognizing text from images of signs in outdoor scenes, incorporating similarity reduces character recognition error by 19 percent, the lexicon reduces word recognition error by 35 percent, and sparse belief propagation reduces the lexicon words considered by 99.9 percent with a 12X speedup and no loss in accuracy.
Jerod J. Weinman, Erik G. Learned-Miller, Allen R. Hanson
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Efficiently Learning Random Fields for Stereo Vision with Sparse Message Passing
Jerod J. Weinman, Lam C. Tran, Christopher Joseph Pal
ECCV (1)1
2008 A discriminative semi-Markov model for robust scene text recognition
abstract
We present a semi-Markov model for recognizing scene text that integrates character and word segmentation with recognition. Using wavelet features, it requires only approximate location of the text baseline and font size; no binarization or prior word segmentation is necessary. Our system is aided by a lexicon, yet it also allows non-lexicon words. To facilitate inference with a large lexicon, we use an approximate Viterbi beam search. Our system performs robustly on low-resolution images of signs containing text in fonts atypical of documents.
Jerod J. Weinman, Erik G. Learned-Miller, Allen R. Hanson
ICPR1
2007 Fast Lexicon-Based Scene Text Recognition with Sparse Belief Propagation
abstract
Using a lexicon can often improve character recognition under challenging conditions, such as poor image quality or unusual fonts. We propose a flexible probabilistic model for character recognition that integrates local language properties, such as bigrams, with lexical decision, having open and closed vocabulary modes that operate simultaneously. Lexical processing is accelerated by performing inference with sparse belief propagation, a bottom-up method for hypothesis pruning. We give experimental results on recognizing text from images of signs in outdoor scenes. Incorporating the lexicon reduces word recognition error by 42% and sparse belief propagation reduces the number of lexicon words considered by 97%.
Jerod J. Weinman, Erik G. Learned-Miller, Allen R. Hanson
ICDAR1
2007 Techniques and Applications for Persistent Backgrounding in a Humanoid Torso Robot
abstract
One of the most basic capabilities for an agent with a vision system is to recognize its own surroundings. Yet surprisingly, despite the ease of doing so, many robots store little or no record of their own visual surroundings. This paper explores the utility of keeping the simplest possible persistent record of the environment of a stationary torso robot, in the form of a collection of images captured from various pan-tilt angles around the robot. We demonstrate that this particularly simple process of storing background images can be useful for a variety of tasks, and can relieve the system designer of certain requirements as well. We explore three uses for such a record: auto-calibration, novel object detection with a moving camera, and developing attentional saliency maps.
David Walker Duhon, Jerod J. Weinman, Erik G. Learned-Miller
ICRA2
2006 Improving Recognition of Novel Input with Similarity
abstract
Many sources of information relevant to computer vision and machine learning tasks are often underused. One example is the similarity between the elements from a novel source, such as a speaker, writer, or printed font. By comparing instances emitted by a source, we help ensure that similar instances are given the same label. Previous approaches have clustered instances prior to recognition. We propose a probabilistic framework that unifies similarity with prior identity and contextual information. By fusing information sources in a single model, we eliminate unrecoverable errors that result from processing the information in separate stages and improve overall accuracy. The framework also naturally integrates dissimilarity information, which has previously been ignored. We demonstrate with an application in printed character recognition from images of signs in natural scenes.
Jerod J. Weinman, Erik G. Learned-Miller
CVPR (1)1
2003 Nonlinear Diffusion Scale-Space and Fast Marching Level Sets for Segmentation of MR Imagery and Volume Estimation of Stroke Lesions
Jerod J. Weinman, George Dean Bissias, Joseph Horowitz, Edward M. Riseman, Allen R. Hanson
MICCAI (2)1