VLDB 2026 Research / reviewers in the wild / expert
Adrian Ulges
dblp:09/7047
· DBLP profile ↗
36ranked-venue papers
10as first author
13since 2021 · last 2026
0009-0001-1915-2464ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 19 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 9 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiteDoc: Distilling Large Document Models into Efficient Task-Specific Encoders
Tayyab Raza, Syed Muhammad Taha Imam, Adrian Ulges, Ulrich Schwanecke, Momina Moetesum, Faisal Shafait |
ICDAR (2) | 3 |
| 2026 | BAFIS: Dataset + Framework to assess occupational Bias and Human Preference in modern Text-to-image ModelsabstractGenerative artificial intelligence has the potential to improve productivity and transform the production of creative content. However, existing research indicates that image generation models are significantly influenced by biases. This work investigates the inherent biases and language-induced biases present in text-to-image models within the context of occupation-related image generation, complementing established metrics with human preference feedback. We present a comprehensive evaluation of five current text-to-image models: Midjourney v6.1, Stable Diffusion 3 Medium, DALL-E 3, Playground v2.5, and FLUX.1-dev, focusing on gender and ethnicity bias, image quality, and prompt alignment. To facilitate this evaluation, we developed the "Battle-Arena for Fair Image Synthesis" (BAFIS), a platform designed to collect human feedback on bias in generated images. Furthermore, we created a dataset comprising 21,140 synthetic images generated using multilingual prompts, which serves as a basis for our analysis. We further place our results within a broader social context by comparing them to official statistics from the German Federal Employment Agency. Our findings reveal systematic biases in text-to-image models, with established evaluation metrics in partial correlation with subjective user ratings. Thus, our research emphasizes the need for including human preferences to develop fairer and more inclusive text-to-image models. Code and dataset are public here. Thomas Klassert, Adrian Ulges, Biying Fu |
WACV | 2 |
| 2026 | Gaze-based Personal Memory: Leveraging Eye Tracking to Improve Relevance in Text Retrieval Systems ETRA013abstractWith the rapid growth of digital textual content, users face increasing challenges in rediscovering information which they have encountered in the past. Traditional search engines lack the ability to prioritize content truly viewed by users over merely visible content. We propose a method that captures both text passages and corresponding gaze data while reading, storing them in a searchable knowledge base. We then explore boosting strategies to improve text retrieval by prioritizing passages that users actually viewed, improving the personalization and relevance of search results. To this end, we repurposed the recent g-Rel-READER dataset to evaluate various gaze-based boosting techniques and address the research gap caused by the lack of combined text, gaze, and relevance data. The evaluation demonstrates the potential of gaze data to serve as a boosting criterion for search, with mean average precision (MAP) increased by over 33% over a purely text-based retrieval. Philipp Gross, Martin Weier, Adrian Ulges |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2025 | Investigating the Configurability of LLMs for the Generation of Knowledge Work Datasets
Desiree Heim, Christian Jilek, Adrian Ulges, Andreas Dengel 0001 |
ICAART (3) | 3 |
| 2025 | SlimDoc: lightweight distillation of document transformer modelsabstractAbstract Deploying state-of-the-art document understanding models remains resource-intensive and impractical in many real-world scenarios, particularly where labeled data is scarce and computational budgets are constrained. To address these challenges, this work proposes a novel approach towards parameter-efficient document understanding models capable of adapting to specific tasks and document types without the need for labeled data. Specifically, we propose an approach coined SlimDoc to distill multimodal document transformer encoder models into smaller student models, using internal signals at different training stages, followed by external signals. Our approach is inspired by TinyBERT and adapted to the domain of document understanding transformers. We demonstrate SlimDoc to outperform both a single-stage distillation and a direct fine-tuning of the student. Experimental results across six document understanding datasets demonstrate our approach’s effectiveness: Our distilled student models achieve on average $$93.0\%$$ 93.0 % of the teacher’s performance, while the fine-tuned students achieve $$87.0\%$$ 87.0 % of the teacher’s performance. Without requiring any labeled data, we create a compact student which achieves $$96.0\%$$ 96.0 % of the performance of its supervised-distilled counterpart and $$86.2\%$$ 86.2 % of the performance of a supervised-fine-tuned teacher model. We demonstrate our distillation approach to pick up on document geometry and to be effective on the two popular document understanding models LiLT and LayoutLMv3. Our implementation and training data is available at https://github.com/marcel-lamott/SlimDoc . Marcel Lamott, Muhammad Armaghan Shakir, Adrian Ulges, Yves-Noel Weweler, Faisal Shafait |
Int. J. Document Anal. Recognit. | 3 |
| 2024 | LAPDoc: Layout-Aware Prompting for Documents
Marcel Lamott, Yves-Noel Weweler, Adrian Ulges, Faisal Shafait, Dirk Krechel, Darko Obradovic |
ICDAR (4) | 3 |
| 2023 | Domain-Specific Knowledge Graph Adaption with Industrial Text Data
Felix Hamann, Adrian Ulges |
IEA/AIE (1) | 2 |
| 2023 | Value Stream Repair Using Graph Structure Learning
Marco Wrzalik, Julian Eversheim, Johannes Villmow, Adrian Ulges, Dirk Krechel, Sven Spieckermann, Robert Forstner |
IEA/AIE (2) | 4 |
| 2023 | How Well Can Masked Language Models Spot Identifiers That Violate Naming Guidelines?abstractUsing meaningful identifiers in source code reduces the risk of errors, the cognitive load of developers, and speeds up the development process. Therefore, recent research has looked into an AI-based analysis of identifiers, for which large-scale language models appear to offer great potential. Based on tokens’ probabilities, such models can suggest identifiers that are likely to appear in a given context. While current research has used language models to predict the most likely identifier names, studies on assessing the quality of given identifiers are scarce. To this end, we explore adherence to identifier naming guidelines as a proxy for identifier quality and propose and evaluate two unsupervised approaches for spotting violations: First, a generative approach, which uses the probability distribution of the language model directly without fine-tuning. Second, a discriminative method, which fine-tunes the model’s encoder to discriminate between original identifiers and similar drop-in replacements suggested by a weak AI. We demonstrate that the proposed approaches can successfully detect violations of common guidelines for identifier naming. To do so, we have developed a dataset built on widely accepted identifier naming guidelines. The manually annotated dataset contains more than 6000 dense annotations of identifiers for 28 common guidelines. Using the data, we show that the generative approach achieves the best results, but that the particular masking strategy and scoring method matter substantially. Also, we demonstrate our approach to outperform other recent code transformers. In a per-guideline analysis, we highlight the potential and limitations of language models, and provide a blue-print for training and evaluating their ability to identify bad identifier names in source code. We make our dataset and models’ implementation publicly available to encourage future research on AI-based identifier quality assessment. Johannes Villmow, Viola Campos, Jean Petry, Amine Abbad-Andaloussi, Adrian Ulges, Barbara Weber |
SCAM | 5 |
| 2022 | Addressing Leakage in Self-Supervised Contextualized Code RetrievalabstractWe address contextualized code retrieval, the search for code snippets helpful to fill gaps in a partial input program. Our approach facilitates a large-scale self-supervised contrastive training by splitting source code randomly into contexts and targets. To combat leakage between the two, we suggest a novel approach based on mutual identifier masking, dedentation, and the selection of syntax-aligned targets. Our second contribution is a new dataset for direct evaluation of contextualized code retrieval, based on a dataset of manually aligned subpassages of code clones. Our experiments demonstrate that the proposed approach improves retrieval substantially, and yields new state-of-the-art results for code clone and defect detection. Johannes Villmow, Viola Campos, Adrian Ulges, Ulrich Schwanecke |
COLING | 3 |
| 2021 | An End-to-end Model for Entity-level Relation Extraction using Multi-instance LearningabstractWe present a joint model for entity-level relation extraction from documents.In contrast to other approaches -which focus on local intra-sentence mention pairs and thus require annotations on mention level -our model operates on entity level.To do so, a multi-task approach is followed that builds upon coreference resolution and gathers relevant signals via multi-instance learning with multi-level representations combining global entity and local mention information.We achieve state-of-theart relation extraction results on the DocRED dataset and report the first entity-level end-toend relation extraction results for future reference.Finally, our experimental results suggest that a joint approach is on par with taskspecific learning, though more efficient due to shared parameters and training steps. Markus Eberts, Adrian Ulges |
EACL | 2 |
| 2021 | Open-World Knowledge Graph Completion Benchmarks for Knowledge Discovery
Felix Hamann, Adrian Ulges, Dirk Krechel, Ralph Bergmann |
IEA/AIE (2) | 2 |
| 2021 | A Structural Transformer with Relative Positions in Trees for Code-to-Sequence TasksabstractWe suggest two approaches to incorporate syntactic information into transformer models encoding trees (e.g. abstract syntax trees) and generating sequences. First, we use self-attention with relative position representations to consider structural relationships between nodes using a representation that encodes movements between any pair of nodes in the tree, and demonstrate how those movements can be computed efficiently on the fly. Second, we suggest an auxiliary loss enforcing the network to predict the lowest common ancestor of node pairs. We apply both methods to source code summarization tasks, where we outperform the state-of-the-art by up to 6 % F1. On natural language machine translation, our models yield competitive results. We also consistently outperform sequence-based transformers, and demonstrate that our method yields representations that are more closely aligned with the AST structure. Johannes Villmow, Adrian Ulges, Ulrich Schwanecke |
IJCNN | 2 |
| 2020 | ManyEnt: A Dataset for Few-shot Entity TypingabstractWe introduce ManyEnt, a benchmark for entity typing models in few-shot scenarios.ManyEnt offers a rich typeset, with a fine-grain variant featuring 256 entity types and a coarse-grain one with 53 entity types.Both versions have been derived from the Wikidata knowledge graph in a semi-automatic fashion.We also report results for two baselines using BERT, reaching up to 70.68% accuracy (10-way 1-shot). Markus Eberts, Kevin Pech, Adrian Ulges |
COLING | 3 |
| 2020 | Span-Based Joint Entity and Relation Extraction with Transformer Pre-TrainingabstractWe introduce SpERT, an attention model for span-based joint entity and relation extraction. Our key contribution is a light-weight reasoning on BERT embeddings, which features entity recognition and filtering, as well as relation classification with a localized, marker-free context representation. The model is trained using strong within-sentence negative samples, which are efficiently extracted in a single BERT pass. These aspects facilitate a search over all spans in the sentence. In ablation studies, we demonstrate the benefits of pre-training, strong negative sampling and localized context. Our model outperforms prior work by up to 2.6% F1 score on several datasets for joint entity and relation extraction. Markus Eberts, Adrian Ulges |
ECAI | 2 |
| 2019 | An Open-World Extension to Knowledge Graph Completion ModelsabstractWe present a novel extension to embedding-based knowledge graph completion models which enables them to perform open-world link prediction, i.e. to predict facts for entities unseen in training based on their textual description. Our model combines a regular link prediction model learned from a knowledge graph with word embeddings learned from a textual corpus. After training both independently, we learn a transformation to map the embeddings of an entity’s name and description to the graph-based embedding space.In experiments on several datasets including FB20k, DBPedia50k and our new dataset FB15k-237-OWE, we demonstrate competitive results. Particularly, our approach exploits the full knowledge graph structure even when textual descriptions are scarce, does not require a joint training on graph and text, and can be applied to any embedding-based link prediction model, such as TransE, ComplEx and DistMult. Haseeb Shah, Johannes Villmow, Adrian Ulges, Ulrich Schwanecke, Faisal Shafait |
AAAI | 3 |
| 2017 | Cross-modal Image-Graphics Retrieval by Neural Transfer Learningabstractresearch-article Share on Cross-modal Image-Graphics Retrieval by Neural Transfer Learning Authors: Fabian Junkert RheinMain University of Applied Sciences, Wiesbaden, Germany RheinMain University of Applied Sciences, Wiesbaden, GermanyView Profile , Markus Eberts RheinMain University of Applied Sciences, Wiesbaden, Germany RheinMain University of Applied Sciences, Wiesbaden, GermanyView Profile , Adrian Ulges RheinMain University of Applied Sciences, Wiesbaden, Germany RheinMain University of Applied Sciences, Wiesbaden, GermanyView Profile , Ulrich Schwanecke RheinMain University of Applied Sciences, Wiesbaden, Germany RheinMain University of Applied Sciences, Wiesbaden, GermanyView Profile Authors Info & Claims ICMR '17: Proceedings of the 2017 ACM on International Conference on Multimedia RetrievalJune 2017 Pages 330–337https://doi.org/10.1145/3078971.3078994Published:06 June 2017Publication History 1citation212DownloadsMetricsTotal Citations1Total Downloads212Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Fabian Junkert, Markus Eberts, Adrian Ulges, Ulrich Schwanecke |
ICMR | 3 |
| 2015 | AMIGO - automatic indexing of lecture footageabstractWe present AMIGO, an automatic indexer for video presentations which - given an e-lecture and supplementary slides - localizes the exact time and position of each slide displayed in the video footage. This offers richer access to viewers, including a slide-accurate navigation and a text-based interaction with the video. AMIGO is based on a matching of local features between video frames and presentation slides. Our key contribution, however, is the combination of local feature matching with two temporal models (a Hidden Markov Model (HMM) and a simple heuristic filter), exploiting the alignment of the presentation with the reading order of its supplementary material. We demonstrate the effectiveness of our approach in quantitative experiments on a dataset of e-lectures and screencasts, which show - with an average accuracy of over 95% - that the approach works under occlusion and camera motion. Markus Eberts, Adrian Ulges, Ulrich Schwanecke |
ICDAR | 2 |
| 2012 | Linking visual concept detection with viewer demographicsabstractThe estimation of demographic target groups for web videos -- with applications in ad targeting -- poses a challenging problem, as the textual description and view statistics available for many clips is extremely sparse. Therefore, the goal of this paper is to link a clip's popularity across different viewer ages and genders on the one hand with the video content on the other: Employing user comments and user profiles on YouTube, we show that there is a strong correlation between demographic target groups and semantic concepts appearing in the video (like "teenage male" and "skateboarding"). Based on this observation, we suggest two approaches: First, the demographic target group of a clip is predicted automatically via a content-based concept detection. Second, should sufficient view statistics already give a good impression of a video's audience, we show that this information can serve as a valuable additional signal to disambiguate concept detection. Adrian Ulges, Markus Koch, Damian Borth |
ICMR | 1 |
| 2012 | Dynamic vocabularies for web-based concept detection by trend discoveryabstractWe present a novel approach towards automatic vocabulary selection for video concept detection. Our key idea is to expand concept vocabularies with trending topics that we mine automatically on other media like Wikipedia or Twitter. We evaluate several strategies for extending concept detection to auto-detect these topics in new videos, either by linking them to a static concept vocabulary, by a visual learning of trends on the fly, or by an expansion of the vocabulary. Damian Borth, Adrian Ulges, Thomas M. Breuel |
ACM Multimedia | 2 |
| 2011 | Automatic detection of child pornography using color visual wordsabstractThis paper addresses the computer-aided detection of child sexual abuse (CSA) images, a challenge of growing importance in multimedia forensics and security. In contrast to previous solutions based on hashsums, file names, or the retrieval of visually similar images, we introduce a system which employs visual recognition techniques to automatically identify suspect material. Our approach is based on color-enhanced visual word features and a statistical classification using SVMs. The detector is adapted to CSA material in a training step. In collaboration with police partners, we have conducted a quantitative evaluation on several datasets (including real-world CSA material). Our results indicate that recognizing child pornography is a challenging problem (more difficult than the detection of regular porn). Yet, while skin detection - a popular approach in pornography detection - fails, our approach can achieve a prioritization of content (equal error 11 – 24%) to improve the efficiency of forensic investigations of child sexual abuse. Examples illustrate that the system employs color cues as key features for discriminating CSA content. Adrian Ulges, Armin Stahl |
ICME | 1 |
| 2011 | Lookapp: interactive construction of web-based concept detectorsabstractWhile online platforms like YouTube and Flickr do provide massive content for training of visual concept detectors, it remains a difficult challenge to retrieve the right training content from such platforms. In this technical demonstration we present lookapp, a system for the interactive construction of web-based concept detectors. It major features are an interactive "concept-to-query" mapping for training data acquisition and an efficient detector construction based on third party cloud computing services. Damian Borth, Adrian Ulges, Thomas M. Breuel |
ICMR | 2 |
| 2011 | Scene-based image retrieval by transitive matchingabstractWe address scene-based image retrieval, the challenge of finding pictures taken at the same location as a given query image, whereas a key challenge lies in the fact that target images may show the same scene but different parts of it. To overcome this lack of direct correspondences with the query image, we study two strategies that exploit the structure of the targeted image collection: first, cluster matching, where pictures are grouped and retrieval is conducted on cluster level. Second, we propose a probabilistically motivated shortest path approach that determines retrieval scores based on the shortest path in a cost graph defined over the image collection. We evaluate both approaches on several datasets including indoor and outdoor locations, demonstrating that the accuracy of scene-based retrieval can be improved distinctly (by up to 40%), particularly by the shortest path approach. Adrian Ulges, Christian Schulze 0001 |
ICMR | 1 |
| 2011 | Third ACM international workshop on multimedia in forensics and intelligence (MiFor 2011)abstractThis paper introduces the context of the workshop and the associated papers. Sebastiano Battiato, Sabu Emmanuel, Adrian Ulges, Marcel Worring |
ACM Multimedia | 3 |
| 2011 | Automatic concept-to-query mapping for web-based concept detector trainingabstractNowadays, online platforms like YouTube provide massive content for training of visual concept detectors. However, it remains a difficult challenge to retrieve the right training content from such platforms since the underlying query construction can be arbitrarily complex. In this paper we present an approach, which offers an automatic concept-to-query mapping for training data acquisition from such platforms. Queries are automatically constructed by a keyword selection and a category assignment using ImageNet and Google Sets as external sources. Our results demonstrate that the proposed method is able to reach retrieval results comparable to queries constructed by humans providing 76% more relevant content for detector training than a one-to-one mapping of concept names to retrieval queries would do. Damian Borth, Adrian Ulges, Thomas M. Breuel |
ACM Multimedia | 2 |
| 2011 | Learning Visual Contexts for Image Annotation From Flickr GroupsabstractWe present an extension of automatic image annotation that takes the context of a picture into account. Our core assumption is that users do not only provide individual images to be tagged, but group their pictures into batches (e.g., all snapshots taken over the same holiday trip), whereas the images within a batch are likely to have a common style. These batches are matched with categories learned from Flickr groups, and an accurate context-specific annotation is performed. Adrian Ulges, Marcel Worring, Thomas M. Breuel |
IEEE Trans. Multim. | 1 |
| 2010 | Can Motion Segmentation Improve Patch-Based Object Recognition?abstractPatch-based methods, which constitute the state of the art in object recognition, are often applied to video data, where motion information provides a valuable clue for separating objects of interest from the background. We show that such motion-based segmentation improves the robustness of patch-based recognition with respect to clutter. Our approach, which employs segmentation information to rule out incorrect correspondences between training and test views, is demonstrated empirically to distinctly outperform baselines operating on unsegmented images. Relative improvements reach 50% for the recognition of specific objects, and 33% for object category retrieval. Adrian Ulges, Thomas M. Breuel |
ICPR | 1 |
| 2010 | Second ACM international workshop on multimedia in forensics, security and intelligence (MiFor 2010)abstractThis paper introduces the context of the workshop and the associated papers. Sebastiano Battiato, Sabu Emmanuel, Adrian Ulges, Marcel Worring |
ACM Multimedia | 3 |
| 2010 | Learning automatic concept detectors from online video
Adrian Ulges, Christian Schulze 0001, Markus Koch, Thomas M. Breuel |
Comput. Vis. Image Underst. | 1 |
| 2009 | Fast Discriminative Linear Models for Scalable Video TaggingabstractWhile video tagging (or "concept detection") is a key building block of research prototypes for video retrieval, its practical use is hindered by the computational effort associated with learning and detecting thousands of concepts. Support vector machines (SVMs), which can be considered the standard approach, scale poorly since the number of support vectors is usually high. In this paper, we propose a novel alternative that offers the benefits of rapid training and detection. This linear-discriminative method is based on the maximization of the area under the ROC. In quantitative experiments on a publicly available dataset of Web videos, we demonstrate that this approach offers a significant speedup at a moderate performance loss compared to SVMs, and also outperforms another well-known linear-discriminative method based on a Passive-Aggressive Online Learning (PAMIR). Roberto Paredes, Adrian Ulges, Thomas M. Breuel |
ICMLA | 2 |
| 2009 | TubeFiler: an automatic web video categorizerabstractWhile hierarchies are powerful tools for organizing content in other application areas, current web video platforms offer only limited support for a taxonomy-based browsing. To overcome this limitation, we present a framework called TubeFiler. Its two key features are an automatic multimodal categorization of videos into a genre hierarchy, and a support of additional fine-grained hierarchy levels based on unsupervised learning. We present experimental results on real-world YouTube clips with a 2-level 46-category genre hierarchy, indicating that - though the problem is clearly challenging - good category suggestions can be achieved. For example, if TubeFiler suggests 5 categories, it hits the right one (or at least its supercategory) in 91.8% of cases. Damian Borth, Jörn Hees, Markus Koch, Adrian Ulges, Christian Schulze 0001, Thomas M. Breuel, Roberto Paredes |
ACM Multimedia | 4 |
| 2009 | Detecting pornographic video content by combining image features with motion informationabstractWith the rise of large-scale digital video collections, the challenge of automatically detecting adult video content has gained significant impact with respect to applications such as content filtering or the detection of illegal material. While most systems represent videos with keyframes and then apply techniques well-known for static images, we investigate motion as another discriminative clue for pornography detection. A framework is presented that combines conventional keyframe-based methods with a statistical analysis of MPEG-4 motion vectors. Two general approaches are followed to describe motion patterns, one based on the detection of periodic motion and one on motion histograms. Our experiments on real-world web video data show that this combination with motion information improves the accuracy of pornography detection significantly (equal error is reduced from 9.9% to 6.0%). Comparing both motion descriptors, histograms outperform periodicity detection. Christian Jansohn, Adrian Ulges, Thomas M. Breuel |
ACM Multimedia | 2 |
| 2008 | Segmentation by combining parametric optical flow with a color modelabstractWe present a simple but efficient model for object segmentation in video scenes that integrates motion and color information in a joint probabilistic framework. Optical flow is modeled using parametric motion with Gaussian noise. The color distribution of foreground and background is described by histograms or Gaussian mixture models. Optimization is carried out using an efficient graph cut algorithm. In quantitative experiments on a variety of video data, we demonstrate that the proposed approach leads to significant reductions in error rates compared to a state-of-the-art motion-only segmentation. Adrian Ulges, Thomas M. Breuel |
ICPR | 1 |
| 2008 | A System That Learns to Tag Videos by Watching Youtube
Adrian Ulges, Christian Schulze 0001, Daniel Keysers, Thomas M. Breuel |
ICVS | 1 |
| 2005 | Document Image Dewarping using Robust Estimation of Curled Text LinesabstractDigital cameras have become almost ubiquitous and their use for fast and casual capturing of natural images is unchallenged. For making images of documents, however, they have not caught up to flatbed scanners yet, mainly because camera images tend to suffer from distortion due to the perspective and are therefore limited in their further use for archival or OCR. For images of non-planar paper surfaces like books, page curl causes additional distortion, which poses an even greater problem due to its nonlinearity. This paper presents a new algorithm for removing both perspective and page curl distortion. It requires only a single camera image as input and relies on a priori layout information instead of additional hardware. Therefore, it is much more user friendly than most previous approaches, and allows for flexible ad hoc document capture. Results are presented showing that the algorithm produces visually pleasing output and increases OCR accuracy, thus having the potential to become a general purpose preprocessing tool for camera based document capture. Adrian Ulges, Christoph H. Lampert, Thomas M. Breuel |
ICDAR | 1 |
| 2004 | Document capture using stereo visionabstractCapturing images of documents using handheld digital cameras has a variety of applications in academia, research, knowledge management, retail, and office settings. The ultimate goal of such systems is to achieve image quality comparable to that currently achieved with flatbed scanners even for curved, warped, or curled pages. This can be achieved by high-accuracy 3D modeling of the page surface, followed by a "flattening" of the surface. A number of previous systems have either assumed only perspective distortions, or used techniques like structured lighting, shading, or side-imaging for obtaining 3D shape. This paper describes a system for handheld camera-based document capture using general purpose stereo vision methods followed by a new document dewarping technique. Examples of shape modeling and dewarping of book images is shown. Adrian Ulges, Christoph H. Lampert, Thomas M. Breuel |
ACM Symposium on Document Engineering | 1 |