VLDB 2026 Research / reviewers in the wild / expert
David S. Doermann
dblp:88/6921
· DBLP profile ↗
235ranked-venue papers
20as first author
54since 2021 · last 2026
0000-0003-1639-4561ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 174 · 17 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 114 · 5 first-author · 24 since 2021Databases, data management, data science and information retrieval · 56 · 8 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Textured Geometry Evaluation: Perceptual 3D Textured Shape Metric via 3D Latent-Geometry NetworkabstractTextured high-fidelity 3D models are crucial for games, AR/VR, and film, but human-aligned evaluation methods still fall behind despite recent advances in 3D reconstruction and generation. Existing metrics, such as Chamfer Distance, often fail to align with how humans evaluate the fidelity of 3D shapes. Recent learning-based metrics attempt to improve this by relying on rendered images and 2D image quality metrics. However, these approaches face limitations due to incomplete structural coverage and sensitivity to viewpoint choices. Moreover, most methods are trained on synthetic distortions, which differ significantly from real-world distortions, resulting in a domain gap. To address these challenges, we propose a new fidelity evaluation method that is based directly on 3D meshes with texture, without relying on rendering. Our method, named Textured Geometry Evaluation TGE, jointly uses the geometry and color information to calculate the fidelity of the input textured mesh with comparison to a reference colored shape. To train and evaluate our metric, we design a human-annotated dataset with real-world distortions. Experiments show that TGE outperforms rendering-based and geometry-only methods on real-world distortion dataset. Tianyu Luan, Xuelu Feng, Zixin Zhu, Phani Nuney, Sheng Liu 0017, David S. Doermann, Chunming Qiao, Junsong Yuan 0001 |
AAAI | 7 |
| 2026 | Preoperative Prediction of Esophageal Cancer Survival in CT via Tumor and Lymph Node Context and Geometry ModelingabstractEsophageal cancer is one of the most lethal cancers, with 5-year survival rate of only 20%. Patient outcomes can vary significantly even though they are at the same cancer stage and receive similar treatments. Accurate prognostic prediction for esophageal cancer patients is highly desired to receive personalized precise treatment. Nevertheless, there are very few automated methods yet to fully exploit the preoperative contrast-enhanced computed tomography (CE-CT) imaging for assessing esophageal cancer prognosis. In addition to image patterns, important prognostic factors should encompass tumor size and location, as well as lymph nodes (LNs) involvement, including features such as LN number, size, spatial distribution, and their proximity to tumor. Considering these complexities, we propose a novel Tumor and LN Context-Geometry network for the preoperative prediction of esophageal cancer survival in CE-CT images. Specifically, we 1) focus on learning survival patterns of CT texture via co-attention context modeling at most informative regions, i.e., automatically segmented tumor, LNs and LN-stations; and 2) integrate tumor and LN anatomical and spatial associations into neural geometry modeling for a comprehensive learning of metastatic involvement and tumor invasion to adjacent structures. Empirical studies show our presented framework can improve overall survival prediction performances compared with existing state-of-the-art survival analysis methods, and evidently suggest that incorporating these findings into the existing esophageal cancer staging system would add its clinical values. Yirui Wang 0002, Haoshen Li, Jiawen Yao, Lianzhen Zhong, Dazhou Guo, Ke Yan 0006, David S. Doermann, Le Lu 0001, Feiran Jiao, Tsung-Ying Ho, Ling Zhang 0002, Abudili Abuduxuku, Xianghua Ye, Dakai Jin |
IEEE Trans. Medical Imaging | 9 |
| 2025 | DFM: Differentiable Feature Matching for Anomaly DetectionabstractFeature matching methods for unsupervised anomaly detection have demonstrated impressive performance. Existing methods primarily rely on self-supervised training and handcrafted matching schemes for task adaptation. However, they can only achieve an inferior feature representation for anomaly detection because the feature extraction and matching modules are separately trained. To address these issues, we propose a Differentiable Feature Matching (DFM) framework for joint optimization of the feature extractor and the matching head. DFM transforms nearest-neighbor matching into a pooling-based module and embeds it within a Feature Matching Network (FMN). This design enables end-to-end feature extraction and feature matching module training, thus providing better feature representation for anomaly detection tasks. DFM is generic and can be incorporated into existing feature-matching methods. We implement DFM with various backbones and conduct extensive experiments across various tasks and datasets, demonstrating its effectiveness. Notably, we achieve state-of-the-art results in the continual anomaly detection task with instance-AUROC improvement of up to 3.9% and pixel-AP improvement of up to 5.5%. Yimi Wang, Yuguang Yang 0007, Runqi Wang, Guodong Guo, David S. Doermann, Baochang Zhang 0001 |
CVPR | 7 |
| 2025 | PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask Conditions
Mahesh Bhosale, Abdul Wasi, Yuanhao Zhai 0001, Yunjie Tian, Samuel P. Border, Nan Xi, Pinaki Sarder, Junsong Yuan 0001, David S. Doermann |
ICCV | 9 |
| 2025 | Text2Outfit: Controllable Outfit Generation With Multimodal Language Models
Yuanhao Zhai 0001, Yen-Liang Lin, Minxu Peng, Larry Davis 0001, Ashwin Chandramouli, Junsong Yuan 0001, David S. Doermann |
ICCV | 7 |
| 2025 | AutoEdit: Automatic Hyperparameter Tuning for Image EditingabstractRecent advances in diffusion models have revolutionized text-guided image editing, yet existing editing methods face critical challenges in hyperparameter identification. To get the reasonable editing performance, these methods often require the user to brute-force tune multiple interdependent hyperparameters, such as inversion timesteps and attention modification, \textit{etc.} This process incurs high computational costs due to the huge hyperparameter search space.
We consider searching optimal editing's hyperparameters as a sequential decision-making task within the diffusion denoising process. Specifically, we propose a reinforcement learning framework, which establishes a Markov Decision Process that dynamically adjusts hyperparameters across denoising steps, integrating editing objectives into a reward function.
The method achieves time efficiency through proximal policy optimization while maintaining optimal hyperparameter configurations. Experiments demonstrate significant reduction in search time and computational overhead compared to existing brute-force approaches, advancing the practical deployment of a diffusion-based image editing framework in the real world. Quan Dao, Mahesh Bhosale, Yunjie Tian, Dimitris N. Metaxas, David S. Doermann |
NeurIPS | 6 |
| 2025 | YOLOv12: Attention-Centric Real-Time Object DetectorsabstractEnhancing the network architecture of the YOLO framework has been crucial for a long time. Still, it has focused on CNN-based improvements despite the proven superiority of attention mechanisms in modeling capabilities. This is because attention-based models cannot match the speed of CNN-based models. This paper proposes an attention-centric YOLO framework, namely YOLOv12, that matches the speed of previous CNN-based ones while harnessing the performance benefits of attention mechanisms. YOLOv12 surpasses popular real-time object detectors in accuracy with competitive speed. For example, YOLOv12-N achieves 40.5% mAP with an inference latency of 1.62 ms on a T4 GPU, outperforming advanced YOLOv10-N / YOLO11-N by 2.0%/1.1% mAP with a comparable speed. This advantage extends to other model scales. YOLOv12 also surpasses end-to-end real-time detectors that improve DETR, such as RT-DETRv2 / RT-DETRv3: YOLOv12-X beats RT-DETRv2-R101 / RT-DETRv3-R101 while running faster with fewer computations and parameters. See more comparisons in Figure 1. Source code is available at https://github.com/sunsmarterjie/yolov12. Yunjie Tian, Qixiang Ye, David S. Doermann |
NeurIPS | 3 |
| 2025 | Scene text recognition: an Indic perspective
Vasanthan P. Vijayan, Sukalpa Chanda, David S. Doermann, Narayanan Chatapuram Krishnan |
Int. J. Document Anal. Recognit. | 3 |
| 2024 | Federated Learning via Input-Output Collaborative DistillationabstractFederated learning (FL) is a machine learning paradigm in which distributed local nodes collaboratively train a central model without sharing individually held private data. Existing FL methods either iteratively share local model parameters or deploy co-distillation. However, the former is highly susceptible to private data leakage, and the latter design relies on the prerequisites of task-relevant real data. Instead, we propose a data-free FL framework based on local-to-central collaborative distillation with direct input and output space exploitation. Our design eliminates any requirement of recursive local parameter exchange or auxiliary task-relevant data to transfer knowledge, thereby giving direct privacy control to local users. In particular, to cope with the inherent data heterogeneity across locals, our technique learns to distill input on which each local model produces consensual yet unique results to represent each expertise. Our proposed FL framework achieves notable privacy-utility trade-offs with extensive experiments on image classification and segmentation tasks under various real-world heterogeneous federated learning settings on both natural and medical images. Code is available at https://github.com/lsl001006/FedIOD. Shanglin Li, Yuxiang Bao, Barry Yao, Yawen Huang, Ziyan Wu 0001, Baochang Zhang 0001, Yefeng Zheng 0001, David S. Doermann |
AAAI | 9 |
| 2024 | IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation
Yuanhao Zhai 0001, Chung-Ching Lin, Zhengyuan Yang, David S. Doermann, Junsong Yuan 0001, Zicheng Liu 0001 |
ECCV (15) | 7 |
| 2024 | ChartReformer: Natural Language-Driven Chart Image Editing
Pengyu Yan, Mahesh Bhosale, Jay Lal, Bikhyat Adhikari, David S. Doermann |
ICDAR (1) | 5 |
| 2024 | Learning 1-Bit Tiny Object Detector with Discriminative Feature Refinementabstract1-bit detectors show impressive performance comparable to their real-valued counterparts when detecting commonly sized objects while exhibiting significant performance degradation on tiny objects. The challenge stems from the fact that high-level features extracted by 1-bit convolutions seem less compelling to reveal the discriminative foreground features. To address these issues, we introduce a Discriminative Feature Refinement method for 1-bit Detectors (DFR-Det), aiming to enhance the discriminative ability of foreground representation for tiny objects in aerial images. This is accomplished by refining the feature representation using an information bottleneck (IB) to achieve a distinctive representation of tiny objects. Specifically, we introduce a new decoder with a foreground mask, aiming to enhance the discriminative ability of high-level features for the target but suppress the background impact. Additionally, our decoder is simple but effective and can be easily mounted on existing detectors without extra burden added to the inference procedure. Extensive experiments on various tiny object detection (TOD) tasks demonstrate DFR-Det’s superiority over state-of-the-art 1-bit detectors. For example, 1-bit FCOS achieved by DFR-Det achieves the 12.8% AP on AI-TOD dataset, approaching the performance of the real-valued counterpart. Sheng Xu 0007, Yanjing Li, Mingbao Lin, Baochang Zhang 0001, David S. Doermann |
ICML | 6 |
| 2024 | Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance DistillationabstractImage diffusion distillation achieves high-fidelity generation with very few sampling steps. However, directly applying these techniques to video models results in unsatisfied frame quality. This issue arises from the limited frame appearance quality in public video datasets, affecting the performance of both teacher and student video diffusion models. Our study aims to improve video diffusion distillation and meanwhile enabling the student model to improve frame appearance using the abundant high-quality image data. To this end, we propose motion consistency models (MCM), a single-stage video diffusion distillation method that disentangles motion and appearance learning. Specifically, MCM involves a video consistency model that distills motion from the video teacher model, and an image discriminator that boosts frame appearance to match high-quality image data. However, directly combining these components leads to two significant challenges: a conflict in frame learning objectives, where video distillation learns from low-quality video frames while the image discriminator targets high-quality images, and training-inference discrepancies due to the differing quality of video samples used during training and inference. To address these challenges, we introduce disentangled motion distillation and mixed trajectory distillation. The former applies the distillation objective solely to the motion representation, while the latter mitigates training-inference discrepancies by mixing distillation trajectories from both the low- and high-quality video domains. Extensive experiments show that our MCM achieves state-of-the-art video diffusion distillation performance. Additionally, our method can enhance frame quality in video diffusion models, producing frames with high aesthetic value or specific styles. Yuanhao Zhai 0001, Zhengyuan Yang, Chung-Ching Lin, David S. Doermann, Junsong Yuan 0001 |
NeurIPS | 7 |
| 2024 | Artemis: Towards Referential Understanding in Complex VideosabstractVideos carry rich visual information including object description, action, interaction, etc., but the existing multimodal large language models (MLLMs) fell short in referential understanding scenarios such as video-based referring. In this paper, we present Artemis, an MLLM that pushes video-based referential understanding to a finer level. Given a video, Artemis receives a natural-language question with a bounding box in any video frame and describes the referred target in the entire video. The key to achieving this goal lies in extracting compact, target-specific video features, where we set a solid baseline by tracking and selecting spatiotemporal features from the video. We train Artemis on the newly established ViderRef45K dataset with 45K video-QA pairs and design a computationally efficient, three-stage training procedure. Results are promising both quantitatively and qualitatively. Additionally, we show that Artemis can be integrated with video grounding and text summarization tools to understand more complex scenarios. Code and data are available at https://github.com/NeurIPS24Artemis/Artemis. Jihao Qiu, Lingxi Xie, Tianren Ma, Pengyu Yan, David S. Doermann, Qixiang Ye, Yunjie Tian |
NeurIPS | 7 |
| 2024 | Future of software development with generative AIabstractAbstract Generative AI is regarded as a major disruption to software development. Platforms, repositories, clouds, and the automation of tools and processes have been proven to improve productivity, cost, and quality. Generative AI, with its rapidly expanding capabilities, is a major step forward in this field. As a new key enabling technology, it can be used for many purposes, from creative dimensions to replacing repetitive and manual tasks. The number of opportunities increases with the capabilities of large-language models (LLMs). This has raised concerns about ethics, education, regulation, intellectual property, and even criminal activities. We analyzed the potential of generative AI and LLM technologies for future software development paths. We propose four primary scenarios, model trajectories for transitions between them, and reflect against relevant software development operations. The motivation for this research is clear: the software development industry needs new tools to understand the potential, limitations, and risks of generative AI, as well as guidelines for using it. Jaakko J. Sauvola, Sasu Tarkoma, Mika Klemettinen, Jukka Riekki, David S. Doermann |
Autom. Softw. Eng. | 5 |
| 2024 | Swin-chart: An efficient approach for chart classification
Anurag Dhote, Mohammed Javed, David S. Doermann |
Pattern Recognit. Lett. | 3 |
| 2023 | Progressive Multi-View Human Mesh Recovery with Self-SupervisionabstractTo date, little attention has been given to multi-view 3D human mesh estimation, despite real-life applicability (e.g., motion capture, sport analysis) and robustness to single-view ambiguities. Existing solutions typically suffer from poor generalization performance to new settings, largely due to the limited diversity of image/3D-mesh pairs in multi-view training data. To address this shortcoming, people have explored the use of synthetic images. But besides the usual impact of visual gap between rendered and target data, synthetic-data-driven multi-view estimators also suffer from overfitting to the camera viewpoint distribution sampled during training which usually differs from real-world distributions. Tackling both challenges, we propose a novel simulation-based training pipeline for multi-view human mesh recovery, which (a) relies on intermediate 2D representations which are more robust to synthetic-to-real domain gap; (b) leverages learnable calibration and triangulation to adapt to more diversified camera setups; and (c) progressively aggregates multi-view information in a canonical 3D space to remove ambiguities in 2D representations. Through extensive benchmarking, we demonstrate the superiority of the proposed solution especially for unseen in-the-wild scenarios. Liangchen Song, Meng Zheng 0002, Benjamin Planche, Terrence Chen, Junsong Yuan 0001, David S. Doermann, Ziyan Wu 0001 |
AAAI | 7 |
| 2023 | Towards Generic Image Manipulation Detection with Weakly-Supervised Self-Consistency LearningabstractAs advanced image manipulation techniques emerge, detecting the manipulation becomes increasingly important. Despite the success of recent learning-based approaches for image manipulation detection, they typically require expensive pixel-level annotations to train, while exhibiting degraded performance when testing on images that are differently manipulated compared with training images. To address these limitations, we propose weakly-supervised image manipulation detection, such that only binary image-level labels (authentic or tampered with) are required for training purpose. Such a weakly-supervised setting can leverage more training images and has the potential to adapt quickly to new manipulation techniques. To improve the generalization ability, we propose weakly-supervised self-consistency learning (WSCL) to leverage the weakly annotated images. Specifically, two consistency properties are learned: multi-source consistency (MSC) and inter-patch consistency (IPC). MSC exploits different content-agnostic information and enables cross-source learning via an online pseudo label generation and refinement process. IPC performs global pair-wise patch-patch relationship reasoning to discover a complete region of manipulation. Extensive experiments validate that our WSCL, even though is weakly supervised, exhibits competitive performance compared with fully-supervised counterpart under both in-distribution and out-of-distribution evaluations, as well as reasonable manipulation localization ability. Yuanhao Zhai 0001, Tianyu Luan, David S. Doermann, Junsong Yuan 0001 |
ICCV | 3 |
| 2023 | SOAR: Scene-debiasing Open-set Action RecognitionabstractDeep learning models have a risk of utilizing spurious clues to make predictions, such as recognizing actions based on the background scene. This issue can severely degrade the open-set action recognition performance when the testing samples have different scene distributions from the training samples. To mitigate this problem, we propose a novel method, called Scene-debiasing Open-set Action Recognition (SOAR), which features an adversarial scene reconstruction module and an adaptive adversarial scene classification module. The former prevents the decoder from reconstructing the video background given video features, and thus helps reduce the background information in feature learning. The latter aims to confuse scene type classification given video features, with a specific emphasis on the action foreground, and helps to learn scene-invariant information. In addition, we design an experiment to quantify the scene bias. The results indicate that the current open-set action recognizers are biased toward the scene, and our proposed SOAR method better mitigates such bias. Furthermore, our extensive experiments demonstrate that our method outperforms state-of-the-art methods, and the ablation studies confirm the effectiveness of our proposed modules. Yuanhao Zhai 0001, Ziyi Liu 0001, Zhenyu Wu 0002, Chunluan Zhou, David S. Doermann, Junsong Yuan 0001, Gang Hua 0001 |
ICCV | 6 |
| 2023 | SpaDen: Sparse and Dense Keypoint Estimation for Real-World Chart Understanding
Saleem Ahmed, Pengyu Yan, David S. Doermann, Srirangaraj Setlur, Venu Govindaraju |
ICDAR (2) | 3 |
| 2023 | LineFormer: Line Chart Data Extraction Using Instance Segmentation
Jay Lal, Aditya Mitkari, Mahesh Bhosale, David S. Doermann |
ICDAR (5) | 4 |
| 2023 | Context-Aware Chart Element Detection
Pengyu Yan, Saleem Ahmed, David S. Doermann |
ICDAR (1) | 3 |
| 2023 | Language-guided Human Motion Synthesis with Atomic ActionsabstractLanguage-guided human motion synthesis has been a challenging task due to the inherent complexity and diversity of human behaviors. Previous methods face limitations in generalization to novel actions, often resulting in unrealistic or incoherent motion sequences. In this paper, we propose ATOM (ATomic mOtion Modeling) to mitigate this problem, by decomposing actions into atomic actions, and employing a curriculum learning strategy to learn atomic action composition. First, we disentangle complex human motions into a set of atomic actions during learning, and then assemble novel actions using the learned atomic actions, which offers better adaptability to new actions. Moreover, we introduce a curriculum learning training strategy that leverages masked motion modeling with a gradual increase in the mask ratio, and thus facilitates atomic action assembly. This approach mitigates the overfitting problem commonly encountered in previous methods while enforcing the model to learn better motion representations. We demonstrate the effectiveness of ATOM through extensive experiments, including text-to-motion and action-to-motion synthesis tasks. We further illustrate its superiority in synthesizing plausible and coherent text-guided human motion sequences. Yuanhao Zhai 0001, Mingzhen Huang, Tianyu Luan, Lu Dong 0004, Ifeoma Nwogu, Siwei Lyu, David S. Doermann, Junsong Yuan 0001 |
ACM Multimedia | 7 |
| 2023 | Exploring the Knowledge Transferred by Response-Based Teacher-Student DistillationabstractResponse-based Knowledge Distillation refers to the technique of supervising the student network with the teacher networks' predictions. The method is motivated by observing that the predicted probabilities reflect the relation among labels, which is the knowledge to be transferred. This paper explores the transferred knowledge from a novel perspective: comparing the knowledge transferred through different teachers. Two intriguing properties are observed. First, higher confidence scores of teachers' predictions lead to better distillation results, and second, teachers' incorrectly predicted training samples should be kept for distillation. We then analyze the phenomenon by studying teachers' decision boundaries, of which some can help the student generalize while some may not. Based on the observations, we further propose an embarrassingly simple distillation framework named Efficient Distillation, which is effective on ImageNet with different teacher-student pairs: When using ResNet34 as the teacher, the student ResNet18 trained from scratch reaches 74.07% Top-1 accuracy within 98 GPU hours (RTX 3090), outperforming current state-of-the-art result (73.19%) by a large margin. Our code is available at https://github.com/lsongx/EffDstl. Liangchen Song, Helong Zhou, Qian Zhang 0009, David S. Doermann, Junsong Yuan 0001 |
ACM Multimedia | 6 |
| 2023 | Defending against Data-Free Model Extraction by Distributionally Robust Defensive TrainingabstractData-Free Model Extraction (DFME) aims to clone a black-box model without knowing its original training data distribution, making it much easier for attackers to steal commercial models. Defense against DFME faces several challenges: (i) effectiveness; (ii) efficiency; (iii) no prior on the attacker's query data distribution and strategy. However, existing defense methods: (1) are highly computation and memory inefficient; or (2) need strong assumptions about attack data distribution; or (3) can only delay the attack or prove a model theft after the model stealing has happened. In this work, we propose a Memory and Computation efficient defense approach, named MeCo, to prevent DFME from happening while maintaining the model utility simultaneously by distributionally robust defensive training on the target victim model. Specifically, we randomize the input so that it: (1) causes a mismatch of the knowledge distillation loss for attackers; (2) disturbs the zeroth-order gradient estimation; (3) changes the label prediction for the attack query data. Therefore, the attacker can only extract misleading information from the black-box model. Extensive experiments on defending against both decision-based and score-based DFME demonstrate that MeCo can significantly reduce the effectiveness of existing DFME methods and substantially improve running efficiency. Zhenyi Wang 0001, Li Shen 0008, Tongliang Liu, Tiehang Duan, Yanjun Zhu, Donglin Zhan, David S. Doermann, Mingchen Gao |
NeurIPS | 7 |
| 2023 | Few-Shot Learning with Complex-Valued Neural Networks and Dependable Learning
Runqi Wang, Baochang Zhang 0001, Guodong Guo, David S. Doermann |
Int. J. Comput. Vis. | 5 |
| 2023 | Anti-Bandit for Neural Architecture Search
Runqi Wang, Linlin Yang 0001, Wei Wang 0016, David S. Doermann, Baochang Zhang 0001 |
Int. J. Comput. Vis. | 5 |
| 2023 | Adaptive Two-Stream Consensus Network for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (W-TAL) aims to classify and localize all action instances in untrimmed videos under only video-level supervision. Without frame-level annotations, it is challenging for W-TAL methods to clearly distinguish actions and background, which severely degrades the action boundary localization and action proposal scoring. In this paper, we present an adaptive two-stream consensus network (A-TSCN) to address this problem. Our A-TSCN features an iterative refinement training scheme: a frame-level pseudo ground truth is generated and iteratively updated from a late-fusion activation sequence, and used to provide frame-level supervision for improved model training. Besides, we introduce an adaptive attention normalization loss, which adaptively selects action and background snippets according to video attention distribution. By differentiating the attention values of the selected action snippets and background snippets, it forces the predicted attention to act as a binary selection and promotes the precise localization of action boundaries. Furthermore, we propose a video-level and a snippet-level uncertainty estimator, and they can mitigate the adverse effect caused by learning from noisy pseudo ground truth. Experiments conducted on the THUMOS14, ActivityNet v1.2, ActivityNet v1.3, and HACS datasets show that our A-TSCN outperforms current state-of-the-art methods, and even achieves comparable performance with several fully-supervised methods. Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Nanning Zheng 0001, David S. Doermann, Junsong Yuan 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Federated Learning With Privacy-Preserving Ensemble Attention DistillationabstractFederated Learning (FL) is a machine learning paradigm where many local nodes collaboratively train a central model while keeping the training data decentralized. This is particularly relevant for clinical applications since patient data are usually not allowed to be transferred out of medical facilities, leading to the need for FL. Existing FL methods typically share model parameters or employ co-distillation to address the issue of unbalanced data distribution. However, they also require numerous rounds of synchronized communication and, more importantly, suffer from a privacy leakage risk. We propose a privacy-preserving FL framework leveraging unlabeled public data for one-way offline knowledge distillation in this work. The central model is learned from local knowledge via ensemble attention distillation. Our technique uses decentralized and heterogeneous local data like existing FL approaches, but more importantly, it significantly reduces the risk of privacy leakage. We demonstrate that our method achieves very competitive performance with more robust privacy preservation based on extensive experiments on image classification, segmentation, and reconstruction tasks. Liangchen Song, Rishi Vedula, Meng Zheng 0002, Benjamin Planche, Arun Innanje, Terrence Chen, Junsong Yuan 0001, David S. Doermann, Ziyan Wu 0001 |
IEEE Trans. Medical Imaging | 10 |
| 2022 | Preserving Privacy in Federated Learning with Ensemble Cross-Domain Knowledge DistillationabstractFederated Learning (FL) is a machine learning paradigm where local nodes collaboratively train a central model while the training data remains decentralized. Existing FL methods typically share model parameters or employ co-distillation to address the issue of unbalanced data distribution. However, they suffer from communication bottlenecks. More importantly, they risk privacy leakage risk. In this work, we develop a privacy preserving and communication efficient method in a FL framework with one-shot offline knowledge distillation using unlabeled, cross-domain, non-sensitive public data. We propose a quantized and noisy ensemble of local predictions from completely trained local models for stronger privacy guarantees without sacrificing accuracy. Based on extensive experiments on image classification and text classification tasks, we show that our method outperforms baseline FL algorithms with superior performance in both accuracy and data privacy preservation. Srikrishna Karanam, Ziyan Wu 0001, Terrence Chen, David S. Doermann, Arun Innanje |
AAAI | 6 |
| 2022 | MagFormer: Hybrid Video Motion Magnification Transformer from Eulerian and Lagrangian Perspectives
Sicheng Gao, Yutang Feng, Linlin Yang 0001, Xuhui Liu, David S. Doermann, Baochang Zhang 0001 |
BMVC | 6 |
| 2022 | Self-supervised Human Mesh Recovery with Cross-Representation Alignment
Meng Zheng 0002, Benjamin Planche, Srikrishna Karanam, Terrence Chen, David S. Doermann, Ziyan Wu 0001 |
ECCV (1) | 6 |
| 2022 | PREF: Predictability Regularized Neural Motion Fields
Liangchen Song, Benjamin Planche, Meng Zheng 0002, David S. Doermann, Junsong Yuan 0001, Terrence Chen, Ziyan Wu 0001 |
ECCV (22) | 5 |
| 2022 | Uncertainty Learning towards Unsupervised Deformable Medical Image RegistrationabstractUncertainty estimation in medical image registration enables surgeons to evaluate the operative risk based on the trustworthiness of the registered image data thus of paramount importance for practical clinical applications. Despite the recent promising results obtained with deep unsupervised learning-based registration methods, reasoning about uncertainty of unsupervised registration models remains largely unexplored. In this work, we propose a predictive module to learn the registration and uncertainty in correspondence simultaneously. Our framework introduces empirical randomness and registration error based uncertainty prediction. We systematically assess the performances on two MRI datasets with different ensemble paradigms. Experimental results highlight that our proposed framework significantly improves the registration accuracy and uncertainty compared with the baseline. Luckyson Khaidem, Wentao Zhu 0001, Baochang Zhang 0001, David S. Doermann |
WACV | 5 |
| 2022 | Towards Compact 1-bit CNNs via Bayesian Learning
Junhe Zhao, Sheng Xu 0007, Baochang Zhang 0001, Jiaxin Gu, David S. Doermann, Guodong Guo |
Int. J. Comput. Vis. | 5 |
| 2022 | A Knowledge Enforcement Network-Based Approach for Classifying a Photographer's ImagesabstractClassification of photos captured by different photographers is an important and challenging problem in knowledge-based and image processing. Monitoring and authenticating images uploaded on social media are essential, and verifying the source is one key piece of evidence. We present a novel framework for classifying photos of different photographers based on the combination of local features and deep learning models. The proposed work uses focused and defocused information in the input images to extract contextual information. The model estimates the weighted gradient and calculates entropy to strengthen context features. The focused and defocused information is fused to estimate cross-covariance and define a linear relationship between them. This relationship results in a feature matrix fed to Knowledge Enforcement Network (KEN) for obtaining representative features. Due to the strong discriminative ability of deep learning models, we employ the lightweight and accurate MobileNetV2. The output of KEN and MobileNetV2 is sent to a classifier for photographer classification. Experimental results of the proposed model on our dataset of 46 photographer classes (46234 images) and publicly available datasets of 41 photographer classes (218303 images) show that the method outperforms the existing techniques by 5%–10% on average. The dataset created for the experimental purpose will be made available upon publication. Palaiahnakote Shivakumara, Pinaki Nath Chowdhury, Umapada Pal 0001, David S. Doermann, Ramachandra Raghavendra, Tong Lu 0002, Michael Blumenstein |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2022 | Data-adaptive binary neural networks for efficient object detection and recognition
Junhe Zhao, Sheng Xu 0007, Runqi Wang, Baochang Zhang 0001, Guodong Guo, David S. Doermann, Dianmin Sun |
Pattern Recognit. Lett. | 6 |
| 2022 | iffDetector: Inference-Aware Feature Filtering for Object DetectionabstractModern convolutional neural network (CNN)-based object detectors focus on feature configuration during training but often ignore feature optimization during inference. In this article, we propose a new feature optimization approach to enhance features and suppress background noise in both the training and inference stages. We introduce a generic inference-aware feature filtering (IFF) module that can be easily combined with existing detectors, resulting in our iffDetector. Unlike conventional open-loop feature calculation approaches without feedback, the proposed IFF module performs the closed-loop feature optimization by leveraging high-level semantics to enhance the convolutional features. By applying the Fourier transform to analyze our detector, we prove that the IFF module acts as a negative feedback that can theoretically guarantee the stability of the feature learning. IFF can be fused with CNN-based object detectors in a plug-and-play manner with little computational cost overhead. Experiments on the PASCAL VOC and MS COCO datasets demonstrate that our iffDetector consistently outperforms state-of-the-art methods with significant margins. Mingyuan Mao, Baochang Zhang 0001, Qixiang Ye, Wanquan Liu, David S. Doermann |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | BiRe-ID: Binary Neural Network for Efficient Person Re-IDabstractPerson re-identification (Re-ID) has been promoted by the significant success of convolutional neural networks (CNNs). However, the application of such CNN-based Re-ID methods depends on the tremendous consumption of computation and memory resources, which affects its development on resource-limited devices such as next generation AI chips. As a result, CNN binarization has attracted increasing attention, which leads to binary neural networks (BNNs). In this article, we propose a new BNN-based framework for efficient person Re-ID (BiRe-ID). In this work, we discover that the significant performance drop of binarized models for Re-ID task is caused by the degraded representation capacity of kernels and features. To address the issues, we propose the kernel and feature refinement based on generative adversarial learning (KR-GAL and FR-GAL) to enhance the representation capacity of BNNs. We first introduce an adversarial attention mechanism to refine the binarized kernels based on their real-valued counterparts. Specifically, we introduce a scale factor to restore the scale of 1-bit convolution. And we employ an effective generative adversarial learning method to train the attention-aware scale factor. Furthermore, we introduce a self-supervised generative adversarial network to refine the low-level features using the corresponding high-level semantic information. Extensive experiments demonstrate that our BiRe-ID can be effectively implemented on various mainstream backbones for the Re-ID task. In terms of the performance, our BiRe-ID surpasses existing binarization methods by significant margins, at the level even comparable with the real-valued counterparts. For example, on Market-1501, BiRe-ID achieves 64.0% mAP on ResNet-18 backbone, with an impressive 12.51× speedup in theory and 11.75× storage saving. In particular, the KR-GAL and FR-GAL methods show strong generalization on multiple tasks such as Re-ID, image classification, object detection, and 3D point cloud processing. Sheng Xu 0007, Baochang Zhang 0001, Jinhu Lü 0001, Guodong Guo, David S. Doermann |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2021 | Layer-Wise Searching for 1-Bit Detectorsabstract1-bit detectors show great promise for resource-constrained embedded devices but often suffer from a significant performance gap compared with their real-valued counterparts. The primary reason lies in the error during binarization. This paper presents a layer-wise searching (LWS) strategy to generate 1-bit detectors that maintain a performance very close to the original real-valued model. The approach introduces angular and amplitude loss functions to increase detector capacity. At 1-bit layers, it exploits a differentiable binarization search (DBS) to minimize the angular error in a student-teacher framework. We also learn the scale factor by minimizing the amplitude loss in the same student-teacher framework. Extensive experiments show that LWS-Det outperforms state-of-the-art 1-bit detectors by a considerable margin on the PASCAL VOC and COCO datasets. For example, the LWS-Det achieves 1-bit Faster-RCNN with ResNet-34 backbone within 2.0% mAP of its real-valued counterpart on the PASCAL VOC dataset. Sheng Xu 0007, Junhe Zhao, Jinhu Lü 0001, Baochang Zhang 0001, Shumin Han, David S. Doermann |
CVPR | 6 |
| 2021 | HH-CompWordNet: Holistic Handwritten Word Recognition in the Compressed DomainabstractHolistic word recognition in handwritten documents is an important research topic in the field of Document Image Analysis. For some applications, given strong language models, it can be more robust and computationally less expensive than character segmentation and recognition. This paper presents HH-CompWordNet, a novel approach to applying a Convolutional Neural Network (CNN) to directly to the DCT coefficients of the compressed domain word images. The efficacy of the HH-CompWordNet is demonstrated with the JPEG compressed version of the CMATERdb2.1.2.1 dataset - standard Handwritten Bangla word images. Experiments show the system obtains state-of-the-art accuracy of 86.80%. Bulla Rajesh, Priyanshu Jain, Mohammed Javed, David S. Doermann |
DCC | 4 |
| 2021 | Ensemble Attention Distillation for Privacy-Preserving Federated LearningabstractWe consider the problem of Federated Learning (FL) where numerous decentralized computational nodes collaborate with each other to train a centralized machine learning model without explicitly sharing their local data samples. Such decentralized training naturally leads to issues of imbalanced or differing data distributions among the local models and challenges in fusing them into a central model. Existing FL methods deal with these issues by either sharing local parameters or fusing models via online distillation. However, such a design leads to multiple rounds of inter-node communication resulting in substantial band-width consumption, while also increasing the risk of data leakage and consequent privacy issues. To address these problems, we propose a new distillation-based FL frame-work that can preserve privacy by design, while also consuming substantially less network communication resources when compared to the current methods. Our framework engages in inter-node communication using only publicly available and approved datasets, thereby giving explicit privacy control to the user. To distill knowledge among the various local models, our framework involves a novel ensemble distillation algorithm that uses both final prediction as well as model attention. This algorithm explicitly considers the diversity among various local nodes while also seeking consensus among them. This results in a comprehensive technique to distill knowledge from various decentralized nodes. We demonstrate the various aspects and the associated benefits of our FL framework through extensive experiments that produce state-of-the-art results on both classification and segmentation tasks on natural and medical images. Srikrishna Karanam, Ziyan Wu 0001, Terrence Chen, David S. Doermann, Arun Innanje |
ICCV | 6 |
| 2021 | IDARTS: Interactive Differentiable Architecture SearchabstractDifferentiable Architecture Search (DARTS) improves the efficiency of architecture search by learning the architecture and network parameters end-to-end. However, the intrinsic relationship between the architecture’s parameters is neglected, leading to a sub-optimal optimization process. The reason lies in the fact that the gradient descent method used in DARTS ignores the coupling relationship of the parameters and therefore degrades the optimization. In this paper, we address this issue by formulating DARTS as a bi-linear optimization problem and introducing an Interactive Differentiable Architecture Search (IDARTS). We first develop a backtracking backpropagation process, which can decouple the relationships of different kinds of parameters and train them in the same framework. The backtracking method coordinates the training of different parameters that fully explore their interaction and optimize training. We present experiments on the CIFAR10 and ImageNet datasets that demonstrate the efficacy of the IDARTS approach by achieving a top-1 accuracy of 76.52% on ImageNet without additional search cost vs. 75.8% with the state-of-the-art PC-DARTS. Runqi Wang, Baochang Zhang 0001, Tian Wang 0002, Guodong Guo, David S. Doermann |
ICCV | 6 |
| 2021 | Scalable Coverage Path Planning of Multi-Robot Teams for Monitoring Non-Convex AreasabstractThis paper presents a novel multi-robot coverage path planning (CPP) algorithm - aka SCoPP - that provides a time-efficient solution, with workload balanced plans for each robot in a multi-robot system, based on their initial states. This algorithm accounts for discontinuities (e.g., no-fly zones) in a specified area of interest, and provides an optimized ordered list of way-points per robot using a discrete, computationally efficient, nearest neighbor path planning algorithm. This algorithm involves five main stages, which include the transformation of the user’s input as a set of vertices in geographical coordinates, discretization, load-balanced partitioning, auctioning of conflict cells in a discretized space, and a path planning procedure. To evaluate the effectiveness of the primary algorithm, a multi-unmanned aerial vehicle (UAV) post-flood assessment application is considered, and the performance of the algorithm is tested on three test maps of varying sizes. Additionally, our method is compared with a state-of-the-art method created by Guasella et al. Further analyses on scalability and computational time of SCoPP are conducted. The results show that SCoPP is superior in terms of mission completion time; its computing time is found to be under 2 mins for a large map covered by a 150-robot team, thereby demonstrating its computationally scalability. Leighton Collins, Payam Ghassemi, Ehsan Tarkesh Esfahani, David S. Doermann, Karthik Dantu, Souma Chowdhury |
ICRA | 4 |
| 2021 | Self-Supervised Learning for Monocular Depth Estimation on Minimally Invasive Surgery ScenesabstractSelf-supervised learning algorithms that compute depth map from monocular videos have achieved remarkable performance on urban scenes and have been applied extensively. These techniques still face significant challenges, however, when applied directly to endoscopic videos because of the brightness variations from frame to frame and inadequate representation learning during the training phase. Inspired by the optical flow for motion alignment between adjacent frames, we design a AFNet with structural stability loss and residual-based smoothness loss to learn the appearance flow across adjacent frames, which handles the brightness inconsistency issue efficaciously. In addition, we propose a novel self-attention mechanism named feature scaling module to alleviate the inadequate representation learning problem. In a comparison study to the current state-of-the-art self-supervised methods explored for urban videos on the SCARED dataset, the developed model surpasses existing methods by a large margin. Shuwei Shao, Zhongcai Pei, Weihai Chen, Baochang Zhang 0001, Xingming Wu, Dianmin Sun, David S. Doermann |
ICRA | 7 |
| 2021 | Uncertainty-aware Binary Neural NetworksabstractBinary Neural Networks (BNN) are promising machine learning solutions for deployment on resource-limited devices. Recent approaches to training BNNs have produced impressive results, but minimizing the drop in accuracy from full precision networks is still challenging. One reason is that conventional BNNs ignore the uncertainty caused by weights that are near zero, resulting in the instability or frequent flip while learning. In this work, we investigate the intrinsic uncertainty of vanishing near-zero weights, making the training vulnerable to instability. We introduce an uncertainty-aware BNN (UaBNN) by leveraging a new mapping function called certainty-sign (c-sign) to reduce these weights' uncertainties. Our c-sign function is the first to train BNNs with a decreasing uncertainty for binarization. The approach leads to a controlled learning process for BNNs. We also introduce a simple but effective method to measure the uncertainty-based on a Gaussian function. Extensive experiments demonstrate that our method improves multiple BNN methods by maintaining stability of training, and achieves a higher performance over prior arts. Junhe Zhao, Linlin Yang 0001, Baochang Zhang 0001, Guodong Guo, David S. Doermann |
IJCAI | 5 |
| 2021 | Using Physiological Information to Classify Task Difficulty in Human-Swarm InteractionabstractHuman-swarm interaction has recently gained attention due to its plethora of new applications in disaster relief, surveillance, rescue, and exploration. However, if the task difficulty increases, the performance of the human operator decreases, thereby decreasing the overall efficacy of the human-swarm team. Thus, it is critical to identify the task difficulty and adaptively allocate the task to the human operator to maintain optimal performance. In this direction, we study the classification of task difficulty in a human-swarm interaction experiment performing a target search mission. The human may control platoons of unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs) to search a partially observable environment during the target search mission. The mission complexity is increased by introducing adversarial teams that humans may only see when the environment is explored. While the human is completing the mission, their brain activity is recorded using an electroencephalogram (EEG), which is used to classify the task difficulty. We have used two different approaches for classification: A feature-based approach using coherence values as input and a deep learning-based approach using raw EEG as input. Both approaches can classify the task difficulty well above the chance. The results showed the importance of the occipital lobe (O1 and O2) coherence feature with the other brain regions. Moreover, we also study individual differences (expert vs. novice) in the classification results. The analysis revealed that the temporal lobe in experts (T4 and T3) is predominant for task difficulty classification compared with novices. Joseph P. Distefano, Hemanth Manjunatha, Souma Chowdhury, Karthik Dantu, David S. Doermann, Ehsan Tarkesh Esfahani |
SMC | 5 |
| 2021 | Deformable Gabor Feature Networks for Biomedical Image ClassificationabstractIn recent years, deep learning has dominated progress in the field of medical image analysis. We find however, that the ability of current deep learning approaches to represent the complex geometric structures of many medical images is insufficient. One limitation is that deep learning models require a tremendous amount of data, and it is very difficult to obtain a sufficient amount with the necessary detail. A second limitation is that there are underlying features of these medical images that are well established, but the black-box nature of existing convolutional neural networks (CNNs) do not allow us to exploit them. In this paper, we revisit Gabor filters and introduce a deformable Gabor convolution (DGConv) to expand deep networks interpretability and enable complex spatial variations. The features are learned at deformable sampling locations with adaptive Gabor convolutions to improve representitiveness and robustness to complex objects. The DGConv replaces standard convolutional layers and is easily trained end-to-end, resulting in deformable Gabor feature network (DGFN) with few additional parameters and minimal additional training cost. We introduce DGFN for addressing deep multi-instance multi-label classification on the INbreast dataset for mammograms and on the ChestX-ray14 dataset for pulmonary x-ray images. Xin Xia 0005, Wentao Zhu 0001, Baochang Zhang 0001, David S. Doermann, Lian Zhuo |
WACV | 5 |
| 2021 | Style Consistent Image Generation for Nuclei Instance SegmentationabstractIn medical image analysis, one limitation of the application of machine learning is the insufficient amount of data with detailed annotation, due primarily to high cost. Another impediment is the domain gap observed between images from different organs and different collections. The differences are even more challenging for the nuclei instance segmentation, where images have significant nuclei stain distribution variations and complex pleomorphisms (sizes and shapes). In this work, we generate style consistent histopathology images for nuclei instance segmentation. We set up a novel instance segmentation framework that integrates a generator and discriminator into the segmentation pipeline with adversarial training to generalize nuclei instances and texture patterns. A segmentation net detects and segments both real nuclei and synthetic nuclei and provides feedback so that the generator can synthesize images that can boost the segmentation performance. Experimental results on three public nuclei datasets indicate that our proposed method outperforms previous nuclei segmentation methods. Baochang Zhang 0001, David S. Doermann |
WACV | 4 |
| 2021 | Binarized Neural Architecture Search for Efficient Object Recognition
Lian Zhuo, Baochang Zhang 0001, Xiawu Zheng, Jianzhuang Liu, Rongrong Ji, David S. Doermann, Guodong Guo |
Int. J. Comput. Vis. | 7 |
| 2021 | Rectified Binary Convolutional Networks with Generative Adversarial Learning
Chunlei Liu 0001, Wenrui Ding, Baochang Zhang 0001, Jianzhuang Liu, Guodong Guo, David S. Doermann |
Int. J. Comput. Vis. | 7 |
| 2021 | MIDPhyNet: Memorized infusion of decomposed physics in neural networks to model dynamic systems
Zhibo Zhang 0003, Rahul Rai, Souma Chowdhury, David S. Doermann |
Neurocomputing | 4 |
| 2021 | Chart Mining: A Survey of Methods for Automated Chart AnalysisabstractCharts are useful communication tools for the presentation of data in a visually appealing format that facilitates comprehension. There have been many studies dedicated to chart mining, which refers to the process of automatic detection, extraction and analysis of charts to reproduce the tabular data that was originally used to create them. By allowing access to data which might not be available in other formats, chart mining facilitates the creation of many downstream applications. This paper presents a comprehensive survey of approaches across all components of the automated chart mining pipeline, such as (i) automated extraction of charts from documents; (ii) processing of multi-panel charts; (iii) automatic image classifiers to collect chart images at scale; (iv) automated extraction of data from each chart image, for popular chart types as well as selected specialized classes; (v) applications of chart mining; and (vi) datasets for training and evaluation, and the methods that were used to build them. Finally, we summarize the main trends found in the literature and provide pointers to areas for further research in chart mining. Kenny Davila, Srirangaraj Setlur, David S. Doermann, Bhargava Urala Kota, Venu Govindaraju |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Learning modulation filter networks for weak signal detection in noise
Duona Zhang, Wenrui Ding, Baochang Zhang 0001, Chunhui Liu 0004, Jungong Han, David S. Doermann |
Pattern Recognit. | 6 |
| 2020 | Binarized Neural Architecture SearchabstractNeural architecture search (NAS) can have a significant impact in computer vision by automatically designing optimal neural network architectures for various tasks. A variant, binarized neural architecture search (BNAS), with a search space of binarized convolutions, can produce extremely compressed models. Unfortunately, this area remains largely unexplored. BNAS is more challenging than NAS due to the learning inefficiency caused by optimization requirements and the huge architecture space. To address these issues, we introduce channel sampling and operation space reduction into a differentiable NAS to significantly reduce the cost of searching. This is accomplished through a performance-based strategy used to abandon less potential operations. Two optimization methods for binarized neural networks are used to validate the effectiveness of our BNAS. Extensive experiments demonstrate that the proposed BNAS achieves a performance comparable to NAS on both CIFAR and ImageNet databases. An accuracy of 96.53% vs. 97.22% is achieved on the CIFAR-10 dataset, but with a significantly compressed model, and a 40% faster search than the state-of-the-art PC-DARTS. Lian Zhuo, Baochang Zhang 0001, Xiawu Zheng, Jianzhuang Liu, David S. Doermann, Rongrong Ji |
AAAI | 6 |
| 2020 | Cogradient Descent for Bilinear OptimizationabstractConventional learning methods simplify the bilinear model by regarding two intrinsically coupled factors independently, which degrades the optimization procedure. One reason lies in the insufficient training due to the asynchronous gradient descent, which results in vanishing gradients for the coupled variables. In this paper, we introduce a Cogradient Descent algorithm (CoGD) to address the bilinear problem, based on a theoretical framework to coordinate the gradient of hidden variables via a projection function. We solve one variable by considering its coupling relationship with the other, leading to a synchronous gradient descent to facilitate the optimization procedure. Our algorithm is applied to solve problems with one variable under the sparsity constraint, which is widely used in the learning paradigm. We validate our CoGD considering an extensive set of applications including image reconstruction, inpainting, and network pruning. Experiments show that it improves the state-of-the-art by a significant margin. Lian Zhuo, Baochang Zhang 0001, Linlin Yang 0001, Qixiang Ye, David S. Doermann, Rongrong Ji, Guodong Guo |
CVPR | 6 |
| 2020 | Anti-bandit Neural Architecture Search for Model Defense
Baochang Zhang 0001, Hong Liu 0009, Rongrong Ji, David S. Doermann |
ECCV (13) | 7 |
| 2020 | NAS-Count: Counting-by-Density with Neural Architecture Search
Yutao Hu 0002, Xuhui Liu, Baochang Zhang 0001, Jungong Han, Xianbin Cao 0001, David S. Doermann |
ECCV (22) | 7 |
| 2020 | What and How? Jointly Forecasting Human Action and PoseabstractForecasting human actions and motion trajectories address the problem of predicting what a person is going to do next and how they will perform it. This is crucial in a wide range of applications, such as assisted living and future co-robotic settings. We propose to simultaneously learn actions and action-related human motion dynamics while existing works perform them independently. This paper presents a method to jointly forecast categories of human action and skeletal joint pose, allowing the two tasks to reinforce each other. As a result, our system can predict future actions and the motion trajectories that will result. To achieve this, we define a task of joint action classification and pose regression. We employ a sequence to sequence encoder-decoder model combined with multi-task learning to forecast future actions and poses progressively before the action happens. Experimental results on two public datasets, IkeaDB and OAD, demonstrate the effectiveness of the proposed method. Yanjun Zhu, David S. Doermann, Qiong Liu 0003, Andreas Girgensohn |
ICPR | 2 |
| 2020 | CP-NAS: Child-Parent Neural Architecture Search for 1-bit CNNsabstractNeural architecture search (NAS) proves to be among the best approaches for many tasks by generating an application-adaptive neural architectures, which are still challenged by high computational cost and memory consumption. At the same time, 1-bit convolutional neural networks (CNNs) with binarized weights and activations show their potential for resource-limited embedded devices. One natural approach is to use 1-bit CNNs to reduce the computation and memory cost of NAS by taking advantage of the strengths of each in a unified framework. To this end, a Child-Parent model is introduced to a differentiable NAS to search the binarized architecture(Child) under the supervision of a full-precision model (Parent). In the search stage, the Child-Parent model uses an indicator generated by the parent and child model accuracy to evaluate the performance and abandon operations with less potential. In the training stage, a kernel level CP loss is introduced to optimize the binarized network. Extensive experiments demonstrate that the proposed CP-NAS achieves a comparable accuracy with traditional NAS on both the CIFAR and ImageNet databases. It achieves an accuracy of 95.27% on CIFAR-10, 64.3% on ImageNet with binarized weights and activations, and a 30% faster search than prior arts. Lian Zhuo, Baochang Zhang 0001, Linlin Yang 0001, Chen Chen 0001, Yanjun Zhu, David S. Doermann |
IJCAI | 7 |
| 2020 | Using Physiological Measurements to Analyze the Tactical Decisions in Human Swarm TeamsabstractHuman-Swarm interaction has attracted a lot of attention for their applications in areas such as exploration, rescue, surveillance, and interplanetary exploration. When humans assume a supervisory or tactician role in managing the robot swarm, the humans' (physiological) state significantly affects the mission performance. In this work, we explore the physiological correlates with the user's tactical decisions in a simulated search and rescue mission. The mission consists of supervising three groups of unmanned aerial vehicles and three groups of unmanned ground vehicles to search for a target building. The mission complexity is increased by introducing static adversarial teams.Due to the adversarial team's presence, the user should employ different tactics to search for a target. While the user interacts with the swarm, brain activity in forms of electroencephalogram (EEG) and eye movements are recorded. 20 participants, with prior experience in playing real-time strategy games, took part in the study. A linear mixed effect model is used to study the correlated physiological features and tactical decisions. Six features are extracted from the physiological data: engagement level, mental workload, Fz-Pz coherence, Fz-O1 coherence, pupil size, and the number of gaze fixations. The results show that mental engagement and Fz-O1 coherence are the important factors in predicting the tactical decisions. Specifically, Fz-O1 coherence in Beta (22.5-30 Hz) and Gamma (38-42 Hz) band is found to be significant. Hemanth Manjunatha, Joseph P. Distefano, Apurv Jani, Payam Ghassemi, Souma Chowdhury, Karthik Dantu, David S. Doermann, Ehsan Tarkesh Esfahani |
SMC | 7 |
| 2020 | Orthogonal Features Fusion Network for Anomaly DetectionabstractGenerative models have been successfully used for anomaly detection, which however need a large number of parameters and computation overheads, especially when training spatial and temporal networks in the same framework. In this paper, we introduce a novel network architecture, Orthogonal Features Fusion Network (OFF-Net), to solve the anomaly detection problem. We show that the convolutional feature maps used for generating future frames are orthogonal with each other, which can improve representation capacity of generative models and strengthen temporal connections between adjacent images. We lead a simple but effective module easily mounted on convolutional neural networks (CNNs) with negligible additional parameters added, which can replace the widely-used optical flow network a nd significantly improve th e performance for anomaly detection. Extensive experiment results demonstrate the effectiveness of OFF-Net that we outperform the state-of-the-art model 1.7% in terms of AUC. We save around 85M-space parameters compared with the prevailing prior arts using optical flow network without comprising the performance. Teli Ma, Jinxin Shao, Baochang Zhang 0001, David S. Doermann |
VCIP | 5 |
| 2019 | Projection Convolutional Neural Networks for 1-bit CNNs via Discrete Back PropagationabstractThe advancement of deep convolutional neural networks (DCNNs) has driven significant improvement in the accuracy of recognition systems for many computer vision tasks. However, their practical applications are often restricted in resource-constrained environments. In this paper, we introduce projection convolutional neural networks (PCNNs) with a discrete back propagation via projection (DBPP) to improve the performance of binarized neural networks (BNNs). The contributions of our paper include: 1) for the first time, the projection function is exploited to efficiently solve the discrete back propagation problem, which leads to a new highly compressed CNNs (termed PCNNs); 2) by exploiting multiple projections, we learn a set of diverse quantized kernels that compress the full-precision kernels in a more efficient way than those proposed previously; 3) PCNNs achieve the best classification performance compared to other state-ofthe-art BNNs on the ImageNet and CIFAR datasets. Jiaxin Gu, Ce Li 0002, Baochang Zhang 0001, Jungong Han, Xianbin Cao 0001, Jianzhuang Liu, David S. Doermann |
AAAI | 7 |
| 2019 | Calibrated Stochastic Gradient Descent for Convolutional Neural NetworksabstractIn stochastic gradient descent (SGD) and its variants, the optimized gradient estimators may be as expensive to compute as the true gradient in many scenarios. This paper introduces a calibrated stochastic gradient descent (CSGD) algorithm for deep neural network optimization. A theorem is developed to prove that an unbiased estimator for the network variables can be obtained in a probabilistic way based on the Lipschitz hypothesis. Our work is significantly distinct from existing gradient optimization methods, by providing a theoretical framework for unbiased variable estimation in the deep learning paradigm to optimize the model parameter calculation. In particular, we develop a generic gradient calibration layer which can be easily used to build convolutional neural networks (CNNs). Experimental results demonstrate that CNNs with our CSGD optimization scheme can improve the stateof-the-art performance for natural image classification, digit recognition, ImageNet object classification, and object detection tasks. This work opens new research directions for developing more efficient SGD updates and analyzing the backpropagation algorithm. Lian Zhuo, Baochang Zhang 0001, Chen Chen 0001, Qixiang Ye, Jianzhuang Liu, David S. Doermann |
AAAI | 6 |
| 2019 | Crowd Counting and Density Estimation by Trellis Encoder-Decoder NetworksabstractCrowd counting has recently attracted increasing interest in computer vision but remains a challenging problem. In this paper, we propose a trellis encoder-decoder network (TEDnet) for crowd counting, which focuses on generating high-quality density estimation maps. The major contributions are four-fold. First, we develop a new trellis architecture that incorporates multiple decoding paths to hierarchically aggregate features at different encoding stages, which improves the representative capability of convolutional features for large variations in objects. Second, we employ dense skip connections interleaved across paths to facilitate sufficient multi-scale feature fusions, which also helps TEDnet to absorb the supervision information. Third, we propose a new combinatorial loss to enforce similarities in local coherence and spatial correlation between maps. By distributedly imposing this combinatorial loss on intermediate outputs, TEDnet can improve the back-propagation process and alleviate the gradient vanishing problem. Finally, on four widely-used benchmarks, our TEDnet achieves the best overall performance in terms of both density map quality and counting accuracy, with an improvement up to 14% in MAE metric. These results validate the effectiveness of TEDnet for crowd counting. Zehao Xiao, Baochang Zhang 0001, Xiantong Zhen, Xianbin Cao 0001, David S. Doermann, Ling Shao 0001 |
CVPR | 6 |
| 2019 | Exploiting Kernel Sparsity and Entropy for Interpretable CNN CompressionabstractCompressing convolutional neural networks (CNNs) has received ever-increasing research focus. However, most existing CNN compression methods do not interpret their inherent structures to distinguish the implicit redundancy. In this paper, we investigate the problem of CNN compression from a novel interpretable perspective. The relationship between the input feature maps and 2D kernels is revealed in a theoretical framework, based on which a kernel sparsity and entropy (KSE) indicator is proposed to quantitate the feature map importance in a feature-agnostic manner to guide model compression. Kernel clustering is further conducted based on the KSE indicator to accomplish high-precision CNN compression. KSE is capable of simultaneously compressing each layer in an efficient way, which is significantly faster compared to previous data-driven feature map pruning methods. We comprehensively evaluate the compression and speedup of the proposed method on CIFAR-10, SVHN and ImageNet 2012. Our method demonstrates superior performance gains over previous ones. In particular, it achieves 4.7× FLOPs reduction and 2.9× compression on ResNet-50 with only a top-5 accuracy drop of 0.35% on ImageNet 2012, which significantly outperforms state-of-the-art methods. Shaohui Lin, Baochang Zhang 0001, Jianzhuang Liu, David S. Doermann, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
CVPR | 5 |
| 2019 | Towards Optimal Structured CNN Pruning via Generative Adversarial LearningabstractStructured pruning of filters or neurons has received increased focus for compressing convolutional neural networks. Most existing methods rely on multi-stage optimizations in a layer-wise manner for iteratively pruning and retraining which may not be optimal and may be computation intensive. Besides, these methods are designed for pruning a specific structure, such as filter or block structures without jointly pruning heterogeneous structures. In this paper, we propose an effective structured pruning approach that jointly prunes filters as well as other structures in an end-to-end manner. To accomplish this, we first introduce a soft mask to scale the output of these structures by defining a new objective function with sparsity regularization to align the output of baseline and network with this mask. We then effectively solve the optimization problem by generative adversarial learning (GAL), which learns a sparse soft mask in a label-free and an end-to-end manner. By forcing more scale factors in the soft mask to zero, the fast iterative shrinkage-thresholding algorithm (FISTA) can be leveraged to fast and reliably remove the corresponding structures. Extensive experiments demonstrate the effectiveness of GAL on different datasets, including MNIST, CIFAR-10 and ImageNet ILSVRC 2012. For example, on ImageNet ILSVRC 2012, the pruned ResNet-50 achieves 10.88% Top-5 error and results in a factor of 3.7x speedup. This significantly outperforms state-of-the-art methods. Shaohui Lin, Rongrong Ji, Chenqian Yan, Baochang Zhang 0001, Liujuan Cao, Qixiang Ye, Feiyue Huang, David S. Doermann |
CVPR | 8 |
| 2019 | Circulant Binary Convolutional Networks: Enhancing the Performance of 1-Bit DCNNs With Circulant Back PropagationabstractThe rapidly decreasing computation and memory cost has recently driven the success of many applications in the field of deep learning. Practical applications of deep learning in resource-limited hardware, such as embedded devices and smart phones, however, remain challenging. For binary convolutional networks, the reason lies in the degraded representation caused by binarizing full-precision filters. To address this problem, we propose new circulant filters (CiFs) and a circulant binary convolution (CBConv) to enhance the capacity of binarized convolutional features via our circulant back propagation (CBP). The CiFs can be easily incorporated into existing deep convolutional neural networks (DCNNs), which leads to new Circulant Binary Convolutional Networks (CBCNs). Extensive experiments confirm that the performance gap between the 1-bit and full-precision DCNNs is minimized by increasing the filter diversity, which further increases the representational ability in our networks. Our experiments on ImageNet show that CBCNs achieve 61.4% top-1 accuracy with ResNet18. Compared to the state-of-the-art such as XNOR, CBCNs can achieve up to 10% higher top-1 accuracy with more powerful representational ability. Chunlei Liu 0001, Wenrui Ding, Xin Xia 0005, Baochang Zhang 0001, Jiaxin Gu, Jianzhuang Liu, Rongrong Ji, David S. Doermann |
CVPR | 8 |
| 2019 | Learning Instance Activation Maps for Weakly Supervised Instance SegmentationabstractDiscriminative region responses residing inside an object instance can be extracted from networks trained with image-level label supervision. However, learning the full extent of pixel-level instance response in a weakly supervised manner remains unexplored. In this work, we tackle this challenging problem by using a novel instance extent filling approach. We first design a process to selectively collect pseudo supervision from noisy segment proposals obtained with previously published techniques. The pseudo supervision is used to learn a differentiable filling module that predicts a class-agnostic activation map for each instance given the image and an incomplete region response. We refer to the above maps as Instance Activation Maps (IAMs), which provide a fine-grained instance-level representation and allow instance masks to be extracted by lightweight CRF. Extensive experiments on the PASCAL VOC12 dataset show that our approach beats the state-of-the-art weakly supervised instance segmentation methods by a significant margin and increases the inference speed by an order of magnitude. Our method also generalizes well across domains and to unseen object categories. Without fine-tuning for the specific tasks, our model trained on VOC12 dataset (20 classes) obtains top performance for weakly supervised object localization on the CUB dataset (200 classes) and achieves competitive results on three widely used salient object detection benchmarks. Yi Zhu 0004, Yanzhao Zhou, Huijuan Xu 0001, Qixiang Ye, David S. Doermann, Jianbin Jiao |
CVPR | 5 |
| 2019 | Planar content selection in images and videos using frontalness
Sungmin Eum, David S. Doermann |
Pattern Recognit. Lett. | 2 |
| 2019 | GiB: A Game Theory Inspired Binarization Technique for Degraded Document ImagesabstractDocument image binarization classifies each pixel in an input document image as either foreground or background under the assumption that the document is pseudo binary in nature. However, noise introduced during acquisition or due to aging or handling of the document can make binarization a challenging task. This paper presents a novel game theory inspired binarization technique for degraded document images. A two-player, non-zero-sum, non-cooperative game is designed at the pixel level to extract the local information, which is then fed to a K-means algorithm to classify a pixel as foreground or background. We also present a preprocessing step that is performed to eliminate the intensity variation that often appears in the background and a post-processing step to refine the results. The method is tested on seven publicly available datasets, namely, DIBCO 2009-14 and 2016. The experimental results show that GiB (Game theory Inspired Binarization) outperforms competing state-of-the-art methods in most cases. Showmik Bhowmik, Ram Sarkar, Bishwadeep Das, David S. Doermann |
IEEE Trans. Image Process. | 4 |
| 2018 | Text and non-text separation in offline document images: a survey
Showmik Bhowmik, Ram Sarkar, Mita Nasipuri, David S. Doermann |
Int. J. Document Anal. Recognit. | 4 |
| 2018 | Application of Structural and Topological Features to Recognize Online Handwritten Bangla CharactersabstractThis article presents a set of novel features for robust online Bangla handwritten character recognition. Two feature extraction methods are presented here. The first describes the transition from background to foreground pixels and vice versa. The second uses a combination of topological features and centre-of-gravity- (CG) based circular features where global information, local information, and Circular Quadrant Mass Distribution information have been extracted. The impact of each along with their combination have also been analyzed. A total of 15,000 isolated online Bangla character samples have been collected and used for the evaluation. A Support Vector Machine classifier records the best recognition rate when the transition count feature, CG-based circular features, and topological features are combined. Shibaprasad Sen, Ankan Bhattacharyya, Pawan Kumar Singh 0001, Ram Sarkar, Kaushik Roy 0004, David S. Doermann |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2017 | IOD-CNN: Integrating object detection networks for event recognitionabstractMany previous methods have showed the importance of considering semantically relevant objects for performing event recognition, yet none of the methods have exploited the power of deep convolutional neural networks to directly integrate relevant object information into a unified network. We present a novel unified deep CNN architecture which integrates architecturally different, yet semantically-related object detection networks to enhance the performance of the event recognition task. Our architecture allows the sharing of the convolutional layers and a fully connected layer which effectively integrates event recognition, rigid object detection and non-rigid object detection. Sungmin Eum, Hyungtae Lee, Heesung Kwon, David S. Doermann |
ICIP | 4 |
| 2016 | No-reference document image quality assessment based on high order image statisticsabstractDocument image quality assessment (DIQA) aims to predict the visual quality of degraded document images. Although the definition of “visual quality” can change based on the specific applications, in this paper, we use OCR accuracy as a metric for quality and develop a novel no-reference DIQA method based on high order image statistics for OCR accuracy prediction. The proposed method consists of three steps. First, normalized local image patches are extracted with regular grid and a comprehensive document image codebook is constructed by K-means clustering. Second, local features are softly assigned to several nearest codewords, and the direct differences between high order statistics of local features and codewords are calculated as global quality aware features. Finally, support vector regression (SVR) is utilized to learn the mapping between extracted image features and OCR accuracies. Experimental results on two document image databases show that the proposed method can accurately predict OCR accuracy and outperforms previous algorithms. Jingtao Xu, Peng Ye 0001, Qiaohong Li, Yong Liu 0027, David S. Doermann |
ICIP | 5 |
| 2016 | Content selection using frontalness evaluation of multiple framesabstractThis paper addresses the problem of selecting instances of a planar object in a video or from a set of images based on an evaluation of its “frontalness”. We introduce the idea of “evaluating the frontalness” by computing how close the object's surface normal aligns with the optical axis of a camera. The unique and novel aspect of our method is that unlike previous planar object pose estimation methods, our method does not require the true frontal image as a reference. The intuition is that a true frontal image can be used to produce other non-frontal images by perspective projection, while the non-frontal images have limited ability to produce other non-frontal images. We show that this intuition of comparing ‘frontal’ and ‘non-frontal’ can be extended to comparing ‘more frontal’ and ‘less frontal’ images. Based on this observation, our method estimates the relative frontalness of an image by exploiting the objective space error. We also propose the usage of K-invariant space to evaluate the frontalness even when the camera intrinsic parameters are unknown (e.g., images/videos from the web). We show that our method outperforms the homography decomposition-based method which also does not require reference images. In addition, a qualitative evaluation is carried out to show that our method can be applied in selecting the most frontal characters from a set of images captured in various viewpoints. Sungmin Eum, David S. Doermann |
ICPR | 2 |
| 2016 | Blind Image Quality Assessment Based on High Order Statistics AggregationabstractBlind image quality assessment (BIQA) research aims to develop a perceptual model to evaluate the quality of distorted images automatically and accurately without access to the non-distorted reference images. The state-of-the-art general purpose BIQA methods can be classified into two categories according to the types of features used. The first includes handcrafted features which rely on the statistical regularities of natural images. These, however, are not suitable for images containing text and artificial graphics. The second includes learning-based features which invariably require large codebook or supervised codebook updating procedures to obtain satisfactory performance. These are time-consuming and not applicable in practice. In this paper, we propose a novel general purpose BIQA method based on high order statistics aggregation (HOSA), requiring only a small codebook. HOSA consists of three steps. First, local normalized image patches are extracted as local features through a regular grid, and a codebook containing 100 codewords is constructed by K-means clustering. In addition to the mean of each cluster, the diagonal covariance and coskewness (i.e., dimension-wise variance and skewness) of clusters are also calculated. Second, each local feature is softly assigned to several nearest clusters and the differences of high order statistics (mean, variance and skewness) between local features and corresponding clusters are softly aggregated to build the global quality aware image representation. Finally, support vector regression is adopted to learn the mapping between perceptual features and subjective opinion scores. The proposed method has been extensively evaluated on ten image databases with both simulated and realistic image distortions, and shows highly competitive performance to the state-of-the-art BIQA methods. Jingtao Xu, Peng Ye 0001, Qiaohong Li, Haiqing Du, Yong Liu 0027, David S. Doermann |
IEEE Trans. Image Process. | 6 |
| 2015 | JH2R: Joint Homography Estimation for Highlight RemovalabstractImagine being in an art museum where there are paintings or pictures held inside glass-frames for protection. There are pieces which you wish to capture using a camera, but you experience difficulties avoiding highlights which are generated by indoor lighting reflected off the glossy surfaces. Similar problems occur when capturing contents off of whiteboards, documents printed on glossy surfaces, objects such as books or CDs with plastic covers. In this work, we address the problem of removing unwanted highlight regions in images generated by reflections of light sources on glossy surfaces. Although there have been efforts made to synthetically fill in the missing regions using the neighboring patterns by applying methods like inpainting [3, 4], it is impossible to recover the missing information in completely saturated regions. Therefore, we need to use multiple images where corresponding regions are not covered by the saturated highlights. Unlike other methods, our method uses the relationship between the highlight regions resulting in more robust removal of saturated highlights. Our method Overview Our method was motivated by a widely acknowledged physical phenomenon referred to as the ‘motion parallax’. Without loss of generality, we can similarly view the relationship between the desired content (e.g., a painting) and the highlights. Since the highlights caused by the light source are the result of the reflection on the glossy surface before they reach the camera, the light source can be modeled to virtually exist on the other side of the content. Note that, the distance from the light source is always larger than the distance from the content (D > d, in Figure 1). Sungmin Eum, Hyungtae Lee, David S. Doermann |
BMVC | 3 |
| 2015 | A graphical model approach for matching partial signaturesabstractIn this paper, we present a novel partial signature matching method using graphical models. Shape context features are extracted from the contour of signatures to capture local variations, and K-means clustering is used to build a visual vocabulary from a set of reference signatures. To describe the signatures, supervised latent Dirichlet allocation is used to learn the latent distributions of the salient regions over the visual vocabulary and hierarchical Dirichlet processes are implemented to infer the number of salient regions needed. Our work is evaluated on three datasets derived from the DS-I Tobacco signature dataset with clean signatures and the DS-II UMD dataset with signatures with different degradations. The results show the effectiveness of the approach for both the partial and full signature matching. Xianzhi Du, David S. Doermann, Wael Abd-Almageed |
CVPR | 2 |
| 2015 | Novel line verification for multiple instance focused retrieval in document collectionsabstractSpatial verification is typically employed to check the spatial consistency among matched local features and to remove outliers. However, when looking for multiple instances of the query within a target image, RANSAC algorithms which are widely applied in many one-to-one matching applications might fail due to the large proportion of “outliers” - correct matches corresponding to other instances. On the other hand, geometrical verification methods are more robust to outliers but usually suffer from high computational costs. In this paper, we introduce a novel two-step line verification method which is more flexible than existing methods and leads to lower computational complexity especially when multiple instances of a query are sought. We study this approach within an information extraction scenario, where the objective is to locate document structures indicative of certain type of information (e.g. different records on invoices). Hongxing Gao, Marçal Rusiñol, Dimosthenis Karatzas, Josep Lladós 0001, Rajiv Jain, David S. Doermann |
ICDAR | 6 |
| 2015 | Localized document image change detectionabstractGiven two versions of a document image, the goal of document image change detection is to automatically determine exactly what content was added, deleted or modified. Typically, one would accomplish this by first performing Optical Character Recognition (OCR) on the two documents and then performing a “diff” to identify the changes. However, this approach can fail due to OCR errors, poor segmentation, or the inability to handle graphical content. We compare the OCR baseline with two techniques based on SIFT features that detect changes in the image at the word level. The first approach performs the “diff” on SIFT features extracted from the center line of the text image. The second approach performs a segmentation free alignment of text blocks using dense SIFT to address the more general cases where segmentation fails or graphical objects are modified. Results on two experimental datasets show the improvement of the segmentation free approach over the baseline approach. Rajiv Jain, David S. Doermann |
ICDAR | 2 |
| 2015 | Word-level script identification for handwritten Indic scriptsabstractAutomatic script identification from handwritten document images facilitates many important applications such as indexing, sorting and triage. A given Optical Character Recognition (OCR) system is typically trained on only a single script but for documents or collections containing different scripts, there must be some way to automatically identify the script prior to OCR. For Indic script research, some results have been reported in the literature but the task is far from solved. In this paper, we propose a word-level script identification technique for six handwritten Indic scripts- Bangla, Devanagari, Gurumukhi, Malayalam, Oriya Telugu and the Roman script. A set of 82 features has been designed using a combination of elliptical and polygonal approximation techniques. Our approach has been evaluated on a dataset of 7000 handwritten text words, using multiple classifiers. A Multi-Layer Perceptron (MLP) classifier was found to be the best classifier resulting in 95.35% accuracy. The result is progressive considering the complexities and shape variations of the Indic scripts. Pawan Kumar Singh 0001, Ram Sarkar, Mita Nasipuri, David S. Doermann |
ICDAR | 4 |
| 2015 | Simultaneous estimation of image quality and distortion via multi-task convolutional neural networksabstractIn this work we describe a compact multi-task Convolutional Neural Network (CNN) for simultaneously estimating image quality and identifying distortions. CNNs are natural choices for multi-task problems because learned convolutional features may be shared by different high level tasks. However, we empirically argue that simply appending additional tasks based on the state of the art structure (e.g., [1]) does not lead to optimal solutions. We design a compact structure with nearly 90% fewer parameters compared to [1], and demonstrate its learning power. Peng Ye 0001, Yi Li 0025, David S. Doermann |
ICIP | 4 |
| 2015 | SHOE: Sibling Hashing with Output EmbeddingsabstractWe present a supervised binary encoding scheme for image retrieval that learns projections by taking into account similarity between classes obtained from output embeddings. Our motivation is that binary hash codes learned in this way improve the visual quality of retrieval results by ranking related (or ``sibling'') class images before unrelated class images. We employ a sequential greedy optimization that learns relationship aware projections by minimizing the difference between inner products of binary codes and output embedding vectors. We develop a joint optimization framework to learn projections which improve the accuracy of supervised hashing over the current state of the art with respect to standard and sibling evaluation metrics. We further obtain discriminative features learned from correlations of kernelized input CNN features and output embeddings, which significantly boosts performance. Experiments are performed on three datasets: CUB-2011, SUN-Attribute and ImageNet ILSVRC 2010, where we show significant improvement in sibling performance metrics over state-of-the-art supervised hashing techniques, while maintaining performance with respect to standard metrics. Sravanthi Bondugula, Varun Manjunatha, Larry Davis 0001, David S. Doermann |
ACM Multimedia | 4 |
| 2015 | Machine-assisted authentication of paper currency: an experiment on Indian banknotes
Ankush Roy, Biswajit Halder, Utpal Garain, David S. Doermann |
Int. J. Document Anal. Recognit. | 4 |
| 2015 | Text Detection and Recognition in Imagery: A SurveyabstractThis paper analyzes, compares, and contrasts technical challenges, methods, and the performance of text detection and recognition research in color imagery. It summarizes the fundamental problems and enumerates factors that should be considered when addressing these problems. Existing techniques are categorized as either stepwise or integrated and sub-problems are highlighted including text localization, verification, segmentation and recognition. Special issues associated with the enhancement of degraded text and the processing of video text, multi-oriented, perspectively distorted and multilingual text are also addressed. The categories and sub-categories of text are illustrated, benchmark datasets are enumerated, and the performance of the most representative approaches is compared. This review provides a fundamental comparison and analysis of the remaining problems in the field. Qixiang Ye, David S. Doermann |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Orientation Robust Text Line Detection in Natural ImagesabstractIn this paper, higher-order correlation clustering (HOCC) is used for text line detection in natural images. We treat text line detection as a graph partitioning problem, where each vertex is represented by a Maximally Stable Extremal Region (MSER). First, weak hypothesises are proposed by coarsely grouping MSERs based on their spatial alignment and appearance consistency. Then, higher-order correlation clustering (HOCC) is used to partition the MSERs into text line candidates, using the hypotheses as soft constraints to enforce long range interactions. We further propose a regularization method to solve the Semidefinite Programming problem in the inference. Finally we use a simple texton-based texture classifier to filter out the non-text areas. This framework allows us to naturally handle multiple orientations, languages and fonts. Experiments show that our approach achieves competitive performance compared to the state of the art. Yi Li 0025, David S. Doermann |
CVPR | 3 |
| 2014 | Convolutional Neural Networks for No-Reference Image Quality AssessmentabstractIn this work we describe a Convolutional Neural Network (CNN) to accurately predict image quality without a reference image. Taking image patches as input, the CNN works in the spatial domain without using hand-crafted features that are employed by most previous methods. The network consists of one convolutional layer with max and min pooling, two fully connected layers and an output node. Within the network structure, feature learning and regression are integrated into one optimization process, which leads to a more effective model for estimating image quality. This approach achieves state of the art performance on the LIVE dataset and shows excellent generalization ability in cross dataset experiments. Further experiments on images with local distortions demonstrate the local quality estimation ability of our CNN, which is rarely reported in previous literature. Peng Ye 0001, Yi Li 0025, David S. Doermann |
CVPR | 4 |
| 2014 | Active Sampling for Subjective Image Quality AssessmentabstractSubjective Image Quality Assessment (IQA) is the most reliable way to evaluate the visual quality of digital images perceived by the end user. It is often used to construct image quality datasets and provide the groundtruth for building and evaluating objective quality measures. Subjective tests based on the Mean Opinion Score (MOS) have been widely used in previous studies, but have many known problems such as an ambiguous scale definition and dissimilar interpretations of the scale among subjects. To overcome these limitations, Paired Comparison (PC) tests have been proposed as an alternative and are expected to yield more reliable results. However, PC tests can be expensive and time consuming, since for n images they require (n2)comparisons. We present a hybrid subjective test which combines MOS and PC tests via a unified probabilistic model and an active sampling method. The proposed method actively constructs a set of queries consisting of MOS and PC tests based on the expected information gain provided by each test and can effectively reduce the number of tests required for achieving a target accuracy. Our method can be used in conventional laboratory studies as well as crowdsourcing experiments. Experimental results show the proposed method outperforms state-of-the-art subjective IQA tests in a crowdsourced setting. Peng Ye 0001, David S. Doermann |
CVPR | 2 |
| 2014 | Beyond Human Opinion Scores: Blind Image Quality Assessment Based on Synthetic ScoresabstractState-of-the-art general purpose Blind Image Quality Assessment (BIQA) models rely on examples of distorted images and corresponding human opinion scores to learn a regression function that maps image features to a quality score. These types of models are considered "opinion-aware" (OA) BIQA models. A large set of human scored training examples is usually required to train a reliable OA-BIQA model. However, obtaining human opinion scores through subjective testing is often expensive and time-consuming. It is therefore desirable to develop "opinion-free" (OF) BIQA models that do not require human opinion scores for training. This paper proposes BLISS (Blind Learning of Image Quality using Synthetic Scores). BLISS is a simple, yet effective method for extending OA-BIQA models to OF-BIQA models. Instead of training on human opinion scores, we propose to train BIQA models on synthetic scores derived from Full-Reference (FR) IQA measures. State-of-the-art FR measures yield high correlation with human opinion scores and can serve as approximations to human opinion scores. Unsupervised rank aggregation is applied to combine different FR measures to generate a synthetic score, which serves as a better "gold standard". Extensive experiments on standard IQA datasets show that BLISS significantly outperforms previous OF-BIQA methods and is comparable to state-of-the-art OA-BIQA methods. Peng Ye 0001, Jayant Kumar, David S. Doermann |
CVPR | 3 |
| 2014 | Combining Local Features for Offline Writer IdentificationabstractSeveral powerful approaches have recently been proposed for writer identification, which rely on local descriptors that capture the texture, shape and curvature properties of the handwriting. In this paper we use combinations of three of these features (K-Adjacent Segments, SURF, and Contour Gradient Descriptors), to address the writer identification problem. Experiments demonstrate that feature combinations outperform individual features, resulting in state-of-the-art performance on three datasets. Rajiv Jain, David S. Doermann |
ICFHR | 2 |
| 2014 | Sharpness-aware document image mosaicing using graphcutsabstractThere are numerous types of documents which are difficult to scan or capture in a single pass due to their physical size or the size of their content. One possible solution that has been proposed is mosaicing multiple overlapping images to capture the complete document. In this paper, we present a novel Graphcut-based document image mosaicing method which seeks to overcome the known limitations of the previous approaches. First, our method does not require any prior knowledge of the content of the given document images, making it more widely applicable and robust. Second, information regarding the geometrical disposition between the overlapping images is exploited to minimize the errors at the boundary regions. Third, our method incorporates a sharpness measure which induces cut generation in a way that results in the mosaic including the sharpest pixels. Our method is shown to outperform previous methods, both quantitatively and qualitatively. Sungmin Eum, David S. Doermann |
ICIP | 2 |
| 2014 | A deep learning approach to document image quality assessmentabstractThis paper proposes a deep learning approach for document image quality assessment. Given a noise corrupted document image, we estimate its quality score as a prediction of OCR accuracy. First the document image is divided into patches and non-informative patches are sifted out using Otsu's binarization technique. Second, quality scores are obtained for all selected patches using a Convolutional Neural Network (CNN), and the patch scores are averaged over the image to obtain the document score. The proposed CNN contains two layers of convolution, location blind max-min pooling, and Rectified Linear Units in the fully connected layers. Experiments on two document quality datasets show our method achieved the state of the art performance. Peng Ye 0001, Yi Li 0025, David S. Doermann |
ICIP | 4 |
| 2014 | No-reference video quality assessment via feature learningabstractIn this paper, we propose a novel “Opinion Free” (OF) No-Reference Video Quality Assessment (NR-VQA) algorithm based on frame-level unsupervised feature learning and hysteresis temporal pooling. The system consists of three components: feature extraction with max-min pooling, frame quality prediction and temporal pooling. Frame level features are first extracted by unsupervised feature learning and used to train a linear Support Vector Regressor (SVR) for predicting quality scores frame by frame. Frame-level quality scores are then combined by temporal pooling to obtain a single video quality score. We tested the proposed method on the LIVE video quality database and experimental results show that without training on human opinion scores the proposed method is comparable to state-of-the-art NR-VQA algorithms. Jingtao Xu, Peng Ye 0001, Yong Liu 0027, David S. Doermann |
ICIP | 4 |
| 2014 | Robust scene text detection using integrated feature discriminationabstractScene text detection in images of cluttered backgrounds and/or multilingual context is very challenging. In this paper, we propose a discriminative approach that integrates appearance and consensus features for robust scene text detection. We propose an integrated discrimination model to perform text classification as well as control component grouping. We design shape, stroke and structural features to describe text component appearance and the consensus among them. Experimental results on three public datasets show that the proposed approach is robust to cluttered backgrounds, and is applicable in multilingual environments. Qixiang Ye, David S. Doermann |
ICIP | 2 |
| 2014 | Signature Matching Using Supervised Topic ModelsabstractIn this paper, we present a novel signature matching method based on supervised topic models. Shape Context features are extracted from signature shape contours which capture the local variations in signature properties. We then use the concept of topic models to learn the shape context features which correspond to individual authors. The approach consists of three primary steps. First, K-means is used to cluster shape context features to form term frequency histograms which correspond to a vocabulary for the set of signatures in the gallery. Second, a supervised topic model is used to construct an observation/author correspondence. Finally, the correspondence is used to classify query signatures and return the corresponding author. Two datasets are used to test our algorithm: DS-I Tobacco signature dataset with clean signatures and DS-II UMD dataset with noisy signatures. We demonstrate considerable improvement over state of the art methods. Xianzhi Du, David S. Doermann, Wael Abd-Almageed |
ICPR | 2 |
| 2014 | Depth Structure Association for RGB-D Multi-target TrackingabstractMulti-target tracking in outdoor scenes plays an important role in many computer vision applications. Most previous work on visual information based multi-target tracking does not incorporate depth information and the absence of depth information often leads to mismatching or tracking failures. In this paper, we propose a Depth Structure Association (DSA) approach for RGB-D data based multi-target tracking. DSA encodes depth information in a chain structure, the structure is used by DSA together with appearance and motion information to address object occlusion issues in outdoor scenes. Additionally, the use of DSA has the advantages of regulating a much smaller solution space, greatly reducing the computational complexity. Experimental results on three datasets demonstrate that our DSA approach can significantly reduce object mismatch and tracking failure for long term occlusions. Shan Gao 0003, Zhenjun Han, David S. Doermann, Jianbin Jiao |
ICPR | 3 |
| 2014 | Convolutional Neural Networks for Document Image ClassificationabstractThis paper presents a Convolutional Neural Network (CNN) for document image classification. In particular, document image classes are defined by the structural similarity. Previous approaches rely on hand-crafted features for capturing structural information. In contrast, we propose to learn features from raw image pixels using CNN. The use of CNN is motivated by the the hierarchical nature of document layout. Equipped with rectified linear units and trained with dropout, our CNN performs well even when document layouts present large inner-class variations. Experiments on public challenging datasets demonstrate the effectiveness of the proposed approach. Jayant Kumar, Peng Ye 0001, Yi Li 0025, David S. Doermann |
ICPR | 5 |
| 2014 | A Belief Based Correlated Topic Model for Trajectory Clustering in Crowded Video ScenesabstractTrajectory clustering in crowded video scenes is very challenging. In this paper, we propose to use a belief based correlated topic model (BCTM) to learn discriminative middle level features for trajectory clustering. By constructing a scene prior based joint Gaussian distribution, the BCTM can uncover relations between trajectory clusters and the middle level features using a parameter estimation procedure. The method has distinct advantages over Correlated Topic Model (CTM) and Random Field Topic (RFT) model previously proposed. The inputs to the BCTM are either full trajectories or trajectory fragments obtained with an existing tracking algorithm. The output BCTM features are input to a hierarchical clustering algorithm to obtain trajectory clusters. Experiments on three benchmark datasets show that the proposed BCTM and trajectory clustering approach improves the state of the art. Jialing Zou, Qixiang Ye, Yanting Cui, David S. Doermann, Jianbin Jiao |
ICPR | 4 |
| 2014 | Structural similarity for document image classification and retrieval
Jayant Kumar, Peng Ye 0001, David S. Doermann |
Pattern Recognit. Lett. | 3 |
| 2013 | Real-Time No-Reference Image Quality Assessment Based on Filter LearningabstractThis paper addresses the problem of general-purpose No-Reference Image Quality Assessment (NR-IQA) with the goal of developing a real-time, cross-domain model that can predict the quality of distorted images without prior knowledge of non-distorted reference images and types of distortions present in these images. The contributions of our work are two-fold: first, the proposed method is highly efficient. NR-IQA measures are often used in real-time imaging or communication systems, therefore it is important to have a fast NR-IQA algorithm that can be used in these real-time applications. Second, the proposed method has the potential to be used in multiple image domains. Previous work on NR-IQA focus primarily on predicting quality of natural scene image with respect to human perception, yet, in other image domains, the final receiver of a digital image may not be a human. The proposed method consists of the following components: (1) a local feature extractor, (2) a global feature extractor and (3) a regression model. While previous approaches usually treat local feature extraction and regression model training independently, we propose a supervised method based on back-projection, which links the two steps by learning a compact set of filters which can be applied to local image patches to obtain discriminative local features. Using a small set of filters, the proposed method is extremely fast. We have tested this method on various natural scene and document image datasets and obtained state-of-the-art results. Peng Ye 0001, Jayant Kumar, David S. Doermann |
CVPR | 4 |
| 2013 | Large-Scale Signature Matching Using Multi-stage HashingabstractIn this paper, we propose a fast large-scale signature matching method based on locality sensitive hashing (LSH). Shape Context features are used to describe the structure of signatures. Two stages of hashing are performed to find the nearest neighbours for query signatures. In the first stage, we use M randomly generated hyper planes to separate shape context feature points into different bins, and compute a term-frequency histogram to represent the feature point distribution as a feature vector. In the second stage we again use LSH to categorize the high-level features into different classes. The experiments are carried out on two datasets - DS-I, a small dataset contains 189 signatures, and DS-II, a large dataset created by our group which contains 26,000 signatures. We show that our algorithm can achieve a high accuracy even when few signatures are collected from one same person and perform fast matching when dealing with a large dataset. Xianzhi Du, Wael Abd-Almageed, David S. Doermann |
ICDAR | 3 |
| 2013 | VisualDiff: Document Image Verification and Change DetectionabstractThis paper explores the related problems of verification and change detection in document images. The goal is to determine if two document images differ, and if so, to determine precisely what content may have been added, deleted, or otherwise modified. This problem has many potential applications, especially for important legal documents such as contractual agreements. These agreements are often edited, shared and stored as scanned or hardcopy documents, where small, undetected changes between edits could create major differences in the contractual language and thus have severe repercussions. One can view the problem of change detection as tracing the revision history of a set of documents. Thus, in order to validate the performance of this approach, we created the "Enron Revisions" dataset. This dataset contains realistic revisions obtained from attachments in the Enron Corpus, and a series of before and after snapshots of the revisions in images with varying levels of noise from resolution, binarization, and blur. The approach taken in this paper utilizes the SIFT descriptor to align two document images without the benefit of OCR and once aligned, to compare dense descriptors to determine changes that have occurred within the image. As a baseline, this "VisualDiff" is compared to a UNIX diff-like approach on text extracted through OCR and results demonstrate the effectiveness of this approach. Rajiv Jain, David S. Doermann |
ICDAR | 2 |
| 2013 | Writer Identification Using an Alphabet of Contour Gradient DescriptorsabstractThis paper presents a new method for writer identification, which emulates the approach taken by forensic document examiners. It combines a novel feature, which uses contour gradients to capture local shape and curvature, with character segmentation to create a pseudo-alphabet for a given handwriting sample. A distance metric is then defined between elements of these alphabets that captures character similarity between two handwriting samples. This approach achieves a Top-1 identification rate of 96.5% on the benchmark IAM dataset, reducing the error rate of previous approaches by 50%. Rajiv Jain, David S. Doermann |
ICDAR | 2 |
| 2013 | Unsupervised Classification of Structurally Similar Document ImagesabstractIn this paper, we present a learning based approach for computing structural similarities among document images for unsupervised exploration in large document collections. The approach is based on multiple levels of content and structure. At a local level, a bag-of-visual words based on SURF features provides an effective way of computing content similarity. The document is then recursively partitioned and a histogram of codewords is computed for each partition. Structural similarity is computed using a random forest classifier trained with these histogram features. We experiment with three diverse datasets of document images varying in size, degree of structural similarity, and types of document images. Our results demonstrate that the proposed approach provides an effective general framework for grouping structurally similar document images. Jayant Kumar, David S. Doermann |
ICDAR | 2 |
| 2013 | Document Image Quality Assessment: A Brief SurveyabstractTo maintain, control and enhance the quality of document images and minimize the negative impact of degradations on various analysis and processing systems, it is critical to understand the types and sources of degradations and develop reliable methods for estimating the levels of degradations. This paper provides a brief survey of research on the topic of document image quality assessment. We first present a detailed analysis of the types and sources of document degradations. We then review techniques for document image degradation modeling. Finally, we discuss objective measures and subjective experiments that are used to characterize document image quality. Peng Ye 0001, David S. Doermann |
ICDAR | 2 |
| 2013 | Clutter noise removal in binary document images
Mudit Agrawal, David S. Doermann |
Int. J. Document Anal. Recognit. | 2 |
| 2012 | Unsupervised feature learning framework for no-reference image quality assessmentabstractIn this paper, we present an efficient general-purpose objective no-reference (NR) image quality assessment (IQA) framework based on unsupervised feature learning. The goal is to build a computational model to automatically predict human perceived image quality without a reference image and without knowing the distortion present in the image. Previous approaches for this problem typically rely on hand-crafted features which are carefully designed based on prior knowledge. In contrast, we use raw-image-patches extracted from a set of unlabeled images to learn a dictionary in an unsupervised manner. We use soft-assignment coding with max pooling to obtain effective image representations for quality estimation. The proposed algorithm is very computationally appealing, using raw image patches as local descriptors and using soft-assignment for encoding. Furthermore, unlike previous methods, our unsupervised feature learning strategy enables our method to adapt to different domains. CORNIA (Codebook Representation for No-Reference Image Assessment) is tested on LIVE database and shown to perform statistically better than the full-reference quality measure, structural similarity index (SSIM) and is shown to be comparable to state-of-the-art general purpose NR-IQA algorithms. Peng Ye 0001, Jayant Kumar, David S. Doermann |
CVPR | 4 |
| 2012 | Logo Retrieval in Document ImagesabstractThis paper presents a scalable algorithm for segmentation free logo retrieval in document images. The contributions include the use of the SURF feature for logo retrieval, a novel indexing algorithm for efficient retrieval and a method to filter results using the orientation of local features and geometric constraints. Results demonstrate that logo retrieval can be performed with high accuracy and efficiently scaled to a large datasets. Rajiv Jain, David S. Doermann |
Document Analysis Systems | 2 |
| 2012 | Local Segmentation of Touching Characters Using Contour Based Shape DecompositionabstractWe propose a contour based shape decomposition approach that provides local segmentation of touching characters. The shape contour is linearized into edge lets and edge lets are merged into boundary fragments. The connection cost between boundary fragments is obtained by considering local smoothness, connection length and a stroke-level property called the Same Stroke Rate. Samples of connections among boundary fragments are randomly generated and the one with the minimum global cost is selected to produce the final segmentation of the shape. To obtain a bipartite segmentation using this approach, we perform an iterative search for the parameters that finally yields two components on a shape. Experimental results on synthetic shape images and the LTP dataset show that this contour based shape decomposition technique is promising and it is effective for providing local segmentation of touching characters. David S. Doermann, Huaigu Cao, Rohit Prasad, Premkumar Natarajan |
Document Analysis Systems | 2 |
| 2012 | Learning Text-Line Segmentation Using Codebooks and Graph PartitioningabstractIn this paper, we present a codebook based method for handwritten text-line segmentation which uses image-patches in the training data to learn a graph-based similarity for clustering. We first construct a codebook of image-patches using K-medoids, and obtain exemplars which encode local evidence. We then obtain the corresponding codewords for all patches extracted from a given image and construct a similarity graph using the learned evidence and partitioned to obtain text-lines. Our learning based approach performs well on a field dataset containing degraded and un-constrained handwritten Arabic document images. Results on ICDAR 2009 segmentation contest dataset show that the method is competitive with previous approaches. Jayant Kumar, Peng Ye 0001, David S. Doermann |
ICFHR | 4 |
| 2012 | Sharpness estimation for document and scene images
Jayant Kumar, Francine Chen 0001, David S. Doermann |
ICPR | 3 |
| 2012 | Learning document structure for retrieval and classification
Jayant Kumar, Peng Ye 0001, David S. Doermann |
ICPR | 3 |
| 2012 | Learning features for predicting OCR accuracy
Peng Ye 0001, David S. Doermann |
ICPR | 2 |
| 2012 | Linguistic Resources for Handwriting Recognition and Translation Evaluation
Zhiyi Song, Safa Ismael, Stephen Grimes, David S. Doermann, Stephanie M. Strassel |
LREC | 4 |
| 2012 | No-Reference Image Quality Assessment Using Visual CodebooksabstractThe goal of no-reference objective image quality assessment (NR-IQA) is to develop a computational model that can predict the human-perceived quality of distorted images accurately and automatically without any prior knowledge of reference images. Most existing NR-IQA approaches are distortion specific and are typically limited to one or two specific types of distortions. In most practical applications, however, information about the distortion type is not really available. In this paper, we propose a general-purpose NR-IQA approach based on visual codebooks. A visual codebook consisting of Gabor-filter-based local features extracted from local image patches is used to capture complex statistics of a natural image. The codebook encodes statistics by quantizing the feature space and accumulating histograms of patch appearances. This method does not assume any specific types of distortions; however, when evaluating images with a particular type of distortion, it does require examples with the same or similar distortion for training. Experimental results demonstrate that the predicted quality score using our method is consistent with human-perceived image quality. The proposed method is comparable to state-of-the-art general-purpose NR-IQA methods and outperforms the full-reference image quality metrics, peak signal-to-noise ratio and structural similarity index on the Laboratory for Image and Video Engineering IQA database. Peng Ye 0001, David S. Doermann |
IEEE Trans. Image Process. | 2 |
| 2011 | Stroke-Like Pattern Noise Removal in Binary Document ImagesabstractThis paper presents a two-phased stroke-like pattern noise (SPN) removal algorithm for binary document images. The proposed approach aims at understanding script-independent prominent text component features using supervised classification as a first step. It then uses their cohesiveness and stroke-width properties to filter and associate smaller text components with them using an unsupervised classification technique. In order to perform text extraction, and hence noise removal, at diacritic-level, this divide-and-conquer technique does not assume the availability of accurate and large amounts of ground-truth data at component-level for training purposes. The method was tested on a collection of degraded and noisy, machine-printed and handwritten binary Arabic text documents. Results show pixel-level precision and recall of 86% and 90% respectively for noise-pixels. Mudit Agrawal, David S. Doermann |
ICDAR | 2 |
| 2011 | Offline Writer Identification Using K-Adjacent SegmentsabstractThis paper presents a method for performing offline writer identification by using K-adjacent segment (KAS) features in a bag-of-features framework to model a user's handwriting. This approach achieves a top 1 recognition rate of 93% on the benchmark IAM English handwriting dataset, which outperforms current state of the art features. Results further demonstrate that identification performance improves as the number of training samples increase, and additionally, that the performance of the KAS features extend to Arabic handwriting found in the MADCAT dataset. Rajiv Jain, David S. Doermann |
ICDAR | 2 |
| 2011 | Template Based Segmentation of Touching Components in Handwritten Text LinesabstractIn this paper, we present a template based approach to the segmentation of touching components in handwritten text lines. Local patches around touching components are identified and a dictionary is created consisting of template patches together with their correct segmentations. We use two shape context based methods to compute similarity between input patches and dictionary templates to find the best match. The template's known segmentation is then transformed to segment the input patch. Experiments are carried on a dataset of touching text lines. David S. Doermann |
ICDAR | 2 |
| 2011 | Fast Rule-Line Removal Using Integral Images and Support Vector MachinesabstractIn this paper, we present a fast and effective method for removing pre-printed rule-lines in handwritten document images. We use an integral-image representation which allows fast computation of features and apply techniques for large scale Support Vector learning using a data selection strategy to sample a small subset of training data. Results on both constructed and real-world data sets show that the method is effective for rule-line removal. We compare our method to a subspace-based method and show that better accuracy can be achieved in considerably less time. The integral-image based features proposed in the paper are generic and can be applied to other problems as well. Jayant Kumar, David S. Doermann |
ICDAR | 2 |
| 2011 | Segmentation of Handwritten Textlines in Presence of Touching ComponentsabstractThis paper presents an approach to text line extraction in handwritten document images which combines local and global techniques. We propose a graph-based technique to detect touching and proximity errors that are common with handwritten text lines. In a refinement step, we use Expectation-Maximization (EM) to iteratively split the error segments to obtain correct text-lines. We show improvement in accuracies using our correction method on datasets of Arabic document images. Results on a set of artificially generated proximity images show that the method is effective for handling touching errors in handwritten document images. Jayant Kumar, David S. Doermann, Wael Abd-Almageed |
ICDAR | 3 |
| 2011 | Document Image Classification and Labeling Using Multiple Instance LearningabstractThe labeling of large sets of images for training or testing analysis systems can be a very costly and time-consuming process. Multiple instance learning (MIL) is a generalization of traditional supervised learning which relaxes the need for exact labels on training instances. Instead, the labels are required only for a set of instances known as bags. In this paper, we apply MIL to the retrieval and localization of signatures and the retrieval of images containing machine-printed text, and show that a gain of 15-20% in performance can be achieved over the supervised learning with weak-labeling. We also compare our approach to supervised learning with fully annotated training data and report a competitive accuracy for MIL. Using our experiments on real-world datasets, we show that MIL is a good alternative when the training data has only document-level annotation. Jayant Kumar, Jaishanker K. Pillai, David S. Doermann |
ICDAR | 3 |
| 2011 | No-reference image quality assessment based on visual codebookabstractIn this paper, we propose a new learning based No-Reference Image Quality Assessment (NR-IQA) algorithm, which uses a visual codebook consisting of robust appearance descriptors extracted from local image patches to capture complex statistics of natural image for quality estimation. We use Gabor filter based local features as appearance descriptors and the codebook method encodes the statistics of natural image classes by vector quantizing the feature space and accumulating histograms of patch appearances based on this coding. This method does not assume any specific types of distortion and experimental results on the LIVE image quality assessment database show that this method provides consistent and reliable performance in quality estimation that exceeds other state-of-the-art NR-IQA approaches and is competitive with the full reference measure PSNR. Peng Ye 0001, David S. Doermann |
ICIP | 2 |
| 2011 | Cross-Language Entity Linking
Paul McNamee, James Mayfield, Dawn J. Lawrie, Douglas W. Oard, David S. Doermann |
IJCNLP | 5 |
| 2010 | Context-aware and content-based dynamic Voronoi page segmentationabstractThis paper presents a dynamic approach to document page segmentation based on inter-component relationships, local patterns and context features. State-of-the art page segmentation algorithms segment zones based on local properties of neighboring connected components such as distance and orientation, and do not typically consider additional properties other than size. Our proposed approach uses a contextually aware and dynamically adaptive page segmentation scheme. The page is first over-segmented using a dynamically adaptive scheme of separation features based on [2] and adapted from [13]. A decision to form zones is then based on the context built from these local separation features and high-level content features. Zone-based evaluation was performed on sets of printed and handwritten documents in English and Arabic scripts with multiple font types, sizes and we achieved an increase of 15% over the accuracy reported in [2]. Mudit Agrawal, David S. Doermann |
Document Analysis Systems | 2 |
| 2010 | Handwritten Arabic text line segmentation using affinity propagationabstractIn this paper, we present a novel graph-based method for extracting handwritten text lines in monochromatic Arabic document images. Our approach consists of two steps - Coarse text line estimation using primary components which define the line and assignment of diacritic components which are more difficult to associate with a given line. We first estimate local orientation at each primary component to build a sparse similarity graph. We then, use a shortest path algorithm to compute similarities between non-neighboring components. From this graph, we obtain coarse text lines using two estimates obtained from Affinity propagation and Breadth-first search. In the second step, we assign secondary components to each text line. The proposed method is very fast and robust to non-uniform skew and character size variations, normally present in handwritten text lines. We evaluate our method using a pixel-matching criteria, and report 96% accuracy on a dataset of 125 Arabic document images. We also present a proximity analysis on datasets generated by artificially decreasing the spacings between text lines to demonstrate the robustness of our approach. Jayant Kumar, Wael Abd-Almageed, David S. Doermann |
Document Analysis Systems | 4 |
| 2010 | The Evolution of Document AuthenticationabstractAuthentication in the document context refers to the ability to trace the origins of a document to a given person or device used to produce it or to a given time or place it was produced. The general approach typically involves comparing physical, visual and/or linguistic properties of a questioned source to reproducible properties of a known or genuine source. The challenges lie in defining acceptable variations between authentic sources and identifying distinguishing characteristics of forgeries or unknown sources. As documents have evolved from physical objects made with primitive devices to manuscripts created by machine to content that lives only in electronic form, methods for authentication have also changed. While there has been considerable work in attempts to automate problems such as signature verification and writer identification in the image domain, and to guarantee authenticity or prove authorship in the electronic text domain, other authentication tasks have continued to rely extensively on human expertise. This talk will overview the general concept of authentication and discuss some of the novel approaches that can be used to authenticate documents and detect forgeries. While technology advances in archeology, antiquities, forensics, security and business are driving new and better ways to perform authentication, they are also enabling more realistic ways to produce counterfeits. As we continue to make progress in automating various analysis and recognition tasks, the question remains as to how well we will be able to automate these highly expert driven authentication tasks. David S. Doermann |
ICFHR | 1 |
| 2010 | Performance Evaluation Tools for Zone Segmentation and Classification (PETS)abstractThis paper describes a set of Performance Evaluation Tools (PETS) for document image zone segmentation and classification. The tools allow researchers and developers to evaluate, optimize and compare their algorithms by providing a variety of quantitative performance metrics. The evaluation of segmentation quality is based on the pixel-based overlaps between two sets of zones proposed by Randriamasy and Vincent. PETS extends the approach by providing a set of metrics for overlap analysis, RLE and polygonal representation of zones and introduces type-matching to evaluate zone classification. The software is available for research use. Wontaek Seo, Mudit Agrawal, David S. Doermann |
ICPR | 3 |
| 2009 | Page Rule-Line Removal Using Linear Subspaces in Monochromatic Handwritten Arabic DocumentsabstractIn this paper we present a novel method for removing page rule lines in monochromatic handwritten Arabic documents using subspace methods with minimal effect on the quality of the foreground text. We use moment and histogram properties to extract features that represent the characteristics of the underlying rule lines. A linear subspace is incrementally built to obtain a line model that can be used to identify rule line pixels. We also introduce a novel scheme for evaluating noise removal algorithms in general and we use it to assess the quality of our rule line removal algorithm. Experimental results presented on a data set of 50 Arabic documents, handwritten by different writers, demonstrate the effectiveness of the proposed method. Wael Abd-Almageed, Jayant Kumar, David S. Doermann |
ICDAR | 3 |
| 2009 | Clutter Noise Removal in Binary Document ImagesabstractThe paper presents a clutter detection and removal algorithm for complex document images. The distance transform based approach is independent of clutter's position, size, shape and connectivity with text. Features are based on a residual image obtained by analysis of the distance transform and clutter elements, if present, are identified with an SVM classifier. Removal is restrictive, so text attached to the clutter is not deleted in the process. The method was tested on a collection of degraded and noisy, machine-printed and handwritten Arabic and English text documents. Results show pixel-level accuracies of 97.5% and 95% for clutter detection and removal respectively. This approach was also extended with a noise detection and removal model for documents having a mix of clutter and salt-n-pepper noise. Mudit Agrawal, David S. Doermann |
ICDAR | 2 |
| 2009 | Voronoi++: A Dynamic Page Segmentation Approach Based on Voronoi and Docstrum FeaturesabstractThis paper presents a dynamic approach to document page segmentation. Current page segmentation algorithms lack the ability to dynamically adapt local variations in the size, orientation and distance of components within a page. Our approach builds upon one of the best algorithms, Kise et. al. work based on Area Voronoi Diagrams, which adapts globally to page content to determine algorithm parameters. In our approach, local thresholds are determined dynamically based on parabolic relations between components, and Docstrum based angular and neighborhood features are integrated to improve accuracy. Zone-based evaluation was performed on four sets of printed and handwritten documents in English and Arabic scripts and an increase of 33% in accuracy is reported. Mudit Agrawal, David S. Doermann |
ICDAR | 2 |
| 2009 | Logo Matching for Document Image RetrievalabstractGraphics detection and recognition are fundamental research problems in document image analysis and retrieval. As one of the most pervasive graphical elements in business and government documents, logos may enable immediate identification of organizational entities and serve extensively as a declaration of a document's source and ownership. In this work, we developed an automatic logo-based document image retrieval system that handles: (1) Logo detection and segmentation by boosting a cascade of classifiers across multiple image scales; and (2) Logo matching using translation, scale, and rotation invariant shape descriptors and matching algorithms. Our approach is segmentation free and layout independent and we address logo retrieval in an unconstrained setting of 2D feature point matching. Finally, we quantitatively evaluate the effectiveness of our approach using large collections of real-world complex document images. Guangyu Zhu 0004, David S. Doermann |
ICDAR | 2 |
| 2009 | Mosaicing of camera-captured document images
Daniel DeMenthon, David S. Doermann |
Comput. Vis. Image Underst. | 3 |
| 2009 | Offline Loop Investigation for Handwriting AnalysisabstractResolution of different types of loops in handwritten script presents a difficult task and is an important step in many classic word recognition systems, writer modeling, and signature verification. When processing a handwritten script, a great deal of ambiguity occurs when strokes overlap, merge, or intersect. This paper presents a novel loop modeling and contour-based handwriting analysis that improves loop investigation. We show excellent results on various loop resolution scenarios, including axial loop understanding and collapsed loop recovery. We demonstrate our approach for loop investigation on several realistic data sets of static binary images and compare with the ground truth of the genuine online signal. Tal Steinherz, David S. Doermann, Ehud Rivlin, Nathan Intrator |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Signature Detection and Matching for Document Image RetrievalabstractAs one of the most pervasive methods of individual identification and document authentication, signatures present convincing evidence and provide an important form of indexing for effective document image processing and retrieval in a broad range of applications. However, detection and segmentation of free-form objects such as signatures from clustered background is currently an open document analysis problem. In this paper, we focus on two fundamental problems in signature-based document image retrieval. First, we propose a novel multiscale approach to jointly detecting and segmenting signatures from document images. Rather than focusing on local features that typically have large variations, our approach captures the structural saliency using a signature production model and computes the dynamic curvature of 2D contour fragments over multiple scales. This detection framework is general and computationally tractable. Second, we treat the problem of signature retrieval in the unconstrained setting of translation, scale, and rotation invariant nonrigid shape matching. We propose two novel measures of shape dissimilarity based on anisotropic scaling and registration residual error and present a supervised learning framework for combining complementary shape information from different dissimilarity metrics using LDA. We quantitatively study state-of-the-art shape representations, shape matching algorithms, measures of dissimilarity, and the use of multiple instances as query in document image retrieval. We further demonstrate our matching techniques in offline signature verification. Extensive experiments using large real-world collections of English and Arabic machine-printed and handwritten documents demonstrate the excellent performance of our approaches. Guangyu Zhu 0004, Yefeng Zheng 0001, David S. Doermann, Stefan Jäger 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Language identification for handwritten document images using a shape codebook
Guangyu Zhu 0004, Xiaodong Yu 0002, Yi Li 0025, David S. Doermann |
Pattern Recognit. | 4 |
| 2008 | Re-targetable OCR with Intelligent Character SegmentationabstractWe have developed a font-model based intelligent character segmentation and recognition system. Using characteristics of structurally similar TrueType fonts, our system automatically builds a model to be used for the segmentation and recognition of the new script, independent of glyph composition. The key is a reliance on known font attributes. In our system three feature extraction methods are used to demonstrate the importance of appropriate features for classification. The methods are tested on both Latin (English) and non-Latin (Khmer) scripts. Results show that the character-level recognition accuracy exceeds 92\% for Khmer and 96\% for English on degraded documents. This work is a step toward the recognition of scripts of low-density languages which typically do not warrant the development of commercial OCR, yet often have complete TrueType font descriptions. Mudit Agrawal, David S. Doermann |
Document Analysis Systems | 2 |
| 2008 | Learning Visual Shape Lexicon for Document Image Content Recognition
Guangyu Zhu 0004, Xiaodong Yu 0002, Yi Li 0025, David S. Doermann |
ECCV (2) | 4 |
| 2008 | Signature-Based Document Image Retrieval
Guangyu Zhu 0004, Yefeng Zheng 0001, David S. Doermann |
ECCV (3) | 3 |
| 2008 | Document-zone classification using partial least squares and hybrid classifiersabstractThis paper introduces a novel document-zone classification algorithm. Low level image features are first extracted from document zones and partial least squares is used on pairs of classes to compute discriminating pairwise features. Rather than using the popular one-against-all and one-against-one voting schemes, we introduce a novel hybrid method which combines the benefits of the two schemes. The algorithm is applied on the University of Washington dataset and 97.3% classification accuracy is obtained. Wael Abd-Almageed, Mudit Agrawal, Wontaek Seo, David S. Doermann |
ICPR | 4 |
| 2008 | Support Vector Data Description for image categorization from Internet imagesabstractTraining a classifier for object category recognition using images on the Internet is an attractive approach due to its scalability. However, a big challenge in this approach is that it is difficult to automatically obtain sets of negative samples that are guaranteed to be free of positive samples. In this paper we propose to address this challenge with a Support Vector Data Description (SVDD) classifier. An SVDD classifier does not need negative images in training. It computes a hypersphere around the potentially good images in the feature space and uses this boundary to distinguish images of target visual category from outliers. Evaluation on standard test sets shows that we are able to achieve competitive classification performance using the contaminated training images from the Internet without the need for large datasets of negative examples. Xiaodong Yu 0002, Daniel DeMenthon, David S. Doermann |
ICPR | 3 |
| 2008 | A camera-based mobile data channel: capacity and analysisabstractIn this paper we propose a novel application, color Video Code (V-Code) and analyze its data transmission capacity through camera-based mobile data channels. Users can use the camera on a mobile device (PDA or camera phone) as a passive and pervasive data channel to download data encoded as a sequence of color visual patterns. The color V-Code is animated on a display, acquired by the camera and decoded by the pre-embedded software in the mobile device. One interesting question is what is the data transmission capacity it can achieve, theoretically and practically. To answer this question we build a camera channel model to measure color degradation using information theory and show that the capacity of the camera channel can be improved with the optimized color selection through color calibration. After initialization color models are learned automatically as downloading proceeds. We address the problem of precise registration, and implemented a fast perspective correction method to accelerate the decoder in real-time on a resource constrained device. With the optimized color set and efficient implementation we achieve a transmission bit rate of 15.4kbps on a common iMate Jamin phone (200MHz CPU). This speed is faster than the average GPRS bit rate (12kbps). Xu Liu 0003, David S. Doermann, Huiping Li 0001 |
ACM Multimedia | 2 |
| 2008 | Mobile Retriever: access to digital documents from their physical source
Xu Liu 0003, David S. Doermann |
Int. J. Document Anal. Recognit. | 2 |
| 2008 | Script-Independent Text Line Segmentation in Freestyle Handwritten DocumentsabstractText line segmentation in freestyle handwritten documents remains an open document analysis problem. Curvilinear text lines and small gaps between neighboring text lines present a challenge to algorithms developed for machine printed or hand-printed documents. In this paper, we propose a novel approach based on density estimation and a state-of-the-art image segmentation technique, the level set method. From an input document image, we estimate a probability map, where each element represents the probability that the underlying pixel belongs to a text line. The level set method is then exploited to determine the boundary of neighboring text lines by evolving an initial estimate. Unlike connected component based methods ( [1], [2] for example), the proposed algorithm does not use any script-specific knowledge. Extensive quantitative experiments on freestyle handwritten documents with diverse scripts, such as Arabic, Chinese, Korean, and Hindi, demonstrate that our algorithm consistently outperforms previous methods [1]-[3]. Further experiments show the proposed algorithm is robust to scale change, rotation, and noise. Yi Li 0025, Yefeng Zheng 0001, David S. Doermann, Stefan Jäger 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | Geometric Rectification of Camera-Captured Document ImagesabstractCompared to typical scanners, handheld cameras offer convenient, flexible, portable, and non-contact image capture, which enables many new applications and breathes new life into existing ones. However, camera-captured documents may suffer from distortions caused by non-planar document shape and perspective projection, which lead to failure of current OCR technologies. We present a geometric rectification framework for restoring the frontal-flat view of a document from a single camera-captured image. Our approach estimates 3D document shape from texture flow information obtained directly from the image without requiring additional 3D/metric data or prior camera calibration. Our framework provides a unified solution for both planar and curved documents and can be applied in many, especially mobile, camera-based document analysis applications. Experiments show that our method produces results that are significantly more OCR compatible than the original images. Daniel DeMenthon, David S. Doermann |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | VCode - Pervasive Data Transfer Using Video BarcodeabstractIn this paper, we describe a novel data transfer scheme that uses the camera in a smart phone as an alternative data channel. The data is encoded as a sequence of 2-D barcode images, displayed on a flat panel display, acquired by the camera, and decoded in real time by the software embedded in device. The decoded data is written to a file. Compared with existing data channels, such as CDMA/GPRS, cables, Bluetooth, and Infrared, our method relies on visual communication and does not require special hardware or data plans. Users only need to point the camera at a monitor displaying the VCode to download. Technical challenges to overcome include correction of perspective distortion, compensation for contrast variation, and efficient implementation of small footprint software into a mobile device. We address these challenges and present our solution in detail. We have implemented a prototype which allows users to download various types of files successfully, including pictures, ring tones and Java games onto camera phones running Symbian and Windows Mobile platforms. We discuss the limitations of our solution and outline future work to overcome these limitations. Xu Liu 0003, David S. Doermann, Huiping Li 0001 |
IEEE Trans. Multim. | 2 |
| 2007 | Simultaneous Appearance Modeling and Segmentation for Matching People Under Occlusion
Zhe Lin 0001, Larry Davis 0001, David S. Doermann, Daniel DeMenthon |
ACCV (2) | 3 |
| 2007 | Object Detection Using Shape Codebook
Xiaodong Yu 0002, Cornelia Fermüller, David S. Doermann |
BMVC | 4 |
| 2007 | Multi-scale Structural Saliency for Signature DetectionabstractDetecting and segmenting free-form objects from cluttered backgrounds is a challenging problem in computer vision. Signature detection in document images is one classic example and as of yet no reasonable solutions have been presented. In this paper, we propose a novel multi-scale approach to jointly detecting and segmenting signatures from documents with diverse layouts and complex backgrounds. Rather than focusing on local features that typically have large variations, our approach aims to capture the structural saliency of a signature by searching over multiple scales. This detection framework is general and computationally tractable. We present a saliency measure based on a signature production model that effectively quantifies the dynamic curvature of 2-D contour fragments. Our evaluation using large real world collections of handwritten and machine printed documents demonstrates the effectiveness of this joint detection and segmentation approach. Guangyu Zhu 0004, Yefeng Zheng 0001, David S. Doermann, Stefan Jäger 0001 |
CVPR | 3 |
| 2007 | Learning Higher-order Transition Models in Medium-scale Camera NetworksabstractWe present a Bayesian framework for learning higher- order transition models in video surveillance networks. Such higher-order models describe object movement between cameras in the network and have a greater predictive power for multi-camera tracking than camera adjacency alone. These models also provide inherent resilience to camera failure, filling in gaps left by single or even multiple non-adjacent camera failures. Our approach to estimating higher-order transition models relies on the accurate assignment of camera observations to the underlying trajectories of objects moving through the network. We addresses this data association problem by gathering the observations and evaluating alternative partitions of the observation set into individual object trajectories. Searching the complete partition space is intractable, so an incremental approach is taken, iteratively adding observations and pruning unlikely partitions. Partition likelihood is determined by the evaluation of a probabilistic graphical model. When the algorithm has considered all observations, the most likely (MAP) partition is taken as the true object trajectories. From these recovered trajectories, the higher-order statistics we seek can be derived and employed for tracking. The partitioning algorithm we present is parallel in nature and can be readily extended to distributed computation in medium-scale smart camera networks. Ryan Farrell, David S. Doermann, Larry Davis 0001 |
ICCV | 2 |
| 2007 | Hierarchical Part-Template Matching for Human Detection and SegmentationabstractLocal part-based human detectors are capable of handling partial occlusions efficiently and modeling shape articulations flexibly, while global shape template-based human detectors are capable of detecting and segmenting human shapes simultaneously. We describe a Bayesian approach to human detection and segmentation combining local part-based and global template-based schemes. The approach relies on the key ideas of matching a part-template tree to images hierarchically to generate a reliable set of detection hypotheses and optimizing it under a Bayesian MAP framework through global likelihood re-evaluation and fine occlusion analysis. In addition to detection, our approach is able to obtain human shapes and poses simultaneously. We applied the approach to human detection and segmentation in crowded scenes with and without background subtraction. Experimental results show that our approach achieves good performance on images and video sequences with severe occlusion. Zhe Lin 0001, Larry Davis 0001, David S. Doermann, Daniel DeMenthon |
ICCV | 3 |
| 2007 | An Interactive Approach to Pose-Assisted and Appearance-based Segmentation of HumansabstractAn interactive human segmentation approach is described. Given regions of interest provided by users, the approach iteratively estimates segmentation via a generalized EM algorithm. Specifically, it encodes both spatial and color information in a nonparametric kernel density estimator, and incorporates local MRF constraints and global pose inferences to propagate beliefs over image space iteratively to determine a coherent segmentation. This ensures the segmented humans resemble the shapes of human poses. Additionally, a layered occlusion model and a probabilistic occlusion reasoning method are proposed to handle segmentation of multiple humans in occlusion. The approach is tested on a wide variety of images containing single or multiple occluded humans, and the segmentation performance is evaluated quantitatively. Zhe Lin 0001, Larry Davis 0001, David S. Doermann, Daniel DeMenthon |
ICCV | 3 |
| 2007 | Automatic Document Logo DetectionabstractAutomatic logo detection and recognition continues to be of great interest to the document retrieval community as it enables effective identification of the source of a document. In this paper, we propose a new approach to logo detection and extraction in document images that robustly classifies and precisely localizes logos using a boosting strategy across multiple image scales. At a coarse scale, a trained Fisher classifier performs initial classification using features from document context and connected components. Each logo candidate region is further classified at successively finer scales by a cascade of simple classifiers, which allows false alarms to be discarded and the detected region to be refined. Our approach is segmentation free and lay-out independent. We define a meaningful evaluation metric to measure the quality of logo detection using labeled groundtruth. We demonstrate the effectiveness of our approach using a large collection of real-world documents. Guangyu Zhu 0004, David S. Doermann |
ICDAR | 2 |
| 2006 | Adaptive Transformation-Based Learning for Improving Dictionary Tagging
Burcu Karagol Ayan, David S. Doermann, Amy Weinberg |
EACL | 2 |
| 2006 | Imaging as an alternative data channel for camera phonesabstractIn this paper, we demonstrate a solution to use cameras to down-load data to cell phones as an alternative to existing wireless (CDMA/GPRS, BlueTooth), infrared or cable connections. In our method the data is encoded as a sequence of images, which can be displayed on any flat display, captured by users with their camera phones, and decoded by pre-embedded software. To solve this problem we need to be able to (1) encode arbitrary data as a sequence of images, (2) process captured images under various lighting variations and perspective distortions while maintaining realtime performance, and (3) decode the processed images robustly even when partial data is lost. In the paper we address these challenges in detail and present our solution. We have implemented a prototype which allows users to successfully download various types of files, including pictures, ring tones and Java programs to the camera phones. We discuss the limitations of our solution, and future works to overcome these limitations. Xu Liu 0003, David S. Doermann, Huiping Li 0001 |
MUM | 2 |
| 2006 | Video retrieval of near-duplicates using kappa-nearest neighbor retrieval of spatio-temporal descriptors
Daniel DeMenthon, David S. Doermann |
Multim. Tools Appl. | 2 |
| 2006 | Robust Point Matching for Nonrigid Shapes by Preserving Local Neighborhood StructuresabstractIn previous work on point matching, a set of points is often treated as an instance of a joint distribution to exploit global relationships in the point set. For nonrigid shapes, however, the local relationship among neighboring points is stronger and more stable than the global one. In this paper, we introduce the lotion of a neighborhood structure for the general point matching problem. We formulate point matching as an optimization problem to preserve local neighborhood structures during matching. Our approach has a simple graph matching interpretation, where each point is a node in the graph, and two nodes are connected by an edge if they are neighbors. The optimal match between two graphs is the one that maximizes the number of matched edges. Existing techniques are leveraged to search for an optimal solution with the shape context distance used to initialize the graph matching, followed by relaxation labeling updates for refinement. Extensive experiments show the robustness of our approach under deformation, noise in point locations, outliers, occlusion, and rotation. It outperforms the shape context and TPS-RPM algorithms on most scenarios. Yefeng Zheng 0001, David S. Doermann |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Flattening Curved Documents in ImagesabstractCompared to scanned images, document pictures captured by camera can suffer from distortions due to perspective and page warping. It is necessary to restore a frontal planar view of the page before other OCR techniques can be applied. In this paper we describe a novel approach for flattening a curved document in a single picture captured by an uncalibrated camera. To our knowledge this is the first reported method able to process general curved documents in images without camera calibration. We propose to model the page surface by a developable surface, and exploit the properties (parallelism and equal line spacing) of the printed textual content on the page to recover the surface shape. Experiments show that the output images are much more OCR friendly than the original ones. While our method is designed to work with any general developable surfaces, it can be adapted for typical special cases including planar pages, scans of thick books, and opened books. Daniel DeMenthon, David S. Doermann |
CVPR (2) | 3 |
| 2005 | Model of Object-Based Coding for Surveillance VideoabstractIn this paper, we explore the model of potential savings of object-based coding for surveillance video. Moving foreground objects in stationary camera surveillance video are detected by a background subtraction technique and encoded with MPEG-4 object-based coding. Experimental results show that compared with frame-based coding, object-based coding can achieve significant savings which are dependent on the video content. We further model the relationship of compression efficiency and the number and size of video objects using a statistical learning method. Simulations show that the model is representative. The model can be used to predict the savings of object-based coding and select coding methods for surveillance video. David S. Doermann |
ICASSP (2) | 2 |
| 2005 | Robust Point Matching for Two-Dimensional Nonrigid ShapesabstractRecently, nonrigid shape matching has received more and more attention. For nonrigid shapes, most neighboring points cannot move independently under deformation due to physical constraints. Furthermore, the rough structure of a shape should be preserved under deformation otherwise even people cannot match shapes reliably. Therefore, though the absolute distance between two points may change significantly, the neighborhood of a point is well preserved in general. Based on this observation, we formulate point matching as a graph matching problem. Each point is a node in the graph, and two nodes are connected by an edge if their Euclidean distance is less than a threshold. The optimal match between two graphs is the one that maximizes the number of matched edges. The shape context distance is used to initialize the graph matching, followed by relaxation labeling for refinement. Nonrigid deformation is overcome by bringing one shape closer to the other in each iteration using deformation parameters estimated from the current point correspondence. Experiments demonstrate the effectiveness of our approach: it outperforms the shape context and TPS-RPM algorithms under nonrigid deformation and noise on a public data set. Yefeng Zheng 0001, David S. Doermann |
ICCV | 2 |
| 2005 | Document Ranking by Layout RelevanceabstractThis paper describes the development of a new document ranking system based on layout similarity. The user has a need represented by a set of "wanted" documents, and the system ranks documents in the collection according to this need. Rather than performing complete document analysis, the system extracts text lines, and models layouts as relationships between pairs of these lines. This paper explores three novel feature sets to support scoring in large document collections. First, pairs of lines are used to form quadrilaterals, which are represented by their turning functions. A non-Euclidean distance is used to measure similarity. Second, the quadrilaterals are represented by 5D Euclidean vectors, and third, each line is represented by a 5D Euclidean vector. We compare the classification performance and computation speed of these three feature sets using a large database of diverse documents including forms, academic papers and handwritten pages in English and Arabic. The approach using quadrilaterals and turning functions produces slightly better results, but the approach using vectors to represent text lines is much faster for large document databases. May Huang, Daniel DeMenthon, David S. Doermann, Lynn Golebiowski, Booz Allen Hamilton |
ICDAR | 3 |
| 2005 | Identifying Script onWord-Level with Informational ConfidencabstractIn this paper, we present a multiple classifier system for script identification. Applying a Gabor filter analysis of textures on word-level, our system identifies Latin and non-Latin words in bilingual printed documents. The classifier system comprises four different architectures based on nearest neighbors, weighted Euclidean distances, Gaussian mixture models, and support vector machines. We report results for Arabic, Chinese, Hindi, and Korean script. Moreover, we show that combining informational confidence values using sum-rule can consistently outperform the best single recognition rate. Stefan Jäger 0001, Huanfeng Ma, David S. Doermann |
ICDAR | 3 |
| 2005 | Selection of Classifiers for the Construction of Multiple Classifier SystemsabstractMost studies on combining multiple classifiers have focused on combination methods, but a few studies have investigated on how to select component classifiers from a classifier pool. Multiple classifier systems performance varies with the component classifiers as well as the combination method. In this paper, methods based on information theory are proposed for selecting component classifiers, provided that the number of component classifiers is fixed in advance. These methods are applied to the classifier pool and examine the possible classifier sets. The system is compared to other multiple classifier systems on the recognition of unconstrained handwritten numerals. Hee-Joong Kang, David S. Doermann |
ICDAR | 2 |
| 2005 | Adaptive OCR with Limited User FeedbackabstractA methodology is proposed for processing noisy printed documents with limited user feedback. Without the support of ground truth, a specific collection of scanned documents can be processed to extract character templates. The adaptiveness of this approach lies in that the extracted templates are used to train an OCR classifier quickly and with limited user feedback. Experimental results show that this approach is extremely useful for the processing of noisy documents with many touching characters. Huanfeng Ma, David S. Doermann |
ICDAR | 2 |
| 2005 | Handwriting Matching and Its Application to Handwriting SynthesisabstractSince it is extremely expensive to collect a large volume of handwriting samples, synthesized data are often used to enlarge the training set. We argue that, in order to generate good handwriting samples, a synthesis algorithm should learn the shape deformation characteristics of handwriting from real samples. In this paper, we present a point matching algorithm to learn the deformation, and apply it to handwriting synthesis. Preliminary experiments show the advantages of our approach. Yefeng Zheng 0001, David S. Doermann |
ICDAR | 2 |
| 2005 | Fast camera motion estimation for hand-held devices and applicationsabstractIn this paper we present an efficient motion estimation algorithm for camera-enabled handheld devices, such as mobile phones and PDAs. Compared to general camera motion estimation, the estimation of ego-motion of handheld devices presents unique challenges because most devices are limited in resources (processing power, memory and battery). The algorithm must be lightweight so it can be efficiently embedded and run fast enough to produce smooth seamless translation. Our solution includes a multi-resolution scheme which searches the match from coarse to fine; and optimization of the search space. As a demonstration, we implemented the algorithm on the Symbian based Nokia 3650/6600 camera phone, and explored several interesting applications, including using camera motion for browsing documents and as a pointing device. A cross-platform implementation is also discussed. Xu Liu 0003, David S. Doermann, Huiping Li 0001 |
MUM | 2 |
| 2005 | Camera-based analysis of text and documents: a survey
David S. Doermann, Huiping Li 0001 |
Int. J. Document Anal. Recognit. | 2 |
| 2005 | A Parallel-Line Detection Algorithm Based on HMM DecodingabstractThe detection of groups of parallel lines is important in applications such as form processing and text (handwriting) extraction from rule lined paper. These tasks can be very challenging in degraded documents where the lines are severely broken. In this paper, we propose a novel model-based method which incorporates high-level context to detect these lines. After preprocessing (such as skew correction and text filtering), we use trained Hidden Markov Models (HMM) to locate the optimal positions of all lines simultaneously on the horizontal or vertical projection profiles, based on the Viterbi decoding. The algorithm is trainable so it can be easily adapted to different application scenarios. The experiments conducted on known form processing and rule line detection show our method is robust, and achieves better results than other widely used line detection methods. Yefeng Zheng 0001, Huiping Li 0001, David S. Doermann |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2004 | Building an information retrieval test collection for spontaneous conversational speechabstractTest collections model use cases in ways that facilitate evaluation of information retrieval systems. This paper describes the use of search-guided relevance assessment to create a test collection for retrieval of spontaneous conversational speech. Approximately 10,000 thematically coherent segments were manually identified in 625 hours of oral history interviews with 246 individuals. Automatic speech recognition results, manually prepared summaries, controlled vocabulary indexing, and name authority control are available for every segment. Those features were leveraged by a team of four relevance assessors to identify topically relevant segments for 28 topics developed from actual user requests. Search-guided assessment yielded sufficient inter-annotator agreement to support formative evaluation during system development. Baseline results for ranked retrieval are presented to illustrate use of the collection. Douglas W. Oard, Dagobert Soergel, David S. Doermann, G. Craig Murray, Jianqiang Wang 0002, Bhuvana Ramabhadran, Martin Franz, Samuel Gustman, James Mayfield, Liliya Kharevych, Stephanie M. Strassel |
SIGIR | 3 |
| 2004 | An appearance-based approach for consistent labeling of humans and objects in video
Martí Balcells, Daniel DeMenthon, David S. Doermann |
Pattern Anal. Appl. | 3 |
| 2004 | Machine Printed Text and Handwriting Identification in Noisy Document ImagesabstractIn this paper, we address the problem of the identification of text in noisy document images. We are especially focused on segmenting and identifying between handwriting and machine printed text because: 1) Handwriting in a document often indicates corrections, additions, or other supplemental information that should be treated differently from the main content and 2) the segmentation and recognition techniques requested for machine printed and handwritten text are significantly different. A novel aspect of our approach is that we treat noise as a separate class and model noise based on selected features. Trained Fisher classifiers are used to identify machine printed text and handwriting from noise and we further exploit context to refine the classification. A Markov Random Field-based (MRF) approach is used to model the geometrical structure of the printed text, handwriting, and noise to rectify misclassifications. Experimental results show that our approach is robust and can significantly improve page segmentation in noisy document collections. Yefeng Zheng 0001, Huiping Li 0001, David S. Doermann |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2004 | Automatic recognition of spontaneous speech for access to multilingual oral history archivesabstractMuch is known about the design of automated systems to search broadcast news, but it has only recently become possible to apply similar techniques to large collections of spontaneous speech. This paper presents initial results from experiments with speech recognition, topic segmentation, topic categorization, and named entity detection using a large collection of recorded oral histories. The work leverages a massive manual annotation effort on 10 000 h of spontaneous speech to evaluate the degree to which automatic speech recognition (ASR)-based segmentation and categorization techniques can be adapted to approximate decisions made by human annotators. ASR word error rates near 40% were achieved for both English and Czech for heavily accented, emotional and elderly spontaneous speech based on 65-84 h of transcribed speech. Topical segmentation based on shifts in the recognized English vocabulary resulted in 80% agreement with manually annotated boundary positions at a 0.35 false alarm rate. Categorization was considerably more challenging, with a nearest-neighbor technique yielding F=0.3. This is less than half the value obtained by the same technique on a standard newswire categorization benchmark, but replication on human-transcribed interviews showed that ASR errors explain little of that difference. The paper concludes with a description of how these capabilities could be used together to search large collections of recorded oral histories. William J. Byrne, David S. Doermann, Martin Franz, Samuel Gustman, Jan Hajic 0001, Douglas W. Oard, Michael Picheny, Josef Psutka, Bhuvana Ramabhadran, Dagobert Soergel, Todd Ward, Wei-Jing Zhu |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | Progress in Camera-Based Document Image AnalysisabstractThe increasing availability of high performance, low priced, portable digital imaging devices has created a tremendous opportunity for supplementing traditional scanning for document image acquisition. Digital cameras attached to cellular phones, PDAs, or as standalone still or video devices are highly mobile and easy to use; they can capture images of any kind of document including very thick books, historical pages too fragile to touch, and text in scenes; and they are much more versatile than desktop scanners. Should robust solutions to the analysis of documents captured with such devices become available, there is clearly a demand from many domains. Traditional scanner-based document analysis techniques provide us with a good reference and starting point, but they cannot be used directly on camera-captured images. Camera captured images can suffer from low resolution, blur, and perspective distortion, as well as complex layout and interaction of the content and background. In this paper we present a survey of application domains, technical challenges and solutions for recognizing documents captured by digital cameras. We begin by describing typical imaging devices and the imaging process. We discuss document analysis from a single camera-captured image as well as multiple frames and highlight some sample applications under development and feasible ideas for future development. David S. Doermann, Huiping Li 0001 |
ICDAR | 1 |
| 2003 | Combining Multiple Classifiers based on Third-Order DependencyabstractWithout an independence assumption, combining multiple classifiers deals with a high order probability distribution composed of classifiers and a class label. Storing and estimating the high order probability distribution is exponentially complex and unmanageable in theoretical analysis, so we rely on an approximation scheme using the dependency. In this paper, as an extension of the second-order dependency approach, the probability distribution is optimally approximated by the third-order dependency and multiple classifiers are combined. The proposed method is evaluated on the recognition of unconstrained handwritten numerals from Concordia University and the University of California, Irvine. Experimental results support the proposed method as a promising approach. Hee-Joong Kang, David S. Doermann |
ICDAR | 2 |
| 2003 | Evaluation of the Information-Theoretic Construction of Multiple Classifier SystemsabstractThe performance of multiple classifier systems varieswith the performance of component classifiers as well asthe method of combination. In this paper, information-theoreticmethods are proposed for constructing multipleclassifier systems, provided that the number of componentclassifiers is constrained in advance. These proposed methodsare applied to a classifier pool and examine the possibleclassifier sets by the selected information-theoretic criteria.One of them is then selected as the candidate and isevaluated together with the other multiple classifier systemson the recognition of unconstrained handwritten numeralsfrom Concordia University and the University of California,Irvine. Experimental results support the approach. Hee-Joong Kang, David S. Doermann |
ICDAR | 2 |
| 2003 | Gabor Filter Based Multi-class Classifier for Scanned Document ImagesabstractWhen scanning documents with a large number of pagessuch as books, it is often feasible to provide a minimalnumber of training samples to personalize the system tocompensate for global shifts in how the document wascreated or in scanning parameters. In this paper, wepresent a supervised multi-class classifier based onGabor filters that is used to classify the scripts, font-faces,and font-styles (bold, italic, normal etc.) in anapplication where the classes are known. Classificationis performed at the word level (glyphs separated by whitespace) given training samples of each class. This methodwas applied to a variety of bilingual dictionaries toidentify different scripts, and simultaneously, to classifyRoman scripts into bold, italic and normal font-styles.Experimental results show the effectiveness of thisapproach in increasing performance over classifierstrained for general documents. Huanfeng Ma, David S. Doermann |
ICDAR | 2 |
| 2003 | A Model-based Line Detection Algorithm in DocumentsabstractIn this paper we present a novel model based approach to detect severely broken parallel lines in noisy textual documents. It is important to detect and remove these lines so the text can be segmented and recognized. We use directional single-connected chain, a vectorization based algorithm, to extract the line segments. We then instantiate a parallel line model with three parameters: the skew angle, the vertical line gap, and the vertical translation. A coarse-to-fine approach is used to improve the estimation accuracy. From the model we can incorporate the high level contextual information to enhance detection results even when lines are severely broken. Our experimental results show our method can detect 94% of the lines in our database with 168 noisy Arabic document images. Yefeng Zheng 0001, Huiping Li 0001, David S. Doermann |
ICDAR | 3 |
| 2003 | Text Identification in Noisy Document Images Using Markov Random FieldabstractIn this paper we address the problem of the identification of text from noisy documents. We segment and identify handwriting from machine printed text because 1) handwriting in a document often indicates corrections, additions or other supplemental information that should be treated differently from the main body or body content, and 2) the segmentation and recognition techniques for machine printed text and handwriting are significantly different. Our novelty is that we treat noise as a separate class and model noise based on selected features. Trained Fisher classifiers are used to identify machine printed text and handwriting from noise. We further exploit context to refine the classification. A Markov random field (MRF) based approach is used to model the geometrical structure of the printed text, handwriting and noise to rectify the mis-classification. Experimental results show our approach is promising and robust, and can significantly improve the page segmentation results in noise documents. Yefeng Zheng 0001, Huiping Li 0001, David S. Doermann |
ICDAR | 3 |
| 2003 | An appearance based approach for human and object trackingabstractA system for tracking humans and detecting human-object interactions in indoor environments is described. A combination of correlogram and histogram information is used to model object and human color distributions. Humans and objects are detected using a background subtraction algorithm. The models are built on the fly and used to track them on a frame by frame basis. The system is able to detect when people merge into groups and segment them during occlusion. Identities are preserved during the sequence, even if a person enters and leaves the scene. The system is also able to detect when a person deposits or removes an object from the scene. In the first case the models are used to track the object retroactively in time. In the second case the objects are tracked for the rest of the sequence. Experimental results using indoor video sequences are presented. Marti Balcells Capellades, David S. Doermann, Daniel DeMenthon, Rama Chellappa |
ICIP (2) | 2 |
| 2003 | Issues in the transmission, analysis, storage and retrieval of surveillance videoabstractIncreased network capabilities and the prevalence of wireless networks have provided an environment where data and compute intensive surveillance applications can be realized as part of existing infrastructures. Nevertheless, careful consideration must be given to how such systems are designed an implemented. In our efforts to consider pervasive surveillance applications, this paper highlights some of the primarily issues related to the transmission, analysis, storage and retrieval of surveillance video and briefly describes our framework for video surveillance from mobile devices. David S. Doermann, Arvind Karunanidhi, Niketu Parekh, Hasan Timucin Ozdemir, M. Miwa, Kuo Chu Lee |
ICME | 1 |
| 2003 | Sports video classification using HMMSabstractIn this paper we address the problem of sports video classification using hidden Markov models (HMMs). For each sports genre, we construct two HMMs representing motion and color features respectively. The observation sequences generated from the principal motion direction and the principal color of each frame are fed to a motion and a color HMM respectively. The outputs are integrated to make a final decision. We tested our scheme on 220 minutes of sports video with four genre types: ice hockey, basketball, football, and soccer, and achieved an overall classification accuracy of 93%. Xavier Gibert, Huiping Li 0001, David S. Doermann |
ICME | 3 |
| 2003 | Video retrieval using spatio-temporal descriptorsabstractThis paper describes a novel methodology for implementing video search functions such as retrieval of near-duplicate videos and recognition of actions in surveillance video. Videos are divided into half-second clips whose stacked frames produce 3D space-time volumes of pixels. Pixel regions with consistent color and motion properties are extracted from these 3D volumes by a threshold-free hierarchical space-time segmentation technique. Each region is then described by a high-dimensional point whose components represent the position, motion and, when possible, color of the region. In the indexing phase for a video database, these points are assigned labels that specify their video clip of origin. All the labeled points for all the clips are stored into a single binary tree for e#cient k-nearest neighbor retrieval. The retrieval phase uses video segments as queries. Half-second clips of these queries are again segmented to produce sets of points, and for each point the labels of its nearest neighbors are retrieved. The labels that receive the largest numbers of votes correspond to the database clips that are the most similar to the query video segment. We illustrate this approach for video indexing and retrieval and for action recognition. First, we describe retrieval experiments for dynamic logos, and for video queries that di#er from the indexed broadcasts by the addition of large overlays. Then we describe experiments in which o#ce actions (such as pulling and closing drawers, taking and storing items, picking up and putting down a phone) are recognized. Color information is ignored to insure independence to people's appearance. One of the distinct advantages of using this approach for action recognition is that there is no need for detection or recognition of body pa... Daniel DeMenthon, David S. Doermann |
ACM Multimedia | 2 |
| 2003 | Acquisition of bilingual MT lexicons from OCRed dictionariesabstractThis paper describes an approach to analyzing the lexical structure of OCRed bilingual dictionaries to construct resources suited for machine translation of low-density languages, where online resources are limited. A rule-based, an HMM-based, and a post-processed HMM-based method are used for rapid construction of MT lexicons based on systematic structural clues provided in the original dictionary. We evaluate the effectiveness of our techniques, concluding that: (1) the rule-based method performs better with dictionaries where the font is not an important distinguishing feature for determining information types; (2) the post-processed stochastic method improves the results of the stochastic method for phrasal entries; and (3) Our resulting bilingual lexicons are comprehensive enough to provide the basis for reasonable translation results when compared to human translations. Burcu Karagol Ayan, David S. Doermann, Bonnie J. Dorr |
MTSummit | 2 |
| 2003 | Desparately Seeking Cebuano
Douglas W. Oard, David S. Doermann, Bonnie J. Dorr, Daqing He, Philip Resnik, Amy Weinberg, William J. Byrne, Sanjeev Khudanpur, David Yarowsky, Anton Leuski, Philipp Koehn, Kevin Knight |
HLT-NAACL | 2 |
| 2003 | Adaptive Hindi OCR using generalized Hausdorff image comparisonabstractWe present an adaptive Hindi OCR implemented as part of a rapidly retargetable language tool effort. The system includes: script identification, character segmentation, training sample creation, and character recognition. In script identification, Hindi words are identified from bilingual or multilingual documents based on features of the Devanagari script or using Support Vector Machines. Identified words are then segmented into individual characters in the next step, where the composite characters are identified and further segmented based on the structural properties of the script and statistical information. Segmented characters are recognized using generalized Hausdorff image comparison (GHIC) and postprocessing is applied to improve the performance. The OCR system, which was designed and implemented in one month, was applied to a complete Hindi--English bilingual dictionary and a set of ideal images extracted from Hindi documents in PDF format. Experimental results show the recognition accuracy can reach 88% for noisy images and 95% for ideal images. The presented method can also be extended to design OCR systems for different scripts. Huanfeng Ma, David S. Doermann |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2002 | Logical Labeling of Document Images Using Layout Graph Matching with Adaptive Learning
David S. Doermann |
Document Analysis Systems | 2 |
| 2002 | The Segmentation and Identification of Handwriting in Noisy Document Images
Yefeng Zheng 0001, Huiping Li 0001, David S. Doermann |
Document Analysis Systems | 3 |
| 2001 | Hybrid independent component analysis and support vector machine learning scheme for face detectionabstractWe propose a new hybrid unsupervised/supervised learning scheme that integrates independent component analysis (ICA) with the support vector machine (SVM) approach and apply this new learning scheme to the face detection problem. In low-level feature extraction, ICA produces independent image bases that emphasize edge information in the image data. In high-level classification, SVM classifies the ICA features as a face or non-faces. Our experimental results show that by using ICA features we obtain a larger margin of separation and fewer support vectors than by training SVM directly on the image data. This indicates better generalization performance, which is verified in our experiments. Yuan Qi 0001, David S. Doermann, Daniel DeMenthon |
ICASSP | 2 |
| 2001 | Classification of document pages using structure-based features
Christian K. Shin, David S. Doermann, Azriel Rosenfeld |
Int. J. Document Anal. Recognit. | 2 |
| 2001 | Forgery Detection by Local CorrespondenceabstractSignatures may be stylish or unconventional and have many personal characteristics that are challenging to reproduce by anyone other than the original author. For this reason, signatures are used and accepted as proof of authorship or consent on personal checks, credit purchases and legal documents. Currently signatures are verified only informally in many environments, but the rapid development of computer technology has stimulated great interest in research on automated signature verification and forgery detection. In this paper, we focus on forgery detection of offline signatures. Although a great deal of work has been done on offline signature verification over the past two decades, the field is not as mature as online verification. Temporal information used in online verification is not available offline and the subtle details necessary for offline verification are embedded at the stroke level and are hard to recover robustly. We approach the offline problem by establishing a local correspondence between a model and a questioned signature. The questioned signature is segmented into consecutive stroke segments that are matched to the stroke segments of the model. The cost of the match is determined by comparing a set of geometric properties of the corresponding substrokes and computing a weighted sum of the property value differences. The least invariant features of the least invariant substrokes are given the biggest weights, thus emphasizing features that are highly writer-dependent. Random forgeries are detected when a good correspondence cannot be found, i.e. the process of making the correspondence yields a high cost. Many simple forgeries can also be identified in this way. The threshold for making these decisions is determined by a Gaussian statistical model. Using the local correspondence between the model and a questioned signature, we perform skilled forgery detection by examining the writer-dependent information embedded at the substroke level and try to capture unballistic motion and tremor information in each stroke segment, rather than as global statistics. Experiments on random, simple and skilled forgery detection are presented. Jinhong Katherine Guo, David S. Doermann, Azriel Rosenfeld |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2000 | Image Distance Using Hidden Markov ModelsabstractWe describe a method for learning statistical models of images using a second-order hidden Markov mesh model. First, an image can be segmented in a way that best matches its statistical model by an approach related to the dynamic programming used for segmenting Markov chains. Next, given an image segmentation, a statistical model (3D state transition matrix and observation distributions within states) can be estimated. These two steps are repeated until convergence to provide both a segmentation and a statistical model of the image. We propose a statistical distance measure between images based on the similarity of their statistical models, for classification and retrieval tasks. Daniel DeMenthon, David S. Doermann, Marc Vuilleumier Stückelberg |
ICPR | 2 |
| 2000 | Tools and Techniques for Video Performance EvaluationabstractWe outline a reconfigurable video performance evaluation resource (ViPER), which provides an interface for ground truth generation, metrics for evaluation and tools for visualization of video analysis results. A key component is that the approach provides the basic infrastructure, and allows users to configure data generation and evaluation. Although ViPER can be used for any type of data, we focus on applications which require video content. David S. Doermann, David Mihalcik |
ICPR | 1 |
| 2000 | Off-Line Skilled Forgery Detection Using Stroke and Sub-Stroke PropertiesabstractResearch has been active in the field of forgery detection, but relatively little work has been done on the detection of skilled forgeries. We present an algorithm for detecting skilled forgeries based on a local correspondence between a questioned signature and a model obtained a priori. Writer-dependent properties are measured at the substroke level and a cost function is trained for each writer. When a candidate signature is presented, the same features are extracted and matched against the model. We present a description of the features and experimental results. Jinhong Katherine Guo, David S. Doermann, Azriel Rosenfeld |
ICPR | 2 |
| 2000 | Superresolution-Based Enhancement of Text in Digital VideoabstractWe present a superresolution-based text enhancement scheme to improve optical character recognition (OCR) and readability of text in digital video. The quality of video text is degraded by factors such as low resolution, antialiasing, clutter and blurring. We use the fact that the same text string often exists in consecutive frames to explore the temporal information and enhance the text image. For graphic text, we register text blocks to subpixel accuracy and fuse them to a new block with high resolution and a cleaner background. We use a projection onto convex sets based method to deblur scene text to improve readability. Experimental results show our scheme can improve OCR for graphic text and readability for scene text significantly. Huiping Li 0001, David S. Doermann |
ICPR | 2 |
| 2000 | A Video Text Detection System Based on Automated TrainingabstractIn this paper we present a video text detection system based on automated neural network training. Compared with previous work which detects only graphical text with fixed parameters, our system (1) provides a training mechanism so the parameters of the system can be adapted to changing environments, (2) can detect both graphical text and scene text located in complex backgrounds, (3) can detect text in any orientation and (4) can perform multilingual text detection. Experiments show the effectiveness of our system in various text detection tasks. Huiping Li 0001, David S. Doermann |
ICPR | 2 |
| 2000 | Event Detection from MPEG Video in the Compressed DomainabstractThe paper describes two techniques for detecting dynamic events using the motion vectors obtained from the MPEG video encoding. In the first technique, feature vectors from motion information form a high dimensional curve for a video clip, and curve simplification allows us to browse the clip along portions where interesting events are likely to happen. In the second technique, the camera motion, pan, tilt, and zoom are computed from the motion vectors, and the residual vectors which are not explained by the camera motion are regarded as generated by moving blobs. Events are detected from these moving blobs. Kyongil Yoon, Daniel DeMenthon, David S. Doermann |
ICPR | 3 |
| 2000 | Residual coding in document image compressionabstractSymbolic document image compression relies on the detection of similar patterns in a document image and construction of a prototype library. Compression is achieved by referencing multiple pattern instances ("components") through a single representative prototype. To provide a lossless compression, however, the residual difference between each component and its assigned prototype must be coded. Since the size of the residual can significantly affect the compression ratio, efficient coding is essential. In this paper, we describe a set of residual coding models for use with symbolic document image compression that exhibit desirable characteristics for compression and rate-distortion and facilitate compressed-domain processing. The first model orders the residual pixels by their distance to the prototype edge. Grouping pixels based on this distance value allows for a more compact coding and lower entropy. This distance model is then extended to a model that defines the structure of the residue and uses it as a basis for continuous and packet reconstruction which provides desired functionality for use in lossy compression and progressive transmission. Omid E. Kia, David S. Doermann |
IEEE Trans. Image Process. | 2 |
| 2000 | Automatic text detection and tracking in digital videoabstractText that appears in a scene or is graphically added to video can provide an important supplemental source of index information as well as clues for decoding the video's structure and for classification. In this work, we present algorithms for detecting and tracking text in digital video. Our system implements a scale-space feature extractor that feeds an artificial neural processor to detect text blocks. Our text tracking scheme consists of two modules: a sum of squared difference (SSD)-based module to find the initial position and a contour-based module to refine the position. Experiments conducted with a variety of video sources show that our scheme can detect and track text robustly. Huiping Li 0001, David S. Doermann, Omid E. Kia |
IEEE Trans. Image Process. | 2 |
| 1999 | On Musical Score Recognition using Probabilistic ReasoningabstractWe present a probabilistic framework for document analysis and recognition and illustrate it on the problem of musical score recognition. Our system uses an explicit descriptive model of the document class to find the most likely interpretation of a scanned document image. In contrast to the traditional pipeline architecture, we carry out all stages of the analysis with a single inference engine, allowing for an end-to-end propagation of the uncertainty. The global modeling structure is similar to a stochastic attribute grammar, and local parameters are estimated using hidden Markov models. Marc Vuilleumier Stückelberg, David S. Doermann |
ICDAR | 2 |
| 1999 | A Robust Method for Unknown Forms AnalysisabstractThis paper proposes a strategy for analyzing unknown, filled forms. First, horizontal and vertical line segments are detected, extracted and filtered. A recursive splitting and merging algorithm eliminates overlapping segments, filters false segments, and groups the segments into lines. Based on the extracted lines, an algorithm for rectangle extraction is proposed. We define the constraints between rectangles and edges. In a process of scanning the horizontal and vertical lines, candidate edges are validated and rectangles are generated if its surrounding edges and their combination are all valid. The process is recursively applied. It can tolerate large breaks in form lines, ignore irrelevant segments and deal with embedded rectangles. Experiments on a collection of forms show that our approach works well on poor quality images. Xingyuan Li 0003, Wen Gao 0001, David S. Doermann, Weon-Geun Oh |
ICDAR | 3 |
| 1999 | Building mosaics from video using MPEG motion vectorsabstractIn this paper we present a novel way of creating mosaics from an MPEG video sequence. Two original aspects of our work are that (1) we explicitly compute camera motion between frames and (2) we deduce the camera motion directly from the motion vectors encoded in the MPEG video stream. This enables us to create mosaics more simply and quickly than with other methods. 1 Introduction The mass digitization of video has elevated automated storage and retrieval to a grand challenge. Video sequences can store a vast amount of useful information, but redundancy between individual frames is a problem when analyzing, browsing, or searching video. Presenting the video sequence in a compact manner is a difficult challenge because eliminating redundancy could also eliminate content. The selection of static keyframes to represent a shot sequence is commonly used for indexing as well as for presentation of retrieval results. This technique is insufficient for revealing much of the content in a vide... Ryan C. Jones, Daniel DeMenthon, David S. Doermann |
ACM Multimedia (2) | 3 |
| 1999 | Text enhancement in digital video using multiple frame integrationabstractIn this paper a multiple frame based technique to enhance text in digital video is presented. After extracting a reference text block, we use an image matching technique to find the corresponding text blocks in consecutive frames. We register these text blocks to subpixel levels by using image interpolation techniques to improve both correspondence and text resolution. The registered text blocks are averaged to obtain a new text block with a clean background and a higher resolution. Experiments conducted on several video sequences show that our enhancement scheme can improve the accuracy of commercial off-the-shelf OCR considerably. Huiping Li 0001, David S. Doermann |
ACM Multimedia (1) | 2 |
| 1999 | Detection of slow-motion replay sequences for identifying sports videosabstractAutomated classification of digital video is emerging as an important piece of the puzzle in the design of content management systems for digital libraries. The ability to classify videos into various genres such as sports, news, movies, or documentaries increases the efficiency of indexing, browsing, and retrieval of video in large databases. In this paper, we present an automated technique for identifying slow-motion replays directly from the compressed domain of MPEG video. It uses the macroblock, motion, and bit-rate information that is readily accessible from MPEG video with very minimal decoding, leading to enormous gains in processing speeds. Vikrant Kobla, Daniel DeMenthon, David S. Doermann |
MMSP | 3 |
| 1998 | Text Extraction, Enhancement and OCR in Digital Video
Huiping Li 0001, David S. Doermann, Omid E. Kia |
Document Analysis Systems | 2 |
| 1998 | Automatic identification of text in digital video key framesabstractScene and graphic text can provide important supplemental index information in video sequences. In this paper we address the problem automatically identifying text regions in digital video key frames. The text contained in video frames is typically very noisy because it is aliased and/or digitized at a much lower resolution than typical document images, making identification, extraction and recognition difficult. The proposed method is based on the use of a hybrid wavelet/neural network segmenter on a series of overlapping small windows to classify regions which contain text. To detect text over a wide range of font sizes, the method is applied to a pyramid of images and the regions identified at each level are integrated. Huiping Li 0001, David S. Doermann |
ICPR | 2 |
| 1998 | Video Summarization by Curve SimplificationabstractA video sequence can be represented as a trajectory curve in a high dimensional feature space. This video curve can be analyzed by tools similar to those developed for planar curves. In particular, the classic binary curve splitting algorithm has been found to be a useful tool for video analysis. With a splitting condition that checks the dimensionality of the curve segment being split, the video curve can be recursively simplified and represented as a tree structure, and the frames that are found to be junctions between curve segments at different levels of the tree can be used as keyframes to summarize the video sequences at different levels of detail. These keyframes can be combined in various spatial and temporal configurations for browsing purposes. We describe a simple video player that displays the keyframes sequentially and lets the user change the summarization level on the fly with a slider. We also describe an approach to automatically selecting a summarization level that provides a concise and representative set of keyframes. Daniel DeMenthon, Vikrant Kobla, David S. Doermann |
ACM Multimedia | 3 |
| 1998 | Automatic text tracking in digital videosabstractWe address the problem of automatically tracking moving text in digital videos. Our scheme consists of two separate processes: monitoring which detects the new text line entering a frame, and tracking which uses prediction techniques to reconcile the text from frame to frame. Temporal consistency allows one to monitor periodically and reduce the computation complexity. The tracking process uses a rapid prediction/search scheme to update the position of the text blocks between frames. We provide details of the implementation and results for text moving in the scene and text which moves as a result of camera motion. Huiping Li 0001, David S. Doermann |
MMSP | 2 |
| 1998 | The Indexing and Retrieval of Document Images: A Survey
David S. Doermann |
Comput. Vis. Image Underst. | 1 |
| 1998 | Symbolic Compression and Processing of Document Images
Omid E. Kia, David S. Doermann, Azriel Rosenfeld, Rama Chellappa |
Comput. Vis. Image Underst. | 2 |
| 1998 | Editorial
David S. Doermann, Seong-Whan Lee, Sargur N. Srihari, Karl Tombre, Azriel Rosenfeld |
Int. J. Document Anal. Recognit. | 1 |
| 1998 | The detection of duplicates in document image databases
David S. Doermann, Huiping Li 0001, Omid E. Kia |
Image Vis. Comput. | 1 |
| 1998 | The function of documents
David S. Doermann, Ehud Rivlin, Azriel Rosenfeld |
Image Vis. Comput. | 1 |
| 1997 | The Retrieval of Document Images: A Brief SurveyabstractThe economic feasibility of creating large databases of document images has left a tremendous need for robust ways to access the information these images contain. Printed documents are often scanned for archiving or an an attempt to move toward a paper-less office and stored as images, but without adequate index information. In order to make full use of the capabilities of traditional database indexing and retrieval techniques, a full conversion of the document may be required. There are many factors, however, which may prohibit complete conversion including its high cost, insufficient document quality, or the fact that parts of the document can simply not be adequately represented in a converted form. In this paper, we provide a survey of methods developed by researchers to access document images without relying on complete and accurate conversion. We briefly discuss traditional text indexing techniques on imperfect data and the retrieval of partially converted documents, followed by a more complete review of techniques for the direct retrieval and characterization of document images including text, drawings and graphics. David S. Doermann |
ICDAR | 1 |
| 1997 | The Detection of Duplicates in Document Image DatabasesabstractWe propose and implement a method for detecting duplicate documents in very large image databases. The method is based on a robust "signature" extracted from each document image which is used to index into a table of previously processed documents. The approach has a number of advantages over OCR or other recognition based methods, including speed and robustness to imaging distortions. To justify the approach and test the scalability, we have developed a simulator which allows us to change parameters of the system and examine performance for millions of document signatures. A complete system is implemented and tested on a test collection of technical articles and memos. David S. Doermann, Huiping Li 0001, Omid E. Kia |
ICDAR | 1 |
| 1997 | The Function of DocumentsabstractThe purpose of a document is to facilitate the transfer of information from its author to its readers. It is the author's job to design the document so that the information it contains can be interpreted accurately and efficiently. To do this, the author can make use of a set of stylistic tools. In this paper, we introduce the concept of document functionality, which attempts to describe the roles of documents and their components in the process of transferring information. A functional description of a document provides insight into the type of the document, into its intended uses, and into strategies for automatic document interpretation and retrieval. To demonstrate these ideas, we define a taxonomy of functional document components and show how functional descriptions can be used to reverse-engineer the intentions of the author, to navigate in document space, and to provide important contextual information to aid in interpretation. David S. Doermann, Azriel Rosenfeld, Ehud Rivlin |
ICDAR | 1 |
| 1997 | Local correspondence for detecting random forgeriesabstractProgress on the problem of signature verification has advanced more rapidly in online applications than offline applications, in part because information which can easily be recorded in online environments, such as pen position and velocity, is lost in static offline data. In offline applications, valuable information which can be used to discriminate between genuine and forged signatures is embedded at the stroke level. We present an approach to segmenting strokes into stylistically meaningful segments and establish a local correspondence between a questioned signature and a reference signature to enable the analysis and comparison of stroke features. Questioned signatures which do not conform to the reference signature are identified as random forgeries. Most simple forgeries can also be identified, as they do not conform to the reference signature's invariant properties such as connections between letters. Since we have access to both local and global information, our approach also shows promise for extension to the identification of skilled forgeries. Jinhong Katherine Guo, David S. Doermann, Azriel Rosenfeld |
ICDAR | 2 |
| 1997 | A distributed management system for testing document image analysis algorithmsabstractWe describe a new approach to manage the testing of document analysis and understanding applications. We propose and present a collection of document images, a set of techniques to prepare the test cases interactively and means to control the testing process. The systems architecture is designed to be distributed, scalable and platform independent utilizing Java, C++ and object-oriented databases. The main features of this system are a basic document categorization and ground truth, degradation models, custom test case creation facilities, a test management module (pipelining, test history), the ability to embed document analysis algorithms into the system, remote usage facilities and robust graphical user interfaces. Jaakko J. Sauvola, Sami Haapakoski, Hannu Kauniskangas, Tapio Seppänen, Matti Pietikäinen, David S. Doermann |
ICDAR | 6 |
| 1997 | OCR-based rate-distortion analysis of residual codingabstractSymbolic compression of document images provides access to symbols found in document images and exploits the redundancy found within them. Document images are highly structured and contain large numbers of repetitive symbols. We have shown that while symbolically compressing a document image we are able to perform compressed-domain processing. Symbolic compression forms representative prototypes for symbols and encode the image by the location of these prototypes and a residual (the difference between symbol and prototype). We analyze the rate-distortion tradeoff by varying the amount of residual used in compression for both distance- and row-order coding. A measure of distortion is based on the performance of an OCR system on the resulting image. The University of Washington document database images, ground truth, and OCR evaluation software are used for experiments. Omid E. Kia, David S. Doermann |
ICIP (3) | 2 |
| 1997 | VideoTrails: Representing and Visualizing Structure in Video SequencesabstractThe problem of determining the physical and semantic structure of an extended video sequence is essential for providing appropriate processing, indexing and retrieval capabilities for video databases.In this paper, we describe a novel technique which reduces a sequence of MPEG encoded video frames to a trail of points in a low dimensional space.In thii space, we can cluster frames, analyze transitions between clusters and compute properties of the resulting trail.By classifying portions of the trail as either stationary or transitional, we are able to detect gradual edits between shots.Furthermore, tracking the interaction of clusters over time, we lay the groundwork for the complete analysis and representation of the video's physical and semantic structure. Vikrant Kobla, David S. Doermann, Christos Faloutsos |
ACM Multimedia | 2 |
| 1997 | The role of compressed document images in transmission and retrievalabstractDocument images belong to a unique class of images where the information content is contained in the language represented by a series of symbols on the page, rather than in the visual objects themselves. For this reason, it is essential to preserve the fidelity of individual components when considering methods of compression. Likewise the component level structure should be a prime consideration when ordering information for lossy or progressive transmission. We refine our work on document image compression as it applies to transmission and retrieval. We first overview the basic compression scheme, then describe a structural hierarchy which provides desirable properties for transmission. We present the results of a rate distortion experiment and discuss the implications for network applications. Omid E. Kia, David S. Doermann |
MMSP | 2 |
| 1997 | Extraction of features for indexing MPEG-compressed videoabstractDevelopment of various multimedia applications is dependent on the availability of fast and efficient storage, browsing, indexing, and retrieval techniques. Given that video is stored efficiently in a compressed format, the costly overhead of decompression can be avoided by analyzing the compressed representation directly. We describe techniques that can be used to extract viable features for indexing shots of video directly from the compressed domain. We develop a type independent representation of frames present in an MPEG video and show how it can be used directly for indexing. Vikrant Kobla, David S. Doermann |
MMSP | 2 |
| 1997 | Multiscale Segmentation of Unstructured Document Pages Using Soft Decision IntegrationabstractWe present an algorithm for layout-independent document page segmentation based on document texture using multiscale feature vectors and fuzzy local decision information. Multiscale feature vectors are classified locally using a neural network to allow soft/fuzzy multi-class membership assignments. Segmentation is performed by integrating soft local decision vectors to reduce their "ambiguities". Kamran Etemad, David S. Doermann, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1996 | The Development of a General Framework for Intelligent Document Image RetrievalabstractWork has recently begun on a joint project between the Universities of Maryland and Oulu on the development of a system for Intelligent Document Image Retrieval (IDIR). The IDIR system will provide close connections with and utilization of document analysis and image processing techniques, advanced computing and networking, and modern approaches to database management. The system design consists of aggressively modularized components to enhance the development of individual parts which are used in the complete solution, including: Interface specifications, multipurpose feature extraction, an integrated efficient query language, physical retrieval from an object-oriented database, and delivery of retrieved objects. In this paper, we introduce the general framework, feature extraction modules, query capabilities, a graphical query interface, and the application interface. We demonstrate each component of the system and how the query mechanisms can be used to handle both content and struc... David S. Doermann, Jaakko J. Sauvola, Hannu Kauniskangas, Christian K. Shin, Matti Pietikäinen, Azriel Rosenfeld |
DAS | 1 |
| 1996 | Structure-preserving document image compressionabstractMaintaining a document in image form is often preferable in order to avoid the high cost of manual conversion or the introduction of large numbers of errors by automatic OCR and/or graphics interpretation. The large volume of data in the image can be greatly reduced by using compression techniques. Text-intensive document images typically have a great deal of redundancy in the bitmap representations of symbols, and we make use of that redundancy for compression by clustering components, representing each cluster by a template and encoding the error. Our method is novel in modeling the error associated with each cluster and in preserving structure, an important component for readability and processing. Omid E. Kia, David S. Doermann |
ICIP (1) | 2 |
| 1996 | Structural compression for document analysisabstractIn this paper we describe a structural compression technique to be used for document text image storage and retrieval. The primary objective is to provide an efficient representation, storage, transmission and display. A secondary objective is to provide an encoding which allows access to specified regions within the image and facilitates traditional document processing operations without requiring complete decoding. We describe an algorithm which symbolically decomposes a document image and structurally orders the error bitmap based on a probabilistic model. The resultant symbol and error representations lend themselves to reasonably high compression ratios and are structured so as to allow operations directly on the compressed image. The compression scheme is implemented and compared to traditional compression methods. Omid E. Kia, David S. Doermann |
ICPR | 2 |
| 1996 | Applying algebraic and differential invariants for logo recognition
David S. Doermann, Ehud Rivlin, Isaac Weiss |
Mach. Vis. Appl. | 1 |
| 1995 | Robust table-form structure analysis based on box-driven reasoningabstractTable form document structure analysis is an important problem in the document processing domain. The paper presents a method called Box Driven Reasoning (BDR) to robustly analyze the structure of table form documents which include touching characters and broken lines. Most previous methods employ a line oriented approach. Real documents are copied repeatedly and overlaid with printed data, resulting in characters which touch cells and lines which are broken. BDR deals with regions directly, in contrast with other previous methods. Experimental tests show that BDR reliably recognizes cells and strings in document images with touching characters and broken lines. Osamu Hori, David S. Doermann |
ICDAR | 2 |
| 1995 | Recovery of temporal information from static images of handwriting
David S. Doermann, Azriel Rosenfeld |
Int. J. Comput. Vis. | 1 |
| 1994 | Page segmentation using decision integration and wavelet packetsabstractA new algorithm for layout-independent document page segmentation is suggested. Text, image and graphics regions in a document image are treated as three different "texture" classes. Soft local decisions on small blocks are made using wavelet packet based feature vectors. Segmentation is performed by propagating and integrating soft local decisions over neighboring blocks, within and across scales. The "uncertainties" associated with local decisions are reduced as more contextual evidence is incorporated in the process of decision integration. The majority, taken over weighted combined votes, determines the final decision. The suggested algorithm is based on parallel independent computations which have low complexity. It can also be applied to other signal and image segmentation tasks. Kamran Etemad, David S. Doermann, Rama Chellappa |
ICPR (2) | 2 |
| 1994 | Instrument grasp: a model and its effects on handwritten strokes
David S. Doermann, Venugopal Varma, Azriel Rosenfeld |
Pattern Recognit. | 1 |
| 1993 | Image based typographic analysis of documentsabstractAn approach to image based typographic analysis of documents is provided. The problem requires a spatial understanding of the document layout as well as knowledge of the proper syntax. The system performs a page synthesis from the stream of formatting commands defined in a DVI file. Since the two-dimensional relationships between document components are not explicit in the page language, the authors develop a representation which preserves the two-dimensional layout, the read-order and the attributes of document components. From this hierarchical representation of the page layout we extract and analyze relevant typographic features such as margins, line and character spacing, and figure placement.> David S. Doermann, Richard Furuta |
ICDAR | 1 |
| 1993 | The processing of form documentsabstractAn overview of an approach to the generic modeling and processing of known forms is presented. The system provides a methodology by which models are generated from regions in the document based on their usage. Automatic extraction of an optimal set of features to be used for registration is proposed, and it is shown how specialized detectors can be designed for each feature based on their position, orientation and width properties. Registration of the form with the model is accomplished using probing to establish correspondence. Form components which are corrupted by markings are detected and isolated, the intersections are interpreted and the properties of the non-form markings are used to reconstruct the strokes through the intersections. The feasibility of these ideas is demonstrated through an implementation of key components of the system.> David S. Doermann, Azriel Rosenfeld |
ICDAR | 1 |
| 1993 | Logo recognition using geometric invariantsabstractThe problem of logo recognition is of great interest in the document domain, especially for databases, because of its potential for identifying the source of the document and its generality as a recognition problem. By recognizing the logo, one obtains semantic information about the document, which may be useful in deciding whether or not to analyze the textual components. A multi-level stages approach to logo recognition which uses global invariants to prune the database and local affine invariants to obtain a more refined match is presented. An invariant signature which can be used for matching under a variety of transformations is obtained. The authors provide a method of computing Euclidean invariants and show how to extend them to capture similarity, affine, and projective invariants when necessary. They implement feature detection, feature extraction, and local invariant algorithms and successfully demonstrate the approach on a small database.> David S. Doermann, Ehud Rivlin, Isaac Weiss |
ICDAR | 1 |
| 1992 | Recovery of temporal information from static images of handwritingabstractA taxonomy of local, regional, and global temporal clues that, along with a detailed examination of the document, allow temporal properties to be recovered from the image is provided. It is shown that this system will benefit from obtaining a comprehensive understanding of the handwriting signal and that it requires a detailed analysis of stroke and sub-stroke properties. It is suggested that this task requires breaking away from traditional thresholding and thinning techniques, and a framework for such analysis is presented. It is shown how the temporal clues can reliably be extracted from this framework and how many of the seemingly ambiguous situations can be resolved by the derived clues and knowledge of the writing process.> David S. Doermann, Azriel Rosenfeld |
CVPR | 1 |
| 1992 | Temporal clues in handwritingabstractHandwritten character recognition is typically classified as online or offline depending on the nature of the input data. Online data consists of a temporal sequence of instrument positions while offline data is in the form of a 2D image of the writing sample. Online recognition techniques have been relatively successful but have the disadvantage of requiring the data to be gathered during the writing process. This paper presents work on the extraction of temporal information from static images of handwriting and its implications for character recognition.> David S. Doermann, Azriel Rosenfeld |
ICPR (2) | 1 |