EDBT 2026 Demo / reviewers in the wild / expert
Xujun Peng
dblp:92/7442 · also Xu-Jun Peng
· DBLP profile ↗
44ranked-venue papers
14as first author
10since 2021 · last 2026
0000-0001-9373-7092ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 20 · 9 first-author · 1 since 2021Databases, data management, data science and information retrieval · 11 · 10 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Stroke-Based Cyclic Amplifier: Image Super-Resolution at Arbitrary Ultra-Large ScalesabstractPrior Arbitrary-Scale Image Super-Resolution (ASISR) methods often experience a significant performance decline when the upsampling factor exceeds the range covered by the training data, introducing substantial blurring. To address this issue, we propose a unified model, Stroke-based Cyclic Amplifier (SbCA), for ultra-large upsampling tasks. The key of SbCA is the stroke vector amplifier, which decomposes the image into a series of strokes represented as vector graphics for magnification. Then, the detail completion module also restores missing details, ensuring high-fidelity image reconstruction. Our cyclic strategy achieves ultra-large upsampling by iteratively refining details with this unified SbCA model, trained only once for all, while keeping sub-scales within the training range. Our approach effectively addresses the distribution drift issue and eliminates artifacts, noise and blurring, producing high-quality, high-resolution super-resolved images. Experimental validations on both synthetic and real-world datasets demonstrate that our approach significantly outperforms existing methods in ultra-large upsampling tasks (e.g. $\times 100$ ), delivering visual quality far superior to state-of-the-art techniques. Wenhao Guo 0003, Peng Lu 0007, Xujun Peng, Zhaoran Zhao, Sheng Li 0008 |
IEEE Trans. Image Process. | 3 |
| 2025 | Learnable adaptive bilateral filter for improved generalization in Single Image Super-Resolution
Wenhao Guo 0003, Peng Lu 0007, Xujun Peng, Zhaoran Zhao |
Pattern Recognit. | 3 |
| 2025 | Self-Supervised Photographic Image Layout Representation LearningabstractImage layout representation learning, which converts layouts into compact vectors, is essential for tasks such as image retrieval, editing, and generation. However, existing methods—especially those applied to photographic images—face several challenges: supervised methods rely on expensive labeled datasets, weakly-supervised methods struggle with generalization, and self-supervised methods are limited in handling the diversity of photographic layouts. To address these issues, we propose a novel heterogeneous layout graph that efficiently captures the layout information in images. The vertices of this graph represent the compositional primitives of the image, capturing their attributes, while the edges encode the relationships between these primitives. We also design effective pretext tasks to guide a layout encoder-decoder in self-supervised training, ultimately generating the layout graph embedding vector. Additionally, we introduce a new layout evaluation dataset—LODB—which features a richer variety of layout categories, significantly better label quality than existing datasets, and a more balanced distribution of semantic scenes across layout categories, providing a comprehensive benchmark for evaluation. Experiments on the LODB dataset demonstrate that our method outperforms existing approaches in representing photographic image layouts. Zhaoran Zhao, Peng Lu 0007, Xujun Peng, Wenhao Guo 0003 |
IEEE Trans. Multim. | 3 |
| 2024 | COLORFLOW: A Conditional Normalizing Flow for Image ColorizationabstractImage colorization is an ill-posed task, as objects within grayscale images can correspond to multiple colors, motivating researchers to establish a one-to-many relationship between objects and colors. Previous work mostly could only create an insufficient deterministic relationship. Normalizing flow can fully capture the color diversity from natural image manifold. However, classical flow often overlooks the color correlations between different objects, resulting in generating unrealistic color. To solve this issue, we propose a conditional normalizing flow, named ColorFlow, to jointly learn the one-to-many relationships between objects and colors, and the color correlations between different objects within image. To represent these color correlations in flow, we design a color distribution predictor to estimate the global color histogram of grayscale image as global tones, which is utilized as the mean value of flow’s latent variables. Experiments results show that ColorFlow outperforms state-of-the-art methods. Wang Yin, Peng Lu 0007, Xujun Peng |
ICASSP | 3 |
| 2024 | BCSCN: Reducing Domain Gap through Bézier Curve basis-based Sparse Coding Network for Single-Image Super-ResolutionabstractSingle Image Super-Resolution (SISR) is a pivotal challenge in computer vision, aiming to restore high-resolution (HR) images from their low-resolution (LR) counterparts. The presence of diverse degradation kernels creates a significant domain gap, limiting the effective generalization of models in real-world scenarios. This study introduces the Bézier Curve basis-based Sparse Coding Network (BCSCN), a preprocessing network designed to mitigate input distribution discrepancies between the training and testing phases of super-resolution networks. BCSCN achieves this by removing visual defects associated with the degradation kernel in LR images, such as artifacts, residual structures, and noise. Additionally, we propose a set of rewards to guide the search for basis coefficients in BCSCN, enhancing the preservation of main content while eliminating information related to degradation. The experimental results highlight the importance of BCSCN, showcasing its capacity to effectively reduce domain gaps and enhance the generalization of super-resolution networks. Wenhao Guo 0003, Peng Lu 0007, Xujun Peng, Zhaoran Zhao, Xiangtao Dong |
ACM Multimedia | 3 |
| 2024 | Learning Realistic Sketching: A Dual-agent Reinforcement Learning ApproachabstractThis paper presents a pioneering method for teaching computer sketching that transforms input images into sequential, parameterized strokes. However, two challenges are raised for this sketching task: weak stimuli during stroke decomposition and maintaining semantic correctness, stylistic consistency, and detail integrity in the final drawings. To tackle the challenge of weak stimuli, our method incorporates an attention agent, which enhances the algorithm's sensitivity to subtle canvas changes by focusing on smaller, magnified areas. Moreover, in enhancing the perceived quality of drawing outcomes, we integrate a sketching style feature extractor to seamlessly capture semantic information and execute style adaptation at feature level, alongside a drawing agent that decomposes strokes under the guidance of a fine-grained reward, thereby ensuring the integrity of sketch details. Based on dual intelligent agents, we have constructed an efficient sketching model. Experimental results attest to the superiority of our approach in both visual effects and perceptual metrics when compared to state-of-the-art techniques, confirming its efficacy in achieving realistic sketching. Peng Lu 0007, Xujun Peng, Wenhao Guo 0003, Zhaoran Zhao, Xiangtao Dong |
ACM Multimedia | 3 |
| 2023 | Learning to Draw Through A Multi-Stage Environment Model Based Reinforcement LearningabstractMachine drawing has gradually become a hot research topic in computer vision and robotics domains recently. However, decomposing a given target image from raster space into an ordered sequence and reconstructing those strokes is a challenging task. In this work, we focus on the drawing task for the images in various styles where the distribution of stroke parameters differs. We propose a multi-stage environment model based reinforcement learning (RL) drawing framework with fine-grained perceptual reward to guide the agent under this framework to draw details and an overall outline of the target image accurately. The experiments show that the visual quality of our method slightly outperforms SOTA method in nature and doodle style, while it outperforms the SOTA approaches by a large margin with high efficiency in sketch style. Peng Lu 0007, Xujun Peng |
ICIP | 3 |
| 2023 | RLSCNet: A Residual Line-Shaped Convolutional Network for Vanishing Point Detection
Peng Lu 0007, Xujun Peng, Wang Yin, Zhaoran Zhao |
MMM (2) | 3 |
| 2021 | Yes, "Attention Is All You Need", for Exemplar based ColorizationabstractConventional exemplar based image colorization tends to transfer colors from reference image only to grayscale image based on the semantic correspondence between them. But their practical capabilities are limited when semantic correspondence can hardly be found. To overcome this issue, additional information, such as colors from the database is normally introduced. However, it's a great challenge to consider color information from reference image and database simultaneously because there lacks a unified framework to model different color information and the multi-modal ambiguity in database cannot be removed easily. Also, it is difficult to fuse different color information effectively. Thus, a general attention based colorization framework is proposed in this work, where the color histogram of reference image is adopted as a prior to eliminate the ambiguity in database. Moreover, a sparse loss is designed to guarantee the success of information fusion. Both qualitative and quantitative experimental results show that the proposed approach achieves better colorization performance compared with the state-of-the-art methods on public databases with different quality metrics. Wang Yin, Peng Lu 0007, Zhaoran Zhao, Xujun Peng |
ACM Multimedia | 4 |
| 2021 | Learning the Relation Between Interested Objects and Aesthetic Region for Image CroppingabstractAs one of the fundamental techniques for image editing, image cropping discards irrelevant contents and remains the pleasing portions of the image to enhance the overall composition and achieve better visual/aesthetic perception. In this paper, we primarily focus on improving the efficiency of automatic image cropping, and on further exploring its potential in public datasets with high accuracy. From this perspective, we propose a deep learning based framework to learn the objects composition from photos with high aesthetic qualities, where an interested object region is detected through a convolutional neural network (CNN) based on the saliency map. The features of the detected interested objects are then fed into a regression network to obtain the final cropping result. Unlike the conventional methods that multiple candidates are proposed and evaluated iteratively, only a single interested object region is produced in our model, which is mapped to the final output directly. Thus, low computational resources are required for the proposed approach. The experimental results on the public datasets show that as a weakly supervised method, the proposed network outperforms the other weakly supervised methods on FLMS and FCD datasets and achieves comparable results to the existing methods on CUHK dataset. Furthermore, the proposed method is more efficient than these methods, where the processing speed is as fast as 20 ms per image. Peng Lu 0007, Xujun Peng, Xiaofu Jin |
IEEE Trans. Multim. | 3 |
| 2020 | Building Super-Resolution Image Generator for OCR Accuracy Improvement
Xujun Peng |
DAS | 1 |
| 2020 | Camera Captured DIQA with Linearity and Monotonicity Constraints
Xujun Peng |
DAS | 1 |
| 2020 | Weakly Supervised Real-time Image Cropping based on Aesthetic DistributionsabstractImage cropping is an effective tool to edit and manipulate images to achieve better aesthetic quality. Most existing cropping approaches rely on the two-step paradigm where multiple candidate cropping areas are proposed initially and the optimal cropping window is determined based on some quality criteria for these candidates afterwards. The obvious disadvantage of this mechanism is its low efficiency due to the huge searching space of candidate crops. In order to tackle this problem, a weakly supervised cropping framework is proposed, where the distribution dissimilarity between high quality images and cropped images is used to guide the coordinate predictor's training and the ground truths of cropping windows are not required by the proposed method. Meanwhile, to improve the cropping performance, a saliency loss is also designed in the proposed framework to force the neural network to focus more on the interested objects in the image. Under this framework, the images can be cropped effectively by the trained coordinate predictor in a one-pass favor without multiple candidates proposals, which ensures the high efficiency of the proposed system . Also, based on the proposed framework, many existing distribution dissimilarity measurements can be applied to train the image cropping system with high flexibility, such as likelihood based and divergence based distribution dissimilarity measure proposed in this work. The experiments on the public databases show that the proposed cropping method achieves the state-of-the-art accuracy, and the high computation efficiency as fast as 285 FPS is also obtained. Peng Lu 0007, Xujun Peng, Xiaojie Wang 0006 |
ACM Multimedia | 3 |
| 2020 | Gray2ColorNet: Transfer More Colors from Reference ImageabstractImage colorization is an effective approach to provide plausible colors for grayscale images, which can achieve better and pleasing visual qualities. Although exemplar based colorization approaches provide promising results, they are relied on semantic colors or global colors only from the reference images. For the former situation, when the correspondence between the input grayscale image and reference image is not established, the colors of the reference image cannot be transferred to the input grayscale image successfully. With the later circumstance, because only global colors are considered, it is hard to produce a color image whose objects have the same color as the reference image when they are semantically related. Thus, an end-to-end colorization network Gray2ColorNet is proposed in this work, where an attention gating mechanism based color fusion network is designed to accomplish the colorization tasks. Relied on the proposed method, the semantic colors and global color distribution from the reference image are fused effectively, which are transferred to the final color images along with the prior knowledge of colors contained in the training data. The experimental results demonstrate the superior colorization performances of the proposed method compared to other state-of-the-art approaches. Peng Lu 0007, Jinbei Yu, Xujun Peng, Zhaoran Zhao, Xiaojie Wang 0006 |
ACM Multimedia | 3 |
| 2019 | Document Binarization via Multi-resolutional Attention Model with DRD LossabstractDocument binarization which separates text from background is a critical pre-processing step for many high level document analysis tasks. Conventional document binarization approaches tend to use hand-craft features and empirical rules to simulate the degradation process of document image and accomplish the binarization task. In this paper, we propose a deep learning framework where the probability of text areas is inferred through a multi-resolutional attention model, which is consequently fed into a convolutional conditional random field (ConvCRF) to obtain the final binarized document image. In the proposed approach, the features of degraded document image are learned by neural networks and the relations between text areas and backgrounds are inferred by ConvCRF, which avoids the dependence of domain knowledge from researchers and has more generalization capabilities. The experimental results on public datasets show that the proposed method has superior binarization performance than the existing state-of-the-art approaches. Xujun Peng, Huaigu Cao |
ICDAR | 1 |
| 2019 | Gated CNN for visual quality assessment based on color perception
Peng Lu 0007, Xujun Peng, Jinbei Yu |
Signal Process. Image Commun. | 2 |
| 2019 | Aesthetic guided deep regression network for image cropping
Peng Lu 0007, Xujun Peng |
Signal Process. Image Commun. | 3 |
| 2018 | Deep Conditional Color Harmony Model for Image Aesthetic AssessmentabstractAs one of the important features, color provides plenty useful information to represent images. Thus, color harmony, which is defined as “two or more colors are sensed together as a single, pleasing, collective impression” [1], can also be served as a fundamental feature and plays a key role to determine the aesthetics quality of images. To reveal the inherent color harmony attribute within patch and the harmonious relations between image patches which construct the pleasing colorful images, we designed a conditional random field (CRF) based color harmony model in this paper to accomplish the image aesthetic assessment tasks. Unlike the previous learning based color harmony models, we used deep neural networks to obtain the coherence properties between original image patch pairs, and embedded these relations along with each patch's own color harmony characteristic into a CRF to measure the harmony scores of the entire image. The experimental results on a public dataset show that the proposed deep conditional color harmony model is superior to the existing color harmony models in respect of the image aesthetic assessment. Peng Lu 0007, Jinbei Yu, Xujun Peng |
ICPR | 3 |
| 2017 | Using Convolutional Encoder-Decoder for Document Image BinarizationabstractDocument image binarization is one of the critical initial steps for document analysis and understanding. Previous work mostly focused on exploiting hand-crafted features to build statistical models for distinguishing text from background. However, these approaches only achieved limited success because: (a) the effectiveness of hand-crafted features is limited by the researcher's domain knowledge and understanding on the documents, and (b) a universal model cannot always capture the complexity of different document degradations. In order to address these challenges, we propose a convolutional encoder-decoder model with deep learning for document image binarization in this paper. In the proposed method, mid-level document image representations are learnt by a stack of convolutional layers, which compose the encoder in this architecture. Then the binarization image is obtained by mapping low resolution representations to the original size through the decoder, which is composed by a series of transposed convolutional layers. We compare the proposed binarization method with other binarization algorithms both qualitatively and quantitatively on the public dataset. The experimental results show that the proposed method has comparable performance to the other hand-crafted binarization approaches and has more generalization capabilities with limited in-domain training data. Xujun Peng, Huaigu Cao, Premkumar Natarajan |
ICDAR | 1 |
| 2016 | Document Image Quality Assessment Using Discriminative Sparse RepresentationabstractThe goal of document image quality assessment (DIQA) is to build a computational model which can predict the degree of degradation for document images. Based on the estimated quality scores, the immediate feedback can be provided by document processing and analysis systems, which helps to maintain, organize, recognize and retrieve the information from document images. Recently, the bag-of-visual-words (BoV) based approaches have gained increasing attention from researchers to fulfill the task of quality assessment, but how to use BoV to represent images more accurately is still a challenging problem. In this paper, we propose to utilize a sparse representation based method to estimate document image's quality with respect to the OCR capability. Unlike the conventional sparse representation approaches, we introduce the target quality scores into the training phase of sparse representation. The proposed method improves the discriminability of the system and ensures the obtained codebook is more suitable for our assessment task. The experimental results on a public dataset show that the proposed method outperforms other hand-crafted and BoV based DIQA approaches. Xujun Peng, Huaigu Cao, Premkumar Natarajan |
DAS | 1 |
| 2016 | Image color harmony modeling through neighbored co-occurrence colors
Peng Lu 0007, Xujun Peng, Caixia Yuan, Ruifan Li, Xiaojie Wang 0006 |
Neurocomputing | 2 |
| 2016 | An EL-LDA based general color harmony model for photo aesthetics assessment
Peng Lu 0007, Xujun Peng, Xinshan Zhu, Ruifan Li |
Signal Process. | 2 |
| 2015 | Document image OCR accuracy prediction via latent Dirichlet allocationabstractOptical character recognition (OCR) accuracy of document images is an important factor for the success of many document processing and analysis tasks, especially for unconstraint captured document images. Although several document image OCR capability assessment methods are proposed, they mostly model the problem based on the empirically defined rules of image degradation, which cause the existing approaches infeasible for predicting the OCR scores. In this paper, a computational model is presented to automatically predict document image quality towards facilitating the OCR accuracy without references. Unlike conventional methods that use heuristically designed features, in our work the raw features are learned from training images and a generative quality model is built based on latent Dirichlet allocation, which is used to assess the document's OCR capability. We present evaluation results on a public dataset which have been captured using digital cameras with different level of blur degradation. The experimental results show that the proposed method outperforms traditional document image quality assessment approaches. Xujun Peng, Huaigu Cao, Premkumar Natarajan |
ICDAR | 1 |
| 2015 | Multiple parameter control for ant colony optimization applied to feature selection problem
Gang Wang 0013, HaiCheng Eric Chu, Huiling Chen 0001, Weitong Hu, Ying Li 0004, Xujun Peng |
Neural Comput. Appl. | 7 |
| 2015 | Towards aesthetics of image: A Bayesian framework for color harmony modeling
Peng Lu 0007, Xujun Peng, Ruifan Li, Xiaojie Wang 0006 |
Signal Process. Image Commun. | 2 |
| 2015 | Finding more relevance: Propagating similarity on Markov random field for object retrieval
Peng Lu 0007, Xujun Peng, Xinshan Zhu, Ruifan Li |
Signal Process. Image Commun. | 2 |
| 2014 | Discovering Harmony: A Hierarchical Colour Harmony Model for Aesthetics Assessment
Peng Lu 0007, Zhijie Kuang, Xujun Peng, Ruifan Li |
ACCV (3) | 3 |
| 2014 | Text Classification via iVector Based Feature RepresentationabstractIn this paper, we address the problem of text classification: classifying modern machine-printed text, handwritten text and historical typewritten text from degraded noisy documents. We propose a novel text classification approach based on iVector, a newly developed concept in speaker verification. To a given text line, the iVector is a fixed-length feature vector representation, transformed from a high-dimensional super vector based on means of Gaussian mixture model (GMM), where the text dependent component is separated from a universal background model (UBM) and can be represented by a low dimensional set of factors. We classify the text lines with a discriminative classifier - support vector machine (SVM) in iVector space. A baseline approach of text classification using GMM in feature space is also presented for evaluation purpose. Experimental results on an Arabic document database show accuracy of 92.04% for text line classification using the proposed method. Furthermore, the relative word error rate (WER) of 9.6% is decreased in optical character recognition (OCR) when coupled with the proposed iVector-SVM classifier. The proposed iVector-SVM approach is language independent, thus, can be applied to other scripts as well. Shengxin Zha, Xujun Peng, Huaigu Cao, Xiaodan Zhuang, Pradeep Natarajan, Premkumar Natarajan |
Document Analysis Systems | 2 |
| 2014 | Text detection and recognition in natural scenes and consumer videosabstractWe propose an end-to-end system for text detection and recognition in natural scenes and consumer videos. Maximally Stable Extremal Regions which are robust to illumination and viewpoint variations are selected as text candidates. Rich shape descriptors such as Histogram of Oriented Gradients, Gabor filter, corners and geometrical features are used to represent the candidates and classified using a support vector machine. Positively labeled candidates serve as anchor regions for word formation. We then group candidate regions based on geometric and color properties to form word boundaries. To speed up the system for practical applications, we use Partial Least Squares approach for dimensionality reduction. The detected words are binarized, filtered and passed to a hidden Markov model based Optical Character Recognition (OCR) system for recognition. We show significant improvement in text detection and recognition tasks over previous approaches on a large consumer video dataset. Furthermore, the event detection system built upon the OCR output of this approach outperformed multiple other OCR-only based submissions in the recently concluded NIST TRECVID 2013 multimedia event detection evaluations. Xujun Peng, Xiaodan Zhuang, Pradeep Natarajan, Huaigu Cao |
ICASSP | 2 |
| 2014 | Progress in the Raytheon BBN Arabic Offline Handwriting Recognition SystemabstractThis paper presents the most recent progress and state of the art result obtained from BBN's Arabic offline handwriting recognition research. Our system is based a left-to-right hidden Markov model and integrates discriminative learning methods including discriminative MPE and n-best rescoring using the scores of glyph classifiers (SVM, DNN) and the RNNLM. Arabic-related features for n-best rescoring are also investigated in this paper. Multi-stage MAP/MLLR and writer verification are applied to adapt the recognizer in all training situations. Consensus network is extensively researched for system combination and improving challenging preprocessing problems. Huaigu Cao, Premkumar Natarajan, Xujun Peng, Krishna Subramanian 0001, David Belanger 0001 |
ICFHR | 3 |
| 2013 | Exploiting Stroke Orientation for CRF Based Binarization of Historical DocumentsabstractWe present a novel binarization method that is especially effective on historical documents with the following characteristics: (a) the documents contain free-form cursive handwritten text with significant but consistent slant, (b) scanning artifacts resulting in the text and background pixels not having uniform intensity even within the same page, and (c) pages containing significant amount of bleeds from the other side of the page. In order to tackle the problem of non-uniform text and background intensity, we use a thresholding algorithm that works equally well for regions of the page containing text and regions of the page containing no text. We then combine this algorithm with a CRF-based framework which handles bleeds using a novel approach to further improve the quality of binarization. We compare the proposed binarization algorithm against other popular binarization algorithms both qualitatively using examples and quantitatively using the word error rate (WER) metric from performing optical character recognition (OCR) on binarized text using the BBN Byblos Offline Handwritten text recognition (OHR) system. Xujun Peng, Huaigu Cao, Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan |
ICDAR | 1 |
| 2013 | Handwritten text separation from annotated machine printed documents using Markov Random Fields
Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
Int. J. Document Anal. Recognit. | 1 |
| 2012 | Extracting information from handwritten content in census forms
Huaigu Cao, Krishna Subramanian 0001, Xujun Peng, Jinying Chen, Rohit Prasad, Premkumar Natarajan |
ICPR | 3 |
| 2012 | Incremental multi-linear discriminant analysis using canonical correlations for action recognition
Chengcheng Jia, Xujun Peng, Wei Pang 0001, Can-Yan Zhang, Chunguang Zhou, Zhezhou Yu |
Neurocomputing | 3 |
| 2012 | Using a boosted tree classifier for text segmentation in hand-annotated documents
Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
Pattern Recognit. Lett. | 1 |
| 2012 | Sparse Tensor Discriminant Color Space for Face VerificationabstractAs one of the fundamental features, color provides useful information and plays an important role for face recognition. Generally, the choice of a color space is different for different visual tasks. How can a color space be sought for the specific face recognition problem? To address this problem, we propose a sparse tensor discriminant color space (STDCS) model that represents a color image as a third-order tensor in this paper. The model cannot only keep the underlying spatial structure of color images but also enhance robustness and give intuitionistic or semantic interpretation. STDCS transforms the eigenvalue problem to a series of regression problems. Then one spare color space transformation matrix and two sparse discriminant projection matrices are obtained by applying lasso or elastic net on the regression problems. The experiments on three color face databases, AR, Georgia Tech, and Labeled Faces in the Wild face databases, show that both the performance and the robustness of the proposed method outperform those of the state-of-the-art TDCS model. Jian Yang 0003, Mingfang Sun, Xujun Peng, Mingming Sun 0006, Chunguang Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2011 | Text Extraction from Video Using Conditional Random FieldsabstractIn this paper, we describe an approach to extract text from broadcast videos. Candidate blocks are detected based on edge extraction results. Corners and geometrical features are used for the purpose of initial classification which is carried out by using a support vector machine (SVM). Considering the spatial inter-dependencies of different regions in the image, we propose a novel conditional random field (CRF) based framework which integrates the outputs of SVM into the system to improve the accuracy of labeling for blocks. The experimental results show that the proposed system achieves reliable performance for text detection/extraction from videos. Xujun Peng, Huaigu Cao, Rohit Prasad, Premkumar Natarajan |
ICDAR | 1 |
| 2011 | Automated image quality assessment for camera-captured OCRabstractCamera-captured optical character recognition (OCR) is a challenging area because of artifacts introduced during image acquisition with consumer-domain hand-held and Smart phone cameras. Critical information is lost if the user does not get immediate feedback on whether the acquired image meets the quality requirements for OCR. To avoid such information loss, we propose a novel automated image quality assessment method that predicts the degree of degradation on OCR. Unlike other image quality assessment algorithms which only deal with blurring, the proposed method quantifies image quality degradation across several artifacts and accurately predicts the impact on OCR error rate. We present evaluation results on a set of machine-printed document images which have been captured using digital cameras with different degradations. Xujun Peng, Huaigu Cao, Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan |
ICIP | 1 |
| 2011 | Exponential locality preserving projections for small sample size problem
Huiling Chen 0001, Xujun Peng, Chunguang Zhou |
Neurocomputing | 3 |
| 2011 | A novel face recognition method based on sub-pattern and tensor
Chunguang Zhou, Xujun Peng, Huiling Chen 0001, Gang Wang 0013 |
Neurocomputing | 4 |
| 2011 | Face recognition using second-order discriminant tensor subspace analysis
Chunguang Zhou, Xujun Peng |
Neurocomputing | 4 |
| 2010 | Overlapped text segmentation using Markov random field and aggregationabstractSeparating machine printed text and handwriting from overlapping text is a challenging problem in the document analysis field and no reliable algorithms have been developed thus far. In this paper, we propose a novel approach for separating handwriting from binary image of overlapped text. Instead of using fixed size training patches, we describe an aggregation method which uses shape context features to extract training samples automatically. We use a Markov Random Field (MRF) to model the overlapped text. The neighbor system is inherited from a coarsening procedure and the prior and likelihood of the MRF is learned based on a distance metric. Experimental results show that the proposed method can achieve 87.97% recall for handwriting and 91.44% recall for machine printed text. Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
Document Analysis Systems | 1 |
| 2010 | Text Separation from Mixed Documents Using a Tree-Structured ClassifierabstractIn this paper, we propose a tree-structured multi-class classifier to identify annotations and overlapping text from machine printed documents. Each node of the tree-structured classifier is a binary weak learner. Unlike normal decision tree(DT) which only considers a subset of training data at each node and is susceptible to over-fitting, we boost the tree using all training data at each node with different weights. The evaluation of the proposed method is presented on a set of machine printed documents which have been annotated by multiple writers in an office/collaborative environment. Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
ICPR | 1 |
| 2009 | Markov Random Field Based Text Identification from Annotated Machine Printed DocumentsabstractIn this paper, we describe an approach to segment handwritten text, machine printed text and noise from annotated machine printed documents. Three categories of word level features are extracted. We use a modified K-Means clustering algorithm for classification followed by a relabeling procedure using Markov Random Field(MRF) based on a concept of neighboring patches and Belief Propagation(BP) rules. Experimental results on an imbalanced data set show that our approach achieves an overall recall of 96.33%. Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram, Kiran Bhuvanagiri |
ICDAR | 1 |