EDBT 2026 Demo / reviewers in the wild / expert
Jun Sun 0004
dblp:s/JunSun4
· DBLP profile ↗
103ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0002-0967-4859ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 16 since 2021Databases, data management, data science and information retrieval · 38 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Following the Teacher's Footsteps: Scheduled Checkpoint Distillation for Domain-Specific LLMs
Chaoliang Zhong, Jun Sun 0004, Yusuke Oishi |
ICPR (5) | 3 |
| 2025 | Channel-Independent Refiner for Multivariate Time Series ForecastingabstractReal-world time series data are usually multivariate with complex channel relations. Some channels are highly related, while others are with limited correlation. Channel independence has shown great significance in multivariate time series forecasting to capture better individual channel characteristics. In order to learn channel patterns better, we propose a channel-independent Refiner as a plug-and-play module. Specifically, we devise a channel-wise Refiner connected to the output of existing methods. Leveraging the concatenation of original input and the coarse prediction from the base model, the Refiner produces a better estimation. Our Refiner also benefits from the channel-independent design and a post-training strategy, achieving significant improvement over base models. Extensive experiments on iTransformer, FEDformer, Autoformer, FreTS, DLinear and TSMixer demonstrate that our Refiner reduces forecasting errors on both Transformer-based and MLP-based models in over 90% of the experimental settings. Our Refiner combined with iTransformer establishes the new state-of-the-art. Jie Wang 0111, Zhongguang Zheng, Chaoliang Zhong, Jun Sun 0004 |
CIKM | 4 |
| 2025 | Proxy-Mamba: Training-Free Architecture Search for Mamba via Gradient-Weight Correlation
Huigang Zhang, Liuan Wang, Jun Sun 0004 |
PRICAI | 3 |
| 2024 | Conditional Past Experience Generation for Dark Continual LearningabstractContinual learning (CL) aims to learn a sequence of tasks without forgetting. Numerous efforts have been made to tackle CL including data-centric, model-centric, and algorithm-centric methods. The more information the algorithm can obtain from previous tasks, e.g., the training data, the easier the CL task will be. However, few studies focus on the most difficult setting, i.e., dark CL (DCL) where only the model of the last task can be obtained. DCL is a typical setting in real-world applications, e.g., Cl tasks based on the models trained on private or privileged data. For solving DCL, we propose a novel recursive generalization bound, which can also be applied to arbitrary Traditional CL (TCL). To minimize the bound proposed, we propose a novel method, i.e., conditional past experience generation (CPEG), which reconstructs the previous conditional training data in the DCL setting. In the experiment, we apply CPEG to a wide range of benchmarks. The experimental results show that CPEG significantly reduces forgetting. On the other hand, CPEG can be used as a regularization term for any CL baseline. We also conduct experiments on the TCL setting. The performance of almost all baselines is improved, especially for the most difficult class-incremental tasks. Chaoliang Zhong, Jie Wang 0111, Jun Sun 0004, Yasuto Yokota |
ICIP | 4 |
| 2024 | NR-CION: Non-rigid Consistent Image Composition Via Diffusion Model
Wei Liu 0169, Liuan Wang, Jun Sun 0004 |
ICPR (25) | 3 |
| 2024 | Temporal Insight Enhancement: Mitigating Temporal Hallucination in Video Understanding by Multimodal Large Language Models
Li Sun 0007, Liuan Wang, Jun Sun 0004, Takayuki Okatani |
ICPR (7) | 3 |
| 2024 | Towards Generated Image Provenance Analysis via Conceptual-Similar-Guided-SLIP RetrievalabstractWith the prevalence of state-of-the-art generative models, photorealistic synthetic images can now be easily generated. However, the generated images may replicate contents from the original training images, which can lead to potential legal issues. In this paper, we propose a novel method calledConceptual-Similar-guided Self-supervised Language-Image Pre-training(CS-SLIP) that leverages both image and text modalities for the generated image provenance. Besides the self-supervised learning branch and contrastive learning branch, a conceptual-similar branch is designed to guide the model to learn a better feature representation of image-text-pairs. We also adopt the re-ranking method to refine the initial matching candidates via the cross-modal bi-directional retrieval. Extensive qualitative and quantitative experiments are conducted, which demonstrate that the replication indeed exists in the generated images, and our proposed method can effectively retrieve the most similar images from the training corpus to achieve the goal of generated image provenance analysis. Xiaojie Xia, Liuan Wang, Jun Sun 0004, Akira Nakagawa |
IEEE Signal Process. Lett. | 3 |
| 2023 | Semantic-Embedded Knowledge Acquisition and Reasoning for Image SegmentationabstractImage segmentation is a difficult and challenging task because of the complex object appearance and diverse object categories. Traditional methods directly use visual features for segmentation but ignore the correlation between objects. We introduce a knowledge reasoning module (KRM) for external knowledge aggregation and leverage a graphic neural network to aggregate the knowledge feature, which is concatenated with a visual feature for semantic segmentation. To this end, we use word embedding of category names as semantic feature and establish the relationship between categories. Through iteration, the aggregated features can be enriched. In experiments, three well known semantic segmentation methods are used as baseline. Our experiment results outperform the baseline methods on the food dataset Food-Seg103 and Cityscapes, and demonstrate the effectiveness of our proposed method. Wei Liu 0169, Huigang Zhang, Xiaojie Xia, Liuan Wang, Jun Sun 0004 |
ICIP | 5 |
| 2023 | Prompt Prototype Learning Based on Ranking Instruction For Few-Shot Visual TasksabstractQuerying large language models (LLMs), such as GPT-3, for high-quality prompts and utilizing pre-trained vision-language models, such as CLIP, to construct a zero-shot visual classification model, offer promising performance across various downstream visual tasks. However, when applied to specific domains, their efficacy is restricted due to the gap between the general prompts they generate and the required domain-specific knowledge. In this paper, we propose a novel, lightweight method for prompt prototype learning through ranking instruction, specifically designed to bridge this gap in the context of few-shot visual classification. We generate domain-specific prompts leveraging the knowledge contained in LLMs and then fine-tune the prompt prototype with effective ranking instructions from several domain images. Our few-shot experiments on facial expression benchmarks demonstrate the efficacy of the prompt prototype. Notably, our method delivers results that are on par with state-of-the-art few-shot image classification techniques and can be integrated with them to further improve performance in the facial expression domain. Our approach provides a promising solution to few-shot visual classification, making use of the knowledge contained in LLMs to generate domain-specific prompts. Li Sun 0007, Liuan Wang, Jun Sun 0004, Takayuki Okatani |
ICIP | 3 |
| 2023 | Zero-shot temporal event localisation: Label-free, training-free, domain-freeabstractAbstract Temporal event localisation (TEL) has recently attracted increasing attention due to the rapid development of video platforms. Existing methods are based on either fully/weakly supervised or unsupervised learning, and thus they rely on expensive data annotation and time‐consuming training. Moreover, these models, which are trained on specific domain data, limit the model generalisation to data distribution shifts. To cope with these difficulties, the authors propose a zero‐shot TEL method that can operate without training data or annotations. Leveraging large‐scale vision and language pre‐trained models, for example, CLIP, we solve the two key problems: (1) how to find the relevant region where the event is likely to occur; (2) how to determine event duration after we find the relevant region. Query guided optimisation for local frame relevance relying on the query‐to‐frame relationship is proposed to find the most relevant frame region where the event is most likely to occur. Proposal generation method relying on the frame‐to‐frame relationship is proposed to determine the event duration. The authors also propose a greedy event sampling strategy to predict multiple durations with high reliability for the given event. The authors’ methodology is unique, offering a label‐free, training‐free, and domain‐free approach. It enables the application of TEL purely at the testing stage. The practical results show it achieves competitive performance on the standard Charades‐STA and ActivityCaptions datasets. Li Sun 0007, Ping Wang 0034, Liuan Wang, Jun Sun 0004, Takayuki Okatani |
IET Comput. Vis. | 4 |
| 2023 | Exploiting spatio-temporal knowledge for video action recognitionabstractAbstract Action recognition has been a popular area of computer vision research in recent years. The goal of this task is to recognise human actions in video frames. Most existing methods often depend on the visual features and their relationships inside the videos. The extracted features only represent the visual information of the current video itself and cannot represent the general knowledge of particular actions beyond the video. Thus, there are some deviations in these features, and the recognition performance still requires improvement. In this sudy, we present a novel spatio‐temporal knowledge module (STKM) to endow the current methods with commonsense knowledge. To this end, we first collect hybrid external knowledge from universal fields, which contains both visual and semantic information. Then graph convolution networks (GCN) are used to represent and aggregate this knowledge. The GCNs involve (i) a spatial graph to capture spatial relations and (ii) a temporal graph to capture serial occurrence relations among actions. By integrating knowledge and visual features, we can get better recognition results. Experiments on AVA, UCF101‐24 and JHMDB datasets show the robustness and generalisation ability of STKM. The results report a new state‐of‐the‐art 32.0 mAP on AVA v2.1. On UCF101‐24 and JHMDB datasets, our method also improves by 1.5 AP and 2.6 AP, respectively, over the baseline method. Huigang Zhang, Liuan Wang, Jun Sun 0004 |
IET Comput. Vis. | 3 |
| 2022 | FBMOT: Flow Bridges the Gap between Detection and Tracking in Multiple Object TrackingabstractThe detection performance is a key bottleneck to the performance of multiple object tracking (MOT) algorithms and the advanced detection algorithms contribute a large portion to the success of MOT algorithms. Even though, the detectors can still make false detections based on only the image frames, which may directly result in the loss or mismatch of tracks. In MOT tasks, we point out that the video frames are temporally correlated and the temporal relationship can be leveraged to further improve the performance of detection and subsequent tracking. In our work, we propose FBMOT which uses optical flow to compute a prior heatmap about the locations of previously tracked objects in the current frame and takes the heatmap as an additional input in detection. We further proposed a novel regularization loss to help our model distinguish useful information in the prior heatmap. As a result, our method improves 0.8, 1.0 and 2.7 MOTA in MOT16, MOT17 and MOT20 test datasets respectively. Lisheng Wu, Liuan Wang, Jun Sun 0004 |
AVSS | 3 |
| 2022 | Discriminative Mutual Learning for Multi-target Domain AdaptationabstractUnsupervised domain adaptation (UDA) has attracted much attention among those seeking to transfer a model from a labeled source domain to an unlabeled target domain. Many effective algorithms for single target domain adaptation (STDA) have been designed, however, STDA cannot satisfy the scenarios of transferring simultaneously to multiple target domains or transferring to a blending target domain. This paper proposes a novel discriminative mutual learning method for multi-target domain adaptation covering both blending target domain adaptation (BTDA) and multiple target domain adaptation (MTDA). Two key points are considered in the proposed method: one is to learn discriminative features for better prediction, and the other is to self-train the model with pseudo-labeled target data based on distance information. These two aspects are integrated through a mutual learning strategy via two different classifiers. According to extensive experiments on three domain adaptation benchmarks, the proposed method demonstrates the state-of-the-art performance in both BTDA and MTDA settings. Jie Wang 0111, Chaoliang Zhong, Ying Zhang 0124, Jun Sun 0004, Yasuto Yokota |
ICPR | 5 |
| 2022 | Learning Unforgotten Domain-Invariant Representations for Online Unsupervised Domain AdaptationabstractExisting unsupervised domain adaptation (UDA) studies focus on transferring knowledge in an offline manner. However, many tasks involve online requirements, especially in real-time systems. In this paper, we discuss Online UDA (OUDA) which assumes that the target samples are arriving sequentially as a small batch. OUDA tasks are challenging for prior UDA methods since online training suffers from catastrophic forgetting which leads to poor generalization. Intuitively, a good memory is a crucial factor in the success of OUDA. We formalize this intuition theoretically with a generalization bound where the OUDA target error can be bounded by the source error, the domain discrepancy distance, and a novel metric on forgetting in continuous online learning. Our theory illustrates the tradeoffs inherent in learning and remembering representations for OUDA. To minimize the proposed forgetting metric, we propose a novel source feature distillation (SFD) method which utilizes the source-only model as a teacher to guide the online training. In the experiment, we modify three UDA algorithms, i.e., DANN, CDAN, and MCC, and evaluate their performance on OUDA tasks with real-world datasets. By applying SFD, the performance of all baselines is significantly improved. Chaoliang Zhong, Jie Wang 0111, Ying Zhang 0124, Jun Sun 0004, Yasuto Yokota |
IJCAI | 5 |
| 2022 | PICA: Point-wise Instance and Centroid Alignment Based Few-shot Domain Adaptive Object Detection with Loose AnnotationsabstractIn this work, we focus on supervised domain adaptation for object detection in few-shot loose annotation setting, where the source images are sufficient and fully labeled but the target images are few-shot and loosely annotated. As annotated objects exist in the target domain, instance level alignment can be utilized to improve the performance. Traditional methods conduct the instance level alignment by semantically aligning the distributions of paired object features with domain adversarial training. Although it is demonstrated that point-wise surrogates of distribution alignment provide a more effective solution in few-shot classification tasks across domains, this point-wise alignment approach has not yet been extended to object detection. In this work, we propose a method that extends the point-wise alignment from classification to object detection. Moreover, in the few-shot loose annotation setting, the background ROIs of target domain suffer from severe label noise problem, which may make the point-wise alignment fail. To this end, we exploit moving average centroids to mitigate the label noise problem of background ROIs. Meanwhile, we exploit point-wise alignment over instances and centroids to tackle the problem of scarcity of labeled target instances. Hence this method is not only robust against label noises of background ROIs but also robust against the scarcity of labeled target objects. Experimental results show that the proposed instance level alignment method brings significant improvement compared with the baseline and is superior to state-of-the-art methods. Chaoliang Zhong, Jie Wang 0111, Ying Zhang 0124, Jun Sun 0004, Yasuto Yokota |
WACV | 5 |
| 2021 | CANN: Coupled Approximation Neural Network for Partial Domain AdaptationabstractUnsupervised domain adaptation (UDA) methods aim to transfer knowledge from a labeled source domain to an unlabeled target domain. Most existing UDA methods try to learn domain-invariant features so that the classifier trained by the source labels can automatically be adapted to the target domain. However, recent works have shown the limitations of these methods when label distributions differ between the source and target domains. Especially, in partial domain adaptation (PDA) where the source domain holds plenty of individual labels (private labels) not appeared in the target domain, the domain-invariant features can cause catastrophic performance degradation. In this paper, based on the originally favorable underlying structures of the two domains, we learn two kinds of target features, i.e., the source-approximate features and target-approximate features instead of the domain-invariant features. The source-approximate features utilize the consistency of the two domains to estimate the distribution of the source private labels. The target-approximate features enhance the feature discrimination in the target domain while detecting the hard (outlier) target samples. A novel Coupled Approximation Neural Network (CANN) has been proposed to co-train the source-approximate and target-approximate features by two parallel sub-networks without sharing the parameters. We apply CANN to three prevalent transfer learning benchmark datasets, Office-Home, Office-31, and Visda2017 with both UDA and PDA settings. The results show that CANN outperforms all baselines by a large margin in PDA and also performs best in UDA. Chaoliang Zhong, Jie Wang 0111, Jun Sun 0004, Yasuto Yokota |
CIKM | 4 |
| 2021 | Key-Guided Identity Document Classification Method by Graph Attention Network
Xiaojie Xia, Wei Liu 0169, Ying Zhang 0124, Liuan Wang, Jun Sun 0004 |
ICDAR (4) | 5 |
| 2021 | EBB: Progressive Optimization For Partial Domain AdaptationabstractUnsupervised domain adaptation (UDA) methods are generally proposed based on the assumption that the source domain and the target domain share an identical group of classes. However, in transfer learning tasks in reality, the target domain often has fewer data with missing classes. Partial domain adaptation (PDA) allows the source domain to have un-shared categories. Anchor points are used to describe the easily identified target samples. It is observed that the shared classes tend to have more anchor points compared with the unshared classes and we introduce a novel progressive optimization method named Ebb PDA tasks. Ebb could resist the negative transfer caused by the category gap and can be applied to any domain adaptation model. Ebb picks the anchor points by analyzing the features of the base model, and it uses the class-wise distribution of anchor points to estimate the category gap. Then Ebb minimizes the errors of shared classes and corrects the error samples caused by blind alignment. To verify the effectiveness of the method, we apply Ebb to three PDA image classification tasks based on three widely used data sets, i.e, Office-Home, Office-31 and ImageCLEF-DA while using three state-of-the-art methods as the base models. The results show that Ebb brings a significant improvement in all tasks and the models optimized by Ebb have stable performance under a wide range of category gaps. Chaoliang Zhong, Jie Wang 0111, Jun Sun 0004, Yasuto Yokota |
ICIP | 4 |
| 2021 | Dual-Consistency Self-Training For Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) is a challenging task characterized by unlabeled target data with domain discrepancy to labeled source data. Many methods have been proposed to learn domain invariant features by marginal distribution alignment, but they ignore the intrinsic structure within target domain, which may lead to insufficient or false alignment. Class-level alignment has been demonstrated to align the features of the same class between source and target domains. These methods rely extensively on the accuracy of predicted pseudo-labels for target data. Here, we develop a novel self-training method that focuses more on accurate pseudo-labels via a dual-consistency strategy involving modelling the intrinsic structure of the target domain. The proposed dual-consistency strategy first improves the accuracy of pseudo-labels through voting consistency, and then reduces the negative effects of incorrect predictions through structure consistency with the relationship of intrinsic structures across domains. Our method has achieved comparable performance to the state-of-the-arts on three standard UDA benchmarks. Jie Wang 0111, Chaoliang Zhong, Jun Sun 0004, Masaru Ide, Yasuto Yokota |
ICIP | 4 |
| 2021 | Feature Disentanglement For Cross-Domain Retina Vessel SegmentationabstractDomain shift is regarded as a key factor affecting the robust-ness of many models. Recently, unsupervised auxiliary learning (e.g., input reconstruction) has been proposed to improve the model’s domain transferability and alleviate cross-domain performance degradation; however, in the paradigm of existing approaches, the features extracted from various tasks are shared, which mixes the domain-invariant features from the main task and domain-specific feature from the auxiliary task, leading to an imperfect learning. To solve this problem, we propose a novel unsupervised domain adaptation method - the Disentangled Reconstruction Neural Network (DRNN) - for cross-domain retina vessel segmentation. DRNN leverages two tandem nets and disentangles the domain-invariant features and the domain-specific features in the multi-task learning process. We perform extensive experiments on public retina datasets and our proposed DRNN outperforms the competitors by a significant margin to achieve state-of-the-art results pertaining to retina vessel segmentation. Jie Wang 0111, Chaoliang Zhong, Jun Sun 0004, Yasuto Yokota |
ICIP | 4 |
| 2021 | Knowledge-Based Reasoning Network For Object DetectionabstractThe mainstream object detection algorithms rely on recognizing object instances individually, but do not consider the high-level relationship among objects in context. This will inevitably lead to biased detection results, due to the lack of commonsense knowledge that humans often use to assist the task for object identification. In this paper, we present a novel reasoning module to endow the current detection systems with the power of commonsense knowledge. Specifically, we use graph attention network (GAT) to represent the knowledge among objects. The knowledge covers visual and semantic relations. Through the iterative update of GAT, the object features can be enriched. Experiments on the COCO detection benchmark indicate that our knowledge-based reasoning network has achieved consistent improvements upon various CNN detectors. We achieved 1.9 and 1.8 points higher Average Precision (AP) than Faster-RCNN and Mask-RCNN respectively, when using ResNet50-FPN as backbone. Huigang Zhang, Liuan Wang, Jun Sun 0004 |
ICIP | 3 |
| 2021 | Exploring Pathologist Knowledge for Automatic Assessment of Breast Cancer Metastases in Whole-slide ImageabstractAutomatic assessment of breast cancer metastases plays an important role to help pathologist reduce the time-consuming work in histopathological whole-slide image diagnosis. From the utilization of knowledge point of view, the low-magnification level and high-magnification level are carefully checked by the pathologists for tumor pattern and cell tumor characteristic. In this paper, we propose a novel automatic patient-level tumor segmentation and classification method, which makes full use of the diagnosis knowledge clues from pathologists. For tumor segmentation, a multi-level view DeepLabV3+ (MLV-DeepLabV3+) is designed to explore the distinguishing features of cell characteristics between tumor and normal tissue. Furthermore, the expert segmentation models are selected and integrated by Pareto-front optimization to imitate the expert consultation to get perfect diagnosis. For wholeslide classification, multi-level magnifications are adaptive checked to focus on the effective features in different magnification. The experimental results demonstrate that our pathologist knowledge-based automatic assessment of whileslide image is effective and robust on the public benchmark dataset. Liuan Wang, Li Sun 0007, Huigang Zhang, Ping Wang 0034, Rong Zhou 0005, Jun Sun 0004 |
ACM Multimedia | 7 |
| 2020 | NAS-EOD: an end-to-end Neural Architecture Search method for Efficient Object DetectionabstractModel efficiency for object detection has become more and more important recently, especially when intelligent mobile devices get more and more popular. Current lightweight object detection model is either migrated from lightweight classification models, or pruned directly from complex object detection models. These pipelines can not match the performance requirements of edge devices. In this work, we propose a neural architecture search (NAS) method to build a detection model automatically that can perform well on edge devices. Specifically, the proposed method supports the search of not only multiscale feature network, but also backbone network. This enables us to search out a global optimal model. To this end, we've made a special design that the backbone and feature network can share the same search space. This method greatly reduces the search time while ensuring the search accuracy. It can find a good architecture in 14GPU days. Additionally, we add latency information into the main objective during performance estimation. Therefore, we can control between the speed and accuracy to better adapt to the edge environment. Experiments on the PASCAL VOC benchmark indicate that the searched architecture (named NAS-EOD) can get good accuracy even if training from scratch. When using a pre-training scheme, our model is superior to state-of-the-art small object detection models. Huigang Zhang, Liuan Wang, Jun Sun 0004, Li Sun 0007, Hiromichi Kobashi, Nobutaka Imamura |
ICPR | 3 |
| 2020 | Vehicle Detection with Partial Anchors in Remote Sensing ImagesabstractVehicle detection in remote sensing(RS) images has been an active topic with the development of computer vision in recent years. However, directly applying conventional horizontal anchor-based detection methods in oriented vehicle detection often acquires poor performance. Although rotated anchors have been used to tackle this problem, this design leads to heavy computational cost because of thousands of rotated anchors generated in each level feature map. In this paper, we propose to detect vehicles with partial anchors, which greatly accelerates detection process. The novel Partial Anchors based Detection Network(PADeN) filter out redundant anchors with semantic information. To boost the performance of PADeN, the centerness mask branch is added into the network. The results demonstrate that PADeN significantly outperforms previous approaches in vehicle detection and achieves the mAP of 76.9%. Fuyan Ma, Bin Sun 0001, Shutao Li 0001, Jun Sun 0004 |
IGARSS | 4 |
| 2019 | Text Line Adjustment Based on Neural NetworkabstractWith the development of artificial intelligence technology, many text generation models have emerged. People generate single words in various ways. For example, different forms of English characters and numbers can be generated through Generative Adversarial Network (GAN). Then, text lines can be synthesized from single words. However, most people only use some well-defined rules to synthesize "text line" from "single word". The result depends on people's experience. Therefore, the generated text lines often seem inconsistent, especially when generating handwriting text line. For example, it's hard to make the adjacent parts of the text look real. Moreover, adjacent words may have different styles. These problems can not be solved by simply modifying the text synthesis program. In this paper, we attempts to use neural networks to adjust the synthesized text lines. The input of our model is a synthetic text line, and then it can output a more coordinated text line. Four different loss and a discriminator ensure the quality of the generated text lines. In experiments, our neural network has successfully corrected many unnatural aspects of synthetic text lines. Considering that no one has tried to adjust the word at the text line level, our experiment offers novel and valuable insights. Xiaojie Xia, Xiaoyi Yu, Jun Sun 0004, Satoshi Naoi |
ICDAR | 5 |
| 2019 | An Efficient off-Line Handwritten Japanese Address Recognition SystemabstractHandwritten address recognition is very practical and popular in ICR (Intelligent Character Recognition) application field, especially for postal and express service. The structure of Japanese address has three components: high level, middle level and low level address. The high level and middle level address is very useful in postal applications since they indicate the destination of the express freight car. However, it's not easy to achieve high recognition accuracy on handwritten address merely by general ICR engine. In this paper, an efficient handwritten Japanese address recognition system is proposed that combines general ICR engine and special processing functions. Rough recognition result is obtained from a general ICR engine based on over-segmentation strategy. Special processing for precise address correction contains three main functions: FCN-based middle level address detection and recognition, BK-Tree-based high level address candidates searching and DTW-based accurate high level address matching. In the experiments on 670 handwritten Japanese address images, the proposed method successfully divided the total address into three levels, and achieved string-level correct rates of 81.6% and 72.5% in high level and middle level address respectively. Xiaojie Xia, Xiaoyi Yu, Wei Liu 0169, Jun Sun 0004, Satoshi Naoi |
ICDAR | 5 |
| 2019 | Automatic Neural Network Search Method for Open Set RecognitionabstractReal-world recognition or classification tasks in computer vision are not apparent in controlled environments and often get involved in open set. Previous research work on real-world recognition problem is knowledge- and labor-intensive to pursue good performance for there are numbers of task domains. Auto Machine Learning (AutoML) approaches supply an easier way to apply advanced machine learning technologies, reduce the demand for experienced human experts and improve classification performance on close set. This paper proposes an automatic neural network search method for designing effective convolution neural network (CNN) models for open set recognition (OSR). Feature distribution information is explicitly incorporated into the main objective. So during the search process, the sampled models will enlarge interclass differences and reduce intra-class variations. We design a flexible search space based on classic CNN models to diversify neural architectures and also add some search principles to limit the size of the search space. Experimental results on CIFAR-10 and Dunhuang historical Chinese datasets show that our approach improves performances on both close and open set. Comparing with the other two OSR algorithms, our method also achieves the best performance. Li Sun 0007, Xiaoyi Yu, Liuan Wang, Jun Sun 0004, Hiroya Inakoshi, Ken Kobayashi, Hiromichi Kobashi |
ICIP | 4 |
| 2019 | Robustness Evaluation of Deep Learning Models Based on Local Prediction ConsistencyabstractIt is important to estimate the performance gap of a given deep learning model on the target data set, since discrepancy or bias between source and target domains is a common and fundamental problem in the practice of machine learning techniques. Without any assumptions on data bias, such as label shift or covariate shift and without target data labels, we propose a robustness estimation method based on prediction consistency evaluation between source and target data in the neighborhood of the source samples. Considering outliers and whether the user provided model is fully trained, a variety of variant methods are also tried, including setting neighborhood threshold to average intra-class distance for each category and relative robustness. Furthermore, the time complexity of this method is O(nlogn), which is applicable for large datasets. Experiments on the handwritten digit recognition and Japanese handwriting recognition show that the proposed methods are effective. Ziqiang Shi, Chaoliang Zhong, Yasuto Yokota, Wensheng Xia, Jun Sun 0004 |
ICMLA | 5 |
| 2019 | Graph-matching-based character recognition for Chinese seal images
Bin Sun 0001, Shaojun Hua, Shutao Li 0001, Jun Sun 0004 |
Sci. China Inf. Sci. | 4 |
| 2018 | Text Line Extraction Based on Integrated K-Shortest Paths OptimizationabstractText in images can be utilized in many image understanding applications due to the exact semantic information. In this paper, we propose a novel integrated k-shortest paths optimization based text line extraction method. Firstly, the candidate text components are extracted by the Maximal Stable Extremal Region (MSER) algorithm on gray, red, green and blue channels. Secondly, one integrated directed graph on red, green, and blue channels are constructed upon the candidate text components, which can effectively incorporate different channels into one framework. Then, the integrated directed graph is transformed guided by the extracted text lines in gray channel to reduced the computational complexity. Finally, we use the k-shortest paths optimization algorithm to extract the text lines by taking advantage of the particular structure of the integrated directed graph. Experimental results demonstrate the effectiveness of the proposed method in comparison with state-of-the-art methods. Liuan Wang, Jun Sun 0004, Seiichi Uchida |
DAS | 2 |
| 2018 | Compact Binary Feature for Open Set RecognitionabstractMost pattern recognition systems are closed set recognition systems in which any input sample is to be classified as belonging to one of the given classes. This paper, however, addresses the open set recognition problem in which a test sample may either come from one of the labeled "known classes" or come from an unknown class. The number of unknown classes could potentially be unlimited. A compact binary feature (CBF) generated by an ensemble binary classifier (EBC) is proposed to solve the open set recognition problem. This method can be regarded as a type of ECOC (Error-Correcting Output Codes) combined with modern CNN (Convolutional Neural Network) techniques and adapted for open set recognition. By randomly partitioning the known classes into two groups and training a binary classifier with CNN to separate them apart, and by repeating such a procedure for many times, rich information is extracted from the training set in the form of an EBC which can associate any test sample with a CBF that can be matched according to Hamming distance which is very efficient to compute. According to the experiments on the Dunhuang ancient Chinese character dataset, EBC can boost the recognition performance significantly compared with a single feedforward CNN. Apart from that, CBF is very efficient for storage and saves lots of time in feature matching at the cost of more computation in the training phase. Jun Sun 0004, Xiaoyi Yu, Liuan Wang |
DAS | 2 |
| 2018 | Harvesting Organization Linked Data from the Web
Zhongguang Zheng, Yingju Xia, Jun Sun 0004 |
KEOD | 5 |
| 2018 | Accelerating and Compressing LSTM Based Model for Online Handwritten Chinese Character RecognitionabstractWith the development of deep learning tools, the online handwritten Chinese character recognition (HCCR) performance has been greatly improved by using deep neural networks (DNNs) especially for long short-term memory (LSTM). However, DNNs suffer from large consumption of computation and storage resources, which may cause problems for service providers, such as server pressure, longer service latency and higher energy consumption. To solve these problems, we propose a framework that combines singular value decomposition (SVD) and adaptive drop weight (ADW) to accelerate and compress LSTM based models. We first build an LSTM based model that achieves an accuracy of 97.83% on the ICDAR2013 online HCCR competition dataset. After restructuring the model with SVD and ADW, it can reduce the FLOPs (floating point operations per second) of the forward process by approximately 10 times and compress the model with 1/30 of the original size with only a 0.5% decrease in accuracy. Finally, integrated with our efficient forward implementation, the recognition of an online character requires only 2.7 ms in average on a CPU with a single thread, while requiring only 0.45 MB for model storage. Yafeng Yang, Kaihuan Liang, Xuefeng Xiao 0001, Zecheng Xie, Jun Sun 0004, Weiying Zhou |
ICFHR | 6 |
| 2018 | Deep Transfer Mapping for Unsupervised Writer AdaptationabstractConvolutional neural network (CNN) has achieved great success in handwriting recognition. However, it relies on large set of labeled data in training and its performance will deteriorate when the data distribution varies. To solve this problem, traditional methods usually consider adaptation of the single top layer of CNN. To better reduce the distribution discrepancy, in this paper, we consider adaptation of all layers of CNN including both convolutional and full layers. Four variations of transformations are designed based on different assumptions about the space relations for adaptation of convolutional layers. In order to make adaptation of multiple layers, we propose to cascade the transformations of different layers to conduct adaptation in a deep manner, and therefore this method is denoted as deep transfer mapping (DTM). DTM can capture the information from different layers and minimize the data divergence under different information abstract levels, thus it is more powerful and flexible for domain adaptation. Experiments on the online Chinese handwriting dataset (OLHWDB) demonstrate the efficiency and effectiveness of the proposed method for unsupervised writer adaptation. Hong-Ming Yang, Xu-Yao Zhang, Jun Sun 0004, Cheng-Lin Liu 0001 |
ICFHR | 4 |
| 2017 | A Convolutional Neural Network Based Two-Stage Document DeblurringabstractBlurring often happens when capturing documents with hand held cameras, which has negative effects on the Optical Character Recognition systems. In this paper, we propose a Convolutional Neural Network (CNN) based two-stage deblurring method. The method can deal with both real motion blur and focal blur situations, while it does not require exact estimation of the blur kernel. To achieve this, the whole blur kernel space is divided into several degradative sub-spaces. Firstly, a CNN classifier is trained to predict which sub-space the blurry image belongs to at the patch level. Then, several patches voting for the specific blur kernel sub-space is developed. Given the strong learning ability of CNN, only one CNN model corresponding to a degradative kernel sub-space is trained to restore the sharp images in the image restoration step. Experimental results show that the proposed approach performs well on the real blurring document images. In addition, we demonstrate that the proposed method could also handle the spatially-varying blurring. Jile Jiao, Jun Sun 0004, Satoshi Naoi |
ICDAR | 2 |
| 2017 | Identifying Machine-Printed and Handwritten Texts Using DropRegion and Deep Convolutional NetworkabstractIn this paper, we propose a deep convolutional neural network to identify machine-printed and handwritten texts. We also propose a novel data augmentation technique called DropRegion to make up for the lack of available data and enhance the generalization of the model. DropRegion increases data diversity by randomly dropping one of the stroke-containing regions in each raw input text-line image. Two parameters are introduced to make DropRegion adjustable for different data. For distinguishing texts of mixture of five languages including English, Chinese, Japanese, Korean and Russian, we have successfully achieved a very promising accuracy of 99.07% after DropRegion is applied, which is a significantly better performance compared to traditional method (97.91%) and our deep convolutional network baseline (98.75%). Zhaoyang Yang, Ziyong Feng, Jun Sun 0004, Weiying Zhou |
ICDAR | 4 |
| 2017 | Webpage cross-browser test from image levelabstractIncompatibility of webpages under different browsers and platforms is a typical technical obstruction for webpage design. To address this issue, a key challenge is to automatically detect the incompatible components and quantitatively assess the distortion extent in cross-browser tests. This paper presents a new algorithm for image pair comparison from webpages, called iterative perceptual hash (IPH), as well as a new distortion evaluation index called structure-color-saliency (SCS). The IPH that operates in an iterative manner is proposed to detect content changes considering both global structure and local content difference. The SCS assesses the distortion extent in both dimensions of image structure and color and is capable of imitating the nonlinear human perception. Experiment results demonstrate the effectiveness of IPH (e.g., F1-score 96%) and the high consistency of SCS with subjective results. P. Lu, Wei Fan 0005, Jun Sun 0004, Satoshi Naoi |
ICME | 3 |
| 2017 | An automated CNN recommendation system for image classification tasksabstractNowadays the CNN is widely used in practical applications for image classification task. However the design of the CNN model is very professional work and which is very difficult for ordinary users. Besides, even for experts of CNN, to select an optimal model for specific task may still need a lot of time (to train many different models). In order to solve this problem, we proposed an automated CNN recommendation system for image classification task. Our system is able to evaluate the complexity of the classification task and the classification ability of the CNN model precisely. By using the evaluation results, the system can recommend the optimal CNN model and which can match the task perfectly. The recommendation process of the system is very fast since we don't need any model training. The experiment results proved that the evaluation methods are very accurate and reliable. Song Wang 0007, Li Sun 0007, Wei Fan 0005, Jun Sun 0004, Satoshi Naoi, Koichi Shirahata, Takuya Fukagai, Yasumoto Tomita, Atsushi Ike, Tetsutaro Hashimoto |
ICME | 4 |
| 2017 | A Width-Variable Window Attention Model for Environmental Sensors
Cuiqin Hou, Yingju Xia, Jun Sun 0004, Ryozo Takasu, Masao Kondo |
ICONIP (2) | 3 |
| 2017 | Building fast and compact convolutional neural networks for offline handwritten Chinese character recognition
Xuefeng Xiao 0001, Yafeng Yang, Jun Sun 0004, Tianhai Chang |
Pattern Recognit. | 5 |
| 2017 | Robust shared feature learning for script and handwritten/machine-printed identification
Ziyong Feng, Zhaoyang Yang, Shuangping Huang, Jun Sun 0004 |
Pattern Recognit. Lett. | 5 |
| 2016 | Globally Optimal Text Line Extraction Based on K-Shortest Paths AlgorithmabstractThe task of text line extraction in images is a crucial prerequisite for content-based image understanding applications. In this paper, we propose a novel text line extraction method based on k-shortest paths global optimization in images. Firstly, the candidate connected components are extracted by reformulating it as Maximal Stable Extremal Region (MSER) results in images. Then, the directed graph is built upon the connected component nodes with edges comprising of unary and pairwise cost function. Finally, the text line extraction problem is solved using the k-shortest paths optimization algorithm by taking advantage of the particular structure of the directed graph. Experimental results on public dataset demonstrate the effectiveness of proposed method in comparison with state-of-the-art methods. Liuan Wang, Seiichi Uchida, Wei Fan 0005, Jun Sun 0004 |
DAS | 4 |
| 2016 | Cascading Training for Relaxation CNN on Handwritten Character RecognitionabstractWith the development of deep learning, many difficult recognition problems can be solved by deep learning models. For handwritten character recognition, the CNN is used the most. In order to improve the performance of CNN, many new models have been proposed and in which the relaxation CNN [35] is widely used. The relaxation CNN has more complicated structure than CNN while the recognition time is the same with which. However, the training of relaxation CNN needs much more time than CNN. In this paper, we propose the cascading training for relaxation CNN. Our method can train a relaxation CNN of better performance while using almost the same training time with normal CNN. The experimental results proved that the relaxation CNN trained by cascading training is able to achieve the state-of-the-art performance on handwritten Chinese character recognition. Li Chen 0017, Song Wang 0007, Wei Fan 0005, Jun Sun 0004, Satoshi Naoi |
ICFHR | 4 |
| 2016 | Deep Knowledge Training and Heterogeneous CNN for Handwritten Chinese Text RecognitionabstractIt is well known that the handwritten Chinese text recognition is a difficult problem since there are a large number of classes. In order to solve this problem, we proposed a whole new framework for unconstrained handwritten Chinese text recognition. The core module of the framework is the heterogeneous CNN trained by deep knowledge. The experimental results showed that our proposed method could achieve much better performance than the state-of-the-art methods (96.28% vs. 91.39% of CR on CASIA test set). Moreover, since the proposed framework is general, it can also be applied to other time sequence problems, such as speech recognition and video analysis. Song Wang 0007, Li Chen 0017, Wei Fan 0005, Jun Sun 0004, Satoshi Naoi |
ICFHR | 5 |
| 2016 | An HMM-based Over-Segmentation Method for Touching Chinese Handwriting RecognitionabstractThe segmentation of touching characters is still a challenging problem in offline Chinese handwriting recognition. One feasible solution is through the over-segmentation strategy which maintains a high recall of correct cuts between adjacent characters and a moderate level of redundant cuts within a single character. Previous redundant cut filtering methods rely on either pure heuristics or learned geometric properties of correct cuts. In this work, we extend learning based cut filtering method from single cut level to cut sequence level by Hidden Markov Model (HMM). As a stochastic sequential modeling tool, HMM can utilize not only properties of individual cuts but also the left-to-right temporal context and spatial dependencies among a sequence of neighboring cuts. The experimental results on a large touching character dataset show that the proposed method is effective for over-segmentation and gives better performance than previous methods. Wei Fan 0005, Jun Sun 0004, Satoshi Naoi |
ICFHR | 3 |
| 2016 | A knowledge-based table recognition method for Chinese bank statement imagesabstractAutomatic processing of large volume scanned Chinese bank statements is a urgent demand recently. Conventional methods can not well handle the following challenges of this problem: various layout styles, noises, and especially requirement of fast speed for large Chinese character set. This paper proposes a knowledge based table recognition method to meet fast speed requirement with good accuracy. Two kinds of knowledge are utilized to accelerate the identification of digit columns and the cell recognition: i) geometric knowledge about column alignment and quasi equal digit width, and ii) semantic knowledge about prior format based on the results from an optical character recognition (OCR) engine of digits. Experimental results on a real dataset show the effectiveness of our method. Wei Fan 0005, Jun Sun 0004, Satoshi Naoi |
ICIP | 3 |
| 2016 | Semi-supervised learning competence of classifiers based on graph for dynamic classifier selectionabstractClassifier competence is critical important for dynamic classifier selection. This study proposes a semi-supervised learning algorithm to learn the competence of classifiers under the proposed optimization framework based on graph. First it constructs a graph based on the training data and some unlabeled data. Then it iteratively learns the competence of classifiers. The learned competence not just reflects the competitiveness of classifiers, but also varies smooth on the neighboring data. Experimental results on five different datasets show the dynamic classifier selection classification systems with the learned classifier competence perform better than the classification systems with local accuracy as the classifier competence. Cuiqin Hou, Yingju Xia, Jun Sun 0004 |
ICPR | 4 |
| 2016 | Character region segmentation based on Stroke Stable RegionsabstractRegion segmentation is the key procedure in various text related image processing tasks. A good region extractor, which separates text area from complex background clutter, will reduce the burden of subsequent text grouping and post-processing functions. This paper propose a character region segmentation method based on a new concept named Stroke Stable Region (SSR) to achieve a better precision than many off-the-shelf region extractors such as MSER in the text segmentation task. Our work presented in this paper is inspired by the structure of MSER. However, instead of evaluating the area variation of each connected component, we proposed a novel parameter, stroke time, to measure the possibility that a pixel belongs to a character or a stroke by analyzing its character affinity. The experiments show that SSR tends to extract the visual objects with prominent text characteristics and is capable of suppressing various background noise. In a text extraction task on the ICDAR 2003 dataset, the SSR based method reduces the extracted noise components to about 1/3 of those obtained by the MSER based method, maintaining the same level of recall rate. The proposed algorithm was successfully applied to a wide range of text segmentation tasks. Hong Shang, Liuan Wang, Tanaka Hiroshi, Wei Fan 0005, Jun Sun 0004, Satoshi Naoi |
ICPR | 5 |
| 2016 | Automatic Identifying Entity Type in Linked Data
Qingliang Miao, Ruiyu Fang, Shuangyong Song, Zhongguang Zheng, Jun Sun 0004 |
PACLIC | 7 |
| 2016 | Blind Bleed-Through Removal for Scanned Historical Document Image With Conditional Random FieldsabstractScanned images of historical documents often suffer from bleed-through, which refers to the ink on one side seeping through the paper and appearing on the other side. In this paper, a new conditional random field (CRF)-based method is proposed to remove the bleed-through from the scanned images of historical images. The proposed method only requires the scanned image of one side, referred as a blind method. In general, the scanned historical document image is composed of three components: foreground, bleed-through, and background. By assuming Gaussian distributions of the three components, the proposed method establishes conditional probability distribution (CPD) models of the three components first. The parameters of the component CPD models are estimated based on an initial segmentation of the input image. Then, CRFs are used to capture the relations between observed pixels in the scanned image and the corresponding labels as well as the spatial relation between the adjacent labels. The belief propagation algorithm is used to calculate the probabilities of different labels for each pixel. Once the labeling is completed by choosing the most possible label for each pixel, the bleed-through component is removed from the input historical image by a random-filling inpainting algorithm. Experimental results on the real data set show that the proposed method preserves the foreground component very well and removes the bleed-through effectively. Bin Sun 0001, Shutao Li 0001, Xiao-Ping Zhang 0002, Jun Sun 0004 |
IEEE Trans. Image Process. | 4 |
| 2015 | Blind bleed-through removal for scanned historical document images with conditional random fieldsabstractDue to the quality of paper and long-time preservation, the ink on one side of the historical documents often seeps through and appears on the other side. In this paper, a new blind ink bleed-through removal method is proposed to deal with the scanned historical document images. The scanned historical document image generally consists of three components: foreground, bleed-through and background. In the proposed method, conditional probability distribution (CPD) models of the three components are firstly established by statistics. Then, conditional random fields (CRFs) are used to model the observed scanned image and the corresponding labels. For each input scanned image, parameters of the component-wise CPD models are estimated and belief propagation is performed on the CRFs model to determine the most possible labels. Once the bleed-through component is found, an inpainting algorithm is proposed to remove the ink bleed-through from the input historical image. Experimental results show that the proposed method preserves the foreground component very well and removes the bleed-through effectively. Bin Sun 0001, Shutao Li 0001, Jun Sun 0004 |
ICASSP | 3 |
| 2015 | Reconstruction combined training for convolutional neural networks on character recognitionabstractRecently, the deep learning methods have achieved great success in pattern recognition tasks. Especially for character recognition, most of the state-of-the-art results belong to the deep learning models. Among those models, the convolutional neural network (CNN) becomes the most popular due to its outstanding performance. Therefore, many trials were made in order to make improvements on CNN. However, most of the trials only focused on the network structure or training skills, the inter-class information is usually ignored. In this paper, we have proposed a novel CNN model with two training feedbacks: the reconstruction feedback and the classification feedback. By using the reconstruction feedback, the inter-class information (for example, shape similarity) of the characters is taken into account. Consequently, without enlarging the network structure, our model can outperform those state-of-the-art improved CNN models, which is proved by the experimental results. Li Chen 0017, Song Wang 0007, Wei Fan 0005, Jun Sun 0004, Satoshi Naoi |
ICDAR | 4 |
| 2015 | Deep learning based language and orientation recognition in document analysisabstractIn practical applications of document understanding, if the documents have multiple languages and orientations, the conventional OCR systems can not be directly applied. This is because those OCR systems are usually designed for texts of single language and normal orientation. To solve this problem, many non-character based recognition approaches were proposed. However, the performance of those methods were not comparable with the mature OCR systems. Consequently, a better idea is to recognize the language type and orientation before the OCR is applied. Besides, the characters of different languages have very ambiguous shape, so it is very difficult to extract stable feature for the recognition. Recently, the convolutional neural networks (CNN) have achieved great success in pattern recognition tasks. Therefore, for such difficult tasks, the CNN is one of the best choice. In this paper, we first applied CNN to the recognition of the document properties. A novel sliding window voting process is proposed to reduce the network scale and fully use the information of the text line. In the experiments, our method had very high recognition rate. The results proved the advantage of the proposed method and which also can be applied to create a document understanding system with OCR systems. Li Chen 0017, Song Wang 0007, Wei Fan 0005, Jun Sun 0004, Satoshi Naoi |
ICDAR | 4 |
| 2015 | Seamless stitching with shape deformation for historical document imagesabstractIn historical document images stitching, the arbitrary placement and the low quality of the historical paper always bring in uneven surface of the document, which might introduce local distortion to the image content. In these situations, some traditional methods may fail to find an optimal seam as they just use the content similarity. Some other methods tried to find two corresponding seams in both images, and the contents of the two images at the opposite sides of seams are transformed and stitched together. Although the image contents are correctly aligned with these methods, the transform may lead to huge content deformation and geometric inconsistency. To solve this problem, we proposed a seamless stitching method with shape deformation. In this method, one of the images is first deformed to approximate the other image, then single or dual optimal seams are estimated using an optimization method to minimize the deformation and to maximize the deformation consistency. With this method, the deformation of image content along the seam is consistent, and geometrically consistent elements such as straight lines can be well preserved. Experimental results show the effectiveness and efficiency of the proposed method. Wei Liu 0169, Wei Fan 0005, Li Chen 0017, Jun Sun 0004, Satoshi Naoi |
ICDAR | 4 |
| 2015 | Text line extraction in document imagesabstractText line extraction in document images is an important prerequisite for many content based image understanding applications. In this paper, we propose an accurate and robust method for generic text line extraction, which can be applied on large categories of document images, diverse languages, and text lines with different orientations. Firstly, the candidate connected components are extracted from document image using Maximal Stable Extremal Region (MSER) with the noises filtered by Adaboost and Convolution Neural Network (CNN). Then, the coarse text lines are generated from hierarchical edges reconstruction and cut by local linearity of text lines in the document spanning tree. Finally, for accurate text line extraction, the cut multi-components are re-connected based on text line energy minimization in terms of text line consistency and the fitting error. Experimental results on multilingual test dataset demonstrate the effectiveness and robust of the proposed method, which yields higher performance compared with state-of-the-art methods. Liuan Wang, Wei Fan 0005, Jun Sun 0004, Satoshi Naoi, Tanaka Hiroshi |
ICDAR | 3 |
| 2015 | Feature Reduction Using Ensemble Approach
Yingju Xia, Cuiqin Hou, Jun Sun 0004 |
PACLIC | 4 |
| 2014 | Real-Time Document Image Super-Resolution by Fast MattingabstractFrom a single low resolution image, a real-time document image super-resolution algorithm is proposed to obtain high resolution document image with sharp text boundaries. First, a highly efficient document image matting algorithm based on local linear modeling is designed to decompose the input image into text, foreground and background layers, which contain the text edge information, the color information of the foreground and background respectively. Then the text layer is up-sampled with Teager filter to increase the sharpness of the text. For efficiency, the foreground and background layers are simply up-sampled through the bi-cubic interpolation. Finally, these three high resolution layers are composed to obtain the high-resolution image. Experiments on real scanned document images demonstrate the effectiveness of the proposed method in both visual perception and OCR performance Xudong Kang, Shutao Li 0001, Yuan He 0001, Jun Sun 0004 |
Document Analysis Systems | 5 |
| 2014 | Paper stitching using maximum tolerant seam under local distortionsabstractPaper stitching technology can reconstruct a whole paper page from two sub-images separately scanned from a camera with limited vision field. Wei Fan 0005, Jun Sun 0004, Satoshi Naoi |
ACM Symposium on Document Engineering | 2 |
| 2014 | Handwritten Character Recognition by Alternately Trained Relaxation Convolutional Neural NetworkabstractDeep learning methods have recently achieved impressive performance in the area of visual recognition and speech recognition. In this paper, we propose a hand- writing recognition method based on relaxation convolutional neural network (R-CNN) and alternately trained relaxation convolutional neural network (ATR-CNN). Previous methods regularize CNN at full-connected layer or spatial-pooling layer, however, we focus on convolutional layer. The relaxation convolution layer adopted in our R-CNN, unlike traditional convolutional layer, does not require neurons within a feature map to share the same convolutional kernel, endowing the neural network with more expressive power. As relaxation convolution sharply increase the total number of parameters, we adopt alternate training in ATR-CNN to regularize the neural network during training procedure. Our previous C- NN took the 1st place in ICDAR'13 Chinese Handwriting Character Recognition Competition, while our latest ATR-CNN outperforms our previous one and achieves the state-of-the-art accuracy with an error rate of 3.94%, further narrowing the gap between machine and human observers (3.87%). Chunpeng Wu, Wei Fan 0005, Yuan He 0001, Jun Sun 0004, Satoshi Naoi |
ICFHR | 4 |
| 2014 | Fast and Accurate Text Detection in Natural Scene Images with User-IntentionabstractText detection in natural scene images plays an important role in content-based image retrieval, especially user-guided text detection for human-computer interaction. In this paper, we propose a fast and accurate text detection method with user-intention in terms of tap gesture. Firstly, a user-intention slice descriptor is designed based on the estimated text property, which contains all the user interested texts, and fast heuristic features and accurate texture feature of decomposed connected components (CCs) are fed into cascade of Gentle Adaboost classifiers to eliminate non-text candidates, finally candidate texts, sharing the same property consistent with the seed CCs, are accumulated to a user-intention text line according to local and global permutation constraint. Experimental results demonstrate the effectiveness and robustness of the proposed method in comparison with the state-of-art methods. Liuan Wang, Wei Fan 0005, Yuan He 0001, Jun Sun 0004, Yutaka Katsuyama, Yoshinobu Hotta |
ICPR | 4 |
| 2014 | Detecting Keyphrases in Micro-blogging with Graph Modeling of Information Diffusion
Shuangyong Song, Jun Sun 0004 |
PRICAI | 3 |
| 2014 | A Fast and Robust Multi-color Object Detection Method with Application to Color Chart Detection
Song Wang 0007, Akihiro Minagawa, Wei Fan 0005, Jun Sun 0004 |
PRICAI | 4 |
| 2014 | Graphical lasso quadratic discriminant function and its application to character recognition
Bo Xu 0008, Kaizhu Huang, Irwin King, Cheng-Lin Liu 0001, Jun Sun 0004, Satoshi Naoi |
Neurocomputing | 5 |
| 2014 | Book inner boundary extraction with modified active shape model
Shufu Xie, Yuan He 0001, Jun Sun 0004, Satoshi Naoi |
Pattern Recognit. Lett. | 4 |
| 2014 | Scanned Image Descreening With Image Redundancy and Adaptive FilteringabstractCurrently, most electrophotographic printers use halftoning technique to print continuous tone images, so scanned images obtained from such hard copies are usually corrupted by screen like artifacts. In this paper, a new model of scanned halftone image is proposed to consider both printing distortions and halftone patterns. Based on this model, an adaptive filtering based descreening method is proposed to recover high quality contone images from the scanned images. Image redundancy based denoising algorithm is first adopted to reduce printing noise and attenuate distortions. Then, screen frequency of the scanned image and local gradient features are used for adaptive filtering. Basic contone estimate is obtained by filtering the denoised scanned image with an anisotropic Gaussian kernel, whose parameters are automatically adjusted with the screen frequency and local gradient information. Finally, an edge-preserving filter is used to further enhance the sharpness of edges to recover a high quality contone image. Experiments on real scanned images demonstrate that the proposed method can recover high quality contone images from the scanned images. Compared with the state-of-the-art methods, the proposed method produces very sharp edges and much cleaner smooth regions. Bin Sun 0001, Shutao Li 0001, Jun Sun 0004 |
IEEE Trans. Image Process. | 3 |
| 2013 | A Book Dewarping System by Boundary-Based 3D Surface ReconstructionabstractNon-contact imaging devices such as digital cameras and overhead scanners can convert hardcopy books to digital images without cutting them to individual pages. However, the captured images have distinct distortions. A book dewarping system is proposed to remove the perspective and geometric distortions automatically from single images. A book boundary model is extracted, and a 3D book surface is reconstructed. And then the horizontal and vertical metrics of each column are restored from it. Experimental results show the good dewarping and speed performance. Since no additional equipments and no restrictions to specific book layouts or contents are needed, the proposed system is very practical in real applications. Yuan He 0001, Shufu Xie, Jun Sun 0004, Satoshi Naoi |
ICDAR | 4 |
| 2013 | Sub-structure Learning Based Handwritten Chinese Text RecognitionabstractThis paper proposed a sub-structure learning based method for handwritten Chinese text recognition. In conventional methods, a standard character recognizer is trained on character classes only. Unreliable recognition results on character segments will decrease final recognition precision. By discovering stable sub-structure patterns from real character segment samples automatically, both character and sub-structure patterns are trained in character recognizer. The judgment reliability of segments being characters is significantly improved. Furthermore, to deal with millions of training segment samples, a two-stage clustering method is proposed for sub-structure learning. Experiment results on HIT-MW database show that the sub-structure learning based method improves performance significantly. The F1-measure evaluation of handwritten Chinese text recognition is improved by 8.84%. Yuanping Zhu, Jun Sun 0004, Satoshi Naoi |
ICDAR | 2 |
| 2013 | Text detection in natural scene images with user-intentionabstractWe propose an accurate and robust coarse-to-fine text detection scheme with user-intention which captures the intrinsic characteristics of natural scene texts. In the coarse detection stage, a double edge detector is designed to estimate the symmetry of stroke and the stroke width, which help segment the foreground. Then the initial user-intention region is extended to generate a coarse bounding box based on the estimated foreground. In the refinement stage, candidate connected components (CCs) from Niblack decomposition, are grouped together by location to form text lines after noise removal and layer selection. Experimental results demonstrate the effectiveness of the proposed method which yields higher performance compared with state-of-the-art methods. Liuan Wang, Yutaka Katsuyama, Wei Fan 0005, Yuan He 0001, Jun Sun 0004, Yoshinobu Hotta |
ICIP | 5 |
| 2013 | Spatially consistent exemplar-based clusteringabstractExemplar-based clustering has drawn much attention in recent years as it produces state-of-the-art results on many practical clustering problems. However, spatial information is missed in the exemplar-based clustering methods, resulting in difficulties in some applications, for example in the image segmentation problem. In this paper, we investigate the issue of integrating spatial information into the exemplar-based clustering through the Markov random field formulation. Two algorithms are proposed to achieve this aim. First, based on the min-sum loopy belief propagation algorithm, a spatially consistent affinity propagation algorithm is proposed. Second, by showing the spatially consistent exemplar-based clustering energy function satisfies the regular property, an efficient minimal s-t graph cut based convergent algorithm is proposed. Experimental results on the image segmentation problem show that the spatially consistent exemplar-based clustering achieves better results than other methods. Pei Chen 0001, Yuan He 0001, Jun Sun 0004, Haifeng Hu 0001 |
ICME | 4 |
| 2012 | Local Consistency Constrained Adaptive Neighbor Embedding for Text Image Super-ResolutionabstractThis paper proposes a robust single-image super-resolution method for enlarging low quality camera captured text image. The contribution of this work is twofold. First, we point out the non-local reconstruction problem in neighbor embedding based super-resolution by statistical analysis on an empirical data set. Second, we introduce a local consistency constraint to explicitly regularize the linear reconstruction process, and adaptively generate the most possible candidates for the high-resolution image patch. For the non-consistent candidates, we rely on its adjacent overlapping patches for capability verification. Experimental results demonstrate that our solution produces visually pleasing enlargements for various text images. Wei Fan 0005, Jun Sun 0004, Satoshi Naoi, Akihiro Minagawa, Yoshinobu Hotta |
Document Analysis Systems | 2 |
| 2012 | A Fast Caption Detection Method for Low Quality Video ImagesabstractCaptions in videos are important and accurate clues for video retrieval. In this paper, we propose a fast and robust video caption detection and localization algorithm to handle low quality video images. First, the stroke response maps from complex background are extracted by a stoke filter. Then, two localization algorithms are used to locate thin stroke and thick stroke caption regions respectively. Finally, a HOG based SVM classifier is carried out on the detected results to further remove noises. Experimental results show the superior performance of our proposed method compared with existing work in terms of accuracy and speed. Tianyi Gui, Jun Sun 0004, Satoshi Naoi, Yutaka Katsuyama, Akihiro Minagawa, Yoshinobu Hotta |
Document Analysis Systems | 2 |
| 2012 | Structured document classification by matching local salient features
Yuan He 0001, Jun Sun 0004, Satoshi Naoi |
ICPR | 3 |
| 2012 | Cascaded heterogeneous convolutional neural networks for handwritten digit recognition
Chunpeng Wu, Wei Fan 0005, Yuan He 0001, Jun Sun 0004, Satoshi Naoi |
ICPR | 4 |
| 2012 | Discriminative normalization method for handwritten Chinese character recognition
Yuanping Zhu, Jun Sun 0004, Satoshi Naoi |
ICPR | 2 |
| 2012 | A Study on Caption Recognition for Multi-color Characters on Complex BackgroundabstractWe propose a caption recognition method for multicolor characters on complex background. Caption characters are used for an efficient search on a large amount of recorded TV programs. In the caption character recognition, the caption appearance section and the area is extracted, the character strokes are extracted from the area, and recognized. This paper focuses on caption character strokes extraction and recognition for multi-color characters on complex background which is a very difficult task for the conventional methods. The proposed method extracts decomposed binary images from input color caption image by color clustering. Then character candidates that are composed of combination of connect components are extracted by using recognition certainty. Finally, characters are selected by beyond-color Dynamic Programming method in which weight on recognition certainty and character alignment are used. In the character recognition evaluation of one-line multi-color character string on a complex background, a great improvement was achieved from a conventional technique that can recognize only one-color characters on complex background image. Yutaka Katsuyama, Akihiro Minagawa, Yoshinobu Hotta, Jun Sun 0004, Shinichiro Omachi |
ISM | 4 |
| 2012 | Mobile-based advertisement information retrieval from images and websitesabstractIn the real world, there are a huge amount of advertisement (ad) boards to make customers have a visual awareness of the products or services easily. However, information appearing in the ad boards is so limited that customers always want to know more ad details in a convenient way. In this paper, we present an mobile-based prototype system to automatically extract web ad information from images and websites. After capturing ad images by smartphones and sending them to a remote server, ad image text is recognized by OCR engine, from where ad phrases and keywords are extracted and combined together as queries. Ad web page candidates are then obtained by specific search engines and clustered to remove noises. OCR results are further used to estimate valid ad topic web pages which are pushed back to end users for searching more detailed ad information. Based on the experiments on a real-world ad image dataset collected by ourselves, true ad topic web pages can be found from top-one and top-ten returned pages in about 51.85% and 83.33% query images respectively, which illustrates the effectiveness of the proposed system. Yi-Feng Pan, Yuan He 0001, Yingju Xia, Jun Sun 0004, Satoshi Naoi |
ACM Multimedia | 6 |
| 2011 | Recognizing Characters with Severe Perspective Distortion Using Hash Tables and Perspective InvariantsabstractIn this paper, we present a novel method to recognize characters with severe perspective distortion using hash tables and perspective invariants. The proposed algorithm consists of storage and voting stages. With the help of perspective invariants, the combinations of 4-tuple bases for the perspective invariant coordinate system are searched out in an efficiently way. The bases are further selected so that the resulting transformation is effective. The characters' features under the perspective invariant coordinate system determine an entry in a one dimensional hash table, which is applied for storage and retrieval. Experimental results show the superior performance of the proposed method in comparison to other existing methods. Yuanping Zhu, Jun Sun 0004, Satoshi Naoi |
ICDAR | 3 |
| 2011 | Improving Scene Text Detection by Scale-Adaptive Segmentation and Weighted CRF VerificationabstractThis paper presents a hybrid method for detecting and localizing texts in natural scene images by stroke segmentation, verification and grouping. To improve system performance, novelties on two aspects are proposed: 1) a scale-adaptive segmentation method is designed for extracting stroke candidates, and 2) a CRF model with pair-wise weight by local line fitting is designed for stroke verification. Moreover, color-based text region estimation is used to guide segmentation and verification more accurately. Experimental results on ICDAR 2005 competition dataset show that the proposed approach can detect and localize scene texts with high accuracy, even under noisy and complex backgrounds. Yi-Feng Pan, Yuanping Zhu, Jun Sun 0004, Satoshi Naoi |
ICDAR | 3 |
| 2011 | Robust Vanishing Point Detection for MobileCam-Based DocumentsabstractDocument images captured by a mobile phone camera often have perspective distortions. In this paper, fast and robust vanishing point detection methods for such perspective documents are presented. Most of previous methods are either slow or unstable. Based on robust detection of text baselines and character tilt orientations, our proposed technology is fast and robust with the following features: (1) quick detection of vanishing point candidates by clustering and voting on the Gaussian sphere space, and (2) precise and efficient detection of the final vanishing points using a hybrid approach, which combines the results from clustering and projection analysis. The rectified image acceptance rate for Mobile Cam-based documents, signboards and posters is more than 98% with an average speed of about 100ms. Xu-Cheng Yin, Hongwei Hao, Jun Sun 0004, Satoshi Naoi |
ICDAR | 3 |
| 2011 | Graphical Lasso Quadratic Discriminant Function for Character Recognition
Bo Xu 0008, Kaizhu Huang, Irwin King, Cheng-Lin Liu 0001, Jun Sun 0004, Satoshi Naoi |
ICONIP (3) | 5 |
| 2010 | Occluded text restoration and recognitionabstractText occlusion is among the most intractable obstacles for OCR engines. A typical example in document images is visible watermark characters, which are often occluded by foreground contents. This paper proposes a solution by restoring watermark characters before recognition. The text restoration process consists a core module as patch-based restoration method, which reconstructs the missing areas by referring to similar patches from undamaged areas. The filling sequence is in a order based on the structure complexity inside each patch, which helps to suppress reconstruction error propagation. Furthermore, the patch size is adaptively selected based on the local character stroke width. Experiments show that the proposed method produces good restoration quality and effectively improves the recognition rate of the following OCR process. Furthermore, the algorithm is optimized based on statistical analysis model and the processing time meets the real-time responding requirement. Lanlan Chang, Jun Sun 0004, Misako Suwa, Hiroaki Takebe, Yuan He 0001, Satoshi Naoi |
Document Analysis Systems | 2 |
| 2010 | Rejection Optimization Based on Threshold Mapping for Offline Handwritten Chinese Character RecognitionabstractIn this paper, a rejection optimization method based on rejection threshold mapping is proposed. Different from conventional rejection methods which use the same rejection threshold for all samples, this technique utilizes the local information of samples to optimize the rejection threshold. The samples with the same class pair in the first two recognition candidates are treated as the same sample category. Based on the rejection distribution of the sample categories, the parameters of rejection threshold mapping of similar character pairs are learned and stored. When performing rejection, the corresponding mapping parameters are searched according to the class pair in the first two recognition candidates, and applied on the input threshold. The transformed threshold is used in final rejection decision. The experiments show that it is able to decrease error rate under same rejection rate on handwritten Chinese character recognition which verify its effectiveness. Yuanping Zhu, Jun Sun 0004, Yoshinobu Hotta, Satoshi Naoi |
ICFHR | 2 |
| 2010 | Separation of overlapped color planes for document imagesabstractColor plane separation is very useful in processing color document images. Many reported methods take it as a multi-class classification problem and work not well in overlapped color regions. This paper proposed a simple but effective linear projection based method for separating overlapped color planes. The separation task is taken as a probability problem, i.e., in the output plane, target color should have high response and the other colors should have low response, or vice versa. Furthermore, it assumes that the number of foreground colors is low, typically one to four, and overlapped areas contain mixed colors instead of opaque covering. Experimental results demonstrate the effectiveness and flexibility of our method. Danian Zheng, Jun Sun 0004, Satoshi Naoi, Misako Suwa, Hiroaki Takebe, Yoshinobu Hotta |
ICIP | 2 |
| 2010 | A Dual Pass Video Stabilization System Using Iterative Motion Estimation and Adaptive Motion SmoothingabstractIn this paper, we propose a novel dual pass video stabilization system using iterative motion estimation and adaptive motion smoothing. In the first pass, the transformation matrix to stabilize each frame is returned. The global motion estimation is carried out by a novel iterative method. The intentional motion is estimated using adaptive window smoothing. Before the beginning of the second pass, we obtain the optimal trim size for a specific video based on the statistics of the transformation parameters. In the second pass, the stabilized video is composed according to the optimal trim size. Experimental results show the superior performance of the proposed method in comparison to other existing methods. Akihiro Minagawa, Jun Sun 0004, Yoshinobu Hotta, Satoshi Naoi |
ICPR | 3 |
| 2010 | Sparse learning for support vector classification
Kaizhu Huang, Danian Zheng, Jun Sun 0004, Yoshinobu Hotta, Katsuhito Fujimoto, Satoshi Naoi |
Pattern Recognit. Lett. | 3 |
| 2009 | Trinary Image Mosaicing Based Watermark String DetectionabstractWatermark string detection is very useful for document security protection. In we proposed an image based document watermark detection system. In this paper, two major modifications are made based on the previous system.First, a trinary image mosaicing algorithm is proposed to merge the broken watermark string images into a high quality mosaic image. Second, we relax the conditions in the Maximum Clique based keyword detection algorithm.The new algorithm can handle watermark string with variant spacing. Therefore, not only English keywords, but also Japanese/Chinese keywords can be processed. The Experimental results show the effectiveness of our algorithm. Jun Sun 0004, Satoshi Naoi, Yusaku Fujii, Hiroaki Takebe, Yoshinobu Hotta |
ICDAR | 1 |
| 2009 | Rejection Strategies with Multiple Classifiers for Handwritten Character RecognitionabstractWith rejection strategies in a handwriting recognition system, we are able to improve the reliability and accuracy of the recognized characters. In this paper, we propose several rejection strategies with multiple classifiers for handwritten character recognition. First, the rejection strategy for the single classifier is introduced, which is composed of three stages: initial scaling, confidence measure calculation, and rejection performing. Then, we analyze rejection strategies for multiple classifiers. We divided our rejection strategies into two categories: (1) for voting combination; and (2) for linear combination with multiple classifiers. In the voting combination style, three rejection strategies, OR, AND, and VOTING, are proposed. And for the linear combination one, rejection strategies for average and weighted combination are analyzed respectively. We also experiment and compare our rejection strategies with handwritten digit recognition. Xu-Cheng Yin, Hongwei Hao, Yun-Feng Tang, Jun Sun 0004, Satoshi Naoi |
ICDAR | 4 |
| 2009 | Separate Chinese Character and English Character by Cascade Classifier and Feature SelectionabstractThe separation of Chinese character and English character is helpful for OCR technique. In this paper, a multi-level cascade classifier combined with feature selection is constructed to identify Chinese character and English character based on individual character. Most of samples are identified by the first node classifier, the remained low classification confidence samples are fed to the next node classifiers to get the final result. For the motivation of utilizing feature complementarity, each node classifier is trained on low classification confidence samples of its previous node classifier with independent feature selection. Furthermore, a confidence bias is utilized to improve the classifier generalization. The experiment results validate the effectiveness of this classifier. Yuanping Zhu, Jun Sun 0004, Akihiro Minagawa, Yoshinobu Hotta, Satoshi Naoi |
ICDAR | 2 |
| 2008 | Unsupervised Decomposition of Color Document Images by Projecting Colors to a Spherical SurfaceabstractA decomposition method for color document images is proposed in this paper. A two dimensional feature surface is constructed from the input color image, and then a novel and unsupervised method based on contour lines is proposed to segment the surface. In detail, colors of the image pixels are firstly projected to a spherical surface whose center is the background color. The projection is used to transform the observed colors to the corresponding 'ideal' foreground colors. Then the spherical surface is segmented into several non-overlapped regions, and each region corresponds to an individual layer of the input color document image. Finally, the image pixels are projected to the spherical surface and classified to the corresponding layers. Yuan He 0001, Jun Sun 0004, Satoshi Naoi, Yusaku Fujii, Katsuhito Fujimoto |
Document Analysis Systems | 2 |
| 2008 | An Image Based Watermark String Detection System for Document Security CheckingabstractDocument security is a very important topic in information management. In this paper, an image based watermark string detection system is proposed to detect the documents that include printed keyword strings as the watermark in the background. Therefore, the disclosure of the sensitive documents can be monitored automatically. Since the documents are represented in image format, the watermark string is detected by a parts based object recognition strategy. The two key contributions of our paper are cross validation based image registration and theMaximum Clique (MC) based object parts recognition. Experiments on PPT and WORD document pages with 5 different watermark keywords show the excellent performance of our system. Jun Sun 0004, Yusaku Fujii, Hiroaki Takebe, Katsuhito Fujimoto, Satoshi Naoi |
Document Analysis Systems | 1 |
| 2008 | Video caption duration extractionabstractCaption detection in the video is an active research topic in recent years. In the conventional methods, one of most difficult problems is to effectively and quickly extract the durations of the different-size captions in the complex background. To solve this problem, a novel and effective method is presented to locate and track the captions in the video. The main contributions are: (1)present a multi-scale Harris-corner based method to detect the initial position of the caption (2)propose the SGF (Steady Global Feature) to determine the caption duration. Extensive experiments demonstrate the effectiveness of the proposed method. Hongliang Bai, Jun Sun 0004, Satoshi Naoi, Yutaka Katsuyama, Yoshinobu Hotta, Katsuhito Fujimoto |
ICPR | 2 |
| 2007 | Curved paper rectification for digital camera document images by shape from parallel geodesics using continuous dynamic programmingabstractMethods to rectify distortion of digital camera document images of curved papers have become important for camera-based image recognition. In this paper we propose a novel distortion rectification method based on "shape from parallel geodesies." This method considers the following features: parallel lines corresponding to character strings or ruled lines of tables on extended surface become parallel geodesies on a curved paper surface and a smoothly curved paper can be modeled by a ruled surface, that is, a sweep surface of rulings. The projected geodesies and the projected rulings exist in the input image derived from perspective transformation. The presented method extracts the projected geodesies, estimates the projected rulings in the input image, estimates the ruled surface that models the curved paper, and generates the corrected image, in this order. It can estimate the ruled surface model directly by numerical operations of differentiation, integration and matrix inversion without any iterative calculation. We also report on experiments that show the effectiveness of the proposed method. Katsuhito Fujimoto, Jun Sun 0004, Hiroaki Takebe, Misako Suwa, Satoshi Naoi |
ICDAR | 2 |
| 2007 | An SVM-Based High-accurate Recognition Approach for Handwritten Numerals by Using Difference FeaturesabstractHandwritten numeral recognition is an important pattern recognition task. It can be widely used in various domains, e.g., bank money recognition, which requires a very high recognition rate. As a state-of-the-art classifier, support vector machine (SVM), has been extensively used in this area. Typically, SVM is trained in a batch model, i.e., all data points are simultaneously input for training the classification boundary. However, some slightly exceptional data, only accounting for a small proportion, are critical for the recognition rates. Training a classifier among all the data may possibly treat such legal but slightly exceptional samples as "noise ". In this paper, we propose a novel approach to attack this problem. This approach exploits a two-stage framework by using difference features. In the first stage, a regular SVM is trained on all the training data; in the second stage, only the samples misclassified in the first stage are specially considered. Therefore, the performance can be lifted. The number of misclassifications is often small because of the good performance of SVM. This will present difficulties in training an accurate SVM engine only for these misclassified samples. We then further propose a multi-way to binary approach using difference features. This approach successfully transforms multi-category classification to binary classification and expands the training samples greatly. In order to evaluate the proposed method, experiments are performed on 10,000 handwritten numeral samples extracted from real banks forms. This new algorithm achieves 99.0% accuracy. In comparison, the traditional SVM only gets 98.4%. Kaizhu Huang, Jun Sun 0004, Yoshinobu Hotta, Katsuhito Fujimoto, Satoshi Naoi |
ICDAR | 2 |
| 2007 | Degraded Character Recognition by Complementary Classifiers CombinationabstractCharacter degradation is a big problem for machine printed character recognition. Two main reasons for degradation are extrinsic image degradation such as blurring and low image dimension, and intrinsic degradation caused by font variations. A recognition method that combines two complementary classifiers is proposed in this paper. The local feature based classifier extracts the local contour direction changes, which is effective for character patterns with less structure deterioration. The global feature based classifier extracts the texture distribution of the character image, which is effective when the character structure is hard to discriminate. The two complementary classifiers are combined by candidate fusion in a coarse-to-fine style. Experiments are carried on degraded Chinese character recognition. The results prove the effectiveness of our method. Jun Sun 0004, Kaizhu Huang, Yoshinobu Hotta, Katsuhito Fujimoto, Satoshi Naoi |
ICDAR | 1 |
| 2007 | A Multi-Stage Strategy to Perspective Rectification for Mobile Phone Camera-Based Document ImagesabstractDocument images captured by a mobile phone camera often have perspective distortions. Efficiency and accuracy are two important issues in designing a rectification system for such perspective documents. In this paper, we propose a new perspective rectification system based on vanishing point detection. This system achieves both the desired ef- ficiency and accuracy using a multi-stage strategy: at the first stage, document boundaries and straight lines are used to compute vanishing points; at the second stage, text base- lines and block aligns are utilized; and at the last stage, character tilt orientations are voted for the vertical vanish- ing point. A profit function is introduced to evaluate the reliability of detected vanishing points at each stage. If van- ishing points at one stage are reliable, then rectification is ended at that stage. Otherwise, our method continues to seek more reliable vanishing points in the next stage. We have tested this method with more than 400 images includ- ing paper documents, signboards and posters. The image acceptance rate is more than 98.5% with an average speed of only about 60ms. Xu-Cheng Yin, Jun Sun 0004, Satoshi Naoi, Katsuhito Fujimoto, Yusaku Fujii, Koji Kurokawa, Hiroaki Takebe |
ICDAR | 2 |
| 2007 | Skew detection using wavelet decomposition and projection profile analysis
Shutao Li 0001, Qinghua Shen, Jun Sun 0004 |
Pattern Recognit. Lett. | 3 |
| 2006 | Robust Chinese Character Recognition by Selection of Binary-Based and Grayscale-Based Classifier
Yoshinobu Hotta, Jun Sun 0004, Yutaka Katsuyama, Satoshi Naoi |
Document Analysis Systems | 2 |
| 2006 | A Hybrid Handwritten Chinese Address Recognition Approach
Kaizhu Huang, Jun Sun 0004, Yoshinobu Hotta, Katsuhito Fujimoto, Satoshi Naoi, Chong Long, Li Zhuang, Xiaoyan Zhu 0001 |
ICONIP (2) | 2 |
| 2005 | Camera based Degraded Text Recognition Using Grayscale FeatureabstractAs the rapid progress of digital imaging technology, camera based character recognition receives more and more attentions. One challenge in camera based OCR is the recognition for degraded text. Conventional OCR engines usually recognize on binary image. However, the performance drops dramatically as the degradation level increases. In this paper, a new recognition method is proposed to recognize degraded character based on dual eigenspace decomposition and synthetic degraded data. Then, the degraded character string is segmented by the combination of binary and grayscale analysis. Experiments on single character and text string recognition prove the effectiveness of our method. Jun Sun 0004, Yoshinobu Hotta, Yutaka Katsuyama, Satoshi Naoi |
ICDAR | 1 |
| 2004 | Video Degradation Model and Its Application to Character Recognition in e-Learning Videos
Jun Sun 0004, Yutaka Katsuyama, Satoshi Naoi |
Document Analysis Systems | 1 |
| 2003 | Effective text extraction and recognition for WWW imagesabstractImages play a very important role in web content delivery. Many WWW images contain text information that can be used for web indexing and searching. A new text extraction and recognition algorithm is proposed in this paper. The character strokes in the image are first extracted by color clustering and connected component analysis. A novel stroke verification algorithm is used to effectively remove non-character strokes. The verified strokes are then used to build the binary text line image, which is segmented and recognized by dynamic programming. Since text in WWW image usually has close relationship with webpage content, approximate string matching is used to revise the recognition result by matching the content in the webpage with the content in the image. This effective post-processing not only improves the recognition performance, but also can be used in other applications such like image - webpage paragraph corresponding. Jun Sun 0004, Zhulong Wang, Hao Yu 0005, Fumihito Nishino, Yutaka Katsuyama, Satoshi Naoi |
ACM Symposium on Document Engineering | 1 |
| 2001 | Sparse image coding with clustering property and its application to face recognition
Jun Sun 0004, Qing Zhuo, Chengyuan Ma |
Pattern Recognit. | 1 |
| 2000 | An Improved Facial Expression Recognition Method
Jun Sun 0004, Qing Zhuo |
ICMI | 1 |