EDBT 2026 Demo / reviewers in the wild / expert
Jien Kato
dblp:97/1875
· DBLP profile ↗
33ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-0196-4405ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NSFF: Noise and semantic features fusion for AI-generated image detection
Ruiqiang Ma, Jien Kato |
Neural Networks | 4 |
| 2026 | LIBS: Instructional Action Quality Assessment via Supervoxel-Based Fine-Grained AttributionabstractThe lack of actionable guidance is a fundamental limitation in Action Quality Assessment (AQA), as traditional methods provide overall scores without offering specific insights for improvement. Moreover, existing interpretable approaches often rely on expensive supervised spatial annotations or yield noisy, unsigned saliency maps. To address these challenges, we propose Learning Interpretability Based Supervoxels (LIBS), a novel framework for generating instructional feedback. Distinguishing itself from fully supervised methods, LIBS employs an unsupervised soft-clustering mechanism to segment videos into coherent supervoxels without requiring pixel-level mask annotations. This allows for scalable, fine-grained spatio-temporal analysis while preserving action continuity. Furthermore, we introduce a sensitivity propensity analysis to quantify the contribution of each supervoxel. Unlike traditional attribution methods, this mechanism explicitly decomposes the quality score into positive (strengths) and negative (flaws) components, enabling the system to decode abstract scores into concrete, actionable instructions. Experimental validation across multiple datasets demonstrates that LIBS achieves superior interpretability and efficiency compared to state-of-the-art baselines, marking an improvement from diagnostic to instructional AQA applications. Xiaoyi Tao, Dongxu Ma, Liangzhi Li 0001, Manisha Verma, Lei Chen 0091, Xin Xie 0001, Sheng Chen 0015, Wenxin Li 0001, Jien Kato, Bing Zhang 0015, Xiulong Liu 0001 |
IEEE Trans. Computers | 9 |
| 2024 | Leveraging discriminative data: A pathway to high-performance, stable One-shot Network Pruning at Initialization
Yinan Yang 0001, Ying Ji 0003, Jien Kato |
Neurocomputing | 3 |
| 2024 | Efficient masked feature and group attention network for stereo image super-resolutionabstractCurrent stereo image super-resolution methods do not fully exploit cross-view and intra-view information, resulting in limited performance. While vision transformers have shown great potential in super-resolution, their application in stereo image super-resolution is hindered by high computational demands and insufficient channel interaction. This paper introduces an efficient masked feature and group attention network for stereo image super-resolution (EMGSSR) designed to integrate the strengths of transformers into stereo super-resolution while addressing their inherent limitations. Specifically, an efficient masked feature block is proposed to extract local features from critical areas within images, guided by sparse masks. A group-weighted cross-attention module consisting of group-weighted cross-view feature interactions along epipolar lines is proposed to fully extract cross-view information from stereo images. Additionally, a group-weighted self-attention module consisting of group-weighted self-attention feature extractions with different local windows is proposed to effectively extract intra-view information from stereo images. Experimental results demonstrate that the proposed EMGSSR outperforms state-of-the-art methods at relatively low computational costs. The proposed EMGSSR offers a robust solution that effectively extracts cross-view and intra-view information for stereo image super-resolution, bringing a promising direction for future research in high-fidelity stereo image super-resolution. Source codes will be released at https://github.com/jianwensong/EMGSSR . Jianwen Song, Arcot Sowmya, Jien Kato, Changming Sun |
Image Vis. Comput. | 3 |
| 2024 | MT-ASM: a multi-task attention strengthening model for fine-grained object recognition
Dichao Liu, Yu Wang 0018, Kenji Mase, Jien Kato |
Multim. Syst. | 4 |
| 2024 | BRT: Buffer Management for RDMA/TCP Mix-Flows in Datacenter NetworksabstractThe coexistence of RDMA and TCP is prevalent in the datacenter. Despite the sound isolation at the end hosts, they share the same switches in the network. Their different networking behaviors (E.g., in hardware demand and transport protocols) lead to huge differentiated buffer demand for switches. However, existing buffer management schemes ignore these dissimilarities and simply treat such RDMA/TCP mix-flows as the typical multi-class traffic, resulting in inferior isolation and degrading networking performances. This paper presents BRT, a first systematic solution for the buffer management of RDMA/TCP mix-flows in the DCN. BRT’s key insight is to allocate buffer with the awareness of traffic’s networking characteristics while minimally impacting the other’s performance. Guided by this insight, it first employs a traffic characteristics-based window to detect whether queues are in the state of persistent long queues. Then, it adjusts the total allocated buffer for each traffic type based on the number of persistent long queues and the normalized dequeue rates to reduce the buffer occupancy of meaningless queuing. Last, it calculates the buffer threshold for RDMA/TCP queues separately and uses a simple yet effective approach to prioritize the absorption of small flows. Our large-scale packet-level evaluations show that BRT can effectively optimize the networking performances for RDMA/TCP mix-flows. For example, compared to current practice, BRT achieves up to 53.5%, 46.7%, and 48.5% lower average FCT for incast flows, RDMA small flows, and TCP small flows, respectively, without sacrificing the overall throughput. Song Zhang 0008, Wenxin Li 0001, Lide Suo, Yulong Li 0001, Jien Kato, Keqiu Li |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2023 | Spatial-temporal Concept based Explanation of 3D ConvNetsabstractConvolutional neural networks (CNNs) have shown remarkable performance on various tasks. Despite its widespread adoption, the decision procedure of the network still lacks transparency and interpretability, making it difficult to enhance the performance further. Hence, there has been considerable interest in providing explanation and interpretability for CNNs over the last few years. Explainable artificial intelligence (XAI) investigates the relationship between input images or videos and output predictions. Recent studies have achieved outstanding success in explaining 2D image classification ConvNets. On the other hand, due to the high computation cost and complexity of video data, the explanation of 3D video recognition ConvNets is relatively less studied. And none of them are able to produce a high-level explanation. In this paper, we propose a STCE (Spatial-temporal Concept-based Explanation) framework for interpreting 3D ConvNets. In our approach: (1) videos are represented with high-level supervoxels, similar supervoxels are clustered as a concept, which is straightforward for human to understand; and (2) the interpreting framework calculates a score for each concept, which reflects its significance in the ConvNet decision procedure. Experiments on diverse 3D ConvNets demonstrate that our method can identify global concepts with different importance levels, allowing us to investigate the impact of the concepts on a target task, such as action recognition, in-depth. The source codes are publicly available at https://github.com/yingji425/STCE. Ying Ji 0003, Yu Wang 0018, Jien Kato |
CVPR | 3 |
| 2023 | Long-Tailed Image Recognition with Dynamic Re-WeightingabstractFor long-tailed image recognition tasks, re-weighting is effective to alleviate data imbalance by assigning higher weights to tail categories. However, existing re-weighting methods typically adopt a static weighting scheme, which usually hurts the accuracy of head categories. To deal with this issue, this paper proposes a progress-relevant weighting scheme called dynamic re-weighting, in which the weight assigned to a particular category first increases and then decreases, proportional to the number of samples that have been used in that category. In addition, we introduce a head-to-tail loss to control the evolving of weights, which makes the model gradually transfer its attention from head categories to tail categories. We conduct experiments on long-tailed CIFAR/ImageNet datasets, and confirm that our method not only outperforms static re-weighting methods, but also improves the accuracy on tail categories without sacrificing the accuracy of head categories. Yu Wang 0018, Jien Kato |
ICASSP | 3 |
| 2023 | Using Classifier Discrepancy for Cross-Domain Image RetrievalabstractIn recent years, cross-domain image retrieval (CDIM) has garnered considerable interest. The primary difficulty of CDIM is the domain gap, which makes it hard for the system to retrieve two photos that belong to the same category but have distinct domains. In this paper, we provide a novel multi-branch network employing the quintuplet structure to minimize retrieval loss and classifier discrepancy to minimize domain loss. Using three public datasets, we test the proposed method for zero-shot sketch-based image retrieval, which is one of CDIM's application tasks. Experiments validated the proposed method's state-of-the-art performance on the majority of datasets. Longjiao Zhao, Yu Wang 0018, Jien Kato |
ICIP | 3 |
| 2023 | Learn from each other to Classify better: Cross-layer mutual attention learning for fine-grained visual classificationabstractFine-grained visual classification (FGVC) is valuable yet challenging. The difficulty of FGVC mainly lies in its intrinsic inter-class similarity, intra-class variation, and limited training data. Moreover, with the popularity of deep convolutional neural networks, researchers have mainly used deep, abstract, semantic information for FGVC, while shallow, detailed information has been neglected. This work proposes a cross-layer mutual attention learning network (CMAL-Net) to solve the above problems. Specifically, this work views the shallow to deep layers of CNNs as “experts” knowledgeable about different perspectives. We let each expert give a category prediction and an attention region indicating the found clues. Attention regions are treated as information carriers among experts, bringing three benefits: (i) helping the model focus on discriminative regions; (ii) providing more training data; (iii) allowing experts to learn from each other to improve the overall performance. CMAL-Net achieves state-of-the-art performance on three competitive datasets: FGVC-Aircraft, Stanford Cars, and Food-11. The source code is available at https://github.com/Dichao-Liu/CMAL Dichao Liu, Longjiao Zhao, Yu Wang 0018, Jien Kato |
Pattern Recognit. | 4 |
| 2023 | Toward Extremely Lightweight Distracted Driver Recognition With Distillation-Based Neural Architecture Search and Knowledge TransferabstractThe number of traffic accidents has been continuously increasing in recent years worldwide. Many accidents are caused by distracted drivers, who take their attention away from driving. Motivated by the success of Convolutional Neural Networks (CNNs) in computer vision, many researchers developed CNN-based algorithms to recognize distracted driving from a dashcam and warn the driver against unsafe behaviors. However, current models have too many parameters, which is unfeasible for vehicle-mounted computing. This work proposes a novel knowledge-distillation-based framework to solve this problem. The proposed framework first constructs a high-performance teacher network by progressively strengthening the robustness to illumination changes from shallow to deep layers of a CNN. Then, the teacher network is used to guide the architecture searching process of a student network through knowledge distillation. After that, we use the teacher network again to transfer knowledge to the student network by knowledge distillation. Experimental results on the Statefarm Distracted Driver Detection Dataset and AUC Distracted Driver Dataset show that the proposed approach is highly effective for recognizing distracted driving behaviors from photos: (i) the teacher network’s accuracy surpasses the previous best accuracy; (ii) the student network achieves very high accuracy with only 0.42M parameters (around 55% of the previous most lightweight model). Furthermore, the student network architecture can be extended to a spatial-temporal 3D CNN for recognizing distracted driving from video clips. The 3D student network largely surpasses the previous best accuracy with only 2.03M parameters on the Drive&Act Dataset. The source code is available athttps://github.com/Dichao-Liu/Lightweight_Distracted_Driver_Recognition_with_Distillation-Based_NAS_and_Knowledge_Transfer Dichao Liu, Toshihiko Yamasaki, Yu Wang 0018, Kenji Mase, Jien Kato |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | One-shot Network Pruning at Initialization with Discriminative Image Patches
Yinan Yang 0001, Yu Wang 0018, Ying Ji 0003, Heng Qi, Jien Kato |
BMVC | 5 |
| 2021 | Rotation Invariance Analysis of Local Convolutional Features in Image RetrievalabstractRecently, local features computed using convolutional neural networks (CNNs) show good performance to image retrieval. However, the local convolutional features obtained by the CNNs (LC features) are inherently sensitive to rotation perturbations. This leads to miss-judgements in retrieval tasks. In this work, our objective is to enhance the robustness of LC features against image rotation. To do this, we conduct a thorough experimental evaluation of two candidate anti-rotation strategies (in-model data augmentation, and post-model feature augmentation), over two kinds of rotation attacks (dataset attack and query attack). We end up a series of good practices with steady quantitative supports, which lead to the best strategy for computing LC features with high rotation invariance in image retrieval. Longjiao Zhao, Yu Wang 0018, Jien Kato |
ICASSP | 3 |
| 2021 | Attention-Based Multi-Task Learning For Fine-Grained Image ClassificationabstractFine-Grained Image Classification is an inherently challenging task because of its inter-class similarity and intra-class variance. Most existing studies solve this problem by localization-and-classification strategies, which, however, always causes the problem of information loss or heavy computational expenses. Instead of localization-and-classification strategy, we propose a novel end-to-end optimization procedure named Multi-Task Attention Learning (MTAL), which reinforces the neural network’ correspondence to attention regions. Experimental results on CUB-Birds and Stanford Cars show that our procedure distinctly outperforms the baselines and is comparable with state-of-the-art studies despite its simplicity*. Dichao Liu, Yu Wang 0018, Kenji Mase, Jien Kato |
ICIP | 4 |
| 2020 | Contrastively-reinforced Attention Convolutional Neural Network for Fine-grained Image Recognition
Dichao Liu, Yu Wang 0018, Jien Kato, Kenji Mase |
BMVC | 3 |
| 2019 | Visual Violence Rating with Pairwise ComparisonabstractChildren's exposure to violence has become a severe problem with the rapid development of Internet. Recognizing violent video and estimating violence extent become crucial. Most researches focus on violent scene or violent action detection, lacking overall violence extent information. In this paper, we propose a violence rating prediction approach and build a novel violent video dataset. Our proposed method has two advantages: (1) videos are represented by features extracted from a learned two-stream network; (2) relationship between different violence extent can be learned and utilized to predict violence rating. To demonstrate the effectiveness of our method, we created a well-labelled dataset which contains 1, 930 violent videos. Each video is labelled with 6 objective violent attributes. Furthermore, we employ pairwise comparison method to obtain ground-truth violence rating for each video. Our proposed approach was evaluated on our dataset. Its results showed that our proposed method outperforms the state-of-art video classification methods. Ying Ji 0003, Yu Wang 0018, Jien Kato |
ICIP | 3 |
| 2019 | Good Choices for Deep Convolutional Feature EncodingabstractDeep convolutional neural networks can be used to produce discriminative image level features. However, when they are used as the feature extractor in a feature encoding pipeline, there are many design choices that are need to be made. In this work, we conduct a comprehensive study on deep convolutional feature encoding, by paying a special attention on its feature extraction aspect. We mainly evaluated the choices of the encoding methods; the choices of the base DCNN models; and the choices of the data augmentation methods. We not only quantitatively confirmed some known and previously unknown good choices for deep convolutional feature encoding, but also found out that some known good choices tune out to be bad. Base on the observations in the experiments, we present a very simple deep feature encoding pipeline, and confirmed its state-of-the-art performances on multiple image recognition datasets. Yu Wang 0018, Jien Kato |
WACV | 2 |
| 2017 | Solving Occlusion Problem in Pedestrian Detection by Constructing Discriminative Part LayersabstractOcclusion handling is one of the most challenging issues for pedestrian detection, and no satisfactory achievement has been found in this issue yet. Using human body parts has been considered as a reasonable way to overcome such an issue. In this paper, we propose a brand new approach based on the fusion of Mid-level body part mining and Convolutional Neural Network (CNN) to solve this problem, named DP-CNN(Discriminative Parts CNN). Two main discussions are included in this paper. First, we take an exhaustive analysis on how to mine useful body parts that contribute to pedestrian detection. Multiple ingredients (e.g. feature representation, pedestrian attributes) are analyzed through a wide range of experiments. Second, we convert the part detectors to the middle layer of CNN and re-train the model to get a better adaption of the dataset. Compare to existing approaches based on fine-tuning CNN models, our method is not only robust to occlusion handling but also has a smaller computational cost. Yu Wang 0018, Jien Kato, Guanwen Zhang, Kenji Mase |
WACV | 3 |
| 2015 | Development of a Real-World Oriented Smartphone AR Supported Learning System for Seasonal Constellation Observation
Ke Tian, Mayu Urata, Mamoru Endo, Katsuhiro Mouri, Takami Yasuda, Jien Kato |
ICCE | 6 |
| 2015 | Action recognition with approximate sparse codingabstractIn this paper, we present a novel feature encoding approach called Approximate Sparse Coding (ASC). ASC computes the sparse codes for a large collection of prototype descriptors in the off-line learning phase with Sparse Coding (SC); and look up the nearest prototype's sparse code for each to-be-encoded descriptor in the encoding phase with Approximate Nearest Neighbour (ANN) search. It shares the low dimensionality of SC and the fast speed of ANN, which are both desired properties for the human action recognition task. We excessively evaluated ASC on the popular HMDB51 dataset, and confirme it is able to encode large number of video features into discriminative low dimensional representations efficiently. Yu Wang 0018, Jien Kato |
ICIP | 2 |
| 2012 | Local Distance Comparison for Multiple-shot People Re-identification
Guanwen Zhang, Yu Wang 0018, Jien Kato, Takafumi Marutani, Kenji Mase |
ACCV (3) | 3 |
| 2012 | A distance metric learning based summarization system for nursery school surveillance videoabstractIn this paper, we present a system for summarizing nursery school surveillance video. The system takes full use of a learned distance metric, which can properly measure the similarity between videos. The metric is combined with supervised classification and unsupervised clustering, to categorize raw video materials into individual events. By selecting representative videos for each event, the system produces short video digests as the summarization output. The digests cover and reflect the children's activities on a daily basis. They are not only of interest to the parents, but also provide easy access to the mass quantity of daily surveillance video data. We implemented the proposed system in a real nursery school environment and confirmed its performance through both quantitative experiment and questionnaire survey. Yu Wang 0018, Jien Kato |
ICIP | 2 |
| 2007 | Extracting Features of Paper-made Objects Recognized from Origami Books Based on Design KnowledgeabstractThis paper proposes an approach to extracting features of origami works which are recognized from facing page images of instructions described as origami drill books. The origami work is presented as a set effaces (polygons). First, in order to associate geometric shapes of objects with knowledge, an origami data model based on design knowledge is proposed. Therefore, the set of faces is classified into clusters which indicate a certain meaning in origami design. The cluster is called "origami molecule ". Next, the relationships among the molecules are defined. The relationships are not only connected relationships but also symmetrical relationships and equivalence relationships. By using these relationships, it is possible to comprehend twin parts of origami work, for example, two forefeet of the four-limbed creature. Finally, a method for extraction of the design knowledge from crease patterns is proposed. Hiroshi Shimanuki, Jien Kato, Toyohide Watanabe |
ICDAR | 2 |
| 2005 | Method for Representing 3-D Virtual OrigamiabstractThis paper proposes a method for representing overlapping-faces of 3D virtual origami in order to support beholder's recognition of its conformation. Generally, an origami model is constructed by planar polygons corresponding to faces of origami. Therefore, when an origami model is displayed, multiple faces (polygons) on the same plane are probably recognized as one face. Our proposed method moves such overlapping-faces apart slightly by rotating polygons along a rotation axis determined from shapes and relationships of faces. As a consequence, beholders can recognize the conformation of origami correctly. This method can be applied to the user interface of a system we develop to recognize folding operations from origami drill books. Takashi Terashima, Hiroshi Shimanuki, Jien Kato, Toyohide Watanabe |
ICDAR | 3 |
| 2004 | An HMM/MRF-based stochastic framework for robust vehicle trackingabstractShadows of moving objects often obstruct robust visual tracking. In this paper, we present a car tracker based on a hidden Markov model/Markov random field (HMM/MRF)-based segmentation method that is capable of classifying each small region of an image into three different categories: vehicles, shadows of vehicles, and background from a traffic-monitoring movie. The temporal continuity of the different categories for one small region location is modeled as a single HMM along the time axis, independently of the neighboring regions. In order to incorporate spatial-dependent information among neighboring regions into the tracking process, at the state-estimation stage, the output from the HMMs is regarded as an MRF and the maximum a posteriori criterion is employed in conjunction with the MRF for optimization. At each time step, the state estimation for the image is equivalent to the optimal configuration of the MRF generated through a stochastic relaxation process. Experimental results show that, using this method, foreground (vehicles) and nonforeground regions including the shadows of moving vehicles can be discriminated with high accuracy. Jien Kato, Toyohide Watanabe, Sébastien Joga, Hiroyuki Hase |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2003 | Recognition of Folding Process from Origami Drill BooksabstractThis paper describes a framework to recognizing and recreating folding process of origami based on illustrations of origami drill books. Illustration images acquired from origami books are motley and not sequenced. Moreover, the information obtained from 2D illustrations is so superficial and incomplete that the folding operations cannot be determined uniquely. To solve these problems, a highly flexible and reliable recognition mechanism is proposed. The paper additionally includes the content as follows. Firstly, an algorithm for revising the positions of folding operations extracted from illustrations is proposed so as to make our recognition approach more reliable. Secondly, the outline of the methods which enable feasible folding operations to be generated based only on superficial and incomplete information extracted from illustrations is described. Finally, some updating procedures are proposed to maintain consistency of data (called internal model) which record the transformation of origami models in 3D virtual space during a folding process. Several examples that prove the validness of proposed algorithms/methods are also given in this paper. Hiroshi Shimanuki, Jien Kato, Toyohide Watanabe |
ICDAR | 2 |
| 2003 | Color segmentation for text extraction
Hiroyuki Hase, Masaaki Yoneda, Shogo Tokai, Jien Kato, Ching Y. Suen |
Int. J. Document Anal. Recognit. | 4 |
| 2002 | An HMM-Based Segmentation Method for Traffic Monitoring MoviesabstractShadows of moving objects often obstruct robust visual tracking. We propose an HMM-based segmentation method which classifies in real time each pixel or region into three categories: shadows, foreground, and background objects. In the case of traffic monitoring movies, the effectiveness of the proposed method has been proven through experimental results. Jien Kato, Toyohide Watanabe, Sébastien Joga, Jens Rittscher, Andrew Blake 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | A proposal for a face planeabstractChanges in facial expressions or differences in individual faces appear not only in distinctive features such as eyes, nose and mouth but also in the bone structure or the movement of facial muscles. We present a new concept that reflects the differences of the bone structure of a face and/or the movement of facial muscles and name this a face plane. The face plane can be described by a few parameters and reduce vague facial features. This paper describes the concept and the way of calculation of the face plane. Hiroyuki Hase, Masaaki Yoneda, Takefumi Kasamatsu, Jien Kato |
ICIP (1) | 4 |
| 2000 | A Probabilistic Background Model for Tracking
Jens Rittscher, Jien Kato, Sébastien Joga, Andrew Blake 0001 |
ECCV (2) | 2 |
| 1999 | Character Extraction from Noisy Background for an Automatic Reference SystemabstractIt is important to provide digitized manuscripts of old literature (in page image form) and their electronic text (in full-text form), with an automatic reference mechanism between the images and the text, on the Internet. As an essential step for creating such an automatic reference system, this paper describes the issue of extracting character areas from page images of old handwritten manuscripts. Page images of old manuscripts are usually terribly dirty and considerable large in size. To overcome the first problem, we propose a new effective method for separating characters from noisy background, since conventional threshold selection techniques are inadequate to cope with the image where the gray levels of the character parts are overlapped by that of the background. To solve the second problem, we propose an approach based on a downscaled image and a recursive labeling method for word extraction. This approach is suitable for large size images because it has the advantage of saving memory and reducing processing time. Hideyuki Negishi, Jien Kato, Hiroyuki Hase, Toyohide Watanabe |
ICDAR | 2 |
| 1998 | A model-based approach for recognizing folding process of origamiabstractThis paper describes an approach for recognition the folding process of origami by interpreting a series of graphics images. Our approach is based on a 3D internal model created during the recognition process. The output of image analysis is matched to a 2D projection of the model such that the contents of folding operations can be identified. The recognized folding operations are used to update the model for further interpretation and recognition. The effectiveness in our approach is illustrated by means of folding process "cicada". Jien Kato, Toyohide Watanabe, Takeshi Nakayama, Lishing Guo, Hitoshi Kato |
ICPR | 1 |
| 1997 | Recognition of Essential Folding Operations: A Step for Interpreting Illustrated Books of OrigamiabstractTo interpret motions by analyzing a series of illustrations is an interesting topic. This paper describes an approach for interpreting essential folding operations in illustrated books of origami. Since a folding operation such as "folding up" can be defined by a rotative surface, axis and direction, we proposed the algorithms to identify arrows and dashed lines which represent rotative surfaces, axes and directions in a typical book of origami. The arrow detection process consists of the arrow-head detection and arc detection. The former is accomplished by applying the distance transformation and border following to an input image, while the latter is composed of two phases: all paths which would correspond to arcs are found out in the thinned and approximated image first, an evaluative function is then applied to them to determine arcs. Dashed lines with different patterns are detected using a nearest-neighbor clustering method. A folding operation is finally interpreted based on the identified graphical primitives and the spatial relationships among them. Experimental results on several test images are presented. The effectiveness of our approach in generating meaningful interpretations of origami diagrams is clear from these results. Jien Kato, Toyohide Watanabe, Takeshi Nakayama |
ICDAR | 1 |