VLDB 2026 Research / reviewers in the wild / expert
Atsuhiro Takasu
dblp:35/2213
· DBLP profile ↗
131ranked-venue papers
22as first author
22since 2021 · last 2026
0000-0002-9061-7949ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 65 · 17 first-author · 11 since 2021Databases, data management, data science and information retrieval · 63 · 9 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 9 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 3 first-authorTheory of computation · 11 · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Security and privacy · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and ChartsabstractWith the growing number of submitted scientific papers, there is an increasing demand for systems that can assist reviewers in evaluating research claims. Experimental results are a core component of scientific work, often presented in varying formats such as tables or charts. Understanding how robust current multimodal large language models (multimodal LLMs) are at verifying scientific claims across different evidence formats remains an important and underexplored challenge. In this paper, we design and conduct a series of experiments to assess the ability of multimodal LLMs to verify scientific claims using both tables and charts as evidence. To enable this evaluation, we adapt two existing datasets of scientific papers by incorporating annotations and structures necessary for a multimodal claim verification task. Using this adapted dataset, we evaluate 12 multimodal LLMs and find that current models perform better with table-based evidence while struggling with chart-based evidence. We further conduct human evaluations and observe that humans maintain strong performance across both formats, unlike the models. Our analysis also reveals that smaller multimodal LLMs (under 8B) show weak correlation in performance between table-based and chart-based tasks, indicating limited cross-modal generalization. These findings highlight a critical gap in current models' multimodal reasoning capabilities. We suggest that future multimodal LLMs should place greater emphasis on improving chart understanding to better support scientific claim verification. Xanh Ho, Yun-Ang Wu, Sunisth Kumar, Florian Boudin, Atsuhiro Takasu, Akiko Aizawa |
AAAI | 5 |
| 2026 | ConvTransAKD: Multi-level Wasserstein Alignment for Efficient Knowledge Distillation in Sensor-Based Activity Recognition
Thi-Hong Vuong, Tung Doan 0001, Atsuhiro Takasu |
DaWaK | 3 |
| 2026 | DDIREC: Domain-Disentanglement on Item Representations for Cross-Domain Recommendation
Pongsakorn Jirachanchaisiri, Saranya Maneeroj, Atsuhiro Takasu |
WSDM | 3 |
| 2026 | MLConvTrans: Multi-Level Convolutional Transformer for Wearable Sensor Based Human Activity RecognitionabstractWearable sensor-based human activity recognition (HAR) methods utilize multimodal data from sensors such as accelerometers, gyroscopes, and magnetometers to infer human activities. Recent advancements employ deep learning techniques, leveraging deep neural networks to automate feature extraction for activity classification. However, existing deep hybrid networks face challenges in distinguishing between similar activities and handling complex scenarios, particularly when integrating data from diverse sensor types. To address these limitations, this study introduces a novel hybrid deep learning architecture, termed MultiLevel Convolutional Transformer (MLConvTrans), aimed at enhancing the efficiency and accuracy of sensor-based HAR systems. MLConvTrans integrates two core components: the Multi-Level Convolutional Network (MLConvNet) and the n-stacked Transformer Encoder (TransEncoder). MLConvNet captures local features from individual sensors and performs multimodal fusion to obtain comprehensive global feature representations, effectively extracting both low-level and high-level information. These features are then processed by the TransEncoder block, which integrates global features and models long-term dependencies for activity classification. The proposed architecture is extensively validated on multiple public datasets, including UCI-HAR, MotionSense, HAPT, KU-HAR, SHL2018, and PAMAP2. Experimental results demonstrate that MLConvTrans significantly outperforms state-of-the-art methods in wearable sensor-based HAR while maintaining computational efficiency. Thi-Hong Vuong, Tung Doan 0001, Atsuhiro Takasu |
Neural Process. Lett. | 3 |
| 2026 | Generalized ordered Wasserstein distance for sequential data
Tung Doan 0001, Hiep To, Thi-Hong Vuong, Muriel Visani, Atsuhiro Takasu, Khoat Than |
Pattern Recognit. | 5 |
| 2025 | RaMen: Multi-Strategy Multi-Modal Learning for Bundle ConstructionabstractExisting studies on bundle construction have relied merely on user feedback via bipartite graphs or enhanced item representations using semantic information. These approaches fail to capture elaborate relations hidden in real-world bundle structures, resulting in suboptimal bundle representations. To overcome this limitation, we propose RaMen, a novel method that provides a holistic multi-strategy approach for bundle construction. RaMen utilizes both intrinsic (characteristics) and extrinsic (collaborative signals) information to model bundle structures through Explicit Strategy-aware Learning (ESL) and Implicit Strategy-aware Learning (ISL). ESL employs task-specific attention mechanisms to encode multi-modal data and direct collaborative relations between items, thereby explicitly capturing essential bundle features. Moreover, ISL computes hyperedge dependencies and hypergraph message passing to uncover shared latent intents among groups of items. Integrating diverse strategies enables RaMen to learn more comprehensive and robust bundle representations. Meanwhile, Multi-strategy Alignment & Discrimination module is employed to facilitate knowledge transfer between learning strategies and ensure discrimination between items/bundles. Extensive experiments demonstrate the effectiveness of RaMen over state-of-the-art models on various domains, justifying valuable insights into complex item set problems. Huy-Son Nguyen, Duc-Hoang Pham, Duc-Trong Le, Hoang-Quynh Le, Padipat Sitkrongwong, Atsuhiro Takasu, Masoud Mansoury |
ECAI | 7 |
| 2025 | TRH2TQA: Table Recognition with Hierarchical Relationships to Table Question-Answering on Business Table ImagesabstractDespite advancements in visual question answering, challenges persist with documents like financial reports, often structured in complicated tabular structures with complex numerical computations. An alternative approach, the pipeline-driven methodology, includes table recognition (TR) and table question-answering (TQA). Recent advancements in TR support this approach with better accuracy and interpretability. However, real-world tables usually represent hierarchical tables. They pose additional challenges due to merged cells and indents, necessitating a specific approach for hierarchical relationship extraction. In this paper, we propose TRH2TQA (Table Recognition with Hierarchical Relationships to Table Question-Answering) for business table images. It consists of three modules on table images with question-answer pairs. First, the TR module extracts structure and textual content from table images into HTML format. Second, post-structure extraction is applied to identify header and hierarchical relationships using predicted column span and bounding box. Finally, this information is combined with natural language questions in the TQA module to generate the answer through the decoder. In extensive experiments, TRH2TQA outperforms in questionanswering performance on the VQAonBD 2023 dataset. Pongsakorn Jirachanchaisiri, Nam Tuan Ly, Atsuhiro Takasu |
WACV | 3 |
| 2025 | COLANet: Cross-Domain Recommender Systems with Latent Overlapping Items on Graph Neural NetworksabstractCross-domain recommender systems (CDRSs) enhance recommendations by transferring knowledge of overlapping users across two domains. Deep canonical correlation analysis (DCCA) shows promising results in CDRSs by maximizing correlations between representations of overlapping users, enabling cross-domain knowledge transfer that depends on the degree of relationship between domains. As a result, DCCA selectively shares only relevant knowledge, alleviating the problem of noisy representation found in traditional CDRSs, where they transfer knowledge regardless of the correlation strength between domains. Although DCCA is used for user transfer, item transfer, referring to the transfer of explicit knowledge of the same items between domains, is impossible due to the absence of overlapping items to facilitate direct knowledge transfer. Meanwhile, graph neural networks (GNNs) embed users and items from separate user and item graphs in each domain. Therefore, better representations are obtained from captured complex relationships and collaborative signals. To construct graphs of overlapping items, latent linkages among items between domains could be discovered by the neural topic model (NTM), forming new graphs representing the latent relationships. Therefore, COLANet, a GNN-based CDRS, is proposed to solve the DCCA limitation on item transfer by proposing the extraction of item representations that do not exist in another domain using latent characteristics. First, user-user graphs are constructed using user similarity, and the item-topic graph is constructed using latent topics learned from item descriptions with NTM. Hence, user and item graphs of each domain are constructed separately, preventing domain relationship misalignment. Second, these graphs are fed to GNN to obtain user and item representations. Third, these representations are fed to DCCA to transfer knowledge between user-user and item-item. Finally, correlated user and item representations of each domain are used to predict ratings. The experiments demonstrate that COLANet outperforms the baselines across four pairs of domains, including both similar and different domains. Pongsakorn Jirachanchaisiri, Saranya Maneeroj, Atsuhiro Takasu |
ACM Trans. Knowl. Discov. Data | 3 |
| 2025 | QUADEN: Discovering Latent Neighbors for Sparse Users and Items across Interaction Quadrants in Recommender SystemabstractMany recommender systems leverage Graph Neural Networks to capture user–item relations for delivering recommendations. However, the representation of nodes heavily relies on neighbors, causing limited neighbor nodes (sparse nodes) to lack expressive representation. Existing works discovered latent neighbors for sparse nodes but ignored node sparsity, resulting in the neighbor misallocation problem due to overlooking quality and quantity aspects. We propose discovering high-quality latent neighbors by progressively transferring knowledge from dense nodes by categorizing user–item interactions into four quadrants (dense user–dense item, dense user–sparse item, sparse user–dense item, and sparse user–sparse item). We leverage the node sparsity to determine the optimal quantity of latent neighbors. We propose a Domain Adaptation Network for transferring knowledge from dense to sparse quadrants without encountering the domain misalignment problem arising from the distinct representations between dense and sparse quadrants. An Enrichment Network is proposed to address the inexpressive representation problem due to limited observed interactions by enriching the sparse node representation. A Heterogeneous Graph Neural Network architecture is proposed to capture multiple relations between dense/sparse users and items. Experimental results on three benchmark datasets demonstrate the superiority of the proposed method over Graph Neural Network baselines, both with and without latent neighbors. Nakarin Sritrakool, Saranya Maneeroj, Atsuhiro Takasu |
ACM Trans. Inf. Syst. | 3 |
| 2024 | A Supervised Contrastive Learning Framework for Aspect-Based Recommendations
Padipat Sitkrongwong, Atsuhiro Takasu |
IDEAS | 2 |
| 2024 | Partial ordered Wasserstein distance for sequential data
Tung Doan 0001, Tuan Phan, Phu Nguyen, Khoat Than, Muriel Visani, Atsuhiro Takasu |
Neurocomputing | 6 |
| 2024 | Controllable Syllable-Level Lyrics Generation From Melody With Prior AttentionabstractMelody-to-lyrics generation, which is based on syllable-level generation, is an intriguing and challenging topic in the interdisciplinary field of music, multimedia, and machine learning. Many previous research projects generate word-level lyrics sequences due to the lack of alignments between syllables and musical notes. Moreover, controllable lyrics generation from melody is also less explored but important for facilitating humans to generate diverse desired lyrics. In this work, we propose a controllable melody-to-lyrics model that is able to generate syllable-level lyrics with user-desired rhythm. An explicit n-gram (EXPLING) loss is proposed to train the Transformer-based model to capture the sequence dependency and alignment relationship between melody and lyrics and predict the lyrics sequences at the syllable level. A prior attention mechanism is proposed to enhance the controllability and diversity of lyrics generation. Experiments and evaluation metrics verified that our proposed model has the ability to generate higher-quality lyrics than previous methods and the feasibility of interacting with users for controllable and diverse lyrics generation. We believe this work provides valuable insights into human-centered AI research in music generation tasks. The source codes for this work will be made publicly available for further reference and exploration. Zhe Zhang 0050, Yi Yu 0001, Atsuhiro Takasu |
IEEE Trans. Multim. | 3 |
| 2023 | An End-to-End Local Attention Based Model for Table Recognition
Nam Tuan Ly, Atsuhiro Takasu |
ICDAR (2) | 2 |
| 2023 | Rethinking Image-Based Table Recognition Using Weakly Supervised MethodsabstractMost of the previous methods for table recognition rely on training datasets containing many richly annotated table images. Detailed table image annotation, e.g., cell or text bounding box annotation, however, is costly and often subjective. In this paper, we propose a weakly supervised model named WSTabNet for table recognition that relies only on HTML (or LaTeX) code-level annotations of table images. The proposed model consists of three main parts: an encoder for feature extraction, a structure decoder for generating table structure, and a cell decoder for predicting the content of each cell in the table. Our system is trained end-to-end by stochastic gradient descent algorithms, requiring only table images and their ground-truth HTML (or LaTeX) representations. To facilitate table recognition with deep learning, we create and release WikiTableSet, the largest publicly available image-based table recognition dataset built from Wikipedia. WikiTableSet contains nearly 4 million English table images, 590K Japanese table images, and 640k French table images with corresponding HTML representation and cell bounding boxes. The extensive experiments on WikiTableSet and two large-scale datasets: FinTabNet and PubTabNet demonstrate that the proposed weakly supervised model achieves better, or similar accuracies compared to the state-of-the-art models on all benchmark datasets. Nam Tuan Ly, Atsuhiro Takasu, Phuc Nguyen 0001, Hideaki Takeda 0001 |
ICPRAM | 2 |
| 2023 | Controllable lyrics-to-melody generation
Zhe Zhang 0050, Yi Yu 0001, Atsuhiro Takasu |
Neural Comput. Appl. | 3 |
| 2022 | Two-Dimensional Encoding Method for Neural Synthesis of Tabular Transformation by Example
Yoshifumi Ujibashi, Atsuhiro Takasu |
ICANN (4) | 2 |
| 2022 | Table-structure Recognition Method Consisting of Plural Neural Network Modules
Hiroyuki Aoyagi, Teruhito Kanazawa, Atsuhiro Takasu, Fumito Uwano, Manabu Ohta |
ICPRAM | 3 |
| 2022 | MEIM: Multi-partition Embedding Interaction Beyond Block Term Format for Efficient and Expressive Link PredictionabstractKnowledge graph embedding aims to predict the missing relations between entities in knowledge graphs. Tensor-decomposition-based models, such as ComplEx, provide a good trade-off between efficiency and expressiveness, that is crucial because of the large size of real world knowledge graphs. The recent multi-partition embedding interaction (MEI) model subsumes these models by using the block term tensor format and provides a systematic solution for the trade-off. However, MEI has several drawbacks, some of which carried from its subsumed tensor-decomposition-based models. In this paper, we address these drawbacks and introduce the Multi-partition Embedding Interaction iMproved beyond block term format (MEIM) model, with independent core tensor for ensemble effects and soft orthogonality for max-rank mapping, in addition to multi-partition embedding. MEIM improves expressiveness while still being highly efficient, helping it to outperform strong baselines and achieve state-of-the-art results on difficult link prediction benchmarks using fairly small embedding sizes. The source code is released at https://github.com/tranhungnghiep/MEIM. Atsuhiro Takasu |
IJCAI | 2 |
| 2021 | Table-structure recognition method using neural networks for implicit ruled line estimation and cell estimationabstractTables are often used to summarize accurate values in academic papers, while graphs are used to show them visually. Automatic graph generation from a table is therefore a topic of research interest. Given that the way tables are written varies depending on the author, in earlier work we proposed a cell-detection-based table-structure recognition method. Our method achieved fair performance in experiments using the ICDAR 2013 table competition dataset, but could not outperform the top-ranked participant in the competition. This paper proposes an improved method using two neural networks: one estimates implicit ruled lines that are necessary to separate cells but are undrawn, and the other estimates cells by merging detected tokens in a table. We demonstrated the effectiveness of the proposed method by experiments using the same ICDAR 2013 dataset. It achieved an F-measure of 0.955, thereby outperforming the other methods including the top-ranked participant. Manabu Ohta, Ryoya Yamada, Teruhito Kanazawa, Atsuhiro Takasu |
DocEng | 4 |
| 2021 | Fully-Neural Approach to Vehicle Weighing and Strain Prediction on Bridges Using Wireless AccelerometersabstractBridge weigh-in-motion (BWIM) is a technique of estimating vehicle loads on bridges and can be used to assess a bridge’s structural fatigue and therefore its life. BWIM can be realized by analyzing the bridge deflection in terms of its response to moving axle loads. To obtain accurate load estimates, current BWIM systems require strain sensors, whose (re-) installation costs have limited their application. In this paper, we propose a new BWIM approach based on a deep neural network using accelerometers, which are easier to install than strain sensors, thus helping the advancement of low-cost BWIM systems. By learning the bridge dynamism, our model estimates axle loads successfully from the noisy acceleration signals sampled on a real bridge in service. Takaya Kawakatsu, Kenro Aihara, Atsuhiro Takasu, Jun Adachi, Tomonori Nagayama |
ICASSP | 3 |
| 2021 | Attentive Hybrid Collaborative Filtering for Rating Conversion in Recommender Systems
Phannakan Tengkiattrakul, Saranya Maneeroj, Atsuhiro Takasu |
ICWE | 3 |
| 2021 | New and improved algorithms for unordered tree inclusion
Tatsuya Akutsu, Jesper Jansson 0001, Ruiming Li, Atsuhiro Takasu, Takeyuki Tamura |
Theor. Comput. Sci. | 4 |
| 2020 | Multi-Partition Embedding Interaction with Block Term Format for Knowledge Graph CompletionabstractKnowledge graph completion is an important task that aims to predict the missing relational link between entities. Knowledge graph embedding methods perform this task by representing entities and relations as embedding vectors and modeling their interactions to compute the matching score of each triple. Previous work has usually treated each embedding as a whole and has modeled the interactions between these whole embeddings, potentially making the model excessively expensive or requiring specially designed interaction mechanisms. In this work, we propose the multi-partition embedding interaction (MEI) model with block term format to systematically address this problem. MEI divides each embedding into a multi-partition vector to efficiently restrict the interactions. Each local interaction is modeled with the Tucker tensor format and the full interaction is modeled with the block term tensor format, enabling MEI to control the trade-off between expressiveness and computational cost, learn the interaction mechanisms from data automatically, and achieve state-of-the-art performance on the link prediction task. In addition, we theoretically study the parameter efficiency problem and derive a simple empirically verified criterion for optimal parameter trade-off. We also apply the framework of MEI to provide a new generalized explanation for several specially designed interaction mechanisms in previous models. Atsuhiro Takasu |
ECAI | 2 |
| 2020 | Fully-Neural Approach to Heavy Vehicle Detection on Bridges Using a Single Strain SensorabstractBridge weigh-in-motion (BWIM) is a technique for detecting heavy vehicles that may cause serious damage to real bridges. BWIM is realized by analyzing the strain signals observed at places on the bridge in terms of bridge-component responses to the axle loads. In current practice, a BWIM system requires multiple strain sensors to collect vehicle properties including speed and axle positions for accurate load estimation, which may limit the system's life-span. Furthermore, BWIM should consider a wide variety of waveforms, which may be caused by vehicle acceleration and/or the various traveling positions in lanes. In this paper, we propose a novel BWIM mechanism, which employs a deep convolutional neural network (CNN). The CNN is able to learn actual traffic conditions and achieve accurate load estimation by using only a single strain sensor. The training dataset is collected from a distant load meter, by consulting traffic surveillance cameras and identifying similar vehicles. After the system initialization, the CNN requires no additional sensors (or cameras) for axle detection, which may reduce the costs of both installation and system maintenance. Takaya Kawakatsu, Kenro Aihara, Atsuhiro Takasu, Jun Adachi |
ICASSP | 3 |
| 2019 | A Cell-detection-based Table-structure Recognition MethodabstractIf tables are automatically recognized to extract the numerical values in them, digital documents containing such tables can be augmented with graphs generated using the recognized tables. In this paper, we propose a cell-detection-based table-structure recognition method for such automatic graph generation from tables. In detecting cells in a table, ruled lines are crucial but do not necessarily surround all cells. We therefore propose a method to detect cells by estimating implicit ruled lines, where necessary, to recognize the table structure. We demonstrate the effectiveness of the proposed method by experiments using the ICDAR 2013 table competition dataset. Manabu Ohta, Ryoya Yamada, Teruhito Kanazawa, Atsuhiro Takasu |
DocEng | 4 |
| 2019 | Exploring Scholarly Data by Semantic Query on Knowledge Graph Embedding Space
Atsuhiro Takasu |
TPDL | 2 |
| 2019 | Sparse Regression-Based Multiple Sequence AlignmentabstractAligning two time series can be performed effectively using a widely used technique - dynamic time warping. However, as more and more data are collected, alignment is simultaneously needed for multiple sequences in many applications. Unfortunately, constructing accurate multiple alignment is a computationally intense and complex task because the problem is well known to be NP-hard. In this paper, we develop a relaxed formulation for the multiple alignment problem based on a regression-coding framework. We then propose an efficient algorithm that employs sparse approximation to reduce computational cost for the relaxation. The goodness of our approach is theoretically analyzed and empirically evaluated on both synthetic and real-world datasets. Tung Doan 0001, Atsuhiro Takasu |
ICME | 2 |
| 2019 | Adversarial Media-fusion Approach to Strain Prediction for Bridges
Takaya Kawakatsu, Kenro Aihara, Atsuhiro Takasu, Jun Adachi |
ICPRAM | 3 |
| 2019 | Unsupervised context extraction via region embedding for context-aware recommendationsabstractMany context-aware recommendation methods extract contexts from reviews using supervised methods. However, this requires the optimal values for contexts to be predefined, which is not a trivial task. Although some approaches have avoided this by utilizing unsupervised methods, the extracted contexts have been limited to a unigram format. Moreover, most methods consider only the influence of context on the entire dataset, ignoring the fact that context might be relevant to individual users or items unequally. This work proposes a novel unsupervised context extraction method that uses predictive models for future ratings. Unlike previous work, we extract context from reviews automatically in the form of skip-grams by applying a region embedding technique. The predictive models utilize the interaction between contexts and users (and items) to model their influence on ratings. Experiments demonstrate that our models can outperform existing review-based recommendations that ignore contexts. Padipat Sitkrongwong, Atsuhiro Takasu |
IDEAS | 2 |
| 2019 | Deep Multi-view Learning from Sequential Data without CorrespondenceabstractMulti-view representation learning has become an active research topic in machine learning and data mining. One underlying assumption of the conventional methods is that training data of the views must be equal in size and sample-wise matching. However, in many real-world applications, such as video analysis, text streaming, and signal processing, data for the views often come in the form of sequences, that are different in length and misaligned, resulting in the failure of directly applying existing methods to such problems. In this paper, we first introduce a novel deep multi-view model that can implicitly discover sample correspondence while learning the representation. It can be shown that our method generalizes deep canonical correlation analysis - a popular multi-view learning method. We then extend our model by integrating the objective function with the reconstruction losses of autoencoders, forming a new variant of the proposed model. Extensively experimental results demonstrate the superior performances of our models over competing methods. Tung Doan 0001, Atsuhiro Takasu |
IJCNN | 2 |
| 2019 | Translation-based Embedding Model for Rating Conversion in Recommender SystemsabstractRatings, which are explicit feedback, are the most popular form that is often used in Recommender System (RSs). However, using the actual ratings from neighbors to predict ratings of target user toward target item often leads to low accuracy prediction due to the improper rating range problem. Rating conversion methods are proposed to solve this problem over the past few years. To propose rating conversion method, each user’s preference or rating pattern is needed. Some studies adopt the idea from translation-based embedding model and represent user’s preference in graph form. Although some studies represent users, items, and relations in embedding vector form, their representation may be improper and inaccurate if the rating pattern of each user is not in the same range. These vectors still suffer from the improper rating range as well. In this work, we propose a translation-based embedding model with rating conversion in RSs. We aim to solve the improper rating range problem in translation-based embedding model. Our challenges are 1) representing the relation (rating) between a pair of user and item in vector form, instead of scalar form and 2) dealing with rating conversion of user’s rating in vector form. The FilmTrust and MovieLens dataset are used in experiments comparing the proposed method with the existing methods. The evaluation showed that the proposed rating conversion method provides better accuracy results in term of both rating prediction and ranking recommendation. Phannakan Tengkiattrakul, Saranya Maneeroj, Atsuhiro Takasu |
WI | 3 |
| 2018 | A Fast Algorithm for Posterior Inference with Latent Dirichlet Allocation
Bui Thi-Thanh-Xuan, Vu Van-Tu, Atsuhiro Takasu, Khoat Than |
ACIIDS (2) | 3 |
| 2018 | Adversarial Learning for Topic Models
Tomonari Masada, Atsuhiro Takasu |
ADMA | 2 |
| 2018 | Adversarial Spiral Learning Approach to Strain Analysis for Bridge Damage Detection
Takaya Kawakatsu, Akira Kinoshita, Kenro Aihara, Atsuhiro Takasu, Jun Adachi |
DaWaK | 4 |
| 2018 | Parallelizing top-k frequent spatiotemporal terms computation on key-value storesabstractWe study an efficiently distributed index structure and parallel processing methods for a top-k frequent spatiotemporal terms query, a basic analytic query on geo-tagged social data. The key challenge is to improve query performance on huge geo-tagged social datasets with minimum storage requirement while guaranteeing query accuracy. We propose a method of distributing the data across a cluster and parallel algorithms to process the aggregating computation. Experimental results on real datasets showed improvement with our proposed methods in both space requirement and query performance compared with baselines. Atsuhiro Takasu |
SIGSPATIAL/GIS | 2 |
| 2018 | Knowledge Graph Embedding via Entities' Type Mapping Matrix
Atsuhiro Takasu |
ICONIP (3) | 2 |
| 2018 | Supervised Deep Polylingual Topic Modeling for Scholarly Information Recommendations
Pannawit Samatthiyadikun, Atsuhiro Takasu |
ICPRAM | 2 |
| 2018 | NPE: Neural Personalized Embedding for Collaborative FilteringabstractMatrix factorization is one of the most efficient approaches in recommender systems. However, such algorithms, which rely on the interactions between users and items, perform poorly for "cold-users" (users with little history of such interactions) and at capturing the relationships between closely related items. To address these problems, we propose a neural personalized embedding (NPE) model, which improves the recommendation performance for cold-users and can learn effective representations of items. It models a user's click to an item in two terms: the personal preference of the user for the item, and the relationships between this item and other items clicked by the user. We show that NPE outperforms competing methods for top-N recommendations, specially for cold-user recommendations. We also performed a qualitative analysis that shows the effectiveness of the representations learned by the model. ThaiBinh Nguyen, Atsuhiro Takasu |
IJCAI | 2 |
| 2018 | New and Improved Algorithms for Unordered Tree InclusionabstractThe tree inclusion problem is, given two node-labeled trees P and T (the "pattern tree" and the "text tree"), to locate every minimal subtree in T (if any) that can be obtained by applying a sequence of node insertion operations to P. Although the ordered tree inclusion problem is solvable in polynomial time, the unordered tree inclusion problem is NP-hard. The currently fastest algorithm for the latter is from 1995 and runs in O(poly(m,n) * 2^{2d}) = O^*(2^{2d}) time, where m and n are the sizes of the pattern and text trees, respectively, and d is the maximum outdegree of the pattern tree. Here, we develop a new algorithm that improves the exponent 2d to d by considering a particular type of ancestor-descendant relationships and applying dynamic programming, thus reducing the time complexity to O^*(2^d). We then study restricted variants of the unordered tree inclusion problem where the number of occurrences of different node labels and/or the input trees' heights are bounded. We show that although the problem remains NP-hard in many such cases, it can be solved in polynomial time for c = 2 and in O^*(1.8^d) time for c = 3 if the leaves of P are distinctly labeled and each label occurs at most c times in T. We also present a randomized O^*(1.883^d)-time algorithm for the case that the heights of P and T are one and two, respectively. Tatsuya Akutsu, Jesper Jansson 0001, Ruiming Li, Atsuhiro Takasu, Takeyuki Tamura |
ISAAC | 4 |
| 2018 | An approach to estimating cited sentences in academic papers using Doc2vecabstractMost academic authors refer to the literature when introducing their proposed methods and the data used in their experiments. These references can be very helpful when trying to understand a paper; however, some authors do not always state clearly the specific part of the referenced work they are referring the reader to and it can be quite laborintensive to have to read the whole document to identify the relevant information. In this paper, we propose a method for estimating the appropriate parts of a referenced work as the "cited parts," with the aim of reducing this burden. We first extract sentences in an academic paper that cites references to the literature as "citing sentences." We then vectorize the citing sentences and all the sentences in the cited papers using doc2vec and estimate the most appropriate cited part as the sentence that has the most similar feature vector to that of the citing sentence. To evaluate the proposed method, we conducted experiments using English-language papers and a questionnaire survey that asked subjects to evaluate the appropriateness of the cited parts estimated by the method. The experiments showed that this approach's success in estimating the appropriate parts of a cited paper as the cited parts depended on the citation intention of the citing sentences. Shunsuke Tanabe, Manabu Ohta, Atsuhiro Takasu, Jun Adachi |
MEDES | 3 |
| 2018 | An automatic graph generation method for scholarly papers based on table structure analysisabstractFor papers reporting on experiments, the types of experiment performed are important, as are the results they generate. Tables are frequently used to show experimental results and are suited to understanding accurate numerical values. However, they are not suited to visually comparing or observing changes in many values at once. In this paper, we propose an automatic graph-generation method for data presented in tabular form. We first convert PDF files of scholarly papers into XML files to analyze the structure of their tables. We then generate graphs based on this table-structure analysis. We also examine the types of graphs that are easy to understand visually. Ryoya Yamada, Manabu Ohta, Atsuhiro Takasu |
MEDES | 3 |
| 2018 | Mini-Batch Variational Inference for Time-Aware Topic Modeling
Tomonari Masada, Atsuhiro Takasu |
PRICAI | 2 |
| 2017 | A Hierarchical Bayesian Factorization Model for Implicit and Explicit Feedback Data
ThaiBinh Nguyen, Atsuhiro Takasu |
ADMA | 2 |
| 2017 | A Unified Approach for Learning Expertise and Authority in Digital Libraries
Baptiste de La Robertie, Liana Ermakova, Yoann Pitarch, Atsuhiro Takasu, Olivier Teste |
DASFAA (2) | 4 |
| 2017 | A Probabilistic Model for the Cold-Start Problem in Rating Prediction Using Click Data
ThaiBinh Nguyen, Atsuhiro Takasu |
ICONIP (5) | 2 |
| 2017 | Collaborative Item Embedding Model for Implicit Feedback Data
ThaiBinh Nguyen, Kenro Aihara, Atsuhiro Takasu |
ICWE | 3 |
| 2017 | Robust vehicle detection from noisy acceleration signal for bridge monitoring systemsabstractThe Internet of Things (IoT) - sensors and actuators connected via internet infrastructure to computing systems - has received enormous research attention. It has found applications in nearly every field, especially, in transportation where bridge monitoring systems are well known examples. This paper proposes a vehicle detection method from acceleration signals acquired from sensors of the bridge monitoring systems. The accelerometers based on micro-electro-mechanical system (MEMs) technology, are sensitive to environmental disturbances such as shock or temperature changes, so that the generated signal is contaminated by various types of noise. Our method first performs noise reduction, retaining only signal components that represent real measurements of forces induced by passing vehicles. The signal is then processed in the second phase using a wavelet-based pattern matching technique. Experimental results on real-world data demonstrate the high performance of both phases of the proposed method. Tung Doan 0001, Atsuhiro Takasu |
iiWAS | 2 |
| 2017 | Entity oriented action recommendations for actionable knowledge graph generationabstractPopular search engines have recently utilized the power of knowledge graphs (KGs) to provide specific answers to queries in a direct way. Search engine result pages (SERPs) are expected to provide facts in response to queries that satisfy semantic meaning. This encourages researchers to propose more influential knowledge graph generation techniques. To achieve and advance the technologies related to actionable knowledge graph presentation, creating action recommendations (ARs) is an essential step and a relatively new research direction to nurture research on generating KGs that are optimized for facilitating an entity's actions. An action represents the physical or mental activity of an entity. For example, for the entity "Donald J. Trump", typical potential actions could be "won the US presidential election" or "targets US journalists". In this paper, we describe the generation of relevant action recommendations based on entity instance and entity type. We propose two models that employ different approaches. Our first model exploits semisupervised learning and we introduce entity context vector (ECV) as an entity's distinguishing features for capturing the context of entities to reveal the similarity between entities, grounded on the prominent word2vec model. The second model is a probabilistic approach based on the Naive Bayes Theorem. We extensively evaluate our proposed models. Our first model significantly outperforms probabilistic and supervised learning-based models. Atsuhiro Takasu |
WI | 2 |
| 2017 | On the parameterized complexity of associative and commutative unificationabstractThis article studies the parameterized complexity of the unification problem with associative, commutative, or associative-commutative functions with respect to the parameter “number of variables”. It is shown that if every variable occurs only once then both of the associative and associative-commutative unification problems can be solved in polynomial time, but that in the general case, both problems are W[1]-hard even when one of the two input terms is variable-free. For commutative unification, an algorithm whose time complexity depends exponentially on the number of variables is presented; moreover, if a certain conjecture is true then the special case where one input term is variable-free belongs to FPT. Some related results are also derived for a natural generalization of the classic string and tree edit distance problems that allows variables. Tatsuya Akutsu, Jesper Jansson 0001, Atsuhiro Takasu, Takeyuki Tamura |
Theor. Comput. Sci. | 3 |
| 2016 | A Simple Stochastic Gradient Variational Bayes for the Correlated Topic Model
Tomonari Masada, Atsuhiro Takasu |
APWeb (2) | 2 |
| 2016 | Frequent Multi-Byte Character Subtring Extraction using a Succinct Data StructureabstractFrequent string mining is widely used in text processing to extract text features. Most researchers have focused on text using single-byte characters. Consequently, their applications have problems when applied to text represented with multibyte characters such as Japanese and Chinese text. The main drawback is huge memory us-age for treating multibyte character strings. To solve this problem,we use wavelet tree-based compressed suffix arrays instead of the normal suffix array to reduce the memory usage, and a novel technique that utilizes the rank operation to improve runtime efficiency.Our experimental evaluation shows that the proposed method reduces the processing time by 45% compared with a method usingonly compressed suffix arrays. The proposed method also reduces the memory usage by 75%. Phanucheep Chotnithi, Atsuhiro Takasu |
DocEng | 2 |
| 2016 | Important Word Organization for Support of Browsing Scholarly Papers Using Author KeywordsabstractWhen new researchers read scholarly papers, they often encounter unfamiliar technical terms, which may require considerable time to investigate. We have been developing a user interface to support the browsing of scholarly papers, which can provide useful links to information about such technical terms. The interface displays "important terms" extracted from a paper on top of the image of the paper. In this study, we organize the important terms extracted from papers by using author keywords. We first identify the important terms and then associate them with author keywords by using a method based on the word2vec model. Experiments showed that our method improved the classification accuracy of important terms compared with a simple baseline method. It associated each author keyword with about 2.5 relevant important terms. Junki Tanijiri, Manabu Ohta, Atsuhiro Takasu, Jun Adachi |
DocEng | 3 |
| 2016 | A Simple Stochastic Gradient Variational Bayes for Latent Dirichlet Allocation
Tomonari Masada, Atsuhiro Takasu |
ICCSA (4) | 2 |
| 2016 | Similar subtree search using extended tree inclusionabstractIn this paper, we have extended the concept of unordered tree inclusion to take the costs of insertions and substitutions into account. The resulting algorithm, MinCostIncl, has the same time complexity as the original algorithm of [4] for unordered tree inclusion (O(22Dmn)). Computational experiments on a large synthetic dataset as well as real datasets showed that our proposed algorithm is fast and scalable. Source codes of the implemented algorithms are available upon request. Tomoya Mori, Atsuhiro Takasu, Jesper Jansson 0001, Jaewook Hwang, Takeyuki Tamura, Tatsuya Akutsu |
ICDE | 2 |
| 2016 | GPS Trajectory Data Enrichment based on a Latent Statistical ModelabstractThis paper proposes a latent statistical model for analyzing global positioning system (GPS) trajectory data.
Because of the rapid spread of GPS-equipped devices, numerous GPS trajectories have become available,
and they are useful for various location-aware systems. To better utilize GPS data, a number of sensor data
mining techniques have been developed. This paper discusses the application of a latent statistical model
to two closely related problems, namely, moving mode estimation and interpolation of the GPS observation.
The proposed model estimates a latent mode of moving objects and represents moving patterns according to
the mode by exploiting a large GPS trajectory dataset. We evaluate the effectiveness of the model through
experiments using the GeoLife GPS Trajectories dataset and show that more than three-quarters of covered
locations were correctly reproduced by interpolation at a fine granularity. Akira Kinoshita, Atsuhiro Takasu, Kenro Aihara, Jun Ishii, Hisashi Kurasawa, Hiroshi Sato 0002, Motonori Nakamura, Jun Adachi |
ICPRAM | 2 |
| 2016 | Optimizing Short Message Text Sentiment Analysis for Mobile Device Forensics
Oluwapelumi Aboluwarin, Panagiotis Andriotis, Atsuhiro Takasu, Theodore Tryfonas |
IFIP Int. Conf. Digital Forensics | 3 |
| 2016 | Applying ant-colony concepts to trust-based recommender systemsabstractCollaborative filtering is a recommender technique that recommends items to an individual user based on the item ratings provided by similar users. However, current systems often do not acquire sufficient ratings to be able to generate recommendations. Trust-based recommender systems have been proposed that use additional trust values in generating recommendations. In this paper, we propose a trust-based ant recommender with two main improvements. First, we achieve better selection of higher-quality raters by our proposed trust-calculation method and an improved pheromone-update mechanism. Second, we can improve the prediction step by converting raters' ratings into a target user's perspective view and considering the influence level of each rater on the active user. The Epinions dataset was used in experiments comparing the proposed method with the ALT-BAR method. The evaluation showed that the proposed method provides better results in term of both accuracy and coverage. Phannakan Tengkiattrakul, Saranya Maneeroj, Atsuhiro Takasu |
iiWAS | 3 |
| 2015 | Highly Efficient Parallel Framework: A Divide-and-Conquer Approach
Takaya Kawakatsu, Akira Kinoshita, Atsuhiro Takasu, Jun Adachi |
DEXA (2) | 3 |
| 2015 | An Efficient Distributed Index for Geospatial Databases
Atsuhiro Takasu |
DEXA (1) | 2 |
| 2015 | Transfer Learning for Bibliographic Information Extraction
Quang-Hong Vuong, Atsuhiro Takasu |
ICPRAM (1) | 2 |
| 2015 | Heuristic Pretraining for Topic Models
Tomonari Masada, Atsuhiro Takasu |
IEA/AIE | 2 |
| 2015 | Bayesian probabilistic model for context-aware recommendationsabstractContext-aware recommender systems that provide better recommendations for users by using their rating history in different situations have been proposed. Because incorporating all contextual information can make the data sparser and degrade the prediction accuracy, most context-aware methods focus on detecting and using only the most effective contextual factors. However, in addition to accuracy, the diversity of the recommendation is also a key to improving users' satisfaction with recommendation results. Moreover, most context-aware techniques have not considered directly the relationships among context, users, and items before predicting the ratings. In the real world, different contextual factors tend to affect users and items differently. This paper proposes a latent probabilistic model to incorporate the contextual information. By adopting a binary particle-swarm optimization technique, the relevant contextual factors for user classes and item classes are identified and incorporated into the model. We optimize our model for two cases, namely considering accuracy alone and considering the trade-off between accuracy and diversity. An evaluation shows that our proposed model performs better than 1) a model that considers only the relation of context to users alone or items alone, 2) a model that exploits all contextual factors, and 3) the traditional context-aware recommendation method. Padipat Sitkrongwong, Saranya Maneeroj, Pannawit Samatthiyadikun, Atsuhiro Takasu |
iiWAS | 4 |
| 2015 | Real-time traffic incident detection using a probabilistic topic modelabstractTraffic congestion occurs frequently in urban settings, and is not always caused by traffic incidents. In this paper, we propose a simple method for detecting traffic incidents from probe-car data by identifying unusual events that distinguish incidents from spontaneous congestion. First, we introduce a traffic state model based on a probabilistic topic model to describe the traffic states for a variety of roads. Formulas for estimating the model parameters are derived, so that the model of usual traffic can be learned using an expectation–maximization algorithm. Next, we propose several divergence functions to evaluate differences between the current and usual traffic states and streaming algorithms that detect high-divergence segments in real time. We conducted an experiment with data collected for the entire Shuto Expressway system in Tokyo during 2010 and 2011. The results showed that our method discriminates successfully between anomalous car trajectories and the more usual, slowly moving traffic patterns. Akira Kinoshita, Atsuhiro Takasu, Jun Adachi |
Inf. Syst. | 2 |
| 2015 | On the complexity of finding a largest common subtree of bounded degree
Tatsuya Akutsu, Takeyuki Tamura, Avraham A. Melkman, Atsuhiro Takasu |
Theor. Comput. Sci. | 4 |
| 2015 | Similar Subtree Search Using Extended Tree InclusionabstractThis paper considers the problem of identifying all locations of subtrees in a large tree or in a large collection of trees that are similar to a specified pattern tree, where all trees are assumed to be rooted and node-labeled. The tree edit distance is a widely-used measure of tree (dis-)similarity, but is NP-hard to compute for unordered trees. To cope with this issue, we propose a new similarity measure which extends the concept of unordered tree inclusion by taking the costs of insertion and substitution operations on the pattern tree into account, and present an algorithm for computing it. Our algorithm has the same time complexity as the original one for unordered tree inclusion, i.e., it runs in O(|T1∥T2|) time, where T1and T2denote the pattern tree and the text tree, respectively, when the maximum outdegree of T1is bounded by a constant. Our experimental evaluation using synthetic and real datasets confirms that the proposed algorithm is fast and scalable and very useful for bibliographic matching, which is a typical entity resolution problem for tree-structured data. Furthermore, we extend our algorithm to also allow a constant number of deletion operations on T1while still running in O(|T1∥T2|) time. Tomoya Mori, Atsuhiro Takasu, Jesper Jansson 0001, Jaewook Hwang, Takeyuki Tamura, Tatsuya Akutsu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | Real-time traffic incident detection using probe-car data on the Tokyo Metropolitan ExpresswayabstractWe have developed a real-time traffic incident detection system for the Tokyo Metropolitan Expressway. This system monitors current traffic using probe-car data and compares actual traffic in real time with the usual traffic, which is estimated in advance using batch processing. Akira Kinoshita, Atsuhiro Takasu, Jun Adachi |
IEEE BigData | 2 |
| 2014 | Empirical Evaluation of CRF-Based Bibliography Extraction from Reference StringsabstractThis paper reports an empirical evaluation of a CRF-based bibliography parser we have developed for reference strings of research papers. The parser uses a conditional random field (CRF) to estimate the correct bibliographic label such as an author's name and a title for each token in a reference string. We applied the parser specifically designed for reference strings to three academic journals, an English one and two Japanese ones, published in Japan. Experiments showed (i) the parser correctly parsed from 90% to 94% of reference strings depending on the kinds of journals used and (ii) segmentation errors induced by tokenization considerably degraded the final parsing accuracies. This paper also discusses some future directions of the bibliography extraction based on a detailed analysis of the experiments. Manabu Ohta, Daiki Arauchi, Atsuhiro Takasu, Jun Adachi |
Document Analysis Systems | 3 |
| 2014 | Rule Management for Information Extraction from Title Pages of Academic PapersabstractThis paper discusses a problem of managing rules for page layout analysis and information extraction. We have been developing an information extraction system from academic papers. The system exploits both page layout and textual information. For this purpose, conditional random field (CRF) are designed according to the layout of objective pages. Since various kinds of layout are used in academic papers, we need to prepare a set of rules for each type of layout to achieve high extraction accuracy. As the number of papers in a system grows, rule management becomes a big problem. For example, when should we make a new set of rules and how can we acquire them efficiently in a process of receiving new articles? This paper examines two scores to measures the fitness of rules and availability of rules learned for another type of layout. We evaluate the scores for bibliographic information extraction from title pages of academic papers and show that they are effective for measuring the fitness. We also examine the availability for sampling training data when learning a new set of rules. Atsuhiro Takasu, Manabu Ohta |
ICPRAM | 1 |
| 2014 | A Topic Model for Traffic Speed Data Analysis
Tomonari Masada, Atsuhiro Takasu |
IEA/AIE (2) | 2 |
| 2014 | Smartphone Message Sentiment Analysis
Panagiotis Andriotis, Atsuhiro Takasu, Theodore Tryfonas |
IFIP Int. Conf. Digital Forensics | 2 |
| 2014 | On the Parameterized Complexity of Associative and Commutative Unification
Tatsuya Akutsu, Jesper Jansson 0001, Atsuhiro Takasu, Takeyuki Tamura |
IPEC | 3 |
| 2014 | ChronoSAGE: Diversifying Topic Modeling Chronologically
Tomonari Masada, Atsuhiro Takasu |
WAIM | 2 |
| 2013 | Timeline adaptation for text classificationabstractIn this paper, we address the text classification problem that a period of time created test data is different from the training data, and present a method for text classification based on temporal adaptation. We first applied lexical chains for the training data to collect terms with semantic relatedness, and created sets (we call these Sem sets). Semantically related terms in the documents are replaced to their representative term. For the results, we identified short terms that are salient for a specific period of time. Finally, we trained SVM classifiers by applying a temporal weighting function to each selected short terms within the training data, and classified test data. Temporal weighting function is weighted each short term in the training data according to the temporal distance between training and test data. The results using MedLine data showed that the method was comparable to the current state-of-the-art biased-SVM method, especially the method is effective when testing on data far from the training data. Fumiyo Fukumoto, Yoshimi Suzuki, Atsuhiro Takasu |
CIKM | 3 |
| 2013 | On the Complexity of Finding a Largest Common Subtree of Bounded Degree
Tatsuya Akutsu, Takeyuki Tamura, Avraham A. Melkman, Atsuhiro Takasu |
FCT | 4 |
| 2013 | Three-Way Nonparametric Bayesian Clustering for Handwritten Digit Image Classification
Tomonari Masada, Atsuhiro Takasu |
ICONIP (3) | 2 |
| 2013 | Page Analysis by 2D Conditional Random Fields
Atsuhiro Takasu |
ICPRAM | 1 |
| 2013 | A Revised Inference for Correlated Topic Model
Tomonari Masada, Atsuhiro Takasu |
ISNN (2) | 2 |
| 2013 | Latent Probabilistic Model for Context-Aware RecommendationsabstractRecommender systems (RS) are software tools that provide personalized recommendations of relevant items to individual users. However, most of them do not take into account additional contextual information that may affect user preferences, such as place, time, or weather. Context-aware recommender systems (CARS) have been proposed to solve this problem by providing recommendations for users based on their rating history in different situations. Although most have tried to identify the contextual variables that have the greatest effect on rating accuracy, they have not directly considered the relationships among context, users, and items before predicting the ratings. In the real world, different contextual factors tend to affect users and items differently. This work proposes a latent probabilistic model for contextual recommendation by extending the flexible mixture model to incorporate different contextual factors. This model has the flexibility to adjust the effects of contextual factors on users and items according to a variety of context-user-item relations to suit specific situations. Our evaluation has shown that the proposed model's recommendations are more accurate than those made by both latent probabilistic models and collaborative filtering-based CARS. Padipat Sitkrongwong, Saranya Maneeroj, Atsuhiro Takasu |
Web Intelligence | 3 |
| 2013 | Approximation and parameterized algorithms for common subtrees and edit distance between unordered trees
Tatsuya Akutsu, Daiji Fukagawa, Magnús M. Halldórsson, Atsuhiro Takasu, Keisuke Tanaka |
Theor. Comput. Sci. | 4 |
| 2012 | Extraction of topic evolutions from references in scientific articles and its GPU accelerationabstractThis paper provides a topic model for extracting topic evolutions as a corpus-wide transition matrix among latent topics. Recent trends in text mining point to a high demand for exploiting metadata. Especially, exploitation of reference relationships among documents induced by hyperlinking Web pages, citing scientific articles, tumblring blog posts, retweeting tweets, etc., is put in the foreground of the effort for an effective mining. We focus on scholarly activities and propose a topic model for obtaining a corpus-wide view on how research topics evolve along citation relationships. Our model, called TERESA, extends latent Dirichlet allocation (LDA) by introducing a corpus-wide topic transition probability matrix, which models reference relationships as transitions among topics. Our approximated variational inference updates LDA posteriors and topic transition posteriors alternately. The main issue is execution time amounting to O(MK2), where K is the number of topics and M is that of links in citation network. Therefore, we accelerate the inference with Nvidia CUDA compatible GPUs. We compare the effectiveness of TERESA with that of LDA by introducing a new measure called diversity plus focusedness (D+F). We also present topic evolution examples our method gives. Tomonari Masada, Atsuhiro Takasu |
CIKM | 2 |
| 2012 | Efficient Exponential Time Algorithms for Edit Distance between Unordered Trees
Tatsuya Akutsu, Takeyuki Tamura, Daiji Fukagawa, Atsuhiro Takasu |
CPM | 4 |
| 2012 | CRF-based Bibliography Extraction from Reference Strings Focusing on Various Token GranularitiesabstractThe references of academic articles include important bibliographic elements such as authors' names and article titles. Automatic extraction of these elements is useful because they can be used for various purposes, including searching. In this paper, a method for automatically extracting bibliographic elements from the text of reference strings is proposed. The proposed method assigns bibliographic labels to reference strings by using linguistic information and conditional random fields. Experimental results indicated that the extraction accuracies of major bibliographies were more than 96%. Manabu Ohta, Daiki Arauchi, Atsuhiro Takasu, Jun Adachi |
Document Analysis Systems | 3 |
| 2012 | Topic and Subject Detection in News Streams for Multi-document Summarization
Fumiyo Fukumoto, Yoshimi Suzuki, Atsuhiro Takasu |
KEOD | 3 |
| 2011 | A recommendation algorithm using positive and negative latent modelsabstractThis paper proposes an algorithm for recommender systems that uses both positive and negative latent user models. In recommending items to a user, recommender systems usually exploit item content information as well as the preferences of similar users. Various types of content information can be attached to items and these are useful for judging user preferences. For example, in movie recommendations, a movie record may include the director, the actors, and reviews. These types of information help systems calculate sophisticated user preferences. We first propose a probabilistic model that maps multi-attributed records into a low-dimensional feature space. The proposed model extends latent Dirichlet allocation to the handling of multi-attributed data. We derive an algorithm for estimating the model's parameters using the Gibbs sampling technique. Next, we propose a probabilistic model to calculate user preferences for items in the feature space. Finally, we develop a recommendation algorithm based on the probabilistic model that works efficiently for large quantities of items and user ratings. We use a publicly available movie corpus to evaluate the proposed algorithm empirically, in terms of both its recommendation accuracy and its processing efficiency. Atsuhiro Takasu, Saranya Maneeroj |
CIDM | 1 |
| 2011 | An Improved Clique-Based Method for Computing Edit Distance between Unordered Trees and Its Application to Comparison of Glycan StructuresabstractThe tree edit distance is one of the most widely used measures for comparison of tree structured data and has been used for analysis of RNA secondary structures, glycan structures, and vascular trees. However, it is known that the tree edit distance problem is NP-hard for unordered trees while it is polynomial time solvable for ordered trees. We have recently proposed a clique-based method for computing the tree edit distance between unordered trees in which each instance of the tree edit distance problem is transformed into an instance of the maximum vertex weighted clique problem and then an existing clique algorithm is applied. In this paper, we propose an improved clique-based method. Different from our previous method, the improved method is basically a dynamic programming algorithm that repeatedly solves instances of the maximum vertex weighted clique problem as sub-problems. Other heuristic techniques, which do not violate the optimality of the solution, are also introduced. When applied to comparison of large glycan structures, our improved method showed significant speed-up in most cases. Tatsuya Akutsu, Tomoya Mori, Takeyuki Tamura, Daiji Fukagawa, Atsuhiro Takasu, Etsuji Tomita |
CISIS | 5 |
| 2011 | Top-k query processing for combinatorial objects using Euclidean distanceabstractConventional search techniques are mainly designed to return a ranked list of single objects that are relevant to a given query. However, they do not meet the criteria for retrieving a combination of objects that is close to the query. This paper presents top-k query processing in which Euclidean distance is used as the scoring function for combinatorial objects. We also propose a pruning method based on clustering and efficiently select object combinations by pruning clusters that do not contain potential candidates for the top-k results. We compared the proposed method with the method that enumerates all the combinatorial objects and calculates the distance to the query. Experimental results revealed that the proposed method improves the processing efficiency to about 95% at maximum. Takanobu Suzuki, Atsuhiro Takasu, Jun Adachi |
IDEAS | 2 |
| 2011 | Steering Time-Dependent Estimation of Posteriors with Hyperparameter Indexing in Bayesian Topic Models
Tomonari Masada, Atsuhiro Takasu, Yuichiro Shibata, Kiyoshi Oguri |
PAKDD (1) | 2 |
| 2011 | A clique-based method for the edit distance between unordered trees and its application to analysis of glycan structuresabstractBACKGROUND: Measuring similarities between tree structured data is important for analysis of RNA secondary structures, phylogenetic trees, glycan structures, and vascular trees. The edit distance is one of the most widely used measures for comparison of tree structured data. However, it is known that computation of the edit distance for rooted unordered trees is NP-hard. Furthermore, there is almost no available software tool that can compute the exact edit distance for unordered trees. RESULTS: In this paper, we present a practical method for computing the edit distance between rooted unordered trees. In this method, the edit distance problem for unordered trees is transformed into the maximum clique problem and then efficient solvers for the maximum clique problem are applied. We applied the proposed method to similar structure search for glycan structures. The result suggests that our proposed method can efficiently compute the edit distance for moderate size unordered trees. It also suggests that the proposed method has the accuracy comparative to those by the edit distance for ordered trees and by an existing method for glycan search. CONCLUSIONS: The proposed method is simple but useful for computation of the edit distance between unordered trees. The object code is available upon request. Daiji Fukagawa, Takeyuki Tamura, Atsuhiro Takasu, Etsuji Tomita, Tatsuya Akutsu |
BMC Bioinform. | 3 |
| 2011 | Exact algorithms for computing the tree edit distance between unordered trees
Tatsuya Akutsu, Daiji Fukagawa, Atsuhiro Takasu, Takeyuki Tamura |
Theor. Comput. Sci. | 3 |
| 2010 | Pivot Selection Method for Optimizing both Pruning and Balancing in Metric Space Indexes
Hisashi Kurasawa, Daiji Fukagawa, Atsuhiro Takasu, Jun Adachi |
DEXA (2) | 3 |
| 2010 | A Variational Bayesian EM Algorithm for Tree SimilarityabstractIn recent times, a vast amount of tree-structured data has been generated. For mining, retrieving, and integrating such data, we need a fine-grained tree similarity measure that can be adapted to objective data. To achieve this goal, this paper (1) proposes a probabilistic generative model that generates pairs of similar trees, and (2) derives a learning algorithm for estimating the parameters of the model based on the variational Bayesian expectation maximization (VBEM) method. This method can handle rooted, ordered, and labeled trees. We show that the tree similarity model obtained via the BEM technique performs better than that obtained via maximum likelihood estimation by tuning the hyper parameters. Atsuhiro Takasu, Daiji Fukagawa, Tatsuya Akutsu |
ICPR | 1 |
| 2010 | Language Model Combination for Community-based Q & A RetrievalabstractThis paper proposes three methods for combining various probabilistic models for retrieving answers from community-based question answering (cQA) archives. We adopt four probabilistic models for these combinations, i.e., (1) the language model measuring similarity between a query and a question stored in the cQA archive, (2) two translation models for measuring the similarity between a query and an answer stored in the cQA archive, and a background language model for smoothing. Then, we developed three parameter estimation methods. Two of them are mixture models of the language models. The remaining model exploits the difference between the models. We apply the proposed methods to a cQA archive and show that they significantly outperform a widely used language model and Okapi BM25. We also show that they achieve a better performance than the recently proposed cQA retrieval method. Atsuhiro Takasu, Jun Adachi |
ICTAI (2) | 2 |
| 2010 | Modeling Topical Trends over Continuous Time with Priors
Tomonari Masada, Daiji Fukagawa, Atsuhiro Takasu, Yuichiro Shibata, Kiyoshi Oguri |
ISNN (2) | 3 |
| 2010 | Approximating Tree Edit Distance through String Edit Distance
Tatsuya Akutsu, Daiji Fukagawa, Atsuhiro Takasu |
Algorithmica | 3 |
| 2009 | Maximal metric margin partitioning for similarity search indexesabstractWe propose a partitioning scheme for similarity search indexes that is called Maximal Metric Margin Partitioning (MMMP). MMMP divides the data on the basis of its distribution pattern, especially for the boundaries of clusters. A partitioning surface created by MMMP is likely to be at maximum distances from the two cluster boundaries. MMMP is the first similarity search index approach to focus on partitioning surfaces and data distribution patterns. We also present an indexing scheme, named the MMMP-Index, which uses MMMP and small ball partitioning. The MMMP-Index prunes many objects that are not relevant to a query, and it reduces the query execution cost. Our experimental results show that MMMP effectively indexes clustered data and reduces the search cost. For clustered vector data, the MMMP-Index reduces the computational cost to less than two thirds that of comparable schemes. Hisashi Kurasawa, Daiji Fukagawa, Atsuhiro Takasu, Jun Adachi |
CIKM | 3 |
| 2009 | Dynamic hyperparameter optimization for bayesian topical trend analysisabstractThis paper presents a new Bayesian topical trend analysis. We regard the parameters of topic Dirichlet priors in latent Dirichlet allocation as a function of document timestamps and optimize the parameters by a gradient-based algorithm. Since our method gives similar hyperparameters to the documents having similar timestamps, topic assignment in collapsed Gibbs sampling is affected by timestamp similarities. We compute TFIDF-based document similarities by using a result of collapsed Gibbs sampling and evaluate our proposal by link detection task of Topic Detection and Tracking. Tomonari Masada, Daiji Fukagawa, Atsuhiro Takasu, Tsuyoshi Hamada, Yuichiro Shibata, Kiyoshi Oguri |
CIKM | 3 |
| 2009 | A Versatile Record Linkage Method by Term Matching Model Using CRF
Quang Minh Vu, Atsuhiro Takasu, Jun Adachi |
DEXA | 2 |
| 2009 | Latent Topic Extraction from Relational Table for Record Matching
Atsuhiro Takasu, Daiji Fukagawa, Tatsuya Akutsu |
Discovery Science | 1 |
| 2009 | Bayesian Similarity Model Estimation for Approximate Recognized Text SearchabstractApproximate text search is a basic technique to handle recognized text that contains recognition errors. This paper proposes an approximate string search for recognized texturing a statistical similarity model focusing on parameter estimation. The main contribution of this paper is to propose a parameter estimation algorithm using variational Bayesian expectation maximization technique. We applied the obtained model to approximate substring detection problem and experimentally showed that the Bayesian estimation is effective. Atsuhiro Takasu |
ICDAR | 1 |
| 2009 | Constant Factor Approximation of Edit Distance of Bounded Height Unordered Trees
Daiji Fukagawa, Tatsuya Akutsu, Atsuhiro Takasu |
SPIRE | 3 |
| 2008 | Information Extraction by Two Dimensional ParserabstractThis paper proposes a learning algorithm for a two dimensional parser. The parser is designed to analyze page layout of documents and extract information using both textual and layout information. The parsing rules are expressed by an extended stochastic context free grammar that decomposes tokens located in two dimensional space both horizontally and vertically. In this paper we focus on the learning aspect of the parser and propose a learning algorithm based on the expectation maximization technique where the dynamic programming (DP) technique is used for efficient process. We apply the proposed algorithm to acquire a stochastic parser for information extraction from scanned document images and show that learned stochastic grammar extracts bibliographic data with high accuracy. Atsuhiro Takasu |
ICTAI (1) | 1 |
| 2008 | Name Disambiguation Boosted by Latent Topics from Web DirectoriesabstractSearch results for personal name queries often contain documents relevant to several people as a personal name is often shared by several people. In order to differentiate people in these search results, it is required to extract contexts relevant to people in documents. However, since Web documents are noisy and the texts related to people might be short, it is difficult to extract contexts of people effectively. We propose a new method that uses web directories as additional information in order to recognize topic terms in documents more easily and to extract contexts of people more effectively. First, we apply latent Dirichlet allocation method to extract latent topics in Web directories. Then, the extracted topics are used to recognize topics contained in name ambiguity documents so that common context measurements can be calculated more effectively. Our experiments, conducted with documents of real people in the Web and several well-known Web directories, show that our approach disambiguates personal names better than some other conventional approaches like vector space model approach and named entity recognition approach. Quang Minh Vu, Atsuhiro Takasu, Jun Adachi |
Web Intelligence | 2 |
| 2008 | Improved approximation of the largest common subtree of two unordered trees of bounded height
Tatsuya Akutsu, Daiji Fukagawa, Atsuhiro Takasu |
Inf. Process. Lett. | 3 |
| 2008 | Improving the performance of personal name disambiguation using web directories
Quang Minh Vu, Atsuhiro Takasu, Jun Adachi |
Inf. Process. Manag. | 2 |
| 2007 | Statistical Learning Algorithm for Tree SimilarityabstractTree edit distance is one of the most frequently used distance measures for comparing trees. When using the tree edit distance, we need to determine the cost of each operation, but this is a labor-intensive and highly skilled task. This paper proposes an algorithm for learning the costs of tree edit operations from training data consisting of pairs of similar trees. To formalize the cost learning problem, we define a probabilistic model for tree alignment that is a variant of tree edit distance. Then, the parameters of the model are estimated using the expectation maximization (EM) technique. In this paper, we develop an algorithm for parameter learning that is polynomial in time (O{mn2d6)) and space (O{n2d4)) where n, d, and m represent the size of the trees, the maximum degree of trees, and the number of training pairs of trees, respectively. Atsuhiro Takasu, Daiji Fukagawa, Tatsuya Akutsu |
ICDM | 1 |
| 2007 | Support Kernel Machine-Based Active Learning to Find Labels and a Proper Kernel Simultaneously
Yasusi Sinohara, Atsuhiro Takasu |
IDEAL | 2 |
| 2006 | An approximate multi-word matching algorithm for robust document retrievalabstractDocument generation from low level data and its utilization is one of the most challenging tasks in document engineering. Word occurrence detection is a fundamental problem in the recognized document utilization obtained by a recognizer, such as OCR and speech recognition. Given a set of words, such as a dictionary, this paper proposes an efficient dynamic programming (DP) algorithm to find the occurrences of each word in a text. In this paper, the string similarity is measured by a statistical similarity model that enables a definition of the similarities in the character level as well as edit operation level. The proposed algorithm uses tree structures to measure similarities in order to avoid measuring similarities of the same substrings appearing in different parts of the text and words. The time complexity of the proposed algorithm is O(|W|⋅|S|⋅|Q|), where |W| (resp. |S|) denote the number of nodes in the trees representing the word set (resp. the text), and |Q| donotes the number of the states of the model used for string similarity. This paper shows the proposed algorithm is experimentally about six times faster than a naive DP algorithm. Atsuhiro Takasu |
CIKM | 1 |
| 2006 | Quality enhancement in information extraction from scanned documentsabstractWhen constructing a large document archive, an important element is the digitizing of printed documents. Although various techniques for document image analysis such as Optical Character Recognition (OCR) have been developed, error handling is required in constructing real document archive systems. This paper discusses the problem from the quality enhancement perspective and proposes a robust reference extraction method for academic articles scanned with OCR mark-up. We applied the proposed method to articles appearing in various journals, and these experiments showed that the proposed method achieved a recognition accuracy of more than 94%. This paper also discusses manual correction and investigates experimentally the relationship between extraction accuracy and cost reduction. Atsuhiro Takasu, Kenro Aihara |
ACM Symposium on Document Engineering | 1 |
| 2006 | A New Measure for Query Disambiguation Using Term Co-occurrences
Hiromi Wakaki, Tomonari Masada, Atsuhiro Takasu, Jun Adachi |
IDEAL | 3 |
| 2006 | Approximating Tree Edit Distance Through String Edit Distance
Tatsuya Akutsu, Daiji Fukagawa, Atsuhiro Takasu |
ISAAC | 3 |
| 2006 | Effect of relationships between words on Japanese information retrievalabstractTwo Japanese-language information retrieval (IR) methods that enhance retrieval effectiveness by utilizing the relationships between words are proposed. The first method uses dependency relationships between words in a sentence. The second method uses proximity relationships, particularly information about the ordered co-occurrence of words in a sentence, to approximate the dependency relationships between them. A Structured Index has been constructed for these two methods, which represents the dependency relationships between words in a sentence as a set of binary trees. The Structured Index is created by morphological analysis and dependency analysis based on simple template matching and compound noun analysis derived from word statistics. Through retrieval experiments using the Japanese test collection for information retrieval systems (NTCIR-1, the NACSIS Test Collection for IR systems), it is shown that these two methods offer superior retrieval effectiveness compared with the TF--IDF method, and are effective with different databases and diverse search topics sets. There is little difference in retrieval effectiveness between these two methods. Atsushi Matsumura, Atsuhiro Takasu, Jun Adachi |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2005 | Relation Analysis Among Patterns on Software Development Process
Hironori Washizaki, Atsuto Kubo, Atsuhiro Takasu, Yoshiaki Fukazawa |
PROFES | 3 |
| 2004 | Web Page Grouping Based on Parameterized Connectivity
Tomonari Masada, Atsuhiro Takasu, Jun Adachi |
DASFAA | 2 |
| 2004 | Feature Word Tracking in Time Series Documents
Atsuhiro Takasu, Katsuaki Tanaka |
IDEAL | 1 |
| 2003 | A Statistical Model for Flexible String Similarity
Atsuhiro Takasu |
IJCAI | 1 |
| 2001 | Document Filtering for Fast Approximate String Matching of Errorneous TextabstractIt is important to utilize retrospective documents. OCR is the most widely applied technology for this purpose; however, error-tolerant methods are essential for utilizing OCR-processed documents. This paper discusses a filtering problem for OCR-processed documents that enables the handling of large numbers of OCR-processed documents in an error-tolerant way. It proposes a systematic index design method for filtering and shows that the filtering method speeds up by about 360 times for a database consisting of about two million records, with little decrease in accuracy. Atsuhiro Takasu |
ICDAR | 1 |
| 2000 | Variance based classifier comparison in text categorizationabstractText categorization is one of the key functions for utilizing vast amount of documents. It can be seen as a classification problem, which has been studied in pattern recognition and machine learning fields for a long time and several classification methods have been developed such as statistical classification, decision tree, support vector machines and so on. Many researchers applied those classification methods to text categorization and reported their performance (e.g., decision tree[3], Bayes classifier[2], support vector machine[l]). Yang conducted comprehensive study of comparison or text categorization and reported that k nearest neighbor and support vector machines works well for text categorization[4].In the previous studies, classification methods were usually compared using single pair of training and test data However, classification method with more complex family of classifiers requires more training data and small training data may result in deriving unreliable classifier, that is, the performance of the derived classifier varies much depending on training data. Therefore, we need to take the size of training data into account when comparing and selecting a classification method. In this paper, we discuss how to select a classifier from those derived by various classification methods and how the size of training data affects the performance of the derived classifier.In order to evaluate the reliability of classification method, we consider the variance of accuracy of derived classifier. We first construct a statistical model. In the text categorization, each document is usually represented with a feature vector that consists of weighted frequencies of terms. In the vector space model, document is a point in high dimensional feature space and a classifier separates the feature space into subspaces each of which is labeled with a category. Atsuhiro Takasu, Kenro Aihara |
SIGIR | 1 |
| 1998 | On the Number of Clusters in Cluster Analysis
Atsuhiro Takasu |
Discovery Science | 1 |
| 1998 | Probabilistic interpage analysis for article extraction from document imagesabstractThe progress of information processing and utilization technologies enables one to handle a huge information space. It is important to utilize the information stored in printed papers for constructing a huge information space. This paper presents an interpage analysis method for processing one issue of a journal and a magazine at one time. The method presented is based on the error correcting parsing technique which enables one to handle errors of recognition processes. The sequence of pages are parsed in two levels, the journal level and the logical component level, to extract logical components of journals and magazines. Atsuhiro Takasu |
ICPR | 1 |
| 1997 | Retrieval methods for English-text with missrecognized OCR charactersabstractThis paper presents three probabilistic text retrieval methods designed to carry out a full-text search of English documents containing OCR errors. By searching for any query term on the premise that there are errors in the recognized text, the methods presented can tolerate such errors, and therefore costly manual post-editing is not required after OCR recognition. In the applied approach, confusion matrices are used to store characters which are likely to be interchanged when a particular character is missrecognized, and the respective probability of each occurrence. Moreover, a 2-gram matrix is used to store probabilities of character connection, i.e., which letter is likely to come after another. Multiple search terms are generated for an input query term by making reference to confusion matrices, after which a full-text search is run for each search term. The validity of retrieved terms is determined based on error-occurrence and character connection probabilities. The performance of these methods is experimentally evaluated by determining retrieval effectiveness, i.e., by calculating recall and precision rates. Results indicate marked improvement in comparison with exact matching. Manabu Ohta, Atsuhiro Takasu, Jun Adachi |
ICDAR | 2 |
| 1997 | An Approximate String Match for Garbled Text with Various AccuracyabstractThis paper presents a fast approximate string matching method. In constructing information spaces such as digital libraries, we have to collect vast amount of information and convert it into uniformly organized data. Since much of the information must be converted from various media automatically, the space contains garbled text with various accuracy. For utilizing these texts, we need to satisfy the three requirements, i.e., high recall, high precision and fast matching process. In order to satisfy these requirements, we have been developing a two-phase matching system. The presented method is used for fast and high recall candidate word selection in the first phase. The key idea of the method is to use a portion of characters of a word and a distance pattern in order to use current index techniques. By experiments, we confirm that the presented method achieves high recall even for the poorly recognized texts. Atsuhiro Takasu |
ICDAR | 1 |
| 1997 | COSPEX: A System for Constructing Private Digital Libraries
Masanori Sugimoto, Norio Katayama, Atsuhiro Takasu |
IJCAI (1) | 3 |
| 1996 | A Query Processing Method for Integrated Access to Multiple Databases
Itaru Nishizawa, Atsuhiro Takasu, Jun Adachi |
DEXA | 2 |
| 1996 | Approximate matching for OCR-processed bibliographic dataabstractThis paper presents a method for matching bibliographies in references of academic papers obtained as document images with records of bibliographic databases. The main subject of this paper is to handle the erroneous bibliographic data obtained by a document understanding methodology. The presented method can find a candidate record set from referral databases in spite of the errors of string by means of approximate matching which is performed as an exact matching of k substrings of length m chosen from the strings of bibliographic data in references and in databases. For the accuracy /spl alpha/ of the OCR, theoretical observation shows that the accuracy of the presented method is 1-(1-/spl alpha//sup m/)/sup k/ under the assumption that the OCR error occurs randomly and independently in the string. The method is applied to references of 187 Japanese articles and achieves accuracy of 94.05%. Atsuhiro Takasu, Norio Katayama, Masaki Yamaoka, Osamu Iwaki, Keizo Oyama, Jun Adachi |
ICPR | 1 |
| 1996 | Sampling Effectiveness in Discovering Functional Relationships in Databases
Atsuhiro Takasu, Tatsuya Akutsu, Moonis Ali |
IEA/AIE | 1 |
| 1995 | An automated generation of an electronic library based on document image understandingabstractThe article describes a framework for electronic library systems incorporating automated generation of electronic library schema from paper printed materials, offering a hypertext style browsing interface. On our framework, document image understanding techniques are applied to table-of-content images of academic journals to automatically acquire bibliographic database schema and hypertext schema. The system should only provide typical hypertext links as implicit links in the form of "functions." Although the basic framework of automatic generation of hypertext schema has been explained before, we describe an overview of an experimental electronic library system CyberMagazine based on this framework to achieve basic electronic library facilities. Shin'ichi Satoh 0001, Atsuhiro Takasu, Eishi Katsura |
ICDAR | 2 |
| 1995 | A rule learning method for academic document image processingabstractA syntactic rule learning method is presented for analyzing document images and constructing a database from them. This method is used in a digital library system named CyberMagazine, where document images are sequentially converted into database tuples by block segmentation, rough classification, and syntactic analysis. The syntactic rule has an ability to analyze symbols located in two dimensional plane, and has a syntax similar to an ordinal context free grammar except for the concatenation of symbols. In the presented learning method, the syntactic rules are generated from a set of parse trees by decomposing the trees according to non terminal symbols, generalizing the decomposed trees to a syntactic rule, and merging them. Atsuhiro Takasu, Shin'ichi Satoh 0001, Eishi Katsura |
ICDAR | 1 |
| 1995 | A Rule Discovery Method Based on Approximate Dependency Inference
Atsuhiro Takasu, Moonis Ali |
IEA/AIE | 1 |
| 1994 | A document understanding method for database construction of an electronic libraryabstractA new method is presented for analyzing document images and constructing a database from them. This method is implemented in an electronic library system named CyberMagazine, where document images are sequentially converted into database tuples by block segmentation, rough classification, and syntactic analysis. CyberMagazine's image understanding method combines the use of decision tree classification and syntactic analysis using a newly presented matrix grammar. In order to combine the classification and the syntactic analysis steps, the authors propose an augmentation technique for carrying out decision tree classification. Experimental results subsequently demonstrate that high understanding accuracy is obtained by adequate augmentation. Atsuhiro Takasu, Shin'ichi Satoh 0001, Eishi Katsura |
ICPR (2) | 1 |
| 1994 | A Database with an Explicit Semantic Representation
Norio Katayama, Atsuhiro Takasu, Jun Adachi |
IEA/AIE | 2 |
| 1993 | A collaborative supporting method between document processing and hypertext constructionabstractA new method of collaborative unification between document image understanding and hypertext construction is presented. Document image understanding is indispensable to electronic library systems, but document understanding technologies are still immature. Moreover, hypertext links are difficult to acquire by hand. In the approach presented, document image understanding is taken as classification of text blocks to classes of bibliographic items which compose the understanding thesaurus. Hypertext links are obtained implicitly with functions corresponding to classes, and thus they are obtained automatically from understanding results. These classes include incompletely recognized classes distinctly to offer utilization of incompletely recognized text blocks as they are. Using this approach, large scale and practical electronic library systems can be offered.> Shin'ichi Satoh 0001, Atsuhiro Takasu, Eishi Katsura |
ICDAR | 2 |