Yasuhisa Fujii

dblp:84/8914 · DBLP profile ↗
← Back
24ranked-venue papers
7as first author
12since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2024 Fine-Tuning Large Language Models for Automatic Font Skeleton Generation: Exploration and Analysis
Yasuhisa Fujii, Xinru Zhu 0001, Kayoko Nohara
ACCV (3)2
2024 Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding
abstract
Table-based reasoning with large language models (LLMs) is a promising direction to tackle many table understanding tasks, such as table-based question answering and fact verification. Compared with generic reasoning, table-based reasoning requires the extraction of underlying semantics from both free-form questions and semi-structured tabular data. Chain-of-Thought and its similar approaches incorporate the reasoning chain in the form of textual context, but it is still an open question how to effectively leverage tabular data in the reasoning chain. We propose the Chain-of-Table framework, where tabular data is explicitly used in the reasoning chain as a proxy for intermediate thoughts. Specifically, we guide LLMs using in-context learning to iteratively generate operations and update the table to represent a tabular reasoning chain. LLMs can therefore dynamically plan the next operation based on the results of the previous ones. This continuous evolution of the table forms a chain, showing the reasoning process for a given tabular problem. The chain carries structured information of the intermediate results, enabling more accurate and reliable predictions. Chain-of-Table achieves new state-of-the-art performance on WikiTQ, FeTaQA, and TabFact benchmarks across multiple LLM choices.
Zilong Wang 0002, Chun-Liang Li, Julian Martin Eisenschlos, Vincent Perot, Zifeng Wang 0002, Lesly Miculicich, Yasuhisa Fujii, Jingbo Shang, Chen-Yu Lee, Tomas Pfister
ICLR8
2024 TableRAG: Million-Token Table Understanding with Language Models
abstract
Recent advancements in language models (LMs) have notably enhanced their ability to reason with tabular data, primarily through program-aided mechanisms that manipulate and analyze tables. However, these methods often require the entire table as input, leading to scalability challenges due to the positional bias or context length constraints. In response to these challenges, we introduce TableRAG, a Retrieval-Augmented Generation (RAG) framework specifically designed for LM-based table understanding. TableRAG leverages query expansion combined with schema and cell retrieval to pinpoint crucial information before providing it to the LMs. This enables more efficient data encoding and precise retrieval, significantly reducing prompt lengths and mitigating information loss. We have developed two new million-token benchmarks from the Arcade and BIRD-SQL datasets to thoroughly evaluate TableRAG's effectiveness at scale. Our results demonstrate that TableRAG's retrieval design achieves the highest retrieval quality, leading to the new state-of-the-art performance on large-scale table understanding.
Si-An Chen, Lesly Miculicich, Julian Martin Eisenschlos, Zifeng Wang 0002, Zilong Wang 0002, Yanfei Chen, Yasuhisa Fujii, Hsuan-Tien Lin, Chen-Yu Lee, Tomas Pfister
NeurIPS7
2024 Hierarchical Text Spotter for Joint Text Spotting and Layout Analysis
abstract
We propose Hierarchical Text Spotter (HTS), a novel method for the joint task of word-level text spotting and geometric layout analysis. HTS can recognize text in an image and identify its 4-level hierarchical structure: characters, words, lines, and paragraphs. The proposed HTS is characterized by two novel components: (1) a Unified-DetectorPolygon (UDP) that produces Bezier Curve polygons of text lines and an affinity matrix for paragraph grouping between detected lines; (2) a Line-to-Character-to-Word (L2C2W) recognizer that splits lines into characters and further merges them back into words. HTS achieves stateof-the-art results on multiple word-level text spotting benchmark datasets as well as geometric layout analysis tasks.
Shangbang Long, Siyang Qin, Yasuhisa Fujii, Alessandro Bissacco, Michalis Raptis
WACV3
2023 FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction
abstract
Chen-Yu Lee, Chun-Liang Li, Hao Zhang, Timothy Dozat, Vincent Perot, Guolong Su, Xiang Zhang, Kihyuk Sohn, Nikolay Glushnev, Renshen Wang, Joshua Ainslie, Shangbang Long, Siyang Qin, Yasuhisa Fujii, Nan Hua, Tomas Pfister. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Chen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Kihyuk Sohn, Nikolay Glushnev, Renshen Wang, Joshua Ainslie, Shangbang Long, Siyang Qin, Yasuhisa Fujii, Nan Hua, Tomas Pfister
ACL (1)14
2023 OCR Language Models with Custom Vocabularies
Peter Garst, R. Reeve Ingle, Yasuhisa Fujii
ICDAR (4)3
2023 ICDAR 2023 Competition on Hierarchical Text Detection and Recognition
Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, Michalis Raptis
ICDAR (2)5
2023 Text Reading Order in Uncontrolled Conditions by Sparse Graph Segmentation
Renshen Wang, Yasuhisa Fujii, Alessandro Bissacco
ICDAR (6)2
2022 FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction
abstract
Chen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Nan Hua, Joshua Ainslie, Renshen Wang, Yasuhisa Fujii, Tomas Pfister. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Chen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Nan Hua, Joshua Ainslie, Renshen Wang, Yasuhisa Fujii, Tomas Pfister
ACL (1)9
2022 Towards End-to-End Unified Scene Text Detection and Layout Analysis
abstract
Scene text detection and document layout analysis have long been treated as two separate tasks in different image domains. In this paper, we bring them together and introduce the task of unified scene text detection and layout analysis. The first hierarchical scene text dataset is introduced to enable this novel research task. We also propose a novel method that is able to simultaneously detect scene text and form text clusters in a unified way. Comprehensive experiments show that our unified model achieves better performance than multiple well-designed baseline methods. Additionally, this model achieves state-of-the-art results on multiple scene text detection datasets without the need of complex post-processing. Dataset and code: https://github.com/google-research-datasets/hiertext.
Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, Michalis Raptis
CVPR5
2022 Unified Line and Paragraph Detection by Graph Convolutional Networks
Renshen Wang, Michalis Raptis, Yasuhisa Fujii
DAS4
2022 Post-OCR Paragraph Recognition by Graph Convolutional Networks
abstract
We propose a new approach for paragraph recognition in document images by spatial graph convolutional networks (GCN) applied on OCR text boxes. Two steps, namely line splitting and line clustering, are performed to extract paragraphs from the lines in OCR results. Each step uses a β-skeleton graph constructed from bounding boxes, where the graph edges provide efficient support for graph convolution operations. With pure layout input features, the GCN model size is 3~4 orders of magnitude smaller compared to RCNN based models, while achieving comparable or better accuracies on PubLayNet and other datasets. Furthermore, the GCN models show good generalization from synthetic training data to real-world images, and good adaptivity for variable document styles.
Renshen Wang, Yasuhisa Fujii, Ashok C. Popat
WACV2
2019 Towards Unconstrained End-to-End Text Spotting
abstract
We propose an end-to-end trainable network that can simultaneously detect and recognize text of arbitrary shape, making substantial progress on the open problem of reading scene text of irregular shape. We formulate arbitrary shape text detection as an instance segmentation problem; an attention model is then used to decode the textual content of each irregularly shaped text region without rectification. To extract useful irregularly shaped text instance features from image scale features, we propose a simple yet effective RoI masking step. Additionally, we show that predictions from an existing multi-step OCR engine can be leveraged as partially labeled training data, which leads to significant improvements in both the detection and recognition accuracy of our model. Our method surpasses the state-of-the-art for end-to-end recognition tasks on the ICDAR15 (straight) benchmark by 4.6%, and on the Total-Text (curved) benchmark by more than 16%.
Siyang Qin, Alessandro Bissacco, Michalis Raptis, Yasuhisa Fujii
ICCV4
2019 A Scalable Handwritten Text Recognition System
abstract
Many studies on (Offline) Handwritten Text Recognition (HTR) systems have focused on building state-of-the-art models for line recognition on small corpora. However, adding HTR capability to a large scale multilingual OCR system poses new challenges. This paper addresses three problems in building such systems: data, efficiency, and integration. Firstly, one of the biggest challenges is obtaining sufficient amounts of high quality training data. We address the problem by using online handwriting data collected for a large scale production online handwriting recognition system. We describe our image data generation pipeline and study how online data can be used to build HTR models. We show that the data improve the models significantly under the condition where only a small number of real images is available, which is usually the case for HTR models. It enables us to support a new script at substantially lower cost. Secondly, we propose a line recognition model based on neural networks without recurrent connections. The model achieves a comparable accuracy with LSTM-based models while allowing for better parallelism in training and inference. Finally, we present a simple way to integrate HTR models into an OCR system. These constitute a solution to bring HTR capability into a large scale OCR system.
R. Reeve Ingle, Yasuhisa Fujii, Thomas Deselaers, Jonathan Baccash, Ashok C. Popat
ICDAR2
2017 Sequence-to-Label Script Identification for Multilingual OCR
abstract
We describe a novel line-level script identification method. Previous work repurposed an OCR model generating per-character script codes, counted to obtain line-level script identification. This has two shortcomings. First, as a sequence-to-sequence model it is more complex than necessary for the sequence-to-label problem of line script identification. This makes it harder to train and inefficient to run. Second, the counting heuristic may be suboptimal compared to a learned model. Therefore we reframe line script identification as a sequence-to-label problem and solve it using two components, trained end-toend: Encoder and Summarizer. The encoder converts a line image into a feature sequence. The summarizer aggregates the sequence to classify the line. We test various summarizers with identical inception-style convolutional networks as encoders. Experiments on scanned books and photos containing 232 languages in 30 scripts show 16% reduction of script identification error rate compared to the baseline. This improved script identification reduces the character error rate attributable to script misidentification by 33%.
Yasuhisa Fujii, Karel Driesen, Jonathan Baccash, Ash Hurst, Ashok C. Popat
ICDAR1
2015 Label transition and selection pruning and automatic decoding parameter optimization for time-synchronous Viterbi decoding
abstract
Hidden Markov Model (HMM)-based classifiers have been successfully used for sequential labeling problems such as speech recognition and optical character recognition for decades. They have been especially successful in the domains where the segmentation is not known or difficult to obtain, since, in principle, all possible segmentation points can be taken into account. However, the benefit comes with a non-negligible computational cost. In this paper, we propose simple yet effective new pruning algorithms to speed up decoding with HMM-based classifiers of up to 95% relative over a baseline. As the number of tunable decoding parameters increases, it becomes more difficult to optimize the parameters for each configuration. We also propose a novel technique to estimate the parameters based on a loss value without relying on a grid search.
Yasuhisa Fujii, Dmitriy Genzel, Ashok C. Popat, Remco Teunen
ICDAR1
2013 A robust/fast spoken term detection method based on a syllable n-gram index with a distance metric
Seiichi Nakagawa, Keisuke Iwami, Yasuhisa Fujii, Kazumasa Yamamoto
Speech Commun.3
2011 Automatic speech recognition using Hidden Conditional Neural Fields
abstract
Hidden Conditional Random Fields(HCRF) is a very promising approach to model speech. However, because HCRF computes the score of a hypothesis by summing up linearly weighted features, it cannot consider non-linearity among features that will be crucial for speech recognition. In this pa per, we extend HCRF by incorporating gate function used in neural networks and propose a new model called Hidden Conditional Neural Fields(HCNF). Differently with conventional approaches, HCNF can be trained without any initial model and incorporate any kinds of features. Experimental results of continuous phoneme recognition on TIMIT core test set and Japanese read speach recognition task using monophone showed that HCNF was superior to HCRF and HMM trained in MPE manner.
Yasuhisa Fujii, Kazumasa Yamamoto, Seiichi Nakagawa
ICASSP1
2011 Efficient out-of-vocabulary term detection by n-gram array indices with distance from a syllable lattice
abstract
For spoken document retrieval, it is very important to con sider Out-of-Vocabulary (OOV) and mis-recognition of spoken words. Therefore, sub-word unit based recognition and retrieval methods have been proposed. This paper describes a Japanese spoken document retrieval system that is robust for considering OOV words and mis-recognition of sub-units. We used individual syllables as sub-word unit in continuous speech recognition and an n-gram sequence of syllables in a recognized syllable-based lattice. We propose an n-gram indexing/retrieval method with distance in the syllable lattice for attacking OOV, recognition errors, and high speed retrieval. We applied this method to academic lecture presentation database of 44 hours, and 0.58(F-value) of the OOV words were detected in less than 2.5 milliseconds.
Keisuke Iwami, Yasuhisa Fujii, Kazumasa Yamamoto, Seiichi Nakagawa
ICASSP2
2011 Hidden Boosted MMI and Hierarchical State Posterior Feature for Automatic Speech Recognition Based on Hidden Conditional Neural Fields
abstract
We have investigated automatic speech recognition using Hidden Conditional Neural Fields (HCNF). In this paper, we propose a new objective function, Hidden Boosted MMI (HBMMI) that considers the number of errors in the training data even if the correct state sequence is not known for training the HCNF. The experimental results show that HB-MMI can improve recognition accuracy if overfitting does not occur. We also present an automatic speech recognition method using a hierarchical state posterior feature where the output from the first stage HCNF is used as input for the second stage HCNF. The experimental results show that the feature improves recognition accuracy. By combining both of the proposed methods, we obtain further improvements. Index Terms: hidden conditional neural fields, automatic speech recognition, hidden boosted MMI, state posterior feature
Yasuhisa Fujii, Kazumasa Yamamoto, Seiichi Nakagawa
INTERSPEECH1
2010 Improving the readability of class lecture ASR results using a confusion network
abstract
This paper presents a method for improving the readability of Automatic Speech Recognition (ASR) results for classroom lectures. Most of the previous research on improving the readability of recognition results focused mainly on manually transcribed texts, and not ASR results. Due to the presence of a large number of domain-dependent words and the casual presentation style, even state-of-the-art recognizers yield a 30-50% word error rate for speech in classroom lectures. Thus, a method for improving the readability of ASR results needs to be robust against recognition errors. In this paper, we propose a novel method for improving the readability based on a machine translation model that uses a confusion network representing multiple hypotheses of the ASR results to achieve robustness against recognition errors. Experimental results show that the proposed method outperforms the baselines in both automatic and manual evaluations. Index Terms: improving readability, confusion network, automatic speech recognition, classroom lecture speech
Yasuhisa Fujii, Kazumasa Yamamoto, Seiichi Nakagawa
INTERSPEECH1
2010 Out-of-vocabulary term detection by n-gram array with distance from continuous syllable recognition results
abstract
For spoken document retrieval, it is very important to consider Out-of-Vocabulary (OOV) and mis-recognition of spoken words. Therefore, sub-word unit based recognition and retrieval methods have been proposed. This paper describes a Japanese spoken document retrieval system that is robust for considering OOV words and mis-recognition of sub-units. To solve the problem of OOV keywords and mis-recognized words, we used individual syllables as sub-word unit in continuous speech recognition and an n-gram sequence of syllables as a retrieval unit. We propose an n-gram indexing/retrieval method with distance in a syllable lattice for attacking OOV, recognition errors, and high speed retrieval. We applied this method to academic lecture presentation database of 44 hours, and 60% of the OOV words were detected in less than 2.5 milliseconds.
Keisuke Iwami, Yasuhisa Fujii, Kazumasa Yamamoto, Seiichi Nakagawa
SLT2
2008 Class lecture summarization taking into account consecutiveness of important sentences
abstract
This paper presents a novel sentence extraction framework that takes into account the consecutiveness of important sentences using a Support Vector Machine (SVM). Generally, most ex-tractive summarizers do not take context information into ac-count, but do take into account the redundancy over the entire summarization. However, there must exist relationships among the extracted sentences. Actually, we can observe these rela-tionships as consecutiveness among the sentences. We deal with this consecutiveness by using dynamic and difference features to decide if a sentence needs to be extracted or not. Since impor-tant sentences tend to be extracted consecutively, we just used the decision made for the previous sentence as the dynamic fea-ture. We used the differences between the current and previ-ous feature values for the difference feature, since adjacent sen-
Yasuhisa Fujii, Kazumasa Yamamoto, Norihide Kitaoka, Seiichi Nakagawa
INTERSPEECH1
2007 Automatic extraction of cue phrases for important sentences in lecture speech and automatic lecture speech summarization
abstract
We automatically extract the summaries of spoken class lectures. This paper presents a novel method for sentence extraction-based automatic speech summarization. We propose a technique that extracts “cue phrases for im-portant sentences (CPs) ” that often appear in important sen-tences. We formulate CP extraction as a labeling problem of word sequences and use Conditional Random Fields (CRF) [1] for labeling. Automatic summarization using CP extraction re-sults as features yields precisions of 0.603 and 0.556 when us-ing manual transcriptions and Automatic Speech Recognition (ASR) results, respectively. Combining the features derived from the CPs and tradi-tional features (including repeated words, words repeated in a slide text, and term frequency (tf), which are surface linguistic information, and speech power and duration, which are prosodic features) [2, 3], we obtained better summarization performance with a κ-value of 0.380, a F-measure of 0.539, and a Rouge-4 of 0.709. Index Terms: automatic speech summarization, sentence ex-traction, speech synthesis
Yasuhisa Fujii, Norihide Kitaoka, Seiichi Nakagawa
INTERSPEECH1