Katsuki Chousa

dblp:224/0249 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Vision and language · 54% Machine translation · 23% Question answering and dialogue systems · 23%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 2 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
span prediction
0.412020
A Supervised Word Alignment Method based on Cross-Language Span Prediction using Multilingual BERT · EMNLP (1) 2020
Natural language and speech › Machine translation › statistical machine translation
word alignment
0.412020
A Supervised Word Alignment Method based on Cross-Language Span Prediction using Multilingual BERT · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

symmetrization · 0.4multilingual BERT · 0.4fine-tuning · 0.4
YearPublicationVenuePosition
2026 One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness
abstract
The hubness problem, in which hub embeddings are close to many unrelated examples, occurs often in high-dimensional embedding spaces and may pose a practical threat for purposes such as information retrieval and automatic evaluation metrics.In particular, since cross-modal similarity between text and images cannot be calculated by direct comparisons, such as string matching, cross-modal encoders that project different modalities into a shared space are helpful for various cross-modal applications, and thus, the existence of hubs may pose practical threats.To reveal the vulnerabilities of cross-modal encoders, we propose a method for identifying the hub embedding and its corresponding hub text.Experiments on image captioning evaluation in MSCOCO and nocaps along with image-to-text retrieval tasks in MSCOCO and Flickr30k showed that our method can identify a single hub text that unreasonably achieves comparable or higher similarity scores than human-written reference captions in many images, thereby revealing the vulnerabilities in cross-modal encoders.
Hiroyuki Deguchi 0002, Katsuki Chousa, Yusuke Sakai 0010
ACL (1)2
2026 JAPAS: A Benchmark and Neural Approach for Japanese Patent Support Relation Extraction
Katsuki Chousa, Ryosuke Sugiura
LREC1
2025 Automatic Evaluation of Language Generation Technology Based on Structure Alignment
abstract
Language generation techniques require automatic evaluation to carry out efficient and reproducible experiments. While n-gram matching is standard, it fails to capture semantic equivalence with different wording. Recent methods have addressed this issue by using contextual embeddings from pre-trained language models to compute the similarity between reference and hypothesis. However, these methods frequently disregard the syntax of sentences, despite its crucial role in determining meaning, and thus assign unjustifiably high scores. This paper proposes an automatic evaluation metric that considers both the words in sentences and their syntactic structures. We integrate syntactic information into the recent embedding-based approach. Experimental results obtained from two NLP tasks show that our method is at least comparable to standard baselines.
Katsuki Chousa, Tsutomu Hirao
COLING1
2024 JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus
abstract
We constructed JaParaPat (Japanese-English Parallel Patent Application Corpus), a bilingual corpus of more than 300 million Japanese-English sentence pairs from patent applications published in Japan and the United States from 2000 to 2021. We obtained the publication of unexamined patent applications from the Japan Patent Office (JPO) and the United States Patent and Trademark Office (USPTO). We also obtained patent family information from the DOCDB, that is a bibliographic database maintained by the European Patent Office (EPO). We extracted approximately 1.4M Japanese-English document pairs, which are translations of each other based on the patent families, and extracted about 350M sentence pairs from the document pairs using a translation-based sentence alignment method whose initial translation model is bootstrapped from a dictionary-based sentence alignment. We experimentally improved the accuracy of the patent translations by 20 bleu points by adding more than 300M sentence pairs obtained from patent applications to 22M sentence pairs obtained from the web.
Masaaki Nagata, Makoto Morishita, Katsuki Chousa, Norihito Yasuda
LREC/COLING3
2024 WikiSplit++: Easy Data Refinement for Split and Rephrase
abstract
The task of Split and Rephrase, which splits a complex sentence into multiple simple sentences with the same meaning, improves readability and enhances the performance of downstream tasks in natural language processing (NLP). However, while Split and Rephrase can be improved using a text-to-text generation approach that applies encoder-decoder models fine-tuned with a large-scale dataset, it still suffers from hallucinations and under-splitting. To address these issues, this paper presents a simple and strong data refinement approach. Here, we create WikiSplit++ by removing instances in WikiSplit where complex sentences do not entail at least one of the simpler sentences and reversing the order of reference simple sentences. Experimental results show that training with WikiSplit++ leads to better performance than training with WikiSplit, even with fewer training instances. In particular, our approach yields significant gains in the number of splits and the entailment ratio, a proxy for measuring hallucinations.
Hayato Tsukagoshi, Tsutomu Hirao, Makoto Morishita, Katsuki Chousa, Ryohei Sasano, Koichi Takeda 0003
LREC/COLING4
2022 JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus
abstract
Most current machine translation models are mainly trained with parallel corpora, and their translation accuracy largely depends on the quality and quantity of the corpora. Although there are billions of parallel sentences for a few language pairs, effectively dealing with most language pairs is difficult due to a lack of publicly available parallel corpora. This paper creates a large parallel corpus for English-Japanese, a language pair for which only limited resources are available, compared to such resource-rich languages as English-German. It introduces a new web-based English-Japanese parallel corpus named JParaCrawl v3.0. Our new corpus contains more than 21 million unique parallel sentence pairs, which is more than twice as many as the previous JParaCrawl v2.0 corpus. Through experiments, we empirically show how our new corpus boosts the accuracy of machine translation models on various domains. The JParaCrawl v3.0 corpus will eventually be publicly available online for research purposes.
Makoto Morishita, Katsuki Chousa, Jun Suzuki 0001, Masaaki Nagata
LREC2
2020 SpanAlign: Sentence Alignment Method based on Cross-Language Span Prediction and ILP
abstract
We propose a novel method of automatic sentence alignment from noisy parallel documents.We first formalize the sentence alignment problem as the independent predictions of spans in the target document from sentences in the source document.We then introduce a total optimization method using integer linear programming to prevent span overlapping and obtain non-monotonic alignments.We implement cross-language span prediction by fine-tuning pre-trained multilingual language models based on BERT architecture and train them using pseudo-labeled data obtained from unsupervised sentence alignment method.While the baseline methods use sentence embeddings and assume monotonic alignment, our method can capture the token-to-token interaction between the tokens of source and target text and handle non-monotonic alignments.In sentence alignment experiments on English-Japanese, our method achieved 70.3 F 1 scores, which are +8.0 points higher than the baseline method.In particular, our method improved by +53.9 F 1 scores for extracting non-parallel sentences.Our method improved the downstream machine translation accuracy by 4.1 BLEU scores when the extracted bilingual sentences are used for fine-tuning a pre-trained Japanese-to-English translation model. 1
Katsuki Chousa, Masaaki Nagata, Masaaki Nishino
COLING1
2020 Incorporating Noisy Length Constraints into Transformer with Length-aware Positional Encodings
abstract
Neural Machine Translation often suffers from an under-translation problem due to its limited modeling of output sequence lengths.In this work, we propose a novel approach to training a Transformer model using length constraints based on length-aware positional encoding (PE).Since length constraints with exact target sentence lengths degrade translation performance, we add random noise within a certain window size to the length constraints in the PE during the training.In the inference step, we predict the output lengths using input sequences and a BERTbased length prediction model.Experimental results in an ASPEC English-to-Japanese translation showed the proposed method produced translations with lengths close to the reference ones and outperformed a vanilla Transformer by 3.22 points in BLEU on short sentences within ten subwords.The average translation results using our length prediction model were also better than another baseline method using input lengths for the length constraints.The proposed noise injection improved robustness for length prediction errors, especially within the window size.
Yui Oka, Katsuki Chousa, Katsuhito Sudoh, Satoshi Nakamura 0001
COLING2
2020 A Supervised Word Alignment Method based on Cross-Language Span Prediction using Multilingual BERT
abstract
We present a novel supervised word alignment method based on cross-language span prediction.We first formalize a word alignment problem as a collection of independent predictions from a token in the source sentence to a span in the target sentence.Since this step is equivalent to a SQuAD v2.0 style question answering task, we solve it using the multilingual BERT, which is fine-tuned on manually created gold word alignment data.It is nontrivial to obtain accurate alignment from a set of independently predicted spans.We greatly improved the word alignment accuracy by adding to the question the source token's context and symmetrizing two directional predictions.In experiments using five word alignment datasets from among Chinese, Japanese, German, Romanian, French, and English, we show that our proposed method significantly outperformed previous supervised and unsupervised word alignment methods without any bitexts for pretraining.For example, we achieved 86.7 F1 score for the Chinese-English data, which is 13.3 points higher than the previous state-of-the-art supervised method. 1
Masaaki Nagata, Katsuki Chousa, Masaaki Nishino
EMNLP (1)2