Kenichiro Ando

dblp:329/6308 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0000-0003-0906-9891ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%
Artificial intelligence
1 paper
Information extraction and text analysis · 100%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Empirical software engineering
mining software repositories
0.812024
WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in Wikipedia · AAAI 2024
Empirical software engineering › mining software repositories
version history analysis
0.812024
WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in Wikipedia · AAAI 2024
Collaborative and social computing
peer production
0.212024
WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in Wikipedia · AAAI 2024
Collaborative and social computing › peer production
wikipedia
0.212024
WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in Wikipedia · AAAI 2024

Methods — techniques the papers use, named apart from their topics

text classification · 2.3
YearPublicationVenuePosition
2026 DEJIMA: A Novel Large-scale Japanese Dataset for Image Captioning and Visual Question Answering
abstract
This work addresses the scarcity of high-quality, large-scale resources for Japanese Vision-and-Language (V&L) modeling. We present a scalable and reproducible pipeline that integrates large-scale web collection with rigorous filtering/deduplication, object-detection-driven evidence extraction, and Large Language Model (LLM)-based refinement under grounding constraints. Using this pipeline, we build two resources: an image-caption dataset (DEJIMA-Cap) and a VQA dataset (DEJIMA-VQA), each containing 3.88M image-text pairs, far exceeding the size of existing Japanese V&L datasets. Human evaluations demonstrate that DEJIMA achieves substantially higher Japaneseness and linguistic naturalness than datasets constructed via translation or manual annotation, while maintaining factual correctness at a level comparable to human-annotated corpora. Quantitative analyses of image feature distributions further confirm that DEJIMA broadly covers diverse visual domains characteristic of Japan, complementing its linguistic and cultural representativeness. Models trained on DEJIMA exhibit consistent improvements across multiple Japanese multimodal benchmarks, confirming that culturally grounded, large-scale resources play a key role in enhancing model performance. All data sources and modules in our pipeline are licensed for commercial use, and we publicly release the resulting dataset and metadata to encourage further research and industrial applications in Japanese V&L modeling.
Toshiki Katsube, Taiga Fukuhara, Kenichiro Ando, Yusuke Mukuta, Kohei Uehara, Tatsuya Harada
LREC3
2024 WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in Wikipedia
abstract
Wikipedia can be edited by anyone and thus contains various quality sentences. Therefore, Wikipedia includes some poor-quality edits, which are often marked up by other editors. While editors' reviews enhance the credibility of Wikipedia, it is hard to check all edited text. Assisting in this process is very important, but a large and comprehensive dataset for studying it does not currently exist. Here, we propose WikiSQE, the first large-scale dataset for sentence quality estimation in Wikipedia. Each sentence is extracted from the entire revision history of English Wikipedia, and the target quality labels were carefully investigated and selected. WikiSQE has about 3.4 M sentences with 153 quality labels. In the experiment with automatic classification using competitive machine learning models, sentences that had problems with citation, syntax/semantics, or propositions were found to be more difficult to detect. In addition, by performing human annotation, we found that the model we developed performed better than the crowdsourced workers. WikiSQE is expected to be a valuable resource for other tasks in NLP.
Kenichiro Ando, Satoshi Sekine, Mamoru Komachi
AAAI1