VLDB 2026 Research / reviewers in the wild / expert
Nao Souma
dblp:341/0982
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0002-8029-0524ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | CL-HumanEval: A Benchmark for Evaluating Cross-lingual Transfer though Code Generation
Miyu Sato, Yui Obara, Nao Souma, Kimio Kuramitsu |
PACLIC | 3 |
| 2024 | Distributed Dataset Framework for Large Language Models Pre-training
Nao Souma, Yui Obara, Yasuhiko Yokote, Yutaka Ishikawa, Kimio Kuramitsu |
PKAW | 1 |
| 2023 | Can ChatGPT Correct Code Based on Logical Steps?abstractChatGPT presents emerging opportunities in software development, yet its capabilities for understanding code remain largely understudied. This study aims to focus on the logical aspect of code comprehension of ChatGPT by examining its performance in detecting and fixing bugs. Our preliminary results suggest that ChatGPT seems to correct code in a different way than human logical steps. Nao Souma, Waka Ito, Momoka Obara, Takako Kawaguchi, Yuka Akinobu, Toshiyuki Kurabayashi, Haruto Tanno, Kimio Kuramitsu |
APSEC | 1 |
| 2022 | An additional approach to pre-trained code model with multilingual natural languagesabstractPre-trained language models have achieved many prominent results in natural language processing. Since software engineering widely includes many natural language documents, the application of pre-trained language models have received much attention in software engineering tasks. However, pretraining a large volume of source code requires a huge amount of computational resources and time. In this study, we propose an additional pre-training approach to a well-trained language model. Our initial results on mT5, multilingual T5 with an additional pretraining of Python code shows improved performance on multiple software engineering tasks including code generation, code summarisation, code repair, and error diagnosis. Teruno Kajiura, Nao Souma, Miyu Sato, Mai Takahashi, Kimio Kuramitsu |
APSEC | 2 |