VLDB 2026 Research / reviewers in the wild / expert
Tomoki Nakamaru
dblp:207/7239
· DBLP profile ↗
12ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-9451-5595ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Annotation-Guided Edit-Aware JIT Compilation for Julia Computational NotebooksabstractIn cell-based computational notebooks, programmers repeatedly edit and execute cells. Recompiling and re-executing an entire cell on every execution is inefficient; ideally, stable code and data should be reused. We present annotation-guided edit-aware JIT compilation for Julia, a technique that leverages two user-provided annotations about code and data stabilities — @hole and @persistent — to avoid unnecessary recompilation and recomputation. It partitions a cell into a stable skeleton and an unstable hole, compiles them separately while folding stable data as inlined constants, and reuses the skeleton across executions. We prototyped this approach as nbjit.jl, an IJulia kernel extension. Across five editing scenarios, cumulative execution time ranges from 3.3 × (lightweight workloads) to 1/167 (compute-intensive workloads) relative to Julia’s standard eval. The prototype currently supports a subset of Julia: basic numeric types, control flow, and composite data structures. Automatic annotation inference and full Julia language coverage remain as future work. Yusuke Izawa, Tomoki Nakamaru, Tetsuro Yamazaki |
MPLR | 2 |
| 2026 | LR parsing for strings with placeholdersabstractThis paper studies a parsing method for strings containing placeholders, each of which may be later replaced by a string derived from the corresponding nonterminal symbol. Such a method potentially applies to parallel/distributed parsing, parsing for templates, modular syntax definitions, and so on. This paper investigates whether the introduction of the placeholder preserves the class of the grammar and proves the following two facts. First, the class of LR( k )grammars is preserved if k ≥ 1 and every nonterminal derives at least one nonempty string; hence, we can apply the standard LR parsing algorithm for parsing strings with placeholders. Second, the class of LR(0) is not. These results extend the preceding study for the LL(1) grammars. Kohei Nakamichi, Akimasa Morihata, Tomoki Nakamaru |
Inf. Process. Lett. | 3 |
| 2025 | Compressing Cell Execution Logs Embedded within Jupyter Notebook FilesabstractJupyter notebooks inevitably become messy through exploratory programming. NBGather is a Jupyter extension that helps programmers manage such messes by logging executed cells and slicing out programs to reproduce outputs from the log. However, its space efficiency is suboptimal because it naively records executed cells in the JSON format. This paper presents a straightforward but effective method for compressing cell execution logs in Jupyter notebooks. In a preliminary experiment using an open dataset, our method achieved a median size reduction of 91% compared to the original $\log$ size and 40% compared to the size achieved by naively applying gzip compression. Tomoki Nakamaru |
APSEC | 1 |
| 2025 | Jupyter Notebook Activity DatasetabstractFine-grained logs of programmers’ activities serve as research material in various fields of computer science. For activities in a traditional environment, there is a publicly available dataset. However, to our knowledge, no comparable dataset exists for activities in a cell-based computational notebook environment. This paper presents an open dataset of activities in the Jupyter Notebook. The dataset comprises the activity logs of 21 programmers who worked on a one-hour data science task in our log collection experiment. The dataset would contribute to understanding programmers’ workflows in cell-based notebook environments and evaluating proposals for those environments. Tomoki Nakamaru, Tomomasa Matsunaga, Tetsuro Yamazaki |
MSR | 1 |
| 2024 | Smells of Misunderstanding in File Path Patterns within DockerignoreabstractWriting file path patterns is a common activity in software development, but the semantics of patterns are not common across developer tools. This fact leads to the following question: Do developers acknowledge such differences in the semantics of file path patterns? This paper presents our preliminary study to answer this question. We built a dataset of 246. dockerignore files collected from GitHub and mined them for smells of misunderstanding of the dockerignore semantics. Our initial analysis found such smells in 42 % of the dataset. This result may motivate further investigation into this problem. Tomoki Nakamaru |
APSEC | 1 |
| 2024 | Multiverse Notebook: Shifting Data Scientists to Time TravelersabstractComputational notebook environments are popular and de facto standard tools for programming in data science, whereas computational notebooks are notorious in software engineering. The criticism there stems from the characteristic of facilitating unrestricted dynamic patching of running programs, which makes exploratory coding quick but the resultant code messy and inconsistent. In this work, we first reveal that dynamic patching is a natural demand rather than a mere bad practice in data science programming on Kaggle. We then develop Multiverse Notebook, a computational notebook engine for time-traveling exploration. It enables users to time-travel to any past state and restart with new code from there under state isolation. We present an approach to efficiently implementing time-traveling exploration. We empirically evaluate Multiverse Notebook on ten real-world tasks from Kaggle. Our experiments show that time-traveling exploration on Multiverse Notebook is reasonably efficient. Shigeyuki Sato 0001, Tomoki Nakamaru |
Proc. ACM Program. Lang. | 2 |
| 2023 | Collecting Cyclic Garbage across Foreign Function Interfaces: Who Takes the Last Piece of Cake?abstractA growing number of libraries written in managed languages, such as Python and JavaScript, are bringing about new demand for a foreign language interface (FFI) between two managed languages. Such an FFI allows a host-language program to seamlessly call a library function written in a foreign language and exchange objects. It is often implemented by a user-level library but such implementation cannot reclaim cyclic garbage, or a group of objects with circular references, across the language boundary. This paper proposes Refgraph GC , which enables FFI implementation that can reclaim cyclic garbage. Refgraph GC coordinates the garbage collectors of two languages and it needs to modify the managed runtime of one language only. It does not modify that of the other language. This paper discusses the soundness and completeness of the proposed algorithm and also shows the results of the experiments with our implementation of FFI with Refgraph GC. This FFI allows a Ruby program to access a JavaScript library. Tetsuro Yamazaki, Tomoki Nakamaru, Ryota Shioya, Tomoharu Ugawa, Shigeru Chiba |
Proc. ACM Program. Lang. | 2 |
| 2022 | An Anomaly-Based Approach for Detecting Modularity Violations on Method PlacementabstractThis paper presents a technique for detecting an anomaly in method placements in Java packages. This anomaly detection helps code reviewers discover a method belonging to an inappropriate package in modularity when developers commit changes in their software development projects. Moving such a method to an appropriate package will contribute to the maintenance of good modularity in their projects. This is particularly beneficial in the later stage of development, where modularity is often violated by adding new features not anticipated in the initial plan. Our technique is based on few-shot classification in machine learning. This paper empirically reveals that our neural network model can detect an anomaly in method placements and a significant portion of the anomalies is considered as inappropriate method placements in modularity. Our model can discover even a method placement that violates a project-specific coding rule that its developers would choose for some reason of maintainability or readability. Our technique is useful for maintaining the consistency in such a project-specific rule. Kazuki Yoda, Tomoki Nakamaru, Soramichi Akiyama, Shigeru Chiba |
QRS | 2 |
| 2022 | Yet Another Generating Method of Fluent Interfaces Supporting Flat- and Sub-chaining StylesabstractResearchers discovered methods to generate fluent interfaces equipped with static checking to verify their calling conventions. This static checking is done by carefully designing classes and method signatures to make type checking to perform a calculation equivalent to syntax checking. In this paper, we propose a method to generate a fluent interface with syntax checking, which accepts both styles of method chaining; flat-chaining style and sub-chaining style. Supporting both styles is worthwhile because it allows programmers to wrap out parts of their method chaining for readability. Our method is based on grammar rewriting so that we could inspect the acceptable grammar. In conclusion, our method succeeds generation when the input grammar is LL(1) and there is no non-terminal symbol that generates either only an empty string or nothing. Tetsuro Yamazaki, Tomoki Nakamaru, Shigeru Chiba |
SLE | 2 |
| 2020 | An Empirical Study of Method Chaining in JavaabstractWhile some promote method chaining as a good practice for improving code readability, others refer to it as a bad practice that worsens code quality. In this paper, we first investigate whether method chaining is a programming style accepted by real-world programmers. To answer this question, we collected 2,814 Java repositories on GitHub and analyzed historical trends in the frequency of method chaining. The results of our analysis revealed the increasing use of method chaining; 23.1% of method invocations were part of method chains in 2018, whereas only 16.0% were such invocations in 2010. We then explore language features that are helpful to the method-chaining style but have not been supported yet in Java. For this aim, we conducted manual inspections of method chains that are randomly sampled from the collected repositories. We also estimated how effective they are to encourage the method-chaining style if they are adopted in Java. Tomoki Nakamaru, Tomomasa Matsunaga, Tetsuro Yamazaki, Soramichi Akiyama, Shigeru Chiba |
MSR | 1 |
| 2019 | Generating a fluent API with syntax checking from an LR grammarabstractThis paper proposes a fluent API generator for Scala, Haskell, and C++. It receives a grammar definition and generates a code skeleton of the library in the host programming language. The generated library is accessed through a chain of method calls; this style of API is called a fluent API. The library uses the host-language type checker to detect an invalid chain of method calls. Each method call is regarded as a lexical token in the embedded domain specific language implemented by that library. A sequence of the lexical tokens is checked and, if the sequence is not acceptable by the grammar, a type error is reported during compilation time. A contribution of this paper is to present an algorithm for generating the code-skeleton for a fluent API that reports a type error when a chain of method calls to the library does not match the given LR grammar. Our algorithm works in Scala, Haskell, and C++. To encode LR parsing, it uses the method/function overloading available in those languages. It does not need an advanced type system, or exponential compilation time or memory consumption. This paper also presents our implementation of the proposed generator. Tetsuro Yamazaki, Tomoki Nakamaru, Kazuhiro Ichikawa, Shigeru Chiba |
Proc. ACM Program. Lang. | 2 |
| 2017 | Silverchain: a fluent API generatorabstractThis paper presents a tool named Silverchain, which generates class definitions for a fluent API from the grammar of the API. A fluent API is an API that is used by method chaining and its grammar is a BNF-like set of rules that defines method chains accepted in type checking. Fluent APIs generated by Silverchain provide two styles of APIs: One is for building a chain by concatenating all method calls in series. The other is for building a chain from partial chains by passing child chains to method calls in the parent chain as their arguments. To generate such a fluent API, Silverchain first translates given grammar into a set of deterministic pushdown automata without ϵ-transitions, then encodes these automata into class definitions. Each constructed automata corresponds to a nonterminal in given grammar and recognizes symbol sequences produced from its corresponding nonterminal. Tomoki Nakamaru, Kazuhiro Ichikawa, Tetsuro Yamazaki, Shigeru Chiba |
GPCE | 1 |