Tetsuro Yamazaki

dblp:207/7224 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-2065-5608ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 14 · 3 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Annotation-Guided Edit-Aware JIT Compilation for Julia Computational Notebooks
abstract
In cell-based computational notebooks, programmers repeatedly edit and execute cells. Recompiling and re-executing an entire cell on every execution is inefficient; ideally, stable code and data should be reused. We present annotation-guided edit-aware JIT compilation for Julia, a technique that leverages two user-provided annotations about code and data stabilities — @hole and @persistent — to avoid unnecessary recompilation and recomputation. It partitions a cell into a stable skeleton and an unstable hole, compiles them separately while folding stable data as inlined constants, and reuses the skeleton across executions. We prototyped this approach as nbjit.jl, an IJulia kernel extension. Across five editing scenarios, cumulative execution time ranges from 3.3 × (lightweight workloads) to 1/167 (compute-intensive workloads) relative to Julia’s standard eval. The prototype currently supports a subset of Julia: basic numeric types, control flow, and composite data structures. Automatic annotation inference and full Julia language coverage remain as future work.
Yusuke Izawa, Tomoki Nakamaru, Tetsuro Yamazaki
MPLR3
2026 LLM-Based Explainable Detection of LLM-Generated Code in Python Programming Courses
Jeonghun Baek, Tetsuro Yamazaki, Akimasa Morihata, Junichiro Mori, Yoko Yamakata, Kenjiro Taura, Shigeru Chiba
SIGCSE (1)2
2026 MaskingAgent: Preventing LLM Tutor from Providing Full Solutions in Python Programming Courses
Jeonghun Baek, Tetsuro Yamazaki, Akimasa Morihata, Junichiro Mori, Yoko Yamakata, Kenjiro Taura, Shigeru Chiba
SIGCSE (2)2
2026 Dynamic Wind for OCaml Effect Handlers with Escaping Continuation Support
abstract
Effect handlers and dynamic wind provide mechanisms for implementing language constructs, such as generators and dynamically scoped variables, as libraries or embedded domain-specific languages (EDSLs). This paper presents a library implementation of dynamic wind in OCaml that is composable with effect handlers in OCaml. Their composition is challenging, especially when a delimited continuation captured by an effect handler escapes the handler's scope. Indeed, a straightforward retrofitting of Voigt's dynamic wind, originally implemented in Effekt, does not behave intuitively. In particular, it fails to implement dynamically scoped variables when a generator created within a dynamic wind scope resumes outside that scope. This is because Voigt's approach captures only the dynamic winds inside delimited continuations. Our key idea is to decorate delimited continuations with the dynamic winds enclosing the effect handler, i.e., the delimited continuation itself. We implement our dynamic wind library on top of OCaml's effect handlers and demonstrate that our design yields intuitive behavior.
Antonino Yann William Gillard, Tetsuro Yamazaki, Tomoharu Ugawa
SLE2
2025 Jupyter Notebook Activity Dataset
abstract
Fine-grained logs of programmers’ activities serve as research material in various fields of computer science. For activities in a traditional environment, there is a publicly available dataset. However, to our knowledge, no comparable dataset exists for activities in a cell-based computational notebook environment. This paper presents an open dataset of activities in the Jupyter Notebook. The dataset comprises the activity logs of 21 programmers who worked on a one-hour data science task in our log collection experiment. The dataset would contribute to understanding programmers’ workflows in cell-based notebook environments and evaluating proposals for those environments.
Tomoki Nakamaru, Tomomasa Matsunaga, Tetsuro Yamazaki
MSR3
2025 Yet Another Trace-Based Approach for the Cause of Software Regressions in JavaScript and Python
abstract
Addressing the cause of software regressions is an important but difficult task, and has not been well studied. Current tools have some limitations, such as low detection accuracy. In this paper, we try to address these limitations and improve the accuracy of locating causes of software regressions by proposing three new techniques. Our techniques are based on tracing source code changes during program execution. Moreover, we extend current benchmark to a new programming language, and compare our techniques against existing techniques on the extended benchmark. Our new techniques outperform existing techniques in accurately locating the root cause of software regressions. Specifically, our LLMpowered technique achieves the state-of-the-art accuracy.
Yuefeng Hu, Tetsuro Yamazaki, Shigeru Chiba
QRS3
2025 Leveraging LLM for Detecting and Explaining LLM-generated Code in Python Programming Courses
Jeonghun Baek, Tetsuro Yamazaki, Akimasa Morihata, Junichiro Mori, Yoko Yamakata, Kenjiro Taura, Shigeru Chiba
SIGCSE (2)2
2024 InferType: A Compiler Toolkit for Implementing Efficient Constraint-Based Type Inference
abstract
Over the last several decades, software has been woven into the fabric of every aspect of our society. As software development surges and code infrastructure of enterprise applications ages, it is now more critical than ever to increase software development productivity and modernize legacy applications. Advances in deep learning and machine learning algorithms have enabled numerous breakthroughs, motivating researchers to leverage AI techniques to improve software development efficiency. Thus, the fast-emerging research area of AI for Code has garnered new interest and gathered momentum. In this paper, we present a large-scale dataset CodeNet, consisting of over 14 million code samples and about 500 million lines of code in 55 different programming languages, which is aimed at teaching AI to code. In addition to its large scale, CodeNet has a rich set of high-quality annotations to benchmark and help accelerate research in AI techniques for a variety of critical coding tasks, including code similarity and classification, code translation between a large variety of programming languages, and code performance (runtime and memory) improvement techniques. Additionally, CodeNet provides sample input and output test sets for 98.5% of the code samples, which can be used as an oracle for determining code correctness and potentially guide reinforcement learning for code quality improvements. As a usability feature, we provide several pre-processing tools in CodeNet to transform source code into representations that can be readily used as inputs into machine learning models. Results of code classification and code similarity experiments using the CodeNet dataset are provided as a reference. We hope that the scale, diversity and rich, high-quality annotations of CodeNet will offer unprecedented research opportunities at the intersection of AI and Software Engineering.
Senxi Li, Tetsuro Yamazaki, Shigeru Chiba
ECOOP2
2024 Interactive Programming for Microcontrollers by Offloading Dynamic Incremental Compilation
abstract
Interactive execution environments are suitable for trial-and-error basis programming for microcontrollers. However, they are mostly implemented as interpreters to meet microcontrollers' limited memory size and demands for portability. Hence, their execution performance is not sufficiently high. In this paper, we propose offloading dynamic incremental compilation and linking to a host computer connected to a microcontroller. Since the computing resources of the host computer are sufficient to execute incremental dynamic compilation, they are used to enhance the relatively poor computing resources of the microcontroller. To show the feasibility of this idea, we design a small programming language named BlueScript and implement its interactive execution environment. Our experiment reveals that BlueScript executes a program one to two orders of magnitude faster than MicroPython, while its interactivity is comparable to that of MicroPython despite using dynamic incremental compilation.
Fumika Mochizuki, Tetsuro Yamazaki, Shigeru Chiba
MPLR2
2024 Dynamic Possible Source Count Analysis for Data Leakage Prevention
abstract
Dynamic Taint Analysis (DTA) is a widely studied technique that can effectively detect various attacks and information leakage. In the context of detecting information leakage, taint is a flag added to data to indicate whether secret data can be inferred from it. DTA tracks the flow of tainted data in a language runtime environment and identifies secret data leakage when tainted data is transmitted externally. We found that existing DTAs can produce false negatives and false positives in complex data flows because of the binary nature of taint. Since taint is binary, meaning either secret data is inferable (=1) or non-inferable (=0), it cannot represent intermediate states that may slightly infer the secret data, and these states are quantized to 0 or 1. As a result of this quantization, existing methods are unable to distinguish between outputs that are practically secure and those that pose a real security threat in complex data flows, resulting in false positives and false negatives. To address this problem, we introduce the concept of Possible Source Count (PSC) and propose Dynamic Possible source Count Analysis (DPCA), which tracks PSC instead of taint. PSC is a metric that indicates how many secrets can be identified by observing the data. DPCA tracks and computes the PSC of each data item using dynamic symbolic execution. By evaluating the PSC of data that reaches the sink point, DPCA can effectively distinguish between data that is practically secure and data that poses a security threat.
Eri Ogawa, Tetsuro Yamazaki, Ryota Shioya
MPLR2
2024 Dynamic Controllability Analysis for Preventing Injection Attacks
abstract
Injection attacks are some of the most serious security threats, and various techniques have been studied to prevent such attacks through program analysis. One of the typical dynamic analysis methods is Dynamic Taint Analysis (DTA), which adds a flag called taint to externally input data and detects an injection attack when these data reach a sink point where the system can be manipulated. However, DTA- based attack detection may produce many false positives and false negatives, especially in complex data flows. We consider that the high rate of false positives and negatives arises because the taint in DTA indicates whether data was controlled, not how much data was controlled. We propose Dynamic Controllability Analysis (DCA), an approach that approximates controllability by generalizing binary taint into natural numbers, indicating the extent of data control. We implemented DCA on a JavaScript runtime and evaluated the controllability computed by DCA. The evaluation results show that the controllability computed by DCA is sensitive to the presence or absence of an injection attack, yielding very low values when the system is safe and very high values when an attack is present.
Eri Ogawa, Tetsuro Yamazaki, Ryota Shioya
PRDC2
2024 Bugfox: A Trace-Based Analyzer for Localizing the Cause of Software Regression in JavaScript
abstract
Software regression has been a persistent issue in software development. Although numerous techniques have been proposed to prevent regression from being introduced before release, few are available to address regression as it occurs post-release. Therefore, identifying the root cause of regression has always been a time-consuming and labor-intensive task. We aim to deliver automated solutions for solving regressions based on tracing. We present Bugfox, a trace-based analyzer that reports functions as the possible cause of regression in JavaScript. The idea is to generate runtime trace with instrumented programs, then extract the differences between clean and regression traces, and apply two heuristic strategies based on invocation order and frequency to identify the suspicious functions among differences. We evaluate our approach on 12 real-world regressions taken from the benchmark BugsJS. First strategy solves 6 regressions, and second strategy solves other 4 regressions, resulting in an overall accuracy of 83% on test cases. Notably, Bugfox solves each regression in under 1 minute with minimal memory overhead (<200 Megabytes). Our findings suggest Bugfox could help developers solve regression in real development.
Yuefeng Hu, Hiromu Ishibe, Tetsuro Yamazaki, Shigeru Chiba
SLE4
2023 Collecting Cyclic Garbage across Foreign Function Interfaces: Who Takes the Last Piece of Cake?
abstract
A growing number of libraries written in managed languages, such as Python and JavaScript, are bringing about new demand for a foreign language interface (FFI) between two managed languages. Such an FFI allows a host-language program to seamlessly call a library function written in a foreign language and exchange objects. It is often implemented by a user-level library but such implementation cannot reclaim cyclic garbage, or a group of objects with circular references, across the language boundary. This paper proposes Refgraph GC , which enables FFI implementation that can reclaim cyclic garbage. Refgraph GC coordinates the garbage collectors of two languages and it needs to modify the managed runtime of one language only. It does not modify that of the other language. This paper discusses the soundness and completeness of the proposed algorithm and also shows the results of the experiments with our implementation of FFI with Refgraph GC. This FFI allows a Ruby program to access a JavaScript library.
Tetsuro Yamazaki, Tomoki Nakamaru, Ryota Shioya, Tomoharu Ugawa, Shigeru Chiba
Proc. ACM Program. Lang.1
2022 Yet Another Generating Method of Fluent Interfaces Supporting Flat- and Sub-chaining Styles
abstract
Researchers discovered methods to generate fluent interfaces equipped with static checking to verify their calling conventions. This static checking is done by carefully designing classes and method signatures to make type checking to perform a calculation equivalent to syntax checking. In this paper, we propose a method to generate a fluent interface with syntax checking, which accepts both styles of method chaining; flat-chaining style and sub-chaining style. Supporting both styles is worthwhile because it allows programmers to wrap out parts of their method chaining for readability. Our method is based on grammar rewriting so that we could inspect the acceptable grammar. In conclusion, our method succeeds generation when the input grammar is LL(1) and there is no non-terminal symbol that generates either only an empty string or nothing.
Tetsuro Yamazaki, Tomoki Nakamaru, Shigeru Chiba
SLE1
2020 An Empirical Study of Method Chaining in Java
abstract
While some promote method chaining as a good practice for improving code readability, others refer to it as a bad practice that worsens code quality. In this paper, we first investigate whether method chaining is a programming style accepted by real-world programmers. To answer this question, we collected 2,814 Java repositories on GitHub and analyzed historical trends in the frequency of method chaining. The results of our analysis revealed the increasing use of method chaining; 23.1% of method invocations were part of method chains in 2018, whereas only 16.0% were such invocations in 2010. We then explore language features that are helpful to the method-chaining style but have not been supported yet in Java. For this aim, we conducted manual inspections of method chains that are randomly sampled from the collected repositories. We also estimated how effective they are to encourage the method-chaining style if they are adopted in Java.
Tomoki Nakamaru, Tomomasa Matsunaga, Tetsuro Yamazaki, Soramichi Akiyama, Shigeru Chiba
MSR3
2019 Generating a fluent API with syntax checking from an LR grammar
abstract
This paper proposes a fluent API generator for Scala, Haskell, and C++. It receives a grammar definition and generates a code skeleton of the library in the host programming language. The generated library is accessed through a chain of method calls; this style of API is called a fluent API. The library uses the host-language type checker to detect an invalid chain of method calls. Each method call is regarded as a lexical token in the embedded domain specific language implemented by that library. A sequence of the lexical tokens is checked and, if the sequence is not acceptable by the grammar, a type error is reported during compilation time. A contribution of this paper is to present an algorithm for generating the code-skeleton for a fluent API that reports a type error when a chain of method calls to the library does not match the given LR grammar. Our algorithm works in Scala, Haskell, and C++. To encode LR parsing, it uses the method/function overloading available in those languages. It does not need an advanced type system, or exponential compilation time or memory consumption. This paper also presents our implementation of the proposed generator.
Tetsuro Yamazaki, Tomoki Nakamaru, Kazuhiro Ichikawa, Shigeru Chiba
Proc. ACM Program. Lang.1
2017 Silverchain: a fluent API generator
abstract
This paper presents a tool named Silverchain, which generates class definitions for a fluent API from the grammar of the API. A fluent API is an API that is used by method chaining and its grammar is a BNF-like set of rules that defines method chains accepted in type checking. Fluent APIs generated by Silverchain provide two styles of APIs: One is for building a chain by concatenating all method calls in series. The other is for building a chain from partial chains by passing child chains to method calls in the parent chain as their arguments. To generate such a fluent API, Silverchain first translates given grammar into a set of deterministic pushdown automata without ϵ-transitions, then encodes these automata into class definitions. Each constructed automata corresponds to a nonterminal in given grammar and recognizes symbol sequences produced from its corresponding nonterminal.
Tomoki Nakamaru, Kazuhiro Ichikawa, Tetsuro Yamazaki, Shigeru Chiba
GPCE3