VLDB 2026 Research / reviewers in the wild / expert
Yoshitaka Arahori
dblp:74/7913
· DBLP profile ↗
13ranked-venue papers
1as first author
2since 2021 · last 2021
0009-0000-7099-6091ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 2 since 2021Security and privacy · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Improving Semantic Consistency of Variable Names with Use-Flow Graph AnalysisabstractConsistency is one of the keys to maintainable source code and hence a successful software project. We propose a novel method of extracting the intent of programmers from source code of a large project (~ 300 kLOC) and checking the semantic consistency of its variable names. Our system learns a project-specific naming convention for variables based on its role solely from source code, and suggest alternatives when it violates its internal consistency. The system can also show the reasoning why a certain variable should be named in a specific way. The system does not rely on any external knowledge. We applied our method to 12 open-source projects and evaluated its results with human reviewers. Our system proposed alternative variable names for 416 out of 1080 (39%) instances that are considered better than ones originally used by the developers. Based on the results, we created patches to correct the inconsistent names and sent them to its developers. Three open-source projects adopted it. Yusuke Shinyama, Yoshitaka Arahori, Katsuhiko Gondow |
APSEC | 2 |
| 2021 | How Do Programmers Express High-Level Concepts using Primitive Data Types?abstractWe investigated how programmers express high-level concepts such as path names and coordinates using primitive data types. While relying too much on primitive data types is sometimes criticized as a bad smell, it is still a common practice among programmers. We propose a novel way to accurately identify expressions for certain predefined concepts by examining API calls. We defined twelve conceptual types used in the Java Standard API. We then obtained expressions for each conceptual type from 26 open source projects. Based on the expressions obtained, we trained a decision tree-based classifier. It achieved 83 % F -score for correctly predicting the conceptual type for a given expression. Our result indicates that it is possible to infer a conceptual type from a source code reasonably well once enough examples are given. The obtained classifier can be used for potential bug detection, test case generation and documentation. Yusuke Shinyama, Yoshitaka Arahori, Katsuhiko Gondow |
APSEC | 2 |
| 2020 | Risk-Aware Leak Detection at Binary LevelabstractMemory leak is a problematic bug that a heap object will never be deallocated after it is last accessed. Memory leaks can be classified into two categories in terms of their risk: high-risk leaks and low-risk ones. High-risk leaks eventually induce failures such as program clashes and performance degradation, whereas low-risk ones may not necessarily cause failures. As promising leak detectors, there have been proposed growth-sensitive or staleness-sensitive leak detectors. However, they have at least one of the major drawbacks: (1) they are not applicable to executable binaries of C/C++ programs, (2) they cannot quickly detect both high-risk and low-risk leaks in distinction, or (3) their run-time overheads are prohibitively high.This paper proposes BIGLeak, a risk-aware, efficient leak detector based on dynamic binary analysis. BIGLeak consists of three components: (1) the BIGLeak algorithm, which enables the accurate detection of high-risk leaks at binary level, (2) the intermittency analysis, which allows for the quick prediction of both high-risk and low-risk leaks in a short period of time, and (3) context-aware execution sampling, which effectively reduces the run-time overheads incurred by the BIGLeak algorithm and the intermittency analysis.Experiments with several synthetic and real programs, including lighttpd and postgres, show that BIGLeak succeeded in accurately detecting both high-risk and low-risk leaks in distinction at binary level in a short period of time. Especially, BIGLeak’s precision and recall of high-risk leak detection were much better than those of an existing staleness detector, SWAT, which works at binary level. Experimental results also indicate that BIGLeak’s run-time overheads were comparable to SWAT, despite BIGLeak’s accuracy outperformed SWAT’s. Yuta Koizumi, Yoshitaka Arahori |
PRDC | 2 |
| 2019 | Quantifying the Limitations of Learning-Assisted Grammar-Based Fuzzing
Yuma Jitsunari, Yoshitaka Arahori, Katsuhiko Gondow |
AINA | 2 |
| 2018 | Space Saving Text Input Method for Head Mounted Display with Virtual 12-key KeyboardabstractHead-Mounted Displays, or HMDs, are rapidly spreading in recent years. However, existing text input methods for HMDs have several problems, hampering their further popularization. In this paper, we focus on the two problems with existing text input methods: (1)the difficulty of its setup and (2) the space needed for its operation. To address these problems, we propose a novel text input system for HMDs. Our system (1) builds on top of a commodity camera device, Leap Motion, and (2) enables effective input in a small physical/virtual space, with a virtual 12-key keyboard. Our experimental results show the advantage of our approach over existing hand-tracking text input systems; especially our method indicates the effectiveness for Japanese text input, with a small virtual keyboard. Based on the experimental results, we believe that our text-input system is effective for any language in HMD environments, if we have a 12-key keyboard tuned for each language (as well as Japanese). Taihei Ogitani, Yoshitaka Arahori, Yusuke Shinyama, Katsuhiko Gondow |
AINA | 2 |
| 2018 | Analyzing Code Comments to Boost Program ComprehensionabstractWe are trying to find source code comments that help programmers understand a nontrivial part of source code. One of such examples would be explaining to assign a zero as a way to "clear" a buffer. Such comments are invaluable to programmers and identifying them correctly would be of great help. Toward this goal, we developed a method to discover explanatory code comments in a source code. We first propose 12 distinct categories of code comments. We then developed a decision-tree based classifier that can identify explanatory comments with 60% precision and 80% recall. We analyzed 2,000 GitHub projects that are written in two languages: Java and Python. This task is novel in that it focuses on a microscopic comment ("local comment") within a method or function, in contrast to the prior efforts that focused on API- or method-level comments. We also investigated how different category of comments is used in different projects. Our key finding is that there are two dominant types of comments: preconditional and postconditional. Our findings also suggest that many English code comments have a certain grammatical structure that are consistent across different projects. Yusuke Shinyama, Yoshitaka Arahori, Katsuhiko Gondow |
APSEC | 2 |
| 2018 | Why Do We Need the C language in Programming Courses?
Katsuhiko Gondow, Yoshitaka Arahori |
ICSOFT | 2 |
| 2018 | TCC (Tracer-Carrying Code): A Hash-based Pinpointable Traceability Tool using Copy&Paste
Katsuhiko Gondow, Yoshitaka Arahori, Koji Yamamoto 0002, Masahiro Fukuyori, Ryuichi Umekawa |
ICSOFT | 2 |
| 2018 | [Research Paper] POI: Skew-Aware Parallel Race DetectionabstractMultithreaded programs are prone to dataraces. Dataraces are known to be hard to detect and reproduce by manual effort, although they often have detrimental effects on program reliability. Automated techniques are thus demanded for detecting dataraces efficiently and precisely. There have been proposed a lot of datarace detectors so far, among which dynamic ones are promising because of their precision. However, existing dynamic race detectors incur high race-checking overheads. Even a state-of-the-art dynamic race detector, called Parallel FastTrack, fails to efficiently detect races under certain conditions, despite its attempt to parallelize race detection for efficiency. In this paper, we propose an efficient and precise parallel race detector. For our proposal, we first experimentally reveal that the load-distribution policy of Parallel FastTrack tends to skew race-checking loads to a few detection threads. We then present a simple but effective technique, called POI, for balancing race-checking loads among detection threads. POI takes race-checking loads of each detection thread into account and reduces the load skew by making each detection thread manage almost the same number of memory addresses to be checked. Experiments on several real multithreaded data-processing applications show that POI succeeded in reducing, on average, about 37% of race detection overheads, which the load-distribution policy of Parallel FastTrack would impose. Yoshitaka Sakurai, Yoshitaka Arahori, Katsuhiko Gondow |
SCAM | 2 |
| 2016 | Sequential pattern mining on electronic medical records with handling time intervals and the efficacy of medicinesabstractIt is useful to employ electronic medical records to improve medical studies. Based on their experience, medical workers conventionally prepare clinical pathways as guidelines for the typical flow for the medical treatment of each disease. In this study, we propose an approach for verifying existing clinical pathways and recommend variants or new pathways by analyzing historical records. We propose a method based on the application of sequential pattern mining to record logs with handling time intervals between treatments. We also focus on the efficacy of medicines instead of their names because various medicines have the same efficacy and they change dynamically. We evaluated the proposed method using actual logs and the results demonstrated that the proposed method is effective. Keishiro Uragaki, Tomoyuki Hosaka, Yoshitaka Arahori, Muneo Kushima, Tomoyoshi Yamazaki, Kenji Araki, Haruo Yokota |
ISCC | 3 |
| 2015 | Investigating the Difficulty of Commercial-level Compiler Warning Messages for Novice Programmers
Yoshitaka Kojima, Yoshitaka Arahori, Katsuhiko Gondow |
CSEDU (2) | 2 |
| 2010 | MieruCompiler: integrated visualization tool with "horizontal slicing" for educational compilersabstractThis paper proposes a novel visualization tool for educational compilers, called MieruCompiler. Educational compilers that generate native assembly code like i386 have many practical and pedagogical advantages, but they also have a disadvantage that the undergraduate students need to acquire a wide range of knowledge on native machine instructions, assembly directives, application binary interface (ABI), so on. To reduce this learning cost, MieruCompiler provides various visualizations as a rich internet application (RIA) including: (1) highlighting all related slices (called "horizontal slicing" after [13], but not implemented in [13]) among the source code, abstract syntax tree, assembly code, symbol table, stack layout and compiler code, when the user hovers the mouse pointer over a piece of them, (2) displaying tooltips for machine instructions, assembly directives, etc., and (3) visualizing stack layouts which are very likely to be implicit. As a preliminary evaluation, MieruCompiler was used in two universities, which produced promising results. Katsuhiko Gondow, Naoki Fukuyasu, Yoshitaka Arahori |
SIGCSE | 3 |
| 2009 | TCBC: Trap Caching Bounds Checking for CabstractIn this paper, we propose a debugging technique for C, which can dynamically find boundary errors on strings in a highly-compatible, accurate and efficient manner. The main idea of our technique is to effectively keep track of hazardous memory bounds (called trap regions) using a small table (called a trap cache) on the static section of the instrumented program. We have implemented our technique as an extension of GCC4.1.1 and conducted experiments. The results show that our technique was easily applicable even to large real programs including Apache 1.3.37 and Linux 2.6.20.4 without requiring significant manual effort, it successfully detected all of ten known boundary errors in them with no false positives, and it incurred low run-time overheads (average 17%) for their benchmarks. Yoshitaka Arahori, Katsuhiko Gondow, Hideo Maejima |
DASC | 1 |