Yukiko Yamaguchi

dblp:09/3383 · DBLP profile ↗
← Back
26ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 6 since 2021Artificial intelligence and machine learning · 10 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorSecurity and privacy · 5 · 1 since 2021
YearPublicationVenuePosition
2026 A Masked-Frozen Approach to Reliable and Secure PUF-Based Authentication
Rizka Reza Pahlevi, Hirokazu Hasegawa, Yukiko Yamaguchi, Hajime Shimada
COMPSAC3
2026 Transferability of the GCG Attack Across LLMs: Correlation with Attention Similarity and Transfer Rate Prediction
Kazuki Takaki, Hirokazu Hasegawa, Yukiko Yamaguchi, Hajime Shimada
COMPSAC3
2025 Privacy-Aware Traffic Log Anonymization Method for Realizing Both Malicious Activity Detection and Privacy
abstract
The number of cyber-attacks is still increasing, and a Network-based Intrusion Detection System (NIDS) plays an important role in countermeasures for the cyber-attacks. However, due to increases in both cyber-attacks and traffic amount, burden of supervisors who check NIDS logs also increases. To alleviate this burden, we propose a method for detecting malicious activity under anonymized traffic logs that can be reviewed by low-privilege staffs. By performing preliminary screening with these anonymized traffic logs, we can reduce the burdens of supervisors. To realize this concept, we propose a privacy-aware traffic log anonymization method. We defined privacy-sensitive features within the traffic logs and explored the importance of Gain metric analysis on LightGBM. Then, we applied simple anonymization to features that have not so large value in metrics and applied complex anonymization such as fine-grained quantization to preserve Gain of the key features. We evaluated the performance of malicious activity detection and found that it achieves over 96% performance in widely accepted metrics. Additionally, we assessed the anonymization performance and confirmed that the number of unique sessions was reduced to less than 1/10 after anonymization. Furthermore, we confirmed that the uniqueness metric after anonymization is usable as a feature in the classifier. It improves widely accepted classification performance metrics by 0.59 to 1.27 percentage points.
Takeshi Ogawa, Hajime Shimada, Hirokazu Hasegawa, Yukiko Yamaguchi
COMPSAC4
2025 An Enhanced Event-Based Dynamic Authentication Protocol Leveraging Arbiter PUF for IoT Devices
abstract
We propose a secure and lightweight authentication protocol tailored for resource-constrained Internet of Things (IoT) environments. Widely adopted approaches that store authentication data on IoT devices are inadequate, particularly in the presence of physical attacks. Building upon the Event-Based Dynamic protocol, we integrate an Enhanced Arbiter Physical Unclonable Function (PUF) to eliminate the need for static key storage and strengthen resistance against physical attacks. The proposed design introduces new Registration and Authentication phases that dynamically generate session keys using hardware-rooted entropy. Formal verification using the Tamarin Prover confirms the protocol’s correctness, mutual authentication, replay resistance, and confidentiality. Informal verification demonstrates robustness against node capture, impersonation, and guessing attacks. Simulation results show that the protocol achieves a throughput of 170–180 authentications per second, with communication overhead of 326–329 bytes and RAM usage of 212–215 bytes per client. These results confirm the protocol’s practicality and scalability for real-world IoT deployment.
Rizka Reza Pahlevi, Hirokazu Hasegawa, Yukiko Yamaguchi, Hajime Shimada
COMPSAC3
2022 Malware Detection using Attributed CFG Generated by Pre-trained Language Model with Graph Isomorphism Network
abstract
Traditional malware detection methods cannot keep up with the massive amount of newly created malware quickly and effectively. Machine learning is a promising method for the detection and classification of large-scale newly created malware according to the features of samples. The current research trend is to use machine learning technology, such as the Gradient Boosting Decision Tree (GBDT) and deep neural network technology, to learn newly created malware rapidly and accurately. We propose Control-Flow Graph (CFG)- and Graph Isomorphism Network (GIN)-based malware classification, where we first extract the CFG from portable executable (PE) files and use the large-scale pre-training language model MiniLM to generate the node features of CFG. The extracted CFG is compressed to a feature vector with GIN and classified with Multi-Layer Perceptron. To evaluate our approach, we made a CFG-based malware detection dataset from PE files of the Dike Dataset, which we call the Malware Geometric Dataset (MGD), and collected the results. The evaluation results show that our proposal demonstrated 0.9977 in the Area Under Curve metric and achieved a 97.44 % detection rate when the False Positive Rate was 0.1 %.
Hirokazu Hasegawa, Yukiko Yamaguchi, Hajime Shimada
COMPSAC3
2022 Cyber Attack Stage Tracing System based on Attack Scenario Comparison
Masahito Kumazaki, Hirokazu Hasegawa, Yukiko Yamaguchi, Hajime Shimada, Hiroki Takakura
ICISSP3
2021 Potential Security Risks of Internationalized Domain Name Processing for Hyperlink
abstract
Domain names and URLs are essential technologies in the current Internet. Thus, a failure in URL processing not only gives an inconvenience to users but also causes serious security vulnerability. If URLs are still organized with only ASCII characters, there may be no problem on URL processing. However, current URLs are further extended and complicated. One of the complexity is coming from Internationalized Domain Name (IDN) related extensions. Thus, there are no simple ways to process URLs due to their characteristics and historical extensions. In this paper, firstly, we introduce possible threats due to wrong IDN processing. Then, we present potential threats due to URL extraction operations in applications with classifying attack surfaces. We examined the above problems with various programming languages and web browsers and confirmed many issues in different environments. Furthermore, we confirmed and demonstrated that the failure pattern are not identical because the issues that come from IDN processing varies. Finally, we conclude the experimental result and propose ways to a comprehensive solution.
Taiga Shirakura, Hirokazu Hasegawa, Yukiko Yamaguchi, Hajime Shimada
COMPSAC3
2020 Quantifying the Significance of Cybersecurity Text through Semantic Similarity and Named Entity Recognition
Otgonpurev Mendsaikhan, Hirokazu Hasegawa, Yukiko Yamaguchi, Hajime Shimada
ICISSP3
2019 Identification of Cybersecurity Specific Content Using the Doc2Vec Language Model
abstract
It has become more challenging for the security analysts to identify cyber threat related content on the Internet because of the vast amount of publicly available digital texts. In this research, we proposed building an autonomous system for extracting cyber threat information from publicly available information sources. We tested a neural embedding method called doc2vec as a natural language filter for the proposed system. With cybersecurity-specific training data and custom preprocessing, we were able to train a doc2vec model and evaluate its performance. According to our evaluation, the natural language filter was able to identify cybersecurity specific natural language text with 83% accuracy.
Otgonpurev Mendsaikhan, Hirokazu Hasegawa, Yukiko Yamaguchi, Hajime Shimada
COMPSAC (1)3
2019 Rogue Wireless AP Detection using Delay Fluctuation in Backbone Network
abstract
Nowadays, wireless LAN service has been taken for granted for everyone. On the other hand, there is an increasing cyber threat in wireless LAN. For example, there is an attack called Evil-Twin Attack which places rogue access point which has the same SSID as legitimate one to make clients unknowingly connect to it. Once attacked, all of the traffic moving across the network will be eavesdropped by attackers. In this paper, we propose a method to detect rogue AP by comparing delay fluctuation of backbone network. We define delay of backbone network as the difference between ICMP travel from client to first gateway and to the Internet Server. By comparing 100 samples of backbone delay evaluation results among 5 different wireless networks, which have different backbone networks, we obtained a perspective to discriminate networks by histogram of the backbone delay.
Ziwei Zhang 0011, Hirokazu Hasegawa, Yukiko Yamaguchi, Hajime Shimada
COMPSAC (1)3
2018 Malware Detection based on HTTPS Characteristic via Machine Learning
Paul Calderon, Hirokazu Hasegawa, Yukiko Yamaguchi, Hajime Shimada
ICISSP3
2016 Evaluation on Malware Classification by Session Sequence of Common Protocols
Shohei Hiruta, Yukiko Yamaguchi, Hajime Shimada, Hiroki Takakura, Takeshi Yagi, Mitsuaki Akiyama
CANS2
2015 Malware Classification Method Based on Sequence of Traffic Flow
abstract
Network-based malware classification plays an important role in improving system security than system-based malware classification. The vast majority of malware needs a network activity in order to accomplish its purpose (e.g., downloading malware, connecting to a C&C server, etc.). Many malware classification approaches based on network behavior have thus been proposed. Nevertheless, they merely rely on either a request URL or payload for signature matching. To classify the network activity of malware, the patterns of network behavior must be understood and the changes in behavior observed. Therefore, the sequence of flows and their correlation caused by the malware should be analysed. In this paper, we present a novel malware classification method based on clustering of flow features and sequence alignment algorithms for computing sequence similarity, which represents network behavior of malware. We focus on analysing the sequence similarity between the sequence patterns of malware traffic flow generated by executing malware on the dynamic analysing system. We also performed an evaluation by using malware traffic collected from a real environment. On the basis of our experimental results, we identified the most appropriate method for classifying malware by similarity of network activity.
Hyoyoung Lim, Yukiko Yamaguchi, Hajime Shimada, Hiroki Takakura
ICISSP2
2014 A Countermeasure Recommendation System against Targeted Attacks with Preserving Continuity of Internal Networks
abstract
Recently, the sophistication of targeted cyber attacks makes conventional countermeasures useless to defend our network. Proper network design, i.e., Moderate segmentation and adequate access control, is one of the most effective countermeasures to prevent stealth activities of the attacks inside the network. By paying attention to the violation of the control, we can be aware of the existence of the attacks. In case that suspicious activities are found, we should adopt more strict design for further analysis and mitigation of damage. However, an organization must assume that its network administrators have full knowledge of its business and enough information of its network structure for selecting the most suitable design. This paper discusses a recommendation system to enhance the ability of a semi-automatic network design system previously proposed by us. Our new system evaluates on the viewpoint of two criteria, the effectiveness against malicious activities and the impact on business. The former takes the infection probability and hazardousness of communication into account and the latter considers the impact of the countermeasure which affects the organization's activities. By reviewing the candidate of the countermeasures with these criteria, the most suitable one to the organization can be selected.
Hirokazu Hasegawa, Yukiko Yamaguchi, Hajime Shimada, Hiroki Takakura
COMPSAC2
2014 Development of a Secure Traffic Analysis System to Trace Malicious Activities on Internal Networks
abstract
In contrast to conventional cyber attacks such as mass infection malware, targeted attacks take a long time to complete their mission. By using a dedicated malware for evading detection at the initial attack, an attacker quietly succeeds in setting up a front-line base in the target organization. Communication between the attacker and the base adopts popular protocols to hide its existence. Because conventional countermeasures deployed on the boundary between the Internet and the internal network will not work adequately, monitoring on the internal network becomes indispensable. In this paper, we propose an integrated sandbox system that deploys a secure and transparent proxy to analyze internal malicious network traffic. The adoption of software defined networking technology makes it possible to redirect any internal traffic from/to a suspicious host to the system for an examination of its insidiousness. When our system finds malicious activity, the traffic is blocked. If the malicious traffic is regarded as mandatory, e.g., For controlled delivery, the system works as a transparent proxy to bypass it. For benign traffic, the system works as a transparent proxy, as well. If binary programs are found in traffic, they are automatically extracted and submitted to a malware analysis module of the sandbox. In this way, we can safely identify the intention of the attackers without making them aware of our surveillance.
Soshi Hirono, Yukiko Yamaguchi, Hajime Shimada, Hiroki Takakura
COMPSAC2
2014 Unknown Attack Detection by Multistage One-Class SVM Focusing on Communication Interval
Shohei Araki, Yukiko Yamaguchi, Hajime Shimada, Hiroki Takakura
ICONIP (3)2
2013 ARIGUMA Code Analyzer: Efficient Variant Detection by Identifying Common Instruction Sequences in Malware Families
abstract
It is required in the first step of malware analysis to determine whether a given malware program is a variant of known ones. If it is surely not a variant, manual analysis against it is required. However, it is impossible to perform manual analysis, the cost of which is very high, over all the enormous number of newly found malware programs. An automatic and accurate malware program classification method should contribute to this situation. Existing methods suffer from such problems as the cost of calculating similarity between every pair of malware programs in a database, and the disability to precisely present the similarity and the difference between programs. In our approach, known malware programs are classified into families. A given malware program is determined to be a variant if it is classified into an existing family. Incremental clustering is then performed for the new one and the family, which reduces the cost of re-training and similarity calculation. Accurate comparison between programs is enabled by evaluating the difference between programs using the longest common subsequences (LCSs) of instructions. To reduce the amount of the costly calculation of LCSs, the numeric features of codes, such as cyclomatic complexity, the number of function calls and so on, are used to filter out dissimilar codes. Subsequences in the LCS of two codes are presented to malware analysts as the similarity between them, while those out of it are given as the difference. Experimental results show that this method can detect the name of APIs used in a malware which existing methods cannot, that it is useful to determine inserted codes which is used for generating variants to avoid pattern detection by anti-virus, and that it actually reduces the time to process malware programs without deteriorating the accuracy of classification.
Hirofumi Yamaki, Yukiko Yamaguchi, Hiroki Takakura
COMPSAC3
2009 Memory Complexity of Automated Trust Negotiation Strategies
Indika H. Katugampala, Hirofumi Yamaki, Yukiko Yamaguchi
PRIMA3
2006 Layered Speech-Act Annotation for Spoken Dialogue Corpus
Yuki Irie, Shigeki Matsubara, Nobuo Kawaguchi, Yukiko Yamaguchi, Yasuyoshi Inagaki
LREC4
2004 Speech understanding, dialogue management and response generation in corpus-based spoken dialogue system
abstract
This paper presents construction of a spoken dialogue system using a large-scale spoken dialogue corpus with intention tags. In this system, all of main components, such as speech understanding, dialogue management, and response generation, are constructed with corpus-based methods. An evaluation experiment using a test set has shown that the performance of the corpus-based dialogue system is improved by adding examples.
Keita Hayashi, Yuki Irie, Yukiko Yamaguchi, Shigeki Matsubara, Nobuo Kawaguchi
INTERSPEECH3
2004 Speech intention understanding based on decision tree learning
abstract
Abstract This paper proposes a method of speech intention understand-ing based on a spoken dialogue corpus to which the intentiontags are given. The intention tag expresses the task-dependentintention of the speaker, and therefore, the proper understand-ing enables a spoken dialogue system to take appropriate ac-tions. We have tagged about 35000 utterances in the CIAIR in-car speech database. In our method, several decision trees forintention understanding are constructed. By constructing deci-sion trees and using them at the same time, the strong amount ofcharacteristic features related to intentions can be retrieved, andit can also be robustly coped with the diversity of the utterances.An experiment on inference of utterance intentions has shown73.1% accuracy. 1. Introduction In order to interact with a user naturally and smoothly, it is nec-essary for a spoken dialogue system to understand the intentionof the user exactly. As a method of speech intention under-standing, example-based approaches have been considered sofar [1, 3, 6]!%In general, these approaches involve comparinga spoken utterance with examples in a correctly-tagged corpus.The intention of the utterance is regarded as the intention tag ofthe most similar example in the corpus. However, it is difficultto infer the intention of the utterance to which any example inthe corpus is not similar.This paper proposes a method of speech intention under-standing based on a spoken dialogue corpus. The method con-structs several decision trees for intention understanding. Byconstructing several decision trees, the strong amount of char-acteristic features related to intentions can be retrieved, and itcan also be robustly coped with the diversity of the utterances.So far, we have designed an organization of the tags which iscalled Layered Intention Tag(LIT). These tags show more de-tailed utterance intention rather than the illocutionary act level,and have built the corpus[2, 3]. LIT is divided into several lay-ers considering the relevance between an intention and variousphenomena relevant to an utterance, such as a style, a keyword,a sentence structure. This method constructs several decisiontrees by using this corpus and infers the intention by combiningthem.In order to evaluate the effectiveness of our method, an ex-periment on inference of the utterance intentions was conductedusing the driver utterances about restaurant search recorded ona large-scale in-car spoken dialogue corpus of CIAIR[4, 5]. Asa result, the effectiveness of the method was confirmed.
Yuki Irie, Shigeki Matsubara, Nobuo Kawaguchi, Yukiko Yamaguchi, Yasuyoshi Inagaki
INTERSPEECH4
2004 CIAIR in-car speech database
Nobuo Kawaguchi, Shigeki Matsubara, Yukiko Yamaguchi, Kazuya Takeda, Fumitada Itakura
INTERSPEECH3
2004 Example-based spoken dialogue system with online example augmentation
abstract
In this paper, we propose a new method to expand an examplebased spoken dialogue system to handle context dependent utterances. The dialogue system refers to the dialogue examples to find an example that is suitable to promote dialogue. Here, the dialogue contexts are expressed in the form of dialogue slots. By constructing dialogue examples with the text of utterances and the dialogue slots, the system handle context dependent dialogue. And we also propose a new framework of spoken dialogue, named “GROW architecture” that consists of the dialogue system and a Wizard-of-OZ (WOZ) system. By using the WOZ system to add dialogue examples via network, it becomes efficient to augment dialogue examples.
Hiroya Murao, Nobuo Kawaguchi, Shigeki Matsubara, Yukiko Yamaguchi, Kazuya Takeda, Yasuyoshi Inagaki
INTERSPEECH4
2003 Construction of an advanced in-car spoken dialogue corpus and its characteristic analysis
abstract
This paper describes an advanced spoken language corpus which has been constructed by enhancing an in-car speech database. The corpus has the following characteristic features: (1) series Advanced tag: Not only linguistic phenomena tags but also advanced discourse tags such as sentential structures, and utterance intentions, have been provided for the transcribed texts. (2) series Large-scale: The sentential structures and the intentions are currently provided for 45,053 phrases and 35,421 utterance units, respectively. (3) series Multi-layer: The corpus consists of different levels of spoken language data such as speech signals, transcribed texts, sentential structures, intentional markers and dialogue structures, moreover, they are related with each other. It allows a very wide variety of analysis of spontaneous spoken dialogue to utilize the multi-layered corpus. This paper also reports the result of investigation of the corpus, especially, focusing on the relations between the syntactic style and the intentional style of spoken utterances.
Itsuki Kishida, Yuki Irie, Yukiko Yamaguchi, Shigeki Matsubara, Nobuo Kawaguchi, Yasuyoshi Inagaki
INTERSPEECH3
2002 Example-based Speech Intention Understanding and Its Application to In-Car Spoken Dialogue System
Shigeki Matsubara, Shinichi Kimura, Nobuo Kawaguchi, Yukiko Yamaguchi, Yasuyoshi Inagaki
COLING4
1990 A neural network approach to multi-language text-to-speech system
Yukiko Yamaguchi, Tatsuro Matsumoto
ICSLP1