Kensuke Fukuda

dblp:66/1199 · DBLP profile ↗
← Back
94ranked-venue papers
17as first author
24since 2021 · last 2026
0000-0001-8372-2807ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 52 · 11 first-author · 11 since 2021Artificial intelligence and machine learning · 9Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Software engineering, systems software and programming languages · 7 · 5 since 2021Security and privacy · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3Graphics, computer vision, multimedia, augmented reality and games · 2
YearPublicationVenuePosition
2026 The Zero Trust IoT (ZT-IoT) Project
Atsuko Takefusa, Atsushi Igarashi, Taro Sekiyama, Kuniyasu Suzaki, Toshihiro Matsui, Atsuya Osaki, Naoki Yamashita, Nobuo Aoki, Sewon Park 0001, Terunobu Inaba, Lélio Brun, Yutaka Ishikawa, Kento Aida, Yasushi Ono, Kensuke Fukuda, Eisaku Sakane, Ichiro Hasuo
COMPSAC16
2026 Towards Inferring Surveillance Locations in IPv6 Networks: A probabilistic approach
Kensuke Fukuda
INFOCOM1
2026 Towards Separating Routing State Exploration from Protocol Semantics in Control Plane Verification
Ryusei Shiiba, Satoru Kobayashi, Osamu Akashi, Kensuke Fukuda
INFOCOM4
2026 Computation-Efficient Multi-Scale Multi-Granularity Framework for Mobile Traffic Forecasting Via Hierarchical Coherence
Yusheng Ji, Kensuke Fukuda
WCNC3
2026 IXP dependencies: Measuring the role of IXPs in the Internet topology
Malte Tashiro, Romain Fontugne, Kensuke Fukuda
Comput. Commun.3
2025 Automatically pinpointing original logging functions from log messages for network troubleshooting
abstract
Modern large-scale computer networks generate massive amounts of log data due to their increasing size, usage, and complexity. At the same time, as cloud-based businesses continue to grow, the need for services and software dedicated to log analysis is more important than ever. Although very useful, log messages often lack the necessary details for efficient troubleshooting, requiring extensive human analysis of the source code. In this paper, we present a new architecture designed with performance in mind, capable of identifying links between software-generated logs and their logging function calls in the source code (referred to as "origins" of the logs). The system we propose uses static code analysis to generate exact log templates, which are used to match log messages efficiently using a combination of a prefix tree and regular expressions. Our implementation SCOLM can pinpoint the origin of log messages with excellent performance and success rate. SCOLM can parse nearly 1 million log lines per minute on a single thread, with a match rate of 90 to 100% on our datasets. It outperforms the speed of traditional regex-based approaches, reducing the speed by about 98.7% in our experiments. The applications of this system are numerous, including live troubleshooting and statistical event analysis.
Gaspard Damoiseau-Malraux, Satoru Kobayashi, Kensuke Fukuda
COMPSAC3
2025 Harnessing the Power of LLMs for Code Smell Detection in Terraform Infrastructure as Code
abstract
Terraform is a widely used Infrastructure as Code (IaC) tool that simplifies cloud resource management through declarative configuration. However, Terraform configurations often exhibit code smells, which can introduce security vulnerabilities, maintainability challenges, and operational inefficiencies. While code smells have been extensively studied in other IaC platforms, research on Terraform remains limited, and existing static analysis tools struggle to detect a broad range of code smells. In this paper, we harness large language models (LLMs) for automated Terraform code smell detection, leveraging their ability to generalize beyond predefined rule-based heuristics. We construct a synthetic benchmark dataset of 42 Terraform configurations covering 14 distinct code smells and evaluate detection performance across traditional linters and LLM-based approaches. Our results show that LLMs significantly outperform static analysis tools, detecting a broader range of code smells, especially logic-related code smells. We extend our analysis to real-world repositories with 71 high-quality Terraform projects from GitHub. Our findings reveal that 84.5% of repositories contain at least one code smell. Notably, some repositories accumulate a high number of unique code smells, with one exhibiting nine distinct issues, underscoring severe quality concerns. Moreover, we find that code smells persist across repositories regardless of popularity, affecting even widely used projects with thousands of stars and forks.
Quoc-Huy Vo, Ha Dao, Kensuke Fukuda
COMPSAC3
2025 Exposed in the Pool: An Analysis of IPv6 NTP Server Scanning Activities
abstract
The NTP Pool project has become a critical infrastructure for Internet time synchronization, and used by major Linux distributions, router firmware, and IoT devices by default. While IPv6 servers are inherently difficult to locate by Internet-wide scanners due to the vast address space, public IPv6 time servers in the NTP Pool can be exposed to the Internet through GeoDNS. As these servers are operated by volunteers, they may have limited security measures, making them particularly vulnerable and attractive targets for IPv6 scanning activities. In this paper, we monitor and analyze scanning activities targeting IPv6 time servers in the NTP Pool. We present an approach to distinguish between benign scans from legitimate NTP clients and suspicious scans from potentially malicious scanners. Our analysis spans several metrics including source, traffic composition, target distribution, and scanning strategies. We find that the suspicious scanners can discover newly deployed NTP servers across different regions within several hours and probe neighboring addresses for reconnaissance. In addition, our honeynet observations suggest that the scanners employ adaptive scanning strategies that focus on responsive targets. Our findings provide comprehensive insights into scanning activities targeting IPv6 NTP Pool servers, contributing to the understanding of IPv6 network security in the critical infrastructure.
Satoru Kobayashi, Kensuke Fukuda
GLOBECOM3
2025 Interpretable Threat Detection with Evidential Classifier: The MQTT Case
abstract
We propose a universal method for threat detection with a set-valued evidential classification based on Dempster-Shafer theory and deep learning. Our approach is designed to handle the inherent uncertainty in threat detection by incorporating extended flow features and packet metadata, making it adaptable to various protocols and environments. To demonstrate its benefit, we apply the method to the threats in the MQTT protocol, which is widely used in the IoT environment for its lightweight character. Our method significantly reduces the number of false security alerts, with a zero false positive rate achieved in most cases. Additionally, the evidential framework provides interpretable outputs that allow for thorough post-event analysis, enabling a clear understanding of the evidence supporting each decision, which is crucial in security-sensitive applications.
Jaroslav Pesek, Kensuke Fukuda, Tomás Cejka
NOMS2
2025 Topology-Driven Configuration of Emulation Networks With Deterministic Templating
abstract
Network emulation is an important component of a digital twin for verifying network behavior without impacting on the service systems. Although we need to repeatedly change network topologies and configuration settings as a part of trial and error for verification, it is not easy to reflect the change without failures because the change affects multiple devices, even if it is as simple as adding a device. We present topology-driven configuration, an idea to separate network topology and generalized configuration to make it easy to change them. Based on this idea, we aim to realize a scalable, simple, and effective configuration platform for emulation networks. We design a configuration generation method using simple and deterministic config templates with a new network parameter data model, and implement it as dot2net. We evaluate three perspectives, scalability, simplicity, and efficacy, of the proposed method using dot2net through measurement and user experiments on existing test network scenarios.
Satoru Kobayashi, Ryusei Shiiba, Shinsuke Miwa, Toshiyuki Miyachi, Kensuke Fukuda
IEEE Trans. Netw. Serv. Manag.5
2024 Practical Anomaly Detection in Internet Services: An ISP centric approach
abstract
Identifying anomalies in a network is a crutial endeavor for Internet Service Providers (ISPs). Anomalies that impact the traffic of the ISP customers can lead to a degradation in the reputation of the company. Moreover, silent anomalies that do not break connectivity can impact the revenue and business of ISPs. Therefore, monitoring and anomaly detection has become essential for ISPs. In this paper, we present an ongoing research project aimed at identifying anomalies in Internet services provided by an ISP. We aim at detecting anomalies within the domain managed by the ISP that impact the customers and the business of the ISP. We propose a rule-based approach designed to promptly detect and provide reporting for such anomalies in near real time, giving information that allows the operator to identify whether a solution can be brought. In this paper, we describe the collected network telemetry metrics and illustrate how they are processed using open-source solutions. We introduce a set of use cases showing that an ISP can monitor Internet services using IETF standard metrics.
Alex Huang Feng, Pierre François, Kensuke Fukuda, Wanting Du, Thomas Graf, Paolo Lucente, Stéphane Frénot
NOMS3
2024 Following the Data Trail: An Analysis of IXP Dependencies
Malte Tashiro, Romain Fontugne, Kensuke Fukuda
PAM (2)3
2024 Exploring the Discovery Process of Fresh IPv6 Prefixes: An Analysis of Scanning Behavior in Darknet and Honeynet
Satoru Kobayashi, Kensuke Fukuda
PAM (1)3
2023 dot2net: A Labeled Graph Approach for Template-Based Configuration of Emulation Networks
abstract
Network emulation is an effective approach to ensure sustainable and reliable network services by verifying the correctness and fault tolerance of them. However, deploying and modifying emulation networks with existing platforms is time-consuming and prone to cause configuration errors because existing emulation platforms do not provide a suitable method for scalable network configuration. To overcome this problem, we propose the design and implementation of dot2net, a template-based platform for simple, scalable, and expressive configuration of emulation networks. The key idea is to separate network configuration into network topology as a labeled graph and label definitions as config template blocks. We evaluate the performance and efficiency of config file generation and show that dot2net is particularly effective at scaling the network topologies. We also demonstrate the expressiveness of dot2net for complicated networks and advanced technologies with test emulation networks of FRR, a widely used router software.
Satoru Kobayashi, Ryusei Shiiba, Ryosuke Miura, Shinsuke Miwa, Toshiyuki Miyachi, Kensuke Fukuda
CNSM6
2023 ASTrack: Automatic Detection and Removal of Web Tracking Code with Minimal Functionality Loss
abstract
Recent advances in web technologies make it more difficult than ever to detect and block web tracking systems. In this work, we propose ASTrack, a novel approach to web tracking detection and removal. ASTrack uses an abstraction of the code structure based on Abstract Syntax Trees to selectively identify web tracking functionality shared across multiple web services. This new methodology allows us to: (i) effectively detect web tracking code even when using evasion techniques (e.g., obfuscation, minification, or webpackaging); and (ii) safely remove those portions of code related to tracking purposes without affecting the legitimate functionality of the website. Our evaluation with the top 10k most popular Internet domains shows that ASTrack can detect web tracking with high precision (98%), while discovering about 50k tracking code pieces and more than 3,400 new tracking URLs not previously recognized by most popular privacy-preserving tools (e.g., uBlock Origin). Moreover, ASTrack achieved a 36% reduction in functionality loss in comparison with the filter lists, one of the safest options available. Using a novel methodology that combines computer vision and manual inspection, we estimate that full functionality is preserved in more than 97% of the websites.
Ismael Castell-Uroz, Kensuke Fukuda, Pere Barlet-Ros
INFOCOM2
2022 Advancement of Logging Monitoring Using ALOS-2/PALSAR-2 for Japanese Forest Management
abstract
Japanese local governments are increasingly utilizing forest cloud systems, and we have developed practical logging detection methodology using ALOS-2/PALSAR-2 data to improve the efficiency of forest management based on these systems. In the study in Ibaraki Prefecture, we established a method to improve both user's and producer's accuracies, and quantitatively compared the characteristics of logging detection by SAR and optical sensors. Currently, the logging information is being used in the Ibaraki Prefecture's forest cloud system.
Masato Hayashi, Takeo Tadono, Osamu Ochiai, Ko Hamamoto, Sota Hirayama, Hideki Saito 0002, Masayoshi Takahashi, Gen Takao, Takashi Yamanobe, Kazushi Matsuura, Kensuke Fukuda, Takuya Itoh
IGARSS11
2022 Characterizing DNS query response sizes through active and passive measurements
abstract
DNS has been one of the most important pieces in the current Internet. As an advanced feature, DNS provides chains of trusts for query responses from authoritative servers (i.e., DNSSEC). However, DNSSEC requires a larger payload size in a DNS query, which could yield packet fragmentation, truncation, and TCP fallback. In this paper, we characterize the DNS query response behavior from client (caching resolver) and server (authoritative server) views. For the client view, we analyze the offered maximum response sizes (EDNS0 size) from resolvers and actual response sizes at the servers of the ccTLD of jp (JP-DNS). For the server view, we characterize DNS response size distributions for different TLDs by actively querying the top 300K popular domain names in the Tranco list. The main findings of our work are as follows: (1) We confirm an increase of an EDNS0 size of 1232B from clients in Apr. 2021 despite not being so evident during the DNS flag day event in Oct. 2020. (2) Query truncation and TCP fallbacks almost occurred in an EDNS0 size of 512B, but its queries were legitimate A/AAAA records. (3) The response size distributions of popular domains are significantly different in TLDs, and median response size does not fit the minimum size (512B) in many TLDs. (4) We clarify that several issues affect the response size distribution: the ratio of signed zones and configurations (e.g., NSEC/NSEC3, signed algorithms).
Kensuke Fukuda, Yoshitaka Aharen, Shinta Sato, Takeshi Mitamura
NOMS1
2022 Comparative Causal Analysis of Network Log Data in Two Large ISPs
abstract
Towards a collaborative analysis of log data obtained from multiple networks, we first need to clarify what kind of information is available as transferable knowledge between different networks. However, we cannot directly compare net-work log data from different sources because the data largely depends on the network architecture and equipment. In this paper, we focus on relational information among network log events that follow standardized network protocols regardless of network environment. We propose a comparative analysis approach relying on causality between log time-series. In this approach, we classify log messages into anonymized log time-series with log templates, reduce the number of log time-series to decrease processing time, and apply causal discovery with the PC algorithm. To decrease the processing time of causal analysis, we propose a new preprocessing method that reduces the number of log time-series without any domain knowledge (i.e., available in any ISPs). We compare log data obtained from two nation-wide ISPs to demonstrate the effectiveness of the causal approach in comparative analysis.
Satoru Kobayashi, Keiichi Shima, Kenjiro Cho, Osamu Akashi, Kensuke Fukuda
NOMS5
2021 A Quantitative Causal Analysis for Network Log Data
abstract
Data logs from network devices are primary data to understand the current status of operational networks. However, since many and heterogeneous devices generate network logs, extracting information on the network status from such logs is not an easy task in network operation, e.g., root cause analysis of network events. Though multi-variate time-series based log analyses extract correlation structure of the logs, identifying causality of the network logs is still a complex and challenging problem. The state of the art algorithm called the PC algorithm had been applied to network log analysis, but it has two fundamental limitations; (1) Generated graphs still have many undirected edges, and (2) Edges have no weight (whether plausible causality or not). To overcome these two limitations, in this paper, we rely on MixedLiNGAM to network log analysis; This algorithm produces weighted DAGs from a set of multivariate log time series. In order to show the effectiveness of the proposed method, we apply MixedLiNGAM to a set of syslog data collected at a research and education network in Japan, and then compare output causal graphs generated by MixedLiNGAM and the PC algorithm. Our result demonstrates that obtained weighted directional edges help better understand the root cause of the network events.
Richard Jarry, Satoru Kobayashi, Kensuke Fukuda
COMPSAC3
2021 Towards Extracting Semantics of Network Config Blocks
abstract
Configuring network devices is a main task of network operators. However, understanding and consistently updating network configuration files (config) is not an easy task especially in a large-scale and complicated networks. In this paper, we propose a semantic approach to provide better understanding of such config files, different from syntax based approaches. The key idea of the work is to extract semantics of blocks of the config files by document embedding techniques in NLP. This extraction enables us to understand context of config blocks with semantic similarity metrics instead of syntax similarity ones. Furthermore, this approach can be naturally extended to additional technical documents such as vendor’s manual documents to add more specific information on the semantics of configs. We first discuss the quality of the obtained semantics for several embedding techniques, by using clustering evaluations. We next demonstrate the effectiveness of our approach with two case studies with real network configs: (1) similar config block detection and (2) automatic labeling of config block with vendor’s documents.
Kazuki Otomo, Satoru Kobayashi, Kensuke Fukuda, Osamu Akashi, Kimihiro Mizutani, Hiroshi Esaki
COMPSAC3
2021 Alternative to third-party cookies: investigating persistent PII leakage-based web tracking
abstract
Many popular websites give users the ability to sign up for their services, which requires personally identifiable information (PII). However, these websites embed third-party tracking and advertising resources, and as a consequence, the authentication flow can intentionally or unintentionally leak PII to these services. Since a user can be identified with PII, trackers can use it for tracking purposes, leading to further privacy leaks when cross-site, cross-browser, and cross-device tracking occur.
Ha Dao, Kensuke Fukuda
CoNEXT2
2021 LogDTL: Network Log Template Generation with Deep Transfer Learning
Thieu Nguyen, Satoru Kobayashi, Kensuke Fukuda
IM3
2021 Latent Semantics Approach for Network Log Analysis: Modeling and its application
Kazuki Otomo, Satoru Kobayashi, Kensuke Fukuda, Hiroshi Esaki
IM3
2021 CNAME Cloaking-Based Tracking on the Web: Characterization, Detection, and Protection
abstract
Third-party tracking on the Web has been used for collecting and correlating user's browsing behavior. Due to the increasing use of ad-blocking and third-party tracking protections, tracking providers introduced a new technique called CNAME cloaking. It misleads Web browsers into believing that a request for a subdomain of the visited website originates from this particular website, while this subdomain uses a CNAME to resolve to a tracking-related third-party domain. This technique thus circumvents the third-party targeting privacy protections. The goals of this paper are to characterize, detect, and protect the end-user against CNAME cloaking based tracking. Firstly, we characterize CNAME cloaking-based tracking by crawling top pages of the Alexa Top 300,000 sites and analyzing the usage of CNAME cloaking with CNAME blocklist, including websites and tracking providers using this technique to track users' activities. We also point out that browsers and privacy protection extensions are largely ineffective to deal with CNAME cloaking-based tracking except for Firefox with a developer's version of the uBlock Origin extension. Secondly, we propose a supervised machine learning-based approach to detect CNAME cloaking-based tracking without the on-demand DNS lookup. We show that the proposed approach outperforms well-known tracking filter lists. Finally, to circumvent the lack of DNS API in Chrome-based browsers, we design and implement a prototype of the supervised machine learning-based browser extension to detect and filter out CNAME cloaking tracking, called CNAMETracking Uncloaker. Our evaluation shows that CNAMETracking Uncloaker is able to filter out CNAME cloaking-based tracking requests without performance degradation when compared with the vanilla setting on the Chrome browser.
Ha Dao, Johan Mazel, Kensuke Fukuda
IEEE Trans. Netw. Serv. Manag.3
2020 amulog: A General Log Analysis Framework for Diverse Template Generation Methods
abstract
One of the ways to analyze unstructured log messages from large-scale IT systems is to classify log messages with log templates generated by template generation methods. However, there is currently no shared knowledge pertained to the comparison and practical use of log template generation methods because they are implemented on the basis of diverse environments. To this end, we design and implement amulog, a general log analysis framework for diverse log template generation methods. There are three key functions of amulog: (1) parsing log messages into headers and segmented messages, (2) classifying the log messages using a scalable template-matching method, and (3) storing the structured data in a database. This framework helps us easily utilize time-series data corresponding to the log templates for further analysis. We evaluate amulog with a log dataset collected from a nation-wide academic network and demonstrate that it works in a reasonable amount of time even with over 100,000 log template candidates.
Satoru Kobayashi, Yuya Yamashiro, Kazuki Otomo, Kensuke Fukuda
CNSM4
2020 Routing and Capacity Optimization Based on Estimated Latent OD Traffic Demand
abstract
This paper introduces a method to estimate latent traffic from its origin to destination based on the link packet loss rate and traffic volume. Using the estimated latent traffic, this paper also shows that we can compute the appropriate link capacity and route of packet transfer. Observed traffic might deviate from the original traffic demand and become latent when the traffic passes through congested links because of transmission control protocol (TCP) congestion control and behavioral change in the users and/or applications owing to a degraded quality of experience (QoE). The latent traffic is actualized when the congested link's capacity is improved. When link provisioning is based on observed traffic, actualized traffic might cause new congestion at other links. Thus, network providers need to estimate the origin-destination (OD) original traffic demand for network planning. Although estimation of original traffic has been researched, the estimation was only for links. In this paper, we propose a method to estimate latent origin-destination traffic by combining and expanding techniques. One approach estimates the actualized OD traffic and loss rate from the actualized traffic and packet loss rate of links. The other method estimates the latent traffic demand. Then, using the estimated value, the link capacity and routing matrix are optimized. We evaluated our method through simulation and confirmed that congestion could be avoided by capacity provisioning based on estimated latent traffic, while provisioning based on observed traffic retained the congestion. The combined method can avoid congestion with a 23% increment of capacity compared to capacity provisioning only.
Takumi Uchida, Keisuke Ishibashi, Kensuke Fukuda
COMPSAC3
2020 A machine learning approach for detecting CNAME cloaking-based tracking on the Web
abstract
Various in-browser privacy protection techniques have been designed to protect end-users from third-party tracking. In an arms race against these counter-measures, the tracking providers developed a new technique called CNAME cloaking based tracking to avoid issues with browsers that block third-party cookies and requests. To detect this tracking technique, browser extensions require on-demand DNS lookup APIs. This feature is however only supported by the Firefox browser.In this paper, we propose a supervised machine learning-based method to detect CNAME cloaking-based tracking without the on-demand DNS lookup. Our goal is to detect both sites and requests linked to CNAME cloaking-related tracking. We crawl a list of target sites and store all HTTP/HTTPS requests with their attributes. Then we label all instances automatically by looking up CNAME record of subdomain, and applying wildcard matching based on well-known tracking filter lists. After extracting features, we build a supervised classification model to distinguish site and request related to CNAME cloaking-based tracking. Our evaluation shows that the proposed approach outperforms well-known tracking filter lists: F1 scores of 0.790 for sites and 0.885 for requests. By analyzing the feature permutation importance, we demonstrate that the number of scripts and the proportion of XMLHttpRequests are discriminative for detecting sites, and the length of URL request is helpful in detecting requests. Finally, we analyze concept drift by using the 2018 dataset to train a model and obtain a reasonable performance on the 2020 dataset for detecting both sites and requests using CNAME cloaking-based tracking.
Ha Dao, Kensuke Fukuda
GLOBECOM2
2020 Towards detecting DNSSEC validation failure with passive measurements
abstract
DNSSEC is a promising technique to enhance the security of the DNS, by authenticating a chain of trusts in the DNS hierarchy. However, the deployment of DNSSEC is still on the way. One of the reasons of this under deployment is due to not enough reliability of DNSSEC validations in real world operations, especially DNSSEC validation failures due to operation misses. Thus, DNS operators require quick and reliable ways to detect such DNSSEC validation failures. The current best practice of this issue is to validate all the DNSSEC available zones periodically. However, this approach has a scalability issue. In this paper, we aim at detecting the validation failures by analyzing passively collected DNS queries at TLD-level DNS authoritative servers in a lightweight way, instead of by actively validating all the zones. In particular, we focus on changes of DNS query patterns at the authoritative server before and after DNSSEC validation failures. We conduct large-scale active and passive controlled measurements with RIPE Atlas probes and a dedicated authoritative server, to show the validity of this approach. We demonstrate that increases in DNSKEY queries are a promising candidate of metrics to detect the failures from passively collected query data. Also, we show that an increase in the number of queries is limited even for short TTL values of DNSSEC queries, thus shorter TTL is beneficial for mitigating caching effect for DNSSEC validation failures.
Kensuke Fukuda, Yoshiro Yoneya, Takeshi Mitamura
NOMS1
2019 Causal analysis of network logs with layered protocols and topology knowledge
abstract
To detect root causes of failures in large-scale networks, we need to extract contextual information from operational data automatically. Correlation-based methods are widely used for this purpose, but they have a problem of spurious correlation, which buries truly important information. In this work, we propose a method for extracting contextual information in network logs by combining a graph-based causal inference algorithm and a pruning method based on domain knowledge (i.e., network protocols and topologies). Applying the proposed method to a set of log data collected from a nation-wide R & E network, we demonstrate that the pruning method reduced processing time by 74% compared with a single-handed causal analysis method, and it detected more useful information for troubleshooting compared with an existing area-based method.
Satoru Kobayashi, Kazuki Otomo, Kensuke Fukuda
CNSM3
2019 BGP Zombies: An Analysis of Beacons Stuck Routes
Romain Fontugne, Esteban Bautista, Colin Petrie, Yutaro Nomura, Patrice Abry, Paulo Gonçalves 0001, Kensuke Fukuda, Emile Aben
PAM7
2019 A comparison of web privacy protection techniques
Johan Mazel, Richard Garnier, Kensuke Fukuda
Comput. Commun.3
2019 Adaptive probabilistic caching technique for caching networks with dynamic content popularity
Saran Tarnoi, Wuttipong Kumwilaisak, Vorapong Suppakitpaisarn, Kensuke Fukuda, Yusheng Ji
Comput. Commun.4
2018 Who Knocks at the IPv6 Door?: Detecting IPv6 Scanning
Kensuke Fukuda, John S. Heidemann
Internet Measurement Conference1
2018 Mining Causality of Network Events in Log Data
abstract
Network log messages (e.g., syslog) are expected to be valuable and useful information to detect unexpected or anomalous behavior in large scale networks. However, because of the huge amount of system log data collected in daily operation, it is not easy to extract pinpoint system failures or to identify their causes. In this paper, we propose a method for extracting the pinpoint failures and identifying their causes from network syslog data. The methodology proposed in this paper relies on causal inference that reconstructs causality of network events from a set of time series of events. Causal inference can filter out accidentally correlated events, thus it outputs more plausible causal events than traditional cross-correlation-based approaches can. We apply our method to 15 months' worth of network syslog data obtained from a nationwide academic network in Japan. The proposed method significantly reduces the number of pseudo correlated events compared with the traditional methods. Also, through three case studies and comparison with trouble ticket data, we demonstrate the effectiveness of the proposed method for practical network operation.
Satoru Kobayashi, Kazuki Otomo, Kensuke Fukuda, Hiroshi Esaki
IEEE Trans. Netw. Serv. Manag.3
2017 GML learning, a generic machine learning model for network measurements analysis
abstract
The application of machine learning models to the analysis of network measurement problems has largely increased in the last decade; however, there is still no clear best-practice or silver bullet approach to address these problems in a general context, and only adhoc and tailored approaches have been evaluated so far. While deep-learning models have provided a major breakthrough in highly-dimensional problems such as image processing, it is difficult to say today which is the best model to address the analysis of large volumes of highly-dimensional data collected in operational networks. In this paper we present a potential solution to fill this gap, exploring the application of ensemble learning models to multiple network measurement problems. We introduce GML Learning, a generic Machine Learning model for the analysis of network measurements. The GML model is a generalization of the well-known stacking approach to ensemble learning, and follows the concepts of the Super Learner model. The Super Learner performs asymptotically as well as the best input base or weak learners, providing a very powerful approach to tackle multiple problems with the same technique. In addition, it defines an approach to minimize over-fitting likelihood during training, using a variant of cross-validation. We deploy the GML model on top of Big-DAMA, a big data analytics framework for network measurement applications. We test the proposed solution in five different and assorted network measurement problems, including detection of network attacks and anomalies, QoE modeling and prediction, and Internet-paths dynamics tracking. Results confirm that the GML model provides better results than any of the single baseline models of the stack, and outperforms traditional bagging and boosting ensemble learning approaches. The GML Learning model opens the door for a generalization of a best-practice technique for the analysis of network measurements.
Pedro Casas, Juan Martin Vanerio, Kensuke Fukuda
CNSM3
2017 Adaptive and distributed monitoring mechanism in software-defined networks
abstract
Network traffic monitoring is an important factor to ensure the controllability and manageability of software-defined network (SDN). The current monitoring mechanism of SDN requires switches to request the controller for instructions to install flow entries for every new incoming flow. For finegrained monitoring, which requires many flow entries in switches' flow tables, this mechanism creates a non-trivial delay in the forwarding of switches and overhead in the control channel. Our previous work presented SDN-Mon, a monitoring framework that supports fine-grained monitoring for SDN. In this paper, we discuss the aspect of monitoring the flows in a distributed manner. We believe that a distributed monitoring capability enhances the monitoring scalability for SDN. We propose a mechanism that supports SDN to distribute the monitoring load over multiple switches in the network, in which it prevents flows monitoring duplication and balances the monitoring load over switches in the network. With the proposed mechanism, each switch handles much less monitoring load; and the overhead at switches, the control channel, and the controller caused by the monitoring duplication is eliminated. We implement the proposal and integrate it to SDN-Mon to enable a scalable and distributed monitoring capability in SDN. Experimental results show that the proposed mechanism significantly reduces the amount of monitoring load per switch, while the monitoring load is well balanced over switches in the network, with only an acceptable polling and processing overhead.
Xuan Thien Phan, Ignacio Dominguez Martinez-Casanueva, Kensuke Fukuda
CNSM3
2017 SINET5: A low-latency and high-bandwidth backbone network for SDN/NFV Era
abstract
SINET5 is a new 100-Gbps-based academic backbone network, which started full-scale operations in April 2016. It uses multi-protocol label switching-transport profile (MPLS-TP) systems and reconfigurable optical add-drop multiplexers (ROADMs) to create a nationwide network and has more than 50 backbone IP routers to provide a wide range of services, such as several virtual private network (VPN) services. It provides end-to-end data communications up to 100 Gbps throughput, minimized-latency, and software-defined networking (SDN)-friendly functions to researchers in every Japanese prefecture. SINET5 is also a platform for dynamic inter-cloud connections and network functions visualization (NFV) services. This paper brief review of the network architecture, and describes new featured services, SDN-oriented layer-2 on-demand VPN services, and NFV functions. Field test results for SINET5 performance are also reported.
Takashi Kurimoto, Shigeo Urushidani, Kenjiro Yamanaka, Motonori Nakamura, Shunji Abe, Kensuke Fukuda, Michihiro Koibuchi, Hiroki Takakura, Shigeki Yamada, Yusheng Ji
ICC7
2017 Mining causes of network events in log data with causal inference
abstract
Network log message (e.g., syslog) is valuable information to detect unexpected or anomalous behavior in a large scale network. However, pinpointing failures and their causes is not an easy problem because of a huge amount of system log data in daily operation. In this study, we propose a method extracting failures and their causes from network syslog data. The main idea of the method relies on causal inference that reconstructs causality of network events from a set of the time series of events. Causal inference allows us to reduce the number of correlated events by chance, thus it outputs more plausible causal events than a traditional cross-correlation based approach. We apply our method to 15 months network syslog data obtained in a nation-wide academic network in Japan. Our method significantly reduces the number of pseudo correlated events compared with the traditional method. Also, through two case studies and comparison with trouble ticket data, we demonstrate the effectiveness of our method for network operation.
Satoru Kobayashi, Kensuke Fukuda, Hiroshi Esaki
IM2
2017 On rate limitation mechanisms for TCP throughput: A longitudinal analysis
João Araújo 0003, Raul Landa, Richard G. Clegg, George Pavlou, Kensuke Fukuda
Comput. Networks5
2017 Scaling in Internet Traffic: A 14 Year and 3 Day Longitudinal Study, With Multiscale Analyses and Random Projections
abstract
In the mid 1990s, it was shown that the statistics of aggregated time series from Internet traffic departed from those of traditional short range-dependent models, and were instead characterized by asymptotic self-similarity. Following this seminal contribution, over the years, many studies have investigated the existence and form of scaling in Internet traffic. This contribution first aims at presenting a methodology, combining multiscale analysis (wavelet and wavelet leaders) and random projections (or sketches), permitting a precise, efficient and robust characterization of scaling, which is capable of seeing through non-stationary anomalies. Second, we apply the methodology to a data set spanning an unusually long period: 14 years, from the MAWI traffic archive, thereby allowing an in-depth longitudinal analysis of the form, nature, and evolutions of scaling in Internet traffic, as well as network mechanisms producing them. We also study a separate three-day long trace to obtain complementary insight into intra-day behavior. We find that a biscaling (two ranges of independent scaling phenomena) regime is systematically observed: long-range dependence over the large scales, and multifractallike scaling over the fine scales. We quantify the actual scaling ranges precisely, verify to high accuracy the expected relationship between the long range dependent parameter and the heavy tail parameter of the flow size distribution, and relate fine scale multifractal scaling to typical IP packet inter-arrival and to round-trip time distributions.
Romain Fontugne, Patrice Abry, Kensuke Fukuda, Darryl Veitch, Kenjiro Cho, Pierre Borgnat, Herwig Wendt
IEEE/ACM Trans. Netw.3
2017 Detecting Malicious Activity With DNS Backscatter Over Time
abstract
Network-wide activity is when one computer (the originator) touches many others (the targets). Motives for activity may be benign (mailing lists, content-delivery networks, and research scanning), malicious (spammers and scanners for security vulnerabilities), or perhaps indeterminate (ad trackers). Knowledge of malicious activity may help anticipate attacks, and understanding benign activity may set a baseline or characterize growth. This paper identifies domain name system (DNS) backscatter as a new source of information about network-wide activity. Backscatter is the reverse DNS queries caused when targets or middleboxes automatically look up the domain name of the originator. Queries are visible to the authoritative DNS servers that handle reverse DNS. While the fraction of backscatter they see depends on the server's location in the DNS hierarchy, we show that activity that touches many targets appear even in sampled observations. We use information about the queriers to classify originator activity using machine-learning. Our algorithm has reasonable accuracy and precision (70-80%) as shown by data from three different organizations operating DNS servers at the root or country level. Using this technique, we examine nine months of activity from one authority to identify trends in scanning, identifying bursts corresponding to Heartbleed, and broad and continuous scanning of secure shell.
Kensuke Fukuda, John S. Heidemann, Abdul Qadeer
IEEE/ACM Trans. Netw.1
2016 Non-linear regression for bivariate self-similarity identification - application to anomaly detection in Internet traffic based on a joint scaling analysis of packet and byte counts
abstract
Internet traffic monitoring is a crucial task for network security. Self-similarity, a key property for a relevant description of internet traffic statistics, has already been massively and successfully involved in anomaly detection. Self-similar analysis was however so far applied either to byte or Packet count time series independently, while both signals are jointly collected and technically deeply related. The present contribution elaborates on a recently proposed multivariate self-similar model, Operator fractional Brownian Motion (OfBm), to analyze jointly self-similarity in bytes and packets. A non-linear regression procedure, based on an original Branch & Bound resolution procedure, is devised for the full identification of bivariate OfBm. The estimation performance is assessed by means of Monte Carlo simulations. Further, an Internet traffic anomaly detection procedure is proposed, that makes use of the vector of Hurst exponents underlying the OfBm based Internet data modeling. Applied to a large set of high quality and modern Internet data from the MAWI repository, proof-of-concept results in anomaly detection are detailed and discussed.
Jordan Frécon, Romain Fontugne, Gustavo Didier, Nelly Pustelnik, Kensuke Fukuda, Patrice Abry
ICASSP5
2016 Machine learning, data mining and Big Data frameworks for network monitoring and troubleshooting
Alessandro D'Alconzo, Pere Barlet-Ros, Kensuke Fukuda, David R. Choffnes
Comput. Networks3
2015 Enhancing the Performance of Mobile Traffic Identification with Communication Patterns
abstract
Traffic classification is important especially for managing and monitoring networks which contain a wide variety of traffic, such as in the mobile network. Using only packet based feature in traditional classification is not enough for classifying mobile application traffic because of the complexity of mobile traffic. Therefore, this study proposes the technique that combines the packet size distribution and communication patterns extracted via graph let for identifying mobile application. The technique is robust to the complexity of mobile traffic and has no privacy concerns. Validation results over five popular mobile applications (Facebook, Line, Skype, You Tube, and Web) demonstrate that our combined method achieves high performance (0.95) of F-measure even using only randomly sampled 50 packets during 3-minute time interval. Moreover, the combination of these features distinguishes various applications with similar characteristics such as Facebook and Web.
Sophon Mongkolluksamee, Vasaka Visoottiviseth, Kensuke Fukuda
COMPSAC3
2015 Random projection and multiscale wavelet leader based anomaly detection and address identification in internet traffic
abstract
We present a new anomaly detector for data traffic, ‘SMS’, based on combining random projections (sketches) with multiscale analysis, which has low computational complexity. The sketches allow ‘normal’ traffic to be automatically and robustly extracted, and anomalies detected, without the need for training data. The multiscale analysis extracts statistical descriptors, using wavelet leader tools developed recently for multifractal analysis, without any need for timescales to be selected a priori. The proposed detector is illustrated using a large recent dataset of Internet backbone traffic from the MAWI archive, and compared against existing detectors.
Romain Fontugne, Patrice Abry, Kensuke Fukuda, Pierre Borgnat, Johan Mazel, Herwig Wendt, Darryl Veitch
ICASSP3
2015 Tracking the Evolution and Diversity in Network Usage of Smartphones
abstract
We analyze the evolution of smartphone usage from a dataset obtained from three, 15-day-long, user-side, measurements with over 1500 recruited smartphone users in the Greater Tokyo area from 2013 to 2015. This dataset shows users across a diverse range of networks; cellular access (3G to LTE), WiFi access (2.4 to 5GHz), deployment of more public WiFi access points (APs), as they use diverse applications such as video, file synchronization, and major software updates.
Kensuke Fukuda, Hirochika Asai, Kenichi Nagami
Internet Measurement Conference1
2015 Detecting Malicious Activity with DNS Backscatter
abstract
Network-wide activity is when one computer (the originator) touches many others (the targets). Motives for activity may be benign (mailing lists, CDNs, and research scanning), malicious (spammers and scanners for security vulnerabilities), or perhaps indeterminate (ad trackers). Knowledge of malicious activity may help anticipate attacks, and understanding benign activity may set a baseline or characterize growth. This paper identifies DNS backscatter as a new source of information about network-wide activity. Backscatter is the reverse DNS queries caused when targets or middleboxes automatically look up the domain name of the originator. Queries are visible to the authoritative DNS servers that handle reverse DNS. While the fraction of backscatter they see depends on the server's location in the DNS hierarchy, we show that activity that touches many targets appear even in sampled observations. We use information about the queriers to classify originator activity using machine-learning. Our algorithm has reasonable precision (70-80%) as shown by data from three different organizations operating DNS servers at the root or country-level. Using this technique we examine nine months of activity from one authority to identify trends in scanning, identifying bursts corresponding to Heartbleed and broad and continuous scanning of ssh.
Kensuke Fukuda, John S. Heidemann
Internet Measurement Conference1
2015 An empirical mixture model for large-scale RTT measurements
abstract
Monitoring delays in the Internet is essential to understand the network condition and ensure the good functioning of time-sensitive applications. Large-scale measurements of round-trip time (RTT) are promising data sources to gain better insights into Internet-wide delays. However, the lack of efficient methodology to model RTTs prevents researchers from leveraging the value of these datasets. In this work, we propose a log-normal mixture model to identify, characterize, and monitor spatial and temporal dynamics of RTTs. This data-driven approach provides a coarse grained view of numerous RTTs in the form of a graph, thus, it enables efficient and systematic analysis of Internet-wide measurements. Using this model, we analyze more than 13 years of RTTs from about 12 millions unique IP addresses in passively measured backbone traffic traces. We evaluate the proposed method by comparison with external data sets, and present examples where the proposed model highlights interesting delay fluctuations due to route changes or congestion. We also introduce an application based on the proposed model to identify hosts deviating from their typical RTTs fluctuations, and we envision various applications for this empirical model.
Romain Fontugne, Johan Mazel, Kensuke Fukuda
INFOCOM3
2014 A longitudinal analysis of Internet rate limitations
abstract
TCP remains the dominant transport protocol for Internet traffic, but the preponderance of its congestion control mechanisms in determining flow throughput is often disputed. This paper analyzes the extent to which network, host and application settings define flow throughput over time and across autonomous systems. Drawing from a longitudinal study spanning five years of passive traces collected from a single transit link, our results show that continuing OS upgrades have reduced the influence of host limitations owing both to windowscale deployment, which by 2011 covered 80% of inbound traffic, and increased socket buffer sizes. On the other hand, we show that for this data set, approximately half of all inbound traffic remains throttled by constraints beyond network capacity, challenging the traditional model of congestion control in TCP traffic as governed primarily by loss and delay.
João Araújo 0003, Raul Landa, Richard G. Clegg, George Pavlou, Kensuke Fukuda
INFOCOM5
2014 Towards a taxonomy of darknet traffic
abstract
Darknets can be used to monitor unexpected network traffic destined for allocated but unused IP address blocks, thus providing an effective traffic measurement technique for viewing certain remote network security events. Past works in this field discussed the possible causes (events) of darknet traffic and applied their classification schemes on short-range traces. Our interest lies, however, in how darknets have evolved since those works and the effectiveness of a darknet taxonomy for real long-range traffic. We thus propose a simple but effective taxonomy of darknet traffic, on the basis of observations, and evaluate it on real darknet traces covering six years. The evaluation results show that we can detect and label anomalous events defined by the taxonomy for over 96% of all sources, making the unlabeled source rate extremely low. We also obtain some interesting findings on the evolution of different anomalous events since 2006 (especially in recent years), determine the most appropriate time bin for traffic analysis of our traces, and highlight the general applicability of our taxonomy on different darknet datasets. Finally, we conclude that most sources in our traces are characterized by just one or two events with simple attack mechanisms.
Kensuke Fukuda
IWCMC2
2014 A taxonomy of anomalies in backbone network traffic
abstract
The potential threat of network anomalies on Internet has led to a constant effort by the research community to design reliable detection methods. Detection is not enough, however, because network administrators need additional information on the nature of events occurring in a network. Several works try to classify detected events or establish a taxonomy of known events. But, these works are non-overlapping in terms of anomaly type coverage. On the one hand, existing classification methods use a limited set of labels. On the other hand, taxonomies often target a single type of anomaly or, when they have wider scope, fail to present the full spectrum of what really happens in the wild. We thus present a new taxonomy of network anomalies with wide coverage of existing work. We also provide a set of signatures that assign taxonomy labels to events. We present a preliminary study applying this taxonomy with six years of real network traffic from the MAWI repository. We classify previously documented anomalous events and draw to main conclusions. First, the taxonomy-based analysis provides new insights regarding events previous classified by heuristic rule labeling. For example, some RST events are now classified as network scan response and the majority of ICMP events are split into network scans and network scan responses. Moreover, some previously unknown events now account for a substantial number of all UDP network scans, network scan responses and port scans. Second, the number of unknown events decreases from 20 to 10% of all events with the proposed taxonomy as compared to the heuristic approach.
Johan Mazel, Romain Fontugne, Kensuke Fukuda
IWCMC3
2014 Towards evaluation of DNS server selection with geodesic distance
abstract
The DNS (Domain Name Service) is the most fundamental service in the current Internet. Investigating macroscopic performance of DNS server selection algorithm is not an easy task while microscopic software implementations are known. This is mainly because of lack of available performance information (e.g., delay) from authoritative server side. In this paper we estimate the global performance of DNS server selection with passively collected DNS query data at authoritative server. A key idea towards the estimation is to use geodesic path length between authoritative server and cache resolver based on their geolocation information as an approximation of end-to-end delay. We introduce two performance metrics to quantify the efficiency of the server selection: RML (the ratio to the minimum path length) and RBL (the ratio to the best path length). The former corresponds to an additional cost of cache resolver deviating from the minimum path length to the server where the resolver actually accessed, and the latter indicates an extra cost deviating from the best path length to the closest (ideal) server ignoring the complexity of the intra and inter AS level topology. We analyze 1-day long DNS queries from over 700 K cache resolvers all over the world to “.jp” TLD servers (JP DNS servers) and demonstrate that 75% of cache resolvers in Japan and 40-60% of cache resolvers in other regions selected the closest server found by their probe, as the most accessed server. Thus, 25% or 60% of cache resolvers could have found a more appropriate authoritative server with a more adaptive selection algorithm. Similarly, 15-35% of resolvers in other regions efficiently accessed the best cost servers as the most accessed server. This means that there is still a room for improvement to reduce the potential delay by change of the infrastructure (i.e., new peering links and deployment of new servers).
Kensuke Fukuda, Shinta Sato, Takeshi Mitamura
NOMS1
2013 An Analysis of Players and Bots Behaviors in MMORPG
abstract
"Bot", automatic robot software, has been one of most serious problems in Massively Multiplayer Online Role Playing Game (MMORPG). Bots earn much more money in a virtual world than human players, and finally collapse balance and fairness of the MMORPG. At present, many techniques have been applied to detect and exterminate bots. However, they have a common problem that the technique effective in a certain MMORPG is not equally effective in other MMORPGs, thus no general method to detect bots has been established. Toward establishing such general technique, we analyze behavioral patterns of human players and bots in server-side game log data with two commonly available features characterizing users (human players and bots): location-based information (i.e., speed of players) and action frequency information (i.e., action count per fixed time slot). The main findings of our analysis are as follows: (1) The variation of the speed of bots is smaller than that of humans (i.e, the movement of bots is more efficient). However, the discriminative power of this location-based feature is not so significant in MMORPG, while it showed good performance in First Person Shooting (FPS) game. (2) Action count and battle count indicate more discriminative power than the speed feature. In particular, the action count is more robust than battle count against the size of the time slot.
Yutaro Mishima, Kensuke Fukuda, Hiroshi Esaki
AINA2
2013 Nine years of observing traffic anomalies: Trending analysis in backbone networks
Youngjoon Won, Romain Fontugne, Kenjiro Cho, Hiroshi Esaki, Kensuke Fukuda
IM5
2013 A technique for counting DNSSEC validators
abstract
The DNS security extensions (DNSSEC) is a new feature of DNS that provides an authentication mechanism that is now being deployed worldwide. However, we do not have enough knowledge about the deployment status of DNSSEC in the wild due to the difficulty of identifying DNSSEC validators (caching validating resolvers). In this paper, a simple and robust method is proposed that estimates DNSSEC validators from DNS query data passively measured at the server side. The key idea of the estimation method relies on the query patterns of the original query and the DNSSEC queries triggered by the original query, which is the ratio of the number of DS queries to the number of total queries per host (DSR: DS ratio). To show the effectiveness of the proposed method, we analyze passive traffic traces measured for all the “.jp” servers and actively send DNSSEC validation requests to caching resolvers that appear in the traces to obtain the ground truth data of DNSSEC validators. Our results of the active measurement reveal that less than 50% of the potential DNSSEC validators were validating caching resolvers in the wild; the remainder was related to stub validators (e.g., browser plugins) behind non-validating caching resolvers. Thus, simple IP address-based counts overestimated the number of DNSSEC validators in an investigation of the deployment of DNSSEC at the organization level (e.g., ISPs). Then, we demonstrate the effectiveness of the DSR by using the active and passive traffic traces. In summary, the ratio of validating caching resolvers in our dataset was estimated to be approximately 70% of the potential DNSSEC validators, and also 15-20% of the ASes sending DNSSEC queries were overestimated as ones with validating caching resolvers. In particular, our results show that some ASes providing public DNS service had few validating caching resolvers though they had a large number of hosts sending DNSSEC queries.
Kensuke Fukuda, Shinta Sato, Takeshi Mitamura
INFOCOM1
2013 Strip, bind, and search: a method for identifying abnormal energy consumption in buildings
abstract
A typical large building contains thousands of sensors, monitoring the HVAC system, lighting, and other operational sub-systems. With the increased push for operational efficiency, operators are relying more on historical data processing to uncover opportunities for energy-savings. However, they are overwhelmed with the deluge of data and seek more efficient ways to identify potential problems. In this paper, we present a new approach called the Strip, Bind and Search (SBS); a method for uncovering abnormal equipment behavior and in-concert usage patterns. SBS uncovers relationships between devices and constructs a model for their usage pattern relative to other devices. It then flags deviations from the model. We run SBS on a set of building sensor traces; each containing hundred sensors reporting data flows over 18 weeks from two separate buildings with fundamentally different infrastructures. We demonstrate that, in many cases, SBS uncovers misbehavior corresponding to inefficient device usage that leads to energy waste. The average waste uncovered is as high as 2500~kWh per device.
Romain Fontugne, Jorge Ortiz 0001, Nicolas Tremblay, Pierre Borgnat, Patrick Flandrin, Kensuke Fukuda, David E. Culler, Hiroshi Esaki
IPSN6
2013 PopCache: Cache more or less based on content popularity for information-centric networking
abstract
Due to a mismatch between downloading and caching content, the network may not gain significant benefit from the sophisticated in-network caching of information-centric networking (ICN) architectures by using a basic caching mechanism. This paper aims to seek an effective caching decision policy to improve the content dissemination in ICN. We propose PopCache-a caching decision policy with respect to the content popularity-that allows an individual ICN router to cache content more or less in accordance with the popularity characteristic of the content. We propose an analytical model to evaluate the performance of different caching decision policies in terms of the server-hit rate and expected round-trip time. The analysis confirmed by simulation results shows that PopCache yields the lowest expected round-trip time compared with three benchmark caching decision policies, i.e., the always, fixed probability and path-capacity-based probability, and PopCache provides the server-hit rate comparable to the lowest ones.
Kalika Suksomboon, Saran Tarnoi, Yusheng Ji, Michihiro Koibuchi, Kensuke Fukuda, Shunji Abe, Motonori Nakamura, Michihiro Aoki, Shigeo Urushidani, Shigeki Yamada
LCN5
2013 A Measurement of Mobile Traffic Offloading
Kensuke Fukuda, Kenichi Nagami
PAM1
2013 ADMIRE: Anomaly detection method using entropy-based PCA with three-step sketches
Yoshiki Kanda, Romain Fontugne, Kensuke Fukuda, Toshiharu Sugawara
Comput. Commun.3
2013 Synoptic Graphlet: Bridging the Gap Between Supervised and Unsupervised Profiling of Host-Level Network Traffic
abstract
End-host profiling by analyzing network traffic comes out as a major stake in traffic engineering. Graphlet constitutes an efficient and common framework for interpreting host behaviors, which essentially consists of a visual representation as a graph. However, graphlet analyses face the issues of choosing between supervised and unsupervised approaches. The former can analyze a priori defined behaviors but is blind to undefined classes, while the latter can discover new behaviors at the cost of difficult a posteriori interpretation. This paper aims at bridging the gap between the two. First, to handle unknown classes, unsupervised clustering is originally revisited by extracting a set of graphlet-inspired attributes for each host. Second, to recover interpretability for each resulting cluster, a synoptic graphlet, defined as a visual graphlet obtained by mapping from a cluster, is newly developed. Comparisons against supervised graphlet-based, port-based, and payload-based classifiers with two datasets demonstrate the effectiveness of the unsupervised clustering of graphlets and the relevance of the a posteriori interpretation through synoptic graphlets. This development is further complemented by studying evolutionary tree of synoptic graphlets, which quantifies the growth of graphlets when increasing the number of inspected packets per host.
Yosuke Himura, Kensuke Fukuda, Kenjiro Cho, Pierre Borgnat, Patrice Abry, Hiroshi Esaki
IEEE/ACM Trans. Netw.2
2012 Efficient query bundling mechanism in a DHT network
abstract
A distributed hash table (DHT) network can be used for many distributed services and systems. In DHT networks, it takes logN look-up steps to search for required data where N is the number of nodes. However, the look-up process is redundant in the IP network because each look-up process generates a lot of communication among nodes. In massive data management such as sensor and web information management, this results in high network load even if each the search process takes only logN look-up steps. To solve this problem, we propose an efficient query bundling mechanism that makes it possible to bundle multiple queries by using range information. Range information consists of ID space information kept by a node. When a source node receives range information from a destination node, the source node matches all queries for the range information and forwards queries matching the range information to the destination node directly. This effectively reduces the number of look-up processes and the network load for the IP network. In addition, our mechanism can be implemented into conventional DHT networks and can easily be combined to effective DHT routing algorithms such as Chord, Kademlia, and Pastry. In evaluation, we implement our mechanism into DHT networks and compare its performance with that of conventional query bundling mechanisms. The results show that our mechanism reduces by up to 75% the total number of forwarding operations to put data compared with other mechanisms. In addition, our mechanism realizes the reduction of the number of forwarding operations per look-up process by up to 85% compared to other mechanisms.
Kimihiro Mizutani, Toru Mano, Osamu Akashi, Kensuke Fukuda
GLOBECOM4
2012 Counting NATted hosts by observing TCP/IP field behaviors
abstract
With the prevalence of Network Address Translation (NAT), identifying a number of Internet users becomes a challenging task because many users share the same public IP address. This paper proposes a passive technique for estimating a number of Internet hosts sharing the same IP address, i.e., NATted hosts. Previous work by Bellovin [1] counted NATted hosts by observing a sequence of IPID fields in IP header. This technique only works on some operating systems with a global counter for the IPID sequence (e.g., Windows). Other operating systems that implement the IPID sequence on a per-flow or a random basis are not detected. The proposed technique overcomes this limitation by observing patterns of the TCP sequence number and the TCP source port, in addition to the IPID sequence. Our technique demonstrates more accurate estimate than the previous work in controlled experiments. Moreover, applying our technique on a collection of longitudinal traffic traces measured at a trans-Pacific link in 2001-2010, we find that the percentage of the NATted hosts is stably less than 2% over years.
Sophon Mongkolluksamee, Kensuke Fukuda, Panita Pongpaibool
ICC2
2011 On the use of weighted syslog time series for anomaly detection
abstract
Finding unusual (anomalous) events in a network is a crucial task for network operation. The collection and analysis of the log messages from the network devices (i.e., router and switch) are a common method. However, the detection of the anomalous events is painful because of a huge amount and many types of log messages; the important event messages are sometimes hidden in a large number of relatively less important and usual event messages. Many researches have been devoted to find out anomalies accurately and quickly by using time series of log message. To construct the log time series for anomaly detection in order to highlight such hidden anomalies, this paper focuses on the effectiveness of using a global weight that is based on a global appearance of a message type in the all data set. We introduce two types of the global weight called residual inverse document frequency (RIDF) and entropy that are well-known method in information retrieval field. Then, we evaluate the performance improvement due to the global weights with a wavelet-based abrupt change detection algorithm through 1-year long collection of core router log messages taken from a Japanese R&E network. Our main findings are (1) the global weights assign a low weight to less important and frequently appeared event messages, (2) they highlight unusual events by removing temporal correlation (i.e., periodic trend), and (3) the wavelet-based detection algorithm more accurately alarms the anomalous time period with the weighted time series.
Kensuke Fukuda
Integrated Network Management1
2010 A PCA Analysis of Daily Unwanted Traffic
abstract
This paper investigates the macroscopic behavior of unwanted traffic (e.g., virus, worm, backscatter of (D)DoS or misconfiguration) passing through the Internet. The data set we used are unwanted packets measured at /18 darknet in Japan from Oct. 2006 to Apr. 2009 that included the recent Conficker outbreak. The traffic behavior is quantified by the entropy of ten packet features (e.g., 5-tuple). Then, we apply PCA (principal component analysis) to a ten dimensional entropy time series matrix to obtain a suitable representation of unwanted traffic. PCA is a well-known and studied method for finding out normal and anomalous behaviors in Internet backbone traffic, however, few studies applied it to darknet traffic. We first demonstrate the high variability nature of the entropy time series for ten packet features. Next, we show that the top four principal components are sufficiently enough to describe the original traffic behavior. In particular, the first component can be interpreted as the type of unwanted traffic (i. e., worm/virus or scanning), and the second one as the difference in communication patterns (e. g., one-to-many or many-to-one). Those two components account for 63.8\% of the original data set in terms of the total variance. On the other hand, the outliers in the higher components indicate the presence of specific anomalies although most of mapped data to the components have less variability. Furthermore, we show that the scatter plot of the first and second principal component scores provides us with a better view of the macroscopic unwanted traffic behavior.
Kensuke Fukuda, Toshio Hirotsu, Osamu Akashi, Toshiharu Sugawara
AINA1
2010 Difficulties of identifying application type in backbone traffic
abstract
A flow identification in backbone traffic is a crucial problem for network management to provide better service to the users. However, application traffic on backbone is in general hard to be identified because of the amount of traffic volume, asymmetric nature of the route, and difficulty of capturing full payload. In this paper, we present our preliminary results of the basic ability and limitation of four well-known flow identification algorithms (port-based heuristics and payload-based one) applied to packet trace called MAWI trace measured at a trans-pacific link. First, we show that the identification ratio of the payload-based algorithms is 40-60% in flow level, however the accuracy of the identification in the payload-based algorithms strongly depends on the algorithm and the definition of the rules themselves. Next, we found that only 3% of the traffic is commonly identified by the payload-based algorithms when we apply them to unknown traffic reported by the port-based heuristics. Finally, we evaluate the dependency of the flow size on the identification ratio. Those results emphasize the need for a more accurate and available identification algorithm.
Kensuke Fukuda
CNSM1
2010 MAWILab: combining diverse anomaly detectors for automated anomaly labeling and performance benchmarking
abstract
Evaluating anomaly detectors is a crucial task in traffic monitoring made particularly difficult due to the lack of ground truth. The goal of the present article is to assist researchers in the evaluation of detectors by providing them with labeled anomaly traffic traces. We aim at automatically finding anomalies in the MAWI archive using a new methodology that combines different and independent detectors. A key challenge is to compare the alarms raised by these detectors, though they operate at different traffic granularities. The main contribution is to propose a reliable graph-based methodology that combines any anomaly detector outputs. We evaluated four unsupervised combination strategies; the best is the one that is based on dimensionality reduction. The synergy between anomaly detectors permits to detect twice as many anomalies as the most accurate detector, and to reject numerous false positive alarms reported by the detectors. Significant anomalous traffic features are extracted from reported alarms, hence the labels assigned to the MAWI archive are concise. The results on the MAWI traffic are publicly available and updated daily. Also, this approach permits to include the results of upcoming anomaly detectors so as to improve over time the quality and variety of labels.
Romain Fontugne, Pierre Borgnat, Patrice Abry, Kensuke Fukuda
CoNEXT4
2010 Adaptive probabilistic task allocation in large-scale multi-agent systems and its evaluation
abstract
In this paper, we introduce the probabilistic awardee selection strategy, under which awardee is selected with a fixed probability, into the award phase of contract net protocol. We then point out that, by changing the probabilities in this strategy according the local workload, the overall performance can be considerably improved.
Toshiharu Sugawara, Kensuke Fukuda, Toshio Hirotsu, Satoshi Kurihara
GECCO2
2010 Evaluation of Anomaly Detection Based on Sketch and PCA
abstract
Using traffic random projections (sketches) and Principal Component Analysis (PCA) for Internet traffic anomaly detection has become popular topics in the anomaly detection fields, but few studies have been undertaken on the subjective and quantitative comparison of multiple methods using the data traces open to the community. In this paper, we propose a new method that combines sketches and PCA to detect and identify the source IP addresses associated with the traffic anomalies in the backbone traces measured at a single link. We compare the results with those of a method incorporating sketches and multi-resolution gamma modeling using the trans-Pacific link traces. The comparison indicates that each method has its own advantages and disadvantages. Our method is good at detecting worm activities with many packets, whereas the gamma method is good at detecting scan activities for peer hosts with only a few packets, but it reports many false positives for traces of worm outbreaks. Therefore, their use in combination would be effective. We also examined the impact of adaptive decision making on a parameter (the number of normal subspaces in PCA) on the basis of the cumulative proportion of each sketched traffic and conclude that it performs at a higher level than the previous method deciding only on one specific value of the parameter for every divided traffics.
Yoshiki Kanda, Kensuke Fukuda, Toshiharu Sugawara
GLOBECOM2
2010 Probabilistic Award Strategy for Contract Net Protocol in Massively Multi-agent Systems
Toshiharu Sugawara, Toshio Hirotsu, Kensuke Fukuda
ICAART (2)3
2010 Estimating Speed of Scanning Activities with a Hough Transform
abstract
In this paper, we propose a method to detect scanning activities in darknet traffic and to estimate their speed of change in time and feature space (e.g., destination address, source port, or destination port). The main idea of the algorithm relies on an image processing technique applied to a two-dimensional image that represents unwanted traffic. Thus, on the two-dimensional image, packets are represented as pixels in the time and feature coordinates, and unwanted activity as a set of pixels. The use of a Progressive Probabilistic Hough Transform (PPHT) that is a known technique to detect edges in an image enables us to detect such unwanted activities as ``lines'' in a traffic trace. We apply our method to darknet traffic traces for three years to investigate the property of such unwanted activities. Our main findings are following: In destination IP address space we confirmed typical host scanning speeds (i.e., a slanted line in the image) although the most of activities are characterized by intensive scans to a specific host (i.e., a horizontal line). Also, we confirmed few port scanning over wide destination port space, meaning that a targeted port attack is dominant in the current network. On the other hand, the consecutive change of source port was also observed; those activities are not tracked by other features. We obtain that 80-90\% of unique source IP addresses appeared in the trace is confirmed by this method. Thus, most unwanted activities is still characterized by some kind of trajectory to be detected in packet feature space, though the rest of them behaves like ``noise''.
Kensuke Fukuda, Romain Fontugne
ICC1
2010 A Flow Analysis for Mining Traffic Anomalies
abstract
Although analyzing anomalous network traffic behavior is a popular research topic, few studies have been undertaken on the analysis of communication pattern per host based on their flows to characterize the anomalous Internet traffic. This paper discusses the possibility of using a flow-based communication pattern per host as a metric to identify anomalies. The key idea underlining our method is that scanning worm-infected hosts reveal the intrinsic characteristics of host's communication pattern and such patterns are distinguishable from those of other hosts. In particular, we found that scanning of worm-infected hosts that generated a lot of flows revealed the intrinsic communication pattern and the pattern could be classified from those of other hosts by k-means clustering. We also found that our flow-based metric could isolate the anomalies that have little influence upon the volumetric information of traffic and flow as "lines", which is remarkable in that the hosts that caused the hidden anomalies were mined out.
Yoshiki Kanda, Kensuke Fukuda, Toshiharu Sugawara
ICC2
2010 Dynamic and distributed routing control for virtualized local area networks
abstract
Advanced Layer-3 (L3) switches achieve high-speed IP packet forwarding by storing parts of the header information from transmitted packets into the flow cache in the switch fabric when relaying the IP packets between subnets. When IP traffic is overloaded on an L3 switch, the flow cache is easily exhausted, decreasing the IP packet forwarding performance. Virtual LAN (VLAN), a virtualization technology at the data-link layer, is widely used for the internal networks of many organizations because it allows network configurations to be changed easily and provides design flexibility. In a VLAN-based local area network with multiple L3 switches, the relaying point for each VLAN can be placed on any of the L3 switches. We developed a new network control scheme called distributed virtual routing, which dynamically controls the packet exchange points for each VLAN to suppress the consumption of flow caches. We describe the basic concept and then evaluate the reduction of the relaying flows through the simulation using the real network data.
Toshio Hirotsu, Satoshi Kurihara, Kensuke Fukuda, Osamu Akashi, Hirotake Abe, Toshiharu Sugawara
LCN3
2010 IM-DB: Information retrieval system for interactive network-status analysis
abstract
To manage the changing conditions of the Internet and diagnose the cause of anomalies within it, network operators must obtain and analyze information from network devices such as routers and network monitoring tools. Of that information, route information exchanged by Border Gateway Protocol (BGP) is essential for understanding the Internet at an inter-autonomous-system (AS) level. However, the information is not easily deployed for actual network management because of the enormous amount of the information and the difficulties in manipulating it. To overcome these problems, we propose IM-DB, a network information retrieval system for the interactive analysis and diagnosis of the Internet. Our IM-DB is based on a relational database and can store large amount of network information. By using the IM-DB, network operators can easily search, integrate, and explore information they require. We developed the first version of our IM-DB for storing BGP routing information. The experimental results using this prototype demonstrate the feasibility of our IM-DB.
Atsushi Terauchi, Osamu Akashi, Kensuke Fukuda
NOMS3
2010 Effect of Alternative Distributed Task Allocation Strategy Based on Local Observations in Contract Net Protocol
Toshiharu Sugawara, Kensuke Fukuda, Toshio Hirotsu, Satoshi Kurihara
PRIMA2
2010 Fluctuated peer selection policy and its performance in large-scale multi-agent systems
abstract
This paper describes how, in large-scale multi-agent systems, each agent's adaptive selection of peer agents for collaborative tasks affects the overall performance and how this performance varies with the workload of the system and with fluctuations
Toshiharu Sugawara, Kensuke Fukuda, Toshio Hirotsu, Shin-ya Sato, Osamu Akashi, Satoshi Kurihara
Web Intell. Agent Syst.2
2009 A Visualization of Internet AS Topology with Valley-free Rules
abstract
In this paper, we propose a new network drawing tool, which considers the direction of traffic on Autonomous System level (AS-level) networks. Almost all traditional network graph drawing tools consider only the connections between nodes and disregard the direction of traffic. Moreover, these tools tend to draw complicated pictures. The reasons for this are that AS-level network structures are very different to random graphs, and have a short diameter, effective connections, and several hubs. We also propose selection of information, and implement the spring algorithm to regulate the positions of nodes for better understanding and some ideas for AS-level network structures in graphic drawing tools. Information traffic is simulated in an AS-level network topology, with dynamic query scheduling to evaluate whether servers balance load well and are suitably located. Given the complexity of this topology, graphical drawing tools are necessary when depicting such networks. Our new network drawing tool draws simple pictures, and can not only express non-directed graphs, but also traffic with direction on AS-level networks, using valley-free rules proposed by Gao. These rules classify links into three policies which are expressed with different colors and shapes by our tool. Our tool clearly shows the differences between traffic routes with valley-free rules and shortest paths in non-directed graphs.
Shingo Nomoto, Kensuke Fukuda, Minoru Uehara, Hideki Mori
CISIS2
2009 An Automatic and Dynamic Parameter Tuning of a Statistics-Based Anomaly Detection Algorithm
abstract
The detection of anomalies in network traffic is a crucial issue affecting the security of Internet users. A statistical network anomaly detection algorithm is a promising way of detecting such anomalies, however, it has to be given appropriate parameters for accurate detection and identification. In general, it is very difficult to obtain appropriate parameter settings a priori, because network traffic is not stable in time or space. Thus, although many anomaly detection methods have been proposed, there has been little discussion about their parameter tunings. In this paper, we investigate an automatic and dynamic parameter tuning of a statistical network traffic anomaly detection method. In particular, we clarify whether one can consistently use the best parameter fixed for a certain instance; this choice clearly depends on the macroscopic and dynamic behavior of Internet traffic anomalies. We ascertain the appropriate learning period for setting a parameter of an anomaly detection algorithm based on a sketch and multi-scale gamma-function model by using real network traces measured in a trans-Pacific link over a period of six months. The main results of our study are as follows: (1) Without learning, the best parameter varies day by day. (2) With a longer learning period, the best parameter setting is affected by significant data during the learning period. (3) The appropriate period of the learning is about 3 days. (4) The performance degradation from introducing dynamic parameter tuning is 17% in the best case.
Yosuke Himura, Kensuke Fukuda, Kenjiro Cho, Hiroshi Esaki
ICC2
2009 Implementation and Evaluation of Layer-1 Bandwidth-on-Demand Capabilities in SINET3
abstract
This paper describes the implementation and evaluation of layer-1 bandwidth-on-demand (BoD) capabilities in the Japanese academic backbone network, called SINET3. The network has a nationwide GMPLS-based layer-1 platform and provides reservation-based and signaling-based BoD services. The overall architecture for providing BoD services including its capabilities, user interface, path calculation, and interface to drive the layer-1 platform are described. Actual examples of BoD services and evaluations of the path setup/release time in the network are also presented.
Shigeo Urushidani, Kensuke Fukuda, Yusheng Ji, Michihiro Koibuchi, Shunji Abe, Motonori Nakamura, Shigeki Yamada, Kaori Shimizu, Rie Hayashi, Ichiro Inoue, Kohei Shiomoto
ICC2
2009 Seven Years and One Day: Sketching the Evolution of Internet Traffic
abstract
This contribution aims at performing a longitudinal study of the evolution of the traffic collected every day for seven years on a trans-Pacific backbone link (the MAWI dataset). Long term characteristics are investigated both at TCP/IP layers (packet and flow attributes) and application usages. The analysis of this unique dataset provides new insights into changes in traffic statistics, notably on the persistence of Long Range Dependence, induced by the on-going increase in link bandwidth. Traffic in the MAWI dataset is subject to bandwidth changes, to congestions, and to a variety of anomalies. This allows the comparison of their impacts on the traffic statistics but at the same time significantly impairs long term evolution characterizations. To account for this difficulty, we show and explain how and why random projection (sketch) based analysis procedures provide practitioners with an efficient and robust tool to disentangle actual long term evolutions from time localized events such as anomalies and link congestions. Our central results consist in showing a strong and persistent long range dependence controlling jointly byte and packet counts. An additional study of a 24-hour trace complements the long-term results with the analysis of intraday variabilities.
Pierre Borgnat, Guillaume Dewaele, Kensuke Fukuda, Patrice Abry, Kenjiro Cho
INFOCOM3
2009 Estimating Relevance of Items on Basis of Proximity of User Groups on Blogspace
abstract
We describe a new method to estimate the relevance of two items (such as products and works of art) on the basis of the relationship between the corresponding user (blogger) groups on a blogspace, where a user group refers to a collection of users interested in an item. We estimated the strength of the relationship between user groups on the basis of their proximity on the blogspace. We validated our approach through experimental studies using actual data. In developing the method for estimating relevance among items, we introduced a new technique for measuring the proximity of two groups of vertices on a network, which can be thought of as an extension of conventional co-occurrence analysis.
Shin-ya Sato, Kensuke Fukuda, Toshio Hirotsu, Satoshi Kurihara, Toshiharu Sugawara
Web Intelligence2
2009 Design of versatile academic infrastructure for multilayer network services
abstract
This paper describes the network design and configurations of the new Japanese academic infrastructure, called SINET3, which provides a rich variety of network services to more than 700 universities and research institutions. Since the start of full-scale operations in June 2007, the network has expanded its services to include multi-layer transfer services (IP, Ethernet, and layer-1), enriched virtual private network services (L3VPN, L2VPN, VPLS, and L1VPN), enhanced QoS services (packet-based and circuit-based), and brand-new layer-1 bandwidth-on-demand (BoD) services. This paper explains how the network provides these various network services on a single network platform by effectively configuring leading-edge networking components, such as high-performance IP routers, layer- 1 switches, and a BoD server. Evaluations of the network design and configurations confirmed that the networking functions were effectively coordinated. The procedures and techniques related to the configuration validation that covered all phases of the network design and construction are also presented.
Shigeo Urushidani, Shunji Abe, Yusheng Ji, Kensuke Fukuda, Michihiro Koibuchi, Motonori Nakamura, Shigeki Yamada, Kaori Shimizu, Rie Hayashi, Ichiro Inoue, Kohei Shiomoto
IEEE J. Sel. Areas Commun.4
2008 Observing slow crustal movement in residential user traffic
abstract
It is often argued that rapidly increasing video content along with the penetration of high-speed access is leading to explosive growth in the Internet traffic. Contrary to this popular claim, technically solid reports show only modest traffic growth worldwide. This paper sheds light on the causes of the apparently slow growth trends by analyzing commercial residential traffic in Japan where the fiber access rate is much higher than other countries. We first report that Japanese residential traffic also has modest growth rates using aggregated measurements from six ISPs. Then, we investigate residential per-customer traffic in one ISP by comparing traffic in 2005 and 2008, before and after the advent of YouTube and other similar services. Although at first glance a small segment of peer-to-peer users still dictate the overall volume, they are slightly decreasing in population and volume share. Meanwhile, the rest of the users are steadily moving towards rich media content with increased diversity. Surely, a huge amount of online data and abundant headroom in access capacity can conceivably lead to a massive traffic growth at some point in the future. The observed trends, however, suggest that video content is unlikely to disastrously overflow the Internet, at least not anytime soon.
Kenjiro Cho, Kensuke Fukuda, Hiroshi Esaki, Akira Kato
CoNEXT2
2008 Correlation Among Piecewise Unwanted Traffic Time Series
abstract
In this paper, we investigate temporal and spatial correlations of time series of unwanted traffic (i.e., darknet or network telescope traffic) in order to estimate statistical behavior of unwanted activities from a small size of darknet address block. First, from the analysis of long-range dependency, we point out that TCP time series has a weak temporal correlation though UDP time series without huge flooding is well-modeled using a Poisson process. Next, we analyze the spatial correlation between two traffic time series divided by different sized darknet address blocks. We confirm that a TCP SYN traffic time series (e.g, virus or worm) has a clear spatial correlation in the arrival of packets between two neighboring address blocks. Indeed, this spatial correlation remains in traffic time series 1,000 addresses far from the target time series, even if a darknet address block is small (e.g., /26). On the other hand, TCP SYNACK traffic (e.g., backscatter) and UDP traffic (e.g., virus or worm) have less spatial correlation between two adjacent large address blocks. Finally, we estimate the average propagation delay of global unwanted activities appearing in TCP SYN traffic by using the generalized inter-correlation coefficient.
Kensuke Fukuda, Toshio Hirotsu, Osamu Akashi, Toshiharu Sugawara
GLOBECOM1
2008 Towards Modeling of Traffic Demand of Node in Large Scale Network
abstract
Abstract—Understanding actual network and traffic proper-ties of the Internet is essential to determine network parame-ters in large-scale network simulations. However, there is little knowledge about the distribution of macroscopic traffic demand for each node, though the topological properties of the network have been focused on. This paper investigates the distribution of traffic volume to and from a node at an organization level. As traffic volume data, we used byte counter data of all interfaces in all backbone routers in a nation-wide research and education (R&E) network in Japan. First, we show that traffic volumes to and from a node in the network are characterized by a lognormal distribution, which has a slower decay than a normal distribution, but a faster decay than a power-law distribution. Thus, an assumption in which the traffic demand is uniformly random or Gaussian distributed is not appropriated to model the traffic demand in large-scale network simulation. This finding implies that one has more possibility to observe an increase of delay or packet drop in simulation, comparing to the result that uses uniformly-random or Gaussian traffic demand, because of the locality of traffic. Moreover, we observed that in 87 % of nodes, a traffic volume from the backbone to the node is 1-10 times larger than that for the opposite direction. This is a similar usage pattern appeared in residential light-user broadband traffic. Finally, we introduce a simple model to explain the distribution of traffic demand, based on a multiplicative growth of traffic volume. We confirm that the multiplicative model can reproduce a lognormal distribution of traffic volume by simple numerical simulation. I.
Kensuke Fukuda
ICC1
2008 Co-occurrence Analysis Focused on Blogger Communities
abstract
We studied the problem of finding a subspace of Web pages that is contextually consistent for co-occurrence analysis. We looked at blogs and proposed blogger-based co-occurrence analysis, which assumes that two items are relevant to each other if they appear in any of the blog entries posted by the same blogger. We show that (1) blogger-based analysis outperforms conventional page-based analysis in solving context-sensitive problems and that (2) analysis focused on bloggers forming a community yields better performance compared with that focused on isolated bloggers.
Shin-ya Sato, Kensuke Fukuda, Toshio Hirotsu, Satoshi Kurihara, Toshiharu Sugawara
Web Intelligence2
2008 Policy-based BGP-control architecture for inter-AS routing adjustment
Osamu Akashi, Kensuke Fukuda, Toshio Hirotsu, Toshiharu Sugawara
Comput. Commun.2
2007 Performance variation due to interference among a large number of self-interested agents
abstract
The performance features of a massively multiagent system (MMAS) when applying the contract net protocol (CNP) are examined. The recent growth in the volume of e-commerce on the Internet is increasing the opportunities for coordinated transactions by agents, concurrently occurring everywhere. Because of limited CPU and network resources, running many interactive tasks among agents can lower the quality or efficiency of MMASs. Although CNP is a widely used negotiation protocol that can allocate tasks and resources to appropriate agents, it is unclear how effectively CNP works in an MMAS where thousands of agents work together and interfere with each other. The performance of CNP in such an MMAS, especially the overall efficiency and the reliability of promised completion times, is investigated by using an MAS simulation environment. The results show that only managerside control of CNP can improve performance in an MMAS.
Toshiharu Sugawara, Toshio Hirotsu, Satoshi Kurihara, Kensuke Fukuda
IEEE Congress on Evolutionary Computation4
2007 Layer-1 Bandwidth on Demand Services in SINET3
abstract
This paper describes brand-new layer-1 bandwidth on demand (BoD) services implemented in the new Japanese academic backbone network, called SINET3. SINET3 is an advanced converged network that provides multi-layer transfer, enriched VPN, enhanced QoS, and layer-1 BoD services. The layer-1 BoD services are dynamic layer-1 resource allocation services directly triggered by users and artfully achieved on the multi-service platform by using a layer-1 BoD server. This paper first explains how the network accommodates a wide variety of network services by effectively combining leading-edge technologies. The paper next describes the overall mechanism for the dynamic layer-1 path setup on the multi-service platform and details the functions of the BoD server in many aspects. The designs focus on the tangible achievement of these services over a nationwide network composed of 75 layer-1 switches and 12 IP/MPLS routers.
Shigeo Urushidani, Jun Matsukata, Kensuke Fukuda, Shunji Abe, Yusheng Ji, Michihiro Koibuchi, Shigeki Yamada, Kaori Shimizu, Tomonori Takeda, Ichiro Inoue, Kohei Shiomoto
GLOBECOM3
2007 Generating Extensional Definitions of Concepts from Ostensive Definitions by Using Web
Shin-ya Sato, Kensuke Fukuda, Satoshi Kurihara, Toshio Hirotsu, Toshiharu Sugawara
WISE2
2006 Adaptive Agent Selection in Large-Scale Multi-Agent Systems
Toshiharu Sugawara, Kensuke Fukuda, Toshio Hirotsu, Shin-ya Sato, Satoshi Kurihara
PRICAI2
2006 The impact and implications of the growth in residential user-to-user traffic
abstract
It has been reported worldwide that peer-to-peer traffic is taking up a significant portion of backbone networks. In particular, it is prominent in Japan because of the high penetration rate of fiber-based broadband access. In this paper, we first report aggregated traffic measurements collected over 21 months from seven ISPs covering 42% of the Japanese backbone traffic. The backbone is dominated by symmetric residential traffic which increased 37%in 2005. We further investigate residential per-customer trafficc in one of the ISPs by comparing DSL and fiber users, heavy-hitters and normal users, and geographic traffic matrices. The results reveal that a small segment of users dictate the overall behavior; 4% of heavy-hitters account for 75% of the inbound volume, and the fiber users account for 86%of the inbound volume. About 63%of the total residential volume is user-to-user traffic. The dominant applications exhibit poor locality and communicate with a wide range and number of peers. The distribution of heavy-hitters is heavy-tailed without a clear boundary between heavy-hitters and normal users, which suggests that users start playing with peer-to-peer applications, become heavy-hitters, and eventually shift from DSL to fiber. We provide conclusive empirical evidence from a large and diverse set of commercial backbone data that the emergence of new attractive applications has drastically affected traffic usage and capacity engineering requirements.
Kenjiro Cho, Kensuke Fukuda, Hiroshi Esaki, Akira Kato
SIGCOMM2
2005 ARTISTE: Agent Organization Management System for Multi-Agent Systems
Atsushi Terauchi, Osamu Akashi, Mitsuru Maruyama, Kensuke Fukuda, Toshiharu Sugawara, Toshio Hirotsu, Satoshi Kurihara
PRIMA4
2003 Dynamics of temporal correlation in daily Internet traffic
abstract
In order to characterize the dynamics of self-similar behavior in daily Internet traffic, we analyze the time series of traffic volume for a 24-hour period in a wide-area Internet, by using detrended fluctuation analysis (DFA) - a well-known method of characterizing nonstationarity in a time series. We show that the estimated scaling exponent (which is directly related to the Hurst parameter) of traffic fluctuations has a dependency on the level of human activity for a time scale greater than 30s. Thus, the temporal correlation for traffic fluctuations is close to 1/f-noise during the day, and becomes weaker at night. This result suggests that Internet traffic cannot be modeled using the unique value of the Hurst parameter.
Kensuke Fukuda, Luis A. Nunes Amaral, Harry Eugene Stanley
GLOBECOM1
2003 Secure and Manageable Virtual Private Networks for End-users
abstract
This paper presents personal networks, which integrate a VPN and the per-VPN execution environments of the hosts included in the VPN. The key point is that each execution environment called a portspace is bound to only one VPN, i.e., single-homed. Using this feature of portspaces, personal networks address several problems at multi-homed hosts that use multiple VPNs. Information flow is separated by personal networks so that it is not mixed at multi-homed hosts. IP addressing in a personal network is independent of the other personal networks, even the base network, and therefore does not conflict with those of other networks at multi-homed hosts. In addition, personal networks provide facilities for easy bootstrapping so that the end-users can construct such isolated networks easily. Inheritance of portspaces supports the creation of new portspaces based on existing portspaces. Self-construction of personal networks enables end-users to construct personal networks without help from the base network.
Kenichi Kourai, Toshio Hirotsu, Koji Sato, Osamu Akashi, Kensuke Fukuda, Toshiharu Sugawara, Shigeru Chiba
LCN5