VLDB 2026 Research / reviewers in the wild / expert
Masato Uchida
dblp:40/4751
· DBLP profile ↗
29ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0002-1998-0788ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Computer networks · 7 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unified Analysis of Continuous Weak Features Learning with Applications to Learning from Missing DataabstractThis paper addresses weak features learning (WFL), focusing on learning scenarios characterized by low-quality input features (weak features; WFs) that arise due to missingness, measurement errors, or ambiguous observations. We present a theoretical formalization and error analysis of WFL for continuous WFs (continuous WFL), which has been insufficiently explored in existing literature. A previous study established formalization and error analysis for WFL with discrete WFs (discrete WFL); however, this analysis does not extend to continuous WFs due to the inherent constraints of discreteness. To address this, we propose a theoretical framework specifically designed for continuous WFL, systematically capturing the interactions between feature estimation models for WFs and label prediction models for downstream tasks. Furthermore, we derive the theoretical conditions necessary for both sequential and iterative learning methods to achieve consistency. By integrating the findings of this study on continuous WFL with the existing theory of discrete WFL, we demonstrate that the WFL framework is universally applicable, providing a robust theoretical foundation for learning with low-quality features across diverse application domains. Kosuke Sugiyama, Masato Uchida |
ICML | 2 |
| 2025 | A Unified Framework for Generalization Error Analysis of Learning with Arbitrary Discrete Weak FeaturesabstractIn many real-world applications, predictive tasks inevitably involve low-quality input features (Weak Features; WFs) which arise due to factors such as misobservations, missingness, or partial observations. While several methods have been proposed to estimate the true values of specific types of WFs and to solve a downstream task, a unified theoretical framework that comprehensively addresses these methods remains underdeveloped. In this paper, we propose a unified framework called Weak Features Learning (WFL), which accommodates arbitrary discrete WFs and a broad range of learning algorithms, and we demonstrate its validity. Furthermore, we introduce a class of algorithms that learn both the estimation model for WFs and the predictive model for a downstream task and perform a generalization error analysis under finite-sample conditions. Our results elucidate the interdependencies between the estimation errors of WFs and the prediction error of a downstream task, as well as the theoretical conditions necessary for the learning approach to achieve consistency. This work establishes a unified theoretical foundation, providing generalization error analysis and performance guarantees, even in scenarios where WFs manifest in diverse forms. Kosuke Sugiyama, Masato Uchida |
ICML | 2 |
| 2025 | Decoding the Mind of Large Language Models: A Quantitative Evaluation of Ideology and BiasesabstractThe widespread integration of Large Language Models (LLMs) across various sectors has highlighted the need for empirical research to understand their biases, thought patterns, and societal implications to ensure ethical and effective use. In this study, we propose a novel framework for evaluating LLMs, focusing on uncovering their ideological biases through a quantitative analysis of 436 binary-choice questions, many of which have no definitive answer. By applying our framework to ChatGPT and Gemini, findings revealed that while LLMs generally maintain consistent opinions on many topics, their ideologies differ across models and languages. Notably, ChatGPT exhibits a tendency to change their opinion to match the questioner’s opinion. Both models also exhibited problematic biases, unethical or unfair claims, which might have negative societal impacts. These results underscore the importance of addressing both ideological and ethical considerations when evaluating LLMs. The proposed framework offers a flexible, quantitative method for assessing LLM behavior, providing valuable insights for the development of more socially aligned AI systems. Manari Hirose, Masato Uchida |
IJCNN | 2 |
| 2024 | Predicting Device Usage Patterns in Patients with Problematic Smartphone Use Through Individualized Hidden Markov ModelsabstractProblematic smartphone use (PSU) is a social issue affecting daily lives without a definitive treatment. Current PSU treatments rely on self-reported data, which can be inaccurate and lead to ineffective treatments. To address this, we propose using a hidden Markov model to objectively analyze smartphone usage patterns from log data. The model is trained on data from various patients and then individualized, allowing for meaningful interpretation and comparison of usage states. We conduct two case studies on patients attending outpatient clinics for Internet addiction due to PSU. These case studies compare the model-estimated usage state of patients with 1) a daily activity log as reported by the patients and 2) a several-month clinical history of the patients as reported by their attending psychiatrists. Through each case study, we validate the appropriateness of the definitions assigned to each usage state and the effectiveness of the state estimations performed using the proposed method. Our approach demonstrated that it is possible to objectively understand and track changes in PSU, which self-reports alone cannot achieve. Shun Furusawa, Toshitaka Hamamura, Akio Yoneyama, Masaru Honjo, Nanase Kobayashi, Daisuke Jitoku, Masato Uchida |
HealthCom | 7 |
| 2024 | Estimating Smartphone Orientation for Visualizing Behavior Around Sleep Periods in PSU PatientsabstractThe widespread use of smartphones has led to Problematic Smartphone Use (PSU), impacting academic performance, causing sleep deprivation, and deteriorating mental health. Assessing PSU often relies on biased self-reports, making it difficult for psychiatrists to accurately understand usage habits through patient interviews. This study proposes a framework using smartphone sensor logs to assess and visualize PSU habits. By analyzing smartphone orientation and spatial movements, we can detect user behavioral patterns. We focus on behaviors around sleep periods, as PSU patients tend to use their smartphones excessively, disrupting sleep and daily rhythms. The framework estimates smartphone orientations by clustering accelerometer data and determines sleep periods from lock state logs. By visualizing usage habits around sleep periods, we aim to identify unique PSU behavioral patterns and support clinical assessments. Case studies using actual log data will demonstrate the framework's effectiveness in detecting problematic behaviors. This visualization aims to be a supportive tool in clinical settings, enhancing the accuracy and reliability of PSU assessments. Kaito Hayashi, Takumi Meguro, Toshitaka Hamamura, Masato Taya, Masaru Honjo, Nanase Kobayashi, Daisuke Jitoku, Masato Uchida |
HealthCom | 8 |
| 2024 | Q&A Label LearningabstractAssigning labels to instances is crucial for supervised machine learning. In this letter, we propose a novel annotation method, Q&A labeling, which involves a question generator that asks questions about the labels of the instances to be assigned and an annotator that answers the questions and assigns the corresponding labels to the instances. We derived a generative model of labels assigned according to two Q&A labeling procedures that differ in the way questions are asked and answered. We showed that in both procedures, the derived model is partially consistent with that assumed in previous studies. The main distinction of this study from previous ones lies in the fact that the label generative model was not assumed but, rather, derived based on the definition of a specific annotation method, Q&A labeling. We also derived a loss function to evaluate the classification risk of ordinary supervised machine learning using instances assigned Q&A labels and evaluated the upper bound of the classification error. The results indicate statistical consistency in learning with Q&A labels. Kota Kawamoto, Masato Uchida |
Neural Comput. | 2 |
| 2023 | Balancing Selection and Diversity in Ensemble Learning with Exponential Mixture Model
Kosuke Sugiyama, Masato Uchida |
ICANN (3) | 2 |
| 2023 | Exploring Hessian Regularization in MixupabstractMixup is a data augmentation technique that gener-ates new samples using a convex combination of samples. Despite its remarkable simplicity, this method significantly improves the generalization performance of Deep Neural Networks (DNNs). Various studies have been conducted to understand Mixup, and it has been theoretically shown that Mixup involves Jacobian regularization. Furthermore, a relationship between Mixup and label smoothing has been suggested. However, these studies are limited to cases where the sample generated by Mixup is close to the original sample, and other cases have not yet been analyzed. In this study, we theoretically proved that applying Mixup to logistic regression results in Hessian regularization when the sample generated by Mixup is far from the original sample, a previously unexplored scenario. Additionally, through numerical experiments, we confirmed that a similar tendency holds true for DNNs as well. Our results provide a novel interpretation that Mixup is a method for collectively approximating several regularization methods, including Jacobian regularization, label smoothing, and Hessian regularization. Kosuke Sugiyama, Masato Uchida |
ICMLA | 2 |
| 2022 | Objection!: Identifying Misclassified Malicious Activities with XAIabstractMany studies have been conducted to detect various malicious activities in cyberspace using classifiers built by machine learning. However, it is natural for any classifier to make mistakes, and hence, human verification is necessary. One method to address this issue is eXplainable AI (XAI), which provides a reason for the classification result. However, when the number of classification results to be verified is large, it is not realistic to check the output of the XAI for all cases. In addition, it is sometimes difficult to interpret the output of XAI. In this study, we propose a machine learning model called classification verifier that verifies the classification results by using the output of XAI as a feature and raises objections when there is doubt about the reliability of the classification results. The results of experiments on malicious website detection and malware detection show that the proposed classification verifier can efficiently identify misclassified malicious activities. Koji Fujita, Toshiki Shibahara, Daiki Chiba 0001, Mitsuaki Akiyama, Masato Uchida |
ICC | 5 |
| 2021 | Analysis of Route Announcements of Unassigned IP AddressesabstractIt is known that some of the unassigned IP addresses are announced as legitimate routes, even though they are never assigned to any end-user. The reason for this is that the current routing system of the Internet with BGP allows unassigned IP addresses to be announced as well. On the other hand, actual situation is still unclear and there is no solid method for such an investigation. Thus, we conducted the very first investigation in Japan and proposed a simple and effective method. In our proposed method, we compare the address pool delegated to Japan and the actual route information that is available online. As a result, we have revealed that some unassigned IP addresses had been announced from both within Japan and overseas for several years because of human-setting-error. These problems affect not only registries, ISPs or carriers but also general end-users as well. Kentaro Goto, Akira Shibuya, Masayuki Okada, Masato Uchida |
COMPSAC | 4 |
| 2021 | Behind The Mask: Masquerading The Reason for PredictionabstractEnsuring the interpretability of the machine-learning-based prediction models is important for gaining users’ trust. The current representative algorithms for ensuring interpretability, LIME and SHAP, explain the prediction of a given black-box machine learning model based on a common perturbation mechanism to input data. In this study, we propose a masquerade layer to nullify this perturbation and hide the reason for prediction. Our proposed masquerade layer can be attached to any prediction models. It can also be attached without altering the prediction model itself, making it possible to manipulate the explanation provided by the interpretability algorithm in a manner that hardly changes the prediction model behavior. The experimental results show that existing representative perturbation-based interpretability algorithms have a critical weaknesses in terms of their reliability. Tomohiro Koide, Masato Uchida |
COMPSAC | 2 |
| 2021 | Auto-creation of Android Malware Family TreeabstractAndroid malware has been a growing threat. For an effective countermeasure against Android malware, we need to not only detect the malware at a certain point in time but also analyze its time-series changes of malware, taking into account that the family of Android malware will increase in number over time. In this paper, we propose a new method for automatically creating a "family tree" of Android malware that can represent how the newly detected Android malware is related to existing Android malware and its families, and how they have changed over time. Our evaluation using 24,474 actual Android malware APKs shows that our proposed family tree is able to accurately represent time-series changes between malware families. Kazuya Nomura, Daiki Chiba 0001, Mitsuaki Akiyama, Masato Uchida |
ICC | 4 |
| 2020 | Bridging Ordinary-Label Learning and Complementary-Label LearningabstractA supervised learning framework has been proposed for the situation where each trainingdata is provided with a complementary label that represents a class to which the pattern does not belong. In the existing literature, complementary-label learning has been studied independently from ordinary-label learning, which assumes that each training data is provided with a label representing the class to which the pattern belongs. However, providing a complementary label should be treated as equivalent to providing the rest of all the labels as the candidates of the one true class. In this paper, we focus on the fact that the loss functions for one-versus-all and pairwise classification corresponding to ordinary-label learning and complementary-label learning satisfy certain additivity and duality, and provide a framework which directly bridge those existing supervised learning frameworks. Further, we derive classification risk and error bound for any loss functions which satisfy additivity and duality. Yasuhiro Katsura, Masato Uchida |
ACML | 2 |
| 2020 | Time-series Measurement of Parked Domain NamesabstractDomain parking is a monetization mechanism for displaying online advertisements in unused domain names. Some domain names used in cyber attacks are known to leverage domain parking services after the attack. However, the temporal relationships between domain parking services and malicious domain names have not been studied well. In this study, we investigated how malicious domain names using domain parking services change over time. We conducted a large-scale measurement study of more than 66.8 million domain names that have used domain parking services in the past 19 months. We reveal the existence of 3,964 domain names that have been malicious after using domain parking. We also reveal the existence of 3.02 million domain names that utilized multiple parking services simultaneously or while switching between them. Our study can contribute to the efficient analysis of malicious domain names using domain parking services. Takayuki Tomatsuri, Daiki Chiba 0001, Mitsuaki Akiyama, Masato Uchida |
GLOBECOM | 4 |
| 2020 | Usage Prediction and Effectiveness Verification of App Restriction Function for Smartphone AddictionabstractIn recent years, there has been a growing problem of smartphone addiction. As the excessive use of smartphones has negatively impacted our daily lives, many apps for reducing smartphone addiction have been developed around the world. In this study, we focus on the app restriction function, which is one of the key features of digital medicines for smartphone addiction, and analyze the usage of the function and verify its effectiveness. The results showed significant differences in both psychological and behavioral aspects between those who used the app restriction function and those who did not. Specifically, we found that the app restriction function was more likely to be used by those who were more aware of their smartphone addiction. We also found that the app restriction function was effective in lessening smartphone usage time, especially when the smartphone addiction is relatively moderate. Katsuki Yasudomi, Toshitaka Hamamura, Masaru Honjo, Akio Yoneyama, Masato Uchida |
HealthCom | 5 |
| 2020 | Tsallis Entropy Based LabellingabstractIn the field of supervised classification, the quality of training data is an essential aspect of accurate learning, along with the selection of a learning algorithm or parameters optimisation. To improve the quality of training data, it is necessary to reflect an annotator's idea on to which class any given instance belongs in the form of a label as flexibly and accurately as possible. However, in conventional problem settings used in machine learning, the number of labels per instance is uniformly fixed at a certain value, and it is implicitly assumed that annotators provide labels under such a constraint. Thus, in this study, we propose an annotation framework; Tsallis entropy based labelling, which models a method that dynamically selects the number of labels for every single given instance depending on the uncertainty regarding the class to which each instance belongs. Using the proposed framework, an annotator's instinctive uncertainty about classification task is expressed based on the Tsallis entropy and Tsallis self-information. In addition, the proposed framework has a well-organised mathematical structure that includes some typical annotation models. We conduct an experiment to evaluate the proposed framework and demonstrate that it outperforms another annotation model in terms of the labels accuracy; In the comparison model, the number of labels per instance is deliberately set at a fixed value for all instances. Moreover, we exemplify that the conventional single labelling scheme is not always the best option, which reveals the fact that increasing the number of labels per instance does not necessarily hinder the labels accuracy. Kentaro Goto, Masato Uchida |
ICMLA | 2 |
| 2019 | Exploration into Gray Area: Efficient Labeling for Malicious Domain Name DetectionabstractThis paper presents a method to reduce the labeling cost when acquiring training data for a system that detects malicious domain names by supervised machine learning. The conventional system requires large quantities of both benign and malicious domain names to be prepared as training data to obtain a classifier with high classification accuracy. In general, malicious domain names are observed less frequently than benign domain names. Therefore, it is difficult to acquire a large number of malicious domain names without a dedicated labeling method. We propose a method based on active learning that labels data around the decision boundary of classification, i.e., in the gray area, and we show that the classification accuracy can be improved by only using approximately 2.5% of the training data used by the conventional system. An additional disadvantage of the conventional system is that, if the classifier is trained with a small amount of training data, its generalization ability cannot be guaranteed. We propose a method based on ensemble learning that integrates multiple classifiers, and we show that the classification accuracy can be stabilized and improved. Naoki Fukushi, Daiki Chiba 0001, Mitsuaki Akiyama, Masato Uchida |
COMPSAC (1) | 4 |
| 2019 | Predicting Network Outages Based on Q-Drop in Optical NetworkabstractThe sudden drop in the quality of an optical signal, called Q-drop, is an important factor for predicting network outages. Herein, we classify sudden drop events into two classes: one the results in a network outage, and one that does not result in a network outage, for the same period of time. Therefore, we build a predictor based on machine learning. The features of the predictor are given by characterizing the Q-drop event based on the optical layer characteristics immediately before the Q-drop event. The predictor is trained for each Q-drop event to adapt to the temporal change of the optical layer characteristics. Additionally, oversampling is applied to training data to avoid overlooking the network outage in the prediction. From the evaluation using real data, we showed that the proposed method is effective for the prediction of network outages in a short period. Furthermore, we found that information regarding the instability of the optical signal is important for the prediction of network outages. The result herein can contribute to improving the availability of the network because the proposed method predicts network outages based on characteristics that are invisible from the IP layer. Yohei Hasegawa, Masato Uchida |
COMPSAC (1) | 2 |
| 2019 | Optimal Hand Sign Selection Using Information Theory for Custom Sign-Based CommunicationabstractImproving the communication abilities of people suffering from speech disorders or hearing impairments and who are struggling to learn sign or spoken language can improve their quality of life. However, methods to assist such people are not varied, and those that consider the degree of physical disability usually fail to attend particular needs. Thus, it is necessary to provide various communication methods according to the characteristics of each physical disability. In this paper, we devise a customized hand sign recognition system according to the degree of physical disability, and propose a method to select a customized set of signs comprising specific hand motions that an individual can effortlessly perform. We consider the optimal set as that providing high reliability and efficiency to realize smooth communication and apply information theory towards their selection. That is, we consider hand sign recognition from myoelectric potentials elicited by finger movement as a communication channel. Then, the optimal hand sign set is determined considering the set with the maximum channel capacity, as it reflects the most reliable and efficient combination. Finally, experimental results obtained from three subjects verify that the proposed method can determine the optimal set of hand signs according to each subject and that increasing the available hand signs or choosing hand signs with high recognition rate do not necessarily contribute to the optimal set. Tokio Takahashi, Masato Uchida |
COMPSAC (1) | 2 |
| 2016 | Human error tolerant anomaly detection based on time-periodic packet sampling
Masato Uchida |
Knowl. Based Syst. | 1 |
| 2011 | An information-theoretic characterization of weighted α-proportional fairness in network resource allocation
Masato Uchida, James F. Kurose |
Inf. Sci. | 1 |
| 2009 | On the Quality of Triangle Inequality Violation Aware Routing Overlay ArchitectureabstractIt is known that Internet routing policies for both intra- and inter-domain routing can naturally give rise to triangle inequality violations (TIVs) with respect to quality of service (QoS) network metrics such as latencies between nodes. This motivates the exploitation of such TIVs phenomenon in network metrics to design TlV-aware routing overlay architecture which is capable of choosing quality overlay routing paths to improve end-to-end QoS without changing the underlying network architecture. Our idea is to find quality overlay routes between node pairs based on TIV optimization in terms of the latency and packet loss ratio, and that can offer near optimal routing quality in cost-effective and scalable manner. Our intuition to do this is to choose these overlay routes from a small set of transit nodes. We propose to assign nodes with transit selection frequency scores that are computed based on previous node usage for transit, and consolidate a small set of highly ranked transit nodes. For every node pair, we choose the best transit node in this small set for overlay routing, based on TIV optimization in latency and packet loss ratio. We analyze the quality of our TlV-aware routing overlay algorithm analytically and using real Internet measurements on latency and packet loss ratio. Our results show good quality performance in improving end-to-end QoS routing. Ryoichi Kawahara, Eng Keong Lua, Masato Uchida, Satoshi Kamei, Hideaki Yoshino |
INFOCOM | 3 |
| 2009 | An Information-Theoretic Characterization of Weighted alpha-Proportional FairnessabstractThis paper provides a novel characterization of fairness concepts in network resource allocation problems from the viewpoint of information theory. The fundamental idea adopted in this paper is to characterize the utility functions used in optimization problems, which motivate fairness concepts, based on a trade-off between user and system satisfaction. Here, user satisfaction is evaluated using information divergence measures that were originally used in information theory to evaluate the difference between two probability distributions. In this paper, information divergence measures are applied to evaluate the difference between the implemented resource allocation and a requested resource allocation. The requested resource allocation is assumed to be ideal in some sense from the user's point of view. Also, system satisfaction is evaluated based on the efficiency of the implemented resource utilization, which is defined as the total amount of resources allocated to each user. The results discussed in this paper indicate that the well-known fairness concept called weighted alpha-proportional fairness can be characterized using the alpha-divergence measure, which is a general class of information divergence measures, as an equilibrium of the trade-off described above. In the process of obtaining these results, we also obtained a new utility function that has a parameter to control the trade-off. This new function is then applied to typical examples to solve resource allocation problems in simple network models such as those for two-link networks and wireless LANs. Masato Uchida, James F. Kurose |
INFOCOM | 1 |
| 2007 | Design of an Unsupervised Weight Parameter Estimation Method in Ensemble Learning
Masato Uchida, Yousuke Maehara, Hiroyuki Shioya |
ICONIP (1) | 1 |
| 2006 | Scheduling algorithms with error rate consideration in HSDPA networksabstractThe probability of transmission failure under a given wireless condition and instantaneous transmission rate is introduced as a new scheduling metric to improve the performance of existing scheduling algorithms. Scheduling algorithms for high-speed downlink packet access (HSDPA) in wideband code division multiple access (WCDMA) networks generally deal with the instantaneous transmission rate of each user as a key parameter reflecting the time-varying and locationdependent characteristics of the wireless channel. As a more detailed measure of the wireless condition of each user, the error rate metric improves the effectiveness of sharing resources without sacrificing a good balance between system performance and user fairness. Simulation results demonstrate that a performance gain of approximately 20% can be achieved for several well-known scheduling algorithms when the error rate is taken into account. Yan Zhang 0026, Masato Uchida, Masato Tsuru 0001, Yuji Oie |
IWCMC | 2 |
| 2006 | Effect of Introducing Predictive CIR on Throughput Performance in HSDPA NetworksabstractPredicting wireless environment parameters as correctly and timely as possible is of practical importance to optimal resource allocation in wireless systems. For example, as the realtime wireless condition information is not available in reality due to delay, the latest historic wireless environment parameters are used to determine the users' instantaneous transmission rate in most of the existing systems in high-speed downlink packet access (HSDPA) networks. However, the wireless condition variation during this delay (e.g., 6 ms) sometimes results in unsuitable determination of the instantaneous transmission rate and eventually influences the user and system performance. It is shown by simulation results that compared with using the ideal realtime wireless environment information, using the latest available historic information will introduce about 10%-15% throughput performance damage averagely. Therefore, in this paper, in order to mitigate such performance degradation, a simple linear prediction of the current wireless environment parameters is introduced. Simulation results demonstrate that nearly 10% improvement can be achieved under several well-known scheduling algorithms in a HSDPA network when the predictive environment parameter is applied to determine the instantaneous transmission rate Yan Zhang 0026, Masato Uchida, Masato Tsuru 0001, Yuji Oie |
PIMRC | 2 |
| 2005 | Replication scheme for traffic load balancing and its parameter tuning in pure P2P communicationabstractA replication scheme is an effective method for maintaining a high hit rate for target files. This paper proposes a new replication scheme that scatters 'replicas' of files to other servants on a P2P network for the purpose of dispersing query and download traffic. First, the details of this proposed scheme are described and its parameter tuning is discussed by introducing rating functions. Second, the effectiveness of this scheme is evaluated by simulation. The simulation results demonstrate that the proposed scheme has superior performance for some characteristic values, such as the hit rate and the number of mean hops to the target servant, compared with other schemes. Finally, the influence of departure rate is discussed. Shinya Nogami, Masato Uchida, Takeo Abe |
ISADS | 2 |
| 2004 | Traffic data analysis based on extreme value theory and its applicationsabstractIt is important to predict serious deterioration of telecommunication quality. The purpose of this paper is to predict such serious events by analyzing only a "short" period of teletraffic data. It presents a method for analyzing the tail distributions (TD) of variables concerning teletraffic states, because TD are suitable to represent serious events. This method is based on extreme value theory (EVT), which provides a firm theoretical foundation for the analysis. To be more precise, we use throughput data measured on an actual network in daily busy hours for 15 min, and use its first 10 s (known data) to analyze the TD. Then, we evaluate how well the obtained TD can predict the TD of the remaining 890 s (unknown data). The result shows that the obtained TD, based on EVT by analyzing the small amount of known data, can predict the TD of the unknown data much better than methods based on an empirical distribution and log-normal distribution. Furthermore, we apply the obtained TD to predict the peak throughput in unknown data. The results of this paper enable us to predict serious events with lower measurement cost. Masato Uchida |
GLOBECOM | 1 |
| 2004 | Identifying elephant flows through periodically sampled packetsabstractIdentifying elephant flows is very important in developing effective and efficient traffic engineering schemes. In addition, obtaining the statistics of these flows is also very useful for network operation and management. On the other hand, with the rapid growth of link speed in recent years, packet sampling has become a very attractive and scalable means to measure flow statistics; however, it also makes identifying elephant flows become much more difficult. Based on Bayes' theorem, this paper develops techniques and schemes to identify elephant flows in periodically sampled packets. We show that our basic framework is very flexible in making appropriate trade-offs between false positives (misidentified flows) and false negatives (missed elephant flows) with regard to a given sampling frequency. We further validate and evaluate our approach by using some publicly available traces. Our schemes are generic and require no per-packet processing; hence, they allow a very cost-effective implementation for being deployed in large-scale high-speed networks. Tatsuya Mori 0003, Masato Uchida, Ryoichi Kawahara, Jianping Pan 0001, Shigeki Goto |
Internet Measurement Conference | 2 |