Fabricio Murai

dblp:30/9221 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-4487-6381ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 9 since 2021Databases, data management, data science and information retrieval · 14 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 6 · 4 since 2021Computer networks · 4 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering
abstract
Trustworthy biomedical question answering (QA) systems must not only provide accurate answers but also justify them with current, verifiable evidence.Retrieval-augmented approaches partially address this gap but lack mechanisms to iteratively refine poor queries, whereas self-reflection methods kick in only after full retrieval is completed.In this context, we introduce PubMed Reasoner, a biomedical QA agent composed of three stages: self-critic query refinement evaluates MeSH terms for coverage, alignment, and redundancy to enhance PubMed queries based on partial (metadata) retrieval; reflective retrieval processes articles in batches until sufficient evidence is gathered; and evidence-grounded response generation produces answers with explicit citations.PubMed Reasoner with a GPT-4o backbone achieves 78.32% accuracy on Pub-MedQA, slightly surpassing human experts, and showing consistent gains on MMLU Clinical Knowledge.Moreover, LLM-as-judge evaluations prefer our responses across: reasoning soundness, evidence grounding, clinical relevance, and trustworthiness.By orchestrating retrieval-first reasoning over authoritative sources, our approach provides practical assistance to clinicians and biomedical researchers while controlling compute and token costs. 1 2 https://github.com/swiftzhang125/ACL_2026_ PubMed_Reasoner, including code, prompt templates, LLM-as-judge evaluation, and baseline systems.
Xiaozhong Liu 0001, Fabricio Murai
ACL (1)3
2025 A Longitudinal Autoethnography of Email Access for a Professional with Chronic Illness and ADHD: Preliminary Insights
Veronica Pimenova, Yotam Sechayk, Fabricio Murai, Andrew Hundt, Shiri Dori-Hacohen
ASSETS3
2025 CLaDMoP: Learning Transferrable Models from Successful Clinical Trials via LLMs
abstract
Many existing models for clinical trial outcome prediction are optimized using task-specific loss functions on trial phase-specific data.While this scheme may boost prediction for common diseases and drugs, it can hinder the learning of generalizable representations, leading to more false positives/negatives.To address this limitation, we introduce CLaDMoP, a new pre-training approach for clinical trial outcome prediction, alongside the Successful Clinical Trials dataset (SCT), specifically designed for this task.CLaDMoP leverages a Large Language Model-to encode trials' eligibility criteria-linked to a lightweight Drug-Molecule branch through a novel multi-level fusion technique.To efficiently fuse long embeddings across levels, we incorporate a grouping block, drastically reducing computational overhead.CLaDMoP avoids reliance on task-specific objectives by pre-training on a "pair matching" proxy task.Compared to established zero-shot and few-shot baselines, our method significantly improves both PR-AUC and ROC-AUC, especially for phase I and phase II trials.We further evaluate and perform ablation on CLaDMoP after Parameter-Efficient Fine-Tuning, comparing it to state-of-the-art supervised baselines, including MEXA-CTP, on the Trial Outcome Prediction (TOP) benchmark.CLaDMoP achieves up to 10.5% improvement in PR-AUC and 3.6% in ROC-AUC, while attaining comparable F1 score to MEXA-CTP, highlighting its potential for clinical trial outcome prediction.Code and SCT dataset can be downloaded from https://github.com/murai-lab/CLaDMoP.
Xiaozhong Liu 0001, Fabricio Murai
KDD (2)3
2025 MEXA-CTP: Mode Experts Cross-Attention for Clinical Trial Outcome Prediction
abstract
Clinical trials are the gold standard for assessing the effectiveness and safety of drugs for treating diseases. Given the vast design space of drug molecules, elevated financial cost, and multi-year timeline of these trials, research on clinical trial outcome prediction has gained immense traction. Accurate predictions must leverage data of diverse modes such as drug molecules, target diseases, and eligibility criteria to infer successes and failures. Previous Deep Learning approaches for this task, such as HINT, often require wet lab data from synthesized molecules and/or rely on prior knowledge to encode interactions as part of the model architecture. To address these limitations, we propose a light-weight attention-based model, MEXA-CTP, to integrate readily-available multi-modal data and generate effective representations via specialized modules dubbed “mode experts”, while avoiding human biases in model design. We optimize MEXA-CTP with the Cauchy loss to capture relevant interactions across modes. Our experiments on the Trial Outcome Prediction (TOP) benchmark demonstrate that MEXA-CTP improves upon existing approaches by, respectively, up to 11.3% in F1 score, 12.2% in PR-AUC, and 2.5% in ROC-AUC, compared to HINT. Ablation studies are provided to quantify the effectiveness of each component in our proposed method. Code can be downloaded from github.com/murai-lab/MEXA-CTP.
Xiaozhong Liu 0001, Fabricio Murai
SDM3
2024 Hidden or Inferred: Fair Learning-To-Rank With Unknown Demographics
abstract
As learning-to-rank models are increasingly deployed for decision-making in areas with profound life implications, the FairML community has been developing fair learning-to-rank (LTR) models. These models rely on the availability of sensitive demographic features such as race or sex. However, in practice, regulatory obstacles and privacy concerns protect this data from collection and use. As a result, practitioners may either need to promote fairness despite the absence of these features or turn to demographic inference tools to attempt to infer them. Given that these tools are fallible, this paper aims to further understand how errors in demographic inference impact the fairness performance of popular fair LTR strategies. In which cases would it be better to keep such demographic attributes hidden from models versus infer them? We examine a spectrum of fair LTR strategies ranging from fair LTR with and without demographic features hidden versus inferred to fairness-unaware LTR followed by fair re-ranking. We conduct a controlled empirical investigation modeling different levels of inference errors by systematically perturbing the inferred sensitive attribute. We also perform three case studies with real-world datasets and popular open-source inference methods. Our findings reveal that as inference noise grows, LTR-based methods that incorporate fairness considerations into the learning process may increase bias. In contrast, fair re-ranking strategies are more robust to inference errors. All source code, data, and experimental artifacts of our experimental study are available here: https://github.com/sewen007/hoiltr.git
Oluseun Olulana, Kathleen Cachel, Fabricio Murai, Elke A. Rundensteiner
AIES (1)3
2024 Reducing Biases towards Minoritized Populations in Medical Curricular Content via Artificial Intelligence for Fairer Health Outcomes
abstract
Biased information (recently termed bisinformation) continues to be taught in medical curricula, often long after having been debunked. In this paper, we introduce bricc, a first-in-class initiative that seeks to mitigate medical bisinformation using machine learning to systematically identify and flag text with potential biases, for subsequent review in an expert-in-the-loop fashion, thus greatly accelerating an otherwise labor-intensive process. We have developed a gold-standard bricc dataset throughout several years containing over 12K pages of instructional materials. Medical experts meticulously annotated these documents for bias according to comprehensive coding guidelines, emphasizing gender, sex, age, geography, ethnicity, and race. Using this labeled dataset, we trained, validated, and tested medical bias classifiers. We test three classifier approaches: a binary type-specific classifier, a general bias classifier; an ensemble combining bias type-specific classifiers independently-trained; and a multi-task learning (MTL) model tasked with predicting both general and type-specific biases. While MTL led to some improvement on race bias detection in terms of F1-score, it did not outperform binary classifiers trained specifically on each task. On general bias detection, the binary classifier achieves up to 0.923 of AUC, a 27.8% improvement over the baseline. This work lays the foundations for debiasing medical curricula by exploring a novel dataset and evaluating different training model strategies. Hence, it offers new pathways for more nuanced and effective mitigation of bisinformation.
Chiman Salavati, Shannon Song, Willmar Sosa Diaz, Scott A. Hale, Roberto E. Montenegro, Fabricio Murai, Shiri Dori-Hacohen
AIES (1)6
2024 Devil in the Noise: Detecting Advanced Persistent Threats with Backbone Extraction
abstract
The use of host intrusion detection systems shows promising results in detecting APT campaigns due to the use of systems logs as source data to get more information about system environment. However, dealing with the increase of logs in time while tracking the execution context is a challenge for security analysts. Therefore, this work presents backbone extraction as a crucial preprocessing step, filtering out irrelevant logs. As the logs are modeled as provenance graphs, we discard spurious edges to detect residuals with distinctive node and edge distributions that indicate security threats. By applying our methodology to state-of-the-art benchmark datasets, we observed an increase in the performance of one-class classifiers by up to 62% on F1-score and 48% on recall in the Streamspot dataset and by up to 40% on F1-score and 33% on recall in the DARPA3 THEIA dataset. Moreover, our results indicate mitigation of the dependency explosion problem and underscore the ability of our methodology to improve the detection landscape by shrinking graph sizes without losing essential aspects to characterize attacks.
Caio M. C. Viana, Carlos Henrique Gomes Ferreira, Fabricio Murai, Aldri Luiz dos Santos, Lourenço Alves Pereira Júnior
ISCC3
2023 Fisher Scoring Method for Neural Networks Optimization
abstract
First-order methods based on the stochastic gradient descent and variants are popularly used in training neural networks. The large dimension of the parameter space prevents the use of second-order methods in current practice. The empirical Fisher information matrix is a readily available estimate of the Hessian matrix that has been used recently to guide informative dropout approaches in deep learning. In this paper, we propose efficient ways to dynamically estimate the empirical Fisher information matrix to speed up the optimization of deep learning loss functions. We propose two different methods, both using rank-1 updates for the empirical Fisher information matrix. The first one is FisherExp and it is based on exponential smoothing using Sherman-Woodbury-Morrison matrix inversion formula. The second one is FisherFIFO, which uses a circular gradient buffer using the Sherman-Woodbury-Morrison formula twice every time a new gradient is replaced. We found that FisherFIFO scales better and we further improve scaling by proposing a partitioning strategy for the empirical Fisher Information matrix. Our methods can be used in conjunction with existing optimizers that leverage momentum-based information to improve them. We compare the performance of our methods with alternative baselines in image classification problems and found that they produce better results. Despite the overhead incurred by using second-order information, the partitioning strategy combined with parallel block updates allows us to reduce the total training time of FisherFIFO relative to the baselines.
Jackson de Faria, Renato Assunção, Fabricio Murai
SDM3
2022 Top-Down Deep Clustering with Multi-Generator GANs
abstract
Deep clustering (DC) leverages the representation power of deep architectures to learn embedding spaces that are optimal for cluster analysis. This approach filters out low-level information irrelevant for clustering and has proven remarkably successful for high dimensional data spaces. Some DC methods employ Generative Adversarial Networks (GANs), motivated by the powerful latent representations these models are able to learn implicitly. In this work, we propose HC-MGAN, a new technique based on GANs with multiple generators (MGANs), which have not been explored for clustering. Our method is inspired by the observation that each generator of a MGAN tends to generate data that correlates with a sub-region of the real data distribution. We use this clustered generation to train a classifier for inferring from which generator a given image came from, thus providing a semantically meaningful clustering for the real distribution. Additionally, we design our method so that it is performed in a top-down hierarchical clustering tree, thus proposing the first hierarchical DC method, to the best of our knowledge. We conduct several experiments to evaluate the proposed method against recent DC methods, obtaining competitive results. Last, we perform an exploratory analysis of the hierarchical clustering tree that highlights how accurately it organizes the data in a hierarchy of semantically coherent patterns.
Daniel P. M. de Mello, Renato Assunção, Fabricio Murai
AAAI3
2022 Uncovering Coordinated Communities on Twitter During the 2020 U.S. Election
abstract
A large volume of content related to claims of election fraud, often associated with hate speech and extremism, was reported on Twitter during the 2020 US election, with evidence that coordinated efforts took place to promote such content on the platform. In response, Twitter announced the suspension of thousands of user accounts allegedly involved in such actions. Motivated by these events, we here propose a novel network-based approach to uncover evidence of coordination in a set of user interactions. Our approach is designed to address the challenges incurred by the often sheer volume of noisy edges in the network (i.e., edges that are unrelated to coordination) and the effects of data sampling. To that end, it exploits the joint use of two network backbone extraction techniques, namely Disparity Filter and Neighborhood Overlap, to reveal strongly tied groups of users (here referred to as communities) exhibiting repeatedly common behavior, consistent with coordination. We employ our strategy to a large dataset of tweets related to the aforementioned fraud claims, in which users were labeled as suspended, deleted or active, according to their accounts status after the election. Our findings reveal well-structured communities, with strong evidence of coordination to promote (i.e., retweet) the aforementioned fraud claims. Moreover, many of those communities are formed not only by suspended and deleted users, but also by users who, despite exhibiting very similar sharing patterns, remained active in the platform. This observation suggests that a significant number of users who were potentially involved in the coordination efforts went unnoticed by the platform, and possibly remained actively spreading this content on the system.
Renan Saldanha Linhares, Jose Martins da Rosa, Carlos Henrique Gomes Ferreira, Fabricio Murai, Gabriel Peres Nobre, Jussara M. Almeida
ASONAM4
2022 DELATOR: Money Laundering Detection via Multi-Task Learning on Large Transaction Graphs
abstract
Money laundering has become one of the most relevant criminal activities in modern societies, as it causes massive financial losses for governments, banks and other institutions. Detecting such activities is among the top priorities when it comes to financial analysis, but current approaches are often costly and labor intensive partly due to the sheer amount of data to be analyzed. Hence, there is a growing need for automatic anti-money laundering systems to assist experts. In this work, we propose DELATOR, a novel framework for detecting money laundering activities based on graph neural networks that learn from large-scale temporal graphs. DELATOR provides an effective and efficient method for learning from heavily imbalanced graph data, by adapting concepts from the GraphSMOTE framework and incorporating elements of multi-task learning to obtain rich node embeddings for node classification. DELATOR outperforms all considered baselines, including an off-the-shelf solution from Amazon AWS by 23% with respect to AUC-ROC. We also conducted real experiments that led to the discovery of 7 new suspicious cases among the 50 analyzed ones, which have been reported to the authorities.
Henrique S. Assumpção, Fabrício R. de Souza, Leandro Lacerda Campos, Vinícius T. de Castro Pires, Paulo M. Laurentys de Almeida, Fabricio Murai
IEEE Big Data6
2021 Mixture Variational Autoencoder of Boltzmann Machines for Text Processing
Bruno Guilherme, Fabricio Murai, Olga Goussevskaia, Ana Paula Couto da Silva
NLDB2
2021 Sequence-Based Word Embeddings for Effective Text Classification
Bruno Guilherme, Fabricio Murai, Olga Goussevskaia, Ana Paula Couto da Silva
NLDB2
2021 Predicting user emotional tone in mental disorder online communities
Bárbara Silveira 0001, Henrique S. Silva, Fabricio Murai, Ana Paula Couto da Silva
Future Gener. Comput. Syst.3
2020 Characterizing (Un)moderated Textual Data in Social Systems
abstract
Despite the valuable social interactions that online media promote, these systems provide space for speech that would be potentially detrimental to different groups of people. The moderation of content imposed by many social media has motivated the emergence of a new social system for free speech named Gab, which lacks moderation of content. This article characterizes and compares moderated textual data from Twitter with a set of unmoderated data from Gab. In particular, we analyze distinguishing characteristics of moderated and unmoderated content in terms of linguistic features, evaluate hate speech and its different forms in both environments. Our work shows that unmoderated content presents different psycholinguistic features, more negative sentiment and higher toxicity. Our findings support that unmoderated environments may have proportionally more online hate speech. We hope our analysis and findings contribute to the debate about hate speech and benefit systems aiming at deploying hate speech detection approaches.
Lucas Lima 0002, Julio C. S. Reis, Philipe F. Melo, Fabricio Murai, Fabrício Benevenuto
ASONAM4
2019 Machine Learning for Performance Prediction of Spark Cloud Applications
abstract
Big data applications and analytics are employed in many sectors for a variety of goals: improving customers satisfaction, predicting market behavior or improving processes in public health. These applications consist of complex software stacks that are often run on cloud systems. Predicting execution times is important for estimating the cost of cloud services and for effectively managing the underlying resources at runtime. Machine Learning (ML), providing black box solutions to model the relationship between application performance and system configuration without requiring in-detail knowledge of the system, has become a popular way of predicting the performance of big data applications. We investigate the cost-benefits of using supervised ML models for predicting the performance of applications on Spark, one of today's most widely used frameworks for big data analysis. We compare our approach with Ernest (an ML-based technique proposed in the literature by the Spark inventors) on a range of scenarios, application workloads, and cloud system configurations. Our experiments show that Ernest can accurately estimate the performance of very regular applications, but it fails when applications exhibit more irregular patterns and/or when extrapolating on bigger data set sizes. Results show that our models match or exceed Ernest's performance, sometimes enabling us to reduce the prediction error from 126-187% to only 5-19%.
Alexandre Maros, Fabricio Murai, Ana Paula Couto da Silva, Jussara M. Almeida, Marco Lattuada 0001, Eugenio Gianniti, Marjan Hosseini, Danilo Ardagna
CLOUD2
2019 Gray-Box Models for Performance Assessment of Spark Applications
abstract
Big data applications are among the most suitable applications to be executed on cluster resources because of their high requirements of computational power and data storage.Correctly sizing the resources devoted to their execution does not guarantee they will be executed as expected.Nevertheless, their execution can be affected by perturbations which can change the expected execution time.Identifying when these types of issue occurred by comparing their actual execution time with the expected one is mandatory to identify potentially critical situations and to take the appropriate steps to prevent them.To fulfill this objective, accurate estimates are necessary.In this paper, machine learning techniques coupled with a posteriori knowledge are exploited to build performance estimation models.Experimental results show how the models built with the proposed approach are able to outperform a reference state-of-the-art method (i.e., Ernest method), reducing in some scenarios the error from the 221.09-167.07%to 13.15-30.58%.
Marco Lattuada 0001, Eugenio Gianniti, Marjan Hosseini, Danilo Ardagna, Alexandre Maros, Fabricio Murai, Ana Paula Couto da Silva, Jussara M. Almeida
CLOSER6
2019 Characterizing Directed and Undirected Networks via Multidimensional Walks with Jumps
abstract
Estimating distributions of node characteristics (labels) such as number of connections or citizenship of users in a social network via edge and node sampling is a vital part of the study of complex networks. Due to its low cost, sampling via a random walk (RW) has been proposed as an attractive solution to this task. Most RW methods assume either that the network is undirected or that walkers can traverse edges regardless of their direction. Some RW methods have been designed for directed networks where edges coming into a node are not directly observable. In this work, we propose Directed Unbiased Frontier Sampling (DUFS), a sampling method based on a large number of coordinated walkers, each starting from a node chosen uniformly at random. It applies to directed networks with invisible incoming edges because it constructs, in real time, an undirected graph consistent with the walkers trajectories, and its use of random jumps to prevent walkers from being trapped. DUFS generalizes previous RW methods and is suited for undirected networks and to directed networks regardless of in-edge visibility. We also propose an improved estimator of node label distribution that combines information from initial walker locations with subsequent RW observations. We evaluate DUFS, compare it to other RW methods, investigate the impact of its parameters on estimation accuracy and provide practical guidelines for choosing them. In estimating out-degree distributions, DUFS yields significantly better estimates of the head of the distribution than other methods, while matching or exceeding estimation accuracy of the tail. Last, we show that DUFS outperforms uniform sampling when estimating distributions of node labels of the top 10% largest degree nodes, even when sampling a node uniformly has the same cost as RW steps.
Fabricio Murai, Bruno Ribeiro 0001, Don Towsley, Pinghui Wang
ACM Trans. Knowl. Discov. Data1
2018 Inside the Right-Leaning Echo Chambers: Characterizing Gab, an Unmoderated Social System
abstract
The moderation of content in many social media systems, such as Twitter and Facebook, motivated the emergence of a new social network system that promotes free speech, named Gab. Soon after that, Gab has been removed from Google Play Store for violating the company's hate speech policy and it has been rejected by Apple for similar reasons. In this paper we characterize Gab, aiming at understanding who are the users who joined it and what kind of content they share in this system. Our findings show that Gab is a very politically oriented system that hosts banned users from other social networks, some of them due to possible cases of hate speech and association with extremism. We provide the first measurement of news dissemination inside a right-leaning echo chamber, investigating a social media where readers are rarely exposed to content that cuts across ideological lines, but rather are fed with content that reinforces their current political or social views.
Lucas Lima 0002, Julio C. S. Reis, Philipe F. Melo, Fabricio Murai, Leandro Araújo, Pantelis Vikatos, Fabrício Benevenuto
ASONAM4
2018 Estimation Errors in Network A/B Testing Due to Sample Variance and Model Misspecification
abstract
Companies that offer services on the Web often rely on randomized experiments known as A/B tests for assessing the impact of development and business decisions. During an experiment, each user is randomly redirected to one of two versions of the website, called treatments. Several response models were proposed to describe the behavior of a user in a social network website as a function of the treatment assigned to her and to her neighbors. However, there is no consensus as to which model should be applied to a given dataset. In this work, we propose a new response model, derive theoretical limits for the estimation error of several models, and obtain empirical results for cases where the response model was misspecified.
Francisco Galuppo Azevedo, Bruno Demattos Nogueira, Fabricio Murai, Ana Paula Couto da Silva
WI3
2018 Reddit Weight Loss Communities: Do They Have What It Takes for Effective Health Interventions?
abstract
Online social networks are an important tool for people to share information and have been extensively used for people to achieve beneficial changes in health. Obesity is a major public health concern that affects about one third of the world's population. In order to alleviate this problem, health professionals are focusing on health interventions, which can be performed online. In this study we analyze three distinct online communities about weight and diet in Reddit. We model our data as 3 directed and weighted graphs of the posts and comments and evaluate the interaction between users of each community. We also analyze specific characteristics of each community, the habits of daily activity of the users and the formation of implicit bonds of friendship through the formation of communities. Our main results show that Reddit is a content-centered social network, in which what matters is what is posted and not who posts. In addition, users tend to create implicit friendship relationships through denser regions of interactions. Our results show that, contrary to expectations, the three communities present the same behavior pattern in a general point of view, which facilitates the development of non-directed online weight loss intervention strategies.
Karen Braga Enes, Pedro Paulo Valadares Brum, Tiago Oliveira Cunha, Fabricio Murai, Ana Paula Couto da Silva, Gisele L. Pappa
WI4
2018 Online Social Networks in Health Care: A Study of Mental Disorders on Reddit
abstract
The alarming increase in the number of people afflicted by mental health disorders has become one of the major public health problems faced by governments worldwide. Traditional face-to-face clinical interventions are costly, and, in many cases, leave out a sizable number of people who are struggling to improve their mental health conditions. Alternative forms of intervention that have a larger reach and allow for continual interaction while reducing costs are being investigated, including those relying on Online Social Networks (OSNs). Initially designed for promoting friendship, OSNs started to connect people willing to share experiences related to mental health disorders. In light of this fact, we investigate four Reddit online communities: Depression, SuicideWatch, Anxiety and Bipolar. We focus on user activities and interactions, and on the discourse pattern analysis of posts and comments made by the community members. We found that (i) interaction patterns are very similar across these subreddits, and interactions are centered around content, rather than users; (ii) most posts that generate the longest discussion trees are requests for help and, more often than not, multiple users offer support; (iii) the four subreddits share a common language and encouragement words are a frequent pattern, for instance. We hope that the insights unveiled by our analyses will help on building successful online interventions to support people in crisis and assist their counselors.
Bárbara Silveira 0001, Ana Paula Couto da Silva, Fabricio Murai
WI3
2018 Selective harvesting over networks
Fabricio Murai, Diogo Rennó, Bruno Ribeiro 0001, Gisele L. Pappa, Don Towsley, Krista Gile
Data Min. Knowl. Discov.1
2013 On Set Size Distribution Estimation and the Characterization of Large Networks via Sampling
abstract
In this work we study the set size distribution estimation problem, where elements are randomly sampled from a collection of non-overlapping sets and we seek to recover the original set size distribution from the samples. This problem has applications to capacity planning and network theory. Examples of real-world applications include characterizing in-degree distributions in large graphs and uncovering TCP/IP flow size distributions on the Internet. We demonstrate that it is difficult to estimate the original set size distribution. The recoverability of original set size distributions presents a sharp threshold with respect to the fraction of elements that remain in the sets. If this fraction lies below the threshold, typically half of the elements in power-law and heavier-than-exponential-tailed distributions, then the original set size distribution is unrecoverable. We also discuss practical implications of our findings.
Fabricio Murai, Bruno Ribeiro 0001, Don Towsley, Pinghui Wang
IEEE J. Sel. Areas Commun.1
2012 Sampling directed graphs with random walks
abstract
Despite recent efforts to characterize complex networks such as citation graphs or online social networks (OSNs), little attention has been given to developing tools that can be used to characterize directed graphs in the wild, where no pre-processed data is available. The presence of hidden incoming edges but observable outgoing edges poses a challenge to characterize large directed graphs through crawling, as existing sampling methods cannot cope with hidden incoming links. The driving principle behind our random walk (RW) sampling method is to construct, in real-time, an undirected graph from the directed graph such that the random walk on the directed graph is consistent with one on the undirected graph. We then use the RW on the undirected graph to estimate the outdegree distribution. Our algorithm accurately estimates outdegree distributions of a variety of real world graphs. We also study the hardness of indegree distribution estimation when indegrees are latent (i.e., incoming links are only observed as outgoing edges). We observe that, in the same scenarios, indegree distribution estimates are highly innacurate unless the directed graph is highly symmetrical.
Bruno Ribeiro 0001, Pinghui Wang, Fabricio Murai, Don Towsley
INFOCOM3
2012 Heterogeneous download times in a homogeneous BitTorrent swarm
Fabricio Murai, Antônio Augusto de Aragão Rocha, Daniel R. Figueiredo 0001, Edmundo de Souza e Silva
Comput. Networks1