Georgos Siganos

dblp:41/5119 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 11 · 4 first-authorSystems, architecture and hardware · 3Security and privacy · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
9 papers
Network measurement and analytics · 33% Content delivery and video streaming · 26% Internet architecture and protocols · 15%
Databases, data mining, and information retrieval
4 papers
Data mining · 58% Graph data management · 39% Web and social media mining · 3%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Distributed systems · 34% High-performance computing · 27% Parallel and multicore computing · 18%
Network and information security
1 paper
Web and mobile security · 100%

Topics — the 30 heaviest of 44, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › structured data mining
graph mining
0.522017
Graph Data Mining with Arabesque · SIGMOD Conference 2017
Arabesque: a system for distributed graph mining · SOSP 2015
Content delivery and video streaming › peer-to-peer file sharing
bittorrent locality
0.432014
BitTorrent Locality and Transit TrafficReduction: When, Why, and at What Cost? · IEEE Trans. Parallel Distributed Syst. 2014
Deep diving into BitTorrent locality · INFOCOM 2011
Deep diving into BitTorrent locality · SIGMETRICS 2010
Data mining › pattern mining › graph pattern mining
clique enumeration
0.312017
Graph Data Mining with Arabesque · SIGMOD Conference 2017
Graph data management
graph partitioning
0.312017
Spinner: Scalable Graph Partitioning in the Cloud · ICDE 2017
Graph data management › graph analytics
large-scale graph analytics
0.312017
Spinner: Scalable Graph Partitioning in the Cloud · ICDE 2017
Graph data management
motif counting
0.312017
Graph Data Mining with Arabesque · SIGMOD Conference 2017
Data mining › structured data mining › graph mining
subgraph mining
0.312017
Graph Data Mining with Arabesque · SIGMOD Conference 2017
Content delivery and video streaming
peer-to-peer content distribution
0.222011
Deep diving into BitTorrent locality · INFOCOM 2011
Deep diving into BitTorrent locality · SIGMETRICS 2010
Data mining › pattern mining › graph pattern mining
frequent subgraph mining
0.212015
Arabesque: a system for distributed graph mining · SOSP 2015
Web and mobile security › online social network security
fake account detection
0.212015
Integro: Leveraging Victim Prediction for Robust Fake Account Detection in OSNs · NDSS 2015
Distributed systems
distributed data processing
0.212015
Arabesque: a system for distributed graph mining · SOSP 2015
High-performance computing › large-scale graph processing
distributed graph mining
0.212015
Arabesque: a system for distributed graph mining · SOSP 2015
Internet architecture and protocols
peer-to-peer networks
0.212014
BitTorrent Locality and Transit TrafficReduction: When, Why, and at What Cost? · IEEE Trans. Parallel Distributed Syst. 2014
Network measurement and analytics › topology discovery
missing link inference
0.222009
Lord of the links: a framework for discovering missing links in the internet topology · IEEE/ACM Trans. Netw. 2009
A Systematic Framework for Unearthing the Missing Links: Measurements and Impact · NSDI 2007
Network measurement and analytics
topology measurement
0.222009
Lord of the links: a framework for discovering missing links in the internet topology · IEEE/ACM Trans. Netw. 2009
A Systematic Framework for Unearthing the Missing Links: Measurements and Impact · NSDI 2007
Distributed systems › distributed data processing
data partitioning and replication
0.112012
The Little Engine(s) That Could: Scaling Online Social Networks · IEEE/ACM Trans. Netw. 2012
Network measurement and analytics › traffic analysis
peer-to-peer traffic analysis
0.112011
On blind mice and the elephant: understanding the network impact of a large distributed system · SIGCOMM 2011
Network measurement and analytics
traffic characterization
0.112011
On blind mice and the elephant: understanding the network impact of a large distributed system · SIGCOMM 2011
Network measurement and analytics
traffic measurement
0.112011
Deep diving into BitTorrent locality · INFOCOM 2011
Routing and switching › inter-domain routing
BGP
0.122007
Neighborhood Watch for Internet Routing: Can We Improve the Robustness of Internet Routing Today? · INFOCOM 2007
Analyzing BGP Policies: Methodology and Tool · INFOCOM 2004
Parallel and multicore computing
data distribution
0.112010
The little engine(s) that could: scaling online social networks · SIGCOMM 2010
Internet architecture and protocols › network topology
internet topology
0.112009
Lord of the links: a framework for discovering missing links in the internet topology · IEEE/ACM Trans. Netw. 2009
Parallel and multicore computing › graph processing
parallel graph analytics
0.112017
Graph Data Mining with Arabesque · SIGMOD Conference 2017
Network management and operations › fault management
fault diagnosis
0.112007
Neighborhood Watch for Internet Routing: Can We Improve the Robustness of Internet Routing Today? · INFOCOM 2007
Network management and operations › network monitoring
route leak detection
0.112007
Neighborhood Watch for Internet Routing: Can We Improve the Robustness of Internet Routing Today? · INFOCOM 2007
Internet of things and sensor networks › wireless sensor network › network diagnosis
routing anomaly detection
0.112007
Neighborhood Watch for Internet Routing: Can We Improve the Robustness of Internet Routing Today? · INFOCOM 2007
Web and social media mining
online social networks
0.112015
Integro: Leveraging Victim Prediction for Robust Fake Account Detection in OSNs · NDSS 2015
High-performance computing
large-scale graph processing
0.112015
Arabesque: a system for distributed graph mining · SOSP 2015
Content delivery and video streaming
peer-to-peer streaming
0.112014
BitTorrent Locality and Transit TrafficReduction: When, Why, and at What Cost? · IEEE Trans. Parallel Distributed Syst. 2014
Network management and operations
configuration verification
0.012004
Analyzing BGP Policies: Methodology and Tool · INFOCOM 2004

Methods — techniques the papers use, named apart from their topics

victim prediction · 0.4machine learning · 0.4vertex-centric pregel abstraction · 0.3distributed hash table · 0.3trace analysis · 0.2measurement · 0.2traffic matrix estimation · 0.2middleware · 0.1large-scale measurement · 0.1dataset analysis · 0.1replication · 0.1registry data analysis · 0.1measurement study · 0.1filter generation · 0.1validation against routing tables · 0.0data inference · 0.0
YearPublicationVenuePosition
2017 QFrag: distributed graph search via subgraph isomorphism
abstract
This paper introduces QFrag, a distributed system for graph search on top of bulk synchronous processing (BSP) systems such as MapReduce and Spark. Searching for patterns in graphs is an important and computationally complex problem. Most current distributed search systems scale to graphs that do not fit in main memory by partitioning the input graph. For analytical queries, however, this approach entails running expensive distributed joins on large intermediate data.
Marco Serafini, Gianmarco De Francisci Morales, Georgos Siganos
SoCC3
2017 Spinner: Scalable Graph Partitioning in the Cloud
abstract
In this paper, we present a graph partitioning algorithm to partition graphs with trillions of edges. To achieve such scale, our solution leverages the vertex-centric Pregel abstraction provided by Giraph, a system for large-scale graph analytics. We designed our algorithm to compute partitions with high locality and fair balance, and focused on the characteristics necessary to reach wide adoption by practitioners in production. Our solution can (i) scale to massive graphs and thousands of compute cores, (ii) efficiently adapt partitions to changes to graphs and compute environments, and (iii) seamlessly integrate in existing systems without additional infrastructure. We evaluate our solution on the Facebook and Instagram graphs, as well as on other large-scale, real-world graphs. We show that it is scalable and computes partitionings with quality comparable, and sometimes outperforming, existing solutions. By integrating the computed partitionings in Giraph, we speedup various real-world applications by up to a factor of 5.6 compared to default hash-partitioning.
Claudio Martella, Dionysios Logothetis, Andreas Loukas, Georgos Siganos
ICDE4
2017 Graph Data Mining with Arabesque
abstract
Graph data mining is defined as searching in an input graph for all subgraphs that satisfy some property that makes them interesting to the user. Examples of graph data mining problems include frequent subgraph mining, counting motifs, and enumerating cliques. These problems differ from other graph processing problems such as PageRank or shortest path in that graph data mining requires searching through an exponential number of subgraphs. Most current parallel graph analytics systems do not provide good support for graph data mining. One notable exception is Arabesque, a system that was built specifically to support graph data mining. Arabesque provides a simple programming model to express graph data mining computations, and a highly scalable and efficient implementation of this model, scaling to billions of subgraphs on hundreds of cores. This demonstration will showcase the Arabesque system, focusing on the end-user experience and showing how Arabesque can be used to simply and efficiently solve practical graph data mining problems that would be difficult with other systems.
Eslam Hussein, Abdurrahman Ghanem, Vinícius Vitor dos Santos Dias, Carlos H. C. Teixeira, Ghadeer AbuOda, Marco Serafini, Georgos Siganos, Gianmarco De Francisci Morales, Ashraf Aboulnaga, Mohammed J. Zaki
SIGMOD Conference7
2016 Íntegro: Leveraging victim prediction for robust fake account detection in large scale OSNs
Yazan Boshmaf, Dionysios Logothetis, Georgos Siganos, Jorge Lería, José Lorenzo, Matei Ripeanu, Konstantin Beznosov, Hassan Halawa
Comput. Secur.3
2015 Integro: Leveraging Victim Prediction for Robust Fake Account Detection in OSNs
Yazan Boshmaf, Dionysios Logothetis, Georgos Siganos, Jorge Lería, José Lorenzo, Matei Ripeanu, Konstantin Beznosov
NDSS3
2015 Arabesque: a system for distributed graph mining
abstract
Distributed data processing platforms such as MapReduce and Pregel have substantially simplified the design and deployment of certain classes of distributed graph analytics algorithms. However, these platforms do not represent a good match for distributed graph mining problems, as for example finding frequent subgraphs in a graph. Given an input graph, these problems require exploring a very large number of subgraphs and finding patterns that match some "interestingness" criteria desired by the user. These algorithms are very important for areas such as social networks, semantic web, and bioinformatics.
Carlos H. C. Teixeira, Alexandre J. Fonseca, Marco Serafini, Georgos Siganos, Mohammed J. Zaki, Ashraf Aboulnaga
SOSP4
2014 BitTorrent Locality and Transit TrafficReduction: When, Why, and at What Cost?
abstract
A substantial amount of work has recently gone into localizing BitTorrent traffic within an ISP in order to avoid excessive and often times unnecessary transit costs. Several architectures and systems have been proposed and the initial results from specific ISPs and a few torrents have been encouraging. In this work we attempt to deepen and scale our understanding of locality and its potential. Looking at specific ISPs, we consider tens of thousands of concurrent torrents, and thus capture ISP-wide implications that cannot be appreciated by looking at only a handful of torrents. Second, we go beyond individual case studies and present results for few thousands ISPs represented in our data set of up to 40K torrents involving more than 3.9M concurrent peers and more than 20M in the course of a day spread in 11K ASes. Finally, we develop scalable methodologies that allow us to process this huge data set and derive accurate traffic matrices of torrents. Using the previous methods we obtain the following main findings: i) Although there are a large number of very small ISPs without enough resources for localizing traffic, by analyzing the 100 largest ISPs we show that Locality policies are expected to significantly reduce the transit traffic with respect to the default random overlay construction method in these ISPs; ii) contrary to the popular belief, increasing the access speed of the clients of an ISP does not necessarily help to localize more traffic; iii) by studying several real ISPs, we have shown that soft speed-aware locality policies guarantee win-win situations for ISPs and end users. Furthermore, the maximum transit traffic savings that an ISP can achieve without limiting the number of inter-ISP overlay links is bounded by “unlocalizable” torrents with few local clients. The application of restrictions in the number of inter-ISP links leads to a higher transit traffic reduction but the QoS of clients downloading “unlocalizable” torrents would be severely harmed.
Rubén Cuevas Rumín, Nikolaos Laoutaris, Xiaoyuan Yang 0001, Georgos Siganos, Pablo Rodriguez 0001
IEEE Trans. Parallel Distributed Syst.4
2012 The Little Engine(s) That Could: Scaling Online Social Networks
abstract
The difficulty of partitioning social graphs has introduced new system design challenges for scaling of online social networks (OSNs). Vertical scaling by resorting to full replication can be a costly proposition. Scaling horizontally by partitioning and distributing data among multiple servers using, for e.g., distributed hash tables (DHTs), can suffer from expensive interserver communication. Such challenges have often caused costly rearchitecting efforts for popular OSNs like Twitter and Facebook. We design, implement, and evaluate SPAR, a Social Partitioning and Replication middleware that mediates transparently between the application and the database layer of an OSN. SPAR leverages the underlying social graph structure in order to minimize the required replication overhead for ensuring that users have their neighbors' data colocated in the same machine. The gains from this are multifold: Application developers can assume local semantics, i.e., develop as they would for a single machine; scalability is achieved by adding commodity machines with low memory and network I/O requirements; and N+K redundancy is achieved at a fraction of the cost. We provide a complete system design, extensive evaluation based on datasets from Twitter, Orkut, and Facebook, and a working implementation. We show that SPAR incurs minimum overhead, can help a well-known Twitter clone reach Twitter's scale without changing a line of its application logic, and achieves higher throughput than Cassandra, a popular key-value store database.
Josep M. Pujol, Vijay Erramilli, Georgos Siganos, Xiaoyuan Yang 0001, Nikolaos Laoutaris, Parminder Chhabra, Pablo Rodriguez 0001
IEEE/ACM Trans. Netw.3
2011 Deep diving into BitTorrent locality
abstract
A substantial amount of work has recently gone into localizing BitTorrent traffic within an ISP in order to avoid excessive and often times unnecessary transit costs. Several architectures and systems have been proposed and the initial results from specific ISPs and a few torrents have been encouraging. In this work we attempt to deepen and scale our understanding of locality and its potential. Looking at specific ISPs, we consider tens of thousands of concurrent torrents, and thus capture ISP-wide implications that cannot be appreciated by looking at only a handful of torrents. Secondly, we go beyond individual case studies and present results for the top 100 ISPs in terms of number of users represented in our dataset of up to 40K torrents involving more than 3.9M concurrent peers and more than 20M in the course of a day spread in 11K ASes. We develop scalable methodologies that allow us to process this huge dataset and get concrete quantitative answers rather than qualitative speculations to questions like: “what is the minimum and the maximum transit traffic reduction across hundreds of ISPs?”, “what are the win-win boundaries for ISPs and their users?”, “what is the maximum amount of transit traffic that can be localized without requiring fine-grained control of inter-AS overlay connections?”.
Rubén Cuevas Rumín, Nikolaos Laoutaris, Xiaoyuan Yang 0001, Georgos Siganos, Pablo Rodriguez 0001
INFOCOM4
2011 On blind mice and the elephant: understanding the network impact of a large distributed system
abstract
A thorough understanding of the network impact of emerging large-scale distributed systems -- where traffic flows and what it costs -- must encompass users' behavior, the traffic they generate and the topology over which that traffic flows. In the case of BitTorrent, however, previous studies have been limited by narrow perspectives that restrict such analysis.
John S. Otto, Mario A. Sánchez, David R. Choffnes, Fabián E. Bustamante, Georgos Siganos
SIGCOMM5
2010 The little engine(s) that could: scaling online social networks
abstract
The difficulty of scaling Online Social Networks (OSNs) has introduced new system design challenges that has often caused costly re-architecting for services like Twitter and Facebook. The complexity of interconnection of users in social networks has introduced new scalability challenges. Conventional vertical scaling by resorting to full replication can be a costly proposition. Horizontal scaling by partitioning and distributing data among multiples servers - e.g. using DHTs - can lead to costly inter-server communication.
Josep M. Pujol, Vijay Erramilli, Georgos Siganos, Xiaoyuan Yang 0001, Nikolaos Laoutaris, Parminder Chhabra, Pablo Rodriguez 0001
SIGCOMM3
2010 Deep diving into BitTorrent locality
abstract
A substantial amount of work has recently gone into localizing BitTorrent traffic within an ISP in order to avoid excessive and often times unnecessary transit costs. In this work we aim to answer yet unanswered questions such as: what is the minimum and the maximum transit traffic reduction across hundreds of ISPs?, what are the win-win boundaries for ISPs and their users?, what is the maximum amount of transit traffic that can be localized without requiring fine-grained control of inter-AS overlay connections?, what is the impact to transit traffic from upgrades of residential broadband speeds?.
Rubén Cuevas Rumín, Nikolaos Laoutaris, Xiaoyuan Yang 0001, Georgos Siganos, Pablo Rodriguez 0001
SIGMETRICS4
2009 Monitoring the Bittorrent Monitors: A Bird's Eye View
Georgos Siganos, Josep M. Pujol, Pablo Rodriguez 0001
PAM1
2009 Lord of the links: a framework for discovering missing links in the internet topology
Yihua He, Georgos Siganos, Michalis Faloutsos, Srikanth V. Krishnamurthy
IEEE/ACM Trans. Netw.2
2007 Neighborhood Watch for Internet Routing: Can We Improve the Robustness of Internet Routing Today?
abstract
Protecting BGP routing from errors and malice is one of the next big challenges for Internet routing. Several approaches have been proposed that attempt to capture and block routing anomalies in a proactive way. In practice, the difficulty of deploying such approaches limits their usefulness. We take a different approach: we start by requiring a solution that can be easily implemented now. With this goal in mind, we consider ourselves situated at an AS, and ask the question: how can I detect erroneous or even suspicious routing behavior? We respond by developing a systematic methodology and a tool to identify such updates by utilizing existing public and local information. Specifically, we process and use the allocation records from the Regional Internet Registries (RIR), the local policy of the AS, and records used to generate filters from Internet Routing Registries (IRR). Using our approach, we can automatically detect routing leaks. Additionally, we identify some simple organizational and procedural issues that would significantly improve the usefulness of the information of the registries. Finally, we propose an initial set of rules with which an ISP can react to routing problems in a way that is systematic, and thus, could be automated.
Georgos Siganos, Michalis Faloutsos
INFOCOM1
2007 A Systematic Framework for Unearthing the Missing Links: Measurements and Impact
Yihua He, Georgos Siganos, Michalis Faloutsos, Srikanth V. Krishnamurthy
NSDI2
2004 Analyzing BGP Policies: Methodology and Tool
abstract
The robustness of the Internet relies heavily on the robustness of BGP routing. BGP is the glue that holds the Internet together: it is the common language of the routers that interconnect networks or autonomous systems (AS). The robustness of BGP and our ability to manage it effectively is hampered by the limited global knowledge and lack of coordination between autonomous systems. One of the few efforts to develop a globally analyzable and secure Internet is the creation of the Internet routing registries (IRRs). IRRs provide a voluntary detailed repository of BGP policy information. The IRR effort has not reached its full potential because of two reasons: a) extracting useful information is far from trivial, and b) its accuracy of the data is uncertain. In this paper, we develop a methodology and a tool (Nemecis) to extract and infer information from IRR and validate it against BGP routing tables. In addition, using our tool, we quantify the accuracy of the information of IRR. We find that IRR has a lot of inaccuracies, but also contains significant and unique information. Finally, we show that our tool can identify and extract the correct information from IRR discarding erroneous data. In conclusion, our methodology and tool close the gap in the IRR vision for an analyzable Internet repository at the BGP level
Georgos Siganos, Michalis Faloutsos
INFOCOM1
2003 Power laws and the AS-level internet topology
abstract
We study and characterize the topology of the Internet at the autonomous system (AS) level. First, we show that the topology can be described efficiently with power laws. The elegance and simplicity of the power laws provide a novel perspective into the seemingly uncontrolled Internet structure. Second, we show that power laws have appeared consistently over the last five years. We also observe that the power laws hold even in the most recent and more complete topology with correlation coefficient above 99% for the degree-based power law. In addition, we study the evolution of the power-law exponents over the five-year interval and observe a variation for the degree-based power law of less than 10%. Thirdly, we provide relationships between the exponents and other topological metrics.
Georgos Siganos, Michalis Faloutsos, Petros Faloutsos, Christos Faloutsos
IEEE/ACM Trans. Netw.1
2002 BGP routing: a study at large time scale
abstract
We conduct a longterm analysis of BGP routing properties, such as path stability. Our work complements previous studies that examine the BGP routing behavior at smaller time scales, such as its convergence to a routing update. We focus on properties that would reflect the BGP evolution due to growth, policy, and business reasons. We use daily snapshots of a number of BGP tables and we study the evolution of the paths for every IP-prefix advertised between two autonomous systems (source AS, destination AS, IP-prefix). In a high level, we observe that BGP routing is characterized by a) fairly robust routing with usually few dominating paths per prefix, b) routing richness in the advertised paths and c) a significant number of short lived IP-prefix advertisements.
Georgos Siganos, Michalis Faloutsos
GLOBECOM1
2001 A simple conceptual model for the Internet topology
abstract
In this paper, we develop a conceptual visual model for the Internet inter-domain topology. Recently, power-laws were used to describe the topology concisely. Despite their success, the power-laws do not help us visualize the topology, ie, draw the topology on paper by hand. In this paper, we deal with the following questions: can we identify a hierarchy in the Internet; how can I represent the network in an abstract graphical way? The focus of this paper is threefold. First, we characterize nodes using three metrics of topological "importance"', which we later use to identify a sense of hierarchy. Second, we identify some new topological properties. We then find that the Internet has a highly connected core and identify layers of nodes in decreasing importance surrounding the core. Finally, we show that our observations suggest an intuitive model. The topology can be seen as a jellyfish, where the core is in the middle of the cap, and one-degree nodes form its legs.
Sudhir Leslie Tauro, Christopher R. Palmer, Georgos Siganos, Michalis Faloutsos
GLOBECOM3