Soheil Hassas Yeganeh

dblp:86/1462 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
2since 2021 · last 2023
0009-0007-7282-5158ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Performance modeling and evaluation · 46% Cloud and datacenter computing · 40% Distributed systems · 12%
Computer networks
2 papers
Network measurement and analytics · 88% Datacenter networks · 12%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%

Topics — the 7 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Network measurement and analytics
passive measurement
0.712023
Fathom: Understanding Datacenter Application Network Performance · SIGCOMM 2023
Performance modeling and evaluation
workload characterization
0.712023
A Cloud-Scale Characterization of Remote Procedure Calls · SOSP 2023
Distributed systems
remote procedure call
0.212023
A Cloud-Scale Characterization of Remote Procedure Calls · SOSP 2023
Data integration and cleaning
schema matching
0.212013
Discovering Linkage Points over Web Data · Proc. VLDB Endow. 2013
Network measurement and analytics
traffic generation
0.112010
Caliper: a tool to generate precise and closed-loop traffic · SIGCOMM 2010
Performance modeling and evaluation
benchmarking
0.112010
Caliper: a tool to generate precise and closed-loop traffic · SIGCOMM 2010
Reconfigurable computing and FPGAs › FPGA-based network processing
FPGA-based packet processing
0.012010
Caliper: a tool to generate precise and closed-loop traffic · SIGCOMM 2010

Methods — techniques the papers use, named apart from their topics

kernel instrumentation · 0.7cloud-scale measurement · 0.7RPC stack instrumentation · 0.7hardware-based packet generation · 0.2
YearPublicationVenuePosition
2023 Fathom: Understanding Datacenter Application Network Performance
abstract
We describe our experience with Fathom, a system for identifying the network performance bottlenecks of any service running in the Google fleet. Fathom passively samples RPCs, the principal unit of work for services. It segments the overall latency into host and network components with kernel and RPC stack instrumentation. It records these detailed latency metrics, along with detailed transport connection state, for every sampled RPC. This lets us determine if the completion is constrained by the client, network or server. To scale while enabling analysis, we also aggregate samples into distributions that retain multi-dimensional breakdowns. This provides us with a macroscopic view of individual services. Fathom runs globally in our datacenters for all production traffic, where it monitors billions of TCP connections 24x7. For five years Fathom has been our primary tool for troubleshooting service network issues and assessing network infrastructure changes. We present case studies to show how it has helped us improve our production services.
Mubashir Adnan Qureshi, Junhua Yan, Yuchung Cheng, Soheil Hassas Yeganeh, Yousuk Seung, Neal Cardwell, Willem de Bruijn, Van Jacobson, Jasleen Kaur 0001, David Wetherall, Amin Vahdat
SIGCOMM4
2023 A Cloud-Scale Characterization of Remote Procedure Calls
abstract
The global scale and challenging requirements of modern cloud applications have led to the development of complex, widely distributed, service-oriented applications. One enabler of such applications is the remote procedure call (RPC), which provides location-independent communication and hides the myriad of cloud communication complexities and requirements within the RPC stack. Understanding RPCs is thus one key to understanding the behavior of cloud applications. While there have been numerous studies of RPCs in distributed systems, as well as attempts to optimize RPC overheads with both software and hardware, there is still a lack of knowledge about the characteristics of RPCs "in the wild" in the modern cloud environment.
Korakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu 0001, Hassan M. G. Wassel, Soheil Hassas Yeganeh, Alex C. Snoeren, Arvind Krishnamurthy, David E. Culler, Henry M. Levy
SOSP6
2014 Beehive: Towards a Simple Abstraction for Scalable Software-Defined Networking
abstract
Simplicity is a prominent advantage of Software-Defined Networking (SDN), and is often exemplified by implementing a complicated control logic as a simple control application on a centralized controller. In practice, however, SDN controllers turn into distributed systems due to performance and reliability limitations, and the supposedly simple control applications transform into complex logics that demand significant effort to design and optimize.
Soheil Hassas Yeganeh, Yashar Ganjali
HotNets1
2013 Discovering Linkage Points over Web Data
abstract
A basic step in integration is the identification of linkage points, i.e., finding attributes that are shared (or related) between data sources, and that can be used to match records or entities across sources. This is usually performed using a match operator, that associates attributes of one database to another. However, the massive growth in the amount and variety of unstructured and semi-structured data on the Web has created new challenges for this task. Such data sources often do not have a fixed pre-defined schema and contain large numbers of diverse attributes. Furthermore, the end goal is not schema alignment as these schemas may be too heterogeneous (and dynamic) to meaningfully align. Rather, the goal is to align any overlapping data shared by these sources. We will show that even attributes with different meanings (that would not qualify as schema matches) can sometimes be useful in aligning data. The solution we propose in this paper replaces the basic schema-matching step with a more complex instance-based schema analysis and linkage discovery. We present a framework consisting of a library of efficient lexical analyzers and similarity functions, and a set of search algorithms for effective and efficient identification of linkage points over Web data. We experimentally evaluate the effectiveness of our proposed algorithms in real-world integration scenarios in several domains.
Oktie Hassanzadeh, Ken Q. Pu, Soheil Hassas Yeganeh, Renée J. Miller, Lucian Popa 0001, Mauricio A. Hernández, C. T. Howard Ho
Proc. VLDB Endow.3
2012 Rethinking end-to-end congestion control in software-defined networks
abstract
TCP is designed to operate in a wide range of networks. Without any knowledge of the underlying network and traffic characteristics, TCP is doomed to continuously increase and decrease its congestion window size to embrace changes in network or traffic. In light of emerging popularity of centrally controlled Software-Defined Networks (SDNs), one might wonder whether we can take advantage of the global network view available at the controller to make faster and more accurate congestion control decisions. In this paper, we identify the need and the underlying requirements for a congestion control adaptation mechanism. To this end, we propose OpenTCP as a TCP adaptation framework that works in SDNs. OpenTCP allows network operators to define rules for tuning TCP as a function of network and traffic conditions. We also present a preliminary implementation of OpenTCP in a ~4000 node data center.
Manya Ghobadi, Soheil Hassas Yeganeh, Yashar Ganjali
HotNets2
2012 CUTE: Traffic Classification Using TErms
abstract
Among different traffic classification approaches, Deep Packet Inspection (DPI) methods are considered as the most accurate. These methods, however, have two drawbacks: (i) they are not efficient since they use complex regular expressions as protocol signatures, and (ii) they require manual intervention to generate and maintain signatures, partly due to the signature complexity. In this paper, we present CUTE, an automatic traffic classification method, which relies on sets of weighted terms as protocol signatures. The key idea behind CUTE is an observation that, given appropriate weights, the occurrence of a specific term is more important than the relative location of terms in a flow. This observation is based on experimental evaluations as well as theoretical analysis, and leads to several key advantages over previous classification techniques: (i) CUTE is extremely faster than other classification schemes since matching flows with weighed terms is significantly faster than matching regular expressions; (ii) CUTE can classify network traffic using only the first few bytes of the flows in most cases; and (iii) Unlike most existing classification techniques, CUTE can be used to classify partial (or even slightly modified) flows. Even though CUTE replaces complex regular expressions with a set of simple terms, using theoretical analysis and experimental evaluations (based on two large packet traces from tier-one ISPs), we show that its accuracy is as good as or better than existing complex classification schemes, i.e. CUTE achieves precision and recall rates of more than 90%. Additionally, CUTE can successfully classify more than half of flows that other DPI methods fail to classify.
Soheil Hassas Yeganeh, Milad Eftekhar, Yashar Ganjali, Ram Keralapura, Antonio Nucci
ICCCN1
2011 Linking Semistructured Data on the Web
Oktie Hassanzadeh, Soheil Hassas Yeganeh, Renée J. Miller
WebDB2
2010 A Novel Method to Find Appropriate epsilon for DBSCAN
Jamshid Esmaelnejad, Jafar Habibi, Soheil Hassas Yeganeh
ACIIDS (1)3
2010 Caliper: a tool to generate precise and closed-loop traffic
abstract
Generating realistic and responsive traffic that reflects different network conditions is a challenging problem associated with performing valid experiments in network testbeds. In this work, we preset Caliper, a highly precise traffic generation tool, built on NetThreads, a flexible platform that we have created for developing packet processing applications on FPGA-based devices and the NetFPGA in particular. We will demonstrate the effect of ad-hoc inter-departure times on a commodity NIC compared to precisely timed inter-departures with Caliper. Both NetThreads and Caliper are available as free software to download.
Manya Ghobadi, Martin Labrecque, Geoffrey Salmon, Kaveh Aasaraai, Soheil Hassas Yeganeh, Yashar Ganjali, J. Gregory Steffan
SIGCOMM5
2009 An approximation algorithm for finding skeletal points for density based clustering approaches
abstract
Clustering is the problem of finding relations in a data set in an supervised manner. These relations can be extracted using the density of a data set, where density of a data point is defined as the number of data points around it. To find the number of data points around another point, region queries are adopted. Region queries are the most expensive construct in density based algorithm, so it should be optimized to enhance the performance of density based clustering algorithms specially on large data sets. Finding the optimum set of region queries to cover all the data points has been proven to be NP-complete. This optimum set is called the skeletal points of a data set. In this paper, we proposed a generic algorithms which fires region queries at most 6 times the optimum number of region queries (has 6 as approximation factor). Also, we have extend this generic algorithm to create a DBSCAN (the most wellknown density based algorithm) derivative, named ADBSCAN. Presented experimental results show that ADBSCAN has a better approximation to DBSCAN than the DBRS-H (the most well-known randomized density based algorithm) in terms of performance and quality of clustering, specially for large data sets.
Soheil Hassas Yeganeh, Jafar Habibi, Hassan Abolhassani, Mahdi Abbaspour Tehrani, Jamshid Esmaelnezhad
CIDM1
2007 Approximation Algorithms for Software Component Selection Problem
abstract
Today's software systems are more frequently composed from preexisting commercial or non-commercial components and connectors. These components provide complex and independent functionality and are engaged in complex interactions. Component-Based Software Engineering (CBSE) is concerned with composing, selecting and designing such components. As the popularity of this approach and hence number of commercially available software components grows, selecting a set of components to satisfy a set of requirements while minimizing cost is becoming more difficult. This problem necessitates the design of efficient algorithms to automate component selection for software developing organizations. We address this challenge through analysis of Component Selection, the NP-complete process of selecting a minimal cost set of components to satisfy a set of objectives. Due to the high order of computational complexity of this problem, we examine approximating solutions that make the component selection process practicable. We adapt a greedy approach and a genetic algorithm to approximate this problem. We examined the performance of studied algorithms on a set of selected ActiveX components. Comparing the results of these two algorithms with the choices made by a group of human experts shows that we obtain better results using these approximation algorithms.
Nima Haghpanah, Shahrouz Moaven, Jafar Habibi, Mehdi Kargar, Soheil Hassas Yeganeh
APSEC5