Roohollah Etemadi

dblp:188/3330 · DBLP profile ↗
← Back
8ranked-venue papers
7as first author
5since 2021 · last 2023
0000-0002-6099-0558ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 1 since 2021
YearPublicationVenuePosition
2023 Embedding-based team formation for community question answering
Roohollah Etemadi, Morteza Zihayat, Kuan Feng, Jason Adelman, Ebrahim Bagheri
Inf. Sci.1
2022 Feature-based question routing in community question answering platforms
Soroosh Sorkhani, Roohollah Etemadi, Amin Bigdeli, Morteza Zihayat, Ebrahim Bagheri
Inf. Sci.2
2022 PES: Priority Edge Sampling in Streaming Triangle Estimation
abstract
The number of triangles (hereafter denoted by$\Delta$) is an important metric to analyze massive graphs. It is also used to compute clustering coefficient in networks. This paper proposes a new algorithm called PES (Priority Edge Sampling) to estimate the number of triangles in the streaming model where we need to minimize the memory window. PES combines edge sampling and reservoir sampling. Compared with the state-of-the-art streaming algorithms, PES outperforms consistently. The results are verified extensively in 48 large real-world networks in different domains and structures. The performance ratio can be as large as 11. More importantly, the ratio grows with data size almost exponentially. This is especially important in the era of big data–while we can tolerate existing algorithms for smaller datasets, our method is indispensable when sampling very large data. In addition to empirical comparisons, we also proved that the estimator is unbiased, and derived the variance.
Roohollah Etemadi, Jianguo Lu
IEEE Trans. Big Data1
2021 Collaborative Experts Discovery in Social Coding Platforms
abstract
The popularity of online social coding (SC) platforms such as GitHub is growing due to their social functionalities and tremendous support during the product development lifecycle. The rich information of experts' contributions on repositories can be leveraged to recruit experts for new/existing projects. In this paper, we define the problem of collaborative experts finding in SC platforms. Given a project, we model an SC platform as an attributed heterogeneous network, learn latent representations of network entities in an end-to-end manner and utilize them to discover collaborative experts to complete a project. Extensive experiments on real-world datasets from GitHub indicate the superiority of the proposed approach over the state-of-the-art in terms of a range of performance measures.
Roohollah Etemadi, Morteza Zihayat, Kuan Feng, Jason Adelman, Ebrahim Bagheri
CIKM1
2021 OpenAttHetRL: An Open Source Toolkit for Attributed Heterogeneous Network Representation Learning
abstract
Learning the latent representations of entities based on their relationships and the data associated with them is an essential task in many applications such as ranking, recommendation systems, graph-based team formation, keyword search, and many more. However, the majority of existing techniques learn the latent representations of either network or textual data. Structural embedding techniques suffer from the sparsity of real-world networks. Attributes of nodes are a source of rich information to ameliorate network embedding vectors which are overlooked in the literature. Thus, most existing network representation learning tools capture structural information. This paper introduces an open-source toolkit called OpenAttHetRL to learn the latent representations of entities based on their both network and textual data in an end-to-end fashion. OpenAttHetRL is easy to employ and adapt for a variety of tasks including ranking, recommendation systems, and expert finding. OpenAttHetRL aims to provide a unified toolkit for data pre-processing, building and training models, and performing predictions for a downstream task. It employs a graph convolution network to capture the relationships among entities and a kernel pooling technique to preserve the similarity of their textual data in the embedding space. We use expert finding in community question answering systems to demonstrate how OpenAttHetRL can be trained to get latent representations of questions, their askers, tags, and answerers and find potential answerers of new questions.
Roohollah Etemadi, Morteza Zihayat, Kuan Feng, Jason Adelman, Ebrahim Bagheri
CIKM1
2017 Bias correction in clustering coefficient estimation
abstract
Clustering coefficient (C) is an important structural property to understand the complex structure of a graph. Calculating C is a computationally intensive task. Thereby, sampling-based methods have attracted substantial research for estimating C, and the closely related metric, the number of triangles. Unfortunately, widely used estimators for C are biased. We quantify the bias using Taylor expansion and find that the bias can be determined by the number of shared wedges and triangles in the sample. Based on the understanding of the bias, we give a new estimator that corrects the bias. The results are derived analytically and verified extensively in 56 networks ranging in different size and structure. The experiments reveal that the bias ranges widely from data to data. The relative bias can be as high as 4% or can be negative. For most of the graphs, the bias is small, although every graph does have a bias as quantified by our analytical results. Negative or small biases occur in online social networks where clustering coefficient is typically high. Positive and large biases typically occur in Web graphs, where there are nodes with high degrees but few neighboring nodes connecting with each other.
Roohollah Etemadi, Jianguo Lu
IEEE BigData1
2016 Identification of discriminative genes for predicting breast cancer subtypes
abstract
Breast cancer is a widespread cancer type in females and accounts for lots of cancer cases and cancer deaths in the world. Identifying the type of breast cancer plays a crucial role in selecting the best treatment. In this paper an optimized hierarchical model is proposed to predict the breast cancer subtype. Suitable filter feature selection methods and new hybrid feature selection methods are utilized in our model to find discriminative genes. The multi-class problem is handled using a proper classifier at each step in the hierarchical model to separate a subtype from the others. The parameters of each classifier are optimized to achieve a better performance. Our proposed model achieves 100% of accuracy for predicting the breast cancer subtypes using the same or even less number of genes.
Roohollah Etemadi, Abedalrhman Alkhateeb, Iman Rezaeian, Luis Rueda 0001
BIBM1
2016 Efficient Estimation of Triangles in Very Large Graphs
abstract
The number of triangles in a graph is an important metric for understanding the graph. It is also directly related to the clustering coefficient of a graph, which is one of the most important indicator for social networks. Counting the number of triangles is computationally expensive for very large graphs. Hence, estimation is necessary for large graphs, particularly for graphs that are hidden behind searchable interfaces where the graphs in their entirety are not available. For instance, user networks in Twitter and Facebook are not available for third parties to explore their properties directly. This paper proposes a new method to estimate the number of triangles based on random edge sampling. It improves the traditional random edge sampling by probing the edges that have a higher probability of forming triangles. The method outperforms the traditional method consistently, and can be better by orders of magnitude when the graph is very large. The result is demonstrated on 20 graphs, including the largest graphs we can find. More importantly, we proved the improvement ratio, and verified our result on all the datasets. The analytical results are achieved by simplifying the variances of the estimators based on the assumption that the graph is very large. We believe that such big data assumption can lead to interesting results not only in triangle estimation, but also in other sampling problems.
Roohollah Etemadi, Jianguo Lu, Yung H. Tsin
CIKM1