EDBT 2026 Demo / reviewers in the wild / expert
Jafar Tanha
dblp:67/8777
· DBLP profile ↗
8ranked-venue papers in the field
4as first author
5since 2021 · last 2026
0000-0002-0779-6027ORCID · verified
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 3 (1 first)Data Mining & Knowledge Discovery · 2 (2 first)Database Systems & Data Management · 1Information Retrieval & Web Search · 1Other / Interdisciplinary · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A2RCMatch: dual-attention framework for reliable sample selection and consistency regularization in semi-supervised learning
Razieh Mohammadi, Jafar Tanha, M. A. Balafar, Mahdi Baghaei Oskouei, Hamed Jalili, Nima Rasi Baghmishe, Pouya Afraz |
Inf. Sci. | 2 |
| 2026 | Beyond Predefined Clusters: A Comprehensive Review of Clustering Methods for Unknown Numbers of ClustersabstractClustering is an unsupervised learning task that groups data points by their inherent similarities. Nonautomatic clustering algorithms face significant challenges when the true number of clusters is unknown or changes dynamically, as they require this number to be predefined. This paper provides a comprehensive review of automatic clustering algorithms specifically designed to handle such uncertainty. In this paper, these algorithms are systematically classified based on three key perspectives: clustering framework (classical vs. deep), clustering strategy (e.g., density-based, model based, graph-theoretic, subspace methods), and the use of labeled data (unsupervised vs. semi-supervised). We analyze each algorithm based on its core principles, key contributions, strengths, and limitations. Furthermore, we address the current challenges in this area and propose future research directions to enhance the scalability, robustness, and effectiveness of automatic clustering algorithms. Nazila Pourhaji Aghayengejeh, M. A. Balafar, Jafar Tanha, M. Alper Selver |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Graph theory-based semi-supervised self-training for data stream classification and emerging class detection
Negin Samadi, Jafar Tanha, Mahdi Jalili |
Inf. Sci. | 2 |
| 2022 | CPSSDS: Conformal prediction for semi-supervised classification on data streams
Jafar Tanha, Negin Samadi, Yousef Abdi, Nazila Razzaghi |
Inf. Sci. | 1 |
| 2021 | A Selection Metric for semi-supervised learning based on neighborhood construction
Mona Emadi, Jafar Tanha, Mohammad Ebrahim Shiri, Mehdi Hosseinzadeh Aghdam |
Inf. Process. Manag. | 2 |
| 2015 | Crossing the lines: making optimal use of context in line-based Handwritten Text RecognitionabstractHand-written text recognition (HTR) is often carried out line-by-line: the decoding of text lines is carried out independently. This approach is known to deteriorate recognition accuracy of words and characters close to the line boundaries. The present study investigates this issue from the point of view of the language modeling component of the HTR system. Obviously, lack of linguistic context may be one of the reasons for loss of accuracy, but it certainly is not the only factor in play. We seek to clarify to which extent the problem can be influenced by the language modeling component of the system. We first discuss how to develop adapted language models which significantly improve HTR performance in general. We then focus on the deployment of methods to improve accuracy at line boundaries. The final result is an efficient approach which significantly improves HTR accuracy without changing the basic HTR system setup. Jafar Tanha, Jesse de Does, Katrien Depuydt, Joan-Andreu Sánchez |
ICDAR | 1 |
| 2013 | Multiclass Semi-Supervised Boosting Using Similarity LearningabstractIn this paper, we consider the multiclass semi-supervised classification problem. A boosting algorithm is proposed to solve the multiclass problem directly. The proposed multiclass approach uses a new multiclass loss function, which includes two terms. The first term is the cost of the multiclass margin and the second term is a regularization term on unlabeled data. The regularization term is used to minimize the inconsistency between the pair wise similarity and the classifier predictions. It assigns the soft labels weighted with the similarity between unlabeled and labeled examples. We then derive a boosting algorithm, named CD-MSSBoost, from the proposed loss function using coordinate gradient descent. The derived algorithm is further used for learning optimal similarity function for a given data. Our experiments on a number of UCI datasets show that CD-MSSBoost outperforms the state-of-the-art methods to multiclass semi-supervised learning. Jafar Tanha, Mohammad J. Saberian, Maarten van Someren |
ICDM | 1 |
| 2012 | An AdaBoost Algorithm for Multiclass Semi-supervised LearningabstractWe present an algorithm for multiclass Semi-Supervised learning which is learning from a limited amount of labeled data and plenty of unlabeled data. Existing semi-supervised algorithms use approaches such as one-versus-all to convert the multiclass problem to several binary classification problems which is not optimal. We propose a multiclass semi-supervised boosting algorithm that solves multiclass classification problems directly. The algorithm is based on a novel multiclass loss function consisting of the margin cost on labeled data and two regularization terms on labeled and unlabeled data. Experimental results on a number of UCI datasets show that the proposed algorithm performs better than the state-of-the-art boosting algorithms for multiclass semi-supervised learning. Jafar Tanha, Maarten van Someren, Hamideh Afsarmanesh |
ICDM | 1 |