VLDB 2026 Research / reviewers in the wild / expert
Joseph Attieh
dblp:338/4572
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0001-6841-9877ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Generative modeling · 67% Machine translation · 33% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › synthetic data generation
LLM-based data generation |
0.9 | 1 | 2025 | Scaling Low-Resource MT via Synthetic Data Generation with LLMs · EMNLP 2025 |
Natural language and speech › Machine translation
low-resource machine translation |
0.9 | 1 | 2025 | Scaling Low-Resource MT via Synthetic Data Generation with LLMs · EMNLP 2025 |
Machine learning › Generative modeling
synthetic data generation |
0.9 | 1 | 2025 | Scaling Low-Resource MT via Synthetic Data Generation with LLMs · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
back-translation · 0.9LLM-based data generation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mitigating the Capacity Gap in Knowledge Distillation via Iterative TutoringabstractLarge language models (LLMs) have resulted in significant improvements in understanding and generating natural language. However, their deployment in resourceconstrained environments is limited by their high computational demands. Hence, Knowledge Distillation (KD) has emerged to address such challenges by enabling the transfer of knowledge from a large, pre-trained model (teacher) to a smaller, more efficient model (student). Yet, some bottlenecks exist in the effectiveness of this technique, such as the “capacity gap” between the teachers’ learning abilities and that of the student models, which may negatively impact the distilled model. We address this limitation by introducing a Tutor-Enhanced Iterative Distillation (TEID) to fill the capacity gap, by adding an intermediate-sized tutor model and selective learning strategy to the traditional distillation setup. To achieve further compression, the TEID is repeated iteratively on the tutor and the previously resultant student, with a new smaller student model. Empirical results on the GLUE benchmark show results in mitigating the model capacity gap, while showcasing the need to improve the efficiency and scalability of the distilled models. Sara Karam, Ralph Aouad, Joseph Attieh, Joe Tekli |
AICCSA | 3 |
| 2025 | Scaling Low-Resource MT via Synthetic Data Generation with LLMsabstractOna de Gibert, Joseph Attieh, Teemu Vahtola, Mikko Aulamo, Zihao Li, Raúl Vázquez, Tiancheng Hu, Jörg Tiedemann. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ona de Gibert Bonet, Joseph Attieh, Teemu Vahtola, Mikko Aulamo, Raúl Vázquez, Tiancheng Hu, Jörg Tiedemann |
EMNLP | 2 |
| 2023 | Optimizing the Performance of Text Classification Models by Improving the Isotropy of the Embeddings Using a Joint Loss Function
Joseph Attieh, Abraham Woubie, Vladimir Vlassov, Adrian Flanagan, Tom Bäckström |
ICDAR (5) | 1 |
| 2023 | Fast Text Classification using Lean Gradient Descent Feed Forward Neural Network for Category Feature AugmentationabstractText classification is a key task of the Natural Language Processing (NLP) field that aims at assigning predefined categories to textual documents. Performing text classification requires features that effectively represent the content and the meaning of textual documents. Selecting a suitable method for term weighting is of central importance and can improve the quality of the classification method. In this paper, we propose to a new text classification solution to perform Category-based Feature Augmentation (CFA) on the document representation. First, a term-category feature matrix is derived from a modified version of the supervised Term-Frequency Inverse-Category-Frequency (TF-ICF) weighting model. This is done by embedding the TF-ICF matrix in a one-layer feed-forward neural network. The latter is trained using the gradient descent algorithm allowing to iteratively update the term-category matrix until reaching convergence. The model produces category-based feature vector representations that are used to augment the document representations and perform the classification task. Experimental results on four benchmark datasets show that our lean model approach improves text classification accuracy and is significantly more efficient compared with its deep model alternatives. Joseph Attieh, Joe Tekli |
TrustCom | 1 |
| 2023 | Supervised term-category feature weighting for improved text classification
Joseph Attieh, Joe Tekli |
Knowl. Based Syst. | 1 |