Joseph Attieh

dblp:338/4572 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0001-6841-9877ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Generative modeling · 67% Machine translation · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › synthetic data generation
LLM-based data generation
0.912025
Scaling Low-Resource MT via Synthetic Data Generation with LLMs · EMNLP 2025
Natural language and speech › Machine translation
low-resource machine translation
0.912025
Scaling Low-Resource MT via Synthetic Data Generation with LLMs · EMNLP 2025
Machine learning › Generative modeling
synthetic data generation
0.912025
Scaling Low-Resource MT via Synthetic Data Generation with LLMs · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

back-translation · 0.9LLM-based data generation · 0.9
YearPublicationVenuePosition
2025 Mitigating the Capacity Gap in Knowledge Distillation via Iterative Tutoring
abstract
Large language models (LLMs) have resulted in significant improvements in understanding and generating natural language. However, their deployment in resourceconstrained environments is limited by their high computational demands. Hence, Knowledge Distillation (KD) has emerged to address such challenges by enabling the transfer of knowledge from a large, pre-trained model (teacher) to a smaller, more efficient model (student). Yet, some bottlenecks exist in the effectiveness of this technique, such as the “capacity gap” between the teachers’ learning abilities and that of the student models, which may negatively impact the distilled model. We address this limitation by introducing a Tutor-Enhanced Iterative Distillation (TEID) to fill the capacity gap, by adding an intermediate-sized tutor model and selective learning strategy to the traditional distillation setup. To achieve further compression, the TEID is repeated iteratively on the tutor and the previously resultant student, with a new smaller student model. Empirical results on the GLUE benchmark show results in mitigating the model capacity gap, while showcasing the need to improve the efficiency and scalability of the distilled models.
Sara Karam, Ralph Aouad, Joseph Attieh, Joe Tekli
AICCSA3
2025 Scaling Low-Resource MT via Synthetic Data Generation with LLMs
abstract
Ona de Gibert, Joseph Attieh, Teemu Vahtola, Mikko Aulamo, Zihao Li, Raúl Vázquez, Tiancheng Hu, Jörg Tiedemann. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Ona de Gibert Bonet, Joseph Attieh, Teemu Vahtola, Mikko Aulamo, Raúl Vázquez, Tiancheng Hu, Jörg Tiedemann
EMNLP2
2023 Optimizing the Performance of Text Classification Models by Improving the Isotropy of the Embeddings Using a Joint Loss Function
Joseph Attieh, Abraham Woubie, Vladimir Vlassov, Adrian Flanagan, Tom Bäckström
ICDAR (5)1
2023 Fast Text Classification using Lean Gradient Descent Feed Forward Neural Network for Category Feature Augmentation
abstract
Text classification is a key task of the Natural Language Processing (NLP) field that aims at assigning predefined categories to textual documents. Performing text classification requires features that effectively represent the content and the meaning of textual documents. Selecting a suitable method for term weighting is of central importance and can improve the quality of the classification method. In this paper, we propose to a new text classification solution to perform Category-based Feature Augmentation (CFA) on the document representation. First, a term-category feature matrix is derived from a modified version of the supervised Term-Frequency Inverse-Category-Frequency (TF-ICF) weighting model. This is done by embedding the TF-ICF matrix in a one-layer feed-forward neural network. The latter is trained using the gradient descent algorithm allowing to iteratively update the term-category matrix until reaching convergence. The model produces category-based feature vector representations that are used to augment the document representations and perform the classification task. Experimental results on four benchmark datasets show that our lean model approach improves text classification accuracy and is significantly more efficient compared with its deep model alternatives.
Joseph Attieh, Joe Tekli
TrustCom1
2023 Supervised term-category feature weighting for improved text classification
Joseph Attieh, Joe Tekli
Knowl. Based Syst.1