Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Nasser Zalmout

dblp:142/9396 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Information extraction and text analysis · 59% Question answering and dialogue systems · 11% Knowledge representation and reasoning · 11%
Databases, data mining, and information retrieval
3 papers
Knowledge graphs · 56% Data mining · 36% Information retrieval · 8%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › morphological analysis
morphological tagging
0.822020
Joint Diacritization, Lemmatization, Normalization, and Fine-Grained Morphological Tagging · ACL 2020
Adversarial Multitask Learning for Joint Multi-Feature and Multi-Dialect Morphological Modeling · ACL (1) 2019
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction
attribute extraction
0.512021
PAM: Understanding Product Images in Cross Product Category Attribute Extraction · KDD 2021
Natural language and speech › Information extraction and text analysis › relation extraction
attribute value extraction
0.512021
AdaTag: Multi-Attribute Value Extraction from Product Profiles with Adaptive Decoding · ACL/IJCNLP (1) 2021
Natural language and speech › Question answering and dialogue systems
conversational search
0.512021
End-to-End Conversational Search for Online Shopping with Utterance Transfer · EMNLP (1) 2021
Data mining › text mining › information extraction
attribute extraction
0.512021
PAM: Understanding Product Images in Cross Product Category Attribute Extraction · KDD 2021
Knowledge graphs
knowledge graph construction
0.512021
PAM: Understanding Product Images in Cross Product Category Attribute Extraction · KDD 2021
Knowledge graphs › domain-specific knowledge graph
product knowledge graph
0.512021
All You Need to Know to Build a Product Knowledge Graph · KDD 2021
Natural language and speech › Information extraction and text analysis › morphological analysis
lemmatization
0.412020
Joint Diacritization, Lemmatization, Normalization, and Fine-Grained Morphological Tagging · ACL 2020
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.412019
Adversarial Multitask Learning for Joint Multi-Feature and Multi-Dialect Morphological Modeling · ACL (1) 2019
Natural language and speech › Information extraction and text analysis
text normalization
0.312018
Utilizing Character and Word Embeddings for Text Normalization with Sequence-to-Sequence Models · EMNLP 2018
Natural language and speech › Information extraction and text analysis
morphological analysis
0.312017
Don't Throw Those Morphological Analyzers Away Just Yet: Neural Morphological Disambiguation for Arabic · EMNLP 2017
Natural language and speech › Information extraction and text analysis › morphological analysis
morphological disambiguation
0.312017
Don't Throw Those Morphological Analyzers Away Just Yet: Neural Morphological Disambiguation for Arabic · EMNLP 2017
Natural language and speech › Language models and text generation › decoding
adaptive decoding
0.112021
AdaTag: Multi-Attribute Value Extraction from Product Profiles with Adaptive Decoding · ACL/IJCNLP (1) 2021
Information retrieval
search engines
0.112021
End-to-End Conversational Search for Online Shopping with Utterance Transfer · EMNLP (1) 2021

Methods — techniques the papers use, named apart from their topics

visual object detection · 1.0utterance transfer · 1.0transformer sequence-to-sequence model · 1.0optical character recognition · 1.0adaptive decoding · 0.5word-level modeling · 0.4character-level modeling · 0.4multi-task learning · 0.4adversarial training · 0.4character embedding · 0.3
YearPublicationVenuePosition
2025 Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-Training
abstract
Yuchen Zhuang, Jingfeng Yang, Haoming Jiang, Xin Liu, Kewei Cheng, Sanket Lokegaonkar, Yifan Gao, Qing Ping, Tianyi Liu, Binxuan Huang, Zheng Li, Zhengyang Wang, Pei Chen, Ruijie Wang, Rongzhi Zhang, Nasser Zalmout, Priyanka Nigam, Bing Yin, Chao Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yuchen Zhuang, Jingfeng Yang 0001, Haoming Jiang, Xin Liu 0039, Kewei Cheng, Sanket Lokegaonkar, Yifan Gao 0001, Qing Ping, Binxuan Huang, Zheng Li 0018, Ruijie Wang 0004, Rongzhi Zhang, Nasser Zalmout, Priyanka Nigam, Chao Zhang 0014
NAACL (Long Papers)16
2025 DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational Capabilities
abstract
Large Language Models (LLMs) are increasingly employed in multi-turn conversational tasks, yet their pre-training data predominantly consists of continuous prose, creating a potential mismatch between required capabilities and training paradigms. We introduce a novel approach to address this discrepancy by synthesizing conversational data from existing text corpora. We present a pipeline that transforms a cluster of multiple related documents into an extended multi-turn, multi-topic information-seeking dialogue. Applying our pipeline to Wikipedia articles, we curate DocTalk, a multi-turn pre-training dialogue corpus consisting of over 730k long conversations. We hypothesize that exposure to such synthesized conversational structures during pre-training can enhance the fundamental multi-turn capabilities of LLMs, such as context memory and understanding. Empirically, we show that incorporating DocTalk during pre-training results in up to 40% gain in context memory and understanding, without compromising base performance. DocTalk is available at https://huggingface.co/datasets/AmazonScience/DocTalk.
Jing Yang Lee, Hamed Bonab, Nasser Zalmout, Ming Zeng 0001, Sanket Lokegaonkar, Colin Lockard, Binxuan Huang, Ritesh Sarkhel
SIGDIAL3
2021 AdaTag: Multi-Attribute Value Extraction from Product Profiles with Adaptive Decoding
abstract
Jun Yan, Nasser Zalmout, Yan Liang, Christan Grant, Xiang Ren, Xin Luna Dong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jun Yan 0012, Nasser Zalmout, Yan Liang 0004, Christan Grant, Xiang Ren 0001, Xin Dong 0001
ACL/IJCNLP (1)2
2021 End-to-End Conversational Search for Online Shopping with Utterance Transfer
abstract
Liqiang Xiao, Jun Ma, Xin Luna Dong, Pascual Martínez-Gómez, Nasser Zalmout, Wei Chen, Tong Zhao, Hao He, Yaohui Jin. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Liqiang Xiao, Jun Ma 0029, Xin Dong 0001, Pascual Martínez-Gómez, Nasser Zalmout, Tong Zhao 0002, Hao He 0007, Yaohui Jin
EMNLP (1)5
2021 PAM: Understanding Product Images in Cross Product Category Attribute Extraction
abstract
Understanding product attributes plays an important role in improving online shopping experience for customers and serves asan integral part for constructing a product knowledge graph. Most existing methods focus on attribute extraction from text description or utilize visual information from product images such as shape and color. Compared to the inputs considered in prior works, a product image in fact contains more information, represented by a rich mixture of words and visual clues with a layout carefully designed to impress customers. This work proposes a more inclusive framework that fully utilizes these different modalities for attribute extraction.Inspired by recent works in visual question answering, we use a transformer based sequence to sequence model to fuse representations of product text, Optical Character Recognition (OCR) tokens and visual objects detected in the product image. The framework is further extended with the capability to extract attribute value across multiple product categories with a single model, by training the decoder to predict both product category and attribute value and conditioning its output on product category. The model provides a unified attribute extraction solution desirable at an e-commerce platform that offers numerous product categories with a diverse body of product attributes. We evaluated the model on two product attributes, one with many possible values and one with a small set of possible values, over 14 product categories and found the model could achieve 15% gain on the Recall and 10% gain on the F1 score compared to existing methods using text-only features.
Rongmei Lin, Xiang He 0007, Nasser Zalmout, Yan Liang 0004, Li Xiong 0001, Xin Dong 0001
KDD4
2021 All You Need to Know to Build a Product Knowledge Graph
abstract
Knowledge graphs have been pivotal in supporting downstream applications like search, recommendation, and question answering, among others. Therefore, knowledge graphs have naturally become key enabling technologies in e-Commerce platforms. Developing a high coverage product knowledge graph is more challenging than generic knowledge graphs. The highly specific and complex domain, the sparsity of training data, along with the dynamic taxonomies and product types, can constrain the resulting knowledge graphs. In this tutorial we present best practices and ML innovations in industry towards building a scalable product knowledge graph. Contributions in this domain benefit from the general literature in areas including information extraction and data mining, tailored to address the specific characteristics of e-Commerce platforms.
Nasser Zalmout, Yan Liang 0004, Xin Dong 0001
KDD1
2020 Joint Diacritization, Lemmatization, Normalization, and Fine-Grained Morphological Tagging
abstract
The written forms of Semitic languages are both highly ambiguous and morphologically rich: a word can have multiple interpretations and is one of many inflected forms of the same concept or lemma.This is further exacerbated for dialectal content, which is more prone to noise and lacks a standard orthography.The morphological features can be lexicalized, like lemmas and diacritized forms, or non-lexicalized, like gender, number, and partof-speech tags, among others.Joint modeling of the lexicalized and non-lexicalized features can identify more intricate morphological patterns, which provide better context modeling, and further disambiguate ambiguous lexical choices.However, the different modeling granularity can make joint modeling more difficult.Our approach models the different features jointly, whether lexicalized (on the characterlevel), or non-lexicalized (on the word-level).We use Arabic as a test case, and achieve stateof-the-art results for Modern Standard Arabic with 20% relative error reduction, and Egyptian Arabic with 11% relative error reduction.
Nasser Zalmout, Nizar Habash
ACL1
2020 Utilizing Subword Entities in Character-Level Sequence-to-Sequence Lemmatization Models
abstract
In this paper we present a character-level sequence-to-sequence lemmatization model, utilizing several subword features in multiple configurations.In addition to generic n-gram embeddings (using FastText), we experiment with concatenative (stems) and templatic (roots and patterns) morphological subwords.We present several architectures that embed these features directly at the encoder side, or learn them jointly at the decoder side with a multitask learning architecture.The results indicate that using the generic n-gram embeddings (through FastText) outperform the other linguistically-driven subwords.We use Modern Standard Arabic and Egyptian Arabic as test cases, with up to 22% and 13% relative error reduction, respectively, from a strong baseline.An error analysis shows that our best system is even able to handle word/lemma pairs that are both unseen in the training data.
Nasser Zalmout, Nizar Habash
COLING1
2020 Morphological Analysis and Disambiguation for Gulf Arabic: The Interplay between Resources and Methods
abstract
In this paper we present the first full morphological analysis and disambiguation system for Gulf Arabic. We use an existing state-of-the-art morphological disambiguation system to investigate the effects of different data sizes and different combinations of morphological analyzers for Modern Standard Arabic, Egyptian Arabic, and Gulf Arabic. We find that in very low settings, morphological analyzers help boost the performance of the full morphological disambiguation task. However, as the size of resources increase, the value of the morphological analyzers decreases.
Salam Khalifa, Nasser Zalmout, Nizar Habash
LREC2
2020 CAMeL Tools: An Open Source Python Toolkit for Arabic Natural Language Processing
abstract
We present CAMeL Tools, a collection of open-source tools for Arabic natural language processing in Python. CAMeL Tools currently provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and Sentiment Analysis. In this paper, we describe the design of CAMeL Tools and the functionalities it provides.
Ossama Obeid, Nasser Zalmout, Salam Khalifa, Dima Taji, Mai Oudah, Bashar Alhafni, Go Inoue, Fadhl Eryani, Alexander Erdmann, Nizar Habash
LREC2
2019 Adversarial Multitask Learning for Joint Multi-Feature and Multi-Dialect Morphological Modeling
abstract
Morphological tagging is challenging for morphologically rich languages due to the large target space and the need for more training data to minimize model sparsity.Dialectal variants of morphologically rich languages suffer more as they tend to be more noisy and have less resources.In this paper we explore the use of multitask learning and adversarial training to address morphological richness and dialectal variations in the context of full morphological tagging.We use multitask learning for joint morphological modeling for the features within two dialects, and as a knowledge-transfer scheme for crossdialectal modeling.We use adversarial training to learn dialect invariant features that can help the knowledge-transfer scheme from the high to low-resource variants.We work with two dialectal variants: Modern Standard Arabic (high-resource "dialect" 1 ) and Egyptian Arabic (low-resource dialect) as a case study.Our models achieve state-of-the-art results for both.Furthermore, adversarial training provides more significant improvement when using smaller training datasets in particular.
Nasser Zalmout, Nizar Habash
ACL (1)1
2018 Utilizing Character and Word Embeddings for Text Normalization with Sequence-to-Sequence Models
abstract
Text normalization is an important enabling technology for several NLP tasks.Recently, neural-network-based approaches have outperformed well-established models in this task.However, in languages other than English, there has been little exploration in this direction.Both the scarcity of annotated data and the complexity of the language increase the difficulty of the problem.To address these challenges, we use a sequence-to-sequence model with character-based attention, which in addition to its self-learned character embeddings, uses word embeddings pre-trained with an approach that also models subword information.This provides the neural model with access to more linguistic information especially suitable for text normalization, without large parallel corpora.We show that providing the model with word-level features bridges the gap for the neural network approach to achieve a state-of-the-art F 1 score on a standard Arabic language correction shared task dataset.
Daniel Watson, Nasser Zalmout, Nizar Habash
EMNLP2
2018 Unified Guidelines and Resources for Arabic Dialect Orthography
Nizar Habash, Fadhl Eryani, Salam Khalifa, Owen Rambow, Dana Abdulrahim, Alexander Erdmann, Reem Faraj, Wajdi Zaghouani, Houda Bouamor, Nasser Zalmout, Sara Hassan, Faisal Al-Shargi, Sakhar B. Alkhereyf, Basma Abdulkareem, Ramy Eskander, Mohammad Salameh, Hind Saddiki
LREC10
2018 Noise-Robust Morphological Disambiguation for Dialectal Arabic
abstract
Nasser Zalmout, Alexander Erdmann, Nizar Habash. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Nasser Zalmout, Alexander Erdmann, Nizar Habash
NAACL-HLT1
2017 Don't Throw Those Morphological Analyzers Away Just Yet: Neural Morphological Disambiguation for Arabic
abstract
This paper presents a model for Arabic morphological disambiguation based on Recurrent Neural Networks (RNN).We train Long Short-Term Memory (LSTM) cells in several configurations and embedding levels to model the various morphological features.Our experiments show that these models outperform state-of-theart systems without explicit use of feature engineering.However, adding learning features from a morphological analyzer to model the space of possible analyses provides additional improvement.We make use of the resulting morphological models for scoring and ranking the analyses of the morphological analyzer for morphological disambiguation.The results show significant gains in accuracy across several evaluation metrics.Our system results in 4.4% absolute increase over the state-of-the-art in full morphological analysis accuracy (30.6% relative error reduction), and 10.6% (31.5% relative error reduction) for out-of-vocabulary words.
Nasser Zalmout, Nizar Habash
EMNLP1