VLDB 2026 Research / reviewers in the wild / expert
Nasser Zalmout
dblp:142/9396
· DBLP profile ↗
15ranked-venue papers
6as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Information extraction and text analysis · 59% Question answering and dialogue systems · 11% Knowledge representation and reasoning · 11% | |
| Databases, data mining, and information retrieval
3 papers |
Knowledge graphs · 56% Data mining · 36% Information retrieval · 8% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › morphological analysis
morphological tagging |
0.8 | 2 | 2020 | Joint Diacritization, Lemmatization, Normalization, and Fine-Grained Morphological Tagging · ACL 2020 Adversarial Multitask Learning for Joint Multi-Feature and Multi-Dialect Morphological Modeling · ACL (1) 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction
attribute extraction |
0.5 | 1 | 2021 | PAM: Understanding Product Images in Cross Product Category Attribute Extraction · KDD 2021 |
Natural language and speech › Information extraction and text analysis › relation extraction
attribute value extraction |
0.5 | 1 | 2021 | AdaTag: Multi-Attribute Value Extraction from Product Profiles with Adaptive Decoding · ACL/IJCNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems
conversational search |
0.5 | 1 | 2021 | End-to-End Conversational Search for Online Shopping with Utterance Transfer · EMNLP (1) 2021 |
Data mining › text mining › information extraction
attribute extraction |
0.5 | 1 | 2021 | PAM: Understanding Product Images in Cross Product Category Attribute Extraction · KDD 2021 |
Knowledge graphs
knowledge graph construction |
0.5 | 1 | 2021 | PAM: Understanding Product Images in Cross Product Category Attribute Extraction · KDD 2021 |
Knowledge graphs › domain-specific knowledge graph
product knowledge graph |
0.5 | 1 | 2021 | All You Need to Know to Build a Product Knowledge Graph · KDD 2021 |
Natural language and speech › Information extraction and text analysis › morphological analysis
lemmatization |
0.4 | 1 | 2020 | Joint Diacritization, Lemmatization, Normalization, and Fine-Grained Morphological Tagging · ACL 2020 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.4 | 1 | 2019 | Adversarial Multitask Learning for Joint Multi-Feature and Multi-Dialect Morphological Modeling · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
text normalization |
0.3 | 1 | 2018 | Utilizing Character and Word Embeddings for Text Normalization with Sequence-to-Sequence Models · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis
morphological analysis |
0.3 | 1 | 2017 | Don't Throw Those Morphological Analyzers Away Just Yet: Neural Morphological Disambiguation for Arabic · EMNLP 2017 |
Natural language and speech › Information extraction and text analysis › morphological analysis
morphological disambiguation |
0.3 | 1 | 2017 | Don't Throw Those Morphological Analyzers Away Just Yet: Neural Morphological Disambiguation for Arabic · EMNLP 2017 |
Natural language and speech › Language models and text generation › decoding
adaptive decoding |
0.1 | 1 | 2021 | AdaTag: Multi-Attribute Value Extraction from Product Profiles with Adaptive Decoding · ACL/IJCNLP (1) 2021 |
Information retrieval
search engines |
0.1 | 1 | 2021 | End-to-End Conversational Search for Online Shopping with Utterance Transfer · EMNLP (1) 2021 |
Methods — techniques the papers use, named apart from their topics
visual object detection · 1.0utterance transfer · 1.0transformer sequence-to-sequence model · 1.0optical character recognition · 1.0adaptive decoding · 0.5word-level modeling · 0.4character-level modeling · 0.4multi-task learning · 0.4adversarial training · 0.4character embedding · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-TrainingabstractYuchen Zhuang, Jingfeng Yang, Haoming Jiang, Xin Liu, Kewei Cheng, Sanket Lokegaonkar, Yifan Gao, Qing Ping, Tianyi Liu, Binxuan Huang, Zheng Li, Zhengyang Wang, Pei Chen, Ruijie Wang, Rongzhi Zhang, Nasser Zalmout, Priyanka Nigam, Bing Yin, Chao Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yuchen Zhuang, Jingfeng Yang 0001, Haoming Jiang, Xin Liu 0039, Kewei Cheng, Sanket Lokegaonkar, Yifan Gao 0001, Qing Ping, Binxuan Huang, Zheng Li 0018, Ruijie Wang 0004, Rongzhi Zhang, Nasser Zalmout, Priyanka Nigam, Chao Zhang 0014 |
NAACL (Long Papers) | 16 |
| 2025 | DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational CapabilitiesabstractLarge Language Models (LLMs) are increasingly employed in multi-turn conversational tasks, yet their pre-training data predominantly consists of continuous prose, creating a potential mismatch between required capabilities and training paradigms. We introduce a novel approach to address this discrepancy by synthesizing conversational data from existing text corpora. We present a pipeline that transforms a cluster of multiple related documents into an extended multi-turn, multi-topic information-seeking dialogue. Applying our pipeline to Wikipedia articles, we curate DocTalk, a multi-turn pre-training dialogue corpus consisting of over 730k long conversations. We hypothesize that exposure to such synthesized conversational structures during pre-training can enhance the fundamental multi-turn capabilities of LLMs, such as context memory and understanding. Empirically, we show that incorporating DocTalk during pre-training results in up to 40% gain in context memory and understanding, without compromising base performance. DocTalk is available at https://huggingface.co/datasets/AmazonScience/DocTalk. Jing Yang Lee, Hamed Bonab, Nasser Zalmout, Ming Zeng 0001, Sanket Lokegaonkar, Colin Lockard, Binxuan Huang, Ritesh Sarkhel |
SIGDIAL | 3 |
| 2021 | AdaTag: Multi-Attribute Value Extraction from Product Profiles with Adaptive DecodingabstractJun Yan, Nasser Zalmout, Yan Liang, Christan Grant, Xiang Ren, Xin Luna Dong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jun Yan 0012, Nasser Zalmout, Yan Liang 0004, Christan Grant, Xiang Ren 0001, Xin Dong 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | End-to-End Conversational Search for Online Shopping with Utterance TransferabstractLiqiang Xiao, Jun Ma, Xin Luna Dong, Pascual Martínez-Gómez, Nasser Zalmout, Wei Chen, Tong Zhao, Hao He, Yaohui Jin. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Liqiang Xiao, Jun Ma 0029, Xin Dong 0001, Pascual Martínez-Gómez, Nasser Zalmout, Tong Zhao 0002, Hao He 0007, Yaohui Jin |
EMNLP (1) | 5 |
| 2021 | PAM: Understanding Product Images in Cross Product Category Attribute ExtractionabstractUnderstanding product attributes plays an important role in improving online shopping experience for customers and serves asan integral part for constructing a product knowledge graph. Most existing methods focus on attribute extraction from text description or utilize visual information from product images such as shape and color. Compared to the inputs considered in prior works, a product image in fact contains more information, represented by a rich mixture of words and visual clues with a layout carefully designed to impress customers. This work proposes a more inclusive framework that fully utilizes these different modalities for attribute extraction.Inspired by recent works in visual question answering, we use a transformer based sequence to sequence model to fuse representations of product text, Optical Character Recognition (OCR) tokens and visual objects detected in the product image. The framework is further extended with the capability to extract attribute value across multiple product categories with a single model, by training the decoder to predict both product category and attribute value and conditioning its output on product category. The model provides a unified attribute extraction solution desirable at an e-commerce platform that offers numerous product categories with a diverse body of product attributes. We evaluated the model on two product attributes, one with many possible values and one with a small set of possible values, over 14 product categories and found the model could achieve 15% gain on the Recall and 10% gain on the F1 score compared to existing methods using text-only features. Rongmei Lin, Xiang He 0007, Nasser Zalmout, Yan Liang 0004, Li Xiong 0001, Xin Dong 0001 |
KDD | 4 |
| 2021 | All You Need to Know to Build a Product Knowledge GraphabstractKnowledge graphs have been pivotal in supporting downstream applications like search, recommendation, and question answering, among others. Therefore, knowledge graphs have naturally become key enabling technologies in e-Commerce platforms. Developing a high coverage product knowledge graph is more challenging than generic knowledge graphs. The highly specific and complex domain, the sparsity of training data, along with the dynamic taxonomies and product types, can constrain the resulting knowledge graphs. In this tutorial we present best practices and ML innovations in industry towards building a scalable product knowledge graph. Contributions in this domain benefit from the general literature in areas including information extraction and data mining, tailored to address the specific characteristics of e-Commerce platforms. Nasser Zalmout, Yan Liang 0004, Xin Dong 0001 |
KDD | 1 |
| 2020 | Joint Diacritization, Lemmatization, Normalization, and Fine-Grained Morphological TaggingabstractThe written forms of Semitic languages are both highly ambiguous and morphologically rich: a word can have multiple interpretations and is one of many inflected forms of the same concept or lemma.This is further exacerbated for dialectal content, which is more prone to noise and lacks a standard orthography.The morphological features can be lexicalized, like lemmas and diacritized forms, or non-lexicalized, like gender, number, and partof-speech tags, among others.Joint modeling of the lexicalized and non-lexicalized features can identify more intricate morphological patterns, which provide better context modeling, and further disambiguate ambiguous lexical choices.However, the different modeling granularity can make joint modeling more difficult.Our approach models the different features jointly, whether lexicalized (on the characterlevel), or non-lexicalized (on the word-level).We use Arabic as a test case, and achieve stateof-the-art results for Modern Standard Arabic with 20% relative error reduction, and Egyptian Arabic with 11% relative error reduction. Nasser Zalmout, Nizar Habash |
ACL | 1 |
| 2020 | Utilizing Subword Entities in Character-Level Sequence-to-Sequence Lemmatization ModelsabstractIn this paper we present a character-level sequence-to-sequence lemmatization model, utilizing several subword features in multiple configurations.In addition to generic n-gram embeddings (using FastText), we experiment with concatenative (stems) and templatic (roots and patterns) morphological subwords.We present several architectures that embed these features directly at the encoder side, or learn them jointly at the decoder side with a multitask learning architecture.The results indicate that using the generic n-gram embeddings (through FastText) outperform the other linguistically-driven subwords.We use Modern Standard Arabic and Egyptian Arabic as test cases, with up to 22% and 13% relative error reduction, respectively, from a strong baseline.An error analysis shows that our best system is even able to handle word/lemma pairs that are both unseen in the training data. Nasser Zalmout, Nizar Habash |
COLING | 1 |
| 2020 | Morphological Analysis and Disambiguation for Gulf Arabic: The Interplay between Resources and MethodsabstractIn this paper we present the first full morphological analysis and disambiguation system for Gulf Arabic. We use an existing state-of-the-art morphological disambiguation system to investigate the effects of different data sizes and different combinations of morphological analyzers for Modern Standard Arabic, Egyptian Arabic, and Gulf Arabic. We find that in very low settings, morphological analyzers help boost the performance of the full morphological disambiguation task. However, as the size of resources increase, the value of the morphological analyzers decreases. Salam Khalifa, Nasser Zalmout, Nizar Habash |
LREC | 2 |
| 2020 | CAMeL Tools: An Open Source Python Toolkit for Arabic Natural Language ProcessingabstractWe present CAMeL Tools, a collection of open-source tools for Arabic natural language processing in Python. CAMeL Tools currently provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and Sentiment Analysis. In this paper, we describe the design of CAMeL Tools and the functionalities it provides. Ossama Obeid, Nasser Zalmout, Salam Khalifa, Dima Taji, Mai Oudah, Bashar Alhafni, Go Inoue, Fadhl Eryani, Alexander Erdmann, Nizar Habash |
LREC | 2 |
| 2019 | Adversarial Multitask Learning for Joint Multi-Feature and Multi-Dialect Morphological ModelingabstractMorphological tagging is challenging for morphologically rich languages due to the large target space and the need for more training data to minimize model sparsity.Dialectal variants of morphologically rich languages suffer more as they tend to be more noisy and have less resources.In this paper we explore the use of multitask learning and adversarial training to address morphological richness and dialectal variations in the context of full morphological tagging.We use multitask learning for joint morphological modeling for the features within two dialects, and as a knowledge-transfer scheme for crossdialectal modeling.We use adversarial training to learn dialect invariant features that can help the knowledge-transfer scheme from the high to low-resource variants.We work with two dialectal variants: Modern Standard Arabic (high-resource "dialect" 1 ) and Egyptian Arabic (low-resource dialect) as a case study.Our models achieve state-of-the-art results for both.Furthermore, adversarial training provides more significant improvement when using smaller training datasets in particular. Nasser Zalmout, Nizar Habash |
ACL (1) | 1 |
| 2018 | Utilizing Character and Word Embeddings for Text Normalization with Sequence-to-Sequence ModelsabstractText normalization is an important enabling technology for several NLP tasks.Recently, neural-network-based approaches have outperformed well-established models in this task.However, in languages other than English, there has been little exploration in this direction.Both the scarcity of annotated data and the complexity of the language increase the difficulty of the problem.To address these challenges, we use a sequence-to-sequence model with character-based attention, which in addition to its self-learned character embeddings, uses word embeddings pre-trained with an approach that also models subword information.This provides the neural model with access to more linguistic information especially suitable for text normalization, without large parallel corpora.We show that providing the model with word-level features bridges the gap for the neural network approach to achieve a state-of-the-art F 1 score on a standard Arabic language correction shared task dataset. Daniel Watson, Nasser Zalmout, Nizar Habash |
EMNLP | 2 |
| 2018 | Unified Guidelines and Resources for Arabic Dialect Orthography
Nizar Habash, Fadhl Eryani, Salam Khalifa, Owen Rambow, Dana Abdulrahim, Alexander Erdmann, Reem Faraj, Wajdi Zaghouani, Houda Bouamor, Nasser Zalmout, Sara Hassan, Faisal Al-Shargi, Sakhar B. Alkhereyf, Basma Abdulkareem, Ramy Eskander, Mohammad Salameh, Hind Saddiki |
LREC | 10 |
| 2018 | Noise-Robust Morphological Disambiguation for Dialectal ArabicabstractNasser Zalmout, Alexander Erdmann, Nizar Habash. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Nasser Zalmout, Alexander Erdmann, Nizar Habash |
NAACL-HLT | 1 |
| 2017 | Don't Throw Those Morphological Analyzers Away Just Yet: Neural Morphological Disambiguation for ArabicabstractThis paper presents a model for Arabic morphological disambiguation based on Recurrent Neural Networks (RNN).We train Long Short-Term Memory (LSTM) cells in several configurations and embedding levels to model the various morphological features.Our experiments show that these models outperform state-of-theart systems without explicit use of feature engineering.However, adding learning features from a morphological analyzer to model the space of possible analyses provides additional improvement.We make use of the resulting morphological models for scoring and ranking the analyses of the morphological analyzer for morphological disambiguation.The results show significant gains in accuracy across several evaluation metrics.Our system results in 4.4% absolute increase over the state-of-the-art in full morphological analysis accuracy (30.6% relative error reduction), and 10.6% (31.5% relative error reduction) for out-of-vocabulary words. Nasser Zalmout, Nizar Habash |
EMNLP | 1 |