Zornitsa Kozareva

dblp:63/5321 · DBLP profile ↗
← Back
38ranked-venue papers
16as first author
9since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 16 first-author · 9 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 Methods for Measuring, Updating, and Visualizing Factual Beliefs in Language Models
abstract
Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, Srinivasan Iyer. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Peter Hase, Mona T. Diab, Asli Celikyilmaz, Xian Li 0003, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, Srinivasan Iyer 0001
EACL5
2022 Fixed Support Tree-Sliced Wasserstein Barycenter
abstract
The Wasserstein barycenter has been widely studied in various fields, including natural language processing, and computer vision. However, it requires a high computational cost to solve the Wasserstein barycenter problem because the computation of the Wasserstein distance requires a quadratic time with respect to the number of supports. By contrast, the Wasserstein distance on a tree, called the tree-Wasserstein distance, can be computed in linear time and allows for the fast comparison of a large number of distributions. In this study, we propose a barycenter under the tree-Wasserstein distance, called the fixed support tree-Wasserstein barycenter (FS-TWB) and its extension, called the fixed support tree-sliced Wasserstein barycenter (FS-TSWB). More specifically, we first show that the FS-TWB and FS-TSWB problems are convex optimization problems and can be solved by using the projected subgradient descent. Moreover, we propose a more efficient algorithm to compute the subgradient and objective function value by using the properties of tree-Wasserstein barycenter problems. Through real-world experiments, we show that, by using the proposed algorithm, the FS-TWB and FS-TSWB can be solved two orders of magnitude faster than the original Wasserstein barycenter.
Yuki Takezawa, Ryoma Sato, Zornitsa Kozareva, Sujith Ravi, Makoto Yamada
AISTATS3
2022 ToKen: Task Decomposition and Knowledge Infusion for Few-Shot Hate Speech Detection
abstract
Badr AlKhamissi, Faisal Ladhak, Srinivasan Iyer, Veselin Stoyanov, Zornitsa Kozareva, Xian Li, Pascale Fung, Lambert Mathias, Asli Celikyilmaz, Mona Diab. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Badr AlKhamissi, Faisal Ladhak, Srinivasan Iyer 0001, Veselin Stoyanov, Zornitsa Kozareva, Xian Li 0003, Pascale Fung, Lambert Mathias, Asli Celikyilmaz, Mona T. Diab
EMNLP5
2022 Efficient Large Scale Language Modeling with Mixtures of Experts
abstract
Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, Giridharan Anantharaman, Xian Li, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Xing Zhou, Punit Singh Koura, Brian O’Horo, Jeffrey Wang, Luke Zettlemoyer, Mona Diab, Zornitsa Kozareva, Veselin Stoyanov. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Mikel Artetxe, Shruti Bhosale, Naman Goyal 0001, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer 0001, Ramakanth Pasunuru, Giri Anantharaman, Xian Li 0003, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Punit Singh Koura, Brian O'Horo, Jeffrey Wang, Luke Zettlemoyer, Mona T. Diab, Zornitsa Kozareva, Veselin Stoyanov
EMNLP23
2022 Few-shot Learning with Multilingual Generative Language Models
abstract
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, Xian Li. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal 0001, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona T. Diab, Veselin Stoyanov, Xian Li 0003
EMNLP18
2022 Improving In-Context Few-Shot Learning via Self-Supervised Training
abstract
Mingda Chen, Jingfei Du, Ramakanth Pasunuru, Todor Mihaylov, Srini Iyer, Veselin Stoyanov, Zornitsa Kozareva. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Mingda Chen, Jingfei Du, Ramakanth Pasunuru, Todor Mihaylov, Srinivasan Iyer 0001, Veselin Stoyanov, Zornitsa Kozareva
NAACL-HLT7
2021 ProFormer: Towards On-Device LSH Projection Based Transformers
abstract
At the heart of text based neural models lay word representations, which are powerful but occupy a lot of memory making it challenging to deploy to devices with memory constraints such as mobile phones, watches and IoT.To surmount these challenges, we introduce ProFormer -a projection based transformer architecture that is faster and lighter making it suitable to deploy to memory constraint devices and preserve user privacy.We use LSH projection layer to dynamically generate word representations on-the-fly without embedding lookup tables leading to significant memory footprint reduction from O(V.d) to O(T ), where V is the vocabulary size, d is the embedding dimension size and T is the dimension of the LSH projection representation.We also propose a local projection attention (LPA) layer, which uses self-attention to transform the input sequence of N LSH word projections into a sequence of N/K representations reducing the computations quadratically by O(K 2 ).We evaluate ProFormer on multiple text classification tasks and observed improvements over prior state-of-the-art on-device approaches for short text classification and comparable performance for long text classification tasks.Pro-Former is also competitive with other popular but highly resource-intensive approaches like BERT and even outperforms small-sized BERT variants with significant resource savings -reduces the embedding memory footprint from 92.16 MB to 1.7 KB and requires 16× less computation overhead, which is very impressive making it the fastest and smallest on-device model.
Chinnadhurai Sankar, Sujith Ravi, Zornitsa Kozareva
EACL3
2021 On-Device Text Representations Robust To Misspellings via Projections
abstract
Recently, there has been a strong interest in developing natural language applications that live on personal devices such as mobile phones, watches and IoT with the objective to preserve user privacy and have low memory.Advances in Locality-Sensitive Hashing (LSH)-based projection networks have demonstrated state-of-the-art performance in various classification tasks without explicit word (or word-piece) embedding lookup tables by computing on-the-fly text representations.In this paper, we show that the projection based neural classifiers are inherently robust to misspellings and perturbations of the input text.We empirically demonstrate that the LSH projection based classifiers are more robust to common misspellings compared to BiL-STMs (with both word-piece & word-only tokenization) and fine-tuned BERT based methods.When subject to misspelling attacks, LSH projection based classifiers had a small average accuracy drop of 2.94% across multiple classifications tasks, while the fine-tuned BERT model accuracy had a significant drop of 11.44%.
Chinnadhurai Sankar, Sujith Ravi, Zornitsa Kozareva
EACL3
2021 SoDA: On-device Conversational Slot Extraction
abstract
We propose a novel on-device neural sequence labeling model which uses embedding-free projections and character information to construct compact word representations to learn a sequence model using a combination of bidirectional LSTM with self-attention and CRF.Unlike typical dialog models that rely on huge, complex neural network architectures and large-scale pre-trained Transformers to achieve state-of-the-art results, our method achieves comparable results to BERT and even outperforms its smaller variant DistilBERT on conversational slot extraction tasks.Our method is faster than BERT models while achieving significant model size reduction-our model requires 135x and 81x fewer model parameters than BERT and DistilBERT, respectively.We conduct experiments on multiple conversational datasets and show significant improvements over existing methods including recent on-device models.Experimental results and ablation studies also show that our neural models preserve tiny memory footprint necessary to operate on smart devices, while still maintaining high performance.
Sujith Ravi, Zornitsa Kozareva
SIGDIAL2
2020 Environment-Agnostic Multitask Learning for Natural Language Grounded Navigation
Xin Wang 0061, Vihan Jain, Eugene Ie, William Yang Wang, Zornitsa Kozareva, Sujith Ravi
ECCV (24)5
2019 On-device Structured and Context Partitioned Projection Networks
abstract
A challenging problem in on-device text classification is to build highly accurate neural models that can fit in small memory footprint and have low latency.To address this challenge, we propose an on-device neural network SGNN++ which dynamically learns compact projection vectors from raw text using structured and context-dependent partition projections.We show that this results in accelerated inference and performance improvements.We conduct extensive evaluation on multiple conversational tasks and languages such as English, Japanese, Spanish and French.Our SGNN++ model significantly outperforms all baselines, improves upon existing on-device neural models and even surpasses RNN, CNN and BiLSTM models on dialog act and intent prediction.Through a series of ablation studies we show the impact of the partitioned projections and structured information leading to 10% improvement.We study the impact of the model size on accuracy and introduce quantization-aware training for SGNN++ to further reduce the model size while preserving the same quality.Finally, we show fast inference on mobile phones.
Sujith Ravi, Zornitsa Kozareva
ACL (1)2
2019 ProSeqo: Projection Sequence Networks for On-Device Text Classification
abstract
Zornitsa Kozareva, Sujith Ravi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zornitsa Kozareva, Sujith Ravi
EMNLP/IJCNLP (1)1
2019 PRADO: Projection Attention Networks for Document Classification On-Device
abstract
Prabhu Kaliamoorthi, Sujith Ravi, Zornitsa Kozareva. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Prabhu Kaliamoorthi, Sujith Ravi, Zornitsa Kozareva
EMNLP/IJCNLP (1)3
2018 Variational Reasoning for Question Answering With Knowledge Graph
abstract
Knowledge graph (KG) is known to be helpful for the task of question answering (QA), since it provides well-structured relational information between entities, and allows one to further infer indirect facts. However, it is challenging to build QA systems which can learn to reason over knowledge graphs based on question-answer pairs alone. First, when people ask questions, their expressions are noisy (for example, typos in texts, or variations in pronunciations), which is non-trivial for the QA system to match those mentioned entities to the knowledge graph. Second, many questions require multi-hop logic reasoning over the knowledge graph to retrieve the answers. To address these challenges, we propose a novel and unified deep learning architecture, and an end-to-end variational learning algorithm which can handle noise in questions, and learn multi-hop reasoning simultaneously. Our method achieves state-of-the-art performance on a recent benchmark dataset in the literature. We also derive a series of new benchmark datasets, including questions for multi-hop reasoning, questions paraphrased by neural translation model, and questions in human voice. Our method yields very promising results on all these challenging datasets.
Yuyu Zhang, Hanjun Dai, Zornitsa Kozareva, Alexander J. Smola
AAAI3
2018 Self-Governing Neural Networks for On-Device Short Text Classification
abstract
Deep neural networks reach state-of-the-art performance for wide range of natural language processing, computer vision and speech applications.Yet, one of the biggest challenges is running these complex networks on devices such as mobile phones or smart watches with tiny memory footprint and low computational capacity.We propose on-device Self-Governing Neural Networks (SGNNs), which learn compact projection vectors with local sensitive hashing.The key advantage of SGNNs over existing work is that they surmount the need for pre-trained word embeddings and complex networks with huge parameters.We conduct extensive evaluation on dialog act classification and show significant improvement over state-of-the-art results.Our findings show that SGNNs are effective at capturing low-dimensional semantic text representations, while maintaining high accuracy.
Sujith Ravi, Zornitsa Kozareva
EMNLP2
2018 Self-Governing Neural Networks for On-Device Short Text Classification
abstract
Deep neural networks reach state-of-the-art performance for wide range of natural language processing, computer vision and speech applications.Yet, one of the biggest challenges is running these complex networks on devices such as mobile phones or smart watches with tiny memory footprint and low computational capacity.We propose on-device Self-Governing Neural Networks (SGNNs), which learn compact projection vectors with local sensitive hashing.The key advantage of SGNNs over existing work is that they surmount the need for pre-trained word embeddings and complex networks with huge parameters.We conduct extensive evaluation on dialog act classification and show significant improvement over state-of-the-art results.Our findings show that SGNNs are effective at capturing low-dimensional semantic text representations, while maintaining high accuracy.
Sujith Ravi, Zornitsa Kozareva
EMNLP2
2018 Learning Steady-States of Iterative Algorithms over Graphs
abstract
Many graph analytics problems can be solved via iterative algorithms where the solutions are often characterized by a set of steady-state conditions. Different algorithms respect to different set of fixed point constraints, so instead of using these traditional algorithms, can we learn an algorithm which can obtain the same steady-state solutions automatically from examples, in an effective and scalable way? How to represent the meta learner for such algorithm and how to carry out the learning? In this paper, we propose an embedding representation for iterative algorithms over graphs, and design a learning method which alternates between updating the embeddings and projecting them onto the steady-state constraints. We demonstrate the effectiveness of our framework using a few commonly used graph algorithms, and show that in some cases, the learned algorithm can handle graphs with more than 100,000,000 nodes in a single machine.
Hanjun Dai, Zornitsa Kozareva, Bo Dai 0001, Alexander J. Smola
ICML2
2016 Query to Knowledge: Unsupervised Entity Extraction from Shopping Queries using Adaptor Grammars
abstract
Web search queries provide a surprisingly large amount of information, which can be potentially organized and converted into a knowledgebase. In this paper, we focus on the problem of automatically identifying brand and product entities from a large collection of web queries in online shopping domain. We propose an unsupervised approach based on adaptor grammars that does not require any human annotation efforts nor rely on any external resources. To reduce the noise and normalize the query patterns, we introduce a query standardization step, which groups multiple search patterns and word orderings together into their most frequent ones. We present three different sets of grammar rules used to infer query structures and extract brand and product entities. To give an objective assessment of the performance of our approach, we conduct experiments on a large collection of online shopping queries and intrinsically evaluate the knowledgebase generated by our method qualitatively and quantitatively. In addition, we also evaluate our framework on extrinsic tasks on query tagging and chunking. Our empirical studies show that the knowledgebase discovered by our approach is highly accurate, has good coverage and significantly improves the performance on the external tasks.
Ke Zhai 0001, Zornitsa Kozareva, Yuening Hu, Weiwei Guo
SIGIR2
2015 Everyone Likes Shopping! Multi-class Product Categorization for e-Commerce
abstract
Online shopping caters the needs of millions of users on a daily basis.To build an accurate system that can retrieve relevant products for a query like "MB252 with travel bags" one requires product and query categorization mechanisms, which classify the text as Home&Garden>Kitchen&Dining>Kitchen Appliances>Blenders.One of the biggest challenges in e-Commerce is that providers like Amazon, e-Bay, Google, Yahoo! and Walmart organize products into different product taxonomies making it hard and time-consuming for sellers to categorize goods for each shopping platform.To address this challenge, we propose an automatic product categorization mechanism, which for a given product title assigns the correct product category from a taxonomy.We conducted an empirical evaluation on 445, 408 product titles and used a rich product taxonomy of 319 categories organized into 6 levels.We compared performance against multiple algorithms and found that the best performing system reaches .88f-score.
Zornitsa Kozareva
HLT-NAACL1
2015 Word from the editors
abstract
Graph structures naturally model connections. In natural language processing (NLP) connections are ubiquitous, on anything between small and web scale. We find them between words – as grammatical, collocation or semantic relations – contributing to the overall meaning, and maintaining the cohesive structure of the text and the discourse unity. We find them between concepts in ontologies or other knowledge repositories – since the early ages of artificial intelligence, associative or semantic networks have been proposed and used as knowledge stores, because they naturally capture the language units and relations between them, and allow for a variety of inference and reasoning processes, simulating some of the functionalities of the human mind. We find them between complete texts or web pages, and between entities in a social network, where they model relations at the web scale. Beyond the more often encountered ‘regular’ graphs, hypergraphs have also appeared in our field to model relations between more than two units.
Zornitsa Kozareva, Vivi Nastase, Rada Mihalcea
Nat. Lang. Eng.1
2013 Multilingual Affect Polarity and Valence Prediction in Metaphor-Rich Texts
Zornitsa Kozareva
ACL (1)1
2013 Sentiment Prediction Using Collaborative Filtering
Jihie Kim, Jae-Bong Yoo, Ho Lim, Huida Qiu, Zornitsa Kozareva, Aram Galstyan
ICWSM5
2011 Insights from Network Structure for Text Mining
Zornitsa Kozareva, Eduard H. Hovy
ACL1
2011 Class Label Enhancement via Related Instances
Zornitsa Kozareva, Konstantin Voevodski, Shang-Hua Teng
EMNLP1
2010 Learning Arguments and Supertypes of Semantic Relations Using Recursive Patterns
Zornitsa Kozareva, Eduard H. Hovy
ACL1
2010 A Semi-Supervised Method to Learn and Construct Taxonomies Using the Web
Zornitsa Kozareva, Eduard H. Hovy
EMNLP1
2010 Not All Seeds Are Equal: Measuring the Quality of Text Mining Seeds
Zornitsa Kozareva, Eduard H. Hovy
HLT-NAACL1
2009 Determining the Polarity and Source of Opinions Expressed in Political Debates
Alexandra Balahur, Zornitsa Kozareva, Andrés Montoyo
CICLing2
2009 Toward Completeness in Concept Extraction and Classification
Eduard H. Hovy, Zornitsa Kozareva, Ellen Riloff
EMNLP2
2008 Semantic Class Learning from the Web with Hyponym Pattern Linkage Graphs
Zornitsa Kozareva, Ellen Riloff, Eduard H. Hovy
ACL1
2008 Domain Information for Fine-Grained Person Name Categorization
Zornitsa Kozareva, Sonia Vázquez, Andrés Montoyo
CICLing1
2007 The Usefulness of Conceptual Representation for the Identification of Semantic Variability Expressions
Zornitsa Kozareva, Sonia Vázquez, Andrés Montoyo
CICLing1
2007 Combining data-driven systems for improving Named Entity Recognition
Zornitsa Kozareva, Óscar Ferrández, Andrés Montoyo, Rafael Muñoz 0001, Armando Suárez, Jaime Gómez
Data Knowl. Eng.1
2006 An Unsupervised Language Independent Method of Name Discrimination Using Second Order Co-occurrence Features
Ted Pedersen, Anagha Kulkarni 0001, Roxana Angheluta, Zornitsa Kozareva, Thamar Solorio
CICLing4
2006 Bootstrapping Named Entity Recognition with Automatically Generated Gazetteer Lists
Zornitsa Kozareva
EACL1
2006 The Role and Resolution of Textual Entailment in Natural Language Processing Applications
Zornitsa Kozareva, Andrés Montoyo
NLDB1
2005 Combining Data-Driven Systems for Improving Named Entity Recognition
Zornitsa Kozareva, Óscar Ferrández, Andrés Montoyo, Rafael Muñoz 0001, Armando Suárez
NLDB1
2004 Cluster Analysis and Classification of Named Entities
Joaquim Ferreira da Silva, Zornitsa Kozareva, José Gabriel Pereira Lopes
LREC2