Thamar Solorio

dblp:79/3530 · DBLP profile ↗
← Back
66ranked-venue papers
11as first author
22since 2021 · last 2026
0000-0002-3541-9405ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 61 · 11 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Fast and Accurate Modeling for Cross-Lingual Label Projection
abstract
Information extraction (IE) systems rely on structured data for training, but such annotated data is highly imbalanced across languages, with low-resource languages receiving little attention.Label projection techniques aim to bridge this gap by transferring structured annotations from high-resource to low-resource languages.However, existing methods are either inaccurate or too slow for large-scale use.This work aims to address this problem by developing a more effective method that remains sufficiently efficient for large-scale projection.In particular, we propose to synthesize alignment sequence pairs and fine-tune an encoder model with span alignment objective, while controlling data influence during training.Experimental results across 50+ languages show that our framework consistently outperforms previous state-of-the-art methods while maintaining fast inference speed.In addition, we introduce EXP -the first benchmark for explicit evaluation of label projection, thereby reducing confounders and non-determinism in method assessment.
Thang Le, Huy Huu Nguyen, Anh Tuan Luu, Thamar Solorio, Thien Huu Nguyen
ACL (1)4
2026 Afri-MCQA: Multimodal Cultural Question Answering for African Languages
abstract
Atnafu Lambebo Tonja, Srija Anand, Emilio Villa-Cueva, Israel Abebe Azime, Jesujoba Oluwadara Alabi, Muhidin A. Mohamed, Debela Desalegn Yadeta, Negasi Haile Abadi, Abigail Oppong, Nnaemeka Casmir Obiefuna, Idris Abdulmumin, Naome A Etori, Eric Peter Wairagala, Kanda Patrick Tshinu, Imanigirimbabazi Emmanuel, Gabofetswe Malema, Alham Fikri Aji, David Ifeoluwa Adelani, Thamar Solorio. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Atnafu Lambebo Tonja, Srija Anand, Emilio Villa-Cueva, Israel Abebe Azime, Jesujoba O. Alabi, Muhidin Mohamed, Debela Desalegn Yadeta, Negasi Haile Abadi, Abigail Oppong, Nnaemeka C. Obiefuna, Idris Abdulmumin, Naome A. Etori, Eric Peter Wairagala, Kanda Patrick Tshinu, Imanigirimbabazi Emmanuel, Gabofetswe Malema, Alham Fikri Aji, David Ifeoluwa Adelani, Thamar Solorio
ACL (1)19
2026 Benchmarking Arabic Authorship Attribution and Style Transfer with Large Language Models
Injy Hamed, Bashar Alhafni, Nizar Habash, Thamar Solorio
LREC4
2025 Why AI Is WEIRD and Shouldn't Be This Way: Towards AI for Everyone, with Everyone, by Everyone
abstract
This paper presents a vision for creating AI systems that are inclusive at every stage of development, from data collection to model design and evaluation. We address key limitations in the current AI pipeline and its WEIRD* representation, such as lack of data diversity, biases in model performance, and narrow evaluation metrics. We also focus on the need for diverse representation among the developers of these systems, as well as incentives that are not skewed toward certain groups. We highlight opportunities to develop AI systems that are for everyone (with diverse stakeholders in mind), with everyone (inclusive of diverse data and annotators), and by everyone (designed and developed by a globally diverse workforce). *WEIRD = an acronym coined by Joseph Henrich to highlight the coverage limitations of many psychological studies, referring to populations that are Western, Educated, Industrialized, Rich, and Democratic; while we do not fully adopt this term for AI, as its current scope does not perfectly align with the WEIRD dimensions, we believe that today's AI has a similarly "weird" coverage, particularly in terms of who is involved in its development and who benefits from it.
Rada Mihalcea, Oana Ignat, Longju Bai, Angana Borah, Luis Chiruzzo, Zhijing Jin 0001, Claude Kwizera, Joan Nwatu, Soujanya Poria, Thamar Solorio
AAAI10
2025 A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions
abstract
Language in the Arab world presents a complex diglossic and multilingual setting, involving the use of Modern Standard Arabic, various dialects and sub-dialects, as well as multiple European languages. This diverse linguistic landscape has given rise to code-switching, both within Arabic varieties and between Arabic and foreign languages. The widespread occurrence of code-switching across the region makes it vital to address these linguistic needs when developing language technologies. In this paper, we provide a review of the current literature in the field of code-switched Arabic NLP, offering a broad perspective on ongoing efforts, challenges, research gaps, and recommendations for future research directions.
Injy Hamed, Caroline Sabty, Slim Abdennadher, Ngoc Thang Vu, Thamar Solorio, Nizar Habash
COLING5
2025 All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages
abstract
Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support low-resource languages, all while effectively integrating corresponding visual cues. In pursuit of culturally diverse global multimodal models, our proposed All Languages Matter Benchmark (ALM-bench) represents the largest and most comprehensive effort to date for evaluating LMMs across 100 languages. ALM-bench challenges existing models by testing their ability to understand and reason about culturally diverse images paired with text in various languages, including many low-resource languages traditionally underrepresented in LMM research. The benchmark offers a robust and nuanced evaluation framework featuring various question formats, including true/false, multiple choice, and open-ended questions, which are further divided into short and long-answer categories. ALM-bench design ensures a comprehensive assessment of a model’s ability to handle varied levels of difficulty in visual and linguistic reasoning. To capture the rich tapestry of global cultures, ALM-bench carefully curates content from 13 distinct cultural aspects, ranging from traditions and rituals to famous personalities and celebrations. Through this, ALM-bench not only provides a rigorous testing ground for state-of-the-art open and closed-source LMMs but also highlights the importance of cultural and linguistic inclusivity, encouraging the development of models that can serve diverse global populations effectively. Our benchmark is publicly available at https://mbzuai-oryx.github.io/ALM-Bench/.
Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana, Noor Ahsan, Nevasini Sasikumar, Omkar Thawakar, Henok Biadglign Ademtew, Yahya Hmaiti, Amandeep Kumar, Kartik Kuckreja, Mykola Maslych, Wafa Al Ghallabi, Mihail Minkov Mihaylov, Abdelrahman M. Shaker, Mike Zhang, Mahardika Krisna Ihsani, Amiel Esplana, Monil Gokani, Shachar Mirkin, Harsh Singh, Ashay Srivastava, Endre Hamerlik, Fathinah Asma Izzati, Fadillah A. Maani, Sebastian Cavada, Jenny Chim, Rohit Gupta 0012, Sanjay Manjunath, Kamila Zhumakhanova, Feno Heriniaina Rabevohitra, Azril Hafizi Amirudin, Muhammad Ridzuan, Daniya Najiha Abdul Kareem, Ketan More, Pramesh Shakya, Amirpouya Ghasemaghaei, Amirbek Djanibekov, Dilshod Azizov, Branislava Jankovic, Naman Bhatia, Alvaro Cabrera, Johan S. Obando-Ceron, Olympiah Otieno, Fabian Farestam, Muztoba Rabbani, Sanoojan Baliah, Santosh Sanjeev, Abduragim Shtanchaev, Maheen Fatima, Amrin Kareem, Toluwani Aremu, Nathan A. Z. Xavier, Amit Bhatkal, Hawau Olamide Toyin, Aman Chadha, Hisham Cholakkal, Rao Muhammad Anwer, Michael Felsberg, Jorma Laaksonen, Thamar Solorio, Monojit Choudhury, Ivan Laptev, Mubarak Shah, Salman Khan 0001, Fahad Shahbaz Khan
CVPR64
2025 CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
Brian Yan, Injy Hamed, Shuichiro Shimizu, Vasista Sai Lodagala, Olga Iakovenko, Bashar Talafha, Amir Hussein, Alexander Polok, Kalvin Chang, Dominik Klement, Sara Althubaiti, Puyuan Peng, Matthew Wiesner, Thamar Solorio, Ahmed Ali 0002, Sanjeev Khudanpur, Shinji Watanabe 0001
INTERSPEECH15
2024 Labeling Comic Mischief Content in Online Videos with a Multimodal Hierarchical-Cross-Attention Model
abstract
We address the challenge of detecting questionable content in online media, specifically the subcategory of comic mischief. This type of content combines elements such as violence, adult content, or sarcasm with humor, making it difficult to detect. Employing a multimodal approach is vital to capture the subtle details inherent in comic mischief content. To tackle this problem, we propose a novel end-to-end multimodal system for the task of comic mischief detection. As part of this contribution, we release a novel dataset for the targeted task consisting of three modalities: video, text (video captions and subtitles), and audio. We also design a HIerarchical Cross-attention model with CAPtions (HICCAP) to capture the intricate relationships among these modalities. The results show that the proposed approach makes a significant improvement over robust baselines and state-of-the-art models for comic mischief detection and its type classification. This emphasizes the potential of our system to empower users, to make informed decisions about the online content they choose to see.
Elaheh Baharlouei, Mahsa Shafaei, Yigeng Zhang, Hugo Jair Escalante, Thamar Solorio
LREC/COLING5
2024 OATS: A Challenge Dataset for Opinion Aspect Target Sentiment Joint Detection for Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) delves into understanding sentiments specific to distinct elements within a user-generated review. It aims to analyze user-generated reviews to determine a) the target entity being reviewed, b) the high-level aspect to which it belongs, c) the sentiment words used to express the opinion, and d) the sentiment expressed toward the targets and the aspects. While various benchmark datasets have fostered advancements in ABSA, they often come with domain limitations and data granularity challenges. Addressing these, we introduce the OATS dataset, which encompasses three fresh domains and consists of 27,470 sentence-level quadruples and 17,092 review-level tuples. Our initiative seeks to bridge specific observed gaps in existing datasets: the recurrent focus on familiar domains like restaurants and laptops, limited data for intricate quadruple extraction tasks, and an occasional oversight of the synergy between sentence and review-level sentiments. Moreover, to elucidate OATS’s potential and shed light on various ABSA subtasks that OATS can solve, we conducted experiments, establishing initial baselines. We hope the OATS dataset augments current resources, paving the way for an encompassing exploration of ABSA (https://github.com/RiTUAL-UH/OATS-ABSA).
Siva Uday Sampreeth Chebolu, Franck Dernoncourt, Nedim Lipka, Thamar Solorio
LREC/COLING4
2024 Interpreting Themes from Educational Stories
abstract
Reading comprehension continues to be a crucial research focus in the NLP community. Recent advances in Machine Reading Comprehension (MRC) have mostly centered on literal comprehension, referring to the surface-level understanding of content. In this work, we focus on the next level - interpretive comprehension, with a particular emphasis on inferring the themes of a narrative text. We introduce the first dataset specifically designed for interpretive comprehension of educational narratives, providing corresponding well-edited theme texts. The dataset spans a variety of genres and cultural origins and includes human-annotated theme keywords with varying levels of granularity. We further formulate NLP tasks under different abstractions of interpretive comprehension toward the main idea of a story. After conducting extensive experiments with state-of-the-art methods, we found the task to be both challenging and significant for NLP research. The dataset and source code have been made publicly available to the research community at https://github.com/RiTUAL-UH/EduStory.
Yigeng Zhang, Fabio A. González 0001, Thamar Solorio
LREC/COLING3
2024 Positive and Risky Message Assessment for Music Products
abstract
In this work, we introduce a pioneering research challenge: evaluating positive and potentially harmful messages within music products. We initiate by setting a multi-faceted, multi-task benchmark for music content assessment. Subsequently, we introduce an efficient multi-task predictive model fortified with ordinality-enforcement to address this challenge. Our findings reveal that the proposed method not only significantly outperforms robust task-specific alternatives but also possesses the capability to assess multiple aspects simultaneously. Furthermore, through detailed case studies, where we employed Large Language Models (LLMs) as surrogates for content assessment, we provide valuable insights to inform and guide future research on this topic. The code for dataset creation and model implementation is publicly available at https://github.com/RiTUAL-UH/music-message-assessment.
Yigeng Zhang, Mahsa Shafaei, Fabio Gonzalez, Thamar Solorio
LREC/COLING4
2024 The Zeno's Paradox of 'Low-Resource' Languages
abstract
The disparity in the languages commonly studied in Natural Language Processing (NLP) is typically reflected by referring to languages as low vs high-resourced.However, there is limited consensus on what exactly qualifies as a 'low-resource language.'To understand how NLP papers define and study 'low resource' languages, we qualitatively analyzed 150 papers from the ACL Anthology and popular speechprocessing conferences that mention the keyword 'low-resource.' Based on our analysis, we show how several interacting axes contribute to 'low-resourcedness' of a language and why that makes it difficult to track progress for each individual language.We hope our work (1) elicits explicit definitions of the terminology when it is used in papers and (2) provides grounding for the different axes to consider when connoting a language as low-resource.
Hellina Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio, Monojit Choudhury
EMNLP4
2024 NLP Progress in Indigenous Latin American Languages
abstract
Atnafu Tonja, Fazlourrahman Balouchzahi, Sabur Butt, Olga Kolesnikova, Hector Ceballos, Alexander Gelbukh, Thamar Solorio. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Atnafu Lambebo Tonja, Fazlourrahman Balouchzahi, Sabur Butt, Olga Kolesnikova, Hector G. Ceballos, Alexander F. Gelbukh, Thamar Solorio
NAACL-HLT7
2024 Adaptive Cross-lingual Text Classification through In-Context One-Shot Demonstrations
abstract
Emilio Cueva, Adrian Lopez Monroy, Fernando Sánchez-Vega, Thamar Solorio. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Emilio Villa-Cueva, Adrián Pastor López-Monroy, Fernando Sánchez-Vega, Thamar Solorio
NAACL-HLT4
2024 CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
abstract
Visual Question Answering~(VQA) is an important task in multimodal AI, which requires models to understand and reason on knowledge present in visual and textual data. However, most of the current VQA datasets and models are primarily focused on English and a few major world languages, with images that are Western-centric. While recent efforts have tried to increase the number of languages covered on VQA datasets, they still lack diversity in low-resource languages. More importantly, some datasets extend the text to other languages, either via translation or some other approaches, but usually keep the same images, resulting in narrow cultural representation. To address these limitations, we create CVQA, a new Culturally-diverse Multilingual Visual Question Answering benchmark dataset, designed to cover a rich set of languages and regions, where we engage native speakers and cultural experts in the data collection process. CVQA includes culturally-driven images and questions from across 28 countries in four continents, covering 26 languages with 11 scripts, providing a total of 9k questions. We benchmark several Multimodal Large Language Models (MLLMs) on CVQA, and we show that the dataset is challenging for the current state-of-the-art models. This benchmark will serve as a probing evaluation suite for assessing the cultural bias of multimodal models and hopefully encourage more research efforts towards increasing cultural awareness and linguistic diversity in this field.
Chenyang Lyu, Haryo Akbarianto Wibowo, Santiago Góngora, Aishik Mandal, Sukannya Purkayastha, Jesús-Germán Ortiz-Barajas, Emilio Villa-Cueva, Jinheon Baek, Soyeong Jeong, Injy Hamed, Zheng Wei Lim, Paula Mónica Silva, Jocelyn Dunstan, Mélanie Jouitteau, David Le Meur, Joan Nwatu, Ganzorig Batnasan, Munkh-Erdene Otgonbold, Munkhjargal Gochoo, Guido Ivetta, Luciana Benotti, Laura Alonso Alemany, Hernán Maina, Jiahui Geng, Tiago Timponi Torrent, Frederico Belcavello, Marcelo Viridiano, Jan Christian Blaise Cruz, Dan John Velasco, Oana Ignat, Zara Burzo, Chenxi Whitehouse, Artem Abzaliev, Teresa Clifford, Grainne Caulfield, Teresa Lynn, Christian Salamea Palacios, Vladimir Araujo, Yova Kementchedjhieva, Mihail Mihaylov, Israel Abebe Azime, Henok Biadglign Ademtew, Bontu Fufa Balcha, Naome A. Etori, David Ifeoluwa Adelani, Rada Mihalcea, Atnafu Lambebo Tonja, Maria Camila Buitrago Cabrera, Gisela Vallejo, Holy Lovenia, Ruochen Zhang 0001, Marcos Estecha-Garitagoitia, Mario Rodríguez-Cantelar, Toqeer Ehsan, Rendi Chevi, Muhammad Farid Adilazuarda, Ryandito Diandaru, Samuel Cahyawijaya, Fajri Koto, Tatsuki Kuribayashi, Haiyue Song, Aditya Khandavally, Thanmay Jayakumar, Raj Dabre, Mohamed Fazli Mohamed Imam, Kumaranage Ravindu Yasas Nagasinghe, Alina Dragonetti, Luis Fernando D'Haro, Olivier Niyomugisha, Jay Gala, Pranjal A. Chitale, Fauzan Farooqui, Thamar Solorio, Alham Fikri Aji
NeurIPS75
2023 A Review of Datasets for Aspect-based Sentiment Analysis
abstract
Siva Uday Sampreeth Chebolu, Franck Dernoncourt, Nedim Lipka, Thamar Solorio. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Siva Uday Sampreeth Chebolu, Franck Dernoncourt, Nedim Lipka, Thamar Solorio
IJCNLP (1)4
2022 Style Transfer as Data Augmentation: A Case Study on Named Entity Recognition
abstract
In this work, we take the named entity recognition task in the English language as a case study and explore style transfer as a data augmentation method to increase the size and diversity of training data in low-resource scenarios.We propose a new method to effectively transform the text from a high-resource domain to a low-resource domain by changing its style-related attributes to generate synthetic data for training.Moreover, we design a constrained decoding algorithm along with a set of key ingredients for data selection to guarantee the generation of valid and coherent data.Experiments and analysis on five different domain pairs under different data regimes demonstrate that our approach can significantly improve results compared to current state-of-the-art data augmentation methods.Our approach is a practical solution to data scarcity, and we expect it to be applicable to other NLP tasks. 1
Shuguang Chen, Leonardo Neves, Thamar Solorio
EMNLP3
2022 Hierarchical attention and transformers for automatic movie rating
L. Fernando Pardo-Sixtos, Adrián Pastor López-Monroy, Mahsa Shafaei, Thamar Solorio
Expert Syst. Appl.4
2021 Data Augmentation for Cross-Domain Named Entity Recognition
abstract
Current work in named entity recognition (NER) shows that data augmentation techniques can produce more robust models.However, most existing techniques focus on augmenting in-domain data in low-resource scenarios where annotated data is quite limited.In contrast, we study cross-domain data augmentation for the NER task.We investigate the possibility of leveraging data from highresource domains by projecting it into the lowresource domains.Specifically, we propose a novel neural architecture to transform the data representation from a high-resource to a low-resource domain by learning the patterns (e.g.style, noise, abbreviations, etc.) in the text that differentiate them and a shared feature space where both domains are aligned.We experiment with diverse datasets and show that transforming the data to the low-resource domain representation achieves significant improvements over only using data from highresource domains. 1
Shuguang Chen, Gustavo Aguilar, Leonardo Neves, Thamar Solorio
EMNLP (1)4
2021 Identifying Keyword Predictors in Lecture Video Screen Text
abstract
Automatic discovery of keywords for lecture video segments is an important component of advanced navigation systems for lecture videos. The suitability of a word or a short phrase to be a keyword depends on various factors, including the frequency in a segment, relative frequency in reference to the full video, font size, time on screen, and the existence in domain and language dictionaries. The research presented in this paper provides a refined understanding of how various factors contribute to predicting keywords based on logistic regression analysis. The analysis employs a real-world dataset consisting of lecture videos from Biology, Computer Science, and Chemistry, hosted on Videopoints, a lecture video management portal. Term frequency, maximum font size, and presence in a domain dictionary were identified as the most important predictors of keywords. The results provide a scientific foundation and valuable insights into the design of future keyword prediction systems.
Farah Naz Chowdhury, Raga Shalini Koka, Mohammad Rajiur Rahman, Thamar Solorio, Jaspal Subhlok
ISM4
2021 Exploring Conditional Text Generation for Aspect-Based Sentiment Analysis
Siva Uday Sampreeth Chebolu, Franck Dernoncourt, Nedim Lipka, Thamar Solorio
PACLIC4
2021 A Human-Centered Systematic Literature Review of the Computational Approaches for Online Sexual Risk Detection
abstract
In the era of big data and artificial intelligence, online risk detection has become a popular research topic. From detecting online harassment to the sexual predation of youth, the state-of-the-art in computational risk detection has the potential to protect particularly vulnerable populations from online victimization. Yet, this is a high-risk, high-reward endeavor that requires a systematic and human-centered approach to synthesize disparate bodies of research across different application domains, so that we can identify best practices, potential gaps, and set a strategic research agenda for leveraging these approaches in a way that betters society. Therefore, we conducted a comprehensive literature review to analyze 73 peer-reviewed articles on computational approaches utilizing text or meta-data/multimedia for online sexual risk detection. We identified sexual grooming (75%), sex trafficking (12%), and sexual harassment and/or abuse (12%) as the three types of sexual risk detection present in the extant literature. Furthermore, we found that the majority (93%) of this work has focused on identifying sexual predators after-the-fact, rather than taking more nuanced approaches to identify potential victims and problematic patterns that could be used to prevent victimization before it occurs. Many studies rely on public datasets (82%) and third-party annotators (33%) to establish ground truth and train their algorithms. Finally, the majority of this work (78%) mostly focused on algorithmic performance evaluation of their model and rarely (4%) evaluate these systems with real users. Thus, we urge computational risk detection researchers to integrate more human-centered approaches to both developing and evaluating sexual risk detection algorithms to ensure the broader societal impacts of this important work.
Afsaneh Razi, Ashwaq Alsoubai, Gianluca Stringhini, Thamar Solorio, Munmun De Choudhury, Pamela J. Wisniewski
Proc. ACM Hum. Comput. Interact.5
2020 From English to Code-Switching: Transfer Learning with Strong Morphological Clues
abstract
Linguistic Code-switching (CS) is still an understudied phenomenon in natural language processing.The NLP community has mostly focused on monolingual and multi-lingual scenarios, but little attention has been given to CS in particular.This is partly because of the lack of resources and annotated data, despite its increasing occurrence in social media platforms.In this paper, we aim at adapting monolingual models to code-switched text in various tasks.Specifically, we transfer English knowledge from a pre-trained ELMo model to different code-switched language pairs (i.e., Nepali-English, Spanish-English, and Hindi-English) using the task of language identification.Our method, CS-ELMo, is an extension of ELMo with a simple yet effective position-aware attention mechanism inside its character convolutions.We show the effectiveness of this transfer learning step by outperforming multilingual BERT and homologous CS-unaware ELMo models and establishing a new state of the art in CS tasks, such as NER and POS tagging.Our technique can be expanded to more English-paired code-switched languages, providing more resources to the CS community.
Gustavo Aguilar, Thamar Solorio
ACL2
2020 Let Me Choose: From Verbal Context to Font Selection
abstract
In this paper, we aim to learn associations between visual attributes of fonts and the verbal context of the texts they are typically applied to. Compared to related work leveraging the surrounding visual context, we choose to focus only on the input text as this can enable new applications for which the text is the only visual element in the document. We introduce a new dataset, containing examples of different topics in social media posts and ads, labeled through crowd-sourcing. Due to the subjective nature of the task, multiple fonts might be perceived as acceptable for an input text, which makes this problem challenging. To this end, we investigate different end-to-end models to learn label distributions on crowd-sourced data and capture inter-subjectivity across all annotations.
Amirreza Shirani, Franck Dernoncourt, Jose Echevarria, Paul Asente, Nedim Lipka, Thamar Solorio
ACL6
2020 Multi-view Story Characterization from Movie Plot Synopses and Reviews
abstract
This paper considers the problem of characterizing stories by inferring properties such as theme and style using written synopses and reviews of movies. We experiment with a multi-label dataset of movie synopses and a tagset representing various attributes of stories (e.g., genre, type of events). Our proposed multi-view model encodes the synopses and reviews using hierarchical attention and shows improvement over methods that only use synopses. Finally, we demonstrate how can we take advantage of such a model to extract a complementary set of story-attributes from reviews without direct supervision. We have made our dataset and source code publicly available at https://ritual.uh.edu/ multiview-tag-2020.
Sudipta Kar, Gustavo Aguilar, Mirella Lapata, Thamar Solorio
EMNLP (1)4
2020 Automatic Identification of Keywords in Lecture Video Segments
abstract
Lecture video is an increasingly important learning resource. However, the challenge of quickly finding the content of interest in a long lecture video is a critical limitation of this format. This paper introduces automatic discovery of keywords (or tags) for lecture video segments to improve navigation. A lecture video is divided into topical segments based on the frame-to-frame similarity of content. A user navigates the lecture video assisted by visual summaries and keywords for the segments. Keywords provide an overview of the content discussed in the segment to improve navigation. The input to the keyword identification algorithm is the text from the video frames extracted by OCR. Automatically discovering keywords is challenging as the suitability of an N-gram to be a keyword depends on a variety of factors including frequency in a segment and relative frequency in reference to the full video, font size, time on screen, and the existence in domain and language dictionaries. This paper explores how these factors are quantified and combined to identify good keywords. The key scientific contribution of this paper is the design, implementation, and evaluation of a keyword selection algorithm for lecture video segments. Evaluation is performed by comparing the keywords generated by the algorithm with the tags chosen by experts on 121 segments of 11 videos from STEM courses.
Raga Shalini Koka, Farah Naz Chowdhury, Mohammad Rajiur Rahman, Thamar Solorio, Jaspal Subhlok
ISM4
2020 LinCE: A Centralized Benchmark for Linguistic Code-switching Evaluation
abstract
Recent trends in NLP research have raised an interest in linguistic code-switching (CS); modern approaches have been proposed to solve a wide range of NLP tasks on multiple language pairs. Unfortunately, these proposed methods are hardly generalizable to different code-switched languages. In addition, it is unclear whether a model architecture is applicable for a different task while still being compatible with the code-switching setting. This is mainly because of the lack of a centralized benchmark and the sparse corpora that researchers employ based on their specific needs and interests. To facilitate research in this direction, we propose a centralized benchmark for Linguistic Code-switching Evaluation (LinCE) that combines eleven corpora covering four different code-switched language pairs (i.e., Spanish-English, Nepali-English, Hindi-English, and Modern Standard Arabic-Egyptian Arabic) and four tasks (i.e., language identification, named entity recognition, part-of-speech tagging, and sentiment analysis). As part of the benchmark centralization effort, we provide an online platform where researchers can submit their results while comparing with others in real-time. In addition, we provide the scores of different popular models, including LSTM, ELMo, and multilingual BERT so that the NLP community can compare against state-of-the-art systems. LinCE is a continuous effort, and we will expand it with more low-resource languages and tasks.
Gustavo Aguilar, Sudipta Kar, Thamar Solorio
LREC3
2020 Age Suitability Rating: Predicting the MPAA Rating Based on Movie Dialogues
abstract
Movies help us learn and inspire societal change. But they can also contain objectionable content that negatively affects viewers’ behaviour, especially children. In this paper, our goal is to predict the suitability of movie content for children and young adults based on scripts. The criterion that we use to measure suitability is the MPAA rating that is specifically designed for this purpose. We create a corpus for movie MPAA ratings and propose an RNN based architecture with attention that jointly models the genre and the emotions in the script to predict the MPAA rating. We achieve 81% weighted F1-score for the classification model that outperforms the traditional machine learning method by 7%.
Mahsa Shafaei, Niloofar Safi Samghabadi, Sudipta Kar, Thamar Solorio
LREC4
2020 Early author profiling on Twitter using profile features with multi-resolution
Adrián Pastor López-Monroy, Fabio A. González 0001, Thamar Solorio
Expert Syst. Appl.3
2020 Gated multimodal networks
John Edison Arevalo Ovalle, Thamar Solorio, Manuel Montes-y-Gómez, Fabio A. González 0001
Neural Comput. Appl.2
2019 Learning Emphasis Selection for Written Text in Visual Media from Crowd-Sourced Label Distributions
abstract
Amirreza Shirani, Franck Dernoncourt, Paul Asente, Nedim Lipka, Seokhwan Kim, Jose Echevarria, Thamar Solorio. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Amirreza Shirani, Franck Dernoncourt, Paul Asente, Nedim Lipka, Seokhwan Kim, Jose Echevarria, Thamar Solorio
ACL (1)7
2018 Folksonomication: Predicting Tags for Movies from Plot Synopses using Emotion Flow Encoded Neural Network
abstract
Folksonomy of movies covers a wide range of heterogeneous information about movies, like the genre, plot structure, visual experiences, soundtracks, metadata, and emotional experiences from watching a movie. Being able to automatically generate or predict tags for movies can help recommendation engines improve retrieval of similar movies, and help viewers know what to expect from a movie in advance. In this work, we explore the problem of creating tags for movies from plot synopses. We propose a novel neural network model that merges information from synopses and emotion flows throughout the plots to predict a set of tags for movies. We compare our system with multiple baselines and found that the addition of emotion flows boosts the performance of the network by learning ≈18% more tags than a traditional machine learning system.
Sudipta Kar, Suraj Maharjan, Thamar Solorio
COLING3
2018 A Genre-Aware Attention Model to Improve the Likability Prediction of Books
abstract
Likability prediction of books has many uses. Readers, writers, as well as the publishing industry, can all benefit from automatic book likability prediction systems. In order to make reliable decisions, these systems need to assimilate information from different aspects of a book in a sensible way. We propose a novel multimodal neural architecture that incorporates genre supervision to assign weights to individual feature types. Our proposed method is capable of dynamically tailoring weights given to feature types based on the characteristics of each book. Our architecture achieves competitive results and even outperforms state-of-the-art for this task.
Suraj Maharjan, Manuel Montes-y-Gómez, Fabio A. González 0001, Thamar Solorio
EMNLP4
2018 MPST: A Corpus of Movie Plot Synopses with Tags
Sudipta Kar, Suraj Maharjan, Adrián Pastor López-Monroy, Thamar Solorio
LREC4
2018 Modeling Noisiness to Recognize Named Entities using Multitask Neural Networks on Social Media
abstract
Gustavo Aguilar, Adrian Pastor López-Monroy, Fabio González, Thamar Solorio. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Gustavo Aguilar, Adrián Pastor López-Monroy, Fabio A. González 0001, Thamar Solorio
NAACL-HLT4
2018 Early Text Classification Using Multi-Resolution Concept Representations
abstract
Adrian Pastor López-Monroy, Fabio A. González, Manuel Montes, Hugo Jair Escalante, Thamar Solorio. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Adrián Pastor López-Monroy, Fabio A. González 0001, Manuel Montes-y-Gómez, Hugo Jair Escalante, Thamar Solorio
NAACL-HLT5
2017 Towards Translating Mixed-Code Comments from Social Media
Thoudam Doren Singh, Thamar Solorio
CICLing (2)2
2017 A Multi-task Approach to Predict Likability of Books
abstract
Suraj Maharjan, John Arevalo, Manuel Montes, Fabio A. González, Thamar Solorio. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Suraj Maharjan, John Edison Arevalo Ovalle, Manuel Montes-y-Gómez, Fabio A. González 0001, Thamar Solorio
EACL (1)5
2016 Domain Adaptation for Authorship Attribution: Improved Structural Correspondence Learning
Upendra Sapkota, Thamar Solorio, Manuel Montes-y-Gómez, Steven Bethard
ACL (1)2
2016 Large Scale Authorship Attribution of Online Reviews
Prasha Shrestha, Arjun Mukherjee, Thamar Solorio
CICLing (2)3
2016 Computational Approaches to Linguistic Code Switching
Mona T. Diab, Pascale Fung, Julia Hirschberg, Thamar Solorio
INTERSPEECH4
2016 Age and Gender Prediction on Health Forum Data
Prasha Shrestha, Nicolas Rey-Villamizar, Farig Sadeque, Ted Pedersen, Steven Bethard, Thamar Solorio
LREC6
2015 Identification of Original Document by Using Textual Similarities
Prasha Shrestha, Thamar Solorio
CICLing (2)2
2015 Not All Character N-grams Are Created Equal: A Study in Authorship Attribution
abstract
Upendra Sapkota, Steven Bethard, Manuel Montes, Thamar Solorio. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Upendra Sapkota, Steven Bethard, Manuel Montes-y-Gómez, Thamar Solorio
HLT-NAACL4
2014 Cross-Topic Authorship Attribution: Will Out-Of-Topic Data Help?
Upendra Sapkota, Thamar Solorio, Manuel Montes-y-Gómez, Steven Bethard, Paolo Rosso
COLING2
2014 Sockpuppet Detection in Wikipedia: A Corpus of Real-World Deceptive Writing for Linking Identities
Thamar Solorio, Ragib Hasan, Mainul Mizan
LREC1
2014 Exploring high-level features for detecting cyberpedophilia
Dasha Bogdanova, Paolo Rosso, Thamar Solorio
Comput. Speech Lang.3
2013 The Use of Orthogonal Similarity Relations in the Prediction of Authorship
Upendra Sapkota, Thamar Solorio, Manuel Montes-y-Gómez, Paolo Rosso
CICLing (2)2
2012 Evaluating NLP Features for Automatic Prediction of Language Impairment Using Child Speech Transcripts
abstract
Language impairment (LI) in children is pervasive in all walks of life. Automatic prediction of LI is useful as a first pass for speech language pathologists in identifying prospective children with LI. Previous work in the automatic prediction of LI has explored various features, mostly shallow and surface level features. In this paper, we evaluate deeper Natural Language Processing (NLP) features such as syntactic, semantic and entity grid model features, along with narrative structure and quality features in the prediction of LI using child language transcripts. Our experiments show that narrative structure and quality features along with a combination of other features are helpful in the prediction of LI in storytelling narratives.
Khairun-nisa Hassanali, Yang Liu 0004, Thamar Solorio
INTERSPEECH3
2011 Local Histograms of Character N-grams for Authorship Attribution
Hugo Jair Escalante, Thamar Solorio, Manuel Montes-y-Gómez
ACL2
2011 Modality Specific Meta Features for Authorship Attribution in Web Forum Posts
Thamar Solorio, Sangita Pillay, Sindhu Raghavan, Manuel Montes-y-Gómez
IJCNLP1
2011 Exploring a corpus-based approach for detecting language impairment in monolingual English-speaking children
Keyur Gabani, Thamar Solorio, Yang Liu 0004, Khairun-nisa Hassanali, Christine A. Dollaghan
Artif. Intell. Medicine2
2011 Analyzing language samples of Spanish-English bilingual children for the automated prediction of language dominance
abstract
Abstract In this work we study how features typically used in natural language processing tasks, together with measures from syntactic complexity, can be adapted to the problem of developing language profiles of bilingual children. Our experiments show that these features can provide high discriminative value for predicting language dominance from story retells in a Spanish–English bilingual population of children. Moreover, some of our proposed features are even more powerful than measures commonly used by clinical researchers and practitioners for analyzing spontaneous language samples of children. This study shows that the field of natural language processing has the potential to make significant contributions to communication disorders and related areas.
Thamar Solorio, Melissa Sherman, Yang Liu 0004, Lisa Bedore, Elizabeth Peña, Aquiles Iglesias
Nat. Lang. Eng.1
2009 A Corpus-Based Approach for the Prediction of Language Impairment in Monolingual English and Spanish-English Bilingual Children
Keyur Gabani, Melissa Sherman, Thamar Solorio, Yang Liu 0004, Lisa Bedore, Elizabeth Peña
HLT-NAACL3
2008 Learning to Predict Code-Switching Points
Thamar Solorio, Yang Liu 0004
EMNLP1
2008 Part-of-Speech Tagging for English-Spanish Code-Switched Text
Thamar Solorio, Yang Liu 0004
EMNLP1
2008 On the Effectiveness of Rebuilding RNA Secondary Structures from Sequence Chunks
abstract
Despite the computing power of emerging technologies, predicting long RNA secondary structures with thermodynamics-based methods is still infeasible, especially if the structures include complex motifs such as pseudoknots. This paper presents preliminary results on rebuilding RNA secondary structures by an extensive and systematic sampling of nucleotide chunks. The rebuilding approach merges the significant motifs found in the secondary structures of the single chunks. The extensive sampling and prediction of nucleotide chunks are supported by grid technology as part of the RNAVLab functionality. Significant motifs are identified in the chunk secondary structures and merged in a single structure based on their recurrences and other statistical insights. A critical analysis of the strengths, weaknesses, and future developments of our method is presented.
Michela Taufer, Thamar Solorio, Abel Licon, David Mireles, Ming-Ying Leung
IPDPS2
2008 RNAVLab: A virtual laboratory for studying RNA secondary structures based on grid computing technology
Michela Taufer, Ming-Ying Leung, Thamar Solorio, Abel Licon, David Mireles, Roberto Araiza, Kyle L. Johnson
Parallel Comput.3
2007 Baby-Steps Towards Building a Spanglish Language Model
Juan Carlos Franco, Thamar Solorio
CICLing2
2006 An Unsupervised Language Independent Method of Name Discrimination Using Second Order Co-occurrence Features
Ted Pedersen, Anagha Kulkarni 0001, Roxana Angheluta, Zornitsa Kozareva, Thamar Solorio
CICLing5
2006 Prosodic feature generation for back-channel prediction
abstract
Using prosodic information to predict when back-channels are appropriate in spontaneous dialogs has become somewhat of a reference problem for automatic discovery techniques. Here we present experiments with two ideas: the use of features derived from randomly generated pitch and energy filters, and the use of instancebased learning, specifically the Locally Weighted Linear Regression (LWLR) algorithm. For the task of predicting possible backchannel locations in Iraqi Arabic [6], we obtain 22 % precision and 51 % recall, which is as good as that obtained using a laboriously developed and hand-tuned rule. 1.
Thamar Solorio, Olac Fuentes, Nigel G. Ward, Yaffa Al Bayyari
INTERSPEECH1
2005 Exploiting Named Entity Taggers in a Second Language
Thamar Solorio
ACL1
2005 Learning Named Entity Recognition in Portuguese from Spanish
Thamar Solorio, Aurelio López-López
CICLing1
2005 Question Classification in Spanish and Portuguese
Thamar Solorio, Manuel Alberto Pérez-Coutiño, Manuel Montes-y-Gómez, Luis Villaseñor-Pineda, Aurelio López-López
CICLing1
2004 Learning Named Entity Classifiers Using Support Vector Machines
Thamar Solorio, Aurelio López-López
CICLing1
2004 A Language Independent Method for Question Classification
Thamar Solorio, Manuel Alberto Pérez-Coutiño, Manuel Montes-y-Gómez, Luis Villaseñor-Pineda, Aurelio López-López
COLING1