Mehwish Fatima

dblp:174/2279 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
1since 2021 · last 2023
0000-0003-3424-2991ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 91% Machine translation · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › text summarization › long document summarization
scientific paper summarization
0.712023
Cross-lingual Science Journalism: Select, Simplify and Rewrite Summaries for Non-expert Readers · ACL (1) 2023
Natural language and speech › Language models and text generation › text generation
text simplification
0.712023
Cross-lingual Science Journalism: Select, Simplify and Rewrite Summaries for Non-expert Readers · ACL (1) 2023
Natural language and speech › Language models and text generation
text summarization
0.712023
Cross-lingual Science Journalism: Select, Simplify and Rewrite Summaries for Non-expert Readers · ACL (1) 2023
Natural language and speech › Machine translation
cross-lingual generation
0.212023
Cross-lingual Science Journalism: Select, Simplify and Rewrite Summaries for Non-expert Readers · ACL (1) 2023
YearPublicationVenuePosition
2023 Cross-lingual Science Journalism: Select, Simplify and Rewrite Summaries for Non-expert Readers
abstract
Automating Cross-lingual Science Journalism (CSJ) aims to generate popular science summaries from English scientific texts for nonexpert readers in their local language.We introduce CSJ as a downstream task of text simplification and cross-lingual scientific summarization to facilitate science journalists' work.We analyze the performance of possible existing solutions as baselines for the CSJ task.Based on these findings, we propose to combine the three components -SELECT, SIMPLIFY and REWRITE (SSR) to produce cross-lingual simplified science summaries for non-expert readers.Our empirical evaluation on the WIKIPEDIA dataset shows that SSR significantly outperforms the baselines for the CSJ task and can serve as a strong baseline for future work.We also perform an ablation study investigating the impact of individual components of SSR.Further, we analyze the performance of SSR on a high-quality, real-world CSJ dataset with human evaluation and in-depth analysis, demonstrating the superior performance of SSR for CSJ. A Scientific and News StructureFigure A.1 presents the difference between a scientific text discourse and a news text discourse.
Mehwish Fatima, Michael Strube 0001
ACL (1)1
2018 Multilingual SMS-based author profiling: Data and methods
abstract
Abstract In the recent years, many benchmark author profiling corpora have been developed for various genres including Twitter, social media, blogs, hotel reviews and e-mail, etc. However, no such standard evaluation resource has been developed for Short Messaging Service (SMS), a popular medium of communication, which is very useful for author profiling. The primary aim of this study is to develop a large multilingual (English and Roman Urdu) benchmark SMS-based author profiling corpus. The proposed corpus contains 810 author profiles, wherein each profile consists of an aggregation of SMS messages as a single document of an author, along with seven demographic traits associated with each author profile: gender, age, native language, native city, qualification, occupation and personality type (introvert/extrovert). The secondary aims of this study include the following: (1) annotating the proposed corpus for code-switching annotations at the lexical level (approximately 0.69 million tokens are manually annotated for code-switching) and (2) applying the stylometry-based method (groups of sixty-four features) and the content-based method (twelve features) for gender identification in order to demonstrate how our proposed corpus can be used for the development and evaluation of various author profiling methods. The results show that the content-based character 5-gram feature outperformed all the other features by obtaining the accuracy score of 0.975 andF1score of 0.947 for gender identification while using the entire corpus. Furthermore, our proposed corpora (SMS–AP–18 and code-switched SMS–AP–18) are freely and publicly available for research purpose.
Mehwish Fatima, Saba Anwar, Amna Naveed, Waqas Arshad
Nat. Lang. Eng.1
2017 Multilingual author profiling on Facebook
Mehwish Fatima, Komal Hasan, Saba Anwar, Rao Muhammad Adeel Nawab
Inf. Process. Manag.1