VLDB 2026 Research / reviewers in the wild / expert
Mehwish Fatima
dblp:174/2279
· DBLP profile ↗
3ranked-venue papers
3as first author
1since 2021 · last 2023
0000-0003-3424-2991ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 91% Machine translation · 9% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › text summarization › long document summarization
scientific paper summarization |
0.7 | 1 | 2023 | Cross-lingual Science Journalism: Select, Simplify and Rewrite Summaries for Non-expert Readers · ACL (1) 2023 |
Natural language and speech › Language models and text generation › text generation
text simplification |
0.7 | 1 | 2023 | Cross-lingual Science Journalism: Select, Simplify and Rewrite Summaries for Non-expert Readers · ACL (1) 2023 |
Natural language and speech › Language models and text generation
text summarization |
0.7 | 1 | 2023 | Cross-lingual Science Journalism: Select, Simplify and Rewrite Summaries for Non-expert Readers · ACL (1) 2023 |
Natural language and speech › Machine translation
cross-lingual generation |
0.2 | 1 | 2023 | Cross-lingual Science Journalism: Select, Simplify and Rewrite Summaries for Non-expert Readers · ACL (1) 2023 |
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Cross-lingual Science Journalism: Select, Simplify and Rewrite Summaries for Non-expert ReadersabstractAutomating Cross-lingual Science Journalism (CSJ) aims to generate popular science summaries from English scientific texts for nonexpert readers in their local language.We introduce CSJ as a downstream task of text simplification and cross-lingual scientific summarization to facilitate science journalists' work.We analyze the performance of possible existing solutions as baselines for the CSJ task.Based on these findings, we propose to combine the three components -SELECT, SIMPLIFY and REWRITE (SSR) to produce cross-lingual simplified science summaries for non-expert readers.Our empirical evaluation on the WIKIPEDIA dataset shows that SSR significantly outperforms the baselines for the CSJ task and can serve as a strong baseline for future work.We also perform an ablation study investigating the impact of individual components of SSR.Further, we analyze the performance of SSR on a high-quality, real-world CSJ dataset with human evaluation and in-depth analysis, demonstrating the superior performance of SSR for CSJ. A Scientific and News StructureFigure A.1 presents the difference between a scientific text discourse and a news text discourse. Mehwish Fatima, Michael Strube 0001 |
ACL (1) | 1 |
| 2018 | Multilingual SMS-based author profiling: Data and methodsabstractAbstract In the recent years, many benchmark author profiling corpora have been developed for various genres including Twitter, social media, blogs, hotel reviews and e-mail, etc. However, no such standard evaluation resource has been developed for Short Messaging Service (SMS), a popular medium of communication, which is very useful for author profiling. The primary aim of this study is to develop a large multilingual (English and Roman Urdu) benchmark SMS-based author profiling corpus. The proposed corpus contains 810 author profiles, wherein each profile consists of an aggregation of SMS messages as a single document of an author, along with seven demographic traits associated with each author profile: gender, age, native language, native city, qualification, occupation and personality type (introvert/extrovert). The secondary aims of this study include the following: (1) annotating the proposed corpus for code-switching annotations at the lexical level (approximately 0.69 million tokens are manually annotated for code-switching) and (2) applying the stylometry-based method (groups of sixty-four features) and the content-based method (twelve features) for gender identification in order to demonstrate how our proposed corpus can be used for the development and evaluation of various author profiling methods. The results show that the content-based character 5-gram feature outperformed all the other features by obtaining the accuracy score of 0.975 andF1score of 0.947 for gender identification while using the entire corpus. Furthermore, our proposed corpora (SMS–AP–18 and code-switched SMS–AP–18) are freely and publicly available for research purpose. Mehwish Fatima, Saba Anwar, Amna Naveed, Waqas Arshad |
Nat. Lang. Eng. | 1 |
| 2017 | Multilingual author profiling on Facebook
Mehwish Fatima, Komal Hasan, Saba Anwar, Rao Muhammad Adeel Nawab |
Inf. Process. Manag. | 1 |