Mehwish Nasim

dblp:91/8376 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-0683-9125ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection
abstract
We investigate the efficacy of Large Language Models (LLMs) in detecting implicit and explicit hate speech, examining how models with minimal safety alignment (uncensored) compare with more heavily aligned (censored) counterparts in a deployed-model setting when deployed using political personas.While uncensored models are often framed as offering a less constrained perspective, our results reveal a trade-off: censored models outperform their uncensored counterparts in both accuracy and robustness, achieving 69.0% versus 64.1% strict accuracy.However, this higher performance is also associated with greater resistance to persona-based influence, while uncensored models are more malleable to ideological framing.Furthermore, we identify critical failures across all models in understanding nuanced language such as irony.We also find alarming fairness disparities in performance across different targeted groups and systemic overconfidence that renders self-reported certainty unreliable.These findings challenge the notion of LLMs as objective arbiters and highlight the need for more sophisticated auditing frameworks that account for fairness, calibration, and ideological consistency.Taken together, these results point to censorship-as-deployed rather than safety alignment in isolation as the more appropriate frame for interpreting model differences.
Sanjeevan Selvaganapathy, Mehwish Nasim
ACL (1)2
2026 They Said Memes Were Harmless - We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural References
abstract
Meme-based social abuse detection is challenging because harmful intent often relies on implicit cultural symbolism and subtle cross-modal incongruence. Prior approaches, from fusion-based methods to in-context learning with Large Vision-Language Models (LVLMs), have made progress but remain limited by three factors: i) cultural blindness (missing symbolic context), ii) boundary ambiguity (satire vs. abuse confusion), and iii) lack of interpretability (opaque model reasoning). We introduce CROSS-ALIGN+, a three-stage framework that systematically addresses these limitations: (1) Stage I mitigates cultural blindness by enriching multimodal representations with structured knowledge from ConceptNet, Wikidata, and Hatebase; (2) Stage II reduces boundary ambiguity through parameter-efficient LoRA adapters that sharpen decision boundaries; and (3) Stage III enhances interpretability by generating cascaded explanations. Extensive experiments on five benchmarks and eight LVLMs demonstrate that CROSS-ALIGN+ consistently outperforms state-of-the-art methods, achieving up to 17% relative F1 improvement while providing interpretable justifications for each decision.
Sahil Tripathi, Gautam Siddharth Kashyap, Mehwish Nasim, Jian Yang 0001, Jiechao Gao, Usman Naseem
WWW3
2026 DiffCom: Decoupled Sparse Priors Guided Diffusion Compression for Point Clouds
abstract
While conventional lossy compression methods predominantly depend on autoencoders to map point clouds into latent representations, they often neglect the intrinsic redundancy within these latent points. To address this limitation, this paper presents a diffusion-based architecture steered by sparse priors, designed to minimize latent redundancy while securing superior reconstruction fidelity, particularly in low-bitrate scenarios. A key feature of the framework is an efficient dual-density data flow that alleviates the stringent size constraints imposed on latent points. By integrating a Probabilistic Attention-based Conditional Denoiser (PACD), the method effectively encapsulates critical reconstruction details within sparse priors, which are hierarchically decoupled into intra- and inter-point components. Specifically, separate encoders are utilized to transform the source point cloud into latent points and decoupled sparse priors, respectively. To dynamically exploit geometric and semantic information, an attention-driven latent denoiser, conditioned on these decoupled priors, is applied across the encoding and decoding layers. Furthermore, inter-point distributions are incorporated into the arithmetic codec to refine local context modeling for sparse points, with the final point cloud recovered via a point decoder. Comprehensive experiments conducted on ShapeNet and standard MPEG PCC datasets demonstrate that the proposed method outperforms state-of-the-art techniques, achieving a superior rate-distortion trade-off.
Xiaoge Zhang 0003, Mingtao Feng, Mehwish Nasim, Saeed Anwar, Ajmal Mian
IEEE Trans. Circuits Syst. Video Technol.4
2025 Simulating Influence Dynamics with LLM Agents
Mehwish Nasim, Syed Muslim M. Gilani, Amin Qasmi, Usman Naseem
IEEE Big Data1
2025 Competing LLM Agents in a Non-Cooperative Game of Opinion Polarisation
Amin Qasmi, Usman Naseem, Mehwish Nasim
IEEE Big Data3
2025 Can LLM Agents Maintain a Persona in Discourse?
abstract
Large Language Models (LLMs) are widely used as conversational agents, exploiting their capabilities in various sectors such as education, law, medicine, and more.However, LLMs are often subjected to context-shifting behaviour, resulting in a lack of consistent and interpretable personality-aligned interactions.Adherence to psychological traits lacks comprehensive analysis, especially in the case of dyadic (pairwise) conversations.We examine this challenge from two viewpoints, initially using two conversation agents to generate a discourse on a certain topic with an assigned personality from the OCEAN framework (Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism) as High/Low for each trait.This is followed by using multiple judge agents to infer the original traits assigned to explore prediction consistency, intermodel agreement, and alignment with the assigned personality.Our findings indicate that while LLMs can be guided toward personalitydriven dialogue, their ability to maintain personality traits varies significantly depending on the combination of models and discourse settings.These inconsistencies emphasise the challenges in achieving stable and interpretable personality-aligned interactions in LLMs.
Pranav Bhandari, Nicolas Fay, Michael J. Wise, Amitava Datta, Stephanie Meek, Usman Naseem, Mehwish Nasim
EMNLP7
2025 Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval
abstract
Current text-driven Video Moment Retrieval (VMR) methods encode all video clips, including irrelevant ones, disrupting multimodal alignment and hindering optimization. To this end, we propose a denoise-then-retrieve paradigm that explicitly filters text-irrelevant clips from videos and then retrieves the target moment using purified multimodal representations. Following this paradigm, we introduce the Denoise-then-Retrieve Network (DRNet), comprising Text-Conditioned Denoising (TCD) and Text-Reconstruction Feedback (TRF) modules. TCD integrates cross-attention and structured state space blocks to dynamically identify noisy clips and produce a noise mask to purify multimodal video representations. TRF further distills a single query embedding from purified video representations and aligns it with the text embedding, serving as auxiliary supervision for denoising during training. Finally, we perform conditional retrieval using text embeddings on purified video representations for accurate VMR. Experiments on Charades-STA and QVHighlights demonstrate that our approach surpasses state-of-the-art methods on all metrics. Furthermore, our denoise-then-retrieve paradigm is adaptable and can be seamlessly integrated into advanced VMR models to boost performance.
Jiuxin Cao, Bo Miao, Zhiheng Fu, Xuelin Zhu, Jiawei Ge 0002, Bo Liu 0004, Mehwish Nasim, Ajmal Mian
IJCAI8
2025 Localization Lens for Improving Medical Vision-Language Models
Hasan Farooq, Murtaza Taj, Mehwish Nasim, Arif Mahmood
MICCAI (9)3
2025 Context-Enhanced Video Moment Retrieval With Large Language Models
abstract
Current methods for Video Moment Retrieval (VMR) struggle to align complex situations involving specific environmental details, character descriptions, and action narratives. To tackle this issue, we propose a Large Language Model-guided Moment Retrieval (LMR) approach that employs the extensive knowledge of Large Language Models (LLMs) to improve video context representation as well as cross-modal alignment, facilitating accurate localization of target moments. Specifically, LMR introduces a context enhancement technique with LLMs to generate crucial target-related context semantics. These semantics are integrated with visual features for producing discriminative video representations. Finally, a language-conditioned transformer is designed to decode free-form language queries, on the fly, using aligned video representations for moment retrieval. Extensive experiments demonstrate that LMR achieves state-of-the-art results, outperforming the nearest competitor by up to 3.28% and 4.06% on the challenging QVHighlights and Charades-STA benchmarks, respectively. More importantly, the performance gains are significantly higher for localization of complex queries.
Bo Miao, Jiuxin Cao, Xuelin Zhu, Jiawei Ge 0002, Bo Liu 0004, Mehwish Nasim, Ajmal Mian
IEEE Trans. Multim.7
2023 NEHATE: Large-Scale Annotated Data Shedding Light on Hate Speech in Nepali Local Election Discourse
abstract
The use of social media during election campaigns has become increasingly popular. However, the unbridled nature of online discourse can lead to the propagation of hate speech, which has far-reaching implications for the democratic process. Natural Language Processing (NLP) techniques are being used to counteract the spread of hate speech and promote healthy online discourse. Despite the increasing need for NLP techniques to combat hate speech, research on low-resource languages such as Nepali is limited, posing a challenge to the realization of the United Nations’ Leave No One Behind principle, which calls for inclusive development that benefits all individuals and communities, regardless of their backgrounds or circumstances. To bridge this gap, we introduce NEHATE, a large-scale manually annotated dataset of hate speech and its targets in Nepali local election discourse. The dataset comprises 13,505 tweets, annotated for hate speech with further sub-categorization of hate speech into targets such as community, individual, and organization. Benchmarking of the dataset with various algorithms has shown potential for performance improvement. We have made the dataset publicly available at https://github.com/shucoll/NEHate to promote further research and development, while also contributing to the UN SDGs aimed at fostering peaceful, inclusive societies, and justice and strong institutions.
Surendrabikram Thapa, Kritesh Rauniyar, Shuvam Shiwakoti, Sweta Poudel, Usman Naseem, Mehwish Nasim
ECAI6
2023 MDKG: Graph-Based Medical Knowledge-Guided Dialogue Generation
Usman Naseem, Surendrabikram Thapa, Qi Zhang 0020, Liang Hu 0004, Mehwish Nasim
SIGIR5
2022 Simulating cyber security management: A gamified approach to executive decision making
abstract
Executive managers are not all equipped with the cyber security expertise necessary to enable them to make business decisions that accurately represent the status and needs of the cyber security side of the business. Unfortunately, the lack of understanding between the business and cyber security domains contribute to structurally endorsed vulnerabilities within a business context, where either the business needs were considered without understanding the impact on cyber security, or alternatively, the cyber security needs were considered without fully understanding the impact this would have on the business strategy and financial stability. To combat this dilemma, a gamified approach to cyber security training for executives is proposed as a solution to not only minimise the realisation of cyber vulnerabilities within a business context, but also to improve business outcomes that are supported by cyber security measures. We developed a serious game software platform, Aurelius, to simulate an executive decision maker’s role in managing the everyday cyber security investment decisions, and linking that to business metrics to incorporate the business and cyber security understanding. Our game includes simulated cyber security attacks that would require the executive decision maker (the player) to respond appropriately. The algorithms underpinning our simulated cyber security game are a product of a complex systems approach, as this most accurately models an executive’s experience. In our design, we set up Aurelius to fulfil eight of the nine criteria specified for a state of the art serious game in the cyber security domain.
Adam Tonkin, William Kosasih, Marthie Grobler, Mehwish Nasim
ASE4
2020 A method to evaluate the reliability of social media data for social network analysis
abstract
In order to study the effects of Online Social Network (OSN) activity on real-world offline events, researchers need access to OSN data, the reliability of which has particular implications for social network analysis. This relates not only to the completeness of any collected dataset, but also to constructing meaningful social and information networks from them. In this multidisciplinary study, we consider the question of constructing traditional social networks from OSN data and then present a measurement case study showing how the reliability of OSN data affects social network analyses. To this end we developed a systematic comparison methodology, which we applied to two parallel datasets we collected from Twitter. We found considerable differences in datasets collected with different tools and that these variations significantly alter the results of subsequent analyses. Our results lead to a set of guidelines for researchers planning to collect online data streams to infer social networks.
Derek Weber, Mehwish Nasim, Lewis Mitchell, Lucia Falzon
ASONAM2
2020 Pachinko Prediction: A Bayesian method for event prediction from social media data
Simon Jonathan Tuke, Andrew Nguyen, Mehwish Nasim, Drew Mellor, Asanga Wickramasinghe, Nigel G. Bean, Lewis Mitchell
Inf. Process. Manag.3
2019 Gathering Cyber Threat Intelligence from Twitter Using Novelty Classification
abstract
The following topics are dealt with: feature extraction; learning (artificial intelligence); virtual reality; solid modelling; user interfaces; neurophysiology; electroencephalography; human computer interaction; augmented reality; medical signal processing.
Ba Dung Le, Mehwish Nasim, Muhammad Ali Babar 0001
CW3
2017 Roman-txt: forms and functions of roman urdu texting
abstract
In this paper, we present a user study conducted on students of a local university in Pakistan and collected a corpus of Roman Urdu text messages. We were interested in forms and functions of Roman Urdu text messages. To this end, we collected a mobile phone usage dataset. The data consists of 116 users and 346, 455 text messages. Roman Urdu text, is the most widely adopted style of writing text messages in Pakistan. Our user study leads to interesting results, for instance, we were able to quantitatively show that a number of words are written using more than one spelling; most participants of our study were not comfortable in English and hence they write their text messages in Roman Urdu; and the choice of language adopted by the participants sometimes varies according to who the message is being sent. Moreover we found that many young students send text messages(SMS) of intimate nature.
Anas Bilal 0001, Aimal Rextin, Ahmad Kakakhel, Mehwish Nasim
MobileHCI4
2017 Data analysis and call prediction on dyadic data from an understudied population
Mehwish Nasim, Aimal Rextin, Shamaila Hayat, Numair Khan, Muhammad Muddassir Malik
Pervasive Mob. Comput.1
2016 Understanding call logs of smartphone users for making future calls
abstract
In this measurement study, we analyze whether mobile phone users exhibit temporal regularity in their mobile communication. To this end, we collected a mobile phone usage dataset from a developing country -- Pakistan. The data consists of 783 users and 229, 450 communication events. We found a number of interesting patterns both at the aggregate level and at dyadic level in the data. Some interesting results include: the number of calls to different alters consistently follow the rank-size rule; a communication event between an ego-alter(user-contact) pair greatly increases the chances of another communication event; certain ego-alter pairs tend to communicate more over weekends; ego-alter pairs exhibit autocorrelation in various time quantum. Identifying such idiosyncrasies in the ego-alter communication can help improve the calling experience of smartphone users by automatically (smartly) sorting the call log without any manual intervention.
Mehwish Nasim, Aimal Rextin, Numair Khan, Muhammad Muddassir Malik
MobileHCI1
2016 Investigating Link Inference in Partially Observable Networks: Friendship Ties and Interaction
abstract
While privacy preserving mechanisms, such as hiding one's friends list, may be available to withhold personal information on online social networking sites, it is not obvious whether to which degree a user's social behavior renders such an attempt futile. In this paper, we study the impact of additional interaction information on the inference of links between nodes in partially covert networks. This investigation is based on the assumption that interaction might be a proxy for connectivity patterns in online social networks. For this purpose, we use data collected from 586 Facebook profiles consisting of friendship ties (conceptualized as the network) and comments on wall posts (serving as interaction information) by a total of 64 000 users. The link-inference problem is formulated as a binary classification problem using a comprehensive set of features and multiple supervised learning algorithms. Our results suggest that interactions reiterate the information contained in friendship ties sufficiently well to serve as a proxy when the majority of a network is unobserved.
Mehwish Nasim, Raphaël Charbey, Ulrik Brandes
IEEE Trans. Comput. Soc. Syst.1
2011 Cooperative Communication for Energy Efficiency in Mobile Wireless Sensor Networks
Mehwish Nasim, Saad B. Qaisar
ICCSA (4)1