Gautam Siddharth Kashyap

dblp:330/2432 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0003-2140-9617ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AlignCultura: Towards Culturally Aligned Large Language Models?
abstract
Cultural alignment in Large Language Models (LLMs) is essential for producing contextually aware, respectful, and trustworthy outputs.Without it, models risk generating stereotyped, insensitive, or misleading responses that fail to reflect cultural diversity w.r.t Helpful, Harmless, and Honest (HHH) paradigm.Existing benchmarks represent early steps toward cultural alignment; yet, no benchmarks currently enables systematic evaluation of cultural alignment in line with UNESCO's 1 principles of cultural diversity w.r.t HHH paradigm.Therefore, to address this gap, we built Align-Cultura 2 , two-stage pipeline for cultural alignment.Stage I constructs CULTURAX, the HHH-English dataset grounded in the UN-ESCO cultural taxonomy, through Query Construction, which reclassifies prompts, expands underrepresented domains (or labels), and prevents data leakage with SimHash.Then, Response Generation pairs prompts with culturally grounded responses via two-stage rejection sampling.The final dataset contains 1,500 samples spanning 30 subdomains of tangible and intangible cultural forms.Stage II benchmarks CULTURAX on general-purpose models, culturally fine-tuned models, and open-weight LLMs (Qwen3-8B and DeepSeek-R1-Distill-Qwen-7B).Empirically, culturally fine-tuned models improve joint HHH by 4%-6%, reduce cultural failures by 18%, achieve 10%-12% efficiency gains, and limit leakage to 0.3%.
Gautam Siddharth Kashyap, Mark Dras, Usman Naseem
ACL (1)1
2026 They Said Memes Were Harmless - We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural References
abstract
Meme-based social abuse detection is challenging because harmful intent often relies on implicit cultural symbolism and subtle cross-modal incongruence. Prior approaches, from fusion-based methods to in-context learning with Large Vision-Language Models (LVLMs), have made progress but remain limited by three factors: i) cultural blindness (missing symbolic context), ii) boundary ambiguity (satire vs. abuse confusion), and iii) lack of interpretability (opaque model reasoning). We introduce CROSS-ALIGN+, a three-stage framework that systematically addresses these limitations: (1) Stage I mitigates cultural blindness by enriching multimodal representations with structured knowledge from ConceptNet, Wikidata, and Hatebase; (2) Stage II reduces boundary ambiguity through parameter-efficient LoRA adapters that sharpen decision boundaries; and (3) Stage III enhances interpretability by generating cascaded explanations. Extensive experiments on five benchmarks and eight LVLMs demonstrate that CROSS-ALIGN+ consistently outperforms state-of-the-art methods, achieving up to 17% relative F1 improvement while providing interpretable justifications for each decision.
Sahil Tripathi, Gautam Siddharth Kashyap, Mehwish Nasim, Jian Yang 0001, Jiechao Gao, Usman Naseem
WWW2
2026 A study of hybrid and evolutionary metaheuristics for single hidden layer feedforward neural network architecture
abstract
Training Artificial Neural Networks (ANNs) with Stochastic Gradient Descent (SGD) frequently encounters difficulties, including substantial computing expense and the risk of converging to local optima, attributable to its dependence on partial weight gradients. Therefore, this work investigates Particle Swarm Optimisation (PSO) and Genetic Algorithms (GAs) – two population-based Metaheuristic Optimisers (MHOs) – as alternatives to Stochastic Gradient Descent (SGD) to mitigate these constraints. A hybrid PSO-SGD strategy is developed to improve local search efficiency. The findings indicate that the Hybrid (PSO-SGD) technique decreases the median training MSE by 90%-95% relative to conventional GA and PSO across various network sizes (e.g. from around 0.02 to approximately 0.001 in the Sphere function). RMHC attains substantial enhancements, reducing MSE by roughly 85%-90% compared to GA. Simultaneously, RS consistently exhibits errors exceeding 0.3, signifying subpar performance. These findings underscore that hybrid and evolutionary procedures significantly improve training efficiency and accuracy compared to conventional optimisation methods and imply that the Building Block Hypothesis (BBH) may still be valid, indicating that advantageous weight structures are retained during evolutionary search.
Gautam Siddharth Kashyap, Md. Tabrez Nafis, Samar Wazir
J. Exp. Theor. Artif. Intell.1
2025 Can We Predict Your Next Move Without Breaking Your Privacy?
Arpita Soni, Sahil Tripathi, Gautam Siddharth Kashyap, Manaswi Kulahara, Mohammad Anas Azeez, Zohaib Hasan Siddiqui, Nipun Joshi, Jiechao Gao
ASONAM (2)3
2025 Too Helpful, Too Harmless, Too Honest or Just Right?
abstract
Large Language Models (LLMs) exhibit strong performance across a wide range of NLP tasks, yet aligning their outputs with the principles of Helpfulness, Harmlessness, and Honesty (HHH) remains a persistent challenge.Existing methods often optimize for individual alignment dimensions in isolation, leading to trade-offs and inconsistent behavior.While Mixture-of-Experts (MoE) architectures offer modularity, they suffer from poorly calibrated routing, limiting their effectiveness in alignment tasks.We propose TrinityX, a modular alignment framework that incorporates a Mixture of Calibrated Experts (Mo-CaE) within the Transformer architecture.Trin-ityX leverages separately trained experts for each HHH dimension, integrating their outputs through a calibrated, task-adaptive routing mechanism that combines expert signals into a unified, alignment-aware representation.Extensive experiments on three standard alignment benchmarks-Alpaca (Helpfulness), Beaver-Tails (Harmlessness), and TruthfulQA (Honesty)-demonstrate that TrinityX outperforms strong baselines, achieving relative improvements of 32.5% in win rate, 33.9% in safety score, and 28.4% in truthfulness.In addition, TrinityX reduces memory usage and inference latency by over 40% compared to prior MoEbased approaches.Ablation studies highlight the importance of calibrated routing, and crossmodel evaluations confirm TrinityX's generalization across diverse LLM backbones.
Gautam Siddharth Kashyap, Mark Dras, Usman Naseem
EMNLP1
2025 Can AI See What We Can't? Leveraging Deep Learning and Multi-Temporal Satellite Data to Revolutionize Crop Type Mapping and Yield Prediction
abstract
Precise mapping of crop types and estimating yields are important in gauging agricultural diversity and yield potential, especially in regions dominated by small-scale farming. Nevertheless, these tasks are challenging due to factors such as small field sizes, inter-cropping, and a lack of sufficient ground truth labels for certain regions. In this paper, we propose an approach that combines advanced deep learning algorithms with Sentinel-2 and MODIS satellite data for improving the accuracy of crop type mapping and yield prediction. We used datasets from the main growing season of 2017 in Kenya (Bungoma, Busia and Siaya) coupled with county level yield data from US, Argentina and Brazil spanning from 2005 to 2016. Our models (CNN, SegNet, MaskRCNN, ResNet, UNet) were evaluated on both tasks i.e., classification of crop types and predicting yields.
Gautam Siddharth Kashyap, Harsh Joshi, Manaswi Kulahara, Rajkumar Dhakar, Atul Sajjanhar, Jiechao Gao, Shahab Saquib Sohail
ICASSP1
2025 Fooling the Forgers: A Multi-Stage Framework for Audio Deepfake Detection
abstract
Audio deepfakes represent a risk to society as they can deteriorate society’s trust in any audio. In this paper, we present a novel approach for audio deepfake detection using Generative Adversarial Networks (GANs) and contrastive learning in a multi-stage detection framework. In our process, we apply the Pre-trained Models (PTM) to extract all suitable audio phonetics, speaker identity, and other spatial prosodic features or contents, which are crucial for the model. We enhance the model’s performance by utilizing a GAN data augmentation strategy in combination with HiFi-GAN. The Contrastive learning approach is then used for improving the model’s ability to discriminate real speech from fake speech. Our experiments demonstrate that this method is superior to existing methodologies in detection and robustness.
Gautam Siddharth Kashyap, Zohaib Hasan Siddiqui, Mohammad Anas Azeez, Rafiq Ali, Shantanu Kumar, Navin Kamuni, Jiechao Gao
ICASSP1
2025 MAGIC-Enhanced Keyword Prompting for Zero-Shot Audio Captioning with CLIP Models
Vijay Govindarajan, Pratik Patel, Sahil Tripathi, Md Azizul Hoque, Gautam Siddharth Kashyap
WISE (2)5
2025 Optimization of the rectangle area inside a concave polygon using PSO and Tabu Search
Gautam Siddharth Kashyap, Karan Malik, Samar Wazir, Alexander E. I. Brownlee
Soft Comput.1
2024 Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
Orchid Chetia Phukan, Gautam Siddharth Kashyap, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH2
2024 Automated Ruleset Generation for "HTTPS Everywhere": Challenges, Implementation, and Insights
abstract
This paper details the implementation of a Web crawler aimed at automating ruleset construction for “HTTPS Everywhere,” with a goal to convert HTTP URLs to secure HTTPS equivalents for enhanced communication security. Developed within a seven-month timeframe, the crawler faced challenges in verifying HTTPS support, varying based on SSL certificate existence and validity. Successful ruleset creation and testing in Firefox and Chrome, adhering to stylistic standards, demonstrated the potential for effective development. The paper explores improving productivity through alternative libraries like Scrapy and Scrapy Cloud. While certain goals, such as in-depth cryptocurrency analysis and web crawler background reading, were unmet due to time constraints, valuable insights were gained. The conclusion underscores the difficulties, successes, and promises of automating ruleset generation through web crawlers for “HTTPS Everywhere,” offering valuable recommendations for advancing web security.
Fares Alharbi, Gautam Siddharth Kashyap, Budoor Ahmad Allehyani
Int. J. Inf. Secur. Priv.2
2022 Using Machine Learning to Quantify the Multimedia Risk Due to Fuzzing
Gautam Siddharth Kashyap, Karan Malik, Samar Wazir, Rijwan Khan
Multim. Tools Appl.1