VLDB 2026 Research / reviewers in the wild / expert
Rakesh M. Verma
dblp:v/RakeshMVerma
· DBLP profile ↗
72ranked-venue papers
33as first author
18since 2021 · last 2026
0000-0002-7466-7823ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 28 · 18 first-authorSecurity and privacy · 18 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 10 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word ImportanceabstractAdversarial text attacks remain a persistent threat to transformer models, yet existing defenses are typically attack-specific or require costly model retraining, leaving a gap for attack-agnostic detection. We introduce Guided Perturbation Sensitivity (GPS), a detection framework that identifies adversarial examples by measuring how embedding representations change when important words are masked. GPS first ranks words using importance heuristics, then measures embedding sensitivity to masking top-k critical words, and processes the resulting patterns with a BiLSTM detector. Experiments show that adversarially perturbed words exhibit disproportionately high masking sensitivity compared to naturally important words. Across three datasets, three attack types, and two victim models, GPS achieves over 85% detection accuracy and demonstrates competitive performance compared to existing state-of-the-art methods, often at lower computational cost. Using Normalized Discounted Cumulative Gain (NDCG) to measure perturbation identification quality, we demonstrate that gradient-based ranking significantly outperforms attention, hybrid, and random selection approaches, with identification quality strongly correlating with detection performance for word-level attacks (ρ = 0.65). GPS generalizes to unseen datasets, attacks, and models without retraining, providing a practical solution for adversarial text detection. Bryan Tuck, Rakesh M. Verma |
AAAI | 2 |
| 2026 | LeTMEMo: Leveraging Topic Modeling for Evaluating (Closed-Vocabulary) Models
Vu Minh Hoang Dang, Rakesh M. Verma |
IDA | 2 |
| 2026 | Orthographic Constraint Satisfaction and Human Difficulty Alignment in Large Language ModelsabstractLarge language models must satisfy hard orthographic constraints during controlled text generation, yet systematic cross-family evaluation remains limited. We evaluate 39 configurations spanning three model families (Qwen3, Claude Haiku 4.5, GPT-5-mini) on 58 word puzzles requiring character-level constraint satisfaction. Cross-family differences produce substantially larger performance gaps (2.0-2.2x, F1 = 0.761 vs. 0.343) than parameter scaling within families (83% gain from 4B to 32B scaling), and a partial-correlation analysis rules out tokenizer design as a confound for within-family scaling. Thinking budget sensitivity proves heterogeneous: high-capacity models show strong returns (+0.102 to +0.136 F1), while mid-sized variants saturate or degrade, showing inconsistent compute benefits. Using difficulty ratings from 10,000 human solvers per puzzle, we establish modest but consistent calibration (\r{ho} = 0.28-0.42) across all families, yet identify systematic failures on common words with unusual orthography ("data", "loll", "acai": 83-91% human success, 94-98% model miss rate). These failures point to over-reliance on distributional plausibility that penalizes orthographically atypical but constraint-valid patterns. Bryan Tuck, Rakesh M. Verma |
LREC | 2 |
| 2025 | Unmasking the Imposters: How Censorship and Domain Adaptation Affect the Detection of Machine-Generated TweetsabstractThe rapid development of large language models (LLMs) has significantly improved the generation of fluent and convincing text, raising concerns about their potential misuse on social media platforms. We present a comprehensive methodology for creating nine Twitter datasets to examine the generative capabilities of four prominent LLMs: Llama 3, Mistral, Qwen2, and GPT4o. These datasets encompass four censored and five uncensored model configurations, including 7B and 8B parameter base-instruction models of the three open-source LLMs. Additionally, we perform a data quality analysis to assess the characteristics of textual outputs from human, “censored,” and “uncensored models,” employing semantic meaning, lexical richness, structural patterns, content characteristics, and detector performance metrics to identify differences and similarities. Our evaluation demonstrates that “uncensored” models significantly undermine the effectiveness of automated detection methods. This study addresses a critical gap by exploring smaller open-source models and the ramifications of “uncensoring,” providing valuable insights into how domain adaptation and content moderation strategies influence both the detectability and structural characteristics of machine-generated text. Bryan Tuck, Rakesh M. Verma |
COLING | 2 |
| 2025 | Vocabulary Quality in NLP Datasets: An Autoencoder-Based Framework Across Domains and Languages
Vu Minh Hoang Dang, Rakesh M. Verma |
IDA | 2 |
| 2025 | Real-Time, Evidence-Based Alerts for Protection From Phishing AttacksabstractDespite two decades of research on automatic filtering systems, phishing attacks remain a serious problem. To alleviate risks from filtering failures, we design and evaluate the effectiveness of a new warning system on users’ susceptibility to phishing. Our proposed technique highlights key sentences based on an analysis of the persuasive techniques used. An online mixed-design study ($n=604$) shows that adding our highlighting technique outperforms existing warning solutions. It also identifies the relative efficacy of different appeals and the characteristics of susceptible users. Results show that adding our highlighting techniqueis useful even with false positives and false negatives. Inspired by this result, we propose an automatic warning generator. We created a small labeled dataset of suspicious sentences and used data augmentation. Our best models achieve F1 score of 99.95% in detecting phishing emails and 88% in detecting suspicious sentences. Shahryar Baki, Fatima Zahra Qachfar, Rakesh M. Verma, Ryan Kennedy, Daniel Jones 0003 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2024 | Domain-Agnostic Adapter Architecture for Deception Detection: Extensive Evaluations with the DIFrauD BenchmarkabstractDespite significant strides in training expansive transformer models, their deployment for niche tasks remains intricate. This paper delves into deception detection, assessing domain adaptation methodologies from a cross-domain lens using transformer Large Language Models (LLMs). We roll out a new corpus with roughly 100,000 honest and misleading statements in seven domains, designed to serve as a benchmark for multidomain deception detection. As a primary contribution, we present a novel parameter-efficient finetuning adapter, PreXIA, which was proposed and implemented as part of this work. The design is model-, domain- and task-agnostic, with broad applications that are not limited by the confines of deception or classification tasks. We comprehensively analyze and rigorously evaluate LLM tuning methods and our original design using the new benchmark, highlighting their strengths, pointing out weaknesses, and suggesting potential areas for improvement. The proposed adapter consistently outperforms all competition on the DIFrauD benchmark used in this study. To the best of our knowledge, it improves on the state-of-the-art in its class for the deception task. In addition, the evaluation process leads to unexpected findings that, at the very least, cast doubt on the conclusions made in some of the recently published research regarding reasoning ability’s unequivocal dominance over representations quality with respect to the relative contribution of each one to a model’s performance and predictions. Dainis Boumber, Fatima Zahra Qachfar, Rakesh M. Verma |
LREC/COLING | 3 |
| 2024 | All Your LLMs Belong to Us: Experiments with a New Extortion Phishing Dataset
Fatima Zahra Qachfar, Rakesh M. Verma |
DBSec | 2 |
| 2024 | Data Quality in NLP: Metrics and a Comprehensive Taxonomy
Vu Minh Hoang Dang, Rakesh M. Verma |
IDA (1) | 2 |
| 2024 | Blue Sky: Multilingual, Multimodal Domain Independent Deception DetectionabstractDeception, a pervasive aspect of communication, has undergone a significant transformation in the digital age. With the globalization of online interactions, individuals are communicating in multiple languages, mixing languages on social media. A variety of data is now available in many languages, while the techniques for detecting deception are similar across the board. Recent studies have shown the possibility of the existence of universal linguistic cues to deception across domains within the English language; however, the existence of such cues in other languages remains unknown. Furthermore, the practical task of deception detection in low-resource languages is not a well-studied problem due to the lack of labeled data. Another dimension of deception is multimodality. For example, in fake news or disinformation, there may be a picture with an altered caption. This paper calls for a comprehensive investigation into the complexities of deceptive language across linguistic boundaries and modalities, and raises the possibility of use of multilingual transformer models and labeled data in a variety of languages to universally address the task of deception detection. Dainis Boumber, Rakesh M. Verma, Fatima Zahra Qachfar |
SDM | 2 |
| 2023 | Enhancement of Twitter event detection using news streamsabstractAbstract A new framework for improving event detection is proposed that employs joint information in news media content and social networks, such as Twitter, to leverage detailed coverage of news media and the timeliness of social media. Specifically, a short text clustering method is employed to detect events from tweets, then the language model representations of the detected events are expanded using another set of events obtained from news articles published simultaneously. The expanded representations of events are employed as a new initialization of the clustering method to run another iteration and consequently enhance the event detection results. The proposed framework is evaluated using two datasets: a tweet dataset with event labels and a news dataset containing news articles published during the same time interval as the tweets. Experimental results show that the proposed framework improves the event detection results in terms of F 1 measure compared to the results obtained from tweets only. Samaneh Karimi, Azadeh Shakery, Rakesh M. Verma |
Nat. Lang. Eng. | 3 |
| 2023 | Sixteen Years of Phishing User Studies: What Have We Learned?abstractSeveral previous studies have investigated user susceptibility to phishing attacks. A thorough meta-analysis or systematic review is required to gain a better understanding of these findings and to assess the strength of evidence for phishing susceptibility of a subpopulation, e.g., older users. We aim to determine whether an effect exists; another aim is to determine whether the effect is positive or negative and to obtain a single summary estimate of the effect.OBJECTIVES:We systematically review the results of previous user studies on phishing susceptibility and conduct a meta-analysis.METHOD:We searched four online databases for English studies on phishing. We included all user studies in phishing detection and prevention, whether they proposed new training techniques or analyzed users’ vulnerability.FINDINGS:A careful analysis reveals some discrepancies between the findings. More than half of the studies that analyzed the effect ofagereported no statistically significant relationship between age and users’ performance. Some studies reported older people performed better while some reported the opposite. A similar finding holds for the gender difference. The meta-analysis shows: 1) a significant relationship between participants’ age and their susceptibility 2) females are more susceptible than males 3) users training significantly improves their detection ability. Shahryar Baki, Rakesh M. Verma |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2022 | Vulnerability Detection via Multimodal Learning: Datasets and AnalysisabstractA vulnerability is a weakness that can be exploited by an attacker, e.g., performing unauthorized actions within a computer system. For example, privilege escalation is a type of vulnerability in software, which can be used to gain elevated access to resources that are normally protected from an application or user. However, most applications contain vulnerabilities, some are fixed over time by patches, but many are discovered only after exploitation, which results in steep costs. Furthermore, program analysis tools are generally quite difficult to use. Security analysts still do manual investigation on software, i.e., using static analysis tools on machine code or source code to find bugs. Multimodal learning has been widely used in image processing, but is rarely seen in software security. We introduce a new dataset for multimodal deep learning, MVDSC-C (Multisource for Vulnerability Detection in Source Code - C/C++). Our preliminary results show that combined modalities perform better than single modalities. Rakesh M. Verma |
AsiaCCS | 2 |
| 2022 | Data Quality and Linguistic Cues for Domain-independent Deception DetectionabstractDeception is pervasive in today’s connected society and is being spread in a multitude of different forms with diverse goals, which we refer to as domains of deception. The most crucial research task in the field of deception is identification of deception, which in most cases involves a machine learning model making the binary classification of Deceptive or Not Deceptive. These classification models are very important as they can help protect the security of an organization by preventing phishing emails from being read, protect online retailers from being flooded with fictitious reviews, and many other tasks depending on the domain of deception they are trained to handle. There has been a fair amount of research focused on the classification of deception, however most research has focused on one domain of deception exclusively. In this work we look at the quality of multiple datasets across different domains of deception, investigate the traces that deception may leave across domains by performing multiple tests using machine learning models, as well as ascertain how using linguistic cues to identify deception performs over multiple domains. Casey Hanks, Rakesh M. Verma |
BDCAT | 2 |
| 2022 | Leveraging Synthetic Data and PU Learning For Phishing Email DetectionabstractImbalanced data classification has always been one of the most challenging problems in data science especially in the cybersecurity field, where we observe an out-of-balance proportion between benign and phishing examples in security datasets. Even though there are many phishing detection methods in literature, most of them neglect the imbalanced nature of phishing email datasets. In this paper, we examine the imbalanced property by varying legitimate to phishing class ratios. We generate new synthetic instances using a generative adversarial network model for long sentences (LeakGAN) to balance out the training process and ameliorate its impact on classification. These synthetic instances are labeled by positive-unlabeled learning and added to the initial imbalanced training set. The resulting dataset is given to the Bidirectional Encoder Representations from Transformers (BERT) model for sequence classification. We compare several state-of-the-art methods from the literature against our approach, which achieves a high performance throughout all the imbalanced ratios reaching an F1-score of 99.6% for the most extreme imbalanced ratio and an F1-score of 99.8% for balanced cases. Fatima Zahra Qachfar, Rakesh M. Verma, Arjun Mukherjee |
CODASPY | 2 |
| 2022 | Does Deception Leave a Content Independent Stylistic Trace?abstractA recent survey claims that there are \em no general linguistic cues for deception. Since Internet societies are plagued with deceptive attacks such as phishing and fake news, this claim means that we must build individual datasets and detectors for each kind of attack. It also implies that when a new scam (e.g., Covid) arrives, we must start the whole process of data collection, annotation, and model building from scratch. In this paper, we put this claim to the test by building a quality domain-independent deception dataset and investigating whether a model can perform well on more than one form of deception. Victor Zeng, Xuting Liu 0003, Rakesh M. Verma |
CODASPY | 3 |
| 2021 | Capacity Expansion in Cybersecurity: Challenges and ProspectsabstractWe focus on the challenges and prospects for cybersecurity capacity expansion via faculty training. We discuss preliminary results from our efforts in this direction in a recent NSF-funded capacity expansion project. We wrap up with some questions Rakesh M. Verma |
SIGCSE | 1 |
| 2021 | What is the Security Mindset? Can it be Developed?abstractWith the proliferation of information technology (IT) in the world, information security has become an important objective. In a networked IT world where transactions are conducted through extensive use of the internet, cybersecurity has become integral with information security. Security is one of the most exciting computing fields [1], since it has something that no other computer science field has, an adversary. The same element also makes it one of the most challenging fields because of the unpredictability and the creativity/imagination required. The dearth of faculty who can teach the theory and practice of security is a serious impediment to offering formal education in this domain. While virtual world environments, game-based learning, cyberwars, and other learning-focused interactions have the potential to introduce security concepts and thinking skills in an engaging way to students and professionals, there is also a need to inculcate a 'security mindset,' [1]. This session will explore the concept of the security mindset including what it is, what tools security educators need to help inculcate it, and how it can be developed and/or taught in the spirit of 'training the trainers.' Rakesh M. Verma, Rangarajan Ray Parthasarathy, Lila Ghemri |
SIGCSE | 1 |
| 2020 | Scam Augmentation and Customization: Identifying Vulnerable Users and Arming DefendersabstractWhy do "classical" attacks such as phishing, IRS scams, etc., still succeed? How do attackers increase their chances of success? How do people reason about scams and frauds they face daily? More research is needed on these questions, which is the focus of this paper. We take a well-known attack, viz. company representative fraud, and study several parameters that bear on its effectiveness with a between-subjects study. We also study the effectiveness of a coherent language generation technique in producing phishing emails. We give ample room for the participants to demonstrate their reasoning and strategies. Shahryar Baki, Rakesh M. Verma, Omprakash Gnawali |
AsiaCCS | 2 |
| 2020 | PhishBench 2.0: A Versatile and Extendable Benchmarking Framework for PhishingabstractWe describe version 2.0 of our benchmarking framework, PhishBench. With the addition of the ability to dynamically load features, metrics, and classifiers, our new and improved framework allows researchers to rapidly evaluate new features and methods for machine-learning based phishing detection. Researchers can compare under identical circumstances their contributions with numerous built-in features, ranking methods, and classifiers used in the literature with the right evaluation metrics. We will demonstrate PhishBench 2.0 and compare it against at least two other automated ML systems. Victor Zeng, Shahryar Baki, Rakesh M. Verma |
CCS | 4 |
| 2020 | Developing A Compelling Vision for Winning the Cybersecurity Arms RaceabstractIn cybersecurity there is a continuous arms race between the attackers and the defenders. In this panel, we investigate three key questions regarding this arms race. First question is whether this arms race is winnable. Second, if the answer to the first question is in the affirmative, what steps we need to take to win this race. Third, if the answer to the first question is negative, what is the justification for this and what steps can we take to improve the state of affairs and increase the bar for the attackers significantly. Elisa Bertino, Anoop Singhal, Srivathsan Srinivasagopalan, Rakesh M. Verma |
CODASPY | 4 |
| 2020 | Poster: A Modular and Innovative Security Analytics CourseabstractTechniques from data science are increasingly being applied by researchers to security challenges. However, requirements unique to the security domain necessitate painstaking care for the models to be valid and robust. In this paper, we outline a novel security analytics course, its modular design, some of its key innovations, and experience in teaching it. Rakesh M. Verma |
SIGCSE | 1 |
| 2019 | Data Quality for Security Challenges: Case Studies of Phishing, Malware and Intrusion Detection DatasetsabstractTechniques from data science are increasingly being applied by researchers to security challenges. However, challenges unique to the security domain necessitate painstaking care for the models to be valid and robust. In this paper, we explain key dimensions of data quality relevant for security, illustrate them with several popular datasets for phishing, intrusion detection and malware, indicate operational methods for assuring data quality and seek to inspire the audience to generate high quality datasets for security challenges. Rakesh M. Verma, Victor Zeng, Houtan Faridi |
CCS | 1 |
| 2019 | Parameter Tuning and Confidence Limits of Malware ClusteringabstractThe growing number of new malware and the sophisticated obfuscation techniques used by malware authors are causing major problems in identifying, managing, and releasing anti-malware products to the consumers. Clustering malware variants based on their behavior has the potential to ease this problem of scale and conveniently lend itself to better, faster, and efficient prioritization of malware analysis. In this paper, we cluster real-world malware and expand on commonly used algorithms through fine grained testing. Results of top performing algorithms are discussed. Houtan Faridi, Srivathsan Srinivasagopalan, Rakesh M. Verma |
CODASPY | 3 |
| 2017 | Scaling and Effectiveness of Email Masquerade Attacks: Exploiting Natural Language GenerationabstractWe focus on email-based attacks, a rich field with well-publicized consequences. We show how current Natural Language Generation (NLG) technology allows an attacker to generate masquerade attacks on scale, and study their effectiveness with a within-subjects study. We also gather insights on what parts of an email do users focus on and how users identify attacks in this realm, by planting signals and also by asking them for their reasoning. We find that: (i) 17% of participants could not identify any of the signals that were inserted in emails, and (ii) Participants were unable to perform better than random guessing on these attacks. The insights gathered and the tools and techniques employed could help defenders in: (i) implementing new, customized anti-phishing solutions for Internet users including training next-generation email filters that go beyond vanilla spam filters and capable of addressing masquerade, (ii) more effectively training and upgrading the skills of email users, and (iii) understanding the dynamics of this novel attack and its ability of tricking humans. Shahryar Baki, Rakesh M. Verma, Arjun Mukherjee, Omprakash Gnawali |
AsiaCCS | 2 |
| 2017 | Comprehensive Method for Detecting Phishing EmailsUsing Correlation-based Analysis and User ParticipationabstractPhishing email has become a popular solution among attackers to steal all kinds of data from people and easily breach organizations' security. Hackers use multiple techniques and tricks to raise the chances of success of their attacks, like using information found on social networking websites to tailor their emails to the target's interests, or targeting employees of an organization who probably can't spot a phishing email or malicious websites and avoid sending emails to IT people or employees from Security department. In this paper we focus on analyzing the coherence of information contained in the different parts of the email: Header, Body, and URLs. After analyzing multiple phishing emails we discovered that there is always incoherence between these different parts. We created a comprehensive method which uses a set of rules that correlates the information collected from analyzing the header, body and URLs of the email and can even include the user in the detection process. We take into account that there is no such thing called perfection, so even if an email is classified as legitimate, our system will still send a warning to the user if the email is suspicious enough. This way even if a phishing email manages to escape our system, the user can still be protected. Rakesh M. Verma, Ayman El Aassal |
CODASPY | 1 |
| 2017 | Uniqueness of Normal Forms for Shallow Term Rewrite SystemsabstractUniqueness of normal forms (UN=) is an important property of term rewrite systems. UN=is decidable for ground (i.e., variable-free) systems and undecidable in general. Recently, it was shown to be decidable for linear, shallow systems. We generalize this previous result and show that this property is decidable for shallow rewrite systems, in contrast to confluence, reachability, and other related properties, which are all undecidable for flat systems. We also prove an upper bound on the complexity of our algorithm. Our decidability result is optimal in a sense, since we prove that the UN=property is undecidable for two classes of linear rewrite systems: left-flat systems in which right-hand sides are of height at most two and right-flat systems in which left-hand sides are of height at most two. Nicholas R. Radcliffe, Luis Felipe Teixeira De Moraes, Rakesh M. Verma |
ACM Trans. Comput. Log. | 3 |
| 2016 | Mining the Web for Collocations: IR Models of Term Associations
Rakesh M. Verma, Vasanthi Vuppuluri, Arjun Mukherjee, Ghita Mammar, Shahryar Baki, Reed Armstrong |
CICLing (1) | 1 |
| 2015 | On the Character of Phishing URLs: Accurate and Robust Statistical Learning ClassifiersabstractPhishing attacks resulted in an estimated $3.2 billion dollars worth of stolen property in 2007, and the success rate for phishing attacks is increasing each year [17]. Phishing attacks are becoming harder to detect and more elusive by using short time windows to launch attacks. In order to combat the increasing effectiveness of phishing attacks, we propose that combining statistical analysis of website URLs with machine learning techniques will give a more accurate classification of phishing URLs. Using a two-sample Kolmogorov-Smirnov test along with other features we were able to accurately classify 99.3% of our dataset, with a false positive rate of less than 0.4%. Thus, accuracy of phishing URL classification can be greatly increased through the use of these statistical measures. Rakesh M. Verma, Keith Dyer |
CODASPY | 1 |
| 2015 | Topic based segmentation of classroom videosabstractVideo of classroom lectures is a valuable and increasingly popular learning resource. A major weakness of the video format is the inability to quickly access the content of interest. The goal of this work is to automatically partition a lecture video into topical segments which are then presented to the user in a customized video player. The approach taken in this work is to identify topics based on text similarities across the video. The paper investigates the use of screen text extracted by Optical Character Recognition tools, as well as the speech text extracted by Automatic Speech Recognition tools. An automatic text-based segmentation algorithm is developed to identify topic changes and evaluated on a set of twenty-five lecture videos. The key conclusions are as follows. Screen text is a better guide to discovering topic changes than speech text, the effectiveness of speech text can be improved significantly with the correction of speech text, and combining screen text and accurate speech text can improve accuracy. Results are presented from surveys showing a high level of satisfaction among student users of automatically segmented videos. The paper also discusses the limits of automatic segmentation and the reasons why it is far from perfect. Tayfun Tuna, Mahima Joshi, Varun Varghese, Rucha Deshpande, Jaspal Subhlok, Rakesh M. Verma |
FIE | 6 |
| 2015 | Phish-IDetector: Message-Id Based Automatic Phishing DetectionabstractPhishing attacks are a well known problem in our age of electronic communication. Sensitive information like credit card details, login credentials for account, etc. are targeted by phishers. Emails are the most common channel for launching phishing attacks. They are made to resemble genuine ones as much as possible to fool recipients into divulging private and sensitive data, causing huge monetary losses every year. This paper presents a novel approach to detect phishing emails, which is simple and effective. It leverages the unique characteristics of the Message-ID field of an email header for successful detection and differentiation of phishing emails from legitimate ones. Using machine learning classifiers on n-gram features extracted from Message-IDs, we obtain over 99% detection rate with low false positives. Rakesh M. Verma, Nirmala Rai |
SECRYPT | 1 |
| 2013 | Modeling and analysis of LEAP, a key management protocol for wireless sensor networksabstractA formal analysis of a key management protocol, called LEAP (Localized Encryption and Authentication Protocol), intended for wireless sensor networks is presented in this paper. LEAP is modeled using the high level formal language HLSPL and checked using the AVISPA tool for attacks on the security and authenticity of the exchanges. We focus on the protocol's establishment of pairwise keys for nearest neighbors and for multi-hop neighbors. We then use this foundation to test the protocol's method of cluster key redistribution. Finally, we check LEAP's use of μTESLA, an authentication protocol utilized a one-way key chain and delayed key disclosure, which LEAP uses for authentication of node revocation messages. Rakesh M. Verma, Bailey E. Basile |
SECON | 1 |
| 2012 | Two-Pronged Phish SnaggingabstractPhishing causes billions of dollars in damage every year and poses a serious threat to the Internet economy. Among the many possible communication channels, electronic mail still remains the most commonly used medium to launch phishing attacks. In this paper, we present a two dimensional approach to detecting phishing emails. We devise two independent, unsupervised classifiers, namely the link and header classifiers, and two combinations of these classifiers. We show that our schemes significantly outperform the previous unsupervised and supervised phishing detection schemes for emails in the literature. We also utilize contextual information, when available, to detect phishing. Finally, our protocol is designed to detect phishing at the email level rather than detecting fraudulent, masqueraded websites. Our implementation framework called PhishSnag, operates between a user's mail transfer agent (MTA) and mail user agent (MUA) and processes each arriving email for phishing attacks even before reaching the inbox. Rakesh M. Verma, Narasimha K. Shashidhar, Nabil Hossain |
ARES | 1 |
| 2012 | Combining Syntax and Semantics for Automatic Extractive Single-Document Summarization
Araly Barrera, Rakesh M. Verma |
CICLing (2) | 2 |
| 2012 | Detecting Phishing Emails the Natural Language Way
Rakesh M. Verma, Narasimha K. Shashidhar, Nabil Hossain |
ESORICS | 1 |
| 2010 | Uniqueness of Normal Forms is Decidable for Shallow Term Rewrite SystemsabstractUniqueness of normal forms (UN=) is an important property of term rewrite systems. UN= is decidable for ground (i.e., variable-free) systems and undecidable in general. Recently it was shown to be decidable for linear, shallow systems. We generalize this previous result and show that this property is decidable for shallow rewrite systems, in contrast to confluence, reachability and other properties, which are all undecidable for flat systems. Our result is also optimal in some sense, since we prove that the UN= property is undecidable for two superclasses of flat systems: left-flat, left-linear systems in which right-hand sides are of depth at most two and right-flat, right-linear systems in which left-hand sides are of depth at most two. Nicholas R. Radcliffe, Rakesh M. Verma |
FSTTCS | 2 |
| 2009 | Complexity of Normal Form Properties and Reductions for Term Rewriting Problems Complexity of Normal Form Properties and Reductions for Term Rewriting ProblemsabstractWe present several new and some significantly improved polynomial-time reductions between basic decision problems of term rewriting systems. We prove two theorems that imply tighter upper bounds for deciding the uniqueness of normal forms (UN $^{=}$ ) and unique normalization (UN $^{→}$ ) properties under certain conditions. From these theorems we derive a new and simpler polynomial-time algorithm for the UN $^{=}$ property of ground rewrite systems, and explicit upper bounds for both UN $^{=}$ and UN $^{→}$ properties of left-linear right-ground systems. We also show that both properties are undecidable for right-ground systems. It was already known that these properties are undecidable for linear systems. Hence, in a sense the decidability results are "close" to optimal. Rakesh M. Verma |
Fundam. Informaticae | 1 |
| 2008 | Improving Techniques for Proving Undecidability of Checking Cryptographic ProtocolsabstractExisting undecidability proofs of checking secrecy of cryptographic protocols have the limitations of not considering protocols common in literature, which are in the form of communication sequences, since only protocols as non- matching roles are considered, and not considering an attacker who is an insider since only an outsider attacker is considered. Therefore the complexity of checking the realistic attacks, such as the attack to the public key Needham-Schroeder protocol, is unknown. The limitations have been observed independently and described similarly by Froschle in a recently published paper, where two open problems are posted. This paper investigates these limitations, and we present a generally applicable approach by reductions with novel features from the reachability problem of 2-counter machines, and we solve the two open problems. We also prove the undecidability of checking authentication which is the first detailed proof to our best knowledge. A unique feature of the proof is to directly address the secrecy and authentication goals as defined for the public key Needham-Schroeder protocol, whose attack has motivated many researches of formal verification of security protocols. Zhiyao Liang, Rakesh M. Verma |
ARES | 2 |
| 2006 | A Query-Based Medical Information Summarization System Using Ontology KnowledgeabstractAs huge amounts of knowledge are created rapidly, effective information access becomes an important issue. Especially for critical domains, such as medical and financial areas, efficient retrieval of concise and relevant information is highly desired. In this paper we propose a new user query based text summarization technique that makes use of unified medical language system, an ontology knowledge source from National Library of Medicine. We compare our method with keyword-only approach, and our ontology-based method performs clearly better. Our method also shows potential to be used in other information retrieval areas Ping Chen 0001, Rakesh M. Verma |
CBMS | 2 |
| 2006 | Automata theory: its relevance to computer science students and course contentsabstractNo abstract available. Michal Armoni, Susan H. Rodger, Moshe Y. Vardi, Rakesh M. Verma |
SIGCSE | 4 |
| 2005 | A visual and interactive automata theory course emphasizing breadth of automataabstractTeaching Theory of Computation and learning it are both challenging tasks. Moreover, students are not sufficiently interested/motivated to learn this material since: (i) they believe that the material is dated and of little use and (ii) it is too abstract and difficult. To counter the first perception, we have developed materials to illustrate the breadth of finite automata concepts. To overcome the second problem we have: enhanced and integrated visualization software and historical background into newly-devloped materials including homeworks and slides for lectures. Most of the materials are available at a web site for the course that we developed. Our preliminary experience is positive overall, but there are some remaining concerns. Rakesh M. Verma |
ITiCSE | 1 |
| 2005 | A new decidability technique for ground term rewriting systems with applicationsabstractProgramming language interpreters, proving equations (e.g. x 3 = x implies the ring is Abelian), abstract data types, program transformation and optimization, and even computation itself (e.g., turing machine) can all be specified by a set of rules, called a rewrite system. Two fundamental properties of a rewrite system are the confluence or Church--Rosser property and the unique normalization property. In this article, we develop a standard form for ground rewrite systems and the concept of standard rewriting. These concepts are then used to: prove a pumping lemma for them, and to derive a new and direct decidability technique for decision problems of ground rewrite systems. To illustrate the usefulness of these concepts, we apply them to prove: (i) polynomial size bounds for witnesses to violations of unique normalization and confluence for ground rewrite systems containing unary symbols and constants, and (ii) polynomial height bounds for witnesses to violations of unique normalization and confluence for arbitrary ground systems. Apart from the fact that our technique is direct in contrast to previous decidability results for both problems, which were indirectly obtained using tree automata techniques, this approach also yields tighter bounds for rewrite systems with unary symbols than the ones that can be derived with the indirect approach. Finally, as part of our results, we give a polynomial-time algorithm for checking whether a rewrite system has the unique normalization property for all subterms in the rules of the system. Rakesh M. Verma, Ara Hayrapetyan |
ACM Trans. Comput. Log. | 1 |
| 2004 | Deciding confluence of certain term rewriting systems in polynomial timeabstractWe present a characterization of confluence for term rewriting systems, which is then refined for special classes of rewriting systems. The refined characterization is used to obtain a polynomial time algorithm for deciding the confluence of ground term rewrite systems. The same approach also shows the decidability of confluence for shallow and linear term rewriting systems. The decision procedure has a polynomial time complexity under the assumption that the maximum arity of a function symbol in the signature is a constant. Guillem Godoy, Ashish Tiwari 0001, Rakesh M. Verma |
Ann. Pure Appl. Log. | 3 |
| 2004 | Remarks on Thatte's transformation of term rewriting systems
Bas Luttik, Pieter Hendrik Rodenburg, Rakesh M. Verma |
Inf. Comput. | 3 |
| 2003 | On the Confluence of Linear Shallow Term Rewrite Systems
Guillem Godoy, Ashish Tiwari 0001, Rakesh M. Verma |
STACS | 3 |
| 2002 | K-tree/forest: efficient indexes for boolean queriesabstractIn Information Retrieval it is well-known that the complexity of processing boolean queries depends on the size of the intermediate results, which could be huge (and are typically on disk) even though the size of the final result may be quite small. In the case of inverted files the most time consuming operation is the merging or intersection of the list of occurrences [1]. We propose, the Keyword tree (K-tree) and forest, efficient structures to handle boolean queries in keyword-based information retrieval. Extensive simulations show that K-tree is orders-of-magnitude faster (i.e., far fewer I/O's) for boolean queries than the usual approach of merging the lists of occurrences and incurs only a small overhead for single keyword queries. The K-tree can be efficiently parallelized as well. The construction cost of K-tree is comparable to the cost of building inverted files. Rakesh M. Verma, Sanjiv Behl |
SIGIR | 1 |
| 2002 | Algorithms and reductions for rewriting problems II
Rakesh M. Verma |
Inf. Process. Lett. | 1 |
| 2001 | Local and Symbolic Bisimulation Using Tabled Constraint Logic Programming
Samik Basu 0001, Madhavan Mukund, C. R. Ramakrishnan 0001, I. V. Ramakrishnan, Rakesh M. Verma |
ICLP | 5 |
| 2001 | Algorithms and Reductions for Rewriting Problems
Rakesh M. Verma, Michaël Rusinowitch, Denis Lugiez |
Fundam. Informaticae | 1 |
| 1999 | LarrowR2: A Laboratory fro Rapid Term Graph Rewriting
Rakesh M. Verma, Shalitha Senanayake |
RTA | 1 |
| 1999 | Tight Bounds for Prefetching and Buffer Management Algorithms for Parallel I/O SystemsabstractThe I/O performance of applications in multiple-disk systems can be improved by overlapping disk accesses. This requires the use of appropriate prefetching and buffer management algorithms that ensure the most useful blocks are accessed and retained in the buffer. In this paper, we answer several fundamental questions on prefetching and buffer management for distributed-buffer parallel I/O systems. First, we derive and prove the optimality of an algorithm, P-min, that minimizes the number of parallel I/Os. Second, we analyze P-con, an algorithm that always matches its replacement decisions with those of the well-known demand-paged MIN algorithm. We show that P-con can become fully sequential in the worst case. Third, we investigate the behavior of on-line algorithms for multiple-disk prefetching and buffer management. We define and analyze P-Iru, a parallel version of the traditional LRU buffer management algorithm. Unexpectedly, we find that the competitive ratio of P-Iru is independent of the number of disks. Finally, we present the practical performance of these algorithms on randomly generated reference strings. These results confirm the conclusions derived from the analysis on worst case inputs. Peter J. Varman, Rakesh M. Verma |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1998 | Algorithms and Reductions for Rewriting Problems
Rakesh M. Verma, Michaël Rusinowitch, Denis Lugiez |
RTA | 1 |
| 1997 | Unique Normal Forms for Nonlinear Term Rewriting Systems: Root Overlaps
Rakesh M. Verma |
FCT | 1 |
| 1997 | On Embedding Rectangular Meshes into Rectangular Meshes of Smaller Aspect Ratio
Shou-Hsuan Stephen Huang, Rakesh M. Verma |
Inf. Process. Lett. | 3 |
| 1997 | General Techniques for Analyzing Recursive Algorithms with ApplicationsabstractThe complexity of divide-and-conquer algorithms is often described by recurrences of various forms. In this paper, we develop general techniques and master theorems for solving several kinds of recurrences, and we give several applications of our results. In particular, almost all of the earlier work on solving the recurrences considered here is subsumed by our work. In the process of solving such recurrences, we establish interesting connections between some elegant mathematics and analysis of recurrences. Using our results and improved bipartite matching algorithms, we also improve existing bounds in the literature for several problems, viz, associative-commutative (AC) matching of linear terms, associative matching of linear terms, rooted subtree isomorphism, and rooted subgraph homeomorphism for trees. Rakesh M. Verma |
SIAM J. Comput. | 1 |
| 1997 | An Efficient Multiversion Access STructureabstractAn efficient multiversion access structure for a transaction-time database is presented. Our method requires optimal storage and query times for several important queries and logarithmic update times. Three version operations-inserts, updates, and deletes-are allowed on the current database, while queries are allowed on any version, present or past. The following query operations are performed in optimal query time: key range search, key history search, and time range view. The key-range query retrieves all records having keys in a specified key range at a specified time; the key history query retrieves all records with a given key in a specified time range; and the time range view query retrieves all records that were current during a specified time interval. Special cases of these queries include the key search query, which retrieves a particular version of a record, and the snapshot query which reconstructs the database at some past time. To the best of our knowledge no previous multiversion access structure simultaneously supports all these query and version operations within these time and space bounds. The bounds on query operations are worst case per operation, while those for storage space and version operations are (worst-case) amortized over a sequence of version operations. Simulation results show that good storage utilization and query performance is obtained. Peter J. Varman, Rakesh M. Verma |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1996 | Tight Bounds for Prefetching and Buffer Management Algorithms for Parallel I/O Systems
Peter J. Varman, Rakesh M. Verma |
FSTTCS | 2 |
| 1996 | A New Combinatorial Approach to Optimal Embeddings of Rectangles
Shou-Hsuan Stephen Huang, Rakesh M. Verma |
Algorithmica | 3 |
| 1995 | Unique Normal Forms and Confluence of Rewrite Systems: Persistence
Rakesh M. Verma |
IJCAI | 1 |
| 1995 | A Theory of Using History for Equational Systems with ApplicationsabstractImplementation of programming language interpreters, proving theorem of the form A=B, implementation of abstract data types, and program optimization are all problems that can be reduced to the problem of finding a normal form for an expression with respect to a finite set of equations. In 1980, Chew proposed an elegant congruence closure based simplifier (CCNS) for computing with regular systems, which stores the history of it computations in a compact data structure. In 1990, Verma and Ramakrishnan showed that it can also be used for noetherian systems with no overlaps. In this paper, we develop a general theory of using CCNS for computing normal forms and present several applications. Our results are more powerful and widely applicable than earlier work. We present an independent set of postulates and prove that CCNS can be used for any system that satisfies them. (This proof is based on the notion of strong closure ). We then show that CCNS can be used for consistent convergent systems and for various kinds of priority rewrite systems. This is the first time that the applicability of CCNS has been shown for priority systems. Finally, we present a new and simpler translation scheme for converting convergent systems into effectively nonoverlapping convergent priority systems. Such a translation scheme has been proposed earlier, but we show that it is incorrect. Because CCNS requires some strong properties of the given system, our demonstration of its wide applicability is both difficult and surprising. The tension between demands imposed by CCNS and our efforts to satisfy them gives our work much general significance. Our results are partly achieved through the idea of effectively simulating “bad” systems by almost-equivalent “good” ones, partly through our theory that substantially weakens the demands, and partly through the design of a powerful and unifying reduction proof method. Rakesh M. Verma |
J. ACM | 1 |
| 1995 | Transformations and Confluence for Rewrite Systems
Rakesh M. Verma |
Theor. Comput. Sci. | 1 |
| 1993 | On Embeddings of Rectangles into Optimal SquaresabstractLet G = h x w be a rectangular grid and H = s x s be the optimal sqaure grid for G, i.e., the least square grid which is no less than G in size. In this paper, an embedding scheme is presented for embedding G into H such that the dilation cost is at most 6. The significance of this result is that optimal expansion is always achieved, regardless of the aspect ratio of the rectangle, while keeping the dilation constant. Shou-Hsuan Stephen Huang, Rakesh M. Verma |
ICPP (3) | 3 |
| 1993 | Smaran: A Congruence-Closure Based System for Equational Computations
Rakesh M. Verma |
RTA | 1 |
| 1992 | Tight Complexity Bounds for Term Matching Problems
Rakesh M. Verma, I. V. Ramakrishnan |
Inf. Comput. | 1 |
| 1992 | Strings, Trees, and Patterns
Rakesh M. Verma |
Inf. Process. Lett. | 1 |
| 1991 | A Theory of Using History for Equational Systems with Applications (Extended Abstract)abstractA general theory of using a congruence closure based simplifier (CCNS) proposed by P. Chew (1980) for computing normal forms is developed, and several applications are presented. An independent set of postulates is given, and it is proved that CCNS can be used for any system that satisfies them. It is then shown that CCNS can be used for consistent convergent systems and for various kinds of priority rewrite systems. A simple translation scheme for converting priority systems into effectively nonoverlapping convergent systems is presented.> Rakesh M. Verma |
FOCS | 1 |
| 1990 | Nonoblivious Normalization Algorithms for Nonlinear Rewrite Systems
Rakesh M. Verma, I. V. Ramakrishnan |
ICALP | 1 |
| 1989 | Some Complexity Theoretic Aspects of AC Rewriting
Rakesh M. Verma, I. V. Ramakrishnan |
STACS | 1 |
| 1989 | An Analysis of a Good Algorithm for the Subtree Problem, CorrectedabstractIt is shown that the proof of the main result in Reyner’s paper, similarly titled, is incorrect. Interestingly, by combining a simple modification of the algorithm with tighter analysis, one can obtain the original result with a minor improvement. Rakesh M. Verma, Steven W. Reyner |
SIAM J. Comput. | 1 |
| 1988 | Optimal Time Bounds for Parallel Term Matching
Rakesh M. Verma, I. V. Ramakrishnan |
CADE | 1 |
| 1987 | Term Matching on Parallel Computers
R. Ramesh 0001, Rakesh M. Verma, Krishnaprasad Thirunarayan, I. V. Ramakrishnan |
ICALP | 2 |
| 1986 | An Efficient Parallel Algorithm for Term Matching
Rakesh M. Verma, Krishnaprasad Thirunarayan, I. V. Ramakrishnan |
FSTTCS | 1 |