Rakesh M. Verma

dblp:v/RakeshMVerma · DBLP profile ↗
← Back
72ranked-venue papers
33as first author
18since 2021 · last 2026
0000-0002-7466-7823ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 28 · 18 first-authorSecurity and privacy · 18 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 10 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word Importance
abstract
Adversarial text attacks remain a persistent threat to transformer models, yet existing defenses are typically attack-specific or require costly model retraining, leaving a gap for attack-agnostic detection. We introduce Guided Perturbation Sensitivity (GPS), a detection framework that identifies adversarial examples by measuring how embedding representations change when important words are masked. GPS first ranks words using importance heuristics, then measures embedding sensitivity to masking top-k critical words, and processes the resulting patterns with a BiLSTM detector. Experiments show that adversarially perturbed words exhibit disproportionately high masking sensitivity compared to naturally important words. Across three datasets, three attack types, and two victim models, GPS achieves over 85% detection accuracy and demonstrates competitive performance compared to existing state-of-the-art methods, often at lower computational cost. Using Normalized Discounted Cumulative Gain (NDCG) to measure perturbation identification quality, we demonstrate that gradient-based ranking significantly outperforms attention, hybrid, and random selection approaches, with identification quality strongly correlating with detection performance for word-level attacks (ρ = 0.65). GPS generalizes to unseen datasets, attacks, and models without retraining, providing a practical solution for adversarial text detection.
Bryan Tuck, Rakesh M. Verma
AAAI2
2026 LeTMEMo: Leveraging Topic Modeling for Evaluating (Closed-Vocabulary) Models
Vu Minh Hoang Dang, Rakesh M. Verma
IDA2
2026 Orthographic Constraint Satisfaction and Human Difficulty Alignment in Large Language Models
abstract
Large language models must satisfy hard orthographic constraints during controlled text generation, yet systematic cross-family evaluation remains limited. We evaluate 39 configurations spanning three model families (Qwen3, Claude Haiku 4.5, GPT-5-mini) on 58 word puzzles requiring character-level constraint satisfaction. Cross-family differences produce substantially larger performance gaps (2.0-2.2x, F1 = 0.761 vs. 0.343) than parameter scaling within families (83% gain from 4B to 32B scaling), and a partial-correlation analysis rules out tokenizer design as a confound for within-family scaling. Thinking budget sensitivity proves heterogeneous: high-capacity models show strong returns (+0.102 to +0.136 F1), while mid-sized variants saturate or degrade, showing inconsistent compute benefits. Using difficulty ratings from 10,000 human solvers per puzzle, we establish modest but consistent calibration (\r{ho} = 0.28-0.42) across all families, yet identify systematic failures on common words with unusual orthography ("data", "loll", "acai": 83-91% human success, 94-98% model miss rate). These failures point to over-reliance on distributional plausibility that penalizes orthographically atypical but constraint-valid patterns.
Bryan Tuck, Rakesh M. Verma
LREC2
2025 Unmasking the Imposters: How Censorship and Domain Adaptation Affect the Detection of Machine-Generated Tweets
abstract
The rapid development of large language models (LLMs) has significantly improved the generation of fluent and convincing text, raising concerns about their potential misuse on social media platforms. We present a comprehensive methodology for creating nine Twitter datasets to examine the generative capabilities of four prominent LLMs: Llama 3, Mistral, Qwen2, and GPT4o. These datasets encompass four censored and five uncensored model configurations, including 7B and 8B parameter base-instruction models of the three open-source LLMs. Additionally, we perform a data quality analysis to assess the characteristics of textual outputs from human, “censored,” and “uncensored models,” employing semantic meaning, lexical richness, structural patterns, content characteristics, and detector performance metrics to identify differences and similarities. Our evaluation demonstrates that “uncensored” models significantly undermine the effectiveness of automated detection methods. This study addresses a critical gap by exploring smaller open-source models and the ramifications of “uncensoring,” providing valuable insights into how domain adaptation and content moderation strategies influence both the detectability and structural characteristics of machine-generated text.
Bryan Tuck, Rakesh M. Verma
COLING2
2025 Vocabulary Quality in NLP Datasets: An Autoencoder-Based Framework Across Domains and Languages
Vu Minh Hoang Dang, Rakesh M. Verma
IDA2
2025 Real-Time, Evidence-Based Alerts for Protection From Phishing Attacks
abstract
Despite two decades of research on automatic filtering systems, phishing attacks remain a serious problem. To alleviate risks from filtering failures, we design and evaluate the effectiveness of a new warning system on users’ susceptibility to phishing. Our proposed technique highlights key sentences based on an analysis of the persuasive techniques used. An online mixed-design study ($n=604$) shows that adding our highlighting technique outperforms existing warning solutions. It also identifies the relative efficacy of different appeals and the characteristics of susceptible users. Results show that adding our highlighting techniqueis useful even with false positives and false negatives. Inspired by this result, we propose an automatic warning generator. We created a small labeled dataset of suspicious sentences and used data augmentation. Our best models achieve F1 score of 99.95% in detecting phishing emails and 88% in detecting suspicious sentences.
Shahryar Baki, Fatima Zahra Qachfar, Rakesh M. Verma, Ryan Kennedy, Daniel Jones 0003
IEEE Trans. Dependable Secur. Comput.3
2024 Domain-Agnostic Adapter Architecture for Deception Detection: Extensive Evaluations with the DIFrauD Benchmark
abstract
Despite significant strides in training expansive transformer models, their deployment for niche tasks remains intricate. This paper delves into deception detection, assessing domain adaptation methodologies from a cross-domain lens using transformer Large Language Models (LLMs). We roll out a new corpus with roughly 100,000 honest and misleading statements in seven domains, designed to serve as a benchmark for multidomain deception detection. As a primary contribution, we present a novel parameter-efficient finetuning adapter, PreXIA, which was proposed and implemented as part of this work. The design is model-, domain- and task-agnostic, with broad applications that are not limited by the confines of deception or classification tasks. We comprehensively analyze and rigorously evaluate LLM tuning methods and our original design using the new benchmark, highlighting their strengths, pointing out weaknesses, and suggesting potential areas for improvement. The proposed adapter consistently outperforms all competition on the DIFrauD benchmark used in this study. To the best of our knowledge, it improves on the state-of-the-art in its class for the deception task. In addition, the evaluation process leads to unexpected findings that, at the very least, cast doubt on the conclusions made in some of the recently published research regarding reasoning ability’s unequivocal dominance over representations quality with respect to the relative contribution of each one to a model’s performance and predictions.
Dainis Boumber, Fatima Zahra Qachfar, Rakesh M. Verma
LREC/COLING3
2024 All Your LLMs Belong to Us: Experiments with a New Extortion Phishing Dataset
Fatima Zahra Qachfar, Rakesh M. Verma
DBSec2
2024 Data Quality in NLP: Metrics and a Comprehensive Taxonomy
Vu Minh Hoang Dang, Rakesh M. Verma
IDA (1)2
2024 Blue Sky: Multilingual, Multimodal Domain Independent Deception Detection
abstract
Deception, a pervasive aspect of communication, has undergone a significant transformation in the digital age. With the globalization of online interactions, individuals are communicating in multiple languages, mixing languages on social media. A variety of data is now available in many languages, while the techniques for detecting deception are similar across the board. Recent studies have shown the possibility of the existence of universal linguistic cues to deception across domains within the English language; however, the existence of such cues in other languages remains unknown. Furthermore, the practical task of deception detection in low-resource languages is not a well-studied problem due to the lack of labeled data. Another dimension of deception is multimodality. For example, in fake news or disinformation, there may be a picture with an altered caption. This paper calls for a comprehensive investigation into the complexities of deceptive language across linguistic boundaries and modalities, and raises the possibility of use of multilingual transformer models and labeled data in a variety of languages to universally address the task of deception detection.
Dainis Boumber, Rakesh M. Verma, Fatima Zahra Qachfar
SDM2
2023 Enhancement of Twitter event detection using news streams
abstract
Abstract A new framework for improving event detection is proposed that employs joint information in news media content and social networks, such as Twitter, to leverage detailed coverage of news media and the timeliness of social media. Specifically, a short text clustering method is employed to detect events from tweets, then the language model representations of the detected events are expanded using another set of events obtained from news articles published simultaneously. The expanded representations of events are employed as a new initialization of the clustering method to run another iteration and consequently enhance the event detection results. The proposed framework is evaluated using two datasets: a tweet dataset with event labels and a news dataset containing news articles published during the same time interval as the tweets. Experimental results show that the proposed framework improves the event detection results in terms of F 1 measure compared to the results obtained from tweets only.
Samaneh Karimi, Azadeh Shakery, Rakesh M. Verma
Nat. Lang. Eng.3
2023 Sixteen Years of Phishing User Studies: What Have We Learned?
abstract
Several previous studies have investigated user susceptibility to phishing attacks. A thorough meta-analysis or systematic review is required to gain a better understanding of these findings and to assess the strength of evidence for phishing susceptibility of a subpopulation, e.g., older users. We aim to determine whether an effect exists; another aim is to determine whether the effect is positive or negative and to obtain a single summary estimate of the effect.OBJECTIVES:We systematically review the results of previous user studies on phishing susceptibility and conduct a meta-analysis.METHOD:We searched four online databases for English studies on phishing. We included all user studies in phishing detection and prevention, whether they proposed new training techniques or analyzed users’ vulnerability.FINDINGS:A careful analysis reveals some discrepancies between the findings. More than half of the studies that analyzed the effect ofagereported no statistically significant relationship between age and users’ performance. Some studies reported older people performed better while some reported the opposite. A similar finding holds for the gender difference. The meta-analysis shows: 1) a significant relationship between participants’ age and their susceptibility 2) females are more susceptible than males 3) users training significantly improves their detection ability.
Shahryar Baki, Rakesh M. Verma
IEEE Trans. Dependable Secur. Comput.2
2022 Vulnerability Detection via Multimodal Learning: Datasets and Analysis
abstract
A vulnerability is a weakness that can be exploited by an attacker, e.g., performing unauthorized actions within a computer system. For example, privilege escalation is a type of vulnerability in software, which can be used to gain elevated access to resources that are normally protected from an application or user. However, most applications contain vulnerabilities, some are fixed over time by patches, but many are discovered only after exploitation, which results in steep costs. Furthermore, program analysis tools are generally quite difficult to use. Security analysts still do manual investigation on software, i.e., using static analysis tools on machine code or source code to find bugs. Multimodal learning has been widely used in image processing, but is rarely seen in software security. We introduce a new dataset for multimodal deep learning, MVDSC-C (Multisource for Vulnerability Detection in Source Code - C/C++). Our preliminary results show that combined modalities perform better than single modalities.
Rakesh M. Verma
AsiaCCS2
2022 Data Quality and Linguistic Cues for Domain-independent Deception Detection
abstract
Deception is pervasive in today’s connected society and is being spread in a multitude of different forms with diverse goals, which we refer to as domains of deception. The most crucial research task in the field of deception is identification of deception, which in most cases involves a machine learning model making the binary classification of Deceptive or Not Deceptive. These classification models are very important as they can help protect the security of an organization by preventing phishing emails from being read, protect online retailers from being flooded with fictitious reviews, and many other tasks depending on the domain of deception they are trained to handle. There has been a fair amount of research focused on the classification of deception, however most research has focused on one domain of deception exclusively. In this work we look at the quality of multiple datasets across different domains of deception, investigate the traces that deception may leave across domains by performing multiple tests using machine learning models, as well as ascertain how using linguistic cues to identify deception performs over multiple domains.
Casey Hanks, Rakesh M. Verma
BDCAT2
2022 Leveraging Synthetic Data and PU Learning For Phishing Email Detection
abstract
Imbalanced data classification has always been one of the most challenging problems in data science especially in the cybersecurity field, where we observe an out-of-balance proportion between benign and phishing examples in security datasets. Even though there are many phishing detection methods in literature, most of them neglect the imbalanced nature of phishing email datasets. In this paper, we examine the imbalanced property by varying legitimate to phishing class ratios. We generate new synthetic instances using a generative adversarial network model for long sentences (LeakGAN) to balance out the training process and ameliorate its impact on classification. These synthetic instances are labeled by positive-unlabeled learning and added to the initial imbalanced training set. The resulting dataset is given to the Bidirectional Encoder Representations from Transformers (BERT) model for sequence classification. We compare several state-of-the-art methods from the literature against our approach, which achieves a high performance throughout all the imbalanced ratios reaching an F1-score of 99.6% for the most extreme imbalanced ratio and an F1-score of 99.8% for balanced cases.
Fatima Zahra Qachfar, Rakesh M. Verma, Arjun Mukherjee
CODASPY2
2022 Does Deception Leave a Content Independent Stylistic Trace?
abstract
A recent survey claims that there are \em no general linguistic cues for deception. Since Internet societies are plagued with deceptive attacks such as phishing and fake news, this claim means that we must build individual datasets and detectors for each kind of attack. It also implies that when a new scam (e.g., Covid) arrives, we must start the whole process of data collection, annotation, and model building from scratch. In this paper, we put this claim to the test by building a quality domain-independent deception dataset and investigating whether a model can perform well on more than one form of deception.
Victor Zeng, Xuting Liu 0003, Rakesh M. Verma
CODASPY3
2021 Capacity Expansion in Cybersecurity: Challenges and Prospects
abstract
We focus on the challenges and prospects for cybersecurity capacity expansion via faculty training. We discuss preliminary results from our efforts in this direction in a recent NSF-funded capacity expansion project. We wrap up with some questions
Rakesh M. Verma
SIGCSE1
2021 What is the Security Mindset? Can it be Developed?
abstract
With the proliferation of information technology (IT) in the world, information security has become an important objective. In a networked IT world where transactions are conducted through extensive use of the internet, cybersecurity has become integral with information security. Security is one of the most exciting computing fields [1], since it has something that no other computer science field has, an adversary. The same element also makes it one of the most challenging fields because of the unpredictability and the creativity/imagination required. The dearth of faculty who can teach the theory and practice of security is a serious impediment to offering formal education in this domain. While virtual world environments, game-based learning, cyberwars, and other learning-focused interactions have the potential to introduce security concepts and thinking skills in an engaging way to students and professionals, there is also a need to inculcate a 'security mindset,' [1]. This session will explore the concept of the security mindset including what it is, what tools security educators need to help inculcate it, and how it can be developed and/or taught in the spirit of 'training the trainers.'
Rakesh M. Verma, Rangarajan Ray Parthasarathy, Lila Ghemri
SIGCSE1
2020 Scam Augmentation and Customization: Identifying Vulnerable Users and Arming Defenders
abstract
Why do "classical" attacks such as phishing, IRS scams, etc., still succeed? How do attackers increase their chances of success? How do people reason about scams and frauds they face daily? More research is needed on these questions, which is the focus of this paper. We take a well-known attack, viz. company representative fraud, and study several parameters that bear on its effectiveness with a between-subjects study. We also study the effectiveness of a coherent language generation technique in producing phishing emails. We give ample room for the participants to demonstrate their reasoning and strategies.
Shahryar Baki, Rakesh M. Verma, Omprakash Gnawali
AsiaCCS2
2020 PhishBench 2.0: A Versatile and Extendable Benchmarking Framework for Phishing
abstract
We describe version 2.0 of our benchmarking framework, PhishBench. With the addition of the ability to dynamically load features, metrics, and classifiers, our new and improved framework allows researchers to rapidly evaluate new features and methods for machine-learning based phishing detection. Researchers can compare under identical circumstances their contributions with numerous built-in features, ranking methods, and classifiers used in the literature with the right evaluation metrics. We will demonstrate PhishBench 2.0 and compare it against at least two other automated ML systems.
Victor Zeng, Shahryar Baki, Rakesh M. Verma
CCS4
2020 Developing A Compelling Vision for Winning the Cybersecurity Arms Race
abstract
In cybersecurity there is a continuous arms race between the attackers and the defenders. In this panel, we investigate three key questions regarding this arms race. First question is whether this arms race is winnable. Second, if the answer to the first question is in the affirmative, what steps we need to take to win this race. Third, if the answer to the first question is negative, what is the justification for this and what steps can we take to improve the state of affairs and increase the bar for the attackers significantly.
Elisa Bertino, Anoop Singhal, Srivathsan Srinivasagopalan, Rakesh M. Verma
CODASPY4
2020 Poster: A Modular and Innovative Security Analytics Course
abstract
Techniques from data science are increasingly being applied by researchers to security challenges. However, requirements unique to the security domain necessitate painstaking care for the models to be valid and robust. In this paper, we outline a novel security analytics course, its modular design, some of its key innovations, and experience in teaching it.
Rakesh M. Verma
SIGCSE1
2019 Data Quality for Security Challenges: Case Studies of Phishing, Malware and Intrusion Detection Datasets
abstract
Techniques from data science are increasingly being applied by researchers to security challenges. However, challenges unique to the security domain necessitate painstaking care for the models to be valid and robust. In this paper, we explain key dimensions of data quality relevant for security, illustrate them with several popular datasets for phishing, intrusion detection and malware, indicate operational methods for assuring data quality and seek to inspire the audience to generate high quality datasets for security challenges.
Rakesh M. Verma, Victor Zeng, Houtan Faridi
CCS1
2019 Parameter Tuning and Confidence Limits of Malware Clustering
abstract
The growing number of new malware and the sophisticated obfuscation techniques used by malware authors are causing major problems in identifying, managing, and releasing anti-malware products to the consumers. Clustering malware variants based on their behavior has the potential to ease this problem of scale and conveniently lend itself to better, faster, and efficient prioritization of malware analysis. In this paper, we cluster real-world malware and expand on commonly used algorithms through fine grained testing. Results of top performing algorithms are discussed.
Houtan Faridi, Srivathsan Srinivasagopalan, Rakesh M. Verma
CODASPY3
2017 Scaling and Effectiveness of Email Masquerade Attacks: Exploiting Natural Language Generation
abstract
We focus on email-based attacks, a rich field with well-publicized consequences. We show how current Natural Language Generation (NLG) technology allows an attacker to generate masquerade attacks on scale, and study their effectiveness with a within-subjects study. We also gather insights on what parts of an email do users focus on and how users identify attacks in this realm, by planting signals and also by asking them for their reasoning. We find that: (i) 17% of participants could not identify any of the signals that were inserted in emails, and (ii) Participants were unable to perform better than random guessing on these attacks. The insights gathered and the tools and techniques employed could help defenders in: (i) implementing new, customized anti-phishing solutions for Internet users including training next-generation email filters that go beyond vanilla spam filters and capable of addressing masquerade, (ii) more effectively training and upgrading the skills of email users, and (iii) understanding the dynamics of this novel attack and its ability of tricking humans.
Shahryar Baki, Rakesh M. Verma, Arjun Mukherjee, Omprakash Gnawali
AsiaCCS2
2017 Comprehensive Method for Detecting Phishing EmailsUsing Correlation-based Analysis and User Participation
abstract
Phishing email has become a popular solution among attackers to steal all kinds of data from people and easily breach organizations' security. Hackers use multiple techniques and tricks to raise the chances of success of their attacks, like using information found on social networking websites to tailor their emails to the target's interests, or targeting employees of an organization who probably can't spot a phishing email or malicious websites and avoid sending emails to IT people or employees from Security department. In this paper we focus on analyzing the coherence of information contained in the different parts of the email: Header, Body, and URLs. After analyzing multiple phishing emails we discovered that there is always incoherence between these different parts. We created a comprehensive method which uses a set of rules that correlates the information collected from analyzing the header, body and URLs of the email and can even include the user in the detection process. We take into account that there is no such thing called perfection, so even if an email is classified as legitimate, our system will still send a warning to the user if the email is suspicious enough. This way even if a phishing email manages to escape our system, the user can still be protected.
Rakesh M. Verma, Ayman El Aassal
CODASPY1
2017 Uniqueness of Normal Forms for Shallow Term Rewrite Systems
abstract
Uniqueness of normal forms (UN=) is an important property of term rewrite systems. UN=is decidable for ground (i.e., variable-free) systems and undecidable in general. Recently, it was shown to be decidable for linear, shallow systems. We generalize this previous result and show that this property is decidable for shallow rewrite systems, in contrast to confluence, reachability, and other related properties, which are all undecidable for flat systems. We also prove an upper bound on the complexity of our algorithm. Our decidability result is optimal in a sense, since we prove that the UN=property is undecidable for two classes of linear rewrite systems: left-flat systems in which right-hand sides are of height at most two and right-flat systems in which left-hand sides are of height at most two.
Nicholas R. Radcliffe, Luis Felipe Teixeira De Moraes, Rakesh M. Verma
ACM Trans. Comput. Log.3
2016 Mining the Web for Collocations: IR Models of Term Associations
Rakesh M. Verma, Vasanthi Vuppuluri, Arjun Mukherjee, Ghita Mammar, Shahryar Baki, Reed Armstrong
CICLing (1)1
2015 On the Character of Phishing URLs: Accurate and Robust Statistical Learning Classifiers
abstract
Phishing attacks resulted in an estimated $3.2 billion dollars worth of stolen property in 2007, and the success rate for phishing attacks is increasing each year [17]. Phishing attacks are becoming harder to detect and more elusive by using short time windows to launch attacks. In order to combat the increasing effectiveness of phishing attacks, we propose that combining statistical analysis of website URLs with machine learning techniques will give a more accurate classification of phishing URLs. Using a two-sample Kolmogorov-Smirnov test along with other features we were able to accurately classify 99.3% of our dataset, with a false positive rate of less than 0.4%. Thus, accuracy of phishing URL classification can be greatly increased through the use of these statistical measures.
Rakesh M. Verma, Keith Dyer
CODASPY1
2015 Topic based segmentation of classroom videos
abstract
Video of classroom lectures is a valuable and increasingly popular learning resource. A major weakness of the video format is the inability to quickly access the content of interest. The goal of this work is to automatically partition a lecture video into topical segments which are then presented to the user in a customized video player. The approach taken in this work is to identify topics based on text similarities across the video. The paper investigates the use of screen text extracted by Optical Character Recognition tools, as well as the speech text extracted by Automatic Speech Recognition tools. An automatic text-based segmentation algorithm is developed to identify topic changes and evaluated on a set of twenty-five lecture videos. The key conclusions are as follows. Screen text is a better guide to discovering topic changes than speech text, the effectiveness of speech text can be improved significantly with the correction of speech text, and combining screen text and accurate speech text can improve accuracy. Results are presented from surveys showing a high level of satisfaction among student users of automatically segmented videos. The paper also discusses the limits of automatic segmentation and the reasons why it is far from perfect.
Tayfun Tuna, Mahima Joshi, Varun Varghese, Rucha Deshpande, Jaspal Subhlok, Rakesh M. Verma
FIE6
2015 Phish-IDetector: Message-Id Based Automatic Phishing Detection
abstract
Phishing attacks are a well known problem in our age of electronic communication. Sensitive information like credit card details, login credentials for account, etc. are targeted by phishers. Emails are the most common channel for launching phishing attacks. They are made to resemble genuine ones as much as possible to fool recipients into divulging private and sensitive data, causing huge monetary losses every year. This paper presents a novel approach to detect phishing emails, which is simple and effective. It leverages the unique characteristics of the Message-ID field of an email header for successful detection and differentiation of phishing emails from legitimate ones. Using machine learning classifiers on n-gram features extracted from Message-IDs, we obtain over 99% detection rate with low false positives.
Rakesh M. Verma, Nirmala Rai
SECRYPT1
2013 Modeling and analysis of LEAP, a key management protocol for wireless sensor networks
abstract
A formal analysis of a key management protocol, called LEAP (Localized Encryption and Authentication Protocol), intended for wireless sensor networks is presented in this paper. LEAP is modeled using the high level formal language HLSPL and checked using the AVISPA tool for attacks on the security and authenticity of the exchanges. We focus on the protocol's establishment of pairwise keys for nearest neighbors and for multi-hop neighbors. We then use this foundation to test the protocol's method of cluster key redistribution. Finally, we check LEAP's use of μTESLA, an authentication protocol utilized a one-way key chain and delayed key disclosure, which LEAP uses for authentication of node revocation messages.
Rakesh M. Verma, Bailey E. Basile
SECON1
2012 Two-Pronged Phish Snagging
abstract
Phishing causes billions of dollars in damage every year and poses a serious threat to the Internet economy. Among the many possible communication channels, electronic mail still remains the most commonly used medium to launch phishing attacks. In this paper, we present a two dimensional approach to detecting phishing emails. We devise two independent, unsupervised classifiers, namely the link and header classifiers, and two combinations of these classifiers. We show that our schemes significantly outperform the previous unsupervised and supervised phishing detection schemes for emails in the literature. We also utilize contextual information, when available, to detect phishing. Finally, our protocol is designed to detect phishing at the email level rather than detecting fraudulent, masqueraded websites. Our implementation framework called PhishSnag, operates between a user's mail transfer agent (MTA) and mail user agent (MUA) and processes each arriving email for phishing attacks even before reaching the inbox.
Rakesh M. Verma, Narasimha K. Shashidhar, Nabil Hossain
ARES1
2012 Combining Syntax and Semantics for Automatic Extractive Single-Document Summarization
Araly Barrera, Rakesh M. Verma
CICLing (2)2
2012 Detecting Phishing Emails the Natural Language Way
Rakesh M. Verma, Narasimha K. Shashidhar, Nabil Hossain
ESORICS1
2010 Uniqueness of Normal Forms is Decidable for Shallow Term Rewrite Systems
abstract
Uniqueness of normal forms (UN=) is an important property of term rewrite systems. UN= is decidable for ground (i.e., variable-free) systems and undecidable in general. Recently it was shown to be decidable for linear, shallow systems. We generalize this previous result and show that this property is decidable for shallow rewrite systems, in contrast to confluence, reachability and other properties, which are all undecidable for flat systems. Our result is also optimal in some sense, since we prove that the UN= property is undecidable for two superclasses of flat systems: left-flat, left-linear systems in which right-hand sides are of depth at most two and right-flat, right-linear systems in which left-hand sides are of depth at most two.
Nicholas R. Radcliffe, Rakesh M. Verma
FSTTCS2
2009 Complexity of Normal Form Properties and Reductions for Term Rewriting Problems Complexity of Normal Form Properties and Reductions for Term Rewriting Problems
abstract
We present several new and some significantly improved polynomial-time reductions between basic decision problems of term rewriting systems. We prove two theorems that imply tighter upper bounds for deciding the uniqueness of normal forms (UN $^{=}$ ) and unique normalization (UN $^{→}$ ) properties under certain conditions. From these theorems we derive a new and simpler polynomial-time algorithm for the UN $^{=}$ property of ground rewrite systems, and explicit upper bounds for both UN $^{=}$ and UN $^{→}$ properties of left-linear right-ground systems. We also show that both properties are undecidable for right-ground systems. It was already known that these properties are undecidable for linear systems. Hence, in a sense the decidability results are "close" to optimal.
Rakesh M. Verma
Fundam. Informaticae1
2008 Improving Techniques for Proving Undecidability of Checking Cryptographic Protocols
abstract
Existing undecidability proofs of checking secrecy of cryptographic protocols have the limitations of not considering protocols common in literature, which are in the form of communication sequences, since only protocols as non- matching roles are considered, and not considering an attacker who is an insider since only an outsider attacker is considered. Therefore the complexity of checking the realistic attacks, such as the attack to the public key Needham-Schroeder protocol, is unknown. The limitations have been observed independently and described similarly by Froschle in a recently published paper, where two open problems are posted. This paper investigates these limitations, and we present a generally applicable approach by reductions with novel features from the reachability problem of 2-counter machines, and we solve the two open problems. We also prove the undecidability of checking authentication which is the first detailed proof to our best knowledge. A unique feature of the proof is to directly address the secrecy and authentication goals as defined for the public key Needham-Schroeder protocol, whose attack has motivated many researches of formal verification of security protocols.
Zhiyao Liang, Rakesh M. Verma
ARES2
2006 A Query-Based Medical Information Summarization System Using Ontology Knowledge
abstract
As huge amounts of knowledge are created rapidly, effective information access becomes an important issue. Especially for critical domains, such as medical and financial areas, efficient retrieval of concise and relevant information is highly desired. In this paper we propose a new user query based text summarization technique that makes use of unified medical language system, an ontology knowledge source from National Library of Medicine. We compare our method with keyword-only approach, and our ontology-based method performs clearly better. Our method also shows potential to be used in other information retrieval areas
Ping Chen 0001, Rakesh M. Verma
CBMS2
2006 Automata theory: its relevance to computer science students and course contents
abstract
No abstract available.
Michal Armoni, Susan H. Rodger, Moshe Y. Vardi, Rakesh M. Verma
SIGCSE4
2005 A visual and interactive automata theory course emphasizing breadth of automata
abstract
Teaching Theory of Computation and learning it are both challenging tasks. Moreover, students are not sufficiently interested/motivated to learn this material since: (i) they believe that the material is dated and of little use and (ii) it is too abstract and difficult. To counter the first perception, we have developed materials to illustrate the breadth of finite automata concepts. To overcome the second problem we have: enhanced and integrated visualization software and historical background into newly-devloped materials including homeworks and slides for lectures. Most of the materials are available at a web site for the course that we developed. Our preliminary experience is positive overall, but there are some remaining concerns.
Rakesh M. Verma
ITiCSE1
2005 A new decidability technique for ground term rewriting systems with applications
abstract
Programming language interpreters, proving equations (e.g. x 3 = x implies the ring is Abelian), abstract data types, program transformation and optimization, and even computation itself (e.g., turing machine) can all be specified by a set of rules, called a rewrite system. Two fundamental properties of a rewrite system are the confluence or Church--Rosser property and the unique normalization property. In this article, we develop a standard form for ground rewrite systems and the concept of standard rewriting. These concepts are then used to: prove a pumping lemma for them, and to derive a new and direct decidability technique for decision problems of ground rewrite systems. To illustrate the usefulness of these concepts, we apply them to prove: (i) polynomial size bounds for witnesses to violations of unique normalization and confluence for ground rewrite systems containing unary symbols and constants, and (ii) polynomial height bounds for witnesses to violations of unique normalization and confluence for arbitrary ground systems. Apart from the fact that our technique is direct in contrast to previous decidability results for both problems, which were indirectly obtained using tree automata techniques, this approach also yields tighter bounds for rewrite systems with unary symbols than the ones that can be derived with the indirect approach. Finally, as part of our results, we give a polynomial-time algorithm for checking whether a rewrite system has the unique normalization property for all subterms in the rules of the system.
Rakesh M. Verma, Ara Hayrapetyan
ACM Trans. Comput. Log.1
2004 Deciding confluence of certain term rewriting systems in polynomial time
abstract
We present a characterization of confluence for term rewriting systems, which is then refined for special classes of rewriting systems. The refined characterization is used to obtain a polynomial time algorithm for deciding the confluence of ground term rewrite systems. The same approach also shows the decidability of confluence for shallow and linear term rewriting systems. The decision procedure has a polynomial time complexity under the assumption that the maximum arity of a function symbol in the signature is a constant.
Guillem Godoy, Ashish Tiwari 0001, Rakesh M. Verma
Ann. Pure Appl. Log.3
2004 Remarks on Thatte's transformation of term rewriting systems
Bas Luttik, Pieter Hendrik Rodenburg, Rakesh M. Verma
Inf. Comput.3
2003 On the Confluence of Linear Shallow Term Rewrite Systems
Guillem Godoy, Ashish Tiwari 0001, Rakesh M. Verma
STACS3
2002 K-tree/forest: efficient indexes for boolean queries
abstract
In Information Retrieval it is well-known that the complexity of processing boolean queries depends on the size of the intermediate results, which could be huge (and are typically on disk) even though the size of the final result may be quite small. In the case of inverted files the most time consuming operation is the merging or intersection of the list of occurrences [1]. We propose, the Keyword tree (K-tree) and forest, efficient structures to handle boolean queries in keyword-based information retrieval. Extensive simulations show that K-tree is orders-of-magnitude faster (i.e., far fewer I/O's) for boolean queries than the usual approach of merging the lists of occurrences and incurs only a small overhead for single keyword queries. The K-tree can be efficiently parallelized as well. The construction cost of K-tree is comparable to the cost of building inverted files.
Rakesh M. Verma, Sanjiv Behl
SIGIR1
2002 Algorithms and reductions for rewriting problems II
Rakesh M. Verma
Inf. Process. Lett.1
2001 Local and Symbolic Bisimulation Using Tabled Constraint Logic Programming
Samik Basu 0001, Madhavan Mukund, C. R. Ramakrishnan 0001, I. V. Ramakrishnan, Rakesh M. Verma
ICLP5
2001 Algorithms and Reductions for Rewriting Problems
Rakesh M. Verma, Michaël Rusinowitch, Denis Lugiez
Fundam. Informaticae1
1999 LarrowR2: A Laboratory fro Rapid Term Graph Rewriting
Rakesh M. Verma, Shalitha Senanayake
RTA1
1999 Tight Bounds for Prefetching and Buffer Management Algorithms for Parallel I/O Systems
abstract
The I/O performance of applications in multiple-disk systems can be improved by overlapping disk accesses. This requires the use of appropriate prefetching and buffer management algorithms that ensure the most useful blocks are accessed and retained in the buffer. In this paper, we answer several fundamental questions on prefetching and buffer management for distributed-buffer parallel I/O systems. First, we derive and prove the optimality of an algorithm, P-min, that minimizes the number of parallel I/Os. Second, we analyze P-con, an algorithm that always matches its replacement decisions with those of the well-known demand-paged MIN algorithm. We show that P-con can become fully sequential in the worst case. Third, we investigate the behavior of on-line algorithms for multiple-disk prefetching and buffer management. We define and analyze P-Iru, a parallel version of the traditional LRU buffer management algorithm. Unexpectedly, we find that the competitive ratio of P-Iru is independent of the number of disks. Finally, we present the practical performance of these algorithms on randomly generated reference strings. These results confirm the conclusions derived from the analysis on worst case inputs.
Peter J. Varman, Rakesh M. Verma
IEEE Trans. Parallel Distributed Syst.2
1998 Algorithms and Reductions for Rewriting Problems
Rakesh M. Verma, Michaël Rusinowitch, Denis Lugiez
RTA1
1997 Unique Normal Forms for Nonlinear Term Rewriting Systems: Root Overlaps
Rakesh M. Verma
FCT1
1997 On Embedding Rectangular Meshes into Rectangular Meshes of Smaller Aspect Ratio
Shou-Hsuan Stephen Huang, Rakesh M. Verma
Inf. Process. Lett.3
1997 General Techniques for Analyzing Recursive Algorithms with Applications
abstract
The complexity of divide-and-conquer algorithms is often described by recurrences of various forms. In this paper, we develop general techniques and master theorems for solving several kinds of recurrences, and we give several applications of our results. In particular, almost all of the earlier work on solving the recurrences considered here is subsumed by our work. In the process of solving such recurrences, we establish interesting connections between some elegant mathematics and analysis of recurrences. Using our results and improved bipartite matching algorithms, we also improve existing bounds in the literature for several problems, viz, associative-commutative (AC) matching of linear terms, associative matching of linear terms, rooted subtree isomorphism, and rooted subgraph homeomorphism for trees.
Rakesh M. Verma
SIAM J. Comput.1
1997 An Efficient Multiversion Access STructure
abstract
An efficient multiversion access structure for a transaction-time database is presented. Our method requires optimal storage and query times for several important queries and logarithmic update times. Three version operations-inserts, updates, and deletes-are allowed on the current database, while queries are allowed on any version, present or past. The following query operations are performed in optimal query time: key range search, key history search, and time range view. The key-range query retrieves all records having keys in a specified key range at a specified time; the key history query retrieves all records with a given key in a specified time range; and the time range view query retrieves all records that were current during a specified time interval. Special cases of these queries include the key search query, which retrieves a particular version of a record, and the snapshot query which reconstructs the database at some past time. To the best of our knowledge no previous multiversion access structure simultaneously supports all these query and version operations within these time and space bounds. The bounds on query operations are worst case per operation, while those for storage space and version operations are (worst-case) amortized over a sequence of version operations. Simulation results show that good storage utilization and query performance is obtained.
Peter J. Varman, Rakesh M. Verma
IEEE Trans. Knowl. Data Eng.2
1996 Tight Bounds for Prefetching and Buffer Management Algorithms for Parallel I/O Systems
Peter J. Varman, Rakesh M. Verma
FSTTCS2
1996 A New Combinatorial Approach to Optimal Embeddings of Rectangles
Shou-Hsuan Stephen Huang, Rakesh M. Verma
Algorithmica3
1995 Unique Normal Forms and Confluence of Rewrite Systems: Persistence
Rakesh M. Verma
IJCAI1
1995 A Theory of Using History for Equational Systems with Applications
abstract
Implementation of programming language interpreters, proving theorem of the form A=B, implementation of abstract data types, and program optimization are all problems that can be reduced to the problem of finding a normal form for an expression with respect to a finite set of equations. In 1980, Chew proposed an elegant congruence closure based simplifier (CCNS) for computing with regular systems, which stores the history of it computations in a compact data structure. In 1990, Verma and Ramakrishnan showed that it can also be used for noetherian systems with no overlaps. In this paper, we develop a general theory of using CCNS for computing normal forms and present several applications. Our results are more powerful and widely applicable than earlier work. We present an independent set of postulates and prove that CCNS can be used for any system that satisfies them. (This proof is based on the notion of strong closure ). We then show that CCNS can be used for consistent convergent systems and for various kinds of priority rewrite systems. This is the first time that the applicability of CCNS has been shown for priority systems. Finally, we present a new and simpler translation scheme for converting convergent systems into effectively nonoverlapping convergent priority systems. Such a translation scheme has been proposed earlier, but we show that it is incorrect. Because CCNS requires some strong properties of the given system, our demonstration of its wide applicability is both difficult and surprising. The tension between demands imposed by CCNS and our efforts to satisfy them gives our work much general significance. Our results are partly achieved through the idea of effectively simulating “bad” systems by almost-equivalent “good” ones, partly through our theory that substantially weakens the demands, and partly through the design of a powerful and unifying reduction proof method.
Rakesh M. Verma
J. ACM1
1995 Transformations and Confluence for Rewrite Systems
Rakesh M. Verma
Theor. Comput. Sci.1
1993 On Embeddings of Rectangles into Optimal Squares
abstract
Let G = h x w be a rectangular grid and H = s x s be the optimal sqaure grid for G, i.e., the least square grid which is no less than G in size. In this paper, an embedding scheme is presented for embedding G into H such that the dilation cost is at most 6. The significance of this result is that optimal expansion is always achieved, regardless of the aspect ratio of the rectangle, while keeping the dilation constant.
Shou-Hsuan Stephen Huang, Rakesh M. Verma
ICPP (3)3
1993 Smaran: A Congruence-Closure Based System for Equational Computations
Rakesh M. Verma
RTA1
1992 Tight Complexity Bounds for Term Matching Problems
Rakesh M. Verma, I. V. Ramakrishnan
Inf. Comput.1
1992 Strings, Trees, and Patterns
Rakesh M. Verma
Inf. Process. Lett.1
1991 A Theory of Using History for Equational Systems with Applications (Extended Abstract)
abstract
A general theory of using a congruence closure based simplifier (CCNS) proposed by P. Chew (1980) for computing normal forms is developed, and several applications are presented. An independent set of postulates is given, and it is proved that CCNS can be used for any system that satisfies them. It is then shown that CCNS can be used for consistent convergent systems and for various kinds of priority rewrite systems. A simple translation scheme for converting priority systems into effectively nonoverlapping convergent systems is presented.>
Rakesh M. Verma
FOCS1
1990 Nonoblivious Normalization Algorithms for Nonlinear Rewrite Systems
Rakesh M. Verma, I. V. Ramakrishnan
ICALP1
1989 Some Complexity Theoretic Aspects of AC Rewriting
Rakesh M. Verma, I. V. Ramakrishnan
STACS1
1989 An Analysis of a Good Algorithm for the Subtree Problem, Corrected
abstract
It is shown that the proof of the main result in Reyner’s paper, similarly titled, is incorrect. Interestingly, by combining a simple modification of the algorithm with tighter analysis, one can obtain the original result with a minor improvement.
Rakesh M. Verma, Steven W. Reyner
SIAM J. Comput.1
1988 Optimal Time Bounds for Parallel Term Matching
Rakesh M. Verma, I. V. Ramakrishnan
CADE1
1987 Term Matching on Parallel Computers
R. Ramesh 0001, Rakesh M. Verma, Krishnaprasad Thirunarayan, I. V. Ramakrishnan
ICALP2
1986 An Efficient Parallel Algorithm for Term Matching
Rakesh M. Verma, Krishnaprasad Thirunarayan, I. V. Ramakrishnan
FSTTCS1