EDBT 2026 Demo / reviewers in the wild / expert
Rakesh M. Verma
dblp:v/RakeshMVerma
· DBLP profile ↗
10ranked-venue papers in the field
3as first author
5since 2021 · last 2026
0000-0002-7466-7823ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 4Other / Interdisciplinary · 3 (2 first)Database Systems & Data Management · 1Information Retrieval & Web Search · 1 (1 first)Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LeTMEMo: Leveraging Topic Modeling for Evaluating (Closed-Vocabulary) Models
Vu Minh Hoang Dang, Rakesh M. Verma |
IDA | 2 |
| 2025 | Vocabulary Quality in NLP Datasets: An Autoencoder-Based Framework Across Domains and Languages
Vu Minh Hoang Dang, Rakesh M. Verma |
IDA | 2 |
| 2024 | Data Quality in NLP: Metrics and a Comprehensive Taxonomy
Vu Minh Hoang Dang, Rakesh M. Verma |
IDA (1) | 2 |
| 2024 | Blue Sky: Multilingual, Multimodal Domain Independent Deception DetectionabstractDeception, a pervasive aspect of communication, has undergone a significant transformation in the digital age. With the globalization of online interactions, individuals are communicating in multiple languages, mixing languages on social media. A variety of data is now available in many languages, while the techniques for detecting deception are similar across the board. Recent studies have shown the possibility of the existence of universal linguistic cues to deception across domains within the English language; however, the existence of such cues in other languages remains unknown. Furthermore, the practical task of deception detection in low-resource languages is not a well-studied problem due to the lack of labeled data. Another dimension of deception is multimodality. For example, in fake news or disinformation, there may be a picture with an altered caption. This paper calls for a comprehensive investigation into the complexities of deceptive language across linguistic boundaries and modalities, and raises the possibility of use of multilingual transformer models and labeled data in a variety of languages to universally address the task of deception detection. Dainis Boumber, Rakesh M. Verma, Fatima Zahra Qachfar |
SDM | 2 |
| 2022 | Data Quality and Linguistic Cues for Domain-independent Deception DetectionabstractDeception is pervasive in today’s connected society and is being spread in a multitude of different forms with diverse goals, which we refer to as domains of deception. The most crucial research task in the field of deception is identification of deception, which in most cases involves a machine learning model making the binary classification of Deceptive or Not Deceptive. These classification models are very important as they can help protect the security of an organization by preventing phishing emails from being read, protect online retailers from being flooded with fictitious reviews, and many other tasks depending on the domain of deception they are trained to handle. There has been a fair amount of research focused on the classification of deception, however most research has focused on one domain of deception exclusively. In this work we look at the quality of multiple datasets across different domains of deception, investigate the traces that deception may leave across domains by performing multiple tests using machine learning models, as well as ascertain how using linguistic cues to identify deception performs over multiple domains. Casey Hanks, Rakesh M. Verma |
BDCAT | 2 |
| 2002 | K-tree/forest: efficient indexes for boolean queriesabstractIn Information Retrieval it is well-known that the complexity of processing boolean queries depends on the size of the intermediate results, which could be huge (and are typically on disk) even though the size of the final result may be quite small. In the case of inverted files the most time consuming operation is the merging or intersection of the list of occurrences [1]. We propose, the Keyword tree (K-tree) and forest, efficient structures to handle boolean queries in keyword-based information retrieval. Extensive simulations show that K-tree is orders-of-magnitude faster (i.e., far fewer I/O's) for boolean queries than the usual approach of merging the lists of occurrences and incurs only a small overhead for single keyword queries. The K-tree can be efficiently parallelized as well. The construction cost of K-tree is comparable to the cost of building inverted files. Rakesh M. Verma, Sanjiv Behl |
SIGIR | 1 |
| 2002 | Algorithms and reductions for rewriting problems II
Rakesh M. Verma |
Inf. Process. Lett. | 1 |
| 1997 | On Embedding Rectangular Meshes into Rectangular Meshes of Smaller Aspect Ratio
Shou-Hsuan Stephen Huang, Rakesh M. Verma |
Inf. Process. Lett. | 3 |
| 1997 | An Efficient Multiversion Access STructureabstractAn efficient multiversion access structure for a transaction-time database is presented. Our method requires optimal storage and query times for several important queries and logarithmic update times. Three version operations-inserts, updates, and deletes-are allowed on the current database, while queries are allowed on any version, present or past. The following query operations are performed in optimal query time: key range search, key history search, and time range view. The key-range query retrieves all records having keys in a specified key range at a specified time; the key history query retrieves all records with a given key in a specified time range; and the time range view query retrieves all records that were current during a specified time interval. Special cases of these queries include the key search query, which retrieves a particular version of a record, and the snapshot query which reconstructs the database at some past time. To the best of our knowledge no previous multiversion access structure simultaneously supports all these query and version operations within these time and space bounds. The bounds on query operations are worst case per operation, while those for storage space and version operations are (worst-case) amortized over a sequence of version operations. Simulation results show that good storage utilization and query performance is obtained. Peter J. Varman, Rakesh M. Verma |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1992 | Strings, Trees, and Patterns
Rakesh M. Verma |
Inf. Process. Lett. | 1 |