VLDB 2026 Research / reviewers in the wild / expert
Mayank Kejriwal
dblp:124/2127
· DBLP profile ↗
38ranked-venue papers
24as first author
15since 2021 · last 2026
0000-0001-5988-8305ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 13 first-author · 12 since 2021Databases, data management, data science and information retrieval · 23 · 16 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 8 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 2Theory of computation · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Deployed Investigative AI Search Engine for Combating Human Trafficking at Web Scale
Mayank Kejriwal |
AAAI | 1 |
| 2025 | HealthEQKG: A Knowledge Graph and Data Model for Health Equity ResearchabstractOwing to the expense of healthcare and aging demographics, health inequity has emerged as a critical challenge in the United States, leading to significant disparities in health outcomes and unequal access to care in different communities. Understanding the complex associations between the distribution of healthcare providers, clinician characteristics, and socioeconomic conditions in communities of practice is essential to developing policies to close health inequity gaps. However, research in this domain is challenging due to the fragmentation of relevant datasets (including by the government), and the difficulty of using these datasets in a semantically unified manner. To address this challenge, we introduce HealthEQKG, an open-source knowledge graph (KG) specifically designed to support both qualitative and quantitative health equity research. Supported by a compact underlying ontology, HealthEQKG integrates two national-level (but independent) government agency datasets containing physician data and socioeconomic data, followed by data augmentation through the use of established Semantic Web resources. The complete KG contains 72,658 physicians and 28,346 Area Deprivation Indices for zip codes across the US. Through a series of fifteen competency questions, and two use-cases, we demonstrate its utility as a queryable resource for public health policymakers and researchers. Navapat Nananukul, Mayank Kejriwal |
ISWC (2) | 2 |
| 2025 | Defining and evaluating decision and composite risk in language models applied to natural language inference
Ke Shen 0003, Mayank Kejriwal |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Multipartite Entity Resolution: Motivating a K-Tuple Perspective (Student Abstract)abstractEntity Resolution (ER) is the problem of algorithmically matching records, mentions, or entries that refer to the same underlying real-world entity. Traditionally, the problem assumes (at most) two datasets, between which records need to be matched. There is considerably less research in ER when k > 2 datasets are involved. The evaluation of such multipartite ER (M-ER) is especially complex, since the usual ER metrics assume (whether implicitly or explicitly) k < 3. This paper takes the first step towards motivating a k-tuple approach for evaluating M-ER. Using standard algorithms and k-tuple versions of metrics like precision and recall, our preliminary results suggest a significant difference compared to aggregated pairwise evaluation, which would first decompose the M-ER problem into independent bipartite problems and then aggregate their metrics. Hence, M-ER may be more challenging and warrant more novel approaches than current decomposition-based pairwise approaches would suggest. Adin Aberbach, Mayank Kejriwal, Ke Shen 0003 |
AAAI | 2 |
| 2024 | Influence Role Recognition and LLM-Based Scholar Recommendation in Academic Social NetworksabstractIdentifying scholars and their relevant publications in interdisciplinary collaborations within an academic social network (ASN) can help drive new scientific knowledge discovery. This involves a challenging and time-consuming process, which requires scholar's influence role recognition in a scholar team for a given research task. In this paper, we propose a novel “ScholarInfluencer” recommendation system that: (a) uses a classification model combined with network analysis on a heterogeneous knowledge graph to recognize the scholar influencers within interdisciplinary teams of collaborators, and (b) features a large language model (LLM) to use influence role recognition results to support user queries to produce pertinent scholar and their publication recommendations. Our novel approach involves building a heterogeneous knowledge graph using diverse ASN datasets involving entities such as scholars, publications, research grants, and the relationship among these entities. We perform an evaluation of ScholarInfluencer using four widely-used ASN datasets (i.e., NSF, DBLP, Cora and CA-HepTh). Our experiment results show that our influence role recognition model outperforms the state-of-the-art models across the different datasets; especially in the case of the NSF dataset, our model outperforms by up to 13.6%. Further, we show how our recommendation model with role recognition outperforms the model without role recognition across the different datasets; especially in the case of the NSF dataset, our model outperforms by 7%. Xiyao Cheng, Lakshmi Srinivas Edara, Yuanxun Zhang, Mayank Kejriwal, Prasad Calyam |
DSAA | 4 |
| 2024 | A Semantic Search Engine for Helping Patients Find Doctors and Locations in a Large Healthcare Organization
Mayank Kejriwal, Hamid Haidarian, Min-Hsueh Chiu, Andy Xiang, Deep Shrestha, Faizan Javed |
SIGIR | 1 |
| 2023 | DeepGraph: Multi-Cluster Interactive Visualization of Complex Networks in a Learned Representation SpaceabstractVisualizing complex networks with many thousands of nodes and interesting community structure remains a challenging problem. This paper introduces DeepGraph, an off-the-shelf package that takes a network (encoded as an edge-list) as input, and uses open-source packages to interactively visualize nodes in the network in a learned representation space on a web browser. At its core, DeepGraph is powered by established unsupervised node embedding and clustering algorithms, allowing it to operate in an end-to-end fashion without requiring technical expertise or algorithmic parameters. More advanced users can 'swap' out the algorithms in a plug-and-play fashion. DeepGraph is especially designed for nontechnical sociologists and subject matter experts looking to explore the data and its community structure before formulating research questions and follow-up studies. We demonstrate the utility and generality of DeepGraph on real-world network datasets spanning domains from digital communication to social media. Mayank Kejriwal |
ASONAM | 2 |
| 2023 | A structural study of Big Tech firm-switching of inventors in the post-recession eraabstractComplex systems research and network science have recently been used to provide novel insights into economic phenomena such as patenting behavior and innovation in firms. Several studies have found that increased mobility of inventors, manifested through firm switching or transitioning, is associated with increased overall productivity. This paper proposes a novel structural study of such transitioning inventors, and the role they play in patent co-authorship networks, in a cohort of highly innovative and economically influential companies such as the five Big Tech firms (Apple, Microsoft, Google, Amazon and Meta) in the post-recession period (2010--2022). We formulate and empirically investigate three research questions using Big Tech patent data. Our results show that transitioning inventors tend to have higher degree centrality than the average Big Tech inventor, and that their removal can lead to greater network fragmentation than would be expected by chance. The rate of transition over the 12-year period of study was found to be highest between 2015--2017, suggesting that the Big Tech innovation ecosystem underwent non-trivial shifts during this time. Finally, transition was associated with higher estimated impact of co-authored patents post-transition. Mayank Kejriwal |
ASONAM | 2 |
| 2023 | Knowledge Graph-based Embedding for Connecting Scholars in Academic Social NetworksabstractIn recent years, research tasks have increasingly involved using multi-disciplinary knowledge through collaborations of scholars from multiple fields. However, identifying a team of suitable collaborators from diverse fields for a given research task is a challenging and time-consuming process. In this paper, we propose a novel “ScholarTeamFinder” model that uses knowledge graph based link prediction to identify collaborators within an academic social network (ASN) to form a research team to address a multi-disciplinary research problem. Our approach involves building a heterogeneous knowledge graph within an ASN using entities such as scholars, publications, research grants, and the relationship among these entities. Following this, we use graph-based deep learning to learn the node embedding from the knowledge graph that can be used for scholar team recommendation. More specifically, we used the classical meth-path2vec as our base graph learning algorithm and improved its performance by considering semantic meaning of entities and encoding edge embeddings in the graph. Finally, we propose a beam-search algorithm for scholar team prediction based on our model embeddings. Our evaluation of ScholarTeamFinder is performed using large ASN datasets including a unique dataset (i.e., NSF award dataset) of federal grant awards collected over the last ten years and the scholars’ publication data, as well as three other widely used datasets (i.e., APS, SCHOLAT and Gowalla). Experiment results show that our model outperforms the state-of-the-art models across the different datasets. Xiyao Cheng, Yuanxun Zhang, Harsh Joshi, Mayank Kejriwal, Prasad Calyam |
DSAA | 4 |
| 2023 | An experimental study measuring the generalization of fine-tuned language representation models across commonsense reasoning benchmarksabstractAbstract In the last 5 years, language representation models, such as BERT and GPT‐3, based on transformer neural networks, have led to enormous progress in natural language processing (NLP). One such NLP task is commonsense reasoning, where performance is usually evaluated through multiple‐choice question answering benchmarks. Till date, many such benchmarks have been proposed, and ‘leaderboards’ tracking state‐of‐the‐art performance on those benchmarks suggest that transformer‐based models are approaching human‐like performance. Because these are commonsense benchmarks, however, such a model should be expected to generalize, that is, at least in aggregate, should not exhibit excessive performance loss across independent commonsense benchmarks regardless of the specific benchmark on (the training set of) which it has been fine‐tuned. In this article, we evaluate this expectation by proposing a methodology and experimental study to measure the generalization ability of language representation models using a rigorous and intuitive metric. Using five established commonsense reasoning benchmarks, our experimental study shows that the models do not generalize well, and may be (potentially) susceptible to issues such as dataset bias. The results therefore suggest that current performance on benchmarks may be an over‐estimate, especially if we want to use such models on novel commonsense problems for which a ‘training’ dataset may not be available, for the language representation model, to fine‐tune on. Ke Shen 0003, Mayank Kejriwal |
Expert Syst. J. Knowl. Eng. | 2 |
| 2022 | Evaluating Language Representation Models on Approximately Rational Decision Making Problems
Mayank Kejriwal, Zhisheng Tang |
CogSci | 1 |
| 2022 | Transfer-based taxonomy induction over concept labels
Mayank Kejriwal, Ke Shen 0003, Chien-Chun Ni, Nicolas Torzec |
Eng. Appl. Artif. Intell. | 1 |
| 2022 | Knowledge Graphs for Social Good: An Entity-Centric Search Engine for the Human Trafficking DomainabstractWeb advertising related to Human Trafficking (HT) activity has been on the rise in recent years. Answering entity-centric questions over crawled HT Web corpora to assist investigators in the real world is an important social problem, involving many technical challenges. This paper describes a recent entity-centric knowledge graph effort that resulted in a semantic search engine to assist analysts and investigative experts in the HT domain. The overall approach takes as input a large corpus of advertisements crawled from the Web, structures it into an indexed knowledge graph, and enables investigators to satisfy their information needs by posing investigative search queries to a special-purpose semantic execution engine. We evaluated the search engine on real-world data collected from over 90,000 webpages, a significant fraction of which correlates with HT activity. Performance on four relevant categories of questions on a mean average precision metric were found to be promising, outperforming a learning-to-rank approach on three of the four categories. The prototype uses open-source components and scales to terabyte-scale corpora. Principles of the prototype have also been independently replicated, with similarly successful results. Mayank Kejriwal, Pedro A. Szekely |
IEEE Trans. Big Data | 1 |
| 2021 | Empirical Best Practices On Using Product-Specific Schema.orgabstractSchema.org has experienced high growth in recent years. Structured descriptions of products embedded in HTML pages are now not uncommon, especially on e-commerce websites. The Web Data Commons (WDC) project has extracted schema.org data at scale from webpages in the Common Crawl and made it available as an RDF `knowledge graph' at scale. The portion of this data that specifically describes products offers a golden opportunity for researchers and small companies to leverage it for analytics and downstream applications. Yet, because of the broad and expansive scope of this data, it is not evident whether the data is usable in its raw form. In this paper, we do a detailed empirical study on the product-specific schema.org data made available by WDC. Rather than simple analysis, the goal of our study is to devise an empirically grounded set of best practices for using and consuming WDC product-specific schema.org data. Our studies reveal five best practices, each of which is justified by experimental data and analysis. Mayank Kejriwal, Ravi Kiran Selvam, Chien-Chun Ni, Nicolas Torzec |
AAAI | 1 |
| 2021 | Unsupervised real-time induction and interactive visualization of taxonomies over domain-specific conceptsabstractGiven a domain-specific set of concept labels, taxonomy induction is the problem of inducing a taxonomy over the concept labels. Despite its importance in problems such as e-commerce, and some algorithmic research as a consequence, practical tools for taxonomy induction and interactive visualization do not currently exist. To be truly useful, such a tool must permit a reasonable solution in a relatively unsupervised setting, and be applicable to general subsets of concept labels. In this paper, we present an unsupervised, end-to-end taxonomy induction system for arbitrary concept-labels from the e-commerce domain. Our system only takes a simple text file as input and yields a tree-like taxonomy that can be rendered on a browser, and that a non-technical user can interact with. Important components of the system can also be customized by a technically experienced user. Mayank Kejriwal, Ke Shen 0003 |
ASONAM | 1 |
| 2020 | Locally Constructing Product Taxonomies from Scratch Using Representation LearningabstractGiven a domain-specific set of concepts, local taxonomy construction (LTC) is the problem of `locally' inducing the neighborhood of a concept (from the set of target concepts) without being given any example links. The problem, despite having practical importance, has received little research attention due to its difficulty (in contrast with link prediction, a problem that resembles it and has undergone broad study). In this paper, we present a formalism and deep empirical study on the LTC problem. In particular, we show that an innovative application of representation learning approaches from the natural language community could be adapted to tackle the problem, often quite effectively. We also present a detailed information retrieval (IR)-based methodology for evaluating these solutions on three real- world product datasets of varying sizes. To the best of our knowledge, this is the first paper to introduce the LTC problem, especially for e-commerce applications, and offer effective, nearly unsupervised, solutions, for addressing it on real-world data. Mayank Kejriwal, Ravi Kiran Selvam, Chien-Chun Ni, Nicolas Torzec |
ASONAM | 1 |
| 2019 | Low-supervision urgency detection and transfer in short crisis messagesabstractHumanitarian disasters have been on the rise in recent years due to the effects of climate change and socio-political situations such as the refugee crisis. Technology can be used to best mobilize resources such as food and water in the event of a natural disaster, by semi-automatically flagging tweets and short messages as indicating an urgent need. The problem is challenging not just because of the sparseness of data in the immediate aftermath of a disaster, but because of the varying characteristics of disasters in developing countries (making it difficult to train just one system) and the noise and quirks in social media. In this paper, we present a robust, low-supervision social media urgency system that adapts to arbitrary crises by leveraging both labeled and unlabeled data in an ensemble setting. The system is also able to adapt to new crises where an unlabeled background corpus may not be available yet by utilizing a simple and effective transfer learning methodology. Experimentally, our transfer learning and low-supervision approaches are found to outperform viable baselines with high significance on myriad disaster datasets. In this short paper, we provide details on the low-supervision approach for urgency detection on Twitter feeds. Mayank Kejriwal, Peilin Zhou |
ASONAM | 1 |
| 2019 | SAVIZ: interactive exploration and visualization of situation labeling classifiers over crisis social media dataabstractDue to climate change and the effects of geopolitical and social challenges like the refugee crisis in Europe, the world is facing an unprecedented set of humanitarian problems. According to the United Nations, there is a projected funding shortfall of more than 20 billion dollars in addressing these needs. Technology can play a vital role in mitigating this burden, especially with the advent of real-time social media and advances in areas like Natural Language Processing and machine learning. An important problem addressed by machine learning in current crisis informatics platforms is situation labeling, which can be intuitively defined as semi-automatically assigning one or more actionable labels (such as food, medicine or water) to tweets or documents from a controlled vocabulary. Despite multiple advances, current situation labeling systems are noisy and do not generalize very well to arbitrary crisis data. Consequentially, consumers of these outputs (which include humanitarian responders) are unwilling to trust these outputs without due diligence or provenance. In this paper, we demonstrate an interactive visualization platform called SAVIZ that provides non-technical first responders with such capabilities. SAVIZ is completely built using open-source technologies, can be rendered on a web browser and is backward-compatible with several pre-existing crisis intelligence platforms. We use two real-world scenarios (the 2015 earthquake in Nepal, and the unfolding Ebola crisis in Africa) to illustrate the potential of SAVIZ. Mayank Kejriwal, Peilin Zhou |
ASONAM | 1 |
| 2019 | Concept drift in bias and sensationalism detection: an experimental studyabstractDue to easy dissemination of news in social media and the Web, there has been an increasing rise of disinformation on important political issues like elections in recent years. Computational solutions for automatic bias and sensationalism detection for news articles can have tremendous impact if used in the right way. Because news is an ever-shifting domain, concept drift is an issue that must be dealt with in any real-world computational news classification system that relies on features and trained machine learning models. Yet, an empirical study of concept drift in such systems, especially popular systems released recently as open-source and used within organizations, has been lacking thus far. This short paper reports results on an empirical study specifically designed to assess concept drift, using an open-source, popular computational news classification system, on real news data crawled from the Web. We find that even a gap of two years (2017 vs. 2019) can lead to significant concept drift, a far narrower gap than observed in traditional machine learning domains, making deployment of pre-trained or openly available computational news classification models an ethically suspect issue. Mayank Kejriwal |
ASONAM | 2 |
| 2019 | Expert-Guided Entity Extraction using Expressive RulesabstractKnowledge Graph Construction (KGC) is an important problem that has many domain-specific applications, including semantic search and predictive analytics. As sophisticated KGC algorithms continue to be proposed, an important, neglected use case is to empower domain experts who do not have much technical background to construct high-fidelity, interpretable knowledge graphs. Such domain experts are a valuable source of input because of their (both formal and learned) knowledge of the domain. In this demonstration paper, we present a system that allows domain experts to construct knowledge graphs by writing sophisticated rule-based entity extractors with minimal training, using a GUI-based editor that offers a range of complex facilities. Mayank Kejriwal, Runqi Shao, Pedro A. Szekely |
SIGIR | 1 |
| 2018 | Constructing Domain-Specific Search Engines With No ProgrammingabstractWe propose a demonstration of myDIG (my Domain-specific Insight Graphs), a system that allows non-technical domain experts, including those with no programming experience, to construct a domain-specific search engine over a raw corpus of webpages. myDIG has been developed and refined over multiple years under the DARPA MEMEX program, and has undergone rigorous user testing with actual domain experts from investigative agencies like the Securities and Exchange Commission (SEC). All components of myDIG are open-source, and the product of fundamental research. Mayank Kejriwal, Pedro A. Szekely |
AAAI | 1 |
| 2018 | Always Lurking: Understanding and Mitigating Bias in Online Human Trafficking DetectionabstractWeb-based human trafficking activity has increased in recent years but it remains sparsely dispersed among escort advertisements and difficult to identify due to its often-latent nature. The use of intelligent systems to detect trafficking can thus have a direct impact on investigative resource allocation and decision-making, and, more broadly, help curb a widespread social problem. Trafficking detection involves assigning a normalized score to a set of escort advertisements crawled from the Web -- a higher score indicates a greater risk of trafficking-related (involuntary) activities. In this paper, we define and study the problem of trafficking detection and present a trafficking detection pipeline architecture developed over three years of research within the DARPA Memex program. Drawing on multi-institutional data, systems, and experiences collected during this time, we also conduct post hoc bias analyses and present a bias mitigation plan. Our findings show that, while automatic trafficking detection is an important application of AI for social good, it also provides cautionary lessons for deploying predictive machine learning algorithms without appropriate de-biasing. This ultimately led to integration of an interpretable solution into a search system that contains over 100 million advertisements and is used by over 200 law enforcement agencies to investigate leads. Kyle Hundman, Thamme Gowda, Mayank Kejriwal, Benedikt Boecking |
AIES | 3 |
| 2018 | Structured Event Entity Resolution in Humanitarian Domains
Mayank Kejriwal, Pedro A. Szekely |
ISWC (1) | 1 |
| 2017 | Neural Embeddings for Populated Geonames Locations
Mayank Kejriwal, Pedro A. Szekely |
ISWC (2) | 1 |
| 2017 | An Investigative Search Engine for the Human Trafficking Domain
Mayank Kejriwal, Pedro A. Szekely |
ISWC (2) | 1 |
| 2017 | Information Extraction in Illicit Web DomainsabstractExtracting useful entities and attribute values from illicit domains such as human trafficking is a challenging problem with the potential for widespread social impact. Such domains employ atypical language models, have 'long tails' and suffer from the problem of concept drift. In this paper, we propose a lightweight, feature-agnostic Information Extraction (IE) paradigm specifically designed for such domains. Our approach uses raw, unlabeled text from an initial corpus, and a few (12-120) seed annotations per domain-specific attribute, to learn robust IE models for unobserved pages and websites. Empirically, we demonstrate that our approach can outperform feature-centric Conditional Random Field baselines by over 18% F-Measure on five annotated sets of real-world human trafficking datasets in both low-supervision and high-supervision settings. We also show that our approach is demonstrably robust to concept drift, and can be efficiently bootstrapped even in a serial computing environment. Mayank Kejriwal, Pedro A. Szekely |
WWW | 1 |
| 2015 | Sorted Neighborhood for the Semantic WebabstractEntity Resolution (ER) concerns identifying logically equivalent entity pairs across databases. To avoid quadratic pairwise comparisons of entities, blocking methods are used. Sorted Neighborhood is an established blocking method for relational databases. It has not been applied on graph-based data models such as the Resource Description Framework (RDF). This poster presents a modular workflow for applying Sorted Neighborhood to RDF. Real-world evaluations demonstrate the workflow's utility against a popular baseline. Mayank Kejriwal |
AAAI | 1 |
| 2015 | Entity Resolution in a Big Data FrameworkabstractEntity Resolution (ER) concerns identifying logically equivalent pairs of entities that may be syntactically disparate. Although ER is a long-standing problem in the artificial intelligence community, the growth of Linked Open Data, a collection of semi-structured datasets published and inter-connected on the Web, mandates a new approach. The thesis is that building a viable Entity Resolution solution for serving Big Data needs requires simultaneously resolving challenges of automation, heterogeneity, scalability and domain independence. The dissertation aims to build such a system and evaluate it on real-world datasets published already as Linked Open Data. Mayank Kejriwal |
AAAI | 1 |
| 2015 | A pipeline for extracting and deduplicating domain-specific knowledge basesabstractBuilding a knowledge base (KB) describing domain-specific entities is an important problem in industry, examples including KBs built over companies (e.g. Dun & Bradstreet), skills (LinkedIn, CareerBuilder) and people (inome). The task involves several engineering challenges, including devising effective procedures for data extraction, aggregation and deduplication. Data extraction involves processing multiple information sources in order to extract domain-specific data instances. The extracted instances must be aggregated and deduplicated; that is, instances referring to the same underlying entity must be identified and merged. This paper describes a pipeline developed at CareerBuilder LLC for building a KB describing employers, by first extracting entities from both global, publicly available data sources (Wikipedia and Freebase) and a proprietary source (Infogroup), and then deduplicating the instances to yield an employer-specific KB. We conduct a range of pilot experiments over three independently labeled datasets sampled from the extracted KB, and comment on some lessons learned. Mayank Kejriwal, Qiaoling Liu, Ferosh Jacob, Faizan Javed |
IEEE BigData | 1 |
| 2015 | Semi-supervised Instance Matching Using Boosted Classifiers
Mayank Kejriwal, Daniel P. Miranker |
ESWC | 1 |
| 2015 | Decision-Making Bias in Instance Matching Model Selection
Mayank Kejriwal, Daniel P. Miranker |
ISWC (1) | 1 |
| 2015 | An unsupervised instance matcher for schema-free RDF data
Mayank Kejriwal, Daniel P. Miranker |
J. Web Semant. | 1 |
| 2014 | Populating Entity Name Systems for Big Data Integration
Mayank Kejriwal |
ISWC (2) | 1 |
| 2014 | Schema matching over relations, attributes, and data valuesabstractAutomatic schema matching algorithms are typically only concerned with finding attribute correspondences. However, real world data integration problems often require matchings whose arguments span all three types of elements in relational databases: relation, attribute and data value. This paper introduces the definitions and semantics of three additional correspondence types concerning both schema and data values. These correspondences cover the higher-order mappings identified in a seminal paper by Krishnamurthy, Litwin, and Kent. It is shown that these correspondences can be automatically translated to tuple generating dependencies (tgds), and thus this research is compatible with data integration applications that leverage tgds. Aibo Tian, Mayank Kejriwal, Daniel P. Miranker |
SSDBM | 2 |
| 2013 | An Unsupervised Algorithm for Learning Blocking SchemesabstractA pair wise comparison of data objects is a requisite step in many data mining applications, but has quadratic complexity. In applications such as record linkage, blocking methods may be applied to reduce the cost. That is, the data is first partitioned into a set of blocks, and pair wise comparisons computed for pairs within each block. To date, blocking methods have required the blocking scheme be given, or the provision of training data enabling supervised learning algorithms to determine a blocking scheme. In either case, a domain expert is required. This paper develops an unsupervised method for learning a blocking scheme for tabular data sets. The method is divided into two phases. First, a weakly labeled training set is generated automatically in time linear in the number of records of the entire dataset. The second phase casts blocking key discovery as a Fisher feature selection problem. The approach is compared to a state-of-the-art supervised blocking key discovery algorithm on three real-world databases and achieves favorable results. Mayank Kejriwal, Daniel P. Miranker |
ICDM | 1 |
| 2013 | Extended scaled neural predictor for improved branch predictionabstractA perceptron-based scaled neural predictor (SNP) was implemented to emphasize the most recent branch histories via the following three approaches: (1) expanding the size of tables that correspond to recent branch histories, (2) scaling the branch histories to increase the weights for the most recent histories but decrease those for the old histories, and (3) expanding most recent branch histories to the whole history path. Furthermore, hash mechanisms, and saturating value for adjusting threshold were tuned to achieve the best prediction accuracy in each case. The resulting extended SNP was tested on well-known floating point and integer benchmarks. Using the SimpleScalar 3.0 simulator, while different features have different impact depending on whether the test is floating point or integer, overall such a well-tuned predictor achieves an improved prediction rate compared to prior approaches. Mayank Kejriwal, Risto Miikkulainen |
IJCNN | 2 |
| 2012 | A framework to access handwritten information within large digitized paper collectionsabstractWe describe our efforts with the National Archives and Records Administration (NARA) to provide a form of automated search of handwritten content within large digitized document archives. With a growing push towards the digitization of paper archives there is an imminent need to develop tools capable of searching the resulting unstructured image data as data from such collections offer valuable historical records that can be mined for information pertinent to a number of fields from the geosciences to the humanities. To carry out the search, we use a Computer Vision technique called Word Spotting. A form of content based image retrieval, it avoids the still difficult task of directly recognizing the text by allowing a user to search using a query image containing handwritten text and ranking a database of images in terms of those that contain more similar looking content. In order to make this search capability available on an archive, three computationally expensive pre-processing steps are required. We describe these steps, the open source framework we have developed, and how it can be used not only on the recently released 1940 Census data containing nearly 4 million high resolution scanned forms, but also on other collections of forms. With a growing demand to digitize our wealth of paper archives we see this type of automated search as a low cost scalable alternative to the costly manual transcription that would otherwise be required. Liana Diesendruck, Luigi Marini, Rob Kooper, Mayank Kejriwal, Kenton McHenry |
eScience | 4 |
| 2012 | Digitization and search: A non-traditional use of HPCabstractAutomated search of handwritten content is a highly interesting and applicative subject, especially important today due to the public availability of large digitized document collections. We describe our efforts with the National Archives (NARA) to provide searchable access to the 1940 Census data and discuss the HPC resources needed to implement the suggested framework. Instead of trying to recognize the handwritten text, a still very difficult task, we use a content based image retrieval technique known as Word Spotting. Through this paradigm, the system is queried by the use of handwritten text images instead of ASCII text and ranked groups of similar looking images are presented to the user. A significant amount of computing power is needed to accomplish the pre-processing of the data so to make this search capability available on an archive. The required preprocessing steps and the open source framework developed are discussed focusing specifically on HPC considerations that are relevant when preparing to provide searchable access to sizeable collections, such as the US Census. Having processed the state of North Carolina from the 1930 Census using 98,000 SUs we estimate the processing of the entire country for 1940 could require up to 2.5 million SUs. The proposed framework can be used to provide an alternative to costly manual transcriptions for a variety of digitized paper archives. Liana Diesendruck, Luigi Marini, Rob Kooper, Mayank Kejriwal, Kenton McHenry |
eScience | 4 |