Saket S. R. Mengle

dblp:53/420 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 first-authorSecurity and privacy · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 72% Data mining · 28%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › query understanding
query classification
0.112010
Context aware query classification using dynamic query window and relationship net · SIGIR 2010
Information retrieval › retrieval-augmented generation
document chunking
0.112008
On document splitting in passage detection · SIGIR 2008
Data mining › text mining
text classification
0.112008
On document splitting in passage detection · SIGIR 2008
Information retrieval
text analysis
0.012008
On document splitting in passage detection · SIGIR 2008

Methods — techniques the papers use, named apart from their topics

hierarchical taxonomy · 0.1conditional random field · 0.1text classification · 0.1dynamic windowing · 0.1
YearPublicationVenuePosition
2010 Context aware query classification using dynamic query window and relationship net
abstract
The context of the user queries, preceding a given query, is utilized to improve the effectiveness of query classification. Earlier efforts utilize fixed number of preceding queries to derive such context information. We propose and evaluate an approach (DQW) that identifies a set of unambiguous preceding queries in a dynamically determined window to utilize in classifying an ambiguous query. Furthermore, utilizing a relationship-net (R-net) that represents relationships among known categories, we improve the classification effectiveness for those ambiguous queries whose predicted category in this relationship-net is related to the category of a query within the window. Our results indicate that the hybrid approach (DQW+R-net) statistically significantly improves the Conditional Random Field (CRF) query classification approach when static query windowing and hierarchical taxonomy are used (SQW+Tax), in terms of precision (10.8%), recall (13.2%), and F1 measure (11.9%).
Nazli Goharian, Saket S. R. Mengle
SIGIR2
2010 Detecting relationships among categories using text classification
abstract
Abstract Discovering relationships among concepts and categories is crucial in various information systems. The authors' objective was to discover such relationships among document categories. Traditionally, such relationships are represented in the form of a concept hierarchy, grouping some categories under the same parent category. Although the nature of hierarchy supports the identification of categories that may share the same parent, not all of these categories have a relationship with each other—other than sharing the same parent. However, some “non‐sibling” relationships exist that although are related to each other are not identified as such. The authors identify and build a relationship network (relationship‐net) with categories as the vertices and relationships as the edges of this network. They demonstrate that using a relationship‐net, some nonobvious category relationships are detected. Their approach capitalizes on the misclassification information generated during the process of text classification to identify potential relationships among categories and automatically generate relationship‐nets. Their results demonstrate a statistically significant improvement over the current approach by up to 73% on 20 News groups 20NG, up to 68% on 17 categories in the Open Directories Project (ODP17), and more than twice on ODP46 and Special Interest Group on Information Retrieval (SIGIR) data sets. Their results also indicate that using misclassification information stemming from passage classification as opposed to document classification statistically significantly improves the results on 20NG (8%), ODP17 (5%), ODP46 (73%), and SIGIR (117%) with respect to F1 measure. By assigning weights to relationships and by performing feature selection, results are further optimized.
Saket S. R. Mengle, Nazli Goharian
J. Assoc. Inf. Sci. Technol.1
2009 Passage detection using text classification
abstract
Abstract Passages can be hidden within a text to circumvent their disallowed transfer. Such release of compartmentalized information is of concern to all corporate and governmental organizations. Passage retrieval is well studied; we posit, however, that passage detection is not. Passage retrieval is the determination of the degree of relevance of blocks of text, namely passages, comprising a document. Rather than determining the relevance of a document in its entirety, passage retrieval determines the relevance of the individual passages. As such, modified traditional information‐retrieval techniques compare terms found in user queries with the individual passages to determine a similarity score for passages of interest. In passage detection, passages are classified into predetermined categories. More often than not, passage detection techniques are deployed to detect hidden paragraphs in documents. That is, to hide information, documents are injected with hidden text into passages. Rather than matching query terms against passages to determine their relevance, using text‐mining techniques, the passages are classified. Those documents with hidden passages are defined as infected. Thus, simply stated, passage retrieval is the search for passages relevant to a user query, while passage detection is the classification of passages. That is, in passage detection, passages are labeled with one or more categories from a set of predetermined categories. We present a keyword‐based dynamic passage approach (KDP) and demonstrate that KDP outperforms statistically significantly (99% confidence) the other document‐splitting approaches by 12% to 18% in the passage detection and passage category‐prediction tasks. Furthermore, we evaluate the effects of the feature selection, passage length, ambiguous passages, and finally training‐data category distribution on passage‐detection accuracy.
Saket S. R. Mengle, Nazli Goharian
J. Assoc. Inf. Sci. Technol.1
2009 Ambiguity measure feature-selection algorithm
abstract
Abstract With the increasing number of digital documents, the ability to automatically classify those documents both efficiently and accurately is becoming more critical and difficult. One of the major problems in text classification is the high dimensionality of feature space. We present the ambiguity measure (AM) feature‐selection algorithm, which selects the most unambiguous features from the feature set. Unambiguous features are those features whose presence in a document indicate a strong degree of confidence that a document belongs to only one specific category. We apply AM feature selection on a naïve Bayes text classifier. We favorably show the effectiveness of our approach in outperforming eight existing feature‐selection methods, using five benchmark datasets with a statistical significance of at least 95% confidence. The support vector machine (SVM) text classifier is shown to perform consistently better than the naïve Bayes text classifier. The drawback, however, is the time complexity in training a model. We further explore the effect of using the AM feature‐selection method on an SVM text classifier. Our results indicate that the training time for the SVM algorithm can be reduced by more than 50%, while still improving the accuracy of the text classifier. We favorably show the effectiveness of our approach by demonstrating that it statistically significantly (99% confidence) outperforms eight existing feature‐selection methods using four standard benchmark datasets.
Saket S. R. Mengle, Nazli Goharian
J. Assoc. Inf. Sci. Technol.1
2008 On document splitting in passage detection
abstract
Passages can be hidden within a text to circumvent their disallowed transfer. Such release of compartmentalized information is of concern to all corporate and governmental organization. We explore the methodology to detect such hidden passages within a document. A document is divided into passages using various document splitting techniques, and a text classifier is used to categorize such passages. We present a novel document splitting technique called dynamic windowing, which significantly improves precision, recall and F1 measure.
Nazli Goharian, Saket S. R. Mengle
SIGIR2
2007 FACT: Fast Algorithm for Categorizing Text
abstract
With the ever-increasing number of digital documents, the ability to automatically classify those documents both quickly and accurately is becoming more critical and difficult. We present Fast Algorithm for Categorizing Text (FACT), which is a statistical based multi-way classifier with our proposed feature selection, Ambiguity Measure (AM), which uses only the most unambiguous keywords to predict the category of a document. Our empirical results show that FACT outperforms the best results on the best performing feature selection for the Naive Bayes classifier namely, Odds Ratio. We empirically show the effectiveness of our approach in outperforming Odds Ratio using four benchmark datasets with a statistical significance of 99% confidence level. Furthermore, the performance of FACT is comparable or better than current non-statistical based classifiers.
Saket S. R. Mengle, Nazli Goharian, Alana Platt
ISI1