Vijay Raghavan 0001

dblp:r/VVRaghavan1 · also Vijay V. Raghavan 0001 · DBLP profile ↗
← Back
111ranked-venue papers
15as first author
4since 2021 · last 2026
0000-0001-7224-7828ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 64 · 13 first-author · 1 since 2021Artificial intelligence and machine learning · 42 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 15 · 1 since 2021Security and privacy · 4 · 1 since 2021Theory of computation · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 SSE-TSR: An Approach to Integrate Secondary Structure Elements Into Triangular Spatial Relationships for Protein Classification
abstract
Protein structures are fundamental to understanding biological function, yet many detailed similarities remain hidden from conventional alignment-based or 3D superposition methods. Triangular Spatial Relationship (TSR) offers an alignment-free encoding of backbone geometry; however, classical TSR ignores the context of secondary structure elements (SSEs), such as helices, strands, and coils. To address this, we introduce SSE-TSR, which enriches each TSR key by categorizing it into one of 18 helix-strand-coil combination labels derived from DSSP-style annotations in PDB HELIX/SHEET records. By mapping the protein representation involving SSE-TSR keys into a sparse tensor, SSE-TSR compactly captures both tertiary geometry and local secondary motifs. We evaluated SSE-TSR on four datasets, two structural (CATH-based, 9.2K; SCOP-based, 7.0K) and two functional (published, 7.8K; new, 7.2K), using a 3D convolutional neural network. On structure-based tasks, SSETSR noticeably boosts accuracy from 96.00% to 98.33% (CATHbased) and from 95.46% to 99.00% (SCOP-based). On functional tasks, it yields modest yet consistent gains (e.g., from 99.41% to 99.50% and 95.83% to 98.83%). Comparisons to Foldseek confirm competitive accuracy across diverse tasks. Additionally, the sparse tensor representation enables memory-efficient handling of large-scale datasets, making SSE-TSR practical for extensive bioinformatics analyses. These results demonstrate SSE-TSR as a scalable, interpretable, and robust method, enhancing protein classification and structural bioinformatics.
Poorya Khajouie, Titli Sarkar, Krishna Rauniyar, Li Chen 0019, Wu Xu, Vijay Raghavan 0001
IEEE Trans. Comput. Biol. Bioinform.6
2025 Detecting Anomalous Communication Behaviors in Dynamically Evolving Networked Systems
Mehedi Hassan, M. Engin Tozal, Vipin Swarup, Steven Noel, Raju N. Gottumukkala, Vijay Raghavan 0001
IEEE Trans. Inf. Forensics Secur.6
2021 KNN Loss and Deep KNN
abstract
The k Nearest Neighbor (KNN) algorithm has been widely applied in various supervised learning tasks due to its simplicity and effectiveness. However, the quality of KNN decision making is directly affected by the quality of the neighborhoods in the modeling space. Efforts have been made to map data to a better feature space either implicitly with kernel functions, or explicitly through learning linear or nonlinear transformations. However, all these methods use pre-determined distance or similarity functions, which may limit their learning capacity. In this paper, we present two loss functions, namely KNN Loss and Fuzzy KNN Loss, to quantify the quality of neighborhoods formed by KNN with respect to supervised learning, such that minimizing the loss function on the training data leads to maximizing KNN decision accuracy on the training data. We further present a deep learning strategy that is able to learn, by minimizing KNN loss, pairwise similarities of data that implicitly maps data to a feature space where the quality of KNN neighborhoods is optimized. Experimental results show that this deep learning strategy (denoted as Deep KNN) outperforms state-of-the-art supervised learning methods on multiple benchmark data sets.
Linh Le, Ying Xie 0001, Vijay Raghavan 0001
Fundam. Informaticae3
2021 Deep Multi-View Learning to Rank
abstract
We study the problem of learning to rank from multiple information sources. Though multi-view learning and learning to rank have been studied extensively leading to a wide range of applications, multi-view learning to rank as a synergy of both topics has received little attention. The aim of the paper is to propose a composite ranking method while keeping a close correlation with the individual rankings simultaneously. We present a generic framework for multi-view subspace learning to rank (MvSL2R), and two novel solutions are introduced under the framework. The first solution captures information of feature mappings from within each view as well as across views using autoencoder-like networks. Novel feature embedding methods are formulated in the optimization of multi-view unsupervised and discriminant autoencoders. Moreover, we introduce an end-to-end solution to learning towards both the joint ranking objective and the individual rankings. The proposed solution enhances the joint ranking with minimum view-specific ranking loss, so that it can achieve the maximum global view agreements in a single optimization process. The proposed method is evaluated on three different ranking problems, i.e., university ranking, multi-view lingual text ranking, and image data ranking, providing superior results compared to related methods.
Guanqun Cao, Alexandros Iosifidis, Moncef Gabbouj, Vijay Raghavan 0001, Raju N. Gottumukkala
IEEE Trans. Knowl. Data Eng.4
2020 Learning with Partial Multi-Outlooks
Yi He 0007, Vijay Raghavan 0001
IJCNN3
2019 Spatio-temporal outlier detection algorithms based on computing behavioral outlierness factor
Maria Bala Duggimpudi, Shaaban Abbady, Jian Chen 0032, Vijay Raghavan 0001
Data Knowl. Eng.4
2018 Distributed Real Time Link Prediction on Graph Streams
abstract
Link prediction refers to estimating the likelihood of a link appearing in the future based on the current status of a graph. Link prediction problem applications in various domains such as bioinformatics, social network analysis, cybersecurity and e-commerce. Some of these graphs are massive and are constantly evolving. Many applications require these graph streams to be processed them in real-time, to predict the link based on the most recent information as the graph features may change over time. Existing approaches to process large graphs for link prediction is non-trivial due to the following reasons: 1) Graphs required to predict the links are too large to be stored in a single RAM. Link prediction on these large graphs is expensive in terms of computation resources and time required to perform link prediction, 2) Sketch-based approaches are not suitable in applications where accuracy is critical (such as analyzing criminal social networks or supply chain networks) and 3) Sketch-based approaches also fail to handle dynamic graphs, where edges are not only added, but also removed. This results in changes to the graph topology, making the features previously computed to be obsolete. Distributed data stream frameworks such as Apache Flink could be potentially used for distributed graph processing. However, there are no techniques to handle link prediction on distributed graph streams. In this paper, we consider three fundamental, neighborhood-based link prediction measures, Jaccard coefficient, Preferential attachment, and common neighbors and enable an accurate measurement of them to address link prediction problem in dynamic graph streams. We propose a neighborhood-centric graph processing approach to handle graphs that exploits the locality, parallelism, and incremental computation of existing distributed frameworks to calculate these graph features with exact results. We perform experimental studies on various real-world graph streams. The results demonstrate that our graph measures are accurate and are more efficient than the existing vertex-centric approaches to graph processing.
Satya Katragadda, Raju N. Gottumukkala, Murali Pusala, Vijay Raghavan 0001, Jessica Wojtkiewicz
IEEE BigData4
2018 Deep Similarity-Enhanced K Nearest Neighbors
abstract
The k Nearest Neighbors (KNN) algorithm has been widely applied in various supervised learning tasks due to its simplicity and effectiveness. However, the quality of KNN decision making is directly affected by the quality of the neighborhoods in the modeling space. Efforts have been made to map data to a better feature space either implicitly with kernel functions, or explicitly through learning linear or nonlinear transformations. However, all these methods use pre-determined distance or similarity functions, which may limit their learning capacity. In this paper, we propose a novel deep learning architecture, which is called the Deep Similarity-Enhanced K Nearest Neighbors (DSE-KNN), to learn an optimized similarity function of the data directly towards the goal of optimizing the KNN decision making. In other words, the type of similarity function that is used in our method is not pre-determined but rather learned to map data to a high-dimensional feature space where the accuracy of the KNN decision making is maximized. Experimental results show that DSE-KNN outperforms other common machine learning methods on classifying different types of disease datasets and predicting daily price direction of different stock ETFs.
Linh Le, Ying Xie 0001, Vijay Raghavan 0001
IEEE BigData3
2018 Unsupervised Learning to Rank Aggregation using Parameterized Function Optimization
abstract
This paper proposes a novel unsupervised rank aggregation method using parameterized function optimization (PFO). This algorithm derives a parameterized rank aggregation model by minimizing the energy of weighted standard deviations of rank lists associated with different rankers or attributes. Parameters, in this problem, are weights representing the impact of rank lists on the final aggregated rank. The proposed learning to rank aggregation method is efficient (linear time complexity) and its accuracy compares favorably with pairwise preference methods (with polynomial time complexity). Two rounds of experiments are run to show the success of PFO in rank aggregation: one on the learning to rank (LETOR) benchmark dataset to show its success in unsupervised rank aggregation and the other on three university ranking datasets to solve a practical problem in education. The experimental results on the LETOR show that PFO significantly outperforms the baseline results and show promising performances in comparison with recent high performance methods developed for unsupervised rank aggregation. The university ranks obtained by our model compare favorably with the ranks reported by well-known organizations. Success of the PFO model for performing unsupervised rank aggregation, specifically on practical problems, supports the use of the algorithm in difficult ranking scenarios without ground truth.
Amirhossein Tavanaei, Raju N. Gottumukkala, Anthony S. Maida, Vijay Raghavan 0001
IJCNN4
2017 Supervised approach to rank predicted links using interestingness measures
abstract
For the last decade, the automatic generation of hypothesis from the literature has been widely studied. One common approach is to model biomedical literature as a concept network; then a prediction model is applied to predict the future relationships (links) between pairs of concept. Typically, this link prediction task can be cast into in one of two forms: (a) predict the future links for a specific concept (node) or (b) predict the future links for the entire network. However, while being able to accurately forecast future relationships is vital, another, equally important question should be addressed: of the predicted links, which will be most important and/or most relevant? Attempts to answer these questions in the past have generally been domain specific. In this paper, we propose a domain-independent, supervised method that predicts the rank of future links utilizing objective interestingness measures. The results, based on analysis of thirteen common interestingness measures, indicate that, while predicting the specific future interestingness values is difficult, our approach allowed us to capture the relative ordering of the links with low error.
Murali Pusala, Ryan G. Benton, Vijay Raghavan 0001, Raju N. Gottumukkala
BIBM3
2017 Descriptor based protein structure representation using triangular spatial relationships in 3-D
abstract
Pairwise protein structure comparison has taken significant scientific research effort in last two decades. Even though it all started with alignment-based comparison methods, recently there are several non-alignment based methods that have shown good potential. One such approach is based on shape descriptors. These methods use histograms or vectors to represent the molecular shapes. They have shown to improve comparison speed but require reference frame transformations, are limited to sequential comparisons as well as lack residue information. Like any sequence independent model, these methods ignore the correspondence of residues in similarity calculation. This work proposes protein structure representation method, Triangular Spatial Relationships in 3D (TSR 3-D). TSR 3-D is local-scale sensitive, reference frame insensitive, protein structure descriptor that incorporates the residue information. It can be used to establish a strict pairwise equivalence that acknowledges the pivotal role played by corresponding residues in determining protein 3-D structure. Protein structures represented by TSR-3D can be used for flexible non-sequential, pairwise structure comparison using local-global equivalences.
Sumi Singh, Wu Xu, Vijay Raghavan 0001
BIBM3
2017 Online mining for association rules and collective anomalies in data streams
abstract
When analyzing streaming data, the results can depreciate in value faster than the analysis can be completed and results deployed. This is certainly the case in the area of anomaly detection, where detecting a potential problem as it is occurring (or in the early stages) can permit corrective behavior. However, most anomaly detection methods focus on point anomalies, whilst many fraudulent behaviors could be detected only through collective analysis of sequences of data in practice. Moreover, anomaly detection systems often stop at detecting anomalies; they typically do not provide information about how the features (attributes) of anomalies relate to each other or to those in normal states. The goal of this research is to create a distributed system that allows for the detection of collective anomalies from streaming data, and to provide a richer context of information about the anomalies besides their presence. To accomplish this, we (a) re-engineered an online sequence anomaly detection algorithm and (b) designed new algorithms for targeted association mining to run on a streaming, distributed environment. Our experiments, conducted on both synthetic and real-world data sets, demonstrated that the proposed framework is able to achieve near real-time response in detecting anomalies and extracting information pertaining to the anomalies.
Shaaban Abbady, Cheng-Yuan Ke, Jennifer Lavergne, Jian Chen 0032, Vijay Raghavan 0001, Ryan G. Benton
IEEE BigData5
2017 Sub-event detection from tweets
abstract
Social media plays an important role in communication between people in recent times. This includes information about news and events that are currently happening. Most of the research on event detection concentrates on identifying events from social media information. These models assume an event to be a single entity and treat it as such during the detection process. This assumption ignores that the composition of an event changes as new information is made available on social media. To capture the change in information over time, we extend an already existing Event Detection at Onset algorithm to study the evolution of an event over time. We introduce the concept of an event life cycle model that tracks various key events in the evolution of an event. The proposed unsupervised sub-event detection method uses a threshold-based approach to identify relationships between sub-events over time. These related events are mapped to an event life cycle to identify sub-events. We evaluate the proposed sub-event detection approach on a large-scale Twitter corpus.
Satya Katragadda, Ryan G. Benton, Vijay Raghavan 0001
IJCNN3
2017 Link prediction based hybrid recommendation system using user-page preference graphs
abstract
The distribution of the amount of preference information across customers is not same in every domain of recommendation problems. It is necessary to treat each user differently based on their available preference information. On the other hand, graph structure can provide better representation of user-item preference information. By exploiting graph structure, recommendation systems could be made more reliable and effective. In this paper, several graph structure based features are used to predict likelihood of a page or item being preferred by a user, or future connection probability of an unconnected user-page node pair in the user-page preference graph. A novel, user-specific, parametric method to integrate page-page content similarity and co-occurrence similarity in the graph context is introduced. Two graphs were created; one using page-page content similarity; other one using page-page co-occurrence similarity. An approach that utilizes features derived from these two graphs to make web page recommendation is introduced. For each user-page pair, one combined feature component is first obtained by making a weighted summation of the eight features extracted from each graph. Use of supervised learning for deriving relative weights of the two eight-feature sets to obtain a combined value of the feature components yielded highly promising results. Finally, the two feature components, from the two graphs are combined in user-specific way to train a model and to make recommendations. Experimental results on Yahoo Front Page Today Module Click Log dataset show better results compared to other approaches in the literature. To the best of our knowledge, ours is the first such effort in the context of graph-based recommendation systems.
Mohammad Amir Sharif, Vijay Raghavan 0001
IJCNN2
2017 Big Data and Data Analytics Research: From Metaphors to Value Space for Collective Wisdom in Human Decision Making and Smart Machines
abstract
The Big Data and Data Analytics is a brand new paradigm, for the integration of Internet Technology in the human and machine context. For the first time in the history of the human mankind we are able to transforming raw data that are massively produced by humans and machines in to knowledge and wisdom capable of supporting smart decision making, innovative services, new business models, innovation, and entrepreneurship. For the Web Science research, this is a new methodological and technological spectrum of advanced methods, frameworks and functionalities never experienced in the past. At the same moment communities out of web science need to realize the potential of this new paradigm with the support of new sound business models and a critical shift in the perception of decision making. In this short visioning article, the authors are analyzing the main aspects of Big Data and Data Analytics Research and they provide their own metaphor for the next years. A number of research directions are outlined as well as a new roadmap towards the evolution of Big Data to Smart Decisions and Cognitive Computing. The authors do hope that the readers would like to react and to propose their own value propositions for the domain initiating a scientific dialogue beyond self-fulfilled expectations.
Miltiadis D. Lytras, Vijay Raghavan 0001, Ernesto Damiani
Int. J. Semantic Web Inf. Syst.2
2016 S3C: An architecture for space-efficient semantic search over encrypted data in the cloud
abstract
The recent rapid growth in Internet speeds and file storage requirements has made cloud storage an appealing option on both a personal and enterprise level. Despite the many benefits offered by cloud storage, many potential users with sensitive data refrain from fully utilizing this service due to valid concerns about information privacy. An established solution to this concern is to perform encryption on the user side with the key stored on a local machine, meaning the cloud will never see the user's plaintext data. However, by encrypting data on the user side data processing capabilities (e.g., searching) are lost. In particular, the ability to semantically search is of the user's interest in large datasets. In this paper, we present S3C, a system that provides a semantic search functionality over encrypted data in the cloud. S3C combines approaches from traditional keyword-based searchable encryption and semantic web searching. It offers a user transparent experience that accepts a simple multi-phrase query and returns a list of documents ranked by semantic relevance to the query. Our proposed approach is space-efficient, which makes it suitable for large scale datasets. Our minimal processing also allows the system to be run on thin clients such as smart-phones or tablets. We evaluate the performance of our system against various real-world datasets, and our results show that it produces accurate search results while maintaining minimal storage overhead (~0.3% of the dataset size).
Jason Woodworth, Mohsen Amini Salehi, Vijay Raghavan 0001
IEEE BigData3
2016 Detection of event onset using Twitter
abstract
Social Media generates information about news and events in real-time. Given the vast amount of data available and the rate of information propagation, reliably identifying events is a challenge. Most state-of-the-art techniques are post hoc techniques that detect an event after it happened. Our goal is to detect onset of an event as it is happening using the user-generated information from Twitter streams. To achieve this goal, we use a discriminative model to identify change in the pattern of conversations over time. We use a topic evolution model to find credible events and eliminate random noise that is prevalent in many of the event detection models. The simplicity of the proposed model allows detect events quickly and efficiently, permitting discovery of events within minutes from the start of conversation about those conversations on Twitter. Our model is evaluated on a large-scale Twitter corpus to detect events in real-time. The proposed model is tested on other datasets to detect change over longer periods of time. The results indicate we can detect real events, within 3 to 8 minutes of it first appearing, with a lower degree of noise compared to other methods.
Satya Katragadda, Shahid Virani, Ryan G. Benton, Vijay Raghavan 0001
IJCNN4
2016 An Ontology-Based Architecture for Providing Insights in Wireless Networks Domain
abstract
Ontology-based approaches have been explored in several domains for knowledge representation and improving accuracy. However, ontology-based approaches for assisting a decision maker by delivering a concrete plan from analyzing the insights extracted from an ontology, have not received much attention. Insights-as-a-service is a technology that aids a decision maker by providing a concrete action plan, involving a comparative analysis of patterns derived from the data and the extraction of insights from such an analysis. In this paper, we propose an ontology-based architecture for mining insights within the Wireless Network Ontology (WNO), an ontology generated for the wireless network domain for delivering better wireless network performance. We present and illustrate: (i) the major components of the architecture together with the algorithms used for summarizing the network performance profiles in the form of rank tables, and (ii) how the insight rules (the action plan) are extracted from these tables. By utilizing the proposed approach, an actionable plan for assisting the decision maker can be obtained as domain knowledge is incorporated in the system. Experimental results on a wireless network dataset show that the proposed model provides an optimal action plan for a wireless network to improve its performance by encoding data-driven rules into the ontology and suggesting changes to its current network configuration.
Maria Bala Duggimpudi, Abdelhamid Moursy, Elshaimaa Ali, Vijay Raghavan 0001
WI4
2015 Detecting adverse drug effects using link classification on twitter data
abstract
Adverse drug events (ADEs) are among the leading causes of death in the United States. Although many ADEs are detected during pharmaceutical drug development and the FDA approval process, all of the possible reactions cannot be identified during this period. Currently, post-consumer drug surveillance relies on voluntary reporting systems, such as the FDA's Adverse Event Reporting System (AERS). With an increase in availability of medical resources and health related data online, interest in medical data mining has grown rapidly. This information coupled with online conversations of people which involve discussions about their health provide a substantial resource for the identification of ADEs. In this work, we propose a method to identify adverse drug effects from tweets by modeling it as a link classification problem in graphs. Drug and symptom mentions are extracted from the tweet history of each user and a drug-symptom graph is built, where nodes represent either drugs or symptoms and edges are labelled positive or negative, for desired or adverse drug effects respectively. A link classification model is then used to identify negative edges i.e. adverse drug effects. We test our model on 864 users using 10-fold cross validation with Sider's dataset as ground truth. Our model was able to achieve an F-Score of 0.77 compared to the best baseline model with an F-Score of 0.58.
Satya Katragadda, Harika Karnati, Murali Pusala, Vijay Raghavan 0001, Ryan G. Benton
BIBM4
2015 Data quality issues in big data
abstract
Though the issues of data quality trace back their origin to the early days of computing, the recent emergence of Big Data has added more dimensions. Furthermore, given the range of Big Data applications, potential consequences of bad data quality can be for more disastrous and widespread. This paper provides a perspective on data quality issues in the Big Data context. it also discusses data integration issues that arise in biological databases and attendant data quality issues.
Dhana Rao, Venkat N. Gudivada, Vijay Raghavan 0001
IEEE BigData3
2015 Extending SKOS: A Wikipedia-Based Unified Annotation Model for Creating Interoperable Domain Ontologies
Elshaimaa Ali, Vijay Raghavan 0001
ISMIS2
2015 Editorial
Vijay Raghavan 0001
Web Intell.2
2014 A clustering based scalable hybrid approach for web page recommendation
abstract
The distribution of the number of items liked by users plays an important role in designing recommender systems. In case of implicit feedback we rarely get some clicking events compared to large item based e-commerce sites, where preference information is not so rare. In this paper we present a novel hybrid recommendation system based on clustering of items using co-occurrence information of pages and content information of pages. These two different types of clusters are used in a parametric form to get aggregated recommendations based on the available preference information of users. Our experimental results on Yahoo! Front Page “Today Module User Click Log” dataset show that the content based clusters plays an important role for users having very less preference information and also the clustering based hybrid approach gives better overall performance compared to other approaches. More-over, clustering of items gives a scalable implementation which minimizes the computational complexity.
Mohammad Amir Sharif, Vijay Raghavan 0001
IEEE BigData2
2014 A Large-Scale, Hybrid Approach for Recommending Pages Based on Previous User Click Pattern and Content
Mohammad Amir Sharif, Vijay Raghavan 0001
ISMIS2
2014 NoSQL Systems for Big Data Management
abstract
The advent of Big Data created a need for out-of-the-box horizontal scalability for data management systems. This ushered in an array of choices for Big Data management under the umbrella term NoSQL. In this paper, we provide a taxonomy and unified perspective on NoSQL systems. Using this perspective, we compare and contrast various NoSQL systems using multiple facets including system architecture, data model, query language, client API, scalability, and availability. We group current NoSQL systems into seven broad categories: Key-Value, Table-type/Column, Document, Graph, Native XML, Native Object, and Hybrid databases. We also describe application scenarios for each category to help the reader in choosing an appropriate NoSQL system for a given application. We conclude the paper by indicating future research directions.
Venkat N. Gudivada, Dhana Rao, Vijay Raghavan 0001
SERVICES3
2013 DynTARM: An In-Memory Data Structure for Targeted Strong and Rare Association Rule Mining over Time-Varying Domains
abstract
Recently, with companies and government agencies saving large repositories of time stream/temporal data, there is a large push for adapting association rule mining algorithms for dynamic, targeted querying. In addition, issues with data processing latency and results depreciating in value with the passage of time, create a need for swifter and more efficient processing. The aim of targeted association mining is to find potentially interesting implications in large repositories of data. Using targeted association mining techniques, specific implications that contain items of user interest can be found faster and before the implications have depreciated in value beyond usefulness. In this paper, the DynTARM algorithm is proposed for the discovery of targeted and rare association rules. DynTARM has the flexibility to discover strong and rare association rules from data streams within the user's sphere of interest. By introducing a measure, called the Volatility Index, to assess the fluctuation in the confidence of rules, rules conforming to different temporal patterns are discovered.
Jennifer Lavergne, Ryan G. Benton, Vijay Raghavan 0001, Alaaeldin M. Hafez
Web Intelligence3
2012 Weighted Fuzzy Aggregation for Metasearch: An Application of Choquet Integral
Elizabeth D. Diaz, Vijay Raghavan 0001
IPMU (1)3
2012 Min-Max Itemset Trees for Dense and Categorical Datasets
Jennifer Lavergne, Ryan G. Benton, Vijay Raghavan 0001
ISMIS3
2012 TRARM-RelSup: Targeted Rare Association Rule Mining Using Itemset Trees and the Relative Support Measure
Jennifer Lavergne, Ryan G. Benton, Vijay Raghavan 0001
ISMIS3
2012 Special issue on advances in web intelligence
Stefan M. Rüger, Vijay Raghavan 0001, Irwin King, Jimmy Huang 0001
Neurocomputing2
2011 Supervised Link Discovery on Large-Scale Biomedical Concept Networks
abstract
Computational approaches to generate hypotheses from biomedical literature have been studied intensively in recent years. Nevertheless, it still remains a challenge to automatically discover novel, cross-silo biomedical hypotheses from large-scale literature repositories. In order to address this challenge, we first model a biomedical literature repository as a comprehensive network of biomedical concepts and formulate hypotheses generation as a process of link discovery on the concept network. We extract the relevant information from the biomedical literature corpus and generate a concept network and concept-author matrix on a cluster using Map-Reduce framework. We extract a set of heterogeneous features such as random walk based features, neighborhood features and common author features. The potential number of links to consider for the possibility of link discovery is large in our concept network and to address the scalability problem, the features from a concept network are extracted using a cluster with Map-Reduce framework. We further model link discovery as a classification problem carried out on two network snapshots taken in two consecutive time frames, such that the classification model that is built on the first snapshot can be tested on the second snapshot. A set of heterogeneous features, which cover both topological and semantic features derived from the concept network, have been studied with respect to their impacts on the accuracy of the proposed supervised link discovery process.
Jayasimha Reddy Katukuri, Ying Xie 0001, Vijay Raghavan 0001
BIBM3
2010 Exploitation of 3D Stereotactic Surface Projection for automated classification of Alzheimer's disease according to dementia levels
abstract
Alzheimer's disease (AD) is one major cause of dementia. Previous studies have indicated that the use of features derived from Positron Emission Tomography (PET) scans lead to more accurate and earlier diagnosis of AD, compared to the traditional approach used for determining dementia ratings, which uses a combination of clinical assessments such as memory tests. In this study, we compare Naïve Bayes (NB), a probabilistic learner, with variations of Support Vector Machines (SVMs), a geometric learner, for the automatic diagnosis of Alzheimer's disease. 3D Stereotactic Surface Projection (3D-SSP) is utilized to extract features from PET scans. At the most detailed level, the dimensionality of the feature space is very high, resulting in 15964 features. Since classifier performance can degrade in the presence of a high number of features, we evaluate the benefits of a correlation-based feature selection method to find a small number of highly relevant features.
Murat Seçkin Ayhan, Ryan G. Benton, Vijay Raghavan 0001, Suresh K. Choubey
BIBM3
2009 Biomedical Relationship Extraction from Literature Based on Bio-semantic Token Subsequences
abstract
Relationship extraction (RE) from biomedical literature is an important and challenging problem in both text mining and bioinformatics. Although various approaches have been proposed to extract protein-protein interaction types, their accuracy rates leave a large room for further exploration of more effective methods. In this paper, two supervised learning algorithms based on newly-defined ldquobio-semantic token subsequencerdquo are proposed for multi-class biomedical relationship extraction. The first approach calculates a ldquobio-semantic token subsequence kernelrdquo, while the second one explicitly extracts weighted features from bio-semantic token subsequences. The proposed structure called ldquobio-semantic token subsequencerdquo is able to capture semantic features from natural language sentences for biomedical RE. Two supervised learning algorithms based on the proposed structure outperform the state-of-the-art biomedical RE methods on multi-class protein-protein interaction extraction.
Jayasimha Reddy Katukuri, Ying Xie 0001, Vijay Raghavan 0001
BIBM3
2007 On Fuzzy Result Merging for Metasearch
abstract
The result merging problem for metasearch engines is to combine multiple result lists, returned by search engines, in response to a query, so as to achieve optimal aggregation. Here we propose three result merging models for metasearch that apply Yager's fuzzy aggregation OWA operator to result merging and extend the OWA model for metasearch proposed by Diaz. These are the importance guided OWA (IGOWA), the algebraic t-norm OWA, and the algebraic t-norm IGOWA models. While the first two are based on Yager's extension of the OWA operator, the third is a combination of the features of the first two. We compare the performance of our models to Diaz's OWA model and the Borda-Fuse and Weighted Borda-Fuse models proposed by Aslam and Montague.
Elizabeth D. Diaz, Vijay Raghavan 0001
FUZZ-IEEE3
2007 AllInOneNews: development and evaluation of a large-scale news metasearch engine
abstract
AllInOneNews is the largest news metasearch engine in the world, connecting to over 1,000 news sites over 150 countries. Implementing a large-scale metasearch engine like AllInOneNews needs to overcome unique challenges not faced by building small metasearch engines such as developing highly scalable search engine selection techniques. In this paper, we discuss these unique challenges and our solutions to these challenges. We also discuss some novel features of AllInOneNews such as highly automated solution and semantic query match. This paper also reports the results of a comparative evaluation of three commercial news search systems, one search engine - Google News and two metasearch engines - Mamma News and AllInOneNews. Several measures such as effectiveness, diversity and time-sensitivity are used to perform the comparison. Another contribution of this paper is that we introduce a novel scheme to compare multiple news search systems in a combined measure that takes both relevance and time-sensitivity of retrieved information into consideration.
King-Lup Liu, Weiyi Meng, Clement T. Yu, Vijay Raghavan 0001, Zonghuan Wu, Yiyao Lu, Hai He, Hongkun Zhao
SIGMOD Conference5
2007 MySearchView: a customized metasearch engine generator
abstract
In this paper, we describe MySearchView, a system for assembling search engines into metasearch engines. With this system, any user can create a metasearch engine by simply letting the system know the URLs of the search engines the user wants to be included and the metasearch engine will be built fully automatically. In this paper, the main steps of building metasearch engines will be sketched. We will also outline our plan to demonstrate all the features of MySearchView.
Yiyao Lu, Zonghuan Wu, Hongkun Zhao, Weiyi Meng, King-Lup Liu, Vijay Raghavan 0001, Clement T. Yu
SIGMOD Conference6
2007 Language-modeling kernel based approach for information retrieval
abstract
Abstract In this presentation, we propose a novel integrated information retrieval approach that provides a unified solution for two challenging problems in the field of information retrieval. The first problem is how to build an optimal vector space corresponding to users' different information needs when applying the vector space model. The second one is how to smoothly incorporate the advantages of machine learning techniques into the language modeling approach. To solve these problems, we designed the language‐modeling kernel function, which has all the modeling powers provided by language modeling techniques. In addition, for each information need, this kernel function automatically determines an optimal vector space, for which a discriminative learning machine, such as the support vector machine, can be applied to find an optimal decision boundary between relevant and nonrelevant documents. Large‐scale experiments on standard test‐beds show that our approach makes significant improvements over other state‐of‐the‐art information retrieval methods.
Ying Xie 0001, Vijay Raghavan 0001
J. Assoc. Inf. Sci. Technol.2
2006 Score Distribution Approach to Automatic Kernel Selection for Image Retrieval Systems
Anca Doloc-Mihu, Vijay Raghavan 0001
ISMIS2
2006 Adaptive relevance feedback method of extended Boolean model using hierarchical clustering techniques
Jongpill Choi, Minkoo Kim, Vijay Raghavan 0001
Inf. Process. Manag.3
2006 Construction of query concepts based on feature clustering of documents
Youjin Chang, Minkoo Kim, Vijay Raghavan 0001
Inf. Retr.3
2006 A cluster-based approach for efficient content-based image retrieval using a similarity-preserving space transformation method
abstract
Abstract The techniques of clustering and space transformation have been successfully used in the past to solve a number of pattern recognition problems. In this article, the authors propose a new approach to content‐based image retrieval (CBIR) that uses (a) a newly proposed similarity‐preserving space transformation method to transform the original low‐level image space into a high‐level vector space that enables efficient query processing, and (b) a clustering scheme that further improves the efficiency of our retrieval system. This combination is unique and the resulting system provides synergistic advantages of using both clustering and space transformation. The proposed space transformation method is shown to preserve the order of the distances in the transformed feature space. This strategy makes this approach to retrieval generic as it can be applied to object types, other than images, and feature spaces more general than metric spaces. The CBIR approach uses the inexpensive “estimated” distance in the transformed space, as opposed to the computationally inefficient “real” distance in the original space, to retrieve the desired results for a given query image. The authors also provide a theoretical analysis of the complexity of their CBIR approach when used for color‐based retrieval, which shows that it is computationally more efficient than other comparable approaches. An extensive set of experiments to test the efficiency and effectiveness of the proposed approach has been performed. The results show that the approach offers superior response time (improvement of 1–2 orders of magnitude compared to retrieval approaches that either use pruning techniques like indexing, clustering, etc., or space transformation, but not both) with sufficiently high retrieval accuracy.
Biren Shah, Vijay Raghavan 0001, Praveen Dhatric, Xiaoquan Zhao
J. Assoc. Inf. Sci. Technol.2
2005 Fully automatic wrapper generation for search engines
abstract
When a query is submitted to a search engine, the search engine returns a dynamically generated result page containing the result records, each of which usually consists of a link to and/or snippet of a retrieved Web page. In addition, such a result page often also contains information irrelevant to the query, such as information related to the hosting site of the search engine and advertisements. In this paper, we present a technique for automatically producing wrappers that can be used to extract search result records from dynamically generated result pages returned by search engines. Automatic search result record extraction is very important for many applications that need to interact with search engines such as automatic construction and maintenance of metasearch engines and deep Web crawling. The novel aspect of the proposed technique is that it utilizes both the visual content features on the result page as displayed on a browser and the HTML tag structures of the HTML source file of the result page. Experimental results indicate that this technique can achieve very high extraction accuracy.
Hongkun Zhao, Weiyi Meng, Zonghuan Wu, Vijay Raghavan 0001, Clement T. Yu
WWW4
2005 A new fuzzy clustering algorithm for optimally finding granular prototypes
Ying Xie 0001, Vijay Raghavan 0001, Praveen Dhatric, Xiaoquan Zhao
Int. J. Approx. Reason.2
2004 Efficient and Effective Content-Based Image Retrieval using Space Transformation
abstract
A promising approach to content-based image retrieval, proposed by Choubey and Raghavan (1997), involves the representation of the original image space, in terms of low-level image features into a feature space, where images are represented as vectors of high-level features. The retrieval system based on that approach consists of three phases: database population; online addition; and image retrieval. Though their framework supports content-based retrieval of images, the issues that arise when database grows dynamically and user queries are different from those in database, were not investigated. In the current work, we: (i) experimentally investigate issues relating to online addition of new images and image retrieval; and (ii) provide a theoretical analysis of the complexity and effectiveness of our retrieval system by comparing it with the conventional approach. We have performed an extensive set of experiments to test the efficiency, effectiveness and scalability of our approach. The experimental results show that our approach is not only efficient but also effective in retrieving images even when the image database is dynamic and user queries are framed with images that are external to the database.
Biren Shah, Vijay Raghavan 0001, Praveen Dhatric
MMM2
2003 A Methodology for Hiding Knowledge in XML Document Collections
abstract
Information marked up as XML data is becoming increasingly pervasive as a part of business-to-business electronic transactions. A possible threat to the continued growth of XML in this domain is that data mining technology may be applied to XML documents in order to reveal sensitive knowledge. This paper presents a methodology for hiding sensitive knowledge in XML documents in the context of association mining algorithms. This methodology involves identifying the sensitive knowledge within the document, formulating an appropriate set of security policies, and finally sanitizing the document to hide the sensitive knowledge.
Tom Johnsten, Robert B. Sweeney, Vijay Raghavan 0001
COMPSAC3
2003 Probability Logic Modeling of Knowledge Discovery in Databases
Jitender S. Deogun, Liying Jiang, Ying Xie 0001, Vijay Raghavan 0001
ISMIS4
2003 Space Transformation Based Approach for Effective Content-Based Image Retrieval
Biren Shah, Vijay Raghavan 0001
ISMIS2
2003 SE-LEGO: creating metasearch engines on demand
abstract
No abstract available.
Zonghuan Wu, Vijay Raghavan 0001, Chun Du, Komanduru Sai C, Weiyi Meng, Hai He, Clement T. Yu
SIGIR2
2003 Creating Customized Metasearch Engines on Demand Using SE-LEGO
Zonghuan Wu, Vijay Raghavan 0001, Weiyi Meng, Hai He, Clement T. Yu, Chun Du
WAIM2
2003 Towards Automatic Incorporation of Search Engines into a Large-Scale Metasearch Engine
abstract
A metasearch engine supports unified access to multiple component search engines. To build a very large-scale metasearch engine that can access up to hundreds of thousands of component search engines, one major challenge is to incorporate large numbers of autonomous search engines in a highly effective manner. To solve this problem, we propose automatic search engine discovery, automatic search engine connection, and automatic search engine result extraction techniques. Experiments indicate that these techniques are highly effective and efficient.
Zonghuan Wu, Vijay Raghavan 0001, Hua Qian, Rama Vuyyuru, Weiyi Meng, Hai He, Clement T. Yu
Web Intelligence2
2003 Itemset Trees for Targeted Association Querying
abstract
Association mining techniques search for groups of frequently co-occurring items in a market-basket type of data and turn these groups into business-oriented rules. Previous research has focused predominantly on how to obtain exhaustive lists of such associations. However, users often prefer a quick response to targeted queries. For instance, they may want to learn about the buying habits of customers that frequently purchase cereals and fruits. To expedite the processing of such queries, we propose an approach that converts the market-basket database into an itemset tree. Experiments indicate that the targeted queries are answered in a time that is roughly linear in the number of market baskets, N. Also, the construction of the itemset tree has O(N) space and time requirements. Some useful theoretical properties are proven.
Miroslav Kubat, Aladdin Hafez, Vijay Raghavan 0001, Jayakrishna R. Lekkala, Wei Kian Chen
IEEE Trans. Knowl. Data Eng.3
2002 On Security and Privacy Risks in Association Mining Algorithms
Tom Johnsten, Vijay Raghavan 0001, Kevin Hill
DBSec2
2002 3M algorithm: finding an optimal fuzzy cluster scheme for proximity data
abstract
In order to find an optimal fuzzy cluster scheme for proximity data, where just pairwise distances among objects are given, two conditions are necessary: A good cluster validity function, which can be applied to proximity data for evaluation of the goodness of cluster schemes for varying number of clusters; a good cluster algorithm that can deal with proximity data and produce an optimal solution for a fixed number of clusters. To satisfy the first condition, a new validity function is proposed, which works well even when the number of clusters is very large. For the second condition, we give a new algorithm called multi-step maxmin and merging algorithm (3M algorithm). Experiments show that, when used in conjunction with the new cluster validity function, the 3M algorithm produces satisfactory results.
Ying Xie 0001, Vijay Raghavan 0001, Xiaoquan Zhao
FUZZ-IEEE2
2002 Visualization of Document Co-Citation Counts
abstract
Visualization can facilitate the understanding of the structures of a collection of documents that are related to each other by links, such as citations in formal publications. We present results of visualizing minimum spanning trees based on document co-citation counts and on document citation correlations.
Steven Noel, Chee-Hung Henry Chu, Vijay Raghavan 0001
IV3
2001 A Theoretical Framework for Association Mining Based on the Boolean Retrieval Model
Peter Bollmann-Sdorra, Aladdin Hafez, Vijay Raghavan 0001
DaWaK3
2001 Security Procedures for Classification Mining Algorithms
Tom Johnsten, Vijay Raghavan 0001
DBSec2
2001 Visualizing Association Mining Results through Hierarchical Clusters
abstract
We propose a new methodology for visualizing association mining results. Inter-item distances are computed from combinations of itemset supports. The new distances retain a simple pairwise structure, and are consistent with important frequently occurring itemsets. Thus standard tools of visualization, e.g. hierarchical clustering dendrograms can still be applied, while the distance information upon which they are based is richer. Our approach is applicable to general association mining applications, as well as applications involving information spaces modeled by directed graphs, e.g. the Web. In the context of collections of hypertext documents, the inter-document distances capture the information inherent in a collection's link structure, a form of link mining. We demonstrate our methodology with document sets extracted from the Science Citation Index, applying a metric that measures consistency between clusters and frequent itemsets.
Steven Noel, Vijay Raghavan 0001, Chee-Hung Henry Chu
ICDM2
2001 BitCube: A Three-Dimensional Bitmap Indexing for XML Documents
abstract
We describe a new bitmap indexing based technique to cluster XML documents. XML is a new standard for exchanging and representing information on the Internet. Documents can be hierarchically represented by XML-elements. XML documents are represented and indexed using a bitmap indexing technique. We define the similarity and popularity operations available in bitmap indexes and propose a method for partitioning a XML document set. Furthermore, a 2-dimensional bitmap index is extended to a 3-dimensional bitmap index, called BitCube. We define statistical measurements in the BitCube: mean, mode, standard derivation, and correlation coefficient. Based on these measurements, we also define the slice, project, and dice operations on a BitCube. BitCube can be manipulated efficiently and improves the performance of document retrieval.
Jong P. Yoon, Vijay Raghavan 0001, Venu Chakilam
SSDBM2
2001 Concept Based Retrieval Using Generalized Retrieval Functions
Minkoo Kim, Jitender S. Deogun, Vijay Raghavan 0001
Fundam. Informaticae3
2001 BitCube: A Three-Dimensional Bitmap Indexing for XML Documents
Jong P. Yoon, Vijay Raghavan 0001, Venu Chakilam, Larry Kerschberg
J. Intell. Inf. Syst.2
2000 Dynamic Data Mining
Vijay Raghavan 0001, Aladdin Hafez
IEA/AIE1
2000 On Modeling of Concept Based Retrieval in Generalized Vector Spaces
Minkoo Kim, Ali H. Alsaffar, Jitender S. Deogun, Vijay Raghavan 0001
ISMIS4
2000 Automatic Construction of Rule-Based Trees for Conceptual Retrieval
abstract
Many intelligent retrieval approaches have been studied to bridge the terminological gap existing between the way in which users specify their information needs and the way in which queries are expressed. One of the approaches, called RUBRIC (RUle-Based Retrieval of Information by Computer), uses production rules to capture user query concepts (or topics). A set of related production rules is represented as an AND/OR tree, called a rule-based tree. One of the main problems in this approach is how to construct such rules that can capture user query concepts. This paper provides a logical framework that is semantically essential to defining the rules for the user query concepts, and proposes a way to automatically construct rule-based trees from typical thesauri. Experiments performed on small collections with a domain-specific thesaurus show that the automatically constructed rules are more effective than hand-made rules in terms of precision.
Minkoo Kim, Fenghua Lu, Vijay Raghavan 0001
SPIRE3
2000 Enhancing Concept-Based Retrieval Based on Minimal Term Sets
Ali H. Alsaffar, Jitender S. Deogun, Vijay Raghavan 0001, Hayri Sever
J. Intell. Inf. Syst.3
1999 The Item-Set Tree: A Data Structure for Data Mining
Aladdin Hafez, Jitender S. Deogun, Vijay Raghavan 0001
DaWaK3
1999 Impact of Decision-Region Based Classification Mining Algorithms on Database Security
Tom Johnsten, Vijay Raghavan 0001
DBSec2
1999 Concept Based Retrieval by Minimal Term Sets
Ali H. Alsaffar, Jitender S. Deogun, Vijay Raghavan 0001, Hayri Sever
ISMIS3
1999 Improving Perceptron Convergence Algorithm for retrieval systems
Nasser Tadayon, Vijay Raghavan 0001
Pattern Recognit. Lett.2
1998 On the Necessity of Term Dependence in a Query Space for Weighted Retrieval
abstract
In recent years, in the context of the vector space model, the view, held by many researchers, that documents, queries, terms, etc., are all elements of a common space has been challenged (Bollmann-Sdorra & Raghavan, 1993). In particular, it was noted that term independence has to be investigated in the context of user preferences and it was shown, through counterexamples, that term independence can hold in the document space, but not in the query space and vice versa. In this article, we continue the investigation of query and document spaces with respect to the property of term independence. We prove, under realistic assumptions, that requiring term independence to hold in the query space is inconsistent with the goal of achieving better performance by means of weighted retrieval. The result that term independence in the query space is undesirable is obtained without making any assumption about whether or not the property of term independence holds in the document space. The results of this article reinforce our position that the properties of document and query spaces must be investigated separately, since the document and query spaces do not necessarily have the same properties.
Peter Bollmann-Sdorra, Vijay Raghavan 0001
J. Am. Soc. Inf. Sci.2
1998 Feature Selection and Effective Classifiers
abstract
In this article, we develop and analyze four algorithms for feature selection in the context of rough set methodology. The initial state and the feasibility criterion of all these algorithms are the same. That is, they start with a given feature set and progressively remove features, while controlling the amount of degradation in classification quality. These algorithms, however, differ in the heuristics used for pruning the search space of features. Our experimental results confirm the expected relationship between the time complexity of these algorithms and the classification accuracy of the resulting upper classifiers. Our experiments demonstrate that a θ-reduct of a given feature set can be found efficiently. Although we have adopted upper classifiers in our investigations, the algorithms presented can, however, be used with any method of deriving a classifier, where the quality of classification is a monotonically decreasing function of the size of the feature set. We compare the performance of upper classifiers with those of lower classifiers. We find that upper classifiers perform better than lower classifiers for a duodenal ulcer data set. This should be generally true when there is a small number of elements in the boundary region. An upper classifier has some important features that make it suitable for data mining applications. In particular, we have shown that the upper classifiers can be summarized at a desired level of abstraction by using extended decision tables. We also point out that an upper classifier results in an inconsistent decision algorithm, which can be interpreted deterministically or non-deterministically to obtain a consistent decision algorithm. © 1998 John Wiley & Sons, Inc.
Jitender S. Deogun, Suresh K. Choubey, Vijay Raghavan 0001, Hayri Sever
J. Am. Soc. Inf. Sci.3
1998 Introduction (Special Topic Issue: Knowledge Discovery and Data Mining)
Vijay Raghavan 0001, Jitender S. Deogun, Hayri Sever
J. Am. Soc. Inf. Sci.1
1997 Generic and Fully Automatic Content-Based Image Retrieval Architecture
Suresh K. Choubey, Vijay Raghavan 0001
ISMIS2
1997 Algorithms for the Boundary Selection Problem
Jay N. Bhuyan, Jitender S. Deogun, Vijay Raghavan 0001
Algorithmica3
1997 Modeling and retrieving images by content
Venkat N. Gudivada, Vijay Raghavan 0001
Inf. Process. Manag.2
1997 Generic and fully automatic content-based image retrieval using color
Suresh K. Choubey, Vijay Raghavan 0001
Pattern Recognit. Lett.2
1996 Measurement in Information Science, by Bert R. Boyce, Charles T. Meadow, and Donald H. Kraft
Vijay Raghavan 0001
J. Am. Soc. Inf. Sci.1
1995 Exploiting Upper Approximation in the Rough Set Methodology
Jitender S. Deogun, Vijay Raghavan 0001, Hayri Sever
KDD2
1995 On the Reuse of Past Optimal Queries
abstract
Retrieval(IR) systems exploit user feedback by generating an optimal query with respect to a particular information need.Since obtaining an optimal query is an expensive process, the need for mechanisms to save and reuse past optimal queries for future queries is obvions.In this article, we propose the use of a query base, a set of persistent past optimal queries, and investigate similarity measures between queries.The query base can be used either to answer user queries or to formulate optimal queries.We justify the former case analytically and the latter case by experiment.
Vijay Raghavan 0001, Hayri Sever
SIGIR1
1995 Conceptual Query Formulation and Retrieval
Sanjiv K. Bhatia, Jitender S. Deogun, Vijay Raghavan 0001
J. Intell. Inf. Syst.3
1995 Design and Evaluation of Algorithms for Image Retrieval by Spatial Similarity
abstract
Similarity-based retrieval of images is an important task in many image database applications. A major class of users' requests requires retrieving those images in the database that are spatially similar to the query image. We propose an algorithm for computing the spatial similarity between two symbolic images. A symbolic image is a logical representation of the original image where the image objects are uniquely labeled with symbolic names. Spatial relationships in a symbolic image are represented as edges in a weighted graph referred to as spatial-orientation graph. Spatial similarity is then quantified in terms of the number of, as well as the extent to which, the edges of the spatial-orientation graph of the database image conform to the corresponding edges of the spatial-orientation graph of the query image. The proposed algorithm is robust in the sense that it can deal with translation, scale, and rotational variances in images. The algorithm has quadratic time complexity in terms of the total number of objects in both the database and query images. We also introduce the idea of quantifying a system's retrieval quality by having an expert specify the expected rank ordering with respect to each query for a set of test queries. This enables us to assess the quality of algorithms comprehensively for retrieval in image databases. The characteristics of the proposed algorithm are compared with those of the previously available algorithms using a testbed of images. The comparison demonstrated that our algorithm is not only more efficient but also provides a rank ordering of images that consistently matches with the expert's expected rank ordering.
Venkat N. Gudivada, Vijay Raghavan 0001
ACM Trans. Inf. Syst.2
1994 Analysis of Common Subexpression Exploitation Models in Multiple-Query Processing
abstract
In multiple-query processing, a subexpression that appears in more than one query is called a common subexpression (CSE). A CSE needs to he evaluated once only to produce a temporary result that can then be used to evaluate all the queries containing the CSE. Therefore, the cost of evaluating the CSE is amortized over the queries requiring its evaluation. Two queries, posed simultaneously to the optimizer, may however contain subexpression that are not equivalent but are, nevertheless related by implication (the extension of one is a proper subset of the other) or intersection (the intersection of the two extensions is a proper subset of both extensions). In order to exploit the opportunity for cost amortization offered by the two latter relationships. the optimizer must rewrite the two queries in such a way that a CSE is induced. This paper compares, empirically and analytically, the performance of the various query execution models that are implied by different approaches to query rewriting.>
Jamal R. Alsabbagh, Vijay Raghavan 0001
ICDE2
1993 On the Delusiveness of Adopting a Common Space for Modeling IR Objects: Are Queries Documents
abstract
Many authors, who adopt the vector space model, take the view that documents, terms, queries, etc., are all elements within the same (conceptual) space. This view seems to be a natural one, given that documents and queries have the same vector notation. We show, however, that the structure of the query space can be very different from that of the document space. To this end, concepts like preference, similarity, term independence, and linearity, both in the document space and in the query space, are discussed. Our conclusion is that a more realistic and complete view of IR is obtained if we do not consider documents and queries to be elements of the same space. This conclusion implies that certain restrictions usually applied in the design of an IR system are obviated. For example, the retrieval function need not be restricted to the ones that have the possibility to be interpreted as a similarity measure. © 1993 John Wiley & Sons, Inc.
Peter Bollmann-Sdorra, Vijay Raghavan 0001
J. Am. Soc. Inf. Sci.2
1992 On Probabilistic Notions of Precision as a Function of Recall
Peter Bollmann-Sdorra, Vijay Raghavan 0001, Gwang S. Jung
Inf. Process. Manag.2
1991 Query Formulation Through Knowledge Acquisition
Sanjiv K. Bhatia, Jitender S. Deogun, Vijay Raghavan 0001
ML3
1991 A Probabilistic Retrieval Scheme for Cluster-based Adaptive Information Retrieval
Jay N. Bhuyan, Vijay Raghavan 0001
ML2
1991 User Profiles for Information Retrieval
Sanjiv K. Bhatia, Jitender S. Deogun, Vijay Raghavan 0001
ISMIS3
1991 An Object-Oriented Modelling of the History of Optimal Retrievals
abstract
Article An object-oriented modeling of the history of optimal retrievals Share on Authors: Yong Zhang Center for Advanced Computer Studies, University of Southwestern Louisiana, Lafayette, LA Center for Advanced Computer Studies, University of Southwestern Louisiana, Lafayette, LAView Profile , Vijay V. Raghavan Center for Advanced Computer Studies, University of Southwestern Louisiana, Lafayette, LA Center for Advanced Computer Studies, University of Southwestern Louisiana, Lafayette, LAView Profile , Jitender S. Deogun Department of Computer Science and Engineering, University of Nebraska, Lincoln, NE Department of Computer Science and Engineering, University of Nebraska, Lincoln, NEView Profile Authors Info & Claims SIGIR '91: Proceedings of the 14th annual international ACM SIGIR conference on Research and development in information retrievalSeptember 1991 Pages 241–250https://doi.org/10.1145/122860.122885Published:01 September 1991 1citation259DownloadsMetricsTotal Citations1Total Downloads259Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Vijay Raghavan 0001, Jitender S. Deogun
SIGIR2
1991 Efficient Algorithms For Selection of Recovery Points in Tree Task Models
abstract
Efficient solutions to the problem of optimally selecting recovery points are developed. The solutions are intended for models of computation in which task precedence has a tree structure and a task may fail due to the presence of faults. An algorithm to minimize the expected computation time of the task system under a uniprocessor environment has been developed for the binary tree model. The algorithm has time complexity of O(N/sub 2/), where N is the number of tasks, while previously reported procedures have exponential time requirements. The results are generalized for an arbitrary tree model.>
Subhada K. Mishra, Vijay Raghavan 0001, Nian-Feng Tzeng
IEEE Trans. Software Eng.2
1990 Design of an Integrated Information Retrieval/Database Management System
abstract
To increase the semantics available in a database system it is useful to integrate it with the well-understood models, leading to ranked response of queries that characterize information retrieval systems. A novel, unified architecture is presented which provides the flexibility of integrating any information retrieval (IR) model with any type of database management system. More importantly, the approach provides for the ability to use 'aggregation' and 'generalization' operations automatically to provide more meaningful responses. In addition, this framework makes possible the systematic investigation of the potential of using IR models as a general tool for supporting management decisions.>
Lawrence V. Saxton, Vijay Raghavan 0001
IEEE Trans. Knowl. Data Eng.2
1989 Retrieval System Evaluation Using Recall and Precision: Problems and Answers
abstract
article Free Access Share on Retrieval system evaluation using recall and precision: problems and answers Authors: V. V. Raghavan The Center for Advanced Computer Studies, University of Southwestern Louisiana,P.O. Box 44330, Lafayette, LA The Center for Advanced Computer Studies, University of Southwestern Louisiana,P.O. Box 44330, Lafayette, LAView Profile , P. Bollmann Technische Universitat Berlin, Fachbereich Informatik, FR 5-l1, Franklinstraβe 28/29, D-1000 Berlin 10 ,West Germany Technische Universitat Berlin, Fachbereich Informatik, FR 5-l1, Franklinstraβe 28/29, D-1000 Berlin 10 ,West GermanyView Profile , G. S. Jung The Center for Advanced Computer Studies, University of Southwestern Louisiana,P.O. Box 44330, Lafayette, LA The Center for Advanced Computer Studies, University of Southwestern Louisiana,P.O. Box 44330, Lafayette, LAView Profile Authors Info & Claims ACM SIGIR ForumVolume 23Issue SIJune 1989 pp 59–68https://doi.org/10.1145/75335.75342Published:01 May 1989Publication History 11citation1,293DownloadsMetricsTotal Citations11Total Downloads1,293Last 12 Months25Last 6 weeks11 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Vijay Raghavan 0001, Peter Bollmann-Sdorra, Gwang S. Jung
SIGIR1
1989 Extended Boolean query processing in the generalized vector space model
S. K. Michael Wong, Wojciech Ziarko, Vijay Raghavan 0001, P. C. N. Wong
Inf. Syst.3
1989 A Critical Investigation of Recall and Precision as Measures of Retrieval System Performance
abstract
Recall and precision are often used to evaluate the effectiveness of information retrieval systems. They are easy to define if there is a single query and if the retrieval result generated for the query is a linear ordering. However, when the retrieval results are weakly ordered, in the sense that several documents have an identical retrieval status value with respect to a query, some probabilistic notion of precision has to be introduced. Relevance probability, expected precision, and so forth, are some alternatives mentioned in the literature for this purpose. Furthermore, when many queries are to be evaluated and the retrieval results averaged over these queries, some method of interpolation of precision values at certain preselected recall levels is needed. The currently popular approaches for handling both a weak ordering and interpolation are found to be inconsistent, and the results obtained are not easy to interpret. Moreover, in cases where some alternatives are available, no comparative analysis that would facilitate the selection of a particular strategy has been provided. In this paper, we systematically investigate the various problems and issues associated with the use of recall and precision as measures of retrieval system performance. Our motivation is to provide a comparative analysis of methods available for defining precision in a probabilistic sense and to promote a better understanding of the various issues involved in retrieval performance evaluation.
Vijay Raghavan 0001, Gwang S. Jung, Peter Bollmann-Sdorra
ACM Trans. Inf. Syst.1
1988 A Utility-Theoretic Analysis of Expected Search Length
abstract
In this paper the expected search length, which is a measure of retrieval system performance, is investigated from the viewpoint of axiomatic utility theory. Necessary and sufficient criteria for the expected search length to be an ordinal scale and sufficient criteria that it is a ratio scale are given.
Peter Bollmann-Sdorra, Vijay Raghavan 0001
SIGIR2
1988 Integration of information retrieval and database management systems
Jitender S. Deogun, Vijay Raghavan 0001
Inf. Process. Manag.2
1987 Optimal Determination of User-Oriented Clusters
abstract
User-oriented clustering schemes enable the classification of documents based upon the user perception of the similarity between documents, rather than on some similarity function presumed by the designer to represent the user criteria. In this paper, an enhancement of such a clustering scheme is presented. This is accomplished by the formulation of the user-oriented clustering as a function-optimization problem. The problem formulated is termed the Boundary Selection Problem (BSP). Heuristic approaches to solve the BSP are proposed and a preliminary for evaluation of these approaches is provided.
Vijay Raghavan 0001, Jitender S. Deogun
SIGIR1
1987 Models of IR (Panel)
Vijay Raghavan 0001, M. Gordon, Robert R. Korfhage, Clement T. Yu
SIGIR1
1987 On Modeling of Information Retrieval Concepts in Vector Space
abstract
The Vector Space Model (VSM) has been adopted in information retrieval as a means of coping with inexact representation of documents and queries, and the resulting difficulties in determining the relevance of a document relative to a given query. The major problem in employing this approach is that the explicit representation of term vectors is not known a priori. Consequently, earlier researchers made the assumption that the vectors corresponding to terms are pairwise orthogonal. Such an assumption is clearly unrealistic. Although attempts have been made to compensate for this assumption by some separate, corrective steps, such methods are ad hoc and, in most cases, formally inconsistent. In this paper, a generalization of the VSM, called the GVSM, is advanced. The developments provide a solution not only for the computation of a measure of similarity (correlation) between terms, but also for the incorporation of these similarities into the retrieval process. The major strength of the GVSM derives from the fact that it is theoretically sound and elegant. Furthermore, experimental evaluation of the model on several test collections indicates that the performance is better than that of the VSM. Experiments have been performed on some variations of the GVSM, and all these results have also been compared to those of the VSM, based on inverse document frequency weighting. These results and some ideas for the efficient implementation of the GVSM are discussed.
S. K. Michael Wong, Wojciech Ziarko, Vijay Raghavan 0001, P. C. N. Wong
ACM Trans. Database Syst.3
1986 User-Oriented Document Clustering: A Framework for Learning in Information Retrieval
abstract
In information retrieval, cluster analysis is an important tool employed to enhance both efficiency and effectiveness of the retrieval process. Most clustering algorithms have difficulty in reflecting the closeness of documents as perceived by the user. A two phase scheme for document clustering, whose results reflect the “conceptual” clusters that are perceived by the user of the retrieval system, is proposed. Since the clusters obtained by this scheme are not characterized in terms of the document representations, a strategy for cluster searching is also developed. Both the proposed document clustering scheme and document searching strategy are experimentally evaluated using a test collection from the SMART system. The preliminary experimental results obtained are very encouraging.
Jitender S. Deogun, Vijay Raghavan 0001
SIGIR2
1986 On Extending the Vector Space Model for Boolean Query Processing
abstract
An information retrieval model, named the Generalized Vector Space Model (GVSM), is extended to handle situations where queries are specified as (extended) Boolean expressions. It is shown that this unified model, unlike currently available alternatives, has the advantage of incorporating term correlations into the retrieval process. The query language extension is attractive in the sense that most of the algebraic properties of the strict Boolean language are still preserved. Although the experimental results for extended Boolean retrieval are not always better than the vector processing method, the developments here are significant in facilitating commercially available retrieval systems to benefit from the vector based methods. The proposed scheme is compared to the p-norm model advanced by Salton and coworkers. An important conclusion is that it is desirable to investigate further extensions that can offer the benefits of both proposals.
S. K. Michael Wong, Wojciech Ziarko, Vijay Raghavan 0001, P. C. N. Wong
SIGIR3
1986 A critical analysis of vector space model for information retrieval
abstract
Notations and definitions necessary to identify the concepts and relationships that are important in modelling information retrieval objects and processes in the context of vector spaces are presented. Earlier work on the use of vector model is evaluated in terms of the concepts introduced and certain problems and inconsistencies are identified. More importantly, this investigation should lead to a clear understanding of the issues and problems in using the vector space model in information retrieval. © 1986 John Wiley & Sons, Inc.
Vijay Raghavan 0001, S. K. Michael Wong
J. Am. Soc. Inf. Sci.1
1984 Vector Space Model of Information Retrieval - A Reevaluation
S. K. Michael Wong, Vijay Raghavan 0001
SIGIR2
1984 Organization of Clustered Files for Consecutive Retrieval
abstract
This paper studies the problem of storing single-level and multilevel clustered files. Necessary and sufficient conditions for a single-level clustered file to have the consecutive retrieval property (CRP) are developed. A linear time algorithm to test the CRP for a given clustered file and to identify the proper arrangement of objects, if CRP exists, is presented. For the single-level clustered files that do not have CRP, it is shown that the problem of identifying a storage organization with minimum redundancy is NP-complete. Consequently, an efficient heuristic algorithm to generate a good storage organization for such files is developed. Furthermore, it is shown that, for certain types of multilevel clustered files, there exists a storage organization such that the objects in each cluster, for all clusters in each level of the clustering, appear in consecutive locations.
Jitender S. Deogun, Vijay Raghavan 0001, Thomas K. W. Tsou
ACM Trans. Database Syst.2
1983 Evaluation of The 2-Poisson Model as a Basis for Using Term Frequency Data in Searching
abstract
The early work on the probabilistic models of retrieval assumed that the document representation is binary, indicating only the presence or absence of index terms. The 2-Poisson (TP) model which was proposed as a model of how the occurrence frequency of specialty words in a collection is distributed, has since been used to develop retrieval strategies that incorporate term frequency information. This work investigates the use of the TP model, in this context, further. It is shown that the search effectiveness, when no relevance information is assumed, can be further enhanced by using this model. Furthermore, when the term weights proposed in this work are used in conjunction with weights known as term significance weights, the results are very encouraging.
Vijay Raghavan 0001, Hong-Pao Shi, Clement T. Yu
SIGIR1
1983 On the Selection of an Optimal Set of Indexes
abstract
A problem of considerable interest in the design of a database is the selection of indexes. In this paper, we present a probabilistic model of transactions (queries, updates, insertions, and deletions) to a file. An evaluation function, which is based on the cost saving (in terms of the number of page accesses) attributable to the use of an index set, is then developed. The maximization of this function would yield an optimal set of indexes. Unfortunately, algorithms known to solve this maximization problem require an order of time exponential in the total number of attributes in the file. Consequently, we develop the theoretical basis which leads to an algorithm that obtains a near optimal solution to the index selection problem in polynomial time. The theoretical result consists of showing that the index selection problem can be solved by solving a properly chosen instance of the knapsack problem. A theoretical bound for the amount by which the solution obtained by this algorithm deviates from the true optimum is provided. This result is then interpreted in the light of evidence gathered through experiments.
Maggie Y. L. Ip, Lawrence V. Saxton, Vijay Raghavan 0001
IEEE Trans. Software Eng.3
1982 Techniques for Measuring the Stability of Clustering: A Comparative Study
Vijay Raghavan 0001, Maggie Y. L. Ip
SIGIR1
1981 A Comparison of the Stability Characteristics of Some Graph Theoretic Clustering Methods
abstract
Assessing the stability of a clustering method involves the measurement of the extent to which the generated clusters are affected by perturbations in the input data. A measure which specifies the disturbance in a set of clusters as the minimum number of operations required to restore the set of modified clusters to the original ones is adopted. A number of well-known graph theoretic clustering methods are compared in terms of their stability as determined by this measure. Specifically, it is shown that among the clustering methods in any of several families of graph theoretic methods, clusters defined as the connected components are the most stable and the clusters specified as the maximal complete subgraphs are the least stable. Furthermore, as one proceeds from the method producing the most narrow clusters (maximal complete subgraphs) to those producing relatively broader clusters, the clustering process is shown to remain at least as stable as any method in the previous stages. Finally, the lower and the upper bounds for the measure of stability, when clusters are defined as the connected components, are derived.
Vijay Raghavan 0001, Clement T. Yu
IEEE Trans. Pattern Anal. Mach. Intell.1
1979 A Clustering Strategy Based on a Formalism of the Reproductive Process in Natural Systems
abstract
Given a set of objects each of which is represented by a finite number of attributes or features and a clustering criterion that associates a value of utility to any classification, the objective of a clustering method is to identify that classification of the objects which optimizes the criterion. A new strategy to solve this problem is developed. The approach is, in essence, a modification of the reproductive plan, a type of adaptive procedure devised by Holland [2], which embodies many principles found in the adaptation of natural systems through evolution. The proposed approach differs from conventional methods in the sense that the search through the space of possible solutions proceeds in a parallel fashion.The adaptive clustering strategy requires the specification of methods for the generation of an initial population of classifications, the parent selection, the modifications and the replacement of current classifications with new ones. The effects of changing several of these features are investigated. Experimental results show that it is possible to devise clustering strategies based on the principles of adaptation in natural systems that are both effective and efficient.
Vijay Raghavan 0001, Kim Birchard
SIGIR1
1979 Experiments on the Determination of the Relationships Between Terms
abstract
The retrieval effectiveness of an automatic method that uses relevance judgments for the determination of positive as well as negative relationships between terms is evaluated. The term relationships are incorporated into the retrieval process by using a generalized similarity function that has a term match component, a positive term relationship component, and a negative term relationship component. Two strategies, query partitioning and query clustering, for the evaluation of the effectiveness of the term relationships are investigated. The latter appears to be more attractive from linguistic as well as economic points of view. The positive and the negative relationships are verified to be effective both when used individually, and in combination. The importance attached to the term relationship components relative to that of term match component is found to have a substantial effect on the retrieval performance. The usefulness of discriminant analysis as a technique for determining the relative importance of these components is investigated.
Vijay Raghavan 0001, Clement T. Yu
ACM Trans. Database Syst.1
1978 Experiments on the Determination of the Reltionships Between Terms
abstract
The retrieval effectiveness of an automatic method that uses relevance judgements for the determination of positive as well as negative relationships between terms is evaluated. The term relationships are incorporated into the retrieval process by using a generalized similarity function that has a term match component, a positive term relationship component, and a negative term relationship component. Two strategies, query partitioning and query clustering, for the evaluation of the effectiveness of the term relationships are investigated. The latter appears to be more attractive from linguistic as well as economic points of view. The positive and the negative relationships are verified to be effective both when used individually, and in combination. The importance attached to the term relationship components relative to that of term match component is found to have a substantial effect on the retrieval performance. The usefulness of discriminant analysis as a technique for determining the relative importance of these components is investigated.
Clement T. Yu, Vijay Raghavan 0001
SIGIR2
1977 A Note on a Multidimensional Searching Problem
Vijay Raghavan 0001, Clement T. Yu
Inf. Process. Lett.1
1977 Single-pass method for determining the semantic relationships between terms
abstract
Abstract A fast single‐pass method for the automatic determination of the semantic relationships between terms is presented. The computing time required for the method is small enough for it to be feasible in a practical environment. The experimental results obtained indicate that the method is effective in the assessment of the relationships between terms. The improvement achieved in retrieval performance over simple keyword matching is quite significant.
Clement T. Yu, Vijay Raghavan 0001
J. Am. Soc. Inf. Sci.2