Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shlomo Geva

dblp:20/3245 · DBLP profile ↗
← Back
47ranked-venue papers
9as first author
1since 2021 · last 2021
0000-0003-1340-2802ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 24 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 17 · 7 first-authorApplied, interdisciplinary, general and emerging computing · 9Graphics, computer vision, multimedia, augmented reality and games · 5Human-computer interaction and ubiquitous computing · 5Systems, architecture and hardware · 1Security and privacy · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Information retrieval · 61% Data mining · 22% Data integration and cleaning · 15%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational social science and digital humanities · 78% Computational finance and economics · 22%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
evaluation
0.422017
A Test Collection for Evaluating Retrieval of Studies for Inclusion in Systematic Reviews · SIGIR 2017
The importance of manual assessment in link discovery · SIGIR 2009
Data mining › clustering
document clustering
0.322015
Parallel Streaming Signature EM-tree: A Clustering Algorithm for Web Scale Applications · WWW 2015
K-tree: large scale document clustering · SIGIR 2009
Information retrieval › document retrieval
domain-specific retrieval
0.312017
A Test Collection for Evaluating Retrieval of Studies for Inclusion in Systematic Reviews · SIGIR 2017
Information retrieval › evaluation
test collection
0.312017
A Test Collection for Evaluating Retrieval of Studies for Inclusion in Systematic Reviews · SIGIR 2017
Computational social science and digital humanities
social media analysis
0.212016
WIMBY: What's in My Backyard? · ACM Multimedia 2016
Information retrieval › image retrieval
content-based image retrieval
0.212016
Improving Retrieval Quality Using Pseudo Relevance Feedback in Content-Based Image Retrieval · SIGIR 2016
Data integration and cleaning › data fusion
multi-source data fusion
0.212016
WIMBY: What's in My Backyard? · ACM Multimedia 2016
Information retrieval › relevance feedback
pseudo-relevance feedback
0.212016
Improving Retrieval Quality Using Pseudo Relevance Feedback in Content-Based Image Retrieval · SIGIR 2016
Information retrieval
retrieval models
0.212016
Improving Retrieval Quality Using Pseudo Relevance Feedback in Content-Based Image Retrieval · SIGIR 2016
Data mining
clustering
0.212015
Parallel Streaming Signature EM-tree: A Clustering Algorithm for Web Scale Applications · WWW 2015
Data integration and cleaning
entity resolution
0.112009
The importance of manual assessment in link discovery · SIGIR 2009
Data mining › clustering
hierarchical clustering
0.112009
K-tree: large scale document clustering · SIGIR 2009
Data integration and cleaning
link discovery
0.112009
The importance of manual assessment in link discovery · SIGIR 2009
Information retrieval › retrieval models
boolean retrieval
0.112017
A Test Collection for Evaluating Retrieval of Studies for Inclusion in Systematic Reviews · SIGIR 2017
Computational finance and economics
financial forecasting
0.112007
Can the Content of Public News Be Used to Forecast Abnormal Stock Market Behaviour? · ICDM 2007
Web and social media mining
web page analysis
0.112015
Parallel Streaming Signature EM-tree: A Clustering Algorithm for Web Scale Applications · WWW 2015
Data mining
text mining
0.012007
Can the Content of Public News Be Used to Forecast Abnormal Stock Market Behaviour? · ICDM 2007

Methods — techniques the papers use, named apart from their topics

spatiotemporal analysis · 0.8data fusion · 0.8boolean retrieval · 0.3binary image signature · 0.2parallel streaming · 0.2compressed document representation · 0.2volatility modeling · 0.1machine learning for link discovery · 0.1k-tree · 0.1
YearPublicationVenuePosition
2021 Mining discriminative itemsets in data streams using the tilted-time window model
Majid Seyfi, Richi Nayak, Yue Xu 0001, Shlomo Geva
Knowl. Inf. Syst.4
2019 High Resolution Change Detection Using Planet Mosaic
abstract
This paper presents a change detection system using Planet's mosaic dataset. This dataset has higher resolution but fewer bands than data captured from Landsat or Sentinel satellites. Here, an object-based random forest regressor is used to detect vegetation change. The mosaic was separated into individual images, which were then ranked in a list. Results indicate adequate performance with a tradeoff in precision and recall.
Alan Woodley, Connor McLaughlin, Holly Hutson, Shlomo Geva, Timothy Chappell, Wayne Kelly, Dimitri Perrin, Wageeh W. Boles, Lance De Vine
IGARSS4
2019 Parallel K-Tree: A multicore, multinode solution to extreme clustering
Alan Woodley, Ling-Xiang Tang, Shlomo Geva, Richi Nayak, Timothy Chappell
Future Gener. Comput. Syst.3
2018 Rapid analysis of metagenomic data using signature-based clustering
abstract
BACKGROUND: Sequencing highly-variable 16S regions is a common and often effective approach to the study of microbial communities, and next-generation sequencing (NGS) technologies provide abundant quantities of data for analysis. However, the speed of existing analysis pipelines may limit our ability to work with these quantities of data. Furthermore, the limited coverage of existing 16S databases may hamper our ability to characterise these communities, particularly in the context of complex or poorly studied environments. RESULTS: In this article we present the SigClust algorithm, a novel clustering method involving the transformation of sequence reads into binary signatures. When compared to other published methods, SigClust yields superior cluster coherence and separation of metagenomic read data, while operating within substantially reduced timeframes. We demonstrate its utility on published Illumina datasets and on a large collection of labelled wound reads sourced from patients in a wound clinic. The temporal analysis is based on tracking the dominant clusters of wound samples over time. The analysis can identify markers of both healing and non-healing wounds in response to treatment. Prominent clusters are found, corresponding to bacterial species known to be associated with unfavourable healing outcomes, including a number of strains of Staphylococcus aureus. CONCLUSIONS: SigClust identifies clusters rapidly and supports an improved understanding of the wound microbiome without reliance on a reference database. The results indicate a promising use for a SigClust-based pipeline in wound analysis and prediction, and a possible novel method for wound management and treatment.
Timothy Chappell, Shlomo Geva, James M. Hogan, Flavia Huygens, Irani U. Rathnayake, Stephen Rudd, Wayne Kelly, Dimitri Perrin
BMC Bioinform.2
2017 Signature-based clustering for analysis of the wound microbiome
abstract
Chronic wounds present a significant risk to the patient and a substantial drain on health budgets, with the problem likely to worsen markedly with increased incidence of type II diabetes. The wound fluid microbiome is known to influence wound healing outcomes, but is poorly characterised. Next Generation Sequencing approaches yield abundant data from wound samples, but progress in understanding these microbial communities may be hampered by the speed of existing analysis pipelines and limitations on coverage by 16S databases. This paper presents SigClust, a novel clustering method based on binary signatures derived from sequence reads. SigClust yields superior cluster coherence and separation of metagenomic read data in timeframes substantially reduced from those of alternative methods. We demonstrate its utility in the wound context on a preliminary set of labelled patient data. We show how a time course analysis based on tracking the dominant clusters over successive wound samples can identify markers of both successful wound healing and wounds refractory to treatment. Clusters prominent in these analyses are found to correspond to bacterial species known to be implicated as a determinant of wound outcomes, notably a number of strains of Staphylococcus aureus. The clusters obtained rapidly via SigClust support improved understanding of the wound microbiome without direct reliance on a reference database, offering the promise of a SigClust-based pipeline for wound analysis and prediction, and potentially novel methods for wound treatment and management.
Timothy Chappell, Shlomo Geva, James M. Hogan, Flavia Huygens, Wayne Kelly, Dimitri Perrin
BIBM2
2017 Integrating the Framing of Clinical Questions via PICO into the Retrieval of Medical Literature for Systematic Reviews
abstract
The PICO process is a technique used in evidence based practice to frame and answer clinical questions. It involves structuring the question around four types of clinical information: population, intervention, control or comparison and outcome. The PICO framework is used extensively in the compilation of systematic reviews as the means of framing research questions. However, when a search strategy (comprising of a large Boolean query) is formulated to retrieve studies for inclusion in the review, PICO is often ignored. This paper evaluates how PICO annotations can be applied and integrated into retrieval to improve the screening of studies for inclusion in systematic reviews. The task is to increase precision while maintaining the high level of recall essential to ensure systematic reviews are representative and unbiased. Our results show that restricting the search strategies to match studies using PICO annotations improves precision, however recall is slightly reduced, when compared to the non-PICO baseline. This can lead to both time and cost savings when compiling systematic reviews.
Harrisen Scells, Guido Zuccon, Bevan Koopman, Anthony Deacon, Leif Azzopardi, Shlomo Geva
CIKM6
2017 Similarity Projection: A Geometric Measure for Comparison of Biological Sequences
abstract
Sequence comparison is a fundamental task in computational biology, traditionally dominated by alignment-based methods such as the Smith-Waterman and Needleman-Wunsch algorithms, or by alignment based heuristics such as BLAST, the ubiquitous Basic Local Alignment Search Tool. For more than a decade researchers have examined a range of alignment-free alternatives to these approaches, citing concerns over scalability in the era of Next Generation Sequencing, the emergence of petascale sequence archives, and a lack of robustness of alignment methods in the face of structural sequence rearrangements. While some of these approaches have proven successful for particular tasks, many continue to exhibit a marked decline in sensitivity as closely related sequence sets diverge. Avoiding the alignment step allows the methods to scale to the challenges of modern sequence collections, but only at the cost of noticeably inferior search. In this paper we re-examine the problem of similarity measures for alignment-free sequence comparison, and introduce a new method which we term Similarity Projection. Similarity Projection offers markedly enhanced sensitivity - comparable to alignment based methods - while retaining the scalability characteristic of alignment-free approaches. As before, we rely on collections of k-mers; overlapping substrings of the molecular sequence of length k, collected without reference to position, but similarity relies on variants of the Hausdorff set distance, allowing similarity to be scored more effectively to the reflect those components which match, while lessening the impact of those which do not. Formally, the algorithm generates a large mutual similarity matrix between sequence pairs based on their component fragments; successive reduction steps yield a final score over the sequences. However, only a small fraction of these underlying comparisons need be performed, and by use of an approximate scheme based on vector quantization, we are able to achieve an order of magnitude improvement in execution time over the naive approach. We evaluate the approach on two large protein collections obtained from UniProtKB, showing that Similarity Projection achieves accuracy rivalling, and at times clearly exceeding, that of BLAST, while exhibiting markedly superior execution speed.
Lawrence Buckingham, Timothy Chappell, James M. Hogan, Shlomo Geva
eScience4
2017 A Test Collection for Evaluating Retrieval of Studies for Inclusion in Systematic Reviews
abstract
This paper introduces a test collection for evaluating the effectiveness of different methods used to retrieve research studies for inclusion in systematic reviews. Systematic reviews appraise and synthesise studies that meet specific inclusion criteria. Systematic reviews intended for a biomedical science audience use boolean queries with many, often complex, search clauses to retrieve studies; these are then manually screened to determine eligibility for inclusion in the review. This process is expensive and time consuming. The development of systems that improve retrieval effectiveness will have an immediate impact by reducing the complexity and resources required for this process. Our test collection consists of approximately 26 million research studies extracted from the freely available MEDLINE database, 94 review (query) topics extracted from Cochrane systematic reviews, and corresponding relevance assessments. Tasks for which the collection can be used for information retrieval system evaluation are described and the use of the collection to evaluate common baselines within one such task is demonstrated. The test collection is available at https://github.com/ielab/SIGIR2017-PICO-Collection.
Harrisen Scells, Guido Zuccon, Bevan Koopman, Anthony Deacon, Leif Azzopardi, Shlomo Geva
SIGIR6
2017 Efficient mining of discriminative itemsets
abstract
Discriminative itemsets can be more useful than frequent itemsets as the former identifies the frequent itemsets in one dataset with much higher frequencies than the same itemsets in other datasets. The discriminative itemsets can distinguish the target dataset from all others. The discriminative itemsets are a small subset of frequent itemsets. The efficient mining of discriminative itemsets is a challenging problem, since the Apriori property of frequent itemsets is not applicable, and the designed algorithms must deal with the exponential number of itemset combinations in more than one dataset. In this paper, a novel algorithm, called DISSparse, is proposed for efficient mining of discriminative itemsets. Two determinative heuristics are proposed for limiting the mining of discriminative itemsets to the potential discriminative itemsets. Our experiments show the efficient time and space usage of the proposed algorithm in the large and complex datasets.
Majid Seyfi, Richi Nayak, Yue Xu 0001, Shlomo Geva
WI4
2017 A greater understanding of social networks privacy requirements: The user perspective
Mohammad Badiul Islam, Jason Watson, Renato Iannella, Shlomo Geva
J. Inf. Secur. Appl.4
2016 Using parallel hierarchical clustering to address spatial big data challenges
abstract
Clustering can help to make large datasets more manageable by grouping together similar objects. However, most clustering approaches are unable to scale to very large datasets (e.g. more than 10 million objects). The K-Tree is a data structure and clustering algorithm that has proven to be scalable with large streaming datasets. Here, we apply the K-Tree to spatial data (satellite images) and extend from a single threaded to a multicore environment. We show that the K-Tree is able to cluster larger dataset more efficiently than baseline approaches.
Alan Woodley, Ling-Xiang Tang, Shlomo Geva, Richi Nayak, Timothy Chappell
IEEE BigData3
2016 WIMBY: What's in My Backyard?
abstract
Location-aware social media is increasing being used to inform decisions in a spatiotemporal context. However, collecting, fusing, processing and merging information from different social media platforms is a challenge because of diversity of information between different platforms. Here, we present the WIMBY, which is able to access multiple social media platforms to help users answer the question "What's in my Backyard?". In doing so, the WIMBY helps to address the challenge of dealing with diverse social media information. It is believed that the WIMBY can be extended to include more information sources (including other social media platforms) to help inform decision makers in a wider array of applications.
Michael Dorkhom, Alan Woodley, Shlomo Geva, Richi Nayak
ACM Multimedia3
2016 Improving Retrieval Quality Using Pseudo Relevance Feedback in Content-Based Image Retrieval
abstract
The increased availability of image capturing devices has enabled collections of digital images to rapidly expand in both size and diversity. This has created a constantly growing need for efficient and effective image browsing, searching, and retrieval tools. Pseudo-relevance feedback (PRF) has proven to be an effective mechanism for improving retrieval accuracy. An original, simple yet effective rank-based PRF mechanism (RB-PRF) that takes into account the initial rank order of each image to improve retrieval accuracy is proposed. This RB-PRF mechanism innovates by making use of binary image signatures to improve retrieval precision by promoting images similar to highly ranked images and demoting images similar to lower ranked images. Empirical evaluations based on standard benchmarks, namely Wang, Oliva & Torralba, and Corel datasets demonstrate the effectiveness of the proposed RB-PRF mechanism in image retrieval.
Dinesha Chathurani Nanayakkara Wasam Uluwitige, Timothy Chappell, Shlomo Geva, Vinod Chandran
SIGIR3
2015 Approximate Nearest-Neighbour Search with Inverted Signature Slice Lists
Timothy Chappell, Shlomo Geva, Guido Zuccon
ECIR2
2015 Learning Higher-Order Interactions for User and Item Profiling Based on Tensor Factorization
abstract
User profiling techniques play a central role in many Recommender Systems (RS). In recent years, multidimensional data are getting increasing attention for making recommendations. Additional metadata help algorithms better understanding users' behaviors and decisions. Existing user/item profiling techniques for Collaborative Filtering (CF) RS in multidimensional environment mostly analyze data through splitting the multidimensional relations. However, this leads to the loss of multidimensionality in user-item interactions; whereas the interactions are naturally multidimensional since users' choices are often affected by contextual information. In this paper, we propose a unified profiling approach which models users/items with latent higher-order interaction factors. We demonstrate that the proposed profiling approach is intimately related to two-dimensional profiling based on Matrix Factorization techniques. We further propose to integrate the profiling approach into three neighborhood-based CF recommenders for item recommendation. Finally, we empirically show on real-world social tagging datasets that the proposed recommenders outperform state-of-the-art CF recommendation approaches in accuracy.
Yue Xu 0001, Shlomo Geva
IUI3
2015 Parallel Streaming Signature EM-tree: A Clustering Algorithm for Web Scale Applications
abstract
The proliferation of the web presents an unsolved problem of automatically analyzing billions of pages of natural language. We introduce a scalable algorithm that clusters hundreds of millions of web pages into hundreds of thousands of clusters. It does this on a single mid-range machine using efficient algorithms and compressed document representations. It is applied to two web-scale crawls covering tens of terabytes. ClueWeb09 and ClueWeb12 contain 500 and 733 million web pages and were clustered into 500,000 to 700,000 clusters. To the best of our knowledge, such fine grained clustering has not been previously demonstrated. Previous approaches clustered a sample that limits the maximum number of discoverable clusters. The proposed EM-tree algorithm uses the entire collection in clustering and produces several orders of magnitude more clusters than the existing algorithms. Fine grained clustering is necessary for meaningful clustering in massive collections where the number of distinct topics grows linearly with collection size. These fine-grained clusters show an improved cluster quality when assessed with two novel evaluations using ad hoc search relevance judgments and spam classifications for external validation. These evaluations solve the problem of assessing the quality of clusters where categorical labeling is unavailable and unfeasible.
Christopher M. De Vries, Lance De Vine, Shlomo Geva, Richi Nayak
WWW3
2014 Refining User and Item Profiles based on Multidimensional Data for Top-N Item Recommendation
abstract
In recommender systems based on multidimensional data, additional metadata provides algorithms with more information for better understanding the interaction between users and items. However, most of the profiling approaches in neighbourhood-based recommendation approaches for multidimensional data merely split or project the dimensional data and lack the consideration of latent interaction between the dimensions of the data. In this paper, we propose a novel user/item profiling approach for Collaborative Filtering (CF) item recommendation on multidimensional data. We further present incremental profiling method for updating the profiles. For item recommendation, we seek to delve into different types of relations in data to understand the interaction between users and items more fully, and propose three multidimensional CF recommendation approaches for top-N item recommendations based on the proposed user/item profiles. The proposed multidimensional CF approaches are capable of incorporating not only localized relations of user-user and/or item-item neighbourhoods but also latent interaction between all dimensions of the data. Experimental results show significant improvements in terms of recommendation accuracy.
Yue Xu 0001, Shlomo Geva
iiWAS3
2014 Mining Discriminative Itemsets in Data Streams
Majid Seyfi, Shlomo Geva, Richi Nayak
WISE (1)2
2014 An evaluation framework for cross-lingual link discovery
Ling-Xiang Tang, Shlomo Geva, Andrew Trotman, Yue Xu 0001, Kelly Y. Itakura
Inf. Process. Manag.2
2013 What Makes an LMS Effective - A Synthesis of Current Literature
abstract
There is a growing number of organizations and universities now utilising e-learning practices in their teaching and learning programs.These systems have allowed for knowledge sharing and provide opportunities for users to have access to learning materials regardless of time and place.However, while the uptake of these systems is quite high, there is little research into the effectiveness of such systems, particularly in higher education.This paper investigates the methods that are used to study the effectiveness of e-learning systems and the factors that are critical for the success of a learning management system (LMS).Five major success categories are identified in this study and explained in depth.These are the teacher, student, LMS design, learning materials and external support.
Nastaran Zanjani, Shaun Nykvist, Shlomo Geva
CSEDU3
2012 Do students and lecturers actively use collaboration tools in learning management systems?
abstract
In recent years there has been a large emphasis placed on the need to use Learning Management Systems (LMS) in the field of higher education, with many universities mandating their use. An important aspect of these systems is their ability to offer collaboration tools to build a community of learners. This paper reports on a study of the effectiveness of an LMS (Blackboard©) in a higher education setting and whether both lecturers and students voluntarily use collaborative tools for teaching and learning. Interviews were conducted with participants (N=67) from the faculties of Science and Technology, Business, Health and Law. Results from this study indicated that participants often use Blackboard© as an online repository of learning materials and that the collaboration tools of Blackboard© are often not utilised. The study also found that several factors have inhibited the use and uptake of the collaboration tools within Blackboard©. These have included structure and user experience, pedagogical practice, response time and a preference for other tools.
Nastaran Zanjani, Shaun Nykvist, Shlomo Geva
ICCE3
2011 TOPSIG: topology preserving document signatures
abstract
Comparisons between file signatures and inverted files for text retrieval have shown the shortcomings of traditional file signatures. It has been widely accepted that traditional file signatures are inferior alternatives to inverted files. This paper describes TopSig, a new approach to the construction of file signatures that extends recent advances in semantic hashing and dimensionality reduction. These were not so far linked to general purpose, signature file based, search engines. We demonstrate significant improvements in the performance of signature file based indexing and retrieval. Performance is comparable to the state of the art inverted file based systems, including language models and BM25. These findings suggest that file signatures offer a viable alternative to inverted files in suitable settings and positions the file signatures model in the class of Vector Space retrieval models.
Shlomo Geva, Christopher M. De Vries
CIKM1
2011 Topical and Structural Linkage in Wikipedia
Kelly Y. Itakura, Charles L. A. Clarke, Shlomo Geva, Andrew Trotman, Wei Chi Huang
ECIR3
2010 Mid-Level Concept Learning with Visual Contextual Ontologies and Probabilistic Inference for Image Annotation
Yuee Liu, Jinglan Zhang, Dian Tjondronegoro, Shlomo Geva, Zhengrong Li
MMM4
2010 Using Association Rules to Solve the Cold-Start Problem in Recommender Systems
Gavin Shaw, Yue Xu 0001, Shlomo Geva
PAKDD (1)3
2010 Ontology-Based Specific and Exhaustive User Profiles for Constraint Information Fusion for Multi-agents
abstract
Intelligent agents are an advanced technology utilized in Web Intelligence. When searching information from a distributed Web environment, information is retrieved by multi-agents on the client site and fused on the broker site. The current information fusion techniques rely on cooperation of agents to provide statistics. Such techniques are computationally expensive and unrealistic in the real world. In this paper, we introduce a model that uses a world ontology constructed from the Dewey Decimal Classification to acquire user profiles. By search using specific and exhaustive user profiles, information fusion techniques no longer rely on the statistics provided by agents. The model has been successfully evaluated using the large INEX data set simulating the distributed Web environment.
Xiaohui Tao 0001, Yuefeng Li 0001, Raymond Y. K. Lau, Shlomo Geva
Web Intelligence4
2010 Current research in focused retrieval and result aggregation
Andrew Trotman, Shlomo Geva, Jaap Kamps, Mounia Lalmas-Roelleke, Vanessa Murdock 0001
Inf. Retr.2
2009 The importance of manual assessment in link discovery
abstract
Using a ground truth extracted from the Wikipedia, and a ground truth created through manual assessment, we show that the apparent performance advantage seen in machine learning approaches to link discovery are an artifact of trivial links that are actively rejected by manual assessors.
Wei Che Huang, Andrew Trotman, Shlomo Geva
SIGIR3
2009 K-tree: large scale document clustering
abstract
We introduce K-tree in an information retrieval context. It is an efficient approximation of the k-means clustering algorithm. Unlike k-means it forms a hierarchy of clusters. It has been extended to address issues with sparse representations. We compare performance and quality to CLUTO using document collections. The K-tree has a low time complexity that is suitable for large document collections. This tree structure allows for efficient disk based implementations where space requirements exceed that of main memory.
Christopher M. De Vries, Shlomo Geva
SIGIR2
2008 Deriving non-redundant approximate association rules from hierarchical datasets
abstract
Association rule mining plays an important job in knowledge and information discovery. However, there are still shortcomings with the quality of the discovered rules and often the number of discovered rules is huge and contain redundancies, especially in the case of multi-level datasets. Previous work has shown that the mining of non-redundant rules is a promising approach to solving this problem, with work by [6,8,9,10] focusing on single level datasets. Recent work by Shaw et. al. [7] has extended the nonredundant approaches presented in [6,8,9] to include the elimination of redundant exact basis rules from multi-level datasets. Here we propose a continuation of the work in [7] that allows for the removal of hierarchically redundant approximate basis rules from multi-level datasets by using a dataset’s hierarchy or taxonomy.
Gavin Shaw, Yue Xu 0001, Shlomo Geva
CIKM3
2008 Extracting Non-redundant Approximate Rules from Multi-level Datasets
abstract
Association rule mining plays an important job in knowledge and information discovery. Often the number of the discovered rules is huge and many of them are redundant, especially for multi-level datasets. Previous work has shown that the mining of non-redundant rules is a promising approach to solving this problem, with work in focusing on single level datasets. Recent work by Shaw et. al. has extended the non-redundant approaches presented in to include the elimination of redundant exact basis rules from multi-level datasets. In this paper, we propose an extension to the work in to allow for the removal of hierarchically redundant approximate basis rules from multi-level datasets through the use of the datasetpsilas hierarchy or taxonomy. Experimentation shows our approach can effectively generate both multi-level and cross level non-redundant rule sets which are lossless.
Gavin Shaw, Yue Xu 0001, Shlomo Geva
ICTAI (2)3
2007 Can the Content of Public News Be Used to Forecast Abnormal Stock Market Behaviour?
abstract
A popular theory of markets is that they are efficient: all available information is deemed to provide an accurate valuation of an asset at any time. In this paper, we consider how the content of market- related news articles contributes to such information. Specifically, we mine news articles for terms of interest, and quantify this degree of interest. We then incorporate this measure into traditional models for market index volatility with a view to forecasting whether the incidence of interesting news is correlated with a shock in the index, and thus if the information can be captured to value the underlying asset. We illustrate the methodology on stock market indices for the USA, the UK, and Australia.
Calum S. Robertson, Shlomo Geva, Rodney C. Wolff
ICDM2
2007 Collection Profiling for Collection Fusion in Distributed Information Retrieval Systems
Chengye Lu, Yue Xu 0001, Shlomo Geva
KSEM3
2005 Secure email-based peer to peer information retrieval
abstract
In this paper, we describe the implementation of a distributed search engine called SEGPX based on email communication. SEGPX uses X.509 public-key and attribute certificate frameworks and utilizes email servers for communications. Advanced features, such as store and forward, document encryption and client side user authentication/authorization etc., are provided to adapt the use for secure group collaboration over the public networks.
Chengye Lu, Shlomo Geva
CW2
2005 ComRank: Metasearch and Automatic Ranking of XML Retrieval System
abstract
Different information retrieval (IR) systems often return very diverse results lists for the same query. This is problematic for users since no one IR system works best for every scenario, and it is difficult for the user to know which system will work best a priori. The challenge of metasearch is to merge results lists from several IR systems, with the goal of outperforming each of the constituent systems. This paper presents ComRank, a metasearch system that discriminates in favour of results that (1) originate by consensus amongst several systems; (2) are highly ranked in their original systems and (3) originate from the better performing systems. Importantly, ComRank determines the `better' performing systems without the need for human judgements. Rather, it uses an automatic assessment process that ranks systems by their pseudo-relevance, as derived from highly ranked results in ComRank's list. We apply our methods to the INEX Collection, which is an unexplored domain for these methods, and show that they are comparable to or better than baseline alternatives
Alan Woodley, Shlomo Geva
CW2
2005 Applying Transformation-Based Error-Driven Learning to Structured Natural Language Queries
abstract
XML information retrieval (XML-IR) systems aim to provide users with highly exhaustive and highly specific results. To interact with XML-IR systems, users must express both their content and structural requirement, in the form of a structured query. Traditionally, these structured queries have been formatted using formal languages such as XPath or NEXI. Unfortunately, formal query languages are very complex and too difficult to be used by experienced, let alone casual users. Therefore, recent research has investigated the idea of specifying users' content and structural needs via natural language queries (NLQs). In previous research we developed NLPX, a natural language interface to an XML-IR system. Here we present additions we have made to NLPX. The additions involve the application of transformation-based error-driven learning (TBL) to structured NLQs, to derive special connotations and group words into an atomic unit of information. TBL has successfully been applied to other areas of natural language processing; however, this paper presents the first time it has been applied to structured NLQs. Here, we investigate the applicability of TBL to NLQs and compare the TBL-based system, with our previous system and a system with a formal language interference. Our results show that TBL is effective for structured NLQs, and that structured NLQs a viable interface tor XML-IR systems
Alan Woodley, Shlomo Geva
CW2
2005 XML Retrieval with a Natural Language Interface
Xavier Tannier, Shlomo Geva
SPIRE2
2002 Rule extraction from local cluster neural nets
Robert Andrews 0001, Shlomo Geva
Neurocomputing2
2001 Boosting the Performance of Nearest Neighbour Methods with Feature Selection
Shlomo Geva
PAKDD1
2000 VQTree: Vector Quantization for Decision Tree Induction
Shlomo Geva, Lawrence Buckingham
PAKDD1
1998 Local cluster neural net: Architecture, training and applications
Shlomo Geva, Kurt Malmstrom, Joaquin Sitte
Neurocomputing1
1997 Refining Expert Knowledge with an Artificial Neural Network
Robert Andrews 0001, Shlomo Geva
ICONIP (2)2
1997 Rule extraction from trained artificial neural network with functional dependency preprocessing
abstract
The paper describes a technique to extract symbolic rules from a trained artificial neural network with functional dependency preprocessing. RULEX (R. Andrews and S. Geva, 1994; 1995), classified as a decompositional technique of rule extraction from trained neural network in a recent survey by R. Andrews et al. (1995), is used to extract symbolic rules from data that have been preprocessed by identification of functional dependency. The identification of functional dependency offers several advantages. It can lead to significant reductions in the computational load, to reduction in the number and complexity of derived rules and to the discovery of alternative solutions that would otherwise be ignored by some methods due to implicit or explicit procedural bias. Benchmark datasets from the UCI repository of machine learning databases are used in the testing. Experimental results indicate that by including functional dependency preprocessing performance of RULEX can be improved. Good rule quality is obtained by applying RULEX with functional dependency preprocessing when compared to symbolic rule extraction technique C4.5.
Shlomo Geva, M. T. Wong, M. Orlowski
KES (2)1
1992 A constructive method for multivariate function approximation by multilayer perceptrons
abstract
Mathematical theorems establish the existence of feedforward multilayered neural networks, based on neurons with sigmoidal transfer functions, that approximate arbitrarily well any continuous multivariate function. However, these theorems do not provide any hint on how to find the network parameters in practice. It is shown how to construct a perceptron with two hidden layers for multivariate function approximation. Such a network can perform function approximation in the same manner as networks based on Gaussian potential functions, by linear combination of local functions.
Shlomo Geva, Joaquin Sitte
IEEE Trans. Neural Networks1
1991 An Exponential Response Neural Net
abstract
By using artificial neurons with exponential transfer functions one can design perfect autoassociative and heteroassociative memory networks, with virtually unlimited storage capacity, for real or binary valued input and output. The autoassociative network has two layers: input and memory, with feedback between the two. The exponential response neurons are in the memory layer. By adding an encoding layer of conventional neurons the network becomes a heteroassociator and classifier. Because for real valued input vectors the dot-product with the weight vector is no longer a measure for similarity, we also consider a euclidean distance based neuron excitation and present Lyapunov functions for both cases. The network has energy minima corresponding only to stored prototype vectors. The exponential neurons make it simpler to build fast adaptive learning directly into classification networks that map real valued input to any class structure at its output.
Shlomo Geva, Joaquin Sitte
Neural Comput.1
1991 Adaptive nearest neighbor pattern classification
abstract
A variant of nearest-neighbor (NN) pattern classification and supervised learning by learning vector quantization (LVQ) is described. The decision surface mapping method (DSM) is a fast supervised learning algorithm and is a member of the LVQ family of algorithms. A relatively small number of prototypes are selected from a training set of correctly classified samples. The training set is then used to adapt these prototypes to map the decision surface separating the classes. This algorithm is compared with NN pattern classification, learning vector quantization, and a two-layer perceptron trained by error backpropagation. When the class boundaries are sharply defined (i.e., no classification error in the training set), the DSM algorithm outperforms these methods with respect to error rates, learning rates, and the number of prototypes required to describe class boundaries.
Shlomo Geva, Joaquin Sitte
IEEE Trans. Neural Networks1
1990 A pseudo-inverse neural net with storage capacity exceeding N
abstract
By limiting the range of interaction between the prototype vectors in the autoassociative neural network and by calculating the pseudoinverse matrix from local clusters of prototypes, it is possible to store more thanNcorrelated prototype vectors, increase the size of the basins of attraction, and include more close neighbors of prototypes in their basins. This conclusion is supported by figures for the sizes and shapes of the basins of attraction obtained from computer simulations. To demonstrate the technique, experiments were conducted with sets of 16 and 20 random prototype vectors and a network of 16 neurons at the input layer. Exhaustive scans of the state space of the input layer were used to get detailed information on the shapes and sizes of the basins. The results are presented and discussed
Shlomo Geva, Joaquin Sitte
IJCNN1