John Shepherd 0001

dblp:s/JohnShepherd · also John A. Shepherd 0001 · DBLP profile ↗
← Back
42ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0003-1241-4182ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 27 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9Artificial intelligence and machine learning · 7 · 3 since 2021Software engineering, systems software and programming languages · 3Theory of computation · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 AEFA: An Ensemble Framework for Fraud Detection in the Forex Market
Weiyuan Wang, Jianke Yu, Zhengyi Yang 0001, Mingchen Ju, Shuyue Yu, Jinglin Wu, Lifan Liu, Yongfei Liu, John Shepherd 0001, Wenjie Zhang 0001
ADMA (3)9
2025 Counting the Number of Hop-Constrained Simple S-T Paths in Large Graphs
Bocheng Han, Weizhang Jiang, John Shepherd 0001, Dong Wen 0001, Zhengyi Yang 0001
WISE (2)6
2022 Popularity Forecasting for Emerging Research Topics at Its Early Stage of Evolution
Yankin Chi, Raymond K. Wong 0001, John Shepherd 0001
ADMA (1)3
2022 Incorporating neighborhood features in RNNs for popularity forecasting for emerging research fields
abstract
Modelling popularity for academic fields have been an ongoing study. This is especially important for emerging fields as accurate models allow efforts and resources to be efficiently utilised on promising research directions. Existing modelling methods mainly face at least one of the following three challenges: Using domain specific binary classifications on topics to be ether emerging or non-emerging leading to low generalizability. Having a biased and restricted scope of investigation due to topic terms requiring manual mining from a limited number of documents. Neglecting the effect of "cold start" when utilising a field’s historical features as inputs for popularity forecasting especially when the field is emerging and possesses limited historical data. In this paper, we build upon existing work to propose a forecasting algorithm addressing all three challenges. Firstly, we define popularity forecasting as a multivariate regression problem. Next, by combining the utilisation of existing academic databases, time specific node embeddings, and dynamic time warping, we extract concurrently trending neighbour fields whose trending pattern are similar to the field of interest. Lastly, multivariate forecasting is conducted using long short-term memory (LSTM) and dual attention recurrent neural networks (DA-RNN). Experimental results on 10 emerging and non-emerging fields of study showcases the existence and various dynamics of "cold start". Additionally, the proposed algorithm is also shown to greatly reduce RMSE, MAE, and MAPE against traditional methods for emerging fields while retaining similar performance for non-emerging fields. This validates the significance of these challenges against existing methods and provides insight on the dependency structure of emerging topics with their historical features.
Yankin Chi, Raymond K. Wong 0001, Hongkuan Wang, John Shepherd 0001
DSAA4
2019 TEXUS: A unified framework for extracting and understanding tables in PDF documents
Roya Rastan, Hye-Young Paik, John Shepherd 0001
Inf. Process. Manag.3
2016 Automated Table Understanding Using Stub Patterns
Roya Rastan, Hye-Young Paik, John Shepherd 0001, Armin Haller
DASFAA (1)3
2016 A PDF Wrapper for Table Processing
abstract
We propose a PDF document wrapper system that is specifically targeted at table processing applications. We (i) review the PDF specifications and identify particular challenges from the table processing point of view, (ii) specify a table-oriented document model containing the required atomic elements for table extraction and understanding applications. Our evaluation showed that the wrapper was able to detect important features such as page columns, bullets and numbering in all measures, recording over 90% accuracy, leading to better table locating and segmenting.
Roya Rastan, Hye-Young Paik, John Shepherd 0001
DocEng3
2016 Building a Process Description Repository with Knowledge Acquisition
Diyin Zhou, Hye-Young Paik, Seung Hwan Ryu, John Shepherd 0001, Paul Compton
PKAW4
2015 TEXUS: A Task-based Approach for Table Extraction and Understanding
abstract
In this paper, we propose a precise, comprehensive model of table processing which aims to remedy some of the problems in the discussion of table processing in the literature. The model targets application-independent, end-to-end table processing, and thus encompasses a large subset of the work in the area. The model can be used to aid the design of table processing systems (We provide an example of such a system), can be considered as a reference framework for evaluating the performance of table processing systems, and can assist in clarifying terminological differences in the table processing literature.
Roya Rastan, Hye-Young Paik, John Shepherd 0001
DocEng3
2015 Robust User Community-Aware Landmark Photo Retrieval
Lin Wu 0001, John Shepherd 0001, Xiaodi Huang 0001, Chunzhi Hu
MMM (2)2
2015 Multi-Query Augmentation-Based Web Landmark Photo Retrieval
abstract
Given a query photo characterizing a location-aware landmark shot by a user, landmark retrieval is about returning a set of photos ranked in their similarities to the query. Existing studies on landmark retrieval focus on conducting a matching process between candidate photos and a query photo by exploiting location-aware visual features. Notwithstanding the good results achieved, these approaches are based on an assumption that a landmark of interest is well-captured and distinctive enough to be distinguished from others. In fact, distinctive landmarks may be badly selected, e.g. changes on viewpoints or angles. This will discourage the recognition results if a biased query photo is issued. In this paper, we present a novel technique that exploits user communities in social media networks. Given a biased query photo containing some landmarks taken by a user, we select multiple users to complement this user for retrieval. Multiple photos are then used to enrich the query photo, constituting a more representative yet robust multi-query set. A pattern mining method is developed to obtain a compact feature representation of photos from the multi-query set. Such a representation is utilized to efficiently query the database so as to improve retrieval results. Extensive experiments on real-world datasets demonstrate the effectiveness and efficiency of our approach.
Lin Wu 0001, Xiaodi Huang 0001, John Shepherd 0001, Yang Wang 0023
Comput. J.3
2015 An efficient framework of Bregman divergence optimization for co-ranking images and tags in a heterogeneous network
Lin Wu 0001, Xiaodi Huang 0001, Chengyuan Zhang 0001, John Shepherd 0001, Yang Wang 0023
Multim. Tools Appl.4
2013 Efficient image and tag co-ranking: a bregman divergence optimization method
abstract
Ranking on image search has attracted considerable attentions. Many graph-based algorithms have been proposed to solve this problem. Despite their remarkable success, these approaches are restricted to their separated image networks. To improve the ranking performance, one effective strategy is to work beyond the separated image graph by leveraging fruitful information from manual semantic labeling (i.e., tags) associated with images, which leads to the technique of co-ranking images and tags, a representative method that aims to explore the reinforcing relationship between image and tag graphs. The idea of co-ranking is implemented by adopting the paradigm of random walks. However, there are two problems hidden in co-ranking remained to be open: the high computational complexity and the problem of out-of-sample. To address the challenges above, in this paper, we cast the co-ranking process into a Bregman divergence optimization framework under which we transform the original random walk into an equivalent optimal kernel matrix learning problem. Enhanced by this new formulation, we derive a novel extension to achieve a better performance for both in-sample and out-of-sample cases. Extensive experiments are conducted to demonstrate the effectiveness and efficiency of our approach.
Lin Wu 0001, Yang Wang 0023, John Shepherd 0001
ACM Multimedia3
2013 Co-ranking Images and Tags via Random Walks on a Heterogeneous Graph
Lin Wu 0001, Yang Wang 0023, John Shepherd 0001
MMM (1)3
2013 An Optimization Method for Proportionally Diversifying Search Results
Lin Wu 0001, Yang Wang 0023, John Shepherd 0001, Xiang Zhao 0002
PAKDD (1)3
2013 Max-sum diversification on image ranking with non-uniform matroid constraints
Lin Wu 0001, Yang Wang 0023, John Shepherd 0001, Xiang Zhao 0002
Neurocomputing3
2009 A novel framework for efficient automated singer identification in large music databases
abstract
Over the past decade, there has been explosive growth in the availability of multimedia data, particularly image, video, and music. Because of this, content-based music retrieval has attracted attention from the multimedia database and information retrieval communities. Content-based music retrieval requires us to be able to automatically identify particular characteristics of music data. One such characteristic, useful in a range of applications, is the identification of the singer in a musical piece. Unfortunately, existing approaches to this problem suffer from either low accuracy or poor scalability. In this article, we propose a novel scheme, calledHybrid Singer Identifier(HSI), for efficient automated singer recognition. HSI uses multiple low-level features extracted from both vocal and nonvocal music segments to enhance the identification process; it achieves this via a hybrid architecture that builds profiles of individual singer characteristics based on statistical mixture models. An extensive experimental study on a large music database demonstrates the superiority of our method over state-of-the-art approaches in terms of effectiveness, efficiency, scalability, and robustness.
Jialie Shen 0001, John Shepherd 0001, Bin Cui 0001, Kian-Lee Tan
ACM Trans. Inf. Syst.2
2006 HSI: A Novel Framework for Efficient Automated Singer Identification in Large Music Database
abstract
The singer’s information is essential in organising, browsing and exploring music data. As an important component of music database systems, the automated artist identification is gaining considerable momentum due to numerous potential applications including music indexing and retrieval, copy right management and music recommendation systems. Unfortunately, the most currently employed approaches are still in their infancy and the performance is by far less satisfactory. Indeed, they suffer from low effectiveness, less robustness and poor scalability to accommodate large scale of data. In this demo, we presents a novel system, called Hybrid Singer Identifier (HSI), for efficient and effective automated singer identification in large music databases.
Jialie Shen 0001, John Shepherd 0001, Bin Cui 0001, Kian-Lee Tan
ICDE2
2006 Efficient benchmarking of content-based image retrieval via resampling
abstract
While content-based image retrieval (CBIR) is an expanding field, and new approaches to ever more effective retrieval are frequently proposed, relatively little attention has so far been paid to the process of evaluating the effectiveness of CBIR methods. Most of the reported evaluations use standard IR evaluation methodologies, with little consideration of their statistical significance or appropriateness for CBIR, which makes it difficult to assess the precise impact of individual methods. In this paper, we present a new approach for evaluating CBIR systems which provides both efficient and statistically-sound performance evaluation. The approach is based on stratified sampling, and provides a significant improvement over existing evaluation approaches. Comprehensive experiments using our approach to evaluate a range of CBIR methods have shown that the approach reduces not only the estimation error, but also reduces the size of the test data set required to achieve specific estimation error levels.
Jialie Shen 0001, John Shepherd 0001
ACM Multimedia2
2006 Towards efficient automated singer identification in large music databases
abstract
Automated singer identification is important in organising, browsing and retrieving data in large music databases. In this paper, we propose a novel scheme, called Hybrid Singer Identifier (HSI), for automated singer recognition. HSI can effectively use multiple low-level features extracted from both vocal and non-vocal music segments to enhance the identification process with a hybrid architecture and build profiles of individual singer characteristics based on statistical mixture models. Extensive experimental results conducted on a large music database demonstrate the superiority of our method over state-of-the-art approaches. Categories and Subject Descriptors
Jialie Shen 0001, Bin Cui 0001, John Shepherd 0001, Kian-Lee Tan
SIGIR3
2006 InMAF: indexing music databases via multiple acoustic features
abstract
Music information processing has become very important due to the ever-growing amount of music data from emerging applications. In this demonstration,we present a novel approach for generating small but comprehensive music descriptors to facilitate efficient content music management (accessing and retrieval, in particular). Unlike previous approaches that rely on low-level spectral features adapted from speech analysis technology, our approach integrates human music perception to enhance the accuracy of the retrieval and classification process via PCA and neural networks. The superiority of our method is demonstrated by comparing it with state-of-the-art approaches in the areas of music classification query effectiveness, and robustness against various audio distortion/alternatives.
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu
SIGMOD Conference2
2006 Towards Effective Content-Based Music Retrieval With Multiple Acoustic Feature Combination
abstract
In this paper, we present a new approach to constructing music descriptors to support efficient content-based music retrieval and classification. The system applies multiple musical properties combined with a hybrid architecture based on principal component analysis (PCA) and a multilayer perceptron neural network. This architecture enables straightforward incorporation of multiple musical feature vectors, based on properties such as timbral texture, pitch, and rhythm structure, into a single low-dimensioned vector that is more effective for classification than the larger individual feature vectors. The use of supervised training enables incorporation of human musical perception that further enhances the classification process. We compare our approach with state of the art techniques and demonstrate its effectiveness on content-based music retrieval. In addition, extensive experimental study illustrates its effectiveness and robustness against various kinds of audio alteration.
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu
IEEE Trans. Multim.2
2005 On Efficient Music Genre Classification
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu
DASFAA2
2005 Semantic-Sensitive Classification for Large Image Libraries
abstract
With advances in multimedia technology, image data with various formats is is becoming available at an explosive rate from various domain applications. How to efficiently organise and access them has been an extremely important issue and enjoying growing attention. In this paper, we present results from experimental studies investigating performance of image classification for a novel dimension reduction scheme with hybrid architecture. We demonstrate that not only can the method provide superior quality of classification accuracy with various machine learning based classifier but also substantially speed up training and categorisation process. Moreover, it is fairly robust against various kinds of visual distortions and noises.
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu
MMM2
2004 Integrating heterogeneous reatures for efficient content based music retrieval
abstract
In this paper, we present a novel feature extraction method facilitating efficient content-based music retrieval and classification, called InMAF. The goal of our approach is to allow straightforward incorporation of multiple musical features, such as timbral texture, pitch and rhythm structure, into a single low dimensional vector that is effective for retrieval and classification. Unlike earlier approaches that used only acoustic properties as the basis for retrieval, our approach can easily incoporate human music perception to improve accuracy of retrieval and classification process. The superiority of our method is demonstrated by comparing it with state-of-the-art approaches in the areas of music classification (using a variety of machine learning algorithms), query effectiveness and robustness against audio distortion.
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu
CIKM2
2004 Improving Query Effectiveness for Large Image Databases with Multiple Visual Feature Combination
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu, Du Q. Huynh
DASFAA2
2004 Information Extraction via Automatic Pattern Discovery in Identified Region
Liping Ma, John Shepherd 0001
DEXA2
2004 Information extraction using two-phase pattern discovery
abstract
This paper presents a new two-phase pattern (2PP) discovery technique for information extraction. 2PP consists of orthographic pattern discovery (OPD) and semantic pattern discovery (SPD) where the OPD determines the structural features from an identified region of a document and the SPD discovers a dominant semantic pattern for the region via inference, apposition and analogy. Then the discovered pattern is applied back into the region to extract required data items through pattern matching. We evaluated 2PP using 6500 data items and obtained effective result.
Liping Ma, John Shepherd 0001
SIGIR2
2004 Query Size Estimation for Joins Using Systematic Sampling
Anne H. H. Ngu, Banchong Harangsri, John Shepherd 0001
Distributed Parallel Databases3
2003 CMVF: A Novel Dimension Reduction Scheme for Efficient Indexing in A Large Image Database
abstract
No abstract available.
Jialie Shen 0001, Anne H. H. Ngu, John Shepherd 0001, Du Q. Huynh, Quan Z. Sheng
SIGMOD Conference3
2003 Enhancing Text Classification Using Synopses Extraction
abstract
This paper describes a novel approach to document classification that uses decision-tree machine learning based on a succinct vector of important terms in each document. The succinct vector itself is generated by a machine-learning approach which builds parsers that can identify significant features in a document by partitioning it into regions based on low-level document characteristics. The fact that the feature vector is succinct overcomes the problem of very large term vectors, which have hindered the application of conventional machine learning to document classification. The fact that the parser can be trained to extract only important terms from documents means that small training sets can be used to achieve the same classification accuracy as with conventional approaches.
Liping Ma, John Shepherd 0001, Yanchun Zhang
WISE2
2002 Extracting Information from Semistructured Data
Liping Ma, John Shepherd 0001, Yanchun Zhang
WAIM2
2001 Modeling and Retrieval of Moving Objects
Mohammad Nabil, Anne H. H. Ngu, John Shepherd 0001
Multim. Tools Appl.3
1998 An Efficient Nearest-Neighbour Search While Varying Euclidean Metrics
abstract
Article An efficient nearest-neighbour search while varying Euclidean metrics Share on Authors: R. Kurniawati University of New South Wales, Sydney 2052, Australia University of New South Wales, Sydney 2052, AustraliaView Profile , J. S. Jin University of New South Wales, Sydney 2052, Australia University of New South Wales, Sydney 2052, AustraliaView Profile , J. A. Shepherd University of New South Wales, Sydney 2052, Australia University of New South Wales, Sydney 2052, AustraliaView Profile Authors Info & Claims MULTIMEDIA '98: Proceedings of the sixth ACM international conference on MultimediaSeptember 1998 Pages 411–418https://doi.org/10.1145/290747.290812Online:01 September 1998Publication History 2citation597DownloadsMetricsTotal Citations2Total Downloads597Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Ruth Kurniawati, Jesse S. Jin, John Shepherd 0001
ACM Multimedia3
1997 Query Size Estimation Using Machine Learning
Banchong Harangsri, John Shepherd 0001, Anne H. H. Ngu
DASFAA2
1997 Modelling Moving Objects in Multimedia Databases
Mohammad Nabil, Anne H. H. Ngu, John Shepherd 0001
DASFAA3
1996 Picture Similarity Retrieval Using 2D Projection Interval Representation
abstract
Spatial relationships are important ingredients for expressing constraints in retrieval systems for pictorial or multimedia databases. We have proposed a unified representation for spatial relationships, 2D Projection Interval Relationships (2D-PIR), that integrates both directional and topological relationships. We develop techniques for similarity retrieval based on the 2D-PIR representation, including a method for dealing with rotated and reflected images.
Mohammad Nabil, Anne H. H. Ngu, John Shepherd 0001
IEEE Trans. Knowl. Data Eng.3
1995 A Two-Phase Approach to Data Allocation in Distributed Databases
John Shepherd 0001, Banchong Harangsri, Hwee Ling Chen, Anne H. H. Ngu
DASFAA1
1989 Partial-match Retrieval using Multiple-Key Hashing with Multiple File Copies
Kotagiri Ramamohanarao, John Shepherd 0001, Ron Sacks-Davis
DASFAA2
1987 Answering Queries in Deductive Database Systems
Kotagiri Ramamohanarao, John Shepherd 0001
ICLP2
1986 A Superimposed Codeword Indexing Scheme for Very Large Prolog Databases
Kotagiri Ramamohanarao, John Shepherd 0001
ICLP2
1981 A critical examination of software science
Jean-Louis Lassez, Dirk van der Knijff, John Shepherd 0001, Catherine Lassez
J. Syst. Softw.3