VLDB 2026 Research / reviewers in the wild / expert
John Shepherd 0001
dblp:s/JohnShepherd · also John A. Shepherd 0001
· DBLP profile ↗
42ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0003-1241-4182ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 27 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9Artificial intelligence and machine learning · 7 · 3 since 2021Software engineering, systems software and programming languages · 3Theory of computation · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AEFA: An Ensemble Framework for Fraud Detection in the Forex Market
Weiyuan Wang, Jianke Yu, Zhengyi Yang 0001, Mingchen Ju, Shuyue Yu, Jinglin Wu, Lifan Liu, Yongfei Liu, John Shepherd 0001, Wenjie Zhang 0001 |
ADMA (3) | 9 |
| 2025 | Counting the Number of Hop-Constrained Simple S-T Paths in Large Graphs
Bocheng Han, Weizhang Jiang, John Shepherd 0001, Dong Wen 0001, Zhengyi Yang 0001 |
WISE (2) | 6 |
| 2022 | Popularity Forecasting for Emerging Research Topics at Its Early Stage of Evolution
Yankin Chi, Raymond K. Wong 0001, John Shepherd 0001 |
ADMA (1) | 3 |
| 2022 | Incorporating neighborhood features in RNNs for popularity forecasting for emerging research fieldsabstractModelling popularity for academic fields have been an ongoing study. This is especially important for emerging fields as accurate models allow efforts and resources to be efficiently utilised on promising research directions. Existing modelling methods mainly face at least one of the following three challenges: Using domain specific binary classifications on topics to be ether emerging or non-emerging leading to low generalizability. Having a biased and restricted scope of investigation due to topic terms requiring manual mining from a limited number of documents. Neglecting the effect of "cold start" when utilising a field’s historical features as inputs for popularity forecasting especially when the field is emerging and possesses limited historical data. In this paper, we build upon existing work to propose a forecasting algorithm addressing all three challenges. Firstly, we define popularity forecasting as a multivariate regression problem. Next, by combining the utilisation of existing academic databases, time specific node embeddings, and dynamic time warping, we extract concurrently trending neighbour fields whose trending pattern are similar to the field of interest. Lastly, multivariate forecasting is conducted using long short-term memory (LSTM) and dual attention recurrent neural networks (DA-RNN). Experimental results on 10 emerging and non-emerging fields of study showcases the existence and various dynamics of "cold start". Additionally, the proposed algorithm is also shown to greatly reduce RMSE, MAE, and MAPE against traditional methods for emerging fields while retaining similar performance for non-emerging fields. This validates the significance of these challenges against existing methods and provides insight on the dependency structure of emerging topics with their historical features. Yankin Chi, Raymond K. Wong 0001, Hongkuan Wang, John Shepherd 0001 |
DSAA | 4 |
| 2019 | TEXUS: A unified framework for extracting and understanding tables in PDF documents
Roya Rastan, Hye-Young Paik, John Shepherd 0001 |
Inf. Process. Manag. | 3 |
| 2016 | Automated Table Understanding Using Stub Patterns
Roya Rastan, Hye-Young Paik, John Shepherd 0001, Armin Haller |
DASFAA (1) | 3 |
| 2016 | A PDF Wrapper for Table ProcessingabstractWe propose a PDF document wrapper system that is specifically targeted at table processing applications. We (i) review the PDF specifications and identify particular challenges from the table processing point of view, (ii) specify a table-oriented document model containing the required atomic elements for table extraction and understanding applications. Our evaluation showed that the wrapper was able to detect important features such as page columns, bullets and numbering in all measures, recording over 90% accuracy, leading to better table locating and segmenting. Roya Rastan, Hye-Young Paik, John Shepherd 0001 |
DocEng | 3 |
| 2016 | Building a Process Description Repository with Knowledge Acquisition
Diyin Zhou, Hye-Young Paik, Seung Hwan Ryu, John Shepherd 0001, Paul Compton |
PKAW | 4 |
| 2015 | TEXUS: A Task-based Approach for Table Extraction and UnderstandingabstractIn this paper, we propose a precise, comprehensive model of table processing which aims to remedy some of the problems in the discussion of table processing in the literature. The model targets application-independent, end-to-end table processing, and thus encompasses a large subset of the work in the area. The model can be used to aid the design of table processing systems (We provide an example of such a system), can be considered as a reference framework for evaluating the performance of table processing systems, and can assist in clarifying terminological differences in the table processing literature. Roya Rastan, Hye-Young Paik, John Shepherd 0001 |
DocEng | 3 |
| 2015 | Robust User Community-Aware Landmark Photo Retrieval
Lin Wu 0001, John Shepherd 0001, Xiaodi Huang 0001, Chunzhi Hu |
MMM (2) | 2 |
| 2015 | Multi-Query Augmentation-Based Web Landmark Photo RetrievalabstractGiven a query photo characterizing a location-aware landmark shot by a user, landmark retrieval is about returning a set of photos ranked in their similarities to the query. Existing studies on landmark retrieval focus on conducting a matching process between candidate photos and a query photo by exploiting location-aware visual features. Notwithstanding the good results achieved, these approaches are based on an assumption that a landmark of interest is well-captured and distinctive enough to be distinguished from others. In fact, distinctive landmarks may be badly selected, e.g. changes on viewpoints or angles. This will discourage the recognition results if a biased query photo is issued. In this paper, we present a novel technique that exploits user communities in social media networks. Given a biased query photo containing some landmarks taken by a user, we select multiple users to complement this user for retrieval. Multiple photos are then used to enrich the query photo, constituting a more representative yet robust multi-query set. A pattern mining method is developed to obtain a compact feature representation of photos from the multi-query set. Such a representation is utilized to efficiently query the database so as to improve retrieval results. Extensive experiments on real-world datasets demonstrate the effectiveness and efficiency of our approach. Lin Wu 0001, Xiaodi Huang 0001, John Shepherd 0001, Yang Wang 0023 |
Comput. J. | 3 |
| 2015 | An efficient framework of Bregman divergence optimization for co-ranking images and tags in a heterogeneous network
Lin Wu 0001, Xiaodi Huang 0001, Chengyuan Zhang 0001, John Shepherd 0001, Yang Wang 0023 |
Multim. Tools Appl. | 4 |
| 2013 | Efficient image and tag co-ranking: a bregman divergence optimization methodabstractRanking on image search has attracted considerable attentions. Many graph-based algorithms have been proposed to solve this problem. Despite their remarkable success, these approaches are restricted to their separated image networks. To improve the ranking performance, one effective strategy is to work beyond the separated image graph by leveraging fruitful information from manual semantic labeling (i.e., tags) associated with images, which leads to the technique of co-ranking images and tags, a representative method that aims to explore the reinforcing relationship between image and tag graphs. The idea of co-ranking is implemented by adopting the paradigm of random walks. However, there are two problems hidden in co-ranking remained to be open: the high computational complexity and the problem of out-of-sample. To address the challenges above, in this paper, we cast the co-ranking process into a Bregman divergence optimization framework under which we transform the original random walk into an equivalent optimal kernel matrix learning problem. Enhanced by this new formulation, we derive a novel extension to achieve a better performance for both in-sample and out-of-sample cases. Extensive experiments are conducted to demonstrate the effectiveness and efficiency of our approach. Lin Wu 0001, Yang Wang 0023, John Shepherd 0001 |
ACM Multimedia | 3 |
| 2013 | Co-ranking Images and Tags via Random Walks on a Heterogeneous Graph
Lin Wu 0001, Yang Wang 0023, John Shepherd 0001 |
MMM (1) | 3 |
| 2013 | An Optimization Method for Proportionally Diversifying Search Results
Lin Wu 0001, Yang Wang 0023, John Shepherd 0001, Xiang Zhao 0002 |
PAKDD (1) | 3 |
| 2013 | Max-sum diversification on image ranking with non-uniform matroid constraints
Lin Wu 0001, Yang Wang 0023, John Shepherd 0001, Xiang Zhao 0002 |
Neurocomputing | 3 |
| 2009 | A novel framework for efficient automated singer identification in large music databasesabstractOver the past decade, there has been explosive growth in the availability of multimedia data, particularly image, video, and music. Because of this, content-based music retrieval has attracted attention from the multimedia database and information retrieval communities. Content-based music retrieval requires us to be able to automatically identify particular characteristics of music data. One such characteristic, useful in a range of applications, is the identification of the singer in a musical piece. Unfortunately, existing approaches to this problem suffer from either low accuracy or poor scalability. In this article, we propose a novel scheme, calledHybrid Singer Identifier(HSI), for efficient automated singer recognition. HSI uses multiple low-level features extracted from both vocal and nonvocal music segments to enhance the identification process; it achieves this via a hybrid architecture that builds profiles of individual singer characteristics based on statistical mixture models. An extensive experimental study on a large music database demonstrates the superiority of our method over state-of-the-art approaches in terms of effectiveness, efficiency, scalability, and robustness. Jialie Shen 0001, John Shepherd 0001, Bin Cui 0001, Kian-Lee Tan |
ACM Trans. Inf. Syst. | 2 |
| 2006 | HSI: A Novel Framework for Efficient Automated Singer Identification in Large Music DatabaseabstractThe singer’s information is essential in organising, browsing and exploring music data. As an important component of music database systems, the automated artist identification is gaining considerable momentum due to numerous potential applications including music indexing and retrieval, copy right management and music recommendation systems. Unfortunately, the most currently employed approaches are still in their infancy and the performance is by far less satisfactory. Indeed, they suffer from low effectiveness, less robustness and poor scalability to accommodate large scale of data. In this demo, we presents a novel system, called Hybrid Singer Identifier (HSI), for efficient and effective automated singer identification in large music databases. Jialie Shen 0001, John Shepherd 0001, Bin Cui 0001, Kian-Lee Tan |
ICDE | 2 |
| 2006 | Efficient benchmarking of content-based image retrieval via resamplingabstractWhile content-based image retrieval (CBIR) is an expanding field, and new approaches to ever more effective retrieval are frequently proposed, relatively little attention has so far been paid to the process of evaluating the effectiveness of CBIR methods. Most of the reported evaluations use standard IR evaluation methodologies, with little consideration of their statistical significance or appropriateness for CBIR, which makes it difficult to assess the precise impact of individual methods. In this paper, we present a new approach for evaluating CBIR systems which provides both efficient and statistically-sound performance evaluation. The approach is based on stratified sampling, and provides a significant improvement over existing evaluation approaches. Comprehensive experiments using our approach to evaluate a range of CBIR methods have shown that the approach reduces not only the estimation error, but also reduces the size of the test data set required to achieve specific estimation error levels. Jialie Shen 0001, John Shepherd 0001 |
ACM Multimedia | 2 |
| 2006 | Towards efficient automated singer identification in large music databasesabstractAutomated singer identification is important in organising, browsing and retrieving data in large music databases. In this paper, we propose a novel scheme, called Hybrid Singer Identifier (HSI), for automated singer recognition. HSI can effectively use multiple low-level features extracted from both vocal and non-vocal music segments to enhance the identification process with a hybrid architecture and build profiles of individual singer characteristics based on statistical mixture models. Extensive experimental results conducted on a large music database demonstrate the superiority of our method over state-of-the-art approaches. Categories and Subject Descriptors Jialie Shen 0001, Bin Cui 0001, John Shepherd 0001, Kian-Lee Tan |
SIGIR | 3 |
| 2006 | InMAF: indexing music databases via multiple acoustic featuresabstractMusic information processing has become very important due to the ever-growing amount of music data from emerging applications. In this demonstration,we present a novel approach for generating small but comprehensive music descriptors to facilitate efficient content music management (accessing and retrieval, in particular). Unlike previous approaches that rely on low-level spectral features adapted from speech analysis technology, our approach integrates human music perception to enhance the accuracy of the retrieval and classification process via PCA and neural networks. The superiority of our method is demonstrated by comparing it with state-of-the-art approaches in the areas of music classification query effectiveness, and robustness against various audio distortion/alternatives. Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu |
SIGMOD Conference | 2 |
| 2006 | Towards Effective Content-Based Music Retrieval With Multiple Acoustic Feature CombinationabstractIn this paper, we present a new approach to constructing music descriptors to support efficient content-based music retrieval and classification. The system applies multiple musical properties combined with a hybrid architecture based on principal component analysis (PCA) and a multilayer perceptron neural network. This architecture enables straightforward incorporation of multiple musical feature vectors, based on properties such as timbral texture, pitch, and rhythm structure, into a single low-dimensioned vector that is more effective for classification than the larger individual feature vectors. The use of supervised training enables incorporation of human musical perception that further enhances the classification process. We compare our approach with state of the art techniques and demonstrate its effectiveness on content-based music retrieval. In addition, extensive experimental study illustrates its effectiveness and robustness against various kinds of audio alteration. Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu |
IEEE Trans. Multim. | 2 |
| 2005 | On Efficient Music Genre Classification
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu |
DASFAA | 2 |
| 2005 | Semantic-Sensitive Classification for Large Image LibrariesabstractWith advances in multimedia technology, image data with various formats is is becoming available at an explosive rate from various domain applications. How to efficiently organise and access them has been an extremely important issue and enjoying growing attention. In this paper, we present results from experimental studies investigating performance of image classification for a novel dimension reduction scheme with hybrid architecture. We demonstrate that not only can the method provide superior quality of classification accuracy with various machine learning based classifier but also substantially speed up training and categorisation process. Moreover, it is fairly robust against various kinds of visual distortions and noises. Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu |
MMM | 2 |
| 2004 | Integrating heterogeneous reatures for efficient content based music retrievalabstractIn this paper, we present a novel feature extraction method facilitating efficient content-based music retrieval and classification, called InMAF. The goal of our approach is to allow straightforward incorporation of multiple musical features, such as timbral texture, pitch and rhythm structure, into a single low dimensional vector that is effective for retrieval and classification. Unlike earlier approaches that used only acoustic properties as the basis for retrieval, our approach can easily incoporate human music perception to improve accuracy of retrieval and classification process. The superiority of our method is demonstrated by comparing it with state-of-the-art approaches in the areas of music classification (using a variety of machine learning algorithms), query effectiveness and robustness against audio distortion. Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu |
CIKM | 2 |
| 2004 | Improving Query Effectiveness for Large Image Databases with Multiple Visual Feature Combination
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu, Du Q. Huynh |
DASFAA | 2 |
| 2004 | Information Extraction via Automatic Pattern Discovery in Identified Region
Liping Ma, John Shepherd 0001 |
DEXA | 2 |
| 2004 | Information extraction using two-phase pattern discoveryabstractThis paper presents a new two-phase pattern (2PP) discovery technique for information extraction. 2PP consists of orthographic pattern discovery (OPD) and semantic pattern discovery (SPD) where the OPD determines the structural features from an identified region of a document and the SPD discovers a dominant semantic pattern for the region via inference, apposition and analogy. Then the discovered pattern is applied back into the region to extract required data items through pattern matching. We evaluated 2PP using 6500 data items and obtained effective result. Liping Ma, John Shepherd 0001 |
SIGIR | 2 |
| 2004 | Query Size Estimation for Joins Using Systematic Sampling
Anne H. H. Ngu, Banchong Harangsri, John Shepherd 0001 |
Distributed Parallel Databases | 3 |
| 2003 | CMVF: A Novel Dimension Reduction Scheme for Efficient Indexing in A Large Image DatabaseabstractNo abstract available. Jialie Shen 0001, Anne H. H. Ngu, John Shepherd 0001, Du Q. Huynh, Quan Z. Sheng |
SIGMOD Conference | 3 |
| 2003 | Enhancing Text Classification Using Synopses ExtractionabstractThis paper describes a novel approach to document classification that uses decision-tree machine learning based on a succinct vector of important terms in each document. The succinct vector itself is generated by a machine-learning approach which builds parsers that can identify significant features in a document by partitioning it into regions based on low-level document characteristics. The fact that the feature vector is succinct overcomes the problem of very large term vectors, which have hindered the application of conventional machine learning to document classification. The fact that the parser can be trained to extract only important terms from documents means that small training sets can be used to achieve the same classification accuracy as with conventional approaches. Liping Ma, John Shepherd 0001, Yanchun Zhang |
WISE | 2 |
| 2002 | Extracting Information from Semistructured Data
Liping Ma, John Shepherd 0001, Yanchun Zhang |
WAIM | 2 |
| 2001 | Modeling and Retrieval of Moving Objects
Mohammad Nabil, Anne H. H. Ngu, John Shepherd 0001 |
Multim. Tools Appl. | 3 |
| 1998 | An Efficient Nearest-Neighbour Search While Varying Euclidean MetricsabstractArticle An efficient nearest-neighbour search while varying Euclidean metrics Share on Authors: R. Kurniawati University of New South Wales, Sydney 2052, Australia University of New South Wales, Sydney 2052, AustraliaView Profile , J. S. Jin University of New South Wales, Sydney 2052, Australia University of New South Wales, Sydney 2052, AustraliaView Profile , J. A. Shepherd University of New South Wales, Sydney 2052, Australia University of New South Wales, Sydney 2052, AustraliaView Profile Authors Info & Claims MULTIMEDIA '98: Proceedings of the sixth ACM international conference on MultimediaSeptember 1998 Pages 411–418https://doi.org/10.1145/290747.290812Online:01 September 1998Publication History 2citation597DownloadsMetricsTotal Citations2Total Downloads597Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Ruth Kurniawati, Jesse S. Jin, John Shepherd 0001 |
ACM Multimedia | 3 |
| 1997 | Query Size Estimation Using Machine Learning
Banchong Harangsri, John Shepherd 0001, Anne H. H. Ngu |
DASFAA | 2 |
| 1997 | Modelling Moving Objects in Multimedia Databases
Mohammad Nabil, Anne H. H. Ngu, John Shepherd 0001 |
DASFAA | 3 |
| 1996 | Picture Similarity Retrieval Using 2D Projection Interval RepresentationabstractSpatial relationships are important ingredients for expressing constraints in retrieval systems for pictorial or multimedia databases. We have proposed a unified representation for spatial relationships, 2D Projection Interval Relationships (2D-PIR), that integrates both directional and topological relationships. We develop techniques for similarity retrieval based on the 2D-PIR representation, including a method for dealing with rotated and reflected images. Mohammad Nabil, Anne H. H. Ngu, John Shepherd 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 1995 | A Two-Phase Approach to Data Allocation in Distributed Databases
John Shepherd 0001, Banchong Harangsri, Hwee Ling Chen, Anne H. H. Ngu |
DASFAA | 1 |
| 1989 | Partial-match Retrieval using Multiple-Key Hashing with Multiple File Copies
Kotagiri Ramamohanarao, John Shepherd 0001, Ron Sacks-Davis |
DASFAA | 2 |
| 1987 | Answering Queries in Deductive Database Systems
Kotagiri Ramamohanarao, John Shepherd 0001 |
ICLP | 2 |
| 1986 | A Superimposed Codeword Indexing Scheme for Very Large Prolog Databases
Kotagiri Ramamohanarao, John Shepherd 0001 |
ICLP | 2 |
| 1981 | A critical examination of software science
Jean-Louis Lassez, Dirk van der Knijff, John Shepherd 0001, Catherine Lassez |
J. Syst. Softw. | 3 |