EDBT 2026 Demo / reviewers in the wild / expert
John Shepherd 0001
dblp:s/JohnShepherd · also John A. Shepherd 0001
· DBLP profile ↗
27ranked-venue papers in the field
1as first author
4since 2021 · last 2025
0000-0003-1241-4182ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 14 (1 first)Information Retrieval & Web Search · 9Data Mining & Knowledge Discovery · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AEFA: An Ensemble Framework for Fraud Detection in the Forex Market
Weiyuan Wang, Jianke Yu, Zhengyi Yang 0001, Mingchen Ju, Shuyue Yu, Jinglin Wu, Lifan Liu, Yongfei Liu, John Shepherd 0001, Wenjie Zhang 0001 |
ADMA (3) | 9 |
| 2025 | Counting the Number of Hop-Constrained Simple S-T Paths in Large Graphs
Bocheng Han, Weizhang Jiang, John Shepherd 0001, Dong Wen 0001, Zhengyi Yang 0001 |
WISE (2) | 6 |
| 2022 | Popularity Forecasting for Emerging Research Topics at Its Early Stage of Evolution
Yankin Chi, Raymond K. Wong 0001, John Shepherd 0001 |
ADMA (1) | 3 |
| 2022 | Incorporating neighborhood features in RNNs for popularity forecasting for emerging research fieldsabstractModelling popularity for academic fields have been an ongoing study. This is especially important for emerging fields as accurate models allow efforts and resources to be efficiently utilised on promising research directions. Existing modelling methods mainly face at least one of the following three challenges: Using domain specific binary classifications on topics to be ether emerging or non-emerging leading to low generalizability. Having a biased and restricted scope of investigation due to topic terms requiring manual mining from a limited number of documents. Neglecting the effect of "cold start" when utilising a field’s historical features as inputs for popularity forecasting especially when the field is emerging and possesses limited historical data. In this paper, we build upon existing work to propose a forecasting algorithm addressing all three challenges. Firstly, we define popularity forecasting as a multivariate regression problem. Next, by combining the utilisation of existing academic databases, time specific node embeddings, and dynamic time warping, we extract concurrently trending neighbour fields whose trending pattern are similar to the field of interest. Lastly, multivariate forecasting is conducted using long short-term memory (LSTM) and dual attention recurrent neural networks (DA-RNN). Experimental results on 10 emerging and non-emerging fields of study showcases the existence and various dynamics of "cold start". Additionally, the proposed algorithm is also shown to greatly reduce RMSE, MAE, and MAPE against traditional methods for emerging fields while retaining similar performance for non-emerging fields. This validates the significance of these challenges against existing methods and provides insight on the dependency structure of emerging topics with their historical features. Yankin Chi, Raymond K. Wong 0001, Hongkuan Wang, John Shepherd 0001 |
DSAA | 4 |
| 2019 | TEXUS: A unified framework for extracting and understanding tables in PDF documents
Roya Rastan, Hye-Young Paik, John Shepherd 0001 |
Inf. Process. Manag. | 3 |
| 2016 | Automated Table Understanding Using Stub Patterns
Roya Rastan, Hye-Young Paik, John Shepherd 0001, Armin Haller |
DASFAA (1) | 3 |
| 2016 | A PDF Wrapper for Table ProcessingabstractWe propose a PDF document wrapper system that is specifically targeted at table processing applications. We (i) review the PDF specifications and identify particular challenges from the table processing point of view, (ii) specify a table-oriented document model containing the required atomic elements for table extraction and understanding applications. Our evaluation showed that the wrapper was able to detect important features such as page columns, bullets and numbering in all measures, recording over 90% accuracy, leading to better table locating and segmenting. Roya Rastan, Hye-Young Paik, John Shepherd 0001 |
DocEng | 3 |
| 2015 | TEXUS: A Task-based Approach for Table Extraction and UnderstandingabstractIn this paper, we propose a precise, comprehensive model of table processing which aims to remedy some of the problems in the discussion of table processing in the literature. The model targets application-independent, end-to-end table processing, and thus encompasses a large subset of the work in the area. The model can be used to aid the design of table processing systems (We provide an example of such a system), can be considered as a reference framework for evaluating the performance of table processing systems, and can assist in clarifying terminological differences in the table processing literature. Roya Rastan, Hye-Young Paik, John Shepherd 0001 |
DocEng | 3 |
| 2013 | An Optimization Method for Proportionally Diversifying Search Results
Lin Wu 0001, Yang Wang 0023, John Shepherd 0001, Xiang Zhao 0002 |
PAKDD (1) | 3 |
| 2009 | A novel framework for efficient automated singer identification in large music databasesabstractOver the past decade, there has been explosive growth in the availability of multimedia data, particularly image, video, and music. Because of this, content-based music retrieval has attracted attention from the multimedia database and information retrieval communities. Content-based music retrieval requires us to be able to automatically identify particular characteristics of music data. One such characteristic, useful in a range of applications, is the identification of the singer in a musical piece. Unfortunately, existing approaches to this problem suffer from either low accuracy or poor scalability. In this article, we propose a novel scheme, calledHybrid Singer Identifier(HSI), for efficient automated singer recognition. HSI uses multiple low-level features extracted from both vocal and nonvocal music segments to enhance the identification process; it achieves this via a hybrid architecture that builds profiles of individual singer characteristics based on statistical mixture models. An extensive experimental study on a large music database demonstrates the superiority of our method over state-of-the-art approaches in terms of effectiveness, efficiency, scalability, and robustness. Jialie Shen 0001, John Shepherd 0001, Bin Cui 0001, Kian-Lee Tan |
ACM Trans. Inf. Syst. | 2 |
| 2006 | HSI: A Novel Framework for Efficient Automated Singer Identification in Large Music DatabaseabstractThe singer’s information is essential in organising, browsing and exploring music data. As an important component of music database systems, the automated artist identification is gaining considerable momentum due to numerous potential applications including music indexing and retrieval, copy right management and music recommendation systems. Unfortunately, the most currently employed approaches are still in their infancy and the performance is by far less satisfactory. Indeed, they suffer from low effectiveness, less robustness and poor scalability to accommodate large scale of data. In this demo, we presents a novel system, called Hybrid Singer Identifier (HSI), for efficient and effective automated singer identification in large music databases. Jialie Shen 0001, John Shepherd 0001, Bin Cui 0001, Kian-Lee Tan |
ICDE | 2 |
| 2006 | Towards efficient automated singer identification in large music databasesabstractAutomated singer identification is important in organising, browsing and retrieving data in large music databases. In this paper, we propose a novel scheme, called Hybrid Singer Identifier (HSI), for automated singer recognition. HSI can effectively use multiple low-level features extracted from both vocal and non-vocal music segments to enhance the identification process with a hybrid architecture and build profiles of individual singer characteristics based on statistical mixture models. Extensive experimental results conducted on a large music database demonstrate the superiority of our method over state-of-the-art approaches. Categories and Subject Descriptors Jialie Shen 0001, Bin Cui 0001, John Shepherd 0001, Kian-Lee Tan |
SIGIR | 3 |
| 2006 | InMAF: indexing music databases via multiple acoustic featuresabstractMusic information processing has become very important due to the ever-growing amount of music data from emerging applications. In this demonstration,we present a novel approach for generating small but comprehensive music descriptors to facilitate efficient content music management (accessing and retrieval, in particular). Unlike previous approaches that rely on low-level spectral features adapted from speech analysis technology, our approach integrates human music perception to enhance the accuracy of the retrieval and classification process via PCA and neural networks. The superiority of our method is demonstrated by comparing it with state-of-the-art approaches in the areas of music classification query effectiveness, and robustness against various audio distortion/alternatives. Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu |
SIGMOD Conference | 2 |
| 2005 | On Efficient Music Genre Classification
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu |
DASFAA | 2 |
| 2004 | Integrating heterogeneous reatures for efficient content based music retrievalabstractIn this paper, we present a novel feature extraction method facilitating efficient content-based music retrieval and classification, called InMAF. The goal of our approach is to allow straightforward incorporation of multiple musical features, such as timbral texture, pitch and rhythm structure, into a single low dimensional vector that is effective for retrieval and classification. Unlike earlier approaches that used only acoustic properties as the basis for retrieval, our approach can easily incoporate human music perception to improve accuracy of retrieval and classification process. The superiority of our method is demonstrated by comparing it with state-of-the-art approaches in the areas of music classification (using a variety of machine learning algorithms), query effectiveness and robustness against audio distortion. Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu |
CIKM | 2 |
| 2004 | Improving Query Effectiveness for Large Image Databases with Multiple Visual Feature Combination
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu, Du Q. Huynh |
DASFAA | 2 |
| 2004 | Information Extraction via Automatic Pattern Discovery in Identified Region
Liping Ma, John Shepherd 0001 |
DEXA | 2 |
| 2004 | Information extraction using two-phase pattern discoveryabstractThis paper presents a new two-phase pattern (2PP) discovery technique for information extraction. 2PP consists of orthographic pattern discovery (OPD) and semantic pattern discovery (SPD) where the OPD determines the structural features from an identified region of a document and the SPD discovers a dominant semantic pattern for the region via inference, apposition and analogy. Then the discovered pattern is applied back into the region to extract required data items through pattern matching. We evaluated 2PP using 6500 data items and obtained effective result. Liping Ma, John Shepherd 0001 |
SIGIR | 2 |
| 2004 | Query Size Estimation for Joins Using Systematic Sampling
Anne H. H. Ngu, Banchong Harangsri, John Shepherd 0001 |
Distributed Parallel Databases | 3 |
| 2003 | CMVF: A Novel Dimension Reduction Scheme for Efficient Indexing in A Large Image DatabaseabstractNo abstract available. Jialie Shen 0001, Anne H. H. Ngu, John Shepherd 0001, Du Q. Huynh, Quan Z. Sheng |
SIGMOD Conference | 3 |
| 2003 | Enhancing Text Classification Using Synopses ExtractionabstractThis paper describes a novel approach to document classification that uses decision-tree machine learning based on a succinct vector of important terms in each document. The succinct vector itself is generated by a machine-learning approach which builds parsers that can identify significant features in a document by partitioning it into regions based on low-level document characteristics. The fact that the feature vector is succinct overcomes the problem of very large term vectors, which have hindered the application of conventional machine learning to document classification. The fact that the parser can be trained to extract only important terms from documents means that small training sets can be used to achieve the same classification accuracy as with conventional approaches. Liping Ma, John Shepherd 0001, Yanchun Zhang |
WISE | 2 |
| 2002 | Extracting Information from Semistructured Data
Liping Ma, John Shepherd 0001, Yanchun Zhang |
WAIM | 2 |
| 1997 | Query Size Estimation Using Machine Learning
Banchong Harangsri, John Shepherd 0001, Anne H. H. Ngu |
DASFAA | 2 |
| 1997 | Modelling Moving Objects in Multimedia Databases
Mohammad Nabil, Anne H. H. Ngu, John Shepherd 0001 |
DASFAA | 3 |
| 1996 | Picture Similarity Retrieval Using 2D Projection Interval RepresentationabstractSpatial relationships are important ingredients for expressing constraints in retrieval systems for pictorial or multimedia databases. We have proposed a unified representation for spatial relationships, 2D Projection Interval Relationships (2D-PIR), that integrates both directional and topological relationships. We develop techniques for similarity retrieval based on the 2D-PIR representation, including a method for dealing with rotated and reflected images. Mohammad Nabil, Anne H. H. Ngu, John Shepherd 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 1995 | A Two-Phase Approach to Data Allocation in Distributed Databases
John Shepherd 0001, Banchong Harangsri, Hwee Ling Chen, Anne H. H. Ngu |
DASFAA | 1 |
| 1989 | Partial-match Retrieval using Multiple-Key Hashing with Multiple File Copies
Kotagiri Ramamohanarao, John Shepherd 0001, Ron Sacks-Davis |
DASFAA | 2 |