Laurence Anthony F. Park

dblp:46/5092 · also Laurence A. F. Park, Laurence Park 0001 · DBLP profile ↗
← Back
43ranked-venue papers
27as first author
5since 2021 · last 2025
0000-0003-0201-4409ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 34 · 26 first-author · 3 since 2021Artificial intelligence and machine learning · 20 · 12 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Determining the Need for Multi-label Classifiers by Measuring Unexplained Covariance
Laurence Anthony F. Park, Jesse Read
PAKDD (6)1
2025 Estimating Multi-Label Expected Accuracy Using Labelset Distributions
abstract
A multi-label classifier estimates the binary label state (relevant/irrelevant) for each of a set of concept labels, for a given instance. Probabilistic multi-label classifiers provide a distribution over all possible labelset combinations of such label states (the powerset of labels), from which we can provide the best estimate by selecting the labelset corresponding to the largest expected accuracy. Providing confidence for predictions is important for real-world application of multi-label models, which provides the practitioner with a sense of the correctness of the prediction. It has been thought that the probability of the chosen labelset is a good measure of the confidence of the prediction, but multi-label accuracy can be measured in many ways and so confidence should align with the expected accuracy of the evaluation method. In this article, we investigate the effectiveness of seven candidate functions for estimating multi-label expected accuracy conditioned on the labelset distribution and the evaluation method. We found most correlate to expected accuracy and have varying levels of robustness. Further, we found that the candidate functions provide high expected accuracy estimates for Hamming similarity, but a combination of the candidates provided an accurate estimate of expected accuracy for Jaccard index and Exact match.
Laurence Anthony F. Park, Jesse Read
IEEE Trans. Knowl. Data Eng.1
2023 JobIQ: Recommending Study Pathways Based on Career Choices
Tomas Trescak, Laurence Anthony F. Park, Mesut Kocyigit
CSEDU (1)2
2023 On the enumeration of integer tetrahedra
James East, Michael Hendriksen, Laurence Anthony F. Park
Comput. Geom.3
2022 Modelling Zeros in Blockmodelling
Laurence Anthony F. Park, Mohadeseh Ganji, Emir Demirovic, Jeffrey Chan, Peter J. Stuckey, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao
PAKDD (2)1
2020 Accelerated Bayesian Optimisation through Weight-Prior Tuning
abstract
Bayesian optimization (BO) is a widely-used method for optimizing expensive (to evaluate) problems. At the core of most BO methods is the modeling of the objective function using a Gaussian Process (GP) whose covariance is selected from a set of standard covariance functions. From a weight-space view, this models the objective as a linear function in a feature space implied by the given covariance $K$, with an arbitrary Gaussian weight prior ${\bf w} \sim ormdist ({\bf 0},{\bf I})$. In many practical applications there is data available that has a similar (covariance) structure to the objective, but which, having different form, cannot be used directly in standard transfer learning. In this paper we show how such auxiliary data may be used to construct a GP covariance corresponding to a more appropriate weight prior for the objective function. Building on this, we show that we may accelerate BO by modeling the objective function using this (learned) weight prior, which we demonstrate on both test functions and a practical application to short-polymer fibre manufacture.
Alistair Shilton, Sunil Gupta 0001, Santu Rana, Pratibha Vellanki, Cheng Li 0003, Svetha Venkatesh, Laurence Anthony F. Park, Alessandra Sutti, David Rubin, Thomas Dorin, Alireza Vahid, Murray Height, Teo Slezak
AISTATS7
2019 Assessing the Multi-labelness of Multi-label Data
Laurence Anthony F. Park, Yi Guo 0001, Jesse Read
ECML/PKDD (2)1
2018 Semi-supervised Blockmodelling with Pairwise Guidance
Mohadeseh Ganji, Jeffrey Chan, Peter J. Stuckey, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao, Laurence Anthony F. Park
ECML/PKDD (2)7
2018 A Blended Metric for Multi-label Optimisation and Evaluation
Laurence Anthony F. Park, Jesse Read
ECML/PKDD (1)1
2016 The Effect on Accuracy of Tweet Sample Size for Hashtag Segmentation Dictionary Construction
Laurence Anthony F. Park, Glenn Stone
PAKDD (1)1
2016 Visual Assessment of Clustering Tendency for Incomplete Data
abstract
The iVAT (asiVAT) algorithms reorder symmetric (asymmetric) dissimilarity data so that an image of the data may reveal cluster substructure. Images formed from incomplete data don't offer a very rich interpretation of cluster structure. In this paper, we examine four methods for completing the input data with imputed values before imaging. We choose a best method using contaminated versions of the complete Iris data, for which the desired results are known. Then, we analyze two real world data sets from social networks that are incomplete using the best imputation method chosen in the juried trials with Iris: (i) Sampson's monastery data, an incomplete, asymmetric relation matrix; and (ii) the karate club data, comprising a symmetric similarity matrix that is about 86 percent incomplete.
Laurence Anthony F. Park, James C. Bezdek, Christopher Leckie, Kotagiri Ramamohanarao, James Bailey 0001, Marimuthu Palaniswami
IEEE Trans. Knowl. Data Eng.1
2015 Using Entropy as a Measure of Acceptance for Multi-label Classification
Laurence Anthony F. Park, Simeon J. Simoff
IDA1
2014 Inducing Controlled Error over Variable Length Ranked Lists
Laurence Anthony F. Park, Glenn Stone
PAKDD (2)1
2014 Second order probabilistic models for within-document novelty detection in academic articles
abstract
It is becoming increasingly difficult to stay aware of the state-of-the-art in any research field due to the exponential increase in the number of academic publications. This problem effects authors and reviewers of submissions to academic journals and conferences, who must be able to identify which portions of an article are novel and which are not. Therefore, having a process to automatically judge the flow of novelty though a document would assist academics in their quest for truth. In this article, we propose the concept of Within Document Novelty Location, a method of identifying locations of novelty and non-novelty within a given document. In this preliminary investigation, we examine if a second order statistical model has any benefit, in terms of accuracy and confidence, over a simpler first order model. Experiments on 928 text sequences taken from three academic articles showed that the second order model provided a significant increase in novelty location accuracy for two of the three documents. There was no significant difference in accuracy for the remaining document, which is likely to be due to the absence of context analysis.
Laurence Anthony F. Park, Simeon J. Simoff
SIGIR1
2013 Automatic detection of retinal vascular landmark features for colour fundus image matching and patient longitudinal study
abstract
Retinal vascular landmark points such as branching points and crossovers are important features for automatic retinal image matching and vascular abnormality detection. These landmark points can enable automatic screening of large dataset through the detection of vascular network abnormalities (i.e., arteriovenous nicking, retinal vein occlusion) which are important for hypertension and cardiovascular disease prediction. Existing methods for crossover point detection use only local information at each image pixel without considering vascular features to detect crossover positions. This leads to the misclassification of very acute crossovers which are represented by two bifurcation points in the skeleton image. In this article, we propose a robust method that utilizes both local information and vascular geometrical features at the crossing to distinguish crossover from non-crossover points in a retinal image. The proposed method was validated on fifteen high resolution retinal images and the results show that our method achieves higher accuracy than any existing methods. In particular, the proposed method can discover more than 74% (recall) of crossovers with a detection accuracy (fraction of detected crossover points that are correct) of 83% (precision). The detected crossovers provide essential results for the automatic detection of vascular network abnormalities, such as arteriovenous nicking, neovascularization, and retinal vein occlusion.
Uyen T. V. Nguyen, Alauddin Bhuiyan, Laurence Anthony F. Park, Ryo Kawasaki, Tien Yin Wong, Kotagiri Ramamohanarao
ICIP3
2013 An effective retinal blood vessel segmentation method using multi-scale line detection
Uyen T. V. Nguyen, Alauddin Bhuiyan, Laurence Anthony F. Park, Kotagiri Ramamohanarao
Pattern Recognit.3
2011 Fast Approximate Text Document Clustering Using Compressive Sampling
Laurence Anthony F. Park
ECML/PKDD (2)1
2011 Clustering ellipses for anomaly detection
Masud Moshtaghi, Timothy C. Havens, James C. Bezdek, Laurence Anthony F. Park, Christopher Leckie, Sutharshan Rajasegarar, James Keller 0001, Marimuthu Palaniswami
Pattern Recognit.4
2011 Multiresolution Web Link Analysis Using Generalized Link Relations
abstract
Web link analysis methods such as PageRank, HITS, and SALSA have focused on obtaining global popularity or authority of the set of Web pages in question. Although global popularity is useful for general queries, we find that global popularity is not as useful for queries in which the global population has less knowledge of. By examining the many different communities that appear within a Web page graph, we are able to compute the popularity or authority from a specific community. Multiresolution popularity lists allow us to observe the popularity of Web pages with respect to communities at different resolutions within the Web. Multiresolution popularity lists have been shown to have high potential when compared against PageRank. In this paper, we generalize the multiresolution popularity analysis to use any form of Web page link relations. We provide results for both the PageRank relations and the In-degree relations. By utilizing the multiresolution popularity lists, we achieve a 13 percent and 25 percent improvement in mean average precision over In-degree and PageRank, respectively.
Laurence Anthony F. Park, Kotagiri Ramamohanarao
IEEE Trans. Knowl. Data Eng.1
2010 Clustering elliptical anomalies in sensor networks
abstract
We model anomalies in wireless sensor networks with ellipsoids that represent node measurements. Elliptical anomalies (EAs) are level sets of ellipsoids, and classify them as type 1, type 2 and higher order anomalies. Three measures of (dis)similarity between pairs of ellipsoids convert model ellipsoids into dissimilarity data. Clusters in the dissimilarity data may correspond to normal and anomalous measurements and nodes in the network. Assessment of (clustering) tendency is facilitated by visual inspection of (VAT/iVAT) images. Two examples illustrate the potential for anomaly detection.
James C. Bezdek, Timothy C. Havens, James Keller 0001, Christopher Leckie, Laurence Anthony F. Park
FUZZ-IEEE5
2010 Click-based evidence for decaying weight distributions in search effectiveness metrics
Yuye Zhang, Laurence Anthony F. Park, Alistair Moffat
Inf. Retr.2
2009 Kernel latent semantic analysis using an information retrieval based kernel
abstract
Hidden term relationships can be found within a document collection using Latent semantic analysis (LSA) and can be used to assist in information retrieval. LSA uses the inner product as its similarity function, which unfortunately introduces bias due to document length and term rarity into the term relationships. In this article, we present the novel kernel based LSA method, which uses separate document and query kernel functions to compute document and query similarities, rather than the inner product. We show that by providing an appropriate kernel function, we are able to provide a better fit of our data and hence produce more effective term relationships.
Laurence Anthony F. Park, Kotagiri Ramamohanarao
CIKM1
2009 Grouped ECOC Conditional Random Fields for Prediction of Web User Behavior
Yong Zhen Guo, Kotagiri Ramamohanarao, Laurence Anthony F. Park
PAKDD3
2009 The Sensitivity of Latent Dirichlet Allocation for Information Retrieval
Laurence Anthony F. Park, Kotagiri Ramamohanarao
ECML/PKDD (2)1
2009 System scoring using partial prior information
abstract
We introduce smoothing of retrieval effectiveness scores, which balances results from prior incomplete query sets against limited additional complete information, in order to obtain more refined system orderings than would be possible on the new queries alone.
Sri Devi Ravana, Laurence Anthony F. Park, Alistair Moffat
SIGIR2
2009 Score adjustment for correction of pooling bias
abstract
Information retrieval systems are evaluated against test collections of topics, documents, and assessments of which documents are relevant to which topics. Documents are chosen for relevance assessment by pooling runs from a set of existing systems. New systems can return unassessed documents, leading to an evaluation bias against them. In this paper, we propose to estimate the degree of bias against an unpooled system, and to adjust the system's score accordingly. Bias estimation can be done via leave-one-out experiments on the existing, pooled systems, but this requires the problematic assumption that the new system is similar to the existing ones. Instead, we propose that all systems, new and pooled, be fully assessed against a common set of topics, and the bias observed against the new system on the common topics be used to adjust scores on the existing topics. We demonstrate using resampling experiments on TREC test sets that our method leads to a marked reduction in error, even with only a relatively small number of common topics, and that the error decreases as the number of topics increases.
William Webber, Laurence Anthony F. Park
SIGIR2
2009 An analysis of latent semantic term self-correlation
abstract
Latent semantic analysis (LSA) is a generalized vector space method that uses dimension reduction to generate term correlations for use during the information retrieval process. We hypothesized that even though the dimension reduction establishes correlations between terms, the dimension reduction is causing a degradation in the correlation of a term to itself (self-correlation). In this article, we have proven that there is a direct relationship to the size of the LSA dimension reduction and the LSA self-correlation. We have also shown that by altering the LSA term self-correlations we gain a substantial increase in precision, while also reducing the computation required during the information retrieval process.
Laurence Anthony F. Park, Kotagiri Ramamohanarao
ACM Trans. Inf. Syst.1
2009 Efficient storage and retrieval of probabilistic latent semantic information for information retrieval
Laurence Anthony F. Park, Kotagiri Ramamohanarao
VLDB J.1
2008 Web Page Prediction Based on Conditional Random Fields
abstract
Web page prefetching is used to reduce the access latency of the Internet. However, if most prefetched Web pages are not visited by the users in their subsequent accesses, the limited network bandwidth and server resources will not be used efficiently and may worsen the access delay problem. Therefore, it is critical that we have an accurate prediction method during prefetching. Conditional Random Fields (CRFs), which are popular sequential learning models, have already been successfully used for many Natural Language Processing (NLP) tasks such as POS tagging, name entity recognition (NER) and segmentation. In this paper, we propose the use of CRFs in the field of Web page prediction. We treat the accessing sessions of previous Web users as observation sequences and label each element of these observation sequences to get the corresponding label sequences, then based on these observation and label sequences we use CRFs to train a prediction model and predict the probable subsequent Web pages for the current users. Our experimental results show that CRFs can produce higher Web page prediction accuracy effectively when compared with other popular techniques like plain Markov Chains and Hidden Markov Models (HMMs).
Yong Zhen Guo, Kotagiri Ramamohanarao, Laurence Anthony F. Park
ECAI3
2008 Query Expansion for the Language Modelling Framework Using the Naïve Bayes Assumption
Laurence Anthony F. Park, Kotagiri Ramamohanarao
PAKDD1
2008 The Effect of Weighted Term Frequencies on Probabilistic Latent Semantic Term Relationships
Laurence Anthony F. Park, Kotagiri Ramamohanarao
SPIRE1
2008 Error Correcting Output Coding-Based Conditional Random Fields for Web Page Prediction
abstract
Web page prefetching has been used efficiently to reduce the access latency problem of the Internet, its success mainly relies on the accuracy of Web page prediction. As powerful sequential learning models, conditional random fields (CRFs) have been used successfully to improve the Web page prediction accuracy when the total number of unique Web pages is small. However, because the training complexity of CRFs is quadratic to the number of labels, when applied to a Web site with a large number of unique pages, the training of CRFs may become very slow and even intractable. In this paper, we decrease the training time and computational resource requirements of CRFs training by integrating error correcting output coding (ECOC) method. Moreover, since the performance of ECOC-based methods crucially depends on the ECOC code matrix in use, we employ a coding method, search coding, to design the code matrix of good quality.
Yong Zhen Guo, Kotagiri Ramamohanarao, Laurence Anthony F. Park
Web Intelligence3
2007 Mining web multi-resolution community-based popularity for information retrieval
abstract
The PageRank algorithm is used in Web information retrieval to calculate a single list of popularity scores for each page in the Web. These popularity scores are used to rank query results when presented to the user. By using the structure of the entire Web to calculate one score per document, we are calculating a general popularity score, not particular to any community. Therefore, the PageRank scores are more suited to general queries. In this paper, we introduce a more general form of PageRank, using Web multi-resolution community-based popularity scores, where each document obtains a popularity score dependent on a given Web community. When a query is related to a specific community, we choose the associated set of popularity scores and order the query results accordingly. Using Web-community based popularity scores, we achieved an 11% increase in precision over PageRank.
Laurence Anthony F. Park, Kotagiri Ramamohanarao
CIKM1
2007 Query Expansion Using a Collection Dependent Probabilistic Latent Semantic Thesaurus
Laurence Anthony F. Park, Kotagiri Ramamohanarao
PAKDD1
2007 Personalized PageRank for Web Page Prediction Based on Access Time-Length and Frequency
abstract
Web page prefetching techniques are used to address the access latency problem of the Internet. To perform successful prefetching, we must be able to predict the next set of pages that will be accessed by users. The PageRank algorithm used by Google is able to compute the popularity of a set of Web pages based on their link structure. In this paper, a novel PageRank-like algorithm is proposed for conducting Web page prediction. Two biasing factors are adopted to personalize PageRank, so that it favors the pages that are more important to users. One factor is the length of time spent on visiting a page and the other is the frequency that a page was visited. The experiments conducted show that using these two factors simultaneously to bias PageRank results in more accurate Web page prediction than other methods that use only one of these two factors.
Yong Zhen Guo, Kotagiri Ramamohanarao, Laurence Anthony F. Park
Web Intelligence3
2005 Broadening Vector Space Schemes for Improving the Quality of Information Retrieval
Kotagiri Ramamohanarao, Laurence Anthony F. Park
APWeb2
2005 A Novel Document Ranking Method Using the Discrete Cosine Transform
abstract
We propose a new Spectral text retrieval method using the Discrete Cosine Transform (DCT). By taking advantage of the properties of the DCT and by employing the fast query and compression techniques found in vector space methods (VSM), we show that we can process queries as fast as VSM and achieve a much higher precision.
Laurence Anthony F. Park, Marimuthu Palaniswami, Kotagiri Ramamohanarao
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 A novel document retrieval method using the discrete wavelet transform
abstract
Current information retrieval methods either ignore the term positions or deal with exact term positions; the former can be seen as coarse document resolution, the latter as fine document resolution. We propose a new spectral-based information retrieval method that is able to utilize many different levels of document resolution by examining the term patterns that occur in the documents. To do this, we take advantage of the multiresolution analysis properties of the wavelet transform. We show that we are able to achieve higher precision when compared to vector space and proximity retrieval methods, while producing fast query times and using a compact index.
Laurence Anthony F. Park, Kotagiri Ramamohanarao, Marimuthu Palaniswami
ACM Trans. Inf. Syst.1
2004 Hybrid Pre-Query Term Expansion using Latent Semantic Analysis
abstract
Latent semantic retrieval methods (unlike vector space methods) take the document and query vectors and map them into a topic space to cluster related terms and documents. This produces a more precise retrieval but also a long query time. We present a new method of document retrieval which allows us to process the latent semantic information into a hybrid latent semantic-vector space query mapping. This mapping automatically expands the users query based on the latent semantic information in the document set. This expanded query is processed using a fast vector space method. Since we have the latent semantic data in a mapping, we are able to store and retrieve vector information in the same fast manner that the vector space method offers. Multiple mappings are combined to produce hybrid latent semantic retrieval which provide precision results 5% greater than the vector space method and fast query times.
Laurence Anthony F. Park, Kotagiri Ramamohanarao
ICDM1
2004 Fourier Domain Scoring: A Novel Document Ranking Method
abstract
Current document retrieval methods use a vector space similarity measure to give scores of relevance to documents when related to a specific query. The central problem with these methods is that they neglect any spatial information within the documents in question. We present a new method, called Fourier Domain Scoring (FDS), which takes advantage of this spatial information, via the Fourier transform, to give a more accurate ordering of relevance to a document set. We show that FDS gives an improvement in precision over the vector space similarity measures for the common case of Web like queries, and it gives similar results to the vector space measures for longer queries.
Laurence Anthony F. Park, Kotagiri Ramamohanarao, Marimuthu Palaniswami
IEEE Trans. Knowl. Data Eng.1
2002 A new implementation technique for fast Spectral based document retrieval systems
abstract
The traditional methods of spectral text retrieval (FDS,CDS) create an index of spatial data and convert the data to its spectral form at query time. We present a new method of implementing and querying an index containing spectral data which will conserve the high precision performance of the spectral methods, reduce the time needed to resolve the query, and maintain an acceptable size for the index. This is done by taking advantage of the properties of the discrete cosine transform and by applying ideas from vector space document ranking methods.
Laurence Anthony F. Park, Marimuthu Palaniswami, Kotagiri Ramamohanarao
ICDM1
2002 A Novel Web Text Mining Method Using the Discrete Cosine Transform
Laurence Anthony F. Park, Marimuthu Palaniswami, Kotagiri Ramamohanarao
PKDD1
2001 Internet Document Filtering Using Fourier Domain Scoring
Laurence Anthony F. Park, Marimuthu Palaniswami, Kotagiri Ramamohanarao
PKDD1