VLDB 2026 Research / reviewers in the wild / expert
Mark Gahegan
dblp:g/MarkGahegan
· DBLP profile ↗
25ranked-venue papers
8as first author
5since 2021 · last 2023
0000-0001-7209-8156ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | LivePublication: The Science Workflow Creates and Updates the PublicationabstractThe uptake of computational methods to support research has led to some remarkable new tools and methods to improve outcomes. But one unintended consequence is that the scientific record ends up being fragmented and distributed amongst several distinct systems. The research we report aims to gather together all of the components of an experiment into a single container—including the publication itself. We describe the architecture of such a system that marries together distributed workflows (Globus) with research object containers (RO-Crate) and adds new methods to describe, update and 'publish: the details of the workflow and its outcomes. Finally, we demonstrate the system with a natural language processing research use case. Augustus Ellerm, Mark Gahegan, Benjamin Adams |
e-Science | 2 |
| 2023 | Multi2Claim: Generating Scientific Claims from Multi-Choice Questions for Scientific Fact-CheckingabstractNeset Tan, Trung Nguyen, Josh Bensemann, Alex Peng, Qiming Bao, Yang Chen, Mark Gahegan, Michael Witbrock. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Neset Tan, Joshua Bensemann, Alex Yuxuan Peng, Qiming Bao 0001, Yang Chen 0028, Mark Gahegan, Michael Witbrock |
EACL | 7 |
| 2022 | Enabling LivePublicationabstractThis paper presents a prototype implementation of LivePublication, a method of publishing which integrates live eScience infrastructure and scientific methods with dynamic natural language research articles. With the maturation of eScience infrastructure, bridging the gap between live computational processes and research articles becomes more tractable and brings long term value to the publication process. The prototype demonstrates the value of integrating computational processes with research articles and provides a first step into LivePublication research and development. Augustus Ellerm, Benjamin Adams, Mark Gahegan, Lukas Trombach |
e-Science | 3 |
| 2022 | DeepPN: a deep parallel neural network based on convolutional neural network and graph convolutional network for predicting RNA-protein binding sitesabstractBACKGROUND: Addressing the laborious nature of traditional biological experiments by using an efficient computational approach to analyze RNA-binding proteins (RBPs) binding sites has always been a challenging task. RBPs play a vital role in post-transcriptional control. Identification of RBPs binding sites is a key step for the anatomy of the essential mechanism of gene regulation by controlling splicing, stability, localization and translation. Traditional methods for detecting RBPs binding sites are time-consuming and computationally-intensive. Recently, the computational method has been incorporated in researches of RBPs. Nevertheless, lots of them not only rely on the sequence data of RNA but also need additional data, for example the secondary structural data of RNA, to improve the performance of prediction, which needs the pre-work to prepare the learnable representation of structural data. RESULTS: To reduce the dependency of those pre-work, in this paper, we introduce DeepPN, a deep parallel neural network that is constructed with a convolutional neural network (CNN) and graph convolutional network (GCN) for detecting RBPs binding sites. It includes a two-layer CNN and GCN in parallel to extract the hidden features, followed by a fully connected layer to make the prediction. DeepPN discriminates the RBP binding sites on learnable representation of RNA sequences, which only uses the sequence data without using other data, for example the secondary or tertiary structure data of RNA. DeepPN is evaluated on 24 datasets of RBPs binding sites with other state-of-the-art methods. The results show that the performance of DeepPN is comparable to the published methods. CONCLUSION: The experimental results show that DeepPN can effectively capture potential hidden features in RBPs and use these features for effective prediction of binding sites. Jidong Zhang, Bo Liu 0024, Zhihan Wang, Klaus Lehnert, Mark Gahegan |
BMC Bioinform. | 5 |
| 2021 | A gastric cancer recognition algorithm on gastric pathological sections based on multistage attention-DenseNetabstractSummary As an important method to diagnose gastric cancer, gastric pathological sections images (GPSI) are hard and time‐consuming to be recognized even by an experienced doctor. An efficient method was designed to detect gastric cancer in magnified (20×) GPSI using deep learning technology. A novel DenseNet architecture was applied, modified with a multistage attention module (MSA‐DenseNet). To develop this model focusing on gastric features, a two‐stage‐input attention module was adopted to select more semantic information of cancer. Moreover, the pretraining process was divided into two steps to improve the effect of the attention mechanism. After training, our method achieved a state‐of‐the‐art performance yielding 0.9947 F1 score and 0.9976 ROC AUC on a test dataset. In line with our expectation in clinical practice, a high recall (0.9929) was produced with high sensitivity to the positive samples. These results indicate that this new model performs better than current artificial detection approaches and its effectiveness is therefore validated in cancer pathological diagnoses. Bo Liu 0024, Yelong Zhao, Bin Yang 0037, Shuangtao Zhao, Rentao Gu, Mark Gahegan |
Concurr. Comput. Pract. Exp. | 6 |
| 2020 | Fourth paradigm GIScience? Prospects for automated discovery and explanation from dataabstractThis article discusses the prospects for automated discovery of explanatory models directly from geospatial data. Rather than taking an approach based on machine learning, which generally leads to models that cannot be understood by humans or related to domain theory, the approach described here suggests we can instead construct models from fragments of domain understanding—such as commonly encountered equation forms, known constants and laws—resulting in discovered models that can both be understood by humans and directly compared with known theory. We then propose a conceptual model of the discovery process by which the various stages and components of discovery and explanation work together to learn models from data. The approach described weaves together ideas for describing models from Harvey’s book ‘Explanation in Geography’ with current thinking on how explanatory models might be ‘discovered’ from data from Inductive Process modeling. On the way, we also highlight: (i) why it is important to have models that explain as well as predict, (ii) how such an approach contrasts with – and goes beyond – current work in deep learning, (iii) how the task of model discovery might be tackled computationally and (iv) how computational model discovery can play a valuable role in creating geographical explanations. Mark Gahegan |
Int. J. Geogr. Inf. Sci. | 1 |
| 2015 | Adventures of Categories: Modelling the Evolution of Categories During Scientific InvestigationabstractCategories are the fundamental components of scientific knowledge and are used in every phase of the scientific process. However, they are often in a state of flux, with new observations, discoveries and changes in our conceptual understanding leading to the birth and death of categories, drift in their identities, as well as merging or splitting. Contemporary research tools rarely support such changes in operationalized categories, neglecting the problem of capturing and utilizing the knowledge lurking behind the process of change. This paper presents a tool -- AdvoCate1 -- that represents the dynamic nature of categories. It allows category evolution to be modelled, while maintaining a category versioning system that captures all the different versions of a category, along with the process of its exploration and evolution through use. This helps us to better understand and communicate different versions of categories and the reasons and decisions behind any changes they undergo. We demonstrate the usefulness of AdvoCate using examples of category evolution from a land cover mapping exercise. Mark Gahegan, Gillian Dobbie |
e-Science | 2 |
| 2015 | Frankenplace: Interactive Thematic Mapping for Ad Hoc Exploratory SearchabstractAd hoc keyword search engines built using modern information retrieval methods do a good job of handling fine-grained queries. However, they perform poorly at facilitating spatial and spatially-embedded thematic exploration of the results, despite the fact that many queries, e.g. "civil war," refer to different documents and topics in different places. This is not for lack of data: geographic information, such as place names, events, and coordinates are common in unstructured document collections on the web. The associations between geographic and thematic contents in these documents can provide a rich groundwork to organize information for exploratory research. In this paper we describe the architecture of an interactive thematic map search engine, Frankenplace, designed to facilitate document exploration at the intersection of theme and place. The map interface enables a user to zoom the geographic context of their query in and out, and quickly explore through thousands of search results in a meaningful way. And by combining topic models with geographically contextualized search results, users can discover related topics based on geographic context. Frankenplace utilizes a novel indexing method called geoboost for boosting terms associated with cells on a discrete global grid. The resulting index factors in the geographic scale of the place or feature mentioned in related text, the relative textual scope of the place reference, and the overall importance of the containing document in the document network. The system is currently indexed with over 5 million documents from the web, including the English Wikipedia and online travel blog entries. We demonstrate that Frankenplace can support four distinct types of exploratory search tasks while being adaptive to scale and location of interest. Benjamin Adams, Grant McKenzie, Mark Gahegan |
WWW | 3 |
| 2014 | An eScience Tool for Understanding Copyright in Data Driven SciencesabstractUnderstanding the impacts of copyright is a challenge for the sharing and reuse of our research data. There is growing recognition of the problem, but the legal knowledge required to navigate through the minefield of restrictions and risks is often too difficult to uncover and understand. As of yet there are no appropriate tools to aid researchers, librarians and research policy makers. To address this gap we present Camden, an automated copyright reasoning tool designed to integrate into existing research workflows. At its core, Camden uses dynamically generated defeasible rules to reason over the legality of a situation of using, combining and publishing data, while additionally suggesting potential licenses by which to safely share derived research outputs. This functionality has been wrapped up into an embedded software library and offered as a web application. In this paper we introduce Camden, describe its model of computational reasoning and discuss how it can be included into existing and future eResearch tools and services. Richard Hosking, Mark Gahegan, Gillian Dobbie |
eScience | 2 |
| 2013 | The Effects of Licensing on Open Data: Computing a Measure of Health for Our Scholarly RecordabstractAs data collections become established in key disciplines, some of the longstanding barriers to data sharing become to dissolve; yet others remain. While metadata and ontologies help overcome the problems of finding and interpreting data, the lack of clarity over licensing remains a real impediment to data reuse. Freedom from legal restriction and uncertainty is essential for the effective sharing, combining and deriving of data from these distributed collections. Reuse and recombination of data will be greatly facilitated by expanding the definition of the semantic web to include the semantics of data licensing. We aim to express licensing terms in a computable manner, within the context of research practice, enabling us to infer the resulting state of rights, obligations and conditions that are inherited by derived and recombined datasets, using a mixed bag of licenses. Building off this we aim to simulate the effects of varying licensing practices within communities, proposing a measure of health of our scholarly record based on compatibility and restrictiveness of the licenses contained therein. Richard Hosking, Mark Gahegan |
ISWC (2) | 2 |
| 2012 | Visual Semiotics & Uncertainty Visualization: An Empirical StudyabstractThis paper presents two linked empirical studies focused on uncertainty visualization. The experiments are framed from two conceptual perspectives. First, a typology of uncertainty is used to delineate kinds of uncertainty matched with space, time, and attribute components of data. Second, concepts from visual semiotics are applied to characterize the kind of visual signification that is appropriate for representing those different categories of uncertainty. This framework guided the two experiments reported here. The first addresses representation intuitiveness, considering both visual variables and iconicity of representation. The second addresses relative performance of the most intuitive abstract and iconic representations of uncertainty on a map reading task. Combined results suggest initial guidelines for representing uncertainty and discussion focuses on practical applicability of results. Alan M. MacEachren, Robert E. Roth, James O'Brien, Bonan Li, Derek Swingley, Mark Gahegan |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2007 | Beyond ontologies: Toward situated representations of scientific knowledge
William Pike, Mark Gahegan |
Int. J. Hum. Comput. Stud. | 2 |
| 2006 | Spatial ordering and encoding for geographic data mining and visualization
Diansheng Guo, Mark Gahegan |
J. Intell. Inf. Syst. | 2 |
| 2004 | Representing, Manipulating and Reasoning with Geographic Semantics within a Knowledge Framework
James O'Brien, Mark Gahegan |
SDH | 2 |
| 2003 | Constructing Semantically Scalable Cognitive Spaces
William Pike, Mark Gahegan |
COSIT | 2 |
| 2003 | ICEAGE: Interactive Clustering and Exploration of Large and High-Dimensional Geodata
Diansheng Guo, Donna J. Peuquet, Mark Gahegan |
GeoInformatica | 3 |
| 2003 | Is inductive machine learning just another wild goose (or might it lay the golden egg)?abstractThe research reported here contrasts the roles, methodologies and capabilities of statistical methods with those of inductive machine learning methods, as they are used inferentially in geographical analysis. To this end, various established problems with statistical inference applied in geographical settings are reviewed, based on Gould's (1970) critique. Possible solutions to the problems outlined by Gould are suggested via reviews of: (i) improved statistical methods, and (ii) recent inductive machine learning techniques. Following this, some newer problems with inference are described, emerging from the increased complexity of geographical datasets and from the analysis tasks to which we put them. Again, some solutions are suggested by pointing to newer methods. By way of results, questions are posed, and answered, relating to the changes brought about by adopting inductive machine learning methods for geographical analysis. Specifically, these questions relate to analysis capabilities, methodologies, the role of the geographer and consequences for teaching and learning. Conclusions argue that there is now a strong need, motivated from many perspectives, to give geographical data a stronger voice, thus favouring techniques that minimize the prior assumptions made of a dataset. Mark Gahegan |
Int. J. Geogr. Inf. Sci. | 1 |
| 1999 | Four barriers to the development of effective exploratory visualisation tools for the geosciencesabstractThis paper outlines four specific problems that appear to represent considerable obstacles to the development of visualisation strategies for use within the domain of geography and the Earth sciences. These are: (1) the speed of graphical rendering, (2) the management of perceptual anomalies due to visual combination effects, (3) the vast range of potential approaches and mappings (the complexity of the visual assignment process), and (4) the orientation of the user into an artificial or virtual reality. Each problem is discussed in terms of the visualisation of geographical data for the purpose of exploratory visual analysis. The specific underlying research issues and questions are described, with particular emphasis to how these relate to the geographical domain. Where possible, some potential solutions are suggested. Specific examples of geographical data visualisation are given to substantiate the arguments presented. The discussion highlights the need for further research in a number of key areas, and stresses the weaknesses of current visualisation theory and technology when applied to non-trivial geographical datasets. Mark Gahegan |
Int. J. Geogr. Inf. Sci. | 1 |
| 1997 | Experiments Using Context and Significance to Enhance the Reporting Capabilities of GIS
Mark Gahegan |
COSIT | 1 |
| 1997 | A Strategy and Architecture for The Visualization of Complex Geographical DatasetsabstractThe use of computer visualization as a means to analyze complex geographic datasets is discussed. Visualization is a valuable tool for conducting exploratory data analysis on geographical data; making good use of the human eye's unparalleled ability to recognize structure and relationships that may be inherent within the data. Traditional GIS are extremely poor at visualization, being limited to a very restricted set of visual attributes with which to convey information (position, size, color). The use of a more sophisticated approach is discussed in detail. Specifically, a system to visualise complex environmental datasets is described, which makes use of knowledge concerning the problem domain as well as knowledge concerning human cognition. In the realizations produced, the most salient attributes in the data, for a particular task, are assigned to the most striking visual attributes. Assignments are controlled by heuristics that may be changed to alter system behavior. Results are presented showing the application of this approach on datasets involving several multi-dimensional thematic layers of environmental data, used in mineral exploration. Mark Gahegan, David O'Brien |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1995 | Proximity Operators for Qualitative Spatial Reasoning
Mark Gahegan |
COSIT | 1 |
| 1993 | Feature-Based Derivation of Drainage NetworksabstractAn approach to the automatic derivation of drainage networks from digital elevation models (DEMs) using regular-gridded data is described. Features such as depressions and plateaus, on which drainage directions are undefined, are handled as entities in their own right, and are identified using operations on connected regions. Such an approach is a departure from the traditional pixel-based (local) approaches and facilitates the construction of a consistent drainage network. A single or multiple flow drainage model can be implemented. Accumulation and watershed delineation of the drainage network are performed by recursively climbing ascending paths from sink points. The algorithms are significantly faster than those of other approaches and operate within an efficient buffering scheme which facilitates handling of large DEMs and multiple data sources. J. E. McCormack, Mark Gahegan, Steven A. Roberts, J. Hogg, B. S. Hoyle |
Int. J. Geogr. Inf. Sci. | 2 |
| 1993 | An object-oriented geographic information system shell
Stephen Andrew Roberts, Mark Gahegan |
Inf. Softw. Technol. | 2 |
| 1989 | An efficient use of quadtrees in a geographical information systemabstractWith the increase in volume of spatial data now available, more effective ways must be found of storing and processing these data. This paper presents a compacted version of the linear quadtree and a spatially-referenced index method that can significantly reduce the storage requirements of a set of images and the time taken to process spatial queries. The index acts as a high-level summary of a regular-sized portion of the underlying image and so can be used to avoid examining areas of the image where none of the required features is present. Some example results are given. A method for the optimization of spatial searches is presented which takes into account the area and distribution of features within an image. Finally, a method for directly associating the edges of features with the individual nodes of a quadtree is reported. This is important since the edges of objects are no longer explicitly present in linear quadtrees and so must be recalculated when they are required for part of a query. Recalculation of object edges or boundaries is expensive; it is best, therefore, to perform the operation once only, and then save the results. Mark Gahegan |
Int. J. Geogr. Inf. Sci. | 1 |
| 1988 | An intelligent, object-oriented geographical information systemabstractThe authors report a novel interface to a spatial analysis system which allows the underlying geographical domain to be represented using a high-level, feature-or object-oriented model. The system incorporates an intelligent, data-driven reporting technique which analyses the domain characteristics in order to highlight significant trends. A second novel feature is the use of data held in our extended data model for the purpose of query optimization. The interface is based on object-oriented techniques, which have been enhanced to capture contextual data and knowledge, including general patterns of data behaviour, general geographical knowledge and database domain-specific characteristics. A geographical information system can offer more flexibility when it can handle searches in both a spatial and object-centred way. Our current work is concerned with increasing the functionality and efficiency of the object-centred approach, and hence increasing the effectiveness of geographical information systems as an aid to analysis and decision-making, especially where very large volumes of data are involved. Mark Gahegan, Stuart A. Roberts |
Int. J. Geogr. Inf. Sci. | 1 |