Jano I. van Hemert

dblp:06/3603 · DBLP profile ↗
← Back
36ranked-venue papers
8as first author
0since 2021 · last 2018
0000-0003-0834-7079ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 7 first-authorSystems, architecture and hardware · 9 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 66% High-performance computing · 34%
Software engineering, system software, and programming languages
1 paper
Services computing and microservices · 100%
Databases, data mining, and information retrieval
1 paper
Data stream processing · 100%
Computer networks
1 paper
Transport protocols and congestion control · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Services computing and microservices
service-oriented architecture
0.112012
Reducing Data Transfer in Service-Oriented Architectures: The Circulate Approach · IEEE Trans. Serv. Comput. 2012
Bioinformatics and computational biology › gene expression analysis
gene expression classification
0.112011
Automatically identifying and annotating mouse embryo gene expression patterns · Bioinform. 2011
Bioinformatics and computational biology › gene expression analysis › gene expression pattern analysis
gene expression pattern annotation
0.112011
Automatically identifying and annotating mouse embryo gene expression patterns · Bioinform. 2011
Bioinformatics and computational biology › bioimage informatics › tissue image analysis
in situ hybridization image analysis
0.112011
Automatically identifying and annotating mouse embryo gene expression patterns · Bioinform. 2011
Data stream processing
streaming graph
0.112010
Towards optimising distributed data streaming graphs using parallel streams · HPDC 2010
High-performance computing
scientific workflow
0.112010
Towards optimising distributed data streaming graphs using parallel streams · HPDC 2010
Transport protocols and congestion control › elastic traffic
file transfer optimization
0.012012
Reducing Data Transfer in Service-Oriented Architectures: The Circulate Approach · IEEE Trans. Serv. Comput. 2012
Distributed systems › workflow management
workflow execution
0.012010
Towards optimising distributed data streaming graphs using parallel streams · HPDC 2010
Distributed systems › distributed coordination
decentralized orchestration
0.012008
Eliminating the middleman: peer-to-peer dataflow · HPDC 2008
Distributed systems
workflow management
0.012008
Eliminating the middleman: peer-to-peer dataflow · HPDC 2008

Methods — techniques the papers use, named apart from their topics

orchestration · 0.4distributed data transport · 0.4choreography model · 0.4pipelining · 0.2parallel streams · 0.2machine learning · 0.1image processing · 0.1
YearPublicationVenuePosition
2018 A Graph Cut Approach to Artery/Vein Classification in Ultra-Widefield Scanning Laser Ophthalmoscopy
abstract
The classification of blood vessels into arterioles and venules is a fundamental step in the automatic investigation of retinal biomarkers for systemic diseases. In this paper, we present a novel technique for vessel classification on ultra-wide-field-of-view images of the retinal fundus acquired with a scanning laser ophthalmoscope. To the best of our knowledge, this is the first time that a fully automated artery/vein classification technique for this type of retinal imaging with no manual intervention has been presented. The proposed method exploits hand-crafted features based on local vessel intensity and vascular morphology to formulate a graph representation from which a globally optimal separation between the arterial and venular networks is computed by graph cut approach. The technique was tested on three different data sets (one publicly available and two local) and achieved an average classification accuracy of 0.883 in the largest data set.
Enrico Pellegrini, Gavin Robertson, Thomas J. MacGillivray, Jano I. van Hemert, Graeme Houston, Emanuele Trucco
IEEE Trans. Medical Imaging4
2015 Retinal Area Detector From Scanning Laser Ophthalmoscope (SLO) Images for Diagnosing Retinal Diseases
abstract
Scanning laser ophthalmoscopes (SLOs) can be used for early detection of retinal diseases. With the advent of latest screening technology, the advantage of using SLO is its wide field of view, which can image a large part of the retina for better diagnosis of the retinal diseases. On the other hand, during the imaging process, artefacts such as eyelashes and eyelids are also imaged along with the retinal area. This brings a big challenge on how to exclude these artefacts. In this paper, we propose a novel approach to automatically extract out true retinal area from an SLO image based on image processing and machine learning approaches. To reduce the complexity of image processing tasks and provide a convenient primitive image pattern, we have grouped pixels into different regions based on the regional size and compactness, called superpixels. The framework then calculates image based features reflecting textural and structural information and classifies between retinal area and artefacts. The experimental evaluation results have shown good performance with an overall accuracy of 92%.
Muhammad Salman Haleem, Liangxiu Han, Jano I. van Hemert, Baihua Li, Alan D. Fleming
IEEE J. Biomed. Health Informatics3
2012 EnzML: multi-label prediction of enzyme classes using InterPro signatures
abstract
BACKGROUND: Manual annotation of enzymatic functions cannot keep up with automatic genome sequencing. In this work we explore the capacity of InterPro sequence signatures to automatically predict enzymatic function. RESULTS: We present EnzML, a multi-label classification method that can efficiently account also for proteins with multiple enzymatic functions: 50,000 in UniProt. EnzML was evaluated using a standard set of 300,747 proteins for which the manually curated Swiss-Prot and KEGG databases have agreeing Enzyme Commission (EC) annotations. EnzML achieved more than 98% subset accuracy (exact match of all correct Enzyme Commission classes of a protein) for the entire dataset and between 87 and 97% subset accuracy in reannotating eight entire proteomes: human, mouse, rat, mouse-ear cress, fruit fly, the S. pombe yeast, the E. coli bacterium and the M. jannaschii archaebacterium. To understand the role played by the dataset size, we compared the cross-evaluation results of smaller datasets, either constructed at random or from specific taxonomic domains such as archaea, bacteria, fungi, invertebrates, plants and vertebrates. The results were confirmed even when the redundancy in the dataset was reduced using UniRef100, UniRef90 or UniRef50 clusters. CONCLUSIONS: InterPro signatures are a compact and powerful attribute space for the prediction of enzymatic function. This representation makes multi-label machine learning feasible in reasonable time (30 minutes to train on 300,747 instances with 10,852 attributes and 2,201 class values) using the Mulan Binary Relevance Nearest Neighbours algorithm implementation (BR-kNN).
Luna De Ferrari, Stuart Aitken, Jano I. van Hemert, Igor Goryanin
BMC Bioinform.3
2012 Reducing Data Transfer in Service-Oriented Architectures: The Circulate Approach
abstract
As the number of services and the size of data involved in workflows increases, centralized orchestration techniques are reaching the limits of scalability. When relying on web services without third-party data transfer, a standard orchestration model needs to pass all data through a centralized engine, which results in unnecessary data transfer and the engine to become a bottleneck to the execution of a workflow. As a solution, this paper presents and evaluates Circulate, an alternative service-oriented architecture which facilitates an orchestration model of central control in combination with a choreography model of optimized distributed data transport. Extensive performance analysis through the PlanetLab framework is conducted on a web service-based implementation over a range of Internet-scale configurations which mirror scientific workflow environments. Performance analysis concludes that our architecture's optimized model of data transport speeds up the execution time of workflows, consistently outperforms standard orchestration and scales with data and node size. Furthermore, Circulate is a less-intrusive solution as individual services do not have to be reconfigured in order to take part in a workflow.
Adam Barker, Jon B. Weissman, Jano I. van Hemert
IEEE Trans. Serv. Comput.3
2011 Automatically identifying and annotating mouse embryo gene expression patterns
abstract
MOTIVATION: Deciphering the regulatory and developmental mechanisms for multicellular organisms requires detailed knowledge of gene interactions and gene expressions. The availability of large datasets with both spatial and ontological annotation of the spatio-temporal patterns of gene expression in mouse embryo provides a powerful resource to discover the biological function of embryo organization. Ontological annotation of gene expressions consists of labelling images with terms from the anatomy ontology for mouse development. If the spatial genes of an anatomical component are expressed in an image, the image is then tagged with a term of that anatomical component. The current annotation is done manually by domain experts, which is both time consuming and costly. In addition, the level of detail is variable, and inevitably errors arise from the tedious nature of the task. In this article, we present a new method to automatically identify and annotate gene expression patterns in the mouse embryo with anatomical terms. RESULTS: The method takes images from in situ hybridization studies and the ontology for the developing mouse embryo, it then combines machine learning and image processing techniques to produce classifiers that automatically identify and annotate gene expression patterns in these images. We evaluate our method on image data from the EURExpress study, where we use it to automatically classify nine anatomical terms: humerus, handplate, fibula, tibia, femur, ribs, petrous part, scapula and head mesenchyme. The accuracy of our method lies between 70% and 80% with few exceptions. We show that other known methods have lower classification performance than ours. We have investigated the images misclassified by our method and found several cases where the original annotation was not correct. This shows our method is robust against this kind of noise. AVAILABILITY: The annotation result and the experimental dataset in the article can be freely accessed at http://www2.docm.mmu.ac.uk/STAFF/L.Han/geneannotation/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Liangxiu Han, Jano I. van Hemert, Richard A. Baldock
Bioinform.2
2011 A user-friendly web portal for T-Coffee on supercomputers
abstract
BACKGROUND: Parallel T-Coffee (PTC) was the first parallel implementation of the T-Coffee multiple sequence alignment tool. It is based on MPI and RMA mechanisms. Its purpose is to reduce the execution time of the large-scale sequence alignments. It can be run on distributed memory clusters allowing users to align data sets consisting of hundreds of proteins within a reasonable time. However, most of the potential users of this tool are not familiar with the use of grids or supercomputers. RESULTS: In this paper we show how PTC can be easily deployed and controlled on a super computer architecture using a web portal developed using Rapid. Rapid is a tool for efficiently generating standardized portlets for a wide range of applications and the approach described here is generic enough to be applied to other applications, or to deploy PTC on different HPC environments. CONCLUSIONS: The PTC portal allows users to upload a large number of sequences to be aligned by the parallel version of TC that cannot be aligned by a single machine due to memory and execution time constraints. The web portal provides a user-friendly solution.
Josep Rius 0002, Fernando Cores, Francesc Solsona Tehàs, Jano I. van Hemert, Jos Koetsier, Cédric Notredame
BMC Bioinform.4
2011 Special Issue: Portals for life sciences - Providing intuitive access to bioinformatic tools
abstract
Abstract The topic ‘Portals for life sciences’ includes various research fields, on the one hand many different topics out of life sciences, e.g. mass spectrometry, on the other hand portal technologies and different aspects of computer science, such as usability of user interfaces and security of systems. The main aspect about portals is to simplify the user's interaction with computational resources that are concerted to a supported application domain. Copyright © 2010 John Wiley & Sons, Ltd.
Sandra Gesing, Jano I. van Hemert, Péter Kacsuk, Oliver Kohlbacher
Concurr. Comput. Pract. Exp.2
2011 Generating web-based user interfaces for computational science
abstract
Abstract Scientific gateways in the form of web portals are becoming the popular approach to share knowledge and resources around a topic in a community of researchers. Unfortunately, the development of web portals is expensive and requires specialists skills. Commercial and more generic web portals have a much larger user base and can afford this kind of development. Here we present two solutions that address this problem in the area of portals for scientific computing; both take the same approach. The whole process of designing, delivering and maintaining a portal can be made more cost‐effective by generating a portal from a description rather than programming in the traditional sense. We show four successful use cases to show how this process works and the results it can deliver. Copyright © 2010 John Wiley & Sons, Ltd.
Jano I. van Hemert, Jos Koetsier, Livia Torterolo, Ivan Porro, Maurizio Melato, Roberto Barbera
Concurr. Comput. Pract. Exp.1
2011 A generic parallel processing model for facilitating data mining and integration
Liangxiu Han, Chee Sun Liew, Jano I. van Hemert, Malcolm P. Atkinson 0001
Parallel Comput.3
2010 TOPP goes Rapid The OpenMS Proteomics Pipeline in a Grid-Enabled Web Portal
abstract
Proteomics, the study of all the proteins contained in a particular sample, e.g., a cell, is a key technology in current biomedical research. The complexity and volume of proteomics data sets produced by mass spectrometric methods clearly suggests the use of grid-based high-performance computing for analysis. TOPP and OpenMS are open-source packages for proteomics data analysis, however, they do not provide support for Grid computing. In this work we present a portal interface for high-throughput data analysis with TOPP. The portal is based on Rapid, a tool for efficiently generating standardized port lets for a wide range of applications. The web-based interface allows the creation and editing of user-defined pipelines and their execution and monitoring on a Grid infrastructure. The portal also supports several file transfer protocols for data staging. It thus provides a simple and complete solution to high-throughput proteomics data analysis for inexperienced users through a convenient portal interface.
Sandra Gesing, Jano I. van Hemert, Jos Koetsier, Andreas Bertsch, Oliver Kohlbacher
CCGRID2
2010 Federated Enactment of Workflow Patterns
Gagarine Yaikhom, Chee Sun Liew, Liangxiu Han, Jano I. van Hemert, Malcolm P. Atkinson 0001, Amy Krause
Euro-Par (1)4
2010 Towards optimising distributed data streaming graphs using parallel streams
abstract
Modern scientific collaborations have opened up the opportunity of solving complex problems that involve multi-disciplinary expertise and large-scale computational experiments. These experiments usually involve large amounts of data that are located in distributed data repositories running various software systems, and managed by different organisations. A common strategy to make the experiments more manageable is executing the processing steps as a workflow. In this paper, we look into the implementation of fine-grained data-flow between computational elements in a scientific workflow as streams. We model the distributed computation as a directed acyclic graph where the nodes represent the processing elements that incrementally implement specific subtasks. The processing elements are connected in a pipelined streaming manner, which allows task executions to overlap. We further optimise the execution by splitting pipelines across processes and by introducing extra parallel streams. We identify performance metrics and design a measurement tool to evaluate each enactment. We conducted experiments to evaluate our optimisation strategies with a real world problem in the Life Sciences---EURExpress-II. The paper presents our distributed data-handling model, the optimisation and instrumentation strategies and the evaluation experiments. We demonstrate linear speed up and argue that this use of data-streaming to enable both overlapped pipeline and parallelised enactment is a generally applicable optimisation strategy.
Chee Sun Liew, Malcolm P. Atkinson 0001, Jano I. van Hemert, Liangxiu Han
HPDC3
2009 Automating Gene Expression Annotation for Mouse Embryo
Liangxiu Han, Jano I. van Hemert, Richard A. Baldock, Malcolm P. Atkinson 0001
ADMA2
2009 An E-infrastructure to Support Collaborative Embryo Research
abstract
Within the context of the EU design study developmental gene expression map, we identify a set of challenges when facilitating collaborative research on early human embryo development. These challenges bring forth requirements, for which we have identified solutions and technology. We summarise our solutions and demonstrate how they integrate to form an e-infrastructure to support collaborative research in this area of developmental biology.
Adam Barker, Jano I. van Hemert, Richard A. Baldock, Malcolm P. Atkinson 0001
CCGRID2
2009 Rapid Chemistry Portals through Engaging Researchers
abstract
In this study, we apply a methodology for rapid development of portlets for scientific computing to the domain of computational chemistry. We report results in terms of the portals delivered, the changes made to our methodology and the experience gained in terms of interaction with domain-specialists. Our major contributions are: several web portals for teaching and research in computational chemistry; a successful transition to having our development tool used by the domain specialist as opposed by us, the developers; and an updated version of our methodology and technology for rapid development of portlets for computational science, which is free for anyone to pick up and use.
Jos Koetsier, Andrew Turner, Patricia Richardson, Jano I. van Hemert
eScience4
2008 Orchestrating Data-Centric Workflows
abstract
When orchestrating data-centric workflows as are commonly found in the sciences, centralised servers can become a bottleneck to the performance of a workflow; output from service invocations are normally transferred via a centralised orchestration engine, when they should be passed directly to where they are needed at the next service in the workflow. To address this performance bottleneck, this paper presents a lightweight hybrid workflow architecture and concrete API, based on a centralised control flow, distributed data flow model. Our architecture maintains the robustness and simplicity of centralised orchestration, but facilitates choreography by allowing services to exchange data directly with one another, reducing data that needs to be transferred through a centralised server. Furthermore our architecture is standards compliment, flexible and is a non-disruptive solution; service definitions do not have to be altered prior to enactment.
Adam Barker, Jon B. Weissman, Jano I. van Hemert
CCGRID3
2008 Graph Colouring Heuristics Guided by Higher Order Graph Properties
István Juhos, Jano I. van Hemert
EvoCOP2
2008 Eliminating the middleman: peer-to-peer dataflow
abstract
Efficiently executing large-scale, data-intensive workflows such as Montage must take into account the volume and pattern of communication. When orchestrating data-centric workflows, centralised servers common to standard workflow systems can become a bottleneck to performance. However, standards-based workflow systems that rely on centralisation, e.g., Web service based frameworks, have many other benefits such as a wide user base and sustained support.
Adam Barker, Jon B. Weissman, Jano I. van Hemert
HPDC3
2006 Improving Graph Colouring Algorithms and Heuristics Using a Novel Representation
István Juhos, Jano I. van Hemert
EvoCOP2
2006 Neighbourhood searches for the bounded diameter minimum spanning tree problem embedded in a VNS, EA, and ACO
abstract
We consider the Bounded Diameter Minimum Spanning Tree problem and describe four neighbourhood searches for it. They are used as local improvement strategies within a variable neighbourhood search (VNS), an evolutionary algorithm (EA) utilising a new encoding of solutions, and an ant colony optimisation (ACO). We compare the performance in terms of effectiveness between these three hybrid methods on a suite of popular benchmark instances, which contains instances too large to solve by current exact methods. Our results show that the EA and the ACO outperform the VNS on almost all used benchmark instances. Furthermore, the ACO yields most of the time better solutions than the EA in long-term runs, whereas the EA dominates when the computation time is strongly restricted.
Martin Gruber, Jano I. van Hemert, Günther R. Raidl
GECCO2
2006 Evolving Combinatorial Problem Instances That Are Difficult to Solve
abstract
This paper demonstrates how evolutionary computation can be used to acquire difficult to solve combinatorial problem instances. As a result of this technique, the corresponding algorithms used to solve these instances are stress-tested. The technique is applied in three important domains of combinatorial optimisation, binary constraint satisfaction, Boolean satisfiability, and the travelling salesman problem. The problem instances acquired through this technique are more difficult than the ones found in popular benchmarks. In this paper, these evolved instances are analysed with the aim to explain their difficulty in terms of structural properties, thereby exposing the weaknesses of corresponding algorithms.
Jano I. van Hemert
Evol. Comput.1
2005 Complexity transitions in evolutionary algorithms: evaluating the impact of the initial population
abstract
This paper proposes an evolutionary approach for the composition of solutions in an incremental way. The approach is based on the metaphor of transitions in complexity discussed in the context of evolutionary biology. Partially defined solutions interact and evolve into aggregations until a full solution for the problem at hand is found. The impact of the initial population on the outcome and the dynamics of the process is evaluated using the domain of binary constraint satisfaction problems.
Anne Defaweux, Tom Lenaerts, Jano I. van Hemert, Johan Parent
Congress on Evolutionary Computation3
2005 Property Analysis of Symmetric Travelling Salesman Problem Instances Acquired Through Evolution
Jano I. van Hemert
EvoCOP1
2005 Heuristic Colour Assignment Strategies for Merge Models in Graph Colouring
István Juhos, Attila Tóth, Jano I. van Hemert
EvoCOP3
2005 Transition models as an incremental approach for problem solving in evolutionary algorithms
abstract
This paper proposes an incremental approach for building solutions using evolutionary computation. It presents a simple evolutionary model called a Transition model in which partial solutions are constructed that interact to provide larger solutions. An evolutionary process is used to merge these partial solutions into a full solution for the problem at hand. The paper provides a preliminary study on the evolutionary dynamics of this model as well as an empirical comparison with other evolutionary techniques on binary constraint satisfaction.
Anne Defaweux, Tom Lenaerts, Jano I. van Hemert, Johan Parent
GECCO3
2004 A Study into Ant Colony Optimisation, Evolutionary Computation and Constraint Programming on Binary Constraint Satisfaction Problems
Jano I. van Hemert, Christine Solnon
EvoCOP1
2004 Binary Merge Model Representation of the Graph Colouring Problem
István Juhos, Attila Tóth, Jano I. van Hemert
EvoCOP3
2004 Dynamic Routing Problems with Fruitful Regions: Models and Evolutionary Computation
Jano I. van Hemert, Han La Poutré
PPSN1
2004 Phase Transition Properties of Clustered Travelling Salesman Problem Instances Generated with Evolutionary Computation
Jano I. van Hemert, Neil Urquhart
PPSN1
2003 Evolving binary constraint satisfaction problem instances that are difficult to solve
abstract
We present a study on the difficulty of solving binary constraint satisfaction problems where an evolutionary algorithm is used to explore the space of problem instances. By directly altering the structure of problem instances and by evaluating the effort it takes to solve them using a complete algorithm we show that the evolutionary algorithm is able to detect problem instances that are harder to solve than those produced with conventional methods. Results from the search of the evolutionary algorithm confirm conjectures about where the most difficult to solve problem instances can be found with respect to the tightness.
Jano I. van Hemert
IEEE Congress on Evolutionary Computation1
2003 Comparing evolutionary algorithms on binary constraint satisfaction problems
abstract
Constraint handling is not straightforward in evolutionary algorithms (EAs) since the usual search operators, mutation and recombination, are 'blind' to constraints. Nevertheless, the issue is highly relevant, for many challenging problems involve constraints. Over the last decade, numerous EAs for solving constraint satisfaction problems (CSP) have been introduced and studied on various problems. The diversity of approaches and the variety of problems used to study the resulting algorithms prevents a fair and accurate comparison of these algorithms. This paper aligns related work by presenting a concise overview and an extensive performance comparison of all these EAs on a systematically generated test suite of random binary CSPs. The random problem instance generator is based on a theoretical model that fixes deficiencies of models and respective generators that have been formerly used in the evolutionary computing field.
Bart G. W. Craenen, A. E. Eiben, Jano I. van Hemert
IEEE Trans. Evol. Comput.3
2002 Measuring the Searched Space to Guide Efficiency: The Principle and Evidence on Constraint Satisfaction
Jano I. van Hemert, Thomas Bäck
PPSN1
2001 Adaptive Genetic Programming Applied to New and Existing Simple Regression Problems
Jeroen Eggermont, Jano I. van Hemert
EuroGP2
1999 Population dynamics and emerging mental features in AEGIS
A. E. Eiben, Donatello Elia, Jano I. van Hemert
GECCO3
1999 A Comparison of Genetic Programming Variants for Data Classification
Jeroen Eggermont, A. E. Eiben, Jano I. van Hemert
IDA3
1998 Solving Binary Constraint Satisfaction Problems Using Evolutionary Algorithms with an Adaptive Fitness Function
A. E. Eiben, Jano I. van Hemert, Elena Marchiori, Adri G. Steenbeek
PPSN2