Jan Lemeire

dblp:86/3718 · DBLP profile ↗
← Back
21ranked-venue papers
8as first author
1since 2021 · last 2023
0000-0002-2106-448XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-authorSystems, architecture and hardware · 6 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Probabilistic and Bayesian machine learning · 87% Language models and text generation · 13%
Theoretical computer science
1 paper
Information theory · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.422016
Conditional Independencies under the Algorithmic Independence of Conditionals · J. Mach. Learn. Res. 2016
Information-geometric approach to inferring causal directions · Artif. Intell. 2012
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian network
0.212016
Conditional Independencies under the Algorithmic Independence of Conditionals · J. Mach. Learn. Res. 2016
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
conditional independence
0.212016
Conditional Independencies under the Algorithmic Independence of Conditionals · J. Mach. Learn. Res. 2016
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability
faithfulness
0.212016
Conditional Independencies under the Algorithmic Independence of Conditionals · J. Mach. Learn. Res. 2016
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning
bayesian network structure learning
0.212013
Algorithms for discovery of multiple Markov boundaries · J. Mach. Learn. Res. 2013
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
0.212013
Algorithms for discovery of multiple Markov boundaries · J. Mach. Learn. Res. 2013
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.212013
Algorithms for discovery of multiple Markov boundaries · J. Mach. Learn. Res. 2013
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
markov blanket discovery
0.212013
Algorithms for discovery of multiple Markov boundaries · J. Mach. Learn. Res. 2013
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
bivariate causal discovery
0.112012
Information-geometric approach to inferring causal directions · Artif. Intell. 2012
Information theory
information geometry
0.112012
Information-geometric approach to inferring causal directions · Artif. Intell. 2012
Reconfigurable computing and FPGAs
FPGA accelerator
0.012013
Performance and toolchain of a combined GPU/FPGA desktop (abstract only) · FPGA 2013

Methods — techniques the papers use, named apart from their topics

information geometry · 0.3algorithmic independence of conditionals · 0.2roofline model · 0.2markov boundary · 0.2high-level synthesis · 0.2constraint-based causal discovery · 0.2OpenCL · 0.2
YearPublicationVenuePosition
2023 Analysis of the analytical performance models for GPUs and extracting the underlying Pipeline model
Jan Lemeire, Jan G. Cornelis, Elias Konstantinidis
J. Parallel Distributed Comput.1
2019 The Pipeline Performance Model: A Generic Executable Performance Model for GPUs
abstract
This paper presents the pipeline performance model, a generic GPU performance model, which helps understand the performance of GPU code by using a code representation that is very close to the source code. The code is represented by a graph in which the nodes correspond to the source code instructions and the edges to data dependences between them. Furthermore, each node is enhanced with two latencies that characterize the instruction's time behavior on the GPU. This graph, together with a simple characterization of the GPU and the execution configuration, is used by a simulator to mimic the execution of the code. We validate the model on the micro-benchmarks used to determine the latencies and on a matrix multiplication kernel, both on an NVIDIA Fermi and an NVIDIA Pascal GPU. Initial results show that the simulated times follow the measured times, with acceptable errors, for a wide occupancy range. We argue that to achieve better accuracies it is necessary to further refine the model to take into account the complexity of memory access and warp scheduling, especially for more recent GPUs.
Jan G. Cornelis, Jan Lemeire
PDP2
2018 Efficiency analysis methodology of FPGAs based on lost frequencies, area and cycles
Jan Lemeire, Bruno da Silva 0001, An Braeken, Jan G. Cornelis, Abdellah Touhafi
J. Parallel Distributed Comput.1
2018 Optimized wavelet-based texture representation and streaming for GPU texture mapping
Bob Andries, Jan Lemeire, Adrian Munteanu 0001
Multim. Tools Appl.2
2017 Scalable texture compression using the wavelet transform
Bob Andries, Jan Lemeire, Adrian Munteanu 0001
Vis. Comput.2
2016 Microbenchmarks for GPU Characteristics: The Occupancy Roofline and the Pipeline Model
abstract
In this paper we present microbenchmarks in OpenCL to measure the most important performance characteristics of GPUs. Microbenchmarks try to measure individual characteristics that influence the performance. First, performance, in operations or bytes per second, is measured with respect to the occupancy and as such provides an occupancy roofline curve. The curve shows at which occupancy level peak performance is reached. Second, when considering the cycles per instruction of each compute unit, we measure the two most important characteristics of an instruction: its issue and completion latency. This is based on modeling each compute unit as a pipeline for computations and a pipeline for the memory access. We also measure some specific characteristics: the influence of independent instructions within a kernel and thread divergence. We argue that these are the most important characteristics for understanding the performance and predicting performance. The results for several Nvidia and AMD GPUs are provided. A free java application containing the microbenchmarks is available on www.gpuperformance.org.
Jan Lemeire, Jan G. Cornelis, Laurent Segers
PDP1
2016 Conditional Independencies under the Algorithmic Independence of Conditionals
abstract
In this paper we analyze the relationship between faithfulness and the more recent condition of algorithmic Independence of Conditionals (IC) with respect to the Conditional Independencies (CIs) they allow. Both conditions have been extensively used for causal inference by refuting factorizations for which the condition does not hold. Violation of faithfulness happens when there are CIs that do not follow from the Markov condition. For those CIs, non-trivial constraints among some parameters of the Conditional Probability Distributions (CPDs) must hold. When such a constraint is defined over parameters of different CPDs, we prove that IC is also violated unless the parameters have a simple description. To understand which non-Markovian CIs are permitted we define a new condition closely related to IC: the Independence from Product Constraints (IPC). The condition reflects that CIs might be the result of specific parameterizations of individual CPDs but not from constraints on parameters of different CPDs. In that sense it is more restrictive than IC: parameters may have a simple description. On the other hand, IC also excludes other forms of algorithmic dependencies between CPDs. Finally, we prove that on top of the CIs permitted by the Markov condition (faithfulness), IPC allows non-minimality, deterministic relations and what we called proportional CPDs. These are the only cases in which a CI follows from a specific parameterization of a single CPD.
Jan Lemeire
J. Mach. Learn. Res.1
2016 The Forward Procedure for HSMMs based on Expected Duration
Jan Lemeire, Francesco Cartella
IEEE Signal Process. Lett.1
2016 A novel MPI reduction algorithm resilient to imbalances in process arrival times
Petar Marendic, Jan Lemeire, Dean Vucinic, Peter Schelkens
J. Supercomput.2
2015 Heterogeneous Acceleration of Volumetric JPEG 2000
abstract
We present the implementation of a volumetric JPEG 2000 codec as a real-world use case of software acceleration with GPUs and multi-core CPUs. We present a generic methodology to accelerate existing code written in C with OpenCL. Furthermore, we account for the volumetric nature of the processed data and formulate associated optimization guidelines. The resulting software can exploit different accelerator types - GPUs and multi-core CPUs - and delivers a decent speedup on a variety of hardware platforms for a relatively small effort.
Jan G. Cornelis, Jan Lemeire, Tim Bruylants, Peter Schelkens
PDP2
2014 Optimized quantization of wavelet subbands for high quality real-time texture compression
abstract
This paper proposes a new wavelet-based system for fixed-rate texture compression in 3D graphics applications. An analysis of the optimized quantization of the wavelet subbands is carried out, focusing on Lloyd-Max quantization of subband blocks as well as on quantization techniques employed in existing texture compression formats on GPUs. The results demonstrate that the proposed compression technique yields higher performance compared to existing GPU texture compression schemes and previously proposed transformed-based systems. Moreover, compared to conventional schemes, the wide range of quantization schemes results in a much wider range of available bitrates. Additionally, the proposed scheme offers real-time execution, being suitable for real time rendering applications.
Bob Andries, Jan Lemeire, Adrian Munteanu 0001
ICIP2
2013 Detecting Marginal and Conditional Independencies between Events and Learning Their Causal Structure
Jan Lemeire, Stijn Meganck, Albrecht Zimmermann, Thomas Dhollander
ECSQARU1
2013 Performance and toolchain of a combined GPU/FPGA desktop (abstract only)
abstract
Low-power, high-performance computing nowadays relies on accelerator cards to speed up the calculations. Combining the power of GPUs with the flexibility of FPGAs enlarges the scope of problems that can be accelerated [2, 3]. We describe the performance analysis of a desktop equipped with a GPU Tesla 2050 and an FPGA Virtex-6 LX240T. First, the balance between the I/O and the raw peak performance is depicted using the roofline model [4]. Next, the performance of a number of image processing algorithms is measured and the results are mapped onto the roofline graph. This allows to compare the GPU and the FPGA and also to optimize the algorithms for both accelerators. A programming toolchain is implemented, consisting of OpenCL for the GPU and several High-Level Synthesis compilers for the FPGA. Our results show that the HLS compilers outperform handwritten code and offer a performance comparable to the GPU. In addition the FPGA compilers reduce the development time by an order of magnitude, at the expense of an increased resource consumption. The roofline model also shows that both accelerators are equally limited by the input/output bandwidth to the host. A well-tuned accelerator-based codesign, identifying the parallelism, the computation and data patterns of different classes of algorithms, will enable to maximize the performance of the combined GPU/FPGA system [1].
Bruno da Silva 0001, An Braeken, Erik H. D'Hollander, Abdellah Touhafi, Jan G. Cornelis, Jan Lemeire
FPGA6
2013 Comparing and combining GPU and FPGA accelerators in an image processing context
abstract
Nowadays, processors alone cannot deliver what computation hungry image processing applications demand. An alternative is to use hardware accelerators such as Graphics Processing Units (GPUs) or Field Programmable Gate Arrays (FPGAs). Applications, however, exhibit different performance characteristics depending on the accelerator. This paper describes the hybrid platform and the programming environment that allows to efficiently create programs on a combined GPU/FPGA desktop. We use the roofline model to identify the most appropriate accelerator for each application and High-Level Synthesis (HLS) tools to reduce the FPGA development time. To introduce our platform and tool chain both accelerators are compared by implementing a basic image operation. Next, a promising algorithm is explored and implemented, splitting and distributing the work between GPU, FPGA and CPU in order to validate the hybrid concept. Our results show that their combination exhibits a higher performance for computational intensive image processing applications than a GPU only.
Bruno da Silva 0001, An Braeken, Erik H. D'Hollander, Abdellah Touhafi, Jan G. Cornelis, Jan Lemeire
FPL6
2013 Real-time texture sampling and reconstruction with wavelet filters
abstract
Currently, the use of the 2D wavelet transform in texture compression for real-time texture mapping on the GPU is limited. The main cause of this is the lack of real-time texture filtering implementations which do not require specialized hardware. This work proposes a novel system to perform 2D wavelet reconstruction and bilinear texture filtering using a high performance GPU shader. The system is able to generate a performant GLSL shader for arbitrary wavelet filter configurations. This goes beyond earlier works in the literature proposing Haar wavelet and Discrete Cosine Transform (DCT) implementations on the GPU. We analyse the shader performance and run-time complexity for several wavelet filters. The experimental results show that filters longer than Haar are deployable on the GPU while maintaining accurate texture filtering and real-time performance.
Bob Andries, Adrian Munteanu 0001, Jan Lemeire, Peter Schelkens
MMSP3
2013 Algorithms for discovery of multiple Markov boundaries
Alexander R. Statnikov, Jan Lemeire, Constantin F. Aliferis
J. Mach. Learn. Res.2
2012 An Investigation into the Performance of Reduction Algorithms under Load Imbalance
Petar Marendic, Jan Lemeire, Tom Haber, Dean Vucinic, Peter Schelkens
Euro-Par2
2012 Information-geometric approach to inferring causal directions
Dominik Janzing, Joris M. Mooij, Kun Zhang 0001, Jan Lemeire, Jakob Zscheischler, Povilas Daniusis, Bastian Steudel, Bernhard Schölkopf
Artif. Intell.4
2012 Conservative independence-based causal structure learning in absence of adjacency faithfulness
Jan Lemeire, Stijn Meganck, Francesco Cartella
Int. J. Approx. Reason.1
2011 Inferring the causal decomposition under the presence of deterministic relations
Jan Lemeire, Stijn Meganck, Francesco Cartella, Alexander R. Statnikov
ESANN1
2008 Mosaicing of Fibered Fluorescence Microscopy Video
Steve De Backer, Frans W. Cornelissen, Jan Lemeire, Rony Nuydens, Theo Meert, Peter Schelkens, Paul Scheunders
ACIVS3