Vincent Danjean

dblp:10/1770 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
0since 2021 · last 2018
0000-0002-1304-8404ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 77% Visualization and visual analytics · 23%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Geometric modeling and processing › surface parameterization
mesh layout
0.112010
Binary Mesh Partitioning for Cache-Efficient Visualization · IEEE Trans. Vis. Comput. Graph. 2010
Memory systems › cache
cache-aware algorithm design
0.112010
Binary Mesh Partitioning for Cache-Efficient Visualization · IEEE Trans. Vis. Comput. Graph. 2010
Memory systems › cache
cache-oblivious algorithms
0.112010
Binary Mesh Partitioning for Cache-Efficient Visualization · IEEE Trans. Vis. Comput. Graph. 2010
Bioinformatics and computational biology
genetic epidemiology
0.112006
ALTree: association detection and localization of susceptibility sites using haplotype phylogenetic trees · Bioinform. 2006
Bioinformatics and computational biology › statistical genetics › haplotype analysis
haplotype-based association testing
0.112006
ALTree: association detection and localization of susceptibility sites using haplotype phylogenetic trees · Bioinform. 2006
Visualization and visual analytics › information visualization
large-scale data visualization
0.012010
Binary Mesh Partitioning for Cache-Efficient Visualization · IEEE Trans. Vis. Comput. Graph. 2010
Bioinformatics and computational biology
phylogenetics
0.012006
ALTree: association detection and localization of susceptibility sites using haplotype phylogenetic trees · Bioinform. 2006

Methods — techniques the papers use, named apart from their topics

cache-oblivious analysis · 0.2BSP tree · 0.2phylogenetic tree construction · 0.1
YearPublicationVenuePosition
2018 A visual performance analysis framework for task-based parallel applications running on hybrid clusters
abstract
Summary Programming paradigms in High‐Performance Computing have been shifting toward task‐based models that are capable of adapting readily to heterogeneous and scalable supercomputers. The performance of task‐based application heavily depends on the runtime scheduling heuristics and on its ability to exploit computing and communication resources. Unfortunately, the traditional performance analysis strategies are unfit to fully understand task‐based runtime systems and applications: they expect a regular behavior with communication and computation phases, while task‐based applications demonstrate no clear phases. Moreover, the finer granularity of task‐based applications typically induces a stochastic behavior that leads to irregular structures that are difficult to analyze. Furthermore, the combination of application structure, scheduler, and hardware information is generally essential to understand performance issues. This paper presents a flexible framework that enables one to combine several sources of information and to create custom visualization panels allowing to understand and pinpoint performance problems incurred by bad scheduling decisions in task‐based applications. Three case‐studies using StarPU‐MPI, a task‐based multi‐node runtime system, are detailed to show how our framework can be used to study the performance of the well‐known Cholesky factorization. Performance improvements include a better task partitioning among the multi‐(GPU, core) to get closer to theoretical lower bounds, improved MPI pipelining in multi‐(node, core, GPU) to reduce the slow start, and changes in the runtime system to increase MPI bandwidth, with gains of up to 13% in the total makespan.
Vinícius Garcia Pinto, Lucas Mello Schnorr, Luka Stanisic, Arnaud Legrand, Samuel Thibault, Vincent Danjean
Concurr. Comput. Pract. Exp.6
2015 Design and analysis of scheduling strategies for multi-CPU and multi-GPU architectures
João V. F. Lima, Vincent Danjean, Bruno Raffin, Nicolas Maillard
Parallel Comput.3
2012 Exploiting Concurrent GPU Operations for Efficient Work Stealing on Multi-GPUs
abstract
The race for Exascale computing has naturally led the current technologies to converge to multi-CPU/multi-GPU computers, based on thousands of CPUs and GPUs interconnected by PCI-Express buses or interconnection networks. To exploit this high computing power, programmers have to solve the issue of scheduling parallel programs on hybrid architectures. And, since the performance of a GPU increases at a much faster rate than the throughput of a PCI bus, data transfers must be managed efficiently by the scheduler. This paper targets multi-GPU compute nodes, where several GPUs are connected to the same machine. To overcome the data transfer limitations on such platforms, the available soft wares compute, usually before the execution, a mapping of the tasks that respects their dependencies and minimizes the global data transfers. Such an approach is too rigid and it cannot adapt the execution to possible variations of the system or to the application's load. We propose a solution that is orthogonal to the above mentioned: extensions of the Xkaapi software stack that enable to exploit full performance of a multi-GPUs system through asynchronous GPU tasks. Xkaapi schedules tasks by using a standard Work Stealing algorithm and the runtime efficiently exploits concurrent GPU operations. The runtime extensions make it possible to overlap the data transfers and the task executions on current generation of GPUs. We demonstrate that the overlapping capability is at least as important as computing a scheduling decision to reduce completion time of a parallel program. Our experiments on two dense linear algebra problems (Matrix Product and Cholesky factorization) show that our solution is highly competitive with other soft wares based on static scheduling. Moreover, we are able to sustain the peak performance (approx. 310 GFlop/s) on DGEMM, even for matrices that cannot be stored entirely in one GPU memory. With eight GPUs, we archive a speed-up of 6.74 with respect to single-GPU. The performance of our Cholesky factorization, with more complex dependencies between tasks, outperforms the state of the art single-GPU MAGMA code.
João V. F. Lima, Nicolas Maillard, Vincent Danjean
SBAC-PAD4
2010 Binary Mesh Partitioning for Cache-Efficient Visualization
abstract
One important bottleneck when visualizing large data sets is the data transfer between processor and memory. Cache-aware (CA) and cache-oblivious (CO) algorithms take into consideration the memory hierarchy to design cache efficient algorithms. CO approaches have the advantage to adapt to unknown and varying memory hierarchies. Recent CA and CO algorithms developed for 3D mesh layouts significantly improve performance of previous approaches, but they lack of theoretical performance guarantees. We present in this paper a {\schmi O}(N\log N) algorithm to compute a CO layout for unstructured but well shaped meshes. We prove that a coherent traversal of a N-size mesh in dimension d induces less than N/B+{\schmi O}(N/M;{1/d}) cache-misses where B and M are the block size and the cache size, respectively. Experiments show that our layout computation is faster and significantly less memory consuming than the best known CO algorithm. Performance is comparable to this algorithm for classical visualization algorithm access patterns, or better when the BSP tree produced while computing the layout is used as an acceleration data structure adjusted to the layout. We also show that cache oblivious approaches lead to significant performance increases on recent GPU architectures.
Marc Tchiboukdjian, Vincent Danjean, Bruno Raffin
IEEE Trans. Vis. Comput. Graph.2
2006 ALTree: association detection and localization of susceptibility sites using haplotype phylogenetic trees
abstract
Abstract Summary: Finding the genes involved in complex diseases susceptibility and among those genes, localizing the variant sites explaining this susceptibility is a major goal of genetic epidemiology. In this context, haplotypic methods that use the joint information on several markers may be of particular interest. When the number of haplotypes is large, a grouping may be required. Phylogenetic trees allow such groupings of haplotypes based on their evolutionary history and may help in the detection and localization of disease susceptibility sites. In this paper, we present a new software to perform phylogeny-based association and localization analysis. Availability: The software package, including all documentation and example files, is freely available at . It is distributed under the GPL license. Contact: [email protected]
Claire Bardel, Vincent Danjean, Emmanuelle Génin
Bioinform.2
2005 An Efficient Multi-level Trace Toolkit for Multi-threaded Applications
Vincent Danjean, Raymond Namyst, Pierre-André Wacrenier
Euro-Par1
2003 Controlling Kernel Scheduling from User Space: An Approach to Enhancing Applications' Reactivity to I/O Events
Vincent Danjean, Raymond Namyst
HiPC1
2002 Improving Reactivity to I/O Events in Multithreaded Environments Using a Uniform, Scheduler-Centric API
Luc Bougé, Vincent Danjean, Raymond Namyst
Euro-Par2