Marc Tchiboukdjian

dblp:74/8300 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 100%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 77% Visualization and visual analytics · 23%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Geometric modeling and processing › surface parameterization
mesh layout
0.112010
Binary Mesh Partitioning for Cache-Efficient Visualization · IEEE Trans. Vis. Comput. Graph. 2010
Memory systems › cache
cache-aware algorithm design
0.112010
Binary Mesh Partitioning for Cache-Efficient Visualization · IEEE Trans. Vis. Comput. Graph. 2010
Memory systems › cache
cache-oblivious algorithms
0.112010
Binary Mesh Partitioning for Cache-Efficient Visualization · IEEE Trans. Vis. Comput. Graph. 2010
Visualization and visual analytics › information visualization
large-scale data visualization
0.012010
Binary Mesh Partitioning for Cache-Efficient Visualization · IEEE Trans. Vis. Comput. Graph. 2010

Methods — techniques the papers use, named apart from their topics

cache-oblivious analysis · 0.2BSP tree · 0.2
YearPublicationVenuePosition
2012 Hierarchical Local Storage: Exploiting Flexible User-Data Sharing Between MPI Tasks
abstract
With the advent of the multicore era, the number of cores per computational node is increasing faster than the amount of memory. This diminishing memory to core ratio sometimes even prevents pure MPI applications to benefit from all cores available on each node. A possible solution is to add a shared memory programming model like Open MP inside the application to share variables between Open MP threads that would otherwise be duplicated for each MPI task. Going to hybrid can thus improve the overall memory consumption, but may be a tedious task on large applications. To allow this data sharing without the overhead of mixing multiple programming models, we propose an MPI extension called Hierarchical Local Storage (HLS) that allows application developers to share common variables between MPI tasks on the same node. HLS is designed as a set of directives that preserve the original parallel semantics of the code and are compatible with C, C++ and Fortran languages and the Open MP programming model. This new mechanism is implemented inside a state-of-the-art MPI 1.3 compliant runtime called MPC. Experiments show that the HLS mechanism can effectively reduce memory consumption of HPC applications. Moreover, by reducing data duplication in the shared cache of modern multicores, the HLS mechanism can also improve performances of memory intensive applications.
Marc Tchiboukdjian, Patrick Carribault, Marc Pérache
IPDPS1
2011 Controlling cache utilization of HPC applications
abstract
This paper discusses the use of software cache partitioning techniques to study and improve cache behavior of HPC applications. Most existing studies use this partitioning to solve quality of service issues, like fair distribution of a shared cache among running processes. We believe that, in the HPC context of a single application being studied/optimized on the system, with a single thread per core, cache partitioning can be used in new and interesting ways.
Swann Perarnau, Marc Tchiboukdjian, Guillaume Huard
ICS2
2010 A Tighter Analysis of Work Stealing
Marc Tchiboukdjian, Nicolas Gast, Denis Trystram, Jean-Louis Roch, Julien Bernard 0001
ISAAC (2)1
2010 Binary Mesh Partitioning for Cache-Efficient Visualization
abstract
One important bottleneck when visualizing large data sets is the data transfer between processor and memory. Cache-aware (CA) and cache-oblivious (CO) algorithms take into consideration the memory hierarchy to design cache efficient algorithms. CO approaches have the advantage to adapt to unknown and varying memory hierarchies. Recent CA and CO algorithms developed for 3D mesh layouts significantly improve performance of previous approaches, but they lack of theoretical performance guarantees. We present in this paper a {\schmi O}(N\log N) algorithm to compute a CO layout for unstructured but well shaped meshes. We prove that a coherent traversal of a N-size mesh in dimension d induces less than N/B+{\schmi O}(N/M;{1/d}) cache-misses where B and M are the block size and the cache size, respectively. Experiments show that our layout computation is faster and significantly less memory consuming than the best known CO algorithm. Performance is comparable to this algorithm for classical visualization algorithm access patterns, or better when the BSP tree produced while computing the layout is used as an acceleration data structure adjusted to the layout. We also show that cache oblivious approaches lead to significant performance increases on recent GPU architectures.
Marc Tchiboukdjian, Vincent Danjean, Bruno Raffin
IEEE Trans. Vis. Comput. Graph.1