EDBT 2026 Demo / reviewers in the wild / expert
Thomas Lambert
dblp:19/11139
· DBLP profile ↗
16ranked-venue papers
4as first author
3since 2021 · last 2026
0000-0002-6517-6598ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 3 since 2021Theory of computation · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ALTOCUMULUS: Enabling Efficient Erasure Coding in IPFS
Mohammad Rizk, Shadi Ibrahim, Thomas Lambert |
CCGrid | 3 |
| 2022 | Stragglers' Detection in Big Data Analytic Systems: The Impact of Heartbeat ArrivalabstractSpeculative execution can significantly improve the performance of Big Data applications by launching other copies of stragglers (slow tasks). Stragglers detection plays an important role in the effectiveness of speculative execution. The methods employed to detect stragglers use the information extracted from the last received heartbeats which may be outdated when triggering detection. This, in turn, can mislead Big Data analytic systems to make wrong detection with high inaccuracy. To shed the light on this issue, we carry out extensive simulations to identify how heartbeat arrival, task starting times, and detection methods impact the accuracy of stragglers detection in Big Data analytic systems. We reveal that the asynchrony in heartbeat arrivals not only lead to marking normal tasks as stragglers (false positives) but can also result in overlooking real stragglers (false negatives). Thomas Lambert, Shadi Ibrahim, Twinkle Jain, David Guyon |
CCGRID | 1 |
| 2022 | PGPregel: an end-to-end system for privacy-preserving graph processing in geo-distributed data centersabstractGraph processing is a popular computing model for big data analytics. Emerging big data applications are often maintained in multiple geographically distributed (geo-distributed) data centers (DCs) to provide low-latency services to global users. Graph processing in geo-distributed DCs suffers from costly inter-DC data communications. Furthermore, due to increasing privacy concerns, geo-distribution imposes diverse, strict, and often asymmetric privacy regulations that constrain geo-distributed graph processing. Existing graph processing systems fail to address these two challenges. In this paper, we design and implement PGPregel, which is an end-to-end system that provides privacy-preserving graph processing in geo-distributed DCs with low latency and high utility. To ensure privacy, PGPregel smartly integrates Differential Privacy into graph processing systems with the help of two core techniques, namely sampling and combiners, to reduce the amount of inter-DC data transfer while preserving good accuracy of graph processing results. We implement our design in Giraph and evaluate it in real cloud DCs. Results show that PGPregel can preserve the privacy of graph data with low overhead and good accuracy. Amelie Chi Zhou, Ruibo Qiu, Thomas Lambert, Tristan Allard, Shadi Ibrahim, Amr El Abbadi |
SoCC | 3 |
| 2020 | Rethinking Operators Placement of Stream Data Application in the EdgeabstractMaximum Sustainable Throughput (MST) refers to the amount of data that a Data Stream Processing (DSP) system can ingest while keeping stable performance. It has been acknowledged as an accurate metric to evaluate the performance of stream data processing. Yet, existing operators placements continue to focus on latency and throughput, not MST, as main performance objective when deploying stream data applications in the Edge. In this paper, we argue that MST should be used as an optimization objective when placing operators. This is specially important in the Edge, where network bandwidth and data streams are highly dynamic. We demonstrate that through the design and evaluation of a MST-driven operators placement (based on constraint programming) for stream data applications. Through simulations, we show how existing placement strategies that target overall communications reduction often fail to keep up with the rate of data streams. Importantly, the constraint programming-based operators placement is able to sustain up to 5x increased data ingestion compared to baseline strategies. Thomas Lambert, David Guyon, Shadi Ibrahim |
CIKM | 1 |
| 2020 | Performance analysis and optimality results for data-locality aware tasks scheduling with replicated inputs
Olivier Beaumont, Thomas Lambert, Loris Marchal, Bastien Thomas |
Future Gener. Comput. Syst. | 2 |
| 2019 | On the Importance of Container Image Placement for Service Provisioning in the EdgeabstractEdge computing promises to extend Clouds by moving computation close to data sources to facilitate short-running and low-latency applications and services. Providing fast and predictable service provisioning time presets a new and mounting challenge, as the scale of Edge-servers grows and the heterogeneity of networks between them increases. This paper is driven by a simple question: can we place container images across Edge-servers in such a way that an image can be retrieved to any Edge-server fast and in a predictable time. To this end, we present KCBP and KCBP-WC, two container image placement algorithms which aim to reduce the maximum retrieval time of container images. KCBP and KCBP-WC are based on k-Center optimization. However, KCBP-WC tries to avoid placing large layers of a container image on the same Edge-server. Evaluations using trace-driven simulations show that KCBP and KCBP-WC can be applied to various network configurations and reduce the maximum retrieval time of container images by 1.1x to 4x compared to state-of-the-art placements (i.e., Best-Fit and Random). Jad Darrous, Thomas Lambert, Shadi Ibrahim |
ICCCN | 2 |
| 2019 | Recent Advances in Matrix Partitioning for Parallel Computing on Heterogeneous PlatformsabstractThe problem of partitioning dense matrices into sets of sub-matrices has received increased attention recently and is crucial when considering dense linear algebra and kernels with similar communication patterns on heterogeneous platforms. The problem of load balancing and minimizing communication is traditionally reducible to an optimization problem that involves partitioning a square into rectangles. This problem has been proven to be NP-Complete for an arbitrary number of partitions. In this paper, we present recent approaches that relax the restriction that all partitions be rectangles. The first approach uses an original mathematical technique to find the exact optimal partitioning. Due to the complexity of the technique, it has been developed for a small number of partitions only. However, even at a small scale, the optimal partitions found by this approach are often non-rectangular and sometimes non-intuitive. The second approach is the study of approximate partitioning methods utilizing recursive partitioning algorithms. In particular we use the work on optimal partitioning to improve pre-existing algorithms. In this paper we discuss the different perspectives this approach opens and present two algorithms, SNRPP which is a$\sqrt{\frac{3}{2}}$approximation, and NRPP which is a$\frac{2}{\sqrt{3}}$approximation. While sub-optimal, the NRRP approach works for an arbitrary number of partitions. We use the first exact approach to analyse how close to the known optimal solutions the NRRP algorithm is for small numbers of partitions. Olivier Beaumont, Brett A. Becker, Ashley M. DeFlumere, Lionel Eyraud-Dubois, Thomas Lambert, Alexey L. Lastovetsky |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | Allocation of Publisher/Subscriber Data Links on a Set of Virtual MachinesabstractThere is an increasing interest in applications where sensors or other devices (acting as publishers) may generate data that are processed and analyzed by specialized software components (acting as subscribers to this data) that extract useful information in a variety of scientific or industrial settings. As a result of the large volumes of data that are often generated, Cloud infrastructures may be used to handle the data links between publishers and subscribers. Assuming that a certain number of virtual machines of some given bandwidth have been booked for this purpose, the problem that this paper considers is how to allocate data links to the virtual machines so that the amount of data received by subscribers is maximized. An Integer Linear Programming formulation of the problem and two heuristics are presented, which are evaluated in a range of experiments. Thomas Lambert, Rizos Sakellariou |
IEEE CLOUD | 1 |
| 2018 | Using Static Allocation Algorithms for Matrix Matrix Multiplication on Multicores and GPUsabstractWe consider the problem of data allocation when performing matrix multiplication on a heterogeneous node, with multicores and GPUs. Classical (cyclic) allocations designed for homogeneous settings are not appropriate, but the advent of task-based runtime systems makes it possible to use more general allocations. Previous theoretical work has proposed square and cube partitioning algorithms aimed at minimizing data movement for matrix multiplication. We propose techniques to adapt these continuous square partitionings to allocating discrete tiles of a matrix, and strategies to adapt the static allocation at runtime. We use these techniques in an implementation of Matrix Multiplication based on the StarPU runtime system, and we show through extensive experiments that this implementation allows to consistently obtain a lower communication volume while improving slightly the execution time, compared to standard state-of-the-art dynamic strategies. Lionel Eyraud-Dubois, Thomas Lambert |
ICPP | 2 |
| 2018 | Scheduling series-parallel task graphs to minimize peak memory
Enver Kayaaslan, Thomas Lambert, Loris Marchal, Bora Uçar |
Theor. Comput. Sci. | 2 |
| 2016 | Cuboid Partitioning for Parallel Matrix Multiplication on Heterogeneous Platforms
Olivier Beaumont, Lionel Eyraud-Dubois, Thomas Lambert |
Euro-Par | 3 |
| 2016 | A New Approximation Algorithm for Matrix Partitioning in Presence of Strongly Heterogeneous ProcessorsabstractIn this paper, we consider the problem of partitioning a square into aset of zones of prescribed areas, while minimizing the overall size oftheir projections onto horizontal and vertical axes. This problemtypically arises when considering the amount of communications inducedwhen partitioning matrices for dense linear algebra kernels onto a setof heterogeneous processors. It has been first introduced for matrixmultiplication in the 2000's, with a best known approximation ratiowas 1.75. Since then, two main new ingredients have beenintroduced. First, Lastovetsky et al. proposed a special partitioningin the case of 2 or 3 strongly heterogeneous processors, as in thecase of a platform made of CPUs and GPUs, relaxing the constraint of arectangular based partitioning. Second, Nagamochi et al. haveintroduced clever recursive partitioning techniques and proved, thanksto a careful analysis, that their algorithm achieves a 1.25approximation ratio. In this paper, we combine both ingredients inorder to obtain a non-rectangular recursive partitioning (NRRP), whoseapproximation ratio is 2/√3 ≃ 1.15. Moreover, we observe on a large set of realistic platforms built from CPUs and GPUs that this proposed NRRP algorithm allows to achieve very efficient partitionings on all considered cases. Olivier Beaumont, Lionel Eyraud-Dubois, Thomas Lambert |
IPDPS | 3 |
| 2015 | Comparison of Static and Runtime Resource Allocation Strategies for Matrix MultiplicationabstractThe tremendous increase in the size and heterogeneity of supercomputers makes it very difficult to predict the performance of a scheduling algorithm. In this context, relying on purely static scheduling and resource allocation strategies, that make scheduling and allocation decisions based on the dependency graph and the platform description, is expected to lead to large and unpredictable make spans whenever the behavior of the platform does not match the predictions. For this reason, the common practice in most runtime libraries is to rely on purely dynamic scheduling strategies, that make short-sighted scheduling decisions at runtime based on the estimations of the duration of the different tasks on the different available resources and on the state of the machine. In this paper, we consider the special case of Matrix Multiplication, for which a number of static allocation algorithms to minimize the amount of communications have been proposed. Through a set of extensive simulations, we analyse the behaviour of static, dynamic, and hybrid strategies, and we assess the possible benefits of introducing more static knowledge and allocation decisions in runtime libraries. Olivier Beaumont, Lionel Eyraud-Dubois, Abdou Guermouche, Thomas Lambert |
SBAC-PAD | 4 |
| 2015 | Comments on the hierarchically structured bin packing problem
Thomas Lambert, Loris Marchal, Bora Uçar |
Inf. Process. Lett. | 1 |
| 2013 | A Linear-Time Algorithm for Computing the Prime Decomposition of a Directed Graph with Regard to the Cartesian Product
Christophe Crespelle, Eric Thierry, Thomas Lambert |
COCOON | 3 |
| 2011 | Tracking Animal Location and Activity with an Automated Radio Telemetry System in a Tropical RainforestabstractHow do animals use their habitat? Where do they go and what do they do? These basic questions are key not only to understanding a species’ ecology and evolution, but also for addressing many of the environmental challenges we currently face, including problems posed by invasive species, the spread of zoonotic diseases and declines in wildlife populations due to anthropogenic climate and land-use changes. Monitoring the movements and activities of wild animals can be difficult, especially when the species in question are small, cryptic or move over large areas. In this paper, we describe an Automated Radio-Telemetry System (ARTS) that we designed and built on Barro Colorado Island (BCI), Panama to overcome these challenges. We describe the hardware and software we used to implement the ARTS, and discuss the scientific successes we have had using the system, as well as the logistical challenges we faced in maintaining the system in real-world, rainforest conditions. The ARTS uses automated radio-telemetry receivers mounted on 40-m towers topped with arrays of directional antennas to track the activity and location of radio-collared study animals, 24 h a day, 7 days a week. These receiving units are connected by a wireless network to a server housed in the laboratory on BCI, making these data available in real time to researchers via a web-accessible database. As long as study animals are within the range of the towers, the ARTS system collects data more frequently than typical animal-borne global positioning system collars (∼12 locations/h) with lower accuracy (approximately 50 m) but at much reduced cost per tag (∼10X less expensive). The geographic range of ARTS, like all VHF telemetry, is affected by the size of the radio-tag as well as its position in the forest (e.g. tags in the canopy transmit farther than those on the forest floor). We present a model of signal propagation based on landscape conditions, which quantifies these effects and identifies sources of interference, including weather events and human activity. ARTS has been used to track 374 individual animals from 38 species, including 17 mammal species, 12 birds, 7 reptiles or amphibians, as well as two species of plant seeds. These data elucidate the spatio-temporal dynamics of animal activity and movement at the site and have produced numerous peer-reviewed publications, student theses, magazine articles, educational programs and film documentaries. These data are also relevant to long-term population monitoring and conservation plans. Both the successes and the failures of the ARTS system are applicable to broader sensor network applications and are valuable for advancing sensor network research. Roland Kays, Sameer Tilak, Margaret Crofoot, Tony Fountain, Daniel Obando, Alejandro Ortega, Franz Kuemmeth, Jamie Mandel, George Swenson, Thomas Lambert, Ben Hirsch, Martin Wikelski |
Comput. J. | 10 |