EDBT 2026 Demo / reviewers in the wild / expert
Thomas Kistler
dblp:98/1048
· DBLP profile ↗
6ranked-venue papers
4as first author
1since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 2 first-authorSystems, architecture and hardware · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 67% Recommender systems · 33% | |
| Artificial intelligence
1 paper |
Language models and text generation · 100% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › retrieval models › neural retrieval › dense retrieval
bi-encoder retrieval |
1.0 | 1 | 2026 | Large Scale Retrieval for the LinkedIn Feed Using Causal Language Models · AAAI 2026 |
Information retrieval › retrieval models › neural retrieval
embedding-based retrieval |
1.0 | 1 | 2026 | Large Scale Retrieval for the LinkedIn Feed Using Causal Language Models · AAAI 2026 |
Natural language and speech › Language models and text generation › pre-trained language model
causal language model |
0.3 | 1 | 2026 | Large Scale Retrieval for the LinkedIn Feed Using Causal Language Models · AAAI 2026 |
Compilers and program optimization
dynamic optimization |
0.1 | 2 | 2003 | Continuous program optimization: A case study · ACM Trans. Program. Lang. Syst. 2003 Continuous Program Optimization: Design and Evaluation · IEEE Trans. Computers 2001 |
Memory systems
cache |
0.1 | 2 | 2003 | Continuous program optimization: A case study · ACM Trans. Program. Lang. Syst. 2003 Automated data-member layout of help objects to improve memory-hierarchy performance · ACM Trans. Program. Lang. Syst. 2000 |
Compilers and program optimization › dynamic optimization
profile-guided optimization |
0.0 | 1 | 2001 | Continuous Program Optimization: Design and Evaluation · IEEE Trans. Computers 2001 |
Compilers and program optimization › memory optimization
data layout optimization |
0.0 | 1 | 2000 | Automated data-member layout of help objects to improve memory-hierarchy performance · ACM Trans. Program. Lang. Syst. 2000 |
Memory systems › memory hierarchy
memory hierarchy performance |
0.0 | 1 | 2000 | Automated data-member layout of help objects to improve memory-hierarchy performance · ACM Trans. Program. Lang. Syst. 2000 |
Methods — techniques the papers use, named apart from their topics
quantization · 2.0fine-tuning · 2.0dual encoder · 2.0live profiling · 0.1instruction rescheduling · 0.1dynamic profiling · 0.1object layout adaptation · 0.0dynamic trace scheduling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Large Scale Retrieval for the LinkedIn Feed Using Causal Language ModelsabstractIn large-scale recommendation systems like LinkedIn’s, the retrieval stage is critical for narrowing billions of potential candidates to a manageable subset for ranking. LinkedIn's feed now serves suggested content based on the topical interests of members, where 2000 candidates are retrieved from several million candidates with a latency budget of a few milliseconds and inbound QPS of several thousand per second. This paper presents a novel retrieval approach that fine tunes a large causal language model (Meta’s LLaMA 3) as a dual encoder to generate high quality embeddings for both users (members) and content (items), using only textual input. We describe the end to end pipeline, including prompt design for embedding generation, techniques for fine tuning at LinkedIn scale, and infrastructure for low latency, cost effective online serving. We share our findings on how quantizing numerical features in the prompt enables the information getting encoded in the embedding facilitating greater alignment between the retrieval and ranking layer. The system was evaluated using offline metrics and an online A/B test, which showed substantial improvements in member engagement. We observed significant gains among newer members, who often lack strong network connections, indicating that high-quality suggested content aids retention. This work demonstrates how generative language models can be effectively adapted for real time, high throughput retrieval in industrial applications. Sudarshan Srinivasa Ramanujam, Antonio Alonso, Saurabh Kataria 0003, Siddharth Dangi, Akhilesh Gupta, Birjodh Singh Tiwana, Manas Haribhai Somaiya, Luke Simon, David Byrne, Sojeong Ha, Sen Zhou, Andrei Akterskii, Zhanglong Liu, Samira Sriram, Zihan Xiong, Zhoutao Pei, Angela Shao, Alex Li, Annie Xiao, Caitlin Kolb, Thomas Kistler, Zach Moore, Hamed Firooz |
AAAI | 21 |
| 2003 | The Transmeta Code Morphing - Software: Using Speculation, Recovery, and Adaptive Retranslation to Address Real-Life ChallengesabstractTransmeta's Crusoe microprocessor is a full, system-level implementation of the x86 architecture, comprising a native VLIW microprocessor with a software layer, the Code Morphing Software (CMS), that combines an interpreter, dynamic binary translator, optimizer, and run-time system. In its general structure, CMS resembles other binary translation systems described in the literature, but it is unique in several respects. The wide range of PC workloads that CMS must handle gracefully in real-life operation, plus the need for full system-level x86 compatibility, expose several issues that have received little or no attention in previous literature, such as exceptions and interrupts, I/O, DMA, and self-modifying code. In this paper we discuss some of the challenges raised by these issues, and present the techniques developed in Crusoe and CMS to meet those challenges. The key to these solutions is the Crusoe paradigm of aggressive speculation, recovery to a consistent x86 state using unique hardware commit-and-rollback support, and adaptive retranslation when exceptions occur too often to be handled efficiently by interpretation. James C. Dehnert, Brian Grant, John P. Banning, Thomas Kistler, Alexander Klaiber, Jim Mattson |
CGO | 5 |
| 2003 | Continuous program optimization: A case studyabstractMuch of the software in everyday operation is not making optimal use of the hardware on which it actually runs. Among the reasons for this discrepancy are hardware/software mismatches, modularization overheads introduced by software engineering considerations, and the inability of systems to adapt to users' behaviors.A solution to these problems is to delay code generation until load time. This is the earliest point at which a piece of software can be fine-tuned to the actual capabilities of the hardware on which it is about to be executed, and also the earliest point at wich modularization overheads can be overcome by global optimization.A still better match between software and hardware can be achieved by replacing the already executing software at regular intervals by new versions constructed on-the-fly using a background code re-optimizer. This not only enables the use of live profiling data to guide optimization decisions, but also facilitates adaptation to changing usage patterns and the late addition of dynamic link libraries.This paper presents a system that provides code generation at load-time and continuous program optimization at run-time. First, the architecture of the system is presented. Then, two optimization techniques are discussed that were developed specifically in the context of continuous optimization. The first of these optimizations continually adjusts the storage layouts of dynamic data structures to maximize data cache locality, while the second performs profile-driven instruction re-scheduling to increase instruction-level parallelism. These two optimizations have very different cost/benefit ratios, presented in a series of benchmarks. The paper concludes with an outlook to future research directions and an enumeration of some remaining research problems.The empirical results presented in this paper make a case in favor of continuous optimization, but indicate that it needs to be applied judiciously. In many situations, the costs of dynamic optimizations outweigh their benefit, so that no break-even point is ever reached. In favorable circumstances, on the other hand, speed-ups of over 120% have been observed. It appears as if the main beneficiaries of continuous optimization are shared libraries, which at different times can be optimized in the context of the currently dominant client application. Thomas Kistler, Michael Franz |
ACM Trans. Program. Lang. Syst. | 1 |
| 2001 | Continuous Program Optimization: Design and EvaluationabstractThis paper presents a system in which the already executing user code is continually and automatically reoptimized in the background, using dynamically collected execution profiles as a guide. Whenever a new code image has been constructed in the background in this manner, it is hot-swapped in place of the previously executing one. Control is then transferred to the new code and construction of yet another code image is initiated in the background. Two new runtime optimization techniques have been implemented in the context of this system: object layout adaptation and dynamic trace scheduling. The former technique constantly improves the storage layout of dynamically allocated data structures to improve data cache locality. The latter increases the instruction-level parallelism by continually adapting the instruction schedule to predominantly executed program paths. The empirical results presented in this paper make a case in favor of continuous optimization, but also indicate some of the pitfalls and current shortcomings of continuous optimization. If not applied judiciously, the costs of dynamic optimizations outweigh their benefit in many situations so that no break-even point is ever reached. In favorable circumstances, however, speed-ups of over 96 percent have been observed. It appears as if the main beneficiaries of continuous optimization are shared libraries in specific application domains which, at different times, can be optimized in the context of the currently dominant client application. Thomas Kistler, Michael Franz |
IEEE Trans. Computers | 1 |
| 2000 | Automated data-member layout of help objects to improve memory-hierarchy performanceabstractWe present and evaluate a simple, yet efficient optimization technique that improves memory-hierarchy performance for pointer-centric applications by up to 24% and reduces cache misses by up to 35%. This is achieved by selecting an improved ordering for the data members of pointer-based data structures. Our optimization is applicable to all type-safe programming languages that completely abstract from physical storage layout; examples of such languages are Java and Oberon. Our technique does not involve programmers in the optimization process, but runs fully automatically, guided by dynamic profiling information that captures which paths through the program are taken with that frequencey. The algorithm first strives to cluster data members that are accessed closely after one another onto the same cache line, increasing spatial locality. Then, the data members that have been mapped to a particular cache line are ordered to minimize load latency in case of a cache miss. Thomas Kistler, Michael Franz |
ACM Trans. Program. Lang. Syst. | 1 |
| 1998 | WebL - A Programming Language for the Web
Thomas Kistler, Hannes Marais |
Comput. Networks | 1 |