Torbjörn Granlund

dblp:79/2489 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 first-authorSystems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
1 paper
Algorithms and data structures · 77% Mathematical optimization · 23%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Processor architecture and microarchitecture · 100%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
instruction set architecture
0.112011
Improved Division by Invariant Integers · IEEE Trans. Computers 2011
Algorithms and data structures › number-theoretic algorithms
integer division
0.112011
Improved Division by Invariant Integers · IEEE Trans. Computers 2011
Compilers and program optimization › loop optimization
strength reduction
0.011994
Division by Invariant Integers using Multiplication · PLDI 1994
Compilers and program optimization › compiler optimization › branch optimization
branch elimination
0.011992
Eliminating Branches using a Superoptimizer and the GNU C Compiler · PLDI 1992
Compilers and program optimization › compiler optimization
superoptimization
0.011992
Eliminating Branches using a Superoptimizer and the GNU C Compiler · PLDI 1992

Methods — techniques the papers use, named apart from their topics

umullo · 0.2umulhi · 0.2umul · 0.2two's complement arithmetic · 0.0integer multiplication · 0.0superoptimizer · 0.0GNU C Compiler · 0.0
YearPublicationVenuePosition
2011 Improved Division by Invariant Integers
abstract
This paper considers the problem of dividing a two-word integer by a single-word integer, together with a few extensions and applications. Due to lack of efficient division instructions in current processors, the division is performed as a multiplication using a precomputed single-word approximation of the reciprocal of the divisor, followed by a couple of adjustment steps. There are three common types of unsigned multiplication instructions: we define full word multiplication (umul), which produces the two-word product of two single-word integers; low multiplication (umullo), which produces only the least significant word of the product; and high multiplication (umulhi), which produces only the most significant word. We describe an algorithm that produces a quotient and remainder using one umul and one umullo. This is an improvement over earlier methods, since the new method uses cheaper multiplication operations. It turns out that we also get some additional savings from simpler adjustment conditions. The algorithm has been implemented in version 4.3 of the gmp library. When applied to the problem of dividing a large integer by a single word, the new algorithm gives a speedup of roughly 30 percent, benchmarked on AMD and Intel processors in the x86_64 family.
Niels Moller, Torbjörn Granlund
IEEE Trans. Computers2
1994 Division by Invariant Integers using Multiplication
abstract
Integer division remains expensive on today's processors as the cost of integer multiplication declines. We present code sequences for division by arbitrary nonzero integer constants and run-time invariants using integer multiplication. The algorithms assume a two's complement architecture. Most also require that the upper half of an integer product be quickly accessible. We treat unsigned division, signed division where the quotient rounds towards zero, signed division where the quotient rounds towards -∞, and division where the result is known a priori to be exact. We give some implementation results using the C compiler GCC.
Torbjörn Granlund, Peter L. Montgomery
PLDI1
1992 Eliminating Branches using a Superoptimizer and the GNU C Compiler
abstract
this paper uses the RS/6000 for all its examples, the techniques described here are applicable to most machines
Torbjörn Granlund, Richard Kenner
PLDI1