VLDB 2026 Research / reviewers in the wild / expert
Lin Ma 0007
dblp:74/3608-7
· DBLP profile ↗
7ranked-venue papers
7as first author
0since 2021 · last 2018
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 50% Parallel and multicore computing · 50% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 2 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation › processor performance modeling › accelerator performance modeling
GPU performance modeling |
0.2 | 1 | 2014 | Theoretical analysis of classic algorithms on highly-threaded many-core GPUs · PPoPP 2014 |
Parallel and multicore computing › parallel algorithms › shared-memory parallel algorithms
many-core algorithms |
0.2 | 1 | 2014 | Theoretical analysis of classic algorithms on highly-threaded many-core GPUs · PPoPP 2014 |
Methods — techniques the papers use, named apart from their topics
threaded many-core memory model · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Analysis of classic algorithms on highly-threaded many-core architectures
Lin Ma 0007, Roger D. Chamberlain, Kunal Agrawal 0001, Chen Tian 0002, Ziang Hu |
Future Gener. Comput. Syst. | 1 |
| 2014 | Performance modeling for highly-threaded many-core GPUsabstractHighly-threaded many-core GPUs can provide high throughput for a wide range of algorithms and applications. Such machines hide memory latencies via the use of a large number of threads and large memory bandwidth. The achieved performance, therefore, depends on the parallelism exploited by the algorithm, the effectiveness of latency hiding, and the utilization of multiprocessors (occupancy). In this paper, we extend previously proposed analytical models, jointly addressing parallelism, latency-hiding, and occupancy. In particular, the model not only helps to explore and reduce the configuration space for tuning kernel execution on GPUs, but also reflects performance bottlenecks and predicts how the runtime will trend as the problem and other parameters scale. The model is validated with empirical experiments. In addition, the model points to at least one circumstance in which the occupancy decisions automatically made by the scheduler are clearly sub-optimal in terms of runtime. Lin Ma 0007, Roger D. Chamberlain, Kunal Agrawal 0001 |
ASAP | 1 |
| 2014 | Theoretical analysis of classic algorithms on highly-threaded many-core GPUsabstractThe Threaded many-core memory (TMM) model provides a framework to analyze the performance of algorithms on GPUs. Here, we investigate the effectiveness of the TMM model by analyzing algorithms for 3 classic problems -- suffix tree/array for string matching, fast Fourier transform, and merge sort -- under this model. Our findings indicate that the TMM model can explain and predict previously unexplained trends and artifacts in experimental data. Lin Ma 0007, Kunal Agrawal 0001, Roger D. Chamberlain |
PPoPP | 1 |
| 2014 | A memory access model for highly-threaded many-core architecturesabstractA number of highly-threaded, many-core architectures hide memory-access latency by low-overhead context switching among a large number of threads. The speedup of a program on these machines depends on how well the latency is hidden. If the number of threads were infinite, theoretically, these machines could provide the performance predicted by the PRAM analysis of these programs. However, the number of threads per processor is not infinite, and is constrained by both hardware and algorithmic limits. In this paper, we introduce the Threaded Many-core Memory (TMM) model which is meant to capture the important characteristics of these highly-threaded, many-core machines. Since we model some important machine parameters of these machines, we expect analysis under this model to provide a more fine-grained and accurate performance prediction than the PRAM analysis. We analyze 4 algorithms for the classic all pairs shortest paths problem under this model. We find that even when two algorithms have the same PRAM performance, our model predicts different performance for some settings of machine parameters. For example, for dense graphs, the dynamic programming algorithm and Johnson’s algorithm have the same performance in the PRAM model. However, our model predicts different performance for large enough memory-access latency and validates the intuition that the dynamic programming algorithm performs better on these machines. We validate several predictions made by our model using empirical measurements on an instantiation of a highly-threaded, many-core machine, namely the NVIDIA GTX 480. Lin Ma 0007, Kunal Agrawal 0001, Roger D. Chamberlain |
Future Gener. Comput. Syst. | 1 |
| 2012 | A Performance Model for Memory Bandwidth Constrained Applications on Graphics EnginesabstractGraphics engines are excellent execution platforms for high-throughput computations that exploit a large degree of available parallelism. The achieved performance is, however, highly dependent on the access patterns that the applicationimposes on the memory subsystem. Here, we propose an analytic model that helps improve the understanding of the performance of memory-limited kernels that employ randommemory access schemes, especially as impacted by cache andvarious configuration parameters that can be used to tunekernel execution, such as the number of blocks and the number of threads per block. The analytic model is first explored through the use of a synthetic micro-benchmark, which is then followed by an empirical validation using a pair of production applications used in computational biology. Lin Ma 0007, Roger D. Chamberlain |
ASAP | 1 |
| 2012 | A Memory Access Model for Highly-threaded Many-core ArchitecturesabstractMany-core architectures are excellent in hiding memory-access latency by low-overhead context switching among a large number of threads. The speedup of algorithms carried out on these machines depends on how well the latency is hidden. If the number of threads were infinite, then theoretically these machines should provide the performance predicted by the PRAM analysis of the programs. However, the number of allowable threads per processor is not infinite. In this paper, we introduce the Threaded Many-core Memory (TMM) model which is meant to capture the important characteristics of these highly-threaded, many-core machines. Since we model some important machine parameters of these machines, we expect analysis under this model to give more fine-grained performance prediction than the PRAM analysis. We analyze 4 algorithms for the classic all pairs shortest paths problem under this model. We find that even when two algorithms have the same PRAM performance, our model predicts different performance for some settings of machine parameters. For example, for dense graphs, the Floyd-Warshall algorithm and Johnson's algorithms have the same performance in the PRAM model. However, our model predicts different performance for large enough memory-access latency and validates the intuition that the Floyd-Warshall algorithm performs better on these machines. Lin Ma 0007, Kunal Agrawal 0001, Roger D. Chamberlain |
ICPADS | 1 |
| 2011 | Bloom Filter Performance on Graphics EnginesabstractBloom filters are a probabilistic technique for large-scale set membership tests. They exhibit no false negative test results but are susceptible to false positive results. They are well-suited to both large sets and large numbers of membership tests. We implement the Bloom filters present in an accelerated version of BLAST, a genome biosequence alignment application, on NVIDIA GPUs and develop an analytic performance model that helps potential users of Bloom filters to quantify the inherent tradeoffs between throughput and false positive rates. Lin Ma 0007, Roger D. Chamberlain, Jeremy Buhler, Mark A. Franklin |
ICPP | 1 |