VLDB 2026 Research / reviewers in the wild / expert
Donald Kline Jr.
dblp:162/9969
· DBLP profile ↗
13ranked-venue papers
5as first author
2since 2021 · last 2021
0000-0002-4414-1513ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 5 first-author · 2 since 2021Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Memory systems · 68% Hardware reliability and fault tolerance · 25% Interconnection networks and networks-on-chip · 6% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
non-volatile memory |
0.8 | 2 | 2021 | A CASTLE With TOWERs for Reliable, Secure Phase-Change Memory · IEEE Trans. Computers 2021 Improving Bit Flip Reduction for Biased and Random Data · IEEE Trans. Computers 2016 |
Memory systems › non-volatile memory
phase change memory |
0.6 | 2 | 2021 | A CASTLE With TOWERs for Reliable, Secure Phase-Change Memory · IEEE Trans. Computers 2021 FLOWER and FaME: A Low Overhead Bit-Level Fault-map and Fault-Tolerance Approach for Deeply Scaled Memories · HPCA 2020 |
Hardware reliability and fault tolerance
error correction |
0.5 | 1 | 2021 | A CASTLE With TOWERs for Reliable, Secure Phase-Change Memory · IEEE Trans. Computers 2021 |
Memory systems
memory encryption |
0.5 | 1 | 2021 | A CASTLE With TOWERs for Reliable, Secure Phase-Change Memory · IEEE Trans. Computers 2021 |
Hardware reliability and fault tolerance
memory fault tolerance |
0.4 | 1 | 2020 | FLOWER and FaME: A Low Overhead Bit-Level Fault-map and Fault-Tolerance Approach for Deeply Scaled Memories · HPCA 2020 |
Memory systems › non-volatile memory
bit-flip reduction |
0.2 | 1 | 2016 | Improving Bit Flip Reduction for Biased and Random Data · IEEE Trans. Computers 2016 |
Interconnection networks and networks-on-chip › switch architecture
buffer design |
0.2 | 1 | 2015 | Domain-wall memory buffer for low-energy NoCs · DAC 2015 |
Memory systems
emerging memory technologies |
0.2 | 1 | 2015 | Domain-wall memory buffer for low-energy NoCs · DAC 2015 |
Memory systems › emerging memory technologies › spintronic memory
racetrack memory |
0.2 | 1 | 2015 | Domain-wall memory buffer for low-energy NoCs · DAC 2015 |
Energy-efficient computing › low-power design
low-power interconnect |
0.1 | 1 | 2015 | Domain-wall memory buffer for low-energy NoCs · DAC 2015 |
Methods — techniques the papers use, named apart from their topics
error-correction pointers · 0.5encoding · 0.5compression · 0.5bloom filter · 0.4MinCI hashing · 0.4differential writing · 0.2coset encoding · 0.2spintronic domain-wall memory · 0.2shift-register control scheme · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Tuning Memory Fault Tolerance on the EdgeabstractError correction and fault tolerance have become pivotal considerations as conventional memories scale and emerging memories come to market. The common thread in these reliability challenges is that deep scaling reveals outliers in the memory system, which are responsible for the vast majority of faults. These cells, which may be attributed to process variation or undetected fabrication defects, tend to be more vulnerable to various forms of crosstalk, read- and write-disturbance, and even radiation-induced faults. By tracking faults in memory cells, identifying the worst offenders, and mitigating their effects accordingly, we can design dramatically improved fault tolerance techniques that are tuned to the fault characteristics of the memory at hand. A critical piece is the development of scalable and fault tolerance registries to track and retain critical information about these faults. The fault registries must be able to function in the faulty memory they protect, operate efficiently at the cell/bit-level, and handle extreme fault rates. Using the knowledge of faulty locations, our fault tolerance techniques applied to conventional main memories like DRAM and endurance-limited memories like flash and phase-change memory improve reliability, endurance, and lifetime by orders of magnitude while maintaining performance and energy efficiency. Alex K. Jones, Stephen Longofono, Sébastien Ollivier, Donald Kline Jr., Jiangwei Zhang, Rami G. Melhem |
ACM Great Lakes Symposium on VLSI | 4 |
| 2021 | A CASTLE With TOWERs for Reliable, Secure Phase-Change MemoryabstractThe use of hardware encryption and new memory technologies such as phase change memory (PCM) are gaining popularity in a variety of server applications such as cloud systems. While PCM provides energy and density advantages over conventional DRAM memory, it faces endurance challenges. Such challenges are exacerbated when employing memory encryption as the stored data is essentially randomized, losing data similarity and reducing or eliminating the effectiveness of energy and endurance oriented encoding techniques. This results in increasing dynamic energy consumption and accelerated wear out. In this article, we propose CASTLE, a technique for in-memory encryption to leverage this encryption process to improve reliability in the presence of endurance faults. We also propose TOWERs for CASTLE that improve reliability as well as energy for encrypted data through a novel application of compression and encoding. CASTLE and TOWERs are compatible with error-correction codes (ECC) and error correction pointers (ECP), the standard for mitigating endurance faults in PCM. When combining CASTLE and TOWERS, we achieve an average lifetime improvement of over 45× compared to SECDED ECC, 7.1× compared to SECRET, and 3.6× compared to the leading partition-and-flip fault-tolerance approach (AEGIS) for the same area overhead. Stephen Longofono, Donald Kline Jr., Rami G. Melhem, Alex K. Jones |
IEEE Trans. Computers | 2 |
| 2020 | FLOWER and FaME: A Low Overhead Bit-Level Fault-map and Fault-Tolerance Approach for Deeply Scaled MemoriesabstractTo maintain appropriate yields in deeply scaled technologies requires fault-tolerance of increasingly high fault rates. These fault rates far exceed traditional general approaches such as ECC, particularly when faults accrue over time. Effective fault tolerance at such high fault rates requires detailed bit-level knowledge of the location of faulty cells. We provide a solution to this problem in the form of a space efficient, bit-level fault map called FLOWER. FLOWER utilizes Bloom filters to provide detailed fault characterization for a relatively small overhead. We demonstrate how FLOWER can enable improved fault tolerance at high fault rates by enhancing existing fault tolerance proposals and yielding 10–100x improvements. Using in-memory processing, FLOWER can maintain a less than 2% performance overhead at 10E-4 fault rates with less than 2% loss of memory density to report bit-level faults with high accuracy. Using a tuned novel hashing technique called MinCI, FLOWER for memory achieves considerably lower false positives than with disk-level hashing techniques at a fraction of the performance overhead. With a new technique to protect against errors during in-memory operations, PETAL bits, FLOWER can remain resilient against random errors while efficiently targeting predictable errors. Furthermore, we propose a new fault tolerance scheme called FaME, which provides ultra-efficient bit-level sparing by using the FLOWER fault map to identify the location of faults. FLOWER+FaME can achieve 14x longer PCM memory lifetime with half the area overhead versus SECDED ECC. Donald Kline Jr., Jiangwei Zhang, Rami G. Melhem, Alex K. Jones |
HPCA | 1 |
| 2019 | Leveraging Transverse Reads to Correct Alignment Faults in Domain Wall MemoriesabstractSpintronic domain wall memories (DWMs) are prone to alignment faults, which cannot be protected by traditional error correction techniques. To solve this problem, we propose a new technique called derived error correction coding (DECC). We construct metadata from the data and shift state of the DWM, on demand, using a novel transverse read (TR). TR reads in an orthogonal direction to the DWM access point and can determine the number of ones in a DWM. Errors in the metadata correspond to shift-faults in the DWM. Rather than storing the metadata, it is created on-demand and protected by storing parity bits. Repairing the metadata with ECC allows restoration of DWM alignment and ensures correct operation. Through these techniques, our shift-aware error correction approaches provide a lifetime of over 15 years with a similar performance, while reducing area and energy by 370% and 52%, versus the state-of-the-art, for a 32-bit nanowire. Sébastien Ollivier, Donald Kline Jr., Kawsher A. Roxy, Rami G. Melhem, Sanjukta Bhanja, Alex K. Jones |
DSN | 2 |
| 2019 | PREMSim: A Resilience Framework for Modeling Traditional and Emerging Memory ReliabilityabstractScaling limitations of conventional and emerging memories has provided the impetus for the increased focus on reliability techniques to overcome associated physical limitations of non-perfect devices. However, despite these reliability advances, critical challenges remain to be solved as new memory types and memory vulnerabilities arise. There continues to be no simulator with extended reliability models for easy comparison of existing and newly developed techniques nor simple integration of innovative new reliability concepts and failure modes. The mission of our simulator, PremSim, is to provide a framework which solves these fundamental limitations. While PremSim can function using memory traces, it was also designed from the ground-up to be fully integrated with several external simulators including the Structural Simulation Toolkit (SST) as a memory backend. It can connect to other detailed memory backends such as DRAMSim2 for detailed energy and timing. Further, it provides modes which give estimated lifetime for endurance-limited memories, as well as the provable correction capability per row for a given fault distribution. To perform these calculations in a reasonable time window and to remain compatible with abstract full-system simulators, we also provide and verify novel abstractions. Additionally, we show case studies of how different fault mitigation strategies can be modeled effectively in PremSim including solutions at the page, row, word, and bit-level granularity of next-generation traditional and emerging faulty memories. Donald Kline Jr., Stephen Longofono, Sébastien Ollivier, Erin Higgins, Rami G. Melhem, Alex K. Jones |
MASCOTS | 1 |
| 2019 | Yielding optimized dependability assurance through bit inversion
Jiangwei Zhang, Donald Kline Jr., Rami G. Melhem, Alex K. Jones |
Integr. | 2 |
| 2018 | Racetrack Queues for Extremely Low-Energy FIFOs
Donald Kline Jr., Rami G. Melhem, Alex K. Jones |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Data Block Partitioning Methods to Mitigate Stuck-At Faults in Limited Endurance Memories
Jiangwei Zhang, Donald Kline Jr., Rami G. Melhem, Alex K. Jones |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Dynamic partitioning to mitigate stuck-at faults in emerging memoriesabstractEmerging non-volatile memories have many advantages over conventional memory. Unfortunately, many are susceptible to write endurance challenges, resulting in stuck-at faults. Existing mitigation methods statically partition and invert data within a block containing such faults (partition-and-flip) to ensure data is written to match stuck-at cells such that they may remain in service. Unfortunately, these schemes have limited fault tolerance capabilities and require the assumption that their auxiliary bits are fault free. We propose a dynamic partitioning scheme that improves the number of tolerated stuck-at faults and simultaneously protects auxiliary bits. Dynamic partitioning can significantly improve the fault tolerance over existing static partitioning approaches with an equal number of auxiliary bits. Moreover, it can often still improve fault tolerance while reducing the number of auxiliary bits. Compared to flip-N-write and Aegis, a leading mitigation scheme, dynamic partitioning can achieve 7-72% and 5-53 x lower write error rates, respectively, for the same capacity overhead with a stuck-at-fault rate of 10-3. Jiangwei Zhang, Donald Kline Jr., Rami G. Melhem, Alex K. Jones |
ICCAD | 2 |
| 2017 | Yoda: Judge Me by My Size, Do You?abstractPhase change memory is a promising alternative to conventional memories such as DRAM due to its density and non-volatility. Unfortunately, reliability is still a challenge as limited write endurance, exacerbated by process variation, leads to increasing numbers of stuck-at faults over the memory's lifetime. Error-correcting Pointers (ECP) is a popular proposal to mitigate stuck-at faults by recording the addresses and the values of faulty bits in order to extend the memory lifetime. In this paper, we propose Yoda, a method to extend ECP with one or a small number of additional encoding bits in order to dramatically improve the effectiveness and guaranteed fault correction capability of ECP. Our simulation results demonstrate that Yoda has a 3.0× improvement in fault coverage compared to a fault-aware ECP with a similar overhead, while also providing a 2.5-3.0× improvement over state-of-the-art schemes with comparable complexity. Jiangwei Zhang, Donald Kline Jr., Rami G. Melhem, Alex K. Jones |
ICCD | 2 |
| 2016 | Improving Bit Flip Reduction for Biased and Random DataabstractNonvolatile memory technologies such as Spin-Transfer Torque Random Access Memory (STT-RAM) and Phase Change Memory (PCM) are emerging as promising replacements to DRAM. Before deploying STT-RAM and PCM into functional systems, a number of challenges still remain must be addressed. Specifically, both require relatively high write energy, STT-RAM suffers from high bit error rates and PCM suffers from low endurance. A common solution to overcome those challenges is to minimize the number of bits changed per write. In this paper, we propose and evaluate the hybrid coset encoder to efficiently improve and balance the bit flip reduction for biased and unbiased data. The main core of the coset encoder consists of biased and unbiased vectors which maps the data input to a larger set of data vectors. Subsequently, the intermediate data vector that yields the least number of differences when compared to the currently stored data is selected. Our evaluation shows that hybrid coset encoder reduces bit flips by up to 25 percent over a baseline differential writing scheme. Further, our proposed scheme reduces bit flips by up to 20 percent over the leading bit-flip minimization scheme for biased data, while achieving very low decoding overhead similar to the Flip-N-Write scheme. Seyed Mohammad Seyedzadeh, Rakan Maddah, Donald Kline Jr., Alex K. Jones, Rami G. Melhem |
IEEE Trans. Computers | 3 |
| 2015 | Domain-wall memory buffer for low-energy NoCsabstractNetworks-on-chip (NoCs) have become a leading energy consumer in modern multi-core processors, with a considerable portion of this energy originating from the large number of virtual channel (FIFO) buffers. While emerging memories have been considered for many architectural components such as caches, the asymmetric access properties and relatively small size of network-FIFOs compared to the required peripheral circuitry has led to few such replacements proposed for NoCs. In this paper, we propose control schemes that leverage the\shift-register" nature of spintronic domain-wall memory (DWM) to replace conventional memory buffers for the NoC. Our results indicate that the best shift-based scheme utilizes a dual-nanowire approach to ensure that reads and writes can be more effectively aligned with access ports for simultaneous access in the same cycle. Our approach provides a 2.93X speedup over a DWM buffer using a traditional FIFO memory control scheme with a 1.16X savings in energy. Compared to a SRAM-FIFO it exhibits an 8% message latency degradation versus a 56% energy reduction. The resulting approach achieves a 53% reduction in energy delay product compared to SRAM and a 42% reduction in energy delay product versus STT-MRAM. Donald Kline Jr., Rami G. Melhem, Alex K. Jones |
DAC | 1 |
| 2015 | MSCS: Multi-hop Segmented Circuit SwitchingabstractNoCs (networks-on-chip) are commonly proposed as scalable on-chip interconnects for current and future CMPs (chip multi-processors) and many-core systems. While scalable, the lack of global control can create routing inefficiencies detrimental to the overall network latency. Recently, NoCs have been proposed that allow flits to traverse multiple network switches in a single cycle. This requires a more global view of control to allow routers along the path of a packet to configure their switches collectively. In this paper, we propose a reservation based circuit-switching design, MSCS, which provides simplified global control and multi-hop traversal while reducing latency. MSCS performs network control once per network dimension for the lifetime of a packet, while the leading methods require multiple arbitration steps depending on contention in the network. Furthermore, MSCS can perform control for a packet prior to the availability of resources through reservations, while previous schemes only perform control on-demand. Overall, MSCS can reduce the buffer size by 50% over the leading multi-hop scheme while maintaining a nominal latency improvement (1.4%). With the same buffer resources per port, MSCS achieves a 12.7% latency improvement. Donald Kline Jr., Rami G. Melhem, Alex K. Jones |
ACM Great Lakes Symposium on VLSI | 1 |