Richard Afoakwa

dblp:228/3609 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
3since 2021 · last 2022
0000-0001-8604-3869ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2022 Increasing ising machine capacity with multi-chip architectures
abstract
Nature has inspired a lot of problem solving techniques over the decades. More recently, researchers have increasingly turned to harnessing nature to solve problems directly. Ising machines are a good example and there are numerous research prototypes as well as many design concepts. They can map a family of NP-complete problems and derive competitive solutions at speeds much greater than conventional algorithms and in some cases, at a fraction of the energy cost of a von Neumann computer.
Anshujit Sharma, Richard Afoakwa, Zeljko Ignjatovic, Michael C. Huang 0001
ISCA2
2022 A CMOS Compatible Bistable Resistively-coupled Ising Machine-BRIM
abstract
Ising machines and other nature-based computing platforms have recently become attractive due to their potential of outperforming conventional computers when solving problems that involve a large number of competing alternatives, such as combinatorial optimizations. In this paper, a newly proposed resistively-coupled Ising machine with bistable nodes (BRIM) is designed and simulated in a 45nm CMOS process. The performance of the proposed machine is evaluated based on its capability of solving the Max-cut graph problem. A spin-fix annealing technique is applied to help escape local minima and improve the solution quality. Simulation result shows that this technique effectively increases the probability of finding the Max-cut solution by 50.5% and reduces the solution error to 1.73 on average.
Yiqiao Zhang, Richard Afoakwa, Uday Kumar Reddy Vengalam, Michael C. Huang 0001, Zeljko Ignjatovic
ISCAS2
2021 BRIM: Bistable Resistively-Coupled Ising Machine
abstract
Physical Ising machines rely on nature to guide a dynamical system towards an optimal state which can be read out as a heuristical solution to a combinatorial optimization problem. Such designs that use nature as a computing mechanism can lead to higher performance and/or lower operation costs. Quantum annealers are a prominent example of such efforts. However, existing Ising machines are generally bulky and energy intensive. Such disadvantages may be acceptable if these designs provide some significant intrinsic advantages at a much larger scale in the future, which remains to be seen. But for now, integrated electronic designs of Ising machines allow more immediate applications. We propose one such design that uses bistable nodes, coupled with programmable and variable strengths. The design is fully CMOS compatible for on-chip applications and demonstrates competitive solution quality and significantly superior execution time and energy.
Richard Afoakwa, Yiqiao Zhang, Uday Kumar Reddy Vengalam, Zeljko Ignjatovic, Michael C. Huang 0001
HPCA1
2019 To Stack or Not To Stack
abstract
3D memory technology, such as Micron's hybrid memory cube (HMC), has re-energized the architectural pursuit of computation very close to, or inside the memory chip. Such a design falls into the broader category of near-data processing (NDP). The motivation for such design is because the current Von Neumann architecture of chip-multiprocessors is thought to make data movement expensive. Current NDP work focuses on the possibility of architecting computation engines, such as accelerators, cores, or graphic processing units right below the memory layers and inside the logic layer of the HMC sub-system. However, such a stacking design does present a number of technical challenges such as heat dissipation, power supply, etc. While these challenges can certainly be overcome, and needs to be addressed, in this work, we seek to answer a related question of whether it is necessary to stack general-purpose computation engines, directly inside the memory unit, in order to achieve the performance potential of NDP system; thus, to stack or not to stack. We show that, with computing models used in current NDP designs, placing the computation engines very close to, but outside the memory system (not stacking) can provide comparable performance without significant energy costs. This can be achieved without inventing any new technology, but utilizing current state-of-the-art high-speed link design practices.
Richard Afoakwa, Lejie Lu, Hui Wu 0007, Michael C. Huang 0001
PACT1
2019 Concurrent Multipoint-to-Multipoint Communication on Interposer Channels
abstract
Chip-to-chip communication for next generation computing will require larger bandwidth density to support ever increasing data traffic between processors, memories and I/O. 3-D integration enables a large number of processor and memory chips to be densely packed on an interposer with fine-pitch interconnect lanes. Advanced signaling techniques such as pulse amplitude modulation (PAM) can be employed to improve bandwidth per lane. Most recent work on interposer-based chip-to-chip interconnects focus primarily on point-to-point serial links. Without adding costly routers, these designs will severely limit the overall system level concurrency. In this paper, we propose an ultrahigh-speed multipoint-to-multipoint link design for interposer channels, which supports PAM signaling. Each node on the link can send, receive, drop, or relay data at line rate without complex routing. This design enables splitting the physical link into segments, and allows multicast/broadcast. A proof-of-concept system prototype with up to 16 nodes integrated on a silicon interposer with up to 22-mm node spacing is designed and evaluated using circuit and system simulations. The PAM-4 transceiver and link interface circuits at each node are implemented using a standard 130-nm SiGe BiCMOS technology. The transceiver can achieve a data rate of 40-Gb/s/lane, with channel loss of -3.5 dB per segment at Nyquist frequency, and energy efficiency between 1.29-pJ/b between two neighboring nodes or 0.21-pJ/b more per additional nodes. Using a cycle-level system simulation, such a high-concurrency communication fabric can improve overall performance between 2% to 18% over baseline.
Lejie Lu, Richard Afoakwa, Michael C. Huang 0001, Hui Wu 0007
ISLPED2
2018 High Swing Pulse-Amplitude Modulation of Transmission Line Links for On-Chip Communication
abstract
With ever increasing core count of chip-multiprocessors (CMPs), the network-on-chip (NoC) fabric continues to be an important component for performance and energy. We propose the use of high voltage swing serial links as the backbone NoC. We designed transmitter drivers to deliver a high output swing and enable high speed Pulse-Amplitude Modulation (PAM-4 and PAM-8) transmissions. We show that with careful circuit-level transceiver design, coupled with system level architectural utilization of such links, it is possible to drive up to 8 cm of on-chip transmission line at diverse adaptive modulations. Using such a design, experimental analysis shows an average of 1.4× performance improvement over baseline. The overall energy-delay product improvement is 1.75×.
Richard Afoakwa, Lejie Lu, Yong Wang 0026, Hui Wu 0007, Michael C. Huang 0001
ISCAS1