Xingwu Dong

dblp:383/3644 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0004-9707-1119ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 61% Emerging computing paradigms · 30% Hardware accelerators and domain-specific architectures · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
non-volatile memory
1.012026
A Transverse-Read-Assisted Fast Valid-Bits Collection in Stochastic Computing MACs for Energy-Efficient in-RTM DNNs · IEEE Trans. Computers 2026
Memory systems › emerging memory technologies › spintronic memory
racetrack memory
1.012026
A Transverse-Read-Assisted Fast Valid-Bits Collection in Stochastic Computing MACs for Energy-Efficient in-RTM DNNs · IEEE Trans. Computers 2026
Emerging computing paradigms › approximate and stochastic computing
stochastic computing
1.012026
A Transverse-Read-Assisted Fast Valid-Bits Collection in Stochastic Computing MACs for Energy-Efficient in-RTM DNNs · IEEE Trans. Computers 2026
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.312026
A Transverse-Read-Assisted Fast Valid-Bits Collection in Stochastic Computing MACs for Energy-Efficient in-RTM DNNs · IEEE Trans. Computers 2026

Methods — techniques the papers use, named apart from their topics

transverse-read · 1.0pseudo-fractal compression · 1.0interleaving data placement · 1.0asynchronous scheduling · 1.0
YearPublicationVenuePosition
2026 A Transverse-Read-Assisted Fast Valid-Bits Collection in Stochastic Computing MACs for Energy-Efficient in-RTM DNNs
abstract
It looks very attractive to coordinate racetrack-memory (RM) and stochastic-computing (SC) jointly to build an ultra-low power neuron-architecture. However, the above combination has always been questioned in a fatal weakness that the heavy valid-bits collection of Racetrack Memory-Magnetic Tunnel Junctions(RM-MTJ), a.k.a. accumulative parallel counters (APCs), cannot physically match the requirement for energy-efficient in-memory DNNs. Fortunately, a recently developed Transverse-Read (TR) provides a lightweight collection of valid-bits by detecting domain-wall resistance between a couple of MTJs on a single nanowire. In this work, we first propose a neuron-architecture that utilizes parallel TRs to build an ultra-fast valid-bits collection scheme specifically targeted at Multiply-Accumulate (MAC) units for in-RTM DNNs, where the multiplication operation inherently involves two operands. To solve the huge storage for full stochastic sequences caused by the limited TR banks, a hybrid coding, pseudo-fractal compression, is designed to generate stochastic sequences by segments. To overcome the misalignment by the parallel early-termination, an asynchronous schedule of TR is further designed to regularize the vectorization, in which the valid-bits from different lanes are merged in multiple RM-stacks for vector-level valid-bits collection. However, an inherent defect of TR, i.e., neighbor parts cannot be accessed simultaneously, could limit the throughput of the parallel vector multiplication, therefore, an interleaving data placement is used for full utilization of the memory bus among different vectors. The experimental results demonstrate that the SC-MAC architecture with TR achieves 2.88×-4.40× speedup over CORUSCANT (the state-of-the-art processin-RM architecture), while simultaneously reducing energy consumption by 1.26×-1.42×.
Xingwu Dong, Danghui Wang
IEEE Trans. Computers3