Ningfeng Yang

dblp:225/9660 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0003-0750-7998ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 27% Embedded and real-time systems · 27% Distributed systems · 23%
Artificial intelligence
2 papers
Efficient and distributed learning · 41% Optimization for machine learning · 41% Motion planning and robot control · 18%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent
0.912025
Improving the Straight-Through Estimator with Zeroth-Order Information · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression › quantization
quantized training
0.912025
Improving the Straight-Through Estimator with Zeroth-Order Information · NeurIPS 2025
Embedded and real-time systems
collision detection
0.712023
Energy-Efficient Realtime Motion Planning · ISCA 2023
Hardware accelerators and domain-specific architectures
robotics accelerator
0.712023
Energy-Efficient Realtime Motion Planning · ISCA 2023
Distributed systems › fault tolerance
byzantine fault tolerance
0.612022
In-ConcReTeS: Interactive Consistency meets Distributed Real-Time Systems, Again! · RTSS 2022
Storage systems
key-value storage
0.612022
In-ConcReTeS: Interactive Consistency meets Distributed Real-Time Systems, Again! · RTSS 2022
Robotics › Motion planning and robot control › motion planning
learning-based motion planning
0.212023
Energy-Efficient Realtime Motion Planning · ISCA 2023
Robotics › Motion planning and robot control
motion planning
0.212023
Energy-Efficient Realtime Motion Planning · ISCA 2023

Methods — techniques the papers use, named apart from their topics

zeroth-order gradient descent · 0.9straight-through estimator · 0.9backpropagation · 0.9replica coordination · 0.6BFT protocol · 0.6
YearPublicationVenuePosition
2025 Improving the Straight-Through Estimator with Zeroth-Order Information
abstract
We study the problem of training neural networks with quantized parameters. Learning low-precision quantized parameters by enabling computation of gradients via the Straight-Through Estimator (STE) can be challenging. While the STE enables back-propagation, which is a first-order method, recent works have explored the use of zeroth-order (ZO) gradient descent for fine-tuning. We note that the STE provides high-quality biased gradients, and ZO gradients are unbiased but can be expensive. We thus propose First-Order-Guided Zeroth-Order Gradient Descent (FOGZO) that reduces STE bias while reducing computations relative to ZO methods. Empirically, we show FOGZO improves the tradeoff between quality and training time in Quantization-Aware Pre-Training. Specifically, versus STE at the same number of iterations, we show a 1-8% accuracy improvement for DeiT Tiny/Small, 1-2% accuracy improvement on ResNet 18/50, and 1-22 perplexity point improvement for LLaMA models with up to 0.3 billion parameters. For the same loss, FOGZO yields a 796$\times$ reduction in computation versus n-SPSA for a 2-layer MLP on MNIST. Code is available at [https://github.com/1733116199/fogzo](https://github.com/1733116199/fogzo).
Ningfeng Yang, Tor M. Aamodt
NeurIPS1
2023 Energy-Efficient Realtime Motion Planning
abstract
Motion planning is a fundamental problem in autonomous robotics with real-time and low-energy requirements for safe navigation through a dynamic environment. More than 90% of computation time in motion planning is spent on collision detection between the robot and the environment. Several motion planning approaches, such as deep learning-based motion planning, have shown significant improvements in motion planning quality and runtime with ample parallelism available in collision detection. However, naive parallelization of collision detection queries significantly increases computation compared to sequential execution. In this work, we investigate the sources of redundant computations in coarsegrained (inter-collision detection) and fine-grained (intracollision detection) parallelism. We find that the physical spatial locality of obstacles results in redundant computation in coarse-grained parallelism. We further show that the primary sources of redundant computation in fine-grained parallelism are easy cases where objects are far apart or significantly overlapping. Based on these insights, we propose MPAccel to improve the energy efficiency of parallelization in motion planning. MPAccel consists of SAS, a Spatially Aware Scheduler for coarse-grained parallelism, and CECDUs, Cascaded Early-exit Collision Detection Units for fine-grained parallelism. SAS results in 7× speedup using 8× parallelization with 6% increase in the computation compared to 3.7× speedup with 83% increase in computation for naive parallelization. CECDU can perform collision detection in 46 -- 154 cycles for a robot with 6 degrees of freedom. We evaluate MPAccel to execute a state-of-the-art learning-based motion planning algorithm. Our simulations suggest MPAccel can achieve real-time motion planning for a robot with 7 degrees of freedom in 0.014ms-0.49ms with an average latency of 0.099ms compared to 1.42ms on a CPU-GPU system.
Deval Shah, Ningfeng Yang, Tor M. Aamodt
ISCA2
2022 In-ConcReTeS: Interactive Consistency meets Distributed Real-Time Systems, Again!
abstract
The problem of replica coordination is fundamental to building Byzantine fault-tolerant (BFT) distributed systems. Seminal BFT architectures for safety-critical real-time systems from the eighties and nineties relied on custom processors and networks, and are hence not readily usable today. Modern-day deployments on cloud platforms do not “scale down” to embedded platforms and are not designed around timeliness. Recent work on real-time BFT protocols focuses on simulations and reliability analyses. In short, there exist no easily programmable BFT libraries that can be conveniently retrofitted onto real-time applications with deadlines and that perform well on embedded platforms. We propose In-ConcReTeS, a BFT key-value store designed for building highly reliable control applications on commodity embedded platforms. At its core, In-ConcReTeS is a real-time friendly redesign and an efficient implementation of a BFT protocol used by seminal fault-tolerant architectures. We evaluated In-ConcReTeS using an inverted pendulum simulation and an automotive benchmark on a cluster of four Raspberry Pis connected over Ethernet. Our results show that, unlike Redis and etcd, In-ConcReTeS can repeatedly synchronize hundreds of key-value pairs, while tolerating faults, every tens of milliseconds.
Arpan Gujarati, Ningfeng Yang, Björn B. Brandenburg
RTSS2
2018 Dynamic Projection Mapping on Multiple Non-rigid Moving Objects for Stage Performance Applications
Ryohei Nakatsu, Ningfeng Yang, Hirokazu Takata, Takashi Nakanishi, Makoto Kitaguchi, Naoko Tosa
ICEC2