Tiantai Deng

dblp:235/9977 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0003-4507-5746ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2026 An Approximate Computing-Based Spiking Neural Networks Neuron Model and STDP Learning
Haihang Xia, Yuqin Zhao, John Goodenough 0001, G. Charith K. Abhayaratne, Sizhao Li, Tiantai Deng
IEEE Trans. Circuits Syst. I Regul. Pap.8
2026 An Adaptive Gradient Recurrent Neural Network for Dynamic Quadratic Programming and its FPGA Implementation
Huang Ouyang, Kaiyuan Yang 0002, Tiantai Deng, Long Jin 0001
IEEE Trans. Sustain. Comput.3
2025 Stochastic Trajectory Prediction via Brownian LSTM Network
abstract
Human trajectory prediction has become an active research area, with applications in various scenarios such as evacuation situation analysis, deployment of intelligent transportation systems, and traffic operations. Pedestrian motions exhibit uncertainty and involve complex interactions with vehicles and the environment. However, some existing methods require substantial resources and the construction of complex models to handle these issues. In contrast, we propose the Brownian LSTM Network (BLN), which combines simple deterministic models with empirical models to achieve efficient prediction. Our results show an improvement over the state of art by 19%/21% on the ADE/FDE metrics, respectively. The improvement are achieved with a 4 times faster inference speed than previously baselines and a significant reduction in parameters. We present a simple and flexible scheme that can seamlessly integrate different networks. In addition, we introduce a Brownian motion module that leverages its irregular properties to simulate movement uncertainty, support multimodal predictions, and optimize the output of deterministic networks for the best prediction results. Extensive experiments on human trajectory prediction benchmarks, including the Stanford Drone and ETH/UCY datasets, demonstrate the superiority of our method.
Sizhao Li, Tiantai Deng, Huosheng Xu
IJCNN4
2025 Enhanced multimodal prediction via feature fusion and momentum buffering
Sizhao Li, Tiantai Deng, Huosheng Xu
Expert Syst. Appl.4
2025 Periodic-noise-tolerant neurodynamic approach for kWTA operation applied to opinions evolution
Jiexing Li, Yongji Guan, Tiantai Deng, Long Jin 0001
Neural Networks3
2025 A 3-D Multi-Precision Scalable Systolic FMA Architecture
abstract
Artificial Intelligence (AI) has almost become the default approach in a wide range of applications, such as computer vision, chatbots, and natural language processing. These AI-based applications require computing large-scale data with sufficient precision, typically in floating-point numbers, within a limited time window. A primary target for AI acceleration is matrix multiplication, mainly involving dot products through Multiply-Accumulate (MAC) operations. Current research employs the Fused Multiply-Add (FMA) operation, based on IEEE-754 Floating Point (FP) standard, to meet these requirements. However, current research focuses more on simplifying the internal digital circuits of the Processing Elements (PEs) performing FMA operations, rather than optimizing the FMA process specifically for MAC tasks. Current PE arrays often use a two-dimensional (2-D) systolic array design, without specific optimization for MAC operations, thus their parallelism is not fully utilized. Additionally, these designs lack reconfigurability and flexibility, leading to suboptimal performance on Field-Programmable Gate Arrays (FPGAs). Moreover, some designs adopt lower precision computing in AI inference for higher performance. However, some AI models still rely on high-precision computing to maintain the accuracy. Thus, multi-precision computing is commonly used in AI accelerators. To address these challenges, this paper proposes a novel Multi-Fused Multiply-Accumulate (MFMA) scheme and a corresponding three-dimensional (3-D) scalable systolic FP computing architecture. The MFMA scheme addresses the problem of the classical FMA scheme. It optimizes FMA for MAC operations with the Fused Multiply-Accumulate (FMAC) operation. Also, it combines multi-precision and mixed-precision FP computing methods for higher accuracy and lower overflow error. The proposed architecture integrates two 2-D systolic arrays into the PE for a 3-D systolic array, achieving higher parallelism and flexibility. The proposed scalable architecture can be customized to suit various FMAC operations. Compared with existing state-of-the-art FP architectures on FPGAs, our proposed architecture achieves 47%, 10%, and 159% energy efficiency improvements in FP32, FP16, and INT8 operations, respectively. Furthermore, our proposed architecture achieves energy efficiency improvements of 105%, 54%, and 262% under efficiency saturation conditions, outperforming the existing state-of-the-art design.
Xicheng Lu, Kaiyuan Yang 0002, Haihang Xia, Sizhao Li, Tiantai Deng
IEEE Trans. Circuits Syst. I Regul. Pap.8