Rendong Ying

dblp:01/8814 · DBLP profile ↗
← Back
44ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0001-6670-149XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 8 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 NeuroUNI: A Unified Event-Driven Multi-Core Architecture Optimizing Neuromorphic Primitives for Brain-Inspired Computing
abstract
The hardware convergence of Artificial Neural Networks (ANNs) and Spiking Neural Networks (SNNs) is hindered by conflicting computational paradigms: dense tensor parallelism versus asynchronous sparse dynamics. Existing unifications typically rely on inefficient spatial partitioning or mode-reconfigurable datapaths, limiting the flexibility needed by heterogeneous ANN-SNN hybrid models requiring frequent cross-domain interaction. To resolve this, we present NeuroUNI, a unified event-driven multi-core architecture. Unlike partitioned designs, NeuroUNI unifies computation at the primitive level using a novel Five-Tuple Event Model, abstracting both continuous activations and discrete spikes to naturally leverage dynamic sparsity. The architecture features a co-optimized hierarchical Macro-Micro-$\mu $OP ISA, a superscalar SIMD-based microarchitecture, and a decentralized multi-core synchronization protocol. Validated in TSMC 28nm technology via post-synthesis simulation and on a Xilinx VCU129 FPGA prototype, NeuroUNI demonstrates competitive cross-paradigm efficiency. It achieves$35.0\times $and$1.21\times $the ANN energy efficiency of the NVIDIA V100 and EyerissV2, respectively, while delivering$5.7\times $the SNN throughput of TrueNorth. In a unified mapless navigation workload, NeuroUNI attains 422.6 GOPS/W (ANN) and 190.8GSOPS/W (SNN), outperforming Loihi 1 with$56.7\times $the throughput and$3.85\times $the energy efficiency, proving the viability of a primitive/ISA-level unified silicon substrate.
Faquan Chen, Qingyang Tian, Ziren Wu, Xiangcheng Shi, Rendong Ying, Fei Wen 0005
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 A Skeleton-Based Topological Planner for Exploration in Complex Unknown Environments
abstract
The capability of autonomous exploration in complex, unknown environments is important in many robotic applications. While recent research on autonomous exploration have achieved much progress, there are still limitations, e.g., existing methods relying on greedy heuristics or optimal path planning are often hindered by repetitive paths and high computational demands. To address such limitations, we propose a novel exploration framework that utilizes the global topology information of observed environment to improve exploration efficiency while reducing computational overhead. Specifically, global information is utilized based on a skeletal topological graph representation of the environment geometry. We first propose an incremental skeleton extraction method based on wavefront propagation, based on which we then design an approach to generate a lightweight topological graph that can effectively capture the environment's structural characteristics. Building upon this, we introduce a finite state machine that leverages the topological structure to efficiently plan coverage paths, which can substantially mitigate the back-and-forth maneuvers (BFMs) problem. Experimental results demonstrate the superiority of our method in comparison with state-of-theart methods. The source code will be made publicly available at: https://github.com/Haochen-Niu/STGPlanner.
Haochen Niu, Xingwu Ji, Lantao Zhang, Fei Wen 0005, Rendong Ying
ICRA5
2025 Low-cost Deployment and Acceleration of Event-based Spiking Convolutional Neural Networks
abstract
By simulating the neurodynamics of biological brains, Spiking Neural Networks (SNNs) leverage sparse spike signal, eliminating the continuous multiply-accumulate operations of traditional Artificial Neural Networks (ANNs). Event-driven SNN processing offers significant advantages in energy efficiency and latency, making it ideal to be deployed on low-end processors. Spiking Convolutional Neural Networks (SC-NNs), which incorporate event-based processing, are increasingly employed for their power efficiency and ability to process spatio-temporal information. Unlike fully-connected networks, which rely on regular vector calculations, convolution operations present challenges for event-driven computation due to their sliding window nature.In this work, we deployed a compact 7-layer SCNN on the ARM Cortex-A9 core of PYNQ Z2 development board, using event-based acceleration. By optimizing data storage and processing sequences, we achieved efficient low-cost deployment. Offline training on the DVS128-Gesture dataset for an object recognition task yielded an accuracy of 93.40%. Following low-precision quantization and deployment, the model maintained a considerable accuracy of 92.36%. Compared to traditional periodic computation, event-based convolution processing achieved a 12.87× speedup. Furthermore, by exploiting the parallelism of feature map data storage along the channel dimension and ARM NEON instruction set, we gained an additional 2.97× speedup.
Qingyang Tian, Faquan Chen, Lisheng Xie, Ziren Wu, Liangshun Wu, Rendong Ying
ISCAS7
2024 SPRCPl: An Efficient Tool for SNN Models Deployment on Multi-Core Neuromorphic Chips via Pilot Running
abstract
This paper introduce SPRCpl, an efficient compiler/toolkit for deploying Spiking Neural Network (SNN) models on multi-core neuromorphic chips. It uses "pilot running" to optimize the deployment process. It includes a front-end compiler, synapse pruning and regeneration optimizer, and a mapping tool. SPRCpl proposes a synapse pruning scheme based on spike firing statistics obtained through pilot running, dynamically reducing model size. It also presents a mapping scheme that minimizes strikes within and between clusters using spike firing statistics and multi-objective optimization. Experimental results demonstrate SPRCpl’s effectiveness in maintaining model accuracy during pruning and outperforming SpiNeMap in terms of communication count, execution time, and memory usage. It achieves lower latency, reduced power consumption, and higher throughput, making it a promising tool for SNN model deployment on multi-core neuromorphic chips.
Liangshun Wu, Lisheng Xie, Jianwei Xue, Faquan Chen, Qingyang Tian, Ziren Wu, Rendong Ying
ISCAS8
2023 SpikeNC: An Accurate and Scalable Simulator for Spiking Neural Network on Multi-Core Neuromorphic Hardware
abstract
Multi-core neuromorphic hardware for spiking neural networks (SNNs) has garnered considerable attention due to its biological plausibility and energy efficiency. However, the performance of SNN applications on such hardware is constrained by the rigid architecture and interconnection among neuron cores. To enable early-stage evaluation of SNN performance on multi-core neuromorphic hardware, we introduce an accurate and scalable simulator, SpikeNC. We present the entire workflow, ranging from SNN model training to simulation, providing comprehensive insights into both model and Network-on-Chip (NoC) related statistics. Moreover, we identify a considerable amount of time wastage in the widely adopted tick-based synchronous scheme. A three-stage agent-based asynchronous scheme is proposed for fast simulation. We evaluate the performance of deep spiking neural networks (DSNNs) with various scales trained on spike-converted datasets using SpikeNC. The results demonstrate that SpikeNC achieves precise and scalable simulation for SNNs on multi-core neuromorphic hardware. Additionally, the proposed asynchronous scheme significantly reduces the simulation cycles and absolute simulation time by approximately 63 % and 56 % respectively, compared to the synchronous scheme. We also delve into the trade-offs between different design parameters and explore the influence of mapping schemes utilizing SpikeNC.
Lisheng Xie, Jianwei Xue, Liangshun Wu, Faquan Chen, Qingyang Tian, Rendong Ying
HiPC7
2023 ParallelNN: A Parallel Octree-based Nearest Neighbor Search Accelerator for 3D Point Clouds
abstract
As Light Detection And Ranging (LiDAR) increasingly becomes an essential component in robotic navigation and autonomous driving, the processing of high throughput 3D point clouds in real time is widely required. This work considers the point cloud k-Nearest Neighbor (kNN) search, which is an important 3D processing kernel. Although applying fine-grained parallelism optimization on internal processing, e.g., using multiple workers, has demonstrated high efficiency, previous accelerators with DDR external memory are fundamentally limited by the external bandwidth bottleneck. To break this bottleneck, this work proposes a highly parallel architecture, namely ParallelNN, for highly efficient kNN search processing of high throughput point clouds. First, we optimize the multichannel cache based on High Bandwidth Memory (HBM) and on-chip memory to provide large external bandwidth. Then, a novel parallel depth-first octree construction algorithm is proposed and mapped onto multiple construction branches with trace-coded construction queues, which can regularize random accesses and perform multi-branch octree construction efficiently. Furthermore, in the search stage, we present algorithm-architecture co-optimization, including parallel keyframe-based scheduling and multi-branch flexible search engines, to provide conflict-free access and maximum reuse opportunities for reference points, which achieves more than 27.0× speedup compared with baseline architectures. We prototype ParallelNN on Virtex HBM FPGA and perform extensive benchmarking on the KITTI dataset. The results demonstrate that ParallelNN achieves up to 107.7× and 12.1× speedup over CPU and GPU implementations, while being more energy efficient, e.g., outperforming CPU and GPU implementations by 73.6× and 31.1×, respectively. Besides, with the proposed algorithm-architecture co-optimization, ParallelNN achieves 11.4× speedup over state-of-the-art architecture. Moreover, ParallelNN is configurable and can be easily generalized to similar octree-based applications.
Faquan Chen, Rendong Ying, Jianwei Xue, Fei Wen 0005
HPCA2
2023 SPAR: An efficient self-attention network using Switching Partition Strategy for skeleton-based action recognition
Rendong Ying, Fei Wen 0005
Neurocomputing2
2023 Domain Generalization for Face Anti-Spoofing via Negative Data Augmentation
abstract
In practical applications, the generalization capability of face anti-spoofing (FAS) models on unseen domains is of paramount importance to adapt to diverse camera sensors, device drift, environmental variation, and unpredictable attack types. Recently, various domain generalization (DG) methods have been developed to improve the generalization capability of FAS models via training on multiple source domains. These DG methods commonly require collecting sufficient real-world attack samples of different attack types for each source domain. This work aims to learn a FAS model without using any real-world attack sample in any source domain but can generalize well to the unseen domain, which can significantly reduce the learning cost. Toward this goal, we draw inspiration from the theoretical error bound of domain generalization to use negative data augmentation instead of real-world attack samples for training. We show that using only a few types of simple synthesized negative samples, e.g., color jitter and color mask, the learned model can achieve competitive performance over state-of-the-art DG methods trained using real-world attack samples. Moreover, a dynamic global common loss and a local contrast loss are proposed to prompt the model to learn a compact and common feature representation for real face samples from different source domains, which can further improve the generalization capability. Experimental results of extensive cross-dataset testing demonstrate that our method can even outperform state-of-the-art DG methods using real-world attack samples for training. The code for reproducing the results of our method is available at https://github.com/WeihangWANG/NDA-FAS.
Weihang Wang 0003, Rendong Ying, Fei Wen 0005
IEEE Trans. Inf. Forensics Secur.4
2023 SFANC: Scalable and Flexible Architecture for Neuromorphic Computing
abstract
Spiking neural networks (SNNs) are recognized as the third generation of neural networks, boasting remarkable computational capabilities, which also require massive computational resources and flexibility to simulate biological neural functions. In this work, we present SFANC, a scalable and flexible neuromorphic architecture for SNNs based on optimized router architecture of network-on-chip (NoC) and highly programmable neuromorphic cores (NCs). SFANC includes 16 NCs, supporting 8000 neurons and four million synapses. The NCs are based on RISC-V with SNN-specific instructions, providing remarkable flexibility and computational speed. The NoC’s spiking routers facilitate multicast routing, flow estimation, and buffer sharing, contributing to increased scalability. Additionally, we propose a spectral cluster mapping approach for efficiently deploying SNNs onto SFANC, ensuring high flexibility and parallelism. We analyze the processing speedup for typical SNN topologies using the leaky-integrate-and-fire (LIF) neuron model within the NCs, achieving up to an$8.6\times $speedup over the RISC-V core on average under different SNN applications. Our enhanced spiking router demonstrates a spike latency reduction of up to 47.4% under various spiking coding schemes. Moreover, when applied to typical SNN topologies, our method exhibits an average spike latency decrease of up to 32.5% compared to the sequential mapping utilized by SpiNNaker. In summary, this work demonstrates high performance, flexibility, and scalability for simulating and accelerating SNNs, showcasing its potential as a promising solution for neuromorphic architectures.
Jianwei Xue, Rendong Ying, Faquan Chen
IEEE Trans. Very Large Scale Integr. Syst.2
2022 MaxEnt Dreamer: Maximum Entropy Reinforcement Learning with World Model
abstract
Model-based reinforcement learning algorithms can alleviate the low sample efficiency problem compared with modelfree methods for control tasks. However, the learned policy's performance often lags behind the best model-free algorithms since its weak exploration ability. Existing model-based reinforcement learning algorithms learn policy by interacting with the learned world model and then use the learned policy to guide a new round of world model learning. Due to weak policy exploration ability, the learned world model has a large bias. As a result, it fails to learn the globally optimal policy on such a world model. This paper improves the learned world model by maximizing both the reward and the corresponding policy entropy in the framework of maximum entropy reinforcement learning. The effectiveness of applying the maximum entropy approach to model-based reinforcement learning is supported by the better performance of our algorithm on several complex mujoco and deepmind control suite tasks.
Hongying Ma, Wuyang Xue, Rendong Ying
IJCNN3
2022 Conv-MLP: A Convolution and MLP Mixed Model for Multimodal Face Anti-Spoofing
abstract
Local features contain crucial clues for face anti-spoofing. Convolutional neural networks (CNNs) are powerful in extracting local features, but the intrinsic inductive bias of CNNs limits the ability to capture long-range dependencies. This paper aims to develop a simple yet effective framework that is versatile in extracting both local information and long-range dependencies for face anti-spoofing. To this end, we propose a novel architecture, namely Conv-MLP, which incorporates local patch convolution with global multi-layer perceptrons (MLP). Conv-MLP breaks the inductive bias limitation of traditional full CNNs and can be expected to better exploit long-range dependencies. Furthermore, we design a new loss specifically for the face anti-spoofing task, namely moat loss. The moat loss benefits discriminative representations learning and can improve the generalization capability on unseen presentation attacks. In this work, multi-modal data are directly fused at the signal level to extract complementary features. Extensive experiments on single and multi-modal datasets demonstrate that Conv-MLP outperforms existing state-of-the-art methods while being more computationally efficient. The code is available at https://github.com/WeihangWANG/Conv-MLP.
Weihang Wang 0003, Fei Wen 0005, Rendong Ying
IEEE Trans. Inf. Forensics Secur.4
2022 R-SDSO: Robust stereo direct sparse odometry
Ruihang Miao, Fei Wen 0005, Wuyang Xue, Rendong Ying
Vis. Comput.6
2021 On Perceptual Lossy Compression: The Cost of Perceptual Reconstruction and An Optimal Training Framework
abstract
Lossy compression algorithms are typically designed to achieve the lowest possible distortion at a given bit rate. However, recent studies show that pursuing high perceptual quality would lead to increase of the lowest achievable distortion (e.g., MSE). This paper provides nontrivial results theoretically revealing that, 1) the cost of achieving perfect perception quality is exactly a doubling of the lowest achievable MSE distortion, 2) an optimal encoder for the “classic” rate-distortion problem is also optimal for the perceptual compression problem, 3) distortion loss is unnecessary for training a perceptual decoder. Further, we propose a novel training framework to achieve the lowest MSE distortion under perfect perception constraint at a given bit rate. This framework uses a GAN with discriminator conditioned on an MSE-optimized encoder, which is superior over the traditional framework using distortion plus adversarial loss. Experiments are provided to verify the theoretical finding and demonstrate the superiority of the proposed training framework.
Fei Wen 0005, Rendong Ying
ICML3
2021 Monocular Semantic Mapping Based on 3D Cuboids Tracking
abstract
Semantic mapping based on information of objects has become a crucial component for the surrounding comprehension and the more robust navigation. In this paper, we propose a system for simultaneous localization and mapping (SLAM) that combines multiple objects tracking and factor graph optimization with semantically meaningful landmarks to achieve accurate monocular semantic mapping. Firstly, the process of object recognition uses a vanishing point sampling-based approach to efficiently infer the class and position of object landmarks from 2D bounding box object detection. Secondly, The semantic frontend utilizes local matching-based data association to track raw cuboid proposals. It can provide semantic constraints to reduce cuboid scale drift and improve its position estimation. Finally, we present a multi-view factor graph optimization which can use motion modal of the camera to optimize stable cuboids. The semantic mapping experiments on our own built virtual scene show better accuracy and robustness over existing approaches. We evaluate the effectiveness of our approach on public KITTI datasets and a real scene.
Xingwu Ji, Ruihang Miao, Wuyang Xue, Rendong Ying
ISCAS5
2021 Adaptive Stereo Direct Visual Odometry with Real-Time Loop Closure Detection and Relocalization
abstract
This paper presents a direct visual odometry method with adaptive stereo coupling factor and relocalization. The adaptive stereo coupling factor is calculated by considering the number of co-visibility frames, which makes stereo direct odometry more robust. The map is built up with keyframes, connecting relationship of keyframes, poses of keyframes and descriptors. The map will be saved automatically when the proposed system shuts down, and will be loaded automatically when the proposed system starts up. Our proposed system always builds a new map until a loop closure is found between new map and loaded map. When loop closure in a loaded map is found, the proposed system will locate camera in the loaded map. Moreover, our proposed system detects loop closure with feature-based bag-of-words (BoW) method on a low resolution level, which makes the loop closure detection achieve realtime. Evaluation results on the KITTI dataset show that, with the aid of the adaptive stereo coupling factor, the stereo visual odometry becomes more accurate and robust. And evaluation results in real world show that the proposed system can successfully detect loop closure and relocate in realtime.
Ruihang Miao, Wuyang Xue, Xingwu Ji, Rendong Ying
ISCAS6
2021 Fast and Positive Definite Estimation of Large Covariance Matrix for High-Dimensional Data Analysis
abstract
Large covariance matrix estimation is a fundamental problem in many high-dimensional statistical analysis applications arises in economics and finance, bioinformatics, social networks, and climate studies. To achieve reliable estimation in the high-dimensional setting, an effective technique is to exploit the intrinsic structure of the covariance matrix, e.g., by sparsity regularization. For sparsity regularization, the lasso penalty is popular and convenient due to its convexity but has a bias problem. A nonconvex penalty can alleviate the bias problem, but the involved nonconvex problem under positive-definiteness constraint is generally difficult to solve. In this work, we propose an efficient algorithm for positive-definiteness constrained covariance estimation by combining the iteratively reweighted method and the alternative direction method of multipliers (ADMM). The iterative reweighting scheme can achieve better sparsity regularization than the lasso method. Meanwhile, the proposed algorithm solves convex subproblems in each iteration and hence is easy to converge. The efficiency and effectiveness of the proposed algorithm has been demonstrated by both simulation study and a gene clustering example for tumor tissues. Code for reproducing the results is available at https://github.com/FWen/pdlc.git.
Fei Wen 0005, Rendong Ying
IEEE Trans. Big Data3
2021 A Simple Local Minimal Intensity Prior and an Improved Algorithm for Blind Image Deblurring
abstract
Blind image deblurring is a long standing challenging problem in image processing and low-level vision. Recently, sophisticated priors such as dark channel prior, extreme channel prior, and local maximum gradient prior, have shown promising effectiveness. However, these methods are computationally expensive. Meanwhile, since these priors involved subproblems cannot be solved explicitly, approximate solution is commonly used, which limits the best exploitation of their capability. To address these problems, this work firstly proposes a simplified sparsity prior of local minimal pixels, namely patch-wise minimal pixels (PMP). The PMP of clear images is much more sparse than that of blurred ones, and hence is very effective in discriminating between clear and blurred images. Then, a novel algorithm is designed to efficiently exploit the sparsity of PMP in deblurring. The new algorithm flexibly imposes sparsity inducing on the PMP under the maximum a posterior (MAP) framework rather than directly uses the half quadratic splitting algorithm. By this, it avoids non-rigorous approximation solution in existing algorithms, while being much more computationally efficient. Extensive experiments demonstrate that the proposed algorithm can achieve better practical stability compared with state-of-the-arts. In terms of deblurring quality, robustness and computational efficiency, the new algorithm is superior to state-of-the-arts. Code for reproducing the results of the new method is available at https://github.com/FWen/deblur-pmp.git.
Fei Wen 0005, Rendong Ying, Yipeng Liu 0001, Trieu-Kien Truong
IEEE Trans. Circuits Syst. Video Technol.2
2020 Restoration of Motion Blur in Time-of-Flight Depth Image Using Data Alignment
abstract
Time-of-flight (ToF) sensors are vulnerable to motion blur in the presence of moving objects. This is due to the principle of ToF camera that it estimates depth from the phase-shift between emitted and received modulated signals. And the phase-shift is measured by four sequential phase-shifted images, which is assumed to be consistent in an integration time. However, object motion would give rise to disparity among the four phase-shifted images, contributing to unreliable depth measurement. In this paper, we propose a novel method that is capable of aligning the four phase-shifted images through investigating the electronic value of each pixel in the phase images. It consists of two steps, motion detecting and deblurring. Furthermore, a refinement utilizing an additional group of phase-shifted images is adopted to further improve the accuracy of depth measurement. Experiment results on a new elaborated dataset with ground-truth demonstrate that the proposed method compares favorably over existing methods in both accuracy and runtime. Particularly, the new method can achieve the best accuracy while being computationally efficient that can support real-time running.
Fei Wen 0005, Jun Wang 0137, Rendong Ying
3DV5
2020 A Platform for Full-Stack Functional Programming
abstract
Traditional CPU design shows signs of fatigue, expressed as overwhelming security vulnerabilities. As we investigate functional programming as an alternative to the insecure, traditional imperative programming model, the inexistence of a stable functional programming system infrastructure and code base acts as the classic chicken-and-egg problem. There is no functional programming basic software because there are no functional programming machines, and vice-versa. In this paper we attempt to break this cycle by designing a baseline platform that enables the research on the practical security properties of architectures under a discipline of full-stack functional programming.
Cecil Accetti, Rendong Ying
ISCAS3
2020 Robust PCA Using Generalized Nonconvex Regularization
abstract
Recently, the robustification of principal component analysis (PCA) has attracted much research attention in numerous areas of science and engineering. The most popular and successful approach is to model the robust PCA problem as a low-rank matrix recovery problem in the presence of sparse corruption. With this model, the nuclear norm and 11-norm penalties are usually used for low-rank and sparsity promotion. Although the nuclear norm and 11-norm are favorable due to their convexity, they have a bias problem. In comparison, nonconvex penalties can be expected to yield better recovery performance. In this paper, we consider a formulation for robust PCA using generalized nonconvex penalties for low-rank and sparsity inducing. This nonconvex formulation is efficiently solved by a multi-block alternative direction method of multipliers (ADMM) algorithm. A sufficient condition for the convergence of this new ADMM algorithm has been derived. Furthermore, to address the important issue of nonconvex penalty selection, we have evaluated the new algorithm via numerical experiments in various low-rank and sparsity conditions. The results indicate that, “exact” recovery of the low-rank principle component can be achieved only by nonconvex regularization. MATLAB code is available at https://github.com/FWen/RPCA.git.
Fei Wen 0005, Rendong Ying, Robert C. Qiu
IEEE Trans. Circuits Syst. Video Technol.2
2020 Efficient Algorithms for Maximum Consensus Robust Fitting
abstract
Maximum consensus robust fitting is a fundamental problem in many computer vision applications, such as vision-based robotic navigation and mapping. While exact search algorithms are computationally demanding, randomized algorithms are cheap but the solution quality is not guaranteed. Deterministic algorithms fill the gap between these two kinds of algorithms, which have better solution quality than randomized algorithms while being much faster than exact algorithms. In this article, we develop two highly efficient deterministic algorithms based on the alternating direction method of multipliers (ADMM) and proximal block coordinate descent (BCD) frameworks. Particularly, the proposed BCD algorithm is guaranteed convergent. Furthermore, on the slack variable in the BCD algorithm, which indicates the inliers and outliers, we establish some meaningful properties, such as support convergence within finite iterations and convergence to restricted strictly local minimizer. Compared with state-of-the-art algorithms, the new algorithms with initialization from a randomized or convex relaxed algorithm can achieve improved solution quality while being much more efficient (e.g., more than an order of magnitude faster). An application of the new ADMM algorithm in simultaneous localization and mapping (SLAM) has also been provided to demonstrate its effectiveness. Code for reproducing the results is available online.
Fei Wen 0005, Rendong Ying
IEEE Trans. Robotics2
2019 Action Recognition Based on 3D Skeleton and RGB Frame Fusion
abstract
Action recognition has wide applications in assisted living, health monitoring, surveillance, and human-computer interaction. In traditional action recognition methods, RGB video-based ones are effective but computationally inefficient, while skeleton-based ones are computationally efficient but do not make use of low-level detail information. This work considers action recognition based on a multimodal fusion between the 3D skeleton and the RGB image. We design a neural network that uses a 3D skeleton sequence and a single middle frame from an RGB video as input. Specifically, our method picks up one frame in a video and extracts spatial features from it using two attention modules, a self-attention module and a skeleton-attention module. Further, temporal features are extracted from the skeleton sequence via a BI-LSTM sub-network. Finally, the spatial features and the temporal features are combined via a feature fusion network for action classification. A distinct feature of our method is that it uses only a single RGB frame rather than an RGB video. Accordingly, it has a light-weighted architecture and is more efficient than RGB video-based methods. Comparative evaluation on two public datasets, NTU-RGBD and SYSU, demonstrates that, our method can achieve competitive performance compared with state-of-the-art methods.
Guiyu Liu, Jiuchao Qian, Fei Wen 0005, Xiaoguang Zhu, Rendong Ying
IROS5
2019 Design of a Reconfigurable Multi-Sensor Testbed for Autonomous Vehicles and Ground Robots
abstract
In the field of autonomous vehicles and robotics research, a well-designed testbed can provide a convenient and safe development environment. In this paper, we propose and implement a reconfigurable testbed equipped with multiple heterogeneous sensors and a compact but powerful local computation unit. The proposed testbed benefits from a self-sustaining design, and can be smoothly reconfigured to different varieties of vehicles. These features make the testbed a novel and ideal platform for testing and verifying navigation, recognition and control algorithms under diverse scenarios.
Wuyang Xue, Ziang Liu 0001, Yimo Zhao, Ruihang Miao, Rendong Ying
ISCAS6
2019 3DTI-Net: Learn 3D Transform-Invariant Feature Using Hierarchical Graph CNN
Guanghua Pan, Jun Wang 0137, Rendong Ying, Fei Wen 0005
PRICAI (2)4
2016 An improved sparse reconstruction algorithm for speech compressive sensing using structured priors
abstract
This work addresses the issue of sparse reconstruction in compressive sensing (CS) for speech signals. We propose a novel sparse reconstruction algorithm based on the approximate message passing (AMP) framework, via exploiting the intrinsic structures of real-life speech signals in the modified discrete cosine transform (MDCT) domain. We use a Gaussian mixture model to characterize the marginal distribution of the MDCT coefficients, and employ a first order Markov chain model to capture the inter-dependencies between neighboring MDCT coefficients. The parameters of these two models are adaptively learned using an expectation-maximization (EM) learning procedure. Compared with several state-of-the-art algorithms, the new algorithm showed significantly better performance in reconstruction experiments on real speech signals.
Xiaobo Jiang, Rendong Ying, Fei Wen 0005, Sumxin Jiang
ICME2
2016 HAVA: Heterogeneous Multicore ASIP for Multichannel Low-Bit-Rate Vocoder Applications
abstract
As are widely used in military and security fields, multiple channels of low-bit-rate vocoders are required to perform on embedded devices efficiently. We propose HAVA, a multicore Application Specific Instruction Set Processor for multichannel low-bit-rate vocoders with real-time performance. To provide both flexibility and efficiency, HAVA integrates two types of processing cores and a shared-memory core on a 2-D-mesh on-chip network. Adopting a single-Instruction Set Architecture heterogeneous multicore architecture, HAVA cuts down the real-time performance requirement of vocoders by over 40% compared with other platforms. By leveraging the on-chip network for intercore communication, HAVA can perform multichannel vocoders with a marginal efficiency loss. The chip implementation of HAVA is finished in a 40-nm CMOS technology and it dissipates 149 mW at 100-MHz operating frequency for four channels of encoders.
Zhenqi Wei, Rongdi Sun, Zunquan Zhou, Xiangming Geng, Rendong Ying
IEEE Trans. Very Large Scale Integr. Syst.7
2015 TAB barrier: Hybrid barrier synchronization for NoC-based processors
abstract
As one of the mostly used synchronization schemes in parallel programming on multi-core processors, barrier synchronization has been extensively studied in former research works. In conventional master-slave barrier or tree barrier, usually one centric core is selected to collect barrier arriving messages and to broadcast barrier releasing messages. Unfortunately the barrier core sometimes is deviated from the center location and may lead to worse synchronization efficiency. We propose a hybrid tree-based all-to-all (TAB) barrier for NoC-based many-core processors to relieve performance degradation caused by the off-centered barrier core. Performance of TAB barrier is compared to canonical algorithms and former solution, and almost 20% time is saved during off-centered scenarios with marginal area and power overhead.
Zhenqi Wei, Rongdi Sun, Rendong Ying
ISCAS4
2015 Finite-state entropy-constrained vector quantiser for audio modified discrete cosine transform coefficients uniform quantisation
abstract
In this paper, an entropy‐constrained vector quantiser (ECVQ) scheme with finite memory, called finite‐state ECVQ (FS‐ECVQ), is presented. This scheme consists of a finite‐state vector quantiser (FSVQ) and multiple component ECVQs. By utilising the FSVQ, the inter‐frame dependencies within source sequence can be effectively exploited and no side information needs to be transmitted. By employing the ECVQs, the total memory requirements of FS‐ECVQ can be efficiently decreased while the coding performance is improved. An FS‐ECVQ, designed for the modified discrete cosine transform coefficients coding, was implemented and evaluated based on the unified speech and audio coding (USAC) scheme. Results showed that the FS‐ECVQ achieved reduction of the total memory requirements by 92.3%, compared with the encoder in USAC working draft 6 (WD6), and over 10%, compared with the encoder in USAC final version (FINAL), while maintaining coding performance similar to FINAL, which was about 4% better than that of WD6.
Sumxin Jiang, Rendong Ying
IET Signal Process.2
2015 Distributed Compressed Sensing off the Grid
abstract
This letter investigates the joint recovery of a frequency-sparse signal ensemble sharing a common frequency-sparse component from the collection of their compressed measurements. Unlike conventional arts in compressed sensing, the frequencies follow an off-the-grid formulation and are continuously valued in [0, 1]. As an extension of atomic norm, the concatenated atomic norm minimization approach is proposed to handle the exact recovery of signals, which is reformulated as a computationally tractable semidefinite program. The optimality of the proposed approach is characterized using a dual certificate. Numerical experiments are performed to illustrate the effectiveness of the proposed approach and its advantage over separate recovery.
Zhenqi Lu, Rendong Ying, Sumxin Jiang, Wenxian Yu
IEEE Signal Process. Lett.2
2014 Spectral compressive sensing with model selection
abstract
The performance of existing approaches to the recovery of frequency-sparse signals from compressed measurements is limited by the coherence of required sparsity dictionaries and the discretization of frequency parameter space. In this paper, we adopt a parametric joint recovery-estimation method based on model selection in spectral compressive sensing. Numerical experiments show that our approach outperforms most state-of-the-art spectral CS recovery approaches in fidelity, tolerance to noise and computation efficiency.
Zhenqi Lu, Rendong Ying, Sumxin Jiang, Zenghui Zhang, Wenxian Yu
ICASSP2
2014 Lossy audio signal compression via structured sparse decomposition and compressed sensing
abstract
In this paper, we propose a method for lossy audio signal compression via structured sparse decomposition and compressed sensing (CS). In this method, a least absolute shrinkage and selection operator (LASSO) is employed to sparse and structured decompose the audio signals into tonal and transient layers, and then, both resulting layers are compressed by a CS method. By employing a new penalty term, which takes advantage of the structure information of transform coefficients, the LASSO is able to achieve a better sparse approximation of the audio signal than traditional methods do. In addition, we propose a sparsity allocation algorithm, which adjusts the sparsity between the two resulting layers, thus improving the performance of CS. Experimental results showed that the new method provided a better compression performance than conventional methods did.
Sumxin Jiang, Rendong Ying, Zhenqi Lu, Zenghui Zhang
ICME2
2014 Efficient Post-Processing Signal Detecting for Distributed Environment Monitoring in Smart Grid
abstract
In harsh and complex electric power system environments of Smart Grid, the observed signal is always disturbed by multiple signal sources, environmental noise or sudden packet loss, generating a significant challenge for reliable remote monitoring. In this paper, we propose a new post-processing detecting scheme for wireless data transmission in the distributed monitoring system to highly tolerate these interferences after transmission procedure and exactly reconstruct the observed signal from rough data in the sink center. Hardware or algorithm implementation to prevent the possible failures is not necessary in sensors, which perfectly solves the problem about changing batteries or updating software for practical application with thousands of energy-constrained sensors. The sink center with high computational capability and abundant energy resource decides the tolerance of interference and retransmission condition, reconstructs the observed signal from the mixing sensed data of sensors by the estimation of frequency domain amplitude via compressed sensing with regarding frequency domain amplitude of the interference signal as 0. In the proposed scheme, we implement compressed sensing with an l1-l2 optimization, where the cost function is defined as a sum of l1 and l2 norms with a mixing parameter, which enables us to control the threshold between the observed signal and the interference ones.
Yujie Liang, Rendong Ying
Intelligent Environments2
2014 Instruction-based high-efficient synchronization in a many-core Network-on-Chip processor
abstract
Parallelized applications running on many-core Network-on-Chip (NoC) processors may consume a great part of execution time to synchronize threads mapped on multiple NoC nodes, if synchronization for NoC processors is not carefully designed. In this paper, we propose an instruction-based synchronization solution applied in a packet-switched many-core NoC processor with 2D mesh grid topology. Return links are added into the on-chip network to transmit acknowledgements of read requests, while a specific instruction SET is designed as instruction set extension to the original pipeline to perform atomic read-modify-write operations. To support various synchronization schemes, a hardware unit SYNC containing globally addressable registers as shared variables is adopted to handle synchronization requests from both local and remote NoC nodes. Additionally, a FIFO located in the SYNC unit can store these synchronization requests to poll on shared variables locally. Thus, network contention due to busy-wait synchronization algorithms is greatly reduced. Synchronization schemes including spinlock, barrier, FIFO spinlock and semaphore are implemented as inline assembly functions. Synthesis results under 55nm process suggest low area and power overhead of the hardware design. Performance of synchronization schemes are evaluated and are compared to results of conventional methods and prior works, showing the proposed solution is of higher efficiency.
Zhenqi Wei, Zhencheng Zeng, Jiangwei Xu, Rendong Ying
ISCAS5
2013 An improved indoor localization method using smartphone inertial sensors
abstract
In this paper, an improved indoor localization method based on smartphone inertial sensors is presented. Pedestrian dead reckoning (PDR), which determines the relative location change of a pedestrian without additional infrastructure supports, is combined with a floor plan for a pedestrian positioning in our work. To address the challenges of low sampling frequency and limited processing power in smartphones, reliable and efficient PDR algorithms have been proposed. A robust step detection technique leaves out the preprocessing of raw signal and reduces complex computation. Given the fact that the precision of the stride length estimation is influenced by different pedestrians and motion modes, an adaptive stride length estimation algorithm based on the motion mode classification is developed. Heading estimation is carried out by applying the principal component analysis (PCA) to acceleration measurements projected to the global horizontal plane, which is independent of the orientation of a smartphone. In addition, to eliminate the sensor drift due to the inaccurate distance and direction estimations, a particle filter is introduced to correct the drift and guarantee the localization accuracy. Extensive field tests have been conducted in a laboratory building to verify the performance of proposed algorithm. A pedestrian held a smartphone with arbitrary orientation in the tests. Test results show that the proposed algorithm can achieve significant performance improvements in terms of efficiency, accuracy and reliability.
Jiuchao Qian, Jiabin Ma, Rendong Ying, Ling Pei
IPIN3
2013 Efficient middleware for network evaluation and optimization in Wireless Sensor Network design
abstract
A computer aided network evaluation and optimization platform is presented to help the development of WSN application. This middleware together with the existing network simulators replace the function of field test during the process of WSN designs. The proposed platform reduces engineering cost and increases development efficiency. Furthermore, the network optimization is implemented in the platform to provide an automatic design of WSN application.
Yujie Liang, Rendong Ying
ISCAS2
2013 A method for optimal SINR under non-i.i.d. interferences
abstract
Interference is a major limiting factor for wireless communications. When interference is independent and identically distributed (i.i.d.), its statistics is invariant with time. The conventional statistical analysis can be used to mitigate the interferences. However, when the interferences are non-i.i.d., few signal processing techniques are available for mitigating the interferences. The matched filter technology is probably the best signal processing technique. In this presentation, a least cross-correlation criterion based rotation-matched filter is presented that can mitigate the interferences and achieve the optimal signal to interference-plus-noise ratio (SINR). Based on the simulation study it is better than the matched filter by 5∼10 dB.
Ruey-Wen Liu, Rendong Ying, Bo Hu 0002
ISCAS2
2013 Optimization of ETSI DSR frontend software on a high-efficient audio DSP
abstract
Server-terminal based distributed speech recognition (DSR) applications are widely adopted on mobile devices. In this paper, we have implemented a power-efficient DSR solution of high performance for real-time speech processing. The DSR frontend algorithms are elaborately optimized in assembly codes utilizing accelerating technics provided by a previously released audio DSP, such as binary scaling operations in a deep instruction pipeline, automatic memory addressing method, and parallel processing of packaged data. The performance of DSR frontend software running on the DSP is greatly improved, and our work is of best efficiency compared with former solutions. The realtime frequency of processing 16 kHz input streams is 124.3 MHz and is only about 30% of what is required on a TI C64x DSP. Based on simulation experiment under SMIC 130 nm process, the power consumed for DSR frontend processing is 23 mW. Besides, the presented implementation of the algorithms is also integrated in a server-terminal demo system, and is proved to be worked well in real speech recognition applications.
Zhenqi Wei, Cun Yu, Hongbin Zhou, Ji Kong, Rendong Ying
ISCAS7
2012 A multiple access for unlicensed spectrum
abstract
The current multiple accesses, such as CDMA, are designed for multi-user wireless communication systems in the licensed spectrum, where uncoordinated interferences can be kept to a minimum. On the other hand, the uncoordinated interferences do exist in the unlicensed spectrum, and they compromise the orthogonal properties and hence degrade the performance of the current multiple accesses. In this presentation, the Autocorrelation-Division Multiple Access (ADMA) is presented, which achieves the best SINR among all passive pre-/post-filters, by eliminating all MAI and CCI, and retaining the signal power when the channel is lossless without enhancing the noise power.
Ruey-Wen Liu, Rendong Ying, Bo Hu 0002
ISCAS2
2011 StreamPoP: Stream programming oriented power-efficient audio DSP
abstract
In stream programming style, the computation and the memory accesses are decoupled as much as possible. Such kind of programming style has brought new profits both on performance and power-efficiency for digital signal processing. In this paper, the architecture of a stream programming oriented power-efficient digital signal processor named as StreamPoP has been presented for consumer audio applications. The instruction and data supply subsystem of StreamPoP has been well designed for flexible stream programming and high power-efficiency. To further reduce the power consumption, an audio computation specific deep-instruction-pipeline (DIP) has been used in the micro-architecture of StreamPoP. To evaluate the performance of StreamPoP for consumer audio applications, a set of audio benchmarks have been used. It has been presented in this paper that the performance of StreamPoP is better than conventional high performance DSPs, while the former architecture is much more power-efficient for audio applications. The simulated power consumption result of StreamPoP under TSMC 90nm process is 5.1mW for AAC real-time decoding.
Ji Kong, Zhenqi Wei, Rendong Ying
ISCAS6
2010 Split Table Extension: A Low Complexity LVQ Extension Scheme in Low Bitrate Audio Coding
abstract
Embedded algebraic vector quantization (EAVQ) is a fast and efficient lattice vector quantization (LVQ) scheme used in low-bitrate audio coding. However, a defect of EAVQ is the overload distortion which causes unpleasant noises in audio coding. To solve this problem, specific base codebook extension schemes should be carefully considered. In this letter, we present a novel EAVQ codebook extension scheme-split table extension (STE), which splits a vector into two smaller vectors: one in the base codebook and the other in the split table. The base codebook and the split table are designed according to the appearance probability of quantized vectors in audio segments. Experiments on encoding multiple audio and speech sequences show that, compared with the existed Voronoi extension scheme, STE greatly reduces computation complexity and storage requirement while achieving similar coding quality.
Ji Kong, Rendong Ying
IEEE Signal Process. Lett.4
2010 Narrowband Interference Mitigation in DS-UWB Systems
abstract
Interference suppression is important for ultra-wideband (UWB) systems to operate over spectrum occupied by pre-existing narrowband systems. In this letter, we first extend the code-aided interference suppression technique to a direct sequence (DS) UWB system which has been proposed for narrowband interference (NBI) suppression in direct sequence spread spectrum (DSSS) systems. Then, we introduce a new type of spreading sequence, which significantly improves the NBI suppression capability of the code-aided technique. We also discuss a preamble-aided synchronization scheme for utilizing the new type of spreading sequence.
Maode Ma, Rendong Ying
IEEE Signal Process. Lett.3
2007 Interference-Blocking Algorithm for OFDM Systems: Insensitive to Time and Frequency Synchronization Error
abstract
A covariance matrix based anti-interference algorithm is proposed for SIMO OFDM systems. This algorithm constructs the interference-blocking filter without the information of the inferences. This algorithm is insensitive to the time synchronization error and to the carrier frequency synchronization error. Therefore, it is especially useful for the case when interferences are strong and unknown, such as the case in the unlicensed spectrum.
Rendong Ying, Ruey-Wen Liu, Guozhi Xu
WCNC1
2007 Covariance-Based Interference Blocking Algorithm for SIMO OFDM System
abstract
A covariance-based algorithm is proposed to find the interference blocking filter for single-input multiple-output orthogonal frequency division multiplexing (SIMO OFDM) systems. The algorithm uses only the auto-covariance matrix of the received signal. It does not require precise carrier synchronization and time synchronization, so it is applicable to an environment with uncooperative high power interferences
Rendong Ying, Ruey-Wen Liu, Guozhi Xu
IEEE Signal Process. Lett.1
2004 Anti-jamming filtering in the autocorrelation domain
abstract
An anti-jamming filtering technique is presented that it fully eliminates the jamming (or noncooperative) signal and its performance is noise-independent. This technique uses multiple receivers; a filtering process explores the nonoverlapping properties of the signals in the autocorrelation domain; and a matching of the input and the output statistics. A limited simulation study supports the theory.
Ruey-Wen Liu, Rendong Ying
IEEE Signal Process. Lett.2