EDBT 2026 Demo / reviewers in the wild / expert
Zhekang Dong
dblp:162/8503
· DBLP profile ↗
33ranked-venue papers
4as first author
26since 2021 · last 2026
0000-0003-4639-3834ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Systems, architecture and hardware · 4 · 2 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CMIF-Calib: A Novel LiDAR-Camera Extrinsic Calibration System With Cross-Modal Information FusionabstractThe complexity of unmanned system operating environments necessitates multi-sensor fusion for robust, long-term perception. Precise extrinsic calibration across heterogeneous sensors is fundamental to effective fusion. While data-driven calibration methods with artificial intelligence offer high accuracy and efficiency, their generalization to unseen scenarios and new sensor configurations remains limited. To address this, we propose CMIF-Calib, a novel LiDAR-Camera extrinsic calibration framework. Our approach employs an encoder-decoder network with cascaded multi-attention modules to learn reliable cross-modal correspondences between LiDAR point clouds and camera images, enabling indirect estimation of extrinsic parameters. By decoupling sensor intrinsics from the calibration network, CMIF-Calib supports simultaneous calibration of diverse sensor types (varying camera intrinsics/LiDAR channels) and multiple sensor combinations (including multi-camera and multi-LiDAR setups). Extensive experiments across diverse datasets and platforms demonstrate that CMIF-Calib achieves higher calibration accuracy and superior generalization compared to existing methods. The implementation of CMIF-Calib will be made publicly available upon publication at https://github.com/LvXudong-HIT/CMIF-Calib. Shuo Wang 0030, Zhekang Dong, Mingyu Gao 0002, Zhiwei He 0001 |
IEEE Internet Things J. | 3 |
| 2026 | A Lightweight Dynamic Gesture Recognition Network for Advanced Driver Assistance SystemsabstractGesture-based human-machine interaction has attracted growing interest in intelligent cockpit systems due to its potential for contactless and intuitive control. However, existing methods based on traditional image processing or wearable sensors lack the flexibility and naturalness required in complex vehicle environments. Although deep learning approaches offer superior representation capabilities, they often suffer from poor real-time performance, limited recognition accuracy, and a lack of validation under real-world driving conditions. To overcome these limitations, a lightweight spatiotemporal fusion network for vehicle-mounted gesture recognition system is proposed. The net integrates a dynamic spatiotemporal fusion network that adaptively aligns inter-frame features via learnable channel-wise weights, enables the model to be more sensitive to the temporal misalignment features caused by changes in gesture speed. A dual-branch spatiotemporal attention module further enhances recognition by capturing both multi-scale spatial features and dynamic gesture trajectories. The complete system is deployed on a resource-constrained Jetson Orin Nano platform and accelerated using TensorRT. Field experiments conducted under real-world driving conditions demonstrate an average recognition accuracy exceeding 92% for commonly used gestures, with inference latency remaining below 30 ms, thereby confirming the practicality of the proposed approach for real-time in-vehicle gesture recognition. Chen Sang, Sihan Gao, Xingwang Zhang, Haixin Zhang, Zhekang Dong, Junfan Wang |
IEEE Internet Things J. | 5 |
| 2026 | Urban scene reconstruction using Geometry-aware Gaussian primitives
Xuepu Zeng, Jinlong Fan 0001, Zhekang Dong, Jing Zhang 0037, Yuxiang Yang 0001 |
Neural Networks | 3 |
| 2026 | FE-SpikeFormer: A Camera-Based Facial Expression Recognition Method for Hospital Health MonitoringabstractFacial expression recognition has emerged as a critical research area in health monitoring, enabling healthcare professionals to assess patients' emotional and psychological states for timely intervention and personalized care. However, existing methods often struggle to balance computational accuracy with energy efficiency. To address this challenge, this paper proposes FE-SpikeFormer - a high-accuracy, low-energy, and deployment-friendly Spiking Neural Network (SNN) for facial emotion recognition. The proposed architecture comprises three key components: the initial convolution module, the spiking extraction block, and the spiking integration block. These three modules collectively support detailed and contextual feature extraction, promote spatial feature integration, and strengthen the representational capacity of spiking signals. Meanwhile, a joint verification is conducted in both controlled laboratory settings and real-world hospital scenarios. Experimental results demonstrate that FE-SpikeFormer achieves top-three recognition accuracy among state-of-the-art methods, while utilizing only 6.93 million parameters. Moreover, it exhibits strong robustness against various noise conditions, underscoring its potential for practical deployment in healthcare environments. Zhekang Dong, Shiqi Zhou, Xiaoyue Ji, Chun Sing Lai, Minjiang Chen, Jiansong Ji |
IEEE J. Biomed. Health Informatics | 1 |
| 2026 | An Efficient Human Activity Recognition In-Memory Computing Architecture Development for Healthcare MonitoringabstractHuman activity recognition has played a crucial role in healthcare information systems due to the fast adoption of artificial intelligence (AI) and the internet of thing (IoT). Most of the existing methods are still limited by computational energy, transmission latency, and computing speed. To address these challenges, we develop an efficient human activity recognition in-memory computing architecture for healthcare monitoring. Specifically, a mechanism-oriented model of Ag/a-Carbon/Ag memristor is designed, serving as the core circuit component of the proposed in-memory computing system. Then, one-transistor-two-memristor (1T2M) crossbar array is proposed to perform high-efficiency multiply-accumulate (MAC) operation and high-density memory in the proposed scheme. To facilitate understanding of the proposed efficient human activity recognition in-memory computing design, self-attention ConvLSTM module, multi-head convolutional attention module, and recognition module are proposed. Furthermore, the proposed system is applied to perform human activity recognition, which contains eleven different human activities, including five different postural falls, and six basic daily activities. The experimental results show that the proposed system has advantages in recognition performance (≥ 0.20% accuracy, ≥ 1.10% F1-score) and time consumption (approximately 8∼10 times speed up) compared to existing methods, indicating an advancement in smart healthcare applications. Xiaoyue Ji, Zhekang Dong, Chenhao Hu, Chun Sing Lai |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | UAWTrack: Universal 3D Single Object Tracking in Adverse Weatherabstract3D single object tracking (3D SOT) in LiDAR point clouds is essential for autonomous driving. Most existing 3D SOT methods focus on clear weather, where point clouds are more defined. However, adverse weather conditions lead to sparser and noisier point clouds, significantly degrading tracking performance and posing safety risks. In this study, we introduce UAWTrack, a universal 3D SOT model designed to perform effectively across diverse real-world weather conditions. UAWTrack comprises three key modules: 1) Voxel Feature Extraction, which mitigates the perturbations in point clouds caused by adverse weather; 2) Motion-centric Spatial-temporal Aggregation and Motion-guided Feature Fusion, capturing motion clues and sampling dense BEV motion features to address the issue of sparsity; and 3) Weather-Specific Tracker, which efficiently handles tracking in various weather conditions. To fill the gap of lacking benchmarks for 3D SOT in adverse weather, we simulate physically valid adverse weather conditions on the KITTI and NuScenes datasets, creating two benchmarks: KITTI-A and NuScenes-A. Extensive experiments demonstrate that UAWTrack achieves state-of-the-art performance under all weather conditions. Yuxiang Yang 0001, Hongjie Gu, Yingqi Deng, Zhekang Dong, Zhiwei He 0001, Jing Zhang 0037 |
AAAI | 4 |
| 2025 | A Dual-Pathway Driver Emotion Classification Network Using Multitask Learning Strategy: A Joint VerificationabstractNegative emotion (e.g., anger, fear) may influence normal driver behavior, resulting in serious traffic accidents. Thus, developing an automatic driver emotion classification method is necessary and urgent. Most of the existing methods are performed in realistic indoor environment and always lack effective utilization of heterogeneous information, resulting in low accuracy and reliability. In this paper, a novel dual-pathway driver emotion classification network using multi-task learning strategy is proposed. To illustrate the design of the proposed driver emotion classification network, three modules are constructed: 1) visual-facial data processing module; 2) driving behavioral data processing module; 3) fusion output module. Meanwhile, considering the influence of emotional states on driving behavior, a comprehensive analysis is conducted to distinguish the positive, neutral, and negative influence on driving behavior. Furthermore, a joint verification in both realistic indoor environment (i.e., laboratory simulation on the PPB-Emo dataset) and real-world outdoor scenario is performed. The experimental results illustrate that the proposed network exhibits superior performance in terms of classification accuracy and response time, achieving good balance between classification accuracy and running speed in internet of things scenarios. Zhekang Dong, Chenhao Hu, Xiaoyue Ji, Chun Sing Lai |
IEEE Internet Things J. | 1 |
| 2025 | CDRP3: Cascade Deep Reinforcement Learning for Urban Driving Safety With Joint Perception, Prediction, and PlanningabstractSafe urban driving is challenging due to the high density of traffic flow and various potential hazards, such as the sudden appearance of unknown objects. Traditional rule-based approaches and imitation learning methods struggle to address the diverse driving scenarios encountered in urban environments. Reinforcement learning (RL), which adapts to a wide range of driving scenarios through continuous interaction with the environment, has demonstrated success in autonomous driving. Making safety decisions when driving in urban environments necessitates a comprehensive perception of the current scene and the ability to predict the evolution of the dynamic scene. In this paper, we present a novel cascade deep reinforcement learning framework, CDRP3, designed to enhance the safety decision-making capabilities of self-driving vehicles in complex scenarios and emergencies. We leverage a multi-modal spatio-temporal perception (MmSTP) module to fuse multi-modal sensor data and introduce temporal perception to capture spatio-temporal information about dynamic driving environments, and a future state prediction (FSP) module to model complex interactions between different traffic participants and explicitly predict their future states. Subsequently, in the PPO-based planning module, we use the comprehensive environmental information obtained from perception and prediction to decode an optimized driving strategy using a lateral and longitudinal separated multi-branch network structure guided by a customized reward function. This approach enables knowledge transfer from the perception and prediction components to planning, and planning-oriented enhancement of safety decision-making capabilities to improve driving safety. Our experiments demonstrate that CDRP3 outperforms state-of-the-art methods, providing superior driving safety in complex urban environments. Yuxiang Yang 0001, Fenglong Ge, Jinlong Fan 0001, Jufeng Zhao, Zhekang Dong |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Time-Frequency Hybrid Neuromorphic Computing Architecture Development for Battery State-of-Health EstimationabstractWith the rapid adoption of Internet of Things (IoT) and artificial intelligence (AI), lithium-ion battery state-of-health (SOH) estimation plays an important role in guaranteeing the secure and stable functioning of various domains. However, the majority of the existing methods are constrained by factors, such as transmission latency, computational energy, and computing speed. To address these challenges, we develop a time-frequency hybrid neuromorphic computing architecture for battery SOH estimation. Specifically, an eco-friendly, biodegradable memristor crossbar array is designed, enabling high-energy efficiency and high-performance density in the proposed system. To improve the understanding of the designed time-frequency hybrid neuromorphic computing system, a local information extraction module, a time-frequency feature fusion module, and a global information perception module are proposed. Furthermore, the proposed system is validated on two publicly available battery ageing data sets (i.e., the CALCE-CS2 data set and the National Aeronautics and Space Administration data set). The experimental results show that the system exhibits superior performance to that of the state-of-the-art (SOTA) methods in terms of estimation accuracy (highest estimation accuracy), time consumption (approximately 8–12 times faster), and transmission latency (approximately 10 times faster). This study is expected to promote the advancement and evolution of next-generation computing systems, enabling the realization of low-power consumption and high-density information processing in IoT scenarios. Xiaoyue Ji, Junfan Wang, Guangdong Zhou, Chun Sing Lai, Zhekang Dong |
IEEE Internet Things J. | 6 |
| 2024 | MSF-SLAM: Multi-Sensor-Fusion-Based Simultaneous Localization and Mapping for Complex Dynamic EnvironmentsabstractWe proposed a multi-sensor fusion-based localization and scene reconstruction method for a complex dynamic scene. The multi-level fusion between multiple sensors was implemented by fusing data collected from different sensors in different system modules. In the front-end of the system, the camera and the LiDAR assisted each other. The LiDAR point clouds provided 3D information for the feature points in the image. The moving objects elimination method based on the image can remove the points on the moving objects in the LiDAR point clouds for localization accuracy improvement and static 3D scene reconstruction. To further improve the localization accuracy, a combination of visual loop closure detection and LiDAR loop closure detection was utilized to ensure the global consistency of scene reconstruction. At the system’s back-end, the observation model of different sensors was integrated to construct a multiple constraint factor graph with nonlinear optimization to obtain the optimal system states. Experimental results demonstrated that the proposed multi-sensor fusion-based localization and scene reconstruction algorithm could operate robustly in multiple complex dynamic scenes. Zhiwei He 0001, Yuxiang Yang 0001, Jiahao Nie 0001, Zhekang Dong, Shuo Wang 0030, Mingyu Gao 0002 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | SpikeTOD: A Biologically Interpretable Spike-Driven Object Detection in Challenging Traffic ScenariosabstractArtificial neural networks (ANN) have shown remarkable performance in intelligent transportation systems (ITS), especially for the traffic object detection. However, as the ITS is applied to a wider range of traffic scenarios, the increasing demand for the trade-off between detection performance and power resources has become inevitable. A biologically interpretable spike-driven traffic object detector for challenging scenarios is proposed in this paper, named SpikeTOD, achieving the trade-off between the accuracy and power consumption. Firstly, the spike neural network (SNN) is employed to realize energy-efficient object detection in traffic scenarios. And a local modulation-based integrate-and-fire (IF) neuron is designed, which provides an efficient way to convert the traffic detection model from ANN to SNN. Secondly, a biology-inspired detail-guided context-aware network (DCNet) is proposed to improve the detection performance. The integration of detail coherence and global priors is leveraged to selectively emphasize object features and improve the detection capabilities within challenging conditions. As far as we know, this is the first application of SNN in traffic object detection tasks. SpikeTOD achieved a mAP@50 of 46.11% on the BDD100K dataset with a power consumption of 4.73E-03J, demonstrating a more efficient trade-off in detection accuracy and power consumption. Notably, SpikeTOD maintained an average missed detection rate of 44.56%, further contributing to its overall efficacy in traffic object detection. Further, we conducted on road test by deploying SpikeTOD on Jetson Xavier NX and Loihi to demonstrate that model achieves a better balance between accuracy and power consumption. Junfan Wang, Xiaoyue Ji, Zhekang Dong, Mingyu Gao 0002, Zhiwei He 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Vehicle-Mounted Adaptive Traffic Sign Detector for Small-Sized Signs in Multiple Working ConditionsabstractTraffic sign detection is of great significance to the development of the Intelligent Transportation System (ITS) as a database for environmental awareness. The main challenges of existing traffic sign detection method are inaccurate small object detection, difficult mobile deployment, and complex working environment. Based on these, a vehicle-mounted adaptive traffic sign detector (VATSD) for small-sized signs in multiple working conditions is proposed in this paper. First, the Backbone of the detector is optimized. A feature tight fusion structure is designed to constitute a new feature extraction module, DCSP, which improves the feature extraction capability and the detection accuracy of small objects with negligible additional parameters. Second, an image enhancement network IENet with an adaptive joint filtering strategy is proposed. The IENet enables the dynamic selection of filters and thus adaptively optimizes low-quality images under multiple conditions to improve the accuracy of subsequent detection tasks. The proposed method has experimented on three traffic sign datasets and the detection accuracy increased by up to 7.6% compared to the original. The proposed detector demonstrates superiority over other state-of-the-art (SOTA) methods in terms of small object detection accuracy, detection speed, and environmental adaptability. Further, we deployed VATSD to Jetson Xavier NX and achieved a detection speed of 21.6 FPS, meeting real-time requirements. Junfan Wang, Xiaoyue Ji, Zhekang Dong, Mingyu Gao 0002, Chun Sing Lai |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Metaverse Meets Intelligent Transportation System: An Efficient and Instructional Visual Perception FrameworkabstractThe combination of the Metaverse and intelligent transportation systems (ITS) holds significant developmental promise, especially for visual perception tasks. However, the acquisition of high-quality scene data poses a challenging and expensive endeavor. Meanwhile, the visual disparity between the Metaverse and the physical world poses an impact on the practical applicability of the visual perception tasks. In this paper, a Metaverse Intelligent Traffic Visual Framework, MITVF, is developed to guide the implementation of visual perception tasks in the physical world. Firstly, a two-stage metadata optimization strategy is proposed that can efficiently provide diverse and high-quality scene data for traffic perception models. Specifically, an element reconfigurability strategy is proposed to flexibly combine dynamic and static traffic elements to enrich the data with a low cost. A diffusion model-based metadata optimization acceleration strategy is proposed to achieve efficient improvement of image resolution. Secondly, a Meta-Physical adaptive learning method is proposed, and further applied to visual perception tasks to compensate for the visual disparity between the Metaverse and the physical world. Experimental results show that MITVF achieves a 10$\times$acceleration in optimization speed, ensuring the image quality and reconstructing diverse. Further, MITVF is applied to the traffic object detection task to verify the effectiveness and validity. The performance of the model trained with 5k real data exceeded that of the model trained with 200k real data, with AP$_{50}$reaching 67.7%. Junfan Wang, Xiaoyue Ji, Zhekang Dong, Mingyu Gao 0002, Chun Sing Lai |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | A lightweight vehicle mounted multi-scale traffic sign detector using attention fusion pyramid
Junfan Wang, Yeting Gu, Yunfeng Yan, Mingyu Gao 0002, Zhekang Dong |
J. Supercomput. | 7 |
| 2024 | MLG-NCS: Multimodal Local-Global Neuromorphic Computing System for Affective Video Content AnalysisabstractDespite neuromorphic computing (NC) technologies offer tremendous potential in executing computationally intensive tasks with high efficiency and low latency, most of existing methods are still difficult to achieve software-comparable accuracy. To address this challenge, we develop a multimodal local–global NC system (MLG-NCS) that can capture local characteristics and exchange global cross-modal information sufficiently. Specifically, a high-density memristor crossbar array is prepared to perform efficient parallel in-memory operations, serving as the fundamental component of the proposed MLG-NCS. To facilitate understanding of the proposed MLG-NCS design, the local feature representation module, the global cross-modal interaction module, and the output module are designed. The experimental results show that the proposed system has advantages in classification accuracy (ranked top three), time consumption (approximately ten times speed up), and latency (about 1.2–15.3 times faster), enabling good inter-related tradeoffs between latency, efficiency, and accuracy. This study is expected to promote the revolution and development of next-generation computing system, which takes a firm step toward artificial general intelligence (AGI). Xiaoyue Ji, Zhekang Dong, Guangdong Zhou, Chun Sing Lai, Donglian Qi |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | Implementation of Multiple-Step Quantized STDP Based on Novel Memristive SynapsesabstractMemristors have been widely studied as artificial synapses in neuromorphic circuits, due to their functional similarity with biological synapses, low operating power, and high integration density. Currently, the synaptic weight symbolic limitation and weight update inaccuracy are two challenging issues to be solved. In this work, a novel memristive synapse and a matched mixed-signal neuron circuit are designed to implement robust yet accurate spike-timing-dependent plasticity learning in excitatory and inhibitory synapses. To break through the weight symbolic limitation, a four memristors and two resistors (4M2R) synapse composed of 4M2R for spiking neural network (SNN) is designed. The proposed synapse can be either excitatory or inhibitory (E/I) by rationally arranging the resistors in the circuit, and it is the first of its kind, enabling Hebbian and anti-Hebbian training without additional adjusting of neural signals. In addition, the high symmetricity, linearity, and stability against device variation of the 4M2R synapse can also greatly improve the weight update accuracy. To further address the inaccurate weight update issue caused by signal complexity, a neuron circuit is designed to generate square-wave pulses for spike transmission and synaptic weight modulation. Simulations are carried out in the MATLAB Simscape as well as Virtuoso using SMIC 0.18$\mu $m process and a specially developed memristor model for SNN synapse simulation. The simulating results show good agreement with the weight change derived from the algorithmic methods, and the influence of weak signal-induced weight variation on circuit performance can be rigorously assessed. Yi-Fan Liu, Dawei Wang 0003, Zhekang Dong, Wen-Sheng Zhao |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | Efficient Path-Space Differentiable Volume Rendering With Respect To ShapesabstractAbstract Differentiable rendering of translucent objects with respect to their shapes has been a long‐standing problem. State‐of‐the‐art methods require detecting object silhouettes or specifying change rates inside translucent objects—both of which can be expensive for translucent objects with complex shapes. In this paper, we address this problem for translucent objects with no refractive or reflective boundaries. By reparameterizing interior components of differential path integrals, our new formulation does not require change rates to be specified in the interior of objects. Further, we introduce new Monte Carlo estimators based on this formulation that do not require explicit detection of object silhouettes. Olivier Maury, Christophe Hery, Zhekang Dong, S. Zhao |
Comput. Graph. Forum | 5 |
| 2023 | FAML-RT: Feature alignment-based multi-level similarity metric learning network for a two-stage robust tracker
Jiahao Nie 0001, Zhekang Dong, Zhiwei He 0001, Mingyu Gao 0002 |
Inf. Sci. | 2 |
| 2023 | SABV-Depth: A biologically inspired deep learning network for monocular depth estimationabstractMonocular depth estimation makes it possible for machines to perceive the real world. The prediction performance of the depth estimation network based on deep learning will be affected due to the depth of the deep network and the locality of convolution operations. The imitation of the biological visual system and its functional structure is becoming a research hotspot. In this paper, we study the interpretability relationship between the biological visual system and the monocular depth estimation network. By concretizing the attention mechanism in biological vision, we propose a monocular depth estimation network based on the self-attention mechanism, named SABV-Depth, which can improve prediction accuracy. Inspired by the biological visual interaction mechanism, we focus on the information transfer between each module of the network and improve the information retention ability, and enable the network to output a depth map with rich object information and detailed information. Further, a decoder module with an inner-connection is proposed to recover depth maps with sharp edge contours. Our method is experimentally validated on the KITTI dataset and NYU Depth V2 dataset. The results show that compared with other works, the proposed method improves prediction accuracy. Meanwhile, the depth map has more object information and detail information, and a better edge information processing effect. Junfan Wang, Zhekang Dong, Mingyu Gao 0002, Huipin Lin, Qiheng Miao |
Knowl. Based Syst. | 3 |
| 2023 | Improved YOLOv5 network for real-time multi-scale traffic sign detection
Junfan Wang, Zhekang Dong, Mingyu Gao 0002 |
Neural Comput. Appl. | 3 |
| 2023 | A Brain-Inspired Hierarchical Interactive In-Memory Computing System and Its Application in Video Sentiment AnalysisabstractVideo sentiment analysis can effectively establish the relationship between the emotion state and the multimodal information, while still suffer from intensive computation and low efficiency, due to the von Neumann computing architecture. Here, we present a brain-inspired hierarchical interactive in-memory computing (IMC) system, which can efficiently solve ‘von Neumann bottleneck’, enabling cross-modal interactions and semantic gap elimination. First, a 1T1M synapse array is fabricated using cost-effective, highly stable, flexible, and eco-friendly carbon materials, offering efficient analog multiply-accumulate operations. To illustrate the complexity of the proposed brain-inspired hierarchical interactive IMC system, three modules are proposed: 1) unimodal extraction module, 2) hierarchical interactive module, 3) output module. Furthermore, the proposed system is validated by applying it to video sentiment analysis. The experimental results demonstrate that the proposed system outperforms the existing state-of-the-art methods with high computational efficiency and good robustness. This work opens up a new way to achieve the deep integration of nanomaterials, deep learning, and modern electronics into IMC. Xiaoyue Ji, Zhekang Dong, Yifeng Han, Chun Sing Lai, Donglian Qi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Learning Localization-Aware Target Confidence for Siamese Visual TrackingabstractSiamese tracking paradigm has achieved great success, providing effective appearance discrimination and size estimation by classification and regression. While such a paradigm typically optimizes the classification and regression independently, leading to task misalignment (accurate prediction boxes have no high target confidence scores). In this paper, to alleviate this misalignment, we propose a novel tracking paradigm, called SiamLA. Within this paradigm, a series of simple, yet effective localization-aware components are introduced to generate localization-aware target confidence scores. Specifically, with the proposedlocalization-aware dynamic label(LADL) loss andlocalization-aware label smoothing(LALS) strategy, collaborative optimization between the classification and regression is achieved, enabling classification scores to be aware of location state, not just appearance similarity. Besides, we propose a separatelocalization-aware quality prediction(LAQP) branch to produce location quality scores to further modify the classification scores. To guide a more reliable modification, a novellocalization-aware feature aggregation(LAFA) module is designed and embedded into this branch. Consequently, the resulting target confidence scores are more discriminative for the location state, allowing accurate prediction boxes tend to be predicted as high scores. Extensive experiments are conducted on six challenging benchmarks, including GOT-10 k, TrackingNet, LaSOT, TNL2K, OTB100 and VOT2018. Our SiamLA achieves competitive performance in terms of both accuracy and efficiency. Furthermore, a stability analysis reveals that our tracking paradigm is relatively stable, implying that the paradigm is potential for real-world applications. Jiahao Nie 0001, Zhiwei He 0001, Yuxiang Yang 0001, Mingyu Gao 0002, Zhekang Dong |
IEEE Trans. Multim. | 5 |
| 2022 | Spreading Fine-Grained Prior Knowledge for Accurate TrackingabstractWith the widespread use of deep learning in single object tracking task, mainstream tracking algorithms treat tracking as a combined classification and regression problem. Classification aims at locating an arbitrary target, and regression aims at estimating the corresponding bounding box. In this paper, we focus on regression and propose a novel box estimation network, which consists of a transformer encoder target pyramid guide (TPG) and transformer decoder target pyramid spread (TPS). Specifically, the transformer encoder TPG is designed to generate fine-grained prior knowledge with explicit representation for template targets. In contrast to the raw transformer encoder, we capture the visual dependence through local-global self-attention and deem the multi-scale target regions as the “local” region. Using this fine-grained prior knowledge, we design the transformer decoder TPS to spread it to the subsequent search regions with high affinity to accurately estimate the bounding boxes. Considering that self-attention fails to model information interaction across channels between the template target and search regions, we develop a channel-wise cross-attention block within the TPS as compensation. Extensive experiments on the OTB100, UAV123, NFS, VOT2020, VOT2021, LaSOT, LaSOT_ext, TrackingNet and GOT-10k benchmarks show that the proposed box estimation network outperforms most existing box estimation methods. Furthermore, our trackers based on this estimation network exhibit a competitive performance against state-of-the-art trackers. Jiahao Nie 0001, Zhiwei He 0001, Mingyu Gao 0002, Zhekang Dong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | A novel kernelized correlation filter by fusing multiple feature response maps, enhanced target re-detection, and improved model updating for visual tracking
Chenjie Du, Zhongping Ji, Zhekang Dong, Mingyu Gao 0002, Zhiwei He 0001 |
Vis. Comput. | 3 |
| 2021 | Reconfigurable logic circuit design for stateful Boolean logic computing
Zhekang Dong, Lidan Wang 0001, Shukai Duan 0001 |
Sci. China Inf. Sci. | 2 |
| 2021 | Neuromorphic extreme learning machines with bimodal memristive synapses
Zhekang Dong, Chun Sing Lai, Donglian Qi, Mingyu Gao 0002, Shukai Duan 0001 |
Neurocomputing | 1 |
| 2020 | Deep Learning Based Strategy for Eye-to-Hand Robotic Tracking and Grabbing
Junwen Zhong, Weijun Sun, Qinyu Cai, Zhekang Dong, Mingyu Gao 0002 |
ICONIP (2) | 5 |
| 2020 | A novel versatile window function for memristor model with application in spiking neural network
Junrui Li, Zhekang Dong, Shukai Duan 0001, Lidan Wang 0001 |
Neurocomputing | 2 |
| 2019 | An Automatic Detection and Sorting System for Valve Core Based on Machine VisionabstractThe air-conditioning energy consumption will indirectly be affected by the machining accuracy of the valve core in throttle valve used to control refrigerant, so each valve core needs to be tested for its intermediate aperture before leaving the factory. Moreover, the repeatability of the general electronic pneumatic measuring instrument cannot meet the requirements. This paper introduces an automatic detection system of the valve core aperture based on machine vision, which is mainly composed of mechanical movement part, visual part, and control part. The system is based on a method of measuring the valve core aperture by pneumatic measuring instrument, and automatically discriminates the values of pneumatic measuring instrument and barometer by machine vision. Then, the system cooperates with mechanical movement part to complete valve core loading and automatic sorting, thus realizing online automatic detection and sorting of the micro-apertures of the air-conditioning valve core. The experimental results show that the online detection system meets the production requirements of enterprises. Mingyu Gao 0002, Zhiping Zhan, Yuxiang Yang 0001, Zouchao Deng, Jiye Huang, Zhekang Dong |
IECON | 6 |
| 2018 | A Novel Convolution Computing Paradigm Based on NOR Flash Array with High Computing Speed and Energy EfficientabstractA novel convolution computing paradigm based on the NOR Flash Array is proposed. Significant improvements both in computing speed and energy consumption are achieved compared to CMOS-based logic computing paradigms. Regarding to the feature extraction task from a 256×256 image, the computing speed of 3.9×104frame per second (fps) and the energy consumption of 0.057nJ/pixel are achieved using the proposed computing paradigm. Runze Han, Peng Huang 0004, Yachen Xiang, Chen Liu 0009, Zhekang Dong, Zhiqiang Su, Yongbo Liu, Jinfeng Kang |
ISCAS | 5 |
| 2018 | A general memristor-based pulse coupled neural network with variable linking coefficient for multi-focus image fusion
Zhekang Dong, Chun Sing Lai, Donglian Qi, Zhao Xu 0002, Chaoyong Li, Shukai Duan 0001 |
Neurocomputing | 1 |
| 2016 | Small-world Hopfield neural networks with weight salience priority and memristor synapses for digit recognition
Shukai Duan 0001, Zhekang Dong, Lidan Wang 0001, Hai Li 0001 |
Neural Comput. Appl. | 2 |
| 2015 | Memristor-Based Cellular Nonlinear/Neural Network: Design, Analysis, and ApplicationsabstractCellular nonlinear/neural network (CNN) has been recognized as a powerful massively parallel architecture capable of solving complex engineering problems by performing trillions of analog operations per second. The memristor was theoretically predicted in the late seventies, but it garnered nascent research interest due to the recent much-acclaimed discovery of nanocrossbar memories by engineers at the Hewlett-Packard Laboratory. The memristor is expected to be co-integrated with nanoscale CMOS technology to revolutionize conventional von Neumann as well as neuromorphic computing. In this paper, a compact CNN model based on memristors is presented along with its performance analysis and applications. In the new CNN design, the memristor bridge circuit acts as the synaptic circuit element and substitutes the complex multiplication circuit used in traditional CNN architectures. In addition, the negative differential resistance and nonlinear current-voltage characteristics of the memristor have been leveraged to replace the linear resistor in conventional CNNs. The proposed CNN design has several merits, for example, high density, nonvolatility, and programmability of synaptic weights. The proposed memristor-based CNN design operations for implementing several image processing functions are illustrated through simulation and contrasted with conventional CNNs. Monte-Carlo simulation has been used to demonstrate the behavior of the proposed CNN due to the variations in memristor synaptic weights. Shukai Duan 0001, Zhekang Dong, Lidan Wang 0001, Pinaki Mazumder |
IEEE Trans. Neural Networks Learn. Syst. | 3 |