EDBT 2026 Demo / reviewers in the wild / expert
Dalin Zhou
dblp:167/1322
· DBLP profile ↗
14ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0003-2363-9125ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Scale Video Demoiréing Based on Selective Time-Domain FusionabstractVideo demoiréing aims to remove Moiré patterns from video footage captured by cameras when recording electronic displays. Considering the limitation that existing methods often fail to effectively exploit temporal information across video frames, thereby hindering their ability to recover details and maintain temporal consistency, this letter proposed a novel Multi-scale Video Demoiréing network (MVDS) based on Selective Temporal Fusion. By employing deformable convolutions for frame alignment and a selective temporal fusion mechanism, MVDS efficiently exploits temporal information to improve demoiréing performance. An embedded multi-scale U-net with attention fusion further refines image details at different scales. Extensive experiments on the VDMoiré dataset demonstrate that MVDS outperforms state-of-the-art methods in terms of both qualitative and quantitative metrics, effectively removing Moiré patterns while preserving image details and temporal consistency. Yinfeng Fang, Xiaohao Pan, Yuxi Wang 0002, Dalin Zhou, Zhaojie Ju |
IEEE Signal Process. Lett. | 4 |
| 2025 | Video Object Detection Considering Dynamic Neighborhood Feature MultiplexingabstractVideo object detection is essential for human-interaction applications, including bimanual manipulation sensing (BMS). The effects of video detection in practical applications still need to be improved, as they are restricted by long-range spatiotemporal dependency analysis. How do humans sense bimanual manipulation in videos, especially for deteriorated clips? We argue that humans analyze the current clips based on earlier memory, namely, long-term spatial and temporal dependencies (LTSTD). However, most existing methods have yet to report significant results, as the limited exploration of these dependencies limits them. Developing an easy-to-integrate module is generally preferred for future applications rather than designing a complex end-to-end framework. Therefore, we propose a dynamic neighborhood feature multiplexing mechanism for online video object detection in this article, which is better at learning LTSTD in flexible and robust ways, boosting existing detection results, called DNFM. Specifically, we develop dynamic memory enhancement neural networks for better long-term feature aggregation with negligible additional computation costs. We multiplex each frame feature to aggregate key enhanced representations under the guidance of dynamic memory recall. The DNFM contributes to various famous detectors in BMS and other challenging detection tasks, and particular attention has been devoted to “low-quality” frame detection. Experimental results show that, while achieving state-of-the-art detection performance, DNFM clearly illustrates the easy-to-integrate operation for boosting the video object detection results. Xuna Wang, Dalin Zhou, Yingke Xu, Zhaojie Ju |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2022 | Modelling EMG driven wrist movements using a bio-inspired neural network
Yinfeng Fang, Jiani Yang, Dalin Zhou, Zhaojie Ju |
Neurocomputing | 3 |
| 2022 | Deep Temporal Model-Based Identity-Aware Hand Detection for Space Human-Robot InteractionabstractHand detection is a crucial technology for space human-robot interaction (SHRI), and the awareness of hand identities is particularly critical. However, most advanced works have three limitations: 1) the low detection accuracy of small-size objects; 2) insufficient temporal feature modeling between frames in videos; and 3) the inability of real-time detection. In the article, a temporal detector (called TA-RSSD) is proposed based on the SSD and spatiotemporal long short-term memory (ST-LSTM) for real-time detection in SHRI applications. Next, based on the online tubelet analysis, a real-time identity-awareness module is designed for multiple hand object identification. Several notable properties are described as follows: 1) the hybrid structure of the Resnet-101 and the SSD improves the detection accuracy of small objects; 2) three-level feature pyramidal structure retains rich semantic information without losing detailed information; 3) a group of the redesigned temporal attentional LSTM (TA-LSTM) is utilized for three-level feature map modeling, which effectively achieves background suppression and scale suppression; 4) low-level attention maps are used to eliminate in-class similarity between hand objects, which improves the accuracy of identity awareness; and 5) a novel association training scheme enhances the temporal coherence between frames. The proposed model is evaluated on the SHRI-VID dataset (collected according to the task requirements), the AU-AIR dataset, and the ImageNet-VID benchmark. Extensive ablation studies and comparisons on detection and identity-awareness capacities show the superiority of the proposed model. Finally, a set of actual testing is conducted on a space robot, and the results show that the proposed model achieves a real-time speed and high accuracy. Hongwei Gao 0002, Dalin Zhou, Jinguo Liu, Qing Gao 0002, Zhaojie Ju |
IEEE Trans. Cybern. | 3 |
| 2022 | Deep Object Detector With Attentional Spatiotemporal LSTM for Space Human-Robot InteractionabstractGlobal temporal information and local semantic information are essential cues for high-performance online object detection in videos. However, despite their promising detection accuracy in most cases, most state-of-the-art approaches have following two limitations: invalid background/scale suppression and inadequate temporal information mining between frames. Many jobs currently focus on temporal information learning based on a single frame. In this article, we propose an attentional global–local information learning network; this is one of the first attempts to fully use both types of information between frames. Attention maps are creatively utilized to transfer temporal contexts between frames. This also effectively alleviates the adverse effects of scale changes. Furthermore, empowered by a detailed framework, a proposed detector effectively uses multilevel feature extraction. Given these contributions, the proposed detector achieves state-of-the-art performance on challenging benchmarks. Finally, practical experiments are conducted on a space human–robot interaction platform. Hongwei Gao 0002, Yongquan Chen, Dalin Zhou, Jinguo Liu, Zhaojie Ju |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2021 | A Novel Curved Gaussian Mixture Model and Its Application in Motion Skill EncodingabstractThe purpose of this paper is to present a novel curved Gaussian Mixture Model (CGMM) and to study the application of it in motion skill encoding. Primarily, Gaussian mixture model (GMM) has been widely applied on many occasions when a probability density function is needed to approximate a complex probability distribution. However, GMM cannot efficiently approach highly non-linear distributions. Thus, the proposed novel CGMM, as a weighted mixture of curved Gaussian models (CGM), is structured with non-linear transfers, which reshapes the flat GMM into a geo-metrically curved one. As a consequence, CGMM has more freedoms and flexibilities than the flat GMM so a CGMM requires fewer number of components in fitting highly non-linear motion trajectories. Moreover, we derive a dedicated iterative parameter estimation algorithm for the CGMM based on maximum likelihood estimation (MLE) theory. To evaluate the performance of the CGMM and its parameter estimation algorithm, a series of quantitative experiments are carried out. We first test the model performance in the data fitting task with the generated synthetic data. Then a motion skill encoding test is carried out on a human motion trajectory dataset built by a Virtual Reality (VR) based motion tracking system. The empirical results support that CGMM outperforms state-of-the-arts in the model performance test. Meanwhile, CGMM has a significant improvement in encoding high dimensional non-linear trajectory data compared to the GMM in motion skill encoding test with its dedicated parameter estimation algorithm. Disi Chen, Gongfa Li, Dalin Zhou, Zhaojie Ju |
IROS | 3 |
| 2021 | Gesture recognition based on multi-modal feature weightabstractSummary With the continuous development of sensor technology, the acquisition cost of RGB‐D images is getting lower and lower, and gesture recognition based on depth images and Red‐Green‐Blue (RGB) images has gradually become a research direction in the field of pattern recognition. However, most of the current processing methods for RGB‐D gesture images are relatively simple, ignoring the relationship and influence between its two modes, and unable to make full use of the correlation factors between different modes. In view of the above problems, this paper optimizes the effect of RGB‐D information processing by considering the independent features and related features of multi‐modal data to construct a weight adaptive algorithm to fuse different features. Simulation experiments show that the method proposed in this paper is better than the traditional RGB‐D gesture image processing method and the gesture recognition rate is higher. Comparing the current more advanced gesture recognition methods, the method proposed in this paper also achieves higher recognition accuracy, which verifies the feasibility and robustness of this method. Haojie Duan, Ying Sun 0004, Du Jiang, Juntong Yun, Ying Liu 0087, Dalin Zhou |
Concurr. Comput. Pract. Exp. | 8 |
| 2021 | Occlusion gesture recognition based on improved SSDabstractSummary Gesture recognition has always been a research hotspot in the field of human‐computer interaction. Its purpose is to realize the natural interaction with the machine by recognizing the semantics expressed by gesture. In the process of gesture recognition, the occlusion of gesture is an inevitable problem. In the process of gesture recognition, some or even all of the gesture features will be lost due to the occlusion of the gesture, resulting in the wrong recognition or even unrecognizability of the gesture. Therefore, it is of great significance to study gesture recognition under occlusion. The single shot multibox detector (SSD) algorithm is analyzed, and the front‐end network is compared. Mobilenets is selected as the front‐end network, and the Mobilenets‐SSD network is improved. In tensorflow environment, based on the improved network model, the self‐occlusion gesture and object occluding gesture are trained in color map, depth map, and color and depth fusion respectively. The recognition models of self‐occlusion gestures and object‐occlusion gestures in color map, depth map, and color and depth fusion are obtained. And compare and analyze the learning rate, loss function, and average accuracy of various models obtained for occlusion gesture recognition. Shangchun Liao, Gongfa Li, Hao Wu 0030, Du Jiang, Ying Liu 0087, Juntong Yun, Dalin Zhou |
Concurr. Comput. Pract. Exp. | 8 |
| 2021 | Enhancement of real-time grasp detection by cascaded deep convolutional neural networksabstractAbstract Robot grasping technology is a hot spot in robotics research. In relatively fixed industrialized scenarios, using robots to perform grabbing tasks is efficient and lasts a long time. However, in an unstructured environment, the items are diverse, the placement posture is random, and multiple objects are stacked and occluded each other, which makes it difficult for the robot to recognize the target when it is grasped and the grasp method is complicated. Therefore, we propose an accurate, real‐time robot grasp detection method based on convolutional neural networks. A cascaded two‐stage convolutional neural network model with course to fine position and attitude was established. The R‐FCN model was used as the extraction of the candidate frame of the picking position for screening and rough angle estimation, and aiming at the insufficient accuracy of the previous methods in pose detection, an Angle‐Net model is proposed to finely estimate the picking angle. Tests on the Cornell dataset and online robot experiment results show that the method can quickly calculate the optimal gripping point and posture for irregular objects with arbitrary poses and different shapes. The accuracy and real‐time performance of the detection have been improved compared to previous methods. Yaoqing Weng, Ying Sun 0004, Du Jiang, Bo Tao 0002, Ying Liu 0087, Juntong Yun, Dalin Zhou |
Concurr. Comput. Pract. Exp. | 7 |
| 2021 | Attribute-Driven Granular Model for EMG-Based Pinch and Fingertip Force Grand RecognitionabstractFine multifunctional prosthetic hand manipulation requires precise control on the pinch-type and the corresponding force, and it is a challenge to decode both aspects from myoelectric signals. This paper proposes an attribute-driven granular model (AGrM) under a machine-learning scheme to solve this problem. The model utilizes the additionally captured attribute as the latent variable for a supervised granulation procedure. It was fulfilled for EMG-based pinch-type classification and the fingertip force grand prediction. In the experiments, 16 channels of surface electromyographic signals (i.e., main attribute) and continuous fingertip force (i.e., subattribute) were simultaneously collected while subjects performing eight types of hand pinches. The use of AGrM improved the pinch-type recognition accuracy to around 97.2% by 1.8% when constructing eight granules for each grasping type and received more than 90% force grand prediction accuracy at any granular level greater than six. Further, sensitivity analysis verified its robustness with respect to different channel combination and interferences. In comparison with other clustering-based granulation methods, AGrM achieved comparable pinch recognition accuracy but was of lowest computational cost and highest force grand prediction accuracy. Yinfeng Fang, Dalin Zhou, Kairu Li, Zhaojie Ju, Honghai Liu 0001 |
IEEE Trans. Cybern. | 2 |
| 2020 | Attention mechanism-based CNN for facial expression recognition
Jing Li 0027, Kan Jin, Dalin Zhou, Naoyuki Kubota, Zhaojie Ju |
Neurocomputing | 3 |
| 2019 | Towards Zero Re-Training for Long-Term Hand Gesture Recognition via Ultrasound SensingabstractWhile myoelectric pattern recognition is a prevailing way for gesture recognition, the inherent nonstationarity of electromyography signals hinders its long-term application. This study aims to prove a hypothesis that morphological information of muscle contraction detected by ultrasound image is potentially suitable for long-term use. A set of ultrasound-based algorithms are proposed to realize robust hand gesture recognition over multiple days, with user training only at the first day. A markerless calibration algorithm is first presented to position the ultrasound probe during donning and doffing; an algorithm combining speeded-up robust features and bag-of-features model being immune to ultrasound probe shift and rotation is then introduced; a self-enhancing classification method is next adopted to update classification model automatically by incorporating useful knowledge from testing data; finally the performance of long-term hand gesture recognition with zero re-training is validated by a six-day experiment of six healthy subjects, whose outcomes strongly support the hypothesis with about 94% of gesture recognition accuracy for each testing day. This study confirms the feasibility of adoption of ultrasound sensing for long-term musculature related applications. Xingchen Yang, Dalin Zhou, Yu Zhou 0013, Youjia Huang, Honghai Liu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | Ultrasound-Based Sensing Models for Finger Motion ClassificationabstractMotions of the fingers are complex since hand grasping and manipulation are conducted by spatial and temporal coordination of forearm muscles and tendons. The dominant methods based on surface electromyography (sEMG) could not offer satisfactory solutions for finger motion classification due to its inherent nature of measuring the electrical activity of motor units at the skin's surface. In order to recognize morphological changes of forearm muscles for accurate hand motion prediction, ultrasound imaging is employed to investigate the feasibility of detecting mechanical deformation of deep muscle compartments in potential clinical applications. In this study, finger motion classification has been represented as subproblems: recognizing the discrete finger motions and predicting the continuous finger angles. Predefined 14 finger motions are presented in both sEMG signals and ultrasound images and captured simultaneously. Linear discriminant analysis classifier shows the ultrasound has better average accuracy (95.88%) than the sEMG (90.14%). On the other hand, the study of predicting the metacarpophalangeal (MCP) joint angle of each finger in nonperiod movements also confirms that classification method based on ultrasound achieves better results (average correlation 0.89 $\pm$ 0.07 and NRMSE 0.15 $\pm$ 0.05) than sEMG (0.81 $\pm$ 0.09 and 0.19 $\pm$ 0.05). The research outcomes evidently demonstrate that the ultrasound can be a feasible solution for muscle-driven machine interface, such as accurate finger motion control of prostheses and wearable robotic devices. Youjia Huang, Xingchen Yang, Dalin Zhou, Keshi He, Honghai Liu 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2017 | A force-driven granular model for EMG based grasp recognitionabstractIt is a challenge to precisely predict hand grasps based on EMG signals given practical scenarios, due to its inherent nature. This paper proposes a solution to tackle the challenge with a force-driven granular model (FDGM). The problem of n-class hand grasp classification has been represented as force-based granular modelling, in which a number of granules are constructed for each class relying on the synchronically captured grasping force. A rule based mechanism is formed for granule generation of each class, and a cross-testing algorithm is proposed to optimise the number of granules. The experiment based on 8-case grasp recognition reveals that the proposed method performs better in terms of motion recognition accuracy of multiple EMG channel combination, and is more insensitive to signal interferences. In comparison with other rules of information granulation, it is confirmed that the force-driven rule is of the most efficiency with comparable classification accuracy. The research outcomes pave the way for real-time prediction of grasps and corresponding force in human-centred environments. Yinfeng Fang, Dalin Zhou, Kairu Li, Zhaojie Ju, Honghai Liu 0001 |
SMC | 2 |