Miaomiao Zhang 0001

dblp:30/3299-1 · also MiaoMiao Zhang 0001 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-9921-6511ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Mobile-DeepRFB: A Lightweight Terrain Classifier for Automatic Mars Rover Navigation
abstract
It requires terrain classification for unmanned Mars Rover to identify the safe areas. The current deep learning-based semantic segmentation and object recognition suffer from a large number of parameters and long training time. In this paper, a lightweight segmentation framework called Mobile-DeepRFB is proposed for the Martian terrain classification. It improves from the DeepLabV3$+$by taking the MobileNetV3 as the backbone module to decrease the parameters and the Receptive Field Block (RFB) module to strengthen the feature extraction capability as well as to enlarge the receptive field. Experimental results on the NASA Mars terrain dataset AI4MARS show that the presented method reduces 94% about the parameter number and improves the mean pixel accuracy by 2% compared to the existing ResNet101 and Xception backbone networks. The deployment of this framework on a low-computing power embedded platform (NVIDIA Jetson Xavier) demonstrates its great potential to apply to Mars rovers.Note to Practitioners—This paper was motivated by the problem of terrain classification of planetary rovers. Existing methods are typically based on semantic segmentation technology to recognize various terrains while suffering from the drawback of a large number of parameters. We propose a lightweight segmentation framework to address this issue. In particular, the lightweight backbone network is applied to significantly reduce the number of parameters. The receptive field module is substantially improved to enhance the feature extraction capability. Eventually, we deploy the framework on a low-computing platform. Experimental tests show that the framework can significantly reduce the number of network parameters and it can be used for planetary rovers with limited computational resources.
Lihang Feng, Sui Wang, Dong Wang 0036, Pengwen Xiong, Jinjin Xie, Miaomiao Zhang 0001, Qi Wu 0003, Aiguo Song
IEEE Trans Autom. Sci. Eng.7
2025 Edge-Assisted Epipolar Transformer for Industrial Scene Reconstruction
abstract
Given a set of calibrated images, Multiple View Stereo (MVS) applies end-to-end depth inference network to recover scene structure. However, previous methods designed pixel-visibility modules to aggregate cross-view cost, ignoring the consistency assumption of 2D contextual features in the 3D depth direction. The current multi-stage depth inference model also relies on intensive depth samples, which requires high memory consumption. To alleviate these problems, this work exploits edge-assisted epipolar Transformer for multi-view depth inference. The improvements of this work are summarized as follows: 1) The epipolar Transformer block is developed for reliable cross-view cost aggregation, and the edge detection branch is designed to constrain the consistency of epipolar geometry and edge features. 2) The dynamic depth range sampling mechanism based on probability volume is adopted to improve the accuracy of uncertain areas. Comprehensive comparisons with the state-of-the-art works indicate that our work can reconstruct dense scene representations with limited memory bottleblockNote to Practitioners—Learning-based MVS can obtain dense point clouds with accurate depth map estimation, which are widely applied in the fields of unmanned driving, battlefield environment perception and robot navigation. MVS-based scene reconstruction technology is the premise of the subsequent planning, decision-making and control of the human-machine system. To obtain dense scene representation with limited memory and runtime, this work proposes a multi-view stereo network with edge-assisted epipolar Transformer. Experiments on public benchmarks verify the feasibility and effectiveness of our model, which has good potential in battlefield environment reconstruction and human-computer interaction fields, and can provide intuitive and dense scene representation for decision-making assistance.
Kevin W. Tong, Xiaorong Guan, Miaomiao Zhang 0001, Ping Li 0044, Qi Wu 0003, Limin Zhu 0001
IEEE Trans Autom. Sci. Eng.3
2025 "Jumpingly" Perceive Time Series: Image Generation Approach to Modeling Functional Brain Activation
abstract
This paper presents a novel Linear Mapping Field (LMF) to map time series into two-dimensional images. The LMF extracts deeper features of fNIRS signals, which makes fNIRS less reliant on some prior. The developed convolution neural networks detect more prominent features than the state-of-the-art methods. The experimental results indicate that different from RNNs which can only perceive the time series in a “sequential” manner, LMF’s characteristic of “jumpingly” perception is the key to achieving excellent results.Note to Practitioners—As an optical and non-invasive technique to obtain the changes of oxyhemoglobin (O2Hb) and deoxyhemoglobin (HHb), fNIRS can be used to measure the changes of cerebral hemodynamics related to brain activities. This work proposes a linear mapping field and other mapping fields to map fNIRS signals to two-dimensional images, and the deep features of these generated images can be further extracted by convolutional neural networks, establishing an end-to-end bridge to detect the activation of brain functions under different tasks. Compared with mainstream methods, this work can be used as an effective mapping for fNIRS with low computational complexity and good performance.
Kevin W. Tong, Miaomiao Zhang 0001, Zhiyi Shi, Yuhong Hou
IEEE Trans Autom. Sci. Eng.4
2024 A parallel neural networks for emotion recognition based on EEG signals
Yuwen Jie, Kevin W. Tong, Miaomiao Zhang 0001, Guangyu Zhu 0001, Qi Wu 0003
Neurocomputing4
2024 Efficient Reinforcement Learning With the Novel N-Step Method and V-Network
abstract
The application of reinforcement learning (RL) in artificial intelligence has become increasingly widespread. However, its drawbacks are also apparent, as it requires a large number of samples for support, making the enhancement of sample efficiency a research focus. To address this issue, we propose a novel N-step method. This method extends the horizon of the agent, enabling it to acquire more long-term effective information, thus resolving the issue of data inefficiency in RL. Additionally, this N-step method can reduce the estimation variance of Q-function, which is one of the factors contributing to estimation errors in Q-function estimation. Apart from high variance, estimation bias in Q-function estimation is another factor leading to estimation errors. To mitigate the estimation bias of Q-function, we design a regularization method based on the V-function, which has been underexplored. The combination of these two methods perfectly addresses the problems of low sample efficiency and inaccurate Q-function estimation in RL. Finally, extensive experiments conducted in discrete and continuous action spaces demonstrate that the proposed novel N-step method, when combined with classical deep Q-network, deep deterministic policy gradient, and TD3 algorithms, is effective, consistently outperforming the classical algorithms.
Miaomiao Zhang 0001, Shuo Zhang 0023, Zhiyi Shi, Xiangyang Deng, Qi Wu 0003, Xin Xu 0001
IEEE Trans. Cybern.1
2024 Robust Depth Estimation Based on Parallax Attention for Aerial Scene Perception
abstract
Given the precalibrated image pairs, stereo matching aims to infer the scene depth information in real-time, which has important research value in the fields of high-precision 3-D reconstruction of the Earth’s surface, automatic driving and unmanned aerial vehicle (UAV) navigation. The cost volume-based stereo matching method adopts a coarse-to-fine manner to construct cascaded cost volume, and applies 3-D convolution to capture the correspondence of feature matching to infer the disparity map, which achieves comparable performance. However, the existing method has difficulty dealing with jitter regions with disparity change, and direct disparity regression easily leads to overfitting of cost volume regularization. To alleviate the above two problems, this work proposes an end-to-end disparity estimation network based on Transformer. Its specific improvements are as follows. 1) The cross-view feature interaction module based on Transformer is introduced to realize the feature interaction of global context information. 2) A parallax attention mechanism is designed to impose global geometric constraints on the epipolar line to improve the reliability of feature matching. 3) Focal loss is applied for the training of the disparity classification model to emphasize one-hot supervision in ambiguous regions. Comprehensive experiments on public datasets Sceneflow, KITTI2015, ETH3D, and aerial WHU datasets validate that the proposed work can effectively enhance the performance of disparity estimation.
Kevin W. Tong, Miaomiao Zhang 0001, Guangyu Zhu 0001, Xin Xu 0001, Qi Wu 0003
IEEE Trans. Ind. Informatics2
2024 SQIX: QMIX Algorithm Activated by General Softmax Operator for Cooperative Multiagent Reinforcement Learning
abstract
Multiagent cooperative systems can be used to conceptualize many real-world problems. Reinforcement learning is a particularly effective tool. The issue of bias in$Q$-function value estimation in single-agent reinforcement learning has garnered a lot of interest and substantial study. Indeed, this challenge endures in multiagent reinforcement learning, primarily owing to the inclusion of maximization operations. The crux of the matter lies in the inability to seamlessly extrapolate single-agent reinforcement learning algorithms to their multiagent counterparts. In this article, we introduce a more encompassing and straightforward principle: the notion of appropriate value correction. We suggest replacing the maximization operation with a monotonically nondecreasing function to obtain more accurate value estimates. We theoretically demonstrate that this operation effectively reduces the potential overestimation bias in the QMIX algorithm. Ultimately, our methodology, dubbed the SMIX algorithm—a fusion of the QMIX algorithm empowered by the Softmax operator, attains state-of-the-art outcomes across diverse multiagent cooperative tasks. This success extends to challenging domains such as StarCraft II, marking it as one of the most formidable games to date.
Miaomiao Zhang 0001, Kevin W. Tong, Guangyu Zhu 0001, Xin Xu 0001, Qi Wu 0003
IEEE Trans. Syst. Man Cybern. Syst.1
2023 Robust Neural Dynamics Method for Redundant Robot Manipulator Control With Physical Constraints
abstract
Redundant robot manipulators play a significant role in modern industry. In this article, we propose a solution scheme to the trajectory tracking problem of the redundant robot manipulator with physical constraints through the Zhang neural dynamics method. Such problem is integrated into a time-varying system consisting of time-varying nonlinear equation (TVNE) and time-varying linear inequality (TVLI) and solved online by the varying-parameter Zhang neural dynamics (VPZND) model. It is ensured that the redundant robot manipulator can still perform the tracking task perfectly under the coexistence of time-varying bounded noise and physical constraints. Theoretical analysis proves that this VPZND model also has an explicit fixed convergence time. Numerical experiments confirm the feasibility of our VPZND model for TVLI. The trajectory tracking problem of the redundant robot manipulator with six or three degrees of freedom under the dual influence of physical constraints and noise is perfectly solved by the VPZND model, which is enough to verify its practical value.
Miaomiao Zhang 0001, Kevin W. Tong, Ping Li 0044, Yuhong Hou, Xin Xu 0001, Limin Zhu 0001, Qi Wu 0003
IEEE Trans. Ind. Informatics1