Zhiyi Mo

dblp:299/5379 · DBLP profile ↗
← Back
23ranked-venue papers
2as first author
23since 2021 · last 2026
0009-0008-6123-363XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 17 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BigCounter: A Bidirectional-Guided Network With Scene-Semantics-Driven Fusion for RGB-Thermal Crowd Counting
abstract
Accurate crowd counting has become increasingly essential for public safety management and Internet of Video Things (IoVT) applications, driven by rapid population growth and urbanization. However, RGB-Thermal (RGB-T) crowd counting remains challenging due to poor recognition of small targets and degraded performance under extreme conditions such as low-light environments. To address these issues, we propose a bidirectional-guided network with scene-semantics driven fusion for RGB-T crowd counting (BigCounter) that enhances robustness and generalization in complex scenes. BigCounter comprises of three parallel branches: a primary branch, a dynamic illumination auxiliary enhancement branch (DIAEB), and a high-resolution auxiliary enhancement branch (HAEB), which respectively improve robustness under illumination variations and accuracy for small target detection. Moreover, a cross-layer scene-driven fusion module (CLSFM) and a cross-modal semantic-driven fusion module (CMSFM) are designed to strengthen structural consistency and explore semantic complementarity between modalities. Through multi-branch collaboration and semantic-aware fusion, BigCounter significantly enhances feature representation. Extensive experiments on two benchmark RGB-T datasets demonstrate that BigCounter achieves superior accuracy and generalization compared with state-of-the-art methods.
Xiaomin Fan, Feng Shao 0001, Baoyang Mu, Xiongli Chai, Zhongjie Zhu, Zhiyi Mo
IEEE Internet Things J.6
2026 PU-TransMamba: A Hybrid Point Cloud Upsampling Framework With Detail-Aware Transformer and Spatially Coherent Mamba
abstract
High-quality 3D point clouds are essential for high-fidelity perception in Internet of Things (IoT)-enabled intelligent systems. While Point Cloud Upsampling (PCU) is widely used to mitigate data sparsity, existing methods often struggle to balance the preservation of fine-grained local details with the maintenance of global topological consistency. Transformer-based approaches frequently suffer from excessive computational overhead and high-frequency detail loss, whereas emerging state space models like Mamba, despite their efficiency, inevitably sacrifice spatial coherence due to the 1D serialization of irregular 3D points. To address these critical bottlenecks, we introduce PU-TransMamba, a hybrid framework that synergistically leverages a Detail-Aware Transformer and a Spatially-Coherent Mamba. Each component is designed to resolve specific PCU limitations: a Complexity-Aware Bilateral Decoder is developed to adaptively recover sharp geometric edges by processing features across dual domains, while a Sequence-Aligned Mamba Encoder utilizes multiple spatial curvature descriptors to compensate for the spatial information loss inherent in serialization. Additionally, a Global Geometry Injector and a Local Neighbor Injector are designed to ensure structural integrity by infusing holistic skeletal priors and neighborhood context, respectively. To minimize feature discrepancies between the hybrid branches, we also propose a Self-Distillation Loss. Extensive experiments on five benchmark datasets demonstrate that PU-TransMamba outperforms state-of-the-art methods in both reconstruction accuracy and computational scalability. The results confirm its ability to recover intricate geometries, indicating significant potential for IoT-driven 3D perception and communication systems.
Feng Shao 0001, Xiongli Chai, Hangwei Chen, Zhongjie Zhu, Zhiyi Mo
IEEE Internet Things J.6
2025 MambaLCT: Boosting Tracking via Long-term Context State Space Model
abstract
Effectively constructing context information with long-term dependencies from video sequences is crucial for object tracking. However, the context length constructed by existing work is limited, only considering object information from adjacent frames or video clips, leading to insufficient utilization of contextual information. To address this issue, we propose MambaLCT, which constructs and utilizes target variation cues from the first frame to the current frame for robust tracking. First, a novel unidirectional Context Mamba module is designed to scan frame features along the temporal dimension, gathering target change cues throughout the entire sequence. Specifically, target-related information in frame features is compressed into a hidden state space through a selective scanning mechanism. The target information across the entire video is continuously aggregated into target variation cues. Next, we inject the target change cues into the attention mechanism, providing temporal information for modeling the relationship between the template and search frames. The advantage of MambaLCT is its ability to continuously extend the length of the context, capturing complete target change cues, which enhances the stability and robustness of the tracker. Extensive experiments show that long-term context information enhances the model's ability to perceive targets in complex scenarios. MambaLCT achieves new SOTA performance on six benchmarks while maintaining real-time runing speeds.
Xiaohai Li, Bineng Zhong 0001, Qihua Liang, Guorong Li, Zhiyi Mo, Shuxiang Song 0001
AAAI5
2025 Robust Tracking via Mamba-based Context-aware Token Learning
abstract
How to make a good trade-off between performance and computational cost is crucial for a tracker. However, current famous methods typically focus on complicated and time-consuming learning that combining temporal and appearance information by input more and more images (or features). Consequently, these methods not only increase the model's computational source and learning burden but also introduce much useless and potentially interfering information. To alleviate the above issues, we propose a simple yet robust tracker that separates temporal information learning from appearance modeling and extracts temporal relations from a set of representative tokens rather than several images (or features). Specifically, we introduce one track token for each frame to collect the target's appearance information in the backbone. Then, we design a mamba-based Temporal Module for track tokens to be aware of context by interacting with other track tokens within a sliding window. This module consists of a mamba layer with autoregressive characteristic and a cross-attention layer with strong global perception ability, ensuring sufficient interaction for track tokens to perceive the appearance changes and movement trends of the target. Finally, track tokens serve as a guidance to adjust the appearance feature for the final prediction in the head. Experiments show our method is effective and achieves competitive performance on multiple benchmarks at a real-time speed.
Jinxia Xie, Bineng Zhong 0001, Qihua Liang, Ning Li 0044, Zhiyi Mo, Shuxiang Song 0001
AAAI5
2025 Dynamic Updates for Language Adaptation in Visual-Language Tracking
abstract
The consistency between the semantic information provided by the multi-modal reference and the tracked object is crucial for visual-language (VL) tracking. However, existing VL tracking frameworks rely on static multi-modal references to locate dynamic objects, which can lead to semantic discrepancies and reduce the robustness of the tracker. To address this issue, we propose a novel vision-language tracking framework, named DUTrack, which captures the latest state of the target by dynamically updating multimodal references to maintain consistency. Specifically, we introduce a Dynamic Language Update Module, which leverages a large language model to generate dynamic language descriptions for the object based on visual features and object category information. Then, we design a Dynamic Template Capture Module, which captures the regions in the image that highly match the dynamic language descriptions. Furthermore, to ensure the efficiency of description generation, we design an update strategy that assesses changes in target displacement, scale, and other factors to decide on updates. Finally, the dynamic template and language descriptions that record the latest state of the target are used to update the multi-modal references, providing more accurate reference information for subsequent inference and enhancing the robustness of the tracker. DUTrack achieves new state-of-the-art performance on five mainstream vision-language and two vision-only tracking benchmarks, including LaSOT, LaSOText, TNL2K, OTB99-Lang, MGIT, GOT-10K, and UAV123. Code and models are available at https://github.com/GXNU-ZhongLab/DUTrack.
Xiaohai Li, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Jian Nong, Shuxiang Song 0001
CVPR4
2025 Robust tracking via rethinking prediction head
Jian Nong, Yongjun Qi, Zhiyi Mo, Yanyan Liang 0001
Image Vis. Comput.3
2025 SIEVL-Track: Exploring Semantic Information Enhancement for Visual-Language Object Tracking
abstract
With the assistance of language descriptions, Visual-Language (VL) object tracking can obtain more accurate semantic information compared to traditional Visual-Only object tracking. However, the ability of current VL trackers to obtain target semantic information has not been fully developed due to limitations such as wasted modeling capabilities and insufficient utilization of historical temporal information. On the one hand, the modeling output from Transformer shallow encoders often does not directly participate in the prediction of tracking results, resulting in a certain degree of model capability waste. On the other hand, the semantic information of historical tracking results has also not been fully utilized in the tracking process, resulting in a certain degree of lack of semantic assistance capability. Therefore, we propose a novel hierarchical multi-stage VL tracker called SIEVL-Track to enhance target semantic information. Specifically, we first design a multi-stage visual language tracking framework for modeling multi-scale semantic information in Visual-Language tracking pipeline. Secondly, we propose a selective deep and shallow semantic information fusion module (S-DSFM) that explicitly integrates shallow output features into deep output features, so to reduce the waste of modeling capabilities and obtain more high-frequency semantic information related to the target. Finally, we design a temporal cue modeling module based on linguistic classification and multi-frame historical information(MHLS-TCM), with the aim of more comprehensive utilization of historical temporal semantic information. Benefit from the above designs, our VL tracker can obtain stronger target semantic information. Competitive performance from extensive experimental results on five popular vision-language tracking benchmarks, including LaSOT, OTB99-Lang, WebUAV-3M, LaSOText and TNL2K, have demonstrated the superiority and effectiveness of our SIEVL-Track.
Ning Li 0044, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Jian Nong, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Mamba Adapter: Efficient Multi-Modal Fusion for Vision-Language Tracking
abstract
Utilizing the high-level semantic information of language to compensate for the limitations of vision information is a highly regarded approach in single-object tracking. However, most existing vision-language (VL) trackers employ full-parameter fine-tuning, which can easily lead to catastrophic forgetting. Therefore, they fail to fully exploit the prior knowledge of pre-trained models from upstream tasks, resulting in unsatisfactory tracking performance. To alleviate the above problem, we propose a simple yet effective Vision-Language Tracking pipeline based on Mamba Adapter, named MAVLT, which adopts the idea of parameter-efficient fine-tuning (PEFT) to realize the interaction between vision-language modalities. This novel approach offers the following advantages: (1)The knowledge of the upstream pre-trained model is efficiently inherited by freezing its parameters. This ensures that the VL tracking framework only learns the modules for vision and language interaction, with a focus on the fusion between modalities. (2)The modal interaction between language and vision encoders is flexibly bridged in each encoder layer via proposed mamba adapter, enabling efficient interaction of visual and language information at multiple levels. Extensive experiments on five popular vision-language tracking benchmarks validate the effectiveness of the proposed MAVLT. Particularly, the MAVLT achieves 73.4% AUC score on the LaSOT benchmarks with only 0.18%(0.32M) of the total parameters updates. Code and models are available at https://github.com/GXNU-ZhongLab/MAVLT.
Liangtao Shi, Bineng Zhong 0001, Qihua Liang, Xiantao Hu, Zhiyi Mo, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Adversarial Robust Salient Object Detection in Optical Remote Sensing Images With Implicit Feature Enhancement
abstract
Deep neural networks (DNNs) have achieved significant progress in optical remote sensing images salient object detection (ORSI-SOD) and are widely applied to various remote sensing image analysis tasks. However, few SOD models demonstrate robust performance under adversarial perturbations, which ultimately leads to a decline in detection accuracy. Moreover, most existing defense methods inject fixed Gaussian noise globally into the image. Although such approaches are easy to implement, they have several limitations in inaccurate uncertainty estimation and neglecting the unique characteristics of local salient regions. Furthermore, existing adversarial defense research rarely addresses the challenges specific to the ORSI-SOD task, leaving a gap in effective defense strategies. To tackle these issues, we propose a novel defense method, which enhances the adversarial robustness of ORSI-SOD models through implicit feature enhancement. The algorithm first proposes a two-stage strategy of reverse local noise search and forward global noise optimization, enhancing generalization ability by implicitly enhancing features to better simulate network uncertainty. Then, the algorithm proposes a global-guided texture information enhancement (GTIE) module for low-level features and a global-guided semantics information enhancement (GSIE) module for high-level features, focusing on strengthening low-level texture information and enhancing the model’s understanding of high-level contextual semantic features, respectively. This dual-module design effectively weakens the impact of adversarial noise, significantly improving the robustness and accuracy of object detection. Extensive experiments on three ORSI-SOD datasets demonstrate that our defense strategy better estimates the uncertainty, resulting in an average performance improvement of 23.2% in$F_{\beta } ^{\mathrm { max}}$and 34.1% in$E_{\xi } ^{\mathrm { max}}$across six ORSI-SOD models under five different adversarial attack methods. Our code will be released in the public repository athttps://github.com/kexi0714/IFe.
Feng Shao 0001, Xiangchao Meng, Hangwei Chen, Xiongli Chai, Zhiyi Mo
IEEE Trans. Geosci. Remote. Sens.6
2025 Robust Multi-Stage Tracking via Multi-Scale and Multi-Level Representation Learning
abstract
How to learn multi-scale and multi-level representations is crucial for robust tracking. However, most current one-stream structure based trackers with visual transformers (dubbed ViTs) cannot effectively capture multi-scale representations due to the structure of their adopted ViTs is non-hierarchical. Meanwhile, they often only use the output features from the final layer for predicting results (i.e., ignoring the utilization of low-level features from the shallow layers) which may result in a certain degree of lacking multi-level representation learning ability. To address these issues, we propose a robust multi-stage tracker that effectively combines the advantages of both hierarchical and one-stream structured ViT as a tracking backbone to improve the multi-scale and multi-level representation learning abilities. Specifically, first of all, we design a hierarchical tracker with a three-stage backbone. In the first two stages of our tracker, we utilize a dual-branch structure to obtain multi-scale features of the template and search region separately. Especially, We design the local scale awareness modules based on simple MLP layers to capture multi-scale features. These modules remove complex operations such as convolutions or shifted window attentions, thus avoiding the performance degradation caused by traditional hierarchical ViTs. In the third stage (i.e. the main stage), we construct a global encoder based on the one-stream ViT to achieve efficient feature extraction and feature interaction for our tracker. Then, we design a multi-level feature integration module in the main stage to explicitly utilize the representation information learned from the shallow layers and fuse them with the features of the final layer to obtain multi-level representation information. Lastly, benefit from the these designs, our tracker can effectively capture more multi-scale and multi-level representations for robust tracking. Comprehensive experiments on GOT-10k, LaSOT, LaSOT$_{ext}$, TNL2K, UAV123, TrackingNet and VOT2020 benchmarks validate the effectiveness and robustness of our method.
Ning Li 0044, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shuxiang Song 0001
IEEE Trans. Multim.4
2025 Uncertainty-Guided Diffusion Model for Camouflaged Object Detection
abstract
Recently, diffusion models have significantly improved the performance of Camouflaged Object Detection (COD) by adding noise to a mask and iteratively denoising it to match the target distributions. Due to the direct extraction of features from noisy masks and the lack of conditional constraints on a prediction area, the diffusion model may deviate from a correct prediction range and produces mispredictions in regions with high uncertainty. To address this issue, we propose an uncertainty-guided diffusion model (UGDNet) for COD, which explicitly quantifies uncertainty and integrates it as an anchor condition into the diffusion models to provide an initialization of the diffusion regions. The core idea is first to utilize a probability representation and transformer to explicitly model uncertainty, aiming to identify areas where a model may generate overconfident mispredictions. Then, we use the uncertainty as an anchor condition to provide a reference prediction range for the diffusion model, guiding each step of the diffusion process. Furthermore, we use uncertainty to guide feature aggregation, prompting the model to pay extra attention to the semantic features of regions with high uncertainty to refine the segmentation results further. The experimental results indicate that our proposed UGDNet achieves higher accuracy than existing state-of-the-art models on five COD benchmarks, including COD10K, NC4K, CAMO, CHAMELEON, and CDS2K.
Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shengping Zhang, Shuxiang Song 0001
IEEE Trans. Multim.4
2024 ODTrack: Online Dense Temporal Token Learning for Visual Tracking
abstract
Online contextual reasoning and association across consecutive video frames are critical to perceive instances in visual tracking. However, most current top-performing trackers persistently lean on sparse temporal relationships between reference and search frames via an offline mode. Consequently, they can only interact independently within each image-pair and establish limited temporal correlations. To alleviate the above problem, we propose a simple, flexible and effective video-level tracking pipeline, named ODTrack, which densely associates the contextual relationships of video frames in an online token propagation manner. ODTrack receives video frames of arbitrary length to capture the spatio-temporal trajectory relationships of an instance, and compresses the discrimination features (localization information) of a target into a token sequence to achieve frame-to-frame association. This new solution brings the following benefits: 1) the purified token sequences can serve as prompts for the inference in the next video frame, whereby past information is leveraged to guide future inference; 2) the complex online update strategies are effectively avoided by the iterative propagation of token sequences, and thus we can achieve more efficient model representation and computation. ODTrack achieves a new SOTA performance on seven benchmarks, while running at real-time speed. Code and models are available at https://github.com/GXNU-ZhongLab/ODTrack.
Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shengping Zhang, Xianxian Li
AAAI4
2024 Autoregressive Queries for Adaptive Tracking with Spatio-Temporal Transformers
abstract
The rich spatio-temporal information is crucial to capture the complicated target appearance variations in visual tracking. However, most top-performing tracking algorithms rely on many hand-crafted components for spatio-temporal information aggregation. Consequently, the spatio-temporal information is far away from being fully explored. To alleviate this issue, we propose an adaptive tracker with spatio-temporal transformers (named AQA-Track), which adopts simple autoregressive queries to effectively learn spatio-temporal information without many hand-designed components. Firstly, we introduce a set of learnable and autoregressive queries to capture the instantaneous target appearance changes in a sliding window fashion. Then, we design a novel attention mechanism for the interaction of existing queries to generate a new query in current frame. Finally, based on the initial target template and learnt autoregressive queries, a spatio-temporal information fusion module (STM) is designed for spatiotemporal formation aggregation to locate a target object. Benefiting from the STM, we can effectively combine the static appearance and instantaneous changes to guide robust tracking. Extensive experiments show that our method significantly improves the tracker's performance on six popular tracking benchmarks: LaSOT, LaSOText, TrackingNet, GOT-10k, TNL2K, and UAV123. Code and models will be https://github.com/orgs/GXNU-ZhongLab.
Jinxia Xie, Bineng Zhong 0001, Zhiyi Mo, Shengping Zhang, Liangtao Shi, Shuxiang Song 0001, Rongrong Ji
CVPR3
2024 Visual Adapt for RGBD Tracking
abstract
Recent RGBD trackers have employed cueing techniques by overlaying Depth modality images as cues onto RGB modality images, which are then fed into the RGB-based model for tracking. However, the direct overlaying interaction method between modalities not only introduces more noise into the feature space but also exhibits the inadaptability of the RGB-based model to mixed-modality inputs. To address these issues, we introduce Visual Adapt for RGBD Tracking (VADT). Specifically, we maintain the input of the RGB-based model as the RGB modality. Additionally, we have devised a fusion module to enable modality interaction between depth and RGB features. Subsequently, a Depth Adapt module has been formulated to facilitate image interaction with the fused features. This module involves cross-attending to the obtained depth-assisted features and the RGB search frame features produced by the RGB-based model’s output. Experimental results indicate that our proposed tracker achieves state-of-the-art results on various RGBD benchmark tests.
Guangtong Zhang, Qihua Liang, Zhiyi Mo, Ning Li 0044, Bineng Zhong 0001
ICASSP3
2024 Diffusion Mask-Driven Visual-language Tracking
Guangtong Zhang, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shuxiang Song 0001
IJCAI4
2024 Dual-stream Multi-modal Interactive Vision-language Tracking
Zhiyi Mo, Guangtong Zhang, Jian Nong, Bineng Zhong 0001, Zhi Li 0017
MMAsia1
2024 Top-Down Cross-Modal Guidance for Robust RGB-T Tracking
abstract
Most RGB-T trackers heavily rely on bottom-up attention and thus overlook top-down cross-modal guidance for learning target features. Consequently, the discriminative power of the learnt target features is weak. To address this issue, we propose a novel RGB-T tracker (called TGTrack) that designs a Top-down Cross-modal Guidance mechanism to learn target features in two stages. In the first stage, our TGTrack effectively generates top-down cross-modal guidance signals with multi-modal encoders-decoders and prior vectors. In the second stage, these signals are transmitted and integrated to improve the discriminative power of our target features by the attention layers of the cross-modal encoders. Moreover, we introduce an Attention-Driven Spatio-Temporal Updater for updating discriminative target features. Through cross-frame attention guidance, it can effectively eliminates irrelevant features within the search region. As a result, our TGTrack can effectively avoid the complex multi-modal fusion modules and thus achieve robust RGB-T tracking. Extensive experiments on three popular RGB-T tracking benchmarks (i.e., LasHeR, RGBT234, and RGBT210) demonstrate that our TGTrack achieves new state-of-the-art performances.
Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Zhiyi Mo, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Robust Tracking via Combing Top-Down and Bottom-Up Attention
abstract
Transformer attention plays an important role in current top-performing trackers. However, it is bottom-up, driven by stimulus and lacks intrinsic prior guidance. This bottom-up attention mechanism leads to an emphasis on all objects in the input images, rather than the task related objects. As a result, the performance of the bottom-up attention based trackers is deteriorated in complicated scenes. To address this issue, we propose a robust tracker that combines bottom-up attention with top-down attention to comply with the existing ViT framework, named TBTrack. TBTrack can not only utilize the existing bottom-up attention mechanisms to model the long-range relationship of input tokens, but also utilize a newly added top-down attention mechanism to pay more attention to task related object and further eliminate interference from similar objects and backgrounds. Specifically, we firstly design a top-down prior generation module using an adaptive learning parameter combined with the template inputs to obtain top-down task guided signals. Then, we inject the prior signals into a bottom-up attention module to obtain a top-down and bottom-up attention combination block (TB-Block). Finally, we stack these TB-Blocks to construct our tracker (TBTrack) with top-down prior guidance capability, which focuses more on the task related object. Through extensive experiments, our TBTrack achieves impressive performance on multiple tracking benchmarks, including GOT-10k, LaSOT, LaSOText, TNL2K, TrackingNet, UAV123 and so on. The code and trained models will be publicly available.
Ning Li 0044, Bineng Zhong 0001, Yaozong Zheng, Qihua Liang, Zhiyi Mo, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 One-Stream Stepwise Decreasing for Vision-Language Tracking
abstract
Based on the fixed language descriptions in the initial frames, a vision-language tracker typically adopts a two-stream model structure to align vision and language features at the feature fusion stages. However, this paradigm may degrade the tracking performance due to inaccurate language descriptions and lacks further modal interaction. To address these issues, we propose a one-stream vision-language model called One-stream Stepwise Decreasing for Vision-Language Tracking (OSDT). Specifically, we first encode the language description using a language encoder. The obtained language features are then combined with visual images and entered jointly into a visual encoder, in which the encoder’s self-attention mechanism is utilized to facilitate more interactions between language and visual features. Moreover, to mitigate the problems caused by inaccurate language descriptions, we design a stepwise decreasing multi-modal interaction framework, in which a Feature Filter Module (FFM) is introduced to select language features that are more relevant to visual information to provide semantic guidance for visual feature extraction. Furthermore, without additional feature fusion modules, our one-stream model framework can efficiently utilize the proposed feature filtering module for feature selection. Consequently, our tracker can achieve fast tracking speed in the vision-language tracking domain compared to existing state-of-the-art methods. We extensively evaluate our tracker on three benchmarks, i.e. TNL2K, LaSOT, and OTB99, demonstrating competing performance compared to state-of-the-art vision-language tracking methods.
Guangtong Zhang, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Ning Li 0044, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Robust Tracking via Unifying Pretrain-Finetuning and Visual Prompt Tuning
abstract
The finetuning paradigm has been a widely used methodology for the supervised training of top-performing trackers. However, the finetuning paradigm faces one key issue: it is unclear how best to perform the finetuning method to adapt a pretrained model to tracking tasks while alleviating the catastrophic forgetting problem. To address this problem, we propose a novel partial finetuning paradigm for visual tracking via unifying pretrain-finetuning and visual prompt tuning (named UPVPT), which can not only efficiently learn knowledge from the tracking task but also reuse the prior knowledge learned by the pre-trained model for effectively handling various challenges in tracking task. Firstly, to maintain the pre-trained prior knowledge, we design a Prompt-style method to freeze some parameters of the pretrained network. Then, to learn knowledge from the tracking task, we update the parameters of the prompt and MLP layers. As a result, we cannot only retain useful prior knowledge of the pre-trained model by freezing the backbone network but also effectively learn target domain knowledge by updating the Prompt and MLP layer. Furthermore, the proposed UPVPT can easily be embedded into existing Transformer trackers (e.g., OSTracker and SwinTracker) by adding only a small number of model parameters (less than 1% of a Backbone network). Extensive experiments on five tracking benchmarks (i.e., UAV123, GOT-10k, LaSOT, TNL2K, and TrackingNet) demonstrate that the proposed UPVPT can improve the robustness and effectiveness of the model, especially in complex scenarios.
Guangtong Zhang, Qihua Liang, Ning Li 0044, Zhiyi Mo, Bineng Zhong 0001
MMAsia4
2023 SpectralTracker: Jointly High and Low-Frequency Modeling for Tracking
Yimin Rong, Qihua Liang, Ning Li 0044, Zhiyi Mo, Bineng Zhong 0001
PRCV (12)4
2021 Application of Bayesian Network Reasoning Algorithm in Emotion Classification
abstract
In the early stage of covid-19 disease transmission, it is easy to lead to public panic and dissatisfaction without timely information feedback. In order to solve this problem, this paper constructs an emotion classification and prediction algorithm based on Bayesian network reasoning by analyzing the variable elimination algorithm, connection tree reasoning algorithm and Gibbs sampling algorithm in Bayesian network reasoning algorithm. The algorithm can quickly identify the emotions of Internet users from the communication text with low computational resources, and provide reference for the relevant departments to formulate the correct public opinion guidance strategy.
Zhiyi Mo, Yutong Xing, Zizhen Peng
TrustCom1
2021 Aspect-Level Sentiment Analysis Approach via BERT and Aspect Feature Location Model
abstract
With the rapid development of Internet social platforms, buyer shows (such as comment text) have become an important basis for consumers to understand products and purchase decisions. The early sentiment analysis methods were mainly text‐level and sentence‐level, which believed that a text had only one sentiment. This phenomenon will cover up the details, and it is difficult to reflect people’s fine‐grained and comprehensive sentiments fully, leading to people’s wrong decisions. Obviously, aspect‐level sentiment analysis can obtain a more comprehensive sentiment classification by mining the sentiment tendencies of different aspects in the comment text. However, the existing aspect‐level sentiment analysis methods mainly focus on attention mechanism and recurrent neural network. They lack emotional sensitivity to the position of aspect words and tend to ignore long‐term dependencies. In order to solve this problem, on the basis of Bidirectional Encoder Representations from Transformers (BERT), this paper proposes an effective aspect‐level sentiment analysis approach (ALM‐BERT) by constructing an aspect feature location model. Specifically, we use the pretrained BERT model first to mine more aspect‐level auxiliary information from the comment context. Secondly, for the sake of learning the expression features of aspect words and the interactive information of aspect words’ context, we construct an aspect‐based sentiment feature extraction method. Finally, we construct evaluation experiments on three benchmark datasets. The experimental results show that the aspect‐level sentiment analysis performance of the ALM‐BERT approach proposed in this paper is significantly better than other comparison methods.
Guangyao Pang, Keda Lu, Zhiyi Mo, Zizhen Peng, Baoxing Pu
Wirel. Commun. Mob. Comput.5