Zhibin Quan

dblp:141/2189 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0003-1748-8586ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 11 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Cross-modal semantic token alignment via contrastive learning for weakly-supervised referring image segmentation
Congwei Zhang, Zhibin Quan, Wankou Yang
Expert Syst. Appl.2
2026 Self-visual prompting for training-free recursive weakly supervised semantic segmentation
Congwei Zhang, Zhibin Quan, Yuncong Yao, Wankou Yang
Multim. Syst.2
2025 Dynamic Compact Consensus Tracking for Aerial Robots
abstract
Existing one-stream trackers have attracted widespread attention. However, they are not applicable in real-time aerial robot tracking systems due to substantial computational overhead, especially when dynamic templates are introduced. To address this issue, we propose a novel Dynamic Compact Consensus Tracker (DC2T), constructed by stacking blocks that each consists of a Compact Token Encoder (CTE) and Dynamic Consensus Attention (DCA). Unlike traditional methods that convert images into a large number of tokens, the CTE, inspired by “superpixel”, extracts a compact set of representative tokens from both initial and dynamic templates, eliminating the need for a large token set. This strategic reduction in the number of compact tokens markedly decreases the computational load of CTE, enhancing the efficiency of subsequent attention operations. To achieve linear complexity of the DCA, compact dynamic template tokens (as keys) are requeried by search tokens (as queries) to perform dynamic consensus on the aggregated tokens (as values). This arrangement seamlessly incorporates dynamic spatio-temporal features into the DCA while avoiding the computational burden typically associated with dynamic templates. With the aim of further enhancing the system's responsiveness and accuracy, a direct control network is crafted to seamlessly incorporate the prediction of high-level control values into the tracking network, ensuring a cohesive and efficient interaction with the controller. Comprehensive experiments and real-world evaluations have proven DC2T's superior performance, accompanied by a significant reduction in FLOPs. Furthermore, we have conducted experiments that demonstrate the tracker's ability to integrate seamlessly with other technologies such as SLAM and detection, enabling precise tracking of arbitrary objects. The tracker code will be released in the github. com/xiaolousun/ refine-pytracking.
Xiaolou Sun, Zhibin Quan, Yuntian Li, Wufei Si, Wenhui Ni, Runwei Guan
ICRA2
2024 CLAT: Convolutional Local Attention Tracker for Real-time UAV Target Tracking System with Feedback Information
abstract
Real-time UAV vision target tracking systems encounter the intricate challenges of striking a trade-off for tracking speed and performance, and the robustness of the following control. In existing tracking systems, the global attention mechanism enhances tracking performance, but it introduces higher computational complexity, impacting target tracking speed; the local attention mechanism can reduce computational complexity but often exhibits limitations in modeling the receptive field. In this paper, we propose a new framework named Convolutional Local Attention Tracker (CLAT) to address these challenges. Firstly, we design a hierarchical convolutional local attention structure as the feature extractor for CLAT. This leverages convolutional projection before local window partitioning, facilitating connections between non-overlapping windows and expanding the receptive field. Secondly, we introduce a streamlined feature fusion network comprising the unshared-weights convolutional layer and a global attention network. The whole design can balance speed and accuracy. Furthermore, to enhance servo control robustness, we have redesigned the upper-level controller by integrating all bounding box information. To capture feedback spatiotemporal information in CLAT, a dynamic template update is implemented by incorporating an IOU head into the predictor. Extensive experiments on visual tracking benchmarks and in the real world demonstrate that CLAT achieves competitive performance. Moreover, we have developed a comprehensive tracking system demonstration capable of precisely tracking targets across various categories. The tracker code will be released on https://github.com/xiaolousun/refine-pytracking.git.
Xiaolou Sun, Zhibin Quan, Wufei Si, Yuntian Li
IROS2
2024 Domain knowledge-powered attention for air traffic management hazardous events classification
Zhibin Quan, Xianghua Tan
Eng. Appl. Artif. Intell.3
2024 SED: Searching Enhanced Decoder with switchable skip connection for semantic segmentation
Zhibin Quan, Qiang Li 0024, Dejun Zhu, Wankou Yang
Pattern Recognit.2
2024 The MorPhEMe Machine: An Addressable Neural Memory for Learning Knowledge-Regularized Deep Contextualized Chinese Embedding
abstract
Deep contextualized embeddings, as learned by large pre-training models, have proven highly effective in various downstream natural language processing tasks. However, the embedding space in these large models lacks explicit regularization, leading to underfitting and substantial costs during large-scale training on huge corpora. In this paper, we present a novel approach to learning deep contextualized embeddings, introducing linguistic knowledge regularization. Specifically, our proposed model, MorPhEMe (Morphology and Phonology Embedding Memory), features an external addressable memory with two additional addressable memories for storing morphology and phonology knowledge. MorPhEMe can be seamlessly stacked into a deep architecture. Notably different from existing pre-training models, MorPhEMe boasts two distinctive features: (1) compositional encoding and decompositional decoding facilitated by a dynamic addressing mechanism; and (2) explicit memory embedding regularization through cross-layer memory sharing. Theoretical analysis suggests that the inclusion of morphology and phonology enables MorPhEMe to reduce the modeling complexity of natural language sequences. We evaluate MorPhEMe across a diverse set of Chinese natural language processing tasks, including language modeling, word similarity computation, word analogy reasoning, relation extraction, and machine reading comprehension. Experimental results demonstrate that MorPhEMe, in contrast to state-of-the-art models, achieves remarkable improvements with fewer parameters and rapid convergence.
Zhibin Quan, Chi-Man Vong, Wankou Yang
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Enhancing motion visual cues for self-supervised video representation learning
Mu Nie, Zhibin Quan, Weiping Ding 0001, Wankou Yang
Eng. Appl. Artif. Intell.2
2023 Detecting and grouping keypoints for multi-person pose estimation using instance-aware attention
Ze Feng, Zhicheng Wang 0001, Shoukui Zhang, Zhibin Quan, Shutao Xia, Wankou Yang
Pattern Recognit.6
2023 Pyramid Geometric Consistency Learning For Semantic Segmentation
Qiang Li 0024, Zhibin Quan, Wankou Yang
Pattern Recognit.3
2022 Siamese Transformer Network: Building an autonomous real-time target tracking system for UAV
Xiaolou Sun, Zhibin Quan, Hao Wang 0144, Yuncong Yao, Wankou Yang
J. Syst. Archit.4
2022 Spatial information enhancement network for 3D object detection from point cloud
Yuncong Yao, Zhibin Quan, Wankou Yang
Pattern Recognit.3
2021 TransPose: Keypoint Localization via Transformer
abstract
While CNN-based models have made remarkable progress on human pose estimation, what spatial dependencies they capture to localize keypoints remains unclear. In this work, we propose a model called Trans-Pose, which introduces Transformer for human pose estimation. The attention layers built in Transformer enable our model to capture long-range relationships efficiently and also can reveal what dependencies the predicted key-points rely on. To predict keypoint heatmaps, the last attention layer acts as an aggregator, which collects contributions from image clues and forms maximum positions of keypoints. Such a heatmap-based localization approach via Transformer conforms to the principle of Activation Maximization [19]. And the revealed dependencies are image-specific and fine-grained, which also can provide evidence of how the model handles special cases, e.g., occlusion. The experiments show that TransPose achieves 75.8 AP and 75.0 AP on COCO validation and test-dev sets, while being more lightweight and faster than mainstream CNN architectures. The TransPose model also transfers very well on MPII benchmark, achieving superior performance on the test set when fine-tuned with small training costs. Code and pre-trained models are publicly available1.
Zhibin Quan, Mu Nie, Wankou Yang
ICCV2
2020 Recurrent Neural Networks With External Addressable Long-Term and Working Memory for Learning Long-Term Dependences
abstract
Learning long-term dependences (LTDs) with recurrent neural networks (RNNs) is challenging due to their limited internal memories. In this paper, we propose a new external memory architecture for RNNs called an external addressable long-term and working memory (EALWM)-augmented RNN. This architecture has two distinct advantages over existing neural external memory architectures, namely the division of the external memory into two parts-long-term memory and working memory-with both addressable and the capability to learn LTDs without suffering from vanishing gradients with necessary assumptions. The experimental results on algorithm learning, language modeling, and question answering demonstrate that the proposed neural memory architecture is promising for practical applications.
Zhibin Quan, Yandong Liu 0002, Yunxiu Yu, Wankou Yang
IEEE Trans. Neural Networks Learn. Syst.1
2015 TBox learning from incomplete data by inference in BelNet+
Man Zhu, Jeff Z. Pan, Zhibin Quan
Knowl. Based Syst.6
2014 Noisy Type Assertion Detection in Semantic Datasets
Man Zhu, Zhibin Quan
ISWC (1)3
2013 Ontology Learning from Incomplete Semantic Web Data by BelNet
abstract
Recent years have seen a dramatic growth of semantic web on the data level, but unfortunately not on the schema level, which contains mostly concept hierarchies. The shortage of schemas makes the semantic web data difficult to be used in many semantic web applications, so schemas learning from semantic web data becomes an increasingly pressing issue. In this paper we propose a novel schemas learning approach -BelNet, which combines description logics (DLs) with Bayesian networks. In this way BelNet is capable to understand and capture the semantics of the data on the one hand, and to handle incompleteness during the learning procedure on the other hand. The main contributions of this work are: (i)we introduce the architecture of BelNet, and corresponding lypropose the ontology learning techniques in it, (ii) we compare the experimental results of our approach with the state-of-the-art ontology learning approaches, and provide discussions from different aspects.
Man Zhu, Jeff Z. Pan, Zhibin Quan
ICTAI6