EDBT 2026 Demo / reviewers in the wild / expert
Bi Zeng
dblp:87/3324
· DBLP profile ↗
37ranked-venue papers
2as first author
30since 2021 · last 2026
0000-0001-8596-8333ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 1 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DSTree: Data-Driven Synchronous Traversals for Decision Forests on GPUs
Bi Zeng, Zhenlin Wu 0001, Hongyuan Liu 0002 |
IPDPS | 1 |
| 2026 | Dynamic comprehensive difficulty knowledge cells based on KAN network and stable learning for knowledge tracing
Qintai Hu, Qingpeng Wen, Bi Zeng, Liangda Fang |
Neural Networks | 4 |
| 2026 | Mitigating single-view constraints for robust 3D reconstruction
Bi Zeng, Song Wen 0008, Qiaoxia Zhong, Zhuangkai Yao, Zhentao Lin |
Pattern Recognit. | 1 |
| 2026 | DDPTA: Zero-Shot Learning for Skeleton-Based Action RecognitionabstractTraditional skeleton-based action recognition methods rely on large labeled datasets, which are costly to collect and unsuitable for hazardous actions, thereby limiting generalization. To overcome these limitations, recent works adopt zero-shot learning by using rich textual descriptions to guide the alignment and recognition of unlabeled skeleton features. However, these methods still struggle with similar actions (e.g., reading vs. writing), due to ambiguity arising from noise in both modalities. We propose the Discriminative Dual-Prototype TextAlignment (DDPTA) framework. Our framework introduces a novel dual-prototype design with tailored refinement strategies to effectively distill these two complementary prototypes. For the Spatial Prototype, our CycleSpatial module first distills the action's core joint form from noisy spatial features, which is then guided by a Sieve-based Alignment. For the Temporal Prototype, our MambaTempo module leverages the Selective State Space Model to extract representations across distinct temporal stages, enabling fine-grained alignment with descriptions of different time periods. Extensive experiments demonstrate the superior performance of our method, showcasing its effectiveness in advancing the field of zero-shot skeleton-based action recognition. Jinjie Wang, Bi Zeng, Shenghong Zhong, Xiaoting Gao |
IEEE Signal Process. Lett. | 2 |
| 2025 | A Refined 3D Gaussian Representation for High-Quality Dynamic Scene ReconstructionabstractIn recent years, Neural Radiance Fields (NeRF) has revolutionized three-dimensional (3D) reconstruction with its implicit representation. Building upon NeRF, 3D Gaussian Splatting (3D-GS) has departed from the implicit representation of neural networks and instead directly represents scenes as point clouds with Gaussian-shaped distributions. While this shift has notably elevated the rendering quality and speed of radiance fields but inevitably led to a significant increase in memory usage. Additionally, effectively rendering dynamic scenes in 3D-GS has emerged as a pressing challenge. To address these concerns, this paper proposes a refined 3D Gaussian representation for high-quality dynamic scene reconstruction. Firstly, we use a deformable multi-layer perceptron (MLP) network to capture the dynamic offset of Gaussian points and express the color features of points through hash encoding and a tiny MLP to reduce storage requirements. Subsequently, we introduce a learnable denoising mask coupled with denoising loss to eliminate noise points from the scene, thereby further compressing 3D Gaussian model. Finally, motion noise of points is mitigated through static constraints and motion consistency constraints. Experimental results demonstrate that our method surpasses existing approaches in rendering quality and speed, while significantly reducing the memory usage associated with 3D-GS, making it highly suitable for various tasks such as novel view synthesis, and dynamic mapping. Bi Zeng, Zexin Peng |
ICRA | 2 |
| 2025 | MAD-GS:3D Gaussian Splatting for Motion and Defocus Images in Robotic VisionabstractRecent advancements in 3D Gaussian Splatting (3DGS) have significantly improved novel view synthesis, playing a crucial role in robotic vision and scene reconstruction. However, 3DGS relies heavily on precise camera poses and sharp images, which are often difficult to obtain in real-world robotic applications due to motion and defocus blur. Directly applying 3DGS to blurred images results in severe degradation, limiting its effectiveness in tasks such as autonomous navigation and manipulation. To address this challenge, we propose MAD-GS, a novel deblurring framework based on 3DGS, specifically designed for robotic vision tasks. MAD-GS effectively mitigates motion and defocus blur while refining imprecise camera poses, enhancing 3D scene reconstruction under real-world uncertainties. Additionally, we introduce a blur segmentation mask to identify and adaptively refine heavily blurred regions, improving visual quality and downstream robotic decision-making. Extensive experiments on synthetic and real-world datasets demonstrate that MAD-GS outperforms existing methods, leading to superior image clarity and fidelity, thereby advancing robust robot perception in dynamic environments. Tianle Zeng, Bi Zeng, Boquan Zhang, Ziqi Zheng |
IROS | 2 |
| 2025 | HGAtt-ARN: A Novel Adversarial Reconstruction Network Based on Higher-order Gate Attention for Incomplete Multimodal Sentiment AnalysisabstractMultimodal Sentiment Analysis (MSA) is a technique for understanding and recognizing human sentiment by learning representations of different modalities. However, most existing MSA models fail to accurately analyze the real sentiment expressed by irony in slang under incomplete modality. Meanwhile, existing modality reconstruction methods suffer from modality confusion problem, which trigger catastrophic errors in sentiment analysis. To address these challenges, we propose a novel Adversarial Reconstruction Network based on Higher-order Gate Attention for incomplete Multimodal Sentiment Analysis (HGAtt-ARN). Specifically, we propose Higher-order Gate Attention (HGAtt), which can model real representations through high-dimensional mapping and gating computations of heterogeneous modalities, and accordingly construct Adversarial Reconstruction Network (ARN), which contains two core components: (a) HGAtt Reconstructor and CA-HGAtt Reconstructor, which reconstruct modality representations of the real sentiment states implied in irony under incomplete modality through HGAtt and Cross Attention; (b) modality Discriminator, which solves the modality confusion problem by maintaining reconstructed modality invariance through adversarial optimization with Reconstructor. Furthermore, we introduce multimodal Contrastive Learning to improve the performance of our model. Experiments on three public datasets show that HGAtt-ARN achieves SOTA performance, and the case study shows it can accurately recognize the real sentiment expressed by the American ironic slang under incomplete modality. Qingpeng Wen, Qintai Hu, Bi Zeng |
ICMR | 5 |
| 2025 | A Two-Stage Multimodal Framework for Real-Time Item Pickup and Return Recognition in Unmanned Retail Stores
Shenghong Zhong, Bi Zeng, Jinjie Wang, Yujun Zhu |
PRCV (7) | 2 |
| 2025 | BTMTrack: Robust RGB-T tracking via dual-template bridging and temporal-modal candidate elimination
Zhongxuan Zhang, Bi Zeng, Xinyu Ni, Yimin Du |
Image Vis. Comput. | 2 |
| 2024 | Learning Motion Priors with DETR for Visual TrackingabstractRecent transformer-based visual tracking models have demonstrated superior performance. However, prior works have been resource-intensive, requiring massive GPU training hours. This resource demand renders them unsuitable for real-world applications. In this paper, we present DETRack, a training-friendly visual object tracking framework that can integrate motion priors by learning. Our framework utilizes an efficient encoder-decoder structure, with the deformable transformer decoder serving as a target head. We introduce a denoising training strategy to simulate historical predictions and enrich the supervision signal during training. Comprehensive experiments confirm the effectiveness and efficiency of our proposed method. Notably, it only takes 11 hours to train DETRack on a single RTX2080Ti, achieving comparable performance on multiple benchmarks to advanced trackers while maintaining a high running speed. Qingmao Wei, Bi Zeng, Guotian Zeng |
ICME | 2 |
| 2024 | LiteTrack: Layer Pruning with Asynchronous Feature Extraction for Lightweight and Efficient Visual TrackingabstractThe recent advancements in transformer-based visual trackers have led to significant progress, attributed to their strong modeling capabilities. However, as performance improves, running latency correspondingly increases, presenting a challenge for real-time robotics applications, especially on edge devices with computational constraints. In response to this, we introduce LiteTrack, an efficient transformer-based tracking model optimized for high-speed operations across various devices. It achieves a more favorable trade-off between accuracy and efficiency than the other lightweight trackers. The main innovations of LiteTrack encompass: 1) asynchronous feature extraction and interaction between the template and search region for better feature fushion and cutting redundant computation, and 2) pruning encoder layers from a heavy tracker to refine the balnace between performance and speed. As an example, our fastest variant, LiteTrack-B4, achieves 65.2% AO on the GOT-10k benchmark, surpassing all preceding efficient trackers, while running over 100 fps with ONNX on the Jetson Orin NX edge device. Moreover, our LiteTrack-B9 reaches competitive 72.2% AO on GOT-10k and 82.4% AUC on TrackingNet, and operates at 171 fps on an NVIDIA 2080Ti GPU. The code and demo materials will be available at https://github.com/TsingWei/LiteTrack. Qingmao Wei, Bi Zeng, Jianqi Liu, Guotian Zeng |
ICRA | 2 |
| 2024 | MSGIM: A Multi-grained Syntactic Graph Interaction Model for Multi-intent Spoken Language UnderstandingabstractCurrent models for multi-intent detection and slot-filling have made some progress, but often overlook the potential of rich syntactic information and rely on window mechanisms that capture only local slot interactions. In this paper, we propose a Multi-grained Syntactic Graph Interaction Model (MSGIM) for joint multi-intent detection and slot filling. The model has two main core components: (1) Multi-grained syntax module. This module utilizes syntactic dependency types and syntactic dependency distances between words, enabling the model to learn syntactic structure and semantic information more comprehensively. (2) Syntactic graph interaction layer. Specifically, we construct a syntactic slot-aware layer and a syntactic intent-slot interaction layer based on dependency tree. These layers can capture more distant slot dependencies and the interactions between intents and slots. Experimental results show that our model achieves a significant performance improvement, with an overall accuracy improvement of 4.4% on the MixATIS dataset compared to the previous best model. Yikai Zheng, Bi Zeng, Yujun Zhu |
IJCNN | 2 |
| 2024 | CGI-MRE: A Comprehensive Genetic-Inspired Model For Multimodal Relation ExtractionabstractMultimodal Relation Extraction (MRE) is an entity relationship extraction method based on multimodal information. Most existing MRE methods have two issues: 1) Weak cross-modal correlation and poor semantic consistency. 2) They do not achieve text-guided fusion of different modalities, resulting in excessive introduction of image noise. To address these issues, we propose an innovative MRE method inspired by genetics-A Comprehensive Genetic-Inspired For Multimodal Relation Extraction (CGI-MRE). It consists of two main modules: Gene Extraction And Recombination Module (GERM) and Text-Guided Fusion Module (TGFM). In the GERM module, we regard the text features and visual features as a feature body respectively, and decompose each feature body into common sub-features and unique sub-features. For these sub-features, we designed a Common Gene Extraction Mechanism (CGEM) to extract common advantageous genes in different modalities, a Unique Gene Extraction Mechanism (UGEM) to extract unique advantageous genes in each modality, and we finally use a Gene Recombination Mechanism (GRM) to obtain recombinant features that highly correlated with different modalities and have strong semantic consistency. In TGFM module, we organically fuse and extract the features in the recombined features that are beneficial to MRE. We use gate to adjust the text-guided original attention score and pooling attention score to obtain the text-guided saliency attention score. We can use this score to strictly extract information that is text-guided and beneficial to MRE from the image recombinant feature. Experimental results on the MNRE dataset show that our model outperforms the state-of-the-art performance and achieves F1-score of 84.62%. Zhaokang Huang, Hongjun Ouyang, Qintai Hu, Bi Zeng |
ICMR | 5 |
| 2024 | VEC-MNER: Hybrid Transformer with Visual-Enhanced Cross-Modal Multi-level Interaction for Multimodal NERabstractMultimodal Named Entity Recognition (MNER) aims to leverage visual information to identify entity boundaries and categories in social media posts. Existing methods mainly adopt heterogeneous architecture, with ResNet (CNN-based) and BERT (Transformer-based) dedicated to modeling visual and textual features, respectively. However, current approaches still face the following issues: (1) Weak cross-modal correlations and poor semantic consistency. (2) Suboptimal fusion results when visual objects and textual entities are inconsistent. To this end, we propose a Hybrid Transformer with Visual-Enhanced Cross-Modal Multi-level Interaction (VEC-MNER) model for MNER. Specifically, compared to heterogeneous architectures, we propose a new homogeneous Hybrid Transformer Architecture, which naturally reduces the heterogeneity. Moreover, we design the Correlation-Aware Alignment (CAA-Encoder) layer and the Correlation-Aware Deep Fusion (CADF-Encoder) layer, combined with contrastive learning, to achieve more effective implicit alignment and deep semantic fusion between modalities, respectively. We also construct a Correlation-Aware (CA) module that can effectively reduce heterogeneity between modalities and alleviate visual deviation. Experimental results demonstrate that our approach achieves SOTA performance, achieving 74.89% and 87.51% F1-score on Twitter-2015 and Twitter-2017, respectively. Hongjun Ouyang, Qintai Hu, Bi Zeng, Qingpeng Wen |
ICMR | 4 |
| 2024 | Transformer-Mamba-Based Trident-Branch RGB-T Tracker
Yimin Du, Bi Zeng, Qingmao Wei, Boquan Zhang, Huiting Hu |
PRICAI (3) | 2 |
| 2024 | PLRTE: Progressive learning for biomedical relation triplet extraction using large language models
Yikai Zheng, Bi Zeng, Yi-Chun Feng |
J. Biomed. Informatics | 2 |
| 2024 | Visual Object Tracking With Mutual Affinity Aligned to Human IntuitionabstractSingle-object tracking generally advances by incrementally determining the tracked target's position through interactions between the search region and the template. However, the template provides less information than does the search region in terms of both temporal cues and spatial resolution. To alleviate this imbalance, we introduce an anthropic tracking framework, MATrack (Mutual Affinity Tracker), which explicitly strengthens weak template information and implicitly reduces background clutter through interactions between multiple templates and the search region. Additionally, we propose a coarse-to-fine localization approach that combines the benefits of corner-based and center-based methods. This approach enables us to simultaneously update the most recent state and background information without two-stage training. MATrack achieves state-of-the-art performance on multiple test benchmarks, including GOT-10k, LASOT, TrackingNet, OTB-100, UAV123, and NFS30. Among these benchmarks, MATrack-320's performance stands out, particularly in the short-term tracking dataset GOT-10k, where it achieves an accuracy overlap (AO) of 77.3. We also conduct comprehensive quantitative and qualitative evaluations to demonstrate that our method significantly outperforms other state-of-the-art approaches. Guotian Zeng, Bi Zeng, Qingmao Wei, Huiting Hu, Hong Zhang 0013 |
IEEE Trans. Multim. | 2 |
| 2023 | BFC-BL: Few-Shot Classification and Segmentation combining Bi-directional Feature Correlation and Boundary constraint
Haibiao Yang, Bi Zeng, Jianqi Liu |
BMVC | 2 |
| 2023 | Co-RGCN: A Bi-path GCN-Based Co-Regression Model for Multi-intent Detection and Slot Filling
Qingpeng Wen, Bi Zeng |
ICANN (4) | 2 |
| 2023 | BIG-FG: A Bi-directional Interaction Graph Framework with Filter Gate Mechanism for Chinese Spoken Language Understanding
Wentao Zhang 0011, Bi Zeng, Huiting Hu |
ICANN (4) | 2 |
| 2023 | A Deep Joint Model of Multi-scale Intent-Slots Interaction with Second-Order Gate for SLU
Qingpeng Wen, Bi Zeng, Huiting Hu |
ICONIP (12) | 2 |
| 2023 | Real-world efficient fall detection: Balancing performance and complexity with FDGA workflow
Guotian Zeng, Bi Zeng, Huiting Hu |
Comput. Vis. Image Underst. | 2 |
| 2023 | Multi-lane detection by combining line anchor and feature shift for urban traffic management
Jianqi Liu, Bin Deng 0005, Caifeng Zou, Bi Zeng, Jianxin Tan |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | PRAT: Accurate object tracking based on progressive attention
Yulin Zeng, Bi Zeng, Huiting Hu |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | SASE: Self-Adaptive noise distribution network for Speech Enhancement with Federated Learning using heterogeneous data
Zhentao Lin, Bi Zeng, Huiting Hu, Linwen Xu, Zhuangze Yao |
Knowl. Based Syst. | 2 |
| 2022 | Taxi demand forecasting based on the temporal multimodal information fusion graph neural network
Wenxiong Liao, Bi Zeng, Jianqi Liu, Pengfei Wei 0001, Xiaochun Cheng |
Appl. Intell. | 2 |
| 2022 | Image-text interaction graph neural network for image-text sentiment analysis
Wenxiong Liao, Bi Zeng, Jianqi Liu, Jiongkun Fang |
Appl. Intell. | 2 |
| 2022 | SiamPCF: siamese point regression with coarse-fine classification network for visual tracking
Yulin Zeng, Bi Zeng, Xiuwen Yin, Guangke Chen |
Appl. Intell. | 2 |
| 2021 | Evolutionary digital twin: A new approach for intelligent industrial product development
Tingyu Lin 0001, Zhengxuan Jia, Chen Yang 0011, Yingying Xiao, Shulin Lan, Guoqiang Shi, Bi Zeng, Heyu Li |
Adv. Eng. Informatics | 7 |
| 2021 | An improved aspect-category sentiment analysis model for text sentiment analysis based on RoBERTa
Wenxiong Liao, Bi Zeng, Xiuwen Yin |
Appl. Intell. | 2 |
| 2016 | A time-recordable cross-layer communication protocol for the positioning of Vehicular Cyber-Physical Systems
Jianqi Liu, Jiafu Wan, Bi Zeng, Shaoliang Fang |
Future Gener. Comput. Syst. | 4 |
| 2007 | Multi-robot Task Allocation Using Compound Emotion Algorithm
Bi Zeng |
APPT | 2 |
| 2004 | The discourse self-adapting fuzzy controller for temperature control processing in disinfecting cupboardabstractThis paper presents a changeable discourse range fuzzy controller, which can narrow down the discourse range and raise up the control precision. To change the range of discourse, the narrowing down factor is used automatically according to the output error of the system. This control method is implemented in electronic disinfecting cupboard one kind of home appliances, and the result shows that it is the useful method to control home appliance, which accurate model is difficult to obtain. Yongquan Yu, Huang Ying, Bi Zeng |
FUZZ-IEEE | 3 |
| 2004 | The dynamic fuzzy method to tune the weight factors of neural fuzzy PID controllerabstractA new method to modify the weight factors of PID neural network (PIDNN) in neural fuzzy PID controller is presented in this paper. The parameter fuzzy inference base (PFIB) is the structure to carry out the weight-value improving. The principle of PFIB is described and the neural fuzzy PID controller has been used in the steel tube pressure detecting system. The result of running shows that the neural fuzzy PID controller with PFIB has the better and satisfactory behavior for real time industrial control processing. Yongquan Yu, Huang Ying, Bi Zeng |
IJCNN | 3 |
| 2003 | Learning fuzzy model for nonlinear system using evolution strategies with adaptive direction mutationabstractThere are many methods to learn rules base for nonlinear system. The special method, the evolution strategies with adaptive direction, is presented in this paper. The actual steps and principle are also given. The simulation on results show this method is the effective one to learn rules base, and the faster generation rate can be obtained. Yongquan Yu, Huang Ying, Bi Zeng, Xianchu Chen |
FUZZ-IEEE | 3 |
| 2003 | A PID neural network controllerabstractIn this paper, the new fuzzy PID controller, which is combined fuzzy controller with PID neural network (PIDNN), is proposed. Its structure is difference from the normal one. The feature of it is to use a PIDNN replace PID parameter loop in controller. And the controller is optimized by the learning processing of PIDNN. The principle of PIDNN is discussed and the learning method based on back-propagation-algorithm is given. The two processes, first and second order systems, are simulated. Results of simulating show that the fuzzy PID controller presented in this paper is a better adaptive controller for linear or nonlinear plant. Yongquan Yu, Huang Ying, Bi Zeng |
IJCNN | 3 |
| 2002 | A real-time method to tune rules base of fuzzy control systemabstractThis paper presented the new method to tune the rule base in fuzzy control system. The consequent part fuzzy values are the important control values, that effect the performance of fuzzy system output. We proposed the real-time method to tune the consequent part fuzzy values using fuzzy arithmetic operations, according to the change status of system error, hence the rule base is regulated. The simulation result shows the method is valuable and easy to implement. Yongquan Yu, Bi Zeng, Guokun Zhong, Haixia Peng |
FUZZ-IEEE | 2 |