Junzhou Chen 0001

dblp:15/4682-1 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0002-3388-3503ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 6 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning Cell-Aware Hierarchical Multi-Modal Representations for Robust Molecular Modeling
abstract
Understanding how chemical perturbations propagate through biological systems is essential for robust molecular property prediction. While most existing methods focus on chemical structures alone, recent advances highlight the crucial role of cellular responses such as morphology and gene expression in shaping drug effects. However, current cell-aware approaches face two key limitations: (1) modality incompleteness in external biological data, and (2) insufficient modeling of hierarchical dependencies across molecular, cellular, and genomic levels. We propose CHMR (Cell-aware Hierarchical Multi-Modal Representations), a robust framework that jointly models local-global dependencies between molecules and cellular responses and captures latent biological hierarchies via a novel tree-structured vector quantization module. Evaluated on public benchmarks spanning 696 tasks, CHMR outperforms state-of-the-art baselines, yielding average improvements of 3.6% on classification and 17.2% on regression tasks. These results demonstrate the advantage of hierarchy-aware, multi-modal learning for reliable and biologically grounded molecular representations, offering a generalizable framework for integrative biomedical modeling.
Mengran Li 0001, Zelin Zang, Wenbin Xing, Junzhou Chen 0001, Jiebo Luo 0001, Stan Z. Li
AAAI4
2026 PGSF: Profile-Guided Semantic Fusion for Decoupled Industrial Pre-Ranking
Chuike Sun, Junzhou Chen 0001
SIGIR6
2026 A survey of large language models for data challenges in graphs
Mengran Li 0001, Wenbin Xing, Klim Zaporojets, Junzhou Chen 0001, Yong Zhang 0029, Siyuan Gong, Jia Hu 0003, Xiaolei Ma, Zhiyuan Liu 0002, Paul Groth, Marcel Worring
Expert Syst. Appl.6
2026 AttriReBoost: A Gradient-Free Propagation Optimization Method for Cold-Start Mitigation in Attribute Missing Graphs
abstract
In real-world graphs, node attributes are often incomplete due to acquisition costs or privacy restrictions, reducing representation quality and harming downstream predictions in graph neural networks (GNNs). A common remedy is feature-propagation-based imputation. However, cold-start effects arising from attribute resetting and low-degree nodes impede effective propagation and convergence in these methods. To address these challenges, we propose AttriReBoost (ARB), a propagation-based method that mitigates cold-start issues in attribute-missing graphs. ARB enhances global feature propagation (FP) by redefining initial boundary conditions and strategically integrating virtual edges, thereby improving node connectivity and ensuring stable and efficient convergence. The method supports gradient-free attribute reconstruction with low computational overhead, and we provide a rigorous convergence analysis. Extensive experiments on several real-world benchmark datasets demonstrate the effectiveness of ARB, achieving an average accuracy improvement of 5.11% over state-of-the-art methods. In addition, ARB exhibits remarkable computational efficiency, processing a large-scale graph with 2.44 million nodes in just 16 s on a single GPU. Our code is available at https://github.com/limengran98/ARB.
Mengran Li 0001, Chaojun Ding, Junzhou Chen 0001, Wenbin Xing, Cong Ye, Songlin Zhuang, Jia Hu 0003, Tony Z. Qiu, Huijun Gao
IEEE Trans. Cybern.3
2026 DSDFormer: An Innovative Transformer-Mamba Framework for Robust High-Precision Driver Distraction Identification
abstract
Driver distraction remains a leading cause of traffic accidents, posing a critical threat to road safety globally. As intelligent transportation systems evolve, accurate and real-time identification of driver distraction has become essential. However, existing methods struggle to capture both global contextual and fine-grained local features while contending with noisy labels in training datasets. To address these challenges, we propose DSDFormer, a novel framework that integrates the strengths of Transformer and Mamba architectures through a Dual State Domain Attention (DSDA) mechanism, enabling a balance between long-range dependencies and detailed feature extraction for robust driver behavior recognition. Additionally, we introduce Temporal Reasoning Confident Learning (TRCL), an unsupervised approach that refines noisy labels by leveraging spatiotemporal correlations in video sequences. Beyond achieving state-of-the-art results on AUC-V1, AUC-V2, and 100-Driver datasets, the proposed model is deployable in real-time on embedded platforms such as NVIDIA Jetson AGX Orin and Xavier. Extensive experimental results confirm that DSDFormer and TRCL significantly improve both the accuracy and robustness of driver distraction detection, offering a scalable solution to enhance road safety. Our code has been released athttps://github.com/zhangzr23/driver-noises-learning
Junzhou Chen 0001, Heqiang Huang, Xuemiao Xu, Bin Sheng 0001, Hong Yan 0001
IEEE Trans. Intell. Transp. Syst.1
2025 Appformer: A novel framework for mobile app usage prediction leveraging progressive multi-modal data fusion and feature extraction
Chuike Sun, Junzhou Chen 0001, Yue Zhao 0040, Ruihai Jing, Guang Tan, Di Wu 0001
Expert Syst. Appl.2
2025 BA-Net: Bridge attention in deep neural networks
Runzong Zou, Yue Zhao 0040, Junzhou Chen 0001, Yue Cao 0002, Chuan Hu 0002, Houbing Song
Expert Syst. Appl.5
2025 Topology-Driven Attribute Recovery for Attribute Missing Graph Learning in Social Internet of Things
abstract
With the advancement of information technology, the Social Internet of Things (SIoT) has fostered the integration of physical devices and social networks, deepening the study of complex interaction patterns. Text attribute graphs (TAGs) capture both topological structures and semantic attributes, enhancing the analysis of complex interactions within the SIoT. However, existing graph learning methods are typically designed for complete attributed graphs, and the common issue of missing attributes in attribute missing graphs (AMGs) increases the difficulty of analysis tasks. To address this, we propose the topology-driven attribute recovery (TDAR) framework, which leverages topological data for AMG learning. TDAR introduces an improved prefilling method for initial attribute recovery using native graph topology. Additionally, it dynamically adjusts propagation weights and incorporates homogeneity strategies within the embedding space to suit AMGs’ unique topological structures, effectively reducing noise during information propagation. Extensive experiments on public datasets demonstrate that TDAR significantly outperforms state-of-the-art methods in attribute reconstruction and downstream tasks, offering a robust solution to the challenges posed by AMGs. The code is available athttps://github.com/limengran98/TDAR.
Mengran Li 0001, Junzhou Chen 0001, Chenyun Yu, Guanying Jiang, Yanming Shen, Houbing Song
IEEE Internet Things J.2
2025 Real-Time Smoke Detection With Split Top-K Transformer and Adaptive Dark Channel Prior in Foggy Environments
abstract
Smoke detection is essential for fire prevention, yet it is significantly hampered by the visual similarities between smoke and fog. To address this challenge, a split top-k attention transformer framework (STKformer) is proposed. The STKformer incorporates split top-k attention (STKA), which partitions the attention map for top-k selection to retain informative self-attention values while capturing long-range dependencies. This approach effectively filters out irrelevant attention scores, preventing information loss. Furthermore, the adaptive dark-channel-prior guidance network (ADGN) is designed to enhance smoke recognition under foggy conditions. ADGN employs pooling operations instead of minimum value filtering, allowing for efficient dark channel extraction with learnable parameters and adaptively reducing the impact of fog. The extracted prior information subsequently guides feature extraction through a priorformer block, improving model robustness. Additionally, a cross-stage fusion module (CSFM) is introduced to aggregate features from different stages efficiently, enabling flexible adaptation to smoke features at various scales and enhancing detection accuracy. Comprehensive experiments demonstrate that the proposed method achieves state-of-the-art performance across multiple datasets, with an accuracy of 89.68% on dataset for smoke detection in fog, 99.76% on CCTV images of smoke, and 99.76% on UAV images of wildfire. The method maintains high speed and lightweight characteristics, validated with an inference speed of 211.46 FPS on an NVIDIA Jetson AGX Orin after TensorRT acceleration, confirming its effectiveness and efficiency for real-world applications. The source code is available athttps://github.com/Jiongze-Yu/STKformerhttps://github.com/Jiongze-Yu/STKformer.
Jiongze Yu, Heqiang Huang, Yuhang Ma 0002, Yueying Wu 0001, Junzhou Chen 0001, Xuemiao Xu, Zhihan Lyu, Guodong Yin
IEEE Internet Things J.5
2025 TDG-Mamba: Advanced Spatiotemporal Embedding for Temporal Dynamic Graph Learning via Bidirectional Information Propagation
abstract
Temporal dynamic graphs (TDGs), representing the dynamic evolution of entities and their relationships over time with intricate temporal features, are widely used in various real-world domains. Existing methods typically rely on mainstream techniques such as transformers and graph neural networks (GNNs) to capture the spatiotemporal information of TDGs. However, despite their advanced capabilities, these methods often struggle with significant computational complexity and limited ability to capture temporal dynamic contextual relationships. Recently, a new model architecture called mamba has emerged, noted for its capability to capture complex dependencies in sequences while significantly reducing computational complexity. Building on this, we propose a novel method, TDG-mamba, which integrates mamba for TDG learning. TDG-mamba introduces deep semantic spatiotemporal embeddings into the mamba architecture through a specially designed spatiotemporal prior tokenization module (SPTM). Furthermore, to better leverage temporal information differences and enhance the modeling of dynamic changes in graph structures, we separately design a bidirectional mamba and a directed GNN for improved spatiotemporal embedding learning. Link prediction experiments on multiple public datasets demonstrate that our method delivers superior performance, with an average improvement of 5.11% over baseline methods across various settings.
Mengran Li 0001, Junzhou Chen 0001, Bo Li 0128, Yong Zhang 0029, Siyuan Gong, Xiaolei Ma, Zhihong Tian 0001
IEEE Trans. Comput. Soc. Syst.2
2025 YOLO-TS: Real-Time Traffic Sign Detection With Enhanced Accuracy Using Optimized Receptive Fields and Anchor-Free Fusion
abstract
Ensuring safety in both autonomous driving and advanced driver-assistance systems (ADAS) depends critically on the efficient deployment of traffic sign recognition technology. While current methods show effectiveness, they often compromise between speed and accuracy. To address this issue, we present a novel real-time and efficient road sign detection network, YOLO-TS. This network significantly improves performance by optimizing the receptive fields of multi-scale feature maps to align more closely with the size distribution of traffic signs in various datasets. Moreover, our innovative feature-fusion strategy, leveraging the flexibility of Anchor-Free methods, allows for multi-scale object detection on a high-resolution feature map abundant in contextual information, achieving remarkable enhancements in both accuracy and speed. To mitigate the adverse effects of the grid pattern caused by dilated convolutions on the detection of smaller objects, we have devised a unique module that not only mitigates this grid effect but also widens the receptive field to encompass an extensive range of spatial contextual information, thus boosting the efficiency of information usage. Moreover, to address the scarcity of traffic sign datasets, especially under adverse weather conditions, we introduce two novel datasets: Generated-TT100K-weather and CAWTSSS. Extensive evaluations conducted on challenging public benchmarks—including TT100K, CCTSDB2021, and GTSDB—as well as on our proposed datasets, demonstrate that YOLO-TS surpasses current state-of-the-art methods in both accuracy and inference speed. The code, datasets and weights are available athttps://github.com/Heqiang-Huang/YOLO-TS
Junzhou Chen 0001, Heqiang Huang, Nengchao Lyu, Yanyong Guo, Hongning Dai, Hong Yan 0001
IEEE Trans. Intell. Transp. Syst.1
2025 MM-STFlowNet: A Transportation Hub-Oriented Multi-Mode Passenger Flow Prediction Method via Spatial-Temporal Dynamic Graph Modeling
abstract
Accurate and refined passenger flow prediction is essential for optimizing the collaborative management of multiple collection and distribution modes in large-scale transportation hubs. Traditional methods often focus only on the overall passenger volume, neglecting the interdependence between different modes within the hub. To address this limitation, we propose MM-STFlowNet, a comprehensive multi-mode prediction framework grounded in dynamic spatial-temporal graph modeling. Initially, an integrated temporal feature processing strategy is implemented using signal decomposition and convolution techniques to address data spikes and high volatility. Subsequently, we introduce the Spatial-Temporal Dynamic Graph Convolutional Recurrent Network (STDGCRN) to capture detailed spatial-temporal dependencies across multiple traffic modes, enhanced by an adaptive channel attention mechanism. Finally, the self-attention mechanism is applied to incorporate various external factors, further enhancing prediction accuracy. Experiments on a real-world dataset from Guangzhounan Railway Station in China demonstrate that MM-STFlowNet achieves state-of-the-art performance, with an average improvement of 52.56% in MSE and 36.38% in MAE. Especially during peak hours, it demonstrates excellent forecasting performance, providing valuable insights for transportation hub management. Our model is also demonstrated strong generalization in low-resource scenarios and different traffic scenarios. Our code is available at https://github.com/BMRETURN/MM-STFlowNet
Wenbin Xing, Mengran Li 0001, Junzhou Chen 0001, Xiaolei Ma, Zhiyuan Liu 0002, Zhengbing He
IEEE Trans. Intell. Transp. Syst.5
2025 AGSENet: A Robust Road Ponding Detection Method for Proactive Traffic Safety
abstract
Road ponding, a prevalent traffic hazard, poses a serious threat to road safety by causing vehicles to lose control and leading to accidents ranging from minor fender benders to severe collisions. Existing technologies struggle to accurately identify road ponding due to complex road textures and variable ponding coloration influenced by reflection characteristics. To address this challenge, we propose a novel approach called Self-Attention-based Global Saliency-Enhanced Network (AGSENet) for proactive road ponding detection and traffic safety improvement. AGSENet incorporates saliency detection techniques through the Channel Saliency Information Focus (CSIF) and Spatial Saliency Information Enhancement (SSIE) modules. The CSIF module, integrated into the encoder, employs self-attention to highlight similar features by fusing spatial and channel information. The SSIE module, embedded in the decoder, refines edge features and reduces noise by leveraging correlations across different feature levels. To ensure accurate and reliable evaluation, we corrected significant mislabeling and missing annotations in the Puddle-1000 dataset. Additionally, we constructed the Foggy-Puddle and Night-Puddle datasets for road ponding detection in low-light and foggy conditions, respectively. Experimental results demonstrate that AGSENet outperforms existing methods, achieving IoU improvements of 2.03%, 0.62%, and 1.06% on the Puddle-1000, Foggy-Puddle, and Night-Puddle datasets, respectively, setting a new state-of-the-art in this field. Finally, we verified the algorithm’s reliability on edge computing devices. This work provides a valuable reference for proactive warning research in road traffic safety. The source code and datasets are placed in thehttps://github.com/Lyu-Dakang/AGSENet.
Shangyu Yang, Dakang Lyu, Junzhou Chen 0001, Yilong Ren, Bolin Gao, Zhihan Lyu
IEEE Trans. Intell. Transp. Syst.5
2025 GrabDAE: An Innovative Framework for Unsupervised Domain Adaptation Utilizing Grab-Mask and Denoise Auto-Encoder
abstract
Existing Unsupervised Domain Adaptation (UDA) methods often fall short in fully leveraging contextual information from the target domain, leading to suboptimal decision boundary separation during source and target domain alignment. To address this, we introduce GrabDAE, an innovative UDA framework designed to tackle domain shift in visual classification tasks. GrabDAE incorporates two key innovations: the Grab-Mask module, which blurs background information in target domain images, enabling the model to focus on essential, domain-relevant features through contrastive learning; and the Denoising Auto-Encoder (DAE), which enhances feature alignment by reconstructing features and filtering noise, ensuring a more robust adaptation to the target domain. These components empower GrabDAE to effectively handle unlabeled target domain data, significantly improving both classification accuracy and robustness. Extensive experiments on benchmark datasets, including VisDA-2017, Office-Home, and Office31, demonstrate that GrabDAE consistently surpasses state-of-the-art UDA methods, setting new performance benchmarks. By tackling UDA's critical challenges with its novel feature masking and denoising approach, GrabDAE offers both significant theoretical and practical advancements in domain adaptation.
Junzhou Chen 0001, Xuan Wen, Bingtao Ren, Di Wu 0001, Zhigang Xu 0001, Danwei Wang
IEEE Trans. Multim.1
2024 A robust and real-time lane detection method in low-light scenarios to advanced driver assistance systems
Jingtao Peng, Wanting Gou, Yuhang Ma 0002, Junzhou Chen 0001, Hongyu Hu, Weihua Li 0004, Guodong Yin, Zhiwu Li 0001
Expert Syst. Appl.5
2024 TA-NET: Empowering Highly Efficient Traffic Anomaly Detection Through Multi-Head Local Self-Attention and Adaptive Hierarchical Feature Reconstruction
abstract
In the realm of road surveillance systems, Automatic Incident Detection (AID) methods have shown promise in swiftly and precisely detecting traffic anomalies. Nevertheless, the paucity of frame-level precisely annotated training data poses substantial challenges. To navigate this issue, we introduce TA-NET, a novel framework designed to highly efficient traffic anomaly detection. TA-NET utilizes a generic dataset pre-training model to facilitate weakly supervised learning. It comprises two main modules: the Adaptive Hierarchical Feature Reconstruction Block (AHFRB) and the Multi-Head Local Self-Attention (MHLSA) mechanism. AHFRB refines the feature extractor by reconstructing the features drawn from the pre-trained model. Concurrently, MHLSA scrutinizes the contextual relationships between contiguous video segments and bolsters anomaly detection accuracy. Our approach was validated using the TAD testing dataset, with the results highlighting the efficacy of TA-NET. Remarkably, it accomplishes an Area Under the Curve (AUC) of 94.47% for the overall dataset and 70.78% for the anomaly subset. These scores surpass the previous state-of-the-art method by 1.54% and 4.96%, respectively, attesting to TA-NET’s superior performance. Consequently, our study offers a fresh perspective and sets a new standard for future video-based traffic incident detection research.
Junzhou Chen 0001, Jiajun Pu, Baiqiao Yin, Jun Jie Wu
IEEE Trans. Intell. Transp. Syst.1
2024 A Prior Guided Wavelet-Spatial Dual Attention Transformer Framework for Heavy Rain Image Restoration
abstract
Heavy rain significantly reduces image visibility, hindering tasks like autonomous driving and video surveillance. Many existing rain removal methods, while effective in light rain, falter under heavy rain due to their reliance on purely spatial features. Recognizing this challenge, we introduce the Wavelet-Spatial Dual Attention Transformer Framework (WSDformer). This innovative architecture adeptly captures both frequency and spatial characteristics, anchored by the wavelet-spatial dual attention (WSDA) mechanism. While the spatial attention zeroes in on intricate local details, the wavelet attention leverages wavelet decomposition to encompass diverse frequency information, augmenting the spatial representations. Furthermore, addressing the persistent issue of incomplete structural detail restoration, we integrate the PriorFormer Block (PFB). This unique module, underpinned by the Prior Fusion Attention (PFA), synergizes residual channel prior features with input features, thereby enhancing background structures and guiding precise rain feature extraction. To navigate the intrinsic constraints of U-shaped transformers, such as semantic discontinuities and subdued multi-scale interactions from skip connections, our Cross Interaction U-Shaped Transformer Network is introduced. This design empowers superior semantic layers to streamline the extraction of their lower-tier counterparts, optimizing network learning. Empirical analysis reveals our method's leading prowess across rainy image datasets and achieves state-of-the-art performance, with notable supremacy in heavy rainfall conditions. This superiority extends to diverse visual challenges and real-world rainy scenarios, affirming its broad applicability and robustness. The source code is available athttps://github.com/Jiongze-Yu/WSDformer.
Jiongze Yu, Junzhou Chen 0001, Guofa Li, Liang Lin 0004, Danwei Wang
IEEE Trans. Multim.3
2022 BA-Net: Bridge Attention for Deep Convolutional Neural Networks
Yue Zhao 0040, Junzhou Chen 0001
ECCV (21)2
2022 A real-time and high-precision method for small traffic-signs recognition
Junzhou Chen 0001, Kunkun Jia, Wenquan Chen, Zhihan Lyu
Neural Comput. Appl.1
2021 Face recognition based on adaptive margin and diversity regularization constraints
abstract
Abstract In recent years, a more robust facial feature can be learned by convolutional neural networks once introducing margins into loss functions. Those methods set a margin for each class manually to squeeze the intra‐class variations within each class equally. However, the internal feature distributions of different persons in the real world are highly unbalanced, and the distance between different identities is not uniform either. As a result, applying the same margin on all classes might not lead to higher inter‐class differences. To address this problem, this paper proposes an adaptive margin based on feature distribution to squeeze the feature interior spaces of different classes. Simultaneously, because the inter‐class margin can adequately represent the distribution of different classes in the feature space, this paper proposes a novel diversity regularization method. The regularization weights of each class are dynamically set depending on their margins. This method proposed in this paper is intuitively interpretable and can be easily applied to other classification scenarios. Experiments on current existing benchmarks have demonstrated the superiority of our method over state‐of‐the‐art competitors.
Zhemin Zhang, Xun Gong 0002, Junzhou Chen 0001
IET Image Process.3
2020 Point clouds learning with attention-based graph convolution networks
Zhuyang Xie, Junzhou Chen 0001, Bo Peng 0006
Neurocomputing2
2016 Automatic Hookworm Detection in Wireless Capsule Endoscopy Images
abstract
Wireless capsule endoscopy (WCE) has become a widely used diagnostic technique to examine inflammatory bowel diseases and disorders. As one of the most common human helminths, hookworm is a kind of small tubular structure with grayish white or pinkish semi-transparent body, which is with a number of 600 million people infection around the world. Automatic hookworm detection is a challenging task due to poor quality of images, presence of extraneous matters, complex structure of gastrointestinal, and diverse appearances in terms of color and texture. This is the first few works to comprehensively explore the automatic hookworm detection for WCE images. To capture the properties of hookworms, the multi scale dual matched filter is first applied to detect the location of tubular structure. Piecewise parallel region detection method is then proposed to identify the potential regions having hookworm bodies. To discriminate the unique visual features for different components of gastrointestinal, the histogram of average intensity is proposed to represent their properties. In order to deal with the problem of imbalance data, Rusboost is deployed to classify WCE images. Experiments on a diverse and large scale dataset with 440 K WCE images demonstrate that the proposed approach achieves a promising performance and outperforms the state-of-the-art methods. Moreover, the high sensitivity in detecting hookworms indicates the potential of our approach for future clinical application.
Xiao Wu 0001, Honghan Chen, Tao Gan, Junzhou Chen 0001, Chong-Wah Ngo, Qiang Peng
IEEE Trans. Medical Imaging4
2015 CSIFT based locality-constrained linear coding for image classification
Junzhou Chen 0001, Qing Li 0058, Qiang Peng, Kin Hong Wong
Pattern Anal. Appl.1
2009 Controlling Virtual Cameras Based on a Robust Model-Free Pose Acquisition Technique
abstract
This paper presents a novel method that acquires camera position and orientation from a stereo image sequence without prior knowledge of the scene. To make the algorithm robust, the interacting multiple model probabilistic data association filter (IMMPDAF) is introduced. The interacting multiple model (IMM) technique allows the existence of more than one dynamic system in the filtering process and in return leads to improved accuracy and stability even under abrupt motion changes. The probabilistic data association (PDA) framework makes the automatic selection of measurement sets possible, resulting in enhanced robustness to occlusions and moving objects. In addition to the IMMPDAF, the trifocal tensor is employed in the computation so that the step of reconstructing the 3-D models can be eliminated. This further guarantees the precision of estimation and computation efficiency. Real stereo image sequences have been used to test the proposed method in the experiment. The recovered 3-D motions are accurate in comparison with the ground truth data and have been applied to control cameras in a virtual environment.
Ying Kin Yu, Kin Hong Wong, Siu-Hang Or, Junzhou Chen 0001
IEEE Trans. Multim.4
2008 Calibration of an Articulated Camera System
abstract
Multiple Camera Systems (MCS) have been widely used in many vision applications and attracted much attention recently. There are two principle types of MCS, one is the Rigid Multiple Camera System (RMCS); the other is the Articulated Camera System (ACS). In a RMCS, the relative poses (relative 3-D position and orientation) between the cameras are invariant. While, in an ACS, the cameras are articulated through movable joints, the relative pose between them may change. Therefore, through calibration of an ACS we want to find not only the relative poses between the cameras but also the positions of the joints in the ACS. Although calibration methods for RMCS have been extensively developed during the past decades, the studies of ACS calibration are still rare. In this paper, two ACS calibration methods are proposed. The first one uses the feature correspondences between the cameras in the ACS. The second one requires only the ego-motion information of the cameras and can be used for the calibration of the non-overlapping view ACS. In both methods, the ACS is assumed to have performed general transformations in a static environment. The efficiency and robustness of the proposed methods are tested by simulation and real experiments. In the real experiment, the intrinsic and extrinsic parameters of the ACS are calibrated using the same image sequences, no extra data capturing step is required. The corresponding trajectory is recovered and illustrated using the calibration results of the ACS. To our knowledge, we are the first to study the calibration of ACS.
Junzhou Chen 0001, Kin Hong Wong
CVPR1
2008 EKF pose estimation: How many filters and cameras to use?
abstract
The extended Kalman filter (EKF) is suitable for real-time pose estimation due its low computational demand and ability to handle the nonlinear perspective camera model. There are many EKF based approaches in the literature; some are very recent while others exist for about two decades. These methods differ in two main aspects: the number and arrangement of cameras, and the number and usage of filters. In this work, we will compare these approaches using simulations and real experiments. As far as we know, it is the first attempt to do this with such details. We will show which is suitable under different motion patterns, and explain the effect of the bas-relief ambiguity upon the accuracy of the different approaches. Additionally, we will discuss how to solve the scale factor ambiguity, and suggest the best strategy to deal with the features fed to the filter.
Mohammad Ehab Ragab, Kin Hong Wong, Junzhou Chen 0001, Michael Ming Yuen Chang
ICIP3
2007 EKF Based Pose Estimation using Two Back-to-Back Stereo Pairs
abstract
In this work, we solve the pose estimation problem for robot motion by placing multiple cameras on the robot. In particular, we use four cameras arranged as two back-to-back stereo pairs combined with the extended Kalman filter (EKF). The reason for using multiple cameras is that the pose estimation problem is more constrained for multiple cameras than for a single camera. Back-to-back cameras are used since they provide more information. Stereo information is used in self initialization and outlier rejection. Different approaches to solve the long-sequence-drift have been suggested. Both the simulations and the real experiments show that our approach is fast, robust, and accurate.
Mohammad Ehab Ragab, Kin Hong Wong, Junzhou Chen 0001, Michael Ming Yuen Chang
ICIP (6)3