VLDB 2026 Research / reviewers in the wild / expert
Sixian Chan 0001
dblp:176/6694-1 · also Chan Si Xian 0001
· DBLP profile ↗
51ranked-venue papers
22as first author
45since 2021 · last 2026
0000-0001-8916-1174ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 21 since 2021Artificial intelligence and machine learning · 22 · 9 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoStruct: Intelligent design system for shear wall building structures
Sixian Chan 0001, Yage Xia, Jiafa Mao, Chao Li 0050 |
Adv. Eng. Informatics | 1 |
| 2026 | Prototype-based latent space distance optimization on vehicle re-identification
Sixian Chan 0001, Jiaao Cui, Zheng Wang 0059, Xiaolong Zhou 0001, Xiaoqin Zhang 0002 |
Expert Syst. Appl. | 1 |
| 2026 | DSSA-depth: Unsupervised monocular depth estimation method based on dual-scale self-attention
Jiafa Mao, Dingkai Yao, Yahong Hu, Sixian Chan 0001, Weiguo Sheng 0001, Hanlin Qin |
Neurocomputing | 4 |
| 2026 | S3Mamba: Region-aware spatio-semantic Mamba for remote sensing change detection
Yuan Wang 0032, Sixian Chan 0001, Guoyu Yang, Tianyang Dong, Xiaoqin Zhang 0002 |
Pattern Recognit. | 2 |
| 2026 | Bi-CFNet: A Multi-Task Audiovisual Framework With Visual Enhancement and Cross-Feedback for Depression DetectionabstractDepressive disorder has evolved into a major public health issue. However, its diagnosis is predominantly based on subjective clinical evaluation, leading to frequent misdiagnosis. Although multimodal methods have demonstrated potential, they still remain limited in their ability to model long-range temporal dependencies, reconcile semantic heterogeneity across modalities (e.g., facial landmarksvs.gaze behavior), and capture deep inter-modal interactions. To address these challenges, we propose a novel multi-task audiovisual network, termed Bi-CFNet, which is built upon a powerful temporal modeling backbone to capture long-range dependencies and integrates visual enhancement and cross-feedback for depression detection. In addition to the primary binary classification task, we establish a multi-task learning framework to predict item-level clinical scores as an auxiliary task. This design provides supplementary supervisory signals, encouraging the model to learn semantically meaningful representations and mitigate the impact of subjective bias. The visual enhancement module (VEM) employs gated attention to integrate gaze vectors with facial features, yielding temporally aligned and semantically enriched visual representations that are more sensitive to depressive indicators. Furthermore, the cross-feedback mechanism (CFM) facilitates the injection of highlevel semantic features from one modality into the early layers of another, thereby integrating global temporal context into local processing and promoting deep cross-modal re-learning. Finally, extensive experiments on the DAIC-WOZ dataset show that our model achieves an F1-score of 0.85 and a recall of 0.87, surpassing several state-of-the-art trimodal baselines. The proposed framework demonstrates the potential of utilizing audiovisual cues for objective depression analysis, providing a technical foundation for future computer-aided clinical screening systems. Sixian Chan 0001, Yuxuan Zhai, Yuan Wang 0032, Jiafa Mao |
IEEE Trans. Affect. Comput. | 1 |
| 2026 | SAM-SS: Straightforward and Efficient Designs Based on Segment Anything Model for Semantic SegmentationabstractImage segmentation is a fundamental task in computer vision and computational social systems, with semantic segmentation aiming to assign each pixel to a corresponding label. Due to the inherent richness of categories and contextual information in images, image segmentation remains a challenging problem. Currently, semantic segmentation models based on the segment anything model have demonstrated promising results. However, they continue to encounter challenges related to training strategies and prompt information generation. To address these issues, we propose a straightforward and efficient design method for semantic segmentation based on a prompt-free model, named SAM-SS. First, we introduce the class prompt encoder, which generates category prompts for the mask decoder to extract category-specific semantic information. Second, we incorporate the deep fusion module to bridge the semantic gap for achieving robust representation. Additionally, we observe that fine-tuning the image encoder via low-rank adaptation often leads to suboptimal convergence. To mitigate this, we propose a learning rate modulation strategy to stabilize training and boost model performance. Finally, we validate our model’s performance on three publicly available datasets. Specifically, on the Cityscapes validation set for natural images, our model achieves a mean intersection over union (mIoU) of 85.28%. Moreover, our model demonstrates strong adaptability to remote sensing imagery, achieving mIoU scores of 54.46% on LoveDA and 80.8% on the ISPRS Potsdam validation sets. These results underscore its potential utility across diverse applications. Yalin Wang 0012, Hong Peng 0003, Weihao Zheng, Zhongfeng Kang, Sixian Chan 0001 |
IEEE Trans. Comput. Soc. Syst. | 7 |
| 2026 | GVLTrack: Global Vision-Language Tracking with Multi-Stage Modal FusionabstractIn general, local Visual-Language (VL) trackers search targets around the previous bounding box by initial VL annotations. However, there is an inherent contradiction between the local searching perspective of the tracker and the orientation descriptions in language conducted under the global perspective. Furthermore, most methods only fuse modality information in a single stage, which tends to an insufficient relation modeling. To address these issues, we propose a Global Vision-Language Tracker (GVLTrack) with multi-stage modal fusion. First, it tracks the target in the entire image instead of local tracking based on previous results to resolve the above contradiction. Second, GVLTrack incorporates three modal interaction modules: Consistent Relationship Modeling (CRM), VL-Guided Query (VLQ) Initialization, and Recurrent Cross-Modal Decoder (RC-Decoder) to fuse vision-language modality and refine the bounding box progressively comprehensively. We conduct extensive experiments on several benchmarks and achieve competitive performance, demonstrating the effectiveness of our approach. The code will be made publicly available as soon as it is accepted. Sixian Chan 0001, Cong Bai, Xiaoqin Zhang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | SGFormer: Semantic-Geometry Fusion Transformer for Multi-modal 3D Panoptic SegmentationabstractModern methods for autonomous driving perception widely adopt multi-modal fusion to enhance 3D scene understanding. However, existing methods suffer from inferior semantic extraction in image encoders that treat all pixels equally, ignoring contextual differences. The generated multi-modal representations also typically lack comprehensive semantic and spatial geometry information, which is crucial for the 3D panoptic segmentation task. In this paper, we propose a novel Semantic-Geometry Fusion Transformer (SGFormer) that extracts adaptive semantic contexts, aggregates geometric information and captures the semantic-geometry fusion. First, in the Image Branch, we tailor semantic contexts for each pixel with context-guided attention and spatial context alignment to refine semantic details. Second, we transform image and voxel features into point-pixel geometry representations, simultaneously learning semantic category priors as embeddings to better represent scene geometry and semantics. Finally, to aggregate semantic information with related geometry, we design a semantic-geometry fusion that combines the transformer, effectively capturing semantic-geometry relationships into multi-modal panoptic representations. Notably, SGFormer achieves the state-of-the-art (SOTA) results on the nuScenes and SemanticPOSS, as well as yielding competitive performance on the SemanticKITTI. Moreover, SGFormer exhibits superior robustness compared to leading methods, marking an improvement of 2% to 10%. Hongqi Yu, Sixian Chan 0001, Xiaolong Zhou 0001, Xiaoqin Zhang 0002 |
AAAI | 2 |
| 2025 | NeighborRetr: Balancing Hub Centrality in Cross-Modal RetrievalabstractCross-modal retrieval aims to bridge the semantic gap between different modalities, such as visual and textual data, enabling accurate retrieval across them. Despite significant advancements with models like CLIP that align cross-modal representations, a persistent challenge remains: the hubness problem, where a small subset of samples (hubs) dominate as nearest neighbors, leading to biased representations and degraded retrieval accuracy. Existing methods often mitigate hubness through post-hoc normalization techniques, relying on prior data distributions that may not be practical in real-world scenarios. In this paper, we directly mitigate hubness during training and introduce NeighborRetr, a novel method that effectively balances the learning of hubs and adaptively adjusts the relations of various kinds of neighbors. Our approach not only mitigates the hubness problem but also enhances retrieval performance, achieving state-of-the-art results on multiple cross-modal retrieval benchmarks. Furthermore, Neighbor-Retr demonstrates robust generalization to new domains with substantial distribution shifts, highlighting its effectiveness in real-world applications. We make our code publicly available at: https://github.com/NeighborRetr. Zengrong Lin, Zheng Wang 0059, Tianwen Qian, Pan Mu, Sixian Chan 0001, Cong Bai |
CVPR | 5 |
| 2025 | IEEnhanceNet: An Implicit and Explicit Enhancement-Based Framework for Malformed Tooth Segmentation in 3D Intraoral Scan DataabstractAccurate segmentation of malformed teeth from 3D intraoral scans is critical for computer-aided orthodontic diagnosis. While deep learning-based methods have advanced dental arch segmentation, their performance deteriorates when applied to malformed teeth, often resulting in incomplete shape segmentation and pronounced speckle noise. To address these limitations, we propose IEEnhanceNet, a novel 3D tooth segmentation framework featuring two key innovations: (1) An Implicit Enhancement Module (IEM) that leverages full-arch segmentation data to implicitly constrain the loss function, synergized with an Adaptive Channel Attention mechanism (ACA) to improve morphological feature capture; and (2) An Explicit Enhancement Module (EEM) that introduces architectural-level constraints through an auxiliary full-arch segmentation branch, effectively suppressing noise while enhancing fine-grained accuracy. Specifically, for method validation, we reannotate the 3D-IOSSeg dataset with specialized labels for malformed teeth, normal teeth, and gingiva—the first such effort focused on malformed tooth segmentation. Finally, extensive experiments demonstrate that IEEn-hanceNet achieves state-of-the-art performance, significantly outperforming existing methods in both Overall Accuracy (OA) and Mean Intersection over Union (mIoU) metrics. This work provides a robust computational tool for orthodontic applications involving dental anomalies. Qitao Shan, Sixian Chan 0001, Jiafa Mao, Jingke Gu, Binyao Kong |
ECAI | 2 |
| 2025 | SMStracker: Tri-Path Score Mask Sigma Fusion for Multi-Modal Tracking
Sixian Chan 0001, Zedong Li, Shijian Lu, Chunhua Shen, Xiaoqin Zhang 0002 |
ICCV | 1 |
| 2025 | Towards Lightweight and Robust MCMT Tracking: Dual-Retrieval Knowledge Distillation and Scene-Aware FusionabstractMulti-Camera Multi-Target Tracking (MCMT) aims to achieve robust identity association of objects across cameras under diverse and challenging real-world conditions, such as varying viewpoints and occlusions. Current mainstream approaches often rely on deploying large-scale models with extensive feature extraction capabilities. However, the high computational demands of these models make them impractical for real-world MCMT scenarios where efficiency is critical. To address these challenges, we propose a novel Dual-Retrieval Knowledge Distillation (DRKD) framework, which enhances the student’s feature representation learning by leveraging multi-teacher guidance and dual-retrieval optimization. Unlike traditional knowledge distillation (KD) methods focusing primarily on feature alignment, DRKD introduces the Cross Triplet Loss, which optimizes the feature space for cross-camera identity association by enhancing intra-class compactness and inter-class separability. This dual-retrieval optimization ensures that the student model learns from the teacher’s feature representations and develops strong retrieval capabilities, which are crucial for robust identity association in MCMT. Additionally, we present the Dynamic Feature Fusion (DFF) module, which integrates short-term and historical features to balance the trade-off between responsiveness in single-camera tracking (SCT) and stability in multi-camera tracking (MCT). To further improve adaptability, we design Scene-Aware Regularization, which dynamically adjusts feature contributions in DFF based on temporal gaps and appearance variations. Extensive experiments on the HST and AI City Challenge S02 datasets demonstrate the effectiveness of our approach. The proposed DRKD framework with DFF achieves state-of-the-art performance, with IDF1 scores of 67.13 and MOTA scores of 60.50, while maintaining high inference speeds of up to 62 FPS. These results highlight the ability of DRKD to combine accuracy, robustness, and efficiency, making it a promising solution for real-world MCMT applications. Sixian Chan 0001, Xiaoxiang Chen, Wei Wang 0307, Jiafa Mao, Jie Hu 0041 |
IJCNN | 1 |
| 2025 | PolypSense3D: A Multi-Source Benchmark Dataset for Depth-Aware Polyp Size Measurement in EndoscopyabstractAccurate polyp sizing during endoscopy is crucial for cancer risk assessment but is hindered by subjective methods and inadequate datasets lacking integrated 2D appearance, 3D structure, and real-world size information. We introduce PolypSense3D, the first multi-source benchmark dataset specifically targeting depth-aware polyp size measurement. It uniquely integrates over 43,000 frames from virtual simulations, physical phantoms, and clinical sequences, providing synchronized RGB, dense/sparse depth, segmentation masks, camera parameters, and millimeter-scale size labels derived via a novel forceps-assisted in-vivo annotation technique. To establish its value, we benchmark state-of-the-art segmentation and depth estimation models. Results quantify significant domain gaps between simulated/phantom and clinical data and reveal substantial error propagation from perception stages to final size estimation, with the best fully automated pipelines achieving an average Mean Absolute Error (MAE) of 0.95 mm on the clinical data subset. Publicly released under CC BY-SA 4.0 with code and evaluation protocols, PolypSense3D offers a standardized platform to accelerate research in robust, clinically relevant quantitative endoscopic vision. The benchmark dataset and code are available at: https://github.com/HNUicda/PolypSense3D and https://doi.org/10.7910/DVN/K13H89. Ruyu Liu, Mingming Zhou, Jianhua Zhang 0002, Xiufeng Liu 0001, Xu Cheng 0003, Sixian Chan 0001, Yanbin Shen, Sheng Dai, Yuping Yan, Yaochu Jin, Lingjuan Lyu |
NeurIPS | 8 |
| 2025 | FocTrack: Focus attention for visual tracking
Sixian Chan 0001, Zhenchao Shi, Cong Bai, Shengyong Chen |
Pattern Recognit. | 2 |
| 2025 | Deformable Blur Sensing and Regression Analysis ReID Feature Fusion for Multitarget Multicamera Tracking Systems in Highway ScenariosabstractIn highway scenarios, the rapid motion of vehicles can cause deformation and blur in camera footage, significantly affecting the accuracy of vehicle detection and re-identification (ReID) in multitarget multicamera tracking (MTMCT) systems. To address this issue, this article develops the deformable and blur sensing and regression analysis ReID feature fusion MTMCT system (DSRF). First, a deformable and blur sensing detection module (DFB) in DSRF is designed to overcome the limitations of cameras in capturing fast-moving objects, thereby accurately detecting vehicles moving at high speeds on highways. Then, a regression-based ReID feature fusion algorithm (RARF) in DSRF is proposed, which enhances ReID features by modeling the relationship between vehicle motion and its features, thereby better associating the detected vehicles in consecutive frames into trajectories and establishing intertrajectory relationships. Finally, extensive experiments are conducted on the highway surveillance traffic (HST) dataset developed by our team and the public dataset (CityFlow). Promising results are achieved, validating the effectiveness of our proposed method. Sixian Chan 0001, Shenghao Ni, Jie Hu 0041, Tinglong Tang, Xiaolong Zhou 0001, Pengyi Hao |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2025 | DBFA-TSNet: A Three-Stage Building Extraction Network Based on Dual-Branch Fusion and Adaptive Enhancement
Jianan Chen 0002, Sixian Chan 0001, Cong Bai |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Adaptive Target-Oriented TrackingabstractThe current one-stream tracking pipelines are early relation modeling in feature extraction. However, insufficient discrimination may result in ambiguous relation modeling during early feature extraction. Moreover, the non-target information occupies most of the search image, rendering most relation modeling futile. To tackle the above issues, we propose tracking via learning adaptive target-oriented representation, named ATOTrack . We design an Untied positional encoding to mark the template token and the search region token separately, which reduces the confused relationship between the template and the search region. Besides, we introduce an Auto-Mask Learner to decouple the target and non-target information in the search region. Interestingly, the Auto-Mask Learner can self-learn and mask the ineffective information to interpret adaptive target-oriented representation. Extensive experiments demonstrate that ATOTrack is superior to existing methods, which achieves the state-of-the-art performance on six tracking benchmarks. In particular, ATOTrack establishes a new record on AViST with 57% AO. The code and models will be released as soon. Sixian Chan 0001, Xianpeng Zeng, Zhoujian Wu, Yu Wang 0254, Xiaolong Zhou 0001, Tinglong Tang, Jie Hu 0041 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2025 | Auxiliary Representation Guided Network for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared Person Re-identification aims to retrieve images of specific identities across modalities. To relieve the large cross-modality discrepancy, researchers introduce the auxiliary modality within the image space to assist modality-invariant representation learning. However, the challenge persists in constraining the inherent quality of generated auxiliary images, further leading to a bottleneck in retrieval performance. In this paper, we propose a novel Auxiliary Representation Guided Network (ARGN) to explore the potential of auxiliary representations, which are directly generated within the modality-shared embedding space. In contrast to the original visible and infrared representations, which contain information solely from their respective modalities, these auxiliary representations integrate cross-modality information by fusing both modalities. In our framework, we utilize these auxiliary representations as modality guidance to reduce the cross-modality discrepancy. First, we propose a High-quality Auxiliary Representation Learning (HARL) framework to generate identity-consistent auxiliary representations. The primary objective of our HARL is to ensure that auxiliary representations capture diverse modality information from both modalities while concurrently preserving identity-related discrimination. Second, guided by these auxiliary representations, we design an Auxiliary Representation Guided Constraint (ARGC) to optimize the modality-shared embedding space. By incorporating this constraint, the modality-shared embedding space is optimized to achieve enhanced intra-identity compactness and inter-identity separability, further improving the retrieval performance. In addition, to improve the robustness of our framework against the modality variation, we introduce a Part-based Adaptive Gaussian Module (PAGM) to adaptively extract discriminative information across modalities. Finally, extensive experiments are conducted to demonstrate the superiority of our method over state-of-the-art approaches on three VI-ReID datasets. Mengzan Qi, Sixian Chan 0001, Chen Hang, Guixu Zhang, Tieyong Zeng, Zhi Li 0080 |
IEEE Trans. Multim. | 2 |
| 2025 | PrimePSegter: Progressively Combined Diffusion for 3D Panoptic Segmentation With Multi-Modal BEV RefinementabstractEffective and robust 3D panoptic segmentation is crucial for scene perception in autonomous driving. Modern methods widely adopt multi-modal fusion based simple feature concatenation to enhance 3D scene understanding, resulting in generated multi-modal representations typically lack comprehensive semantic and geometry information. These methods focused on panoptic prediction in a single step also limit the capability to progressively refine panoptic predictions under varying noise levels, which is essential for enhancing model robustness. To address these limitations, we first utilize BEV space to unify semantic-geometry perceptual representation, allowing for a more effective integration of LiDAR and camera data. Then, we propose PrimePSegter, a progressively combined diffusion 3D panoptic segmentation model that is conditioned on BEV maps to iteratively refine predictions by denoising samples generated from Gaussian distribution. PrimePSegter adopts a conditional encoder-decoder architecture for fine-grained panoptic predictions. Specifically, a multi-modal conditional encoder is equipped with BEV fusion network to integrate semantic and geometric information from LiDAR and camera streams into unified BEV space. Additionally, a diffusion transformer decoder operates on multi-modal BEV features with varying noise levels to guide the training of diffusion model, refining the BEV panoptic representations enriched with semantics and geometry in a progressive way. PrimePSegter achieves state-of-the-art performance on the nuScenes and competitive results on the SemanticKITTI, respectively. Moreover, PrimePSegter demonstrates superior robustness towards various scenarios, outperforming leading methods. Hongqi Yu, Sixian Chan 0001, Xiaolong Zhou 0001, Xiaoqin Zhang 0002 |
IEEE Trans. Multim. | 2 |
| 2025 | SyNet: A Synergistic Network for 3D Object Detection Through Geometric-Semantic-Based Multi-Interaction FusionabstractDriven by rising demands in autonomous driving, robotics,etc., 3D object detection has recently achieved great advancement by fusing optical images and LiDAR point data. On the other hand, most existing optical-LiDAR fusion methods straightly overlay RGB images and point clouds without adequately exploiting the synergy between them, leading to suboptimal fusion and 3D detection performance. Additionally, they often suffer from limited localization accuracy without proper balancing of global and local object information. To address this issue, we design a synergistic network (SyNet) that fuses geometric information, semantic information, as well as global and local information of objects for robust and accurate 3D detection. The SyNet captures synergies between optical images and LiDAR point clouds from three perspectives. The first is geometric, which derives high-quality depth by projecting point clouds onto multi-view images, enriching optical RGB images with 3D spatial information for a more accurate interpretation of image semantics. The second is semantic, which voxelizes point clouds and establishes correspondences between the derived voxels and image pixels, enriching 3D point clouds with semantic information for more accurate 3D detection. The third is balancing local and global object information, which introduces deformable self-attention and cross-attention to process the two types of complementary information in parallel for more accurate object localization. Extensive experiments show that SyNet achieves 70.7% mAP and 73.5% NDS on the nuScenes test set, demonstrating its effectiveness and superiority as compared with the state-of-the-art. Xiaoqin Zhang 0002, Kenan Bi, Sixian Chan 0001, Shijian Lu, Xiaolong Zhou 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Attention-Enhanced Multi-View Stereo with Probabilistic Depth Variance RefinementabstractMulti-view stereo (MVS) reconstruction is a fundamental task in computer vision. While learning-based multi-view stereo methods have demonstrated excellent performance, insufficient attention has been paid to areas with significant depth estimation uncertainty. To address this issue, we propose an attention-enhanced network with probabilistic depth variance refinement for multi-view stereo reconstruction called AE-DR-MVSNet. Specifically, it contains two kernel modules: a diverse attention fusion module(DAFM) and a probability depth variance refinement module(PDVRM). The DAFM is designed by fusing the Bi-Level routing attention and the efficient multi-scale attention to achieve robust multi-scale features. The PDVRM is advanced to enhance depth estimation accuracy and adaptability by dynamically adjusting the depth hypothesis range based on depth distribution variance. Finally, extensive experiments conducted on benchmark datasets, including DTU and Blend-edMVS, demonstrate that AE-DR-MVSNet achieves competitive performance compared to lots of traditional and learning-based methods. Sixian Chan 0001, Aofeng Qiu, Jiafa Mao |
CSCWD | 1 |
| 2024 | WRIM-Net: Wide-Ranging Information Mining Network for Visible-Infrared Person Re-identification
Yonggan Wu, Ling-Chao Meng, Yuan Zichao, Sixian Chan 0001 |
ECCV (53) | 4 |
| 2024 | CurSegNet: 3D Dental Model Segmentation Network Based on Curve Feature Aggregation
Jiafa Mao, Jingke Gu, Sixian Chan 0001 |
ICANN (8) | 5 |
| 2024 | EE-MVSNet: Deep Learning-Based Cascaded High-Precision Multi-View Stereo Network with ECA & EVCabstractMulti-view stereo (MVS) has emerged as a pivotal algorithm in 3D reconstruction, garnering significant research attention over the past several decades. While recent coarse-to-fine methods have demonstrated promising results in enhancing the reconstruction quality of traditional algorithms, they often neglect the crucial aspect of feature layer refinement. Additionally, these methods face the challenge of low-cost feature matching. To address these limitations, we propose a novel learning-based MVS framework(EE-MVSNet). Firstly, we propose a novel approach incorporating an explicit visual center (EVC) module within the feature pyramid network (FPN), strengthening the adjustment within feature layers and improving model accuracy. Furthermore, we introduce the ECA+3DCNN module, which utilizes channel attention to alleviate the problem of low-cost feature matching. Finally, our model achieves competitive performance through extensive experimentation on the DTU dataset, showcasing its high-quality 3D reconstruction. Changfei Kong, Jiafa Mao, Xu Cheng 0003, Sixian Chan 0001 |
SMC | 5 |
| 2024 | Multi-scale feature correspondence and restriction mechanism for visible X-ray baggage re-Identification
Sixian Chan 0001, Jiaao Cui, Yonggan Wu, Cong Bai |
Multim. Syst. | 1 |
| 2024 | An efficient and real-time steel surface defect detection method based on single-stage detection algorithm
Qiqi Miao, Suqiang Li, Sixian Chan 0001, Jie Hu 0041, Cong Bai |
Multim. Tools Appl. | 5 |
| 2024 | DECNet: Dense embedding contrast for unsupervised semantic segmentation
Xiaoqin Zhang 0002, Xiaolong Zhou 0001, Sixian Chan 0001 |
Neural Networks | 4 |
| 2024 | Revisiting single-step adversarial training for robustness and generalization
Zhuorong Li, Daiwei Yu, Sixian Chan 0001, Hongchuan Yu, Zhike Han |
Pattern Recognit. | 4 |
| 2024 | Video-Based Multi-Camera Vehicle Tracking via Appearance-Parsing Spatio-Temporal Trajectory Matching NetworkabstractMulti-camera vehicle tracking is a fundamental task for city traffic management to count traffic flow or monitor roads. This paper focuses on multi-camera tracking on the highway, which is more challenging compared with city streets in some problems such as fast-moving vehicles, tiny similar vehicles in appearance, longer tracking distance, and lighting intensity changes in the dark tunnels. In this paper, we propose a practical Appearance-Parsing Spatio-Temporal Trajectory Matching Network (ASTM-Net) based on the global appearance matching of local trajectory for addressing the cross-camera tracking tasks on the highway. Specifically, considering that the environmental disturbance and small vehicles have a similar appearance, we propose a multiple appearance-attribute parsing (MAP) module consisting of a Bi-propagation top-down (Bi-TD) block and appearance re-identification (ARe-ID) block to obtain salient global appearance-attribute features through given a video sequence. To address discrete tracking fragments caused by occlusion, we develop an appearance-joint-tracking (AJT) mechanism to merge the isolated tracklets with target interaction and occlusion handling. We then exploit an appearance-informed spatio-temporal matching (ASTM) module to achieve multi-camera tracklet-totarget assignment, which employs spatio-temporal consistency relation for intra-camera trajectory correction and coarse intercamera tracklet correlation and aggregate appearance matrix of local trajectories for assigning global trajectory ID. Finally, in order to evaluate our proposed ASTM-Net, a new dataset, named HST, collected on the highway is established.We verify the ASTM-Net on the HST and the other three public datasets,i.e., CityFlow, UA-DETRAC, and Synthehicle, whose experimental results demonstrate the effectiveness and robustness of the proposed method. Xiaoqin Zhang 0002, Hongqi Yu, Xiaolong Zhou 0001, Sixian Chan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Diverse-Feature Collaborative Progressive Learning for Visible-Infrared Person Re-IdentificationabstractVisible–infrared person reidentification (VI-ReID) aims to search for pedestrian identities in different spectra. The major challenge is the modality differences between infrared and visible images for the VI-ReID task. Existing approaches try to design networks based on a single-stage training strategy to extract features. However, they often excessively rely on a particular feature, such as modality-specific features or modality-independent features, and overlook the significance of the diverse features obtained by combining them. To address this problem, we propose a diverse-feature collaborative progressive learning network (DCPLNet) for VI-ReID in this article. With the benefit of diverse information, our DCPLNet can effectively learn informative representations for reducing the modality differences. Specifically, we propose a novel three-stage progressive learning strategy (t-PLS) to progressively learn diverse features. For the proposed t-PLS, we design a contour feature enhancement module to mine human contour features and raise a perceptual contour feature loss for supervised feature extraction. Finally, we advance a batch adaptation module to establish feature links between samples. Extensive experiments on SYSU-MM01, RegDB, and LLCM datasets demonstrate that our proposed model performs better than most state-of-the-art methods. Sixian Chan 0001, Weihao Meng, Cong Bai, Jie Hu 0041, Shenyong Chen |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Auxiliary Feature Fusion and Noise Suppression for HOI DetectionabstractIn recent years, one-stage HOI (Human–Object Interaction) detection methods tend to divide the original task into multiple sub-tasks by using a multi-branch network structure. However, there is no sufficient attention to information communication between these branches. The inference approach in the cascaded structure is singular, while fully parallel methods will disrupt the associations between different pieces of information. Besides, noise interference may occur during the fusion of different features and thus affect the detection performance. To address these issues, this article proposes a one-stage three-branch parallel HOI detection method, which treats HOI as three separate sub-tasks (human detection, object detection, and interaction detection) and leverages three distinct reasoning relationships to generate richer relational information. Firstly , an auxiliary feature fusion (AFF) module is introduced, which integrates features originally extracted independently to form fused features enriched with supplementary information. This approach strengthens communication between branches in the network while handling the three sub-tasks concurrently, thereby facilitating the exchange of more contextual information. Secondly , to mitigate noise interference generated during the fusion process, a fusion noise suppression (FNS) module is introduced, which effectively suppresses noise and enhances the model’s performance in interaction detection tasks. Finally , experiments are conducted on two major benchmark datasets, and experimental results show that our HOI detection method is superior to previous methods. Also, ablation studies confirm the effectiveness of all the components in our proposed method. Sixian Chan 0001, Xianpeng Zeng, Jie Hu 0041, Cong Bai |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | MGTCF: Multi-Generator Tropical Cyclone Forecasting with Heterogeneous Meteorological DataabstractAccurate forecasting of tropical cyclone (TC) plays a critical role in the prevention and defense of TC disasters. We must explore a more accurate method for TC prediction. Deep learning methods are increasingly being implemented to make TC prediction more accurate. However, most existing methods lack a generic framework for adapting heterogeneous meteorological data and do not focus on the importance of the environment. Therefore, we propose a Multi-Generator Tropical Cyclone Forecasting model (MGTCF), a generic, extensible, multi-modal TC prediction model with the key modules of Generator Chooser Network (GC-Net) and Environment Net (Env-Net). The proposed method can utilize heterogeneous meteorologic data efficiently and mine environmental factors. In addition, the Multi-generator with Generator Chooser Net is proposed to tackle the drawbacks of single-generator TC prediction methods: the prediction of undesired out-of-distribution samples and the problems stemming from insufficient learning ability. To prove the effectiveness of MGTCF, we conduct extensive experiments on the China Meteorological Administration Tropical Cyclone Best Track Dataset. MGTCF obtains better performance compared with other deep learning methods and outperforms the official prediction method of the China Central Meteorological Observatory in most indexes. Cong Bai, Sixian Chan 0001, Yuquan Wu |
AAAI | 3 |
| 2023 | LE-MVSNet: Lightweight Efficient Multi-view Stereo Network
Changfei Kong, Jiafa Mao, Sixian Chan 0001, Weigou Sheng |
ICANN (8) | 4 |
| 2023 | Towards General and Fast Video Derain via Knowledge DistillationabstractAs a common natural weather condition, rain can obscure video frames and thus affect the performance of the visual system, so video derain receives a lot of attention. In natural environments, rain has a wide variety of streak types, which increases the difficulty of the rain removal task. In this paper, we propose a Rain Review-based General video derain Network via knowledge distillation (named RRGNet) that handles different rain streak types with one pre-training weight. Specifically, we design a frame grouping-based encoder-decoder network that makes full use of the temporal information of the video. Further, we use the old task model to guide the current model in learning new rain streak types while avoiding forgetting. To consolidate the network’s ability to derain, we design a rain review module to play back data from old tasks for the current model. The experimental results show that our developed general method achieves the best results in terms of running speed and derain effect. Defang Cai, Pan Mu, Sixian Chan 0001, Zhanpeng Shao, Cong Bai |
ICME | 3 |
| 2023 | Visible-Xray Cross-Modality Package Re-IdentificationabstractDuring the package inspection, once prohibited articles are checked under the X-ray, the inspector needs to find the corresponding package in time for further confirmation. When the number of passengers increases, this process is time-consuming. Recently, prohibited articles detection as a detection task has attracted much attention, but few studies have focused on Visible-Xray package re-identification (VX-ReID) task. In this paper, we mainly explore the VX-ReID task. Firstly, we establish the first VX-ReID dataset RX01, which includes 55883 Visible and 29174 X-ray package images. Furthermore, we introduce a baseline model that includes a cross-modality channel attention module (CMCA) and momentum contrast mAP (MoCoAP). CMCA is used to enhance channels that contain modality-invariant information. MoCoAP is a differentiable mAP approximation strategy that directly optimizes the retrieve performance of the model. By combining these two strategies, we achieve competitive performance on the RX01 and SYSU-MM01 datasets. Code will be released at https://github.com/cjjjao/VX-ReID. Sixian Chan 0001, Jiaao Cui, Yonggan Wu, Cong Bai |
ICME | 1 |
| 2023 | Fine-grained Learning for Visible-Infrared Person Re-identificationabstractVisible-Infrared Person Re-identification aims to retrieve specific identities from different modalities. In order to relieve the modality discrepancy, previous works mainly concentrate on aligning the distribution of high-level features, while disregarding the exploration of fine-grained information. In this paper, we propose a novel Fine-grained Information Exploration Network (FIENet) to implement discriminative representation, further alleviating the modality discrepancy. Firstly, we propose a Progressive Feature Aggregation Module (PFAM) to progressively aggregate mid-level features, and a Multi-Perception Interaction Module (MPIM) to achieve the interaction with diverse perceptions. Additionally, combined with PFAM and MPIM, more fine-grained information can be extracted, which is beneficial for FIENet to focus on discriminative human parts in both modalities effectively. Secondly, in terms of the feature center, we introduce an Identity-Guided Center Loss (IGCL) to supervise identity representation with intra-identity and inter-identity information. Finally, extensive experiments are conducted to demonstrate that our method achieves state-of-the-art performance. Mengzan Qi, Sixian Chan 0001, Chen Hang, Guixu Zhang, Zhi Li 0080 |
ICME | 2 |
| 2023 | SGPT: The Secondary Path Guides the Primary Path in Transformers for HOI DetectionabstractHOI detection is essential for human-computer interaction, especially in behavior detection and robot manipulation. Existing mainstream transformer methods of HOI detection are focused on single-stream detection only, e.g.,$image \rightarrow HOI(\mathcal{P}_{1})$, or$image \rightarrow HO\rightarrow I(\mathcal{P}_{2})$. Both paths have their own characteristics of concern, so we propose a novel method, using the Secondary path$(\mathcal{P}_{2})$Guides the Primary path$(\mathcal{P}_{1})$in Transformers (SGPT). SGPT contains two core modules: the Dual-Path Consistency (DPC) module and the Instance Interaction Attention (IIA) module. DPC keeps human, object and interaction consistent on the dual-path and lets$\mathcal{P}_{2}$guide$\mathcal{P}_{1}$to learn more meaningful features. IIA fuses human and object to enhance interaction in$\mathcal{P}_{2}$, which allows instance to constrain interaction. Our proposed dual-path are employed during training, and only the$\mathcal{P}_{1}$path is used for inference. Hence, SGPT improves generalization without increasing model capacity in HICO-DET and V-COCO datasets compared to the state-of-the-arts. The code of this work is available at https://github.com/visualVk/sgpt.git. Sixian Chan 0001, Weixiang Wang, Zhanpeng Shao, Cong Bai |
ICRA | 1 |
| 2023 | A Generalized Physical-knowledge-guided Dynamic Model for Underwater Image EnhancementabstractUnderwater images often suffer from color distortion and low contrast resulting in various image types, due to the scattering and absorption of light by water. While it is difficult to obtain high-quality paired training samples with a generalized model. To tackle these challenges, we design a Generalized Underwater image enhancement method via a Physical-knowledge-guided Dynamic Model (short for GUPDM). In particular, to cover complex underwater scenes, this study changes the global atmosphere light and the transmission to simulate various underwater image types through the formation model. We then design an Atmosphere-based Dynamic Structure (ADS) and Transmission-guided Dynamic Structure (TDS) that use dynamic convolutions to adaptively extract prior information from underwater images and generate parameters for Prior-based Multi-scale Structure (PMS). These two modules enable the network to select appropriate parameters for various water types adaptively. Besides, the multi-scale feature extraction module in PMS uses convolution blocks with different kernel sizes and obtains weights for each feature map via channel attention block. The source code will be available at https://github.com/shiningZZ/GUPDM Pan Mu, Hanning Xu, Zheyuan Liu 0009, Zheng Wang 0059, Sixian Chan 0001, Cong Bai |
ACM Multimedia | 5 |
| 2023 | Asymmetric Cascade Fusion Network for Building ExtractionabstractThe U-Net-like model has been widely studied in the field of building extraction. However, most of these models are based on locally sensed Convolutional Neural Networks(CNNs) designed with symmetric structure and single feature processing, which cannot accurately identify buildings with different sizes, shapes, and colors in remote sensing images. To overcome these problems, we propose the asymmetric cascade fusion network(ACFN), based on the Vision Transformer(ViT), to design a novel asymmetric architecture to recognize buildings of different sizes and shapes by processing multi-granularity features by different means. First, the asymmetric architecture obtains multi-granularity features with global contextual information by embedding different types of attention in encoder-decoders of different sizes. This architecture can identify densely distributed and occluded buildings by semantic reasoning in remote sensing images with complex information. Second, we design a multi-branch weighted pyramid pooling module, which sets different branch weights to offset the background noise introduced in introducing global contextual information. Our ACFN significantly improves the Beijing buildings, ISPRS-Vaihingen, and LoveDA datasets. Sixian Chan 0001, Yuan Wang 0032, Yanjing Lei, Xu Cheng 0003, Wei Wu 0029 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | MLPT: Multilayer Perceptron based TrackingabstractThe global receptive field plays a critical role in visual object tracking. In most popular tracking paradigms, we find that the local receptive field introduced by the convolutional neural network prevents the tracker from focusing on the long-range dependency. Although the Vision transformer brings the global receptive field in downstream tasks, its computational burden remains unaffordable. In this paper, we present a simple yet effective Multilayer Perceptron-based Tracking (MLPT), including the global receptive field. The MLPT contains three Components: Feature Correlation (FC) module, Global Information Encoder (GIE) module and Corner Head(CH). Firstly, the FC module is proposed to effectively converge the template and search region features for generating delicate features. Secondly, the GIE is designed to integrate the channel and spatial-information dealed with channel-encoding the token-encoding, separately. Specially, the same kernel is utilized for token-encoding in all channels so that our model has the global receptive field. Then, the CH is applied to establish a simple flexible way via computing the box corners coordinates for tracking. Finally, the MLPT, to our knowledge, is the first baseline of MLP-based architecture for object tracking. Extensive experiments are conducted on four challenging datasets, including GOT-10K, LaSOT, UAV123, and TrackingNet. The results show that the proposed method achieves state-of-the-art performance. Sixian Chan 0001, Yu Wang 0254, Xiaolong Zhou 0001, Qike Shao |
SMC | 1 |
| 2022 | Res2-UNeXt: a novel deep learning framework for few-shot cell image segmentation
Sixian Chan 0001, Cong Bai, Shengyong Chen |
Multim. Tools Appl. | 1 |
| 2022 | Online multiple object tracking using joint detection and embedding network
Sixian Chan 0001, Yangwei Jia, Xiaolong Zhou 0001, Cong Bai, Shengyong Chen, Xiaoqin Zhang 0002 |
Pattern Recognit. | 1 |
| 2022 | Siamese Implicit Region Proposal Network With Compound Attention for Visual TrackingabstractRecently, siamese-based trackers have achieved significant successes. However, those trackers are restricted by the difficulty of learning consistent feature representation with the object. To address the above challenge, this paper proposes a novel siamese implicit region proposal network with compound attention for visual tracking. First, an implicit region proposal (IRP) module is designed by combining a novel pixel-wise correlation method. This module can aggregate feature information of different regions that are similar to the pre-defined anchor boxes in Region Proposal Network. To this end, the adaptive feature receptive fields then can be obtained by linear fusion of features from different regions. Second, a compound attention module including a channel and non-local attention is raised to assist the IRP module to perform a better perception of the scale and shape of the object. The channel attention is applied for mining the discriminative information of the object to handle the background clutters of the template, while non-local attention is trained to aggregate the contextual information to learn the semantic range of the object. Finally, experimental results demonstrate that the proposed tracker achieves state-of-the-art performance on six challenging benchmark tests, including VOT-2018, VOT-2019, OTB-100, GOT-10k, LaSOT, and TrackingNet. Further, our obtained results demonstrate that the proposed approach can be run at an average speed of 72 FPS in real time. Sixian Chan 0001, Xiaolong Zhou 0001, Cong Bai, Xiaoqin Zhang 0002 |
IEEE Trans. Image Process. | 1 |
| 2021 | Tropical Cyclones Tracking Based on Satellite Cloud Images: Database and Comprehensive Study
Sixian Chan 0001, Cong Bai |
MMM (2) | 2 |
| 2021 | Box Regression-Guided Anchor-free for Robust Visual TrackingabstractThe Siamese tracker-based approach has achieved significant success in recent years. However, these approaches do not consider the different requirements for input feature in classification and regression branches. The regression branch needs feature information slightly larger than the object region, while the classification branch needs to avoid classification failure caused by the introduction of background information. In this paper, we present a novel Box Regression-Guided Anchor-free for Robust Visual Tracking. Firstly, a scale-aware regression module is designed to satisfy the feature requirements of the regression branch, which can capture feature information of various scales. Secondly, regression-guided classification module is applied to aligning the feature between the regression result and correlation feature, thereby avoiding the introduction of background information to classification branch. In addition, the new correlation operation is introduced to gain more superb correlation feature. Comparsion experimental exhibits that the proposed tracker achieves promising results in five challenging benchmark tests, including GOT-10K, OTB-2015, VOT-2018, VOT-2019 and TrackingNet, and run at an average speed of 60 FPS in real-time. Sixian Chan 0001, Xiaolong Zhou 0001, Cong Bai, Hua Gao, Shengyong Chen |
SMC | 2 |
| 2019 | Dictionary Learning and Confidence Map Estimation-Based Tracker for Robot-Assisted Therapy System
Xiaolong Zhou 0001, Sixian Chan 0001, Shengyong Chen, Honghai Liu 0001 |
PRCV (1) | 2 |
| 2018 | Online classification for object tracking based on superpixel
Sixian Chan 0001, Xiaolong Zhou 0001, Shengyong Chen |
Neurocomputing | 1 |
| 2017 | Compressive tracking with locality sensitive histograms featuresabstractCurrently, Compressive Tracking (CT) method has drawn great attention because of its high efficiency. However, it cannot well deal with some appearance variations due to its limitations of feature expression and it only uses a fixed parameter to update the appearance model. In order to handle such matters, we propose an adaptive CT method that combines the predicted target position with CT based on Locality Sensitive Histograms (LSH) features. Our method significantly improves CT in four aspects. First, the efficient illumination invariant features extracted based on LSH are used to represent an effective appearance model that is robust to illumination changes. Second, the color attributes tracker is adopted to predict the target position for re-building the new weighted discriminant function which brings in the color information to make up for the inadequacy of Haar-like characteristics. Third, a new model update mechanism is proposed to preserve the stable features while avoid the noisy appearance variations during tracking. Fourth, a trajectory rectification method is employed to refine the tracking location when possible inaccurate tracking occurs. Finally, we show that our tracker achieves state-of-the-art performance in a comprehensive evaluation over 47 challenging color sequences. Sixian Chan 0001, Xiaolong Zhou 0001, Zhuo Zhang 0012, Shengyong Chen |
ICRA | 1 |
| 2017 | Object tracking using a convolutional network and a structured output SVMabstractObject tracking has been a challenge in computer vision. In this paper, we present a novel method to model target appearance and combine it with structured output learning for robust online tracking within a tracking-by-detection framework. We take both convolutional features and handcrafted features into account to robustly encode the target appearance. First, we extract convolutional features of the target by kernels generated from the initial annotated frame. To capture appearance variation during tracking, we propose a new strategy to update the target and background kernel pool. Secondly, we employ a structured output SVM for refining the target’s location to mitigate uncertainty in labeling samples as positive or negative. Compared with existing state-of-the-art trackers, our tracking method not only enhances the robustness of the feature representation, but also uses structured output prediction to avoid relying on heuristic intermediate steps to produce labelled binary samples. Extensive experimental evaluation on the challenging OTB-50 video sequences shows competitive results in terms of both success and precision rate, demonstrating the merits of the proposed tracking method. Xiaolong Zhou 0001, Sixian Chan 0001, Shengyong Chen |
Comput. Vis. Media | 3 |
| 2017 | Adaptive Compressive Tracking based on Locality Sensitive Histograms
Sixian Chan 0001, Xiaolong Zhou 0001, Shengyong Chen |
Pattern Recognit. | 1 |
| 2016 | A pipeline using multi-layer Tumors Automata for interactive multi-label image segmentationabstractIn this paper, we investigate a novel algorithm to the problem of interactive image segmentation. We propose an extension of the Growcut framework using the Tumors Automata (TA) formed from the superpixel. The proposed TA is similar to Cellular Automata but can directly deal with superpixel. The superpixels (image segments) can provide powerful boundary cues to guide segmentation, where superpixels can be collected easily by over-segmenting the image using any reasonable existing segmentation algorithms. Given a small number of user-labelled superpixels, the rest of the image is segmented automatically by a TA. When the automaton labels the image, the segmentation evolution is faster than Growcut because of the iterative process. Moreover, a level set method and multi-layer TA are employed to further improve the performance. Experiments conducted on the Berkeley Segmentation Database demonstrate the superior performance of our method over the state-of-the-art methods. Sixian Chan 0001, Xiaolong Zhou 0001, Zhuo Zhang 0012, Shengyong Chen |
HSI | 1 |