VLDB 2026 Research / reviewers in the wild / expert
Xiankai Lu
dblp:153/2122
· DBLP profile ↗
69ranked-venue papers
12as first author
50since 2021 · last 2026
0000-0002-9543-6960ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 9 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 7 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OVFormer+: Improved Open-Vocabulary Video Instance Segmentation via Text-Guided Unified Embedding Alignment
Hao Fang 0010, Xiankai Lu, Henghui Ding, Yunchao Wei, Yawei Li 0001, Runmin Cong |
Int. J. Comput. Vis. | 2 |
| 2026 | SeeGait: Synergistic Co-Evolving Representations for Multimodal Gait Recognition via Hierarchical Multi-Stage FusionabstractGait recognition offers non-contact, long-distance identification but struggles with robustness against covariates like clothing variations, carrying conditions, and viewpoint changes. Existing methods predominantly rely on single modalities (e.g., silhouettes or skeletons) or employ shallow multimodal fusion, such as simple concatenation, which treats modalities as independent and static, failing to exploit their complementary strengths, shape cues from silhouettes and structural kinematics from skele-tons. To address these limitations, we introduce the Synergistic co-evolving representations (See) principle, enabling modalities to iteratively interact, guide, and refine each other across semantic hierarchies, fostering a unified, robust identity representation resilient to complex environments. This is realized through SeeGait, a novel multimodal framework featuring hierarchical multi-stage fusion. At its core, the Bidirectional Hierarchical Cross-Attention Synergy Module (BiHCASM) employs adaptive cross-modal attention to dynamically align and reweight features bidirectionally, allowing structural insights to enhance appearance focus and vice versa. Complementing this, the Hierarchical Spatiotemporal Transformer Encoder (HSTE) captures long-range skeleton dynamics, overcoming GCN limitations, while the Hierarchical Convolutional Silhouette Encoder (HCSE) extracts multi-scale silhouette pyramids for rich shape priors. Finally, a Holistic Feature Aggregation (HFA) strategy consolidates features from all stages for deep supervision, ensuring comprehensive optimization. By promoting mutual refinement, SeeGait mitigates covariate disruptions through enhanced complementarity, yielding superior discriminability. Extensive experiments show state-of-the-art performance, with 97.1% average Rank-1 accuracy on CASIA-B, and top results on CCPG and SUSTech1K. Hanyue Du, Xianye Ben, Xiankai Lu, Zunxiao Xu, Qingshuo Gao, Qiang Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Decoupled Motion Expression Video SegmentationabstractMotion expression video segmentation aims to segment objects based on input motion descriptions. Compared with traditional referring video object segmentation, it focuses on motion and multi-object expressions and is more challenging. Previous works achieved it by simply injecting text information into the video instance segmentation (VIS) model. However, this requires retraining the entire model and optimization is difficult. In this work, we propose DMVS, a simple framework constructed on the existing query-based VIS model, emphasizing decoupling the task into video instance segmentation and motion expression understanding. Firstly, we use a frozen video instance segmenter to extract object-specific contexts and convert them into frame-level and video-level queries. Secondly, we interact two levels of queries with static and motion cues, respectively, to further encode visually enhanced motion expressions. Furthermore, we propose a novel query initialization strategy that uses video queries guided by classification priors to initialize motion queries, greatly reducing the difficulty of optimization. Without bells and whistles, DMVS achieves state-of-the-art performance on the MeViS dataset at a lower training cost. Extensive experiments verify the effectiveness and efficiency of our framework. Hao Fang 0010, Runmin Cong, Xiankai Lu, Xiaofei Zhou 0003, Sam Kwong, Wei Zhang 0021 |
CVPR | 3 |
| 2025 | Semantic and Sequential Alignment for Referring Video Object SegmentationabstractReferring video object segmentation (RVOS) seeks to segment the objects within a video referred by linguistic expressions. Existing RVOS solutions follow a "fuse then select" paradigm: establishing semantic correlation between visual and linguistic feature, and performing frame-level query interaction to select the instance mask per frame with instance segmentation module. This paradigm overlooks the challenge of semantic gap between the linguistic descriptor and the video object as well as the underlying clutters in the video. This paper proposes a novel Semantic and Sequential Alignment (SSA) paradigm to handle these challenges. We first insert a lightweight adapter after the vision language model (VLM) to perform the semantic alignment. Then, prior to selecting mask per frame, we exploit the trajectory-to-instance enhancement for each frame via sequential alignment. This paradigm leverages the visual-language alignment inherent in VLM during adaptation and tries to capture global information by ensembling trajectories. This helps understand videos and the corresponding descriptors by mitigating the discrepancy with intricate activity semantics, particularly when facing occlusion or similar interference. SSA demonstrates competitive performance while maintaining fewer learnable parameters. Feiyu Pan, Hao Fang 0010, Fangkai Li, Yawei Li 0001, Luca Benini, Xiankai Lu |
CVPR | 7 |
| 2025 | LOGICZSL: Exploring Logic-induced Representation for Compositional Zero-shot LearningabstractCompositional zero-shot learning (CZSL) aims to recognize unseen attribute-object compositions by learning the primitive concepts (i.e., attribute and object) from the training set. While recent works achieve impressive results in CZSL by leveraging large vision-language models like CLIP, they ignore the rich semantic relationships between primitive concepts and their compositions. In this work, we propose LogiCzsl, a novel logic-induced learning framework to explicitly model the semantic relationships. Our logic-induced learning framework formulates the relational knowledge constructed from large language models as a set of logic rules, and grounds them onto the training data. Our logic-induced losses are complementary to the widely used CZSL losses, therefore can be employed to inject the semantic information into any existing CZSL methods. Extensive experimental results show that our method brings significant performance improvements across diverse datasets (i.e., CGQA, UT-Zappos50K, MIT-States) with strong CLIP-based methods and settings (i.e., Close World, Open World). Peng Wu 0014, Xiankai Lu, Yongqin Xian, Jianbing Shen, Wenguan Wang |
CVPR | 2 |
| 2025 | A Conditional Probability Framework for Compositional Zero-Shot LearningabstractCompositional Zero-Shot Learning (CZSL) aims to recognize unseen combinations of known objects and attributes by leveraging knowledge from previously seen compositions. Traditional approaches primarily focus on disentangling attributes and objects, treating them as independent entities during learning. However, this assumption overlooks the semantic constraints and contextual dependencies inside a composition. For example, certain attributes naturally pair with specific objects (e.g., "striped" applies to "zebra" or "shirts" but not "sky" or "water"), while the same attribute can manifest differently depending on context (e.g., "young" in "young tree" vs. "young dog"). Thus, capturing attribute-object interdependence remains a fundamental yet long-ignored challenge in CZSL. In this paper, we adopt a Conditional Probability Framework (CPF) to explicitly model attribute-object dependencies. We decompose the probability of a composition into two components: the likelihood of an object and the conditional likelihood of its attribute. To enhance object feature learning, we incorporate textual descriptors to highlight semantically relevant image regions. These enhanced object features then guide attribute learning through a cross-attention mechanism, ensuring better contextual alignment. By jointly optimizing object likelihood and conditional attribute likelihood, our method effectively captures compositional dependencies and generalizes well to unseen compositions. Extensive experiments on multiple CZSL benchmarks demonstrate the superiority of our approach. Code is available at here. Peng Wu 0014, Qiuxia Lai, Hao Fang 0010, Guosen Xie, Yilong Yin, Xiankai Lu, Wenguan Wang |
ICCV | 6 |
| 2025 | Context-Enhanced Zero-Shot Video Temporal Grounding with Adaptive Boundary RefinementabstractIn this paper, we introduce a novel training-free framework for Video Temporal Grounding (VTG) that combines pre-trained Visual Language Models (VLMs) and Large Language Models (LLMs). Existing methods often struggle with capturing the semantics of natural language queries and identifying the dynamic transitions at event boundaries. To address these challenges, our approach uses VLMs to generate detailed contextual descriptions of video content, providing richer prompts for LLMs to understand and reason about event temporal relations. Furthermore, we introduce an adaptive event boundary refinement strategy, ensuring better coverage of the full event phases. Our framework demonstrates superior performance in zero-shot settings on several benchmark datasets, including Charades-STA and ActivityNet Captions, and exhibits remarkable robustness in out-of-distribution (OOD) scenarios. Fangkai Li, Feiyu Pan, Yiyou Guo, Xiankai Lu |
ICME | 6 |
| 2025 | Pixel-wise Single Image Reflection Removal Method Based on Reinforcement LearningabstractSingle image reflection removal is particularly important in improving image quality. However, existing single image reflection removal methods cannot remove reflection in a pixel-by-pixel manner, significantly reducing their effectiveness. To address this issue, we propose a pixel-wise single image reflection removal method based on reinforcement learning. Specifically, we formalize the image reflection removal process as a sequential decision-making process and remove the single image reflection pixel by pixel. We first introduce the single image reflection removal method at each time step. We then describe the reinforcement learning setup for image reflection, including state, action, reward, and agent network design. Finally, we conducted experiments on benchmark datasets, and the results show that our method outperforms other single image reflection removal methods. Xueshi Yu, Zhengzhe Zhang, Xiankai Lu, Yilong Yin, Wenjia Meng |
ICME | 4 |
| 2025 | CAN-ST: Clustering Adaptive Normalization for Spatio-temporal OOD LearningabstractSpatio-temporal data mining is crucial for decision-making and planning in diverse domains. However, in real-world scenarios, training and testing data are often not independent or identically distributed due to rapid changes in data distributions over time and space, resulting in spatio-temporal out-of-distribution (OOD) challenges. This non-stationarity complicates accurate predictions and has motivated research efforts focused on mitigating non-stationarity through normalization operations. Existing methods, nonetheless, often address individual time series in isolation, neglecting correlations across series, which limits their capacity to handle complex spatio-temporal dynamics and results in suboptimal solutions. To overcome these challenges, we propose Clustering Adaptive Normalization (CAN-ST), a general and model-agnostic method that mitigates non-stationarity by capturing both localized distributional changes and shared patterns across nodes via adaptive clustering and a parameter register. As a plugin, CAN-ST can be easily integrated into various spatio-temporal prediction models. Extensive experiments on multiple datasets with diverse forecasting models demonstrate that CAN-ST consistently improves performance by over 20% on average and outperforms state-of-the-art normalization methods. Min Yang 0006, Jinliang Deng, Ji Zhong, Xiankai Lu, Yongshun Gong |
IJCAI | 7 |
| 2025 | Towards Region-Adaptive Feature Disentanglement and Enhancement for Small Object DetectionabstractCurrent feature fusion strategies often fail to adequately account for the influence of activation intensity across different scales on small object features, which impedes the effective detection of small objects. To address this limitation, we propose the Region-Adaptive Feature Disentanglement and Enhancement (RAFDE) strategy, which improves both downsampling and feature fusion by leveraging activation intensity variations at multiple scales. First, we introduce the Boundary Transitional Region-enhanced Downsampling (BTRD) module, which enhances boundary transitional regions containing both strongly and weakly activated features, thereby mitigating the loss of crucial boundary information for small objects. Second, we present the Regional-Adaptive Feature Fusion (RAFF) module, which adaptively disentangles and fuses co-activated and uni-activated regions from adjacent levels into the current level, effectively reducing the risk of small objects being overwhelmed. Extensive experiments on several public datasets demonstrate that the RAFDE strategy is highly effective and outperforms state-of-the-art methods. The code is available at https://github.com/b-yanchao/RAFDE.git. Yanchao Bi, Yang Ning, Xiushan Nie, Xiankai Lu, Yongshun Gong, Leida Li |
IJCAI | 4 |
| 2025 | Cross-graph meta matching correction for noisy graph matching
Fangkai Li, Feiyu Pan, Wenjia Meng, Haoliang Sun, Xiushan Nie, Yilong Yin, Xiankai Lu |
Comput. Vis. Image Underst. | 7 |
| 2025 | FGHDet: Delving into Fine-Grained Features with Head Selection for UAV Object Detection
Yanchao Bi, Yang Ning, Xiu-Shan Nie, Xiankai Lu, Rui-Heng Zhang, Huan-Long Zhang |
J. Comput. Sci. Technol. | 4 |
| 2025 | A contrast-invariant feature extraction framework for single-domain generalization in infrared small target detection
Weijian Chi, Xiankai Lu, Yue Ni |
Knowl. Based Syst. | 3 |
| 2025 | LAC-PS: A Light Direction Selection Policy Under the Accuracy Constraint for Photometric StereoabstractPhotometric stereo (PS) methods recover surface normals from appearance changes under varying light directions, excelling in tasks like 3D surface reconstruction and defect inspection. However, collecting the illumination images is expensive, and current PS methods cannot obtain the light direction set that satisfies the pre-defined accuracy constraint, limiting their adaptability to various applications with varying accuracy requirements. To address this issue, we propose the LAC-PS, a light direction selection policy under the accuracy constraint for photometric stereo, which optimizes the light direction set to meet target reconstruction accuracy. In our method, we develop an accuracy assessment network that estimates reconstruction accuracy without ground truth. With this estimated accuracy, we put forward a reinforcement learning-based method that can utilize policy to sequentially select light directions and obtain the light directions satisfying the desired PS recovery accuracy constraint. Experimental results on real and synthetic datasets demonstrate that our method effectively selects light directions that satisfy accuracy constraints. Wenjia Meng, Huimin Han, Xiankai Lu, Yilong Yin, Gang Pan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Remote Sensing Scene Classification via Pseudo-Category-Relationand Orthogonal Feature LearningabstractRemote sensing (RS) scene classification is a crucial component in the analysis of Earth observation data, aiding in a deeper understanding and monitoring of our dynamic planet. Its applications extend across various fields, including land management, urban analysis, and environmental monitoring. The complex semantic information in RS scene images and the relationships between different scene categories present significant challenges to improving scene classification tasks. Unlike previous methods that only focus on network structure or feature encoding, our approach emphasizes the association of scene categories, integrating feature learning and knowledge transfer together to enhance the analysis of scenes at a higher semantic level. To this end, we propose an RS scene classification scheme based on pseudo-scene category-relation reasoning and orthogonal feature (OF) learning modules, capturing the inherent semantic connections among diverse scene classes. Additionally, we introduce cascaded attention (CA) and selected separation modules to strategically optimize the network, targeting challenging classes with high feature similarities. Knowledge is then distilled across different branches, guiding to enhance the model’s robustness and prediction accuracy. Experiments are conducted on three challenging RS scene datasets of AID30, UCMerced21, and NWPU-RESISC45 to validate the effectiveness of the learned pseudo-category relationships. The results demonstrate that the proposed framework outperforms existing hierarchical approaches in leveraging the hierarchical structure of RS scene images. Jinsheng Ji, Xiankai Lu, Tao Zhang 0027, Yiyou Guo, Gongping Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Detail-Aware Network for Infrared Image EnhancementabstractInfrared (IR) images inherently face the dual challenges of noise contamination and reduced contrast. However, existing image enhancement methods often overlook the intrinsic correlations between these factors—noise and low contrast—during multistage enhancement processes. Consequently, this oversight leads to a significant reduction in the fidelity of intricate details in IR images. In this article, we present a synergistic IR image enhancement network that simultaneously achieves denoising, contrast improvement, and detail preservation (DCDNet), which breaks down the overall enhancement process into more manageable steps. DCDNet is comprised of a detail awareness unit (DAU), a deep denoising prior (DDP), and a contrast improvement module (CIM). To maintain the details in the IR image, DAU is developed to extract the original detail feature information in DDP and integrate them into the CIM during contrast improvement to improve the final result. The detail information is derived from the encoder of the DDP, which focuses on denoising. The preserved detail features are subsequently incorporated into the decoder of the CIM, which is dedicated to enhancing contrast. Experimental results validate that our proposed approach surpasses other state-of-the-art methods for enhancing IR images in terms of the peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), visual information fidelity (VIF), and performance in downstream tasks. The code and dataset are publicly available athttps://github.com/ChickenEating/IR-Enhancement. Ruiheng Zhang 0001, Guanyu Liu, Qi Zhang 0004, Xiankai Lu, Renwei Dian, Yang Yang 0074, Lixin Xu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | FasterSal: Robust and Real-Time Single-Stream Architecture for RGB-D Salient Object DetectionabstractRGB-D Salient Object Detection (SOD) aims to segment the most prominent areas and objects in a given pair of RGB and depth images. Most current models adopt a dual-stream structure to extract information from both RGB and depth images. However, this leads to an exponential increase in the number of parameters and computations in the model. Moreover, the discrepancy between RGB pretrained and the 3D geometric relationships in depth maps present a challenge for the encoder in capturing spatial structural details. These issues impact the model's accuracy in locating salient objects and distinguishing edge details. To address these, we propose a novel early feature fusion network, named FasterSal, which enables more efficient RGB-D SOD. FasterSal uses a single stream structure to receive RGB images and depth maps, extracting features based on the 3D geometric relationships in the depth map while fully leveraging the pretrained RGB encoder. This approach effectively avoids the inconsistencies between depth modality and the RGB pretrained encoder. It also significantly reduces the number of network parameters while maintaining efficient feature encoding capabilities. To achieve finer edge learning, the detail-aware loss and texture enhancement module are introduced. These modules are designed to extract latent details in high-frequency component features and to enhance the edge learning capability of the model using distance information. Experimental results on several benchmark datasets confirm the effectiveness and superiority of our method over the state-of-the-art approaches, achieving a good balance between performance and speed with only 3.4 million parameters and a CPU operating speed of 63 FPS. Jing Zhang 0037, Ruiheng Zhang 0001, Lixin Xu 0001, Xiankai Lu, Yushu Yu, Min Xu 0001, He Zhao 0002 |
IEEE Trans. Multim. | 4 |
| 2024 | Unified Embedding Alignment for Open-Vocabulary Video Instance Segmentation
Hao Fang 0010, Peng Wu 0014, Yawei Li 0001, Xinxin Zhang 0004, Xiankai Lu |
ECCV (70) | 5 |
| 2024 | Two-phase Parametric Registration for Retinal ImagesabstractWe propose a two-phase parametric registration algorithm for retinal images. Our algorithm focuses on dealing with the geometric transformation and the intensity transformation in the retinal image registration problem. In the first phase, we efficiently detect only one pair of feature points in the source and the target retinal images to estimate a translation transformation and get a warped source image. In the second phase, we estimate both the intensity and the geometric transformations between the target image and the warped source image by fitting parametric expressions. The displacement field is generated by a super fast and accurate coarse-to-fine elastic registration algorithm—local all-pass filters algorithm (LAP). At each iteration of the LAP, the elastic displacement field and the intensity difference take turns being fitted by two different low-order polynomial functions. The fitting steps are performed by solving linear systems of equations efficiently. Experiments on real retinal image datasets demonstrated the high accuracy and computational efficiency of the proposed retinal image registration method. Xinxin Zhang 0004, Xiankai Lu, Jizhou Li, Yongshun Gong, Qiangchang Wang, Yilong Yin |
ICME | 2 |
| 2024 | Structural Transformer with Region Strip Attention for Video Object Segmentation
Qingfeng Guan 0002, Hao Fang 0010, Chenchen Han, Zhicheng Wang 0017, Ruiheng Zhang 0001, Xiankai Lu |
Neurocomputing | 7 |
| 2024 | Video Corpus Moment Retrieval via Deformable Multigranularity Feature Fusion and Adversarial TrainingabstractAs a new emerging task, video corpus moment retrieval (VCMR) aims to find the video segments relevant to a given natural language query from a large number of untrimmed videos. It mainly includes two subtasks, finding the most relevant video based on the query text (video retrieval), and locating the segment most relevant to a given query in a video (moment localization). At the same time, since videos often contain rich multi-modal information such as audio, text, and images, how to align and interact with the multi-modal information of videos and the text information of natural language queries across modalities is the core issue of this task. This article proposes a Deformable Multigranularity Feature Fusion with Adversarial Training Network (DMFAT), first inputs the subtitle and frame multi-modal information of the video into our Multi-Scale Deformable Attention module and performs multi-granularity feature fusion through Deformable Attention respectively. Then, guided by the query, adaptive weights are generated to fuse the two multi-granularity modality features of the video. Finally, the cross-modal representation of the query and video features is obtained through a bidirectional attention module, and an adversarial contrastive learning objective is introduced to enhance more precise moment localization. Our model is evaluated on two representative video corpus moment retrieval benchmarks: TVR and DiDeMo. Extensive experiments have been conducted to demonstrate that our method outperforms existing work. Peng Zhao 0016, Jinsheng Ji, Xiankai Lu, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Quality-Aware Selective Fusion Network for V-D-T Salient Object DetectionabstractDepth images and thermal images contain the spatial geometry information and surface temperature information, which can act as complementary information for the RGB modality. However, the quality of the depth and thermal images is often unreliable in some challenging scenarios, which will result in the performance degradation of the two-modal based salient object detection (SOD). Meanwhile, some researchers pay attention to the triple-modal SOD task, namely the visible-depth-thermal (VDT) SOD, where they attempt to explore the complementarity of the RGB image, the depth image, and the thermal image. However, existing triple-modal SOD methods fail to perceive the quality of depth maps and thermal images, which leads to performance degradation when dealing with scenes with low-quality depth and thermal images. Therefore, in this paper, we propose a quality-aware selective fusion network (QSF-Net) to conduct VDT salient object detection, which contains three subnets including the initial feature extraction subnet, the quality-aware region selection subnet, and the region-guided selective fusion subnet. Firstly, except for extracting features, the initial feature extraction subnet can generate a preliminary prediction map from each modality via a shrinkage pyramid architecture, which is equipped with the multi-scale fusion (MSF) module. Then, we design the weakly-supervised quality-aware region selection subnet to generate the quality-aware maps. Concretely, we first find the high-quality and low-quality regions by using the preliminary predictions, which further constitute the pseudo label that can be used to train this subnet. Finally, the region-guided selective fusion subnet purifies the initial features under the guidance of the quality-aware maps, and then fuses the triple-modal features and refines the edge details of prediction maps through the intra-modality and inter-modality attention (IIA) module and the edge refinement (ER) module, respectively. Extensive experiments are performed on VDT-2048 dataset, and the results show that our saliency model consistently outperforms 13 state-of-the-art methods with a large margin. Our code and results are available at https://github.com/Lx-Bao/QSFNet. Liuxin Bao, Xiaofei Zhou 0003, Xiankai Lu, Yaoqi Sun, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | SAFER-STUDENT for Safe Deep Semi-Supervised Learning With Unseen-Class Unlabeled DataabstractDeep semi-supervised learning (SSL) methods aim to utilize abundant unlabeled data to improve the seen-class classification. However, in the open-world scenario, collected unlabeled data tend to contain unseen-class data, which would degrade the generalization to seen-class classification. Formally, we define the problem as safe deep semi-supervised learning with unseen-class unlabeled data. One intuitive solution is removing these unseen-class instances after detecting them during the SSL process. Nevertheless, the performance of unseen-class identification is limited by the lack of suitable score function, the uncalibrated model, and the small number of labeled data. To this end, we propose a safe SSL method called SAFER-STUDENT from the teacher-student view. First, to enhance the ability of teacher model to identify seen and unseen classes, we propose a general scoring framework calledDiscrepancy withRaw (DR). Second, based on unseen-class data mined by teacher model from unlabeled data, we calibrate student model by newly proposedUnseen-classEnergy-boundedCalibration (UEC) loss. Third, based on seen-class data mined by teacher model from unlabeled data, we proposeWeightedConfirmationBiasElimination (WCBE) loss to boost seen-class classification of student model. Extensive studies show that SAFER-STUDENT remarkably outperforms the state-of-the-art, verifying the effectiveness of our method in the under-explored problem. Rundong He, Zhongyi Han, Xiankai Lu, Yilong Yin |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Learning Feature Semantic Matching for Spatio-Temporal Video GroundingabstractSpatio-temporal video grounding (STVG) aims to localize a spatio-temporal tube, including temporal boundaries and object bounding boxes, that semantically corresponds to a given language description in an untrimmed video. The existing onestage solutions in this task face two significant challenges, namely, vision-text semantic misalignment and spatial mislocalization, which limit their performance in grounding. These two limitations are mainly caused by neglect of fine-grained alignment in crossmodality fusion and the reliance on a text-agnostic query in sequentially spatial localization. To address these issues, we propose an effective model with a newly designed Feature Semantic Matching (FSM) module based on a Transformer architecture to address the above issues. Our method introduces a crossmodal feature matching module to achieve multi-granularity alignment between video and text while preventing the weakening of important features during the feature fusion stage. Additionally, we design a query-modulated matching module to facilitate text-relevant tube construction by multiple query generation and tubulet sequence matching. To ensure the quality of tube construction, we employ a novel mismatching rectify contrastive loss to rectify the mismatching between the learnable query and the objects corresponding to the text descriptions by restricting the generated spatial query. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods on two challenging STVG benchmarks. Hao Fang 0010, Hao Zhang 0048, Jialin Gao, Xiankai Lu, Xiushan Nie, Yilong Yin |
IEEE Trans. Multim. | 5 |
| 2024 | Relational Network via Cascade CRF for Video Language GroundingabstractVideo Language Grounding is one of the most challenging cross-modal video understanding tasks. This task aims to localize a target moment semantically corresponding to a given language query in an untrimmed video. Many existing VLG methods rely on the proposal-based framework, despite the dominant performance achieved, they usually focus on interacting a few internal frames with the query to score segment proposals, trapping in the long-range dependencies when the proposal feature is limited. Meanwhile, adjacent proposals share similar visual semantics, making VLG models hard to align the accurate semantics of video-query contents and degenerating the ranking performance. To remedy the above limitations, we propose VLG-CRF by introducing the conditional random fields (CRFs) to handle the discrete yet indistinguishable proposals. Specifically, VLG-CRF consists of two cascade CRF-based modules. The AttentiveCRFs is developed for multi-modal feature fusion to better integrate temporal and semantic relation between modalities. We also devise a new variant of ConvCRFs to capture the relation of discrete segments and rectify the predicting scores to make relatively high prediction scores clustered in a range. Experiments on three benchmark datasets,i.e., Charades-STA, ActivityNet-Caption, and TACoS, show the superiority of our method and the state-of-the-art performance is achieved. Xiankai Lu, Hao Zhang 0048, Xiushan Nie, Yilong Yin, Jianbing Shen |
IEEE Trans. Multim. | 2 |
| 2023 | Exposing the Self-Supervised Space-Time Correspondence Learning via Graph KernelsabstractSelf-supervised space-time correspondence learning is emerging as a promising way of leveraging unlabeled video. Currently, most methods adapt contrastive learning with mining negative samples or reconstruction adapted from the image domain, which requires dense affinity across multiple frames or optical flow constraints. Moreover, video correspondence predictive models require mining more inherent properties in videos, such as structural information. In this work, we propose the VideoHiGraph, a space-time correspondence framework based on a learnable graph kernel. Concerning the video as the spatial-temporal graph, the learning objectives of VideoHiGraph are emanated in a self-supervised manner for predicting unobserved hidden graphs via graph kernel manner. We learn a representation of the temporal coherence across frames in which pairwise similarity defines the structured hidden graph, such that a biased random walk graph kernel along the sub-graph can predict long-range correspondence. Then, we learn a refined representation across frames on the node-level via a dense graph kernel. The self-supervision of the model training is formed by the structural and temporal consistency of the graph. VideoHiGraph achieves superior performance and demonstrates its robustness across the benchmark of label propagation tasks involving objects, semantic parts, keypoints, and instances. Our algorithm implementations have been made publicly available at https://github.com/zyqin19/VideoHiGraph. Zheyun Qin, Xiankai Lu, Xiushan Nie, Yilong Yin, Jianbing Shen |
AAAI | 2 |
| 2023 | From Coarse to Fine: Learning Semantic Relations for Hyperspectral Image ClassificationabstractHyperspectral images (HSIs) is consisted of many narrow spectral bands which are capable of recording abundant features including both the spectral and spatial signatures information which have been widely used in various fields, such as urban planning, disaster monitoring. Due to the large number and high similarity of the collected spectral brands, many methods are developed to handle the problem of extracting effective features. Recently, many CNN-based methods have been proposed by exploiting the spectral-spatial signatures of the HSIs data and achieved promising results. Although many methods adopt patch-based input pattern to emphasize the importance of the spatial neighbor information of each pixel, the relations are still limited to a small area around the pixel and the latent relations among the pixels belonging to different semantic categories at the boundary are still not well exploited. To explore the relationships between pixels from a more global perspective, a neighbor-based relation mining framework is proposed to explore the long-range relations among different local regions. Experiments are conducted on two hyperspectral image classification datasets and the results demonstrate the effectiveness of the proposed long-range relations mining scheme by comparison with some state-of-the-art methods. Jinsheng Ji, Xiankai Lu, Tao Zhang 0027, Yiyou Guo, Huan Xie 0001 |
IGARSS | 2 |
| 2023 | From Coarse to Fine: Knowledge Distillation for Remote Sensing Scene ClassificationabstractScene classification is one of the most commonly studied areas of parsing the earth observation data. How to effectively interpreting the remote sensing images and extracting informative features are the great challenges for remote sensing image classification. Many important applications, such as land management and urban analysis, are based on the performance of remote sensing classification model. Recently, a lot of CNN based methods have been proposed and achieve promising results. Inspired by the success of knowledge distillation which transfers the learned information from a teacher model to a student model, a knowledge distillation based framework is proposed in this paper to handle the task of remote sensing scene classification from coarse to fine. Specifically, the learned knowledge from the teacher network is transformed into the coarse soft label and fine output mask to better guiding the student network to learn more informative features. Experiments are conducted on two widely used remote sensing scene datasets to evaluate the effectiveness of the proposed method and achieve comparable results compared with some state-of-the-art methods. Jinsheng Ji, Xiaoming Xi, Xiankai Lu, Yiyou Guo, Huan Xie 0001 |
IGARSS | 3 |
| 2023 | Clip Fusion with Bi-level Optimization for Human Mesh Reconstruction from Monocular VideosabstractHuman mesh reconstruction (HMR) from monocular video is the key step to many mixed reality and robotic applications. Although existing methods show promising results by capturing frames' temporal information, these methods predict human mesh with the design of implicit temporal learning modules in a sequence to frame manner. To mine more temporal information from the video, we present a bi-level clip inference network for HMR, which leverages both local motion and global context explicitly for dense 3D reconstruction. Specifically, we propose a novel bi-level temporal fusion strategy that takes both neighboring and long-range relations into consideration. In addition, different from traditional frame-wise operation, we investigate an alternative perspective by treating video-based HMR as clip-wise inference. We evaluate the proposed method on multiple datasets (3DPW, Human3.6M, and MPI-INF-3DHP) quantitatively and qualitatively, demonstrating a significant improvement over existing methods (in terms of PA-MPJPE, ACC-Error etc). Furthermore, we extend the proposed method on more challenging Multiple Shots HMR task to demonstrate its generalizability. Some visual demos can be seen https://github.com/bicf0/bicf_demo. Peng Wu 0014, Xiankai Lu, Jianbing Shen, Yilong Yin |
ACM Multimedia | 2 |
| 2023 | Unified 3D Segmenter As Prototypical ClassifiersabstractThe task of point cloud segmentation, comprising semantic, instance, and panoptic segmentation, has been mainly tackled by designing task-specific network architectures, which often lack the flexibility to generalize across tasks, thus resulting in a fragmented research landscape. In this paper, we introduce ProtoSEG, a prototype-based model that unifies semantic, instance, and panoptic segmentation tasks. Our approach treats these three homogeneous tasks as a classification problem with different levels of granularity. By leveraging a Transformer architecture, we extract point embeddings to optimize prototype-class distances and dynamically learn class prototypes to accommodate the end tasks. Our prototypical design enjoys simplicity and transparency, powerful representational learning, and ad-hoc explainability. Empirical results demonstrate that ProtoSEG outperforms concurrent well-known specialized architectures on 3D point cloud benchmarks, achieving 72.3%, 76.4% and 74.2% mIoU for semantic segmentation on S3DIS, ScanNet V2 and SemanticKITTI, 66.8% mCov and 51.2% mAP for instance segmentation on S3DIS and ScanNet V2, 62.4% PQ for panoptic segmentation on SemanticKITTI, validating the strength of our concept and the effectiveness of our algorithm. The code and models are available at https://github.com/zyqin19/PROTOSEG. Zheyun Qin, Cheng Han 0001, Qifan Wang 0001, Xiushan Nie, Yilong Yin, Xiankai Lu |
NeurIPS | 6 |
| 2023 | Missingness-Pattern-Adaptive Learning With Incomplete DataabstractMany real-world problems deal with collections of data with missing values, e.g., RNA sequential analytics, image completion, video processing, etc. Usually, such missing data is a serious impediment to a good learning achievement. Existing methods tend to use a universal model for all incomplete data, resulting in a suboptimal model for each missingness pattern. In this paper, we present a general model for learning with incomplete data. The proposed model can be appropriately adjusted with different missingness patterns, alleviating competitions between data. Our model is based on observable features only, so it does not incur errors from data imputation. We further introduce a low-rank constraint to promote the generalization ability of our model. Analysis of the generalization error justifies our idea theoretically. In additional, a subgradient method is proposed to optimize our model with a proven convergence rate. Experiments on different types of data show that our method compares favorably with typical imputation strategies and other state-of-the-art models for incomplete data. More importantly, our method can be seamlessly incorporated into the neural networks with the best results achieved. The source code is released at https://github.com/YS-GONG/missingness-patterns. Yongshun Gong, Zhibin Li 0002, Wei Liu 0007, Xiankai Lu, Xinwang Liu 0002, Ivor W. Tsang, Yilong Yin |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Attentional prototype inference for few-shot segmentation
Haoliang Sun, Xiankai Lu, Yilong Yin, Xiantong Zhen, Cees Snoek, Ling Shao 0001 |
Pattern Recognit. | 2 |
| 2023 | Reformulating Graph Kernels for Self-Supervised Space-Time Correspondence LearningabstractSelf-supervised space-time correspondence learning utilizing unlabeled videos holds great potential in computer vision. Most existing methods rely on contrastive learning with mining negative samples or adapting reconstruction from the image domain, which requires dense affinity across multiple frames or optical flow constraints. Moreover, video correspondence prediction models need to uncover more inherent properties of the video, such as structural information. In this work, we propose HiGraph+, a sophisticated space-time correspondence framework based on learnable graph kernels. By treating videos as a spatial-temporal graph, the learning objective of HiGraph+ is issued in a self-supervised manner, predicting the unobserved hidden graph via graph kernel methods. First, we learn the structural consistency of sub-graphs in graph-level correspondence learning. Furthermore, we introduce a spatio-temporal hidden graph loss through contrastive learning that facilitates learning temporal coherence across frames of sub-graphs and spatial diversity within the same frame. Therefore, we can predict long-term correspondences and drive the hidden graph to acquire distinct local structural representations. Then, we learn a refined representation across frames on the node-level via a dense graph kernel. The structural and temporal consistency of the graph forms the self-supervision of model training. HiGraph+ achieves excellent performance and demonstrates robustness in benchmark tests involving object, semantic part, keypoint, and instance labeling propagation tasks. Our algorithm implementations have been made publicly available at https://github.com/zyqin19/HiGraph. Zheyun Qin, Xiankai Lu, Dongfang Liu, Xiushan Nie, Yilong Yin, Jianbing Shen, Alexander C. Loui |
IEEE Trans. Image Process. | 2 |
| 2022 | Safe-Student for Safe Deep Semi-Supervised Learning with Unseen-Class Unlabeled DataabstractDeep semi-supervised learning (SSL) methods aim to take advantage of abundant unlabeled data to improve the algorithm performance. In this paper, we consider the problem of safe SSL scenario where unseen-class instances appear in the unlabeled data. This setting is essential and commonly appears in a variety of real applications. One intuitive solution is removing these unseen-class instances after detecting them during the SSL process. Nevertheless, the performance of unseen-class identification is limited by the small number of labeled data and ignoring the availability of unlabeled data. To take advantage of these unseen-class data and ensure performance, we propose a safe SSL method called SAFE-STUDENT from the teacher-student view. Firstly, a new scoring function called energy-discrepancy (ED) is proposed to help the teacher model improve the security of instances selection. Then, a novel unseen-class label distribution learning mechanism mitigates the unseen-class perturbation by calibrating the unseen-class label distribution. Finally, we propose an iterative optimization strategy to facilitate teacher-student network learning. Extensive studies on several representative datasets show that SAFE-STUDENT remarkably outperforms the state-of-the-art, verifying the feasibility and robustness of our method in the under-explored problem. Rundong He, Zhongyi Han, Xiankai Lu, Yilong Yin |
CVPR | 3 |
| 2022 | Self-Filtering: A Noise-Aware Sample Selection for Label Noise with Confidence Penalization
Qi Wei 0004, Haoliang Sun, Xiankai Lu, Yilong Yin |
ECCV (30) | 3 |
| 2022 | Dpnet: end-to-end Aerial Image Segmentation Via Deformable Point NetworkabstractAerial image Segmentation segmentation faces intrinsic foreground-background imbalance and background clutter distraction. To guide the segmentation model to learn more discriminative foreground ability and more invariant back-ground representation features, we design a Deformable Point Network (DPNet). It is an end-to-end segmentation network and consists of a multi-head deformable attention module that simultaneously considers foreground object information and background suppression. Specifically, we first employ a feature pyramid network to aggregate multiple-layer features to handle scale variants. And then, we further investigate deformable convolution to select some representative points for each layer and propose a differential module to implement it automatically instead of traditional dense fusion. Moreover, we incorporate the multi-head mechanism in the feature fusion to focus on the key contents from different representation regions. Experimental results on the representative iSAID, Vaihingen, and Postdam datasets demonstrate that our DPNet achieves competitive performance. Also, the multiple-head deformable attention facilitates the network convergence significantly. Yiyou Guo, Zheyun Qin, Yongtai Yang, Xiankai Lu, Huan Xie 0001, Xiaohua Tong |
IGARSS | 5 |
| 2022 | RONF: Reliable Outlier Synthesis under Noisy Feature Space for Out-of-Distribution DetectionabstractOut-of-distribution~(OOD) detection is fundamental to guaranteeing the reliability of multimedia applications during deployment in the open world. However, due to the lack of supervision signals from OOD data, the current model easily outputs overconfident predictions to OOD data during the inference phase. Several previous methods rely on large-scale auxiliary OOD datasets for model regularization. However, obtaining suitable and clean large-scale auxiliary OOD datasets is usually challenging. In this paper, we present Reliable Outlier synthesis under Noisy Feature space (RONF), which synthesizes reliable virtual outliers in noisy feature space to provide supervision signals for model regularization. Specifically, RONF first introduces a novel virtual outlier synthesis strategy Boundary Feature Mixup (BFM), which mixes up samples from the low-likelihood region of the class-conditional distribution in the feature space. However, the feature space is noisy due to the spurious features, which cause unreliable outlier synthesizing. To mitigate this problem, RONF then introduces Optimal Parameter Learning (OPL) to obtain desirable features and remove spurious features. Alongside, RONF proposes a provable and effective scoring function called Energy with Energy Discrepancy (EED) for the uncertainty measurement of OOD data. Extensive studies on several representative datasets of multimedia applications show that RONF outperforms the state-of-the-arts remarkably Rundong He, Zhongyi Han, Xiankai Lu, Yilong Yin |
ACM Multimedia | 3 |
| 2022 | Learning disentangled representation for self-supervised video object segmentation
Wenjie Hou, Zheyun Qin, Xiaoming Xi, Xiankai Lu, Yilong Yin |
Neurocomputing | 4 |
| 2022 | Deep Object Tracking With Shrinkage LossabstractIn this paper, we address the issue of data imbalance in learning deep models for visual object tracking. Although it is well known that data distribution plays a crucial role in learning and inference models, considerably less attention has been paid to data imbalance in visual tracking. For the deep regression trackers that directly learn a dense mapping from input images of target objects to soft response maps, we identify their performance is limited by the extremely imbalanced pixel-to-pixel differences when computing regression loss. This prevents existing end-to-end learnable deep regression trackers from performing as well as discriminative correlation filters (DCFs) trackers. For the deep classification trackers that draw positive and negative samples to learn discriminative classifiers, there exists heavy class imbalance due to a limited number of positive samples when compared to the number of negative samples. To balance training data, we propose a novel shrinkage loss to penalize the importance of easy training data mostly coming from the background, which facilitates both deep regression and classification trackers to better distinguish target objects from the background. We extensively validate the proposed shrinkage loss function on six benchmark datasets, including the OTB-2013, OTB-2015, UAV-123, VOT-2016, VOT-2018 and LaSOT. Equipped with our shrinkage loss, the proposed one-stage deep regression tracker achieves favorable results against state-of-the-art methods, especially in comparison with DCFs trackers. Meanwhile, our shrinkage loss generalizes well to deep classification trackers. When replacing the original binary cross entropy loss with our shrinkage loss, three representative baseline trackers achieve large performance gains, even setting new state-of-the-art results. Xiankai Lu, Chao Ma 0004, Jianbing Shen, Xiaokang Yang 0001, Ian D. Reid 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Segmenting Objects From Relational Visual DataabstractIn this article, we model a set of pixelwise object segmentation tasks - automatic video segmentation (AVS), image co-segmentation (ICS) and few-shot semantic segmentation (FSS) - in a unified view of segmenting objects from relational visual data. To this end, we propose an attentive graph neural network (AGNN) that addresses these tasks in a holistic fashion, by formulating them as a process of iterative information fusion over data graphs. It builds a fully-connected graph to efficiently represent visual data as nodes and relations between data instances as edges. The underlying relations are described by a differentiable attention mechanism, which thoroughly examines fine-grained semantic similarities between all the possible location pairs in two data instances. Through parametric message passing, AGNN is able to capture knowledge from the relational visual data, enabling more accurate object discovery and segmentation. Experiments show that AGNN can automatically highlight primary foreground objects from video sequences (i.e., automatic video segmentation), and extract common objects from noisy collections of semantically related images (i.e., image co-segmentation). AGNN can even generalize segment new categories with little annotated data (i.e., few-shot semantic segmentation). Taken together, our results demonstrate that AGNN provides a powerful tool that is applicable to a wide range of pixel-wise object pattern understanding tasks with relational visual data. Our algorithm implementations have been made publicly available at https://github.com/carrierlxk/AGNN. Xiankai Lu, Wenguan Wang, Jianbing Shen, David Crandall, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Zero-Shot Video Object Segmentation With Co-Attention Siamese NetworksabstractWe introduce a novel network, called CO-attention siamese network (COSNet), to address the zero-shot video object segmentation task in a holistic fashion. We exploit the inherent correlation among video frames and incorporate a global co-attention mechanism to further improve the state-of-the-art deep learning based solutions that primarily focus on learning discriminative foreground representations over appearance and motion in short-term temporal segments. The co-attention layers in COSNet provide efficient and competent stages for capturing global correlations and scene context by jointly computing and appending co-attention responses into a joint feature space. COSNet is a unified and end-to-end trainable framework where different co-attention variants can be derived for capturing diverse properties of the learned joint feature space. We train COSNet with pairs (or groups) of video frames, and this naturally augments training data and allows increased learning capacity. During the segmentation stage, the co-attention model encodes useful information by processing multiple reference frames together, which is leveraged to infer the frequently reappearing and salient foreground objects better. Our extensive experiments over three large benchmarks demonstrate that COSNet outperforms the current alternatives by a large margin. Our implementations are available at https://github.com/carrierlxk/COSNet. Xiankai Lu, Wenguan Wang, Jianbing Shen, David Crandall, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Distilled Siamese Networks for Visual TrackingabstractIn recent years, Siamese network based trackers have significantly advanced the state-of-the-art in real-time tracking. Despite their success, Siamese trackers tend to suffer from high memory costs, which restrict their applicability to mobile devices with tight memory budgets. To address this issue, we propose a distilled Siamese tracking framework to learn small, fast and accurate trackers (students), which capture critical knowledge from large Siamese trackers (teachers) by a teacher-students knowledge distillation model. This model is intuitively inspired by the one teacher versus multiple students learning method typically employed in schools. In particular, our model contains a single teacher-student distillation module and a student-student knowledge sharing mechanism. The former is designed using a tracking-specific distillation strategy to transfer knowledge from a teacher to students. The latter is utilized for mutual learning between students to enable in-depth knowledge understanding. Extensive empirical evaluations on several popular Siamese trackers demonstrate the generality and effectiveness of our framework. Moreover, the results on five tracking benchmarks show that the proposed distilled trackers achieve compression rates of up to 18× and frame-rates of 265 FPS, while obtaining comparable tracking accuracy compared to base models. Jianbing Shen, Yuanpei Liu, Xingping Dong, Xiankai Lu, Fahad Shahbaz Khan, Steven C. H. Hoi |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | CubeNet: X-shape connection for camouflaged object detection
Mingchen Zhuge, Xiankai Lu, Yiyou Guo, Zhihua Cai, Shuhan Chen |
Pattern Recognit. | 2 |
| 2022 | Regularized Two Granularity Loss Function for Weakly Supervised Video Moment RetrievalabstractWeakly supervised video moment retrieval or weakly supervised language moment retrieval aims to search the most relevant moment given a language query. In order to guide the model to capture the most matching video segments with the text description, we design a two-granularity loss function that simultaneously considers both video-level and instance-level relationships. Specifically, we first generate coarse video segments and regard each video segment as an instance. For video-level regularized multiple instance loss (MIL), we leverage the latent alignment between all intra-video segments (ie., positive bag) and text descriptions. Then, we classify these segments by regarding this procedure as a supervised learning task under noisy labels. With the instance-level regularized loss function, our model can learn to correct noisy instance-level labels so as to locate the more accurate frame boundary from all the positive instances. Comprehensive experimental results onActivityNetandDiDeModemonstrate that the proposed loss function sets a new state-of-the-art. Junya Teng, Xiankai Lu, Yongshun Gong, Xinfang Liu, Xiushan Nie, Yilong Yin |
IEEE Trans. Multim. | 2 |
| 2021 | Video Object Segmentation Using Global and Instance Embedding LearningabstractIn this paper, we propose a feature embedding based video object segmentation (VOS) method which is simple, fast and effective. The current VOS task involves two main challenges: object instance differentiation and cross-frame instance alignment. Most state-of-the-art matching based VOS methods simplify this task into a binary segmentation task and tackle each instance independently. In contrast, we decompose the VOS task into two subtasks: global embedding learning that segments foreground objects of each frame in a pixel-to-pixel manner, and instance feature embedding learning that separates instances. The outputs of these two subtasks are fused to obtain the final instance masks quickly and accurately. Through using the relation among different instances per-frame as well as temporal relation across different frames, the proposed network learns to differentiate multiple instances and associate them properly in one feed-forward manner. Extensive experimental results on the challenging DAVIS[34] and Youtube-VOS [57] datasets show that our method achieves better performances than most counterparts in each case. Wenbin Ge, Xiankai Lu, Jianbing Shen |
CVPR | 2 |
| 2021 | Learning Hierarchical Embedding for Video Instance SegmentationabstractIn this paper, we address video instance segmentation using a new generative model that learns effective representations of the target and background appearance. We propose to exploit hierarchical structural embedding over spatio-temporal space, which is compact, powerful, and flexible in contrast to current tracking-by-detection methods. Specifically, our model segments and tracks instances across space and time in a single forward pass, which is formulated as hierarchical embedding learning. The model is trained to locate the pixels belonging to specific instances over a video clip. We firstly take advantage of a novel mixing function to better fuse spatio-temporal embeddings. Moreover, we introduce normalizing flows to further improve the robustness of the learned appearance embedding, which theoretically extends conventional generative flows to a factorized conditional scheme. Comprehensive experiments on the video instance segmentation benchmark, i.e., YouTube-VIS, demonstrate the effectiveness of the proposed approach. Furthermore, we evaluate our method on an unsupervised video object segmentation dataset to demonstrate its generalizability. Zheyun Qin, Xiankai Lu, Xiushan Nie, Xiantong Zhen, Yilong Yin |
ACM Multimedia | 2 |
| 2021 | Multi-level dictionary learning for fine-grained images categorization with attention model
Jinsheng Ji, Yiyou Guo, Zhen Yang 0012, Tao Zhang 0027, Xiankai Lu |
Neurocomputing | 5 |
| 2021 | Paying Attention to Video Object Pattern Understandingabstract) with dynamic eye-tracking data in the unsupervised video object segmentation (UVOS) setting. For the first time, we quantitatively verified the high consistency of visual attention behavior among human observers, and found strong correlation between human attention and explicit primary object judgments during dynamic, task-driven viewing. Such novel observations provide an in-depth insight of the underlying rationale behind video object pattens. Inspired by these findings, we decouple UVOS into two sub-tasks: UVOS-driven Dynamic Visual Attention Prediction (DVAP) in spatiotemporal domain, and Attention-Guided Object Segmentation (AGOS) in spatial domain. Our UVOS solution enjoys three major advantages: 1) modular training without using expensive video segmentation annotations, instead, using more affordable dynamic fixation data to train the initial video attention module and using existing fixation-segmentation paired static/image data to train the subsequent segmentation module; 2) comprehensive foreground understanding through multi-source learning; and 3) additional interpretability from the biologically-inspired and assessable attention. Experiments on four popular benchmarks show that, even without using expensive video object mask annotations, our model achieves compelling performance compared with state-of-the-arts and enjoys fast processing speed (10 fps on a single GPU). Our collected eye-tracking data and algorithm implementations have been made publicly available at https://github.com/wenguanwang/AGS. Wenguan Wang, Jianbing Shen, Xiankai Lu, Steven C. H. Hoi, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Adaptive Region Proposal With Channel Regularization for Robust Object TrackingabstractIn this paper, we propose an adaptive region proposal scheme with feature channel regularization to facilitate robust object tracking. We consider tracking as a linear regression problem and an ensemble of correlation filters is trained on-line to distinguish the foreground target from the background. Further, we integrate adaptively learned region proposals into an enhanced two-stream tracking framework based on correlation filters. For the tracking stream, we learn two-stage cascade correlation filters on deep convolutional features to ensure competitive tracking performance. For the detection stream, we employ adaptive region proposals, which are effective in recovering target objects from tracking failures caused by heavy occlusion or out-of-view movement. In contrast to traditional tracking-by-detection methods using random samples or sliding windows, we perform target re-detection over adaptively learned region proposals. Since region proposals naturally take the objectness information into account, we show that the proposed adaptive region proposals can handle the challenging scale estimation problem as well. In addition, we observe the channel redundancy and noisy of feature representation, especially for the convolutional features. Thus, we apply a channel regularization to the correlation filter learning. Extensive experimental validations on OTB, VOT and UAV-123 datasets demonstrate that the proposed method performs favorably against state-of-the-art tracking algorithms. Xiankai Lu, Chao Ma 0004, Bingbing Ni, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Deep Neural Network Based Vehicle and Pedestrian Detection for Autonomous Driving: A SurveyabstractVehicle and pedestrian detection is one of the critical tasks in autonomous driving. Since heterogeneous techniques have been proposed, the selection of a detection system with an appropriate balance among detection accuracy, speed and memory consumption for a specific task has become very challenging. To deal with this issue and to provide guidance for model selection, this paper analyzes several mainstream object detection architectures, including Faster R-CNN, R-FCN, and SSD, along with several typical feature extractors, such as ResNet50, ResNet101, MobileNet_V1, MobileNet_V2, Inception_V2 and Inception_ResNet_V2. By conducting extensive experiments using the KITTI benchmark, which is a commonly used street dataset, we demonstrate that Faster R-CNN ResNet50 obtains the best average precision (AP) (58%) for vehicle and pedestrian detection, with a speed of 8.6 FPS. Faster R-CNN Inception_V2 performs best for detecting cars and detecting pedestrians respectively (74.5% and 47.3%). ResNet101 consumes the highest memory (9907 MB) and has the largest number of parameters (64.42 millions), and Inception_ResNet_V2 is the slowest model (3.05 FPS). SSD MobileNet_V2 is the fastest model (70 FPS), and SSD MobileNet_V1 is the lightest model in terms of memory usage (875 MB), both of which are suitable for applications on mobile and embedded devices. Long Chen 0005, Shaobo Lin, Xiankai Lu, Dongpu Cao, Hangbin Wu, Chi Guo, Chun Liu 0003, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Learning Video Object Segmentation From Unlabeled VideosabstractWe propose a new method for video object segmentation (VOS) that addresses object pattern learning from unlabeled videos, unlike most existing methods which rely heavily on extensive annotated data. We introduce a unified unsupervised/weakly supervised learning framework, called MuG, that comprehensively captures intrinsic properties of VOS at multiple granularities. Our approach can help advance understanding of visual patterns in VOS and significantly reduce annotation burden. With a carefully-designed architecture and strong representation learning ability, our learned model can be applied to diverse VOS settings, including object-level zero-shot VOS, instance-level zero-shot VOS, and one-shot VOS. Experiments demonstrate promising performance in these settings, as well as the potential of MuG in leveraging unlabeled data to further improve the segmentation accuracy. Xiankai Lu, Wenguan Wang, Jianbing Shen, Yu-Wing Tai, David Crandall, Steven C. H. Hoi |
CVPR | 1 |
| 2020 | Video Object Segmentation with Episodic Graph Memory Networks
Xiankai Lu, Wenguan Wang, Martin Danelljan, Tianfei Zhou, Jianbing Shen, Luc Van Gool |
ECCV (3) | 1 |
| 2020 | Geospatial Object Detection with Single Shot Anchor-Free NetworkabstractGeospatial object detection has made considerable progress with the use of anchor-based object detectors. In such a situation, the detection performance relies heavily on the parameter settings of anchor boxes. We present a Single Shot Anchor-Free Network (SSAFNet) to tackle with this problem. By eliminating the anchor boxes, the SSAFNet completely avoids the carefully predefined anchor boxes parameters and the computation for adapting the huge scale variation of geospatial objects. A compositional attention network is further introduced to enhance the saliency of foreground objects. We evaluate the SSAFNet on the representative NWPU VHR-10 and RSOD datasets, achieving competitive performance with state-of-the-art anchor-based detection methods. Yiyou Guo, Jinsheng Ji, Xiankai Lu, Huan Xie 0001, Xiaohua Tong |
IGARSS | 3 |
| 2020 | M2 Net: Multi-modal Multi-channel Network for Overall Survival Time Prediction of Brain Tumor Patients
Tao Zhou 0002, Huazhu Fu, Yu Zhang 0009, Changqing Zhang 0002, Xiankai Lu, Jianbing Shen, Ling Shao 0001 |
MICCAI (2) | 5 |
| 2020 | Novel visual tracking approach via ant lion optimiserabstractAnt lion optimiser (ALO) is a new nature‐inspired swarm intelligence optimisation algorithm that mimics the hunting mechanism of antlions in nature. ALO has been proved to have the merits of high exploitation and convergence speed benefiting from adaptive boundary shrinking mechanism and elitism. In this work, visual tracking is expressed as searching for object in whole search space by interaction between antlions and ants. A novel ALO‐based visual tracking framework is proposed and the adaptation and sensitivity of the parameters in ALO are discussed to improve tracking performance. In addition, considering that ALO tracker needs a lot of iteration consumption, kernel correlation filter with deep feature is integrated into the ALO tracking framework (ALOKCF) to improve track efficiency. Extensive experimental results prove that the ALO tracker is very competitive compared to other trackers, especially for abrupt motion tracking. At the same time, two visual tracking benchmarks are used to verify ALOKCF tracker achieves state‐of‐the‐art performance. Huanlong Zhang, Zeng Gao, Jie Zhang 0066, Xiankai Lu, Jian Chen 0038, Guohao Nie, Xiaoliang Qian |
IET Image Process. | 4 |
| 2020 | Learning Driving Models From Parallel End-to-End Driving Data SetabstractParallel end-to-end driving aims to improve the performance of end-to-end driving models using both simulated- and real-world data. However, how to efficiently utilize the data from both the simulated world and the real world remains a difficult issue, since these data are usually not well aligned. In this article, we build a data set called the parallel end-to-end driving data set (PED) for parallel end-to-end driving research. PED consists of 13 000 images from the simulated world and 13 000 images from the real world that are used to train the model, as well as 2700 images from the real world that are used to test the model. The simulated-world data in PED are constructed according to the real world, and each simulated-world image corresponds to a real-world image. PED also contains the vehicle measurement data (GPS, speed, steering angle, and heading direction of the vehicle) related to both the simulated- and real-world images, which are not available in some other data sets. We conduct two types of experiments to illustrate the effectiveness and the superiority of PED and explore some ways to mix the simulated-world data with the real-world data to improve the performance of end-to-end driving models. Long Chen 0005, Qing Wang 0025, Xiankai Lu, Dongpu Cao, Fei-Yue Wang 0001 |
Proc. IEEE | 3 |
| 2019 | See More, Know More: Unsupervised Video Object Segmentation With Co-Attention Siamese NetworksabstractWe introduce a novel network, called as CO-attention Siamese Network (COSNet), to address the unsupervised video object segmentation task from a holistic view. We emphasize the importance of inherent correlation among video frames and incorporate a global co-attention mechanism to improve further the state-of-the-art deep learning based solutions that primarily focus on learning discriminative foreground representations over appearance and motion in short-term temporal segments. The co-attention layers in our network provide efficient and competent stages for capturing global correlations and scene context by jointly computing and appending co-attention responses into a joint feature space. We train COSNet with pairs of video frames, which naturally augments training data and allows increased learning capacity. During the segmentation stage, the co-attention model encodes useful information by processing multiple reference frames together, which is leveraged to infer the frequently reappearing and salient foreground objects better. We propose a unified and end-to-end trainable framework where different co-attention variants can be derived for mining the rich context within videos. Our extensive experiments over three large benchmarks manifest that COSNet outperforms the current alternatives by a large margin. We will publicly release our implementation and models. Xiankai Lu, Wenguan Wang, Chao Ma 0004, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
CVPR | 1 |
| 2019 | Human-Aware Motion DeblurringabstractThis paper proposes a human-aware deblurring model that disentangles the motion blur between foreground (FG) humans and background (BG). The proposed model is based on a triple-branch encoder-decoder architecture. The first two branches are learned for sharpening FG humans and BG details, respectively; while the third one produces global, harmonious results by comprehensively fusing multi-scale deblurring information from the two domains. The proposed model is further endowed with a supervised, human-aware attention mechanism in an end-to-end fashion. It learns a soft mask that encodes FG human information and explicitly drives the FG/BG decoder-branches to focus on their specific domains. Above designs lead to a fully differentiable motion deblurring network, which can be trained end-to-end. To further benefit the research towards Human-aware Image Deblurring, we introduce a large-scale dataset, named HIDE, which consists of 8,422 blurry and sharp image pairs with 65,784 densely annotated FG human bounding boxes. HIDE is specifically built to span a broad range of scenes, human object sizes, motion patterns, and background complexities. Extensive experiments on public benchmarks and our dataset demonstrate that our model performs favorably against the state-of-the-art motion deblurring methods, especially in capturing semantic details. Ziyi Shen, Wenguan Wang, Xiankai Lu, Jianbing Shen, Haibin Ling, Tingfa Xu, Ling Shao 0001 |
ICCV | 3 |
| 2019 | Zero-Shot Video Object Segmentation via Attentive Graph Neural NetworksabstractThis work proposes a novel attentive graph neural network (AGNN) for zero-shot video object segmentation (ZVOS). The suggested AGNN recasts this task as a process of iterative information fusion over video graphs. Specifically, AGNN builds a fully connected graph to efficiently represent frames as nodes, and relations between arbitrary frame pairs as edges. The underlying pair-wise relations are described by a differentiable attention mechanism. Through parametric message passing, AGNN is able to efficiently capture and mine much richer and higher-order relations between video frames, thus enabling a more complete understanding of video content and more accurate foreground estimation. Experimental results on three video segmentation datasets show that AGNN sets a new state-of-the-art in each case. To further demonstrate the generalizability of our framework, we extend AGNN to an additional task: image object co-segmentation (IOCS). We perform experiments on two famous IOCS datasets and observe again the superiority of our AGNN model. The extensive experiments verify that AGNN is able to learn the underlying semantic/appearance relationships among video frames or related images, and discover the common objects. Wenguan Wang, Xiankai Lu, Jianbing Shen, David Crandall, Ling Shao 0001 |
ICCV | 2 |
| 2019 | AggregationNet: Identifying Multiple Changes Based on Convolutional Neural Network in Bitemporal Optical Remote Sensing Images
Qiankun Ye, Xiankai Lu, Lihong Wan, Yiyou Guo |
PAKDD (3) | 2 |
| 2019 | Adaptive convolutional layer selection based on historical retrospect for visual trackingabstractVisual tracking has recently gained a great advance with the use of the convolutional neural network (CNN). Usually, existing CNN‐based trackers exploit the features from a single layer or a certain combination of multiple layers. However, these features only characterise an object from an invariable aspect and cannot adapt to scene variation, which limits the performance of such trackers. To overcome this limitation, the authors study the problem from a new perspective and propose a novel convolutional layer selection method. To obtain robust appearance representation, they investigate the advantages of features extracted from different convolutional layers. To determine the correctness of the tracking prediction and updated model, they design a verification mechanism based on historical retrospect, which can estimate the deviation for each layer by bidirectionally locating the target. Meanwhile, the deviation works as the layer‐wise selection criteria. Extensive evaluations on the OTB‐2013, visual object tracking (VOT)‐2016 and VOT‐2017 benchmarks demonstrate that the proposed tracker performs favourably against several state‐of‐the‐art trackers. Fuhui Tang, Xiankai Lu, Lingkun Luo, Shiqiang Hu, Huanlong Zhang |
IET Comput. Vis. | 2 |
| 2019 | Learning transform-aware attentive network for object tracking
Xiankai Lu, Bingbing Ni, Chao Ma 0004, Xiaokang Yang 0001 |
Neurocomputing | 1 |
| 2019 | Deep feature tracking based on interactive multiple model
Fuhui Tang, Xiankai Lu, Shiqiang Hu, Huanlong Zhang |
Neurocomputing | 2 |
| 2019 | Robust visual tracking based on spatial context pyramid
Fuhui Tang, Xiankai Lu, Shiqiang Hu, Huanlong Zhang |
Multim. Tools Appl. | 3 |
| 2019 | Learning channel-aware deep regression for object tracking
Xiankai Lu, Fuhui Tang |
Pattern Recognit. Lett. | 1 |
| 2018 | Deep Regression Tracking with Shrinkage Loss
Xiankai Lu, Chao Ma 0004, Bingbing Ni, Xiaokang Yang 0001, Ian D. Reid 0001, Ming-Hsuan Yang 0001 |
ECCV (14) | 1 |
| 2018 | Non-convex joint bilateral guided depth upsampling
Xiankai Lu, Yiyou Guo, Na Liu 0007, Lihong Wan |
Multim. Tools Appl. | 1 |
| 2017 | SIFT flow for abrupt motion tracking via adaptive samples selection with sparse representation
Huanlong Zhang, Yanfeng Wang 0002, Lingkun Luo, Xiankai Lu |
Neurocomputing | 4 |
| 2015 | Efficient image categorization with sparse Fisher vectorabstractIn object recognition, Fisher vector (FV) representation is one of the state-of-the-art image representations ways at the expense of dense, high dimensional features and increased computation time. A simplification of FV is attractive, so we propose Sparse Fisher vector (SFV). By incorporating locality strategy, we can accelerate the Fisher coding step in image categorization which is implemented from a collective of local descriptors. Combining with pooling step, we explore the relationship between coding step and pooling step to give a theoretical explanation about SFV. Experiments on benchmark datasets have shown that SFV leads to a speedup of several-fold of magnitude compares with FV, while maintaining the categorization performance. In addition, we demonstrate how SFV preserves the consistence in representation of similar local features. Xiankai Lu, Haiting Zhang, Hongya Tuo |
ICASSP | 1 |