Xiangyang Li 0002

dblp:80/4579-2 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-3944-4704ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Continual novel class discovery under domain shift with entropy-based selection and representation evolution
Feifei Shi, Xiangyang Li 0002, Shuqiang Jiang, Yong Rui
Multim. Syst.2
2024 Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language Navigation
abstract
Vision-and-language navigation (VLN) enables the agent to navigate to a remote location following the natural language instruction in 3D environments. At each navigation step, the agent selects from possible candidate locations and then makes the move. For better navigation planning, the lookahead exploration strategy aims to effectively evaluate the agent's next action by accurately anticipating the future environment of candidate locations. To this end, some existing works predict RGB images for future environments, while this strategy suffers from image distortion and high computational cost. To address these issues, we propose the pre-trained hierarchical neural radiance representation model (HNR) to produce multi-level semantic features for future environments, which are more robust and efficient than pixel-wise RGB reconstruction. Furthermore, with the predicted future environmental representations, our lookahead VLN model is able to construct the navigable future path tree and select the optimal path via efficient parallel evaluation. Extensive experiments on the VLN-CE datasets confirm the effectiveness of our method. The code is available at https://github.com/MrZihan/HNR-VLN
Xiangyang Li 0002, Yeqi Liu, Junjie Hu 0001, Ming Jiang 0018, Shuqiang Jiang
CVPR2
2023 KERM: Knowledge Enhanced Reasoning for Vision-and-Language Navigation
abstract
Vision-and-language navigation (VLN) is the task to enable an embodied agent to navigate to a remote location following the natural language instruction in real scenes. Most of the previous approaches utilize the entire features or object-centric features to represent navigable candidates. However, these representations are not efficient enough for an agent to perform actions to arrive the target location. As knowledge provides crucial information which is complementary to visible content, in this paper, we propose a Knowledge Enhanced Reasoning Model (KERM) to leverage knowledge to improve agent navigation ability. Specifically, we first retrieve facts (i.e., knowledge described by language descriptions) for the navigation views based on local regions from the constructed knowledge base. The re-trieved facts range from properties of a single object (e.g., color, shape) to relationships between objects (e.g., action, spatial position), providing crucial information for VLN. We further present the KERM which contains the purification, fact-aware interaction, and instruction-guided aggregation modules to integrate visual, history, instruction, and fact features. The proposed KERM can automatically select and gather crucial and relevant cues, obtaining more accurate action prediction. Experimental results on the REVERIE, R2R, and SOON datasets demonstrate the effectiveness of the proposed method. The source code is available at https://github.com/XiangyangLi20/KERM.
Xiangyang Li 0002, Yaowei Wang 0001, Shuqiang Jiang
CVPR1
2023 GridMM: Grid Memory Map for Vision-and-Language Navigation
abstract
Vision-and-language navigation (VLN) enables the agent to navigate to a remote location following the natural language instruction in 3D environments. To represent the previously visited environment, most approaches for VLN implement memory using recurrent states, topological maps, or top-down semantic maps. In contrast to these approaches, we build the top-down egocentric and dynamically growing Grid Memory Map (i.e., GridMM) to structure the visited environment. From a global perspective, historical observations are projected into a unified grid map in a top-down view, which can better represent the spatial relations of the environment. From a local perspective, we further propose an instruction relevance aggregation method to capture fine-grained visual clues in each grid region. Extensive experiments are conducted on both the REVERIE, R2R, SOON datasets in the discrete environments, and the R2R-CE dataset in the continuous environments, showing the superiority of our proposed method. The source code is available at https://github.com/MrZihan/GridMM.
Xiangyang Li 0002, Yeqi Liu, Shuqiang Jiang
ICCV2
2023 Dataset Bias in Few-Shot Image Recognition
abstract
The goal of few-shot image recognition (FSIR) is to identify novel categories with a small number of annotated samples by exploiting transferable knowledge from training data (base categories). Most current studies assume that the transferable knowledge can be well used to identify novel categories. However, such transferable capability may be impacted by the dataset bias, and this problem has rarely been investigated before. Besides, most of few-shot learning methods are biased to different datasets, which is also an important issue that needs to be investigated deeply. In this paper, we first investigate the impact of transferable capabilities learned from base categories. Specifically, we use the relevance to measure relationships between base categories and novel categories. Distributions of base categories are depicted via the instance density and category diversity. The FSIR model learns better transferable knowledge from relevant training data. In the relevant data, dense instances or diverse categories can further enrich the learned knowledge. Experimental results on different sub-datasets of Imagenet demonstrate category relevance, instance density and category diversity can depict transferable bias from distributions of base categories. Second, we investigate performance differences on different datasets from the aspects of dataset structures and different few-shot learning methods. Specifically, we introduce image complexity, intra-concept visual consistency, and inter-concept visual similarity to quantify characteristics of dataset structures. We use these quantitative characteristics and eight few-shot learning methods to analyze performance differences on multiple datasets. Based on the experimental analysis, some insightful observations are obtained from the perspective of both dataset structures and few-shot learning methods. We hope these observations are useful to guide future few-shot learning research on new datasets or tasks. Our data is available at http://123.57.42.89/dataset-bias/dataset-bias.html.
Shuqiang Jiang, Chenlong Liu, Xinhang Song, Xiangyang Li 0002, Weiqing Min
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 TransWeaver: Weave Image Pairs for Class Agnostic Common Object Detection
abstract
Measuring the similarity of two images is of crucial importance in computer vision. Class agnostic common object detection is a nascent research topic about mining image similarity, which aims to detect common object pairs from two images without category information. This task is general and less restrictive which explores the similarity between objects and can further describe the commonality of image pairs at the object level. However, previous works suffer from features with low discrimination caused by the lack of category information. Moreover, most existing methods compare objects extracted from two images in a simple and direct way, ignoring the internal relationships between objects in the two images. To overcome these limitations, in this paper, we propose a new framework called TransWeaver, which learns intrinsic relationships between objects. Our TransWeaver takes image pairs as input and flexibly captures the inherent correlation between candidate objects from two images. It consists of two modules (i.e., the representation-encoder and the weave-decoder) and captures efficient context information by weaving image pairs to make them interact with each other. The representation-encoder is used for representation learning, which can obtain more discriminative representations for candidate proposals. Furthermore, the weave-decoder weaves the objects from two images and is able to explore the inter-image and intra-image context information at the same time, bringing a better object matching ability. We reorganize the PASCAL VOC, COCO, and Visual Genome datasets to obtain training and testing image pairs. Extensive experiments demonstrate the effectiveness of the proposed TransWeaver which achieves state-of-the-art performance on all datasets.
Xiaoqian Guo, Xiangyang Li 0002, Yaowei Wang 0001, Shuqiang Jiang
IEEE Trans. Image Process.2
2023 MemBridge: Video-Language Pre-Training With Memory-Augmented Inter-Modality Bridge
abstract
Video-language pre-training has attracted considerable attention recently for its promising performance on various downstream tasks. Most existing methods utilize the modality-specific or modality-joint representation architectures for the cross-modality pre-training. Different from previous methods, this paper presents a novel architecture named Memory-augmented Inter-Modality Bridge (MemBridge), which uses the learnable intermediate modality representations as the bridge for the interaction between videos and language. Specifically, in the transformer-based cross-modality encoder, we introduce the learnable bridge tokens as the interaction approach, which means the video and language tokens can only perceive information from bridge tokens and themselves. Moreover, a memory bank is proposed to store abundant modality interaction information for adaptively generating bridge tokens according to different cases, enhancing the capacity and robustness of the inter-modality bridge. Through pre-training, MemBridge explicitly models the representations for more sufficient inter-modality interaction. Comprehensive experiments show that our approach achieves competitive performance with previous methods on various downstream tasks including video-text retrieval, video captioning, and video question answering on multiple datasets, demonstrating the effectiveness of the proposed method. The code has been available at https://github.com/jahhaoyang/MemBridge.
Xiangyang Li 0002, Mao Zheng, Xiaoqian Guo, Yuchen Yuan, Zifeng Chai, Shuqiang Jiang
IEEE Trans. Image Process.2
2023 Focus and Align: Learning Tube Tokens for Video-Language Pre-Training
abstract
Video-language pre-training (VLP) has attracted increasing attention for cross-modality understanding tasks. To enhance visual representations, recent works attempt to adopt transformer-based architectures as video encoders. These works usually focus on the visual representations of the sampled frames. Compared with frame representations, frame patches incorporate more fine-grained spatio-temporal information, which could lead to a better understanding of video contents. However, how to exploit the spatio-temporal information within frame patches for VLP has been less investigated. In this work, we propose a method to learn tube tokens to model the key spatio-temporal information from frame patches. To this end, multiple semantic centers are introduced to focus on the underlying patterns of frame patches. Based on each semantic center, the spatio-temporal information within frame patches is integrated into a unique tube token. Complementary to frame representations, tube tokens provide detailed clues of video contents. Furthermore, to better align the generated tube tokens and the contents of descriptions, a local alignment mechanism is introduced. The experiments based on a variety of downstream tasks demonstrate the effectiveness of the proposed method.
Xiangyang Li 0002, Mao Zheng, Xiaoqian Guo, Zifeng Chai, Yuchen Yuan, Shuqiang Jiang
IEEE Trans. Multim.2
2020 Expressional Region Retrieval
abstract
Image retrieval is a long-standing topic in the multimedia community due to its various applications, e.g., product search and artworks retrieval in museum. The regions in images contain a wealth of information. Users may be interested in the objects presented in the image regions or the relationships between them. But previous retrieval methods are either limited to the single object of images, or tend to the entire visual scene. In this paper, we introduce a new task called expressional region retrieval, in which the query is formulated as a region of image with the associated description. The goal is to find images containing the similar content with the query and localize the regions within them. As far as we know, this task has not been explored yet. We propose a framework to address this issue. The region proposals are first generated based on region detectors and language features are extracted. Then the Gated Residual Network (GRN) takes language information as a gate to control the transformation of visual features. In this way, the combined visual and language representation is more specific and discriminative for expressional region retrieval. We evaluate our method on a new established benchmark which is constructed based on the Visual Genome dataset. Experimental results demonstrate that our model effectively utilizes both visual and language information, outperforming the baseline methods.
Xiaoqian Guo, Xiangyang Li 0002, Shuqiang Jiang
ACM Multimedia2
2019 Learning Object Context for Dense Captioning
abstract
Dense captioning is a challenging task which not only detects visual elements in images but also generates natural language sentences to describe them. Previous approaches do not leverage object information in images for this task. However, objects provide valuable cues to help predict the locations of caption regions as caption regions often highly overlap with objects (i.e. caption regions are usually parts of objects or combinations of them). Meanwhile, objects also provide important information for describing a target caption region as the corresponding description not only depicts its properties, but also involves its interactions with objects in the image. In this work, we propose a novel scheme with an object context encoding Long Short-Term Memory (LSTM) network to automatically learn complementary object context for each caption region, transferring knowledge from objects to caption regions. All contextual objects are arranged as a sequence and progressively fed into the context encoding module to obtain context features. Then both the learned object context features and region features are used to predict the bounding box offsets and generate the descriptions. The context learning procedure is in conjunction with the optimization of both location prediction and caption generation, thus enabling the object context encoding LSTM to capture and aggregate useful object context. Experiments on benchmark datasets demonstrate the superiority of our proposed approach over the state-of-the-art methods.
Xiangyang Li 0002, Shuqiang Jiang, Jungong Han
AAAI1
2019 Class Agnostic Image Common Object Detection
abstract
Learning similarity of two images is an important problem in computer vision and has many potential applications. Most of previous works focus on generating image similarities in three aspects: global feature distance computing, local feature matching and image concepts comparison. However, the task of directly detecting class agnostic common objects from two images has not been studied before, which goes one step further to capture image similarities at region level. In this paper, we propose an end-to-end Image Common Object Detection Network (CODN) to detect class agnostic common objects from two images. The proposed method consists of two main modules: locating module and matching module. The locating module generates candidate proposals of each two images. The matching module learns the similarities of the candidate proposal pairs from two images, and refines the bounding boxes of the candidate proposals. The learning procedure of CODN is implemented in an integrated way and a multi-task loss is designed to guarantee both region localization and common object matching. Experiments are conducted on PASCAL VOC 2007 and COCO 2014 datasets. Experimental results validate the effectiveness of the proposed method.
Shuqiang Jiang, Sisi Liang, Chengpeng Chen, Xiangyang Li 0002
IEEE Trans. Image Process.5
2019 Know More Say Less: Image Captioning Based on Scene Graphs
abstract
Automatically describing the content of an image has been attracting considerable research attention in the multimedia field. To represent the content of an image, many approaches directly utilize convolutional neural networks (CNNs) to extract visual representations, which are fed into recurrent neural networks to generate natural language. Recently, some approaches have detected semantic concepts from images and then encoded them into high-level representations. Although substantial progress has been achieved, most of the previous methods treat entities in images individually, thus lacking structured information that provides important cues for image captioning. In this paper, we propose a framework based on scene graphs for image captioning. Scene graphs contain abundant structured information because they not only depict object entities in images but also present pairwise relationships. To leverage both visual features and semantic knowledge in structured scene graphs, we extract CNN features from the bounding box offsets of object entities for visual representations, and extract semantic relationship features from triples (e.g.,man riding bike) for semantic representations. After obtaining these features, we introduce a hierarchical-attention-based module to learn discriminative features for word generation at each time step. The experimental results on benchmark datasets demonstrate the superiority of our method compared with several state-of-the-art methods.
Xiangyang Li 0002, Shuqiang Jiang
IEEE Trans. Multim.1
2018 Bundled Object Context for Referring Expressions
abstract
Referring expressions are natural language descriptions of objects within a given scene. Context is of crucial importance for a referring expression, as the description not only depicts the properties of the object but also involves the relationships of the referred object with other ones. Most of previous work uses either the whole image or one particular contextual object as the context. However, the context of these approaches is holistic and insufficient, as a referring expression often describes relationships of multiple objects in an image. To leverage rich context information from all objects in an image, in this paper, we propose a novel scheme that is composed of a visual context long short-term memory (LSTM) module and a sentence LSTM module to model bundled object context for referring expressions. All contextual objects are arranged with their spatial locations and progressively fed into the visual context LSTM module to acquire and aggregate the context features. Then the concatenation of the learned context features and the features of the referred object are put into the sentence LSTM module to learn the probability of a referring expression. The feedback connections and internal gating mechanism of the LSTM cells enable our model to selectively propagate relevant contextual information through the whole network. Experiments on three benchmark datasets show that our methods can achieve promising results compared to state-of-the-art methods. Moreover, visualization of the internal states of the visual context LSTM cells also shows that our method can automatically select the pertinent context objects.
Xiangyang Li 0002, Shuqiang Jiang
IEEE Trans. Multim.1
2017 Visual relationship detection with object spatial distribution
abstract
Recently, object recognition techniques have been rapidly developed. Most of existing object recognition focused on recognizing several independent concepts. The relationship of objects is also an important problem, which shows in-depth semantic information of images. In this work, toward general visual relationship detection, we propose a method to integrate spatial distribution of object to facilitate visual relation detection. Spatial distribution can not only reflect positional relation of object but also describe structural information between objects. Spatial distributions are described with different features such as positional relation, size relation, shape relation, and so on. By combing spatial distribution features with visual and concept features, we establish a modeling method to make these three aspects working together to facilitate visual relationship detection. To evaluate the proposed method, we conduct experiments on two datasets, which are the Stanford VRD dataset, and a newly proposed larger new dataset which contains 15k images. Experimental results demonstrate that our approach is effective.
Shuqiang Jiang, Xiangyang Li 0002
ICME3
2017 Modality-specific and hierarchical feature learning for RGB-D hand-held object recognition
Xiong Lv, Xinda Liu, Xiangyang Li 0002, Shuqiang Jiang, Zhiqiang He 0002
Multim. Tools Appl.3
2016 Scene Recognition with CNNs: Objects, Scales and Dataset Bias
abstract
Since scenes are composed in part of objects, accurate recognition of scenes requires knowledge about both scenes and objects. In this paper we address two related problems: 1) scale induced dataset bias in multi-scale convolutional neural network (CNN) architectures, and 2) how to combine effectively scene-centric and object-centric knowledge (i.e. Places and ImageNet) in CNNs. An earlier attempt, Hybrid-CNN[23], showed that incorporating ImageNet did not help much. Here we propose an alternative method taking the scale into account, resulting in significant recognition gains. By analyzing the response of ImageNet-CNNs and Places-CNNs at different scales we find that both operate in different scale ranges, so using the same network for all the scales induces dataset bias resulting in limited performance. Thus, adapting the feature extractor to each particular scale (i.e. scale-specific CNNs) is crucial to improve recognition, since the objects in the scenes have their specific range of scales. Experimental results show that the recognition accuracy highly depends on the scale, and that simple yet carefully chosen multi-scale combinations of ImageNet-CNNs and Places-CNNs, can push the stateof-the-art recognition accuracy in SUN397 up to 66.26% (and even 70.17% with deeper architectures, comparable to human performance).
Luis Herranz, Shuqiang Jiang, Xiangyang Li 0002
CVPR3
2016 Image Captioning with both Object and Scene Information
abstract
Recently, automatic generation of image captions has attracted great interest not only because of its extensive applications but also because it connects computer vision and natural language processing. By combining convolutional neural networks (CNNs), which learn visual representations from images, and recurrent neural networks (RNNs), which translate the learned features into text sequences, the content of a image can be transformed into linguistic sequences. Existing approaches typically focus on visual features extracted form an object-oriented CNN (train on ImageNet) and then decode them into natural language. In this paper, we propose a novel model using not only object-related, but also scene-related information extracted from the images. To make full use of both object and scene information, we first combine object information and scene information (extracted from a scene-oriented CNN), and then using as inputs to RNNs. Both types of information provide complementary aspects that help in generating a more complete description of the image. Qualitative and quantitative evaluation results validate the effectiveness of our method.
Xiangyang Li 0002, Xinhang Song, Luis Herranz, Shuqiang Jiang
ACM Multimedia1