Ye Xiang

dblp:21/2560 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
15since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2025 Multi-scale motion-based relational reasoning for group activity recognition
Yihao Zheng 0002, Zhuming Wang, Lifang Wu, Zun Li 0001, Ye Xiang
Eng. Appl. Artif. Intell.6
2024 Exploring Spatio-Temporal Discriminative Cues for Group Activity Recognition Via Contrastive Learning
abstract
Group activity recognition is a challenging task that involves multiple moving actors within a cluttered scene. Existing methods often rely on object detector to avoid individual bounding box labeling during testing, but are prone to false detections due to factors such as occlusion and background clutter. In addition, existing detector-free method based on Transformer attends to attention map that is too sparse, resulting in the loss of some important foreground information. In this paper, we introduce foreground-background contrast loss (FB-Loss) to help accurately seek discriminative cues in the foreground and eliminate noise interference in the background. Neither ground-truth bounding boxes nor object detectors are required during both training and testing. Experimental results on public datasets show that our proposed method achieves the state-of-the-art performance.
Ye Xiang, Lifang Wu
ICASSP2
2024 Learning to compose diversified prompts for image emotion classification
abstract
Image emotion classification (IEC) aims to extract the abstract emotions evoked in images. Recently, language-supervised methods such as contrastive language-image pretraining (CLIP) have demonstrated superior performance in image understanding. However, the underexplored task of IEC presents three major challenges: a tremendous training objective gap between pretraining and IEC, shared suboptimal prompts, and invariant prompts for all instances. In this study, we propose a general framework that effectively exploits the language-supervised CLIP method for the IEC task. First, a prompt-tuning method that mimics the pretraining objective of CLIP is introduced, to exploit the rich image and text semantics associated with CLIP. Subsequently, instance-specific prompts are automatically composed, conditioning them on the categories and image content of instances, diversifying the prompts, and thus avoiding suboptimal problems. Evaluations on six widely used affective datasets show that the proposed method significantly outperforms state-of-the-art methods (up to 9.29% accuracy gain on the EmotionROI dataset) on IEC tasks with only a few trained parameters. The code is publicly available at https://github.com/dsn0w/PT-DPC/for research purposes .
Sinuo Deng, Lifang Wu, Ge Shi 0002, Lehao Xing, Meng Jian, Ye Xiang, Ruihai Dong
Comput. Vis. Media6
2024 GLOCAL: A self-supervised learning framework for global and local motion estimation
Yihao Zheng 0002, Kunming Luo, Shuaicheng Liu, Zun Li 0001, Ye Xiang, Lifang Wu, Bing Zeng 0001, Chang Wen Chen
Pattern Recognit. Lett.5
2024 Learning Label Semantics for Weakly Supervised Group Activity Recognition
abstract
Weakly supervised group activity recognition deals with the dependence on individual-level annotations during understanding scenes involving multiple individuals, which is a challenging task. Existing methods either take the trained detectors to extract individual features or utilize the attention mechanisms for partial context encoding, followed by integration to form the final group-level representations. However, the detectors require individual-level annotations during the training phase and have a mis-detection issue, and the partial contexts extracted immediately from the whole complex scene are too ambiguous without the guidance of concrete semantics. In this paper, we investigate the hierarchical structure inherent in group-level labels to extract the fine-grained semantics without using detectors for weakly supervised group activity recognition. A multi-hot encoding strategy combined with a semantic encoder is first adopted to get the label semantics embeddings. The semantic and visual scene information are then fused through a semantic decoder to obtain activity-specific features. Lastly, we employ the multi-label classification and integrate the scores of hierarchical activity labels. Experimental results show that our proposed method achieves the state-of-the-art performance on three benchmarks, and the accuracy on the Volleyball dataset exceeds the second-best method by 2%.
Lifang Wu, Ye Xiang, Ke Gu 0001, Ge Shi 0002
IEEE Trans. Multim.3
2023 Graph Contrastive Learning on Complementary Embedding for Recommendation
abstract
Previous works build interest learning via mining deeply on interactions. However, the interactions come incomplete and insufficient to support interest modeling, even bringing severe bias into recommendations. To address the interaction sparsity and the consequent bias challenges, we propose a graph contrastive learning on complementary embedding (GCCE), which introduces negative interests to assist positive interests of interactions for interest modeling. To embed interest, we design a perturbed graph convolution by preventing embedding distribution from bias. Since negative samples are not available in the general scenario of implicit feedback, we elaborate a complementary embedding generation to depict users’ negative interests. Finally, we develop a new contrastive task to contrastively learn from the positive and negative interests to promote recommendation. We validate the effectiveness of GCCE on two real datasets, where it outperforms the state-of-the-art models for recommendation.
Meishan Liu, Meng Jian, Ge Shi 0002, Ye Xiang, Lifang Wu
ICMR4
2023 Simple But Powerful, a Language-Supervised Method for Image Emotion Classification
abstract
Image emotion classification is an important computer vision task to extract emotions from images. The methods for image emotion classification (IEC) are primarily based on label or distribution as a supervision signal, which neither has enough accessibility nor diversity, limiting the development of IEC research. Inspired by psychology research and the recent booming of large-scale pretrained language models. We figure out a language-supervised paradigm, which can cleverly combine the features of language and visual emotion to drive the visual model to gain stronger emotional discernment with language prompts. To practice the paradigm, we present a conceptually simple while empirically powerful framework for image emotion classification, SimEmotion. That we propose a prompt-based fine-tuning strategy to learn task-specific representations by composing a template with the emotion-level concept and entity-level information. Evaluations on four widely-used affective datasets, namely, Flickr and Instagram (FI), EmotionROI, Twitter I, and Twitter II, demonstrate that the proposed algorithm outperforms the state-of-the-art methods with a large margin (i.e.,$8.42\%$absolute accuracy gain on EmotionROI) on image emotion classification tasks. Our codes will be publicly available for research purposes.
Sinuo Deng, Lifang Wu, Ge Shi 0002, Lehao Xing, Wenjin Hu 0002, Ye Xiang
IEEE Trans. Affect. Comput.7
2023 Active Spatial Positions Based Hierarchical Relation Inference for Group Activity Recognition
abstract
Group activity recognition aims to recognize behaviors characterized by multiple individuals within a scene. Existing schemes rely on individual relation inference and usually take the individuals as tokens. Essentially they select the most relevant region of the group activity from the entire image while filtering out irrelevant background noises. However, these schemes require individual bounding box labeling in both training and testing stages. Since individuals have usually been presented at one scale, multi-scale individuals cannot be combined in an effective way. In this paper, we present a novel end-to-end hierarchical relation inference framework based on active spatial positions for group activity recognition. This framework is designed to locate active spatial positions and use them as visual tokens to infer the relations for token embeddings. It requires individual bounding box labeling only in the training stage while automatically eliminating the background after locating active spatial positions from the entire scene. The hierarchical relations can be naturally inferred based on the visual tokens at different scales, contributing to further performance improvement. Experimental results demonstrate that the proposed framework is competitive against existing schemes that require more laboring and computation to generate labels in both the training and testing stage.
Lifang Wu, Xianglong Lang, Ye Xiang, Chang Wen Chen, Zun Li 0001, Zhuming Wang
IEEE Trans. Circuits Syst. Video Technol.3
2022 SimEmotion: A Simple Knowledgeable Prompt Tuning Method for Image Emotion Classification
Sinuo Deng, Ge Shi 0002, Lifang Wu, Lehao Xing, Wenjin Hu 0002, Ye Xiang
DASFAA (3)7
2022 Part Based Interaction Learning for Group Activity Recognition
abstract
Group activity recognition is a subject with broad applications, and its main challenge is to model the interactions between individuals. Existing algorithms mostly model the interactions merely based on holistic features of persons, which completely ignore the local details and local interactions that could be significant for recognition. In this paper, we propose a novel part based interaction learning algorithm for group activity recognition. Our proposed algorithm introduces both the physical structural information and fine-grained contextual information into representations, through exploring the intraand inter-actor part interactions. Specifically, a dual-branch framework is adopted to extract the appearance and motion features respectively. For each branch, we utilize the key point detection technique for proper part division and then extract the part features. The part features are further enhanced by the transformers for intra- and inter-actor part interactions, and are lastly used for group activity recognition. Comparison with the state-of-the-arts on two public datasets demonstrate the effectiveness of our proposed algorithm.
Xianglong Lang, Ye Xiang, Yan Huang 0008, Lifang Wu
MMSP3
2021 Semantic Guided Multi-directional Mixed-Color 3D Printing
Lifang Wu, Tianqin Yang, Yupeng Guan, Ge Shi 0002, Ye Xiang, Yisong Gao
ICIG (3)5
2021 GLM-Net: Global and Local Motion Estimation via Task-Oriented Encoder-Decoder Structure
abstract
In this work, we study the problem of separating the global camera motion and the local dynamic motion from an optical flow. Previous methods either estimate global motions by a parametric model, such as a homography, or estimate both of them by an optical flow field. However, none of these methods can directly estimate global and local motions through an end-to-end manner. In addition, separating the two motions accurately from a hybrid flow field is challenging. Because one motion can easily confuse the estimate of the other one when they are compounded together. To this end, we propose an end-to-end global and local motion estimation network GLM-Net. We design two encoder-decoder structures for the motion separation in the optical flow based on different task orientations. One structure adopts a mask autoencoder to extract the global motion, while the other one uses attention U-net for the local motion refinement. We further designed two effective training methods to overcome the problem of lacking supervisions. We apply our method on the action recognition datasets NCAA and UCF-101 to verify the accuracy of the local motion, and the homography estimation dataset DHE for the accuracy of the global motion. Experimental results show that our method can achieve competitive performance in both tasks at the same time, validating the effectiveness of the motion separation.
Ye Xiang, Shuaicheng Liu, Lifang Wu, Boxuan Zhao, Bing Zeng 0001
ACM Multimedia2
2021 Relation-Guided Actor Attention for Group Activity Recognition
Lifang Wu, Qi Wang 0076, Ye Xiang, Xianglong Lang
PRCV (1)4
2021 Attribute-Level Interest Matching Network for Personalized Recommendation
Meng Jian, Ge Shi 0002, Lifang Wu, Ye Xiang
PRCV (2)5
2021 Latent label mining for group activity recognition in basketball videos
abstract
Abstract Motion information has been widely exploited for group activity recognition in sports video. However, in order to model and extract the various motion information between the adjacent frames, existing algorithms only use the coarse video‐level labels as supervision cues. This may lead to the ambiguity of extracted features and the omission of changing rules of motion patterns that are also important sports video recognition. In this paper, a latent label mining strategy for group activity recognition in basketball videos is proposed. The authors' novel strategy allows them to obtain the latent labels set for marking different frames in an unsupervised way, and build the frame‐level and video‐level representations with two separate levels of supervision signal. Firstly, the latent labels of motion patterns are digged using the unsupervised hierarchical clustering technique. The generated latent labels are then taken as the frame‐level supervision signal to train a deep CNN for the frame‐level features extraction. Lastly, the frame‐level features are fed into an LSTM network to build the spatio‐temporal representation for group activity recognition. Experimental results on the public NCAA dataset demonstrate that the proposed algorithm achieves state‐of‐the‐art performance.
Lifang Wu, Ye Xiang, Meng Jian, Jialie Shen 0001
IET Image Process.3
2020 Global Topology Constraint Network for Fine-Grained Vehicle Recognition
abstract
Fine-grained vehicle recognition is challenging due to the large intra-class variation in vehicle pose and viewpoint. Many existing methods, especially convolution neural network (CNN)-based methods, solve this problem via detecting and aligning parts individually, and do not consider the interaction between parts, which is very important for effective part detection and vehicle recognition. In this paper, we propose a global topology constraint network for fine-grained vehicle recognition, which adopts the constraint of global topology relationship to depict the interaction between parts and integrates it into CNN in an efficient way. Different CNNs for image classification truncated at intermediate layer can be used for part detection. The global topology relationship between parts is encoded into kernel of depthwise convolution layer, and can be learned from training images. The response under global topology constraint reflects the probability of topology relationship between parts existing in input image. All responses build the description for classification with translation invariance. Through training the whole network, the back-propagation of gradient information of kernel for global topology relationship will guide former layers to better detect useful parts, and thereby improve vehicle recognition. Our proposed method does not require additional annotation, such as bounding box or part annotation, and can be trained in an end-to-end way. We conduct comparison experiments on public Stanford Cars and CompCars datasets, which both show that our method achieves the state-of-the-art performance.
Ye Xiang, Ying Fu 0001, Hua Huang 0001
IEEE Trans. Intell. Transp. Syst.1
2019 Incremental Learning Using Conditional Adversarial Networks
abstract
Incremental learning using Deep Neural Networks (DNNs) suffers from catastrophic forgetting. Existing methods mitigate it by either storing old image examples or only updating a few fully connected layers of DNNs, which, however, requires large memory footprints or hurts the plasticity of models. In this paper, we propose a new incremental learning strategy based on conditional adversarial networks. Our new strategy allows us to use memory-efficient statistical information to store old knowledge, and fine-tune both convolutional layers and fully connected layers to consolidate new knowledge. Specifically, we propose a model consisting of three parts, i.e., a base sub-net, a generator, and a discriminator. The base sub-net works as a feature extractor which can be pre-trained on large scale datasets and shared across multiple image recognition tasks. The generator conditioned on labeled embeddings aims to construct pseudo-examples with the same distribution as the old data. The discriminator combines real-examples from new data and pseudo-examples generated from the old data distribution to learn representation for both old and new classes. Through adversarial training of the discriminator and generator, we accomplish the multiple continuous incremental learning. Comparison with the state-of-the-arts on public CIFAR-100 and CUB-200 datasets shows that our method achieves the best accuracies on both old and new classes while requiring relatively less memory storage.
Ye Xiang, Ying Fu 0001, Pan Ji, Hua Huang 0001
ICCV1
2019 Global relative position space based pooling for fine-grained vehicle recognition
Ye Xiang, Ying Fu 0001, Hua Huang 0001
Neurocomputing1
2006 Research on IT outsourcing based on IT systems management
abstract
To adapt to changing market conditions and intense competition more quickly than ever before, enterprises can not help equipping more and more new IT applications continuously. But new challenges emerged. Most of enterprises lack for IT technique and experience which result to ruin. In this environment, IT outsourcing is a new solution for the development and management of IT applications. However, once lack of a comprehensive IT systems management, IT outsourcing is hard to come true. This paper first briefly introduced concepts, functions and significances of IT outsourcing and IT systems management. Then it was proposed that IT systems management was the key factors to determine whether IT outsourcing succeed or not. The execution of IT Outsourcing based on IT Systems Management was emphasized by a workflow consist of three main performable phases. Present research challenges and future research trends were discussed at last.
Hongxun Jiang, Du Honglu, Ye Xiang, Su Jun
ICEC3