EDBT 2026 Demo / reviewers in the wild / expert
Yanli Ji
dblp:79/8728
· DBLP profile ↗
53ranked-venue papers
14as first author
30since 2021 · last 2026
0000-0001-9122-6141ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 10 first-author · 21 since 2021Artificial intelligence and machine learning · 16 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vision-Language Collaborative Representation Learning for Action Quality AssessmentabstractAction Quality Assessment (AQA) has gained significant attention due to its potential real-world applications, which require a fine-grained understanding of action sequences. Recent works have attempted to utilize multimodal video features and address some existing challenges. However, these approaches primarily focus on leveraging textual information from language models only, leading to instability and suboptimal performance due to directional bias in a vision-language joint embedding space. To tackle these issues, we propose a Vision-Language Collaboration Representation Learning approach (VLC-Net) to understand fine-grained action sequences and create a unified feature representation along with their temporal dependencies for accurate AQA score prediction. Specifically, we design a bidirectional knowledge distillation operation to perform collaboration learning between vision-language pre-trained knowledge and visual action knowledge for fine-grained action feature learning. Furthermore, we design vision-language alignment guidance to explicitly align action features with the same action semantics across modalities, thereby unifying their joint representation. Leveraging these aligned features, we propose multimodal contrastive learning to relate modalities and align subactions with textual descriptions, ensuring accurate action representation. We conduct experiments on the FineDiving, MTL-AQA, FineFS, and Fis-V datasets, demonstrating the effectiveness of our approach, which outperforms state-of-the-art methods. Kumie Gedamu, Yanli Ji, Wangmeng Zuo, Jamal Bentahar, Yang Yang 0002, Jie Shao 0001, Heng Tao Shen |
IEEE Trans. Image Process. | 2 |
| 2025 | ReMP-AD: Retrieval-Enhanced Multi-Modal Prompt Fusion for Few-Shot Industrial Visual Anomaly Detection
Hongchi Ma, Guanglei Yang, Debin Zhao, Yanli Ji, Wangmeng Zuo |
ICCV | 4 |
| 2025 | Visual-Semantic Alignment Temporal Parsing for Action Quality AssessmentabstractAction Quality Assessment (AQA) is a challenging task involving analyzing fine-grained technical subactions, aligning high-level visual-semantic representations, and exploring internal temporal structures that capture the overall meaning of given action sequences. To address these challenges, we propose a Visual-semantic Alignment Temporal Parsing Network (VATP-Net) to understand the high-level visual semantics of subaction sequences and internal temporal structures without explicit supervision for action quality assessment. The proposed approach designs a self-supervised temporal parsing module to generate subaction sequences from the given video by aligning the visual and semantic action features. It captures high-level semantics and the internal temporal dynamics of subaction sequences. Furthermore, a multimodal interaction module is proposed to capture the interaction between different modalities of action features, enabling a comprehensive assessment of fine-grained and scene-invariant action details. The proposed module captures the intricate relationships and encourages interactions between different modalities within an action sequence, enhancing the overall understanding of action assessment. We exhaustively evaluate our proposed approach on the MTL-AQA, Rhythmic Gymnastics (RG), FineFS, and Fis-V datasets. Extensive experimental results demonstrate the effectiveness and feasibility of our proposed approach, which outperforms state-of-the-art methods by a significant margin. Kumie Gedamu, Yanli Ji, Yang Yang 0002, Jie Shao 0001, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Independency Adversarial Learning for Cross-Modal Sound SeparationabstractThe sound mixture separation is still challenging due to heavy sound overlapping and disturbance from noise. Unsupervised separation would significantly increase the difficulty. As sound overlapping always hinders accurate sound separation, we propose an Independency Adversarial Learning based Cross-Modal Sound Separation (IAL-CMS) approach, where IAL employs adversarial learning to minimize the correlation of separated sound elements, exploring high sound independence; CMS performs cross-modal sound separation, incorporating audio-visual consistent feature learning and interactive cross-attention learning to emphasize the semantic consistency among cross-modal features. Both audio-visual consistency and audio consistency are kept to guarantee accurate separation. The consistency and sound independence ensure the decomposition of overlapping mixtures into unrelated and distinguishable sound elements. The proposed approach is evaluated on MUSIC, VGGSound, and AudioSet. Extensive experiments certify that our approach outperforms existing approaches in supervised and unsupervised scenarios. Zhenkai Lin, Yanli Ji, Yang Yang 0002 |
AAAI | 2 |
| 2024 | Learning with noisy labels using collaborative sample selection and contrastive semi-supervised learning
Xiaohe Wu, Chao Xu 0003, Yanli Ji, Wangmeng Zuo, Yiwen Guo, Zhaopeng Meng |
Knowl. Based Syst. | 4 |
| 2024 | Corrigendum to "Learning with Noisy Labels Using Collaborative Sample Selection and Contrastive Semi-Supervised Learning" [Knowledge-Based Systems 296 (2024) 111860]
Xiaohe Wu, Chao Xu 0003, Yanli Ji, Wangmeng Zuo, Yiwen Guo, Zhaopeng Meng |
Knowl. Based Syst. | 4 |
| 2024 | Self-Supervised Sub-Action Parsing Network for Semi-Supervised Action Quality AssessmentabstractSemi-supervised Action Quality Assessment (AQA) using limited labeled and massive unlabeled samples to achieve high-quality assessment is an attractive but challenging task. The main challenge relies on how to exploit solid and consistent representations of action sequences for building a bridge between labeled and unlabeled samples in the semi-supervised AQA. To address the issue, we propose a Self-supervised sub-Action Parsing Network (SAP-Net) that employs a teacher-student network structure to learn consistent semantic representations between labeled and unlabeled samples for semi-supervised AQA. We perform actor-centric region detection and generate high-quality pseudo-labels in the teacher branch and assists the student branch in learning discriminative action features. We further design a self-supervised sub-action parsing solution to locate and parse fine-grained sub-action sequences. Then, we present the group contrastive learning with pseudo-labels to capture consistent motion-oriented action features in the two branches. We evaluate our proposed SAP-Net on four public datasets: the MTL-AQA, FineDiving, Rhythmic Gymnastics, and FineFS datasets. The experiment results show that our approach outperforms state-of-the-art semi-supervised methods by a significant margin. Kumie Gedamu, Yanli Ji, Yang Yang 0002, Jie Shao 0001, Heng Tao Shen |
IEEE Trans. Image Process. | 2 |
| 2024 | SV-Learner: Support-Vector Contrastive Learning for Robust Learning With Noisy LabelsabstractNoisy-label data inevitably gives rise to confusion in various perception applications. In this work, we revisit the theory of support vector machines (SVM) which mines support vectors to build the maximum-margin hyperplane for robust classification, and propose a robust-to-noise deep learning framework, SV-Learner, including the Support Vector Contrastive Learning (SVCL) and Support Vector-based Noise Screening (SVNS). The SV-Learner mines support vectors to solve the learning problem with noisy labels (LNL) reliably. Support Vector Contrastive Learning (SVCL) adopts support vectors as positive and negative samples, driving robust contrastive learning to enlarge the feature distribution margin for learning convergent feature distributions. Support Vector-based Noise Screening (SVNS) uses support vectors with valid labels to assist in screening noisy ones from confusable samples for reliable clean-noisy sample screening. Finally, Semi-Supervised classification is performed to realize the recognition of noisy samples. Extensive experiments are evaluated on CIFAR-10, CIFAR-100, Clothing1M, and Webvision datasets, and results demonstrate the effectiveness of our proposed approach. The source code is availablehttps://github.com/yanliji/SV-Learner. Yanli Ji, Wei-Shi Zheng 0001, Wangmeng Zuo, Xiaofeng Zhu 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Dominant SIngle-Modal SUpplementary Fusion (SIMSUF) for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis remains a big challenge due to the lack of effective fusion solutions. An effective fusion is expected to obtain the correct semantic representation for all modalities, and simultaneously thoroughly explore the contribution of each modality. In this paper, we propose a dominant SIngle-Modal SUpplementary Fusion (SIMSUF) approach to perform effective multimodal fusion for sentiment analysis. The SIMSUF is composed of three major components, a dominant modality supplementary module, a modality enhancement module, and a multimodal fusion module. The dominant modality supplementary module realizes dominant modality determination by estimating mutual dependence between every two modalities, and then the dominant modality is adopted to supplement other modalities for representative feature learning. To further explore the modality contribution, we propose a two-branch modality enhancement module, where one branch learns common representation distribution for multiple modalities, and simultaneously a specific modality enhancement branch is presented to perform semantic difference enhancement and distribution difference enhancement for each modality. Finally, a dominant modality leading fusion module is designed to fuse multimodal representations of two branches for sentiment analysis. Extensive experiments are evaluated on the CMU-MOSEI and CMU-MOSI datasets. Experiment results certify that our approach is superior to the state-of-the-art approaches. The source code of this work is available athttps://github.com/HumanCenteredUndestanding/SIMSUF. Yanli Ji, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 2 |
| 2023 | Unsupervised Sounding Pixel LearningabstractSounding source localization is a challenging cross-modal task due to the difficulty of crossmodal alignment.Although supervised crossmodal methods achieve encouraging performance, heavy manual annotations are expensive and inefficient.Thus it is valuable and meaningful to develop unsupervised solutions.In this paper, we propose an Unsupervised Sounding Pixel Learning (USPL) approach which enables a pixel-level sounding source localization in unsupervised paradigm.We first design a mask augmentation based multiinstance contrastive learning to realize unsupervised cross-modal coarse localization, which aligns audio-visual features to obtain coarse sounding maps.Secondly, we present an Unsupervised Sounding Map Refinement (SMR) module which employs the visual semantic affinity learning to explore inter-pixel relations of adjacent coordinate features.It contributes to recovering the boundary of coarse sounding maps and obtaining fine sounding maps.Finally, a Sounding Pixel Segmentation (SPS) module is presented to realize audio-supervised semantic segmentation.Extensive experiments are performed on the AVSBench-S4 and VG-GSound datasets, exhibiting encouraging results compared with previous SOTA methods. Yanli Ji, Yang Yang 0002 |
EMNLP | 2 |
| 2023 | Cross-modality Representation Interactive Learning for Multimodal Sentiment AnalysisabstractEffective alignment and fusion of multimodal features remain a significant challenge for multimodal sentiment analysis. In various multimodal applications, the text modal exhibits a significant advantage of compact yet expressive representation ability. In this paper, we propose a Cross-modality Representation Interactive Learning (CRIL) approach, which adopts the text modality to guide other modalities for learning representative feature tokens, contributing to effective multimodal fusion in multimodal sentiment analysis. We propose a semantic representation interactive learning module to learn concise semantic representation tokens for audio and video modalities under the guidance of the text modality, ensuring semantic alignment of representations among multiple modalities. Furthermore, we design a semantic relationship interactive learning module, which calculates a self-attention matrix for each modality and controls their consistency to enable the semantic relationship alignment for multiple modalities. Finally, we present a two-stage interactive fusion solution to bridge the modality gap for multimodal fusion and sentiment analysis. Extensive experiments are performed on the CMU-MOSEI, CMU-MOSI, and UR-FUNNY datasets, and experiment results demonstrate the effectiveness of our proposed approach. Yanli Ji, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 2 |
| 2023 | Localization-assisted Uncertainty Score Disentanglement Network for Action Quality AssessmentabstractAction Quality Assessment (AQA) has wide applications in various scenarios. Regarding the AQA of long-term figure skating, the big challenge lies in semantic context feature learning for Program Component Score (PCS) prediction and fine-grained technical subaction analysis for Technical Element Score (TES) prediction. In this paper, we propose a Localization-assisted Uncertainty Score Disentanglement Network (LUSD-Net) to deal with PCS and TES two predictions. In the LUSD-Net, we design an uncertainty score disentanglement solution, including score disentanglement and uncertainty regression, to decouple PCS-oriented and TES-oriented representations from skating sequences, ensuring learning differential representations for two types of score prediction. For long-term feature learning, a temporal interaction encoder is presented to build temporal context relation learning on PCS-oriented and TES-oriented features. To address subactions in TES prediction, a weakly-supervised temporal subaction localization is adopted to locate technical subactions in long sequences. For evaluation, we collect a large-scale Fine-grained Figure Skating dataset (FineFS) involving RGB videos and estimated skeleton sequences, providing rich annotations for multiple downstream action analysis tasks. The extensive experiments illustrate that our proposed LUSD-Net significantly improves the AQA performance, and the FineFS dataset provides a quantity data source for the AQA. The source code of LUSD-Net and the FineFS dataset is released at https://github.com/yanliji/FineFS-dataset. Yanli Ji, Lingfeng Ye, Huili Huang, Lijing Mao, Lingling Gao |
ACM Multimedia | 1 |
| 2023 | Layer-fusion for online mutual knowledge distillation
Gan Hu, Yanli Ji, Xingzhu Liang, Yuexing Han |
Multim. Syst. | 2 |
| 2023 | Relation-mining self-attention network for skeleton-based human action recognition
Kumie Gedamu, Yanli Ji, Lingling Gao, Yang Yang 0002, Heng Tao Shen |
Pattern Recognit. | 2 |
| 2023 | Fine-Grained Spatio-Temporal Parsing Network for Action Quality Assessment
Kumie Gedamu, Yanli Ji, Yang Yang 0002, Jie Shao 0001, Heng Tao Shen |
IEEE Trans. Image Process. | 2 |
| 2023 | Region Attention Enhanced Unsupervised Cross-Domain Facial Emotion RecognitionabstractThe visual emotion recognition from facial expressions easily suffers barrier problems of varying brightness, head pose change, various image scales when the recognition is performed in different domains. Therefore, it is required to erase such domain barriers. Considering that the human expresses their emotions always relying on the muscle motion near five sense organs of face, local features around them are typically crucial. In this paper, we propose a Region Attention eNhanced Domain Adaptation (RANDA) approach for unsupervised cross-domain facial expression recognition (FER). In RANDA, we design an unsupervised domain adaptation solution that adopts an iterative pseudo label assignment method to provide pseudo labels in the target domain, then employs adversarial learning to confuse feature representation of facial expressions in the source and target domains. Furthermore, a facial landmark guided fine-grained region attention learning module is designed to enhance significant emotion features and simultaneously weaken domain discrepancy. The proposed RANDA is adopted for cross-domain emotion recognition, and extensive evaluations are performed on multiple datasets, i.e., CK+, MMI, SFEW, RAF-DB, AffectNet. Results indicate that the RANDA outperforms the state-of-the-art approaches. It provides an effective solution for the cross-domain FER. Yanli Ji, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Self-Supervised Fine-Grained Cycle-Separation Network (FSCN) for Visual-Audio SeparationabstractAudio mixture separation is still challenging due to heavy overlaps and interactions. To correctly separate audio mixtures, we propose a novel self-supervised Fine-grained Cycle-Separation Network (FCSN) for vision-guided audio mixture separation. In the proposed approach, we design a two-stage procedure to perform self-supervised separation on audio mixtures. Using visual information as guidance, a primary-stage separation is realized via a U-net network, then the residual spectrogram is calculated by removing separated spectrograms from the original audio mixture. At the second-stage separation, a cycle-separation module is proposed to refine separation using separated results and the residual spectrogram. Self-supervision learning between vision and audio modalities is presented to push the cycle separation until the residual spectrogram becomes empty. Extensive experiments are evaluated on three large-scale datasets, MUSIC (MUSIC-21), AudioSet, and VGGSound. Experiment results certify that our approach outperforms the state-of-the-art approaches, and demonstrate the effectiveness for separating audio mixtures with overlap and interaction. Yanli Ji, Xing Xu 0001, Xuelong Li 0001, Heng Tao Shen |
IEEE Trans. Multim. | 1 |
| 2022 | Selective Hypergraph Convolutional Networks for Skeleton-based Action RecognitionabstractIn skeleton-based action recognition, Graph Convolutional Networks (GCNs) have achieved remarkable performance since the skeleton representation of human action can be naturally modeled by the graph structure. Most of the existing GCN-based methods extract skeleton features by exploiting single-scale joint information, while neglecting the valuable multi-scale contextual information. Besides, the commonly used strided convolution in temporal dimension could evenly filters out the keyframes we expect to preserve and leads to the loss of keyframe information. To address these issues, we propose a novel Selective Hypergraph Convolution Network, dubbed Selective-HCN, which stacks two key modules: Selective-scale Hypergraph Convolution (SHC) and Selective-frame Temporal Convolution (STC). The SHC module represents the human skeleton as the graph and hypergraph to fully extract multi-scale information, and selectively fuse features at various scales. Instead of traditional strided temporal convolution, the STC module can adaptively select keyframes and filter redundant frames according to the importance of the frames. Extensive experiments on two challenging skeleton action benchmarks, i.e., NTU-RGB+D and Skeleton-Kinetics, demonstrate the superiority and effectiveness of our proposed method. Yiran Zhu, Guangji Huang, Xing Xu 0001, Yanli Ji, Fumin Shen |
ICMR | 4 |
| 2022 | Global-Local Cross-View Fisher Discrimination for View-Invariant Action RecognitionabstractView change brings a significant challenge to action representation and recognition due to pose occlusion and deformation. We propose a Global-Local Cross-View Fisher Discrimination (GL-CVFD) algorithm to tackle this problem. In the GL-CVFD approach, we firstly capture the motion trajectory of body joints in action sequences as feature input to weaken the effect of view change. Secondly, we design a Global-Local Cross-View Representation (CVR) learning module, which builds global-level and local-level graphs to link body parts and joints between different views. It can enhance the cross-view information interaction and obtain an effective view-common action representation. Thirdly, we present a Cross-View Fisher Discrimination (CVFD) module, which performs a view-differential operation to separate view-specific action features and modifies the Fisher discriminator to implement view-semantic Fisher contrastive learning. It operates by pulling and pushing on view-specific and view-common action features in the view term to guarantee the validity of the CVR module, then distinguishes view-common action features in the semantic term for view-invariant recognition. Extensive and fair evaluations are implemented in the UESTC, NTU 60, and NTU 120 datasets. Experiment results illustrate that our proposed approach achieves encouraging performance in skeleton-based view-invariant action recognition. Lingling Gao, Yanli Ji, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 2 |
| 2022 | Answer Again: Improving VQA With Cascaded-Answering ModelabstractVisual Question Answering (VQA) is a very challenging task, which requires to understand visual images and natural language questions simultaneously. In the open-ended VQA task, most previous solutions focus on understanding the question and image contents, as well as their correlations. However, they mostly reason the answers in a one-stage way, which results in that the generated answers are significantly ignored. In this paper, we propose a novel approach, termed Cascaded-Answering Model (CAM), which extends the conventional one-stage VQA model to a two-stage model. Hence, the proposed model can fully explore the semantics embedded in the predicted answers. Specifically, CAM is composed of two cascaded answering modules: Candidate Answer Generation (CAG) module and Final Answer Prediction (FAP) module. In CAG module, we select multiple relevant candidates from the generated answers using a typical VQA approach with Co-Attention. While in FAP module, we integrate the information of question and image, together with the semantics explored from the selected candidate answers to predict the final answer. Experimental results demonstrate that the proposed model produces high-quality candidate answers and achieves the state-of-the-art performance on five large benchmark datasets, VQA-1.0, VQA-2.0, VQA-CP v2, TDIUC and COCO-QA. Yang Yang 0002, Xiaopeng Zhang 0008, Yanli Ji, Huimin Lu 0001, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | View-Invariant Human Action Recognition Via View Transformation Network (VTN)abstractSince the human body is non-rigid, actions captured in different views always involve action occlusion and information loss. Recently, view-variation-related human action recognition is still a challenging problem. To address the problem, we propose a View Transformation Network (VTN) that realizes the view normalization by transforming arbitrary-view action samples to a base view to seek for a view-invariant representation. an attention learning module is designed to learn a co-attention for action samples of different views, that contributes to output a similar feature representation to erase the view diversity in different views. Extensive and fair evaluations are performed on the UESTC varying-view RGB-D dataset, the NTU RGB-D 60 dataset, and the NTU RGB-D 120 dataset, where three evaluation types,i.e.X-subject, X-view, and A-view recognition, are performed. Experiments illustrate that our VTN model achieves outstanding performance. Lingling Gao, Yanli Ji, Kumie Gedamu, Xiaofeng Zhu 0001, Xing Xu 0001, Heng Tao Shen |
IEEE Trans. Multim. | 2 |
| 2022 | Cross-Modal Dynamic Networks for Video Moment Retrieval With Text QueryabstractVideo moment retrieval with text query aims to retrieve the most relevant segment from the whole video based on the given text query. It is a challenging cross-modal alignment task due to the huge gap between visual and linguistic modalities and the noise generated by manual labeling of time segments. Most of the existing works only use language information in the cross-modal fusion stage, neglecting that language information plays an important role in the retrieval stage. Besides, these works roughly compress the visual information in the video clips to reduce the computation cost which loses subtle video information in the long video. In this paper, we propose a novel model termed Cross-modal Dynamic Networks (CDN) which dynamically generates convolution kernel by visual and language features. In the feature extraction stage, we also propose a frame selection module to capture the subtle video information in the video segment. By this approach, the CDN can reduce the impact of the visual noise without significantly increasing the computation cost and leads to a better video moment retrieval result. The experiments on two challenge datasets,i.e., Charades-STA and TACoS, show that our proposed CDN method outperforms a bundle of state-of-the-art methods with more accurately retrieved moment video clips. The implementation code and extensive instruction of our proposed CDN method are provided athttps://github.com/CFM-MSG/Code_CDN. Gongmian Wang, Xing Xu 0001, Fumin Shen, Huimin Lu 0001, Yanli Ji, Heng Tao Shen |
IEEE Trans. Multim. | 5 |
| 2021 | Partial Feature Selection and Alignment for Multi-Source Domain AdaptationabstractMulti-Source Domain Adaptation (MSDA), which dedicates to transfer the knowledge learned from multiple source domains to an unlabeled target domain, has drawn increasing attention in the research community. By assuming that the source and target domains share consistent key feature representations and identical label space, existing studies on MSDA typically utilize the entire union set of features from both the source and target domains to obtain the feature map and align the map for each category and domain. However, the default setting of MSDA may neglect the issue of "partialness", i.e., 1) a part of the features contained in the union set of multiple source domains may not present in the target domain; 2) the label space of the target domain may not completely overlap with the multiple source domains. In this paper, we unify the above two cases to a more generalized MSDA task as Multi-Source Partial Domain Adaptation (MSPDA). We propose a novel model termed Partial Feature Selection and Alignment (PFSA) to jointly cope with both MSDA and MSPDA tasks. Specifically, we firstly employ a feature selection vector based on the correlation among the features of multiple sources and target domains. We then design three effective feature alignment losses to jointly align the selected features by preserving the domain information of the data sample clusters in the same category and the discrimination between different classes. Extensive experiments on various benchmark datasets for both MSDA and MSPDA tasks demonstrate that our proposed PFSA approach remarkably outperforms the state-of-the-art MSDA and unimodal PDA methods. Yangye Fu, Xing Xu 0001, Zuo Cao, Yanli Ji, Kai Zuo, Huimin Lu 0001 |
CVPR | 6 |
| 2021 | Multi-Stage Aggregated Transformer Network for Temporal Language Localization in VideosabstractWe address the problem of localizing a specific moment from an untrimmed video by a language sentence query. Generally, previous methods mainly exist two problems that are not fully solved: 1) How to effectively model the fine-grained visual-language alignment between video and language query? 2) How to accurately localize the moment in the original video length? In this paper, we streamline the temporal language localization as a novel multi-stage aggregated transformer network. Specifically, we first intro-duce a new visual-language transformer backbone, which enables iterations and alignments among all elements in visual and language sequences. Different from previous multi-modal transformers, our backbone keeps both structure unified and modality specific. Moreover, we also pro-pose a multi-stage aggregation module topped on the trans-former backbone. In this module, we compute three stage-specific representations corresponding to different moment stages respectively, i.e. starting, middle and ending stages, for each video element. Then for a moment candidate, we concatenate the starting/middle/ending representations of its starting/middle/ending elements respectively to form the final moment representation. Because the obtained moment representation captures the stage specific information, it is very discriminative for accurate localization. Extensive experiments on ActivityNet Captions and TACoS datasets demonstrate our proposed method achieves significant improvements compared with all other methods. Yang Yang 0002, Xinghan Chen, Yanli Ji, Xing Xu 0001, Jingjing Li 0001, Heng Tao Shen |
CVPR | 4 |
| 2021 | Graph Convolutional Hourglass Networks for Skeleton-Based Action RecognitionabstractGraph convolution networks (GCNs) have become the mainstream framework for the skeleton-based action recognition task, since the skeleton representation of human action can be naturally modeled by the graph structure. Generally, most of the existing GCN based models extract and aggregate skeleton features by exploiting single-scale joint information, while neglecting the valuable multi-scale information such as part and body features in the skeleton. To address this issue, we propose a novel Graph Convolutional Hourglass Network (GCHN) model, which is scalable by stacking several basic modules of Graph Convolutional Hourglass Block (GCHB). Each GCHB module consists of the sequential operations of graph convolution, graph pooling and graph unpooling, which can promote the interaction of multi-scale information in the skeleton and effectively improve the recognition performance. Extensive experiments on the challenging NTU-RGB+D and Kinetics-Skeleton datasets demonstrate that the proposed GCHN model achieves state-of-the-art performance. Yiran Zhu, Xing Xu 0001, Yanli Ji, Fumin Shen, Heng Tao Shen, Huimin Lu 0001 |
ICME | 3 |
| 2021 | PoseGTAC: Graph Transformer Encoder-Decoder with Atrous Convolution for 3D Human Pose EstimationabstractGraph neural networks (GNNs) have been widely used in the 3D human pose estimation task, since the pose representation of a human body can be naturally modeled by the graph structure. Generally, most of the existing GNN-based models utilize the restricted receptive fields of filters and single-scale information, while neglecting the valuable multi-scale contextual information. To tackle this issue, we propose a novel Graph Transformer Encoder-Decoder with Atrous Convolution, named PoseGTAC, to effectively extract multi-scale context and long-range information. In our proposed PoseGTAC model, Graph Atrous Convolution (GAC) and Graph Transformer Layer (GTL), respectively for the extraction of local multi-scale and global long-range information, are combined and stacked in an encoder-decoder structure, where graph pooling and unpooling are adopted for the interaction of multi-scale information from local to global (e.g., part-scale and body-scale). Extensive experiments on the Human3.6M and MPI-INF-3DHP datasets demonstrate that the proposed PoseGTAC model exceeds all previous methods and achieves state-of-the-art performance. Yiran Zhu, Xing Xu 0001, Fumin Shen, Yanli Ji, Lianli Gao, Heng Tao Shen |
IJCAI | 4 |
| 2021 | Vision-guided Music Source Separation via a Fine-grained Cycle-Separation NetworkabstractMusic source separation from a sound mixture remains a big challenge because there often exist heavy overlaps and interactions among similar music signals. In order to correctly separate mixed sources, we propose a novel Fine-grained Cycle-Separation Network (FCSN) for vision-guided music source separation. With the guidance of visual features, the proposed FCSN approach preliminarily separated music sources by minimizing the residual spectrogram which is calculated by removing preliminarily separated music spectrograms from the original music mixture. The separation is repeated several times until the residual spectrogram becomes empty or leaves only noise. Extensive experiments are performed on three large-scale datasets, the MUSIC (MUSIC-21), the AudioSet, and the VGGSound. Our approach outperforms state-of-the-art approaches in all datasets, and both separation accuracies and visualization results demonstrate its effectiveness for solving the problem of overlap and interaction in music source separation. Yanli Ji, Xing Xu 0001, Xiaofeng Zhu 0001 |
ACM Multimedia | 2 |
| 2021 | Arbitrary-view human action recognition via novel-view action generation
Kumie Gedamu, Yanli Ji, Yang Yang 0002, Lingling Gao, Heng Tao Shen |
Pattern Recognit. | 2 |
| 2021 | View-invariant action recognition via Unsupervised AttentioN Transfer (UANT)
Yanli Ji, Yang Yang 0002, Heng Tao Shen, Tatsuya Harada |
Pattern Recognit. | 1 |
| 2021 | Arbitrary-View Human Action Recognition: A Varying-View RGB-D Action DatasetabstractCurrent researches of action recognition which focus on single-view and multi-view recognition can hardly satisfy the requirements of human-robot interaction (HRI) applications for recognizing human actions from arbitrary views. Arbitrary-view recognition is still a challenging issue due to view changes and visual occlusions. In addition, the lack of datasets also sets up barriers. To provide data for arbitrary-view action recognition, we collect a new large-scale RGB-D action dataset for arbitrary-view action analysis, including RGB videos, depth and skeleton sequences. The dataset includes action samples captured in 8 fixed viewpoints and varying-view sequences which cover the entire 360° view angles. In total, 118 persons are invited to act 40 action categories. Our dataset involves more participants, more viewpoints and a large number of samples. More importantly, it is the first dataset containing the entire 360° varying-view sequences. The dataset provides sufficient data for multi-view, cross-view and arbitrary-view action analysis. Besides, we propose a View-guided Skeleton CNN (VS-CNN) to tackle the problem of arbitrary-view action recognition. Experiment results show that the VS-CNN achieves superior performance, and our dataset provides valuable but challenging data for the evaluation of arbitrary-view recognition. Yanli Ji, Yang Yang 0002, Fumin Shen, Heng Tao Shen, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Learning to Optimize Non-Rigid TrackingabstractOne of the widespread solutions for non-rigid tracking has a nested-loop structure: with Gauss-Newton to minimize a tracking objective in the outer loop, and Preconditioned Conjugate Gradient (PCG) to solve a sparse linear system in the inner loop. In this paper, we employ learnable optimizations to improve tracking robustness and speed up solver convergence. First, we upgrade the tracking objective by integrating an alignment data term on deep features which are learned end-to-end through CNN. The new tracking objective can capture the global deformation which helps Gauss-Newton to jump over local minimum, leading to robust tracking on large non-rigid motions. Second, we bridge the gap between the preconditioning technique and learning method by introducing a ConditionNet which is trained to generate a preconditioner such that PCG can converge within a small number of steps. Experimental results indicate that the proposed learning method converges faster than the original PCG by a large margin. Yang Li 0143, Aljaz Bozic, Tianwei Zhang 0002, Yanli Ji, Tatsuya Harada, Matthias Nießner |
CVPR | 4 |
| 2020 | Universal Weighting Metric Learning for Cross-Modal MatchingabstractCross-modal matching has been a highlighted research topic in both vision and language areas. Learning appropriate mining strategy to sample and weight informative pairs is crucial for the cross-modal matching performance. However, most existing metric learning methods are developed for unimodal matching, which is unsuitable for cross-modal matching on multimodal data with heterogeneous features. To address this problem, we propose a simple and interpretable universal weighting framework for cross-modal matching, which provides a tool to analyze the interpretability of various loss functions. Furthermore, we introduce a new polynomial loss under the universal weighting framework, which defines a weight function for the positive and negative informative pairs respectively. Experimental results on two image-text matching benchmarks and two video-text matching benchmarks validate the efficacy of the proposed method. Jiwei Wei, Xing Xu 0001, Yang Yang 0002, Yanli Ji, Zheng Wang 0044, Heng Tao Shen |
CVPR | 4 |
| 2020 | Graph-based variational auto-encoder for generalized zero-shot learningabstractZero-shot learning has been a highlighted research topic in both vision and language areas. Recently, generative methods have emerged as a new trend of zero-shot learning, which synthesizes unseen categories samples via generative models. However, the lack of fine-grained information in the synthesized samples makes it difficult to improve classification accuracy. It is also time-consuming and inefficient to synthesize samples and using them to train classifiers. To address such issues, we propose a novel Graph-based Variational Auto-Encoder for zero-shot learning. Specifically, we adopt knowledge graph to model the explicit inter-class relationships, and design a full graph convolution auto-encoder framework to generate the classifier from the distribution of the class-level semantic features on individual nodes. The encoder learns the latent representations of individual nodes, and the decoder generates the classifiers from latent representations of individual nodes. In contrast to synthesize samples, our proposed method directly generates classifiers from the distribution of the class-level semantic features for both seen and unseen categories, which is more straightforward, accurate and computationally efficient. We conduct extensive experiments and evaluate our method on the widely used large-scale ImageNet-21K dataset. Experimental results validate the efficacy of the proposed approach. Jiwei Wei, Yang Yang 0002, Xing Xu 0001, Yanli Ji, Xiaofeng Zhu 0001, Heng Tao Shen |
MMAsia | 4 |
| 2020 | A Survey of Human Action Analysis in HRI ApplicationsabstractThe human action is an important information source for human social interaction, and it simultaneously plays a crucial role in human-robot interaction (HRI). For a natural and fluent interaction, robots are required to understand human actions and have the capacity to predict action intentions and to imitate human actions for an appropriate response. Currently, existing survey papers for the action recognition mainly summarize algorithms that perform action recognition in experimental scenarios, and survey papers of the HRI mainly introduced various interaction interfaces in the HRI. Different from these surveys, we focus on the human action analysis on robot platforms for the HRI application, including the body motion and gestures. We review the existing HRI related references involving the action recognition, prediction, and the robot imitation of the human action. Moreover, we give a summary of robot platforms and action datasets that are frequently used in the study of HRI. Finally, we give an analysis on the development trend and future research directions of action analysis for the HRI applications. Yanli Ji, Yang Yang 0002, Fumin Shen, Heng Tao Shen, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | A Context Knowledge Map Guided Coarse-to-Fine Action RecognitionabstractHuman actions involve a wide variety and a large number of categories, which leads to a big challenge in action recognition. However, according to similarities on human body poses, scenes, interactive objects, human actions can be grouped into some semantic groups, i.e sports, cooking, etc. Therefore, in this paper, we propose a novel approach which recognizes human actions from coarse to fine. Taking full advantage of contributions from high-level semantic contexts, a context knowledge map guided recognition method is designed to realize the coarse-to-fine procedure. In the approach, we define semantic contexts with interactive objects, scenes and body motions in action videos, and build a context knowledge map to automatically define coarse-grained groups. Then fine-grained classifiers are proposed to realize accurate action recognition. The coarse-to-fine procedure narrows action categories in target classifiers, so it is beneficial to improving recognition performance. We evaluate the proposed approach on the CCV, the HMDB-51, and the UCF101 database. Experiments verify its significant effectiveness, on average, improving more than 5% of recognition precisions than current approaches. Compared with the state-of-the-art, it also obtains outstanding performance. The proposed approach achieves higher accuracies of 93.1%, 95.4% and 74.5% in the CCV, the UCF-101 and the HMDB51 database, respectively. Yanli Ji, Yue Zhan, Yang Yang 0002, Xing Xu 0001, Fumin Shen, Heng Tao Shen |
IEEE Trans. Image Process. | 1 |
| 2019 | Attention Transfer (ANT) Network for View-invariant Action RecognitionabstractWith wide applications in surveillance and human-robot interaction, view-invariant human action recognition is critical, however, challenging, due to the action occlusion and information loss caused by view change. Current methods mainly seek for a common feature space for different views. However, such solutions become invalid when there exist few common features, e.g. large view change. To tackle the problem, we propose an AttentioN Transfer (ANT) Network for view-invariant action recognition. Other than transferring features, ANT transfers attention from the reference view to arbitrary views, which correctly emphasize crucial body joints and their relations for view-invariant representation. In addition, the attention calculation method taking into account both recognition contribution and reliability of skeleton joints generates effective attention. Experiments showed its effectiveness for correctly locating crucial body joints in action sequences. We exhaustively evaluate our approach on the UESTC and the NTU dataset with three types of view-invariant evaluations, i.e. X-view, X-sub, and Arbitrary-view evaluation. Experiment results demonstrate its superiority in view-invariant representation and recognition. Yanli Ji, Feixiang Xu, Yang Yang 0002, Ning Xie 0003, Heng Tao Shen, Tatsuya Harada |
ACM Multimedia | 1 |
| 2019 | Learning one-to-many stylised Chinese character transformation and generation by generative adversarial networksabstractOwing to the complex structure of Chinese characters and the huge number of Chinese characters, it is very challenging and time consuming for artists to design a new font of Chinese characters. Therefore, the generation of Chinese characters and the transformation of font styles have become research hotspots. At present, most of the models on Chinese character transformation cannot generate multiple fonts, and they are not doing well in faking fonts. In this article, the authors propose a novel method of Chinese character fonts transformation and generation based on generative adversarial networks. The authors’ model is able to generate multiple fonts at once through font style‐specifying mechanism and it can generate a new font at the same time if the authors combine the characteristics of existing fonts. Jiefu Chen, Yanli Ji, Xing Xu 0001 |
IET Image Process. | 2 |
| 2019 | Cross-domain facial expression recognition via an intra-category common feature and inter-category Distinction feature fusion network
Yanli Ji, Yang Yang 0002, Fumin Shen, Heng Tao Shen |
Neurocomputing | 1 |
| 2019 | Word-to-region attention network for visual question answering
Yang Yang 0002, Yi Bin, Ning Xie 0003, Fumin Shen, Yanli Ji, Xing Xu 0001 |
Multim. Tools Appl. | 6 |
| 2019 | More is Better: Precise and Detailed Image Captioning Using Online Positive Recall and Missing Concepts MiningabstractRecently, a great progress in automatic image captioning has been achieved by using semantic concepts detected from the image. However, we argue that existing concepts-to-caption framework, in which the concept detector is trained using the image-caption pairs to minimize the vocabulary discrepancy, suffers from the deficiency of insufficient concepts. The reasons are two-fold: 1) the extreme imbalance between the number of occurrence positive and negative samples of the concept and 2) the incomplete labeling in training captions caused by the biased annotation and usage of synonyms. In this paper, we propose a method, termed online positive recall and missing concepts mining, to overcome those problems. Our method adaptively re-weights the loss of different samples according to their predictions for online positive recall and uses a two-stage optimization strategy for missing concepts mining. In this way, more semantic concepts can be detected and a high accuracy will be expected. On the caption generation stage, we explore an element-wise selection process to automatically choose the most suitable concepts at each time step. Thus, our method can generate more precise and detailed caption to describe the image. We conduct extensive experiments on the MSCOCO image captioning data set and the MSCOCO online test server, which shows that our method achieves superior image captioning performance compared with other competitive methods. Yang Yang 0002, Hanwang Zhang, Yanli Ji, Heng Tao Shen, Tat-Seng Chua |
IEEE Trans. Image Process. | 4 |
| 2019 | Deep adversarial metric learning for cross-modal retrieval
Xing Xu 0001, Li He 0001, Huimin Lu 0001, Lianli Gao, Yanli Ji |
World Wide Web | 5 |
| 2018 | A Large-scale RGB-D Database for Arbitrary-view Human Action RecognitionabstractCurrent researches mainly focus on single-view and multiview human action recognition, which can hardly satisfy the requirements of human-robot interaction (HRI) applications to recognize actions from arbitrary views. The lack of databases also sets up barriers. In this paper, we newly collect a large-scale RGB-D action database for arbitrary-view action analysis, including RGB videos, depth and skeleton sequences. The database includes action samples captured in 8 fixed viewpoints and varying-view sequences which covers the entire 360 view angles. In total, 118 persons are invited to act 40 action categories, and 25,600 video samples are collected. Our database involves more articipants, more viewpoints and a large number of samples. More importantly, it is the first database containing the entire 360? varying-view sequences. The database provides sufficient data for cross-view and arbitrary-view action analysis. Besides, we propose a View-guided Skeleton CNN (VS-CNN) to tackle the problem of arbitrary-view action recognition. Experiment results show that the VS-CNN achieves superior performance. Yanli Ji, Feixiang Xu, Yang Yang 0002, Fumin Shen, Heng Tao Shen, Wei-Shi Zheng 0001 |
ACM Multimedia | 1 |
| 2018 | Domain Invariant Subspace Learning for Cross-Modal Retrieval
Chenlu Liu, Xing Xu 0001, Yang Yang 0002, Huimin Lu 0001, Fumin Shen, Yanli Ji |
MMM (2) | 6 |
| 2018 | Hierarchical topology based hand pose estimation from a single depth image
Yanli Ji, Haoxin Li, Yang Yang 0002 |
Multim. Tools Appl. | 1 |
| 2018 | Semantic binary coding for visual recognition via joint concept-attribute modelling
Xing Xu 0001, Haiping Wu, Yang Yang 0002, Fumin Shen, Ning Xie 0003, Yanli Ji |
Multim. Tools Appl. | 6 |
| 2018 | One-shot learning based pattern transition map for action early recognition
Yanli Ji, Yang Yang 0002, Xing Xu 0001, Heng Tao Shen |
Signal Process. | 1 |
| 2018 | Recurrent attention network using spatial-temporal relations for action recognition
Yang Yang 0002, Yanli Ji, Ning Xie 0003, Fumin Shen |
Signal Process. | 3 |
| 2018 | Recognition and Detection of Two-Person Interactive Actions Using Automatically Selected Skeleton FeaturesabstractRecognition and detection of interactive actions performed by multiple persons have a wide range of real-world applications. Existing studies on the human activity analysis focus mainly on classifying video clips of simple actions performed by a single person, whereas the problem of understanding complex human activities with causal relationships between two people has not been sufficiently addressed yet. In this paper, we employ systematically organized skeleton features enhanced with directional features, and utilize sparse-group lasso to automatically choose discriminative factors that help in dealing with interactive action recognition and real-time detection tasks. Experiments on two person interaction datasets demonstrate the superiority of our approach to the state-of-the-art methods. Huimin Wu 0001, Jie Shao 0001, Xing Xu 0001, Yanli Ji, Fumin Shen, Heng Tao Shen |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2018 | Video Captioning by Adversarial LSTMabstractIn this paper, we propose a novel approach to video captioning based on adversarial learning and Long-Short Term Memory (LSTM). With this solution concept we aim at compensating for the deficiencies of LSTM-based video captioning methods that generally show potential to effectively handle temporal nature of video data when generating captions, but that also typically suffer from exponential error accumulation. Specifically, we adopt a standard Generative Adversarial Network (GAN) architecture, characterized by an interplay of two competing processes: a "generator", which generates textual sentences given the visual content of a video, and a "discriminator" which controls the accuracy of the generated sentences. The discriminator acts as an "adversary" towards the generator and with its controlling mechanism helps the generator to become more accurate. For the generator module, we take an existing video captioning concept using LSTM network. For the discriminator, we propose a novel realization specifically tuned for the video captioning problem and taking both the sentences and video features as input. This leads to our proposed LSTM-GAN system architecture, for which we show experimentally to significantly outperform the existing methods on standard public datasets. Yang Yang 0002, Jie Zhou 0001, Jiangbo Ai, Yi Bin, Alan Hanjalic, Heng Tao Shen, Yanli Ji |
IEEE Trans. Image Process. | 7 |
| 2017 | Gazing point dependent eye gaze estimation
Hong Cheng 0002, Yanli Ji, Lu Yang 0002, Yang Zhao 0024, Jie Yang 0001 |
Pattern Recognit. | 4 |
| 2015 | Learning contrastive feature distribution model for interaction recognition
Yanli Ji, Hong Cheng 0002, Haoxin Li |
J. Vis. Commun. Image Represent. | 1 |
| 2013 | A cooperative spectrum sensing scheme based on Detecting Reliability Statistics in cognitive radioabstractCooperative sensing can improve the performance of spectrum sensing in cognitive radio. This paper focuses on the design of hard information fusion strategy in cooperative sensing when detecting nodes work under varying degrees of reliability. A Cooperative Spectrum Sensing Scheme based on Detecting Reliability Statistics (CSS-DRS scheme) is proposed in this paper, in which fusion center estimates the reliability of each detecting nodes by means of probability statistics. The reliability estimation is based on the historical sensing results and does not need extra data transmission in report channel. Simulation results show that CSS-DRS scheme has significant advantage over traditional K-out-of-N strategy when detecting nodes work under varying degrees of reliability. Ben Wang 0004, Yanli Ji, Weidong Wang 0001, Yinghai Zhang |
PIMRC | 2 |
| 2010 | Human Action Recognition by SOM Considering the Probability of Spatio-temporal Features
Yanli Ji, Atsushi Shimada 0001, Rin-Ichiro Taniguchi |
ICONIP (2) | 1 |