Siqian Yang

dblp:93/3647 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-6100-3414ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Computer networks · 5 · 2 first-authorSystems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Exploring the Frontiers of Animation Video Generation in the Sora Era: Method, Dataset and Benchmark
abstract
Animation has gained significant interest in the recent film and TV industry. Despite the success of advanced video generation models like Sora, Kling, and CogVideoX in generating natural videos, they lack the same effectiveness in handling animation videos. Evaluating animation video generation is also a great challenge due to its unique artist styles, violating the laws of physics and exaggerated motions. In this paper, we present a comprehensive system, AniSora, designed for animation video generation, which includes a data processing pipeline, a controllable generation model, and an evaluation benchmark. Supported by the data processing pipeline with over 10M high-quality data, the generation model incorporates a spatiotemporal mask module to facilitate key animation production functions such as image-to-video generation, frame interpolation, and localized image-guided animation. We also collect an evaluation benchmark of 948 various animation videos, with specifically developed metrics for animation video generation. Our entire project is publicly available on https://github.com/bilibili/Index-anisora/tree/main
Yudong Jiang, Baohan Xu, Siqian Yang, Mingyu Ying, Yidi Wu 0002, Bingwen Zhu, Jinlong Hou, Huyang Sun
IJCAI3
2023 SpatialFormer: Semantic and Target Aware Attentions for Few-Shot Learning
abstract
Recent Few-Shot Learning (FSL) methods put emphasis on generating a discriminative embedding features to precisely measure the similarity between support and query sets. Current CNN-based cross-attention approaches generate discriminative representations via enhancing the mutually semantic similar regions of support and query pairs. However, it suffers from two problems: CNN structure produces inaccurate attention map based on local features, and mutually similar backgrounds cause distraction. To alleviate these problems, we design a novel SpatialFormer structure to generate more accurate attention regions based on global features. Different from the traditional Transformer modeling intrinsic instance-level similarity which causes accuracy degradation in FSL, our SpatialFormer explores the semantic-level similarity between pair inputs to boost the performance. Then we derive two specific attention modules, named SpatialFormer Semantic Attention (SFSA) and SpatialFormer Target Attention (SFTA), to enhance the target object regions while reduce the background distraction. Particularly, SFSA highlights the regions with same semantic information between pair features, and SFTA finds potential foreground object regions of novel feature that are similar to base categories. Extensive experiments show that our methods are effective and achieve new state-of-the-art results on few-shot classification benchmarks.
Jinxiang Lai, Siqian Yang, Guannan Jiang, Jun Liu 0116, Bin-Bin Gao, Wei Zhang 0217, Yuan Xie 0006, Chengjie Wang 0001
AAAI2
2023 Clustered-patch Element Connection for Few-shot Learning
abstract
Weak feature representation problem has influenced the performance of few-shot classification task for a long time. To alleviate this problem, recent researchers build connections between support and query instances through embedding patch features to generate discriminative representations. However, we observe that there exists semantic mismatches (foreground/ background) among these local patches, because the location and size of the target object are not fixed. What is worse, these mismatches result in unreliable similarity confidences, and complex dense connection exacerbates the problem. According to this, we propose a novel Clustered-patch Element Connection (CEC) layer to correct the mismatch problem. The CEC layer leverages Patch Cluster and Element Connection operations to collect and establish reliable connections with high similarity patch features, respectively. Moreover, we propose a CECNet, including CEC layer based attention module and distance metric. The former is utilized to generate a more discriminative representation benefiting from the global clustered-patch features, and the latter is introduced to reliably measure the similarity between pair-features. Extensive experiments demonstrate that our CECNet outperforms the state-of-the-art methods on classification benchmark. Furthermore, our CEC approach can be extended into few-shot segmentation and detection tasks, which achieves competitive performances.
Jinxiang Lai, Siqian Yang, Junhong Zhou, Xiaochen Chen, Jun Liu 0116, Bin-Bin Gao, Chengjie Wang 0001
IJCAI2
2023 PatchMix Augmentation to Identify Causal Features in Few-Shot Learning
abstract
The task of Few-shot learning (FSL) aims to transfer the knowledge learned from base categories with sufficient labelled data to novel categories with scarce known information. It is currently an important research question and has great practical values in the real-world applications. Despite extensive previous efforts are made on few-shot learning tasks, we emphasize that most existing methods did not take into account the distributional shift caused by sample selection bias in the FSL scenario. Such a selection bias can induce spurious correlation between the semantic causal features, that are causally and semantically related to the class label, and the other non-causal features. Critically, the former ones should be invariant across changes in distributions, highly related to the classes of interest, and thus well generalizable to novel classes, while the latter ones are not stable to changes in the distribution. To resolve this problem, we propose a novel data augmentation strategy dubbed as PatchMix that can break this spurious dependency by replacing the patch-level information and supervision of the query images with random gallery images from different classes from the query ones. We theoretically show that such an augmentation mechanism, different from existing ones, is able to identify the causal features. To further make these features to be discriminative enough for classification, we propose Correlation-guided Reconstruction (CGR) and Hardness-Aware module for instance discrimination and easier discrimination between similar classes. Moreover, such a framework can be adapted to the unsupervised FSL scenario. The utility of our method is demonstrated on the state-of-the-art results consistently achieved on several benchmarks including miniImageNet, tieredImageNet, CIFAR-FS, CUB, Cars, Places and Plantae, in all settings of single-domain, cross-domain and unsupervised FSL. By studying the intra-variance property of learned features and visualizing the learned features, we further quantitatively and qualitatively show that such a promising result is due to the effectiveness in learning causal features.
Chengming Xu 0001, Chen Liu 0030, Xinwei Sun 0001, Siqian Yang, Yabiao Wang, Chengjie Wang 0001, Yanwei Fu 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 tSF: Transformer-Based Semantic Filter for Few-Shot Learning
Jinxiang Lai, Siqian Yang, Yi Zeng 0006, Jun Liu 0116, Bin-Bin Gao, Chengjie Wang 0001
ECCV (20)2
2022 Split-PU: Hardness-aware Training Strategy for Positive-Unlabeled Learning
abstract
Positive-Unlabeled (PU) learning aims to learn a model with rare positive samples and abundant unlabeled samples. Compared with classical binary classification, the task of PU learning is much more challenging due to the existence of many incompletely-annotated data instances. Since only part of the most confident positive samples are available and evidence is not enough to categorize the rest samples, many of these unlabeled data may also be the positive samples. Research on this topic is particularly useful and essential to many real-world tasks which demand very expensive labelling cost. For example, the recognition tasks in disease diagnosis, recommendation system and satellite image recognition may only have few positive samples that can be annotated by the experts. While this problem is receiving increasing attention, most of the efforts have been dedicated to the design of trustworthy risk estimators such as uPU and nnPU and direct knowledge distillation, e.g., Self-PU. These methods mainly omit the intrinsic hardness of some unlabeled data, which can result in sub-optimal performance as a consequence of fitting the easy noisy data and not sufficiently utilizing the hard data. In this paper, we focus on improving the commonly-used nnPU with a novel training pipeline. We highlight the intrinsic difference of hardness of samples in the dataset and the proper learning strategies for easy and hard data. By considering this fact, we propose first splitting the unlabeled dataset with an early-stop strategy. The samples that have inconsistent predictions between the temporary and base model are considered as hard samples. Then the model utilizes a noise-tolerant Jensen-Shannon divergence loss for easy data; and a dual-source consistency regularization for hard data which includes a cross-consistency between student and base model for low-level features and self-consistency for high-level features and predictions, respectively. Our method achieves much better results compared with existing methods on CIFAR10 and two medical datasets of liver cancer survival time prediction, and low blood pressure diagnosis of pregnant, individually. The experimental results validates the efficacy of our proposed method.
Chengming Xu 0001, Chen Liu 0030, Siqian Yang, Yabiao Wang, Lijie Jia, Yanwei Fu 0001
ACM Multimedia3
2022 Rethinking the Metric in Few-shot Learning: From an Adaptive Multi-Distance Perspective
abstract
Few-shot learning problem focuses on recognizing unseen classes given a few labeled images. In recent effort, more attention is paid to fine-grained feature embedding, ignoring the relationship among different distance metrics. In this paper, for the first time, we investigate the contributions of different distance metrics, and propose an adaptive fusion scheme, bringing significant improvements in few-shot classification. We start from a naive baseline of confidence summation and demonstrate the necessity of exploiting the complementary property of different distance metrics. By finding the competition problem among them, built upon the baseline, we propose an Adaptive Metrics Module (AMM) to decouple metrics fusion into metric-prediction fusion and metric-losses fusion. The former encourages mutual complementary, while the latter alleviates metric competition via multi-task collaborative learning. Based on AMM, we design a few-shot classification framework AMTNet, including the AMM and the Global Adaptive Loss (GAL), to jointly optimize the few-shot task and auxiliary self-supervised task, making the embedding features more robust. In the experiment, the proposed AMM achieves 2% higher performance than the naive metrics fusion module, and our AMTNet outperforms the state-of-the-arts on multiple benchmark datasets.
Jinxiang Lai, Siqian Yang, Guannan Jiang, Yuxi Li 0009, Zihui Jia, Xiaochen Chen, Jun Liu 0116, Bin-Bin Gao, Wei Zhang 0217, Yuan Xie 0006, Chengjie Wang 0001
ACM Multimedia2
2021 Learning a Few-shot Embedding Model with Contrastive Learning
abstract
Few-shot learning (FSL) aims to recognize target classes by adapting the prior knowledge learned from source classes. Such knowledge usually resides in a deep embedding model for a general matching purpose of the support and query image pairs. The objective of this paper is to repurpose the contrastive learning for such matching to learn a few-shot embedding model. We make the following contributions: (i) We investigate the contrastive learning with Noise Contrastive Estimation (NCE) in a supervised manner for training a few-shot embedding model; (ii) We propose a novel contrastive training scheme dubbed infoPatch, exploiting the patch-wise relationship to substantially improve the popular infoNCE; (iii) We show that the embedding learned by the proposed infoPatch is more effective; (iv) Our model is thoroughly evaluated on few-shot recognition task; and demonstrates state-of-the-art results on miniImageNet and appealing performance on tieredImageNet, Fewshot-CIFAR100 (FC-100).
Chen Liu 0030, Yanwei Fu 0001, Chengming Xu 0001, Siqian Yang, Chengjie Wang 0001, Li Zhang 0040
AAAI4
2019 APP: Augmented Proactive Perception for Driving Hazards with Sparse GPS Trace
abstract
Driving safety is a persistent concern for urban dwellers who spend hours driving on road in ordinary daily life. Traditional driving hazard detection solutions heavily rely on onboard sensors (e.g., front and rear radars, cameras) with limited sensing range. In this article, we propose a proactive hazard warning system, called APP, which aims to alert drivers when there are vehicles with dangerous behaviors nearby. To this end, APP incorporates several basic techniques (e.g, tensor decomposition, similarity comparison) to estimate behavioral data of a driver based on sparse sampled GPS trace at first. Then, with the estimated unlabelled data, potential dangerous behaviors of a particular vehicle are identified and recognized with a Gaussian Mixture Model (GMM) based approach. We have implemented and evaluated our system with a dataset collected for 30 days from over 13,676 taxicabs. Our method shows on average 81% accuracy in potential dangerous behavior recognition.
Siqian Yang, Cheng Wang 0001, Hongzi Zhu, Changjun Jiang 0002
MobiHoc1
2019 iLogBook: Enabling Text-Searchable Event Query Using Sparse Vehicle-Mounted GPS Data
abstract
Querying an incident (i.e., an occurrence of seemingly minor importance) from coarse-grained driving log (i.e., GPS trace) has been a daunting task. For example, “Which restaurant did I drive by at exactly 4 pm yesterday?” The question seems very simple but is nontrivial, because the question is semantics-driven while the actual log data are GPS coordinate-based. Especially, the practical GPS log is very sparse and inaccurate for high-speed mobile objects such as vehicles. This paper seeks to answer any fuzzy query over sparse vehicles GPS data. Our system, called iLogBook, achieves these two goals by leveraging tensor technique and latent semantic analysis to high-precision trajectory recovery and similarity matching. We have implemented and evaluated the iLogBook with the GPS data of over 13, 798 taxicabs collected in eight days in Shenzhen, China. Our results show about 97% accuracy in trajectory inference. Moreover, the system handles about 90% daily queries among 10 marked drivers.
Siqian Yang, Cheng Wang 0001, Lei Yang 0025, Changjun Jiang 0002
IEEE Trans. Intell. Transp. Syst.1
2019 Passive neighbor discovery with social recognition for mobile ad hoc social networking applications
Hao Ling, Siqian Yang
Wirel. Networks2
2018 Centron: Cooperative neighbor discovery in mobile Ad-hoc networks
Siqian Yang, Cheng Wang 0001, Changjun Jiang 0002
Comput. Networks1
2015 LASS: Local-Activity and Social-Similarity Based Data Forwarding in Mobile Social Networks
abstract
This paper aims to design an efficient data forwarding scheme based on local activity and social similarity(LASS) for mobile social networks (MSNs). Various definitions of social similarity have been proposed as the criterion for relay selection, which results in various forwarding schemes. The appropriateness and practicality of various definitions determine the performances of these forwarding schemes. A popular definition has recently been proven to be more efficient than other existing ones, i.e., the more common interests between two nodes, the larger social similarity between them. In this work, we show that schemes based on such definition ignore the fact that members within the same community, i.e., with the same interest, usually have different levels of local activity, which will result in a low efficiency of data delivery. To address this, in this paper, we design a new data forwarding scheme for MSNs based on community detection in dynamic weighted networks, called Local-Activity and Social-Similarity, taking into account the difference of members' internal activity within each community, i.e., local activity. To the best of our knowledge, the proposed scheme is the first one that utilizes different levels of local activity within communities. Through extensive simulations, we demonstrate that LASS achieves better performance than state-of-the-art protocols.
Zhong Li 0006, Cheng Wang 0001, Siqian Yang, Changjun Jiang 0002, Xiang-Yang Li 0001
IEEE Trans. Parallel Distributed Syst.3
2015 Space-Crossing: Community-Based Data Forwarding in Mobile Social Networks Under the Hybrid Communication Architecture
abstract
In this paper, we study two tightly coupled issues, space-crossing community detection and its influence on data forwarding in mobile social networks (MSNs). We propose a communication framework containing the hybrid underlying network with access point (AP) support for data forwarding and the base stations for managing most of control traffic. The concept of physical proximity community can be extended to be one across the geographical space, because APs can facilitate the communication among long-distance nodes. Space-crossing communities are obtained by merging some pairs of physical proximity communities. Based on the space-crossing community, we define two cases of node local activity and use them as the input of inner product similarity measurement. We design a novel data forwarding algorithm Social Attraction and Infrastructure Support (SAIS), which applies similarity attraction to route to neighbor more similar to destination, and infrastructure support phase to route the message to other APs within common connected components. We evaluate our SAIS algorithm on real-life datasets from MIT Reality Mining and University of Illinois Movement (UIM). Results show that space-crossing community plays a positive role in data forwarding in MSNs. Based on this new type of community, SAIS achieves a better performance than existing popular social community-based data forwarding algorithms in practice, including Simbet, Bubble Rap and Nguyen's Routing algorithms.
Zhong Li 0006, Cheng Wang 0001, Siqian Yang, Changjun Jiang 0002, Ivan Stojmenovic
IEEE Trans. Wirel. Commun.3
2014 Improving data forwarding in Mobile Social Networks with infrastructure support: A space-crossing community approach
abstract
In this paper, we study two tightly coupled issues: space-crossing community detection and its influence on data forwarding in Mobile Social Networks (MSNs) by taking the hybrid underlying networks with infrastructure support into consideration. The hybrid underlying network is composed of large numbers of mobile users and a small portion of Access Points (APs). Because APs can facilitate the communication among long-distance nodes, the concept of physical proximity community can be extended to be one across the geographical space. In this work, we first investigate a space-crossing community detection method for MSNs. Based on the detection results, we design a novel data forwarding algorithm SAAS (Social Attraction and AP Spreading), and show how to exploit the space-crossing communities to improve the data forwarding efficiency. We evaluate our SAAS algorithm on real-life data from MIT Reality Mining and University of Illinois Movement (UIM). Results show that space-crossing community plays a positive role in data forwarding in MSNs in terms of delivery ratio and delay. Based on this new type of community, SAAS achieves a better performance than existing social community-based data forwarding algorithms in practice, including Bubble Rap and Nguyen's Routing algorithms.
Zhong Li 0006, Cheng Wang 0001, Siqian Yang, Changjun Jiang 0002, Ivan Stojmenovic
INFOCOM3