Wanqi Yang

dblp:117/3040 · DBLP profile ↗
← Back
52ranked-venue papers
13as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 6 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Decomposing and Composing: Towards Efficient Vision-Language Continual Learning via Rank-1 Expert Pool in a Single LoRA
abstract
Continual learning (CL) in vision-language models (VLMs) faces significant challenges in improving task adaptation and avoiding catastrophic forgetting. Existing methods usually have heavy inference burden or rely on external knowledge, while Low-Rank Adaptation (LoRA) has shown potential in reducing these issues by enabling parameter-efficient tuning. However, considering directly using LoRA to alleviate the catastrophic forgetting problem is non-trivial, we introduce a novel framework that restructures a single LoRA module as a decomposable Rank-1 Expert Pool. Our method learns to dynamically compose a sparse, task-specific update by selecting from this expert pool, guided by the semantics of the [CLS] token. In addition, we propose an Activation-Guided Orthogonal (AGO) loss that orthogonalizes critical parts of LoRA weights across tasks. This sparse composition and orthogonalization enable fewer parameter updates, resulting in domain-aware learning while minimizing inter-task interference and maintaining downstream task performance. Extensive experiments across multiple settings demonstrate state-of-the-art results in all metrics, surpassing zero-shot upper bounds in generalization. Notably, it reduces trainable parameters by 96.7% compared to the baseline method, eliminating reliance on external datasets or task-ID discriminators. The merged LoRAs retain less weights and incur no inference latency, making our method computationally lightweight.
Zhan Fa, Yue Duan, Jian Zhang 0090, Lei Qi 0001, Wanqi Yang, Yinghuan Shi
AAAI5
2026 ProfRCA: LLM-Enabled Fine-Grained Root Cause Analysis with Continuous Profiling Data
Siyuan Ye, Gou Tan, Wanqi Yang, Pengfei Chen 0002
SANER3
2026 Cluster-aware contrastive learning for partially view-aligned clustering
abstract
Multi-view clustering aims to discover shared semantic information from different perspectives, but real-world time and space factors often lead to missing pairing relationships, resulting in partial data alignment between views and affecting clustering performance. Clustering on such data is referred to as the Partial View-aligned clustering Problem (PVP). Most existing PVP methods mainly capture shared semantic features in a common subspace to infer missing pairing relationships of samples between views. However, excessive reliance on shared feature representations between views could affect the enforcement of the clustering structures. To address this, we propose a novel Cluster-aware cOntrastive Learning method for partially view-Aligned clustering (COLA), by performing intra-view cluster contrastive learning to compact the clustering structure within a view and aligning the soft-label and distributions between views. In this way, the consistent clustering structure between views can be captured and then used to help recover the missing pairing relationships. Extensive experiments demonstrate that COLA outperforms state-of-the-art methods, improving NMI by 14.61% on average and improving ACC by 6.26% on average across eight datasets.
Wanqi Yang, Like Xin, Ming Yang 0014
Neurocomputing2
2026 Global-and-local guidance with synthesized view for unpaired multi-view clustering
Like Xin, Wanqi Yang, Lei Wang 0001, Ming Yang 0014
Inf. Process. Manag.2
2026 A unified and efficient training framework for open-ended non-transitive games
Shaokang Dong, Shangdong Yang, Hongye Cao, Wanqi Yang, Yang Gao 0001
Neural Networks5
2025 TraceWizard: End-to-End Distributed Tracing Across Host and Network Devices in Cloud
abstract
The rise of microservice architecture in cloud computing has introduced additional complexities in diagnosing faults, as traditional end-to-end tracing systems often fail to address issues beyond the application layer, such as network devices. To overcome this limitation, we introduce TraceWizard, an enhanced end-to-end tracing system that integrates eBPF and SDN (Software Defined Network) technologies to track requests across host and network devices. By enabling full life-cycle tracing and maintaining consistent trace contexts, TraceWizard provides fine-grained insights into faults across applications, the OS kernel, and network devices. Our evaluations show that its data enables more effective fault detection, achieving an average accuracy of 91.6 % across different algorithms—significantly outperforming application-layer monitoring tools. Additionally, it helps operators identify root causes with minimal overhead, reducing QPS by 2.2%, increasing QCT by 2.2%, and adding 3.41 % CPU and 2.43% memory usage.
Kuangyuan Li, Jingrun Zhang, Pengfei Chen 0002, Hongyang Chen 0002, Ruipeng Hong, Wanqi Yang, Chen Sun 0005
CLOUD6
2025 Beyond Mandatory Federations: Balancing Egoism, Utilitarianism and Egalitarianism in Mixed-Motive Games
abstract
In the field of mixed-motive games, extensive multi-agent learning studies have explored the balance between egoism (individual interest), utilitarianism (collective interest), and egalitarianism (fairness). Traditional approaches often rely on manually designed reward functions, social norms, and alliance/federation mechanisms to transition agents from individualistic behaviors toward cooperative strategies. However, these methods typically require all agents to share private local information or to mandatorily participate in federations, which is impractical in real-world applications. To address these issues, this paper proposes a Flexible-Participation Federation (FPF) framework that allows agents to participate in the federation voluntarily. Furthermore, we extend the federation from a global to a Local Multi-Federation (LMF) framework, enabling agents to form multiple localized federations, thereby promoting more efficient and adaptive cooperation. Theoretical evidence demonstrates that the global FPF model, along with the discrepancy between decentralized egoistic policies and federated utilitarian policies, achieves an O(1/T) convergence rate. Agents in the LMF framework also reach consensus within a sublinear gap. Extensive experiments show that agents opting out of federation participation experience a reduction in egoism, and our approach outperforms multiple baselines in terms of both utilitarianism and egalitarianism.
Shaokang Dong, Shangdong Yang, Hongye Cao, Wanqi Yang, Yang Gao 0001
AAAI5
2025 Robust Seizure Prediction Based on Riemannian Manifold Enhanced Denoising Adversarial Autoencoder
abstract
The seizure early warning devices based on multichannel EEG signals is one of the most used assisted-living strategies for drug-resistant epileptic patients. One of the challenges in the development of these devices is that existing algorithms cannot avoid the effects of electrode loosening. To alleviate such problem, a seizure prediction model robust to corrupted EEG recordings is proposed in this paper. This method intends to learn a stable feature space adapted to different corruption versions via constructing a jointly-optimized autoencoder. Robust information is captured during the reconstruction process by developing a corruption module, while an adversarial training procedure circumvents overfitting using variational inference. Moreover, a regularization term is designed to minimize the distance among various corruption versions in the embedding space with the Riemannian manifold-based distribution alignment. Both raw EEG data and corrupted EEG data are utilized to evaluate prediction performance. Experimental results indicate that this method can predict seizures effectively and stably under the condition of electrode looseness.
Peizhen Peng, Yansheng Wu, Wanqi Yang, Wenxin Wei
ICASSP3
2025 Hybrid Feature Collaborative Reconstruction Network for Few-Shot Fine-Grained Image Classification
abstract
Our research focuses on few-shot fine-grained image classification (FS-FGIC), which faces two main challenges: the similarity of fine-grained objects and a limited number of samples. Traditional feature reconstruction networks enhance key features through spatial reconstruction and error minimization but often fail to capture interclass differences with limited samples. We propose a Hybrid Feature Collaborative Reconstruction Network (HFCR-Net) with two key components: the Hybrid Feature Fusion Process (HFFP) and the Hybrid Feature Reconstruction Process (HFRP). In HFRP, dynamic weight adjustment is employed to enhance spatial dependencies and channel correlations, increasing inter-class differences. In HFRP, we introduce channel dimension reconstruction to improve the processes of support-to-query and query-to-support reconstruction, further increasing interclass and reducing intra-class differences. Extensive experiments on three widely used fine-grained datasets confirm the effectiveness and superiority of our approach.
Shulei Qiu, Wanqi Yang, Ming Yang 0014
ICASSP2
2025 Adaptive Mobile Agent for Dynamic Interactions
abstract
With the rise of Multimodal Large Language Models (MLLM), LLM-driven visual agents are transforming software interfaces, especially those with graphical user interfaces. However, existing methods often struggle with diverse and complex mobile environments, such as rapidly changing app interfaces or non-standard UI components, limiting their adaptability and precision. This work presents a novel LLM-based multimodal agent framework for mobile devices, designed to enhance interaction and adaptive capabilities in dynamic mobile environments. By autonomously navigating devices and emulating human-like behaviors, the agent integrates parsing, text, and vision descriptions to construct a flexible action space. During the exploration phase, functionalities of user interface elements are documented into a customized structured knowledge base. In the deployment phase, RAG technology enables efficient retrieval and updates from this knowledge base. Experimental results across multiple benchmarks validate the framework's superior performance and practical effectiveness.
Yanda Li, Chi Zhang 0007, Wenjia Jiang, Wanqi Yang, Xin Chen 0040, Ling Chen 0006, Yunchao Wei
ICME4
2025 MicroGuard: Non-Intrusive Dynamic Analysis for Inter-Service Access Control of Microservices
abstract
Cloud-native systems enable high-scalibility for application development and deployment with loosely coupled microservices that interact over the network, but they also introduce security risks, potentially leading to unauthorized inter-service access.To mitigate such risks, existing approaches rely on manual policy configuration or static code analysis.However, these methods are time-consuming for policy maintainance, require code avalibility and fail to support access control for hidden service invocations.To address these limitations, we propose MicroGuard, an access control system that automatically generates and enforces access control policies in microservices.During microservice development and testing, MicroGuard captures and analyzes inter-service communication packets to identify hidden runtime service invocations.MicroGuard leverages a bidirectional prefix tree (Trie) and pretrained language models to extract comprehensive access control policies.During execution of microservice systems, MicroGuard enforces these policies to detect and block unauthorized access requests.Our experimental results show that MicroGuard effectively captures and rejects all unauthorized attempts, introducing an average processing delay of only 2% to the original response time of microservice systems.
Haoming Luo, Wanqi Yang, Pengfei Chen 0002
Internetware2
2025 Label-Semantics-Guided Multi-View Multi-Label Learning via High-Order Semantic Fusion
abstract
Incomplete multi-view multi-label learning faces significant challenges arising from semantic heterogeneity across modalities and incomplete modality availability. Traditional fusion approaches typically emphasize superficial feature alignment, neglecting high-order semantic interactions among modalities and labels, thus resulting in redundant or conflicting information integration. To address these limitations, we propose a novel Label Semantic Guided Adaptive Fusion framework. Specifically, we leverage pretrained language models to generate semantic embeddings for both multi-view data and associated labels, facilitating unified semantic understanding. Subsequently, we construct dual-domain hypergraphs separately within the modality and label semantic spaces to explicitly model complex high-order semantic correlations. Based on these hypergraphs, we employ hypergraph neural networks to mine intrinsic semantic relationships and dynamically assess semantic consistency between each modality and the label space. Finally, an adaptive weighting strategy guided by this semantic consistency measure is introduced to fuse modalities effectively, assigning high weights to modalities with greater semantic alignment. Extensive experiments demonstrate that our LSGMM improves fusion accuracy and robustness over state-of-the-art IMvML methods, confirming the effectiveness of integrating label semantics and high-order semantic relationships into adaptive multi-view fusion.
Kaixiang Wang 0001, Xiaojian Ding, Wanqi Yang, Ming Yang 0014
ACM Multimedia3
2025 CloudHeal: A Lightweight Online Learning Based Self-Healing Framework for Cloud-Native Systems
abstract
With the increasing scale and complexity of cloud-native systems, faults have become common and impact the end-user experience. Therefore, it is necessary to propose efficient self-healing mechanisms to ensure reliability and reduce service disruptions. However, traditional self-healing approaches, including rule-based and offline machine learning methods, suffer from inefficiency and limited adaptability. In this paper, we propose CloudHeal, a lightweight online learning-based self-healing framework that dynamically optimizes fault mitigation strategies for cloud-native environments. CloudHeal leverages Spectrum-Based Fault Localization (SBFL) and feedback-enhanced PageRank for root cause identification, significantly reducing unnecessary self-healing actions. To support adaptive decision-making, CloudHeal employs a context-aware online learning algorithm that continuously refines fault recovery strategies based on real-time system feedback. Experimental evaluations on microservice benchmarks demonstrate that CloudHeal achieves a 94.8% fault recovery rate, outperforming existing fault self-healing methods. Additionally, the Root Cause Analyzer in CloudHeal improves fault localization accuracy by$\mathbf{1 6. 9 \%}$compared to state-of-the-art approaches, further enhancing self-healing effectiveness. These results highlight CloudHeal's potential in advancing self-healing capabilities for large-scale cloud-native systems.
Junquan Yi, Wanqi Yang, Pengfei Chen 0002
SRDS3
2025 E2MPL: An Enduring and Efficient Meta Prompt Learning Framework for Few-Shot Unsupervised Domain Adaptation
abstract
Few-shot unsupervised domain adaptation (FS-UDA) leverages a limited amount of labeled data from a source domain to enable accurate classification in an unlabeled target domain. Despite recent advancements, current approaches of FS-UDA continue to confront a major challenge: models often demonstrate instability when adapted to new FS-UDA tasks and necessitate considerable time investment. To address these challenges, we put forward a novel framework called Enduring and Efficient Meta-Prompt Learning (E2MPL) for FS-UDA. Within this framework, we utilize the pre-trained CLIP model as the backbone of feature learning. Firstly, we design domain-shared prompts, consisting of virtual tokens, which primarily capture meta-knowledge from a wide range of meta-tasks to mitigate the domain gaps. Secondly, we develop a task prompt learning network that adaptively learns task-specific prompts with the goal of achieving fast and stable task generalization. Thirdly, we formulate the meta-prompt learning process as a bilevel optimization problem, consisting of (outer) meta-prompt learner and (inner) task-specific classifier and domain adapter. Also, the inner objective of each meta-task has the closed-form solution, which enables efficient prompt learning and adaptation to new tasks in a single step. Extensive experimental studies demonstrate the promising performance of our framework in a domain adaptation benchmark dataset DomainNet. Compared with state-of-the-art methods, our approach has improved the average accuracy by at least 15 percentage points and reduces the average time by 64.67% in the 5-way 1-shot task; in the 5-way 5-shot task, it achieves at least a 9-percentage-point improvement in average accuracy and reduces the average time by 63.18%. Moreover, our method exhibits more enduring and stable performance than the other methods, i.e., reducing the average IQR value by over 40.80% and 25.35% in the 5-way 1-shot and 5-shot task, respectively.
Wanqi Yang, Lei Wang 0001, Ming Yang 0014, Yang Gao 0001
IEEE Trans. Image Process.1
2025 Selective Contrastive Learning for Unpaired Multi-View Clustering
abstract
In this article, we investigate a novel but insufficiently studied issue, unpaired multi-view clustering (UMC), where no paired observed samples exist in multi-view data, and the goal is to leverage the unpaired observed samples in all views for effective joint clustering. Existing methods in incomplete multi-view clustering usually utilize the sample pairing relationship between views to connect the views for joint clustering, but unfortunately, it is invalid for the UMC case. Therefore, we strive to mine a consistent cluster structure between views and propose an effective method, namely selective contrastive learning for UMC (scl-UMC), which needs to solve the following two challenging issues: 1) uncertain clustering structure under no supervision information and 2) uncertain pairing relationship between the clusters of views. Specifically, for the first one, we design an inner-view (IV) selective contrastive learning module to enhance the clustering structures and alleviate the uncertainty, which selects confident samples near the cluster centroids to perform contrastive learning in each view. For the second one, we design a cross-view (CV) selective contrastive learning module to first iteratively match the clusters between views and then tighten the matched clusters. Also, we utilize mutual information to further enhance the correlation of the matched clusters between views. Extensive experiments show the efficiency of our methods for UMC, compared with the state-of-the-art methods.
Like Xin, Wanqi Yang, Lei Wang 0001, Ming Yang 0014
IEEE Trans. Neural Networks Learn. Syst.2
2025 Unpaired Multiview Clustering via Reliable View Guidance
abstract
This article focuses on unpaired multiview clustering (UMC), a challenging problem, where paired observed samples are unavailable across multiple views. The goal is to perform effective joint clustering using the unpaired observed samples in all views. In incomplete multiview clustering (IMC), existing methods typically rely on sample pairing between views to capture their complementary. However, this is not applicable in the case of UMC. Hence, we aim to extract the consistent cluster structure across views. In UMC, two challenging issues arise: the uncertain cluster structure due to the lack of labels and the uncertain pairing relationship due to the absence of paired samples. We assume that the view with a good cluster structure is the reliable view, which acts as a supervisor to guide the clustering of the other views. With the guidance of reliable views, a more certain cluster structure of these views is obtained while achieving alignment between the reliable views and the other views. Then, we propose reliable view guided UMC with one reliable view (RG-UMC) and reliable view guided UMC with multiple reliable views (RGs-UMC). Specifically, we design alignment modules with one reliable view and multiple reliable views, respectively, to adaptively guide the optimization process. Also, we utilize the compactness module to enhance the relationship of samples within the same cluster. Meanwhile, an orthogonal constraint is applied to the latent representation to obtain discriminate features. Extensive experiments show that both RG-UMC and RGs-UMC outperform the best state-of-the-art method by an average of 24.14% and 29.42% in normalized mutual information (NMI), respectively.
Like Xin, Wanqi Yang, Lei Wang 0001, Ming Yang 0014
IEEE Trans. Neural Networks Learn. Syst.2
2025 Multilevel Reliable Guidance for Unpaired Multiview Clustering
abstract
In this article, we address the challenging problem of unpaired multiview clustering (UMC), which aims to achieve effective joint clustering using unpaired samples observed across multiple views. Traditional incomplete multiview clustering (IMC) methods typically rely on paired samples to capture complementary information between views. However, such strategies become impractical in the UMC due to the absence of paired samples. Although some researchers have attempted to address this issue by preserving consistent cluster structures across views, effectively mining such consistency remains challenging when the cluster structures with low confidence. Therefore, we propose a novel method, multilevel reliable guidance for UMC (MRG-UMC), which integrates multilevel clustering and reliable view guidance to learn consistent and confident cluster structures from three perspectives. Specifically, inner view multilevel clustering exploits high-confidence sample pairs across different levels to reduce the impact of boundary samples, resulting in more confident cluster structures. Synthesized-view alignment leverages a synthesized view to mitigate cross-view discrepancies and promote consistency. Cross-view guidance employs a reliable view guidance strategy to enhance the clustering confidence of poorly clustered views. These three modules are jointly optimized across multiple levels to achieve consistent and confident cluster structures. Furthermore, theoretical analyses verify the effectiveness of MRG-UMC in enhancing clustering confidence. Extensive experimental results show that MRG-UMC outperforms state-of-the-art UMC methods, achieving an average NMI improvement of 12.95% on multiview datasets. The source code is available at https://anonymous.4open.science/r/MRG-UMC-5E20.
Like Xin, Wanqi Yang, Lei Wang 0001, Ming Yang 0014
IEEE Trans. Neural Networks Learn. Syst.2
2025 ZeroTracer: In-Band eBPF-Based Trace Generator With Zero Instrumentation for Microservice Systems
abstract
Microservice enables agility in modern cloud-native applications but introduces challenges in fault troubleshooting due to its complex service coordination and cooperation. To tackle these challenges, distributed tracing has emerged for end-to-end request tracing and system understanding. However, existing tracing solutions often suffer from code instrumentation, trace loss and inaccuracy. To overcome these limitations, we introduce ZeroTracer, an in-kernel online distributed tracing system equipped with an eBPF-based (extended Berkeley Packet Filter) trace generator. ZeroTracer tailors for tracking HTTP requests due to its popularity in microservice systems. In our evaluations, ZeroTracer achieves remarkable trace accuracy (i.e., over 91%) and maintains stable performance under different workload concurrency. Moreover, ZeroTracer outperforms other non-invasive approaches which fail to reconcile accurate request causality. Notably, ZeroTracer effectively tracks end-to-end requests in multi-threaded microservice applications, which is absent in existing invasive distributed tracing systems with third-party library instrumentation. Moreover, ZeroTracer introduces a negligible overhead, with latency increasing by only 0.5%–1.2% and a modest 3%–5.8% increase in CPU and memory consumption.
Wanqi Yang, Pengfei Chen 0002, Huxing Zhang
IEEE Trans. Parallel Distributed Syst.1
2025 eProbe: eBPF-Enhanced Accurate Container Status Probing in Cloud-Native Systems
abstract
Cloud-native systems enhance scalability and availability by leveraging containers. For better container management and scheduling, existing container management systems (e.g., Kubernetes) employ readiness and liveness probes to probe container status. However, failures in these probing systems can lead to traffic scheduling errors, application failures, or even cluster crashes, due to real-time probing issues or implementation bugs. To address these shortcomings, we propose a non-intrusive, real-time container status probing system based on eBPF (extended Berkeley Packet Filter), called eProbe. eProbe intercepts packets passing through containers to probe container status in real time within the operating system kernel, without additional instrumentation or modifications to existing container management systems. Our evaluations show that eProbe can accurately probe container status, whereas Kubernetes probes can only partially achieve this. Moreover, eProbe incurs a low overhead, consuming only 0.03% to 0.26% additional memory and 0.22% to 0.33% additional CPU, while introducing a 2.48% increase in response time for cloud-native applications.
Wanqi Yang, Pengfei Chen 0002
IEEE Trans. Serv. Comput.1
2024 Continual Learning for Temporal-Sensitive Question Answering
abstract
In this study, we explore an emerging research area of Continual Learning for Temporal Sensitive Question Answering (CLTSQA). Previous research has primarily focused on Temporal Sensitive Question Answering (TSQA), often overlooking the unpredictable nature of future events. In real-world applications, it’s crucial for models to continually acquire knowledge over time, rather than relying on a static, complete dataset. Our paper investigates strategies that enable models to adapt to the ever-evolving information landscape, thereby addressing the challenges inherent in CLTSQA. To support our research, we first create a novel dataset, divided into five subsets, designed specifically for various stages of continual learning. We then propose a training framework for CLTSQA that integrates temporal memory replay and temporal contrastive learning. Our experimental results highlight two significant insights: First, the CLTSQA task introduces unique challenges for existing models. Second, our proposed framework effectively navigates these challenges, resulting in improved performance.
Wanqi Yang, Yunqiu Xu, Yanda Li, Kunze Wang, Binbin Huang 0006, Ling Chen 0006
IJCNN1
2024 Network shortcut in data plane of service mesh with eBPF
Wanqi Yang, Pengfei Chen 0002, Guangba Yu, Huxing Zhang
J. Netw. Comput. Appl.1
2024 Siamada: visual tracking based on Siamese adaptive learning network
Wanqi Yang
Neural Comput. Appl.3
2024 Iterative Multiview Subspace Learning for Unpaired Multiview Clustering
abstract
In real applications, several unpredictable or uncertain factors could result in unpaired multiview data, i.e., the observed samples between views cannot be matched. Since joint clustering among views is more effective than individual clustering in each view, we investigate unpaired multiview clustering (UMC), which is a valuable but insufficiently studied problem. Due to lack of matched samples between views, we could fail to build the connection between views. Therefore, we aim to learn the latent subspace shared by views. However, existing multiview subspace learning methods usually rely on the matched samples between views. To address this issue, we propose an iterative multiview subspace learning strategy [iterative unpaired multiview clustering (IUMC)], aiming to learn a complete and consistent subspace representation among views for UMC. Moreover, based on IUMC, we design two effective UMC methods: 1) Iterative unpaired multiview clustering via covariance matrix alignment (IUMC-CA) that further aligns the covariance matrix of subspace representations and then performs clustering on the subspace and 2) iterative unpaired multiview clustering via one-stage clustering assignments (IUMC-CY) that performs one-stage multiview clustering (MVC) by replacing the subspace representations with clustering assignments. Extensive experiments show the excellent performance of our methods for UMC, compared with the state-of-the-art methods. Also, the clustering performance of observed samples in each view can be considerably improved by those observed samples from the other views. In addition, our methods have good applicability in incomplete MVC.
Wanqi Yang, Like Xin, Lei Wang 0001, Ming Yang 0014, Wenzhu Yan, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 High-Level Semantic Feature Matters Few-Shot Unsupervised Domain Adaptation
abstract
In few-shot unsupervised domain adaptation (FS-UDA), most existing methods followed the few-shot learning (FSL) methods to leverage the low-level local features (learned from conventional convolutional models, e.g., ResNet) for classification. However, the goal of FS-UDA and FSL are relevant yet distinct, since FS-UDA aims to classify the samples in target domain rather than source domain. We found that the local features are insufficient to FS-UDA, which could introduce noise or bias against classification, and not be used to effectively align the domains. To address the above issues, we aim to refine the local features to be more discriminative and relevant to classification. Thus, we propose a novel task-specific semantic feature learning method (TSECS) for FS-UDA. TSECS learns high-level semantic features for image-to-class similarity measurement. Based on the high-level features, we design a cross-domain self-training strategy to leverage the few labeled samples in source domain to build the classifier in target domain. In addition, we minimize the KL divergence of the high-level feature distributions between source and target domains to shorten the distance of the samples between the two domains. Extensive experiments on DomainNet show that the proposed method significantly outperforms SOTA methods in FS-UDA by a large margin (i.e., ~10%).
Wanqi Yang, Shengqi Huang, Lei Wang 0001, Ming Yang 0014
AAAI2
2023 A robust tracking architecture using tracking failure detection in Siamese trackers
Yanchun Zhao, Wanqi Yang
Appl. Intell.4
2022 Few-Shot Unsupervised Domain Adaptation via Meta Learning
abstract
Unsupervised domain adaptation (UDA) has raised a lot of interests in recent years. However, current UDA methods are still not capable enough in dealing with two issues: 1) the scarcity of labeled data in source domain and 2) the need of a general model that can quickly adapt to solve new UDA tasks. To address this situation, we investigate available but rarely-studied setting called few-shot unsupervised domain adaptation (FS-UDA), in which the data of source domain is few-shot per category and the data of target domain remains unlabeled. To realize effective adaptation for FS-UDA tasks in the same source and target domains, we propose a novel meta learning method namely meta-FUDA, which leverages meta learning to perform task-level transfer and domain-level transfer jointly. Extensive experiments demonstrate the promising performance of our method on multiple benchmark data sets.
Wanqi Yang, Chengmei Yang, Shengqi Huang, Lei Wang 0001, Ming Yang 0014
ICME1
2022 Design, Modeling and Control of a Composable and Extensible Drone with Tilting Rotors
abstract
In this paper, we introduce a composable and extensible drone with tilting rotors (CEDTR). We aimed for a function that could optimally match the load capacity, degree of freedom (DOF), speed and endurance with diverse mission requirements by changing the quantity and form of combinations. First, we propose a decentralized modular controller to allow a team of physically connected modules to fly cooperatively. Second, we divide all the combinations into three categories according to the different control methods. Three generalized control strategies are proposed to control both position and attitude independently by tilting the directions of the propellers. We carried out experiments to demonstrate the feasibility of this mechanical design and control method. The experiment video is available at https://youtu.be/7RvxiV4FPq4.
Zegui Wu, Ruqing Zhao, Mengfan Yu, Yanchun Zhao, Wanqi Yang, Weiye Zhang
IROS5
2022 Dual-scale correlation analysis for robust multi-label classification
Kaixiang Wang 0001, Ming Yang 0014, Wanqi Yang, Lei Wang 0001
Appl. Intell.3
2021 Few-shot Unsupervised Domain Adaptation with Image-to-Class Sparse Similarity Encoding
abstract
This paper investigates a valuable setting called few-shot unsupervised domain adaptation (FS-UDA), which has not been sufficiently studied in the literature. In this setting, the source domain data are labelled, but with few-shot per category, while the target domain data are unlabelled. To address the FS-UDA setting, we develop a general UDA model to solve the following two key issues: the few-shot labeled data per category and the domain adaptation between support and query sets. Our model is general in that once trained it will be able to be applied to various FS-UDA tasks from the same source and target domains. Inspired by the recent local descriptor based few-shot learning (FSL), our general UDA model is fully built upon local descriptors (LDs) for image classification and domain adaptation. By proposing a novel concept called similarity patterns (SPs), our model not only effectively considers the spatial relationship of LDs that was ignored in previous FSL methods, but also makes the learned image similarity better serve the required domain alignment. Specifically, we propose a novel IMage-to-class sparse Similarity Encoding (IMSE) method. It learns SPs to extract the local discriminative information for classification and meanwhile aligns the covariance matrix of the SPs for domain adaptation. Also, domain adversarial training and multi-scale local feature matching are performed upon LDs. Extensive experiments conducted on a multi-domain benchmark dataset DomainNet demonstrates the state-of-the-art performance of our IMSE for the novel setting of FS-UDA. In addition, for FSL, our IMSE can also show better performance than most of recent FSL methods on miniImageNet.
Shengqi Huang, Wanqi Yang, Lei Wang 0001, Luping Zhou, Ming Yang 0014
ACM Multimedia2
2021 Learning-Based Computer-Aided Prescription Model for Parkinson's Disease: A Data-Driven Perspective
abstract
In this article, we study a novel problem: "automatic prescription recommendation for PD patients." To realize this goal, we first build a dataset by collecting 1) symptoms of PD patients, and 2) their prescription drug provided by neurologists. Then, we build a novel computer-aided prescription model by learning the relation between observed symptoms and prescription drug. Finally, for the new coming patients, we could recommend (predict) suitable prescription drug on their observed symptoms by our prescription model. From the methodology part, our proposed model, namely Prescription viA Learning lAtent Symptoms (PALAS), could recommend prescription using the multi-modality representation of the data. In PALAS, a latent symptom space is learned to better model the relationship between symptoms and prescription drug, as there is a large semantic gap between them. Moreover, we present an efficient alternating optimization method for PALAS. We evaluated our method using the data collected from 136 PD patients at Nanjing Brain Hospital, which can be regarded as a large dataset in PD research community. The experimental results demonstrate the effectiveness and clinical potential of our method in this recommendation task, if compared with other competing methods.
Yinghuan Shi, Wanqi Yang, Kim-Han Thung, Hao Wang 0013, Yang Gao 0001, Dinggang Shen
IEEE J. Biomed. Health Informatics2
2021 SA-LuT-Nets: Learning Sample-Adaptive Intensity Lookup Tables for Brain Tumor Segmentation
abstract
In clinics, the information about the appearance and location of brain tumors is essential to assist doctors in diagnosis and treatment. Automatic brain tumor segmentation on the images acquired by magnetic resonance imaging (MRI) is a common way to attain this information. However, MR images are not quantitative and can exhibit significant variation in signal depending on a range of factors, which increases the difficulty of training an automatic segmentation network and applying it to new MR images. To deal with this issue, this paper proposes to learn a sample-adaptive intensity lookup table (LuT) that dynamically transforms the intensity contrast of each input MR image to adapt to the following segmentation task. Specifically, the proposed deep SA-LuT-Net framework consists of a LuT module and a segmentation module, trained in an end-to-end manner: the LuT module learns a sample-specific nonlinear intensity mapping function through communication with the segmentation module, aiming at improving the final segmentation performance. In order to make the LuT learning sample-adaptive, we parameterize the intensity mapping function by exploring two families of non-linear functions (i.e., piece-wise linear and power functions) and predict the function parameters for each given sample. These sample-specific parameters make the intensity mapping adaptive to samples. We develop our SA-LuT-Nets separately based on two backbone networks for segmentation, i.e., DMFNet and the modified 3D Unet, and validate them on BRATS2018 and BRATS2019 datasets for brain tumor segmentation. Our experimental results clearly demonstrate the superior performance of the proposed SA-LuT-Nets using either single or multiple MR modalities. It not only significantly improves the two baselines (DMFNet and the modified 3D Unet), but also wins a set of state-of-the-art segmentation methods. Moreover, we show that, the LuTs learnt using one segmentation model could also be applied to improving the performance of another segmentation model, indicating the general segmentation information captured by LuTs.
Biting Yu, Luping Zhou, Lei Wang 0001, Wanqi Yang, Ming Yang 0014, Pierrick Bourgeat, Jurgen Fripp
IEEE Trans. Medical Imaging4
2020 Learning Sample-Adaptive Intensity Lookup Table for Brain Tumor Segmentation
Biting Yu, Luping Zhou, Lei Wang 0001, Wanqi Yang, Ming Yang 0014, Pierrick Bourgeat, Jurgen Fripp
MICCAI (4)4
2020 A novel spectral-spatial based adaptive minimum spanning forest for hyperspectral image classification
Jing Lv, Ming Yang 0014, Wanqi Yang
GeoInformatica4
2020 An Effective MR-Guided CT Network Training for Segmenting Prostate in CT Images
abstract
Segmentation of prostate in medical imaging data (e.g., CT, MRI, TRUS) is often considered as a critical yet challenging task for radiotherapy treatment. It is relatively easier to segment prostate from MR images than from CT images, due to better soft tissue contrast of the MR images. For segmenting prostate from CT images, most previous methods mainly used CT alone, and thus their performances are often limited by low tissue contrast in the CT images. In this article, we explore the possibility of using indirect guidance from MR images for improving prostate segmentation in the CT images. In particular, we propose a novel deep transfer learning approach, i.e., MR-guided CT network training (namely MICS-NET), which can employ MR images to help better learning of features in CT images for prostate segmentation. In MICS-NET, the guidance from MRI consists of two steps: (1) learning informative and transferable features from MRI and then transferring them to CT images in a cascade manner, and (2) adaptively transferring the prostate likelihood of MRI model (i.e., well-trained convnet by purely using MR images) with a view consistency constraint. To illustrate the effectiveness of our approach, we evaluate MICS-NET on a real CT prostate image set, with the manual delineations available as the ground truth for evaluation. Our methods generate promising segmentation results which achieve (1) six percentages higher Dice Ratio than the CT model purely using CT images and (2) comparable performance with the MRI model purely using MR images.
Wanqi Yang, Yinghuan Shi, Sanghyun Park 0004, Ming Yang 0014, Yang Gao 0001, Dinggang Shen
IEEE J. Biomed. Health Informatics1
2018 Attributes Consistent Faces Generation Under Arbitrary Poses
Fengyi Song, Jinhui Tang 0001, Ming Yang 0014, Weiling Cai, Wanqi Yang
ACCV (2)5
2018 Deep Correlation Structure Preserved Label Space Embedding for Multi-label Classification
abstract
Label embedding is an effective and efficient method which can jointly extract the information of all labels for better performance of multi-label classification. However, most existing embedding methods ignore information of feature space or intrinsic structure of previous label space, such that their learned latent space will not have strong predictability and discriminant ability. We propose a novel deep neural network (DNN) based model, namely Deep Correlation Structure Preserved Label Space Embedding (DCSPE). Specifically, DCSPE derives a deep latent space by performing feature-aware label space embedding with deep canonical correlation analysis (DCCA) and preserving the intrinsic structure of the previous label space with proposed deep multidimensional scaling (DMDS). Our DCSPE is achieved by integrating the DNN architectures of the two DNN based models and can learn a feature-aware structure preserved deep latent space. Furthermore, extensive experimental results on datasets with many labels demonstrate that our proposed approach is significantly better than the existing label embedding algorithms.
Kaixiang Wang 0001, Ming Yang 0014, Wanqi Yang, Yilong Yin
ACML3
2018 Deep Cross-View Label Embedding with Correlation and Structure Preserved for Multi-Label Classification
abstract
Label embedding is an important family of multi-label classification algorithms which can jointly extract the information of all labels for better performance. However, few works have been done on label embedding methods which consider the structure information of original feature and label space simultaneously. We propose a novel deep neural network (DNN) based model for learning an effective deep latent space, namely Deep Cross-view label space Embedding with Correlation and Structure preserved (DCECS). In DCECS, the latent space correlates with feature and label spaces closely by virtue of the deep cross-view embedding. Meanwhile, the latent space is also learned under the guidance of label correlation and local structure of feature space which are exploited by hypergraph and graph regularizations. The overall framework achieves the complementarity and correspondence between information of feature and label space, therefore the feature-aware deep latent space we learned has strong predictability and discriminant ability. Extensive experimental results on datasets with many labels demonstrate that our proposed approach is significantly better than the existing label embedding algorithms.
Kaixiang Wang 0001, Ming Yang 0014, Wanqi Yang, Yilong Yin
ICTAI3
2018 Feature Integration with Adaptive Importance Maps for Visual Tracking
abstract
Discriminative correlation filters have recently achieved excellent performance for visual object tracking. The key to success is to make full use of dense sampling and specific properties of circulant matrices in the Fourier domain. However, previous studies don't take into consideration the importance and complementary information of different features, simply concatenating them. This paper investigates an effective method of feature integration for correlation filters, which jointly learns filters, as well as importance maps in each frame. These importance maps borrow the advantages of different features, aiming to achieve complementary traits and improve robustness. Moreover, for each feature, an importance map is shared by its all channels to avoid overfitting. In addition, we introduce a regularization term for the importance maps and use the penalty factor to control the significance of features. Based on handcrafted and CNN features, we implement two trackers, which achieve a competitive performance compared with several state-of-the-art trackers.
Aishi Li, Ming Yang 0014, Wanqi Yang
IJCAI3
2018 Online multi-view subspace learning via group structure analysis for visual object tracking
Wanqi Yang, Yinghuan Shi, Yang Gao 0001, Ming Yang 0014
Distributed Parallel Databases1
2018 Heterogeneous Face Recognition by Margin-Based Cross-Modality Metric Learning
abstract
Heterogeneous face recognition deals with matching face images from different modalities or sources. The main challenge lies in cross-modal differences and variations and the goal is to make cross-modality separation among subjects. A margin-based cross-modality metric learning (MCM2L) method is proposed to address the problem. A cross-modality metric is defined in a common subspace where samples of two different modalities are mapped and measured. The objective is to learn such metrics that satisfy the following two constraints. The first minimizes pairwise, intrapersonal cross-modality distances. The second forces a margin between subject specific intrapersonal and interpersonal cross-modality distances. This is achieved by defining a hinge loss on triplet-based distance constraints for efficient optimization. It allows the proposed method to focus more on optimizing distances of those subjects whose intrapersonal and interpersonal distances are hard to separate. The proposed method is further extended to a kernelized MCM2L (KMCM2L). Both methods have been evaluated on an ID card face dataset and two other cross-modality benchmark datasets. Various feature extraction methods have also been incorporated in the study, including recent deep learned features. In extensive experiments and comparisons with the state-of-the-art methods, the MCM2L and KMCM2L methods achieved marked improvements in most cases.
Jing Huo, Yang Gao 0001, Yinghuan Shi, Wanqi Yang, Hujun Yin
IEEE Trans. Cybern.4
2018 Incomplete-Data Oriented Multiview Dimension Reduction via Sparse Low-Rank Representation
abstract
For dimension reduction on multiview data, most of the previous studies implicitly take an assumption that all samples are completed in all views. Nevertheless, this assumption could often be violated in real applications due to the presence of noise, limited access to data, equipment malfunction, and so on. Most of the previous methods will cease to work when missing values in one or multiple views occur, thus an incomplete-data oriented dimension reduction becomes an important issue. To this end, we mathematically formulate the above-mentioned issue as sparse low-rank representation through multiview subspace (SRRS) learning to impute missing values, by jointly measuring intraview relations (via sparse low-rank representation) and interview relations (through common subspace representation). Moreover, by exploiting various subspace priors in the proposed SRRS formulation, we develop three novel dimension reduction methods for incomplete multiview data: 1) multiview subspace learning via graph embedding; 2) multiview subspace learning via structured sparsity; and 3) sparse multiview feature selection via rank minimization. For each of them, the objective function and the algorithm to solve the resulting optimization problem are elaborated, respectively. We perform extensive experiments to investigate their performance on three types of tasks including data recovery, clustering, and classification. Both two toy examples (i.e., Swiss roll and -curve) and four real-world data sets (i.e., face images, multisource news, multicamera activity, and multimodality neuroimaging data) are systematically tested. As demonstrated, our methods achieve the performance superior to that of the state-of-the-art comparable methods. Also, the results clearly show the advantage of integrating the sparsity and low-rankness over using each of them separately.
Wanqi Yang, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Ming Yang 0014
IEEE Trans. Neural Networks Learn. Syst.1
2017 Does Manual Delineation only Provide the Side Information in CT Prostate Segmentation?
Yinghuan Shi, Wanqi Yang, Yang Gao 0001, Dinggang Shen
MICCAI (3)2
2016 Multi-view Subspace Clustering via a Global Low-Rank Affinity Matrix
Lei Qi 0001, Yinghuan Shi, Wanqi Yang, Yang Gao 0001
IDEAL4
2016 Ensemble of Sparse Cross-Modal Metrics for Heterogeneous Face Recognition
abstract
Heterogeneous face recognition aims to identify or verify person identity by matching facial images of different modalities. In practice, it is known that its performance is highly influenced by modality inconsistency, appearance occlusions, illumination variations and expressions. In this paper, a new method named as ensemble of sparse cross-modal metrics is proposed for tackling these challenging issues. In particular, a weak sparse cross-modal metric learning method is firstly developed to measure distances between samples of two modalities. It learns to adjust rank-one cross-modal metrics to satisfy two sets of triplet based cross-modal distance constraints in a compact form. Meanwhile, a group based feature selection is performed to enforce that features in the same position of two modalities are selected simultaneously. By neglecting features that attribute to "noise" in the face regions (eye glasses, expressions and so on), the performance of learned weak metrics can be markedly improved. Finally, an ensemble framework is incorporated to combine the results of differently learned sparse metrics into a strong one. Extensive experiments on various face datasets demonstrate the benefit of such feature selection especially when heavy occlusions exist. The proposed ensemble metric learning has been shown superiority over several state-of-the-art methods in heterogeneous face recognition.
Jing Huo, Yang Gao 0001, Yinghuan Shi, Wanqi Yang, Hujun Yin
ACM Multimedia4
2015 Interactive image segmentation via cascaded metric learning
abstract
In this paper, we propose an interactive image segmentation method from a novel perspective of cascaded metric learning. Given an image with user-marked scribbles that are essentially uncertain and noisy, our method completes the segmentation task by solving a binary classification problem. Starting from the initial training samples with known class labels (i.e., regions of the image that are believed with high confidence to be foreground or background), we first find an optimal metric that can best describe the classification of these samples. After that, we classify the unlabeled samples using the learnt metric. Samples classified with high confidence are used as new training samples to refine the metric. This cycle of metric learning and classification repeats until the accomplishment of the image segmentation task. The proposed method is extensively evaluated on the MSRC image set. Experiment results show that our method outperforms the state-of-the-art methods.
Wenbin Li 0006, Yinghuan Shi, Wanqi Yang, Hao Wang 0013, Yang Gao 0001
ICIP3
2015 MRM-Lasso: A Sparse Multiview Feature Selection Method via Low-Rank Analysis
abstract
Learning about multiview data involves many applications, such as video understanding, image classification, and social media. However, when the data dimension increases dramatically, it is important but very challenging to remove redundant features in multiview feature selection. In this paper, we propose a novel feature selection algorithm, multiview rank minimization-based Lasso (MRM-Lasso), which jointly utilizes Lasso for sparse feature selection and rank minimization for learning relevant patterns across views. Instead of simply integrating multiple Lasso from view level, we focus on the performance of sample-level (sample significance) and introduce pattern-specific weights into MRM-Lasso. The weights are utilized to measure the contribution of each sample to the labels in the current view. In addition, the latent correlation across different views is successfully captured by learning a low-rank matrix consisting of pattern-specific weights. The alternating direction method of multipliers is applied to optimize the proposed MRM-Lasso. Experiments on four real-life data sets show that features selected by MRM-Lasso have better multiview classification performance than the baselines. Moreover, pattern-specific weights are demonstrated to be significant for learning about multiview data, compared with view-specific weights.
Wanqi Yang, Yang Gao 0001, Yinghuan Shi, Longbing Cao
IEEE Trans. Neural Networks Learn. Syst.1
2014 A Novel Ego-Centered Academic Community Detection Approach via Factor Graph Model
Yusheng Jia, Yang Gao 0001, Wanqi Yang, Jing Huo, Yinghuan Shi
IDEAL3
2014 mPadal: a joint local-and-global multi-view feature selection method for activity recognition
Wanqi Yang, Yang Gao 0001, Longbing Cao, Ming Yang 0014, Yinghuan Shi
Appl. Intell.1
2014 Multi-Instance Dictionary Learning for Detecting Abnormal Events in Surveillance Videos
abstract
In this paper, a novel method termed Multi-Instance Dictionary Learning (MIDL) is presented for detecting abnormal events in crowded video scenes. With respect to multi-instance learning, each event (video clip) in videos is modeled as a bag containing several sub-events (local observations); while each sub-event is regarded as an instance. The MIDL jointly learns a dictionary for sparse representations of sub-events (instances) and multi-instance classifiers for classifying events into normal or abnormal. We further adopt three different multi-instance models, yielding the Max-Pooling-based MIDL (MP-MIDL), Instance-based MIDL (Inst-MIDL) and Bag-based MIDL (Bag-MIDL), for detecting both global and local abnormalities. The MP-MIDL classifies observed events by using bag features extracted via max-pooling over sparse representations. The Inst-MIDL and Bag-MIDL classify observed events by the predicted values of corresponding instances. The proposed MIDL is evaluated and compared with the state-of-the-art methods for abnormal event detection on the UMN (for global abnormalities) and the UCSD (for local abnormalities) datasets and results show that the proposed MP-MIDL and Bag-MIDL achieve either comparable or improved detection performances. The proposed MIDL method is also compared with other multi-instance learning methods on the task and superior results are obtained by the MP-MIDL scheme.
Jing Huo, Yang Gao 0001, Wanqi Yang, Hujun Yin
Int. J. Neural Syst.3
2013 Image Super Resolution via Visual Prior Based Digital Image Characteristics
Yusheng Jia, Wanqi Yang, Yang Gao 0001, Hujun Yin, Yinghuan Shi
IDEAL2
2013 TRASMIL: A local anomaly detection framework based on trajectory segmentation and multi-instance learning
Wanqi Yang, Yang Gao 0001, Longbing Cao
Comput. Vis. Image Underst.1
2012 Abnormal Event Detection via Multi-Instance Dictionary Learning
Jing Huo, Yang Gao 0001, Wanqi Yang, Hujun Yin
IDEAL3