Yang Gao 0001

dblp:89/4402-1 · DBLP profile ↗
← Back
292ranked-venue papers
7as first author
174since 2021 · last 2026
0000-0002-2488-1813ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 186 · 3 first-author · 112 since 2021Graphics, computer vision, multimedia, augmented reality and games · 108 · 68 since 2021Applied, interdisciplinary, general and emerging computing · 39 · 2 first-author · 22 since 2021Databases, data management, data science and information retrieval · 22 · 2 first-author · 9 since 2021Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Computer networks · 3 · 2 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Causality-Aware Efficient Exploration for Cooperative Multi-Agent Reinforcement Learning
abstract
Exploration is critical for cooperative multi agent reinforcement learning (MARL) to improve sample efficiency. However, existing intrinsic motivation based exploration strategies in MARL overlook the causal relationships among agents, global states, and rewards, suffering from interference by irrelevant factors and resulting in sample inefficiency. To address this issue, we propose Causality aware Efficient Exploration (CEE), a novel framework that enhances sample efficiency by inferring causal relationships between agents, global states with respect to rewards, thereby enabling causality guided exploration. Specifically, CEE operates through two components. First, CEE identifies causal relationships between global states and rewards, filtering out causally irrelevant state features that do not have a high impact on rewards to keep decision critical state information. Second, CEE discovers causal relationships between agents' behaviors and rewards to quantify each agent's contribution to collective performance. To achieve this, we introduce a causal entropy objective that promotes exploration aligned with decision critical aspects of the underlying causal structure. We provide comprehensive validation through experiments on 21 challenging tasks spanning SMAC, SMAC v2, and Google Research Football (GRF) environments. Our results demonstrate that CEE achieves superior performance in terms of sample efficiency and asymptotic performance compared to existing MARL methods.
Hongye Cao, Tianpei Yang, Hammadi Rafik Ouariachi, Yali Du 0001, Jing Huo, Yang Gao 0001
AAAI8
2026 ManiLong-Shot: Interaction-Aware One-Shot Imitation Learning for Long-Horizon Manipulation
abstract
One-shot imitation learning (OSIL) offers a promising way to teach robots new skills without large-scale data collection. However, current OSIL methods are primarily limited to short-horizon tasks, thus limiting their applicability to complex, long-horizon manipulations. To address this limitation, we propose ManiLong-Shot, a novel framework that enables effective OSIL for long-horizon prehensile manipulation tasks. ManiLong-Shot structures long-horizon tasks around physical interaction events, reframing the problem as sequencing interaction-aware primitives instead of directly imitating continuous trajectories. This primitive decomposition can be driven by high-level reasoning from a vision-language model (VLM) or by rule-based heuristics derived from robot state changes. For each primitive, ManiLong-Shot predicts invariant regions critical to the interaction, establishes correspondences between the demonstration and the current observation, and computes the target end-effector pose, enabling effective task execution. Extensive simulation experiments show that ManiLong-Shot, trained on only 10 short-horizon tasks, generalizes to 20 unseen long-horizon tasks across three difficulty levels via one-shot imitation, achieving a 22.8% relative improvement over the SOTA. Additionally, real-robot experiments validate ManiLong-Shot’s ability to robustly execute three long-horizon manipulation tasks via OSIL, confirming its practical applicability.
Chongkai Gao, Lin Shao 0002, Jieqi Shi, Jing Huo, Yang Gao 0001
AAAI6
2026 Faster Game Solving via Asymmetry of Step Sizes
abstract
Counterfactual Regret Minimization (CFR) algorithms are widely used to compute a Nash equilibrium (NE) in two-player zero-sum imperfect-information extensive-form games (IIGs). Among them, Predictive CFR+ (PCFR+) is particularly powerful, achieving an exceptionally fast empirical convergence rate via the prediction in many games. However, the empirical convergence rate of PCFR+ would significantly degrade if the prediction is inaccurate, leading to unstable performance on certain IIGs. To enhance the robustness of PCFR+, we propose Asymmetric PCFR+ (APCFR+), which employs an adaptive asymmetry of step sizes between the updates of implicit and explicit accumulated counterfactual regrets to mitigate the impact of the prediction inaccuracy on convergence. We present a theoretical analysis demonstrating why APCFR+ can enhance the robustness. To the best of our knowledge, we are the first to propose the asymmetry of step sizes, a simple yet novel technique that effectively improves the robustness of PCFR+. Then, to reduce the difficulty of implementing APCFR+ caused by the adaptive asymmetry, we propose a simplified version of APCFR+ called Simple APCFR+ (SAPCFR+), which uses a fixed asymmetry of step sizes to enable only a single-line modification compared to original PCFR+. Experimental results on five standard IIG benchmarks and two heads-up no-limit Texas Hold’em (HUNL) Subagems show that (i) both APCFR+ and SAPCFR+ outperform PCFR+ in most of the tested games, (ii) SAPCFR+ achieves a comparable empirical convergence rate with APCFR+, and (iii) our approach can be generalized to improve other CFR algorithms, e.g., Discount CFR (DCFR).
Linjian Meng, Tianpei Yang, Youzhi Zhang 0001, Zhenxing Ge, Yang Gao 0001
AAAI5
2026 Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning
abstract
Siyuan Gan, Jiaheng Liu, Boyan Wang, Tianpei Yang, Runqing Miao, Yuyao Zhang, Fanyu Meng, Junlan Feng, Linjian Meng, Jing Huo, Yang Gao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Siyuan Gan, Tianpei Yang, Runqing Miao, Junlan Feng, Linjian Meng, Jing Huo, Yang Gao 0001
ACL (1)11
2026 MDTace: Agentic Context Engineering for Multi-disciplinary Team Medical Consultation
Kai Chen 0026, Yang Gao 0001
ICIC (8)2
2026 MedMentor: Teacher-Student Collaboration with Decoupled Experience for Clinical Diagnosis
Kai Chen 0026, Yang Gao 0001
ICIC (8)2
2026 Diffusion-Based Data Augmentation for Image Recognition: A Systematic Analysis and Evaluation
Zekun Li 0010, Yinghuan Shi, Yang Gao 0001
Int. J. Comput. Vis.3
2026 Efficient privacy-preserving sparse matrix-vector multiplication using homomorphic encryption
Yang Gao 0001, Gang Quan, Wujie Wen, Scott Piersall, Qian Lou, Liqiang Wang 0001
Inf. Sci.1
2026 Improving MPI error detection and repair with large language models and bug references
Scott Piersall, Yang Gao 0001, Shenyang Liu, Liqiang Wang 0001
J. Parallel Distributed Comput.2
2026 Mitigating Security Risks in Large Language Models: A Full Lifecycle Perspective
Yanming Wang, Zhixin Bai, Jing Huo, Hongye Cao, Yang Gao 0001
Mach. Learn.9
2026 A unified and efficient training framework for open-ended non-transitive games
Shaokang Dong, Shangdong Yang, Hongye Cao, Wanqi Yang, Yang Gao 0001
Neural Networks6
2026 A continual learning framework with long-term and multiple short-term memory networks
Shangge Liu, Lei Wang 0001, Rui Yan 0005, Jing Huo, Wenbin Li 0006, Yang Gao 0001
Neural Networks6
2026 AMPL: An adaptive meta-prompt learner for few-shot image classification
Zhiping Wu, Lian Huai, Zeyu Shangguan, Lei Wang 0001, Jing Huo, Wenbin Li 0006, Yang Gao 0001, Xingqun Jiang
Neural Networks8
2026 Toward the Connection Between Activation Sparsity and Flat Minima
abstract
The observation that activation sparsity emerges in MLP blocks of standardly trained Transformers offers an opportunity to drastically reduce computation costs without sacrificing performance. To theoretically explain this phenomenon, existing works have shown that activation sparsity does not result from the data properties or data fitting but from the implicit bias of the training process. However, these connections are obtained with strong assumptions (e.g., shallow networks, a small number of training steps, and special training techniques), which cannot be applied to deep models standardly trained with a large number of steps. Different from these works, we find that the flatness of loss landscapes is also closely related to the MLP activation sparsity and can serve as a weaker assumption because it naturally emerges in the standard training of deep networks without the above strong assumptions. Specifically, we find that 1) the MLP activation sparsity equals a ratio between "augmented flatness" (a weighted sum of flatness measures) and the product of the input norm and activation gradient of the MLP. We empirically find that this ratio decreases during training, leading to sparse activations. 2) We also propose the notion of derivative sparsity, which reduces to activation sparsity under $\operatorname{ReLU}$ReLU, but further enables pruning in the backward propagation and is more stable than activation sparsity. With the theoretical findings, we can further encourage activation sparsity by decreasing the numerator and increasing the denominator of the ratio: 1) To improve (lower) the flatness, we add different bias vectors to input tokens of MLP blocks to strengthen stochastic gradient noise that drives the model to a flat area. 2) We restrict the lower bound of affine parameters in LayerNorm to increase the input norm of MLPs. 3) To increase the activation sparsity, we propose an activation function $\operatorname{JSReLU}$JSReLU to encourage the search of parameters with sparse derivatives and sparse activations. These plug-and-play modifications can effectively reduce the ratio and produce sparser activations. Experiments on ImageNet-1K and C4 demonstrate relative improvements of at least 36% on inference sparsity and at least 50% on training sparsity over vanilla Transformers, indicating further potential cost reduction in both inference and training.
Ze Peng 0001, Jian Zhang 0090, Lei Qi 0001, Yang Gao 0001, Yinghuan Shi
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Unlocking the candidates: Beam-aware reasoning for audio-visual speech recognition
Shao Zeng, Tianjun Gu, Shouhong Ding, Xin Tan 0002, Yang Gao 0001
Pattern Recognit.11
2026 From static to adaptive multi-view: Nuanced expert prompt tuning for Fine-Grained Image Retrieval
Ke-Yue Zhang, Jingyu Gong, Yang Gao 0001, Xin Tan 0002, Lizhuang Ma
Pattern Recognit.5
2026 Unsupervised few-shot learning with object-aware and attribute-consistent augmentation
Zhiping Wu, Lian Huai, Zeyu Shangguan, Wenbin Li 0006, Yang Gao 0001, Xingqun Jiang
Pattern Recognit.6
2026 Retrieval-augmented diffusion with acoustic priors for high-fidelity sonar image generation
Shaocong Yang, Zheng Gu 0001, Hongye Cao, Xiaolong Qi, Jing Huo, Yang Gao 0001
Pattern Recognit.7
2026 Double Buffer Vaccination: Bolstering Immunity Against Catastrophic Forgetting in Continual Medical Image Segmentation
Kai Chen 0026, Hewei Wang 0001, Pinzhuo Tian, Jing Huo, Taihang Zhen, Yang Gao 0001
IEEE Trans. Circuits Syst. Video Technol.6
2026 Diversity-Enhanced Collaborative Mamba for Semi-Supervised Medical Image Segmentation
abstract
Acquiring high-quality annotated data for medical image segmentation is tedious and costly. Semi-supervised segmentation techniques alleviate this burden by leveraging unlabeled data to generate pseudo labels. Recently, advanced state space models, represented by Mamba, have shown efficient handling of long-range dependencies. This drives us to explore their potential in semi-supervised medical image segmentation. In this paper, we propose a novel Diversity-enhanced Collaborative Mamba framework (namely DCMamba) for semi-supervised medical image segmentation, which explores and utilizes the diversity from data, network, and feature perspectives. Firstly, from the data perspective, we develop patch-level weak-strong mixing augmentation with Mamba's scanning modeling characteristics. Moreover, from the network perspective, we introduce a diverse-scan collaboration module, which could benefit from the prediction discrepancies arising from different scanning directions. Furthermore, from the feature perspective, we adopt an uncertainty-weighted contrastive learning mechanism to enhance the diversity of feature representation. Experiments demonstrate that our DCMamba significantly outperforms other semi-supervised medical image segmentation methods, e.g., yielding the latest SSM-based method by 6.69% on the Synapse dataset with 20% labeled data. The code is available at https://github.com/ShumengLI/DCMamba.
Shumeng Li, Jian Zhang 0090, Lei Qi 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Medical Imaging6
2026 Model-Based Offline Reinforcement Learning With Adversarial Data Augmentation
abstract
Model-based offline reinforcement learning (RL) constructs environment models from offline datasets to perform conservative policy optimization. Existing approaches focus on learning state transitions through ensemble models, rolling out conservative estimation to mitigate extrapolation errors. However, the static data makes it challenging to develop a robust policy, and offline agents cannot access the environment to gather new data. To address these challenges, we introduce Model-based Offline Reinforcement learning with AdversariaL data augmentation (MORAL). In MORAL, we replace the fixed horizon rollout by employing adversarial data augmentation to execute alternating sampling with ensemble models to enrich training data. Specifically, this adversarial process dynamically selects ensemble models against policy for biased sampling, mitigating the optimistic estimation of fixed models, thus robustly expanding the training data for policy optimization. Moreover, a differential factor (DF) is integrated into the adversarial process for regularization, ensuring error minimization in extrapolations. This data-augmented optimization adapts to diverse offline tasks without rollout horizon tuning, showing remarkable applicability. Extensive experiments on the D4RL benchmark demonstrate that MORAL outperforms other model-based offline RL methods in terms of policy learning and sample efficiency.
Hongye Cao, Jing Huo, Shangdong Yang, Tianpei Yang, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.7
2026 Pushing Physical Limits and Uncovering Motion Templates of Spine-Based Quadruped Locomotion via Reinforcement Learning
abstract
Flexible spines are critical to the remarkable agility and speed of animals. Translating this biological advantage to quadruped robots presents a significant control challenge, particularly in coordinating the spine and limbs for maximal velocity. In this work, we utilize reinforcement learning (RL) to develop high-speed locomotion for a bioinspired mouse robot with a lateral flexible spine. The resulting controller achieves motor performance that demonstrably surpasses non-spined and model-based methods. More importantly, our analysis reveals the principles behind this performance: the emergence of two distinct motion templates. For high-speed walking, the robot learns a “whip-like” spinal oscillation to increase leg swing frequency, while for agile turning, it adopts a dynamic “bend-and-straighten” pattern. These findings demonstrate the capability of RL to not only generate high-performance controllers but also to produce emergent strategies that, upon analysis, reveal underlying principles of high-speed, spine-driven locomotion.
Zhenshan Bing, Yulong Xiao, Yuhong Huang, Long Cheng 0007, Biao Hu 0001, Gang Chen 0023, Yang Gao 0001, Fuchun Sun 0001, Kai Huang 0001, Alois C. Knoll
IEEE Trans. Robotics8
2026 FISN: FInding Spatial Neighborhoods for Generalizable Novel View Synthesis
abstract
We present FISN, a generalizable novel view synthesis algorithm that enables feedforward inference of Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS) from reference images. Unlike existing work that either separately model the 3D feature space on each view or process multiview reference features by 3D-point-based view aggregation, FISN integrates multi-reference 3D cost volumes into a unified high-dimensional entity. Specifically, we reconceptualize the generalizable novel view synthesis task as a feedforward process of FInding Spatial Neighborhoods across this unified 4D feature space, comprising both view and spatial dimensions, and introduce View-Spatial Convolutions for direct 4D feature aggregation. This enhances the correlation among multiview neighboring points in a window-to-window manner and incorporates 3D spatial awareness. However, this approach poses two intertwined challenges: high computational expense for high-dimensional features and degraded rendering performance with low-resolution features. To address these challenges, FISN constructs a new efficient convolution paradigm, Decomposable View-Spatial Convolution, which includes a Spatial Cross Decomposition strategy as well as a Feature Compression and Upscaling module. This paradigm maintains multiview geometric consistency better than existing decomposition methods and achieves a balance between efficiency and fine-grained spatial features. Furthermore, by integrating Depth Refinement modules based on this paradigm, FISN further improves global depth understanding. Comprehensive evaluations on mainstream datasets and benchmarks demonstrate that FISN achieves state-of-the-art performance for both NeRF and 3DGS, and remains robust in challenging scenarios where existing 3DGS-based methods struggle, such as those with noisy poses or dense references. The code will be released soon.
Yanqi Bao, Tianyu Ding, Jing Huo, Wenbin Li 0006, Yang Gao 0001
IEEE Trans. Vis. Comput. Graph.5
2025 Beyond Mandatory Federations: Balancing Egoism, Utilitarianism and Egalitarianism in Mixed-Motive Games
abstract
In the field of mixed-motive games, extensive multi-agent learning studies have explored the balance between egoism (individual interest), utilitarianism (collective interest), and egalitarianism (fairness). Traditional approaches often rely on manually designed reward functions, social norms, and alliance/federation mechanisms to transition agents from individualistic behaviors toward cooperative strategies. However, these methods typically require all agents to share private local information or to mandatorily participate in federations, which is impractical in real-world applications. To address these issues, this paper proposes a Flexible-Participation Federation (FPF) framework that allows agents to participate in the federation voluntarily. Furthermore, we extend the federation from a global to a Local Multi-Federation (LMF) framework, enabling agents to form multiple localized federations, thereby promoting more efficient and adaptive cooperation. Theoretical evidence demonstrates that the global FPF model, along with the discrepancy between decentralized egoistic policies and federated utilitarian policies, achieves an O(1/T) convergence rate. Agents in the LMF framework also reach consensus within a sublinear gap. Extensive experiments show that agents opting out of federation participation experience a reduction in egoism, and our approach outperforms multiple baselines in terms of both utilitarianism and egalitarianism.
Shaokang Dong, Shangdong Yang, Hongye Cao, Wanqi Yang, Yang Gao 0001
AAAI6
2025 Enhancing Few-Shot Class-Incremental Learning via Training-Free Bi-Level Modality Calibration
abstract
Few-shot Class-Incremental Learning (FSCIL) challenges models to adapt to new classes with limited samples, presenting greater difficulties than traditional class-incremental learning. While existing approaches rely heavily on visual models and require additional training during base or incremental phases, we propose a training-free framework that leverages pre-trained visual-language models like CLIP. At the core of our approach is a novel Bi-level Modality Calibration (BiMC) strategy. Our framework initially performs intra-modal calibration, combining LLM-generated fine-grained category descriptions with visual prototypes from the base session to achieve precise classifier estimation. This is further complemented by inter-modal calibration that fuses pre-trained linguistic knowledge with task-specific visual priors to mitigate modality-specific biases. To enhance prediction robustness, we introduce additional metrics and strategies that maximize the utilization of limited data. Extensive experimental results demonstrate that our approach significantly outperforms existing methods. Code is available at: https://github.com/yychen016/BiMC.
Tianyu Ding, Lei Wang 0001, Jing Huo, Yang Gao 0001, Wenbin Li 0006
CVPR5
2025 Enhancing Trust-Region Bayesian Optimization via Newton Methods
abstract
Bayesian Optimization (BO) has been widely applied to optimize expensive black-box functions while retaining sample efficiency. However, scaling BO to high-dimensional spaces remains challenging. Existing literature proposes performing standard BO in multiple local trust regions (TuRBO) for heterogeneous modeling of the objective function and avoiding over-exploration. Despite its advantages, using local Gaussian Processes (GPs) reduces sampling efficiency compared to a global GP. To enhance sampling efficiency while preserving heterogeneous modeling, we propose to construct multiple local quadratic models using gradients and Hessians from a global GP, and select new sample points by solving the bound-constrained quadratic program. Additionally, we address the issue of vanishing gradients of GPs in high-dimensional spaces. We provide a convergence analysis and demonstrate through experimental results that our method enhances the efficacy of TuRBO and outperforms a wide range of high-dimensional BO techniques on synthetic functions and real-world applications.
Quanlin Chen, Jing Huo, Tianyu Ding, Yang Gao 0001, Yuetong Chen
ECAI5
2025 A Self-Evolving Framework for Multi-Agent Medical Consultation Based on Large Language Models
abstract
We propose a multi-agent approach (SeM-Agents) based on large language models for medical consultations. This framework incorporates various doctor roles and auxiliary roles, with agents communicating through natural language. Using a residual structure, the system conducts multi-round medical consultations based on the patient’s treatment background and symptoms. In the final summary and output stage of the consultation, it utilizes two experience databases—the Correct Consultation Experience Database and the Chain of Thought (CoT) Experience Database—which evolve with accumulated experience during consultations. This evolution drives the framework’s self-improvement, significantly enhancing the rationality and accuracy of the consultations. To ensure that the conclusions are safe, reliable, and aligned with human values, the final decisions undergo a safety review before being provided to the patient. This framework achieved accuracy rates of 89.2% and 83.1% on the MedQA and PubMedQA datasets, respectively.
Kai Chen 0026, Jing Huo, Pinzhuo Tian, Yang Gao 0001
ICASSP7
2025 Adapting In-Domain Few-Shot Segmentation to New Domains Without Source Domain Retraining
abstract
Cross-domain few-shot segmentation (CD-FSS) aims to segment objects of novel classes in new domains, which is often challenging due to the diverse characteristics of target domains and the limited availability of support data. Most CD-FSS methods redesign and retrain in-domain FSS models using abundant base data from the source domain, which are effective but costly to train. To address these issues, we propose adapting informative model structures of the well-trained FSS model for target domains by learning domain characteristics from few-shot labeled support samples during inference, thereby eliminating the need for source domain retraining. Specifically, we first adaptively identify domain-specific model structures by measuring parameter importance using a novel structure Fisher score in a data-dependent manner. Then, we progressively train the selected informative model structures with hierarchically constructed training samples, progressing from fewer to more support shots. The resulting Informative Structure Adaptation (ISA) method effectively addresses domain shifts and equips existing well-trained in-domain FSS models with flexible adaptation capabilities for new domains, eliminating the need to redesign or retrain CD-FSS models on base data. Extensive experiments validate the effectiveness of our method, demonstrating superior performance across multiple CD-FSS benchmarks. Codes are at https://github.com/fanq15/ISA.
Nian Liu 0002, Hisham Cholakkal, Rao Muhammad Anwer, Wenbin Li 0006, Yang Gao 0001
ICCV7
2025 Dynamic Multi-Layer Null Space Projection for Vision-Language Continual Learning
Borui Kang, Lei Wang 0001, Zhiping Wu, Yawen Li 0001, Yang Gao 0001, Wenbin Li 0006
ICCV6
2025 Towards Empowerment Gain through Causal Structure Learning in Model-Based Reinforcement Learning
abstract
In Model-Based Reinforcement Learning (MBRL), incorporating causal structures into dynamics models provides agents with a structured understanding of the environments, enabling efficient decision. Empowerment as an intrinsic motivation enhances the ability of agents to actively control their environments by maximizing the mutual information between future states and actions. We posit that empowerment coupled with causal understanding can improve controllability, while enhanced empowerment gain can further facilitate causal reasoning in MBRL. To improve learning efficiency and controllability, we propose a novel framework, Empowerment through Causal Learning (ECL), where an agent with the awareness of causal dynamics models achieves empowerment-driven exploration and optimizes its causal structure for task learning. Specifically, ECL operates by first training a causal dynamics model of the environment based on collected data. We then maximize empowerment under the causal structure for exploration, simultaneously using data gathered through exploration to update causal dynamics model to be more controllable than dense dynamics model without causal structure. In downstream task learning, an intrinsic curiosity reward is included to balance the causality, mitigating overfitting. Importantly, ECL is method-agnostic and is capable of integrating various causal discovery methods. We evaluate ECL combined with $3$ causal discovery methods across $6$ environments including pixel-based tasks, demonstrating its superior performance compared to other causal MBRL methods, in terms of causal discovery, sample efficiency, and asymptotic performance.
Hongye Cao, Shaokang Dong, Tianpei Yang, Jing Huo, Yang Gao 0001
ICLR7
2025 Causal Information Prioritization for Efficient Reinforcement Learning
abstract
Current Reinforcement Learning (RL) methods often suffer from sample-inefficiency, resulting from blind exploration strategies that neglect causal relationships among states, actions, and rewards. Although recent causal approaches aim to address this problem, they lack grounded modeling of reward-guided causal understanding of states and actions for goal-orientation, thus impairing learning efficiency. To tackle this issue, we propose a novel method named Causal Information Prioritization (CIP) that improves sample efficiency by leveraging factored MDPs to infer causal relationships between different dimensions of states and actions with respect to rewards, enabling the prioritization of causal information. Specifically, CIP identifies and leverages causal relationships between states and rewards to execute counterfactual data augmentation to prioritize high-impact state features under the causal understanding of the environments. Moreover, CIP integrates a causality-aware empowerment learning objective, which significantly enhances the agent's execution of reward-guided actions for more efficient exploration in complex environments. To fully assess the effectiveness of CIP, we conduct extensive experiments across $39$ tasks in $5$ diverse continuous control environments, encompassing both locomotion and manipulation skills learning with pixel-based and sparse reward settings. Experimental results demonstrate that CIP consistently outperforms existing RL methods across a wide range of scenarios.
Hongye Cao, Tianpei Yang, Jing Huo, Yang Gao 0001
ICLR5
2025 GravMAD: Grounded Spatial Value Maps Guided Action Diffusion for Generalized 3D Manipulation
abstract
Robots' ability to follow language instructions and execute diverse 3D manipulation tasks is vital in robot learning. Traditional imitation learning-based methods perform well on seen tasks but struggle with novel, unseen ones due to variability. Recent approaches leverage large foundation models to assist in understanding novel tasks, thereby mitigating this issue. However, these methods lack a task-specific learning process, which is essential for an accurate understanding of 3D environments, often leading to execution failures. In this paper, we introduce GravMAD, a sub-goal-driven, language-conditioned action diffusion framework that combines the strengths of imitation learning and foundation models. Our approach breaks tasks into sub-goals based on language instructions, allowing auxiliary guidance during both training and inference. During training, we introduce Sub-goal Keypose Discovery to identify key sub-goals from demonstrations. Inference differs from training, as there are no demonstrations available, so we use pre-trained foundation models to bridge the gap and identify sub-goals for the current task. In both phases, GravMaps are generated from sub-goals, providing GravMAD with more flexible 3D spatial guidance compared to fixed 3D positions. Empirical evaluations on RLBench show that GravMAD significantly outperforms state-of-the-art methods, with a 28.63\% improvement on novel tasks and a 13.36\% gain on tasks encountered during training. Evaluations on real-world robotic tasks further show that GravMAD can reason about real-world tasks, associate them with relevant visual information, and generalize to novel tasks. These results demonstrate GravMAD's strong multi-task learning and generalization in 3D manipulation. Video demonstrations are available at: https://gravmad.github.io.
Yangtao Chen, Junhui Yin, Jing Huo, Pinzhuo Tian, Jieqi Shi, Yang Gao 0001
ICLR7
2025 Leveraging Flatness to Improve Information-Theoretic Generalization Bounds for SGD
abstract
Information-theoretic (IT) generalization bounds have been used to study the generalization of learning algorithms. These bounds are intrinsically data- and algorithm-dependent so that one can exploit the properties of data and algorithm to derive tighter bounds. However, we observe that although the flatness bias is crucial for SGD’s generalization, these bounds fail to capture the improved generalization under better flatness and are also numerically loose. This is caused by the inadequate leverage of SGD's flatness bias in existing IT bounds. This paper derives a more flatness-leveraging IT bound for the flatness-favoring SGD. The bound indicates the learned models generalize better if the large-variance directions of the final weight covariance have small local curvatures in the loss landscape. Experiments on deep neural networks show our bound not only correctly reflects the better generalization when flatness is improved, but is also numerically much tighter. This is achieved by a flexible technique called "omniscient trajectory". When applied to Gradient Descent’s minimax excess risk on convex-Lipschitz-Bounded problems, it improves representative IT bounds’ $\Omega(1)$ rates to $O(1/\sqrt{n})$. It also implies a by-pass of memorization-generalization trade-offs. Codes are available at [https://github.com/peng-ze/omniscient-bounds](https://github.com/peng-ze/omniscient-bounds).
Ze Peng 0001, Jian Zhang 0090, Yisen Wang 0001, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
ICLR6
2025 Advancing Safe Language Generation: Exploring Alternative Constrained RLHF
abstract
As Large Language Models (LLMs) have been increasingly deployed for various language generation tasks, ensuring the safety and appropriateness of generated content has become a critical challenge. While these models excel at tasks such as dialogue generation, text completion, and content creation, they may inadvertently generate harmful or inappropriate responses that compromise user safety and model reliability. Recent advancements have attempted to address this problem by incorporating safety constraints within the Reinforcement Learning with Human Feedback (RLHF) framework. However, these approaches struggle with the complexity of selecting appropriate compromise parameters, leading to suboptimal performance in safe language generation. To address this problem, we explore advanced alternative first-order constrained optimization strategies that dynamically adjust the balance of helpfulness and harmlessness during training. Specifically, we choose FOCOPS and P3O, which have demonstrated strong performance in various reinforcement learning tasks, and integrate them into the RLHF framework to balance the quality and safety of generated responses. Then we evaluate the performance of different constrained optimization methods for language generation. The results indicate that the P3O algorithm shows significant potential to advance safe language generation while maintaining the helpfulness of model responses.
Zhixin Bai, Yanming Wang, Jing Huo, Yang Gao 0001
ICME7
2025 Reducing Variance of Stochastic Optimization for Approximating Nash Equilibria in Normal-Form Games
abstract
Nash equilibrium (NE) plays an important role in game theory. How to efficiently compute an NE in NFGs is challenging due to its complexity and non-convex optimization property. Machine Learning (ML), the cornerstone of modern artificial intelligence, has demonstrated remarkable empirical performance across various applications including non-convex optimization. To leverage non-convex stochastic optimization techniques from ML for approximating an NE, various loss functions have been proposed. Among these, only one loss function is unbiased, allowing for unbiased estimation under the sampled play. Unfortunately, this loss function suffers from high variance, which degrades the convergence rate. To improve the convergence rate by mitigating the high variance associated with the existing unbiased loss function, we propose a novel surrogate loss function named Nash Advantage Loss (NAL). NAL is theoretically proved unbiased and exhibits significantly lower variance than the existing unbiased loss function. Experimental results demonstrate that the algorithm minimizing NAL achieves a significantly faster empirical convergence rates compared to other algorithms, while also reducing the variance of estimated loss value by several orders of magnitude.
Linjian Meng, Wubing Chen, Wenbin Li 0006, Tianpei Yang, Youzhi Zhang 0001, Yang Gao 0001
ICML6
2025 HDBO-B: On Benchmarking High-Dimensional Bayesian Optimization
abstract
Bayesian optimization (BO) has been extensively studied and applied as a sample-efficient, black-box optimization method in neural network learning, particularly for hyperparameter optimization. However, high-dimensional Bayesian optimization (HDBO) remains a challenging research direction. While existing BO benchmarks focus on practical low-dimensional tasks and specific high-dimensional settings, current HDBO studies often rely on custom experimental settings, resulting in a lack of standardized benchmarks for comprehensive evaluation. To address this, we introduce the first standardized HDBO benchmark HDBO-B. HDBO-B encompasses diverse high-dimensional optimization tasks, synthetic functions tailored for common research settings, and a robust framework with standardized testing, method implementation, and extensive examples. By benchmarking nine representative optimization methods, we demonstrate HDBO-B’s validity and highlight new research directions in neural network methods. Codes are available at: https://github.com/Yiyuiii/HDBO-B.
Hongye Cao, Quanlin Chen, Jing Huo, Dong Li 0016, Yang Gao 0001
IJCNN7
2025 Towards Perfection: Building Inter-component Mutual Correction for Retinex-based Low-light Image Enhancement
abstract
In low-light image enhancement, Retinex-based deep learning methods have garnered significant attention due to their exceptional interpretability. These methods decompose images into mutually independent illumination and reflectance components, allows each component to be enhanced separately. In fact, achieving perfect decomposition of illumination and reflectance components proves to be quite challenging, with some residuals still existing after decomposition. In this paper, we formally name these residuals as inter-component residuals (ICR), which has been largely underestimated by previous methods. In our investigation, ICR not only affects the accuracy of the decomposition but also causes enhanced components to deviate from the ideal outcome, ultimately reducing the final synthesized image quality. To address this issue, we propose a novel Inter-correction Retinex model (IRetinex) to alleviate ICR during the decomposition and enhancement stage. In the decomposition stage, we leverage inter-component residual reduction module to reduce the feature similarity between illumination and reflectance components. In the enhancement stage, we utilize the feature similarity between the two components to detect and mitigate the impact of ICR within each enhancement unit. Extensive experiments on three low-light benchmark datasets demonstrated that by reducing ICR, our method outperforms state-of-the-art approaches both qualitatively and quantitatively. Our code is available at: https://github.com/caoluyang0830/IRetinex.git.
Luyang Cao, Han Xu 0001, Jian Zhang 0090, Lei Qi 0001, Jiayi Ma 0001, Yinghuan Shi, Yang Gao 0001
ACM Multimedia7
2025 In-Context Fully Decentralized Cooperative Multi-Agent Reinforcement Learning
abstract
In this paper, we consider fully decentralized cooperative multi-agent reinforcement learning, where each agent has access only to the states, its local actions, and the shared rewards. The absence of information about other agents' actions typically leads to the non-stationarity problem during per-agent value function updates, and the relative overgeneralization issue during value function estimation. However, existing works fail to address both issues simultaneously, as they lack the capability to model the agents' joint policy in a fully decentralized setting. To overcome this limitation, we propose a simple yet effective method named Return-Aware Context (RAC). RAC formalizes the dynamically changing task, as locally perceived by each agent, as a contextual Markov Decision Process (MDP), and addresses both non-stationarity and relative overgeneralization through return-aware context modeling. Specifically, the contextual MDP attributes the non-stationary local dynamics of each agent to switches between contexts, each corresponding to a distinct joint policy. Then, based on the assumption that the joint policy changes only between episodes, RAC distinguishes different joint policies by the training episodic return and constructs contexts using discretized episodic return values. Accordingly, RAC learns a context-based value function for each agent to address the non-stationarity issue during value function updates. For value function estimation, an individual optimistic marginal value is constructed to encourage the selection of optimal joint actions, thereby mitigating the relative overgeneralization problem. Experimentally, we evaluate RAC on various cooperative tasks (including matrix game, predator and prey, and SMAC), and its significant performance validates its effectiveness.
Bing-Kun Bao, Yang Gao 0001
NeurIPS3
2025 DON'T NEED RETRAINING: A Mixture of DETR and Vision Foundation Models for Cross-Domain Few-Shot Object Detection
abstract
Cross-Domain Few-Shot Object Detection (CD-FSOD) aims to generalize to unseen domains by leveraging a few annotated samples of the target domain, requiring models to exhibit both strong generalization and localization capabilities. However, existing well-trained detectors typically have strong localization capabilities but lack generalization, whereas vision foundation models (VFMs) generally exhibit better generalization but lack accurate localization capabilities. In this paper, we propose a novel Mixture-of-Experts (MoE) structure that integrates the detector's localization capability and the VFM's generalization by using VFM features to improve detector features. Specifically, we propose Expert-wise Router (ER) that selects the most relevant VFM experts for each backbone layer, and Region-wise Router (RR) that emphasizes foreground and suppress background. To bridge representation gaps, we further propose Shared Expert Projection (SEP) module and Private Expert Projection (PEP) module, which align VFM features to the detector feature space while decoupling shared image feature from private image feature in the VFM feature map. Finally, we propose MoE module to transfer the VFM’s generalization to the detector without altering the detector original architecture. Furthermore, our method extend well-trained detectors for detecting novel classes in unseen domains without re-training on the base classes. Experimental results on multiple cross-domain datasets validate the effectiveness of our method.
Changhan Liu, Xunzhi Xiang, Zixuan Duan, Yang Gao 0001
NeurIPS6
2025 Efficient Last-Iterate Convergence in Solving Extensive-Form Games
abstract
To establish last-iterate convergence for Counterfactual Regret Minimization (CFR) algorithms in learning a Nash equilibrium (NE) of extensive-form games (EFGs), recent studies reformulate learning an NE of the original EFG as learning the NEs of a sequence of (perturbed) regularized EFGs. Hence, proving last-iterate convergence in solving the original EFG reduces to proving last-iterate convergence in solving (perturbed) regularized EFGs. However, these studies only establish last-iterate convergence for Online Mirror Descent (OMD)-based CFR algorithms instead of Regret Matching (RM)-based CFR algorithms in solving perturbed regularized EFGs, resulting in a poor empirical convergence rate, as RM-based CFR algorithms typically outperform OMD-based CFR algorithms. In addition, as solving multiple perturbed regularized EFGs is required, fine-tuning across multiple perturbed regularized EFGs is infeasible, making parameter-free algorithms highly desirable. This paper show that CFR$^+$, a classical parameter-free RM-based CFR algorithm, achieves last-iterate convergence in learning an NE of perturbed regularized EFGs. This is the first parameter-free last-iterate convergence for RM-based CFR algorithms in perturbed regularized EFGs. Leveraging CFR$^+$ to solve perturbed regularized EFGs, we get Reward Transformation CFR$^+$ (RTCFR$^+$). Importantly, we extend prior work on the parameter-free property of CFR$^+$, enhancing its stability, which is vital for the empirical convergence of RTCFR$^+$. Experiments show that RTCFR$^+$ exhibits a significantly faster empirical convergence rate than existing algorithms that achieve theoretical last-iterate convergence. Interestingly, RTCFR$^+$ show performance no worse than average-iterate convergence CFR algorithms. It is the first last-iterate convergence algorithm to achieve such performance. Our code is available at https://github.com/menglinjian/NeurIPS-2025-RTCFR.
Linjian Meng, Tianpei Yang, Youzhi Zhang 0001, Zhenxing Ge, Shangdong Yang, Tianyu Ding, Wenbin Li 0006, Bo An 0001, Yang Gao 0001
NeurIPS9
2025 Last-Iterate Convergence of Smooth Regret Matching$^+$ Variants in Learning Nash Equilibria
abstract
Regret Matching$^+$ (RM$^+$) variants are widely used to build superhuman Poker AIs, yet few studies investigate their last-iterate convergence in learning a Nash equilibrium (NE). Although their last-iterate convergence is established for games satisfying the Minty Variational Inequality (MVI), no studies have demonstrated that these algorithms achieve such convergence in the broader class of games satisfying the weak MVI. A key challenge in proving last-iterate convergence for RM$^+$ variants in games satisfying the weak MVI is that even if the game's loss gradient satisfies the weak MVI, RM$^+$ variants operate on a transformed loss feedback which does not satisfy the weak MVI. To provide last-iterate convergence for RM$^+$ variants, we introduce a concise yet novel proof paradigm that involves: (i) transforming an RM$^+$ variant into an Online Mirror Descent (OMD) instance that updates within the original strategy space of the game to recover the weak MVI, and (ii) showing last-iterate convergence by proving the distance between accumulated regrets converges to zero via the recovered weak MVI of the feedback. Inspired by our proof paradigm, we propose Smooth Optimistic Gradient Based RM$^+$ (SOGRM$^+$) and show that it achieves last-iterate and finite-time best-iterate convergence in learning an NE of games satisfying the weak MVI, the weakest condition among all known RM$^+$ variants. Experiments show that SOGRM$^+$ significantly outperforms other algorithms. Our code is available at https://github.com/menglinjian/NeurIPS-2025-SOGRM.
Linjian Meng, Youzhi Zhang 0001, Zhenxing Ge, Tianyu Ding, Shangdong Yang, Wenbin Li 0006, Yang Gao 0001
NeurIPS8
2025 Association-Focused Path Aggregation for Graph Fraud Detection
abstract
Fraudulent activities have caused substantial negative social impacts and are exhibiting emerging characteristics such as intelligence and industrialization, posing challenges of high-order interactions, intricate dependencies, and the sparse yet concealed nature of fraudulent entities. Existing graph fraud detectors are limited by their narrow "receptive fields", as they focus only on the relations between an entity and its neighbors while neglecting longer-range structural associations hidden between entities. To address this issue, we propose a novel fraud detector based on Graph Path Aggregation (GPA). It operates through variable-length path sampling, semantic-associated path encoding, path interaction and aggregation, and aggregation-enhanced fraud detection. To further facilitate interpretable association analysis, we synthesize G-Internet, the first benchmark dataset in the field of internet fraud detection. Extensive experiments across datasets in multiple fraud scenarios demonstrate that the proposed GPA outperforms mainstream fraud detectors by up to +15% in Average Precision (AP). Additionally, GPA exhibits enhanced robustness to noisy labels and provides excellent interpretability by uncovering implicit fraudulent patterns across broader contexts. Code is available at https://github.com/horrible-dong/GPA.
Zunlei Feng, Jie Lei 0002, Mingli Song, Yang Gao 0001
NeurIPS8
2025 Multi-Agent Reinforcement Learning with Communication-Constrained Priors
abstract
Communication is one of the effective means to improve the learning of cooperative policy in multi-agent systems. However, in most real-world scenarios, lossy communication is a prevalent issue. Existing multi-agent reinforcement learning with communication, due to their limited scalability and robustness, struggles to apply to complex and dynamic real-world environments. To address these challenges, we propose a generalized communication-constrained model to uniformly characterize communication conditions across different scenarios. Based on this, we utilize it as a learning prior to distinguish between lossy and lossless messages for specific scenarios. Additionally, we decouple the impact of lossy and lossless messages on distributed decision-making, drawing on a dual mutual information estimatior, and introduce a communication-constrained multi-agent reinforcement learning framework, quantifying the impact of communication messages into the global reward. Finally, we validate the effectiveness of our approach across several communication-constrained benchmarks.
Guang Yang 0066, Tianpei Yang, Jingwen Qiao, Yanqing Wu, Jing Huo, Xingguo Chen, Yang Gao 0001
NeurIPS7
2025 REST: A resolution preserving network for photorealistic style transfer via semantic distillation
Jing Huo, Zheng Gu 0001, Jiulin Zhang, Xiangde Liu, Shiyin Jin, Pinzhuo Tian, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001
Comput. Vis. Image Underst.10
2025 Semi-supervised cross-modality person re-identification based on pseudo label learning
Fei Wu 0004, Ruixuan Zhou, Yang Gao 0001, Yujian Feng, Qinghua Huang, Xiaoyuan Jing
Image Vis. Comput.3
2025 Coordinating Multi-Agent Reinforcement Learning via Dual Collaborative Constraints
Shaokang Dong, Shangdong Yang, Yujing Hu, Wenbin Li 0006, Yang Gao 0001
Neural Networks6
2025 MIXRTs: Toward Interpretable Multi-Agent Reinforcement Learning via Mixing Recurrent Soft Decision Trees
abstract
While achieving tremendous success in various fields, existing multi-agent reinforcement learning (MARL) with a black-box neural network makes decisions in an opaque manner that hinders humans from understanding the learned knowledge and how input observations influence decisions. In contrast, existing interpretable approaches usually suffer from weak expressivity and low performance. To bridge this gap, we propose MIXing Recurrent soft decision Trees (MIXRTs), a novel interpretable architecture that can represent explicit decision processes via the root-to-leaf path and reflect each agent's contribution to the team. Specifically, we construct a novel soft decision tree using a recurrent structure and demonstrate which features influence the decision-making process. Then, based on the value decomposition framework, we linearly assign credit to each agent by explicitly mixing individual action values to estimate the joint action value using only local observations, providing new insights into interpreting the cooperation mechanism. Theoretical analysis confirms that MIXRTs guarantee additivity and monotonicity in the factorization of joint action values. Evaluations on complex tasks like Spread and StarCraft II demonstrate that MIXRTs compete with existing methods while providing clear explanations, paving the way for interpretable and high-performing MARL systems.
Zichuan Liu, Yuanyang Zhu, Zhi Wang 0001, Yang Gao 0001, Chunlin Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 ONNXPruner: ONNX-Based General Model Pruning Adapter
abstract
Recent advancements in model pruning have focused on developing new algorithms and improving upon benchmarks. However, the practical application of these algorithms across various models and platforms remains a significant challenge. To address this challenge, we propose ONNXPruner, a versatile pruning adapter designed for the ONNX format models. ONNXPruner streamlines the adaptation process across diverse deep learning frameworks and hardware platforms. A novel aspect of ONNXPruner is its use of node association trees, which automatically adapt to various model architectures. These trees clarify the structural relationships between nodes, guiding the pruning process, particularly highlighting the impact on interconnected nodes. Furthermore, we introduce a tree-level evaluation method. By leveraging node association trees, this method allows for a comprehensive analysis beyond traditional single-node evaluations, enhancing pruning performance without the need for extra operations. Experiments across multiple models and datasets confirm ONNXPruner's strong adaptability and increased efficacy. Our work aims to advance the practical application of model pruning.
Dongdong Ren, Wenbin Li 0006, Tianyu Ding, Lei Wang 0001, Jing Huo, Hongbing Pan, Yang Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 3D Gaussian Splatting: Survey, Technologies, Challenges, and Opportunities
abstract
3D Gaussian Splatting (3DGS) has emerged as a prominent technique with the potential to become a mainstream method for 3D representations. It can effectively transform multi-view images into explicit 3D Gaussian through efficient training, and achieve real-time rendering of novel views. This survey aims to analyze existing 3DGS-related works from multiple intersecting perspectives, including related tasks, technologies, challenges, and opportunities. The primary objective is to provide newcomers with a rapid understanding of the field and to assist researchers in methodically organizing existing technologies and challenges. Specifically, we delve into the optimization, application, and extension of 3DGS, categorizing them based on their focuses or motivations. Additionally, we summarize and classify nine types of technical modules and corresponding improvements identified in existing works. Based on these analyses, we further examine the common challenges and technologies across various tasks, proposing potential research opportunities.
Yanqi Bao, Tianyu Ding, Jing Huo, Yaoli Liu, Wenbin Li 0006, Yang Gao 0001, Jiebo Luo 0001
IEEE Trans. Circuits Syst. Video Technol.7
2025 An Adaptor for Triggering Semi-Supervised Learning to Out-of-Box Serve Deep Image Clustering
abstract
Recently, some works integrate SSL techniques into deep clustering frameworks to enhance image clustering performance. However, they all need pretraining, clustering learning, or a trained clustering model as prerequisites, limiting the flexible and out-of-box application of SSL learners in the image clustering task. This work introduces ASD, an adaptor that enables the cold-start of SSL learners for deep image clustering without any prerequisites. Specifically, we first randomly sample pseudo-labeled data from all unlabeled data, and set an instance-level classifier to learn them with semantically aligned instance-level labels. With the ability of instance-level classification, we track the class transitions of predictions on unlabeled data to extract high-level similarities of instance-level classes, which can be utilized to assign cluster-level labels to pseudo-labeled data. Finally, we use the pseudo-labeled data with assigned cluster-level labels to trigger a general SSL learner trained on the unlabeled data for image clustering. We show the superior performance of ASD across various benchmarks against the latest deep image clustering approaches and very slight accuracy gaps compared to SSL methods using ground-truth, e.g., only 1.33% on CIFAR-10. Moreover, ASD can also further boost the performance of existing SSL-embedded deep image clustering methods.
Yue Duan, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Image Process.4
2025 Leveraging Frequency Analysis for Image Denoising Network Pruning
abstract
As a common model compression technique, network pruning is widely used to reduce storage and computational cost of deep models in the resource-constrained regime. However, most current pruning methods are designed for high-level vision tasks, with few developed for low-level vision tasks. We observed that the norm-based pruning criterion, originally designed for high-level vision tasks, is highly unsuitable for low-level image denoising networks. This difference arises because image denoising networks pursue distinct feature granularities and goals compared to typical high-level vision tasks. To address this issue, we propose a novel filter evaluation method, termed High-Frequency Components Pruning (HFCP), specifically tailored for image denoising network pruning. HFCP assesses filter importance based on high-frequency components. To the best of our knowledge, this is the first pruning method designed specifically for image denoising tasks, straightforward and applicable to various types of noise. Furthermore, HFCP enhances the pruned model's high-frequency information content with high reliability and interpretability. This facilitates the network's ability to distinguish high-frequency signals from noise. We comprehensively analyzed multiple image denoising networks and validated HFCP's effectiveness across four mainstream networks.
Dongdong Ren, Wenbin Li 0006, Jing Huo, Lei Wang 0001, Hongbing Pan, Yang Gao 0001
IEEE Trans. Image Process.6
2025 E2MPL: An Enduring and Efficient Meta Prompt Learning Framework for Few-Shot Unsupervised Domain Adaptation
abstract
Few-shot unsupervised domain adaptation (FS-UDA) leverages a limited amount of labeled data from a source domain to enable accurate classification in an unlabeled target domain. Despite recent advancements, current approaches of FS-UDA continue to confront a major challenge: models often demonstrate instability when adapted to new FS-UDA tasks and necessitate considerable time investment. To address these challenges, we put forward a novel framework called Enduring and Efficient Meta-Prompt Learning (E2MPL) for FS-UDA. Within this framework, we utilize the pre-trained CLIP model as the backbone of feature learning. Firstly, we design domain-shared prompts, consisting of virtual tokens, which primarily capture meta-knowledge from a wide range of meta-tasks to mitigate the domain gaps. Secondly, we develop a task prompt learning network that adaptively learns task-specific prompts with the goal of achieving fast and stable task generalization. Thirdly, we formulate the meta-prompt learning process as a bilevel optimization problem, consisting of (outer) meta-prompt learner and (inner) task-specific classifier and domain adapter. Also, the inner objective of each meta-task has the closed-form solution, which enables efficient prompt learning and adaptation to new tasks in a single step. Extensive experimental studies demonstrate the promising performance of our framework in a domain adaptation benchmark dataset DomainNet. Compared with state-of-the-art methods, our approach has improved the average accuracy by at least 15 percentage points and reduces the average time by 64.67% in the 5-way 1-shot task; in the 5-way 5-shot task, it achieves at least a 9-percentage-point improvement in average accuracy and reduces the average time by 63.18%. Moreover, our method exhibits more enduring and stable performance than the other methods, i.e., reducing the average IQR value by over 40.80% and 25.35% in the 5-way 1-shot and 5-shot task, respectively.
Wanqi Yang, Lei Wang 0001, Ming Yang 0014, Yang Gao 0001
IEEE Trans. Image Process.7
2025 Label Distribution Guided Hashing for Cross-Modal Retrieval
abstract
Hashing methods have recently attracted extensive attention in cross-modal retrieval. Most supervised hashing methods attempt to preserve the semantic information into hash codes by leveraging the original logical label matrix. However, they generally treat all labels equally, and ignore the relative significance of different labels due to the variety of data features. In this article, we argue that exploring the relative importance of labels benefits the enhancement of semantic information, and we propose a novel LAbel Distribution Guided Hashing (LADH) method for cross-modal retrieval. In particular, LADH first learns a feature-induced label distribution for each sample to weigh different labels, which leverages the multi-modal feature information to enrich the semantic label information. By jointly using the learned label distributions and multi-modal features, the latent representation and hash codes are obtained with multi-modal feature selection and enhanced semantic similarities embedded. An efficient algorithm is designed to solve the proposed method whose time complexity is linear to the number of the training instances. Experimental results on several public benchmark datasets verify the effectiveness and efficiency of our method compared with the state-of-the-art methods.
Fatang Lei, Chao Zhang 0078, Huaxiong Li, Yang Gao 0001, Chunlin Chen 0001
ACM Trans. Knowl. Discov. Data4
2025 Mamba-Sea: A Mamba-Based Framework With Global-to-Local Sequence Augmentation for Generalizable Medical Image Segmentation
abstract
To segment medical images with distribution shifts, domain generalization (DG) has emerged as a promising setting to train models on source domains that can generalize to unseen target domains. Existing DG methods are mainly based on CNN or ViT architectures. Recently, advanced state space models, represented by Mamba, have shown promising results in various supervised medical image segmentation. The success of Mamba is primarily owing to its ability to capture long-range dependencies while keeping linear complexity with input sequence length, making it a promising alternative to CNNs and ViTs. Inspired by the success, in the paper, we explore the potential of the Mamba architecture to address distribution shifts in DG for medical image segmentation. Specifically, we propose a novel Mamba-based framework, Mamba-Sea, incorporating global-to-local sequence augmentation to improve the model's generalizability under domain shift issues. Our Mamba-Sea introduces a global augmentation mechanism designed to simulate potential variations in appearance across different sites, aiming to suppress the model's learning of domain-specific information. At the local level, we propose a sequence-wise augmentation along input sequences, which perturbs the style of tokens within random continuous sub-sequences by modeling and resampling style statistics associated with domain shifts. To our best knowledge, Mamba-Sea is the first work to explore the generalization of Mamba for medical image segmentation, providing an advanced and promising Mamba-based architecture with strong robustness to domain shifts. Remarkably, our proposed method is the first to surpass a Dice coefficient of 90% on the Prostate dataset, which exceeds previous SOTA of 88.61%. The code is available at https://github.com/orange-czh/Mamba-Sea.
Zihan Cheng 0001, Jintao Guo, Jian Zhang 0090, Lei Qi 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Medical Imaging7
2025 Stitching, Fine-Tuning, and Re-Training: A SAM-Enabled Framework for Semi-Supervised 3D Medical Image Segmentation
abstract
Segment Anything Model (SAM) fine-tuning has shown remarkable performance in medical image segmentation in a fully supervised manner, but requires precise annotations. To reduce the annotation cost and maintain satisfactory performance, in this work, we leverage the capabilities of SAM for establishing semi-supervised medical image segmentation models. Rethinking the requirements of effectiveness, efficiency, and compatibility, we propose a three-stage framework, i.e., Stitching, Fine-tuning, and Re-training (SFR). The current fine-tuning approaches mostly involve 2D slice-wise fine-tuning that disregards the contextual information between adjacent slices. Our stitching strategy mitigates the mismatch between natural and 3D medical images. The stitched images are then used for fine-tuning SAM, providing robust initialization of pseudo-labels. Afterwards, we train a 3D semi-supervised segmentation model while maintaining the same parameter size as the conventional segmenter such as V-Net. Our SFR framework is plug-and-play, and easily compatible with various popular semi-supervised methods. We also develop an extended framework SFR+ with selective fine-tuning and re-training through confidence estimation. Extensive experiments validate that our SFR and SFR+ achieve significant improvements in both moderate annotation and scarce annotation across five datasets. In particular, SFR framework improves the Dice score of Mean Teacher from 29.68% to 74.40% with only one labeled data of LA dataset. The code is available at https://github.com/ShumengLI/SFR.
Shumeng Li, Lei Qi 0001, Qian Yu 0007, Jing Huo, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Medical Imaging6
2025 Unleashing the Power of Intermediate Domains for Mixed Domain Semi-Supervised Medical Image Segmentation
abstract
Both limited annotation and domain shift are prevalent challenges in medical image segmentation. Traditional semi-supervised segmentation and unsupervised domain adaptation methods address one of these issues separately. However, the coexistence of limited annotation and domain shift is quite common, which motivates us to introduce a novel and challenging scenario: Mixed Domain Semi-supervised medical image Segmentation (MiDSS), where limited labeled data from a single domain and a large amount of unlabeled data from multiple domains. To tackle this issue, we propose the UST-RUN framework, which fully leverages intermediate domain information to facilitate knowledge transfer. We employ Unified Copy-paste (UCP) to construct intermediate domains, and propose a Symmetric GuiDance training strategy (SymGD) to supervise unlabeled data by merging pseudo-labels from intermediate samples. Subsequently, we introduce a Training Process aware Random Amplitude MixUp (TP-RAM) to progressively incorporate style-transition components into intermediate samples. To generate more diverse intermediate samples, we further select reliable samples with high-quality pseudo-labels, which are then mixed with other unlabeled data. Additionally, we generate sophisticated intermediate samples with high-quality pseudo-labels for unreliable samples, ensuring effective knowledge transfer for them. Extensive experiments on four public datasets demonstrate the superiority of UST-RUN. Notably, UST-RUN achieves a 12.94% improvement in Dice score on the Prostate dataset. Our code is available at https://github.com/MQinghe/UST-RUN.
Qinghe Ma, Jian Zhang 0090, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Medical Imaging6
2025 Balancing Multi-Target Semi-Supervised Medical Image Segmentation With Collaborative Generalist and Specialists
abstract
Despite the promising performance achieved by current semi-supervised models in segmenting individual medical targets, many of these models suffer a notable decrease in performance when tasked with the simultaneous segmentation of multiple targets. A vital factor could be attributed to the imbalanced scales among different targets: during simultaneously segmenting multiple targets, large targets dominate the loss, leading to small targets being misclassified as larger ones. To this end, we propose a novel method, which consists of a Collaborative Generalist and several Specialists, termed CGS. It is centered around the idea of employing a specialist for each target class, thus avoiding the dominance of larger targets. The generalist performs conventional multi-target segmentation, while each specialist is dedicated to distinguishing a specific target class from the remaining target classes and the background. Based on a theoretical insight, we demonstrate that CGS can achieve a more balanced training. Moreover, we develop cross-consistency losses to foster collaborative learning between the generalist and the specialists. Lastly, regarding their intrinsic relation that the target class of any specialized head should belong to the remaining classes of the other heads, we introduce an inter-head error detection module to further enhance the quality of pseudo-labels. Experimental results on three popular benchmarks showcase its superior performance compared to state-of-the-art methods. Our code is available at https://github.com/wangyou0804/CGS.
Zekun Li 0010, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Medical Imaging6
2025 Dictionary Based Generative Adversarial Network for Multi-Collection Style Transfer
abstract
Most collection-based style transfer methods require training a separate model for each individual collection of styles, making the extension to multiple collections of styles less flexible. Besides, the existing collection-based methods are also less flexible in extending to new style collections in a continual manner. To address these issues, we propose a novelMultI-Dictionary Generative Adversarial Network framework (MID-GAN)for multi-collection style transfer. Specifically, we design a multi-dictionary architecture within a GAN, with each dictionary consisting of a set of local style codes for a specific style collection. Benefiting from the local style codes used in the dictionary, a stylization module with aligned skip connections is further proposed, which can better preserve both the local details and the overall image structure. The dictionary design allows a flexible extension to new style collections by readily adding new dictionaries and we propose a continual training strategy that can both preserve the style transfer ability of old styles and achieve good transfer results for newly added styles. Extensive experiments are performed to show that the proposed method is better than existing collection-based style transfer methods. We also demonstrate the proposed method can generate diverse meaningful style transfer results of the same style collection.
Jing Huo, Shiyin Jin, Jiashen Li, Pinzhuo Tian, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001
IEEE Trans. Multim.8
2025 Multi-Task Multi-Agent Reinforcement Learning With Interaction and Task Representations
abstract
Multi-task multi-agent reinforcement learning (MT-MARL) is capable of leveraging useful knowledge across multiple related tasks to improve performance on any single task. While recent studies have tentatively achieved this by learning independent policies on a shared representation space, we pinpoint that further advancements can be realized by explicitly characterizing agent interactions within these multi-agent tasks and identifying task relations for selective reuse. To this end, this article proposes Representing Interactions and Tasks (RIT), a novel MT-MARL algorithm that characterizes both intra-task agent interactions and inter-task task relations. Specifically, for characterizing agent interactions, RIT presents the interactive value decomposition to explicitly take the dependency among agents into policy learning. Theoretical analysis demonstrates that the learned utility value of each agent approximates its Shapley value, thus representing agent interactions. Moreover, we learn task representations based on per-agent local trajectories, which assess task similarities and accordingly identify task relations. As a result, RIT facilitates the effective transfer of interaction knowledge across similar multi-agent tasks. Structurally, RIT develops universal policy structure for scalable multi-task policy learning. We evaluate RIT against multiple state-of-the-art baselines in various cooperative tasks, and its significant performance under both multi-task and zero-shot settings demonstrates its effectiveness.
Shaokang Dong, Shangdong Yang, Yujing Hu, Tianyu Ding, Wenbin Li 0006, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.7
2025 State Abstraction via Deep Supervised Hash Learning
abstract
State abstraction is a widely used technique in reinforcement learning (RL) that compresses the state space to accelerate learning algorithms. However, designing an effective abstraction function in large-scale or high-dimensional state space problems remains a significant challenge. In this brief, we present a novel state abstraction method based on deep supervised hash learning (DSH) and provide a theoretical analysis of its near-optimal property. Furthermore, by leveraging the DSH-based representation as the optimization objective, we propose a direct and concise optimization method based on the target value. In addition, we construct an auxiliary learning task for state abstraction that can be combined with various RL algorithms. In particular, we apply the DSH-based state abstraction to both deep Q-learning (DQN) and soft actor-critic (SAC). Extensive experiments are conducted on Atari and several classic control benchmarks to evaluate the effectiveness of the DSH-based state abstraction method, showing that our method surpasses existing state abstraction algorithms in performance.
Guang Yang 0066, Jing Huo, Shangdong Yang, Tianyu Ding, Xingguo Chen, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.7
2024 PG-LBO: Enhancing High-Dimensional Bayesian Optimization with Pseudo-Label and Gaussian Process Guidance
abstract
Variational Autoencoder based Bayesian Optimization (VAE-BO) has demonstrated its excellent performance in addressing high-dimensional structured optimization problems. However, current mainstream methods overlook the potential of utilizing a pool of unlabeled data to construct the latent space, while only concentrating on designing sophisticated models to leverage the labeled data. Despite their effective usage of labeled data, these methods often require extra network structures, additional procedure, resulting in computational inefficiency. To address this issue, we propose a novel method to effectively utilize unlabeled data with the guidance of labeled data. Specifically, we tailor the pseudo-labeling technique from semi-supervised learning to explicitly reveal the relative magnitudes of optimization objective values hidden within the unlabeled data. Based on this technique, we assign appropriate training weights to unlabeled data to enhance the construction of a discriminative latent space. Furthermore, we treat the VAE encoder and the Gaussian Process (GP) in Bayesian optimization as a unified deep kernel learning process, allowing the direct utilization of labeled data, which we term as Gaussian Process guidance. This directly and effectively integrates the goal of improving GP accuracy into the VAE training, thereby guiding the construction of the latent space. The extensive experiments demonstrate that our proposed method outperforms existing VAE-BO algorithms in various optimization scenarios. Our code will be published at https://github.com/TaicaiChen/PG-LBO.
Taicai Chen, Yue Duan, Dong Li 0016, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
AAAI6
2024 ViT-Calibrator: Decision Stream Calibration for Vision Transformer
abstract
A surge of interest has emerged in utilizing Transformers in diverse vision tasks owing to its formidable performance. However, existing approaches primarily focus on optimizing internal model architecture designs that often entail significant trial and error with high burdens. In this work, we propose a new paradigm dubbed Decision Stream Calibration that boosts the performance of general Vision Transformers. To achieve this, we shed light on the information propagation mechanism in the learning procedure by exploring the correlation between different tokens and the relevance coefficient of multiple dimensions. Upon further analysis, it was discovered that 1) the final decision is associated with tokens of foreground targets, while token features of foreground target will be transmitted into the next layer as much as possible, and the useless token features of background area will be eliminated gradually in the forward propagation. 2) Each category is solely associated with specific sparse dimensions in the tokens. Based on the discoveries mentioned above, we designed a two-stage calibration scheme, namely ViT-Calibrator, including token propagation calibration stage and dimension propagation calibration stage. Extensive experiments on commonly used datasets show that the proposed approach can achieve promising results.
Zhijie Jia, Lechao Cheng, Yang Gao 0001, Jie Lei 0002, Yijun Bei, Zunlei Feng
AAAI4
2024 DGA-GNN: Dynamic Grouping Aggregation GNN for Fraud Detection
abstract
Fraud detection has increasingly become a prominent research field due to the dramatically increased incidents of fraud. The complex connections involving thousands, or even millions of nodes, present challenges for fraud detection tasks. Many researchers have developed various graph-based methods to detect fraud from these intricate graphs. However, those methods neglect two distinct characteristics of the fraud graph: the non-additivity of certain attributes and the distinguishability of grouped messages from neighbor nodes. This paper introduces the Dynamic Grouping Aggregation Graph Neural Network (DGA-GNN) for fraud detection, which addresses these two characteristics by dynamically grouping attribute value ranges and neighbor nodes. In DGA-GNN, we initially propose the decision tree binning encoding to transform non-additive node attributes into bin vectors. This approach aligns well with the GNN’s aggregation operation and avoids nonsensical feature generation. Furthermore, we devise a feedback dynamic grouping strategy to classify graph nodes into two distinct groups and then employ a hierarchical aggregation. This method extracts more discriminative features for fraud detection tasks. Extensive experiments on five datasets suggest that our proposed method achieves a 3% ~ 16% improvement over existing SOTA methods. Code is available at https://github.com/AtwoodDuan/DGA-GNN.
Mingjiang Duan, Tongya Zheng, Yang Gao 0001, Zunlei Feng, Xinyu Wang 0001
AAAI3
2024 Optimistic Value Instructors for Cooperative Multi-Agent Reinforcement Learning
abstract
In cooperative multi-agent reinforcement learning, decentralized agents hold the promise of overcoming the combinatorial explosion of joint action space and enabling greater scalability. However, they are susceptible to a game-theoretic pathology called relative overgeneralization that shadows the optimal joint action. Although recent value-decomposition algorithms guide decentralized agents by learning a factored global action value function, the representational limitation and the inaccurate sampling of optimal joint actions during the learning process make this problem still. To address this limitation, this paper proposes a novel algorithm called Optimistic Value Instructors (OVI). The main idea behind OVI is to introduce multiple optimistic instructors into the value-decomposition paradigm, which are capable of suggesting potentially optimal joint actions and rectifying the factored global action value function to recover these optimal actions. Specifically, the instructors maintain optimistic value estimations of per-agent local actions and thus eliminate the negative effects caused by other agents' exploratory or sub-optimal non-cooperation, enabling accurate identification and suggestion of optimal joint actions. Based on the instructors' suggestions, the paper further presents two instructive constraints to rectify the factored global action value function to recover these optimal joint actions, thus overcoming the RO problem. Experimental evaluation of OVI on various cooperative multi-agent tasks demonstrates its superior performance against multiple baselines, highlighting its effectiveness.
Jianqi Wang, Yujing Hu, Shaokang Dong, Wenbin Li 0012, Tangjie Lv, Changjie Fan, Yang Gao 0001
AAAI9
2024 Angle Robustness Unmanned Aerial Vehicle Navigation in GNSS-Denied Scenarios
abstract
Due to the inability to receive signals from the Global Navigation Satellite System (GNSS) in extreme conditions, achieving accurate and robust navigation for Unmanned Aerial Vehicles (UAVs) is a challenging task. Recently emerged, vision-based navigation has been a promising and feasible alternative to GNSS-based navigation. However, existing vision-based techniques are inadequate in addressing flight deviation caused by environmental disturbances and inaccurate position predictions in practical settings. In this paper, we present a novel angle robustness navigation paradigm to deal with flight deviation in point-to-point navigation tasks. Additionally, we propose a model that includes the Adaptive Feature Enhance Module, Cross-knowledge Attention-guided Module and Robust Task-oriented Head Module to accurately predict direction angles for high-precision navigation. To evaluate the vision-based navigation methods, we collect a new dataset termed as UAV_AR368. Furthermore, we design the Simulation Flight Testing Instrument (SFTI) using Google Earth to simulate different flight environments, thereby reducing the expenses associated with real flight testing. Experiment results demonstrate that the proposed model outperforms the state-of-the-art by achieving improvements of 26.0% and 45.6% in the success rate of arrival under ideal and disturbed circumstances, respectively.
Zunlei Feng, Haofei Zhang, Yang Gao 0001, Jie Lei 0002, Mingli Song
AAAI4
2024 Exploiting Inter-sample and Inter-feature Relations in Dataset Distillation
abstract
Dataset distillation has emerged as a promising approach in deep learning, enabling efficient training with small synthetic datasets derived from larger real ones. Particularly, distribution matching-based distillation methods attract attention thanks to its effectiveness and low computational cost. However, these methods face two primary limitations: the dispersed feature distribution within the same class in synthetic datasets, reducing class discrim-ination, and an exclusive focus on mean feature consistency, lacking precision and comprehensiveness. To address these challenges, we introduce two novel constraints: a class centralization constraint and a covariance matching constraint. The class centralization constraint aims to enhance class discrimination by more closely clustering samples within classes. The covariance matching constraint seeks to achieve more accurate feature distribution matching between real and synthetic datasets through local feature covariance matrices, particularly beneficial when sample sizes are much smaller than the number of features. Experiments demonstrate notable improvements with these constraints, yielding performance boosts of up to 6.6% on CIFAR10, 2.9% on SVHN, 2.5% on CIFAR100, and 2.5% on TinyImageNet, compared to the state-of-the-art relevant methods. In addition, our method maintains robust performance in cross-architecture settings, with a maximum performance drop of 1.7% on four architectures. Code is avail-able at https://github.com/VincenDen/IID.
Wenxiao Deng, Wenbin Li 0006, Tianyu Ding, Lei Wang 0001, Kuihua Huang, Jing Huo, Yang Gao 0001
CVPR8
2024 Constructing and Exploring Intermediate Domains in Mixed Domain Semi-supervised Medical Image Segmentation
abstract
Both limited annotation and domain shift are preva-lent challenges in medical image segmentation. Traditional semi-supervised segmentation and unsupervised do-main adaptation methods address one of these issues sepa-rately. However, the coexistence of limited annotation and domain shift is quite common, which motivates us to in-troduce a novel and challenging scenario: Mixed Domain Semi-supervised medical image Segmentation (MiDSS). In this scenario, we handle data from multiple medical cen-ters, with limited annotations available for a single do-main and a large amount of unlabeled data from multi-ple domains. We found that the key to solving the prob-lem lies in how to generate reliable pseudo labels for the unlabeled data in the presence of domain shift with la-beled data. To tackle this issue, we employ Unified Copy-Paste (UCP) between images to construct intermediate do-mains, facilitating the knowledge transfer from the do-main of labeled data to the domains of unlabeled data. To fully utilize the information within the intermediate do-main, we propose a symmetric Guidance training strategy (SymGD), which additionally offers direct guidance to un-labeled data by merging pseudo labels from intermediate samples. Subsequently, we introduce a Training Process aware Random Amplitude MixUp (TP-RAM) to progres-sively incorporate style-transition components into inter-mediate samples. Compared with existing state-of-the-art approaches, our method achieves a notable 13.57% im-provement in Dice score on Prostate dataset, as demon-strated on three public datasets. Our code is available at https://github.com/MQinghe/MiDSS
Qinghe Ma, Jian Zhang 0090, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001
CVPR6
2024 Learn to Preserve and Diversify: Parameter-Efficient Group with Orthogonal Regularization for Domain Generalization
Jiajun Hu, Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
ECCV (51)5
2024 The Devil Is in the Statistics: Mitigating and Exploiting Statistics Difference for Generalizable Semi-supervised Medical Image Segmentation
Muyang Qiu, Jian Zhang 0090, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001
ECCV (54)6
2024 Multi-Agent Exploration via Self-Learning and Social Learning
abstract
Self-learning and social learning stand as two pivotal constituents in multi-agent exploration. Inspired by the fact that animals and humans explore unfamiliar environments to learn survival skills by training themselves using unlabeled data and replicating others’ successful experiences, we propose a multi-agent reinforcement learning method, named Self-Learning and Social Learning (S2L), which aims to address the complex tasks caused by sparse rewards and intricate sequential structures. Specifically, in Self-Learning, we incorporate both task-specific and task-agnostic intrinsic rewards. These incentives steer individual agents towards exploration and comprehension of the environment. Furthermore, in Social Learning, different independent agents can implicitly share the successful experience by observing others in view and without additional communication or parameter-sharing overhead. Finally, experimental evaluation of S2L on the complex task characterized by sparse rewards and intricate sequential structures demonstrates its superior performance against other competing exploration baselines.
Shaokang Dong, Wubing Chen, Hongye Cao, Yang Gao 0001
ICASSP6
2024 Multi-Agent Sparse Interaction Modeling is an Anomaly Detection Problem
abstract
Most real-world multi-agent tasks exhibit the characteristic of sparse interaction, wherein agents interact with each other in a limited number of crucial states while largely acting independently. Effectively modeling the sparse interaction and leveraging the learned interaction structure to instruct agents’ learning processes can enhance the efficiency of multi-agent reinforcement learning algorithms. However, it remains unclear how to identify these specific interactive states solely through trials and errors within current multi-agent tasks. To address this challenge, this paper introduces a novel algorithm called Sparse Interaction as Anomaly (SIA), which innovatively casts the sparse interaction modeling into an anomaly detection problem. The underlying intuition is that interactive states appear rarely in agents’ trajectories and exhibit distinct dynamics compared to other commonplace states. Building upon this insight, SIA first employs variational inference to model the latent dynamics of agents’ trajectories. It then designates states with anomalous dynamics as the elusive interactive states and subsequently instructs agents to explore these states more extensively. This facilitates the emergence of interactive behaviors and promotes the learning of multi-agent policies. Experimental evaluation of SIA across various multi-agent tasks demonstrates its superior performance against multiple baselines, highlighting its effectiveness.
Shaokang Dong, Shangdong Yang, Hongye Cao, Yang Gao 0001
ICASSP6
2024 Dynamic Replay Training for Class-Incremental Learning
abstract
Replay-based methods for Class-Incremental Learning (CIL) typically employ new classes and a limited subset of old classes stored in memory to facilitate the model training. However, these methods often lead to class imbalance and catastrophic forgetting, where the model forgets previously learned tasks. While some studies have attempted to address such a class imbalance issue, they do not fully consider the dynamic nature of forgetting in the model. In this paper, we propose a novel method called Dynamic Replay Training (DRT) to address the dynamic forgetting of previously learned tasks by the model. DRT replays memory data with dynamically changing frequencies, offering a novel perspective to tackle catastrophic forgetting and class imbalance. The proposed method is evaluated on CIFAR-100 and ImageNet-100 in various settings, showing significant improvements of 8.28% and 4.53% in terms of classification accuracy compared to the baseline method on the two datasets, respectively.
Dongdong Ren, Chenglei Peng, Jing Huo, Wenbin Li 0006, Yang Gao 0001
ICASSP6
2024 InsertNeRF: Instilling Generalizability into NeRF with HyperNet Modules
abstract
Generalizing Neural Radiance Fields (NeRF) to new scenes is a significant challenge that existing approaches struggle to address without extensive modifications to vanilla NeRF framework. We introduce **InsertNeRF**, a method for **INS**tilling g**E**ne**R**alizabili**T**y into **NeRF**. By utilizing multiple plug-and-play HyperNet modules, InsertNeRF dynamically tailors NeRF's weights to specific reference scenes, transforming multi-scale sampling-aware features into scene-specific representations. This novel design allows for more accurate and efficient representations of complex appearances and geometries. Experiments show that this method not only achieves superior generalization performance but also provides a flexible pathway for integration with other NeRF-like systems, even in sparse input settings. Code will be available at: https://github.com/bbbbby-99/InsertNeRF.
Yanqi Bao, Tianyu Ding, Jing Huo, Wenbin Li 0006, Yang Gao 0001
ICLR6
2024 Diving Segmentation Model into Pixels
abstract
More distinguishable and consistent pixel features for each category will benefit the semantic segmentation under various settings. Existing efforts to mine better pixel-level features attempt to explicitly model the categorical distribution, which fails to achieve optimal due to the significant pixel feature variance. Moreover, prior research endeavors have scarcely delved into the thorough analysis and meticulous handling of pixel-level variance, leaving semantic segmentation at a coarse granularity. In this work, we analyze the causes of pixel-level variance and introduce the concept of $\textbf{pixel learning}$ to concentrate on the tailored learning process of pixels to handle the pixel-level variance, enhancing the per-pixel recognition capability of segmentation models. Under the context of the pixel learning scheme, each image is viewed as a distribution of pixels, and pixel learning aims to pursue consistent pixel representation inside an image, continuously align pixels from different images (distributions), and eventually achieve consistent pixel representation for each category, even cross-domains. We proposed a pure pixel-level learning framework, namely PiXL, which consists of a pixel partition module to divide pixels into sub-domains, a prototype generation, a selection module to prepare targets for subsequent alignment, and a pixel alignment module to guarantee pixel feature consistency intra-/inter-images, and inter-domains. Extensive evaluations of multiple learning paradigms, including unsupervised domain adaptation and semi-/fully-supervised segmentation, show that PiXL outperforms state-of-the-art performances, especially when annotated images are scarce. Visualization of the embedding space further demonstrates that pixel learning attains a superior representation of pixel features. The code is available at https://github.com/ChenGan-JS/PiXL.
Chen Gan, Kelei He, Yang Gao 0001
ICLR4
2024 Hybrid Sharing for Multi-Label Image Classification
abstract
Existing multi-label classification methods have long suffered from label heterogeneity, where learning a label obscures another. By modeling multi-label classification as a multi-task problem, this issue can be regarded as a negative transfer, which indicates challenges to achieve simultaneously satisfied performance across multiple tasks. In this work, we propose the Hybrid Sharing Query (HSQ), a transformer-based model that introduces the mixture-of-experts architecture to image multi-label classification. HSQ is designed to leverage label correlations while mitigating heterogeneity effectively. To this end, HSQ is incorporated with a fusion expert framework that enables it to optimally combine the strengths of task-specialized experts with shared experts, ultimately enhancing multi-label classification performance across most labels. Extensive experiments are conducted on two benchmark datasets, with the results demonstrating that the proposed method achieves state-of-the-art performance and yields simultaneous improvements across most labels. The code is available at https://github.com/zihao-yin/HSQ
Chen Gan, Kelei He, Yang Gao 0001
ICLR4
2024 Safe and Robust Subgame Exploitation in Imperfect Information Games
abstract
Opponent exploitation is an important task for players to exploit the weaknesses of others in games. Existing approaches mainly focus on balancing between exploitation and exploitability but are often vulnerable to modeling errors and deceptive adversaries. To address this problem, our paper offers a novel perspective on the safety of opponent exploitation, named Adaptation Safety. This concept leverages the insight that strategies, even those not explicitly aimed at opponent exploitation, may inherently be exploitable due to computational complexities, rendering traditional safety overly rigorous. In contrast, adaptation safety requires that the strategy should not be more exploitable than it would be in scenarios where opponent exploitation is not considered. Building on such adaptation safety, we further propose an Opponent eXploitation Search (OX-Search) framework by incorporating real-time search techniques for efficient online opponent exploitation. Moreover, we provide theoretical analyses to show the adaptation safety and robust exploitation of OX-Search, even with inaccurate opponent models. Empirical evaluations in popular poker games demonstrate OX-Search’s superiority in both exploitability and exploitation compared to previous methods.
Zhenxing Ge, Tianyu Ding, Linjian Meng, Bo An 0001, Wenbin Li 0006, Yang Gao 0001
ICML7
2024 Discriminative Feature Decoupling Enhancement for Speech Forgery Detection
Yijun Bei, Erteng Liu, Yang Gao 0001, Kewei Gao, Zunlei Feng
IJCAI4
2024 STAR: Spatio-Temporal State Compression for Multi-Agent Tasks with Rich Observations
Yujing Hu, Shangdong Yang, Tangjie Lv, Changjie Fan, Wenbin Li 0012, Chongjie Zhang, Yang Gao 0001
IJCAI8
2024 SCaR: Refining Skill Chaining for Long-Horizon Robotic Manipulation via Dual Regularization
abstract
Long-horizon robotic manipulation tasks typically involve a series of interrelated sub-tasks spanning multiple execution stages. Skill chaining offers a feasible solution for these tasks by pre-training the skills for each sub-task and linking them sequentially. However, imperfections in skill learning or disturbances during execution can lead to the accumulation of errors in skill chaining process, resulting in execution failures. In this paper, we investigate how to achieve stable and smooth skill chaining for long-horizon robotic manipulation tasks. Specifically, we propose a novel skill chaining framework called Skill Chaining via Dual Regularization (SCaR). This framework applies dual regularization to sub-task skill pre-training and fine-tuning, which not only enhances the intra-skill dependencies within each sub-task skill but also reinforces the inter-skill dependencies between sequential sub-task skills, thus ensuring smooth skill chaining and stable long-horizon execution. We evaluate the SCaR framework on two representative long-horizon robotic manipulation simulation benchmarks: IKEA furniture assembly and kitchen organization. Additionally, we conduct a simple real-world validation in tabletop robot pick-and-place tasks. The experimental results show that, with the support of SCaR, the robot achieves a higher success rate in long-horizon tasks compared to relevant baselines and demonstrates greater robustness to perturbations.
Ze Ji, Jing Huo, Yang Gao 0001
NeurIPS4
2024 START: A Generalized State Space Model with Saliency-Driven Token-Aware Transformation
abstract
Domain Generalization (DG) aims to enable models to generalize to unseen target domains by learning from multiple source domains. Existing DG methods primarily rely on convolutional neural networks (CNNs), which inherently learn texture biases due to their limited receptive fields, making them prone to overfitting source domains. While some works have introduced transformer-based methods (ViTs) for DG to leverage the global receptive field, these methods incur high computational costs due to the quadratic complexity of self-attention. Recently, advanced state space models (SSMs), represented by Mamba, have shown promising results in supervised learning tasks by achieving linear complexity in sequence length during training and fast RNN-like computation during inference. Inspired by this, we investigate the generalization ability of the Mamba model under domain shifts and find that input-dependent matrices within SSMs could accumulate and amplify domain-specific features, thus hindering model generalization. To address this issue, we propose a novel SSM-based architecture with saliency-based token-aware transformation (namely START), which achieves state-of-the-art (SOTA) performances and offers a competitive alternative to CNNs and ViTs. Our START can selectively perturb and suppress domain-specific features in salient tokens within the input-dependent matrices of SSMs, thus effectively reducing the discrepancy between different domains. Extensive experiments on five benchmarks demonstrate that START outperforms existing SOTA DG methods with efficient linear complexity. Our code is available at https://github.com/lingeringlight/START.
Jintao Guo, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
NeurIPS4
2024 Transformer Doctor: Diagnosing and Treating Vision Transformers
abstract
Due to its powerful representational capabilities, Transformers have gradually become the mainstream model in the field of machine vision. However, the vast and complex parameters of Transformers impede researchers from gaining a deep understanding of their internal mechanisms, especially error mechanisms. Existing methods for interpreting Transformers mainly focus on understanding them from the perspectives of the importance of input tokens or internal modules, as well as the formation and meaning of features. In contrast, inspired by research on information integration mechanisms and conjunctive errors in the biological visual system, this paper conducts an in-depth exploration of the internal error mechanisms of Transformers. We first propose an information integration hypothesis for Transformers in the machine vision domain and provide substantial experimental evidence to support this hypothesis. This includes the dynamic integration of information among tokens and the static integration of information within tokens in Transformers, as well as the presence of conjunctive errors therein. Addressing these errors, we further propose heuristic dynamic integration constraint methods and rule-based static integration constraint methods to rectify errors and ultimately improve model performance. The entire methodology framework is termed as Transformer Doctor, designed for diagnosing and treating internal errors within transformers. Through a plethora of quantitative and qualitative experiments, it has been demonstrated that Transformer Doctor can effectively address internal errors in transformers, thereby enhancing model performance.
Jiacong Hu, Hao Chen 0041, Kejia Chen 0007, Yang Gao 0001, Jingwen Ye, Xingen Wang, Mingli Song, Zunlei Feng
NeurIPS4
2024 Model LEGO: Creating Models Like Disassembling and Assembling Building Blocks
abstract
With the rapid development of deep learning, the increasing complexity and scale of parameters make training a new model increasingly resource-intensive. In this paper, we start from the classic convolutional neural network (CNN) and explore a paradigm that does not require training to obtain new models. Similar to the birth of CNN inspired by receptive fields in the biological visual system, we draw inspiration from the information subsystem pathways in the biological visual system and propose Model Disassembling and Assembling (MDA). During model disassembling, we introduce the concept of relative contribution and propose a component locating technique to extract task-aware components from trained CNN classifiers. For model assembling, we present the alignment padding strategy and parameter scaling strategy to construct a new model tailored for a specific task, utilizing the disassembled task-aware components. The entire process is akin to playing with LEGO bricks, enabling arbitrary assembly of new models, and providing a novel perspective for model creation and reuse. Extensive experiments showcase that task-aware components disassembled from CNN classifiers or new models assembled using these components closely match or even surpass the performance of the baseline, demonstrating its promising results for model reuse. Furthermore, MDA exhibits diverse potential applications, with comprehensive experiments exploring model decision route analysis, model compression, knowledge distillation, and more.
Jiacong Hu, Jingwen Ye, Yang Gao 0001, Xingen Wang, Zunlei Feng, Mingli Song
NeurIPS4
2024 Re-examining Supervised Dimension Reduction for High-Dimensional Bayesian Optimization
Quanlin Chen, Jing Huo, Tianyu Ding, Yang Gao 0001, Dong Li 0016
PPSN (2)5
2024 Task-Aware Few-Shot Image Generation via Dynamic Local Distribution Estimation and Sampling
Zheng Gu 0001, Wenbin Li 0006, Tianyu Ding, Jing Huo, Kuihua Huang, Yang Gao 0001
PRCV (2)7
2024 Making the Primary Task Primary: Boosting Few-Shot Classification by Gradient-Biased Multi-task Learning
Yunchen Wu, Boyao Shi, Jing Huo, Wenbin Li 0006, Yang Gao 0001, Tinghao Yu
PRCV (1)5
2024 Distributionally Robust Graph-based Recommendation System
abstract
With the capacity to capture high-order collaborative signals, Graph Neural Networks (GNNs) have emerged as powerful methods in Recommender Systems (RS). However, their efficacy often hinges on the assumption that training and testing data share the same distribution (\aka IID assumption), and exhibits significant declines under distribution shifts. Distribution shifts commonly arises in RS, often attributed to the dynamic nature of user preferences or ubiquitous biases during data collection in RS. Despite its significance, researches on GNN-based recommendation against distribution shift are still sparse. To bridge this gap, we propose Distributionally Robust GNN (DR-GNN) that incorporates Distributional Robust Optimization (DRO) into the GNN-based recommendation. DR-GNN addresses two core challenges: 1) To enable DRO to cater to graph data intertwined with GNN, we reinterpret GNN as a graph smoothing regularizer, thereby facilitating the nuanced application of DRO; 2) Given the typically sparse nature of recommendation data, which might impede robust optimization, we introduce slight perturbations in the training distribution to expand its support. Notably, while DR-GNN involves complex optimization, it can be implemented easily and efficiently. Our extensive experiments validate the effectiveness of DR-GNN against three typical distribution shifts. The code is available at https://github.com/WANGBohaO-jpg/DR-GNN.
Bohao Wang 0001, Jiawei Chen 0007, Changdong Li, Sheng Zhou 0004, Qihao Shi, Yang Gao 0001, Chun Chen 0001, Can Wang 0001
WWW6
2024 Towards efficient image and video style transfer via distillation and learnable feature transformation
Jing Huo, Meihao Kong, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001
Comput. Vis. Image Underst.6
2024 Decentralized Counterfactual Value with Threat Detection for Multi-Agent Reinforcement Learning in mixed cooperative and competitive environments
Shaokang Dong, Shangdong Yang, Yang Gao 0001
Expert Syst. Appl.5
2024 CAT: A Simple yet Effective Cross-Attention Transformer for One-Shot Object Detection
Yuyan Deng, Yang Gao 0001, Ning Wang 0020, Lingqiao Liu, Lei Zhang 0054, Peng Wang 0015
J. Comput. Sci. Technol.3
2024 Selective policy transfer in multi-agent systems with sparse interactions
Yunkai Zhuang, Yong Liu 0007, Shangdong Yang, Yang Gao 0001
Knowl. Based Syst.4
2024 A Novel Contrastive Learning Model for Aerial Images
abstract
In the field of remote sensing, tasks related to images often require a large amount of manually annotated data for model training, thus requiring a large amount of human resources with professional knowledge. To alleviate this problem, some self-supervised methods have been proposed for learning image representations for downstream tasks. As a specific form, contrastive learning can be effectively used for unlabeled image data. Using this method, we achieve the ability to learn informative feature representations from extensive unlabeled data, thereby establishing a robust initialization model for subsequent tasks. But during the model learning process, the abundance and quality of negative samples significantly impact the model’s overall performance. Consequently, we improve the existing contrastive learning architecture of the conventional two-pathway network by incorporating an additional branch equipped with a hybrid encoder. This modification facilitates the fusion of both global and local features. Furthermore, we introduce other optimization techniques, including image reconstruction and indicator vector transformation. Notably, our approach emphasizes maintaining diversity among negative samples stored in the historical feature storage queue. Our model provides competitive results on aerial image dataset (AID) classification. Besides, our model outperforms other self-supervised baselines under different proportions of labeled data fine-tuning.
Taihang Zhen, Kai Chen 0026, Yang Gao 0001
IEEE Geosci. Remote. Sens. Lett.3
2024 Egoism, utilitarianism and egalitarianism in multi-agent reinforcement learning
Shaokang Dong, Shangdong Yang, Bo An 0001, Wenbin Li 0006, Yang Gao 0001
Neural Networks6
2024 Attention-based investigation and solution to the trade-off issue of adversarial training
Chang-Bin Shao, Wenbin Li 0006, Jing Huo, Zhenhua Feng 0001, Yang Gao 0001
Neural Networks5
2024 WToE: Learning When to Explore in Multiagent Reinforcement Learning
abstract
Existing multiagent exploration works focus on how to explore in the fully cooperative task, which is insufficient in the environment with nonstationarity induced by agent interactions. To tackle this issue, we propose When to Explore (WToE), a simple yet effective variational exploration method to learn WToE under nonstationary environments. WToE employs an interaction-oriented adaptive exploration mechanism to adapt to environmental changes. We first propose a novel graphical model that uses a latent random variable to model the step-level environmental change resulting from interaction effects. Leveraging this graphical model, we employ the supervised variational auto-encoder (VAE) framework to derive a short-term inferred policy from historical trajectories to deal with the nonstationarity. Finally, agents engage in exploration when the short-term inferred policy diverges from the current actor policy. The proposed approach theoretically guarantees the convergence of the Q -value function. In our experiments, we validate our exploration mechanism in grid examples, multiagent particle environments and the battle game of MAgent environments. The results demonstrate the superiority of WToE over multiple baselines and existing exploration methods, such as MAEXQ, NoisyNets, EITI, and PR2.
Shaokang Dong, Hangyu Mao, Shangdong Yang, Shengyu Zhu 0001, Wenbin Li 0006, Jianye Hao, Yang Gao 0001
IEEE Trans. Cybern.7
2024 Modeling Rationality: Toward Better Performance Against Unknown Agents in Sequential Games
abstract
Opponent modeling is necessary for autonomous agents to capture the intents of others during strategic interactions. Most previous works assume that they can access enough interaction history to build the model. However, it may not be realistic. To solve this problem, we present a novel rationality-consistent opponent modeling (ROM) method for games with imperfect information. In our approach, a game-theoretical concept of consistence about rationality is proposed to take advantage of the characteristic of imperfect information sequential games that rational behavior at disjoint information sets is correlated through anticipated opponent's behavior. With the correlation between different information sets, agents could infer the opponents' strategies at information sets correlated to observed behavior. To exploit the correlation, ROM attempts to conduct reasoning from the opponent's perspective and rationalize its past behavior. In this way, ROM acquires the ability to better adapt to different opponents and achieves a more accurate opponent model with insufficient observation history, which is verified by experiments in different settings. A heuristic adaptation approach is also applied in ROM, which updates the opponent model in an online manner and significantly reduces the computation cost. We evaluate ROM in both a grid world game and a poker game. Compared with other opponent modeling methods, ROM shows better performance and has more accurate predictions in both games against different types of opponents with limited action interactions. Experimental results also show that ROM's time cost is significantly reduced through heuristic adaptation.
Zhenxing Ge, Shangdong Yang, Pinzhuo Tian, Yang Gao 0001
IEEE Trans. Cybern.5
2024 SETA: Semantic-Aware Edge-Guided Token Augmentation for Domain Generalization
abstract
Domain generalization (DG) aims to enhance the model robustness against domain shifts without accessing target domains. A prevalent category of methods for DG is data augmentation, which focuses on generating virtual samples to simulate domain shifts. However, existing augmentation techniques in DG are mainly tailored for convolutional neural networks (CNNs), with limited exploration in token-based architectures, i.e., vision transformer (ViT) and multi-layer perceptrons (MLP) models. In this paper, we study the impact of prior CNN-based augmentation methods on token-based models, revealing their performance is suboptimal due to the lack of incentivizing the model to learn holistic shape information. To tackle the issue, we propose the Semantic-aware Edge-guided Token Augmentation (SETA) method. SETA transforms token features by perturbing local edge cues while preserving global shape features, thereby enhancing the model learning of shape information. To further enhance the generalization ability of the model, we introduce two stylized variants of our method combined with two state-of-the-art (SOTA) style augmentation methods in DG. We provide a theoretical insight into our method, demonstrating its effectiveness in reducing the generalization risk bound. Comprehensive experiments on five benchmarks prove that our method achieves SOTA performances across various ViT and MLP architectures. Our code is available at https://github.com/lingeringlight/SETA.
Jintao Guo, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Image Process.4
2024 Learning Generalizable Models via Disentangling Spurious and Enhancing Potential Correlations
abstract
Domain generalization (DG) intends to train a model on multiple source domains to ensure that it can generalize well to an arbitrary unseen target domain. The acquisition of domain-invariant representations is pivotal for DG as they possess the ability to capture the inherent semantic information of the data, mitigate the influence of domain shift, and enhance the generalization capability of the model. Adopting multiple perspectives, such as the sample and the feature, proves to be effective. The sample perspective facilitates data augmentation through data manipulation techniques, whereas the feature perspective enables the extraction of meaningful generalization features. In this paper, we focus on improving the generalization ability of the model by compelling it to acquire domain-invariant representations from both the sample and feature perspectives by disentangling spurious correlations and enhancing potential correlations. 1) From the sample perspective, we develop a frequency restriction module, guiding the model to focus on the relevant correlations between object features and labels, thereby disentangling spurious correlations. 2) From the feature perspective, the simple Tail Interaction module implicitly enhances potential correlations among all samples from all source domains, facilitating the acquisition of domain-invariant representations across multiple domains for the model. The experimental results show that Convolutional Neural Networks (CNNs) or Multi-Layer Perceptrons (MLPs) with a strong baseline embedded with these two modules can achieve superior results, e.g., an average accuracy of 92.30% on Digits-DG. Source code is available at https://github.com/RubyHoho/DGeneralization.
Lei Qi 0001, Jintao Guo, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Image Process.5
2024 Learning Multi-Intersection Traffic Signal Control via Coevolutionary Multi-Agent Reinforcement Learning
abstract
Effective management of multi-intersection traffic signal control (MTSC) is vital for intelligent transportation systems. Multi-agent reinforcement learning (MARL) has shown promise in achieving MTSC. However, existing MARL-based MTSC algorithms have primarily focused on capturing the spatial relationship between multi-intersection traffic signals but have overlooking the importance of the temporally stable traffic pattern. This pattern refers to the fixed positions and relatively stable traffic flow between intersections over short periods in real-world MTSC scenarios, which indicates that the learned spatial relationships between traffic signals should co-evolve over time. To this end, we propose a novel algorithm calledCoevolutionaryMulti-AgentReinforcementLearning (CoevoMARL). CoevoMARL employs a graph neural network to capture the complex spatial interaction network among traffic signals. Furthermore, we propose a relationship-driven progressive LSTM (RDP-LSTM) that dynamically evolves the learned spatial interaction network over time by leveraging insights from the temporally stable traffic pattern. To accelerate convergence, we also propose the mutual information reward optimization (MIRO) technique, which strengthens the correlation between policy learning and high-performance samples by using a mutual information-based intrinsic reward. Experimental results on both synthetic and realistic datasets demonstrate the superiority of CoevoMARL over existing MTSC algorithms, providing valuable insights into incorporating the temporally stable traffic pattern.
Wubing Chen, Shangdong Yang, Wenbin Li 0006, Yujing Hu, Yang Gao 0001
IEEE Trans. Intell. Transp. Syst.6
2024 Secure and efficient general matrix multiplication on cloud using homomorphic encryption
Yang Gao 0001, Gang Quan, Soamar Homsi, Wujie Wen, Liqiang Wang 0001
J. Supercomput.1
2024 Open-Domain Semi-Supervised Learning via Glocal Cluster Structure Exploitation
abstract
Semi-supervised learning (SSL) aims to reduce the heavy reliance of current deep models on costly manual annotation by leveraging a large amount of unlabeled data in combination with a much smaller set of labeled data. However, most existing SSL methods assume that all labeled and unlabeled data are drawn from the same feature distribution, which can be impractical in real-world applications. In this study, we take the initial step to systematically investigate the open-domain semi-supervised learning setting, where a feature distribution mismatch exists between labeled and unlabeled data. In pursuit of an effective solution for open-domain SSL, we propose a novel framework calledGlocalMatch, which aims to exploit bothglobal and local(i.e., glocal) cluster structure of open-domain unlabeled data. The glocal cluster structure is utilized in two complementary ways. First, GlocalMatch optimizes a Glocal Cluster Compacting (GCC) objective, that encourages feature representations of the same class, whether with in the same domain or across different domains, to become closer to each other. Second, GlocalMatch incorporates a Glocal Semantic Aggregation (GSA) strategy to produce more reliable pseudo-labels by aggregating predictions from neighboring clusters. Extensive experiments demonstrate that GlocalMatch outperforms the state-of-the-art SSL methods significantly, achieving superior performance for both in-domain and out-of-domain generalization.
Zekun Li 0010, Lei Qi 0001, Yawen Li 0001, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Knowl. Data Eng.5
2024 Exploring Flat Minima for Domain Generalization With Large Learning Rates
abstract
Domain Generalization (DG) aims to generalize to arbitrary unseen domains. A promising approach to improve model generalization in DG is the identification of flat minima. One typical method for this task is SWAD, which involves averaging weights along the training trajectory. However, the success of weight averaging depends on the diversity of weights, which is limited when training with a small learning rate. Instead, we observe that leveraging a large learning rate can simultaneously promote weight diversity and facilitate the identification of flat regions in the loss landscape. However, employing a large learning rate suffers from the convergence problem, which cannot be resolved by simply averaging the training weights. To address this issue, we introduce a training strategy called Lookahead which involves the weight interpolation, instead of average, between fast and slow weights. The fast weight explores the weight space with a large learning rate, which is not converged while the slow weight interpolates with it to ensure the convergence. Besides, weight interpolation also helps identify flat minima by implicitly optimizing the local entropy loss that measures flatness. To further prevent overfitting during training, we propose two variants to regularize the training weight with weighted averaged weight or with accumulated history weight. Taking advantage of this new perspective, our methods achieve state-of-the-art performance on both classification and semantic segmentation domain generalization benchmarks. The code is available athttps://github.com/koncle/DG-with-Large-LR.
Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Knowl. Data Eng.4
2024 MutexMatch: Semi-Supervised Learning With Mutex-Based Consistency Regularization
abstract
The core issue in semi-supervised learning (SSL) lies in how to effectively leverage unlabeled data, whereas most existing methods tend to put a great emphasis on the utilization of high-confidence samples yet seldom fully explore the usage of low-confidence samples. In this article, we aim to utilize low-confidence samples in a novel way with our proposed mutex-based consistency regularization, namely MutexMatch. Specifically, the high-confidence samples are required to exactly predict "what it is" by the conventional true-positive classifier (TPC), while low-confidence samples are employed to achieve a simpler goal-to predict with ease "what it is not" by the true-negative classifier (TNC). In this sense, we not only mitigate the pseudo-labeling errors but also make full use of the low-confidence unlabeled data by the consistency of dissimilarity degree. MutexMatch achieves superior performance on multiple benchmark datasets, i.e., Canadian Institute for Advanced Research (CIFAR)-10, CIFAR-100, street view house numbers (SVHN), self-taught learning 10 (STL-10), and mini-ImageNet. More importantly, our method further shows superiority when the amount of labeled data is scarce, e.g., 92.23% accuracy with only 20 labeled data on CIFAR-10. Code has been released at https://github.com/NJUyued/MutexMatch4SSL.
Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Lei Wang 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.7
2024 Iterative Multiview Subspace Learning for Unpaired Multiview Clustering
abstract
In real applications, several unpredictable or uncertain factors could result in unpaired multiview data, i.e., the observed samples between views cannot be matched. Since joint clustering among views is more effective than individual clustering in each view, we investigate unpaired multiview clustering (UMC), which is a valuable but insufficiently studied problem. Due to lack of matched samples between views, we could fail to build the connection between views. Therefore, we aim to learn the latent subspace shared by views. However, existing multiview subspace learning methods usually rely on the matched samples between views. To address this issue, we propose an iterative multiview subspace learning strategy [iterative unpaired multiview clustering (IUMC)], aiming to learn a complete and consistent subspace representation among views for UMC. Moreover, based on IUMC, we design two effective UMC methods: 1) Iterative unpaired multiview clustering via covariance matrix alignment (IUMC-CA) that further aligns the covariance matrix of subspace representations and then performs clustering on the subspace and 2) iterative unpaired multiview clustering via one-stage clustering assignments (IUMC-CY) that performs one-stage multiview clustering (MVC) by replacing the subspace representations with clustering assignments. Extensive experiments show the excellent performance of our methods for UMC, compared with the state-of-the-art methods. Also, the clustering performance of observed samples in each view can be considerably improved by those observed samples from the other views. In addition, our methods have good applicability in incomplete MVC.
Wanqi Yang, Like Xin, Lei Wang 0001, Ming Yang 0014, Wenzhu Yan, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Analogist: Out-of-the-box Visual In-Context Learning with Image Diffusion Model
abstract
Visual In-Context Learning (ICL) has emerged as a promising research area due to its capability to accomplish various tasks with limited example pairs through analogical reasoning. However, training-based visual ICL has limitations in its ability to generalize to unseen tasks and requires the collection of a diverse task dataset. On the other hand, existing methods in the inference-based visual ICL category solely rely on textual prompts, which fail to capture fine-grained contextual information from given examples and can be time-consuming when converting from images to text prompts. To address these challenges, we propose Analogist, a novel inference-based visual ICL approach that exploits both visual and textual prompting techniques using a text-to-image diffusion model pretrained for image inpainting. For visual prompting, we propose a self-attention cloning (SAC) method to guide the fine-grained structural-level analogy between image examples. For textual prompting, we leverage GPT-4V's visual reasoning capability to efficiently generate text prompts and introduce a cross-attention masking (CAM) operation to enhance the accuracy of semantic-level analogy guided by text prompts. Our method is out-of-the-box and does not require fine-tuning or optimization. It is also generic and flexible, enabling a wide range of visual tasks to be performed in an in-context manner. Extensive experiments demonstrate the superiority of our method over existing approaches, both qualitatively and quantitatively. Our project webpage is available at https://analogist2d.github.io.
Zheng Gu 0001, Jing Liao 0001, Jing Huo, Yang Gao 0001
ACM Trans. Graph.5
2024 PLACE Dropout: A Progressive Layer-wise and Channel-wise Dropout for Domain Generalization
abstract
Domain generalization (DG) aims to learn a generic model from multiple observed source domains that generalizes well to arbitrary unseen target domains without further training. The major challenge in DG is that the model inevitably faces a severe overfitting issue due to the domain gap between source and target domains. To mitigate this problem, some dropout-based methods have been proposed to resist overfitting by discarding part of the representation of the intermediate layers. However, we observe that most of these methods only conduct the dropout operation in some specific layers, leading to an insufficient regularization effect on the model. We argue that applying dropout at multiple layers can produce stronger regularization effects, which could alleviate the overfitting problem on source domains more adequately than previous layer-specific dropout methods. In this article, we develop a novel layer-wise and channel-wise dropout for DG, which randomly selects one layer and then randomly selects its channels to conduct dropout. Particularly, the proposed method can generate a variety of data variants to better deal with the overfitting issue. We also provide theoretical analysis for our dropout method and prove that it can effectively reduce the generalization error bound. Besides, we leverage the progressive scheme to increase the dropout ratio with the training progress, which can gradually boost the difficulty of training the model to enhance its robustness. Extensive experiments on three standard benchmark datasets have demonstrated that our method outperforms several state-of-the-art DG methods. Our code is available at https://github.com/lingeringlight/PLACEdropout .
Jintao Guo, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Learning Explicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning via Polarization Policy Gradient
abstract
Cooperative multi-agent policy gradient (MAPG) algorithms have recently attracted wide attention and are regarded as a general scheme for the multi-agent system. Credit assignment plays an important role in MAPG and can induce cooperation among multiple agents. However, most MAPG algorithms cannot achieve good credit assignment because of the game-theoretic pathology known as centralized-decentralized mismatch. To address this issue, this paper presents a novel method, Multi-Agent Polarization Policy Gradient (MAPPG). MAPPG takes a simple but efficient polarization function to transform the optimal consistency of joint and individual actions into easily realized constraints, thus enabling efficient credit assignment in MAPPG. Theoretically, we prove that individual policies of MAPPG can converge to the global optimum. Empirically, we evaluate MAPPG on the well-known matrix game and differential game, and verify that MAPPG can converge to the global optimum for both discrete and continuous action spaces. We also evaluate MAPPG on a set of StarCraft II micromanagement tasks and demonstrate that MAPPG outperforms the state-of-the-art MAPG algorithms.
Wubing Chen, Wenbin Li 0006, Shangdong Yang, Yang Gao 0001
AAAI5
2023 An Efficient Deep Reinforcement Learning Algorithm for Solving Imperfect Information Extensive-Form Games
abstract
One of the most popular methods for learning Nash equilibrium (NE) in large-scale imperfect information extensive-form games (IIEFGs) is the neural variants of counterfactual regret minimization (CFR). CFR is a special case of Follow-The-Regularized-Leader (FTRL). At each iteration, the neural variants of CFR update the agent's strategy via the estimated counterfactual regrets. Then, they use neural networks to approximate the new strategy, which incurs an approximation error. These approximation errors will accumulate since the counterfactual regrets at iteration t are estimated using the agent's past approximated strategies. Such accumulated approximation error causes poor performance. To address this accumulated approximation error, we propose a novel FTRL algorithm called FTRL-ORW, which does not utilize the agent's past strategies to pick the next iteration strategy. More importantly, FTRL-ORW can update its strategy via the trajectories sampled from the game, which is suitable to solve large-scale IIEFGs since sampling multiple actions for each information set is too expensive in such games. However, it remains unclear which algorithm to use to compute the next iteration strategy for FTRL-ORW when only such sampled trajectories are revealed at iteration t. To address this problem and scale FTRL-ORW to large-scale games, we provide a model-free method called Deep FTRL-ORW, which computes the next iteration strategy using model-free Maximum Entropy Deep Reinforcement Learning. Experimental results on two-player zero-sum IIEFGs show that Deep FTRL-ORW significantly outperforms existing model-free neural methods and OS-MCCFR.
Linjian Meng, Zhenxing Ge, Pinzhuo Tian, Bo An 0001, Yang Gao 0001
AAAI5
2023 Enhanced Tensor Low-Rank and Sparse Representation Recovery for Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering (IMVC) has attracted remarkable attention due to the emergence of multi-view data with missing views in real applications. Recent methods attempt to recover the missing information to address the IMVC problem. However, they generally cannot fully explore the underlying properties and correlations of data similarities across views. This paper proposes a novel Enhanced Tensor Low-rank and Sparse Representation Recovery (ETLSRR) method, which reformulates the IMVC problem as a joint incomplete similarity graphs learning and complete tensor representation recovery problem. Specifically, ETLSRR learns the intra-view similarity graphs and constructs a 3-way tensor by stacking the graphs to explore the inter-view correlations. To alleviate the negative influence of missing views and data noise, ETLSRR decomposes the tensor into two parts: a sparse tensor and an intrinsic tensor, which models the noise and underlying true data similarities, respectively. Both global low-rank and local structured sparse characteristics of the intrinsic tensor are considered, which enhances the discrimination of similarity matrix. Moreover, instead of using the convex tensor nuclear norm, ETLSRR introduces a generalized non-convex tensor low-rank regularization to alleviate the biased approximation. Experiments on several datasets demonstrate the effectiveness of our method compared with the state-of-the-art methods.
Chao Zhang 0078, Huaxiong Li, Zizheng Huang, Yang Gao 0001, Chunlin Chen 0001
AAAI5
2023 Drift-aware Anomaly Detection for Non-stationary Time Series
abstract
Anomaly detection of time series is vital in various scenarios with explosively growing time series data. However, the non-stationary time series degrade the performance of current anomaly detection methods, where data drift causes unpredictable changes. This paper proposes a Drift-aware Anomaly Detection (DAD) method for detecting anomalies in non-stationary time series. DAD adopts a self-attention mechanism to learn an embedding, distinguishing the anomaly embeddings from the normal embeddings. Next, the KL divergence calculates the drift deviation between two data segments at adjacent periods. Then, the drift deviation module combined with the latent vector which is used to reconstruct the original vector. During the encoding stage of the time series, the latent code is modeled using different Gaussian mixture distributions and the data reconstruction error at each time tick is regarded as an anomaly metric. Furthermore, we propose a new metric to measure the degree of drift deviation for a dataset used for a fair experiment comparison. Experimental results on several public datasets and a newly collected sensor dataset demonstrate that for the non-stationary time series anomaly detection task, DAD outperforms state-of-the-art anomaly detection models up to 11.5% on the F1score.
Yang Gao 0001, Ying Li 0097, Zunlei Feng, Mingli Song, Chun Chen 0001
IEEE Big Data1
2023 Orthogonal Annotation Benefits Barely-supervised Medical Image Segmentation
abstract
Recent trends in semi-supervised learning have significantly boosted the performance of 3D semi-supervised medical image segmentation. Compared with 2D images, 3D medical volumes involve information from different directions, e.g., transverse, sagittal, and coronal planes, so as to naturally provide complementary views. These complementary views and the intrinsic similarity among adjacent 3D slices inspire us to develop a novel annotation way and its corresponding semi-supervised model for effective segmentation. Specifically, we firstly propose the orthogonal annotation by only labeling two orthogonal slices in a labeled volume, which significantly relieves the burden of annotation. Then, we perform registration to obtain the initial pseudo labels for sparsely labeled volumes. Subsequently, by introducing unlabeled volumes, we propose a dual-network paradigm named Dense-Sparse Co-training (DeSCO) that exploits dense pseudo labels in early stage and sparse labels in later stage and meanwhile forces consistent output of two networks. Experimental results on three benchmark datasets validated our effectiveness in performance and efficiency in annotation. For example, with only 10 annotated slices, our method reaches a Dice up to 86.93% on KiTS19 dataset. Our code and models are available at https://github.com/HengCai-NJU/DeSCO.
Heng Cai, Shumeng Li, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001
CVPR6
2023 Modeling Inter-Class and Intra-Class Constraints in Novel Class Discovery
abstract
Novel class discovery (NCD) aims at learning a model that transfers the common knowledge from a class-disjoint labelled dataset to another unlabelled dataset and discovers new classes (clusters) within it. Many methods, as well as elaborate training pipelines and appropriate objectives, have been proposed and considerably boosted performance on NCD tasks. Despite all this, we find that the existing methods do not sufficiently take advantage of the essence of the NCD setting. To this end, in this paper, we propose to model both inter-class and intra-class constraints in NCD based on the symmetric Kullback-Leibler divergence (sKLD). Specifically, we propose an inter-class sKLD constraint to effectively exploit the disjoint relationship between labelled and unlabelled classes, enforcing the separability for different classes in the embedding space. In addition, we present an intra-class sKLD constraint to explicitly constrain the intra-relationship between a sample and its augmentations and ensure the stability of the training process at the same time. We conduct extensive experiments on the popular CIFAR10, CIFAR100 and ImageNet benchmarks and successfully demonstrate that our method can establish a new state of the art and can achieve significant performance improvements, e.g., 3.5%/3.7% clustering accuracy improvements on CIFAR100-50 dataset split under the task-aware/-agnostic evaluation protocol, over previous state-of-the-art methods. Code is available at https://github.com/FanZhichen/NCD-IIC.
Wenbin Li 0006, Zhichen Fan, Jing Huo, Yang Gao 0001
CVPR4
2023 Backtracking Exploration for Reinforcement Learning
abstract
Exploration of the behavior policy plays an important role in reinforcement learning as it helps learning algorithms escape local optima. Taking linear value function approximation as an example, exploration directly affects the sampling of states, thereby altering the distribution of states. This distribution is a component of the key matrix, and the magnitude of the smallest eigenvalue of the key matrix is proportional to the convergence speed. However, existing exploration methods are constrained by the MDP chain and require step-by-step backtracking to reach the target policy distribution. This paper breaks the assumption that the action settings of the training environment must be identical to that of the testing environment by introducing state resetting in the training environment and proposes a backtracking exploration algorithm with time window and punishment. This algorithm can be directly combined with existing exploration strategies and value function update rules, and it has the potential to become a new paradigm for the training process in reinforcement learning. Experimental results validate the effectiveness of the proposed algorithm.
Xingguo Chen, Zening Chen, Dingyuanhao Sun, Yang Gao 0001
DAI4
2023 Enhancing OOD Generalization in Offline Reinforcement Learning with Energy-Based Policy Optimization
abstract
Offline Reinforcement Learning (RL) is an important research domain for real-world applications because it can avert expensive and dangerous online exploration. Offline RL is prone to extrapolation errors caused by the distribution shift between offline datasets and states visited by behavior policy. Existing offline RL methods constrain the policy to offline behavior to prevent extrapolation errors. But these methods limit the generalization potential of agents in Out-Of-Distribution (OOD) regions and cannot effectively evaluate OOD generalization behavior. To improve the generalization of the policy in OOD regions while avoiding extrapolation errors, we propose an Energy-Based Policy Optimization (EBPO) method for OOD generalization. An energy function based on the distribution of offline data is proposed for the evaluation of OOD generalization behavior, instead of relying on model discrepancies to constrain the policy. The way of quantifying exploration behavior in terms of energy values can balance the return and risk. To improve the stability of generalization and solve the problem of sparse reward in complex environment, episodic memory is applied to store successful experiences that can improve sample efficiency. Extensive experiments on the D4RL datasets demonstrate that EBPO outperforms the state-of-the-art methods and achieves robust performance on challenging tasks that require OOD generalization.
Hongye Cao, Shangdong Yang, Jing Huo, Xingguo Chen, Yang Gao 0001
ECAI5
2023 LoSS: Local Structural Separation Hypergraph Convolutional Neural Network
abstract
Graph classification is a classic problem with practical applications in many real-life scenes. Existing graph neural networks, including GCN, GAT, and GIN, are proposed to extract useful features from complex graph structures. However, most existing methods’ feature extraction and aggregation inevitably mix the useful and redundant features, which will disturb the final classification performance. In this paper, to handle the above drawback, we put forward the Local Structural Separation Hypergraph Convolutional Neural Network (LoSS) based on two discoveries: most graph classification tasks only focus on a few groups of adjacent nodes, and different categories have their specific high response bits in graph embeddings. In LoSS, we first decouple the original graph into different hypergraphs and aggregate the features in each substructure, which aims to find useful features for the final classification. Next, the low-correlation feature suppression strategy is devised to suppress the irrelevant node-level and bit-level features in the forward inference process, effectively reducing the disturbance of redundant features. Experiments on five datasets show that the proposed LoSS can effectively locate and aggregate useful hypergraph features and achieve SOTA performance compared with existing methods.
Bingde Hu, Yang Gao 0001, Zunlei Feng, Mingli Song, Xinyu Wang 0001, Ying Li 0097
ECAI2
2023 Convergence Analysis of Graphical Game-Based Nash Q-Learning using the Interaction Detection Signal of N-Step Return
abstract
The graphical game provides an effective method for modeling different kinds of sparse interactions in multi-agent reinforcement learning. Most previous work on game abstraction lacks theoretical guarantees of convergence. In this paper, we adopt the ${\mathcal{N}}$-step return signal to detect interactions between agents and build the Markov graphical game based on it. We analyze that the solution of the Markov graphical game is an ϵ-Nash equilibrium which guarantees the convergence of the proposed NSR-G2NashQ algorithm theoretically. Also, we have done experiments in different multi-agent reinforcement learning tasks with both tabular and function approximation solutions. The results show the NSR-G2NashQ algorithm accelerates the convergence of agents to the optimal policy.
Yunkai Zhuang, Shangdong Yang, Wenbin Li 0006, Yang Gao 0001
ICASSP4
2023 DomainAdaptor: A Novel Approach to Test-time Adaptation
abstract
To deal with the domain shift between training and test samples, current methods have primarily focused on learning generalizable features during training and ignore the specificity of unseen samples that are also critical during the test. In this paper, we investigate a more challenging task that aims to adapt a trained CNN model to unseen domains during the test. To maximumly mine the information in the test data, we propose a unified method called DomainAdaptor for the test-time adaptation, which consists of an AdaMixBN module and a Generalized Entropy Minimization (GEM) loss. Specifically, AdaMixBN addresses the domain shift by adaptively fusing training and test statistics in the normalization layer via a dynamic mixture co-efficient and a statistic transformation operation. To further enhance the adaptation ability of AdaMixBN, we design a GEM loss that extends the Entropy Minimization loss to better exploit the information in the test data. Extensive experiments show that DomainAdaptor consistently outperforms the state-of-the-art methods on four benchmarks. Furthermore, our method brings more remarkable improvement against existing methods on the few-data unseen domain. The code is available at https://github.com/koncle/DomainAdaptor.
Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
ICCV4
2023 IOMatch: Simplifying Open-Set Semi-Supervised Learning with Joint Inliers and Outliers Utilization
abstract
Semi-supervised learning (SSL) aims to leverage massive unlabeled data when labels are expensive to obtain. Unfortunately, in many real-world applications, the collected unlabeled data will inevitably contain unseen-class outliers not belonging to any of the labeled classes. To deal with the challenging open-set SSL task, the mainstream methods tend to first detect outliers and then filter them out. However, we observe a surprising fact that such approach could result in more severe performance degradation when labels are extremely scarce, as the unreliable outlier detector may wrongly exclude a considerable portion of valuable inliers. To tackle with this issue, we introduce a novel open-set SSL framework, IOMatch, which can jointly utilize inliers and outliers, even when it is difficult to distinguish exactly between them. Specifically, we propose to employ a multi-binary classifier in combination with the standard closed-set classifier for producing unified open-set classification targets, which regard all outliers as a single new class. By adopting these targets as open-set pseudo-labels, we optimize an open-set classifier with all unlabeled samples including both inliers and outliers. Extensive experiments have shown that IOMatch significantly outperforms the baseline methods across different benchmark datasets and different settings despite its remarkable simplicity. Our code and models are available at https://github.com/nukezil/IOMatch.
Zekun Li 0010, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
ICCV4
2023 3D Medical Image Segmentation with Sparse Annotation via Cross-Teaching Between 3D and 2D Networks
Heng Cai, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001
MICCAI (3)5
2023 Where and How: Mitigating Confusion in Neural Radiance Fields from Sparse Inputs
abstract
Neural Radiance Fields from Sparse inputs (NeRF-S) have shown great potential in synthesizing novel views with a limited number of observed viewpoints. However, due to the inherent limitations of sparse inputs and the gap between non-adjacent views, rendering results often suffer from over-fitting and foggy surfaces, a phenomenon we refer to as "CONFUSION" during volume rendering. In this paper, we analyze the root cause of this confusion and attribute it to two fundamental questions: "WHERE" and "HOW". To this end, we present a novel learning framework, WaH-NeRF, which effectively mitigates confusion by tackling the following challenges: (i) "WHERE" to Sample? in NeRF-S-we introduce a Deformable Sampling strategy and a Weight-based Mutual Information Loss to address sample-position confusion arising from the limited number of viewpoints; and (ii) "HOW" to Predict? in NeRF-S-we propose a Semi-Supervised NeRF learning Paradigm based on pose perturbation and a Pixel-Patch Correspondence Loss to alleviate prediction confusion caused by the disparity between training and testing viewpoints. By integrating our proposed modules and loss functions, WaH-NeRF outperforms previous methods under the NeRF-S setting. Code is available https://github.com/bbbbby-99/WaH-NeRF.
Yanqi Bao, Jing Huo, Tianyu Ding, Wenbin Li 0006, Yang Gao 0001
ACM Multimedia7
2023 Efficient Subgame Refinement for Extensive-form Games
abstract
Subgame solving is an essential technique in addressing large imperfect information games, with various approaches developed to enhance the performance of refined strategies in the abstraction of the target subgame. However, directly applying existing subgame solving techniques may be difficult, due to the intricate nature and substantial size of many real-world games. To overcome this issue, recent subgame solving methods allow for subgame solving on limited knowledge order subgames, increasing their applicability in large games; yet this may still face obstacles due to extensive information set sizes. To address this challenge, we propose a generative subgame solving (GS2) framework, which utilizes a generation function to identify a subset of the earliest-reached nodes, reducing the size of the subgame. Our method is supported by a theoretical analysis and employs a diversity-based generation function to enhance safety. Experiments conducted on medium-sized games as well as the challenging large game of GuanDan demonstrate a significant improvement over the blueprint.
Zhenxing Ge, Tianyu Ding, Wenbin Li 0006, Yang Gao 0001
NeurIPS5
2023 Modified Retrace for Off-Policy Temporal Difference Learning
abstract
Off-policy learning is a key to extend reinforcement learning as it allows to learn a target policy from a different behavior policy that generates the data. However, it is well known as “the deadly triad” when combined with bootstrapping and function approximation. Retrace is an efficient and convergent off-policy algorithm with tabular value functions which employs truncated importance sampling ratios. Unfortunately, Retrace is known to be unstable with linear function approximation. In this paper, we propose modified Retrace to correct the off-policy return, derive a new off-policy temporal difference learning algorithm (TD-MRetrace) with linear function approximation, and obtain a convergence guarantee under standard assumptions. Experimental results on counterexamples and control tasks validate the effectiveness of the proposed algorithm compared with traditional algorithms.
Xingguo Chen, Xingzhou Ma, Guang Yang 0066, Shangdong Yang, Yang Gao 0001
UAI6
2023 ASN: action semantics network for multiagent reinforcement learning
Tianpei Yang, Weixun Wang, Jianye Hao, Matthew E. Taylor, Yong Liu 0007, Xiaotian Hao, Yujing Hu, Changjie Fan, Chunxu Ren, Jiangcheng Zhu, Yang Gao 0001
Auton. Agents Multi Agent Syst.13
2023 Reinforcement learning based web crawler detection for diversity and dynamics
Yang Gao 0001, Zunlei Feng, Mingli Song, Xingen Wang, Xinyu Wang 0001, Chun Chen 0001
Neurocomputing1
2023 Online attentive kernel-based temporal difference learning
Xingguo Chen, Guang Yang 0066, Shangdong Yang, Shaokang Dong, Yang Gao 0001
Knowl. Based Syst.6
2023 LibFewShot: A Comprehensive Library for Few-Shot Learning
abstract
Few-shot learning, especially few-shot image classification, has received increasing attention and witnessed significant advances in recent years. Some recent studies implicitly show that many generic techniques or "tricks", such as data augmentation, pre-training, knowledge distillation, and self-supervision, may greatly boost the performance of a few-shot learning method. Moreover, different works may employ different software platforms, backbone architectures and input image sizes, making fair comparisons difficult and practitioners struggle with reproducibility. To address these situations, we propose a comprehensive library for few-shot learning (LibFewShot) by re-implementing eighteen state-of-the-art few-shot learning methods in a unified framework with the same single codebase in PyTorch. Furthermore, based on LibFewShot, we provide comprehensive evaluations on multiple benchmarks with various backbone architectures to evaluate common pitfalls and effects of different training tricks. In addition, with respect to the recent doubts on the necessity of meta- or episodic-training mechanism, our evaluation results confirm that such a mechanism is still necessary especially when combined with pre-training. We hope our work can not only lower the barriers for beginners to enter the area of few-shot learning but also elucidate the effects of nontrivial tricks to facilitate intrinsic research on few-shot learning.
Wenbin Li 0006, Xuesong Yang, Chuanqi Dong, Pinzhuo Tian, Tiexin Qin, Jing Huo, Yinghuan Shi, Lei Wang 0001, Yang Gao 0001, Jiebo Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.10
2023 Defensive Few-Shot Learning
abstract
This article investigates a new challenging problem called defensive few-shot learning in order to learn a robust few-shot model against adversarial attacks. Simply applying the existing adversarial defense methods to few-shot learning cannot effectively solve this problem. This is because the commonly assumed sample-level distribution consistency between the training and test sets can no longer be met in the few-shot setting. To address this situation, we develop a general defensive few-shot learning (DFSL) framework to answer the following two key questions: (1) how to transfer adversarial defense knowledge from one sample distribution to another? (2) how to narrow the distribution gap between clean and adversarial examples under the few-shot setting? To answer the first question, we propose an episode-based adversarial training mechanism by assuming a task-level distribution consistency to better transfer the adversarial defense knowledge. As for the second question, within each few-shot task, we design two kinds of distribution consistency criteria to narrow the distribution gap between clean and adversarial examples from the feature-wise and prediction-wise perspectives, respectively. Extensive experiments demonstrate that the proposed framework can effectively make the existing few-shot models robust against adversarial attacks. Code is available at https://github.com/WenbinLee/DefensiveFSL.git.
Wenbin Li 0006, Lei Wang 0001, Xingxing Zhang 0001, Lei Qi 0001, Jing Huo, Yang Gao 0001, Jiebo Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Global- and local-aware feature augmentation with semantic orthogonality for few-shot image classification
Boyao Shi, Wenbin Li 0006, Jing Huo, Lei Wang 0001, Yang Gao 0001
Pattern Recognit.6
2023 Better pseudo-label: Joint domain-aware label and dual-classifier for semi-supervised domain generalization
Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
Pattern Recognit.4
2023 Multi-Level Cascade Sparse Representation Learning for Small Data Classification
abstract
Deep learning (DL) methods have recently captured much attention for image classification. However, such methods may lead to a suboptimal solution for small-scale data since the lack of training samples. Sparse representation stands out with its efficiency and interpretability, but its precision is not so competitive. We develop a Multi-Level Cascade Sparse Representation (ML-CSR) learning method to combine both advantages when processing small-scale data. ML-CSR is proposed using a pyramid structure to expand the training data size. It adopts two core modules, the Error-To-Feature (ETF) module, and the Generate-Adaptive-Weight (GAW) module, to further improve the precision. ML-CSR calculates the inter-layer differences by the ETF module to increase the diversity of samples and obtains adaptive weights based on the layer accuracy in the GAW module. This helps ML-CSR learn more discriminative features. State-of-the-art results on the benchmark face databases validate the effectiveness of the proposed ML-CSR. Ablation experiments demonstrate that the proposed pyramid structure, ETF, and GAW module can improve the performance of ML-CSR. The code is available athttps://github.com/Zhongwenyuan98/ML-CSR.
Wenyuan Zhong, Huaxiong Li, Qinghua Hu, Yang Gao 0001, Chunlin Chen 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Adaptive Label Correlation Based Asymmetric Discrete Hashing for Cross-Modal Retrieval
abstract
Hashing methods have captured much attention for cross-modal retrieval in recent years. Most existing approaches mainly focus on preserving the semantic similarity across heterogeneous modalities in a shared Hamming subspace, while the label information and potential correlations of multi-label semantics are not fully excavated. In this article, a novel Adaptive Label correlation based asymmEtric Cross-modal Hashing method, i.e., ALECH, is proposed for cross-modal retrieval. ALECH decomposes hash learning into two steps, hash codes learning and hash functions learning. For hash codes learning, the high-order semantic label correlations are adaptively exploited to guide the latent feature learning, while simultaneously generating the binary codes in a discrete manner. The asymmetric strategy is utilized to connect the latent feature space and Hamming space, and preserve the pairwise semantic similarity. Different from other two-step methods that directly adopt simple least-squares regression to learn hash functions based on binary codes, ALECH leverages both hash codes and semantic labels for hash functions learning which further preserves the similarity. Experiments on several benchmark datasets demonstrate that the proposed ALECH method outperforms the state-of-the-art cross-hashing methods.
Huaxiong Li, Chao Zhang 0078, Xiuyi Jia, Yang Gao 0001, Chunlin Chen 0001
IEEE Trans. Knowl. Data Eng.4
2023 Weakly-Supervised Enhanced Semantic-Aware Hashing for Cross-Modal Retrieval
abstract
Owing to its query and storage efficiency, hash learning has sparked much interest for Cross-Modal Retrieval (CMR) task. Previous literatures have proved the superiority of supervised Cross-Modal Hashing (CMH) methods over unsupervised ones. Nevertheless, most existing supervised CMH methods still suffer from some limitations: 1) it is assumed that the observed labels of training data are complete and accurate, which may be impractical due to the missing and wrong class assignments in real applications, and 2) the semantic information is not fully excavated, especially for the semantic correlations among labels. To address these issues, this paper proposes a Weakly-supervised enhAnced Semantic-aware Hashing (WASH) method which simultaneously estimates the label noises and performs enhanced semantic-aware hash learning. WASH employs the low-rank and sparse decomposition to alleviate the label noises, and a high-level semantic factor as well as a semantic correlation matrix is obtained by low-rank factorization on the noise-reduced labels. The low-rank semantic factors and multi-modal features are jointly factorized into a common subspace to reduce the heterogeneity gaps, so as to enhance the semantic awareness of shared representation. In this way, the hash codes can be obtained by binarizing the shared representation with pairwise semantic similarity preserved. Experiments on several benchmark datasets verify the effectiveness of the proposed method in comparison with the state-of-the-art CMH approaches.
Chao Zhang 0078, Huaxiong Li, Yang Gao 0001, Chunlin Chen 0001
IEEE Trans. Knowl. Data Eng.3
2023 PLN: Parasitic-Like Network for Barely Supervised Medical Image Segmentation
abstract
It is known that annotations for 3D medical image segmentation tasks are laborious, time-consuming and expensive. Considering the similarities existing in inter-slice and inter-volume, we believe that the delineation way and the model architecture should be tightly coupled. In this paper, by introducing an extremely sparse annotation way of labeling only one slice per 3D image, we investigate a novel barely-supervised segmentation setting with only a few sparsely-labeled images along with a large amount of unlabeled images. To achieve this goal, we present a new parasitic-like network including a registration module (as host) and a semi-supervised segmentation module (as parasite) to deal with inter-slice label propagation and inter-volume segmentation prediction, respectively. Specifically, our parasitism mechanism effectively achieves the collaboration of these two modules through three stages of infection, development and eclosion, providing accurate pseudo-labels for training. Extensive results demonstrate that our framework is capable of achieving high performance on extremely sparse annotation tasks, e.g., we achieve Dice of 84.83% on LA dataset with only 16 labeled slices. The code is available athttps://github.com/ShumengLI/PLN.
Shumeng Li, Heng Cai, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Medical Imaging6
2023 Occluded Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) aims to match person images between the visible and near-infrared modalities. Previous VI-ReID methods are based on holistic pedestrian images and achieve excellent performance. However, in real-world scenarios, images captured by visible and near-infrared cameras usually contain occlusions. The performance of these methods degrades significantly due to the loss of information of discriminative features from the occlusion of the images. We define visible-infrared person re-identification in this occlusion scene as Occluded VI-ReID, where only partial content information of pedestrian images can be used to match images of different modalities from different cameras. In this paper, we propose a matching framework for occlusion scenes, which contains a local feature enhance module (LFEM) and a modality information fusion module (MIFM). LFEM adopts Transformer to learn features of each modality, and adjusts the importance of patches to enhance the representation ability of local features of the non-occluded areas. MIFM utilizes a co-attention mechanism to infer the correlation between each image for reducing the difference between modalities. We construct two occluded VI-ReID datasets, namely Occluded-SYSU-MM01 and Occluded-RegDB datasets. Our approach outperforms existing state-of-the-art methods on two occlusion datasets, while remains top performance on two holistic datasets.
Yujian Feng, Yimu Ji 0001, Fei Wu 0004, Guangwei Gao, Yang Gao 0001, Tianliang Liu, Shangdong Liu, Xiaoyuan Jing, Jiebo Luo 0001
IEEE Trans. Multim.5
2023 A Multilayer Framework for Online Metric Learning
abstract
Online metric learning (OML) has been widely applied in classification and retrieval. It can automatically learn a suitable metric from data by restricting similar instances to be separated from dissimilar instances with a given margin. However, the existing OML algorithms have limited performance in real-world classifications, especially, when data distributions are complex. To this end, this article proposes a multilayer framework for OML to capture the nonlinear similarities among instances. Different from the traditional OML, which can only learn one metric space, the proposed multilayer OML (MLOML) takes an OML algorithm as a metric layer and learns multiple hierarchical metric spaces, where each metric layer follows a nonlinear layer for the complicated data distribution. Moreover, the forward propagation (FP) strategy and backward propagation (BP) strategy are employed to train the hierarchical metric layers. To build a metric layer of the proposed MLOML, a new Mahalanobis-based OML (MOML) algorithm is presented based on the passive-aggressive strategy and one-pass triplet construction strategy. Furthermore, in a progressively and nonlinearly learning way, MLOML has a stronger learning ability than traditional OML in the case of limited available training data. To make the learning process more explainable and theoretically guaranteed, theoretical analysis is provided. The proposed MLOML enjoys several nice properties, indeed learns a metric progressively, and performs better on the benchmark datasets. Extensive experiments with different settings have been conducted to verify these properties of the proposed MLOML.
Wenbin Li 0006, Jing Huo, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Jiebo Luo 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 Online Passive-Aggressive Active Learning for Trapezoidal Data Streams
abstract
The idea of combining the active query strategy and the passive-aggressive (PA) update strategy in online learning can be credited to the PA active (PAA) algorithm, which has proven to be effective in learning linear classifiers from datasets with a fixed feature space. We propose a novel family of online active learning algorithms, named PAA learning for trapezoidal data streams (PAATS) and multiclass PAATS (MPAATS) (and their variants), for binary and multiclass online classification tasks on trapezoidal data streams where the feature space may expand over time. Under the context of an ever-changing feature space, we provide the theoretical analysis of the mistake bounds for both PAATS and MPAATS. Our experiments on a wide variety of benchmark datasets have confirm that the combination of the instance-regulated active query strategy and the PA update strategy is much more effective in learning from trapezoidal data streams. We have also compared PAATS with online learning with streaming features (OLSF)-the state-of-the-art approach in learning linear classifiers from trapezoidal data streams. PAATS could achieve much better classification accuracy, especially for large-scale real-world data streams.
Xiaocong Fan, Wenbin Li 0006, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2022 LaSSL: Label-Guided Self-Training for Semi-supervised Learning
abstract
The key to semi-supervised learning (SSL) is to explore adequate information to leverage the unlabeled data. Current dominant approaches aim to generate pseudo-labels on weakly augmented instances and train models on their corresponding strongly augmented variants with high-confidence results. However, such methods are limited in excluding samples with low-confidence pseudo-labels and under-utilization of the label information. In this paper, we emphasize the cruciality of the label information and propose a Label-guided Self-training approach to Semi-supervised Learning (LaSSL), which improves pseudo-label generations from two mutually boosted strategies. First, with the ground-truth labels and iteratively-polished pseudo-labels, we explore instance relations among all samples and then minimize a class-aware contrastive loss to learn discriminative feature representations that make same-class samples gathered and different-class samples scattered. Second, on top of improved feature representations, we propagate the label information to the unlabeled samples across the potential data manifold at the feature-embedding level, which can further improve the labelling of samples with reference to their neighbours. These two strategies are seamlessly integrated and mutually promoted across the whole training process. We evaluate LaSSL on several classification benchmarks under partially labeled settings and demonstrate its superiority over the state-of-the-art approaches.
Zhen Zhao 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi, Yang Gao 0001
AAAI5
2022 Few-shot Semantic Segmentation with Support-induced Graph Convolutional Network
Jie Liu 0043, Yanqi Bao, Wenzhe Yin, Yang Gao 0001, Jan-Jakob Sonke, Efstratios Gavves
BMVC5
2022 ST++: Make Self-trainingWork Better for Semi-supervised Semantic Segmentation
abstract
Self-training via pseudo labeling is a conventional, simple, and popular pipeline to leverage unlabeled data. In this work, we first construct a strong baseline of self-training (namely ST) for semi-supervised semantic segmentation via injecting strong data augmentations (SDA) on unlabeled images to alleviate overfitting noisy labels as well as decouple similar predictions between the teacher and student. With this simple mechanism, our ST outperforms all existing methods without any bells and whistles, e.g., iterative retraining. Inspired by the impressive results, we thoroughly investigate the SDA and provide some empirical analysis. Nevertheless, incorrect pseudo labels are still prone to accumulate and degrade the performance. To this end, we further propose an advanced self-training framework (namely ST++), that performs selective re-training via prioritizing reliable unlabeled images based on holistic prediction-level stability. Concretely, several model checkpoints are saved in the first stage supervised training, and the discrepancy of their predictions on the unlabeled image serves as a measurement for reliability. Our image-level selection offers holistic contextual information for learning. We demonstrate that it is more suitable for segmentation than common pixel-wise selection. As a result, ST+ further boosts the performance of our ST. Code is available at https://github.com/LiheYoung/ST-PlusPlus.
Lihe Yang, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
CVPR5
2022 MVDG: A Unified Multi-view Framework for Domain Generalization
Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
ECCV (27)4
2022 Individual Reward Assisted Multi-Agent Reinforcement Learning
abstract
In many real-world multi-agent systems, the sparsity of team rewards often makes it difficult for an algorithm to successfully learn a cooperative team policy. At present, the common way for solving this problem is to design some dense individual rewards for the agents to guide the cooperation. However, most existing works utilize individual rewards in ways that do not always promote teamwork and sometimes are even counterproductive. In this paper, we propose Individual Reward Assisted Team Policy Learning (IRAT), which learns two policies for each agent from the dense individual reward and the sparse team reward with discrepancy constraints for updating the two policies mutually. Experimental results in different scenarios, such as the Multi-Agent Particle Environment and the Google Research Football Environment, show that IRAT significantly outperforms the baseline methods and can greatly promote team policy learning without deviating from the original team objective, even when the individual rewards are misleading or conflict with the team rewards.
Yujing Hu, Weixun Wang, Chongjie Zhang, Yang Gao 0001, Jianye Hao, Tangjie Lv, Changjie Fan
ICML6
2022 HiSA: Facilitating Efficient Multi-Agent Coordination and Cooperation by Hierarchical Policy with Shared Attention
Zhirui Zhu, Guang Yang 0066, Yang Gao 0001
PRICAI (3)4
2022 DDMA: Discrepancy-Driven Multi-agent Reinforcement Learning
Yujing Hu, Pinzhuo Tian, Shaokang Dong, Yang Gao 0001
PRICAI (3)5
2022 Improving meta-learning model via meta-contrastive loss
Pinzhuo Tian, Yang Gao 0001
Frontiers Comput. Sci.2
2022 Part-facial relational and modality-style attention networks for heterogeneous face recognition
Jian Yu 0007, Yujian Feng, Yang Gao 0001
Neurocomputing4
2022 Online active classification via margin-based and feature-based label queries
Tingting Zhai, Frédéric Koriche, Yang Gao 0001, Junwu Zhu, Bin Li 0006
Mach. Learn.3
2022 Generalizable model-agnostic semantic segmentation via target-specific normalization
Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
Pattern Recognit.4
2022 Adversarial Camera Alignment Network for Unsupervised Cross-Camera Person Re-Identification
abstract
In person re-identification (Re-ID), supervised methods usually need a large amount of expensive label information, while unsupervised ones are still unable to deliver satisfactory identification performance. In this paper, we introduce a novel person Re-ID task called unsupervised cross-camera person Re-ID, which only needs the within-camera (intra-camera) label information but not cross-camera (inter-camera) labels which are more expensive to obtain. In real-world applications, the intra-camera label information can be easily captured by tracking algorithms and few manual annotations. In this situation, the main challenge becomes the distribution discrepancy across different camera views, caused by the various body pose, occlusion, image resolution, illumination conditions, and background noises in different cameras. To address this situation, we propose a novel Adversarial Camera Alignment Network (ACAN) for unsupervised cross-camera person Re-ID. It consists of the camera-alignment task and the supervised within-camera learning task. To achieve the camera alignment, we develop a Multi-Camera Adversarial Learning (MCAL) to map images of different cameras into a shared subspace. Particularly, we investigate two different schemes, including the existing GRL (i.e., gradient reversal layer) scheme and the proposed scheme called “other camera equiprobability” (OCE), to conduct the multi-camera adversarial task. Based on this shared subspace, we then leverage the within-camera labels to train the network. Extensive experiments on five large-scale datasets demonstrate the superiority of ACAN over multiple state-of-the-art unsupervised methods that take advantage of labeled source domains and generated images by GAN-based models. In particular, we verify that the proposed multi-camera adversarial task does contribute to the significant improvement.
Lei Qi 0001, Lei Wang 0001, Jing Huo, Yinghuan Shi, Xin Geng 0001, Yang Gao 0001
IEEE Trans. Circuits Syst. Video Technol.6
2022 Feature-Based Style Randomization for Domain Generalization
abstract
As a recent noticeable topic, domain generalization (DG) aims to first learn a generic model on multiple source domains and then directly generalize to an arbitrary unseen target domain without any additional adaption. In previous DG models, by generating virtual data to supplement observed source domains, the data augmentation based methods have shown its effectiveness. To simulate the possible unseen domains, most of them enrich the diversity of original data via image-level style transformation. However, we argue that the potential styles are hard to be exhaustively illustrated and fully augmented due to the limited referred styles, leading the diversity could not be always guaranteed. Unlike image-level augmentation, we in this paper develop a simple yet effective feature-based style randomization module to achieve feature-level augmentation, which can produce random styles via integrating random noise into the original style. Compared with existing image-level augmentation, our feature-level augmentation favors a more goal-oriented and sample-diverse way. Furthermore, to sufficiently explore the efficacy of the proposed module, we design a novel progressive training strategy to enable all parameters of the network to be fully trained. Extensive experiments on three standard benchmark datasets,i.e., PACS, VLCS and Office-Home, highlight the superiority of our method compared to the state-of-the-art methods.
Yue Wang 0076, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 CAST: Learning Both Geometric and Texture Style Transfers for Effective Caricature Generation
abstract
Given a photo of a subject, ability to generate a caricature image that captures distinct characteristics of the subject but with certain exaggeration of their prominent features is of fundamental importance to image processing and facial recognition. There are two main challenges in this task: shape exaggeration and style transfer. The former morphs and exaggerates key facial features of the subject, while the latter generates caricature images in a certain artistic style. In this paper, we propose a CAricature Style Transfer (CAST) framework for caricature generation. There are two modules in the proposed framework. The first is a geometric warping module. Different from the existing style transfer methods, we incorporate the Whitening and Coloring Transformation (WCT) in the geometric style transfer. The WCT is learned on photo and caricature landmarks or the caricature landmark space of a specific artist and is capable of transforming input photo landmarks to caricature landmarks. The second module is a texture style rendering module. We propose a new style transfer method by considering a semantic region-aligned style transfer via affinity constraint. Given a reference caricature image as the style reference, this module is capable of transferring styles between the same or similar semantic regions in caricatures and photos. Furthermore, it can transfer visual attributes of the reference caricatures (such as mouth shape and expressions) to the output caricatures. Experiments have shown desirable effects of the proposed method in transferring both the geometric and artistic texture styles of caricatures. Both qualitative and quantitative results show that the CAST framework is more effective compared than the state-of-the-art caricature generation methods.
Jing Huo, Xiangde Liu, Wenbin Li 0006, Yang Gao 0001, Hujun Yin, Jiebo Luo 0001
IEEE Trans. Image Process.4
2022 Crosslink-Net: Double-Branch Encoder Network via Fusing Vertical and Horizontal Convolutions for Medical Image Segmentation
abstract
Accurate image segmentation plays a crucial role in medical image analysis, yet it faces great challenges caused by various shapes, diverse sizes, and blurry boundaries. To address these difficulties, square kernel-based encoder-decoder architectures have been proposed and widely used, but their performance remains unsatisfactory. To further address these challenges, we present a novel double-branch encoder architecture. Our architecture is inspired by two observations. (1) Since the discrimination of the features learned via square convolutional kernels needs to be further improved, we propose utilizing nonsquare vertical and horizontal convolutional kernels in a double-branch encoder so that the features learned by both branches can be expected to complement each other. (2) Considering that spatial attention can help models to better focus on the target region in a large-sized image, we develop an attention loss to further emphasize the segmentation of small-sized targets. With the above two schemes, we develop a novel double-branch encoder-based segmentation framework for medical image segmentation, namely, Crosslink-Net, and validate its effectiveness on five datasets with experiments. The code is released at https://github.com/Qianyu1226/Crosslink-Net.
Qian Yu 0007, Lei Qi 0001, Yang Gao 0001, Wuzhang Wang, Yinghuan Shi
IEEE Trans. Image Process.3
2022 Deep Symmetric Adaptation Network for Cross-Modality Medical Image Segmentation
abstract
Unsupervised domain adaptation (UDA) methods have shown their promising performance in the cross-modality medical image segmentation tasks. These typical methods usually utilize a translation network to transform images from the source domain to target domain or train the pixel-level classifier merely using translated source images and original target images. However, when there exists a large domain shift between source and target domains, we argue that this asymmetric structure, to some extent, could not fully eliminate the domain gap. In this paper, we present a novel deep symmetric architecture of UDA for medical image segmentation, which consists of a segmentation sub-network, and two symmetric source and target domain translation sub-networks. To be specific, based on two translation sub-networks, we introduce a bidirectional alignment scheme via a shared encoder and two private decoders to simultaneously align features 1) from source to target domain and 2) from target to source domain, which is able to effectively mitigate the discrepancy between domains. Furthermore, for the segmentation sub-network, we train a pixel-level classifier using not only original target images and translated source images, but also original source images and translated target images, which could sufficiently leverage the semantic information from the images with different styles. Extensive experiments demonstrate that our method has remarkable advantages compared to the state-of-the-art methods in three segmentation tasks, such as cross-modality cardiac, BraTS, and abdominal multi-organ segmentation.
Xiaoting Han, Lei Qi 0001, Qian Yu 0007, Yefeng Zheng 0001, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Medical Imaging7
2022 Inconsistency-Aware Uncertainty Estimation for Semi-Supervised Medical Image Segmentation
abstract
In semi-supervised medical image segmentation, most previous works draw on the common assumption that higher entropy means higher uncertainty. In this paper, we investigate a novel method of estimating uncertainty. We observe that, when assigned different misclassification costs in a certain degree, if the segmentation result of a pixel becomes inconsistent, this pixel shows a relative uncertainty in its segmentation. Therefore, we present a new semi-supervised segmentation model, namely, conservative-radical network (CoraNet in short) based on our uncertainty estimation and separate self-training strategy. In particular, our CoraNet model consists of three major components: a conservative-radical module (CRM), a certain region segmentation network (C-SN), and an uncertain region segmentation network (UC-SN) that could be alternatively trained in an end-to-end manner. We have extensively evaluated our method on various segmentation tasks with publicly available benchmark datasets, including CT pancreas, MR endocardium, and MR multi-structures segmentation on the ACDC dataset. Compared with the current state of the art, our CoraNet has demonstrated superior performance. In addition, we have also analyzed its connection with and difference from conventional methods of uncertainty estimation in semi-supervised medical image segmentation.
Yinghuan Shi, Jian Zhang 0090, Tong Ling, Jiwen Lu, Yefeng Zheng 0001, Qian Yu 0007, Lei Qi 0001, Yang Gao 0001
IEEE Trans. Medical Imaging8
2022 CariMe: Unpaired Caricature Generation With Multiple Exaggerations
abstract
Caricature generation aims to translate real photos into caricatures with artistic styles and shape exaggerations while maintaining the identity of the subject. Different from generic image-to-image translation, drawing caricatures automatically is a more challenging task due to the existence of various spatial deformations. Previous caricature generation methods are obsessed with predicting definite image warping from a given photo while ignoring the intrinsic representation and distribution of geometric exaggerations in caricatures. This limits their ability on diverse exaggeration generation. In this paper, we generalize the caricature generation problem from instance-level warping prediction to distribution-level deformation modeling. Based on this assumption, we present the first exploration forunpaired CARIcature generation with Multiple Exaggerations (CariMe). Technically, we propose a Multi-exaggeration Warper network to learn the distribution-level mapping from photos to facial exaggerations. This makes it possible to generate diverse and reasonable exaggerations from randomly sampled warp codes given one input photo. To better represent the facial exaggeration and produce fine-grained warping, a deformation-field-based warping method is also proposed, which captures more detailed exaggerations than previous point-based warping methods. Experiments and two perceptual studies prove the superiority of our method comparing with other state-of-the-art methods, showing the improvement of our work on caricature generation. The source code is available athttps://github.com/edward3862/CariMe-pytorch.
Zheng Gu 0001, Chuanqi Dong, Jing Huo, Wenbin Li 0006, Yang Gao 0001
IEEE Trans. Multim.5
2022 Consistent Meta-Regularization for Better Meta-Knowledge in Few-Shot Learning
abstract
Recently, meta-learning provides a powerful paradigm to deal with the few-shot learning problem. However, existing meta-learning approaches ignore the prior fact that good meta-knowledge should alleviate the data inconsistency between training and test data, caused by the extremely limited data, in each few-shot learning task. Moreover, legitimately utilizing the prior understanding of meta-knowledge can lead us to design an efficient method to improve the meta-learning model. Under this circumstance, we consider the data inconsistency from the distribution perspective, making it convenient to bring in the prior fact, and propose a new consistent meta-regularization (Con-MetaReg) to help the meta-learning model learn how to reduce the data-distribution discrepancy between the training and test data. In this way, the ability of meta-knowledge on keeping the training and test data consistent is enhanced, and the performance of the meta-learning model can be further improved. The extensive analyses and experiments demonstrate that our method can indeed improve the performances of different meta-learning models in few-shot regression, classification, and fine-grained classification.
Pinzhuo Tian, Wenbin Li 0006, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2021 LoFGAN: Fusing Local Representations for Few-shot Image Generation
abstract
Given only a few available images for a novel unseen category, few-shot image generation aims to generate more data for this category. Previous works attempt to globally fuse these images by using adjustable weighted coefficients. However, there is a serious semantic misalignment between different images from a global perspective, making these works suffer from poor generation quality and diversity. To tackle this problem, we propose a novel Local-Fusion Generative Adversarial Network (LoFGAN) for fewshot image generation. Instead of using these available images as a whole, we first randomly divide them into a base image and several reference images. Next, LoFGAN matches local representations between the base and reference images based on semantic similarities, and replaces the local features with the closest related local features. In this way, LoFGAN can produce more realistic and diverse images at a more fine-grained level, and simultaneously enjoy the characteristic of semantic alignment. Furthermore, a local reconstruction loss is also proposed, which can provide better training stability and generation quality. We conduct extensive experiments on three datasets, which successfully demonstrates the effectiveness of our proposed method for few-shot image generation and downstream visual applications with limited data. Code is available at https://github.com/edward3862/LoFGAN-pytorch.
Zheng Gu 0001, Wenbin Li 0006, Jing Huo, Lei Wang 0001, Yang Gao 0001
ICCV5
2021 Manifold Alignment for Semantically Aligned Style Transfer
abstract
Most existing style transfer methods follow the assumption that styles can be represented with global statistics (e.g., Gram matrices or covariance matrices), and thus address the problem by forcing the output and style images to have similar global statistics. An alternative is the assumption of local style patterns, where algorithms are designed to swap similar local features of content and style images. However, the limitation of these existing methods is that they neglect the semantic structure of the content image which may lead to corrupted content structure in the output. In this paper, we make a new assumption that image features from the same semantic region form a manifold and an image with multiple semantic regions follows a multi-manifold distribution. Based on this assumption, the style transfer problem is formulated as aligning two multi-manifold distributions and a Manifold Alignment based Style Transfer (MAST) framework is proposed. The proposed frame-work allows semantically similar regions between the output and the style image share similar style patterns. Moreover, the proposed manifold alignment method is flexible to allow user editing or using semantic segmentation maps as guidance for style transfer. To allow the method to be applicable to photorealistic style transfer, we propose a new adaptive weight skip connection network structure to preserve the content details. Extensive experiments verify the effectiveness of the proposed framework for both artistic and photorealistic style transfer. Code is available at https://github.com/NJUHuoJing/MAST.
Jing Huo, Shiyin Jin, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yinghuan Shi, Yang Gao 0001
ICCV7
2021 Mining Latent Classes for Few-shot Segmentation
abstract
Few-shot segmentation (FSS) aims to segment unseen classes given only a few annotated samples. Existing methods suffer the problem of feature undermining, i.e., potential novel classes are treated as background during training phase. Our method aims to alleviate this problem and enhance the feature embedding on latent novel classes. In our work, we propose a novel joint-training framework. Based on conventional episodic training on support-query pairs, we introduce an additional mining branch that exploits latent novel classes via transferable sub-clusters, and a new rectification technique on both background and fore-ground categories to enforce more stable prototypes. Over and above that, our transferable sub-cluster has the ability to leverage extra unlabeled data for further feature enhancement. Extensive experiments on two FSS benchmarks demonstrate that our method outperforms previous state-of-the-art by a large margin of 3.7% mIOU on PASCAL-5iand 7.0% mIOU on COCO-20iat the cost of 74% fewer parameters and 2.5x faster inference speed. The source code is available at https://github.com/LiheYoung/MiningFSS.
Lihe Yang, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001
ICCV5
2021 Episodic Multi-agent Reinforcement Learning with Curiosity-driven Exploration
abstract
Efficient exploration in deep cooperative multi-agent reinforcement learning (MARL) still remains challenging in complex coordination problems. In this paper, we introduce a novel Episodic Multi-agent reinforcement learning with Curiosity-driven exploration, called EMC. We leverage an insight of popular factorized MARL algorithms that the ``induced" individual Q-values, i.e., the individual utility functions used for local execution, are the embeddings of local action-observation histories, and can capture the interaction between agents due to reward backpropagation during centralized training. Therefore, we use prediction errors of individual Q-values as intrinsic rewards for coordinated exploration and utilize episodic memory to exploit explored informative experience to boost policy training. As the dynamics of an agent's individual Q-value function captures the novelty of states and the influence from other agents, our intrinsic reward can induce coordinated exploration to new or promising states. We illustrate the advantages of our method by didactic examples, and demonstrate its significant outperformance over state-of-the-art MARL baselines on challenging tasks in the StarCraft II micromanagement benchmark.
Lulu Zheng, Jiarui Chen, Jiamin He, Yujing Hu, Changjie Fan, Yang Gao 0001, Chongjie Zhang
NeurIPS8
2021 Interactive medical image segmentation via a point-based interaction
Jian Zhang 0090, Yinghuan Shi, Jinquan Sun, Lei Wang 0001, Luping Zhou, Yang Gao 0001, Dinggang Shen
Artif. Intell. Medicine6
2021 Learning 3D face reconstruction from a single sketch
Jing Wu 0004, Jing Huo, Yukun Lai, Yang Gao 0001
Graph. Model.5
2021 NAS-FCOS: Efficient Search for Object Detection Architectures
Ning Wang 0020, Yang Gao 0001, Hao Chen 0041, Peng Wang 0015, Zhi Tian, Chunhua Shen, Yanning Zhang 0001
Int. J. Comput. Vis.2
2021 Mutual-information-inspired heuristics for constraint-based causal structure learning
Xiaolong Qi, Xiaocong Fan, Yang Gao 0001
Inf. Sci.5
2021 MetricUNet: Synergistic image- and voxel-level learning for precise prostate segmentation via online sampling
Kelei He, Chunfeng Lian, Ehsan Adeli-Mosabbeb, Jing Huo, Yang Gao 0001, Bing Zhang 0012, Dinggang Shen
Medical Image Anal.5
2021 A novel multiple instance learning framework for COVID-19 severity assessment via data augmentation and self-supervised learning
Zekun Li 0010, Wei Zhao 0040, Feng Shi 0001, Lei Qi 0001, Xingzhi Xie, Ying Wei 0009, Zhongxiang Ding, Yang Gao 0001, Shangjie Wu, Jun Liu 0075, Yinghuan Shi, Dinggang Shen
Medical Image Anal.8
2021 Multiset Feature Learning for Highly Imbalanced Data Classification
abstract
With the expansion of data, increasing imbalanced data has emerged. When the imbalance ratio (IR) of data is high, most existing imbalanced learning methods decline seriously in classification performance. In this paper, we systematically investigate the highly imbalanced data classification problem, and propose an uncorrelated cost-sensitive multiset learning (UCML) approach for it. Specifically, UCML first constructs multiple balanced subsets through random partition, and then employs the multiset feature learning (MFL) to learn discriminant features from the constructed multiset. To enhance the usability of each subset and deal with the non-linearity issue existed in each subset, we further propose a deep metric based UCML (DM-UCML) approach. DM-UCML introduces the generative adversarial network technique into the multiset constructing process, such that each subset can own similar distribution with the original dataset. To cope with the non-linearity issue, DM-UCML integrates deep metric learning with MFL, such that more favorable performance can be achieved. In addition, DM-UCML designs a new discriminant term to enhance the discriminability of learned metrics. Experiments on eight traditional highly class-imbalanced datasets and two large-scale datasets indicate that: the proposed approaches outperform state-of-the-art highly imbalanced learning methods and are more robust to high IR.
Xiaoyuan Jing, Xinyu Zhang 0012, Xiaoke Zhu, Fei Wu 0004, Xinge You, Yang Gao 0001, Shiguang Shan, Jing-Yu Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2021 Synergistic learning of lung lobe segmentation and hierarchical multi-instance classification for automated severity assessment of COVID-19 in CT images
Kelei He, Wei Zhao 0040, Xingzhi Xie, Mingxia Liu 0001, Zhenyu Tang 0002, Yinghuan Shi, Feng Shi 0001, Yang Gao 0001, Jun Liu 0075, Dinggang Shen
Pattern Recognit.9
2021 Local descriptor-based multi-prototype network for few-shot Learning
Zhangkai Wu, Wenbin Li 0006, Jing Huo, Yang Gao 0001
Pattern Recognit.5
2021 Crossover-Net: Leveraging vertical-horizontal crossover relation for robust medical image segmentation
Qian Yu 0007, Yang Gao 0001, Yefeng Zheng 0001, Jianbing Zhu, Yakang Dai, Yinghuan Shi
Pattern Recognit.2
2021 Pairwise Relations Oriented Discriminative Regression
abstract
Linear Regression (LR) is a popular and effective technique in pattern recognition area, which aims to find a transform matrix between source data and target data (usually label matrix). However, a binary zero-one label matrix may be too strict and inappropriate for regression. Besides, directly projecting source data to target data by one transform matrix may lose some intrinsic data information. To address these issues, this paper proposes a novel Pairwise Relations oriented Discriminative Regression (PRDR) method. In PRDR, the source data is regressed into a latent space instead of label space. To supervise the discriminative projection learning, the pairwise relations in source data space and label space are exploited in the latent space simultaneously. The pairwise label relations are transferred into the latent subspace by solving a distance-distance difference minimization problem, and the intraclass instance relations are also preserved in latent space. These two constraints ensure the pairwise similarity of data points after transformation which is beneficial for classification. By further enlarging the margins between true and false classes, PRDR is extended to a robust version, i.e., R-PRDR. An efficient algorithm is presented to solve the PRDR model. Extensive experiments on several popular image datasets demonstrate the effectiveness and efficiency of the proposed method compared with some state-of-the-art regression approaches.
Chao Zhang 0078, Huaxiong Li, Chunlin Chen 0001, Yang Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2021 MW-GAN: Multi-Warping GAN for Caricature Generation With Multi-Style Geometric Exaggeration
abstract
Given an input face photo, the goal of caricature generation is to produce stylized, exaggerated caricatures that share the same identity as the photo. It requires simultaneous style transfer and shape exaggeration with rich diversity, and meanwhile preserving the identity of the input. To address this challenging problem, we propose a novel framework called Multi-Warping GAN (MW-GAN), including a style network and a geometric network that are designed to conduct style transfer and geometric exaggeration respectively. We bridge the gap between the style/landmark space and their corresponding latent code spaces by a dual way design, so as to generate caricatures with arbitrary styles and geometric exaggeration, which can be specified either through random sampling of latent code or from a given caricature sample. Besides, we apply identity preserving loss to both image space and landmark space, leading to a great improvement in quality of generated caricatures. Experiments show that caricatures generated by MW-GAN have better quality than existing methods.
Haodi Hou, Jing Huo, Jing Wu 0004, Yukun Lai, Yang Gao 0001
IEEE Trans. Image Process.5
2021 Learning-Based Computer-Aided Prescription Model for Parkinson's Disease: A Data-Driven Perspective
abstract
In this article, we study a novel problem: "automatic prescription recommendation for PD patients." To realize this goal, we first build a dataset by collecting 1) symptoms of PD patients, and 2) their prescription drug provided by neurologists. Then, we build a novel computer-aided prescription model by learning the relation between observed symptoms and prescription drug. Finally, for the new coming patients, we could recommend (predict) suitable prescription drug on their observed symptoms by our prescription model. From the methodology part, our proposed model, namely Prescription viA Learning lAtent Symptoms (PALAS), could recommend prescription using the multi-modality representation of the data. In PALAS, a latent symptom space is learned to better model the relationship between symptoms and prescription drug, as there is a large semantic gap between them. Moreover, we present an efficient alternating optimization method for PALAS. We evaluated our method using the data collected from 136 PD patients at Nanjing Brain Hospital, which can be regarded as a large dataset in PD research community. The experimental results demonstrate the effectiveness and clinical potential of our method in this recommendation task, if compared with other competing methods.
Yinghuan Shi, Wanqi Yang, Kim-Han Thung, Hao Wang 0013, Yang Gao 0001, Dinggang Shen
IEEE J. Biomed. Health Informatics5
2021 HF-UNet: Learning Hierarchically Inter-Task Relevance in Multi-Task U-Net for Accurate Prostate Segmentation in CT Images
abstract
Accurate segmentation of the prostate is a key step in external beam radiation therapy treatments. In this paper, we tackle the challenging task of prostate segmentation in CT images by a two-stage network with 1) the first stage to fast localize, and 2) the second stage to accurately segment the prostate. To precisely segment the prostate in the second stage, we formulate prostate segmentation into a multi-task learning framework, which includes a main task to segment the prostate, and an auxiliary task to delineate the prostate boundary. Here, the second task is applied to provide additional guidance of unclear prostate boundary in CT images. Besides, the conventional multi-task deep networks typically share most of the parameters (i.e., feature representations) across all tasks, which may limit their data fitting ability, as the specificity of different tasks are inevitably ignored. By contrast, we solve them by a hierarchically-fused U-Net structure, namely HF-UNet. The HF-UNet has two complementary branches for two tasks, with the novel proposed attention-based task consistency learning block to communicate at each level between the two decoding branches. Therefore, HF-UNet endows the ability to learn hierarchically the shared representations for different tasks, and preserve the specificity of learned representations for different tasks simultaneously. We did extensive evaluations of the proposed method on a large planning CT image dataset and a benchmark prostate zonal dataset. The experimental results show HF-UNet outperforms the conventional multi-task network architectures and the state-of-the-art methods.
Kelei He, Chunfeng Lian, Bing Zhang 0012, Xin Zhang 0013, Xiaohuan Cao, Dong Nie, Yang Gao 0001, Dinggang Shen
IEEE Trans. Medical Imaging7
2021 An Optimal Algorithm for the Stochastic Bandits While Knowing the Near-Optimal Mean Reward
abstract
This brief studies a variation of the stochastic multiarmed bandit (MAB) problems, where the agent knows the a priori knowledge named the near-optimal mean reward (NoMR). In common MAB problems, an agent tries to find the optimal arm without knowing the optimal mean reward. However, in more practical applications, the agent can usually get an estimation of the optimal mean reward defined as NoMR. For instance, in an online Web advertising system based on MAB methods, a user's near-optimal average click rate (NoMR) can be roughly estimated from his/her demographic characteristics. As a result, application of the NoMR is efficient at improving the algorithm's performance. First, we formalize the stochastic MAB problem by knowing the NoMR that is in between the suboptimal mean reward and the optimal mean reward. Second, we use the cumulative regret as the performance metric for our problem, and we get that this problem's lower bound of the cumulative regret is Ω(1/∆) , where ∆ is the difference between the suboptimal mean reward and the optimal mean reward. Compared with the conventional MAB problem with the increasing logarithmic lower bound of the regret, our regret lower bound is uniform with the learning step. Third, a novel algorithm, NoMR-BANDIT, is set forth to solve this problem. In NoMR-BANDIT, the NoMR is used to design an efficient exploration strategy. In addition, we analyzed the regret's upper bound in NoMR-BANDIT and concluded that it also has a uniform upper bound of O(1/∆) , which is in the same order as the lower bound. Consequently, NoMR-BANDIT is an optimal algorithm of this problem. To enhance our method's generalization, CASCADE-BANDIT based on NoMR-BANDIT is proposed to solve the problem, where NoMR is less than the suboptimal mean reward. CASCADE-BANDIT has an upper bound of O(∆logn) , where n represents the learning step, and the order of O(∆logn) is the same with that of the conventional MAB methods. Finally, extensive experimental results demonstrated that the established NoMR-BANDIT is more efficient than the compared bandit solutions. After sufficient iterations, NOMR-BANDIT saved 10%-80% more cumulative regret than the state of the art.
Shangdong Yang, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2021 GreyReID: A Novel Two-stream Deep Framework with RGB-grey Information for Person Re-identification
abstract
In this article, we observe that most false positive images (i.e., different identities with query images) in the top ranking list usually have the similar color information with the query image in person re-identification (Re-ID). Meanwhile, when we use the greyscale images generated from RGB images to conduct the person Re-ID task, some hard query images can obtain better performance compared with using RGB images. Therefore, RGB and greyscale images seem to be complementary to each other for person Re-ID. In this article, we aim to utilize both RGB and greyscale images to improve the person Re-ID performance. To this end, we propose a novel two-stream deep neural network with RGB-grey information, which can effectively fuse RGB and greyscale feature representations to enhance the generalization ability of Re-ID. First, we convert RGB images to greyscale images in each training batch. Based on these RGB and greyscale images, we train the RGB and greyscale branches, respectively. Second, to build up connections between RGB and greyscale branches, we merge the RGB and greyscale branches into a new joint branch. Finally, we concatenate the features of all three branches as the final feature representation for Re-ID. Moreover, in the training process, we adopt the joint learning scheme to simultaneously train each branch by the independent loss function, which can enhance the generalization ability of each branch. Besides, a global loss function is utilized to further fine-tune the final concatenated feature. The extensive experiments on multiple benchmark datasets fully show that the proposed method can outperform the state-of-the-art person Re-ID methods. Furthermore, using greyscale images can indeed improve the person Re-ID performance in the proposed deep framework.
Lei Qi 0001, Lei Wang 0001, Jing Huo, Yinghuan Shi, Yang Gao 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2020 Layerwise Sparse Coding for Pruned Deep Neural Networks with Extreme Compression Ratio
abstract
Deep neural network compression is important and increasingly developed especially in resource-constrained environments, such as autonomous drones and wearable devices. Basically, we can easily and largely reduce the number of weights of a trained deep model by adopting a widely used model compression technique, e.g., pruning. In this way, two kinds of data are usually preserved for this compressed model, i.e., non-zero weights and meta-data, where meta-data is employed to help encode and decode these non-zero weights. Although we can obtain an ideally small number of non-zero weights through pruning, existing sparse matrix coding methods still need a much larger amount of meta-data (may several times larger than non-zero weights), which will be a severe bottleneck of the deploying of very deep models. To tackle this issue, we propose a layerwise sparse coding (LSC) method to maximize the compression ratio by extremely reducing the amount of meta-data. We first divide a sparse matrix into multiple small blocks and remove zero blocks, and then propose a novel signed relative index (SRI) algorithm to encode the remaining non-zero blocks (with much less meta-data). In addition, the proposed LSC performs parallel matrix multiplication without full decoding, while traditional methods cannot. Through extensive experiments, we demonstrate that LSC achieves substantial gains in pruned DNN compression (e.g., 51.03x compression ratio on ADMM-Lenet) and inference computation (i.e., time reduction and extremely less memory bandwidth), over state-of-the-art baselines.
Wenbin Li 0006, Jing Huo, Lili Yao, Yang Gao 0001
AAAI5
2020 Multi-Agent Game Abstraction via Graph Attention Neural Network
abstract
In large-scale multi-agent systems, the large number of agents and complex game relationship cause great difficulty for policy learning. Therefore, simplifying the learning process is an important research issue. In many multi-agent systems, the interactions between agents often happen locally, which means that agents neither need to coordinate with all other agents nor need to coordinate with others all the time. Traditional methods attempt to use pre-defined rules to capture the interaction relationship between agents. However, the methods cannot be directly used in a large-scale environment due to the difficulty of transforming the complex interactions between agents into rules. In this paper, we model the relationship between agents by a complete graph and propose a novel game abstraction mechanism based on two-stage attention network (G2ANet), which can indicate whether there is an interaction between two agents and the importance of the interaction. We integrate this detection mechanism into graph neural network-based multi-agent reinforcement learning for conducting game abstraction and propose two novel learning algorithms GA-Comm and GA-AC. We conduct experiments in Traffic Junction and Predator-Prey. The results indicate that the proposed methods can simplify the learning process and meanwhile get better asymptotic performance compared with state-of-the-art algorithms.
Yong Liu 0007, Weixun Wang, Yujing Hu, Jianye Hao, Xingguo Chen, Yang Gao 0001
AAAI6
2020 Differentiable Meta-Learning Model for Few-Shot Semantic Segmentation
abstract
To address the annotation scarcity issue in some cases of semantic segmentation, there have been a few attempts to develop the segmentation model in the few-shot learning paradigm. However, most existing methods only focus on the traditional 1-way segmentation setting (i.e., one image only contains a single object). This is far away from practical semantic segmentation tasks where the K-way setting (K > 1) is usually required by performing the accurate multi-object segmentation. To deal with this issue, we formulate the few-shot semantic segmentation task as a learning-based pixel classification problem, and propose a novel framework called MetaSegNet based on meta-learning. In MetaSegNet, an architecture of embedding module consisting of the global and local feature branches is developed to extract the appropriate meta-knowledge for the few-shot segmentation. Moreover, we incorporate a linear model into MetaSegNet as a base learner to directly predict the label of each pixel for the multi-object segmentation. Furthermore, our MetaSegNet can be trained by the episodic training mechanism in an end-to-end manner from scratch. Experiments on two popular semantic segmentation datasets, i.e., PASCAL VOC and COCO, reveal the effectiveness of the proposed MetaSegNet in the K-way few-shot semantic segmentation task.
Pinzhuo Tian, Zhangkai Wu, Lei Qi 0001, Lei Wang 0001, Yinghuan Shi, Yang Gao 0001
AAAI6
2020 From Few to More: Large-Scale Dynamic Multiagent Curriculum Learning
abstract
A lot of efforts have been devoted to investigating how agents can learn effectively and achieve coordination in multiagent systems. However, it is still challenging in large-scale multiagent settings due to the complex dynamics between the environment and agents and the explosion of state-action space. In this paper, we design a novel Dynamic Multiagent Curriculum Learning (DyMA-CL) to solve large-scale problems by starting from learning on a multiagent scenario with a small size and progressively increasing the number of agents. We propose three transfer mechanisms across curricula to accelerate the learning process. Moreover, due to the fact that the state dimension varies across curricula, and existing network structures cannot be applied in such a transfer setting since their network input sizes are fixed. Therefore, we design a novel network structure called Dynamic Agent-number Network (DyAN) to handle the dynamic size of the network input. Experimental results show that DyMA-CL using DyAN greatly improves the performance of large-scale multiagent learning compared with state-of-the-art deep reinforcement learning approaches. We also investigate the influence of three transfer mechanisms across curricula through extensive simulations.
Weixun Wang, Tianpei Yang, Yong Liu 0007, Jianye Hao, Xiaotian Hao, Yujing Hu, Changjie Fan, Yang Gao 0001
AAAI9
2020 NAS-FCOS: Fast Neural Architecture Search for Object Detection
abstract
The success of deep neural networks relies on significant architecture engineering. Recently neural architecture search (NAS) has emerged as a promise to greatly reduce manual effort in network design by automatically searching for optimal architectures, although typically such algorithms need an excessive amount of computational resources, e.g., a few thousand GPU-days. To date, on challenging vision tasks such as object detection, NAS, especially fast versions of NAS, is less studied. Here we propose to search for the decoder structure of object detectors with search efficiency being taken into consideration. To be more specific, we aim to efficiently search for the feature pyramid network (FPN) as well as the prediction head of a simple anchor-free object detector, namely FCOS, using a tailored reinforcement learning paradigm. With carefully designed search space, search algorithms and strategies for evaluating network quality, we are able to efficiently search a top-performing detection architecture within 4 days using 8 V100 GPUs. The discovered architecture surpasses state-of-the-art object detection models (such as Faster R-CNN, RetinaNet and FCOS) by 1.5 to 3.5 points in AP on the COCO dataset, with comparable computation complexity and memory footprint, demonstrating the efficacy of the proposed NAS for object detection.
Ning Wang 0020, Yang Gao 0001, Hao Chen 0041, Peng Wang 0015, Zhi Tian, Chunhua Shen, Yanning Zhang 0001
CVPR2
2020 Unsupervised Domain Attention Adaptation Network for Caricature Attribute Recognition
Kelei He, Jing Huo, Zheng Gu 0001, Yang Gao 0001
ECCV (8)5
2020 Automatic Data Augmentation Via Deep Reinforcement Learning for Effective Kidney Tumor Segmentation
abstract
Conventional data augmentation realized by performing simple pre-processing operations (e.g., rotation, crop, etc.) has been validated for its advantage in enhancing the performance for medical image segmentation. However, the data generated by these conventional augmentation methods are random and sometimes harmful to the subsequent segmentation. In this paper, we developed a novel automatic learning-based data augmentation method for medical image segmentation which models the augmentation task as a trial-and-error procedure using deep reinforcement learning (DRL). In our method, we innovatively combine the data augmentation module and the subsequent segmentation module in an end-to-end training manner with a consistent loss. Specifically, the best sequential combination of different basic operations is automatically learned by directly maximizing the performance improvement (i.e., Dice ratio) on the available validation set. We extensively evaluated our method on CT kidney tumor segmentation which validated the promising results of our method.
Tiexin Qin, Kelei He, Yinghuan Shi, Yang Gao 0001, Dinggang Shen
ICASSP5
2020 Action Semantics Network: Considering the Effects of Actions in Multiagent Systems
Weixun Wang, Tianpei Yang, Yong Liu 0007, Jianye Hao, Xiaotian Hao, Yujing Hu, Changjie Fan, Yang Gao 0001
ICLR9
2020 Learning Task-aware Local Representations for Few-shot Learning
abstract
Few-shot learning for visual recognition aims to adapt to novel unseen classes with only a few images. Recent work, especially the work based on low-level information, has achieved great progress. In these work, local representations (LRs) are typically employed, because LRs are more consistent among the seen and unseen classes. However, most of them are limited to an individual image-to-image or image-to-class measure manner, which cannot fully exploit the capabilities of LRs, especially in the context of a certain task. This paper proposes an Adaptive Task-aware Local Representations Network (ATL-Net) to address this limitation by introducing episodic attention, which can adaptively select the important local patches among the entire task, as the process of human recognition. We achieve much superior results on multiple benchmarks. On the miniImagenet, ATL-Net gains 0.93% and 0.88% improvements over the compared methods under the 5-way 1-shot and 5-shot settings. Moreover, ATL-Net can naturally tackle the problem that how to adaptively identify and weight the importance of different key local parts, which is the major concern of fine-grained recognition. Specifically, on the fine-grained dataset Stanford Dogs, ATL-Net outperforms the second best method with 5.39% and 9.69% gains under the 5-way 1-shot and 5-shot settings.
Chuanqi Dong, Wenbin Li 0006, Jing Huo, Zheng Gu 0001, Yang Gao 0001
IJCAI5
2020 Asymmetric Distribution Measure for Few-shot Learning
abstract
The core idea of metric-based few-shot image classification is to directly measure the relations between query images and support classes to learn transferable feature embeddings. Previous work mainly focuses on image-level feature representations, which actually cannot effectively estimate a class's distribution due to the scarcity of samples. Some recent work shows that local descriptor based representations can achieve richer representations than image-level based representations. However, such works are still based on a less effective instance-level metric, especially a symmetric metric, to measure the relation between a query image and a support class. Given the natural asymmetric relation between a query image and a support class, we argue that an asymmetric measure is more suitable for metric-based few-shot learning. To that end, we propose a novel Asymmetric Distribution Measure (ADM) network for few-shot learning by calculating a joint local and global asymmetric measure between two multivariate local distributions of a query and a class. Moreover, a task-aware Contrastive Measure Strategy (CMS) is proposed to further enhance the measure function. On popular miniImageNet and tieredImageNet, ADM can achieve the state-of-the-art results, validating our innovative design of asymmetric distribution measures for few-shot learning. The source code can be downloaded from https://github.com/WenbinLee/ADM.git.
Wenbin Li 0006, Lei Wang 0001, Jing Huo, Yinghuan Shi, Yang Gao 0001, Jiebo Luo 0001
IJCAI5
2020 Biased Feature Learning for Occlusion Invariant Face Recognition
abstract
To address the challenges posed by unknown occlusions, we propose a Biased Feature Learning (BFL) framework for occlusion-invariant face recognition. We first construct an extended dataset using a multi-scale data augmentation method. For model training, we modify the label loss to adjust the impact of normal and occluded samples. Further, we propose a biased guidance strategy to manipulate the optimization of a network so that the feature embedding space is dominated by non-occluded faces. BFL not only enhances the robustness of a network to unknown occlusions but also maintains or even improves its performance for normal faces. Experimental results demonstrate its superiority as well as the generalization capability with different network architectures and loss functions.
Chang-Bin Shao, Jing Huo, Lei Qi 0001, Zhenhua Feng 0001, Wenbin Li 0006, Chuanqi Dong, Yang Gao 0001
IJCAI7
2020 Consistent MetaReg: Alleviating Intra-task Discrepancy for Better Meta-knowledge
abstract
In the few-shot learning scenario, the data-distribution discrepancy between training data and test data in a task usually exists due to the limited data. However, most existing meta-learning approaches seldom consider this intra-task discrepancy in the meta-training phase which might deteriorate the performance. To overcome this limitation, we develop a new consistent meta-regularization method to reduce the intra-task data-distribution discrepancy. Moreover, the proposed meta-regularization method could be readily inserted into existing optimization-based meta-learning models to learn better meta-knowledge. Particularly, we provide the theoretical analysis to prove that using the proposed meta-regularization, the conventional gradient-based meta-learning method can reach the lower regret bound. The extensive experiments also demonstrate the effectiveness of our method, which indeed improves the performances of the state-of-the-art gradient-based meta-learning models in the few-shot classification task.
Pinzhuo Tian, Lei Qi 0001, Shaokang Dong, Yinghuan Shi, Yang Gao 0001
IJCAI5
2020 Cross-Domain Adversarial Autoencoder for Fine Grained Category Preserving Image Translation
abstract
Cross-domain image translation attempt to translate images from one domain to another domain, with the content of images preserved. Current approaches treat image's content as the underlying spatial structure, and translation only change image's style of color and texture. These methods can generate realistic results, but may not be able to preserve image's fine grained semantic category information and suffer from the lack of diversity in objects' shapes and viewing angles. In this paper, we propose the problem of fine grained category preserving image translation that aims at preserving image's fine grained category information in cross-domain translation. A novel framework called Cross-Domain Adversarial AutoEncoder (CDAAE) is proposed to solve the problem. CDAAE assumes that cross-domain images have shared content-latent-code space and separate style-latent-code spaces. The content latent code encodes image's basic category information, while the style latent code represents other domain-specific properties, including color, texture, shape, etc. Our experiments evaluate models from aspects of image's quality, diversity as well as category preserving ability, showing CDAAE's advantages over current methods. We also design an algorithm to apply CDAAE to domain adaptation. Experiments on benchmark datasets demonstrate that the proposed method achieves state-of-the-art results.
Haodi Hou, Jing Huo, Yang Gao 0001
IJCNN3
2020 Robust neighborhood embedding for unsupervised feature selection
Dongyi Ye, Wenbin Li 0006, Yang Gao 0001
Knowl. Based Syst.5
2020 CariGAN: Caricature generation through weakly paired adversarial learning
Wenbin Li 0006, Wei Xiong 0008, Haofu Liao, Jing Huo, Yang Gao 0001, Jiebo Luo 0001
Neural Networks5
2020 Progressive Cross-Camera Soft-Label Learning for Semi-Supervised Person Re-Identification
abstract
In this paper, we focus on the semi-supervised person re-identification (Re-ID) case, which only has the intra-camera (within-camera) labels but not inter-camera (cross-camera) labels. In real-world applications, these intra-camera labels can be readily captured by tracking algorithms or few manual annotations, when compared with cross-camera labels. In this case, it is very difficult to explore the relationships between cross-camera persons in the training stage due to the lack of cross-camera label information. To deal with this issue, we propose a novel Progressive Cross-camera Soft-label Learning (PCSL) framework for the semi-supervised person Re-ID task, which can generate cross-camera soft-labels and utilize them to optimize the network. Concretely, we calculate an affinity matrix based on person-level features and adapt them to produce the similarities between cross-camera persons (i.e., cross-camera soft-labels). To exploit these soft-labels to train the network, we investigate the weighted cross-entropy loss and the weighted triplet loss from the classification and discrimination perspectives, respectively. Particularly, the proposed framework alternately generates progressive cross-camera soft-labels and gradually improves feature representations in the whole learning course. Extensive experiments on five large-scale benchmark datasets show that PCSL significantly outperforms the state-of-the-art unsupervised methods that employ labeled source domains or the images generated by the GANs-based models. Furthermore, the proposed method even has a competitive performance with respect to deep supervised Re-ID methods.
Lei Qi 0001, Lei Wang 0001, Jing Huo, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2020 An Effective MR-Guided CT Network Training for Segmenting Prostate in CT Images
abstract
Segmentation of prostate in medical imaging data (e.g., CT, MRI, TRUS) is often considered as a critical yet challenging task for radiotherapy treatment. It is relatively easier to segment prostate from MR images than from CT images, due to better soft tissue contrast of the MR images. For segmenting prostate from CT images, most previous methods mainly used CT alone, and thus their performances are often limited by low tissue contrast in the CT images. In this article, we explore the possibility of using indirect guidance from MR images for improving prostate segmentation in the CT images. In particular, we propose a novel deep transfer learning approach, i.e., MR-guided CT network training (namely MICS-NET), which can employ MR images to help better learning of features in CT images for prostate segmentation. In MICS-NET, the guidance from MRI consists of two steps: (1) learning informative and transferable features from MRI and then transferring them to CT images in a cascade manner, and (2) adaptively transferring the prostate likelihood of MRI model (i.e., well-trained convnet by purely using MR images) with a view consistency constraint. To illustrate the effectiveness of our approach, we evaluate MICS-NET on a real CT prostate image set, with the manual delineations available as the ground truth for evaluation. Our methods generate promising segmentation results which achieve (1) six percentages higher Dice Ratio than the CT model purely using CT images and (2) comparable performance with the MRI model purely using MR images.
Wanqi Yang, Yinghuan Shi, Sanghyun Park 0004, Ming Yang 0014, Yang Gao 0001, Dinggang Shen
IEEE J. Biomed. Health Informatics5
2020 Leveraging Coupled Interaction for Multimodal Alzheimer's Disease Diagnosis
abstract
As the population becomes older worldwide, accurate computer-aided diagnosis for Alzheimer's disease (AD) in the early stage has been regarded as a crucial step for neurodegeneration care in recent years. Since it extracts the low-level features from the neuroimaging data, previous methods regarded this computer-aided diagnosis as a classification problem that ignored latent featurewise relation. However, it is known that multiple brain regions in the human brain are anatomically and functionally interlinked according to the current neuroscience perspective. Thus, it is reasonable to assume that the extracted features from different brain regions are related to each other to some extent. Also, the complementary information between different neuroimaging modalities could benefit multimodal fusion. To this end, we consider leveraging the coupled interactions in the feature level and modality level for diagnosis in this paper. First, we propose capturing the feature-level coupled interaction using a coupled feature representation. Then, to model the modality-level coupled interaction, we present two novel methods: 1) the coupled boosting (CB) that models the correlation of pairwise coupled-diversity on both inconsistently and incorrectly classified samples between different modalities and 2) the coupled metric ensemble (CME) that learns an informative feature projection from different modalities by integrating the intrarelation and interrelation of training samples. We systematically evaluated our methods with the AD neuroimaging initiative data set. By comparison with the baseline learning-based methods and the state-of-the-art methods that are specially developed for AD/MCI (mild cognitive impairment) diagnosis, our methods achieved the best performance with accuracy of 95.0% and 80.7% (CB), 94.9% and 79.9% (CME) for AD/NC (normal control), and MCI/NC identification, respectively.
Yinghuan Shi, Heung-Il Suk, Yang Gao 0001, Seong-Whan Lee, Dinggang Shen
IEEE Trans. Neural Networks Learn. Syst.3
2019 Distribution Consistency Based Covariance Metric Networks for Few-Shot Learning
abstract
Few-shot learning aims to recognize new concepts from very few examples. However, most of the existing few-shot learning methods mainly concentrate on the first-order statistic of concept representation or a fixed metric on the relation between a sample and a concept. In this work, we propose a novel end-to-end deep architecture, named Covariance Metric Networks (CovaMNet). The CovaMNet is designed to exploit both the covariance representation and covariance metric based on the distribution consistency for the few-shot classification tasks. Specifically, we construct an embedded local covariance representation to extract the second-order statistic information of each concept and describe the underlying distribution of this concept. Upon the covariance representation, we further define a new deep covariance metric to measure the consistency of distributions between query samples and new concepts. Furthermore, we employ the episodic training mechanism to train the entire network in an end-to-end manner from scratch. Extensive experiments in two tasks, generic few-shot image classification and fine-grained fewshot image classification, demonstrate the superiority of the proposed CovaMNet. The source code can be available from https://github.com/WenbinLee/CovaMNet.git.
Wenbin Li 0006, Jinglin Xu, Jing Huo, Lei Wang 0001, Yang Gao 0001, Jiebo Luo 0001
AAAI5
2019 Revisiting Local Descriptor Based Image-To-Class Measure for Few-Shot Learning
abstract
Few-shot learning in image classification aims to learn a classifier to classify images when only few training examples are available for each class. Recent work has achieved promising classification performance, where an image-level feature based measure is usually used. In this paper, we argue that a measure at such a level may not be effective enough in light of the scarcity of examples in few-shot learning. Instead, we think a local descriptor based image-to-class measure should be taken, inspired by its surprising success in the heydays of local invariant features. Specifically, building upon the recent episodic training mechanism, we propose a Deep Nearest Neighbor Neural Network (DN4 in short) and train it in an end-to-end manner. Its key difference from the literature is the replacement of the image-level feature based measure in the final layer by a local descriptor based image-to-class measure. This measure is conducted online via a k-nearest neighbor search over the deep local descriptors of convolutional feature maps. The proposed DN4 not only learns the optimal deep local descriptors for the image-to-class measure, but also utilizes the higher efficiency of such a measure in the case of example scarcity, thanks to the exchangeability of visual patterns across the images in the same class. Our work leads to a simple, effective, and computationally efficient framework for few-shot learning. Experimental study on benchmark datasets consistently shows its superiority over the related state-of-the-art, with the largest absolute improvement of 17% over the next best. The source code can be available from https://github.com/WenbinLee/DN4.git.
Wenbin Li 0006, Lei Wang 0001, Jinglin Xu, Jing Huo, Yang Gao 0001, Jiebo Luo 0001
CVPR5
2019 A Novel Unsupervised Camera-Aware Domain Adaptation Framework for Person Re-Identification
abstract
Unsupervised cross-domain person re-identification (Re-ID) faces two key issues. One is the data distribution discrepancy between source and target domains, and the other is the lack of discriminative information in target domain. From the perspective of representation learning, this paper proposes a novel end-to-end deep domain adaptation framework to address them. For the first issue, we highlight the presence of camera-level sub-domains as a unique characteristic in person Re-ID, and develop a “camera-aware” domain adaptation method via adversarial learning. With this method, the learned representation reduces distribution discrepancy not only between source and target domains but also across all cameras. For the second issue, we exploit the temporal continuity in each camera of target domain to create discriminative information. This is implemented by dynamically generating online triplets within each batch, in order to maximally take advantage of the steadily improved representation in training process. Together, the above two methods give rise to a new unsupervised domain adaptation framework for person Re-ID. Extensive experiments and ablation studies conducted on benchmark datasets demonstrate its superiority and interesting properties.
Lei Qi 0001, Lei Wang 0001, Jing Huo, Luping Zhou, Yinghuan Shi, Yang Gao 0001
ICCV6
2019 A Mask Based Deep Ranking Neural Network for Person Retrieval
abstract
Person retrieval faces many challenges including cluttered background, appearance variations (e.g., illumination, pose, occlusion) among different camera views and the similarity among different person's images. To address these issues, we put forward a novel mask based deep ranking neural network with a skipped fusing layer. Firstly, to alleviate the problem of cluttered background, masked images with only the foreground regions are incorporated as input in the proposed neural network. Secondly, to reduce the impact of the appearance variations, the multi-layer fusion scheme is developed to obtain more discriminative fine-grained information. Lastly, considering person retrieval is a special image retrieval task, we propose a novel ranking loss to optimize the whole network. The proposed ranking loss can further mitigate the interference problem of similar negative samples when producing ranking results. The extensive experiments validate the superiority of the proposed method compared with the state-of-the-art methods on many benchmark datasets.
Lei Qi 0001, Jing Huo, Lei Wang 0001, Yinghuan Shi, Yang Gao 0001
ICME5
2019 Feature-Selected and -Preserved Sampling for High-Dimensional Stream Data Summary
abstract
Along with the prosperity of the Mobile Internet, a large amount of stream data has emerged. Stream data cannot be completely stored in memory because of its massive volume and continuous arrival. Moreover, it should be accessed only once and handled in time due to the high cost of multiple accesses. Therefore, the intrinsic nature of stream data calls facilitates the development of a summary in the main memory to enable fast incremental learning and to allow working in limited time and memory. Sampling techniques are one of the commonly used methods for constructing data stream summaries. Given that the traditional random sampling algorithm deviates from the real data distribution and does not consider the true distribution of the stream data attributes, we propose a novel sampling algorithm based on feature-selected and -preserved algorithm. We first use matrix approximation to select important features in stream data. Then, the feature-preserved sampling algorithm is used to generate high-quality representative samples over a sliding window. The sampling quality of our algorithm could guarantee a high degree of consistency between the distribution of attribute values in the population (the entire data) and that in the sample. Experiments on real datasets show that the proposed algorithm can select a representative sample with high efficiency.
Qian Yu 0007, Yang Gao 0001
ICTAI4
2019 Accelerating Nash Q-Learning with Graphical Game Representation and Equilibrium Solving
abstract
Traditional Nash Q-learning algorithm generally accepts a fact that agents are tightly coupled, which brings huge computing burden. However, many multi-agent systems in the real world have sparse interactions between agents. In this paper, sparse interactions are divided into two categories: intra-group sparse interactions and inter-group sparse interactions. Previous methods can only deal with one specific type of sparse interactions. Aiming at characterizing the two categories of sparse interactions, we use a novel mathematical model called Markov graphical game. On this basis, graphical game-based Nash Q-learning is proposed to deal with different types of interactions. Experimental results show that our algorithm takes less time per episode and acquires a good policy.
Yunkai Zhuang, Xingguo Chen, Yang Gao 0001, Yujing Hu
ICTAI3
2019 Value Function Transfer for Deep Multi-Agent Reinforcement Learning Based on N-Step Returns
abstract
Many real-world problems, such as robot control and soccer game, are naturally modeled as sparse-interaction multi-agent systems. Reutilizing single-agent knowledge in multi-agent systems with sparse interactions can greatly accelerate the multi-agent learning process. Previous works rely on bisimulation metric to define Markov decision process (MDP) similarity for controlling knowledge transfer. However, bisimulation metric is costly to compute and is not suitable for high-dimensional state space problems. In this work, we propose more scalable transfer learning methods based on a novel MDP similarity concept. We start by defining the MDP similarity based on the N-step return (NSR) values of an MDP. Then, we propose two knowledge transfer methods based on deep neural networks called direct value function transfer and NSR-based value function transfer. We conduct experiments in image-based grid world, multi-agent particle environment (MPE) and Ms. Pac-Man game. The results indicate that the proposed methods can significantly accelerate multi-agent reinforcement learning and meanwhile get better asymptotic performance.
Yong Liu 0007, Yujing Hu, Yang Gao 0001, Changjie Fan
IJCAI3
2019 A Novel Deep Multi-Modal Feature Fusion Method for Celebrity Video Identification
abstract
In this paper, we develop a novel multi-modal feature fusion method for the 2019 iQIYI Celebrity Video Identification Challenge, which is held in conjunction with ACM MM 2019. The purpose of this challenge is to retrieve all the video clips of a given identity in the testing set. In this challenge, the multi-modal features of a celebrity are encouraged to be combined for a promising performance, such as face features, head features, body features, and audio features. As we know, the features from different modalities usually have their own influences on the results. To achieve better results, a novel weighted multi-modal feature fusion method is designed to obtain the final feature representation. After many experimental verification, we found that different feature fusion weights for training and testing make the method robust to multi-modal person identification. Experiments on the iQIYI-VID-2019 dataset show that our multi-modal feature fusion strategy effectively improves the accuracy of person identification. Specifically, for competition, we use a single model to get the result of 0.8952 in mAP, which ranks TOP-5 among all the competitive results.
Jianrong Chen, Jing Huo, Yinghuan Shi, Yang Gao 0001
ACM Multimedia6
2019 DeepMEF: A Deep Model Ensemble Framework for Video Based Multi-modal Person Identification
abstract
The goal of video based multi-modal person identification is to identify a person of interest using multi-modal video features, such as person's face, body, audio or head features. This task is challenging due to many factors, for example, variant body or face poses, poor face image quality, low frame resolution, etc. To address these problems, we propose a deep model ensemble framework, namely DeepMEF. Specifically, the proposed framework includes three novel modules, i.e., the video feature fusion module, the multi-modal feature fusion module and the model ensemble module. The first and second module form the basic deep model for ensemble, with the video feature fusion module fuses facial features from different frames as one. Then the multi-modal feature fusion module further fuses the face feature and features of other modalities for identification. In this work, we adopt the scene feature extracted by ourselves as the additional input of the multi-modal module. At last, the model ensemble module promotes the overall performance by combining the predictions of multiple multi-modal learners. The proposed method achieves a competitive result of 89.86% in mAP on the iQIYI-VID-2019 dataset, which helps us win the third place in the 2019 iQIYI Celebrity Video Identification Challenge.
Chuanqi Dong, Zheng Gu 0001, Zhonghao Huang, Jing Huo, Yang Gao 0001
ACM Multimedia6
2019 Concept Drift Based Multi-dimensional Data Streams Sampling Method
Xiaolong Qi, Zhirui Zhu, Yang Gao 0001
PAKDD (1)4
2019 NeoLOD: A Novel Generalized Coupled Local Outlier Detection Model Embedded Non-IID Similarity Metric
Yang Gao 0001, Jing Huo, Xiaolong Qi
PAKDD (1)2
2019 A Contextual Bandit Approach to Personalized Online Recommendation via Sparse Interactions
Hao Wang 0013, Shangdong Yang, Yang Gao 0001
PAKDD (2)4
2019 MIDCN: A Multiple Instance Deep Convolutional Network for Image Classification
Kelei He, Jing Huo, Yinghuan Shi, Yang Gao 0001, Dinggang Shen
PRICAI (1)4
2019 Learning Bayesian network structures using weakest mutual-information-first strategy
Xiaolong Qi, Xiaocong Fan, Yang Gao 0001
Int. J. Approx. Reason.3
2019 Joint multi-label classification and label correlations with missing labels and feature selection
Zhifen He, Ming Yang 0014, Yang Gao 0001, Hui-Dong Liu, Yilong Yin
Knowl. Based Syst.3
2019 Crossbar-Net: A Novel Convolutional Neural Network for Kidney Tumor Segmentation in CT Images
abstract
Due to the unpredictable location, fuzzy texture and diverse shape, accurate segmentation of the kidney tumor in CT images is an important yet challenging task. To this end, we in this paper present a cascaded trainable segmentation model termed as Crossbar-Net. Our method combines two novel schemes: (1) we originally proposed the crossbar patches, which consists of two orthogonal non-squared patches (i.e., the vertical patch and horizontal patch). The crossbar patches are able to capture both the global and local appearance information of the kidney tumors from both the vertical and horizontal directions simultaneously. (2) With the obtained crossbar patches, we iteratively train two sub-models (i.e., horizontal sub-model and vertical sub-model) in a cascaded training manner. During the training, the trained sub-models are encouraged to become more focus on the difficult parts of the tumor automatically (i.e., mis-segmented regions). Specifically, the vertical (horizontal) sub-model is required to help segment the mis-segmented regions for the horizontal (vertical) sub-model. Thus, the two sub-models could complement each other to achieve the self-improvement until convergence. In the experiment, we evaluate our method on a real CT kidney tumor dataset which is collected from 94 different patients including 3,500 CT slices. Compared with the state-of-the-art segmentation methods, the results demonstrate the superior performance of our method on the Dice similarity coefficient, true positive fraction, centroid distance and Hausdorff distance. Moreover, to exploit the generalization to other segmentation tasks, we also extend our Crossbar-Net to two related segmentation tasks: (1) cardiac segmentation in MR images and (2) breast mass segmentation in X-ray images, showing the promising results for these two tasks. Our implementation is released at https: //github.com/Qianyu1226/Crossbar-Net.
Qian Yu 0007, Yinghuan Shi, Jinquan Sun, Yang Gao 0001, Jianbing Zhu, Yakang Dai
IEEE Trans. Image Process.4
2019 Pelvic Organ Segmentation Using Distinctive Curve Guided Fully Convolutional Networks
abstract
Accurate segmentation of pelvic organs (i.e., prostate, bladder, and rectum) from CT image is crucial for effective prostate cancer radiotherapy. However, it is a challenging task due to: 1) low soft tissue contrast in CT images and 2) large shape and appearance variations of pelvic organs. In this paper, we employ a two-stage deep learning-based method, with a novel distinctive curve-guided fully convolutional network (FCN), to solve the aforementioned challenges. Specifically, the first stage is for fast and robust organ detection in the raw CT images. It is designed as a coarse segmentation network to provide region proposals for three pelvic organs. The second stage is for fine segmentation of each organ, based on the region proposal results. To better identify those indistinguishable pelvic organ boundaries, a novel morphological representation, namely, distinctive curve, is also introduced to help better conduct the precise segmentation. To implement this, in this second stage, a multi-task FCN is initially utilized to learn the distinctive curve and the segmentation map separately and then combine these two tasks to produce accurate segmentation map. The final segmentation results of all three pelvic organs are generated by a weighted max-voting strategy. We have conducted exhaustive experiments on a large and diverse pelvic CT data set for evaluating our proposed method. The experimental results demonstrate that our proposed method is accurate and robust for this challenging segmentation task, by also outperforming the state-of-the-art segmentation methods.
Kelei He, Xiaohuan Cao, Yinghuan Shi, Dong Nie, Yang Gao 0001, Dinggang Shen
IEEE Trans. Medical Imaging5
2019 Tracking Sparse Linear Classifiers
abstract
In this paper, we investigate the problem of sparse online linear classification in changing environments. We first analyze the tracking performance of standard online linear classifiers, which use gradient descent for minimizing the regularized hinge loss. The derived shifting bounds highlight the importance of choosing appropriate step sizes in the presence of concept drifts. Notably, we show that a better adaptability to concept drifts can be achieved using constant step sizes rather than the state-of-the-art decreasing step sizes. Based on these observations, we then propose a novel sparse approximated linear classifier, called sparse approximated linear classification (SALC), which uses a constant step size. In essence, SALC simply rounds small weights to zero for achieving sparsity and controls the truncation error in a principled way for achieving a low tracking regret. The degree of sparsity obtained by SALC is continuous and can be controlled by a parameter which captures the tradeoff between the sparsity of the model and the regret performance of the algorithm. Experiments on nine stationary data sets show that SALC is superior to the state-of-the-art sparse online learning algorithms, especially when the solution is required to be sparse; on seven groups of nonstationary data sets with various total shifting amounts, SALC also presents a good ability to track drifts. When wrapped with a drift detector, SALC achieves a remarkable tracking performance regardless of the total shifting amount.
Tingting Zhai, Frédéric Koriche, Hao Wang 0013, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2018 A Joint Local and Global Deep Metric Learning Method for Caricature Recognition
Wenbin Li 0006, Jing Huo, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Jiebo Luo 0001
ACCV (4)4
2018 WebCaricature: a benchmark for caricature recognition
Jing Huo, Wenbin Li 0006, Yinghuan Shi, Yang Gao 0001, Hujun Yin
BMVC4
2018 Modelling Diffusion Process by Deep Neural Networks for Image Retrieval
Yan Zhao 0019, Lei Wang 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001
BMVC5
2018 A Novel Image-Specific Transfer Approach for Prostate Segmentation in MR Images
abstract
Prostate segmentation in Magnetic Resonance (MR) Images is a significant yet challenging task for prostate cancer treatment. Most of the existing works attempted to design a global classifier for all MR images, which neglect the discrepancy of images across different patients. To this end, we propose a novel transfer approach for prostate segmentation in MR images. Firstly, an image-specific classifier is built for each training image. Secondly, a pair of dictionaries and a mapping matrix are jointly obtained by a novel Semi-Coupled Dictionary Transfer Learning (SCDTL). Finally, the classifiers on the source domain could be selectively transferred to the target domain (i.e. testing images) by the dictionaries and the mapping matrix. The evaluation demonstrates that our approach has a competitive performance compared with the state-of-the-art transfer learning methods. Moreover, the proposed transfer approach outperforms the conventional deep neural network based method.
Pinzhuo Tian, Lei Qi 0001, Yinghuan Shi, Luping Zhou, Yang Gao 0001, Dinggang Sheri
ICASSP5
2018 Feature Learning and Transfer Performance Prediction for Video Reinforcement Learning Tasks via a Siamese Convolutional Neural Network
Jinhua Song, Yang Gao 0001, Hao Wang 0013
ICONIP (1)2
2018 A Novel Two-Stage Deep Method for Mitosis Detection in Breast Cancer Histology Images
abstract
The accurate detection and counting of mitosis in breast cancer histology images is very important for computer-aided diagnosis, which is manually completed by the pathologist according to her or his clinic experience. However, this procedure is extremely time consuming and tedious. Moreover, it always results in low agreement among different pathologists. Although several computer-aided detection methods have been developed recently, they suffer from high FN (false negative) and FP (false positive) with simply treating the detection task as a binary classification problem. In this paper, we present a novel two-stage detection method with multi-scale and similarity learning convnets (MSSN). Firstly, large amount of possible candidates will be generated in the first stage in order to reduce FN (i.e., prevent treating mitosis as non-mitosis), by using the different square and non-square filters, to capture the spatial relation from different scales. Secondly, a similarity prediction model is subsequently performed on the obtained candidates for the final detection to reduce FP, which is realized by imposing a large margin constraint. On both 2014 and 2012 ICPR MITOSIS datasets, our MSSN achieved a promising result with a highest Recall (outperforming other methods by a large margin) and a comparable F-score.
Minglin Ma, Yinghuan Shi, Wenbin Li 0006, Yang Gao 0001
ICPR4
2018 A Novel Automatic Context-Based Similarity Metric for Local Outlier Detection Tasks
abstract
Local outlier detection is able to capture local behavior to improve detection performance compared to traditional global outlier detection techniques. Most existing local outlier detection methods have the fundamental assumption that attributes and attribute values are independent and identically distributed (IID). However, in many situations, since the attributes usually have an inner structure, they should not be handled equally. To address the issue above, we propose a novel automatic context-based similarity metric for local outlier detection tasks. This paper mainly includes three aspects: (i) to propose a novel approach to automatically detect the contextual attributes by capturing the attribute intra-coupling and inter-coupling; (ii) to introduce a Non-IID similarity metric to derive the kNN set and reachability distance of an object based on the attribute structure and incorporate it into local outlier detection tasks; (iii) to build a data set called EG-Permission, which is a real-world data set from an E-Government Information System for context-based local outlier detection. Results obtained from 10 data sets show the proposed approach can identify the attribute structure effectively and improve the performance in local outlier detection tasks.
Yang Gao 0001, Ruili Wang 0001
ICTAI2
2018 Online Feature Selection by Adaptive Sub-gradient Methods
Tingting Zhai, Hao Wang 0013, Frédéric Koriche, Yang Gao 0001
ECML/PKDD (2)4
2018 Joint Multi-field Siamese Recurrent Neural Network for Entity Resolution
Lei Qi 0001, Jing Huo, Hao Wang 0013, Yang Gao 0001
PRICAI5
2018 Online multi-view subspace learning via group structure analysis for visual object tracking
Wanqi Yang, Yinghuan Shi, Yang Gao 0001, Ming Yang 0014
Distributed Parallel Databases3
2018 Active learning with confidence-based answers for crowdsourcing labeling tasks
Jinhua Song, Hao Wang 0013, Yang Gao 0001, Bo An 0001
Knowl. Based Syst.3
2018 OPML: A one-pass closed-form solution for online metric learning
Wenbin Li 0006, Yang Gao 0001, Lei Wang 0001, Luping Zhou, Jing Huo, Yinghuan Shi
Pattern Recognit.2
2018 Heterogeneous Face Recognition by Margin-Based Cross-Modality Metric Learning
abstract
Heterogeneous face recognition deals with matching face images from different modalities or sources. The main challenge lies in cross-modal differences and variations and the goal is to make cross-modality separation among subjects. A margin-based cross-modality metric learning (MCM2L) method is proposed to address the problem. A cross-modality metric is defined in a common subspace where samples of two different modalities are mapped and measured. The objective is to learn such metrics that satisfy the following two constraints. The first minimizes pairwise, intrapersonal cross-modality distances. The second forces a margin between subject specific intrapersonal and interpersonal cross-modality distances. This is achieved by defining a hinge loss on triplet-based distance constraints for efficient optimization. It allows the proposed method to focus more on optimizing distances of those subjects whose intrapersonal and interpersonal distances are hard to separate. The proposed method is further extended to a kernelized MCM2L (KMCM2L). Both methods have been evaluated on an ID card face dataset and two other cross-modality benchmark datasets. Various feature extraction methods have also been incorporated in the study, including recent deep learned features. In extensive experiments and comparisons with the state-of-the-art methods, the MCM2L and KMCM2L methods achieved marked improvements in most cases.
Jing Huo, Yang Gao 0001, Yinghuan Shi, Wanqi Yang, Hujun Yin
IEEE Trans. Cybern.2
2018 Cross-Modal Metric Learning for AUC Optimization
abstract
Cross-modal metric learning (CML) deals with learning distance functions for cross-modal data matching. The existing methods mostly focus on minimizing a loss defined on sample pairs. However, the numbers of intraclass and interclass sample pairs can be highly imbalanced in many applications, and this can lead to deteriorating or unsatisfactory performances. The area under the receiver operating characteristic curve (AUC) is a more meaningful performance measure for the imbalanced distribution problem. To tackle the problem as well as to make samples from different modalities directly comparable, a CML method is presented by directly maximizing AUC. The method can be further extended to focus on optimizing partial AUC (pAUC), which is the AUC between two specific false positive rates (FPRs). This is particularly useful in certain applications where only the performances assessed within predefined false positive ranges are critical. The proposed method is formulated as a log-determinant regularized semidefinite optimization problem. For efficient optimization, a minibatch proximal point algorithm is developed. The algorithm is experimentally verified stable with the size of sampled pairs that form a minibatch at each iteration. Several data sets have been used in evaluation, including three cross-modal data sets on face recognition under various scenarios and a single modal data set, the Labeled Faces in the Wild. Results demonstrate the effectiveness of the proposed methods and marked improvements over the existing methods. Specifically, pAUC-optimized CML proves to be more competitive for performance measures such as Rank-1 and verification rate at FPR = 0.1%.
Jing Huo, Yang Gao 0001, Yinghuan Shi, Hujun Yin
IEEE Trans. Neural Networks Learn. Syst.2
2018 Incomplete-Data Oriented Multiview Dimension Reduction via Sparse Low-Rank Representation
abstract
For dimension reduction on multiview data, most of the previous studies implicitly take an assumption that all samples are completed in all views. Nevertheless, this assumption could often be violated in real applications due to the presence of noise, limited access to data, equipment malfunction, and so on. Most of the previous methods will cease to work when missing values in one or multiple views occur, thus an incomplete-data oriented dimension reduction becomes an important issue. To this end, we mathematically formulate the above-mentioned issue as sparse low-rank representation through multiview subspace (SRRS) learning to impute missing values, by jointly measuring intraview relations (via sparse low-rank representation) and interview relations (through common subspace representation). Moreover, by exploiting various subspace priors in the proposed SRRS formulation, we develop three novel dimension reduction methods for incomplete multiview data: 1) multiview subspace learning via graph embedding; 2) multiview subspace learning via structured sparsity; and 3) sparse multiview feature selection via rank minimization. For each of them, the objective function and the algorithm to solve the resulting optimization problem are elaborated, respectively. We perform extensive experiments to investigate their performance on three types of tasks including data recovery, clustering, and classification. Both two toy examples (i.e., Swiss roll and -curve) and four real-world data sets (i.e., face images, multisource news, multicamera activity, and multimodality neuroimaging data) are systematically tested. As demonstrated, our methods achieve the performance superior to that of the state-of-the-art comparable methods. Also, the results clearly show the advantage of integrating the sparsity and low-rankness over using each of them separately.
Wanqi Yang, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Ming Yang 0014
IEEE Trans. Neural Networks Learn. Syst.3
2018 Joint user knowledge and matrix factorization for recommender systems
Yonghong Yu, Yang Gao 0001, Hao Wang 0013, Ruili Wang 0001
World Wide Web2
2017 Beyond IID: Learning to Combine Non-IID Metrics for Vision Tasks
abstract
Metric learning has been widely employed, especially in various computer vision tasks, with the fundamental assumption that all samples (e.g., regions/superpixels in images/videos) are independent and identically distributed (IID). However, since the samples are usually spatially-connected or temporally-correlated with their physically-connected neighbours, they are not IID (non-IID for short), which cannot be directly handled by existing methods. Thus, we propose to learn and integrate non-IID metrics (NIME). To incorporate the non-IID spatial/temporal relations, instead of directly using non-IID features and metric learning as previous methods, NIME first builds several non-IID representations on original (non-IID) features by various graph kernel functions, and then automatically learns the metric under the best combination of various non-IID representations. NIME is applied to solve two typical computer vision tasks: interactive image segmentation and histology image identification. The results show that learning and integrating non-IID metrics improves the performance, compared to the IID methods. Moreover, our method achieves results comparable or better than that of the state-of-the-arts.
Yinghuan Shi, Wenbin Li 0006, Yang Gao 0001, Longbing Cao, Dinggang Shen
AAAI3
2017 Revisiting Metric Learning for SPD Matrix Based Visual Representation
abstract
The success of many visual recognition tasks largely depends on a good similarity measure, and distance metric learning plays an important role in this regard. Meanwhile, Symmetric Positive Definite (SPD) matrix is receiving increased attention for feature representation in multiple computer vision applications. However, distance metric learning on SPD matrices has not been sufficiently researched. A few existing works approached this by learning either d2× p or d × k transformation matrix for d× d SPD matrices. Different from these methods, this paper proposes a new member to the family of distance metric learning for SPD matrices. It learns only d parameters to adjust the eigenvalues of the SPD matrices through an efficient optimisation scheme. Also, it is shown that the proposed method can be interpreted as learning a sample-specific transformation matrix, instead of the fixed transformation matrix learned for all the samples in the existing works. The optimised d parameters can be used to massage the SPD matrices for better discrimination while still keeping them in the original space. From this perspective, the proposed method complements, rather than competes with, the existing linear-transformation-based methods, as the latter can always be applied to the output of the former to perform distance metric learning in further. The proposed method has been tested on multiple SPD-based visual representation data sets used in the literature, and the results demonstrate its interesting properties and attractive performance.
Luping Zhou, Lei Wang 0001, Jianjia Zhang, Yinghuan Shi, Yang Gao 0001
CVPR5
2017 Cost-Sensitive Alternating Direction Method of Multipliers for Large-Scale Classification
Yinghuan Shi, Xingguo Chen, Yang Gao 0001
IDEAL4
2017 Does Manual Delineation only Provide the Side Information in CT Prostate Segmentation?
Yinghuan Shi, Wanqi Yang, Yang Gao 0001, Dinggang Shen
MICCAI (3)3
2017 Exploiting Location Significance and User Authority for Point-of-Interest Recommendation
Yonghong Yu, Hao Wang 0013, Shuanzhu Sun, Yang Gao 0001
PAKDD (2)4
2017 Attributes coupling based matrix factorization for item recommendation
Yonghong Yu, Can Wang 0004, Hao Wang 0013, Yang Gao 0001
Appl. Intell.4
2017 Classification of high-dimensional evolving data streams via a resource-efficient online ensemble
Tingting Zhai, Yang Gao 0001, Hao Wang 0013, Longbing Cao
Data Min. Knowl. Discov.2
2017 Group-Based Alternating Direction Method of Multipliers for Distributed Linear Classification
abstract
The alternating direction method of multipliers (ADMM) algorithm has been widely employed for distributed machine learning tasks. However, it suffers from several limitations, e.g., a relative low convergence speed, and an expensive time cost. To this end, in this paper, a novel method, namely the group-based ADMM (GADMM), is proposed for distributed linear classification. In particular, to accelerate the convergence speed and improve global consensus, a group layer is first utilized in GADMM to divide all the slave nodes into several groups. Then, all the local variables (from the slave nodes) are gathered in the group layer to generate different group variables. Finally, by using a weighted average method, the group variables are coordinated to update the global variable (from the master node) until the solution of the global problem is reached. According to the theoretical analysis, we found that: 1) GADMM can mathematically converge at the rate , where is the number of outer iterations and 2) by using the grouping methods, GADMM can improve the convergence speed compared with the distributed ADMM framework without grouping methods. Moreover, we systematically evaluate GADMM on four publicly available LIBSVM datasets. Compared with disADMM and stochastic dual coordinate ascent with alternating direction method of multipliers-ADMM, for distributed classification, GADMM is able to reduce the number of outer iterations, which leads to faster convergence speed and better global consensus. In particular, the statistical significance test has been experimentally conducted and the results validate that GADMM can significantly save up to 30% of the total time cost (with less than 0.6% accuracy loss) compared with disADMM on large-scale datasets, e.g., webspam and epsilon.
Yang Gao 0001, Yinghuan Shi, Ruili Wang 0001
IEEE Trans. Cybern.2
2017 Multiagent Reinforcement Learning With Sparse Interactions by Negotiation and Knowledge Transfer
abstract
Reinforcement learning has significant applications for multiagent systems, especially in unknown dynamic environments. However, most multiagent reinforcement learning (MARL) algorithms suffer from such problems as exponential computation complexity in the joint state-action space, which makes it difficult to scale up to realistic multiagent problems. In this paper, a novel algorithm named negotiation-based MARL with sparse interactions (NegoSIs) is presented. In contrast to traditional sparse-interaction-based MARL algorithms, NegoSI adopts the equilibrium concept and makes it possible for agents to select the nonstrict equilibrium-dominating strategy profile (nonstrict EDSP) or meta equilibrium for their joint actions. The presented NegoSI algorithm consists of four parts: 1) the equilibrium-based framework for sparse interactions; 2) the negotiation for the equilibrium set; 3) the minimum variance method for selecting one joint action; and 4) the knowledge transfer of local Q -values. In this integrated algorithm, three techniques, i.e., unshared value functions, equilibrium solutions, and sparse interactions are adopted to achieve privacy protection, better coordination and lower computational complexity, respectively. To evaluate the performance of the presented NegoSI algorithm, two groups of experiments are carried out regarding three criteria: 1) steps of each episode; 2) rewards of each episode; and 3) average runtime. The first group of experiments is conducted using six grid world games and shows fast convergence and high scalability of the presented algorithm. Then in the second group of experiments NegoSI is applied to an intelligent warehouse problem and simulated results demonstrate the effectiveness of the presented NegoSI algorithm compared with other state-of-the-art MARL algorithms.
Luowei Zhou, Chunlin Chen 0001, Yang Gao 0001
IEEE Trans. Cybern.4
2016 Efficient Average Reward Reinforcement Learning Using Constant Shifting Values
abstract
There are two classes of average reward reinforcement learning (RL) algorithms: model-based ones that explicitly maintain MDP models and model-free ones that do not learn such models. Though model-free algorithms are known to be more efficient, they often cannot converge to optimal policies due to the perturbation of parameters. In this paper, a novel model-free algorithm is proposed, which makes use of constant shifting values (CSVs) estimated from prior knowledge. To encourage exploration during the learning process, the algorithm constantly subtracts the CSV from the rewards. A terminating condition is proposed to handle the unboundedness of Q-values caused by such substraction. The convergence of the proposed algorithm is proved under very mild assumptions. Furthermore, linear function approximation is investigated to generalize our method to handle large-scale tasks. Extensive experiments on representative MDPs and the popular game Tetris show that the proposed algorithms significantly outperform the state-of-the-art ones.
Shangdong Yang, Yang Gao 0001, Bo An 0001, Hao Wang 0013, Xingguo Chen
AAAI2
2016 A Fast Distributed Classification Algorithm for Large-Scale Imbalanced Data
abstract
The Alternating Direction Method of Multipliers (ADMM) has been developed recently for distributed classification. Nevertheless, the widely-existing class imbalance problem has not been well investigated. Furthermore, previous imbalanced classification methods lack of efforts in studying the complex imbalance problem in a distributed environment. In this paper, we consider the imbalance problem as distributed data imbalance which includes three imbalance issues: (i) within-node class imbalance, (ii)between-node class imbalance, and (iii) between-node structure imbalance. In order to adequately deal with imbalanced data as well as improve time efficiency, a novel distributed Cost-Sensitive classification algorithm via Group-based ADMM (CS-GADMM) is proposed. Briefly, CS-GADMM derives the classification problem as a series of sub-problems with within-node class imbalance. To alleviate the time delay caused by between-node class imbalance, we propose a extension of dual coordinate descent method for the sub-problem optimization. Meanwhile, for between-node structure imbalance, we discreetly study the relationship between local functions, and combine the resulting local variables intra-group to update the global variables for prediction. The experimental results on various imbalanced datasets validate that CS-GADMM could be a efficient algorithm for imbalanced classification.
Yang Gao 0001, Yinghuan Shi, Hao Wang 0013
ICDM2
2016 A Novel Multi-objective Bionic Algorithm Based on Plant Root System Growth Mechanism
Yang Gao 0001
ICIC (3)4
2016 Grouping Parallel Bayesian Network Structure Learning Algorithm Based on Variable Ordering
Xiaolong Qi, Yinghuan Shi, Hao Wang 0013, Yang Gao 0001
IDEAL4
2016 Multi-view Subspace Clustering via a Global Low-Rank Affinity Matrix
Lei Qi 0001, Yinghuan Shi, Wanqi Yang, Yang Gao 0001
IDEAL5
2016 Incremental Nonnegative Matrix Factorization Based on Matrix Sketching and k-means Clustering
Hao Wang 0013, Shangdong Yang, Yang Gao 0001
IDEAL4
2016 Ensemble of Sparse Cross-Modal Metrics for Heterogeneous Face Recognition
abstract
Heterogeneous face recognition aims to identify or verify person identity by matching facial images of different modalities. In practice, it is known that its performance is highly influenced by modality inconsistency, appearance occlusions, illumination variations and expressions. In this paper, a new method named as ensemble of sparse cross-modal metrics is proposed for tackling these challenging issues. In particular, a weak sparse cross-modal metric learning method is firstly developed to measure distances between samples of two modalities. It learns to adjust rank-one cross-modal metrics to satisfy two sets of triplet based cross-modal distance constraints in a compact form. Meanwhile, a group based feature selection is performed to enforce that features in the same position of two modalities are selected simultaneously. By neglecting features that attribute to "noise" in the face regions (eye glasses, expressions and so on), the performance of learned weak metrics can be markedly improved. Finally, an ensemble framework is incorporated to combine the results of differently learned sparse metrics into a strong one. Extensive experiments on various face datasets demonstrate the benefit of such feature selection especially when heavy occlusions exist. The proposed ensemble metric learning has been shown superiority over several state-of-the-art methods in heterogeneous face recognition.
Jing Huo, Yang Gao 0001, Yinghuan Shi, Wanqi Yang, Hujun Yin
ACM Multimedia2
2016 Joint User Knowledge and Matrix Factorization for Recommender Systems
Yonghong Yu, Yang Gao 0001, Hao Wang 0013, Ruili Wang 0001
WISE (1)2
2016 A learning-based CT prostate segmentation method via joint transductive feature selection and regression
Yinghuan Shi, Yaozong Gao, Shu Liao, Daoqiang Zhang, Yang Gao 0001, Dinggang Shen
Neurocomputing5
2015 Interactive image segmentation via cascaded metric learning
abstract
In this paper, we propose an interactive image segmentation method from a novel perspective of cascaded metric learning. Given an image with user-marked scribbles that are essentially uncertain and noisy, our method completes the segmentation task by solving a binary classification problem. Starting from the initial training samples with known class labels (i.e., regions of the image that are believed with high confidence to be foreground or background), we first find an optimal metric that can best describe the classification of these samples. After that, we classify the unlabeled samples using the learnt metric. Samples classified with high confidence are used as new training samples to refine the metric. This cycle of metric learning and classification repeats until the accomplishment of the image segmentation task. The proposed method is extensively evaluated on the MSRC image set. Experiment results show that our method outperforms the state-of-the-art methods.
Wenbin Li 0006, Yinghuan Shi, Wanqi Yang, Hao Wang 0013, Yang Gao 0001
ICIP5
2015 Semi-Automatic Segmentation of Prostate in CT Images via Coupled Feature Representation and Spatial-Constrained Transductive Lasso
abstract
Conventional learning-based methods for segmenting prostate in CT images ignore the relations among the low-level features by assuming all these features are independent. Also, their feature selection steps usually neglect the image appearance changes in different local regions of CT images. To this end, we present a novel semi-automatic learning-based prostate segmentation method in this article. For segmenting the prostate in a certain treatment image, the radiation oncologist will be first asked to take a few seconds to manually specify the first and last slices of the prostate. Then, prostate is segmented with the following two steps: (i) Estimation of 3D prostate-likelihood map to predict the likelihood of each voxel being prostate by employing the coupled feature representation, and the proposed Spatial-COnstrained Transductive LassO (SCOTO); (ii) Multi-atlases based label fusion to generate the final segmentation result by using the prostate shape information obtained from both planning and previous treatment images. The major contribution of the proposed method mainly includes: (i) incorporating radiation oncologist's manual specification to aid segmentation, (ii) adopting coupled features to relax previous assumption of feature independency for voxel representation, and (iii) developing SCOTO for joint feature selection across different local regions. The experimental result shows that the proposed method outperforms the state-of-the-art methods in a real-world prostate CT dataset, consisting of 24 patients with totally 330 images, all of which were manually delineated by the radiation oncologist for performance evaluation. Moreover, our method is also clinically feasible, since the segmentation performance can be improved by just requiring the radiation oncologist to spend only a few seconds for manual specification of ending slices in the current treatment CT image.
Yinghuan Shi, Yaozong Gao, Shu Liao, Daoqiang Zhang, Yang Gao 0001, Dinggang Shen
IEEE Trans. Pattern Anal. Mach. Intell.5
2015 Multiagent Reinforcement Learning With Unshared Value Functions
abstract
One important approach of multiagent reinforcement learning (MARL) is equilibrium-based MARL, which is a combination of reinforcement learning and game theory. Most existing algorithms involve computationally expensive calculation of mixed strategy equilibria and require agents to replicate the other agents' value functions for equilibrium computing in each state. This is unrealistic since agents may not be willing to share such information due to privacy or safety concerns. This paper aims to develop novel and efficient MARL algorithms without the need for agents to share value functions. First, we adopt pure strategy equilibrium solution concepts instead of mixed strategy equilibria given that a mixed strategy equilibrium is often computationally expensive. In this paper, three types of pure strategy profiles are utilized as equilibrium solution concepts: pure strategy Nash equilibrium, equilibrium-dominating strategy profile, and nonstrict equilibrium-dominating strategy profile. The latter two solution concepts are strategy profiles from which agents can gain higher payoffs than one or more pure strategy Nash equilibria. Theoretical analysis shows that these strategy profiles are symmetric meta equilibria. Second, we propose a multistep negotiation process for finding pure strategy equilibria since value functions are not shared among agents. By putting these together, we propose a novel MARL algorithm called negotiation-based Q-learning (NegoQ). Experiments are first conducted in grid-world games, which are widely used to evaluate MARL algorithms. In these games, NegoQ learns equilibrium policies and runs significantly faster than existing MARL algorithms (correlated Q-learning and Nash Q-learning). Surprisingly, we find that NegoQ also performs well in team Markov games such as pursuit games, as compared with team-task-oriented MARL algorithms (such as friend Q-learning and distributed Q-learning).
Yujing Hu, Yang Gao 0001, Bo An 0001
IEEE Trans. Cybern.2
2015 Accelerating Multiagent Reinforcement Learning by Equilibrium Transfer
abstract
An important approach in multiagent reinforcement learning (MARL) is equilibrium-based MARL, which adopts equilibrium solution concepts in game theory and requires agents to play equilibrium strategies at each state. However, most existing equilibrium-based MARL algorithms cannot scale due to a large number of computationally expensive equilibrium computations (e.g., computing Nash equilibria is PPAD-hard) during learning. For the first time, this paper finds that during the learning process of equilibrium-based MARL, the one-shot games corresponding to each state's successive visits often have the same or similar equilibria (for some states more than 90% of games corresponding to successive visits have similar equilibria). Inspired by this observation, this paper proposes to use equilibrium transfer to accelerate equilibrium-based MARL. The key idea of equilibrium transfer is to reuse previously computed equilibria when each agent has a small incentive to deviate. By introducing transfer loss and transfer condition, a novel framework called equilibrium transfer-based MARL is proposed. We prove that although equilibrium transfer brings transfer loss, equilibrium-based MARL algorithms can still converge to an equilibrium policy under certain assumptions. Experimental results in widely used benchmarks (e.g., grid world game, soccer game, and wall game) show that the proposed framework: 1) not only significantly accelerates equilibrium-based MARL (up to 96.7% reduction in learning time), but also achieves higher average rewards than algorithms without equilibrium transfer and 2) scales significantly better than algorithms without equilibrium transfer when the state/action space grows and the number of agents increases.
Yujing Hu, Yang Gao 0001, Bo An 0001
IEEE Trans. Cybern.2
2015 MRM-Lasso: A Sparse Multiview Feature Selection Method via Low-Rank Analysis
abstract
Learning about multiview data involves many applications, such as video understanding, image classification, and social media. However, when the data dimension increases dramatically, it is important but very challenging to remove redundant features in multiview feature selection. In this paper, we propose a novel feature selection algorithm, multiview rank minimization-based Lasso (MRM-Lasso), which jointly utilizes Lasso for sparse feature selection and rank minimization for learning relevant patterns across views. Instead of simply integrating multiple Lasso from view level, we focus on the performance of sample-level (sample significance) and introduce pattern-specific weights into MRM-Lasso. The weights are utilized to measure the contribution of each sample to the labels in the current view. In addition, the latent correlation across different views is successfully captured by learning a low-rank matrix consisting of pattern-specific weights. The alternating direction method of multipliers is applied to optimize the proposed MRM-Lasso. Experiments on four real-life data sets show that features selected by MRM-Lasso have better multiview classification performance than the baselines. Moreover, pattern-specific weights are demonstrated to be significant for learning about multiview data, compared with view-specific weights.
Wanqi Yang, Yang Gao 0001, Yinghuan Shi, Longbing Cao
IEEE Trans. Neural Networks Learn. Syst.2
2014 Joint Coupled-Feature Representation and Coupled Boosting for AD Diagnosis
abstract
Recently, there has been a great interest in computer-aided Alzheimer's Disease (AD) and Mild Cognitive Impairment (MCI) diagnosis. Previous learning based methods defined the diagnosis process as a classification task and directly used the low-level features extracted from neuroimaging data without considering relations among them. However, from a neuroscience point of view, it's well known that a human brain is a complex system that multiple brain regions are anatomically connected and functionally interact with each other. Therefore, it is natural to hypothesize that the low-level features extracted from neuroimaging data are related to each other in some ways. To this end, in this paper, we first devise a coupled feature representation by utilizing intra-coupled and inter-coupled interaction relationship. Regarding multi-modal data fusion, we propose a novel coupled boosting algorithm that analyzes the pairwise coupled-diversity correlation between modalities. Specifically, we formulate a new weight updating function, which considers both incorrectly and inconsistently classified samples. In our experiments on the ADNI dataset, the proposed method presented the best performance with accuracies of 94.7% and 80.1% for AD vs. Normal Control (NC) and MCI vs. NC classifications, respectively, outperforming the competing methods and the state-of-the-art methods.
Yinghuan Shi, Heung-Il Suk, Yang Gao 0001, Dinggang Shen
CVPR3
2014 A Novel Ego-Centered Academic Community Detection Approach via Factor Graph Model
Yusheng Jia, Yang Gao 0001, Wanqi Yang, Jing Huo, Yinghuan Shi
IDEAL2
2014 mPadal: a joint local-and-global multi-view feature selection method for activity recognition
Wanqi Yang, Yang Gao 0001, Longbing Cao, Ming Yang 0014, Yinghuan Shi
Appl. Intell.2
2014 Multi-Instance Dictionary Learning for Detecting Abnormal Events in Surveillance Videos
abstract
In this paper, a novel method termed Multi-Instance Dictionary Learning (MIDL) is presented for detecting abnormal events in crowded video scenes. With respect to multi-instance learning, each event (video clip) in videos is modeled as a bag containing several sub-events (local observations); while each sub-event is regarded as an instance. The MIDL jointly learns a dictionary for sparse representations of sub-events (instances) and multi-instance classifiers for classifying events into normal or abnormal. We further adopt three different multi-instance models, yielding the Max-Pooling-based MIDL (MP-MIDL), Instance-based MIDL (Inst-MIDL) and Bag-based MIDL (Bag-MIDL), for detecting both global and local abnormalities. The MP-MIDL classifies observed events by using bag features extracted via max-pooling over sparse representations. The Inst-MIDL and Bag-MIDL classify observed events by the predicted values of corresponding instances. The proposed MIDL is evaluated and compared with the state-of-the-art methods for abnormal event detection on the UMN (for global abnormalities) and the UCSD (for local abnormalities) datasets and results show that the proposed MP-MIDL and Bag-MIDL achieve either comparable or improved detection performances. The proposed MIDL method is also compared with other multi-instance learning methods on the task and superior results are obtained by the MP-MIDL scheme.
Jing Huo, Yang Gao 0001, Wanqi Yang, Hujun Yin
Int. J. Neural Syst.2
2014 Structurally Enhanced Incremental Neural Learning for Image Classification with Subgraph Extraction
abstract
In this paper, a structurally enhanced incremental neural learning technique is proposed to learn a discriminative codebook representation of images for effective image classification applications. In order to accommodate the relationships such as structures and distributions among visual words into the codebook learning process, we develop an online codebook graph learning method based on a novel structurally enhanced incremental learning technique, called as "visualization-induced self-organized incremental neural network (ViSOINN)". The hidden structural information in the images is embedded into the graph representation evolving dynamically with the adaptive and competitive learning mechanism. Afterwards, image features can be coded using a sub-graph extraction process based on the learned codebook graph, and a classifier is subsequently used to complete the image classification task. Compared with other codebook learning algorithms originated from the classical Bag-of-Features (BoF) model, ViSOINN holds the following advantages: (1) it learns codebook efficiently and effectively from a small training set; (2) it models the relationships among visual words in metric scaling fashion, so preserving high discriminative power; (3) it automatically learns the codebook without a fixed pre-defined size; and (4) it enhances and preserves better the structure of the data. These characteristics help to improve image classification performance and make it more suitable for handling large-scale image classification tasks. Experimental results on the widely used Caltech-101 and Caltech-256 benchmark datasets demonstrate that ViSOINN achieves markedly improved performance and reduces the computational cost considerably.
Yang Gao 0001, Hujun Yin
Int. J. Neural Syst.3
2014 Local histogram specification for face recognition under varying lighting conditions
Hui-Dong Liu, Ming Yang 0014, Yang Gao 0001, Chunyan Cui
Image Vis. Comput.3
2014 Bilinear discriminative dictionary learning for face recognition
Hui-Dong Liu, Ming Yang 0014, Yang Gao 0001, Yilong Yin
Pattern Recognit.3
2014 Fast Local Histogram Specification
abstract
Local histogram specification (LHS) is a useful technique for image processing. However, LHS faces a critical computational challenge when it is applied to high-resolution high-precision images. The calculation of the values in the cumulative distribution function (CDF) and the mapped value for the central pixel in each sliding window is time consuming with the computational complexity O(s + L) of the state-of-theart techniques, where s is the side length of the square window and L is the number of gray levels. In this paper, we propose a fast algorithm for LHS, called fast local histogram specification (FLHS). FLHS reduces the complexity of calculating the CDF value for the central pixel in each sliding window to O(s + √L), and the time complexity for the mapping procedure in each window to O(log L). This results in the overall time complexity of LHS reduced from O(s+L) to O(s+√L) in each sliding window. Theoretical analysis shows that the newly developed algorithm is efficient. Experimental results on the 8-bit and high-resolution high-precision (16-bit) images demonstrate the efficiency of our proposed algorithm.
Hui-Dong Liu, Ming Yang 0014, Yang Gao 0001, Longbing Cao
IEEE Trans. Circuits Syst. Video Technol.3
2014 Pairwise Costs in Semisupervised Discriminant Analysis for Face Recognition
abstract
In recent years, face recognition is being recognized as a cost-sensitive learning problem. Many cost-sensitive classifiers have been proposed. However, no sufficient attention is paid to the research on cost-sensitive dimensionality reduction, especially on the cost-sensitive semisupervised dimensionality reduction. To the best of our knowledge, cost sensitive semisupervised discriminant analysis (CS3DA) may be the first work. CS3DA first uses the sparse representation to infer a soft label for unlabeled sample and then learns the projection direction by incorporating misclassification costs into both labeled and unlabeled data. Although CS3DA reduces the loss of misclassification, it has two major drawbacks: 1) the sparsity is not a feature of face recognition, and therefore sparse approximations may not deliver the robustness or performance desired and 2) CS3DA is not proven to satisfy the minimal misclassification loss criterion. In this paper, we embed pairwise costs in semisupervised discriminant analysis (PCSDA) for face recognition. PCSDA first uses a simple l2approach to predict the label of unlabeled data, and then learns the projection direction by embedding pairwise costs in both labeled and unlabeled data. Compared with CS3DA, PCSDA has three major advantages: 1) l2approach is more accurate and robust than sparse representation for face recognition; 2) we prove that CS3DA approximates the pairwise Bayesian risk only when the classes are balanced and without outliers in face data sets; and 3) PCSDA approximates the pairwise Bayesian risk considering the class imbalance problem and outliers in face recognition. Hence, the projection direction obtained by using PCSDA can be more discriminative, immunes to outliers and class imbalance problem. The experimental results on AR, PIE, ORL, and extended Yale B data sets demonstrate the effectiveness of PCSDA.
Jianwu Wan, Ming Yang 0014, Yang Gao 0001, Yin-Juan Chen
IEEE Trans. Inf. Forensics Secur.3
2013 Prostate Segmentation in CT Images via Spatial-Constrained Transductive Lasso
abstract
Accurate prostate segmentation in CT images is a significant yet challenging task for image guided radiotherapy. In this paper, a novel semi-automated prostate segmentation method is presented. Specifically, to segment the prostate in the current treatment image, the physician first takes a few seconds to manually specify the first and last slices of the prostate in the image space. Then, the prostate is segmented automatically by the proposed two steps: (i) The first step of prostate-likelihood estimation to predict the prostate likelihood for each voxel in the current treatment image, aiming to generate the 3-D prostate-likelihood map by the proposed Spatial-COnstrained Transductive LassO (SCOTO), (ii) The second step of multi-atlases based label fusion to generate the final segmentation result by using the prostate shape information obtained from the planning and previous treatment images. The experimental result shows that the proposed method outperforms several state-of-the-art methods on prostate segmentation in a real prostate CT dataset, consisting of 24 patients with 330 images. Moreover, it is also clinically feasible since our method just requires the physician to spend a few seconds on manual specification of the first and last slices of the prostate.
Yinghuan Shi, Shu Liao, Yaozong Gao, Daoqiang Zhang, Yang Gao 0001, Dinggang Shen
CVPR5
2013 Image Super Resolution via Visual Prior Based Digital Image Characteristics
Yusheng Jia, Wanqi Yang, Yang Gao 0001, Hujun Yin, Yinghuan Shi
IDEAL3
2013 Voting-XCSc: A Consensus Clustering Method via Learning Classifier System
Liqiang Qian, Yinghuan Shi, Yang Gao 0001, Hujun Yin
IDEAL3
2013 A Coupled Clustering Approach for Items Recommendation
Yonghong Yu, Can Wang 0004, Yang Gao 0001, Longbing Cao, Xixi Chen
PAKDD (2)3
2013 Erratum: A Coupled Clustering Approach for Items Recommendation
Yonghong Yu, Can Wang 0004, Yang Gao 0001, Longbing Cao
PAKDD (2)3
2013 Transductive cost-sensitive lung cancer image classification
Yinghuan Shi, Yang Gao 0001, Ruili Wang 0001, Dong Wang 0038
Appl. Intell.2
2013 TRASMIL: A local anomaly detection framework based on trajectory segmentation and multi-instance learning
Wanqi Yang, Yang Gao 0001, Longbing Cao
Comput. Vis. Image Underst.2
2013 Visual word coding based on difference maximization
Ling-Yan Pan, Yang Gao 0001, Guang-Nan He, Yao Zhang 0002
Neurocomputing3
2013 Online Selective Kernel-Based Temporal Difference Learning
abstract
In this paper, an online selective kernel-based temporal difference (OSKTD) learning algorithm is proposed to deal with large scale and/or continuous reinforcement learning problems. OSKTD includes two online procedures: online sparsification and parameter updating for the selective kernel-based value function. A new sparsification method (i.e., a kernel distance-based online sparsification method) is proposed based on selective ensemble learning, which is computationally less complex compared with other sparsification methods. With the proposed sparsification method, the sparsified dictionary of samples is constructed online by checking if a sample needs to be added to the sparsified dictionary. In addition, based on local validity, a selective kernel-based value function is proposed to select the best samples from the sample dictionary for the selective kernel-based value function approximator. The parameters of the selective kernel-based value function are iteratively updated by using the temporal difference (TD) learning algorithm combined with the gradient descent technique. The complexity of the online sparsification procedure in the OSKTD algorithm is O(n). In addition, two typical experiments (Maze and Mountain Car) are used to compare with both traditional and up-to-date O(n) algorithms (GTD, GTD2, and TDC using the kernel-based value function), and the results demonstrate the effectiveness of our proposed algorithm. In the Maze problem, OSKTD converges to an optimal policy and converges faster than both traditional and up-to-date algorithms. In the Mountain Car problem, OSKTD converges, requires less computation time compared with other sparsification methods, gets a better local optima than the traditional algorithms, and converges much faster than the up-to-date algorithms. In addition, OSKTD can reach a competitive ultimate optima compared with the up-to-date algorithms.
Xingguo Chen, Yang Gao 0001, Ruili Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2013 Adaptive support framework for wisdom web of things
Yang Gao 0001, Mufeng Lin, Ruili Wang 0001
World Wide Web1
2012 Abnormal Event Detection via Multi-Instance Dictionary Learning
Jing Huo, Yang Gao 0001, Wanqi Yang, Hujun Yin
IDEAL2
2012 Codebook Quantization for Image Classification Using Incremental Neural Learning and Subgraph Extraction
Yang Gao 0001, Yao Zhang 0002, Ying-Chun Cao
IDEAL3
2012 Self-paced dictionary learning for image classification
abstract
Image classification is an important research task in multimedia content analysis and processing. Learning a compact dictionary easying to derive sparse representation is one of the focused issues in the state-of-the-art image classification framework. Most existing dictionary learning approaches assign equal importance to all training samples, which in fact have different complexity in terms of sparse representation. Meanwhile, the contextual information "hidden" in different samples is ignored as well. In this paper, we propose a self-paced dictionary learning algorithm in order to accommodate the "hidden" information of the samples into the learning procedure, which uses the easy samples to train the dictionary first, and then iteratively introduces more complex samples in the remaining training procedure until the entire training data are all easy samples. The algorithm adaptively chooses the easy samples in each iteration, while the learned dictionary in the previous iteration is in turn used as a basis for the current iteration. This strategy implicitly takes advantage of the contextual relationships among training samples. The number of the chosen samples in each iteration is determined by an adaptive threshold function proposed in this paper. Experimental results on benchmark datasets, including Caltech-101 and 15-Scene, show that our algorithm leads to better dictionary representation and classification performance than the baseline methods.
Yang Gao 0001
ACM Multimedia3
2011 P2LSA and P2LSA+: Two Paralleled Probabilistic Latent Semantic Analysis Algorithms Based on the MapReduce Model
Yang Gao 0001, Yinghuan Shi, Lin Shang 0001, Ruili Wang 0001
IDEAL2
2011 Image Region Segmentation Based on Color Coherence Quantization
Guang-Nan He, Yao Zhang 0002, Yang Gao 0001, Lin Shang 0001
IEA/AIE (1)4
2011 Xcsc: a Novel Approach to Clustering with Extended Classifier System
abstract
In this paper, we propose a novel approach to clustering noisy and complex data sets based on the eXtend Classifier Systems (XCS). The proposed approach, termed XCSc, has three main processes: (a) a learning process to evolve the rule population, (b) a rule compacting process to remove redundant rules after the learning process, and (c) a rule merging process to deal with the overlapping rules that commonly occur between the clusters. In the first process, we have modified the clustering mechanisms of the current available XCS and developed a new accelerate learning method to improve the quality of the evolved rule population. In the second process, an effective rule compacting algorithm is utilized. The rule merging process is based on our newly proposed agglomerative hierarchical rule merging algorithm, which comprises the following steps: (i) all the generated rules are modeled by a graph, with each rule representing a node; (ii) the vertices in the graph are merged to form a number of sub-graphs (i.e. rule clusters) under some pre-defined criteria, which generates the final rule set to represent the clusters; (iii) each data is re-checked and assigned to a cluster that it belongs to, guided by the final rule set. In our experiments, we compared the proposed XCSc with CHAMELEON, a benchmark algorithm well known for its excellent performance, on a number of challenging data sets. The results show that the proposed approach outperforms CHAMELEON in the successful rate, and also demonstrates good stability.
Liangdong Shi, Yinghuan Shi, Yang Gao 0001, Lin Shang 0001
Int. J. Neural Syst.3
2010 Learning classifier system using both labeled and unlabeled data
abstract
In this paper, we propose a Semi-UCS, which is an extension of the classical sUpervised Classifier System (UCS) [1] for semi-supervised learning tasks. A UCS works under a supervised learning scheme and uses only labeled data to train the system. In the Semi-UCS, we add an additional semi-supervised learning component to the original UCS, enablingthe LCS to learn from both labeled and unlabeled data. We provide three methods of how this semi-supervised learning component can be implemented: self-learning method, k-NN distance measure method and tri-training method. The Semi-UCS enlarges UCS's' application domains into semi-supervised settings and is a great addition to the LCS's model family. Experimental results on benchmark data sets of UCI repository have shown that Semi-UCS reaches a good performance for semi-supervised learning tasks.
Chi Su, Yang Gao 0001, Chun Cao
GECCO2
2010 Real-Time Abnormal Event Detection in Complicated Scenes
abstract
In this paper, we proposed a novel real-time abnormal event detection framework that requires a short training period and has a fast processing speed. Our approach is based on phase correlation and our newly developed spatial-temporal co-occurrence Gaussian mixture models (STCOG)with the following steps: (i) a frame is divided into non-overlapping local regions; (ii) phase correlation is used to estimate the motion vectors between successive two frames for all corresponding local regions, and (iii) STCOG is used to model normal events and detect abnormal events if any deviation from the trained STCOG is found. Our proposed approach is also able to update the parameters incrementally and can be applied in complicated scenes. The proposed approach outperforms previous ones in terms of shorter training periods and lower computational complexity.
Yinghuan Shi, Yang Gao 0001, Ruili Wang 0001
ICPR2
2010 CASTLE: a computer-assisted stress teaching and learning environment for learners of English as a second language
Jingli Lu, Ruili Wang 0001, Liyanage C. De Silva, Yang Gao 0001
INTERSPEECH4
2010 RL-DOT: A Reinforcement Learning NPC Team for Playing Domination Games
abstract
In this paper, we describe the design of reinforcement-learning-based domination team (RL-DOT), a nonplayer character (NPC) team for playing Unreal Tournament (UT) Domination games. In RL-DOT, there is a commander NPC and several soldier NPCs. The running process of RL-DOT consists of several decision cycles. In each decision cycle, the commander NPC makes a decision of troop distribution and, according to that decision, sends action orders to other soldier NPCs. Each soldier NPC tries to accomplish its task in a goal-directed way, i.e., decomposing the final ultimate task (attacking or defending a domination point) into basic actions (such as running and shooting) that are directly supported by UT application programming interfaces (APIs). We use a Q-learning-style algorithm to learn the optimal decision-making policy. We carefully choose some opponent policies for our illustrative experiments. In these experiments, RL-DOT shows a distinct learning characteristic, which illustrates its efficiency in playing UT Domination games.
Hao Wang 0013, Yang Gao 0001, Xingguo Chen
IEEE Trans. Comput. Intell. AI Games2
2009 Apply ant colony optimization to Tetris
abstract
Tetris is a falling block game where the player's objective is to arrange a sequence of different shaped tetrominoes smoothly in order to survive. In the intelligence games, agent imitates the real player and chooses the best move based on a linear value function. In this paper, we apply Ant Colony Optimization (ACO) method to learn the weights of the function, trying to search an optimal weight-path in the weight graph. We use dynamic heuristic to prevent premature convergence to local optima. Our experimental result is better than most of traditional reinforcement learning methods.
Xingguo Chen, Hao Wang 0013, Weiwei Wang 0002, Yinghuan Shi, Yang Gao 0001
GECCO5
2009 Syllable nucleus Durations Estimation using Linear Regression based ensemble model
abstract
Unlike conventional automatic continuous speech segmentation models that deal with each boundary time-mark individually, in this paper, we propose an interval-data-based linear regression model for syllable nucleus durations estimation (LRM-DE), which treats syllable boundary time-marks in pairs. This characteristic of LRM-DE makes it more suitable for estimating syllable durations for English sentences, which can be used for sentence stress detection. LRM-DE combines the outcomes of multiple base automatic speech segmentation machines (ASMs) to generate final boundary time-marks that miminize the average distance of the predicted and reference boundary-pairs of syllable nuclei. Experimental results show that on TIMIT dataset, LRM-DE reduces the average difference between the predicted syllable nucleus durations and their reference ones from 13.64 ms (the best result of a single ASM) to 11.81 ms. Also, LRM-DE improves the syllable nucleus segmentation accuracy from 81.59% to 83.98% within a tolerance of 20 ms.
Jingli Lu, Ruili Wang 0001, Liyanage C. De Silva, Yang Gao 0001
ICASSP4
2009 Prediction-Based Prefetching to Support VCR-like Operations in Gossip-Based P2P VoD Systems
abstract
Supporting free VCR-like operations in P2P VoD streaming systems is challenging. The uncertainty of frequent VCR operations makes it difficult to provide high quality realtime streaming services over distributed self-organized P2P overlay networks. Recently, prefetching has emerged as a promising approach to smooth the streaming quality. However, how to efficiently and effectively prefetch suitable segments is still an open issue. In this paper, we propose PREP, a PREdiction-based Prefetching scheme to support VCR-like operations over gossip-based P2P on-demand streaming systems. By employing the reinforcement learning technique, PREP transforms users' streaming service procedure into a set of abstract states and presents an online prediction model to predict a user's VCR behavior via analyzing the large volumes of user viewing logs collected on the tracker. We further present a distributed data scheduling algorithm to proactively prefetch segments according to the predicted VCR behavior. Moreover, PREP takes advantage of the inherent peer collaboration of gossip protocol to optimize the response latency. Through comprehensive simulations, we demonstrate the efficiency of PREP by gaining the accumulated hit ratio close to 75% while reducing the response latency close to 70% with only less than 15% extra stress on the server side.
Tianyin Xu, Weiwei Wang 0002, Sanglu Lu, Yang Gao 0001
ICPADS6
2009 Clustering with XCS and Agglomerative Rule Merging
Liangdong Shi, Yinghuan Shi, Yang Gao 0001
IDEAL3
2009 3D Scene Analysis Using UIMA Framework
Yang Gao 0001, Yao Zhang 0002, Chunsheng Yang
IEA/AIE4
2009 Detecting Abnormal Events via Hierarchical Dirichlet Processes
Xian-Xing Zhang, Yang Gao 0001, Derek Hao Hu
PAKDD3
2007 Learning classifier system ensemble and compact rule set
abstract
This paper presents a learning classifier system ensemble for knowledge discovery from incremental data. The new ensemble was designed with a two-level architecture to improve the generalization ability. The new incoming cases are first bootstrapped to generate data as inputs to the first level classical learning classifier systems. The second level contains a plurality-vote module to determine the final classification by combining the classification results of the first level learning classifier systems. Each learning classifier system in the first level consists of two major modules, a genetic algorithm module for facilitating rule-discovery and a reinforcement learning module for adjusting the strength of the corresponding rules when rewards are received from the environment. We propose a revised Wilson's compact rule algorithm for generation of the compact rule set from the population set to improve the readability of the model. Two experiments were conducted. One was data mining of medical data and the other was steganalysis of images. The experimental results have shown that the new ensemble produced better performance on incremental data mining and better generalization than the single learning classifier system and other supervised learning methods. The results also showed that the compact rules were more interpretable.
Yang Gao 0001, Joshua Zhexue Huang
Connect. Sci.1
2007 A two-layered multi-agent reinforcement learning model and algorithm
Ben-Nian Wang, Yang Gao 0001, Zhaoqian Chen, Junyuan Xie, Shifu Chen
J. Netw. Comput. Appl.2
2005 An Efficient Adaptive Focused Crawler Based on Ontology Learning
abstract
The enormous growth of the World Wide Web has made it important to perform resource discovery efficiently. Consequently, several new ideas have been proposed; among them a key technique is focused crawling which is able to crawl particular topical portions of the World Wide Web quickly without having to explore all Web pages. In this paper, we present an intelligent focused crawler algorithm in which we embed ontology to evaluate the page's relevance to the topic. Compared with other algorithms using domain knowledge, our algorithm can evolve the ontology automatically during crawl process. Considering the instinct characteristics of the ontology, propagation has also been imported to accelerate the evolution of the ontology. We applied our approaches in several tasks and provided an empirical evaluation which has shown promising results.
Yang Gao 0001, Jianmei Yang, Bin Luo 0003
HIS2
2005 Towards Efficient Selection of Web Services with Reinforcement Learning Process
abstract
As an emerging technology for implementing Web services over the Internet, mobile agent model has several advantages over the traditional RFC model. However, with the popularity of distributed networks (e.g. Internet), Web service providers tend to rely on external resources to complete certain tasks. This definitely increases the difficulty in locating appropriate service providers according to clients' requirements in the new scenario. To address this issue, we propose a reinforcement learning process based on the mobile agent model, which makes agents more efficient and intelligent in selecting Web service providers. Finally, an implementation of our prototype is presented
Dongjun Cai, Zongwei Luo, Yang Gao 0001
ICTAI4
2005 Applying Neural Network to Reinforcement Learning in Continuous Spaces
Dongli Wang, Yang Gao 0001
ISNN (1)2
2005 Adaptive grid job scheduling with genetic algorithms
Yang Gao 0001, Hongqiang Rong, Joshua Zhexue Huang
Future Gener. Comput. Syst.1
2004 Mining Web Sequential Patterns Using Reinforcement Learning
Yang Gao 0001, Guifeng Tang, Shifu Chen
APWeb2
2004 Believability Based Iterated Belief Revision
Yang Gao 0001, Zhaoqian Chen, Shifu Chen
PRICAI2