Wenbin Li 0006

dblp:27/1736-6 · DBLP profile ↗
← Back
66ranked-venue papers
11as first author
49since 2021 · last 2026
0000-0002-0935-7124ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 10 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 6 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
YearPublicationVenuePosition
2026 A continual learning framework with long-term and multiple short-term memory networks
Shangge Liu, Lei Wang 0001, Rui Yan 0005, Jing Huo, Wenbin Li 0006, Yang Gao 0001
Neural Networks5
2026 AMPL: An adaptive meta-prompt learner for few-shot image classification
Zhiping Wu, Lian Huai, Zeyu Shangguan, Lei Wang 0001, Jing Huo, Wenbin Li 0006, Yang Gao 0001, Xingqun Jiang
Neural Networks7
2026 Unsupervised few-shot learning with object-aware and attribute-consistent augmentation
Zhiping Wu, Lian Huai, Zeyu Shangguan, Wenbin Li 0006, Yang Gao 0001, Xingqun Jiang
Pattern Recognit.5
2026 FISN: FInding Spatial Neighborhoods for Generalizable Novel View Synthesis
abstract
We present FISN, a generalizable novel view synthesis algorithm that enables feedforward inference of Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS) from reference images. Unlike existing work that either separately model the 3D feature space on each view or process multiview reference features by 3D-point-based view aggregation, FISN integrates multi-reference 3D cost volumes into a unified high-dimensional entity. Specifically, we reconceptualize the generalizable novel view synthesis task as a feedforward process of FInding Spatial Neighborhoods across this unified 4D feature space, comprising both view and spatial dimensions, and introduce View-Spatial Convolutions for direct 4D feature aggregation. This enhances the correlation among multiview neighboring points in a window-to-window manner and incorporates 3D spatial awareness. However, this approach poses two intertwined challenges: high computational expense for high-dimensional features and degraded rendering performance with low-resolution features. To address these challenges, FISN constructs a new efficient convolution paradigm, Decomposable View-Spatial Convolution, which includes a Spatial Cross Decomposition strategy as well as a Feature Compression and Upscaling module. This paradigm maintains multiview geometric consistency better than existing decomposition methods and achieves a balance between efficiency and fine-grained spatial features. Furthermore, by integrating Depth Refinement modules based on this paradigm, FISN further improves global depth understanding. Comprehensive evaluations on mainstream datasets and benchmarks demonstrate that FISN achieves state-of-the-art performance for both NeRF and 3DGS, and remains robust in challenging scenarios where existing 3DGS-based methods struggle, such as those with noisy poses or dense references. The code will be released soon.
Yanqi Bao, Tianyu Ding, Jing Huo, Wenbin Li 0006, Yang Gao 0001
IEEE Trans. Vis. Comput. Graph.4
2025 Text and Image Are Mutually Beneficial: Enhancing Training-Free Few-Shot Classification with CLIP
abstract
Contrastive Language-Image Pretraining (CLIP) has been widely used in vision tasks. Notably, CLIP has demonstrated promising performance in few-shot learning (FSL). However, existing CLIP-based methods in training-free FSL (i.e., without the requirement of additional training) mainly learn different modalities independently, leading to two essential issues: 1) severe anomalous match in image modality; 2) varying quality of generated text prompts. To address these issues, we build a mutual guidance mechanism, that introduces an Image-Guided-Text (IGT) component to rectify varying quality of text prompts through image representations, and a Text-Guided-Image (TGI) component to mitigate the anomalous match of image modality through text representations. By integrating IGT and TGI, we adopt a perspective of Text-Image Mutual guidance Optimization, proposing TIMO. Extensive experiments show that TIMO significantly outperforms the state-of-the-art (SOTA) training-free method. Additionally, by exploring the extent of mutual guidance, we propose an enhanced variant, TIMO-S, which even surpasses the best training-required methods by 0.33% with approximately ×100 less time cost.
Yayuan Li, Jintao Guo, Lei Qi 0001, Wenbin Li 0006, Yinghuan Shi
AAAI4
2025 Enhancing Few-Shot Class-Incremental Learning via Training-Free Bi-Level Modality Calibration
abstract
Few-shot Class-Incremental Learning (FSCIL) challenges models to adapt to new classes with limited samples, presenting greater difficulties than traditional class-incremental learning. While existing approaches rely heavily on visual models and require additional training during base or incremental phases, we propose a training-free framework that leverages pre-trained visual-language models like CLIP. At the core of our approach is a novel Bi-level Modality Calibration (BiMC) strategy. Our framework initially performs intra-modal calibration, combining LLM-generated fine-grained category descriptions with visual prototypes from the base session to achieve precise classifier estimation. This is further complemented by inter-modal calibration that fuses pre-trained linguistic knowledge with task-specific visual priors to mitigate modality-specific biases. To enhance prediction robustness, we introduce additional metrics and strategies that maximize the utilization of limited data. Extensive experimental results demonstrate that our approach significantly outperforms existing methods. Code is available at: https://github.com/yychen016/BiMC.
Tianyu Ding, Lei Wang 0001, Jing Huo, Yang Gao 0001, Wenbin Li 0006
CVPR6
2025 Adapting In-Domain Few-Shot Segmentation to New Domains Without Source Domain Retraining
abstract
Cross-domain few-shot segmentation (CD-FSS) aims to segment objects of novel classes in new domains, which is often challenging due to the diverse characteristics of target domains and the limited availability of support data. Most CD-FSS methods redesign and retrain in-domain FSS models using abundant base data from the source domain, which are effective but costly to train. To address these issues, we propose adapting informative model structures of the well-trained FSS model for target domains by learning domain characteristics from few-shot labeled support samples during inference, thereby eliminating the need for source domain retraining. Specifically, we first adaptively identify domain-specific model structures by measuring parameter importance using a novel structure Fisher score in a data-dependent manner. Then, we progressively train the selected informative model structures with hierarchically constructed training samples, progressing from fewer to more support shots. The resulting Informative Structure Adaptation (ISA) method effectively addresses domain shifts and equips existing well-trained in-domain FSS models with flexible adaptation capabilities for new domains, eliminating the need to redesign or retrain CD-FSS models on base data. Extensive experiments validate the effectiveness of our method, demonstrating superior performance across multiple CD-FSS benchmarks. Codes are at https://github.com/fanq15/ISA.
Nian Liu 0002, Hisham Cholakkal, Rao Muhammad Anwer, Wenbin Li 0006, Yang Gao 0001
ICCV6
2025 Dynamic Multi-Layer Null Space Projection for Vision-Language Continual Learning
Borui Kang, Lei Wang 0001, Zhiping Wu, Yawen Li 0001, Yang Gao 0001, Wenbin Li 0006
ICCV7
2025 Reducing Variance of Stochastic Optimization for Approximating Nash Equilibria in Normal-Form Games
abstract
Nash equilibrium (NE) plays an important role in game theory. How to efficiently compute an NE in NFGs is challenging due to its complexity and non-convex optimization property. Machine Learning (ML), the cornerstone of modern artificial intelligence, has demonstrated remarkable empirical performance across various applications including non-convex optimization. To leverage non-convex stochastic optimization techniques from ML for approximating an NE, various loss functions have been proposed. Among these, only one loss function is unbiased, allowing for unbiased estimation under the sampled play. Unfortunately, this loss function suffers from high variance, which degrades the convergence rate. To improve the convergence rate by mitigating the high variance associated with the existing unbiased loss function, we propose a novel surrogate loss function named Nash Advantage Loss (NAL). NAL is theoretically proved unbiased and exhibits significantly lower variance than the existing unbiased loss function. Experimental results demonstrate that the algorithm minimizing NAL achieves a significantly faster empirical convergence rates compared to other algorithms, while also reducing the variance of estimated loss value by several orders of magnitude.
Linjian Meng, Wubing Chen, Wenbin Li 0006, Tianpei Yang, Youzhi Zhang 0001, Yang Gao 0001
ICML3
2025 Efficient Last-Iterate Convergence in Solving Extensive-Form Games
abstract
To establish last-iterate convergence for Counterfactual Regret Minimization (CFR) algorithms in learning a Nash equilibrium (NE) of extensive-form games (EFGs), recent studies reformulate learning an NE of the original EFG as learning the NEs of a sequence of (perturbed) regularized EFGs. Hence, proving last-iterate convergence in solving the original EFG reduces to proving last-iterate convergence in solving (perturbed) regularized EFGs. However, these studies only establish last-iterate convergence for Online Mirror Descent (OMD)-based CFR algorithms instead of Regret Matching (RM)-based CFR algorithms in solving perturbed regularized EFGs, resulting in a poor empirical convergence rate, as RM-based CFR algorithms typically outperform OMD-based CFR algorithms. In addition, as solving multiple perturbed regularized EFGs is required, fine-tuning across multiple perturbed regularized EFGs is infeasible, making parameter-free algorithms highly desirable. This paper show that CFR$^+$, a classical parameter-free RM-based CFR algorithm, achieves last-iterate convergence in learning an NE of perturbed regularized EFGs. This is the first parameter-free last-iterate convergence for RM-based CFR algorithms in perturbed regularized EFGs. Leveraging CFR$^+$ to solve perturbed regularized EFGs, we get Reward Transformation CFR$^+$ (RTCFR$^+$). Importantly, we extend prior work on the parameter-free property of CFR$^+$, enhancing its stability, which is vital for the empirical convergence of RTCFR$^+$. Experiments show that RTCFR$^+$ exhibits a significantly faster empirical convergence rate than existing algorithms that achieve theoretical last-iterate convergence. Interestingly, RTCFR$^+$ show performance no worse than average-iterate convergence CFR algorithms. It is the first last-iterate convergence algorithm to achieve such performance. Our code is available at https://github.com/menglinjian/NeurIPS-2025-RTCFR.
Linjian Meng, Tianpei Yang, Youzhi Zhang 0001, Zhenxing Ge, Shangdong Yang, Tianyu Ding, Wenbin Li 0006, Bo An 0001, Yang Gao 0001
NeurIPS7
2025 Last-Iterate Convergence of Smooth Regret Matching$^+$ Variants in Learning Nash Equilibria
abstract
Regret Matching$^+$ (RM$^+$) variants are widely used to build superhuman Poker AIs, yet few studies investigate their last-iterate convergence in learning a Nash equilibrium (NE). Although their last-iterate convergence is established for games satisfying the Minty Variational Inequality (MVI), no studies have demonstrated that these algorithms achieve such convergence in the broader class of games satisfying the weak MVI. A key challenge in proving last-iterate convergence for RM$^+$ variants in games satisfying the weak MVI is that even if the game's loss gradient satisfies the weak MVI, RM$^+$ variants operate on a transformed loss feedback which does not satisfy the weak MVI. To provide last-iterate convergence for RM$^+$ variants, we introduce a concise yet novel proof paradigm that involves: (i) transforming an RM$^+$ variant into an Online Mirror Descent (OMD) instance that updates within the original strategy space of the game to recover the weak MVI, and (ii) showing last-iterate convergence by proving the distance between accumulated regrets converges to zero via the recovered weak MVI of the feedback. Inspired by our proof paradigm, we propose Smooth Optimistic Gradient Based RM$^+$ (SOGRM$^+$) and show that it achieves last-iterate and finite-time best-iterate convergence in learning an NE of games satisfying the weak MVI, the weakest condition among all known RM$^+$ variants. Experiments show that SOGRM$^+$ significantly outperforms other algorithms. Our code is available at https://github.com/menglinjian/NeurIPS-2025-SOGRM.
Linjian Meng, Youzhi Zhang 0001, Zhenxing Ge, Tianyu Ding, Shangdong Yang, Wenbin Li 0006, Yang Gao 0001
NeurIPS7
2025 REST: A resolution preserving network for photorealistic style transfer via semantic distillation
Jing Huo, Zheng Gu 0001, Jiulin Zhang, Xiangde Liu, Shiyin Jin, Pinzhuo Tian, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001
Comput. Vis. Image Underst.7
2025 Coordinating Multi-Agent Reinforcement Learning via Dual Collaborative Constraints
Shaokang Dong, Shangdong Yang, Yujing Hu, Wenbin Li 0006, Yang Gao 0001
Neural Networks5
2025 ONNXPruner: ONNX-Based General Model Pruning Adapter
abstract
Recent advancements in model pruning have focused on developing new algorithms and improving upon benchmarks. However, the practical application of these algorithms across various models and platforms remains a significant challenge. To address this challenge, we propose ONNXPruner, a versatile pruning adapter designed for the ONNX format models. ONNXPruner streamlines the adaptation process across diverse deep learning frameworks and hardware platforms. A novel aspect of ONNXPruner is its use of node association trees, which automatically adapt to various model architectures. These trees clarify the structural relationships between nodes, guiding the pruning process, particularly highlighting the impact on interconnected nodes. Furthermore, we introduce a tree-level evaluation method. By leveraging node association trees, this method allows for a comprehensive analysis beyond traditional single-node evaluations, enhancing pruning performance without the need for extra operations. Experiments across multiple models and datasets confirm ONNXPruner's strong adaptability and increased efficacy. Our work aims to advance the practical application of model pruning.
Dongdong Ren, Wenbin Li 0006, Tianyu Ding, Lei Wang 0001, Jing Huo, Hongbing Pan, Yang Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 3D Gaussian Splatting: Survey, Technologies, Challenges, and Opportunities
abstract
3D Gaussian Splatting (3DGS) has emerged as a prominent technique with the potential to become a mainstream method for 3D representations. It can effectively transform multi-view images into explicit 3D Gaussian through efficient training, and achieve real-time rendering of novel views. This survey aims to analyze existing 3DGS-related works from multiple intersecting perspectives, including related tasks, technologies, challenges, and opportunities. The primary objective is to provide newcomers with a rapid understanding of the field and to assist researchers in methodically organizing existing technologies and challenges. Specifically, we delve into the optimization, application, and extension of 3DGS, categorizing them based on their focuses or motivations. Additionally, we summarize and classify nine types of technical modules and corresponding improvements identified in existing works. Based on these analyses, we further examine the common challenges and technologies across various tasks, proposing potential research opportunities.
Yanqi Bao, Tianyu Ding, Jing Huo, Yaoli Liu, Wenbin Li 0006, Yang Gao 0001, Jiebo Luo 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Leveraging Frequency Analysis for Image Denoising Network Pruning
abstract
As a common model compression technique, network pruning is widely used to reduce storage and computational cost of deep models in the resource-constrained regime. However, most current pruning methods are designed for high-level vision tasks, with few developed for low-level vision tasks. We observed that the norm-based pruning criterion, originally designed for high-level vision tasks, is highly unsuitable for low-level image denoising networks. This difference arises because image denoising networks pursue distinct feature granularities and goals compared to typical high-level vision tasks. To address this issue, we propose a novel filter evaluation method, termed High-Frequency Components Pruning (HFCP), specifically tailored for image denoising network pruning. HFCP assesses filter importance based on high-frequency components. To the best of our knowledge, this is the first pruning method designed specifically for image denoising tasks, straightforward and applicable to various types of noise. Furthermore, HFCP enhances the pruned model's high-frequency information content with high reliability and interpretability. This facilitates the network's ability to distinguish high-frequency signals from noise. We comprehensively analyzed multiple image denoising networks and validated HFCP's effectiveness across four mainstream networks.
Dongdong Ren, Wenbin Li 0006, Jing Huo, Lei Wang 0001, Hongbing Pan, Yang Gao 0001
IEEE Trans. Image Process.2
2025 WeakMedSAM: Weakly-Supervised Medical Image Segmentation via SAM With Sub-Class Exploration and Prompt Affinity Mining
abstract
We have witnessed remarkable progress in foundation models in vision tasks. Currently, several recent works have utilized the segmenting anything model (SAM) to boost the segmentation performance in medical images, where most of them focus on training an adaptor for fine-tuning a large amount of pixel-wise annotated medical images following a fully supervised manner. In this paper, to reduce the labeling cost, we investigate a novel weakly-supervised SAM-based segmentation model, namely WeakMedSAM. Specifically, our proposed WeakMedSAM contains two modules: 1) to mitigate severe co-occurrence in medical images, a sub-class exploration module is introduced to learn accurate feature representations. 2) to improve the quality of the class activation maps, our prompt affinity mining module utilizes the prompt capability of SAM to obtain an affinity map for random-walk refinement. Our method can be applied to any SAM-like backbone, and we conduct experiments with SAMUS and EfficientSAM. The experimental results on three popularly-used benchmark datasets, i.e., BraTS 2019, AbdomenCT-1K, and MSD Cardiac dataset, show the promising results of our proposed WeakMedSAM. Our code is available at https://github.com/wanghr64/WeakMedSAM.
Lian Huai, Wenbin Li 0006, Lei Qi 0001, Xingqun Jiang, Yinghuan Shi
IEEE Trans. Medical Imaging3
2025 Dictionary Based Generative Adversarial Network for Multi-Collection Style Transfer
abstract
Most collection-based style transfer methods require training a separate model for each individual collection of styles, making the extension to multiple collections of styles less flexible. Besides, the existing collection-based methods are also less flexible in extending to new style collections in a continual manner. To address these issues, we propose a novelMultI-Dictionary Generative Adversarial Network framework (MID-GAN)for multi-collection style transfer. Specifically, we design a multi-dictionary architecture within a GAN, with each dictionary consisting of a set of local style codes for a specific style collection. Benefiting from the local style codes used in the dictionary, a stylization module with aligned skip connections is further proposed, which can better preserve both the local details and the overall image structure. The dictionary design allows a flexible extension to new style collections by readily adding new dictionaries and we propose a continual training strategy that can both preserve the style transfer ability of old styles and achieve good transfer results for newly added styles. Extensive experiments are performed to show that the proposed method is better than existing collection-based style transfer methods. We also demonstrate the proposed method can generate diverse meaningful style transfer results of the same style collection.
Jing Huo, Shiyin Jin, Jiashen Li, Pinzhuo Tian, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001
IEEE Trans. Multim.5
2025 Multi-Task Multi-Agent Reinforcement Learning With Interaction and Task Representations
abstract
Multi-task multi-agent reinforcement learning (MT-MARL) is capable of leveraging useful knowledge across multiple related tasks to improve performance on any single task. While recent studies have tentatively achieved this by learning independent policies on a shared representation space, we pinpoint that further advancements can be realized by explicitly characterizing agent interactions within these multi-agent tasks and identifying task relations for selective reuse. To this end, this article proposes Representing Interactions and Tasks (RIT), a novel MT-MARL algorithm that characterizes both intra-task agent interactions and inter-task task relations. Specifically, for characterizing agent interactions, RIT presents the interactive value decomposition to explicitly take the dependency among agents into policy learning. Theoretical analysis demonstrates that the learned utility value of each agent approximates its Shapley value, thus representing agent interactions. Moreover, we learn task representations based on per-agent local trajectories, which assess task similarities and accordingly identify task relations. As a result, RIT facilitates the effective transfer of interaction knowledge across similar multi-agent tasks. Structurally, RIT develops universal policy structure for scalable multi-task policy learning. We evaluate RIT against multiple state-of-the-art baselines in various cooperative tasks, and its significant performance under both multi-task and zero-shot settings demonstrates its effectiveness.
Shaokang Dong, Shangdong Yang, Yujing Hu, Tianyu Ding, Wenbin Li 0006, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Exploiting Inter-sample and Inter-feature Relations in Dataset Distillation
abstract
Dataset distillation has emerged as a promising approach in deep learning, enabling efficient training with small synthetic datasets derived from larger real ones. Particularly, distribution matching-based distillation methods attract attention thanks to its effectiveness and low computational cost. However, these methods face two primary limitations: the dispersed feature distribution within the same class in synthetic datasets, reducing class discrim-ination, and an exclusive focus on mean feature consistency, lacking precision and comprehensiveness. To address these challenges, we introduce two novel constraints: a class centralization constraint and a covariance matching constraint. The class centralization constraint aims to enhance class discrimination by more closely clustering samples within classes. The covariance matching constraint seeks to achieve more accurate feature distribution matching between real and synthetic datasets through local feature covariance matrices, particularly beneficial when sample sizes are much smaller than the number of features. Experiments demonstrate notable improvements with these constraints, yielding performance boosts of up to 6.6% on CIFAR10, 2.9% on SVHN, 2.5% on CIFAR100, and 2.5% on TinyImageNet, compared to the state-of-the-art relevant methods. In addition, our method maintains robust performance in cross-architecture settings, with a maximum performance drop of 1.7% on four architectures. Code is avail-able at https://github.com/VincenDen/IID.
Wenxiao Deng, Wenbin Li 0006, Tianyu Ding, Lei Wang 0001, Kuihua Huang, Jing Huo, Yang Gao 0001
CVPR2
2024 Dynamic Replay Training for Class-Incremental Learning
abstract
Replay-based methods for Class-Incremental Learning (CIL) typically employ new classes and a limited subset of old classes stored in memory to facilitate the model training. However, these methods often lead to class imbalance and catastrophic forgetting, where the model forgets previously learned tasks. While some studies have attempted to address such a class imbalance issue, they do not fully consider the dynamic nature of forgetting in the model. In this paper, we propose a novel method called Dynamic Replay Training (DRT) to address the dynamic forgetting of previously learned tasks by the model. DRT replays memory data with dynamically changing frequencies, offering a novel perspective to tackle catastrophic forgetting and class imbalance. The proposed method is evaluated on CIFAR-100 and ImageNet-100 in various settings, showing significant improvements of 8.28% and 4.53% in terms of classification accuracy compared to the baseline method on the two datasets, respectively.
Dongdong Ren, Chenglei Peng, Jing Huo, Wenbin Li 0006, Yang Gao 0001
ICASSP5
2024 InsertNeRF: Instilling Generalizability into NeRF with HyperNet Modules
abstract
Generalizing Neural Radiance Fields (NeRF) to new scenes is a significant challenge that existing approaches struggle to address without extensive modifications to vanilla NeRF framework. We introduce **InsertNeRF**, a method for **INS**tilling g**E**ne**R**alizabili**T**y into **NeRF**. By utilizing multiple plug-and-play HyperNet modules, InsertNeRF dynamically tailors NeRF's weights to specific reference scenes, transforming multi-scale sampling-aware features into scene-specific representations. This novel design allows for more accurate and efficient representations of complex appearances and geometries. Experiments show that this method not only achieves superior generalization performance but also provides a flexible pathway for integration with other NeRF-like systems, even in sparse input settings. Code will be available at: https://github.com/bbbbby-99/InsertNeRF.
Yanqi Bao, Tianyu Ding, Jing Huo, Wenbin Li 0006, Yang Gao 0001
ICLR4
2024 Safe and Robust Subgame Exploitation in Imperfect Information Games
abstract
Opponent exploitation is an important task for players to exploit the weaknesses of others in games. Existing approaches mainly focus on balancing between exploitation and exploitability but are often vulnerable to modeling errors and deceptive adversaries. To address this problem, our paper offers a novel perspective on the safety of opponent exploitation, named Adaptation Safety. This concept leverages the insight that strategies, even those not explicitly aimed at opponent exploitation, may inherently be exploitable due to computational complexities, rendering traditional safety overly rigorous. In contrast, adaptation safety requires that the strategy should not be more exploitable than it would be in scenarios where opponent exploitation is not considered. Building on such adaptation safety, we further propose an Opponent eXploitation Search (OX-Search) framework by incorporating real-time search techniques for efficient online opponent exploitation. Moreover, we provide theoretical analyses to show the adaptation safety and robust exploitation of OX-Search, even with inaccurate opponent models. Empirical evaluations in popular poker games demonstrate OX-Search’s superiority in both exploitability and exploitation compared to previous methods.
Zhenxing Ge, Tianyu Ding, Linjian Meng, Bo An 0001, Wenbin Li 0006, Yang Gao 0001
ICML6
2024 Rare Fungi Image Classification Based on Few-Shot Learning and Data Augmentation
Jiayi Hao, Yulin Feng, Wenbin Li 0006, Jiebo Luo 0001
ICPR (16)3
2024 Task-Aware Few-Shot Image Generation via Dynamic Local Distribution Estimation and Sampling
Zheng Gu 0001, Wenbin Li 0006, Tianyu Ding, Jing Huo, Kuihua Huang, Yang Gao 0001
PRCV (2)2
2024 Making the Primary Task Primary: Boosting Few-Shot Classification by Gradient-Biased Multi-task Learning
Yunchen Wu, Boyao Shi, Jing Huo, Wenbin Li 0006, Yang Gao 0001, Tinghao Yu
PRCV (1)4
2024 Towards efficient image and video style transfer via distillation and learnable feature transformation
Jing Huo, Meihao Kong, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001
Comput. Vis. Image Underst.3
2024 Egoism, utilitarianism and egalitarianism in multi-agent reinforcement learning
Shaokang Dong, Shangdong Yang, Bo An 0001, Wenbin Li 0006, Yang Gao 0001
Neural Networks5
2024 Attention-based investigation and solution to the trade-off issue of adversarial training
Chang-Bin Shao, Wenbin Li 0006, Jing Huo, Zhenhua Feng 0001, Yang Gao 0001
Neural Networks2
2024 WToE: Learning When to Explore in Multiagent Reinforcement Learning
abstract
Existing multiagent exploration works focus on how to explore in the fully cooperative task, which is insufficient in the environment with nonstationarity induced by agent interactions. To tackle this issue, we propose When to Explore (WToE), a simple yet effective variational exploration method to learn WToE under nonstationary environments. WToE employs an interaction-oriented adaptive exploration mechanism to adapt to environmental changes. We first propose a novel graphical model that uses a latent random variable to model the step-level environmental change resulting from interaction effects. Leveraging this graphical model, we employ the supervised variational auto-encoder (VAE) framework to derive a short-term inferred policy from historical trajectories to deal with the nonstationarity. Finally, agents engage in exploration when the short-term inferred policy diverges from the current actor policy. The proposed approach theoretically guarantees the convergence of the Q -value function. In our experiments, we validate our exploration mechanism in grid examples, multiagent particle environments and the battle game of MAgent environments. The results demonstrate the superiority of WToE over multiple baselines and existing exploration methods, such as MAEXQ, NoisyNets, EITI, and PR2.
Shaokang Dong, Hangyu Mao, Shangdong Yang, Shengyu Zhu 0001, Wenbin Li 0006, Jianye Hao, Yang Gao 0001
IEEE Trans. Cybern.5
2024 Learning Multi-Intersection Traffic Signal Control via Coevolutionary Multi-Agent Reinforcement Learning
abstract
Effective management of multi-intersection traffic signal control (MTSC) is vital for intelligent transportation systems. Multi-agent reinforcement learning (MARL) has shown promise in achieving MTSC. However, existing MARL-based MTSC algorithms have primarily focused on capturing the spatial relationship between multi-intersection traffic signals but have overlooking the importance of the temporally stable traffic pattern. This pattern refers to the fixed positions and relatively stable traffic flow between intersections over short periods in real-world MTSC scenarios, which indicates that the learned spatial relationships between traffic signals should co-evolve over time. To this end, we propose a novel algorithm calledCoevolutionaryMulti-AgentReinforcementLearning (CoevoMARL). CoevoMARL employs a graph neural network to capture the complex spatial interaction network among traffic signals. Furthermore, we propose a relationship-driven progressive LSTM (RDP-LSTM) that dynamically evolves the learned spatial interaction network over time by leveraging insights from the temporally stable traffic pattern. To accelerate convergence, we also propose the mutual information reward optimization (MIRO) technique, which strengthens the correlation between policy learning and high-performance samples by using a mutual information-based intrinsic reward. Experimental results on both synthetic and realistic datasets demonstrate the superiority of CoevoMARL over existing MTSC algorithms, providing valuable insights into incorporating the temporally stable traffic pattern.
Wubing Chen, Shangdong Yang, Wenbin Li 0006, Yujing Hu, Yang Gao 0001
IEEE Trans. Intell. Transp. Syst.3
2023 Learning Explicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning via Polarization Policy Gradient
abstract
Cooperative multi-agent policy gradient (MAPG) algorithms have recently attracted wide attention and are regarded as a general scheme for the multi-agent system. Credit assignment plays an important role in MAPG and can induce cooperation among multiple agents. However, most MAPG algorithms cannot achieve good credit assignment because of the game-theoretic pathology known as centralized-decentralized mismatch. To address this issue, this paper presents a novel method, Multi-Agent Polarization Policy Gradient (MAPPG). MAPPG takes a simple but efficient polarization function to transform the optimal consistency of joint and individual actions into easily realized constraints, thus enabling efficient credit assignment in MAPPG. Theoretically, we prove that individual policies of MAPPG can converge to the global optimum. Empirically, we evaluate MAPPG on the well-known matrix game and differential game, and verify that MAPPG can converge to the global optimum for both discrete and continuous action spaces. We also evaluate MAPPG on a set of StarCraft II micromanagement tasks and demonstrate that MAPPG outperforms the state-of-the-art MAPG algorithms.
Wubing Chen, Wenbin Li 0006, Shangdong Yang, Yang Gao 0001
AAAI2
2023 Modeling Inter-Class and Intra-Class Constraints in Novel Class Discovery
abstract
Novel class discovery (NCD) aims at learning a model that transfers the common knowledge from a class-disjoint labelled dataset to another unlabelled dataset and discovers new classes (clusters) within it. Many methods, as well as elaborate training pipelines and appropriate objectives, have been proposed and considerably boosted performance on NCD tasks. Despite all this, we find that the existing methods do not sufficiently take advantage of the essence of the NCD setting. To this end, in this paper, we propose to model both inter-class and intra-class constraints in NCD based on the symmetric Kullback-Leibler divergence (sKLD). Specifically, we propose an inter-class sKLD constraint to effectively exploit the disjoint relationship between labelled and unlabelled classes, enforcing the separability for different classes in the embedding space. In addition, we present an intra-class sKLD constraint to explicitly constrain the intra-relationship between a sample and its augmentations and ensure the stability of the training process at the same time. We conduct extensive experiments on the popular CIFAR10, CIFAR100 and ImageNet benchmarks and successfully demonstrate that our method can establish a new state of the art and can achieve significant performance improvements, e.g., 3.5%/3.7% clustering accuracy improvements on CIFAR100-50 dataset split under the task-aware/-agnostic evaluation protocol, over previous state-of-the-art methods. Code is available at https://github.com/FanZhichen/NCD-IIC.
Wenbin Li 0006, Zhichen Fan, Jing Huo, Yang Gao 0001
CVPR1
2023 Convergence Analysis of Graphical Game-Based Nash Q-Learning using the Interaction Detection Signal of N-Step Return
abstract
The graphical game provides an effective method for modeling different kinds of sparse interactions in multi-agent reinforcement learning. Most previous work on game abstraction lacks theoretical guarantees of convergence. In this paper, we adopt the ${\mathcal{N}}$-step return signal to detect interactions between agents and build the Markov graphical game based on it. We analyze that the solution of the Markov graphical game is an ϵ-Nash equilibrium which guarantees the convergence of the proposed NSR-G2NashQ algorithm theoretically. Also, we have done experiments in different multi-agent reinforcement learning tasks with both tabular and function approximation solutions. The results show the NSR-G2NashQ algorithm accelerates the convergence of agents to the optimal policy.
Yunkai Zhuang, Shangdong Yang, Wenbin Li 0006, Yang Gao 0001
ICASSP3
2023 Where and How: Mitigating Confusion in Neural Radiance Fields from Sparse Inputs
abstract
Neural Radiance Fields from Sparse inputs (NeRF-S) have shown great potential in synthesizing novel views with a limited number of observed viewpoints. However, due to the inherent limitations of sparse inputs and the gap between non-adjacent views, rendering results often suffer from over-fitting and foggy surfaces, a phenomenon we refer to as "CONFUSION" during volume rendering. In this paper, we analyze the root cause of this confusion and attribute it to two fundamental questions: "WHERE" and "HOW". To this end, we present a novel learning framework, WaH-NeRF, which effectively mitigates confusion by tackling the following challenges: (i) "WHERE" to Sample? in NeRF-S-we introduce a Deformable Sampling strategy and a Weight-based Mutual Information Loss to address sample-position confusion arising from the limited number of viewpoints; and (ii) "HOW" to Predict? in NeRF-S-we propose a Semi-Supervised NeRF learning Paradigm based on pose perturbation and a Pixel-Patch Correspondence Loss to alleviate prediction confusion caused by the disparity between training and testing viewpoints. By integrating our proposed modules and loss functions, WaH-NeRF outperforms previous methods under the NeRF-S setting. Code is available https://github.com/bbbbby-99/WaH-NeRF.
Yanqi Bao, Jing Huo, Tianyu Ding, Wenbin Li 0006, Yang Gao 0001
ACM Multimedia6
2023 Efficient Subgame Refinement for Extensive-form Games
abstract
Subgame solving is an essential technique in addressing large imperfect information games, with various approaches developed to enhance the performance of refined strategies in the abstraction of the target subgame. However, directly applying existing subgame solving techniques may be difficult, due to the intricate nature and substantial size of many real-world games. To overcome this issue, recent subgame solving methods allow for subgame solving on limited knowledge order subgames, increasing their applicability in large games; yet this may still face obstacles due to extensive information set sizes. To address this challenge, we propose a generative subgame solving (GS2) framework, which utilizes a generation function to identify a subset of the earliest-reached nodes, reducing the size of the subgame. Our method is supported by a theoretical analysis and employs a diversity-based generation function to enhance safety. Experiments conducted on medium-sized games as well as the challenging large game of GuanDan demonstrate a significant improvement over the blueprint.
Zhenxing Ge, Tianyu Ding, Wenbin Li 0006, Yang Gao 0001
NeurIPS4
2023 LibFewShot: A Comprehensive Library for Few-Shot Learning
abstract
Few-shot learning, especially few-shot image classification, has received increasing attention and witnessed significant advances in recent years. Some recent studies implicitly show that many generic techniques or "tricks", such as data augmentation, pre-training, knowledge distillation, and self-supervision, may greatly boost the performance of a few-shot learning method. Moreover, different works may employ different software platforms, backbone architectures and input image sizes, making fair comparisons difficult and practitioners struggle with reproducibility. To address these situations, we propose a comprehensive library for few-shot learning (LibFewShot) by re-implementing eighteen state-of-the-art few-shot learning methods in a unified framework with the same single codebase in PyTorch. Furthermore, based on LibFewShot, we provide comprehensive evaluations on multiple benchmarks with various backbone architectures to evaluate common pitfalls and effects of different training tricks. In addition, with respect to the recent doubts on the necessity of meta- or episodic-training mechanism, our evaluation results confirm that such a mechanism is still necessary especially when combined with pre-training. We hope our work can not only lower the barriers for beginners to enter the area of few-shot learning but also elucidate the effects of nontrivial tricks to facilitate intrinsic research on few-shot learning.
Wenbin Li 0006, Xuesong Yang, Chuanqi Dong, Pinzhuo Tian, Tiexin Qin, Jing Huo, Yinghuan Shi, Lei Wang 0001, Yang Gao 0001, Jiebo Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Defensive Few-Shot Learning
abstract
This article investigates a new challenging problem called defensive few-shot learning in order to learn a robust few-shot model against adversarial attacks. Simply applying the existing adversarial defense methods to few-shot learning cannot effectively solve this problem. This is because the commonly assumed sample-level distribution consistency between the training and test sets can no longer be met in the few-shot setting. To address this situation, we develop a general defensive few-shot learning (DFSL) framework to answer the following two key questions: (1) how to transfer adversarial defense knowledge from one sample distribution to another? (2) how to narrow the distribution gap between clean and adversarial examples under the few-shot setting? To answer the first question, we propose an episode-based adversarial training mechanism by assuming a task-level distribution consistency to better transfer the adversarial defense knowledge. As for the second question, within each few-shot task, we design two kinds of distribution consistency criteria to narrow the distribution gap between clean and adversarial examples from the feature-wise and prediction-wise perspectives, respectively. Extensive experiments demonstrate that the proposed framework can effectively make the existing few-shot models robust against adversarial attacks. Code is available at https://github.com/WenbinLee/DefensiveFSL.git.
Wenbin Li 0006, Lei Wang 0001, Xingxing Zhang 0001, Lei Qi 0001, Jing Huo, Yang Gao 0001, Jiebo Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Global- and local-aware feature augmentation with semantic orthogonality for few-shot image classification
Boyao Shi, Wenbin Li 0006, Jing Huo, Lei Wang 0001, Yang Gao 0001
Pattern Recognit.2
2023 A Multilayer Framework for Online Metric Learning
abstract
Online metric learning (OML) has been widely applied in classification and retrieval. It can automatically learn a suitable metric from data by restricting similar instances to be separated from dissimilar instances with a given margin. However, the existing OML algorithms have limited performance in real-world classifications, especially, when data distributions are complex. To this end, this article proposes a multilayer framework for OML to capture the nonlinear similarities among instances. Different from the traditional OML, which can only learn one metric space, the proposed multilayer OML (MLOML) takes an OML algorithm as a metric layer and learns multiple hierarchical metric spaces, where each metric layer follows a nonlinear layer for the complicated data distribution. Moreover, the forward propagation (FP) strategy and backward propagation (BP) strategy are employed to train the hierarchical metric layers. To build a metric layer of the proposed MLOML, a new Mahalanobis-based OML (MOML) algorithm is presented based on the passive-aggressive strategy and one-pass triplet construction strategy. Furthermore, in a progressively and nonlinearly learning way, MLOML has a stronger learning ability than traditional OML in the case of limited available training data. To make the learning process more explainable and theoretically guaranteed, theoretical analysis is provided. The proposed MLOML enjoys several nice properties, indeed learns a metric progressively, and performs better on the benchmark datasets. Extensive experiments with different settings have been conducted to verify these properties of the proposed MLOML.
Wenbin Li 0006, Jing Huo, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Jiebo Luo 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Online Passive-Aggressive Active Learning for Trapezoidal Data Streams
abstract
The idea of combining the active query strategy and the passive-aggressive (PA) update strategy in online learning can be credited to the PA active (PAA) algorithm, which has proven to be effective in learning linear classifiers from datasets with a fixed feature space. We propose a novel family of online active learning algorithms, named PAA learning for trapezoidal data streams (PAATS) and multiclass PAATS (MPAATS) (and their variants), for binary and multiclass online classification tasks on trapezoidal data streams where the feature space may expand over time. Under the context of an ever-changing feature space, we provide the theoretical analysis of the mistake bounds for both PAATS and MPAATS. Our experiments on a wide variety of benchmark datasets have confirm that the combination of the instance-regulated active query strategy and the PA update strategy is much more effective in learning from trapezoidal data streams. We have also compared PAATS with online learning with streaming features (OLSF)-the state-of-the-art approach in learning linear classifiers from trapezoidal data streams. PAATS could achieve much better classification accuracy, especially for large-scale real-world data streams.
Xiaocong Fan, Wenbin Li 0006, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Uncertainty-Based Network for Few-Shot Image Classification
abstract
The transductive inference is an effective technique in the few-shot learning task, where query sets update prototypes to improve themselves. However, these methods optimize the model by considering only the classification scores of the query instances as confidence while ignoring the uncertainty of these classification scores. In this paper, we propose a novel method called Uncertainty-Based Network, which models the uncertainty of classification results with the help of mutual information. Specifically, we first data augment and classify the query instance and calculate the mutual information of these classification scores. Then, mutual information is used as uncertainty to assign weights to classification scores, and the iterative update strategy based on classification scores and uncertainties assigns the optimal weights to query instances in prototype optimization. Extensive results on four benchmarks show that Uncertainty-Based Network achieves comparable performance in classification accuracy compared to state-of-the-art methods.
Minglei Yuan, Chunhao Cai, Yin-Dong Zheng, Tao Wang 0052, Tong Lu 0002, Wenbin Li 0006
ICME7
2022 CAST: Learning Both Geometric and Texture Style Transfers for Effective Caricature Generation
abstract
Given a photo of a subject, ability to generate a caricature image that captures distinct characteristics of the subject but with certain exaggeration of their prominent features is of fundamental importance to image processing and facial recognition. There are two main challenges in this task: shape exaggeration and style transfer. The former morphs and exaggerates key facial features of the subject, while the latter generates caricature images in a certain artistic style. In this paper, we propose a CAricature Style Transfer (CAST) framework for caricature generation. There are two modules in the proposed framework. The first is a geometric warping module. Different from the existing style transfer methods, we incorporate the Whitening and Coloring Transformation (WCT) in the geometric style transfer. The WCT is learned on photo and caricature landmarks or the caricature landmark space of a specific artist and is capable of transforming input photo landmarks to caricature landmarks. The second module is a texture style rendering module. We propose a new style transfer method by considering a semantic region-aligned style transfer via affinity constraint. Given a reference caricature image as the style reference, this module is capable of transferring styles between the same or similar semantic regions in caricatures and photos. Furthermore, it can transfer visual attributes of the reference caricatures (such as mouth shape and expressions) to the output caricatures. Experiments have shown desirable effects of the proposed method in transferring both the geometric and artistic texture styles of caricatures. Both qualitative and quantitative results show that the CAST framework is more effective compared than the state-of-the-art caricature generation methods.
Jing Huo, Xiangde Liu, Wenbin Li 0006, Yang Gao 0001, Hujun Yin, Jiebo Luo 0001
IEEE Trans. Image Process.3
2022 CariMe: Unpaired Caricature Generation With Multiple Exaggerations
abstract
Caricature generation aims to translate real photos into caricatures with artistic styles and shape exaggerations while maintaining the identity of the subject. Different from generic image-to-image translation, drawing caricatures automatically is a more challenging task due to the existence of various spatial deformations. Previous caricature generation methods are obsessed with predicting definite image warping from a given photo while ignoring the intrinsic representation and distribution of geometric exaggerations in caricatures. This limits their ability on diverse exaggeration generation. In this paper, we generalize the caricature generation problem from instance-level warping prediction to distribution-level deformation modeling. Based on this assumption, we present the first exploration forunpaired CARIcature generation with Multiple Exaggerations (CariMe). Technically, we propose a Multi-exaggeration Warper network to learn the distribution-level mapping from photos to facial exaggerations. This makes it possible to generate diverse and reasonable exaggerations from randomly sampled warp codes given one input photo. To better represent the facial exaggeration and produce fine-grained warping, a deformation-field-based warping method is also proposed, which captures more detailed exaggerations than previous point-based warping methods. Experiments and two perceptual studies prove the superiority of our method comparing with other state-of-the-art methods, showing the improvement of our work on caricature generation. The source code is available athttps://github.com/edward3862/CariMe-pytorch.
Zheng Gu 0001, Chuanqi Dong, Jing Huo, Wenbin Li 0006, Yang Gao 0001
IEEE Trans. Multim.4
2022 Consistent Meta-Regularization for Better Meta-Knowledge in Few-Shot Learning
abstract
Recently, meta-learning provides a powerful paradigm to deal with the few-shot learning problem. However, existing meta-learning approaches ignore the prior fact that good meta-knowledge should alleviate the data inconsistency between training and test data, caused by the extremely limited data, in each few-shot learning task. Moreover, legitimately utilizing the prior understanding of meta-knowledge can lead us to design an efficient method to improve the meta-learning model. Under this circumstance, we consider the data inconsistency from the distribution perspective, making it convenient to bring in the prior fact, and propose a new consistent meta-regularization (Con-MetaReg) to help the meta-learning model learn how to reduce the data-distribution discrepancy between the training and test data. In this way, the ability of meta-knowledge on keeping the training and test data consistent is enhanced, and the performance of the meta-learning model can be further improved. The extensive analyses and experiments demonstrate that our method can indeed improve the performances of different meta-learning models in few-shot regression, classification, and fine-grained classification.
Pinzhuo Tian, Wenbin Li 0006, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2021 LoFGAN: Fusing Local Representations for Few-shot Image Generation
abstract
Given only a few available images for a novel unseen category, few-shot image generation aims to generate more data for this category. Previous works attempt to globally fuse these images by using adjustable weighted coefficients. However, there is a serious semantic misalignment between different images from a global perspective, making these works suffer from poor generation quality and diversity. To tackle this problem, we propose a novel Local-Fusion Generative Adversarial Network (LoFGAN) for fewshot image generation. Instead of using these available images as a whole, we first randomly divide them into a base image and several reference images. Next, LoFGAN matches local representations between the base and reference images based on semantic similarities, and replaces the local features with the closest related local features. In this way, LoFGAN can produce more realistic and diverse images at a more fine-grained level, and simultaneously enjoy the characteristic of semantic alignment. Furthermore, a local reconstruction loss is also proposed, which can provide better training stability and generation quality. We conduct extensive experiments on three datasets, which successfully demonstrates the effectiveness of our proposed method for few-shot image generation and downstream visual applications with limited data. Code is available at https://github.com/edward3862/LoFGAN-pytorch.
Zheng Gu 0001, Wenbin Li 0006, Jing Huo, Lei Wang 0001, Yang Gao 0001
ICCV2
2021 Manifold Alignment for Semantically Aligned Style Transfer
abstract
Most existing style transfer methods follow the assumption that styles can be represented with global statistics (e.g., Gram matrices or covariance matrices), and thus address the problem by forcing the output and style images to have similar global statistics. An alternative is the assumption of local style patterns, where algorithms are designed to swap similar local features of content and style images. However, the limitation of these existing methods is that they neglect the semantic structure of the content image which may lead to corrupted content structure in the output. In this paper, we make a new assumption that image features from the same semantic region form a manifold and an image with multiple semantic regions follows a multi-manifold distribution. Based on this assumption, the style transfer problem is formulated as aligning two multi-manifold distributions and a Manifold Alignment based Style Transfer (MAST) framework is proposed. The proposed frame-work allows semantically similar regions between the output and the style image share similar style patterns. Moreover, the proposed manifold alignment method is flexible to allow user editing or using semantic segmentation maps as guidance for style transfer. To allow the method to be applicable to photorealistic style transfer, we propose a new adaptive weight skip connection network structure to preserve the content details. Extensive experiments verify the effectiveness of the proposed framework for both artistic and photorealistic style transfer. Code is available at https://github.com/NJUHuoJing/MAST.
Jing Huo, Shiyin Jin, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yinghuan Shi, Yang Gao 0001
ICCV3
2021 Local descriptor-based multi-prototype network for few-shot Learning
Zhangkai Wu, Wenbin Li 0006, Jing Huo, Yang Gao 0001
Pattern Recognit.3
2021 Temperature network for few-shot learning with distribution-aware large-margin metric
Wei Zhu 0015, Wenbin Li 0006, Haofu Liao, Jiebo Luo 0001
Pattern Recognit.2
2020 Layerwise Sparse Coding for Pruned Deep Neural Networks with Extreme Compression Ratio
abstract
Deep neural network compression is important and increasingly developed especially in resource-constrained environments, such as autonomous drones and wearable devices. Basically, we can easily and largely reduce the number of weights of a trained deep model by adopting a widely used model compression technique, e.g., pruning. In this way, two kinds of data are usually preserved for this compressed model, i.e., non-zero weights and meta-data, where meta-data is employed to help encode and decode these non-zero weights. Although we can obtain an ideally small number of non-zero weights through pruning, existing sparse matrix coding methods still need a much larger amount of meta-data (may several times larger than non-zero weights), which will be a severe bottleneck of the deploying of very deep models. To tackle this issue, we propose a layerwise sparse coding (LSC) method to maximize the compression ratio by extremely reducing the amount of meta-data. We first divide a sparse matrix into multiple small blocks and remove zero blocks, and then propose a novel signed relative index (SRI) algorithm to encode the remaining non-zero blocks (with much less meta-data). In addition, the proposed LSC performs parallel matrix multiplication without full decoding, while traditional methods cannot. Through extensive experiments, we demonstrate that LSC achieves substantial gains in pruned DNN compression (e.g., 51.03x compression ratio on ADMM-Lenet) and inference computation (i.e., time reduction and extremely less memory bandwidth), over state-of-the-art baselines.
Wenbin Li 0006, Jing Huo, Lili Yao, Yang Gao 0001
AAAI2
2020 Deep Embedded Complementary and Interactive Information for Multi-View Classification
abstract
Multi-view classification optimally integrates various features from different views to improve classification tasks. Though most of the existing works demonstrate promising performance in various computer vision applications, we observe that they can be further improved by sufficiently utilizing complementary view-specific information, deep interactive information between different views, and the strategy of fusing various views. In this work, we propose a novel multi-view learning framework that seamlessly embeds various view-specific information and deep interactive information and introduces a novel multi-view fusion strategy to make a joint decision during the optimization for classification. Specifically, we utilize different deep neural networks to learn multiple view-specific representations, and model deep interactive information through a shared interactive network using the cross-correlations between attributes of these representations. After that, we adaptively integrate multiple neural networks by flexibly tuning the power exponent of weight, which not only avoids the trivial solution of weight but also provides a new approach to fuse outputs from different deterministic neural networks. Extensive experiments on several public datasets demonstrate the rationality and effectiveness of our method.
Jinglin Xu, Wenbin Li 0006, Dingwen Zhang, Junwei Han 0001
AAAI2
2020 Learning Task-aware Local Representations for Few-shot Learning
abstract
Few-shot learning for visual recognition aims to adapt to novel unseen classes with only a few images. Recent work, especially the work based on low-level information, has achieved great progress. In these work, local representations (LRs) are typically employed, because LRs are more consistent among the seen and unseen classes. However, most of them are limited to an individual image-to-image or image-to-class measure manner, which cannot fully exploit the capabilities of LRs, especially in the context of a certain task. This paper proposes an Adaptive Task-aware Local Representations Network (ATL-Net) to address this limitation by introducing episodic attention, which can adaptively select the important local patches among the entire task, as the process of human recognition. We achieve much superior results on multiple benchmarks. On the miniImagenet, ATL-Net gains 0.93% and 0.88% improvements over the compared methods under the 5-way 1-shot and 5-shot settings. Moreover, ATL-Net can naturally tackle the problem that how to adaptively identify and weight the importance of different key local parts, which is the major concern of fine-grained recognition. Specifically, on the fine-grained dataset Stanford Dogs, ATL-Net outperforms the second best method with 5.39% and 9.69% gains under the 5-way 1-shot and 5-shot settings.
Chuanqi Dong, Wenbin Li 0006, Jing Huo, Zheng Gu 0001, Yang Gao 0001
IJCAI2
2020 Asymmetric Distribution Measure for Few-shot Learning
abstract
The core idea of metric-based few-shot image classification is to directly measure the relations between query images and support classes to learn transferable feature embeddings. Previous work mainly focuses on image-level feature representations, which actually cannot effectively estimate a class's distribution due to the scarcity of samples. Some recent work shows that local descriptor based representations can achieve richer representations than image-level based representations. However, such works are still based on a less effective instance-level metric, especially a symmetric metric, to measure the relation between a query image and a support class. Given the natural asymmetric relation between a query image and a support class, we argue that an asymmetric measure is more suitable for metric-based few-shot learning. To that end, we propose a novel Asymmetric Distribution Measure (ADM) network for few-shot learning by calculating a joint local and global asymmetric measure between two multivariate local distributions of a query and a class. Moreover, a task-aware Contrastive Measure Strategy (CMS) is proposed to further enhance the measure function. On popular miniImageNet and tieredImageNet, ADM can achieve the state-of-the-art results, validating our innovative design of asymmetric distribution measures for few-shot learning. The source code can be downloaded from https://github.com/WenbinLee/ADM.git.
Wenbin Li 0006, Lei Wang 0001, Jing Huo, Yinghuan Shi, Yang Gao 0001, Jiebo Luo 0001
IJCAI1
2020 Biased Feature Learning for Occlusion Invariant Face Recognition
abstract
To address the challenges posed by unknown occlusions, we propose a Biased Feature Learning (BFL) framework for occlusion-invariant face recognition. We first construct an extended dataset using a multi-scale data augmentation method. For model training, we modify the label loss to adjust the impact of normal and occluded samples. Further, we propose a biased guidance strategy to manipulate the optimization of a network so that the feature embedding space is dominated by non-occluded faces. BFL not only enhances the robustness of a network to unknown occlusions but also maintains or even improves its performance for normal faces. Experimental results demonstrate its superiority as well as the generalization capability with different network architectures and loss functions.
Chang-Bin Shao, Jing Huo, Lei Qi 0001, Zhenhua Feng 0001, Wenbin Li 0006, Chuanqi Dong, Yang Gao 0001
IJCAI5
2020 Joint Multi-view 2D Convolutional Neural Networks for 3D Object Classification
abstract
Three-dimensional (3D) object classification is widely involved in various computer vision applications, e.g., autonomous driving, simultaneous localization and mapping, which has attracted lots of attention in the committee. However, solving 3D object classification by directly employing the 3D convolutional neural networks (CNNs) generally suffers from high computational cost. Besides, existing view-based methods cannot better explore the content relationships between views. To this end, this work proposes a novel multi-view framework by jointly using multiple 2D-CNNs to capture discriminative information with relationships as well as a new multi-view loss fusion strategy, in an end-to-end manner. Specifically, we utilize multiple 2D views of a 3D object as input and integrate the intra-view and inter-view information of each view through the view-specific 2D-CNN and a series of modules (outer product, view pair pooling, 1D convolution, and fully connected transformation). Furthermore, we design a novel view ensemble mechanism that selects several discriminative and informative views to jointly infer the category of a 3D object. Extensive experiments demonstrate that the proposed method is able to outperform current state-of-the-art methods on 3D object classification. More importantly, this work provides a new way to improve 3D object classification from the perspective of fully utilizing well-established 2D-CNNs.
Jinglin Xu, Xiangsen Zhang, Wenbin Li 0006, Junwei Han 0001
IJCAI3
2020 Alleviating the Incompatibility Between Cross Entropy Loss and Episode Training for Few-Shot Skin Disease Classification
Wei Zhu 0015, Haofu Liao, Wenbin Li 0006, Weijian Li 0001, Jiebo Luo 0001
MICCAI (6)3
2020 Robust neighborhood embedding for unsupervised feature selection
Dongyi Ye, Wenbin Li 0006, Yang Gao 0001
Knowl. Based Syst.3
2020 CariGAN: Caricature generation through weakly paired adversarial learning
Wenbin Li 0006, Wei Xiong 0008, Haofu Liao, Jing Huo, Yang Gao 0001, Jiebo Luo 0001
Neural Networks1
2019 Distribution Consistency Based Covariance Metric Networks for Few-Shot Learning
abstract
Few-shot learning aims to recognize new concepts from very few examples. However, most of the existing few-shot learning methods mainly concentrate on the first-order statistic of concept representation or a fixed metric on the relation between a sample and a concept. In this work, we propose a novel end-to-end deep architecture, named Covariance Metric Networks (CovaMNet). The CovaMNet is designed to exploit both the covariance representation and covariance metric based on the distribution consistency for the few-shot classification tasks. Specifically, we construct an embedded local covariance representation to extract the second-order statistic information of each concept and describe the underlying distribution of this concept. Upon the covariance representation, we further define a new deep covariance metric to measure the consistency of distributions between query samples and new concepts. Furthermore, we employ the episodic training mechanism to train the entire network in an end-to-end manner from scratch. Extensive experiments in two tasks, generic few-shot image classification and fine-grained fewshot image classification, demonstrate the superiority of the proposed CovaMNet. The source code can be available from https://github.com/WenbinLee/CovaMNet.git.
Wenbin Li 0006, Jinglin Xu, Jing Huo, Lei Wang 0001, Yang Gao 0001, Jiebo Luo 0001
AAAI1
2019 Revisiting Local Descriptor Based Image-To-Class Measure for Few-Shot Learning
abstract
Few-shot learning in image classification aims to learn a classifier to classify images when only few training examples are available for each class. Recent work has achieved promising classification performance, where an image-level feature based measure is usually used. In this paper, we argue that a measure at such a level may not be effective enough in light of the scarcity of examples in few-shot learning. Instead, we think a local descriptor based image-to-class measure should be taken, inspired by its surprising success in the heydays of local invariant features. Specifically, building upon the recent episodic training mechanism, we propose a Deep Nearest Neighbor Neural Network (DN4 in short) and train it in an end-to-end manner. Its key difference from the literature is the replacement of the image-level feature based measure in the final layer by a local descriptor based image-to-class measure. This measure is conducted online via a k-nearest neighbor search over the deep local descriptors of convolutional feature maps. The proposed DN4 not only learns the optimal deep local descriptors for the image-to-class measure, but also utilizes the higher efficiency of such a measure in the case of example scarcity, thanks to the exchangeability of visual patterns across the images in the same class. Our work leads to a simple, effective, and computationally efficient framework for few-shot learning. Experimental study on benchmark datasets consistently shows its superiority over the related state-of-the-art, with the largest absolute improvement of 17% over the next best. The source code can be available from https://github.com/WenbinLee/DN4.git.
Wenbin Li 0006, Lei Wang 0001, Jinglin Xu, Jing Huo, Yang Gao 0001, Jiebo Luo 0001
CVPR1
2018 A Joint Local and Global Deep Metric Learning Method for Caricature Recognition
Wenbin Li 0006, Jing Huo, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Jiebo Luo 0001
ACCV (4)1
2018 WebCaricature: a benchmark for caricature recognition
Jing Huo, Wenbin Li 0006, Yinghuan Shi, Yang Gao 0001, Hujun Yin
BMVC2
2018 A Novel Two-Stage Deep Method for Mitosis Detection in Breast Cancer Histology Images
abstract
The accurate detection and counting of mitosis in breast cancer histology images is very important for computer-aided diagnosis, which is manually completed by the pathologist according to her or his clinic experience. However, this procedure is extremely time consuming and tedious. Moreover, it always results in low agreement among different pathologists. Although several computer-aided detection methods have been developed recently, they suffer from high FN (false negative) and FP (false positive) with simply treating the detection task as a binary classification problem. In this paper, we present a novel two-stage detection method with multi-scale and similarity learning convnets (MSSN). Firstly, large amount of possible candidates will be generated in the first stage in order to reduce FN (i.e., prevent treating mitosis as non-mitosis), by using the different square and non-square filters, to capture the spatial relation from different scales. Secondly, a similarity prediction model is subsequently performed on the obtained candidates for the final detection to reduce FP, which is realized by imposing a large margin constraint. On both 2014 and 2012 ICPR MITOSIS datasets, our MSSN achieved a promising result with a highest Recall (outperforming other methods by a large margin) and a comparable F-score.
Minglin Ma, Yinghuan Shi, Wenbin Li 0006, Yang Gao 0001
ICPR3
2018 OPML: A one-pass closed-form solution for online metric learning
Wenbin Li 0006, Yang Gao 0001, Lei Wang 0001, Luping Zhou, Jing Huo, Yinghuan Shi
Pattern Recognit.1
2017 Beyond IID: Learning to Combine Non-IID Metrics for Vision Tasks
abstract
Metric learning has been widely employed, especially in various computer vision tasks, with the fundamental assumption that all samples (e.g., regions/superpixels in images/videos) are independent and identically distributed (IID). However, since the samples are usually spatially-connected or temporally-correlated with their physically-connected neighbours, they are not IID (non-IID for short), which cannot be directly handled by existing methods. Thus, we propose to learn and integrate non-IID metrics (NIME). To incorporate the non-IID spatial/temporal relations, instead of directly using non-IID features and metric learning as previous methods, NIME first builds several non-IID representations on original (non-IID) features by various graph kernel functions, and then automatically learns the metric under the best combination of various non-IID representations. NIME is applied to solve two typical computer vision tasks: interactive image segmentation and histology image identification. The results show that learning and integrating non-IID metrics improves the performance, compared to the IID methods. Moreover, our method achieves results comparable or better than that of the state-of-the-arts.
Yinghuan Shi, Wenbin Li 0006, Yang Gao 0001, Longbing Cao, Dinggang Shen
AAAI2
2015 Interactive image segmentation via cascaded metric learning
abstract
In this paper, we propose an interactive image segmentation method from a novel perspective of cascaded metric learning. Given an image with user-marked scribbles that are essentially uncertain and noisy, our method completes the segmentation task by solving a binary classification problem. Starting from the initial training samples with known class labels (i.e., regions of the image that are believed with high confidence to be foreground or background), we first find an optimal metric that can best describe the classification of these samples. After that, we classify the unlabeled samples using the learnt metric. Samples classified with high confidence are used as new training samples to refine the metric. This cycle of metric learning and classification repeats until the accomplishment of the image segmentation task. The proposed method is extensively evaluated on the MSRC image set. Experiment results show that our method outperforms the state-of-the-art methods.
Wenbin Li 0006, Yinghuan Shi, Wanqi Yang, Hao Wang 0013, Yang Gao 0001
ICIP1