Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Boyu Yang 0002

dblp:63/6922-2 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0003-3799-6625ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Transfer learning and domain adaptation · 38% Vision and language · 30% Image recognition and object detection · 12%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot class-incremental learning
1.322023
Dynamic Support Network for Few-Shot Class Incremental Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Learnable Distribution Calibration for Few-Shot Class-Incremental Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Transfer learning and domain adaptation
few-shot learning
1.322023
Dynamic Support Network for Few-Shot Class Incremental Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Learnable Distribution Calibration for Few-Shot Class-Incremental Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension · ICLR 2025
Computer vision › Vision and language
multimodal referring expressions
0.912025
ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension · ICLR 2025
Computer vision › Vision and language › visual grounding
referring and grounding
0.912025
ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension · ICLR 2025
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.712023
Dynamic Support Network for Few-Shot Class Incremental Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Transfer learning and domain adaptation › few-shot learning
distribution calibration
0.712023
Learnable Distribution Calibration for Few-Shot Class-Incremental Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.712023
Learnable Distribution Calibration for Few-Shot Class-Incremental Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › Image recognition and object detection › object detection
few-shot object detection
0.512021
Beyond Max-Margin: Class Margin Equilibrium for Few-Shot Object Detection · CVPR 2021
Computer vision › Image recognition and object detection
object detection
0.512021
Beyond Max-Margin: Class Margin Equilibrium for Few-Shot Object Detection · CVPR 2021
Computer vision › Segmentation and scene understanding › semantic segmentation
few-shot segmentation
0.412020
Prototype Mixture Models for Few-Shot Semantic Segmentation · ECCV (8) 2020

Methods — techniques the papers use, named apart from their topics

token collectives · 0.9hybrid perception · 0.9variational inference · 0.7parameterized calibration unit · 0.7dynamic network expansion · 0.7covariance matrix sharing · 0.7class distribution recalling · 0.7class margin loss · 0.5adversarial min-max training · 0.5mixture model · 0.4
YearPublicationVenuePosition
2025 ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension
abstract
Aligning vision and language concepts at a finer level remains an essential topic of multimodal large language models (MLLMs), particularly for tasks such as referring and grounding. Existing methods, such as *proxy encoding* and *geometry encoding* genres, incorporate additional syntax to encode spatial information, imposing extra burdens when communicating between language with vision modules. In this study, we propose ClawMachine, offering a new methodology that explicitly notates each entity using **token collectives**—groups of visual tokens that collaboratively represent higher-level semantics. A hybrid perception mechanism is also explored to perceive and understand scenes from both discrete and continuous spaces. Our method unifies the prompt and answer of visual referential tasks without using additional syntax. By leveraging a joint vision-language vocabulary, ClawMachine integrates referring and grounding in an auto-regressive manner, demonstrating great potential with scaled up pre-training data. Experiments show that ClawMachine achieves superior performance on scene-level and referential understanding tasks with higher efficiency. It also exhibits the potential to integrate multi-source information for complex visual reasoning, which is beyond the capability of many MLLMs. Our code is available at https://github.com/martian422/ClawMachine.
Tianren Ma, Lingxi Xie, Yunjie Tian, Boyu Yang 0002, Qixiang Ye
ICLR4
2023 Learnable Distribution Calibration for Few-Shot Class-Incremental Learning
abstract
Few-shot class-incremental learning (FSCIL) faces the challenges of memorizing old class distributions and estimating new class distributions given few training samples. In this study, we propose a learnable distribution calibration (LDC) approach, to systematically solve these two challenges using a unified framework. LDC is built upon a parameterized calibration unit (PCU), which initializes biased distributions for all classes based on classifier vectors (memory-free) and a single covariance matrix. The covariance matrix is shared by all classes, so that the memory costs are fixed. During base training, PCU is endowed with the ability to calibrate biased distributions by recurrently updating sampled features under supervision of real distributions. During incremental learning, PCU recovers distributions for old classes to avoid 'forgetting', as well as estimating distributions and augmenting samples for new classes to alleviate 'over-fitting' caused by the biased distributions of few-shot samples. LDC is theoretically plausible by formatting a variational inference procedure. It improves FSCIL's flexibility as the training procedure requires no class similarity priori. Experiments on CUB200, CIFAR100, and mini-ImageNet datasets show that LDC respectively outperforms the state-of-the-arts by 4.64%, 1.98%, and 3.97%. LDC's effectiveness is also validated on few-shot learning scenarios.
Binghao Liu, Boyu Yang 0002, Lingxi Xie, Qi Tian 0001, Qixiang Ye
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Dynamic Support Network for Few-Shot Class Incremental Learning
abstract
Few-shot class-incremental learning (FSCIL) is challenged by catastrophically forgetting old classes and over-fitting new classes. Revealed by our analyses, the problems are caused by feature distribution crumbling, which leads to class confusion when continuously embedding few samples to a fixed feature space. In this study, we propose a Dynamic Support Network (DSN), which refers to an adaptively updating network with compressive node expansion to "support" the feature space. In each training session, DSN tentatively expands network nodes to enlarge feature representation capacity for incremental classes. It then dynamically compresses the expanded network by node self-activation to pursue compact feature representation, which alleviates over-fitting. Simultaneously, DSN selectively recalls old class distributions during incremental learning to support feature distributions and avoid confusion between classes. DSN with compressive node expansion and class distribution recalling provides a systematic solution for the problems of catastrophic forgetting and overfitting. Experiments on CUB, CIFAR-100, and miniImage datasets show that DSN significantly improves upon the baseline approach, achieving new state-of-the-arts.
Boyu Yang 0002, Mingbao Lin, Binghao Liu, Xiaodan Liang, Rongrong Ji, Qixiang Ye
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Part-Based Semantic Transform for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation remains an open problem for the lack of an effective method to handle the semantic misalignment between objects. In this article, we propose part-based semantic transform (PST) and target at aligning object semantics in support images with those in query images by semantic decomposition-and-match. The semantic decomposition process is implemented with prototype mixture models (PMMs), which use an expectation-maximization (EM) algorithm to decompose object semantics into multiple prototypes corresponding to object parts. The semantic match between prototypes is performed with a min-cost flow module, which encourages correct correspondence while depressing mismatches between object parts. With semantic decomposition-and-match, PST enforces the network's tolerance to objects' appearance and/or pose variation and facilities channelwise and spatial semantic activation of objects in query images. Extensive experiments on Pascal VOC and MS-COCO datasets show that PST significantly improves upon state-of-the-arts. In particular, on MS-COCO, it improves the performance of five-shot semantic segmentation by up to 7.79% with a moderate cost of inference speed and model size. Code for PST is released at https://github.com/Yang-Bob/PST.
Boyu Yang 0002, Fang Wan 0001, Chang Liu 0047, Xiangyang Ji, Qixiang Ye
IEEE Trans. Neural Networks Learn. Syst.1
2021 Beyond Max-Margin: Class Margin Equilibrium for Few-Shot Object Detection
abstract
Few-shot object detection has made substantial progress by representing novel class objects using the feature representation learned upon a set of base class objects. However, an implicit contradiction between novel class classification and representation is unfortunately ignored. On the one hand, to achieve accurate novel class classification, the distributions of either two base classes must be far away from each other (max-margin). On the other hand, to precisely represent novel classes, the distributions of base classes should be close to each other to reduce the intra-class distance of novel classes (min-margin). In this paper, we propose a class margin equilibrium (CME) approach, with the aim to optimize both feature space partition and novel class reconstruction in a systematic way. CME first converts the few-shot detection problem to the few-shot classification problem by using a fully connected layer to decouple localization features. CME then reserves adequate margin space for novel classes by introducing simple-yet-effective class margin loss during feature learning. Finally, CME pursues margin equilibrium by disturbing the features of novel class instances in an adversarial min-max fashion. Experiments on Pascal VOC and MS-COCO datasets show that CME significantly improves upon two baseline detectors (up to 3 ~ 5% in average), achieving state-of-the-art performance. Code is available at https://github.com/BohaoLee/CME.
Boyu Yang 0002, Chang Liu 0042, Feng Liu 0050, Rongrong Ji, Qixiang Ye
CVPR2
2020 Prototype Mixture Models for Few-Shot Semantic Segmentation
Boyu Yang 0002, Chang Liu 0042, Jianbin Jiao, Qixiang Ye
ECCV (8)1