EDBT 2026 Demo / reviewers in the wild / expert
Boyu Yang 0002
dblp:63/6922-2
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0003-3799-6625ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Transfer learning and domain adaptation · 38% Vision and language · 30% Image recognition and object detection · 12% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot class-incremental learning |
1.3 | 2 | 2023 | Dynamic Support Network for Few-Shot Class Incremental Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Learnable Distribution Calibration for Few-Shot Class-Incremental Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
1.3 | 2 | 2023 | Dynamic Support Network for Few-Shot Class Incremental Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Learnable Distribution Calibration for Few-Shot Class-Incremental Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension · ICLR 2025 |
Computer vision › Vision and language
multimodal referring expressions |
0.9 | 1 | 2025 | ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension · ICLR 2025 |
Computer vision › Vision and language › visual grounding
referring and grounding |
0.9 | 1 | 2025 | ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension · ICLR 2025 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.7 | 1 | 2023 | Dynamic Support Network for Few-Shot Class Incremental Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Transfer learning and domain adaptation › few-shot learning
distribution calibration |
0.7 | 1 | 2023 | Learnable Distribution Calibration for Few-Shot Class-Incremental Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.7 | 1 | 2023 | Learnable Distribution Calibration for Few-Shot Class-Incremental Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › Image recognition and object detection › object detection
few-shot object detection |
0.5 | 1 | 2021 | Beyond Max-Margin: Class Margin Equilibrium for Few-Shot Object Detection · CVPR 2021 |
Computer vision › Image recognition and object detection
object detection |
0.5 | 1 | 2021 | Beyond Max-Margin: Class Margin Equilibrium for Few-Shot Object Detection · CVPR 2021 |
Computer vision › Segmentation and scene understanding › semantic segmentation
few-shot segmentation |
0.4 | 1 | 2020 | Prototype Mixture Models for Few-Shot Semantic Segmentation · ECCV (8) 2020 |
Methods — techniques the papers use, named apart from their topics
token collectives · 0.9hybrid perception · 0.9variational inference · 0.7parameterized calibration unit · 0.7dynamic network expansion · 0.7covariance matrix sharing · 0.7class distribution recalling · 0.7class margin loss · 0.5adversarial min-max training · 0.5mixture model · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ClawMachine: Learning to Fetch Visual Tokens for Referential ComprehensionabstractAligning vision and language concepts at a finer level remains an essential topic of multimodal large language models (MLLMs), particularly for tasks such as referring and grounding. Existing methods, such as *proxy encoding* and *geometry encoding* genres, incorporate additional syntax to encode spatial information, imposing extra burdens when communicating between language with vision modules. In this study, we propose ClawMachine, offering a new methodology that explicitly notates each entity using **token collectives**—groups of visual tokens that collaboratively represent higher-level semantics. A hybrid perception mechanism is also explored to perceive and understand scenes from both discrete and continuous spaces. Our method unifies the prompt and answer of visual referential tasks without using additional syntax. By leveraging a joint vision-language vocabulary, ClawMachine integrates referring and grounding in an auto-regressive manner, demonstrating great potential with scaled up pre-training data. Experiments show that ClawMachine achieves superior performance on scene-level and referential understanding tasks with higher efficiency. It also exhibits the potential to integrate multi-source information for complex visual reasoning, which is beyond the capability of many MLLMs. Our code is available at https://github.com/martian422/ClawMachine. Tianren Ma, Lingxi Xie, Yunjie Tian, Boyu Yang 0002, Qixiang Ye |
ICLR | 4 |
| 2023 | Learnable Distribution Calibration for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) faces the challenges of memorizing old class distributions and estimating new class distributions given few training samples. In this study, we propose a learnable distribution calibration (LDC) approach, to systematically solve these two challenges using a unified framework. LDC is built upon a parameterized calibration unit (PCU), which initializes biased distributions for all classes based on classifier vectors (memory-free) and a single covariance matrix. The covariance matrix is shared by all classes, so that the memory costs are fixed. During base training, PCU is endowed with the ability to calibrate biased distributions by recurrently updating sampled features under supervision of real distributions. During incremental learning, PCU recovers distributions for old classes to avoid 'forgetting', as well as estimating distributions and augmenting samples for new classes to alleviate 'over-fitting' caused by the biased distributions of few-shot samples. LDC is theoretically plausible by formatting a variational inference procedure. It improves FSCIL's flexibility as the training procedure requires no class similarity priori. Experiments on CUB200, CIFAR100, and mini-ImageNet datasets show that LDC respectively outperforms the state-of-the-arts by 4.64%, 1.98%, and 3.97%. LDC's effectiveness is also validated on few-shot learning scenarios. Binghao Liu, Boyu Yang 0002, Lingxi Xie, Qi Tian 0001, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Dynamic Support Network for Few-Shot Class Incremental LearningabstractFew-shot class-incremental learning (FSCIL) is challenged by catastrophically forgetting old classes and over-fitting new classes. Revealed by our analyses, the problems are caused by feature distribution crumbling, which leads to class confusion when continuously embedding few samples to a fixed feature space. In this study, we propose a Dynamic Support Network (DSN), which refers to an adaptively updating network with compressive node expansion to "support" the feature space. In each training session, DSN tentatively expands network nodes to enlarge feature representation capacity for incremental classes. It then dynamically compresses the expanded network by node self-activation to pursue compact feature representation, which alleviates over-fitting. Simultaneously, DSN selectively recalls old class distributions during incremental learning to support feature distributions and avoid confusion between classes. DSN with compressive node expansion and class distribution recalling provides a systematic solution for the problems of catastrophic forgetting and overfitting. Experiments on CUB, CIFAR-100, and miniImage datasets show that DSN significantly improves upon the baseline approach, achieving new state-of-the-arts. Boyu Yang 0002, Mingbao Lin, Binghao Liu, Xiaodan Liang, Rongrong Ji, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Part-Based Semantic Transform for Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation remains an open problem for the lack of an effective method to handle the semantic misalignment between objects. In this article, we propose part-based semantic transform (PST) and target at aligning object semantics in support images with those in query images by semantic decomposition-and-match. The semantic decomposition process is implemented with prototype mixture models (PMMs), which use an expectation-maximization (EM) algorithm to decompose object semantics into multiple prototypes corresponding to object parts. The semantic match between prototypes is performed with a min-cost flow module, which encourages correct correspondence while depressing mismatches between object parts. With semantic decomposition-and-match, PST enforces the network's tolerance to objects' appearance and/or pose variation and facilities channelwise and spatial semantic activation of objects in query images. Extensive experiments on Pascal VOC and MS-COCO datasets show that PST significantly improves upon state-of-the-arts. In particular, on MS-COCO, it improves the performance of five-shot semantic segmentation by up to 7.79% with a moderate cost of inference speed and model size. Code for PST is released at https://github.com/Yang-Bob/PST. Boyu Yang 0002, Fang Wan 0001, Chang Liu 0047, Xiangyang Ji, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Beyond Max-Margin: Class Margin Equilibrium for Few-Shot Object DetectionabstractFew-shot object detection has made substantial progress by representing novel class objects using the feature representation learned upon a set of base class objects. However, an implicit contradiction between novel class classification and representation is unfortunately ignored. On the one hand, to achieve accurate novel class classification, the distributions of either two base classes must be far away from each other (max-margin). On the other hand, to precisely represent novel classes, the distributions of base classes should be close to each other to reduce the intra-class distance of novel classes (min-margin). In this paper, we propose a class margin equilibrium (CME) approach, with the aim to optimize both feature space partition and novel class reconstruction in a systematic way. CME first converts the few-shot detection problem to the few-shot classification problem by using a fully connected layer to decouple localization features. CME then reserves adequate margin space for novel classes by introducing simple-yet-effective class margin loss during feature learning. Finally, CME pursues margin equilibrium by disturbing the features of novel class instances in an adversarial min-max fashion. Experiments on Pascal VOC and MS-COCO datasets show that CME significantly improves upon two baseline detectors (up to 3 ~ 5% in average), achieving state-of-the-art performance. Code is available at https://github.com/BohaoLee/CME. Boyu Yang 0002, Chang Liu 0042, Feng Liu 0050, Rongrong Ji, Qixiang Ye |
CVPR | 2 |
| 2020 | Prototype Mixture Models for Few-Shot Semantic Segmentation
Boyu Yang 0002, Chang Liu 0042, Jianbin Jiao, Qixiang Ye |
ECCV (8) | 1 |