EDBT 2026 Demo / reviewers in the wild / expert
Yang Shen 0006
dblp:95/5308-6
· DBLP profile ↗
17ranked-venue papers
6as first author
17since 2021 · last 2025
0000-0002-6344-9951ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LaTexBlend: Scaling Multi-concept Customized Generation with Latent Textual BlendingabstractCustomized text-to-image generation renders user-specified concepts into novel contexts based on textual prompts. Scaling the number of concepts in customized generation meets a broader demand for user creation, whereas existing methods face challenges with generation quality and computational efficiency. In this paper, we propose LaTexBlend, a novel framework for effectively and efficiently scaling multi-concept customized generation. The core idea of LaTexBlend is to represent single concepts and blend multiple concepts within a Latent Textual space, which is positioned after the text encoder and a linear projection. LaTexBlend customizes each concept individually, storing them in a concept bank with a compact representation of latent textual features that captures sufficient concept information to ensure high fidelity. At inference, concepts from the bank can be freely and seamlessly combined in the latent textual space, offering two key merits for multi-concept generation: 1) excellent scalability, and 2) significant reduction of denoising deviation, preserving coherent layouts. Extensive experiments demonstrate that LaTexBlend can flexibly integrate multiple customized concepts with harmonious structures and high subject fidelity, substantially outperforming baselines in both generation quality and computational efficiency. Project page: https://jinjianrick.github.io/latexblend/ Zhenbo Yu, Yang Shen 0006, Zhenyong Fu, Jian Yang 0003 |
CVPR | 3 |
| 2025 | Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot GeneralizationabstractComputer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language Processing (NLP), despite following many of the milestones established in NLP, such as large transformer models, extensive pre-training, and the auto-regression paradigm, among others. In this paper, we rethink the reality that CV adopts discrete and terminological task definitions (e.g., "image segmentation"), and conjecture it is a key barrier that hampers zero-shot task generalization. Our hypothesis is that without truly understanding previously-seen tasks—due to these terminological definitions—deep models struggle to generalize to novel tasks. To verify this, we introduce Explanatory Instructions, which provide an intuitive way to define CV task objectives through detailed linguistic transformations from input images to outputs. We create a large-scale dataset comprising 12 million "image input $\to$ explanatory instruction $\to$ output" triplets, and train an auto-regressive-based vision-language model (AR-based VLM) that takes both images and explanatory instructions as input. By learning to follow these instructions, the AR-based VLM achieves instruction-level zero-shot capabilities for previously-seen tasks and demonstrates strong zero-shot generalization for unseen CV tasks. Code and dataset will be open-sourced. Yang Shen 0006, Xiu-Shen Wei, Yifan Sun 0003, YuXin Song 0001, Heyang Xu, Yazhou Yao, Errui Ding |
ICML | 1 |
| 2025 | UniCanvas: Affordance-Aware Unified Real Image Editing via Customized Text-to-Image Generation
Yang Shen 0006, Zhenyong Fu, Jian Yang 0003 |
Int. J. Comput. Vis. | 2 |
| 2025 | Equiangular Basis Vectors: A Novel Paradigm for Classification Tasks
Yang Shen 0006, Xuhao Sun, Xiu-Shen Wei, Lingyan Gao |
Int. J. Comput. Vis. | 1 |
| 2025 | An Empirical Study on Training Paradigms for Deep Supervised Hashing
Yang Shen 0006, Peng Wang 0023, Xiu-Shen Wei, Yazhou Yao |
Int. J. Comput. Vis. | 1 |
| 2025 | Delving Deep into Simplicity Bias for Long-Tailed Image Recognition
Xiu-Shen Wei, Xuhao Sun, Yang Shen 0006, Peng Wang 0023 |
Int. J. Comput. Vis. | 3 |
| 2025 | Prune and Merge: Efficient Token Compression for Vision Transformer With Spatial Information PreservedabstractToken compression is essential for reducing the computational and memory requirements of transformer models, enabling their deployment in resource-constrained environments. In this work, we propose an efficient and hardware-compatible token compression method called Prune and Merge. Our approach integrates token pruning and merging operations within transformer models to achieve layer-wise token compression. By introducing trainable merge and reconstruct matrices and utilizing shortcut connections, we efficiently merge tokens while preserving important information and enabling the restoration of pruned tokens. Additionally, we introduce a novel gradient-weighted attention scoring mechanism that computes token importance scores during the training phase, eliminating the need for separate computations during inference and enhancing compression efficiency. We also leverage gradient information to capture the global impact of tokens and automatically identify optimal compression structures. Extensive experiments on the ImageNet-1 k and ADE20 K datasets validate the effectiveness of our approach, achieving significant speed-ups with minimal accuracy degradation compared to state-of-the-art methods. For instance, on DeiT-Small, we achieve a 1.64× speed-up with only a 0.2% drop in accuracy on ImageNet-1k. Moreover, by compressing segmenter models and comparing with existing methods, we demonstrate the superior performance of our approach in terms of efficiency and effectiveness. Junzhu Mao, Yang Shen 0006, Jinyang Guo 0002, Yazhou Yao, Xian-Sheng Hua 0001, Heng Tao Shen |
IEEE Trans. Multim. | 2 |
| 2024 | Customized Generation Reimagined: Fidelity and Editability Harmonized
Yang Shen 0006, Zhenyong Fu, Jian Yang 0003 |
ECCV (50) | 2 |
| 2024 | Few-shot open-set recognition via pairwise discriminant aggregation
Yang Shen 0006, Zhenyong Fu, Jian Yang 0003 |
Neurocomputing | 2 |
| 2023 | Equiangular Basis VectorsabstractWe propose Equiangular Basis Vectors (EBVs) for classification tasks. In deep neural networks, models usually end with a k-way fully connected layer with softmax to handle different classification tasks. The learning objective of these methods can be summarized as mapping the learned feature representations to the samples' label space. While in metric learning approaches, the main objective is to learn a transformation function that maps training data points from the original space to a new space where similar points are closer while dissimilar points become farther apart. Different from previous methods, our EBVs generate normalized vector embeddings as “predefined classifiers” which are required to not only be with the equal status between each other, but also be as orthogonal as possible. By minimizing the spherical distance of the embedding of an input between its categorical EBV in training, the predictions can be obtained by identifying the categorical EBV with the smallest distance during inference. Various experiments on the ImageNet-1K dataset and other downstream tasks demonstrate that our method outperforms the general fully connected classifier while it does not introduce huge additional computation compared with classical metric learning methods. Our EBVs won the first place in the 2022 DIGIX Global AI Challenge, and our code is open-source and available at https://github.com/NJUST-VIPGroup/Equiangular-Basis-Vectors. Yang Shen 0006, Xuhao Sun, Xiu-Shen Wei |
CVPR | 1 |
| 2023 | Hawkeye: A PyTorch-based Library for Fine-Grained Image Recognition with Deep LearningabstractFine-Grained Image Recognition (FGIR) is a fundamental and challenging task in computer vision and multimedia that plays a crucial role in Intellectual Economy and Industrial Internet applications. However, the absence of a unified open-source software library covering various paradigms in FGIR poses a significant challenge for researchers and practitioners in the field. To address this gap, we present Hawkeye, a PyTorch-based library for FGIR with deep learning. Hawkeye is designed with a modular architecture, emphasizing high-quality code and human-readable configuration, providing a comprehensive solution for FGIR tasks. In Hawkeye, we have implemented 16 state-of-the-art fine-grained methods, covering 6 different paradigms, enabling users to explore various approaches for FGIR. To the best of our knowledge, Hawkeye represents the first open-source PyTorch-based library dedicated to FGIR. It is publicly available at https://github.com/Hawkeye-FineGrained/Hawkeye/, providing researchers and practitioners with a powerful tool to advance their research and development in the field of FGIR. Yang Shen 0006, Xiu-Shen Wei, Ye Wu 0001 |
ACM Multimedia | 2 |
| 2023 | Attribute-Aware Deep Hashing With Self-Consistency for Large-Scale Fine-Grained Image RetrievalabstractOur work focuses on tackling large-scale fine-grained image retrieval as ranking the images depicting the concept of interests (i.e., the same sub-category labels) highest based on the fine-grained details in the query. It is desirable to alleviate the challenges of both fine-grained nature of small inter-class variations with large intra-class variations and explosive growth of fine-grained data for such a practical task. In this paper, we propose attribute-aware hashing networks with self-consistency for generating attribute-aware hash codes to not only make the retrieval process efficient, but also establish explicit correspondences between hash codes and visual attributes. Specifically, based on the captured visual representations by attention, we develop an encoder-decoder structure network of a reconstruction task to unsupervisedly distill high-level attribute-specific vectors from the appearance-specific visual representations without attribute annotations. Our models are also equipped with a feature decorrelation constraint upon these attribute vectors to strengthen their representative abilities. Then, driven by preserving original entities' similarity, the required hash codes can be generated from these attribute-specific vectors and thus become attribute-aware. Furthermore, to combat simplicity bias in deep hashing, we consider the model design from the perspective of the self-consistency principle and propose to further enhance models' self-consistency by equipping an additional image reconstruction path. Comprehensive quantitative experiments under diverse empirical settings on six fine-grained retrieval datasets and two generic retrieval datasets show the superiority of our models over competing methods. Moreover, qualitative results demonstrate that not only the obtained hash codes can strongly correspond to certain kinds of crucial properties of fine-grained objects, but also our self-consistency designs can effectively overcome simplicity bias in fine-grained hashing. Xiu-Shen Wei, Yang Shen 0006, Xuhao Sun, Peng Wang 0023, Yuxin Peng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Automatic Check-Out via Prototype-Based Classifier Learning from Single-Product Exemplars
Hao Chen 0052, Xiu-Shen Wei, Faen Zhang, Yang Shen 0006, Liang Xiao 0001 |
ECCV (25) | 4 |
| 2022 | SEMICON: A Learning-to-Hash Solution for Large-Scale Fine-Grained Image Retrieval
Yang Shen 0006, Xuhao Sun, Xiu-Shen Wei, Qing-Yuan Jiang, Jian Yang 0003 |
ECCV (14) | 1 |
| 2022 | A Channel Mix Method for Fine-Grained Cross-Modal RetrievalabstractIn this paper, we propose a simple but effective method for dealing with the challenging fine-grained cross-modal retrieval task where it aims to enable flexible retrieval among subor-dinate categories across different modalities. Specifically, in order to enhance information interaction in different modalities for fine-grained objects, a channel mix method is developed and performed upon the channels of deep activations across dif-ferent modalities. After that, a 1 x 1 convolution is employed to aggregate the mixed channels into a unified feature vector. Moreover, equipped with a novel fine-grained cross-modal cen-ter loss, our method can further improve the intra-class separa-bility as well as inter-class compactness for multi-modalities. Experiments are conducted on the fine-grained cross-modal benchmark dataset and show our superiority over competing methods. Meanwhile, ablation studies also demonstrate the effectiveness of our proposals. Yang Shen 0006, Xuhao Sun, Xiu-Shen Wei, Hanxu Hu |
ICME | 1 |
| 2022 | Webly-Supervised Fine-Grained Recognition with Partial Label LearningabstractThe task of webly-supervised fine-grained recognition is to boost recognition accuracy of classifying subordinate categories (e.g., different bird species) by utilizing freely available but noisy web data. As the label noises significantly hurt the network training, it is desirable to distinguish and eliminate noisy images. In this paper, we propose two strategies, i.e., open-set noise removal and closed-set noise correction, to both remove such two kinds of web noises w.r.t. fine-grained recognition. Specifically, for open-set noise removal, we utilize a pre-trained deep model to perform deep descriptor transformation to estimate the positive correlation between these web images, and detect the open-set noises based on the correlation values. Regarding closed-set noise correction, we develop a top-k recall optimization loss for firstly assigning a label set towards each web image to reduce the impact of hard label assignment for closed-set noises. Then, we further propose to correct the sample with its label set as the true single label from a partial label learning perspective. Experiments on several webly-supervised fine-grained benchmark datasets show that our method obviously outperforms other existing state-of-the-art methods. Yu-Yan Xu, Yang Shen 0006, Xiu-Shen Wei, Jian Yang 0003 |
IJCAI | 2 |
| 2021 | A$^2$-Net: Learning Attribute-Aware Hash Codes for Large-Scale Fine-Grained Image RetrievalabstractOur work focuses on tackling large-scale fine-grained image retrieval as ranking the images depicting the concept of interests (i.e., the same sub-category labels) highest based on the fine-grained details in the query. It is desirable to alleviate the challenges of both fine-grained nature of small inter-class variations with large intra-class variations and explosive growth of fine-grained data for such a practical task. In this paper, we propose an Attribute-Aware hashing Network (A$^2$-Net) for generating attribute-aware hash codes to not only make the retrieval process efficient, but also establish explicit correspondences between hash codes and visual attributes. Specifically, based on the captured visual representations by attention, we develop an encoder-decoder structure network of a reconstruction task to unsupervisedly distill high-level attribute-specific vectors from the appearance-specific visual representations without attribute annotations. A$^2$-Net is also equipped with a feature decorrelation constraint upon these attribute vectors to enhance their representation abilities. Finally, the required hash codes are generated by the attribute vectors driven by preserving original similarities. Qualitative experiments on five benchmark fine-grained datasets show our superiority over competing methods. More importantly, quantitative results demonstrate the obtained hash codes can strongly correspond to certain kinds of crucial properties of fine-grained objects. Xiu-Shen Wei, Yang Shen 0006, Xuhao Sun, Han-Jia Ye, Jian Yang 0003 |
NeurIPS | 2 |