VLDB 2026 Research / reviewers in the wild / expert
Yantao Shen 0002
dblp:86/3372-2
· DBLP profile ↗
15ranked-venue papers
8as first author
9since 2021 · last 2024
0000-0001-5413-2445ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Face, body and person analysis · 18% Efficient and distributed learning · 14% Representation and self-supervised learning · 13% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 24 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis
person re-identification |
1.8 | 5 | 2021 | Person Re-Identification With Deep Kronecker-Product Matching and Group-Shuffling Random Walk · IEEE Trans. Pattern Anal. Mach. Intell. 2021 Person Re-identification with Deep Similarity-Guided Graph Neural Network · ECCV (15) 2018 Improving Deep Visual Representation for Person Re-identification by Global and Local Image-language Association · ECCV (16) 2018 |
Computer vision › Segmentation and scene understanding
medical image segmentation |
1.1 | 2 | 2022 | Multiple Consistency Supervision based Semi-supervised OCT Segmentation using Very Limited Annotations · ICRA 2022 Multibranch Learning for Angiodysplasia Segmentation with Attention-Guided Networks and Domain Adaptation · ICRA 2021 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
ensemble distillation |
0.8 | 1 | 2024 | Elodi: Ensemble Logit Difference Inhibition for Positive-Congruent Training · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.8 | 1 | 2024 | Elodi: Ensemble Logit Difference Inhibition for Positive-Congruent Training · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Deep learning architectures and training
memory mechanism |
0.8 | 1 | 2024 | B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory · NeurIPS 2024 |
Natural language and speech › Language models and text generation › knowledge editing
negative flip reduction |
0.8 | 1 | 2024 | Elodi: Ensemble Logit Difference Inhibition for Positive-Congruent Training · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
open-world learning |
0.8 | 1 | 2024 | Open-World Dynamic Prompt and Continual Visual Representation Learning · ECCV (49) 2024 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.8 | 1 | 2024 | B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory · NeurIPS 2024 |
Machine learning › Learning paradigms › semi-supervised learning
consistency regularization |
0.6 | 1 | 2022 | Multiple Consistency Supervision based Semi-supervised OCT Segmentation using Very Limited Annotations · ICRA 2022 |
Machine learning › Learning paradigms
semi-supervised learning |
0.6 | 1 | 2022 | Multiple Consistency Supervision based Semi-supervised OCT Segmentation using Very Limited Annotations · ICRA 2022 |
Computer vision › Image recognition and object detection
image retrieval |
0.5 | 1 | 2021 | Person Re-Identification With Deep Kronecker-Product Matching and Group-Shuffling Random Walk · IEEE Trans. Pattern Anal. Mach. Intell. 2021 |
Medical and health informatics
computer-aided diagnosis |
0.5 | 1 | 2021 | Multibranch Learning for Angiodysplasia Segmentation with Attention-Guided Networks and Domain Adaptation · ICRA 2021 |
Machine learning › Representation and self-supervised learning › transferable representation › transferable representation learning
backward-compatible representation learning |
0.4 | 1 | 2020 | Towards Backward-Compatible Representation Learning · CVPR 2020 |
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning |
0.4 | 1 | 2020 | Towards Backward-Compatible Representation Learning · CVPR 2020 |
Computer vision › Face, body and person analysis › person re-identification
deep person re-identification |
0.3 | 1 | 2018 | End-to-End Deep Kronecker-Product Matching for Person Re-Identification · CVPR 2018 |
Computer vision › 3D vision
feature matching |
0.3 | 1 | 2018 | End-to-End Deep Kronecker-Product Matching for Person Re-Identification · CVPR 2018 |
Machine learning › Graph learning
graph neural network |
0.3 | 1 | 2018 | Person Re-identification with Deep Similarity-Guided Graph Neural Network · ECCV (15) 2018 |
Information retrieval
image retrieval |
0.3 | 1 | 2018 | Deep Group-Shuffling Random Walk for Person Re-Identification · CVPR 2018 |
Computer vision › Face, body and person analysis › person re-identification
vehicle re-identification |
0.3 | 1 | 2017 | Learning Deep Neural Networks for Vehicle Re-ID with Visual-spatio-Temporal Path Proposals · ICCV 2017 |
Computer vision › Image recognition and object detection
image classification |
0.2 | 1 | 2024 | Elodi: Ensemble Logit Difference Inhibition for Positive-Congruent Training · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Natural language and speech › Language models and text generation
language modeling |
0.2 | 1 | 2024 | B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation learning › metric learning
deep metric learning |
0.1 | 1 | 2021 | Person Re-Identification With Deep Kronecker-Product Matching and Group-Shuffling Random Walk · IEEE Trans. Pattern Anal. Mach. Intell. 2021 |
Computer vision › Face, body and person analysis
face recognition |
0.1 | 1 | 2020 | Towards Backward-Compatible Representation Learning · CVPR 2020 |
Computer vision › Image recognition and object detection
image similarity |
0.1 | 1 | 2017 | Learning Deep Neural Networks for Vehicle Re-ID with Visual-spatio-Temporal Path Proposals · ICCV 2017 |
Methods — techniques the papers use, named apart from their topics
kronecker product matching · 0.8stochastic realization theory · 0.8state space model · 0.8logit difference inhibition · 0.8ensemble distillation · 0.8dynamic prompting · 0.8continual learning · 0.8scaling transformation · 0.6feature perturbation · 0.6encoder-decoder network · 0.6domain adversarial training · 0.5domain adaptation · 0.5convolutional neural network · 0.5attention mechanism · 0.5random walk · 0.3group shuffling · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Open-World Dynamic Prompt and Continual Visual Representation Learning
Youngeun Kim, Zhaowei Cai, Yantao Shen 0002, Rahul Duggal, Dripta S. Raychaudhuri, Zhuowen Tu, Yifan Xing, Onkar Dabeer |
ECCV (49) | 5 |
| 2024 | B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading MemoryabstractWe describe a family of architectures to support transductive inference by allowing memory to grow to a finite but a-priori unknown bound while making efficient use of finite resources for inference. Current architectures use such resources to represent data either eidetically over a finite span ('context' in Transformers), or fading over an infinite span (in State Space Models, or SSMs). Recent hybrid architectures have combined eidetic and fading memory, but with limitations that do not allow the designer or the learning process to seamlessly modulate the two, nor to extend the eidetic memory span. We leverage ideas from Stochastic Realization Theory to develop a class of models called B'MOJO to seamlessly combine eidetic and fading memory within an elementary composable module. The overall architecture can be used to implement models that can access short-term eidetic memory 'in-context,' permanent structural memory 'in-weights,' fading memory 'in-state,' and long-term eidetic memory 'in-storage' by natively incorporating retrieval from an asynchronously updated memory. We show that Transformers, existing SSMs such as Mamba, and hybrid architectures such as Jamba are special cases of B'MOJO and describe a basic implementation that can be stacked and scaled efficiently in hardware. We test B'MOJO on transductive inference tasks, such as associative recall, where it outperforms existing SSMs and Hybrid models; as a baseline, we test ordinary language modeling where B'MOJO achieves perplexity comparable to similarly-sized Transformers and SSMs up to 1.4B parameters, while being up to 10% faster to train. Finally, we test whether models trained inductively on a-priori bounded sequences (up to 8K tokens) can still perform transductive inference on sequences many-fold longer. B'MOJO's ability to modulate eidetic and fading memory results in better inference on longer sequences tested up to 32K tokens, four-fold the length of the longest sequences seen during training. Luca Zancato, Arjun Seshadri, Yonatan Dukler, Aditya Golatkar, Yantao Shen 0002, Benjamin Bowman, Matthew Trager, Alessandro Achille, Stefano Soatto |
NeurIPS | 5 |
| 2024 | Elodi: Ensemble Logit Difference Inhibition for Positive-Congruent TrainingabstractNegative flips are errors introduced in a classification system when a legacy model is updated. Existing methods to reduce the negative flip rate (NFR) either do so at the expense of overall accuracy by forcing a new model to imitate the old models, or use ensembles, which multiply inference cost prohibitively. We analyze the role of ensembles in reducing NFR and observe that they remove negative flips that are typically not close to the decision boundary, but often exhibit large deviations in the distance among their logits. Based on the observation, we present a method, called Ensemble Logit Difference Inhibition (ELODI), to train a classification system that achieves paragon performance in both error rate and NFR, at the inference cost of a single model. The method distills a homogeneous ensemble to a single student model which is used to update the classification system. ELODI also introduces a generalized distillation objective, Logit Difference Inhibition (LDI), which only penalizes the logit difference of a subset of classes with the highest logit values. On multiple image classification benchmarks, model updates with ELODI demonstrate superior accuracy retention and NFR reduction. Yue Zhao 0006, Yantao Shen 0002, Yuanjun Xiong, Shuo Yang 0003, Wei Xia 0009, Zhuowen Tu, Bernt Schiele, Stefano Soatto |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Multiple Consistency Supervision based Semi-supervised OCT Segmentation using Very Limited AnnotationsabstractOptical Coherence Tomography (OCT) is a rapidly growing and promising imaging technique, enabling non-invasive high-resolution visualization of biological tissues. Segmentation of tissue structures from OCT scans is essen-tial for disease diagnosis but remains challenging for the blurry boundaries and large volumes. Deep learning-based OCT segmentation algorithms always require large numbers of annotations for satisfying performance, which is hard to meet since manually labeling is time-consuming and labor-intensive. Therefore, we propose a novel semi-supervised OCT segmentation framework utilizing very few labeled scans, i.e., 5 samples, and abundant unlabeled data. Specifically, our framework con-sists of one shared encoder and two different decoder branches. For the two branches, we design a strong augmentation-consistent supervision module and a scaling transformation-consistent supervision module respectively to improve their generalization ability. Besides, cross consistency supervision with feature perturbations between two branches is proposed to incorporate their advantages for further regularization. With such multiple consistency supervision, we aim to enrich the diversity of unsupervised information so as to make full use of labeled and unlabeled data. Experimental results on a public retinal OCT dataset demonstrate the effectiveness of our method, achieving an average dice score of 87.25% in the case of only 5 labeled samples used. It outperforms the supervised baseline by 3.46% and the best semi-supervised model by 1.42% in our experiments. Yantao Shen 0002, Xiaohan Xing, Max Q.-H. Meng |
ICRA | 2 |
| 2022 | Discrepancy-Based Active Learning for Weakly Supervised Bleeding Segmentation in Wireless Capsule Endoscopy Images
Fan Bai 0008, Xiaohan Xing, Yantao Shen 0002, Max Q.-H. Meng |
MICCAI (8) | 3 |
| 2022 | Task-Relevant Feature Replenishment for Cross-Centre Polyp Segmentation
Yantao Shen 0002, Xiao Jia 0005, Fan Bai 0008, Max Q.-H. Meng |
MICCAI (4) | 1 |
| 2021 | Multibranch Learning for Angiodysplasia Segmentation with Attention-Guided Networks and Domain AdaptationabstractAs a common cause of anemia and gastrointestinal bleeding, angiodysplasia (AD) diagnosis in wireless capsule endoscopy (WCE) images is important in clinical. Current manual review requires undivided concentration of the gastroenterologists, which is laborious and time-consuming. The development of computational methods that can assist automated diagnosis of angiodysplasia is highly desirable. In this paper, we present a new approach, ADNet, for angiodysplasia segmentation using convolutional neural networks (CNNs). Compared with previous learning strategies, ADNet gains accuracy from attentionguided and domain-adversarial training via a multibranch CNN architecture. Specifically, the core branch is constructed for AD segmentation in a fully convolutional manner. Then we propose an attention module embedded in the attention branch to enhance network feature learning, which allows ADNet to focus on the most informative and AD relevant regions while processing. Furthermore, an adaptation branch is built to learn domain-invariant features by adversarial training, aiming to improve the performance when datasets are expanded while preventing the degradation induced by the variations in WCE image acquisition. ADNet is evaluated using two WCE datasets with angiodysplasia and the results show the accuracy gains we obtain, where the state-of-the-art segmentation performance on the public dataset of GIANA’17 is achieved. Xiao Jia 0005, Xiaochun Mai, Xiaohan Xing, Yantao Shen 0002, Jiankun Wang 0001, Max Q.-H. Meng |
ICRA | 4 |
| 2021 | HRENet: A Hard Region Enhancement Network for Polyp Segmentation
Yantao Shen 0002, Xiao Jia 0005, Max Q.-H. Meng |
MICCAI (1) | 1 |
| 2021 | Person Re-Identification With Deep Kronecker-Product Matching and Group-Shuffling Random WalkabstractPerson re-identification (re-ID) aims to robustly measure visual affinities between person images. It has wide applications in intelligent surveillance by associating same persons' images across multiple cameras. It is generally treated as an image retrieval problem: given a probe person image, the affinities between the probe image and gallery images (P2G affinities) are used to rank the retrieved gallery images. There exist two main challenges for effectively solving this problem. 1) Person images usually show significant variations because of different person poses and viewing angles. The spatial layouts and correspondences between person images are therefore vital information for tackling this problem. State-of-the-art methods either ignore such spatial variation or utilize extra pose information for handling the challenge. 2) Most existing person re-ID methods rank gallery images considering only P2G affinities but ignore the affinities between the gallery images (G2G affinity). Such affinities could provide important clues for accurate gallery image ranking but were only utilized in post-processing stages by current methods. In this article, we propose a unified end-to-end deep learning framework to tackle the two challenges. For handling viewpoint and pose variations between compared person images, we propose a novel Kronecker Product Matching operation to match and warp feature maps of different persons. Comparing warped feature maps results in more accurate P2G affinities. To fully utilize all available P2G and G2G affinities for accurately ranking gallery person images, a novel group-shuffling random walk operation is proposed. Both Kronecker Product Matching and Group-shuffling Random Walk operations are end-to-end trainable and are shown to improve the learned visual features if integrated in the deep learning framework. The proposed approach outperforms state-of-the-art methods on Market-1501, CUHK03 and DukeMTMC datasets, which demonstrates the effectiveness and generalization ability of our proposed approach. Code is available at https://github.com/YantaoShen/kpm_rw_person_reid. Yantao Shen 0002, Tong Xiao 0003, Shuai Yi, Dapeng Chen, Xiaogang Wang 0001, Hongsheng Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Towards Backward-Compatible Representation LearningabstractWe propose a way to learn visual features that are compatible with previously computed ones even when they have different dimensions and are learned via different neural network architectures and loss functions. Compatible means that, if such features are used to compare images, then ``new'' features can be compared directly to ``old'' features, so they can be used interchangeably. This enables visual search systems to bypass computing new features for all previously seen images when updating the embedding models, a process known as backfilling. Backward compatibility is critical to quickly deploy new embedding models that leverage ever-growing large-scale training datasets and improvements in deep learning architectures and training methods. We propose a framework to train embedding models, called backward-compatible training (BCT), as a first step towards backward compatible representation learning. In experiments on learning embeddings for face recognition, models trained with BCT successfully achieve backward compatibility without sacrificing accuracy, thus enabling backfill-free model updates of visual embeddings. Yantao Shen 0002, Yuanjun Xiong, Wei Xia 0009, Stefano Soatto |
CVPR | 1 |
| 2018 | Deep Group-Shuffling Random Walk for Person Re-IdentificationabstractPerson re-identification aims at finding a person of interest in an image gallery by comparing the probe image of this person with all the gallery images. It is generally treated as a retrieval problem, where the affinities between the probe image and gallery images (P2G affinities) are used to rank the retrieved gallery images. However, most existing methods only consider P2G affinities but ignore the affinities between all the gallery images (G2G affinity). Some frameworks incorporated G2G affinities into the testing process, which is not end-to-end trainable for deep neural networks. In this paper, we propose a novel group-shuffling random walk network for fully utilizing the affinity information between gallery images in both the training and testing processes. The proposed approach aims at end-to-end refining the P2G affinities based on G2G affinity information with a simple yet effective matrix operation, which can be integrated into deep neural networks. Feature grouping and group shuffle are also proposed to apply rich supervisions for learning better person features. The proposed approach outperforms state-of-the-art methods on the Market-1501, CUHK03, and DukeMTMC datasets by large margins, which demonstrate the effectiveness of our approach. Yantao Shen 0002, Hongsheng Li 0001, Tong Xiao 0003, Shuai Yi, Dapeng Chen, Xiaogang Wang 0001 |
CVPR | 1 |
| 2018 | End-to-End Deep Kronecker-Product Matching for Person Re-IdentificationabstractPerson re-identification aims to robustly measure similarities between person images. The significant variation of person poses and viewing angles challenges for accurate person re-identification. The spatial layout and correspondences between query person images are vital information for tackling this problem but are ignored by most state-of-the-art methods. In this paper, we propose a novel Kronecker Product Matching module to match feature maps of different persons in an end-to-end trainable deep neural network. A novel feature soft warping scheme is designed for aligning the feature maps based on matching results, which is shown to be crucial for achieving superior accuracy. The multi-scale features based on hourglass-like networks and self residual attention are also exploited to further boost the re-identification performance. The proposed approach outperforms state-of-the-art methods on the Market-1501, CUHK03, and DukeMTMC datasets, which demonstrates the effectiveness and generalization ability of our proposed approach. Yantao Shen 0002, Tong Xiao 0003, Hongsheng Li 0001, Shuai Yi, Xiaogang Wang 0001 |
CVPR | 1 |
| 2018 | Improving Deep Visual Representation for Person Re-identification by Global and Local Image-language Association
Dapeng Chen, Hongsheng Li 0001, Xihui Liu, Yantao Shen 0002, Zejian Yuan, Xiaogang Wang 0001 |
ECCV (16) | 4 |
| 2018 | Person Re-identification with Deep Similarity-Guided Graph Neural Network
Yantao Shen 0002, Hongsheng Li 0001, Shuai Yi, Dapeng Chen, Xiaogang Wang 0001 |
ECCV (15) | 1 |
| 2017 | Learning Deep Neural Networks for Vehicle Re-ID with Visual-spatio-Temporal Path ProposalsabstractVehicle re-identification is an important problem and has many applications in video surveillance and intelligent transportation. It gains increasing attention because of the recent advances of person re-identification techniques. However, unlike person re-identification, the visual differences between pairs of vehicle images are usually subtle and even challenging for humans to distinguish. Incorporating additional spatio-temporal information is vital for solving the challenging re-identification task. Existing vehicle re-identification methods ignored or used oversimplified models for the spatio-temporal relations between vehicle images. In this paper, we propose a two-stage framework that incorporates complex spatio-temporal information for effectively regularizing the re-identification results. Given a pair of vehicle images with their spatiotemporal information, a candidate visual-spatio-temporal path is first generated by a chain MRF model with a deeply learned potential function, where each visual-spatiotemporal state corresponds to an actual vehicle image with its spatio-temporal information. A Siamese-CNN+Path- LSTM model takes the candidate path as well as the pairwise queries to generate their similarity score. Extensive experiments and analysis show the effectiveness of our proposed method and individual components. Yantao Shen 0002, Tong Xiao 0003, Hongsheng Li 0001, Shuai Yi, Xiaogang Wang 0001 |
ICCV | 1 |