EDBT 2026 Demo / reviewers in the wild / expert
Xinyue Huo
dblp:268/5780
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0003-1724-9438ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SAM-CP: Marrying SAM with Composable Prompts for Versatile SegmentationabstractThe Segment Anything model (SAM) has shown a generalized ability to group image pixels into patches, but applying it to semantic-aware segmentation still faces major challenges. This paper presents SAM-CP, a simple approach that establishes two types of composable prompts beyond SAM and composes them for versatile segmentation. Specifically, given a set of classes (in texts) and a set of SAM patches, the Type-I prompt judges whether a SAM patch aligns with a text label, and the Type-II prompt judges whether two SAM patches with the same text label also belong to the same instance. To decrease the complexity in dealing with a large number of semantic classes and patches, we establish a unified framework that calculates the affinity between (semantic and instance) queries and SAM patches, and then merges patches with high affinity to the query. Experiments show that SAM-CP achieves semantic, instance, and panoptic segmentation in both open and closed domains. In particular, it achieves state-of-the-art performance in open-vocabulary segmentation. Our research offers a novel and generalized methodology for equipping vision foundation models like SAM with multi-grained semantic perception abilities. Codes are released on https://github.com/ucas-vg/SAM-CP. Pengfei Chen 0004, Lingxi Xie, Xinyue Huo, Xuehui Yu, Xiaopeng Zhang 0008, Yingfei Sun, Zhenjun Han, Qi Tian 0001 |
ICLR | 3 |
| 2024 | Decoding Matters: Addressing Amplification Bias and Homogeneity Issue in Recommendations for Large Language ModelsabstractAdapting Large Language Models (LLMs) for recommendation requires careful consideration of the decoding process, given the inherent differences between generating items and natural language.Existing approaches often directly apply LLMs' original decoding methods.However, we find these methods encounter significant challenges: 1) amplification bias-where standard length normalization inflates scores for items containing tokens with generation probabilities close to 1 (termed ghost tokens), and 2) homogeneity issue-generating multiple similar or repetitive items for a user.To tackle these challenges, we introduce a new decoding approach named Debiasing-Diversifying Decoding (D 3 ).D 3 disables length normalization for ghost tokens to alleviate amplification bias, and it incorporates a text-free assistant model to encourage tokens less frequently generated by LLMs for counteracting recommendation homogeneity.Extensive experiments on real-world datasets demonstrate the method's effectiveness in enhancing accuracy and diversity.The code is available at https://github. com/SAI990323/DecodingMatters. Keqin Bao, Jizhi Zhang, Yang Zhang 0072, Xinyue Huo, Chong Chen 0001, Fuli Feng |
EMNLP | 4 |
| 2024 | Domain-Agnostic Priors for Semantic Segmentation Under Unsupervised Domain Adaptation and Domain Generalization
Xinyue Huo, Lingxi Xie, Hengtong Hu, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001 |
Int. J. Comput. Vis. | 1 |
| 2024 | One-Bit Supervision for Image Classification: Problem, Solution, and BeyondabstractThis article presents one-bit supervision, a novel setting of learning with fewer labels, for image classification. Instead of the training model using the accurate label of each sample, our setting requires the model to interact with the system by predicting the class label of each sample and learn from the answer whether the guess is correct, which provides one bit (yes or no) of information. An intriguing property of the setting is that the burden of annotation largely is alleviated in comparison to offering the accurate label. There are two keys to one-bit supervision: (i) improving the guess accuracy and (ii) making good use of the incorrect guesses. To achieve these goals, we propose a multi-stage training paradigm and incorporate negative label suppression into an off-the-shelf semi-supervised learning algorithm. Theoretical analysis shows that one-bit annotation is more efficient than full-bit annotation in most cases and gives the conditions of combining our approach with active learning. Inspired by this, we further integrate the one-bit supervision framework into the self-supervised learning algorithm, which yields an even more efficient training schedule. Different from training from scratch, when self-supervised learning is used for initialization, both hard example mining and class balance are verified to be effective in boosting the learning performance. However, these two frameworks still need full-bit labels in the initial stage. To cast off this burden, we utilize unsupervised domain adaptation to train the initial model and conduct pure one-bit annotations on the target dataset. In multiple benchmarks, the learning efficiency of the proposed approach surpasses that using full-bit, semi-supervised supervision. Hengtong Hu, Lingxi Xie, Xinyue Huo, Richang Hong, Qi Tian 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Focus on Your Target: A Dual Teacher-Student Framework for Domain-adaptive Semantic SegmentationabstractWe study unsupervised domain adaptation (UDA) for semantic segmentation. Currently, a popular UDA framework lies in self-training which endows the model with two-fold abilities: (i) learning reliable semantics from the labeled images in the source domain, and (ii) adapting to the target domain via generating pseudo labels on the unlabeled images. We find that, by decreasing/increasing the proportion of training samples from the target domain, the ‘learning ability’ is strengthened/weakened while the ‘adapting ability’ goes in the opposite direction, implying a conflict between these two abilities, especially for a single model. To alleviate the issue, we propose a novel dual teacher-student (DTS) framework and equip it with a bidirectional learning strategy. By increasing the proportion of target-domain data, the second teacher-student model learns to ‘Focus on Your Target’ while the first model is not affected. DTS is easily plugged into existing self-training approaches. In a standard UDA scenario (training on synthetic, labeled data and real, unlabeled data), DTS shows consistent gains over the baselines and sets new state-of-the-art results of 76.5% and 75.1% mIoUs on GTAv→Cityscapes and SYNTHIA→Cityscapes, respectively. The implementation is available at https://github.com/xinyuehuo/DTS. Xinyue Huo, Lingxi Xie, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001 |
ICCV | 1 |
| 2023 | VoxSeP: semi-positive voxels assist self-supervised 3D medical segmentation
Zijie Yang, Lingxi Xie, Xinyue Huo, Longhui Wei, Qi Tian 0001, Sheng Tang |
Multim. Syst. | 4 |
| 2022 | Domain-Agnostic Prior for Transfer Semantic SegmentationabstractUnsupervised domain adaptation (UDA) is an important topic in the computer vision community. The key difficulty lies in defining a common property between the source and target domains so that the source-domain features can align with the target-domain semantics. In this paper, we present a simple and effective mechanism that regularizes cross-domain representation learning with a domain-agnostic prior (DAP) that constrains the features extracted from source and target domains to align with a domain-agnostic space. In practice, this is easily implemented as an extra loss term that requires a little extra costs. In the standard evaluation protocol of transferring synthesized data to real data, we validate the effectiveness of different types of DAP, especially that borrowed from a text embedding model that shows favorable performance beyond the state-of-the-art UDA approaches in terms of segmentation accuracy. Our research reveals that UDA benefits much from better proxies, possibly from other data modalities. Xinyue Huo, Lingxi Xie, Hengtong Hu, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001 |
CVPR | 1 |
| 2022 | Vibration-Based Uncertainty Estimation for Learning from Limited Supervision
Hengtong Hu, Lingxi Xie, Xinyue Huo, Richang Hong, Qi Tian 0001 |
ECCV (30) | 3 |
| 2022 | Finding the Host from the Lesion by Iteratively Mining the Registration GraphabstractVoxel-level annotation has always been a burden of training medical image segmentation models. This paper investigates an interesting problem that finds the host organ of a lesion without actually labeling the organ. To remedy the missing annotation, we construct a graph using an off-the-shelf registration algorithm, on which lesion labels over the training set are accumulated to obtain the pseudo organ for each case. These pseudo labels are used to train a deep network, whose predictions determine the affinity of each lesion on the registration graph. We iteratively update the pseudo labels with the affinity until the training convergence. Our method is evaluated on the MSD Liver and KiTS datasets, without seeing any organ annotation, we achieve the test Dice score of 93% for liver and 92% for kidney, and boosts the accuracy of tumor segmentation to a considerable degree, $3%$, which even surpasses the model trained with ground-truth of both organ and tumor. Zijie Yang, Lingxi Xie, Xinyue Huo, Sheng Tang, Qi Tian 0001, Yongdong Zhang 0001 |
ACM Multimedia | 3 |
| 2022 | Heterogeneous Contrastive Learning: Encoding Spatial Information for Compact Visual RepresentationsabstractUnsupervised pretraining is of great significance for visual representation. Especially, contrastive learning has achieved great success recently, but existing approaches mostly ignored spatial information which is often crucial for visual representation. Strong semantic embedding has an inherent advantage for classification, but dense prediction tasks require more spatial and low-level representation. This paper presentsheterogeneous contrastive learning(HCL), an effective approach that adds spatial information to the encoding stage to alleviate the learning inconsistency between the contrastive objective and strong data augmentation operations. We demonstrate the effectiveness of HCL by showing that (i) it achieves higher accuracy in instance discrimination, (ii) it surpasses existing pre-training methods in a series of downstream tasks (iii) and it shrinks the pre-training costs by half for almost 800 GPU-hours. More importantly, we show that our approach achieves higher efficiency in visual representations, and thus delivers a key message to inspire the future research of self-supervised visual representation learning. Xinyue Huo, Lingxi Xie, Longhui Wei, Xiaopeng Zhang 0008, Xin Chen 0033, Hao Li 0090, Zijie Yang, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | ATSO: Asynchronous Teacher-Student Optimization for Semi-Supervised Image SegmentationabstractSemi-supervised learning is a useful tool for image segmentation, mainly due to its ability in extracting knowledge from unlabeled data to assist learning from labeled data. This paper focuses on a popular pipeline known as self-learning, where we point out a weakness named lazy mimicking that refers to the inertia that a model retains the prediction from itself and thus resists updates. To alleviate this issue, we propose the Asynchronous Teacher-Student Optimization (ATSO) algorithm that (i) breaks up continual learning from teacher to student and (ii) partitions the unlabeled training data into two subsets and alternately uses one subset to fine-tune the model which updates the labels on the other. We show the ability of ATSO on medical and natural image segmentation. In both scenarios, our method reports competitive performance, on par with the state-of-the-arts, in either using partial labeled data in the same dataset or transferring the trained model to an unlabeled dataset. Xinyue Huo, Lingxi Xie, Zijie Yang, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001 |
CVPR | 1 |