EDBT 2026 Demo / reviewers in the wild / expert
Xuan Jin
dblp:80/5681
· DBLP profile ↗
12ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Semi-VFL: Communication-efficient Few-label Vertical Federated Learning with Stacked Generalization and Model-level ConsistencyabstractVertical federated learning (VFL) is a collaborative learning scheme where clients share some overlapping samples but have different feature spaces. Existing VFL schemes are restricted in model performance and deployment feasibility due to the scarcity of overlapping labeled samples and high communication costs. To tackle these issues, we propose a practical VFL scheme Semi-VFL using stacked generalization, which can effectively improve model performance with limited aligned labeled samples and only two communication times. A built-in local semi-supervised learning strategy FewMatch with model-level consistency for few-label VFL setting is designed in our scheme. Extensive experiments indicate the superiority of Semi-VFL in both image and tabular datasets. Specifically, Semi-VFL can achieve accuracy improvement by more than 10.9% and communication cost reduction by more than 380× over the state-of-the-art few-shot VFL scheme on CIFAR-10 with 128 aligned samples. Xuan Jin, Yuanzhi Yao, Caihong Kai, Rui Wang 0043, Nenghai Yu |
ICASSP | 1 |
| 2024 | One-dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing ApplicationsabstractThe prevalent use of commercial and open-source diffusion models (DMs) for text-to-image generation prompts risk mitigation to prevent undesired behaviors. Existing concept erasing methods in academia are all based on full parameter or specification-based fine-tuning, from which we observe the following issues: 1) Generation alteration towards erosion: Parameter drift during target elimination causes alterations and potential deformations across all generations, even eroding other concepts at varying degrees, which is more evident with multi-concept erased; 2) Transfer in-ability & deployment inefficiency: Previous model-specific erasure impedes the flexible combination of concepts and the training-free transfer towards other models, resulting in linear cost growth as the deployment scenarios increase. To achieve non-invasive, precise, customizable, and transferable elimination, we ground our erasing framework on one-dimensional adapters to erase multiple concepts from most DMs at once across versatile erasing applications. The concept-SemiPermeable structure is injected as a Membrane (SPM) into any DM to learn targeted erasing, and mean-time the alteration and erosion phenomenon is effectively mitigated via a novel Latent Anchoring fine-tuning strategy. Once obtained, SPMs can be flexibly combined and plug-and-play for other DMs without specific re-tuning, enabling timely and efficient adaptation to diverse scenarios. During generation, our Facilitated Transport mechanism dynamically regulates the permeability of each SPM to re-spond to different input prompts, further minimizing the impact on other concepts. Quantitative and qualitative results across ~40 concepts, 7 DMs and 4 erasing applications have demonstrated the superior erasing of SPM. Our code and pre-tuned SPMs are available on the project page https:/lyumengyao.github.io/projects/spm. Mengyao Lyu, Yuhong Yang 0008, Haiwen Hong, Hui Chen 0013, Xuan Jin, Yuan He 0011, Hui Xue 0001, Jungong Han, Guiguang Ding |
CVPR | 5 |
| 2023 | Combat COVID-19 at National Level using Risk Stratification with Appropriate InterventionabstractIn the national battle against COVID-19, harnessing population-level big data is imperative, enabling authorities to devise effective care policies, allocate healthcare resources efficiently, and enact targeted interventions. Singapore adopted the Home Recovery Programme (HRP) in September 2021, diverting low-risk COVID-19 patients to home care to ease hospital burdens amid high vaccination rates and mild symptoms. While a patient’s suitability for HRP could be assessed using broad-based criteria, integrating machine learning (ML) model becomes invaluable for identifying high-risk patients prone to severe illness, facilitating early medical assessment. Most prior studies have traditionally depended on clinical and laboratory data, necessitating initial clinic or hospital evaluations. None of these studies incorporated vaccination status, a crucial variable in a well-vaccinated population. This paper proposes a machine learning approach to nationwide risk stratification, offering intervention recommendations by harnessing nationwide datasets. Our best-performing ML model, XGBoost achieves an AUROC of 0.930 utilizing data from multiple data sources including patients’ demographic information, vaccination status and medical history. For broader applicability, we also propose a parsimonious XGBoost model with an AUROC of 0.885 with a selection of five commonly collected variables, namely age, number of vaccine doses taken and number of days since the first, second and booster doses. Importantly, both of our proposed models achieve robust predictive performance without requiring the collection of clinical or laboratory data from patients. We believe that the parsimonious model, leveraging easily attainable data, has the potential for broader adoption across diverse nations, ultimately delivering paramount value to their populations. Xuan Jin, Kar Way Tan |
IEEE Big Data | 1 |
| 2023 | Dual Collaborative Visual-Semantic Mapping for Multi-Label Zero-Shot Image RecognitionabstractMulti-label zero-shot learning (ML-ZSL), with the difficulty of both multi-label learning and zero-shot learning, aims to recognize various unseen objects that are not observed during training. Previous methods mainly use a single directional visual-semantic mapping to associate the visual and semantic embedding space, which is not sufficient to adequately realize knowledge transfer from seen to unseen classes. In this paper, we propose a novel dual collaborative visual-semantic mapping framework, constructing abundant connection relationships by exploring two aspects of mapping streams, i.e., the visual-to-semantic (V2S) mapping and the semantic-to-visual (S2V) mapping. Through the collaborative learning of these two effective mappings, our method achieves state-of-the-art performance on the MS-COCO and PASCAL-VOC, two benchmarks for ML-ZSL. Yunqing Hu, Xuan Jin, Yin Zhang 0006 |
ICASSP | 2 |
| 2023 | Open-Vocabulary Object Detection With an Open CorpusabstractExisting open vocabulary object detection (OVD) works expand the object detector toward open categories by replacing the classifier with the category text embeddings and optimizing the region-text alignment on data of the base categories. However, both the class-agnostic proposal generator and the classifier are biased to the seen classes as demonstrated by the gaps of objectness and accuracy assessment between base and novel classes. In this paper, an open corpus, composed of a set of external object concepts and clustered to several centroids, is introduced to improve the generalization ability in the detector. We propose the generalized objectness assessment (GOAT) in the proposal generator based on the visual-text alignment, where the similarities of visual feature to the cluster centroids are summarized as the objectness. This simple heuristic evaluates objectness with concepts in open corpus and is thus generalized to open categories. We further propose category expanding (CE) with open corpus in two training tasks, which enables the detector to perceive more categories in the feature space and get more reasonable optimization direction. For the classification task, we introduce an open corpus classifier by reconstructing original classifier with similar words in text space. For the image-caption alignment task, the open corpus centroids are incorporated to enlarge the negative samples in the contrastive loss. Extensive experiments demonstrate the effectiveness of GOAT and CE, which greatly improve the performance on novel classes and get new state-of-the-art on the OVD benchmarks. Haiwen Hong, Xuan Jin, Yuan He 0011, Hui Xue 0001, Zhou Zhao 0001 |
ICCV | 4 |
| 2023 | Budgeted Training for Vision Transformer
Zhuofan Xia, Xuran Pan, Xuan Jin, Yuan He 0011, Hui Xue 0001, Shiji Song, Gao Huang 0001 |
ICLR | 3 |
| 2022 | MAKD: MULTIPLE Auxiliary Knowledge DistillationabstractKnowledge distillation aims to learn a small student model by leveraging knowledge from a larger teacher model. The gap between these heterogeneous models hinder their knowledge transfer and it would be more challenging when the teacher model is from another task. Previous methods view the teacher model as a perfect feature extractor and train the student model to mimic it. While we notice that the teacher model has defect in extracting features of another task samples. To improve knowledge distillation under such situation, we propose Multiple Auxiliary Subspaces (MAS). Most previous methods improve distillation performance by representation alignment, while we resort to the promotion of the teacher model which is more suitable for cross-task distillation. The MAS distills the knowledge in a mutual learning way based on an auxiliary network. Along with the training procedure, the teacher model is improved by the auxiliary network which works as a trainable part of the teacher model and learn the features distribution of target samples from the student model. And this promotion of the teacher model will benefit the student model via the following distillation procedure. We adopt the representation alignment technique, multiple auxiliary networks to further enhance the proposed method. The MAS works well with limited or sufficient labeled target data. If the source data is available, based on it the MAS can construct another auxiliary network to further improve the distillation. Experiments are conducted to validate that the MAS outperforms baseline methods and achieves state-of-the-art results on several standard benchmarks. Zehan Chen, Xuan Jin, Yuan He 0011, Hui Xue 0001 |
ICASSP | 2 |
| 2022 | Diverse Instance Discovery: Vision-Transformer for Instance-Aware Multi-Label Image RecognitionabstractPrevious works on multi-label image recognition (MLIR) usually use CNNs as a starting point for research. In this paper, we take pure Vision Transformer (ViT) as the research base and make full use of the advantages of Transformer with long-range dependency modeling to circumvent the disadvantages of CNNs limited to local receptive field. However, for multi-label images containing multiple objects from different categories, scales, and spatial relations, it is not optimal to use global information alone. Our goal is to leverage ViT's patch tokens and self-attention mechanism to mine rich instances in multi-label images, named diverse instance discovery (DiD). To this end, we propose a semantic category-aware module and a spatial relationship-aware module, respectively, and then combine the two by a re-constraint strategy to obtain instance-aware attention maps. Finally, we propose a weakly supervised object localization-based approach to extract multi-scale local features, to form a multi-view pipeline. Our method requires only weakly supervised information at the label level, no additional knowledge injection or other strongly supervised information is required. Experiments on three benchmark datasets show that our method significantly outperforms previous works and achieves state-of-the-art results under fair experimental comparisons. Yunqing Hu, Xuan Jin, Yin Zhang 0006, Haiwen Hong, Jingfeng Zhang, Feihu Yan, Yuan He 0011, Hui Xue 0001 |
ICME | 2 |
| 2021 | DRDF: Determining the Importance of Different Multimodal Information with Dual-Router Dynamic FrameworkabstractIn multimodal tasks, the importance of text and image modal information often varies for different input cases. To model the difference of importance of different modal information, we propose a high-performance and highly general Dual-Router Dynamic Framework (DRDF), consisting of Dual-Router, MWF-Layer, experts and expert fusion unit. The text router and image router in Dual-Router take text modal information and image modal information respectively, and MWF-Layer is responsible to determine the importance of modal information. Based on the result of the determination, MWF-Layer generates fused weights for the subsequent experts fusion. Experts can adopt a variety of backbones that match the current multimodal or unimodal task. DRDF features high generality and modularity, and we test 12 backbones such as Visual BERT and their corresponding DRDF instances on the multimodal dataset Hateful memes, and unimodal datasets CIFAR10, CIFAR100, and TinyImagenet. Our DRDF instance outperforms those backbones. We also validate the effectiveness of components of DRDF by ablation studies, and discuss the reasons and ideas of DRDF design. Haiwen Hong, Xuan Jin, Yin Zhang 0006, Yunqing Hu, Jingfeng Zhang, Yuan He 0011, Hui Xue 0001 |
ACM Multimedia | 2 |
| 2021 | RAMS-Trans: Recurrent Attention Multi-scale Transformer for Fine-grained Image RecognitionabstractIn fine-grained image recognition (FGIR), the localization and amplification of region attention is an important factor, which has been explored extensively convolutional neural networks (CNNs) based approaches. The recently developed vision transformer (ViT) has achieved promising results in computer vision tasks. Compared with CNNs, Image sequentialization is a brand new manner. However, ViT is limited in its receptive field size and thus lacks local attention like CNNs due to the fixed size of its patches, and is unable to generate multi-scale features to learn discriminative region attention. To facilitate the learning of discriminative region attention without box/part annotations, we use the strength of the attention weights to measure the importance of the patch tokens corresponding to the raw images. We propose the recurrent attention multi-scale transformer (RAMS-Trans), which uses the transformer's self-attention to recursively learn discriminative region attention in a multi-scale manner. Specifically, at the core of our approach lies the dynamic patch proposal module (DPPM) responsible for guiding region amplification to complete the integration of multi-scale image patches. The DPPM starts with the full-size image patches and iteratively scales up the region attention to generate new patches from global to local by the intensity of the attention weights generated at each scale as an indicator. Our approach requires only the attention weights that come with ViT itself and can be easily trained end-to-end. Extensive experiments demonstrate that RAMS-Trans performs better than exising works, in addition to efficient CNN models, achieving state-of-the-art results on three benchmark datasets. Yunqing Hu, Xuan Jin, Yin Zhang 0006, Haiwen Hong, Jingfeng Zhang, Yuan He 0011, Hui Xue 0001 |
ACM Multimedia | 2 |
| 2020 | The Open Brands Dataset: Unified Brand Detection and Recognition at ScaleabstractIntellectual property protection(IPP) have received more and more attention recently due to the development of the global e-commerce platforms. brand recognition plays a significant role in IPP. Recent studies for brand recognition and detection are based on small-scale datasets that are not comprehensive enough when exploring emerging deep learning techniques. Moreover, it is challenging to evaluate the true performance of brand detection methods in realistic and open scenes. In order to tackle these problems, we first define the special issues of brand detection and recognition compared with generic object detection. Second, a novel brands benchmark called "Open Brands" is established. The dataset contains 1,437,812 images which have brands and 50,000 images without any brand. The part with brands in Open Brands contains 3,113,828 instances annotated in 3 dimensions: 4 types, 559 brands and 1216 logos. To the best of our knowledge, it is the largest dataset for brand detection and recognition with rich annotations. We provide in-depth comprehensive statistics about the dataset, validate the quality of the annotations and study how the performance of many modern models evolves with an increasing amount of training data. Third, we design a network called "Brand Net" to handle brand recognition. Brand Net gets state-of-art mAP on Open Brand compared with existing detection methods. Xuan Jin, Rong Zhang 0006, Yuan He 0011, Hui Xue 0001 |
ICASSP | 1 |
| 2020 | Heuristic Domain AdaptationabstractIn visual domain adaptation (DA), separating the domain-specific characteristics from the domain-invariant representations is an ill-posed problem. Existing methods apply different kinds of priors or directly minimize the domain discrepancy to address this problem, which lack flexibility in handling real-world situations. Another research pipeline expresses the domain-specific information as a gradual transferring process, which tends to be suboptimal in accurately removing the domain-specific properties. In this paper, we address the modeling of domain-invariant and domain-specific information from the heuristic search perspective. We identify the characteristics in the existing representations that lead to larger domain discrepancy as the heuristic representations. With the guidance of heuristic representations, we formulate a principled framework of Heuristic Domain Adaptation (HDA) with well-founded theoretical guarantees. To perform HDA, the cosine similarity scores and independence measurements between domain-invariant and domain-specific representations are cast into the constraints at the initial and final states during the learning procedure. Similar to the final condition of heuristic search, we further derive a constraint enforcing the final range of heuristic network output to be small. Accordingly, we propose Heuristic Domain Adaptation Network (HDAN), which explicitly learns the domain-invariant and domain-specific representations with the above mentioned constraints. Extensive experiments show that HDAN has exceeded state-of-the-art on unsupervised DA, multi-source DA and semi-supervised DA. The code is available at https://github.com/cuishuhao/HDA. Shuhao Cui, Xuan Jin, Shuhui Wang, Yuan He 0011, Qingming Huang |
NeurIPS | 2 |