EDBT 2026 Demo / reviewers in the wild / expert
Suncheng Xiang
dblp:223/1663
· DBLP profile ↗
35ranked-venue papers
13as first author
30since 2021 · last 2026
0000-0002-9141-6460ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 10 first-author · 22 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DiverSeed: Integrating Active Learning for Target Domain Data Generation in Instruction Tuning
Jingsheng Gao, Mengnan Qi, Suncheng Xiang, Ke Ji, Jiacheng Ruan, Ting Liu 0016, Yuzhuo Fu |
Mach. Learn. | 4 |
| 2026 | Ortho-OPD: An Automatic Osteotomy Planes Design Model for Orthognathic Surgery Based on Deep LearningabstractOrthognathic surgery is applied to restore esthetical facial profile and functional occlusion for patients with dentofacial deformity. Virtual surgical planning (VSP) is indispensable for precise and individualized treatment. Manually designing osteotomy planes is time-consuming and highly experience-dependent. This study aimed to develop and validate an automatic osteotomy plane design method based on deep learning. Methods: A deep learning model, Ortho-OPD (orthognathic osteotomy planes designer), was proposed, consisting of a segmentation network and the random sample consensus (RANSAC) algorithm. The segmentation network was based on a convolutional neural network (CNN), orthognathic segmenting the craniomaxillofacial (CMF) CT data. Osteotomy planes were then defined by the RANSAC algorithm. Ortho-OPD was trained on 71 samples and tested on 31 cases. The performance was evaluated quantitatively and qualitatively. Results: Ortho-OPD functioned smoothly and all cases were successfully performed. The 3D boundary-sensitive loss was employed to optimize precision. Evaluation metrics included accuracy and clinical efficiency.The mean dice similarity coefficient (DSC) was 0.92 $\pm$ 0.032 in CMF segmentation. Ortho-OPD showcased excellent productivity, taking an average of about 9 seconds to complete virtual bimaxillary osteotomy compared to manual work. The angular errors between the predicted planes and ground truth planes, plus the shortest distance from the neural tube or the adjacent apical points to predicted planes, were examined, indicating no significant difference and reliability for preserving vital anatomical structures. Overall, the automatic osteotomy plane design from raw CT data was realized using Ortho-OPD, composed of CNN and RANSAC, providing an efficient and ideal alternative in orthognathic osteotomy planning. Xiangmin Li, Mengjia Cheng, Jiahao Bao, Hongjun Qian, Zhihan Wu, Suncheng Xiang, Yaofeng Wen |
IEEE J. Biomed. Health Informatics | 10 |
| 2025 | TTE: Two Tokens Are Enough to Improve Parameter-Efficient TuningabstractExisting fine-tuning paradigms are predominantly characterized by Full Parameter Tuning (FPT) and Parameter-Efficient Tuning (PET). FPT fine-tunes all parameters of a pre-trained model on downstream tasks, whereas PET freezes the pre-trained model and employs only a minimal number of learnable parameters for fine-tuning. However, both approaches face issues of overfitting, especially in scenarios where downstream samples are limited. This issue has been thoroughly explored in FPT, but less so in PET. To this end, this paper investigates overfitting in PET, representing a pioneering study in the field. Specifically, across 19 image classification datasets, we employ three classic PET methods (e.g., VPT, Adapter/Adaptformer, and LoRA) and explore various regularization techniques to mitigate overfitting. Regrettably, the results suggest that existing regularization techniques are incompatible with the PET process and may even lead to performance degradation. Consequently, we introduce a new framework named TTE (Two Tokens are Enough), which effectively alleviates overfitting in PET through a novel constraint function based on the learnable tokens. Experiments conducted on 24 datasets across image and few-shot classification tasks demonstrate that our fine-tuning framework not only mitigates overfitting but also significantly enhances PET's performance. Notably, our TTE framework surpasses the highest-performing FPT framework (DR-Tune), utilizing significantly fewer parameters (0.15M vs. 85.84M) and achieving an improvement of 1%. Jiacheng Ruan, Mingye Xie, Jingsheng Gao, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
AAAI | 5 |
| 2025 | A Training-Free Correlation-Weighted Model for Zero-/Few-Shot Industrial Anomaly Detection with Retrieval AugmentationabstractObtaining labeled data in the field of industrial anomaly detection is challenging, which necessitates the development of label-free frameworks. However, current methods mainly focus on the unsupervised paradigm, which uses a large number of normal samples of the same category to train the model, and distinguish anomalies during testing. This training approach necessitates retraining when new datasets or object categories are encountered. Recently, studies have suggested using large pre-trained multimodal vision-language models, such as CLIP, for zero-shot and few-shot anomaly detection, yielding promising outcomes. However, the lack of spatial awareness of these models results in less effectiveness in dense prediction tasks such as anomaly localization. To mitigate this issue, various fine-tuning methods using additional labeled anomaly data have been employed. In other words, substantial data and extensive training efforts are still necessary to ensure optimal model performance on specific datasets. In this paper, we introduce a training-free, CLIP-based model that utilizes patch correlations and prototype guidance to enable zero-shot and few-shot anomaly detection. Specifically, we first use a self-supervised pre-trained model to capture patch correlations within a single image, enhancing the model's regional awareness of defects. Then, we dynamically construct prototypes using a retrieval-enhanced method to alleviate domain gap in general domain models for anomaly detection. Extensive experiments on the popular benchmarks MVTec and VisA demonstrate that our approach achieves state-of-the-art performance across nearly all metrics. Furthermore, we validate the generalization of our method on collected real industrial data. Wei Ran, Zefang Yu, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 3 |
| 2025 | GPA: Enhancing Generalizable Physical Adversarial Attacks Across Multiple Vision TasksabstractAdversarial attacks pose a significant challenge in deep learning, as carefully crafted perturbations can severely degrade even the most advanced models. In real-world scenarios, where the target models are often unknown, previous works often focus on creating adversarial patterns for specific known models, with the goal of generalizing these patterns to other models. However, such attacks rely heavily on prior model information, leading to poor generalization. To overcome this, we propose a novel method called GPA. Our solution includes an attention extraction module based on a pre-trained vision encoder, which captures precise and generalizable features of model attention on objects. We also introduce attack loss functions that divert attention away from target objects. Compared to state-of-the-art methods, our approach achieves superior attack performance across various downstream vision tasks, including object detection, instance segmentation, and depth estimation. Moreover, the adversarial patterns generated by GPA maintain their effectiveness in real-world scenarios. Mingye Xie, Suncheng Xiang, Jiacheng Ruan, Zefang Yu, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 2 |
| 2025 | Visual Encoders for Generalized Chromosome RecognitionabstractChromosome recognition is a vital task in karyotyping, crucial for birth defect diagnosis and advancing biomedical research. However, developing generalizable classification models faces significant challenges due to the inter-class similarities, intra-class variations, and stark distribution discrepancies across multi-center data. To bridge this gap, we propose a supervised contrastive learning strategy aimed at training robust domain-generalized encoders for accurate chromosome classification. Our model trained using over 3,700,000 chromosome images from multiple centers, excels at extracting fine-grained chromosomal embeddings. These embeddings effectively widen inter-class margins and minimize intra-class variations, thereby enhancing the distinctiveness crucial for precise chromosome type recognition. We comprehensively validate our domain-generalized encoders on two additional large-scale datasets, demonstrating their substantial ability to improve generalization performance. We release our code and pre-trained model weights at https://github.com/RuijiaChang/Chromosome-SCL-Encoder. Ruijia Chang, Suncheng Xiang, Kui Su, Dahong Qian, Jun Wang 0072 |
ICIP | 3 |
| 2025 | Bridging the Gap: Balancing Human Perception and Detector Attention in Adversarial AttacksabstractAdversarial attacks on deep neural networks have garnered significant attention, with recent studies focusing on generating intricate, colorful patterns designed to mislead models. However, such patterns are often easily discernible to human observers. To address this limitation, we propose a novel adversarial camouflage framework BHD that simultaneously mitigates human perceptibility and reduces detector attention. Our approach defines a base pattern by leveraging the surroundings of target object, making it less conspicuous to the human eye. We further design a loss function, integrated with a pre-trained vision encoder with fine-tuned projector, to optimize the adversarial pattern. This allows for effective attacks on detectors while ensuring robust camouflage across diverse environments. In addition, we incorporate human-aligned large multimodal models as objective metrics to quantify the impact on human perception. Extensive experimental results demonstrate that our framework achieves a superior balance between human imperceptibility and model deception, outperforming state-of-the-art methods. Mingye Xie, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICME | 2 |
| 2025 | Learning multi-axis representation in frequency domain for medical image segmentation
Jiacheng Ruan, Jingsheng Gao, Mingye Xie, Suncheng Xiang |
Mach. Learn. | 4 |
| 2025 | Learning Visual-Semantic Embedding for Generalizable Person Re-Identification: A Unified PerspectiveabstractGeneralizable person Re-Identification (Re-ID) is a very hot research topic in machine learning and computer vision, which plays a significant role in realistic scenarios due to its various applications in public security and video surveillance. However, previous methods mainly focus on the visual representation learning, while neglect to explore the potential of semantic features during training, which easily leads to poor generalization capability when adapted to the new domain. In this article, we present a unified perspective called MMET for more robust visual-semantic embedding learning on generalizable Re-ID. To further enhance the robust feature learning in the context of transformer, a dynamic masking mechanism called Masked Multimodal Modeling (MMM) strategy is introduced to mask both the image patches and the text tokens, which can jointly work on multimodal or unimodal data and significantly boost the performance of generalizable person Re-ID. Extensive experiments on benchmark datasets demonstrate the competitive performance of our method over previous approaches. We hope this method could advance the research towards visual-semantic representation learning. Our source code is also publicly available at https://github.com/JeremyXSC/MMET . Suncheng Xiang, Jingsheng Gao, Mingye Xie, Mengyuan Guan, Jiacheng Ruan, Yuzhuo Fu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | LAMM: Label Alignment for Multi-Modal Prompt LearningabstractWith the success of pre-trained visual-language (VL) models such as CLIP in visual representation tasks, transferring pre-trained models to downstream tasks has become a crucial paradigm. Recently, the prompt tuning paradigm, which draws inspiration from natural language processing (NLP), has made significant progress in VL field. However, preceding methods mainly focus on constructing prompt templates for text and visual inputs, neglecting the gap in class label representations between the VL models and downstream tasks. To address this challenge, we introduce an innovative label alignment method named \textbf{LAMM}, which can dynamically adjust the category embeddings of downstream datasets through end-to-end training. Moreover, to achieve a more appropriate label distribution, we propose a hierarchical loss, encompassing the alignment of the parameter space, feature space, and logits space. We conduct experiments on 11 downstream vision datasets and demonstrate that our method significantly improves the performance of existing multi-modal prompt learning models in few-shot scenarios, exhibiting an average accuracy improvement of 2.31(\%) compared to the state-of-the-art methods on 16 shots. Moreover, our methodology exhibits the preeminence in continual learning compared to other prompt tuning methods. Importantly, our method is synergistic with existing prompt tuning methods and can boost the performance on top of them. Our code and dataset will be publicly available at https://github.com/gaojingsheng/LAMM. Jingsheng Gao, Jiacheng Ruan, Suncheng Xiang, Zefang Yu, Ke Ji, Mingye Xie, Ting Liu 0016, Yuzhuo Fu |
AAAI | 3 |
| 2024 | VT-ReID: Learning Discriminative Visual-Text Representation for Polyp Re-IdentificationabstractColonoscopic Polyp Re-Identification (ReID) aims to match a specific polyp in a large gallery with different cameras and views, which plays a key role in the prevention and treatment of colorectal cancer in the computer-aided diagnosis. However, traditional methods mainly focus on the visual representation learning, while neglecting to explore the potential of semantic features during training, which may easily lead to poor generalization capability when adapting the pre-trained model to the new scenarios. To relieve this dilemma, we propose a simple but effective training method named VT-ReID, which can remarkably enrich the representation of polyp videos with the interchange of high-level semantic information. Moreover, a dynamic mechanism named DCM is introduced to leverage contrastive learning to promote better separation between different categories. Empirical results show that our method significantly outperforms current state-of-the art methods with a clear margin. Suncheng Xiang, Cang Liu, Jiacheng Ruan, Shilun Cai, Sijia Du, Dahong Qian |
ICASSP | 1 |
| 2024 | Oceanship: A Large-Scale Dataset for Underwater Audio Target Recognition
Suncheng Xiang, Jingsheng Gao, Jiacheng Ruan, Yanping Hu, Ting Liu 0016, Yuzhuo Fu |
ICIC (4) | 2 |
| 2024 | iDAT: inverse Distillation Adapter-TuningabstractAdapter-Tuning (AT) method involves freezing a pre-trained model and introducing trainable adapter modules to acquire downstream knowledge, thereby calibrating the model for better adaptation to downstream tasks. This paper proposes a distillation framework for the AT method instead of crafting a carefully designed adapter module, which aims to improve fine-tuning performance. For the first time, we explore the possibility of combining the AT method with knowledge distillation. Via statistical analysis, we observe significant differences in the knowledge acquisition between adapter modules of different models. Leveraging these differences, we propose a simple yet effective framework called inverse Distillation Adapter-Tuning (iDAT). Specifically, we designate the smaller model as the teacher and the larger model as the student. The two are jointly trained, and online knowledge distillation is applied to inject knowledge of different perspective to student model, and significantly enhance the fine-tuning performance on downstream tasks. Extensive experiments on the VTAB-1K benchmark with 19 image classification tasks demonstrate the effectiveness of iDAT. The results show that using existing AT method within our iDAT framework can further yield a 2.66% performance gain, with only an additional 0.07M trainable parameters. Our approach compares favorably with state-of-the-arts without bells and whistles. Our code is available at https://github.com/JCruan519/iDAT. Jiacheng Ruan, Jingsheng Gao, Mingye Xie, Daize Dong, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICME | 5 |
| 2024 | GIST: Improving Parameter Efficient Fine-Tuning via Knowledge InteractionabstractRecently, the Parameter Efficient Fine-Tuning (PEFT) method, which adjusts or introduces fewer trainable parameters to calibrate pre-trained models on downstream tasks, has been a hot research topic. However, existing PEFT methods within the traditional fine-tuning framework have two main shortcomings: 1) They overlook the explicit association between trainable parameters and downstream knowledge. 2) They neglect the interaction between the intrinsic task-agnostic knowledge of pre-trained models and the task-specific knowledge of downstream tasks. These oversights lead to insufficient utilization of knowledge and suboptimal performance. To address these issues, we propose a novel fine-tuning framework, named GIST, that can be seamlessly integrated into the current PEFT methods in a plug-and-play manner. Specifically, our framework first introduces a trainable token, called the Gist token, when applying PEFT methods on downstream tasks. This token serves as an aggregator of the task-specific knowledge learned by the PEFT methods and builds an explicit association with downstream tasks. Furthermore, to facilitate explicit interaction between task-agnostic and task-specific knowledge, we introduce the concept of knowledge interaction via a Bidirectional Kullback-Leibler Divergence objective. As a result, PEFT methods within our framework can enable the pre-trained model to understand downstream tasks more comprehensively by fully leveraging both types of knowledge. Extensive experiments on the 35 datasets demonstrate the universality and scalability of our framework. Notably, the PEFT method within our GIST framework achieves up to a 2.25% increase on the VTAB-1K benchmark with an addition of just 0.8K parameters (0.009 of ViT-B/16). The code is available at https://github.com/JCruan519/GIST. Jiacheng Ruan, Jingsheng Gao, Mingye Xie, Suncheng Xiang, Zefang Yu, Ting Liu 0016, Yuzhuo Fu, Xiaoye Qu |
ACM Multimedia | 4 |
| 2024 | Deep multimodal representation learning for generalizable person re-identification
Suncheng Xiang, Wei Ran, Zefang Yu, Ting Liu 0016, Dahong Qian, Yuzhuo Fu |
Mach. Learn. | 1 |
| 2024 | Toward an end-to-end implicit addressee modeling for dialogue disentanglement
Jingsheng Gao, Suncheng Xiang, Zhuowei Wang 0003, Ting Liu 0016, Yuzhuo Fu |
Multim. Tools Appl. | 3 |
| 2024 | SubFace: learning with softmax approximation for face recognition
Suncheng Xiang, Mingye Xie, Dahong Qian |
Multim. Tools Appl. | 1 |
| 2024 | Editing outdoor scenes with a large annotated synthetic dataset
Mingye Xie, Zongwei Liu, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
Multim. Tools Appl. | 3 |
| 2024 | A Simple Normalization Technique Using Window Statistics to Improve the Out-of-Distribution Generalization on Medical ImagesabstractSince data scarcity and data heterogeneity are prevailing for medical images, well-trained Convolutional Neural Networks (CNNs) using previous normalization methods may perform poorly when deployed to a new site. However, a reliable model for real-world clinical applications should generalize well both on in-distribution (IND) and out-of-distribution (OOD) data (e.g., the new site data). In this study, we present a novel normalization technique called window normalization (WIN) to improve the model generalization on heterogeneous medical images, which offers a simple yet effective alternative to existing normalization methods. Specifically, WIN perturbs the normalizing statistics with the local statistics computed within a window. This feature-level augmentation technique regularizes the models well and improves their OOD generalization significantly. Leveraging its advantage, we propose a novel self-distillation method called WIN-WIN. WIN-WIN can be easily implemented with two forward passes and a consistency constraint, serving as a simple extension to existing methods. Extensive experimental results on various tasks (6 tasks) and datasets (24 datasets) demonstrate the generality and effectiveness of our methods. Chengfeng Zhou, Jun Wang 0072, Suncheng Xiang, Hefeng Huang, Dahong Qian |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Rethinking Person Re-Identification via Semantic-based PretrainingabstractPretraining is a dominant paradigm in computer vision. Generally, supervised ImageNet pretraining is commonly used to initialize the backbones of person re-identification (Re-ID) models. However, recent works show a surprising result that CNN-based pretraining on ImageNet has limited impacts on Re-ID system due to the large domain gap between ImageNet and person Re-ID data. To seek an alternative to traditional pretraining, here we investigate semantic-based pretraining as another method to utilize additional textual data against ImageNet pretraining. Specifically, we manually construct a diversified FineGPR-C caption dataset for the first time on person Re-ID events. Based on it, a pure semantic-based pretraining approach named VTBR is proposed to adopt dense captions to learn visual representations with fewer images. We train convolutional neural networks from scratch on the captions of FineGPR-C dataset, and then transfer them to downstream Re-ID tasks. Comprehensive experiments conducted on benchmark datasets show that our VTBR can achieve competitive performance compared with ImageNet pretraining—despite using up to 1.4× fewer images, revealing its potential in Re-ID pretraining. Our source code is also publicly available at https://github.com/JeremyXSC/VTBR . Suncheng Xiang, Dahong Qian, Jingsheng Gao, Ting Liu 0016, Yuzhuo Fu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | MTDL-NET: Morphological and Temporal Discriminative Learning for Heartbeat ClassificationabstractHeartbeat classification based on Electrocardiogram (ECG) signal is crucial to the clinical diagnosis of heart diseases, which has attracted special interest both industrially and scientifically. However, previous methods on ECG mainly lay emphasis on extracting the optimal hand-crafted or deep features, while ignore to explore the potential of morphological and temporal representation to further boost the performance of heartbeat classification task. To address this challenge, in this work, we propose two main modules: (1) Masked attention embedding for extracting discriminative morphological feature; (2) Temporal feature enhanced mechanism for enhancing temporal representation of heartbeat. We combine two modules with transformer encoder architecture and obtain a simple yet effective signal classification model dubbed as MTDL-Net. Comprehensive experiments on benchmark dataset demonstrate that our method can surpass the previous methods by a clear margin quantitatively. Qualitative analysis also validate that MTDL-Net has strong feature extraction capacity and interpretability in the heartbeat classification task. Can Han, Suncheng Xiang, Dahong Qian |
ICASSP | 2 |
| 2023 | AV-TAD: Audio-Visual Temporal Action Detection With TransformerabstractAs an important and challenging task in video understanding, Temporal Action Detection (TAD) has been deeply studied in recent years. However, current works mainly tackle this task with visual information, while neglecting to explore the potential of the audio modality. To address this challenge, in this paper, we propose a simple yet effective AudioVisual Temporal Action Detection Transformer named AV- TAD, which performs early fusion on audio and visual modalities in an end-to-end fashion. On top of it, a novel query formulation is introduced by directly adopting temporal segment coordinates as queries in Transformer decoder, thus allowing us to perform dynamic segment update layer-by-layer. To the best of our knowledge, this is the first attempt to investigate both audio and video feature with a multi-modal Transformer in TAD task. Extensive experiments on THUMOS14 dataset demonstrate that our proposed AV-TAD can outperform the previous methods by a clear margin. Yangcheng Li, Zefang Yu, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 3 |
| 2023 | CC-PoseNet: Towards Human Pose Estimation in Crowded ClassroomsabstractHuman pose estimation has long been motivated for its application in human behavior understanding and activity recognition. Despite recent advances in multi-person pose estimation, existing solutions remain challenging in crowded scenes, especially in classroom scenarios where students are extremely overlapped and have different poses. In this paper, we focus on improving human pose estimation in crowded classrooms from the perspective of crowd detection and pose refinement. Specifically, we first follow a top-down strategy to detect persons in a multi-instance prediction manner and perform single-person pose estimation on each detected human region. Then, the pose estimation is refined with Transformer blocks by capturing the interactions among multiple persons in the image. Importantly, we replace self-attention in Transformer with a lightweight attention mechanism to reduce computational complexity. Quantitative and qualitative experiments demonstrate that our method remarkably outperforms previous methods with a clear margin on both standard benchmarks and self-collected classroom images. Zefang Yu, Yanping Hu, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 3 |
| 2023 | Colo-SCRL: Self-Supervised Contrastive Representation Learning for Colonoscopic Video RetrievalabstractColonoscopic video retrieval, which is a critical part of polyp treatment, has great clinical significance for the prevention and treatment of colorectal cancer. However, retrieval models trained on action recognition datasets usually produce unsatisfactory retrieval results on colonoscopic datasets due to the large domain gap between them. To seek a solution to this problem, we construct a large-scale colonoscopic dataset named Colo-Pair for medical practice. Based on this dataset, a simple yet effective training method called Colo-SCRL is proposed for more robust representation learning. It aims to refine general knowledge from colonoscopies through masked autoencoder-based reconstruction and momentum contrast to improve retrieval performance. To the best of our knowledge, this is the first attempt to employ the contrastive learning paradigm for medical video retrieval. Empirical results show that our method significantly outperforms current state-of-the-art methods in the colonoscopic video retrieval task. Qingzhong Chen, Shilun Cai, Crystal Cai, Zefang Yu, Dahong Qian, Suncheng Xiang |
ICME | 6 |
| 2023 | AutoKary2022: A Large-Scale Densely Annotated Dataset for Chromosome Instance SegmentationabstractAutomated chromosome instance segmentation from metaphase cell microscopic images is critical for the diagnosis of chromosomal disorders (i.e., karyotype analysis). However, it is still a challenging task due to lacking of densely annotated datasets and the complicated morphologies of chromosomes, e.g., dense distribution, arbitrary orientations, and wide range of lengths. To facilitate the development of this area, we take a big step forward and manually construct a large-scale densely annotated dataset named AutoKary2022, which contains over 27,000 chromosome instances in 612 microscopic images from 50 patients. Specifically, each instance is annotated with a polygonal mask and a class label to assist in precise chromosome detection and segmentation. On top of it, we systematically investigate representative methods on this dataset and obtain a number of interesting findings, which helps us have a deeper understanding of the fundamental problems in chromosome instance segmentation. We hope this dataset could advance research towards medical understanding. The dataset can be available at:https://github.com/wangjuncongyu/chromosome-instance-segmentation-dataset. Dan You, Qiuzhu Chen, Minghui Wu 0001, Suncheng Xiang, Jun Wang 0072 |
ICME | 5 |
| 2023 | Learning from self-discrepancy via multiple co-teaching for cross-domain person re-identification
Suncheng Xiang, Yuzhuo Fu, Mengyuan Guan, Ting Liu 0016 |
Mach. Learn. | 1 |
| 2023 | Less Is More: Learning from Synthetic Data with Fine-Grained Attributes for Person Re-IdentificationabstractPerson re-identification (ReID) plays an important role in applications such as public security and video surveillance. Recently, learning from synthetic data [ 9 ], which benefits from the popularity of the synthetic data engine, has attracted great attention from the public. However, existing datasets are limited in quantity, diversity, and realisticity, and cannot be efficiently used for the ReID problem. To address this challenge, we manually construct a large-scale person dataset named FineGPR with fine-grained attribute annotations. Moreover, aiming to fully exploit the potential of FineGPR and promote the efficient training from millions of synthetic data, we propose an attribute analysis pipeline called AOST based on the traditional machine learning algorithm, which dynamically learns attribute distribution in a real domain, then eliminates the gap between synthetic and real-world data and thus is freely deployed to new scenarios. Experiments conducted on benchmarks demonstrate that FineGPR with AOST outperforms (or is on par with) existing real and synthetic datasets, which suggests its feasibility for the ReID task and proves the proverbial less-is-more principle. Our synthetic FineGPR dataset is publicly available at https://github.com/JeremyXSC/FineGPR . Suncheng Xiang, Dahong Qian, Mengyuan Guan, Binjie Yan, Ting Liu 0016, Yuzhuo Fu, Guanjie You |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | MALUNet: A Multi-Attention and Light-weight UNet for Skin Lesion SegmentationabstractRecently, some pioneering works have preferred applying more complex modules to improve segmentation performances. However, it is not friendly for actual clinical environments due to limited computing resources. To address this challenge, we propose a light-weight model to achieve competitive performances for skin lesion segmentation at the lowest cost of parameters and computational complexity so far. Briefly, we propose four modules: (1) DGA consists of dilated convolution and gated attention mechanisms to extract global and local feature information; (2) IEA, which is based on external attention to characterize the overall datasets and enhance the connection between samples; (3) CAB is composed of 1D convolution and fully connected layers to perform a global and local fusion of multi-stage features to generate attention maps at channel axis; (4) SAB, which operates on multi-stage features by a shared 2D convolution to generate attention maps at spatial axis. We combine four modules with our U-shape architecture and obtain a light-weight medical image segmentation model dubbed as MALUNet. Compared with UNet, our model improves the mIoU and DSC metrics by 2.39% and 1.49%, respectively, with a 44x and 166x reduction in the number of parameters and computational complexity. In addition, we conduct comparison experiments on two skin lesion segmentation datasets (ISIC2017 and ISIC2018). Experimental results show that our model achieves state-of-the-art in balancing the number of parameters, computational complexity and segmentation performances. Code is available at https://github.com/JCruan519/MALUNet. Jiacheng Ruan, Suncheng Xiang, Mingye Xie, Ting Liu 0016, Yuzhuo Fu |
BIBM | 2 |
| 2022 | Spatial Attention Guided Local Facial Attribute EditingabstractFacial attribute manipulation has attracted great attention from the public due to its wide range of applications. Aiming to smoothly manipulate the attributes of real facial images, it is critical to search for a proper latent code that aligns with the domain of pre-trained GAN for faithful inversion and controls the transformation within the scope of the attribute for precise editing. Previous methods mainly focused on improving the quality of reconstruction but often ignored the editing effect. To address this issue, we first propose a mapping network to manipulate latent code which is effective for diverse situations, and design a spatial attention network to predict binary mask of the certain attribute which encourages to only alter the relevant region of images and suppress irrelevant changes. In addition, we introduce a novel latent space into the GAN inversion framework which achieves high reconstruction quality especially preserving identity features and retains the ability to edit face attributes. Our methods pave the way to semantically meaningful and disentangled manipulations on both generated images and real images. Ex-perimental results indicate a clear improvement over the cur-rent state-of-the-art methods in various metrics. Mingye Xie, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICME | 2 |
| 2021 | Taking A Closer Look at Synthesis: Fine-Grained Attribute Analysis for Person Re-IdentificationabstractPerson re-identification (re-ID) plays an important role in applications such as public security and video surveillance. Recently, learning from synthetic data, which benefits from the popularity of synthetic data engine, has achieved remarkable performance. However, in pursuit of high accuracy, researchers in the academic always focus on training with large-scale datasets at a high cost of time and label expenses, while neglect to explore the potential of performing efficient training from millions of synthetic data. To facilitate development in this field, we reviewed the previously developed synthetic dataset GPR and built an improved one (GPR+) with larger number of identities and distinguished attributes. Based on it, we quantitatively analyze the influence of dataset attribute on re-ID system. To our best knowledge, we are among the first attempts to explicitly dissect person re-ID from the aspect of attribute on synthetic dataset. This research helps us have a deeper understanding of the fundamental problems in person re-ID, which also provides useful insights for dataset building and future practical usage. Suncheng Xiang, Yuzhuo Fu, Guanjie You, Ting Liu 0016 |
ICASSP | 1 |
| 2020 | Unsupervised Domain Adaptation Through Synthesis For Person Re-IdentificationabstractPerson re-identification is a hot topic because of its widespread applications in video surveillance and public security. However, it remains a challenging task because of drastic variations in illumination or background across surveillance cameras, which causes the current methods can not work well in real-world scenarios. In addition, due to the scarce dataset, many methods suffer from over-fitting to a different extent. To remedy the above two problems, firstly, we develop a data collector and labeler, which can generate the synthetic random scenes and simultaneously annotate them without any manpower. Based on it, we build a large-scale, diverse synthetic dataset. Secondly, we propose a novel unsupervised Re-ID method via domain adaptation, which can exploit the synthetic data to boost the performance of re-identification in a completely unsupervised way, and free humans from heavy data annotations. Extensive experiments show that our proposed method achieves the state-of-the-art performance on two benchmark datasets, and is very competitive with current cross-domain Re-ID method. Suncheng Xiang, Yuzhuo Fu, Guanjie You, Ting Liu 0016 |
ICME | 1 |
| 2020 | Progressive learning with style transfer for distant domain adaptationabstractThis article studies a novel transfer learning problem termed distant domain transfer learning. Different from traditional transfer learning which assumes there is a close relation between source and target data, in this study, the objective is to execute an unseen and unrelated task based on a labelled data set training previously without any samples from intermediate domains. To this end, the authors propose deep unsupervised progressive learning (DUPL) framework and its upgraded version, end‐to‐end DUPL (eDUPL). eDUPL consists of two components, i.e. (i) translating the style of labelled images from irrelevant source domain to the target domain and (ii) learning a domain adaptation model with progressive learning for testing on the target domain. In comparison, eDUPL can integrate the two components of the framework seamlessly. In general, the proposed method is easy to be implemented and can be viewed as a strong convolutional baseline for distant domain adaptation task. Comprehensive experiments based on VeRi Vehicle, CUB‐200‐2011 Birds and Oxford5k Buildings data sets are conducted and the results indicate that the proposed method robustly achieves state‐of‐the‐art performances compared with existing approaches, which demonstrates the effectiveness and superiority of the proposed algorithm. Suncheng Xiang, Yuzhuo Fu, Ting Liu 0016 |
IET Image Process. | 1 |
| 2020 | Multi-level feature learning with attention for person re-identification
Suncheng Xiang, Yuzhuo Fu, Wei Ran, Ting Liu 0016 |
Multim. Tools Appl. | 1 |
| 2020 | Unsupervised person re-identification by hierarchical cluster and domain transfer
Suncheng Xiang, Yuzhuo Fu, Mingye Xie, Zefang Yu, Ting Liu 0016 |
Multim. Tools Appl. | 1 |
| 2019 | Deep Unsupervised Progressive Learning for Distant Domain AdaptationabstractThe superiority of deeply learned representation has been reported in very recent literature of re-identification (Re-ID) task. In this paper, we study a novel transfer learning problem termed Distant Domain Transfer Learning (DDTL) for Re-ID task. Different from existing transfer learning problems which assume that there is a close relation between source domain and target domain, in the DDTL problem, target domain can be totally different from source domain. For example, the source domain classifies pedestrian images but the target domain distinguishes vehicle images. In this work, our goal is to execute an unseen and unrelated task based on a labeled dataset training previously without any samples from intermediate domains. Particularly, we consider the more pragmatic issue of learning a deep feature with no labels, and propose a Deep Unsupervised Progressive Learning (DUPL) method to transfer pretrained deep representations to unseen domains. Specifically, our work performs clustering and fine-tuning of the CNN to improve the performance of original model trained on the irrelevant labeled dataset. Empirical studies on distant domain adaptation task (pedestrian -> vehicle) demonstrate the effectiveness of the proposed method, and the improvement in terms of the mAP accuracy is up to 15% over "non-transfer" methods. Suncheng Xiang, Yuzhuo Fu, Ting Liu 0016 |
ICTAI | 1 |