EDBT 2026 Demo / reviewers in the wild / expert
Xinxing Xu
dblp:15/10654
· DBLP profile ↗
42ranked-venue papers
5as first author
31since 2021 · last 2026
0000-0003-1449-3072ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 18 since 2021Artificial intelligence and machine learning · 20 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 16 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorComputer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Annotation-efficient medical image segmentation via cross-latent graphs and vector-quantized memory
Yanyu Xu 0001, Menghan Zhou, Xinxing Xu, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026, Li-Zhen Cui 0001 |
Medical Image Anal. | 3 |
| 2025 | Ad2Mix: Adversarial and Adaptive Mixup for Unsupervised Domain AdaptationabstractTransformer has recently gained tremendous popularity in unsupervised domain adaptation tasks due to its superior generalization ability. State-of-the-art methods leverage mixup to build an intermediate domain to reduce domain gap. However, such strategy becomes less effective when the domain gap becomes large, as the domain gap between intermediate domain and source domain is not minimized and the constructed intermediate domain is non informative. How to address the adaptation problem when domain gap becomes large is an important research problem in domain adaptation. In this paper, we propose an adversarial and adaptive mixup (Ad2mix) framework which gradually aligns the intermediate domain towards source domain to fully unleash the potential of both the transformer architecture and mixup to address the large domain gap problem. Specifically, we formulate a general framework for intermediate domain learning with mixup. We propose adversarial mixup with a specially designed mixup alike adversarial adaptation operation to reduce the domain gap between the intermediate domain and source domain. To construct an informative intermediate domain, unlike existing methods which utilize a Beta distribution to generate mixup coefficients to interpolate source and target data, we adaptively assign mixup coefficient for each target data instance based on their transferability and discriminativity information. Our framework creates a natural curriculum of intermediate domains from near source domain to near target domain for gradual adaptation. Extensive experimental studies and evaluations on three public domain adaptation benchmark datasets and one medical domain adaptation task demonstrate the superiority of our framework. Lei Zhu 0003, Yanyu Xu 0001, Yong Liu 0026, Rick Siow Mong Goh, Xinxing Xu |
WACV | 5 |
| 2025 | MDDIP: Efficient single-layer pixel-based metasurface denoising via deep image prior
Manna Dai, Feng Yang 0011, Joyjit Chattoraj, Yingzhi Xia, Xinxing Xu, Weijiang Zhao, My Ha Dao, Yong Liu 0026 |
Knowl. Based Syst. | 6 |
| 2025 | Text to Image for Multi-Label Image Recognition With Joint Prompt-Adapter LearningabstractBenefited from image-text contrastive learning, pre-trained vision-language models, e.g., CLIP, allow to direct leverage texts as images (TaI) for parameter-efficient fine-tuning (PEFT). While CLIP is capable of making image features to be similar to the corresponding text features, the modality gap remains a nontrivial issue and limits image recognition performance of TaI. Using multi-label image recognition (MLR) as an example, we present a novel method, called T2I-PAL to tackle the modality gap issue when using only text captions for PEFT. The core design of T2I-PAL is to leverage pre-trained text-to-image generation models to generate photo-realistic and diverse images from text captions, thereby reducing the modality gap. To further enhance MLR, T2I-PAL incorporates a class-wise heatmap and learnable prototypes. This aggregates local similarities, making the representation of local visual features more robust and informative for multi-label recognition. For better PEFT, we further combine both prompt tuning and adapter learning to enhance classification performance. T2I-PAL offers significant advantages: it eliminates the need for fully semantically annotated training images, thereby reducing the manual annotation workload, and it preserves the intrinsic mode of the CLIP model, allowing for seamless integration with any existing CLIP framework. Extensive experiments on multiple benchmarks, including MS-COCO, VOC2007, and NUS-WIDE, show that our T2I-PAL can boost recognition performance by 3.47% in average above the top-ranked state-of-the-art methods. Chun-Mei Feng 0001, Kai Yu 0009, Xinxing Xu, Salman Khan 0001, Rick Siow Mong Goh, Wangmeng Zuo, Yong Liu 0026 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Reliable Federated Disentangling Network for Non-IID Domain FeatureabstractFederated Learning (FL), as an efficient decentralized distributed learning approach, enables multiple institutions to collaboratively train a model without sharing their local data. Despite its advantages, the performance of FL models is substantially impacted by the domain feature shift arising from different acquisition devices/clients. Moreover, existing FL methods often prioritize accuracy without considering reliability factors such as confidence or uncertainty, leading to unreliable predictions in safety-critical applications. Thus, our goal is to enhance FL performance by addressing non-domain feature issues and ensuring model reliability. In this study, we introduce a novel approach named RFedDis (Reliable Federated Disentangling Network). RFedDis leverages feature disentangling to capture a global domain-invariant cross-client representation while preserving local client-specific feature learning. Additionally, we incorporate an uncertainty-aware decision fusion mechanism to effectively integrate the decoupled features. This ensures dynamic integration at the evidence level, producing reliable predictions accompanied by estimated uncertainties. Therefore, RFedDis is the FL approach to combine evidential uncertainty with feature disentangling, enhancing both performance and reliability in handling non-IID domain features. Extensive experimental results demonstrate that RFedDis outperforms other state-of-the-art FL approaches, providing outstanding performance coupled with a high degree of reliability. Meng Wang 0038, Kai Yu 0009, Chun-Mei Feng 0001, Yiming Qian, Ke Zou, Lianyu Wang, Rick Siow Mong Goh, Xinxing Xu, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Big Data | 8 |
| 2024 | RLPeri: Accelerating Visual Perimetry Test with Reinforcement Learning and Convolutional Feature ExtractionabstractVisual perimetry is an important eye examination that helps detect vision problems caused by ocular or neurological conditions. During the test, a patient's gaze is fixed at a specific location while light stimuli of varying intensities are presented in central and peripheral vision. Based on the patient's responses to the stimuli, the visual field mapping and sensitivity are determined. However, maintaining high levels of concentration throughout the test can be challenging for patients, leading to increased examination times and decreased accuracy. In this work, we present RLPeri, a reinforcement learning-based approach to optimize visual perimetry testing. By determining the optimal sequence of locations and initial stimulus values, we aim to reduce the examination time without compromising accuracy. Additionally, we incorporate reward shaping techniques to further improve the testing performance. To monitor the patient's responses over time during testing, we represent the test's state as a pair of 3D matrices. We apply two different convolutional kernels to extract spatial features across locations as well as features across different stimulus values for each location. Through experiments, we demonstrate that our approach results in a 10-20% reduction in examination time while maintaining the accuracy as compared to state-of-the-art methods. With the presented approach, we aim to make visual perimetry testing more efficient and patient-friendly, while still providing accurate results. Tanvi Verma, Linh Le Dinh, Nicholas Tan, Xinxing Xu, Ching Yu Cheng, Yong Liu 0026 |
AAAI | 4 |
| 2024 | Sentence-level Prompts Benefit Composed Image RetrievalabstractComposed image retrieval (CIR) is the task of retrieving specific images by using a query that involves both a reference image and a relative caption. Most existing CIR models adopt the late-fusion strategy to combine visual and language features. Besides, several approaches have also been suggested to generate a pseudo-word token from the reference image, which is further integrated into the relative caption for CIR. However, these pseudo-word-based prompting methods have limitations when target image encompasses complex changes on reference image, e.g., object removal and attribute modification. In this work, we demonstrate that learning an appropriate sentence-level prompt for the relative caption (SPRC) is sufficient for achieving effective composed image retrieval. Instead of relying on pseudo- word-based prompts, we propose to leverage pretrained V-L models, e.g., BLIP-2, to generate sentence-level prompts. By concatenating the learned sentence-level prompt with the relative caption, one can readily use existing text-based image retrieval models to enhance CIR performance. Furthermore, we introduce both image-text contrastive loss and text prompt alignment loss to enforce the learning of suitable sentence-level prompts. Experiments show that our proposed method performs favorably against the state-of-the-art CIR methods on the Fashion-IQ and CIRR datasets. Yang Bai 0011, Xinxing Xu, Yong Liu 0026, Salman Khan 0001, Fahad Shahbaz Khan, Wangmeng Zuo, Rick Siow Mong Goh, Chun-Mei Feng 0001 |
ICLR | 2 |
| 2024 | Multi-Scale Region-Aware Implicit Neural Network for Medical Images Matting
Yanyu Xu 0001, Yingzhi Xia, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026, Xinxing Xu |
MICCAI (9) | 6 |
| 2024 | A New Perspective to Boost Performance Fairness For Medical Federated Learning
Yunlu Yan, Lei Zhu 0003, Yuexiang Li, Xinxing Xu, Rick Siow Mong Goh, Yong Liu 0026, Salman Khan 0001, Chun-Mei Feng 0001 |
MICCAI (10) | 4 |
| 2024 | UrFound: Towards Universal Retinal Foundation Models via Knowledge-Guided Masked Modeling
Kai Yu 0009, Yang Zhou 0017, Yang Bai 0011, Zhi Da Soh, Xinxing Xu, Rick Siow Mong Goh, Ching Yu Cheng, Yong Liu 0026 |
MICCAI (12) | 5 |
| 2024 | BenchX: A Unified Benchmark Framework for Medical Vision-Language Pretraining on Chest X-RaysabstractMedical Vision-Language Pretraining (MedVLP) shows promise in learning generalizable and transferable visual representations from paired and unpaired medical images and reports. MedVLP can provide useful features to downstream tasks and facilitate adapting task-specific models to new setups using fewer examples. However, existing MedVLP methods often differ in terms of datasets, preprocessing, and finetuning implementations. This pose great challenges in evaluating how well a MedVLP method generalizes to various clinically-relevant tasks due to the lack of unified, standardized, and comprehensive benchmark. To fill this gap, we propose BenchX, a unified benchmark framework that enables head-to-head comparison and systematical analysis between MedVLP methods using public chest X-ray datasets. Specifically, BenchX is composed of three components: 1) Comprehensive datasets covering nine datasets and four medical tasks; 2) Benchmark suites to standardize data preprocessing, train-test splits, and parameter selection; 3) Unified finetuning protocols that accommodate heterogeneous MedVLP methods for consistent task adaptation in classification, segmentation, and report generation, respectively. Utilizing BenchX, we establish baselines for nine state-of-the-art MedVLP methods and found that the performance of some early MedVLP methods can be enhanced to surpass more recent ones, prompting a revisiting of the developments and conclusions from prior works in MedVLP. Our code are available at https://github.com/yangzhou12/BenchX. Yang Zhou 0017, Tan Li Hui Faith, Yanyu Xu 0001, Sicong Leng, Xinxing Xu, Yong Liu 0026, Rick Siow Mong Goh |
NeurIPS | 5 |
| 2024 | Clinical domain knowledge-derived template improves post hoc AI explanations in pneumothorax classification
Chuan Hong, Pengtao Jiang, Gangming Zhao, Nguyen Tuan Anh Tran, Xinxing Xu, Yet Yen Yan, Nan Liu 0003 |
J. Biomed. Informatics | 6 |
| 2024 | A surrogate-assisted extended generative adversarial network for parameter optimization in free-form metasurface design
Manna Dai, Feng Yang 0011, Joyjit Chattoraj, Yingzhi Xia, Xinxing Xu, Weijiang Zhao, My Ha Dao, Yong Liu 0026 |
Neural Networks | 6 |
| 2024 | DS-Depth: Dynamic and Static Depth Estimation via a Fusion Cost VolumeabstractSelf-supervised monocular depth estimation methods typically rely on the reprojection error to capture geometric relationships between successive frames in static environments. However, this assumption does not hold in dynamic objects in scenarios, leading to errors during the view synthesis stage, such as feature mismatch and occlusion, which can significantly reduce the accuracy of the generated depth maps. To address this problem, we propose a novel dynamic cost volume that exploits residual optical flow to describe moving objects, improving incorrectly occluded regions in static cost volumes used in previous work. Nevertheless, the dynamic cost volume inevitably generates extra occlusions and noise, thus we alleviate this by designing a fusion module that makes static and dynamic cost volumes compensate for each other. In other words, occlusion from the static volume is refined by the dynamic volume, and incorrect information from the dynamic volume is eliminated by the static volume. Furthermore, we propose a pyramid distillation loss to reduce photometric error inaccuracy at low resolutions and an adaptive photometric error loss to alleviate the flow direction of the large gradient in the occlusion regions. We conducted extensive experiments on the KITTI and Cityscapes datasets, and the results demonstrate that our model outperforms previously published baselines for self-supervised monocular depth estimation. Xingyu Miao, Yang Bai 0011, Haoran Duan 0001, Yawen Huang, Fan Wan, Xinxing Xu, Yang Long 0001, Yefeng Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Geometric Correspondence-Based Multimodal Learning for Ophthalmic Image AnalysisabstractColor fundus photography (CFP) and Optical coherence tomography (OCT) images are two of the most widely used modalities in the clinical diagnosis and management of retinal diseases. Despite the widespread use of multimodal imaging in clinical practice, few methods for automated diagnosis of eye diseases utilize correlated and complementary information from multiple modalities effectively. This paper explores how to leverage the information from CFP and OCT images to improve the automated diagnosis of retinal diseases. We propose a novel multimodal learning method, named geometric correspondence-based multimodal learning network (GeCoM-Net), to achieve the fusion of CFP and OCT images. Specifically, inspired by clinical observations, we consider the geometric correspondence between the OCT slice and the CFP region to learn the correlated features of the two modalities for robust fusion. Furthermore, we design a new feature selection strategy to extract discriminative OCT representations by automatically selecting the important feature maps from OCT slices. Unlike the existing multimodal learning methods, GeCoM-Net is the first method that formulates the geometric relationships between the OCT slice and the corresponding region of the CFP image explicitly for CFP and OCT fusion. Experiments have been conducted on a large-scale private dataset and a publicly available dataset to evaluate the effectiveness of GeCoM-Net for diagnosing diabetic macular edema (DME), impaired visual acuity (VA) and glaucoma. The empirical results show that our method outperforms the current state-of-the-art multimodal learning methods by improving the AUROC score 0.4%, 1.9% and 2.9% for DME, VA and glaucoma detection, respectively. Yan Wang 0015, Liangli Zhen, Tien-En Tan, Huazhu Fu, Yangqin Feng, Zizhou Wang, Xinxing Xu, Rick Siow Mong Goh, Yipin Ng, Claire Calhoun, Gavin Siew Wei Tan, Jennifer K. Sun, Yong Liu 0026, Daniel S. W. Ting |
IEEE Trans. Medical Imaging | 7 |
| 2023 | Learning Federated Visual Prompt in Null Space for MRI ReconstructionabstractFederated Magnetic Resonance Imaging (MRI) reconstruction enables multiple hospitals to collaborate distributedly without aggregating local data, thereby protecting patient privacy. However, the data heterogeneity caused by different MRI protocols, insufficient local training data, and limited communication bandwidth inevitably impair global model convergence and updating. In this paper, we propose a new algorithm, FedPR, to learn federated visual prompts in the null space of global prompt for MRI reconstruction. FedPR is a new federated paradigm that adopts a powerful pre-trained model while only learning and communicating the prompts with few learnable parameters, thereby significantly reducing communication costs and achieving competitive performance on limited local data. Moreover, to deal with catastrophic forgetting caused by data heterogeneity, FedPR also updates efficient federated visual prompts that project the local prompts into an approximate null space of the global prompt, thereby suppressing the interference of gradients on the server performance. Extensive experiments on federated MRI show that FedPR significantly outperforms state-of-the-art FL algorithms with < 6% of communication costs when given the limited amount of local training data. Chun-Mei Feng 0001, Bangjun Li, Xinxing Xu, Yong Liu 0026, Huazhu Fu, Wangmeng Zuo |
CVPR | 3 |
| 2023 | Towards Instance-adaptive Inference for Federated LearningabstractFederated learning (FL) is a distributed learning paradigm that enables multiple clients to learn a powerful global model by aggregating local training. However, the performance of the global model is often hampered by non-i.i.d. distribution among the clients, requiring extensive efforts to mitigate inter-client data heterogeneity. Going beyond inter-client data heterogeneity, we note that intra-client heterogeneity can also be observed on complex real-world data and seriously deteriorate FL performance. In this paper, we present a novel FL algorithm, i.e., FedIns, to handle intra-client data heterogeneity by enabling instance-adaptive inference in the FL framework. Instead of huge instance-adaptive models, we resort to a parameter-efficient fine-tuning method, i.e., scale and shift deep features (SSF), upon a pre-trained model. Specifically, we first train an SSF pool for each client, and aggregate these SSF pools on the server side, thus still maintaining a low communication cost. To enable instance-adaptive inference, for a given instance, we dynamically find the best-matched SSF subsets from the pool and aggregate them to generate an adaptive SSF specified for the instance, thereby reducing the intra-client as well as the inter-client heterogeneity. Extensive experiments show that our FedIns outperforms state-of-the-art FL algorithms, e.g., a 6.64% improvement against the top-performing method with less than 15% communication cost on Tiny-ImageNet. Chun-Mei Feng 0001, Kai Yu 0009, Nian Liu 0002, Xinxing Xu, Salman Khan 0001, Wangmeng Zuo |
ICCV | 4 |
| 2023 | Medical Phrase Grounding with Region-Phrase Context Contrastive Alignment
Zhihao Chen 0004, Yang Zhou 0017, Junting Zhao, Gideon Ooi, Lionel Tim-Ee Cheng, Choon Hua Thng, Xinxing Xu, Yong Liu 0026, Huazhu Fu |
MICCAI (7) | 9 |
| 2023 | Category-Independent Visual Explanation for Medical Deep Network Understanding
Yiming Qian, Liangzhi Li 0004, Huazhu Fu, Meng Wang 0001, Qingsheng Peng, Ching Yu Cheng, Yong Liu 0026, Rick Siow Mong Goh, Xinxing Xu |
MICCAI (2) | 10 |
| 2023 | Federated Uncertainty-Aware Aggregation for Fundus Diabetic Retinopathy Staging
Meng Wang 0001, Lianyu Wang, Xinxing Xu, Ke Zou, Yiming Qian, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu |
MICCAI (2) | 3 |
| 2023 | Minimal-Supervised Medical Image Segmentation via Vector Quantization Memory
Yanyu Xu 0001, Menghan Zhou, Yangqin Feng, Xinxing Xu, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026 |
MICCAI (3) | 4 |
| 2023 | Contrastive domain adaptation with consistency match for automated pneumonia diagnosis
Yangqin Feng, Zizhou Wang, Xinxing Xu, Yan Wang 0015, Huazhu Fu, Shaohua Li 0003, Liangli Zhen, Xiaofeng Lei, Yingnan Cui, Jordan Zheng Ting Sim, Yonghan Ting, Joey Tianyi Zhou, Yong Liu 0026, Rick Siow Mong Goh, Cher Heng Tan |
Medical Image Anal. | 3 |
| 2023 | GAMMA challenge: Glaucoma grAding from Multi-Modality imAges
Huihui Fang, Fei Li 0021, Huazhu Fu, Fengbin Lin, Jiongcheng Li, Yue Huang 0001, Qinji Yu, Sifan Song, Xinxing Xu, Yanyu Xu 0001, Wensai Wang, Shuai Lu 0003, Huiqi Li, Shihua Huang, Zhichao Lu, Chubin Ou, Xifei Wei, Bingyuan Liu, Riadh Kobbi, Xiaoying Tang 0001, Li Lin 0006, Hrvoje Bogunovic, José Ignacio Orlando, Xiulan Zhang, Yanwu Xu 0001 |
Medical Image Anal. | 10 |
| 2022 | CRAFT: Cross-Attentional Flow Transformer for Robust Optical FlowabstractOptical flow estimation aims to find the 2D motion field by identifying corresponding pixels between two images. Despite the tremendous progress of deep learning-based optical flow methods, it remains a challenge to accurately estimate large displacements with motion blur. This is mainly because the correlation volume, the basis of pixel matching, is computed as the dot product of the convolutional features of the two images. The locality of convolutional features makes the computed correlations susceptible to various noises. On large displacements with motion blur, noisy correlations could cause severe errors in the estimated flow. To overcome this challenge, we propose a new architecture “CRoss-Attentional Flow Trans-former” (CRAFT), aiming to revitalize the correlation volume computation. In CRAFT, a Semantic Smoothing Trans-former layer transforms the features of one frame, making them more global and semantically stable. In addition, the dot-product correlations are replaced with trans-former Cross-Frame Attention. This layer filters out feature noises through the Query and Key projections, and computes more accurate correlations. On Sintel (Final) and KITTI (foreground) benchmarks, CRAFT has achieved new state-of-the-art performance. Moreover, to test the robust-ness of different models on large motions, we designed an image shifting attack that shifts input images to generate large artificial motions. Under this attack, CRAFT per-forms much more robustly than two representative meth-ods, RAFT and GMA. The code of CRAFT is is available at https://github.com/askerlee/craft. Xiuchao Sui, Shaohua Li 0003, Xue Geng, Yan Wu 0002, Xinxing Xu, Yong Liu 0026, Rick Siow Mong Goh, Hongyuan Zhu 0002 |
CVPR | 5 |
| 2022 | Airfoil Inverse Design using Conditional Generative Adversarial NetworksabstractCreating an aerodynamic shape, like an airfoil wing, requires many factors to be considered, especially aerodynamic properties such as its lift-to-drag ratio (L/D). Currently, generating feasible airfoil shapes usually requires computationally expensive tools, such as Computational Fluid Dynamics (CFD). In recent years, increasing work has been directed to utilizing machine learning algorithms to synthesize accurate airfoil shapes while reducing the required computational cost. Generative Adversarial Network (GAN) is one of many algorithms to see success in airfoil shape optimization and is shown to generate good airfoils given a small set of training examples. This paper focuses on implementing a conditional GAN (cGAN) based framework with various filters for airfoil inverse design problem. By labelling the training dataset with aerodynamic characteristics separated by pre-defined thresholds to lift-to-drag ratio (L/D) and shape area, the class labels will be able to guide the network to generate different classes of airfoils influenced by these characteristics. Together with layers of Savitzky-Golay (SG) filter and B-Spline Interpolation, the developed model was shown to achieve good performance in generating new airfoils. In addition, we explored the viability of adding Wasserstein loss from Wasserstein GAN into the network architecture, forming a cWGAN-GP. Testing results showed that cWGAN-GP was able to achieve better performance for a specific airfoil class. Xavier Tan, Manna Dai, Joyjit Chattoraj, Yong Liu 0026, Xinxing Xu, My Ha Dao, Feng Yang 0011 |
ICARCV | 5 |
| 2022 | Research on cluster operation optimization of multi-port energy routersabstractAbstract The penetration rate of distributed new energy and DC load increase in the distribution network improve the new energy consumption rate and the reliability of power supply. This paper proposes an AC/DC hybrid system structure with multi‐port energy router (ER) cluster, and compares the advantages of this structure over the AC/DC hybrid microgrid with multiple conversion links. In order to realize the stable, economic, and reliable operation of this structure, the multi‐layer and multi‐time scale control method and communication mode are established. According to the non‐linear characteristics of energy consumption of ER port, the optimization model of the cluster control layer and the control mode of each port in the cluster system are studied. By setting several examples, the rationality of the proposed ER cluster operation control mode and the optimization method and the effectiveness in improving the total energy efficiency of the ER cluster system are verified; the advantages of ER cluster operation in improving the operation energy efficiency of key devices in the distribution network are illustrated. Xinxing Xu, Nengling Tai, Chen Qing |
IET Commun. | 1 |
| 2022 | Deep Supervised Domain Adaptation for Pneumonia Diagnosis From Chest X-Ray ImagesabstractPneumonia is one of the most common treatable causes of death, and early diagnosis allows for early intervention. Automated diagnosis of pneumonia can therefore improve outcomes. However, it is challenging to develop high-performance deep learning models due to the lack of well-annotated data for training. This paper proposes a novel method, called Deep Supervised Domain Adaptation (DSDA), to automatically diagnose pneumonia from chest X-ray images. Specifically, we propose to transfer the knowledge from a publicly available large-scale source dataset (ChestX-ray14) to a well-annotated but small-scale target dataset (the TTSH dataset). DSDA aligns the distributions of the source domain and the target domain according to the underlying semantics of the training samples. It includes two task-specific sub-networks for the source domain and the target domain, respectively. These two sub-networks share the feature extraction layers and are trained in an end-to-end manner. Unlike most existing domain adaptation approaches that perform the same tasks in the source domain and the target domain, we attempt to transfer the knowledge from a multi-label classification task in the source domain to a binary classification task in the target domain. To evaluate the effectiveness of our method, we compare it with several existing peer methods. The experimental results show that our method can achieve promising performance for automated pneumonia diagnosis. Yangqin Feng, Xinxing Xu, Yan Wang 0015, Xiaofeng Lei, Soo Kng Teo, Jordan Zheng Ting Sim, Yonghan Ting, Liangli Zhen, Joey Tianyi Zhou, Yong Liu 0026, Cher Heng Tan |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Crowd Counting With Partial Annotations in an ImageabstractTo fully leverage the data captured from different scenes with different view angles while reducing the annotation cost, this paper studies a novel crowd counting setting, i.e. only using partial annotations in each image as training data. Inspired by the repetitive patterns in the annotated and unannotated regions as well as the ones between them, we design a network with three components to tackle those unannotated regions: i) in an Unannotated Regions Characterization (URC) module, we employ a memory bank to only store the annotated features, which could help the visual features extracted from these annotated regions flow to these unannotated regions; ii) For each image, Feature Distribution Consistency (FDC) regularizes the feature distributions of annotated head and unannotated head regions to be consistent; iii) a Cross-regressor Consistency Regularization (CCR) module is designed to learn the visual features of unannotated regions in a self-supervised style. The experimental results validate the effectiveness of our proposed model under the partial annotation setting for several datasets, such as ShanghaiTech, UCF-CC-50, UCF-QNRF, NWPU-Crowd and JHU-CROWD++. With only 10% annotated regions in each image, our proposed model achieves better performance than the recent methods and baselines under semi-supervised or active learning settings on all datasets. The code is https://github.com/svip-lab/CrwodCountingPAL. Yanyu Xu 0001, Ziming Zhong, Dongze Lian, Jing Li 0117, Xinxing Xu, Shenghua Gao |
ICCV | 6 |
| 2021 | Medical Image Segmentation using Squeeze-and-Expansion TransformersabstractMedical image segmentation is important for computer-aided diagnosis. Good segmentation demands the model to see the big picture and fine details simultaneously, i.e., to learn image features that incorporate large context while keep high spatial resolutions. To approach this goal, the most widely used methods -- U-Net and variants, extract and fuse multi-scale features. However, the fused features still have small "effective receptive fields" with a focus on local image cues, limiting their performance. In this work, we propose Segtran, an alternative segmentation framework based on transformers, which have unlimited "effective receptive fields" even at high feature resolutions. The core of Segtran is a novel Squeeze-and-Expansion transformer: a squeezed attention block regularizes the self attention of transformers, and an expansion block learns diversified representations. Additionally, we propose a new positional encoding scheme for transformers, imposing a continuity inductive bias for images. Experiments were performed on 2D and 3D medical image segmentation tasks: optic disc/cup segmentation in fundus images (REFUGE'20 challenge), polyp segmentation in colonoscopy images, and brain tumor segmentation in MRI scans (BraTS'19 challenge). Compared with representative existing methods, Segtran consistently achieved the highest segmentation accuracy, and exhibited good cross-domain generalization capabilities. Shaohua Li 0003, Xiuchao Sui, Xiangde Luo, Xinxing Xu, Yong Liu 0026, Rick Siow Mong Goh |
IJCAI | 4 |
| 2021 | Few-Shot Domain Adaptation with Polymorphic Transformers
Shaohua Li 0003, Xiuchao Sui, Huazhu Fu, Xiangde Luo, Yangqin Feng, Xinxing Xu, Yong Liu 0026, Daniel S. W. Ting, Rick Siow Mong Goh |
MICCAI (2) | 7 |
| 2021 | Partially-Supervised Learning for Vessel Segmentation in Ocular Images
Yanyu Xu 0001, Xinxing Xu, Shenghua Gao, Rick Siow Mong Goh, Daniel S. W. Ting, Yong Liu 0026 |
MICCAI (1) | 2 |
| 2017 | Laplacian-Steered Neural Style TransferabstractNeural Style Transfer based on Convolutional Neural Networks (CNN) aims to synthesize a new image that retains the high-level structure of a content image, rendered in the low-level texture of a style image. This is achieved by constraining the new image to have high-level CNN features similar to the content image, and lower-level CNN features similar to the style image. However in the traditional optimization objective, low-level features of the content image are absent, and the low-level features of the style image dominate the low-level detail structures of the new image. Hence in the synthesized image, many details of the content image are lost, and a lot of inconsistent and unpleasing artifacts appear. As a remedy, we propose to steer image synthesis with a novel loss function: the Laplacian loss. The Laplacian matrix ("Laplacian" in short), produced by a Laplacian operator, is widely used in computer vision to detect edges and contours. The Laplacian loss measures the difference of the Laplacians, and correspondingly the difference of the detail structures, between the content image and a new image. It is flexible and compatible with the traditional style transfer constraints. By incorporating the Laplacian loss, we obtain a new optimization objective for neural style transfer named Lapstyle. Minimizing this objective will produce a stylized image that better preserves the detail structures of the content image and eliminates the artifacts. Experiments show that Lapstyle produces more appealing stylized images with less artifacts, without compromising their "stylishness". Shaohua Li 0003, Xinxing Xu, Liqiang Nie, Tat-Seng Chua |
ACM Multimedia | 2 |
| 2017 | Action and Event Recognition in Videos by Learning From Heterogeneous Web SourcesabstractIn this paper, we propose new approaches for action and event recognition by leveraging a large number of freely available Web videos (e.g., from Flickr video search engine) and Web images (e.g., from Bing and Google image search engines). We address this problem by formulating it as a new multi-domain adaptation problem, in which heterogeneous Web sources are provided. Specifically, we are given different types of visual features (e.g., the DeCAF features from Bing/Google images and the trajectory-based features from Flickr videos) from heterogeneous source domains and all types of visual features from the target domain. Considering the target domain is more relevant to some source domains, we propose a new approach named multi-domain adaptation with heterogeneous sources (MDA-HS) to effectively make use of the heterogeneous sources. In MDA-HS, we simultaneously seek for the optimal weights of multiple source domains, infer the labels of target domain samples, and learn an optimal target classifier. Moreover, as textual descriptions are often available for both Web videos and images, we propose a novel approach called MDA-HS using privileged information (MDA-HS+) to effectively incorporate the valuable textual information into our MDA-HS method, based on the recent learning using privileged information paradigm. MDA-HS+ can be further extended by using a new elastic-net-like regularization. We solve our MDA-HS and MDA-HS+ methods by using the cutting-plane algorithm, in which a multiple kernel learning problem is derived and solved. Extensive experiments on three benchmark data sets demonstrate that our proposed approaches are effective for action and event recognition without requiring any labeled samples from the target domain. Li Niu 0002, Xinxing Xu, Lin Chen 0021, Lixin Duan, Dong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | Transfer Hashing with Privileged Information
Joey Tianyi Zhou, Xinxing Xu, Sinno Jialin Pan, Ivor W. Tsang, Zheng Qin 0004, Rick Siow Mong Goh |
IJCAI | 2 |
| 2016 | Co-Labeling for Multi-View Weakly Labeled LearningabstractIt is often expensive and time consuming to collect labeled training samples in many real-world applications. To reduce human effort on annotating training samples, many machine learning techniques (e.g., semi-supervised learning (SSL), multi-instance learning (MIL), etc.) have been studied to exploit weakly labeled training samples. Meanwhile, when the training data is represented with multiple types of features, many multi-view learning methods have shown that classifiers trained on different views can help each other to better utilize the unlabeled training samples for the SSL task. In this paper, we study a new learning problem called multi-view weakly labeled learning, in which we aim to develop a unified approach to learn robust classifiers by effectively utilizing different types of weakly labeled multi-view data from a broad range of tasks including SSL, MIL and relative outlier detection (ROD). We propose an effective approach called co-labeling to solve the multi-view weakly labeled learning problem. Specifically, we model the learning problem on each view as a weakly labeled learning problem, which aims to learn an optimal classifier from a set of pseudo-label vectors generated by using the classifiers trained from other views. Unlike traditional co-training approaches using a single pseudo-label vector for training each classifier, our co-labeling approach explores different strategies to utilize the predictions from different views, biases and iterations for generating the pseudo-label vectors, making our approach more robust for real-world applications. Moreover, to further improve the weakly labeled learning on each view, we also exploit the inherent group structure in the pseudo-label vectors generated from different strategies, which leads to a new multi-layer multiple kernel learning problem. Promising results for text-based image retrieval on the NUS-WIDE dataset as well as news classification and text categorization on several real-world multi-view datasets clearly demonstrate that our proposed co-labeling approach achieves state-of-the-art performance for various multi-view weakly labeled learning problems including multi-view SSL, multi-view MIL and multi-view ROD. Xinxing Xu, Wen Li 0001, Dong Xu 0001, Ivor W. Tsang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Image Classification With Densely Sampled Image Windows and Generalized Adaptive Multiple Kernel LearningabstractWe present a framework for image classification that extends beyond the window sampling of fixed spatial pyramids and is supported by a new learning algorithm. Based on the observation that fixed spatial pyramids sample a rather limited subset of the possible image windows, we propose a method that accounts for a comprehensive set of windows densely sampled over location, size, and aspect ratio. A concise high-level image feature is derived to effectively deal with this large set of windows, and this higher level of abstraction offers both efficient handling of the dense samples and reduced sensitivity to misalignment. In addition to dense window sampling, we introduce generalized adaptive l(p)-norm multiple kernel learning (GA-MKL) to learn a robust classifier based on multiple base kernels constructed from the new image features and multiple sets of prelearned classifiers from other classes. With GA-MKL, multiple levels of image features are effectively fused, and information is shared among different classifiers. Extensive evaluation on benchmark datasets for object recognition (Caltech256 and Caltech101) and scene recognition (15Scenes) demonstrate that the proposed method outperforms the state-of-the-art under a broad range of settings. Shengye Yan, Xinxing Xu, Dong Xu 0001, Stephen Lin 0001, Xuelong Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2015 | Distance Metric Learning Using Privileged Information for Face Verification and Person Re-IdentificationabstractIn this paper, we propose a new approach to improve face verification and person re-identification in the RGB images by leveraging a set of RGB-D data, in which we have additional depth images in the training data captured using depth cameras such as Kinect. In particular, we extract visual features and depth features from the RGB images and depth images, respectively. As the depth features are available only in the training data, we treat the depth features as privileged information, and we formulate this task as a distance metric learning with privileged information problem. Unlike the traditional face verification and person re-identification tasks that only use visual features, we further employ the extra depth features in the training data to improve the learning of distance metric in the training process. Based on the information-theoretic metric learning (ITML) method, we propose a new formulation called ITML with privileged information (ITML+) for this task. We also present an efficient algorithm based on the cyclic projection method for solving the proposed ITML+ formulation. Extensive experiments on the challenging faces data sets EUROCOM and CurtinFaces for face verification as well as the BIWI RGBD-ID data set for person re-identification demonstrate the effectiveness of our proposed approach. Xinxing Xu, Wen Li 0001, Dong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | Learning the object location, scale and view for image categorization with adapted classifier
Shengye Yan, Xinxing Xu, Qingshan Liu 0001 |
Inf. Sci. | 2 |
| 2013 | Soft Margin Multiple Kernel LearningabstractMultiple kernel learning (MKL) has been proposed for kernel methods by learning the optimal kernel from a set of predefined base kernels. However, the traditional L1MKL method often achieves worse results than the simplest method using the average of base kernels (i.e., average kernel) in some practical applications. In order to improve the effectiveness of MKL, this paper presents a novel soft margin perspective for MKL. Specifically, we introduce an additional slack variable called kernel slack variable to each quadratic constraint of MKL, which corresponds to one support vector machine model using a single base kernel. We first show that L1MKL can be deemed as hard margin MKL, and then we propose a novel soft margin framework for MKL. Three commonly used loss functions, including the hinge loss, the square hinge loss, and the square loss, can be readily incorporated into this framework, leading to the new soft margin MKL objective functions. Many existing MKL methods can be shown as special cases under our soft margin framework. For example, the hinge loss soft margin MKL leads to a new box constraint for kernel combination coefficients. Using different hyper-parameter values for this formulation, we can inherently bridge the method using average kernel, L1MKL, and the hinge loss soft margin MKL. The square hinge loss soft margin MKL unifies the family of elastic net constraint/regularizer based approaches; and the square loss soft margin MKL incorporates L2MKL naturally. Moreover, we also develop efficient algorithms for solving both the hinge loss and square hinge loss soft margin MKL. Comprehensive experimental studies for various MKL algorithms on several benchmark data sets and two real world applications, including video action recognition and event recognition demonstrate that our proposed algorithms can efficiently achieve an effective yet sparse solution for MKL. Xinxing Xu, Ivor W. Tsang, Dong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2012 | Beyond Spatial Pyramids: A New Feature Extraction Framework with Dense Spatial Sampling for Image Classification
Shengye Yan, Xinxing Xu, Dong Xu 0001, Stephen Lin 0001, Xuelong Li 0001 |
ECCV (4) | 2 |
| 2012 | Handling Ambiguity via Input-Output Kernel LearningabstractData ambiguities exist in many data mining and machine learning applications such as text categorization and image retrieval. For instance, it is generally beneficial to utilize the ambiguous unlabeled documents to learn a more robust classifier for text categorization under the semi-supervised learning setting. To handle general data ambiguities, we present a unified kernel learning framework named Input-Output Kernel Learning (IOKL). Based on our framework, we further propose a novel soft margin group sparse Multiple Kernel Learning (MKL) formulation by introducing a group kernel slack variable to each group of base input-output kernels. Moreover, an efficient block-wise coordinate descent algorithm with an analytical solution for the kernel combination coefficients is developed to solve the proposed formulation. We conduct comprehensive experiments on benchmark datasets for both semi-supervised learning and multiple instance learning tasks, and also apply our IOKL framework to a computer vision application called text-based image retrieval on the NUS-WIDE dataset. Promising results demonstrate the effectiveness of our proposed IOKL framework. Xinxing Xu, Ivor W. Tsang, Dong Xu 0001 |
ICDM | 1 |
| 2012 | Human Gait Recognition Using Patch Distribution Feature and Locality-Constrained Group Sparse RepresentationabstractIn this paper, we propose a new patch distribution feature (PDF) (i.e., referred to as Gabor-PDF) for human gait recognition. We represent each gait energy image (GEI) as a set of local augmented Gabor features, which concatenate the Gabor features extracted from different scales and different orientations together with the X-Y coordinates. We learn a global Gaussian mixture model (GMM) (i.e., referred to as the universal background model) with the local augmented Gabor features from all the gallery GEIs; then, each gallery or probe GEI is further expressed as the normalized parameters of an image-specific GMM adapted from the global GMM. Observing that one video is naturally represented as a group of GEIs, we also propose a new classification method called locality-constrained group sparse representation (LGSR) to classify each probe video by minimizing the weighted l(1, 2) mixed-norm-regularized reconstruction error with respect to the gallery videos. In contrast to the standard group sparse representation method that is a special case of LGSR, the group sparsity and local smooth sparsity constraints are both enforced in LGSR. Our comprehensive experiments on the benchmark USF HumanID database demonstrate the effectiveness of the newly proposed feature Gabor-PDF and the new classification method LGSR for human gait recognition. Moreover, LGSR using the new feature Gabor-PDF achieves the best average Rank-1 and Rank-5 recognition rates on this database among all gait recognition algorithms proposed to date. Dong Xu 0001, Yi Huang 0014, Xinxing Xu |
IEEE Trans. Image Process. | 4 |