VLDB 2026 Research / reviewers in the wild / expert
Xiaofeng Liu 0001
dblp:95/6332-1
· DBLP profile ↗
70ranked-venue papers
40as first author
48since 2021 · last 2026
0000-0002-4514-2016ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 27 first-author · 21 since 2021Artificial intelligence and machine learning · 36 · 21 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 15 first-author · 20 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Power Battery Detection
Xiaoqi Zhao 0003, Peiqian Cao, Chenyang Yu, Zonglei Feng, Lihe Zhang, Hanqi Liu, Jiaming Zuo, Youwei Pang, Jinsong Ouyang, Weisi Lin, Georges El Fakhri, Huchuan Lu, Xiaofeng Liu 0001 |
Int. J. Comput. Vis. | 13 |
| 2026 | A speech-to-video synthesis approach using spatio-temporal diffusion for vocal tract MRI
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Fangxu Xing, Xiaofeng Liu 0001, Maureen Stone 0001, Jiachen Zhuo, Juan Rafael Orozco-Arroyave, Elmar Nöth, Jana Hutter, Jerry L. Prince, Andreas K. Maier, Jonghye Woo |
Medical Image Anal. | 4 |
| 2026 | Causal Graph Learning for Face-Based Interpretable Hierarchical Diagnosis of DepressionabstractDepression has become one of the most serious mental illnesses, leading to a substantial decline in quality of life, an elevated risk of suicide, and significant societal challenges. Despite significant progress in the application of deep learning for depression diagnosis, most prevalent methods rely on correlative rather than causal features, limiting their accuracy and interpretability. Here, we propose a causal graph learning (CGL) method for the hierarchical diagnosis of depression. Specifically, we first construct a novel depression facial graph (DFGraph) structure based on a prior knowledge, which collects information about subjects’ facial cues. Our CGL model leverages the DFGraph structure and incorporates a built-in masking mechanism, which is designed to effectively differentiate causal features from confounding ones. It employs backdoor adjustment techniques, which control for confounding variables by blocking noncausal paths, to identify and select pertinent causal features, thereby enhancing the accuracy of the hierarchical diagnosis of depression. We conducted extensive experiments on the collected depression dataset. Our results show that the proposed method provides better results and interpretability is further improved compared to the publicly available baseline. Baoliang Zhang, Dixin Wang, Mingmei Cheng, Xiaofeng Liu 0001, Yanzhong Wang, Feng Zhu 0004, Zhixiong Lin, Chuan Shi 0006, Wanqing Xie |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2026 | Ultra-High-Definition Image Restoration: New Benchmarks and a Dual Interaction Prior-Driven SolutionabstractUltra-High-Definition (UHD) image restoration has acquired remarkable attention due to its practical demand. In this paper, we construct UHD snow and rain benchmarks, named UHD-Snow and UHD-Rain, to remedy the deficiency in this field. The UHD-Snow/UHD-Rain is established by simulating the physics process of rain/snow into consideration and each benchmark contains 3200 degraded/clear image pairs of 4K resolution. Furthermore, we propose an effective UHD image restoration solution by considering gradient and normal priors in model design, thanks to these priors’ spatial and detail contributions. Specifically, our method contains two branches: (a) feature fusion and reconstruction branch in high-resolution space and (b) prior feature interaction branch in low-resolution space. The former learns high-resolution features and fuses prior-guided low-resolution features to reconstruct clear images, while the latter utilizes normal and gradient priors to mine useful spatial features and detail features to guide high-resolution recovery better. To better utilize these priors, we introduce single prior feature interaction and dual prior feature interaction, where the former respectively fuses normal and gradient priors with high-resolution features to enhance prior ones, while the latter calculates the similarity between enhanced prior ones and further exploits dual guided filtering to boost the feature interaction of dual priors. We conduct experiments on both new and existing public datasets and demonstrate the state-of-the-art performance of our method on UHD image low-light enhancement, dehazing, deblurring, desnowing, and deraining. The source codes and benchmarks are available at https://github.com/wlydlut/UHDDIP. Cong Wang 0018, Jinshan Pan, Xiaofeng Liu 0001, Weixiang Zhou, Xiaoran Sun, Wei Wang 0335, Zhixun Su |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Variance Extrapolated Class-Imbalance-Aware Domain Adaptive Myocardial Segmentation in Multi-Sequence Cardiac MRIabstractFully automated myocardial segmentation from cardiac magnetic resonance imaging (MRI) is vital for efficient diagnosis and treatment planning. Although numerous automated methods have been proposed, they typically focus on single MRI sequences and therefore have difficulties in generalizing across vendors and across cardiac MRI protocols. Simultaneous analysis of complementary cardiac MRI sequences, such as cine, T1 mapping, and late gadolinium enhancement (LGE) MRI, remains challenging due to their distinct image characteristics and scanner-specific variations. To address these issues, we propose an unsupervised domain adaptation approach that allows robust myocardial segmentation across multi-vendor cine, T1, and LGE MRI data. In particular, we introduce a class-imbalance self-training framework to transfer information learned from a source domain with labels to any unlabeled target domain, while maintaining consistent performance across different MRI sequences. Our framework iteratively refines segmentation accuracy by generating pseudo-labels for target data using a hardness-aware strategy, thus effectively addressing the problem of class imbalance in cardiac MRI segmentation. To mitigate data scarcity following pseudo-label selection, we employ a variance-guided vicinal feature extrapolation, which expands data points in the feature space into a probabilistic distribution. This, in turn, facilitates joint source-target training by generating a larger intersection in the feature space. Experimental results demonstrate that our framework outperforms existing methods when assessed using the Dice coefficient and Hausdorff distance. Our framework enables cardiac evaluation across MRI protocols without sequence-specific manual annotations. Fangxu Xing, Xiaofeng Liu 0001, Iman Aganj, Georges El Fakhri, Panki Kim, Jonghye Woo |
IEEE J. Biomed. Health Informatics | 2 |
| 2026 | Three-Dimensional MRI Reconstruction With 3D Gaussian Representations: Tackling the Undersampling ProblemabstractThree-Dimensional Gaussian representation (3DGS) has shown substantial promise in the field of computer vision, but remains unexplored in the field of magnetic resonance imaging (MRI). This study explores its potential for the reconstruction of isotropic resolution 3D MRI from undersampled k-space data. We introduce a novel framework termed 3D Gaussian MRI (3DGSMR), which employs 3D Gaussian distributions as an explicit representation for MR volumes. Experimental evaluations indicate that this method can effectively reconstruct voxelized MR images, achieving a quality on par with that of well-established 3D MRI reconstruction techniques found in the literature. Notably, the 3DGSMR scheme operates under a self-supervised framework, obviating the need for extensive training datasets or prior model training. This approach introduces significant innovations to the domain, notably the adaptation of 3DGS to MRI reconstruction and the novel application of the existing 3DGS methodology to decompose MR signals, which are presented in a complex-valued format. Tengya Peng, Ruyi Zha, Xiaofeng Liu 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | GlioSurvNet: Multimodal Survival Prediction for Glioblastoma Using Deep Learning and Clinical Variables from Brain MRIabstractAccurate survival prediction using multimodal magnetic resonance imaging (MRI) plays a crucial role in clinical decision-making for patients with glioblastoma (GBM). In this work, we propose a multimodal framework, GlioSurvNet, that integrates deep learning features extracted from Swin UNETR and clinical variables to predict patient survival. Our framework makes use of multiple MRI sequences, including T1, T1 with contrast enhancement, T2-weighted, and FLAIR MRI, to capture diverse tumor characteristics. The Swin UNETR architecture simultaneously carries out tumor segmentation and extracts hierarchical features from multimodal MRI data. These deep learning features are then combined with clinical variables, which are input into a multi-layer perceptron network to yield survival probabilities. We evaluated our framework on a cohort of 287 patients from two independent databases, UPENN-GBM and UCSF-PDGM, demonstrating superior survival prediction performance when compared with existing methods. Our framework achieved a time-dependent concordance index of 0.693 and an integrated brier score of 0.14 with improved risk stratification. GlioSurvNet offers a robust tool for personalized prognosis and treatment planning in GBM patients. Gihyeon Kim, Fangxu Xing, Hyoun-Joong Kong, Emiliano Santarnecchi, Helen A. Shih, Thomas Bortfeld, Georges El Fakhri, Xiaofeng Liu 0001, Jang Hwan Choi 0001, Jonghye Woo |
ICIP | 8 |
| 2025 | Towards Robust Deterministic and Probabilistic Modeling for Predictive LearningabstractPredictive modeling of unannotated spatiotemporal data presents inherent challenges, primarily due to the highly entangled visual dynamics in real-world scenes. To tackle these complexities, we introduce a novel insight through Disentangling Deterministic and Probabilistic (DDP) modeling. We note a key observation in spatiotemporal data where low-level details typically remain stable, whereas high-level motion frequently exhibits dynamic variations. The core motivation involves constructing two distinct pathways in the latent space: a deterministic path and a probabilistic path. The probabilistic path begins by defining the motion flow, which explicitly describes complex many-to-many motion patterns between patches, and models its probabilistic distribution using a motion diffuser. The deterministic path incorporates a spectral-aware enhancer to retain and amplify visual details in the frequency domain. These designs ensure visual consistency while also capturing intricate long-term motion dynamics. Extensive experiments demonstrate the superiority of DDP across diverse scenario evaluations. Xuesong Nie, Haoyuan Jin, B. V. K. Vijaya Kumar, Xiaofeng Liu 0001 |
IJCAI | 4 |
| 2025 | TNT-GS: Truncated and Tailored Gaussian SplattingabstractGaussian Splatting (GS) is widely used for efficient 3D scene representation and rendering by modeling scenes as continuous Gaussian distributions. However, GS struggles with high-frequency details and sharp transitions due to its low-pass filtering effect, often requiring multiple Gaussian stacking, which increases computational and memory costs. To overcome these limitations, we propose Truncated and Tailored Gaussian Splatting (TNT-GS), a novel approach that enhances shape complexity and preserves sharp boundaries. Our method truncates Gaussians to generate sharp edges and flexible shapes without excessive stacking, improving efficiency. We also introduce learnable parameters to dynamically tailor the receptive field of the primitives, optimizing the balance between high-frequency details and smooth regions. Furthermore, we employ specialized densification strategies to further improve efficiency during tile computation. Experimental results show that TNT-GS outperforms state-of-the-art methods in storage efficiency and rendering speed, offering a robust solution for real-time rendering. The code of TNT-GS is available at https://github.com/GoogolplexGoodenough/TNT-GS. Xiaofeng Liu 0001, Guanchen Meng, Chongyang Feng, Risheng Liu, Zhongxuan Luo, Xin Fan 0001 |
ACM Multimedia | 1 |
| 2025 | Rethinking Evaluation of Infrared Small Target DetectionabstractAs an essential vision task, infrared small target detection (IRSTD) has seen significant advancements through deep learning. However, critical limitations in current evaluation protocols impede further progress. First, existing methods rely on fragmented pixel- and target-level specific metrics, which fails to provide a comprehensive view of model capabilities. Second, an excessive emphasis on overall performance scores obscures crucial error analysis, which is vital for identifying failure modes and improving real-world system performance. Third, the field predominantly adopts dataset-specific training-testing paradigms, hindering the understanding of model robustness and generalization across diverse infrared scenarios. This paper addresses these issues by introducing a hybrid-level metric incorporating pixel- and target-level performance, proposing a systematic error analysis method, and emphasizing the importance of cross-dataset evaluation. These aim to offer a more thorough and rational hierarchical analysis framework, ultimately fostering the development of more effective and robust IRSTD models. An open-source toolkit has be released to facilitate standardized benchmarking. Youwei Pang, Xiaoqi Zhao 0003, Lihe Zhang, Huchuan Lu, Georges El Fakhri, Xiaofeng Liu 0001, Shijian Lu |
NeurIPS | 6 |
| 2025 | UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised CompensationabstractMulti-modal image segmentation faces real-world deployment challenges from incomplete/corrupted modalities degrading performance. While existing methods address training-inference modality gaps via specialized per-combination models, they introduce high deployment costs by requiring exhaustive model subsets and model-modality matching. In this work, we propose a unified modality-relax segmentation network (UniMRSeg) through hierarchical self-supervised compensation (HSSC). Our approach hierarchically bridges representation gaps between complete and incomplete modalities across input, feature and output levels.
First, we adopt modality reconstruction with the hybrid shuffled-masking augmentation, encouraging the model to learn the intrinsic modality characteristics and generate meaningful representations for missing modalities through cross-modal fusion.
Next, modality-invariant contrastive learning implicitly compensates the feature space distance among incomplete-complete modality pairs. Furthermore, the proposed lightweight reverse attention adapter explicitly compensates for the weak perceptual semantics in the frozen encoder. Last, UniMRSeg is fine-tuned under the hybrid consistency constraint to ensure stable prediction under all modality combinations without large performance fluctuations. Without bells and whistles, UniMRSeg significantly outperforms the state-of-the-art methods under diverse missing modality scenarios on MRI-based brain tumor segmentation, RGB-D semantic segmentation, RGB-D/T salient object segmentation. The code will be released at \url{https://github.com/Xiaoqi-Zhao-DLUT/UniMRSeg}. Xiaoqi Zhao 0003, Youwei Pang, Chenyang Yu, Lihe Zhang, Huchuan Lu, Shijian Lu, Georges El Fakhri, Xiaofeng Liu 0001 |
NeurIPS | 8 |
| 2025 | Ordinal Unsupervised Domain Adaptation With Recursively Conditional Gaussian Imposed Variational DisentanglementabstractThere has been a growing interest in unsupervised domain adaptation (UDA) to alleviate the data scalability issue, while the existing works usually focus on classifying independently discrete labels. However, in many tasks (e.g., medical diagnosis), the labels are discrete and successively distributed. The UDA for ordinal classification requires inducing non-trivial ordinal distribution prior to the latent space. Target for this, the partially ordered set (poset) is defined for constraining the latent vector. Instead of the typically i.i.d. Gaussian latent prior, in this work, a recursively conditional Gaussian (RCG) set is proposed for ordered constraint modeling, which admits a tractable joint distribution prior. Furthermore, we are able to control the density of content vectors that violate the poset constraint by a simple "three-sigma rule." We explicitly disentangle the cross-domain images into a shared ordinal prior induced ordinal content space and two separate source/target ordinal-unrelated spaces, and the self-training is worked on the shared space exclusively for ordinal-aware domain alignment. Extensive experiments on UDA medical diagnoses and facial age estimation demonstrate its effectiveness. Xiaofeng Liu 0001, Site Li, Yubin Ge, Pengyi Ye, Jane You, Jun Lu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Texture and noise dual adaptation for infrared image super-resolutionabstractRecent efforts have explored leveraging visible light images to enrich texture details in infrared (IR) super-resolution. However, this direct adaptation approach often becomes a double-edged sword, as it improves texture at the cost of introducing noise and blurring artifacts. Such imperfections are inherent in the spatial domain of visible images and are accentuated during the imaging process. Enhancing IR image quality by integrating rich texture details from visible images, while minimizing noise transfer, presents a challenging research avenue. To address these challenges, we propose the Texture and Noise Dual Adaptation SRGAN (DASRGAN), an innovative framework specifically engineered for robust IR super-resolution model adaptation. DASRGAN operates on the synergy of two key components: (1) Texture-Oriented Adaptation (TOA) to refine texture details meticulously, and (2) Noise-Oriented Adaptation (NOA), dedicated to minimizing noise transfer. Specifically, TOA uniquely integrates a specialized discriminator, incorporating a prior extraction branch, and employs a Sobel-guided adversarial loss to align texture distributions effectively. Concurrently, NOA utilizes a noise adversarial loss to distinctly separate the generative and Gaussian noise pattern distributions during adversarial training. Our extensive experiments confirm DASRGAN’s superiority. Comparative analyses against leading methods across multiple benchmarks and upsampling factors reveal that DASRGAN sets new state-of-the-art performance standards. Code are available at https://github.com/yongsongH/DASRGAN . • DASRGAN improves infrared image super-resolution using visible textures and noise control through dual adaptations. • Texture-Oriented Adaptation uses Sobel-based discriminator and texture alignment loss for detail enhancement. • Noise-Oriented Adaptation employs domain-specific loss to suppress noise propagation between modalities. • DASRGAN achieves state-of-the-art performance in infrared super-resolution benchmarks across metrics. Yongsong Huang, Tomo Miyazaki, Xiaofeng Liu 0001, Yafei Dong, Shinichiro Omachi |
Pattern Recognit. | 3 |
| 2025 | Vision-language foundation model for generalizable nasal disease diagnosis using unlabeled endoscopic recordsabstractMedical artificial intelligence (AI) holds significant potential in identifying signs of health conditions in nasal endoscopic images, thereby accelerating the diagnosis of diseases and systemic disorders. However, the performance of AI models heavily relies on expert annotations, and these models are usually task-specific with limited generalization performance across various clinical applications. In this paper, we introduce NasVLM, a Nasal Vision-Language foundation Model designed to extract universal representations from unlabeled nasal endoscopic data. Additionally, we construct a large-scale nasal endoscopic pre-training dataset and three downstream validation datasets from routine diagnostic records. The core strength of NasVLM lies in its ability to learn cross-modal semantic representations and perform multi-granular report-image alignment without depending on expert annotations. Furthermore, to the best of our knowledge, it is the first medical foundation model that effectively aligns medical report with multiple images of different anatomic regions, facilitated by a well-designed hierarchical report-supervised learning framework. The experimental results demonstrate that NasVLM has superior generalization performance across diverse diagnostic tasks and surpasses state-of-the-art self- and report-supervised methods in disease classification and lesion localization, especially in scenarios requiring label-efficient fine-tuning. For instance, NasVLM can distinguish normal nasopharynx (NOR) from abnormalities (benign hyperplasia, BH, and nasopharyngeal carcinoma, NPC) with an accuracy of 91.38% (95% CI, 90.59 to 92.17) and differentiate NPC from BH and NOR with an accuracy of 81.45% (95% CI, 80.21 to 82.67) on the multi-center NPC-Screen dataset using only 1% labeled data, on par with the performance of traditional supervised methods using 100% labeled data. Wentao Gong, Yinlong Liu, Xicai Sun, Xiaofeng Liu 0001, Xinrong Chen, Hongmeng Yu |
Pattern Recognit. | 9 |
| 2025 | IRSRMamba: Infrared Image Super-Resolution via Mamba-Based Wavelet Transform Feature Modulation ModelabstractInfrared image super-resolution (IRSR) is challenging due to weak structures and textures. While Mamba-based state-space models (SSMs) efficiently model long-range dependencies, their inherent block-wise processing disrupts spatial consistency, limiting direct IRSR applicability. We propose IRSRMamba, a novel Mamba-based framework overcoming this limitation via tailored structural and textural preservation. Integrated into a Mamba backbone, our key innovations are: 1) Wavelet Transform Feature Modulation (WTFM), enhancing multi-scale frequency-aware feature extraction to mitigate block-induced coherence loss; and 2) an SSMs-based Semantic Consistency Loss, enforcing cross-block alignment to restore fragmented context. IRSRMamba achieves superior global-local fusion, structural coherence, and fine-detail preservation. Experiments show state-of-the-art PSNR, SSIM, and perceptual quality on IR benchmarks, as well as robust generalization to remote sensing. This work establishes Mamba-based architectures as highly promising for high-fidelity IR image enhancement. Code is available at https://github.com/yongsongH/IRSRMamba. Yongsong Huang, Tomo Miyazaki, Xiaofeng Liu 0001, Shinichiro Omachi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Label Space-Induced Pseudo Label Refinement for Multi-Source Black-Box Domain AdaptationabstractConventional unsupervised domain adaptation (UDA) requires access to source data and/or source model parameters, prohibiting its practical application in terms of privacy, security, and intellectual property. Recent black-box UDA (BDA) reduces such constraints by defining a pseudo label from a single encapsulated source application programming interface (API) prediction, which allows for self-training of the target model. Nonetheless, existing methods have limited consideration for multi-source settings, in which multiple source domain APIs are available to generate pseudo labels. In this work, we introduce a novel training framework for multi-source BDA (MSBDA), dubbed Label Space-Induced Pseudo Label Refinement (LPR). Specifically, LPR incorporates a Pseudo label Refinery Network (PRN) that learns the relationship among source domains conditioned by the target domain only utilizing source API's prediction. The target model is adapted by our dual phases PRN. First, a warm-up phase targets to avoid failure due to noisy samples and provide an initial pseudo-label, which is followed by a label refinement phase with domain relationship exploration. We provide theoretical support for the mechanism of the LPR. Experimental results on four benchmark datasets demonstrate that MSBDA using LPR achieves competitive performance compared to state-of-the-art approaches with different DA settings. Chae Hwa Yoo, Xiaofeng Liu 0001, Fangxu Xing, Jonghye Woo, Je-Won Kang |
IEEE Trans. Image Process. | 2 |
| 2025 | Bayesian Posterior Distribution Estimation of Kinetic Parameters in Dynamic Brain PET Using Generative Deep Learning ModelsabstractPositron Emission Tomography (PET) is a valuable imaging method for studying molecular-level processes in the body, such as hyperphosphorylated tau (p-tau) protein aggregates, a hallmark of several neurodegenerative diseases including Alzheimer's disease. P-tau density and cerebral perfusion can be quantified from dynamic PET images using tracer kinetic modeling techniques. However, noise in PET images leads to uncertainty in the estimated kinetic parameters, which can be quantified by estimating the posterior distribution of kinetic parameters using Bayesian inference (BI). Markov Chain Monte Carlo (MCMC) techniques are commonly used for posterior estimation but with significant computational needs. This work proposes an Improved Denoising Diffusion Probabilistic Model (iDDPM)-based method to estimate the posterior distribution of kinetic parameters in dynamic PET, leveraging the high computational efficiency of deep learning. The performance of the proposed method was evaluated on a [18F]MK6240 study and compared to a Conditional Variational Autoencoder with dual decoder (CVAE-DD)-based method and a Wasserstein GAN with gradient penalty (WGAN-GP)-based method. Posterior distributions inferred from Metropolis-Hasting MCMC were used as reference. Our approach consistently outperformed the CVAE-DD and WGAN-GP methods and offered significant reduction in computation time than the MCMC method (over 230 times faster), inferring accurate ( $\lt {0}.{67}\,\%$ mean error) and precise ( $\lt {7}.{23}\,\%$ standard deviation error) posterior distributions. Yanis Djebra, Xiaofeng Liu 0001, Thibault Marin, Amal Tiss, Maëva Dhaynaut, Nicolas J. Guehl, Keith A. Johnson, Georges El Fakhri, Chao Ma 0018, Jinsong Ouyang |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Contrastive Learning Approach for Assessment of Phonological Precision in Patients with Tongue Cancer Using MRI DataabstractMagnetic Resonance Imaging (MRI) allows analyzing speech production by capturing high-resolution images of the dynamic processes in the vocal tract. In clinical applications, combining MRI with synchronized speech recordings leads to improved patient outcomes, especially if a phonological-based approach is used for assessment. However, when audio signals are unavailable, the recognition accuracy of sounds is decreased when using only MRI data. We propose a contrastive learning approach to improve the detection of phonological classes from MRI data when acoustic signals are not available at inference time. We demonstrate that frame-wise recognition of phonological classes improves from an f1 of 0.74 to 0.85 when the contrastive loss approach is implemented. Furthermore, we show the utility of our approach in the clinical application of using such phonological classes to assess speech disorders in patients with tongue cancer, yielding promising results in the recognition task. Tomás Arias-Vergara, Paula Andrea Pérez-Toro, Xiaofeng Liu 0001, Fangxu Xing, Maureen Stone 0001, Jiachen Zhuo, Jerry L. Prince, Maria Schuster, Elmar Nöth, Jonghye Woo, Andreas K. Maier |
INTERSPEECH | 3 |
| 2024 | Tagged-to-Cine MRI Sequence Synthesis via Light Spatial-Temporal Transformer
Xiaofeng Liu 0001, Fangxu Xing, Zhangxing Bian, Tomás Arias-Vergara, Paula Andrea Pérez-Toro, Andreas K. Maier, Maureen Stone 0001, Jiachen Zhuo, Jerry L. Prince, Jonghye Woo |
MICCAI (7) | 1 |
| 2024 | Nuanced Multi-class Detection of Machine-Generated Scientific Text
Yubin Ge, Xiaofeng Liu 0001 |
PACLIC | 3 |
| 2024 | Subtype-Aware Dynamic Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) has been successfully applied to transfer knowledge from a labeled source domain to target domains without their labels. Recently introduced transferable prototypical networks (TPNs) further address class-wise conditional alignment. In TPN, while the closeness of class centers between source and target domains is explicitly enforced in a latent space, the underlying fine-grained subtype structure and the cross-domain within-class compactness have not been fully investigated. To counter this, we propose a new approach to adaptively perform a fine-grained subtype-aware alignment to improve the performance in the target domain without the subtype label in both domains. The insight of our approach is that the unlabeled subtypes in a class have the local proximity within a subtype while exhibiting disparate characteristics because of different conditional and label shifts. Specifically, we propose to simultaneously enforce subtype-wise compactness and class-wise separation, by utilizing intermediate pseudo-labels. In addition, we systematically investigate various scenarios with and without prior knowledge of subtype numbers and propose to exploit the underlying subtype structure. Furthermore, a dynamic queue framework is developed to evolve the subtype cluster centroids steadily using an alternative processing scheme. Experimental results, carried out with multiview congenital heart disease data and VisDA and DomainNet, show the effectiveness and validity of our subtype-aware UDA, compared with state-of-the-art UDA methods. Xiaofeng Liu 0001, Fangxu Xing, Jane You, Jun Lu 0002, C.-C. Jay Kuo, Georges El Fakhri, Jonghye Woo |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Motion-Scenario Decoupling for Rat-Aware Video Position Prediction: Strategy and Benchmark
Xiaofeng Liu 0001, Jiaxin Gao 0001, Nenggan Zheng, Risheng Liu |
ICIG (2) | 1 |
| 2023 | Incremental Learning for Heterogeneous Structure Segmentation in Brain Tumor MRI
Xiaofeng Liu 0001, Helen A. Shih, Fangxu Xing, Emiliano Santarnecchi, Georges El Fakhri, Jonghye Woo |
MICCAI (2) | 1 |
| 2023 | Speech Audio Synthesis from Tagged MRI and Non-negative Matrix Factorization via Plastic Transformer
Xiaofeng Liu 0001, Fangxu Xing, Maureen Stone 0001, Jiachen Zhuo, Sidney S. Fels, Jerry L. Prince, Georges El Fakhri, Jonghye Woo |
MICCAI (7) | 1 |
| 2023 | Source-free domain adaptive segmentation with class-balanced complementary self-training
Yongsong Huang, Wanqing Xie, Ethan Xiao, Jane You, Xiaofeng Liu 0001 |
Artif. Intell. Medicine | 6 |
| 2023 | Inducing semantic hierarchy structure in empirical risk minimization with optimal transport measures
Wanqing Xie, Yubin Ge, Site Li, Zhenhua Guo 0001, Xiaofeng Liu 0001 |
Neurocomputing | 6 |
| 2023 | Attentive continuous generative self-training for unsupervised domain adaptive medical image translation
Xiaofeng Liu 0001, Jerry L. Prince, Fangxu Xing, Jiachen Zhuo, Timothy G. Reese, Maureen Stone 0001, Georges El Fakhri, Jonghye Woo |
Medical Image Anal. | 1 |
| 2023 | Memory consistent unsupervised off-the-shelf model adaptation for source-relaxed medical image segmentation
Xiaofeng Liu 0001, Fangxu Xing, Georges El Fakhri, Jonghye Woo |
Medical Image Anal. | 1 |
| 2022 | Cmri2spec: Cine MRI Sequence to Spectrogram Synthesis via A Pairwise Heterogeneous TranslatorabstractMultimodal representation learning using visual movements from cine magnetic resonance imaging (MRI) and their acoustics has shown great potential to learn shared representation and to predict one modality from another. Here, we propose a new synthesis framework to translate from cine MRI sequences to spectrograms with a limited dataset size. Our framework hinges on a novel fully convolutional heterogeneous translator, with a 3D CNN encoder for efficient sequence encoding and a 2D transpose convolution decoder. In addition, a pairwise correlation of the samples with the same speech word is utilized with a latent space representation disentanglement scheme. Furthermore, an adversarial training approach with generative adversarial networks is incorporated to provide enhanced realism on our generated spectrograms. Our experimental results, carried out with a total of 63 cine MRI sequences alongside speech acoustics, show that our framework improves synthesis accuracy, compared with competing methods. Our framework thereby has shown the potential to aid in better understanding the relationship between the two modalities. Xiaofeng Liu 0001, Fangxu Xing, Maureen Stone 0001, Jerry L. Prince, Jangwon Kim, Georges El Fakhri, Jonghye Woo |
ICASSP | 1 |
| 2022 | Tagged-MRI Sequence to Audio Synthesis via Self Residual Attention Guided Heterogeneous Translator
Xiaofeng Liu 0001, Fangxu Xing, Jerry L. Prince, Jiachen Zhuo, Maureen Stone 0001, Georges El Fakhri, Jonghye Woo |
MICCAI (6) | 1 |
| 2022 | ACT: Semi-supervised Domain-Adaptive Medical Image Segmentation with Asymmetric Co-training
Xiaofeng Liu 0001, Fangxu Xing, Nadya Shusharina, Ruth Lim, C.-C. Jay Kuo, Georges El Fakhri, Jonghye Woo |
MICCAI (5) | 1 |
| 2022 | Constraining pseudo-label in self-training unsupervised domain adaptation with energy-based modelabstractDeep learning is usually data starved, and the unsupervised domain adaptation (UDA) is developed to introduce the knowledge in the labeled source domain to the unlabeled target domain. Recently, deep self-training presents a powerful means for UDA, involving an iterative process of predicting the target domain and then taking the confident predictions as hard pseudo-labels for retraining. However, the pseudo-labels are usually unreliable, thus easily leading to deviated solutions with propagated errors. In this paper, we resort to the energy-based model and constrain the training of the unlabeled target sample with an energy function minimization objective. It can be achieved via a simple additional regularization or an energy-based loss. This framework allows us to gain the benefits of the energy-based model, while retaining strong discriminative performance following a plug-and-play fashion. The convergence property and its connection with classification expectation minimization are investigated. We deliver extensive experiments on the most popular and large-scale UDA benchmarks of image classification as well as semantic segmentation to demonstrate its generality and effectiveness. Lingsheng Kong, Xiongchang Liu, Jun Lu 0002, Jane You, Xiaofeng Liu 0001 |
Int. J. Intell. Syst. | 6 |
| 2022 | Mutual Information Regularized Feature-Level Frankenstein for Discriminative RecognitionabstractDeep learning recognition approaches can potentially perform better if we can extract a discriminative representation that controllably separates nuisance factors. In this paper, we propose a novel approach to explicitly enforce the extracted discriminative representation d, extracted latent variation l (e,g., background, unlabeled nuisance attributes), and semantic variation label vector s (e.g., labeled expressions/pose) to be independent and complementary to each other. We can cast this problem as an adversarial game in the latent space of an auto-encoder. Specifically, with the to-be-disentangled s, we propose to equip an end-to-end conditional adversarial network with the ability to decompose an input sample into d and l. However, we argue that maximizing the cross-entropy loss of semantic variation prediction from d is not sufficient to remove the impact of s from d, and that the uniform-target and entropy regularization are necessary. A collaborative mutual information regularization framework is further proposed to avoid unstable adversarial training. It is able to minimize the differentiable mutual information between the variables to enforce independence. The proposed discriminative representation inherits the desired tolerance property guided by prior knowledge of the task. Our proposed framework achieves top performance on diverse recognition tasks, including digits classification, large-scale face recognition on LFW and IJB-A datasets, and face recognition tolerant to changes in lighting, makeup, disguise, etc. Xiaofeng Liu 0001, Chao Yang 0011, Jane You, C.-C. Jay Kuo, B. V. K. Vijaya Kumar |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | VoxelHop: Successive Subspace Learning for ALS Disease Classification Using Structural MRIabstractDeep learning has great potential for accurate detection and classification of diseases with medical imaging data, but the performance is often limited by the number of training datasets and memory requirements. In addition, many deep learning models are considered a "black-box," thereby often limiting their adoption in clinical applications. To address this, we present a successive subspace learning model, termed VoxelHop, for accurate classification of Amyotrophic Lateral Sclerosis (ALS) using T2-weighted structural MRI data. Compared with popular convolutional neural network (CNN) architectures, VoxelHop has modular and transparent structures with fewer parameters without any backpropagation, so it is well-suited to small dataset size and 3D imaging data. Our VoxelHop has four key components, including (1) sequential expansion of near-to-far neighborhood for multi-channel 3D data; (2) subspace approximation for unsupervised dimension reduction; (3) label-assisted regression for supervised dimension reduction; and (4) concatenation of features and classification between controls and patients. Our experimental results demonstrate that our framework using a total of 20 controls and 26 patients achieves an accuracy of 93.48 % and an AUC score of 0.9394 in differentiating patients from controls, even with a relatively small number of datasets, showing its robustness and effectiveness. Our thorough evaluations also show its validity and superiority to the state-of-the-art 3D CNN classification approaches. Our framework can easily be generalized to other classification tasks using different imaging modalities. Xiaofeng Liu 0001, Fangxu Xing, Chao Yang 0011, C.-C. Jay Kuo, Suma Babu, Georges El Fakhri, Thomas Jenkins, Jonghye Woo |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Interpreting Depression From Question-Wise Long-Term Video Recording of SDS EvaluationabstractSelf-Rating Depression Scale (SDS) questionnaire has frequently been used for efficient depression preliminary screening. However, the uncontrollable self-administered measure can be easily affected by insouciantly or deceptively answering, and producing the different results with the clinician-administered Hamilton Depression Rating Scale (HDRS) and the final diagnosis. Clinically, facial expression (FE) and actions play a vital role in clinician-administered evaluation, while FE and action are underexplored for self-administered evaluations. In this work, we collect a novel dataset of 200 subjects to evidence the validity of self-rating questionnaires with their corresponding question-wise video recording. To automatically interpret depression from the SDS evaluation and the paired video, we propose an end-to-end hierarchical framework for the long-term variable-length video, which is also conditioned on the questionnaire results and the answering time. Specifically, we resort to a hierarchical model which utilizes a 3D CNN for local temporal pattern exploration and a redundancy-aware self-attention (RAS) scheme for question-wise global feature aggregation. Targeting for the redundant long-term FE video processing, our RAS is able to effectively exploit the correlations of each video clip within a question set to emphasize the discriminative information and eliminate the redundancy based on feature pair-wise affinity. Then, the question-wise video feature is concatenated with the questionnaire scores for final depression detection. Our thorough evaluations also show the validity of fusing SDS evaluation and its video recording, and the superiority of our framework to the conventional state-of-the-art temporal modeling methods. Wanqing Xie, Lizhong Liang, Jihong Shen, Xiaofeng Liu 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | Brain MR Atlas Construction Using Symmetric Deep Neural InpaintingabstractModeling statistical properties of anatomical structures using magnetic resonance imaging is essential for revealing common information of a target population and unique properties of specific subjects. In brain imaging, a statistical brain atlas is often constructed using a number of healthy subjects. When tumors are present, however, it is difficult to either provide a common space for various subjects or align their imaging data due to the unpredictable distribution of lesions. Here we propose a deep learning-based image inpainting method to replace the tumor regions with normal tissue intensities using only a patient population. Our framework has three major innovations: 1) incompletely distributed datasets with random tumor locations can be used for training; 2) irregularly-shaped tumor regions are properly learned, identified, and corrected; and 3) a symmetry constraint between the two brain hemispheres is applied to regularize inpainted regions. Henceforth, regular atlas construction and image registration methods can be applied using inpainted data to obtain tissue deformation, thereby achieving group-specific statistical atlases and patient-to-atlas registration. Our framework was tested using the public database from the Multimodal Brain Tumor Segmentation challenge. Results showed increased similarity scores as well as reduced reconstruction errors compared with three existing image inpainting methods. Patient-to-atlas registration also yielded better results with improved normalized cross-correlation and mutual information and a reduced amount of deformation over the tumor regions. Fangxu Xing, Xiaofeng Liu 0001, C.-C. Jay Kuo, Georges El Fakhri, Jonghye Woo |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Wasserstein Loss With Alternative Reinforcement Learning for Severity-Aware Semantic SegmentationabstractSemantic segmentation is important for many real-world systems, e.g., autonomous vehicles, which predict the class of each pixel. Recently, deep networks achieved significant progress w.r.t. the mean Intersection-over Union (mIoU) with the cross-entropy loss. However, the cross entropy loss can essentially ignore the difference of severity for an autonomous car with different wrong prediction mistakes. For example, predicting the car to the road is much more servery than recognize it as the bus. Targeting for this difficulty, we develop a Wasserstein training framework to explore the inter-class correlation by defining its ground metric as misclassification severity. The ground metric of Wasserstein distance can be pre-defined following the experience on a specific task. From the optimization perspective, we further propose to set the ground metric as an increasing function of the pre-defined ground metric. Furthermore, an adaptively learning scheme of the ground matrix is proposed to utilize the high-fidelity CARLA simulator. Specifically, we follow a reinforcement alternative learning scheme. The experiments on both CamVid and Cityscapes datasets evidenced the effectiveness of our Wasserstein loss. The SegNet, ENet, FCN and Deeplab networks can be adapted following a plug in manner. We achieve significant improves on the predefined important classes, and much longer continuous play time in our simulator. Xiaofeng Liu 0001, Yunhong Lu, Xiongchang Liu, Site Li, Jane You |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative ModelsabstractAI Safety is a major concern in many deep learning applications such as autonomous driving. Given a trained deep learning model, an important natural problem is how to reliably verify the model's prediction. In this paper, we propose a novel framework --- deep verifier networks (DVN) to detect unreliable inputs or predictions of deep discriminative models, using separately trained deep generative models. Our proposed model is based on conditional variational auto-encoders with disentanglement constraints to separate the label information from the latent representation. We give both intuitive and theoretical justifications for the model. Our verifier network is trained independently with the prediction model, which eliminates the need of retraining the verifier network for a new model. We test the verifier network on both out-of-distribution detection and adversarial example detection problems, as well as anomaly detection problems in structured prediction tasks such as image caption generation. We achieve state-of-the-art results in all of these problems. Tong Che, Xiaofeng Liu 0001, Site Li, Yubin Ge, Ruixiang Zhang, Caiming Xiong, Yoshua Bengio |
AAAI | 2 |
| 2021 | Subtype-aware Unsupervised Domain Adaptation for Medical DiagnosisabstractRecent advances in unsupervised domain adaptation (UDA) show that transferable prototypical learning presents a powerful means for class conditional alignment, which encourages the closeness of cross-domain class centroids. However, the cross-domain inner-class compactness and the underlying fine-grained subtype structure remained largely underexplored. In this work, we propose to adaptively carry out the fine-grained subtype-aware alignment by explicitly enforcing the class-wise separation and subtype-wise compactness with intermediate pseudo labels. Our key insight is that the unlabeled subtypes of a class can be divergent to one another with different conditional and label shifts, while inheriting the local proximity within a subtype. The cases with or without the prior information on subtype numbers are investigated to discover the underlying subtype structure in an online fashion. The proposed subtype-aware dynamic UDA achieves promising results on a medical diagnosis task. Xiaofeng Liu 0001, Xiongchang Liu, Wenxuan Ji, Fangxu Xing, Jun Lu 0002, Jane You, C.-C. Jay Kuo, Georges El Fakhri, Jonghye Woo |
AAAI | 1 |
| 2021 | BACO: A Background Knowledge- and Content-Based Framework for Citing Sentence GenerationabstractYubin Ge, Ly Dinh, Xiaofeng Liu, Jinsong Su, Ziyao Lu, Ante Wang, Jana Diesner. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yubin Ge, Ly Dinh, Xiaofeng Liu 0001, Jinsong Su, Ziyao Lu, Ante Wang, Jana Diesner |
ACL/IJCNLP (1) | 3 |
| 2021 | Embedding Semantic Hierarchy in Discrete Optimal Transport for Risk MinimizationabstractThe widely-used cross-entropy (CE) loss-based deep networks achieved significant progress w.r.t. the classification accuracy. However, the CE loss can essentially ignore the risk of misclassification which is usually measured by the distance between the prediction and label in a semantic hierarchical tree. In this paper, we propose to incorporate the risk-aware inter-class correlation in a discrete optimal transport (DOT) training framework by configuring its ground distance matrix. The ground distance matrix can be pre-defined following a priori of hierarchical semantic risk. Specifically, we define the tree induced error (TIE) on a hierarchical semantic tree and extend it to its increasing function from the optimization perspective. The semantic similarity in each level of a tree is integrated with the information gain. We achieve promising results on several large scale image classification tasks with a semantic tree structure in a plug and play manner. Yubin Ge, Site Li, Wanqing Xie, Jane You, Xiaofeng Liu 0001 |
ICASSP | 7 |
| 2021 | Adversarial Unsupervised Domain Adaptation with Conditional and Label Shift: Infer, Align and IterateabstractIn this work, we propose an adversarial unsupervised domain adaptation (UDA) method under inherent conditional and label shifts, in which we aim to align the distributions w.r.t. both p(x|y) and p(y). Since labels are inaccessible in a target domain, conventional adversarial UDA methods assume that p(y) is invariant across domains and rely on aligning p(x) as an alternative to the p(x|y) alignment. To address this, we provide a thorough theoretical and empirical analysis of the conventional adversarial UDA methods under both conditional and label shifts, and propose a novel and practical alternative optimization scheme for adversarial UDA. Specifically, we infer the marginal p(y) and align p(x|y) iteratively at the training stage, and precisely align the posterior p(y|x) at the testing stage. Our experimental results demonstrate its effectiveness on both classification and segmentation UDA and partial UDA. Xiaofeng Liu 0001, Zhenhua Guo 0001, Site Li, Fangxu Xing, Jane You, C.-C. Jay Kuo, Georges El Fakhri, Jonghye Woo |
ICCV | 1 |
| 2021 | Recursively Conditional Gaussian for Ordinal Unsupervised Domain AdaptationabstractThe unsupervised domain adaptation (UDA) has been widely adopted to alleviate the data scalability issue, while the existing works usually focus on classifying independently discrete labels. However, in many tasks (e.g., medical diagnosis), the labels are discrete and successively distributed. The UDA for ordinal classification requires inducing non-trivial ordinal distribution prior to the latent space. Target for this, the partially ordered set (poset) is defined for constraining the latent vector Instead of the typically i.i.d. Gaussian latent prior, in this work, a recursively conditional Gaussian (RCG) set is adapted for ordered constraint modeling, which admits a tractable joint distribution prior Furthermore, we are able to control the density of content vector that violates the poset constraints by a simple "three-sigma rule". We explicitly disentangle the cross-domain images into a shared ordinal prior induced ordinal content space and two separate source/target ordinal-unrelated spaces, and the self-training is worked on the shared space exclusively for ordinal-aware domain alignment. Extensive experiments on UDA medical diagnoses and facial age estimation demonstrate its effectiveness. Xiaofeng Liu 0001, Site Li, Yubin Ge, Pengyi Ye, Jane You, Jun Lu 0002 |
ICCV | 1 |
| 2021 | Domain Generalization under Conditional and Label Shifts via Variational Bayesian InferenceabstractIn this work, we propose a domain generalization (DG) approach to learn on several labeled source domains and transfer knowledge to a target domain that is inaccessible in training. Considering the inherent conditional and label shifts, we would expect the alignment of p(x|y) and p(y). However, the widely used domain invariant feature learning (IFL) methods relies on aligning the marginal concept shift w.r.t. p(x), which rests on an unrealistic assumption that p(y) is invariant across domains. We thereby propose a novel variational Bayesian inference framework to enforce the conditional distribution alignment w.r.t. p(x|y) via the prior distribution matching in a latent space, which also takes the marginal label shift w.r.t. p(y) into consideration with the posterior alignment. Extensive experiments on various benchmarks demonstrate that our framework is robust to the label shift and the cross-domain accuracy is significantly improved, thereby achieving superior performance over the conventional IFL counterparts. Xiaofeng Liu 0001, Linghao Jin, Fangxu Xing, Jinsong Ouyang, Jun Lu 0002, Georges El Fakhri, Jonghye Woo |
IJCAI | 1 |
| 2021 | Generative Self-training for Cross-Domain Unsupervised Tagged-to-Cine MRI Synthesis
Xiaofeng Liu 0001, Fangxu Xing, Maureen Stone 0001, Jiachen Zhuo, Timothy G. Reese, Jerry L. Prince, Georges El Fakhri, Jonghye Woo |
MICCAI (3) | 1 |
| 2021 | Adapting Off-the-Shelf Source Segmenter for Target Medical Image Segmentation
Xiaofeng Liu 0001, Fangxu Xing, Chao Yang 0011, Georges El Fakhri, Jonghye Woo |
MICCAI (2) | 1 |
| 2021 | Automated interpretation of congenital heart disease from multi-view echocardiograms
Xiaofeng Liu 0001, Fangyun Wang, Fengqiao Gao, Hanwen Zhang 0013, Wanqing Xie |
Medical Image Anal. | 2 |
| 2021 | Mutual information regularized identity-aware facial expression recognition in compressed video
Xiaofeng Liu 0001, Linghao Jin, Jane You |
Pattern Recognit. | 1 |
| 2020 | Importance-Aware Semantic Segmentation in Self-Driving with Discrete Wasserstein TrainingabstractSemantic segmentation (SS) is an important perception manner for self-driving cars and robotics, which classifies each pixel into a pre-determined class. The widely-used cross entropy (CE) loss-based deep networks has achieved significant progress w.r.t. the mean Intersection-over Union (mIoU). However, the cross entropy loss can not take the different importance of each class in an self-driving system into account. For example, pedestrians in the image should be much more important than the surrounding buildings when make a decisions in the driving, so their segmentation results are expected to be as accurate as possible. In this paper, we propose to incorporate the importance-aware inter-class correlation in a Wasserstein training framework by configuring its ground distance matrix. The ground distance matrix can be pre-defined following a priori in a specific task, and the previous importance-ignored methods can be the particular cases. From an optimization perspective, we also extend our ground metric to a linear, convex or concave increasing function w.r.t. pre-defined ground distance. We evaluate our method on CamVid and Cityscapes datasets with different backbones (SegNet, ENet, FCN and Deeplab) in a plug and play fashion. In our extenssive experiments, Wasserstein loss demonstrates superior segmentation performance on the predefined critical classes for safe-driving. Xiaofeng Liu 0001, Yuzhuo Han, Yi Ge, Tianxing Wang 0003, Site Li, Jane You, Jun Lu 0002 |
AAAI | 1 |
| 2020 | Severity-Aware Semantic Segmentation With Reinforced Wasserstein TrainingabstractSemantic segmentation is a class of methods to classify each pixel in an image into semantic classes, which is critical for autonomous vehicles and surgery systems. Cross-entropy (CE) loss-based deep neural networks (DNN) achieved great success w.r.t. the accuracy-based metrics, e.g., mean Intersection-over Union. However, the CE loss has a limitation in that it ignores varying degrees of severity of pair-wise misclassified results. For instance, classifying a car into the road is much more terrible than recognizing it as a bus. To sidestep this, in this work, we propose to incorporate the severity-aware inter-class correlation into our Wasserstein training framework by configuring its ground distance matrix. In addition, our method can adaptively learn the ground metric in a high-fidelity simulator, following a reinforcement alternative optimization scheme. We evaluate our method using the CARLA simulator with the Deeplab backbone, demonstraing that our method significantly improves the survival time in the CARLA simulator. In addition, our method can be readily applied to existing DNN architectures and algorithms while yielding superior performance. We report results from experiments carried out with the CamVid and Cityscapes datasets. Xiaofeng Liu 0001, Wenxuan Ji, Jane You, Georges El Fakhri, Jonghye Woo |
CVPR | 1 |
| 2020 | AUTO3D: Novel View Synthesis Through Unsupervisely Learned Variational Viewpoint and Global 3D Representation
Xiaofeng Liu 0001, Tong Che, Yiqun Lu, Chao Yang 0011, Site Li, Jane You |
ECCV (9) | 1 |
| 2020 | Energy-constrained Self-training for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims to transfer the knowledge on a labeled source domain distribution to perform well on an unlabeled target domain. Recently, the deep self-training involves an iterative process of predicting on the target domain and then taking the confident predictions as hard pseudo-labels for retraining. However, the pseudo-labels are usually unreliable, and easily leading to deviated solutions with propagated errors. In this paper, we resort to the energy-based model and constrain the training of the unlabeled target sample with the energy function minimization objective. It can be applied as a simple additional regularization. In this framework, it is possible to gain the benefits of the energy-based model, while retaining strong discriminative performance following a plug-and-play fashion. We deliver extensive experiments on the most popular and large scale UDA benchmarks of image classification as well as semantic segmentation to demonstrate its generality and effectiveness. Xiaofeng Liu 0001, Xiongchang Liu, Jun Lu 0002, Jane You, Lingsheng Kong |
ICPR | 1 |
| 2020 | Identity-aware Facial Expression Recognition in Compressed VideoabstractThis paper targets to explore the inter-subject variations eliminated facial expression representation in the compressed video domain. Most of the previous methods process the RGB images of a sequence, while the off-the-shelf and valuable expression-related muscle movement already embedded in the compression format. In the up to two orders of magnitude compressed domain, we can explicitly infer the expression from the residual frames and possible to extract identity factors from the I frame with a pre-trained face recognition network. By enforcing the marginal independent of them, the expression feature is expected to be purer for the expression and be robust to identity shifts. We do not need the identity label or multiple expression samples from the same person for identity elimination. Moreover, when the apex frame is annotated in the dataset, the complementary constraint can be further added to regularize the feature-level game. In testing, only the compressed residual frames are required to achieve expression prediction. Our solution can achieve comparable or better performance than the recent decoded image based methods on the typical FER benchmarks with about 3× faster inference with compressed data. Xiaofeng Liu 0001, Linghao Jin, Jun Lu 0002, Jane You, Lingsheng Kong |
ICPR | 1 |
| 2020 | Unimodal regularized neuron stick-breaking for ordinal classification
Xiaofeng Liu 0001, Lingsheng Kong, Zhihui Diao, Wanqing Xie, Jun Lu 0002, Jane You |
Neurocomputing | 1 |
| 2020 | Dependency-Aware Attention Control for Image Set-Based Face RecognitionabstractThis paper considers the problem of image set-based face verification and identification. Unlike traditional single sample (an image or a video) setting, this situation assumes the availability of a set of heterogeneous collection of orderless images and videos. The samples can be taken at different check points, different identity documents $etc$ . The importance of each image is usually considered either equal or based on a quality assessment of that image independent of other images and/or videos in that image set. How to model the relationship of orderless images within a set remains a challenge. We address this problem by formulating it as a Markov Decision Process (MDP) in a latent space. Specifically, we first propose a dependency-aware attention control (DAC) network, which uses actor-critic reinforcement learning for attention decision of each image to exploit the correlations among the unordered images. An off-policy experience replay is introduced to speed up the learning process. Moreover, the DAC is combined with a temporal model for videos using divide and conquer strategies. We also introduce a pose-guided representation (PGR) scheme that can further boost the performance at extreme poses. We propose a parameter-free PGR without the need for training as well as a novel metric learning-based PGR for pose alignment without the need for pose detection in testing stage. Extensive evaluations on IJB-A/B/C, YTF, Celebrity-1000 datasets demonstrate that our method outperforms many state-of-art approaches on the set-based as well as video-based face recognition databases. Xiaofeng Liu 0001, Zhenhua Guo 0001, Jane You, B. V. K. Vijaya Kumar |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2019 | Feature-Level Frankenstein: Eliminating Variations for Discriminative RecognitionabstractRecent successes of deep learning-based recognition rely on maintaining the content related to the main-task label. However, how to explicitly dispel the noisy signals for better generalization remains an open issue. We systematically summarize the detrimental factors as task-relevant/irrelevant semantic variations and unspecified latent variation. In this paper, we cast these problems as an adversarial minimax game in the latent space. Specifically, we propose equipping an end-to-end conditional adversarial network with the ability to decompose an input sample into three complementary parts. The discriminative representation inherits the desired invariance property guided by prior knowledge of the task, which is marginally independent to the task-relevant/irrelevant semantic and latent variations. Our proposed framework achieves top performance on a serial of tasks, including digits recognition, lighting, makeup, disguise-tolerant face recognition, and facial attributes recognition. Xiaofeng Liu 0001, Site Li, Lingsheng Kong, Wanqing Xie, Ping Jia, Jane You, B. V. K. Vijaya Kumar |
CVPR | 1 |
| 2019 | Permutation-Invariant Feature Restructuring for Correlation-Aware Image Set-Based RecognitionabstractWe consider the problem of comparing the similarity of image sets with variable-quantity, quality and un-ordered heterogeneous images. We use feature restructuring to exploit the correlations of both inner&inter-set images. Specifically, the residual self-attention can effectively restructure the features using the other features within a set to emphasize the discriminative images and eliminate the redundancy. Then, a sparse/collaborative learning-based dependency-guided representation scheme reconstructs the probe features conditional to the gallery features in order to adaptively align the two sets. This enables our framework to be compatible with both verification and open-set identification. We show that the parametric self-attention network and non-parametric dictionary learning can be trained end-to-end by a unified alternative optimization scheme, and that the full framework is permutation-invariant. In the numerical experiments we conducted, our method achieves top performance on competitive image set/video-based face recognition and person re-identification benchmarks. Xiaofeng Liu 0001, Zhenhua Guo 0001, Site Li, Ping Jia, Lingsheng Kong, Jane You, B. V. K. Vijaya Kumar |
ICCV | 1 |
| 2019 | Conservative Wasserstein Training for Pose EstimationabstractThis paper targets the task with discrete and periodic class labels (e.g., pose/orientation estimation) in the context of deep learning. The commonly used cross-entropy or regression loss is not well matched to this problem as they ignore the periodic nature of the labels and the class similarity, or assume labels are continuous value. We propose to incorporate inter-class correlations in a Wasserstein training framework by pre-defining (i.e., using arc length of a circle) or adaptively learning the ground metric. We extend the ground metric as a linear, convex or concave increasing function w.r.t. arc length from an optimization perspective. We also propose to construct the conservative target labels which model the inlier and outlier noises using a wrapped unimodal-uniform mixture distribution. Unlike the one-hot setting, the conservative label makes the computation of Wasserstein distance more challenging. We systematically conclude the practical closed-form solution of Wasserstein distance for pose data with either one-hot or conservative target label. We evaluate our method on head, body, vehicle and 3D object pose benchmarks with exhaustive ablation studies. The Wasserstein loss obtaining superior performance over the current methods, especially using convex mapping function for ground metric, conservative label, and closed-form solution. Xiaofeng Liu 0001, Yang Zou 0003, Tong Che, Ping Jia, Jane You, B. V. K. Vijaya Kumar |
ICCV | 1 |
| 2019 | Confidence Regularized Self-TrainingabstractRecent advances in domain adaptation show that deep self-training presents a powerful means for unsupervised domain adaptation. These methods often involve an iterative process of predicting on target domain and then taking the confident predictions as pseudo-labels for retraining. However, since pseudo-labels can be noisy, self-training can put overconfident label belief on wrong classes, leading to deviated solutions with propagated errors. To address the problem, we propose a confidence regularized self-training (CRST) framework, formulated as regularized self-training. Our method treats pseudo-labels as continuous latent variables jointly optimized via alternating optimization. We propose two types of confidence regularization: label regularization (LR) and model regularization (MR). CRST-LR generates soft pseudo-labels while CRST-MR encourages the smoothness on network output. Extensive experiments on image classification and semantic segmentation show that CRSTs outperform their non-regularized counterpart with state-of-the-art performance. The code and models of this work are available at https://github.com/yzou2/CRST. Yang Zou 0003, Zhiding Yu, Xiaofeng Liu 0001, B. V. K. Vijaya Kumar |
ICCV | 3 |
| 2019 | Hard negative generation for identity-disentangled facial expression recognition
Xiaofeng Liu 0001, B. V. K. Vijaya Kumar, Ping Jia, Jane You |
Pattern Recognit. | 1 |
| 2018 | Dependency-Aware Attention Control for Unconstrained Face Recognition with Image Sets
Xiaofeng Liu 0001, B. V. K. Vijaya Kumar, Chao Yang 0011, Qingming Tang, Jane You |
ECCV (11) | 1 |
| 2018 | Contextual-Based Image Inpainting: Infer, Match, and Translate
Yuhang Song 0003, Chao Yang 0011, Zhe Lin 0001, Xiaofeng Liu 0001, Qin Huang 0006, Hao Li 0015, C.-C. Jay Kuo |
ECCV (2) | 4 |
| 2018 | A joint optimization framework of low-dimensional projection and collaborative representation for discriminative classificationabstractVarious representation-based methods have been developed and shown great potential for pattern classification. To further improve their discriminability, we propose a Bi-level optimization framework in terms of both low-dimensional projection and collaborative representation. Specifically, during the projection phase, we try to minimize the intra-class similarity and inter-class dissimilarity, while in the representation phase, our goal is to achieve the lowest correlation of the representation results. Solving this joint optimization mutually reinforces both aspects of feature projection and representation. Experiments on face recognition, object categorization and scene classification dataset demonstrate remarkable performance improvements led by the proposed framework. Xiaofeng Liu 0001, Zhaofeng Li 0002, Lingsheng Kong, Zhihui Diao, Junliang Yan, Yang Zou 0003, Chao Yang 0011, Ping Jia, Jane You |
ICPR | 1 |
| 2018 | Data Augmentation via Latent Space Interpolation for Image ClassificationabstractEffective training of the deep neural networks requires much data to avoid underdetermined and poor generalization. Data Augmentation alleviates this by using existing data more effectively. However standard data augmentation produces only limited plausible alternative data by for example, flipping, distorting, adding noise to, cropping a patch from the original samples. In this paper, we introduce the adversarial autoencoder (AAE) to impose the feature representations with uniform distribution and apply the linear interpolation on latent space, which is potential to generate a much broader set of augmentations for image classification. As a possible “recognition via generation” framework, it has potentials for several other classification tasks. Our experiments on the ILSVRC 2012, CIFAR-10 datasets show that the latent space interpolation (LSI) improves the generalization and performance of state-of-the-art deep neural networks. Xiaofeng Liu 0001, Yang Zou 0003, Lingsheng Kong, Zhihui Diao, Junliang Yan, Site Li, Ping Jia, Jane You |
ICPR | 1 |
| 2012 | Incompressible Deformation Estimation Algorithm (IDEA) From Tagged MR ImagesabstractMeasuring the 3D motion of muscular tissues, e.g., the heart or the tongue, using magnetic resonance (MR) tagging is typically carried out by interpolating the 2D motion information measured on orthogonal stacks of images. The incompressibility of muscle tissue is an important constraint on the reconstructed motion field and can significantly help to counter the sparsity and incompleteness of the available motion information. Previous methods utilizing this fact produced incompressible motions with limited accuracy. In this paper, we present an incompressible deformation estimation algorithm (IDEA) that reconstructs a dense representation of the 3D displacement field from tagged MR images and the estimated motion field is incompressible to high precision. At each imaged time frame, the tagged images are first processed to determine components of the displacement vector at each pixel relative to the reference time. IDEA then applies a smoothing, divergence-free, vector spline to interpolate velocity fields at intermediate discrete times such that the collection of velocity fields integrate over time to match the observed displacement components. Through this process, IDEA yields a dense estimate of a 3D displacement field that matches our observations and also corresponds to an incompressible motion. The method was validated with both numerical simulation and in vivo human experiments on the heart and the tongue. Xiaofeng Liu 0001, Khaled Z. Abd-Elmoniem, Maureen Stone 0001, Emi Z. Murano, Jiachen Zhuo, Rao P. Gullapalli, Jerry L. Prince |
IEEE Trans. Medical Imaging | 1 |
| 2010 | Shortest Path Refinement for Motion Estimation From Tagged MR ImagesabstractMagnetic resonance tagging makes it possible to measure the motion of tissues such as muscles in the heart and tongue. The harmonic phase (HARP) method largely automates the process of tracking points within tagged MR images, permitting many motion properties to be computed. However, HARP tracking can yield erroneous motion estimates due to 1) large deformations between image frames, 2) through-plane motion, and 3) tissue boundaries. Methods that incorporate the spatial continuity of motion--so-called refinement or flood-filling methods--have previously been reported to reduce tracking errors. This paper presents a new refinement method based on shortest path computations. The method uses a graph representation of the image and seeks an optimal tracking order from a specified seed to each point in the image by solving a single source shortest path problem. This minimizes the potential errors for those path dependent solutions that are found in other refinement methods. In addition to this, tracking in the presence of through-plane motion is improved by introducing synthetic tags at the reference time (when the tissue is not deformed). Experimental results on both tongue and cardiac images show that the proposed method can track the whole tissue more robustly and is also computationally efficient. Xiaofeng Liu 0001, Jerry L. Prince |
IEEE Trans. Medical Imaging | 1 |
| 2009 | Incompressible Cardiac Motion Estimation of the Left Ventricle Using Tagged MR Images
Xiaofeng Liu 0001, Khaled Z. Abd-Elmoniem, Jerry L. Prince |
MICCAI (1) | 1 |
| 2009 | Prostate Brachytherapy Seed Reconstruction With Gaussian Blurring and Optimal Coverage CostabstractIntraoperative dosimetry in prostate brachytherapy requires localization of the implanted radioactive seeds. A tomosynthesis-based seed reconstruction method is proposed. A three-dimensional volume is reconstructed from Gaussian-blurred projection images and candidate seed locations are computed from the reconstructed volume. A false positive seed removal process, formulated as an optimal coverage problem, iteratively removes "ghost" seeds that are created by tomosynthesis reconstruction. In an effort to minimize pose errors that are common in conventional C-arms, initial pose parameter estimates are iteratively corrected by using the detected candidate seeds as fiducials, which automatically "focuses" the collected images and improves successive reconstructed volumes. Simulation results imply that the implanted seed locations can be estimated with a detection rate of > or = 97.9% and > or = 99.3% from three and four images, respectively, when the C-arm is calibrated and the pose of the C-arm is known. The algorithm was also validated on phantom data sets successfully localizing the implanted seeds from four or five images. In a Phase-1 clinical trial, we were able to localize the implanted seeds from five intraoperative fluoroscopy images with 98.8% (STD=1.6) overall detection rate. Xiaofeng Liu 0001, Ameet K. Jain, Danny Y. Song, Everette Clif Burdette, Jerry L. Prince, Gabor Fichtinger |
IEEE Trans. Medical Imaging | 2 |
| 2008 | Prostate Brachytherapy Seed Localization with Gaussian Blurring and Camera Self-calibration
Xiaofeng Liu 0001, Jerry L. Prince, Gabor Fichtinger |
MICCAI (2) | 2 |
| 2007 | Prostate Implant Reconstruction with Discrete Tomography
Xiaofeng Liu 0001, Ameet K. Jain, Gabor Fichtinger |
MICCAI (1) | 1 |