EDBT 2026 Demo / reviewers in the wild / expert
Jian Sun 0009
dblp:68/4942-9
· DBLP profile ↗
99ranked-venue papers
13as first author
54since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 61 · 10 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 47 · 9 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 10 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAME: Safety-Aware Model Editing Guided by Safety TransformationabstractEditing large language models is challenging as incorporating new knowledge often requires sequential parameter updates while maintaining model capability.In this work, we experimentally observe that sequential knowledge updating under the locate-then-edit framework can introduce safety risks, regardless of whether the knowledge being edited is benign or malicious.We propose a novel model editing approach that estimates safety transforms and identifies corresponding safety direction in the neural activation space, and then aligns neural activation updates and network parameter updates under the safety constraints, resulting in a safety-aware model editing approach.We evaluate our approach on open-source LLMs, Llama-3-8B-Instruct, Qwen3-4B-Instruct and Qwen2.5-14B-Instruct,using the benchmark datasets ZsRE and COUNTERFACT, as well as the malicious dataset Mal-KSet.Experimental results demonstrate that our approach effectively reduces unsafe responses to malicious queries while preserving the effectiveness of model editing. Shipeng Wang 0002, Jian Sun 0009 |
ACL (1) | 4 |
| 2026 | Generative Diffusion-Based Bayesian Modeling for Universal Channel EstimationabstractThe growth of frequency bandwidths in the new generation of wireless networks gives rise to the multitude of wireless communication scenarios and highlights the challenges of generalized capability of the wireless communication system in different scenarios, especially the channel estimation module. In this paper, we propose a large model dubbed Conditional Latent Diffusion Channel Generation Model (C-LCGM) to learn the distributions of channel state information (CSI) in different wireless communication scenarios to form Bayesian Modeling based Channel Estimation Scheme (BMCE) for universal channel estimation. BMCE conducts universal channel estimation by generating reference CSIs from C-LCGM and mapping the reference CSIs to the optimal channel estimation neural network for each scenario. Specifically in C-LCGM, we propose to compress the CSIs into latent codes and design a conditional diffusion model to model the distribution of the latent codes given the large-scale parameters (LSP) of CSIs as the condition. Further in BMCE, we propose to deploy C-LCGM on the server center and design a hyper-network dubbed Parameters Generating Module (PGM) to map the generated CSIs of C-LCGM to the channel estimation networks for the base stations (BS) according to the reported LSPs. The design rationale and training loss of C-LCGM and BMCE are derived theoretically in this paper. We also conduct extensive simulations to verify the performance of C-LCGM and BMCE. The simulation results show that BMCE can achieve optimal channel estimation performance in different and novel scenarios with C-LCGM generating high-quality CSIs approximating the real CSIs in each scenario. Complexity analysis shows BMCE can fit the delay requirement of wireless communication systems. Runhua Li, Jian Sun 0009, Jiang Xue 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2026 | Learning Continuous Wasserstein Barycenter Space for Generalized All-in-One Image RestorationabstractDespite substantial advances in all-in-one image restoration for addressing diverse degradations within a unified model, existing methods remain vulnerable to out-of-distribution degradations, thereby limiting their generalization in real-world scenarios. To tackle the challenge, this work is motivated by the intuition that multisource degraded feature distributions are induced by different degradation-specific shifts from an underlying degradation-agnostic distribution, and recovering such a shared distribution is thus crucial for achieving generalization across degradations. With this insight, we propose BaryIR, a representation learning framework that aligns multisource degraded features in the Wasserstein barycenter (WB) space, which models a degradation-agnostic distribution by minimizing the average of Wasserstein distances to multisource degraded distributions. We further introduce residual subspaces, whose embeddings are mutually contrasted while remaining orthogonal to the WB embeddings. Consequently, BaryIR explicitly decouples two orthogonal spaces: a WB space that encodes the degradation-agnostic invariant contents shared across degradations, and residual subspaces that adaptively preserve the degradation-specific knowledge. This disentanglement mitigates overfitting to in-distribution degradations and enables adaptive restoration grounded on the degradation-agnostic shared invariance. Extensive experiments demonstrate that BaryIR performs competitively against state-of-the-art all-in-one methods. Notably, BaryIR generalizes well to unseen degradations (e.g., types and levels) and shows remarkable robustness in learning generalized features, even when trained on limited degradation types and evaluated on real-world data with mixed degradations. Xiaole Tang, Xiaoyi He, Xiang Gu 0005, Jian Sun 0009 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Variational masking generative model for anomaly detection on incomplete tabular data
Yannan Pu, Xiang Gu 0005, Niansheng Tang, Jian Sun 0009 |
Pattern Recognit. | 5 |
| 2026 | Re-Visible Dual-Domain Self-Supervised Deep Unfolding Network for MRI ReconstructionabstractMagnetic Resonance Imaging (MRI) is widely used in clinical practice, but suffers from prolonged acquisition time. Although deep learning methods have been proposed to accelerate acquisition and demonstrate promising performance, they rely on high-quality fully-sampled datasets for training in a supervised manner. However, such datasets are time-consuming and expensive-to-collect, which constrains their broader applications. On the other hand, self-supervised methods offer an alternative by enabling learning from under-sampled data alone, but most existing methods rely on further partitioned under-sampled k-space data as model's input for training, which causes an input distribution shift between the the training stage and the inference stage. Additionally, their models have not effectively incorporated comprehensive image priors, leading to degraded reconstruction performance. In this paper, we propose a novel re-visible dual-domain self-supervised deep unfolding network to address these issues when only under-sampled datasets are available. Specifically, by incorporating re-visible dual-domain loss, all under-sampled k-space data are utilized during training to mitigate the input distribution shift caused by further partitioning. This design enables the model to implicitly adapt to all under-sampled k-space data as input. Additionally, we design a Deep Unfolding Network based on Chambolle and Pock Proximal Point Algorithm (DUN-CP-PPA) to achieve end-to-end reconstruction. By employing a Spatial-Frequency Feature Extraction (SFFE) block to capture both global and local representations, the model effectively integrates imaging physics with comprehensive image priors to enhance reconstruction performance. Experiments on both single-coil and multi-coil datasets demonstrate that our method outperforms state-of-the-art approaches in terms of reconstruction performance and generalization capability. Hao Zhang 0026, Qi Wang 0128, Jian Sun 0009, Zhijie Wen, Jun Shi 0004, Shihui Ying |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Wasserstein Style Distribution Analysis and Transform for Stylized Image Generation
Xiang Gu 0005, Zhihao Shi, Jian Sun 0009 |
ICCV | 4 |
| 2025 | Efficient Multi-Person Motion Prediction by Lightweight Spatial and Temporal Interactionsabstract3D multi-person motion prediction is a highly complex task, primarily due to the dependencies on both individual past movements and the interactions between agents. Moreover, effectively modeling these interactions often incurs substantial computational costs. In this work, we propose a computationally efficient model for multi-person motion prediction by simplifying spatial and temporal interactions. Our approach begins with the design of lightweight dual branches that learn local and global representations for individual and multiple persons separately. Additionally, we introduce a novel cross-level interaction block to integrate the spatial and temporal representations from both branches. To further enhance interaction modeling, we explicitly incorporate the spatial inter-person distance embedding. With above efficient temporal and spatial design, we achieve state-of-the-art performance for multiple metrics on standard datasets of CMU-Mocap, MuPoTS-3D, and 3DPW, while significantly reducing the computational cost. Code is available at https://github.com/Yuanhong-Zheng/EMPMP. Yuanhong Zheng, Ruixuan Yu, Jian Sun 0009 |
ICCV | 3 |
| 2025 | Improved Baselines with Synchronized Encoding for Universal Medical Image Segmentation
Jiadong Feng, Xuande Mi, Haixia Bi, Jian Sun 0009 |
MICCAI (2) | 6 |
| 2025 | SpiderSolver: A Geometry-Aware Transformer for Solving PDEs on Complex GeometriesabstractTransformers have demonstrated effectiveness in solving partial differential equations (PDEs). However, extending them to solve PDEs on complex geometries remains a challenge. In this work, we propose SpiderSolver, a geometry-aware transformer that introduces spiderweb tokenization for handling complex domain geometry and irregularly discretized points. Our method partitions the irregular spatial domain into spiderweb-like patches, guided by the domain boundary geometry. SpiderSolver leverages a coarse-grained attention mechanism to capture global interactions across spiderweb tokens and a fine-grained attention mechanism to refine feature interactions between the domain boundary and its neighboring interior points. We evaluate SpiderSolver on PDEs with diverse domain geometries across seven datasets, including cars, airfoils, blood flow in the human thoracic aorta, as well as canonical cases governed by the Navier-Stokes, Darcy flow, elasticity, and plasticity equations. Experimental results demonstrate that SpiderSolver consistently achieves state-of-the-art performance across different datasets and metrics, with better generalization ability in the OOD setting. The code is available at https://github.com/Kai-Qi/SpiderSolver. Kai Qi, Zhewen Dong, Jian Sun 0009 |
NeurIPS | 4 |
| 2025 | Towards Prospective Medical Image Reconstruction via Knowledge-Informed Dynamic Optimal TransportabstractMedical image reconstruction from measurement data is a vital but challenging inverse problem. Deep learning approaches have achieved promising results, but often requires paired measurement and high-quality images, which is typically simulated through a forward model, i.e., retrospective reconstruction. However, training on simulated pairs commonly leads to performance degradation on real prospective data due to the retrospective-to-prospective gap caused by incomplete imaging knowledge in simulation. To address this challenge, this paper introduces imaging Knowledge-Informed Dynamic Optimal Transport (KIDOT), a novel dynamic optimal transport framework with optimality in the sense of preserving consistency with imaging physics in transport, that conceptualizes reconstruction as finding a dynamic transport path. KIDOT learns from unpaired data by modeling reconstruction as a continuous evolution path from measurements to images, guided by an imaging knowledge-informed cost function and transport equation. This dynamic and knowledge-aware approach enhances robustness and better leverages unpaired data while respecting acquisition physics. Theoretically, we demonstrate that KIDOT naturally generalizes dynamic optimal transport, ensuring its mathematical rationale and solution existence. Extensive experiments on MRI and CT reconstruction demonstrate KIDOT's superior performance. Code is available at https://github.com/TaoranZheng717/KIDOT. Taoran Zheng, Yan Yang 0007, Xing Li 0027, Xiang Gu 0005, Jian Sun 0009, Zongben Xu |
NeurIPS | 5 |
| 2025 | Self-supervised distributional and contrastive learning model for image anomaly detection
Yannan Pu, Jian Sun 0009, Niansheng Tang, Zongben Xu |
Knowl. Based Syst. | 2 |
| 2025 | Degradation-Aware Residual-Conditioned Optimal Transport for Unified Image RestorationabstractUnified, or more formally, all-in-one image restoration has emerged as a practical and promising low-level vision task for real-world applications. In this context, the key issue lies in how to deal with different types of degraded images simultaneously. Existing methods fit joint regression models over multi-domain degraded-clean image pairs of different degradations. However, due to the severe ill-posedness of inverting heterogeneous degradations, they often struggle with thoroughly perceiving the degradation semantics and rely on paired data for supervised training, yielding suboptimal restoration maps with structurally compromised results and lacking practicality for real-world or unpaired data. To break the barriers, we present a Degradation-Aware Residual-Conditioned Optimal Transport (DA-RCOT) approach that models (all-in-one) image restoration as an optimal transport (OT) problem for unpaired and paired settings, introducing the transport residual as a degradation-specific cue for both the transport cost and the transport map. Specifically, we formalize image restoration with a residual-guided OT objective by exploiting the degradation-specific patterns of the Fourier residual in the transport cost. More crucially, we design the transport map for restoration as a two-pass DA-RCOT map, in which the transport residual is computed in the first pass and then encoded as multi-scale residual embeddings to condition the second-pass restoration. This conditioning process injects intrinsic degradation knowledge (e.g., degradation type and level) and structural information from the multi-scale residual embeddings into the OT map, which thereby can dynamically adjust its behaviors for all-in-one restoration. Extensive experiments across five degradations demonstrate the favorable performance of DA-RCOT as compared to state-of-the-art methods, in terms of distortion measures, perceptual quality, and image structure preservation. Notably, DA-RCOT delivers superior adaptability to real-world scenarios even with mixed degradations and shows distinctive robustness to both degradation levels and the number of degradations. Xiaole Tang, Xiang Gu 0005, Xiaoyi He, Jian Sun 0009 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Training Networks in Null Space of Feature Covariance With Self-Supervision for Incremental LearningabstractIn the context of incremental learning, a network is sequentially trained on a stream of tasks, where data from previous tasks are particularly assumed to be inaccessible. The major challenge is how to overcome the stability-plasticity dilemma, i.e., learning knowledge from new tasks without forgetting the knowledge of previous tasks. To this end, we propose two mathematical conditions for guaranteeing network stability and plasticity with theoretical analysis. The conditions demonstrate that we can restrict the parameter update in the null space of uncentered feature covariance at each linear layer to overcome the stability-plasticity dilemma, which can be realized by layerwise projecting gradient into the null space. Inspired by it, we develop two algorithms, dubbed Adam-NSCL and Adam-SFCL respectively, for incremental learning. Adam-NSCL and Adam-SFCL provide different ways to compute the projection matrix. The projection matrix in Adam-NSCL is constructed by singular vectors associated with the smallest singular values of the uncentered feature covariance matrix, while the projection matrix in Adam-SFCL is constructed by all singular vectors associated with adaptive scaling factors. Additionally, we explore adopting self-supervised techniques, including self-supervised label augmentation and a newly proposed contrastive loss, to improve the performance of incremental learning. These self-supervised techniques are orthogonal to Adam-NSCL and Adam-SFCL and can be incorporated with them seamlessly, leading to Adam-NSCL-SSL and Adam-SFCL-SSL respectively. The proposed algorithms are applied to task-incremental and class-incremental learning on various benchmark datasets with multiple backbones, and the results show that they outperform the compared incremental learning methods. Shipeng Wang 0002, Xiaorong Li, Jian Sun 0009, Zongben Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Multi-Scale Part-Based Feature Representation for 3D Domain Generalization and AdaptationabstractDeep networks for 3D point clouds have achieved remarkable success in classification task but remain vulnerable to geometric variations resulting from inconsistent data acquisition procedures. This leads to significant performance degradation when models trained on a source domain are tested on out-of-distribution target domains, highlighting the challenges of 3D domain generalization and adaptation. In this paper, we introduce a novel Multi-Scale Part-based feature Representation, dubbed MSPR, as a generalizable representation for point cloud domain generalization and adaptation. Rather than relying on global shape feature, we align the part-level features of shapes at different scales to a set of learnable part-template features that encode local geometric structures shared between the source and the target domains. Specifically, shapes from different domains are organized into part-level features at various scales and then aligned to the part-template features. To balance the generalization and discrimination abilities of parts at different scales, we further design a cross-scale feature fusion module to exchange information between aligned part-based features at different scales. The fused part-based representations are finally aggregated by a part-based feature aggregation module. To improve the robustness of the aligned part-based representations and global shape representation to geometry variations, we further propose a Contrastive Learning framework on Shape Representation (CLSR). Experiments are conducted on 3D domain generalization and adaptation benchmarks for point cloud classification. Extensive experiments on 3D domain generalization and adaptation benchmarks demonstrate that proposed approach outperforms previous state-of-the-art methods in both tasks. Ablation studies confirm the effectiveness of the components in our model. Xiang Gu 0005, Jian Sun 0009 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Adaptive representation learning and sample weighting for low-quality 3D face recognition
Cuican Yu, Fengxun Sun, Huibin Li 0001, Liming Chen 0002, Jian Sun 0009, Zongben Xu |
Pattern Recognit. | 6 |
| 2025 | In-Context Learning Implicit Representation for Time Series Forecasting
Xiang Gu 0005, Jian Sun 0009 |
IEEE Signal Process. Lett. | 3 |
| 2025 | UniSTAD: An Unified Triple-Tower Student-Teacher Model for Multi-Class Anomaly Detection and LocalizationabstractDespite the rapid advancements in the unsupervised anomaly detection and localization, most existing methods require to train different models for different categories, leading to increased computational and memory demands for real application with the number of classes grows. A more practical task is to detect anomalies from different categories using one unified model. However, this unified setting is challenging for modeling the multi-class normal feature representation due to the diversity of data categories, and the existing methods often drop in performance under this setting. In this work, we propose UniSTAD, a novel and effective unified method for multi-class anomaly detection and localization, using a transformer-based triple-tower students-teacher model. The triple-tower design contains global and local student models, respectively predicting features from global and local context features. UniSTAD learns the feature representation of normal data by joint distilling features to pre-trained teacher model, and enforcing the global/local context-based feature reconstruction and consistency. In the inference stage, UniSTAD identifies anomalous regions where expected feature consistencies are broken. Additionally, we integrate an untrained, category-agnostic localization refinement module, further improving multi-class anomaly detection and localization performance. Evaluated on real-world industrial datasets, UniSTAD demonstrates the state-of-the-art performance, validating its efficacy for multi-class anomaly detection and localization. Jian Sun 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Bidirectional Projection-Based Multi-Modal Fusion Transformer for Early Detection of Cerebral Palsy in InfantsabstractPeriventricular white matter injury (PWMI) is the most frequent magnetic resonance imaging (MRI) finding in infants with Cerebral Palsy (CP). We aim to detect CP and identify subtle, sparse PWMI lesions in infants under two years of age with immature brain structures. Based on the characteristic that the responsible lesions are located within five target regions, we first construct a multi-modal dataset including 243 cases with the mask annotations of five target regions for delineating anatomical structures on T1-Weighted Imaging (T1WI) images, masks for lesions on T2-Weighted Imaging (T2WI) images, and categories (CP or Non-CP). Furthermore, we develop a bidirectional projection-based multi-modal fusion transformer (BiP-MFT), incorporating a Bidirectional Projection Fusion Module (BPFM) for integrating the features between five target regions on T1WI images and lesions on T2WI images. Our BiP-MFT achieves subject-level classification accuracy of 0.90, specificity of 0.87, and sensitivity of 0.94. It surpasses the best results of nine comparative methods, with 0.10, 0.08, and 0.09 improvements in classification accuracy, specificity and sensitivity respectively. Our BPFM outperforms eight compared feature fusion strategies using Transformer and U-Net backbones on our dataset. Ablation studies on the dataset annotations and model components justify the effectiveness of our annotation method and the model rationality. The proposed dataset and codes are available at https://github.com/Kai-Qi/BiP-MFT. Kai Qi, Yizhe Yang, Shihui Ying, Jian Sun 0009 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Domain-Generalized Discrete Diffusion Model for Cross-Domain Medical Image SegmentationabstractDomain shift is a significant challenge in medical image segmentation, primarily due to variations in image acquisition protocols, modalities, etc. Domain shift often causes models trained on a source domain to perform poorly on unseen target domains. In this work, we introduce the Domain-Generalized Discrete Diffusion Model for Segmentation (DG-DDM-Seg), a diffusion-based generative model designed for single-source domain generalization in medical image segmentation. DG-DDM-Seg generates discrete conditional distributions of segmentation masks. To ensure domain independence, we employ two key strategies: 1) We extract robust features from conditional images to enhance the domain independence of diffusion model. 2) We use both conditional images and pseudo-labels as inputs to improve cross-domain segmentation performance. Along this idea, we propose a two-path reverse diffusion process during training, utilizing Robust Feature Extraction Subnet and Mask-Generation Transformer to learn a domain-generalized discrete conditional distribution based on robust image features and pseudo-labels. This learned distribution is then used to generate segmentation masks for unseen target domains. Experimental results demonstrate that DG-DDM-Seg achieves state-of-the-art performance in cross-domain medical image segmentation, with domain shifts in modality, sequence, and site. The code is available at https://github.com/HeranYang/DG-DDM-Seg. Heran Yang, Wenbo Hua, Zongben Xu, Jian Sun 0009 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Pose-Transformed Equivariant Network for 3D Point Trajectory PredictionabstractPredicting 3D point trajectory is a fundamental learning task which commonly should be equivariant under Eu-clidean transformation, e.g., SE(3). The existing equivari-ant models are commonly based on the group equivariant convolution, equivariant message passing, vector neuron, frame averaging, etc. In this paper, we propose a novel pose-transformed equivariant network, in which the points are firstly uniquely normalized and then transformed by the learned pose transformations, upon which the points after motion are predicted and aggregated. Under each trans-formed pose, we design the point position predictor consisting of multiple Pose- Transformed Points Prediction blocks, in which the global and local motions are estimated and aggregated. This framework can be proven to be equiv-ariant to SE(3) transformation over 3D points. We eval-uate the pose-transformed equivariant network on exten-sive datasets including human motion capture, molecular dynamics modeling and dynamics simulation. Extensive experimental comparisons demonstrated our SOTA performance compared with the existing equivariant networks for 3D point trajectory prediction. Ruixuan Yu, Jian Sun 0009 |
CVPR | 2 |
| 2024 | Residual-Conditioned Optimal Transport: Towards Structure-Preserving Unpaired and Paired Image RestorationabstractDeep learning-based image restoration methods generally struggle with faithfully preserving the structures of the original image. In this work, we propose a novel Residual-Conditioned Optimal Transport (RCOT) approach, which models image restoration as an optimal transport (OT) problem for both unpaired and paired settings, introducing the transport residual as a unique degradation-specific cue for both the transport cost and the transport map. Specifically, we first formalize a Fourier residual-guided OT objective by incorporating the degradation-specific information of the residual into the transport cost. We further design the transport map as a two-pass RCOT map that comprises a base model and a refinement process, in which the transport residual is computed by the base model in the first pass and then encoded as a degradation-specific embedding to condition the second-pass restoration. By duality, the RCOT problem is transformed into a minimax optimization problem, which can be solved by adversarially training neural networks. Extensive experiments on multiple restoration tasks show that RCOT achieves competitive performance in terms of both distortion measures and perceptual quality, restoring images with more faithful structures as compared with state-of-the-art methods. Xiaole Tang, Xiang Gu 0005, Jian Sun 0009 |
ICML | 4 |
| 2024 | Adversarial data splitting for domain generalization
Xiang Gu 0005, Jian Sun 0009, Zongben Xu |
Sci. China Inf. Sci. | 2 |
| 2024 | Adversarial Reweighting with α-Power Maximization for Domain Adaptation
Xiang Gu 0005, Yan Yang 0007, Jian Sun 0009, Zongben Xu |
Int. J. Comput. Vis. | 4 |
| 2024 | Unsupervised and Semi-Supervised Robust Spherical Space Domain AdaptationabstractAdversarial domain adaptation has been an effective approach for learning domain-invariant features by adversarial training. In this paper, we propose a novel adversarial domain adaptation approach defined in the spherical feature space, in which we define spherical classifier for label prediction and spherical domain discriminator for discriminating domain labels. In the spherical feature space, we develop a spherical robust pseudo-label loss to utilize pseudo-labels robustly, which weights the importance of the estimated labels of target domain data by the posterior probability of correct labeling, modeled by the Gaussian-uniform mixture model in the spherical space. Our proposed approach can be generally applied to both unsupervised and semi-supervised domain adaptation settings. In particular, to tackle the semi-supervised domain adaptation setting where a few labeled target domain data are available for training, we propose a novel reweighted adversarial training strategy for effectively reducing the intra-domain discrepancy within the target domain. We also present theoretical analysis for the proposed method based on the domain adaptation theory. Extensive experiments are conducted on multiple benchmarks for object recognition, digit recognition, and face recognition. The results show that our method either surpasses or is competitive compared with the recent methods for both unsupervised and semi-supervised domain adaptation. Ablation studies also confirm the effectiveness of the spherical classifier, spherical discriminator, spherical robust pseudo-label loss, and reweighted adversarial training strategy. Xiang Gu 0005, Jian Sun 0009, Zongben Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | AMMD: Attentive maximum mean discrepancy for few-shot image classification
Shipeng Wang 0002, Jian Sun 0009 |
Pattern Recognit. | 3 |
| 2024 | Contrasting augmented features for domain adaptation with limited target domain data
Xiang Gu 0005, Jian Sun 0009 |
Pattern Recognit. | 3 |
| 2024 | Scenario-Aware Learning Approaches to Adaptive Channel EstimationabstractThe growth of frequency bandwidths and applications with the forthcoming generations of wireless networks will give rise to a multitude of wireless transmission scenarios, topologies and channel structures. In this work, we go beyond existing learning-based channel estimation methods tailored for specific scenarios, to develop an adaptive learning-based channel state information (CSI) estimation approach. We offer the adaptivity in the learning approach through extracting the scenario embeddings of CSI and adjusting the channel estimation method with the extracted information automatically in each scenario. Specifically, Learning-Based Scenario-Adaptive Channel Estimation Algorithm (LACE) is designed. LACE is based on a Scenario-Aware Hyper-Network (SAH-Net) that incorporates the embedding loss to make the Convolutional Neural Network (CNN) based encoder learn to extract the effective scenario embeddings from the time-space two dimensional features of the CSI. The extracted embeddings are utilized by a Multi-Layer Perceptron (MLP) based tuning module to tune the parameters of the channel estimation method. Our learning design is complemented with analysis to verify that the theoretical performance of LACE is strictly superior to that of the mix-training method, which involves conventionally training the deep network-based channel estimation method using samples from all scenarios. Our results show that the performance of LACE trained in finite scenarios is comparable to that of the deep network-based channel estimation method trained in each scenario, while having lower complexity. Further more, the performance of LACE trained in infinite scenarios is demonstrated to be superior to that of the mix-training method in all test scenarios. Runhua Li, Jian Sun 0009, Jiang Xue 0001, Christos Masouros |
IEEE Trans. Commun. | 2 |
| 2024 | A Deep Learning Approach for Universal NPRACH Detection With Inter-Cell InterferenceabstractThis paper works on the detection of physical random access channel (NPRACH) in Narrowband Internet of Things (NB-IoT) system. The frequency hopping preamble design and increasing number of IoT terminals lead to inter-cell interference among different cells, resulting in inevitable increase of false alarm rate. Due to the ambiguity between preamble and interference, it is a great challenge for NPRACH detection methods to achieve low false alarm rate when having strong interference. In this paper, we analyze the difference between preamble and interference in the propagation environments of NPRACH signals in the 2-dimensional Fast Fourier Transformation (2-D FFT) domain. Then we propose a deep learning-based NPRACH detection method, dubbed Mask Assisted Anti-Interference Universal Detection Scheme (MIUS), in the 2-D FFT domain for preamble detection with inter-cell interference in different repetition cases. In the proposed MIUS, the Mask-ResNet Block is designed as a building block to extract features distinguishing the preamble and interference based on masking operations. Our proposed MIUS utilizes the Mask-ResNet Block in a separate manner to detect the preambles in sequential repetitions across different repetition cases. Simulation results show that MIUS can simultaneously maintain the low false alarm rate and achieve high detection accuracy in low Signal to Interference and Noise Ratio (SINR) regime in all repetition cases. Runhua Li, Jiang Xue 0001, Jian Sun 0009, Symeon Chatzinotas |
IEEE Trans. Commun. | 3 |
| 2024 | Polarimetry-Inspired Contrastive Learning for Class-Imbalanced PolSAR Image ClassificationabstractIn recent years, deep neural networks have significantly boosted the performance of polarimetric synthetic aperture radar (PolSAR) image classification. However, existing deep learning-based approaches still suffer from the following limitations. First, the performance of them is subject to the availability of massive annotations which are difficult to acquire for PolSAR images. Secondly, the class imbalance in PolSAR data greatly hinders the correct classification of minority yet equally pivotal classes. To overcome the above shortcomings, we propose a polarimetry-inspired contrastive learning PolSAR image classification approach, in the hope of elevating the classification accuracy by taking advantage of the polarimetric domain knowledge. Firstly, a complex-valued contrastive learning framework is designed, via which powerful polarimetric representations are learnt without any manual annotations. Specifically, we innovatively design two distribution-inspired positive sample generation strategies, i.e., WishartPSG and NoisePSG, to enable discriminative and domain-specific representation learning. A novel hybrid anti-imbalance scheme is further devised to tackle the class imbalance issue, which combines a contextual consistency-based pseudo-label generation and a weighted feature-level synthetic data over-sampling technique. It should be highlighted that the domain knowledge of PolSAR, including the data and noise distributions, complex-valued characteristics and the spatial consistency prior, is fully exploited throughout our model design. Extensive experiments on four benchmark datasets demonstrated the effectiveness of the proposed model. For the Flevoland 1989 dataset, our method improves the overall accuracy, average accuracy and Kappa metrics by 3.54%, 6.81% and 7.29% respectively, compared to existing state-of-the-art method. Our code will be available at https://github.com/HaixiaBi1982/PiCL. Zuzheng Kuang, Haixia Bi, Fan Li 0003, Jian Sun 0009 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Generalized Semantic Segmentation by Self-Supervised Source Domain Projection and Multi-Level Contrastive LearningabstractDeep networks trained on the source domain show degraded performance when tested on unseen target domain data. To enhance the model's generalization ability, most existing domain generalization methods learn domain invariant features by suppressing domain sensitive features. Different from them, we propose a Domain Projection and Contrastive Learning (DPCL) approach for generalized semantic segmentation, which includes two modules: Self-supervised Source Domain Projection (SSDP) and Multi-Level Contrastive Learning (MLCL). SSDP aims to reduce domain gap by projecting data to the source domain, while MLCL is a learning scheme to learn discriminative and generalizable features on the projected data. During test time, we first project the target data by SSDP to mitigate domain shift, then generate the segmentation results by the learned segmentation network based on MLCL. At test time, we can update the projected data by minimizing our proposed pixel-to-pixel contrastive loss to obtain better results. Extensive experiments for semantic segmentation demonstrate the favorable generalization capability of our method on benchmark datasets. Xiang Gu 0005, Jian Sun 0009 |
AAAI | 3 |
| 2023 | Prototypical Partial Optimal Transport for Universal Domain AdaptationabstractUniversal domain adaptation (UniDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain without requiring the same label sets of both domains. The existence of domain and category shift makes the task challenging and requires us to distinguish “known” samples (i.e., samples whose labels exist in both domains) and “unknown” samples (i.e., samples whose labels exist in only one domain) in both domains before reducing the domain gap. In this paper, we consider the problem from the point of view of distribution matching which we only need to align two distributions partially. A novel approach, dubbed mini-batch Prototypical Partial Optimal Transport (m-PPOT), is proposed to conduct partial distribution alignment for UniDA. In training phase, besides minimizing m-PPOT, we also leverage the transport plan of m-PPOT to reweight source prototypes and target samples, and design reweighted entropy loss and reweighted cross-entropy loss to distinguish “known” and “unknown” samples. Experiments on four benchmarks show that our method outperforms the previous state-of-the-art UniDA methods. Yucheng Yang 0004, Xiang Gu 0005, Jian Sun 0009 |
AAAI | 3 |
| 2023 | Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only ImagesabstractGenerating 3D faces from textual descriptions has a multitude of applications, such as gaming, movie, and robotics. Recent progresses have demonstrated the success of unconditional 3D face generation and text-to-3D shape generation. However, due to the limited text-3D face data pairs, text-driven 3D face generation remains an open problem. In this paper, we propose a text-guided 3D faces generation method, refer as TG-3DFace, for generating realistic 3D faces using text guidance. Specifically, we adopt an unconditional 3D face generation framework and equip it with text conditions, which learns the text-guided 3D face generation with only text-2D face data. On top of that, we propose two text-to-face cross-modal alignment techniques, including the global contrastive learning and the fine-grained alignment module, to facilitate high semantic consistency between generated 3D faces and input texts. Besides, we present directional classifier guidance during the inference process, which encourages creativity for out-of-domain generations. Compared to the existing methods, TG-3DFace creates more realistic and aesthetically pleasing 3D faces, boosting 9% multi-view consistency (MVIC) over Latent3D. The rendered face images generated by TG-3DFace achieve higher FID and CLIP score than text-to-2D face/image generation models, demonstrating our superiority in generating realistic and semantic-consistent textures. Cuican Yu, Guansong Lu, Yihan Zeng, Jian Sun 0009, Xiaodan Liang, Huibin Li 0001, Zongben Xu, Songcen Xu, Wei Zhang 0196, Hang Xu 0004 |
ICCV | 4 |
| 2023 | Optimal Transport-Guided Conditional Score-Based Diffusion ModelabstractConditional score-based diffusion model (SBDM) is for conditional generation of target data with paired data as condition, and has achieved great success in image translation. However, it requires the paired data as condition, and there would be insufficient paired data provided in real-world applications. To tackle the applications with partially paired or even unpaired dataset, we propose a novel Optimal Transport-guided Conditional Score-based diffusion model (OTCS) in this paper. We build the coupling relationship for the unpaired or partially paired dataset based on $L_2$-regularized unsupervised or semi-supervised optimal transport, respectively. Based on the coupling relationship, we develop the objective for training the conditional score-based model for unpaired or partially paired settings, which is based on a reformulation and generalization of the conditional SBDM for paired setting. With the estimated coupling relationship, we effectively train the conditional score-based model by designing a ``resampling-by-compatibility'' strategy to choose the sampled data with high compatibility as guidance. Extensive experiments on unpaired super-resolution and semi-paired image-to-image translation demonstrated the effectiveness of the proposed OTCS model. From the viewpoint of optimal transport, OTCS provides an approach to transport data across distributions, which is a challenge for OT on large-scale datasets. We theoretically prove that OTCS realizes the data transport in OT with a theoretical bound. Xiang Gu 0005, Jian Sun 0009, Zongben Xu |
NeurIPS | 3 |
| 2023 | Constructing Non-isotropic Gaussian Diffusion Model Using Isotropic Gaussian Diffusion Model for Image EditingabstractScore-based diffusion models (SBDMs) have achieved state-of-the-art results in image generation. In this paper, we propose a Non-isotropic Gaussian Diffusion Model (NGDM) for image editing, which requires editing the source image while preserving the image regions irrelevant to the editing task. We construct NGDM by adding independent Gaussian noises with different variances to different image pixels. Instead of specifically training the NGDM, we rectify the NGDM into an isotropic Gaussian diffusion model with different pixels having different total forward diffusion time. We propose to reverse the diffusion by designing a sampling method that starts at different time for different pixels for denoising to generate images using the pre-trained isotropic Gaussian diffusion model. Experimental results show that NGDM achieves state-of-the-art performance for image editing tasks, considering the trade-off between the fidelity to the source image and alignment with the desired editing target. Xiang Gu 0005, Haozhi Liu, Jian Sun 0009 |
NeurIPS | 4 |
| 2023 | Deep expectation-maximization network for unsupervised image segmentation and clustering
Yannan Pu, Jian Sun 0009, Niansheng Tang, Zongben Xu |
Image Vis. Comput. | 2 |
| 2023 | Variational Data-Free Knowledge Distillation for Continual LearningabstractDeep neural networks suffer from catastrophic forgetting when trained on sequential tasks in continual learning. Various methods rely on storing data of previous tasks to mitigate catastrophic forgetting, which is prohibited in real-world applications considering privacy and security issues. In this paper, we consider a realistic setting of continual learning, where training data of previous tasks are unavailable and memory resources are limited. We contribute a novel knowledge distillation-based method in an information-theoretic framework by maximizing mutual information between outputs of previously learned and current networks. Due to the intractability of computation of mutual information, we instead maximize its variational lower bound, where the covariance of variational distribution is modeled by a graph convolutional network. The inaccessibility of data of previous tasks is tackled by Taylor expansion, yielding a novel regularizer in network training loss for continual learning. The regularizer relies on compressed gradients of network parameters. It avoids storing previous task data and previously learned networks. Additionally, we employ self-supervised learning technique for learning effective features, which improves the performance of continual learning. We conduct extensive experiments including image classification and semantic segmentation, and the results show that our method achieves state-of-the-art performance on continual learning benchmarks. Xiaorong Li, Shipeng Wang 0002, Jian Sun 0009, Zongben Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Learning View-Based Graph Convolutional Network for Multi-View 3D Shape AnalysisabstractView-based approach that recognizes 3D shape through its projected 2D images has achieved state-of-the-art results for 3D shape recognition. The major challenges are how to aggregate multi-view features and deal with 3D shapes in arbitrary poses. We propose two versions of a novel view-based Graph Convolutional Network, dubbed view-GCN and view-GCN++, to recognize 3D shape based on graph representation of multiple views. We first construct view-graph with multiple views as graph nodes, then design two graph convolutional networks over the view-graph to hierarchically learn discriminative shape descriptor considering relations of multiple views. Specifically, view-GCN is a hierarchical network based on two pivotal operations, i.e., feature transform based on local positional and non-local graph convolution, and graph coarsening based on a selective view-sampling operation. To deal with rotation sensitivity, we further propose view-GCN++ with local attentional graph convolution operation and rotation robust view-sampling operation for graph coarsening. By these designs, view-GCN++ achieves invariance to transformations under the finite subgroup of rotation group SO(3). Extensive experiments on benchmark datasets (i.e., ModelNet40, ScanObjectNN, RGBD and ShapeNet Core55) show that view-GCN and view-GCN++ achieve state-of-the-art results for 3D shape classification and retrieval tasks under aligned and rotated settings. Ruixuan Yu, Jian Sun 0009 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | A Noising-Denoising Framework for Point Cloud Upsampling via Normalizing Flows
Jian Sun 0009 |
Pattern Recognit. | 3 |
| 2023 | Memory efficient data-free distillation for continual learning
Xiaorong Li, Shipeng Wang 0002, Jian Sun 0009, Zongben Xu |
Pattern Recognit. | 3 |
| 2023 | Meta-learning-based adversarial training for deep 3D face recognition on point clouds
Cuican Yu, Huibin Li 0001, Jian Sun 0009, Zongben Xu |
Pattern Recognit. | 4 |
| 2023 | An Unrolled Implicit Regularization Network for Joint Image and Sensitivity Estimation in Parallel MR Imaging with Convergence GuaranteeabstractAbstract. Parallel imaging (PI), relying on multicoils to sense [Formula: see text]-space data, is an effective technique to accelerate magnetic resonance imaging by exploiting spatial sensitivity coding of multiple coils, with an integrated compressive sensing (CS) technology to achieve higher acceleration. In this paper, we propose a novel nonconvex reconstruction model and its proximal alternating linearized minimization (PALM) algorithm for PI in a blind setting that MR image and multichannel sensitivity maps are jointly estimated, regularized by image and sensitivity regularizers. Instead of hand-crafting the image and sensitivity regularizers, we propose unrolling the PALM algorithm to be a deep network for Blind Parallel MRI, dubbed as BPMRI-Net, with two learnable subnetworks to substitute the proximal operators of the image and sensitivity regularizers. We theoretically prove the linear convergence of BPMRI-Net as an iterative algorithm, which alternately updates two variables based on the learnable proximal operators. The learned BPMRI-Net can simultaneously output the MR image and sensitivity maps from undersampled multichannel [Formula: see text]-space data even when the number of low-frequency sampling lines in the center of [Formula: see text]-space is small. Numerical results demonstrate the effectiveness of our method with state-of-the-art reconstruction accuracy. Yan Yang 0007, Yizhou Wang 0006, Jiazhen Wang, Jian Sun 0009, Zongben Xu |
SIAM J. Imaging Sci. | 4 |
| 2023 | Promoting Monocular Depth Estimation by Multi-Scale Residual Laplacian Pyramid FusionabstractDeep learning approach has achieved great success in monocular depth estimation. However, the learned deep network may produce a depth map with fewer details and incorrect global depth layout, especially when the learned network is applied to a high-resolution image. In order to generate a high-quality depth map with better global structure and richer details, we propose a multi-scale residual Laplacian pyramid fusion net (MS-RLap-FNet), to fuse the multi-scale depth maps estimated by the existing depth estimation models, for depth refinement. Our approach relies on a proposed multi-scale residual Laplacian pyramid decomposition of the multi-scale depth maps, and the fusion network modules to gradually refine the depth maps based on the decomposition from low to high resolution. Comprehensive experiments show that our method, by refining the depth maps based on three popular monocular depth estimation models (DPT, MiDas, SGR), outperforms the existing state-of-the-art methods both in quantity and quality on three public datasets with different image resolutions. The depth map refined by our method has better global depth layout with richer fine details. Anmei Zhang, Yunchao Ma, Jiangyu Liu, Jian Sun 0009 |
IEEE Signal Process. Lett. | 4 |
| 2023 | Learning Unified Hyper-Network for Multi-Modal MR Image Synthesis and Tumor Segmentation With Missing ModalitiesabstractAccurate segmentation of brain tumors is of critical importance in clinical assessment and treatment planning, which requires multiple MR modalities providing complementary information. However, due to practical limits, one or more modalities may be missing in real scenarios. To tackle this problem, existing methods need to train multiple networks or a unified but fixed network for various possible missing modality cases, which leads to high computational burdens or sub-optimal performance. In this paper, we propose a unified and adaptive multi-modal MR image synthesis method, and further apply it to tumor segmentation with missing modalities. Based on the decomposition of multi-modal MR images into common and modality-specific features, we design a shared hyper-encoder for embedding each available modality into the feature space, a graph-attention-based fusion block to aggregate the features of available modalities to the fused features, and a shared hyper-decoder for image reconstruction. We also propose an adversarial common feature constraint to enforce the fused features to be in a common space. As for missing modality segmentation, we first conduct the feature-level and image-level completion using our synthesis method and then segment the tumors based on the completed MR images together with the extracted common features. Moreover, we design a hypernet-based modulation module to adaptively utilize the real and synthetic modalities. Experimental results suggest that our method can not only synthesize reasonable multi-modal MR images, but also achieve state-of-the-art performance on brain tumor segmentation with missing modalities. Heran Yang, Jian Sun 0009, Zongben Xu |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Modality-Adaptive Feature Interaction for Brain Tumor Segmentation with Missing Modalities
Zechen Zhao, Heran Yang, Jian Sun 0009 |
MICCAI (5) | 3 |
| 2022 | Keypoint-Guided Optimal Transport with Applications in Heterogeneous Domain AdaptationabstractExisting Optimal Transport (OT) methods mainly derive the optimal transport plan/matching under the criterion of transport cost/distance minimization, which may cause incorrect matching in some cases. In many applications, annotating a few matched keypoints across domains is reasonable or even effortless in annotation burden. It is valuable to investigate how to leverage the annotated keypoints to guide the correct matching in OT. In this paper, we propose a novel KeyPoint-Guided model by ReLation preservation (KPG-RL) that searches for the matching guided by the keypoints in OT. To impose the keypoints in OT, first, we propose a mask-based constraint of the transport plan that preserves the matching of keypoint pairs. Second, we propose to preserve the relation of each data point to the keypoints to guide the matching. The proposed KPG-RL model can be solved by the Sinkhorn's algorithm and is applicable even when distributions are supported in different spaces. We further utilize the relation preservation constraint in the Kantorovich Problem and Gromov-Wasserstein model to impose the guidance of keypoints in them. Meanwhile, the proposed KPG-RL model is extended to partial OT setting. As an application, we apply the proposed KPG-RL model to the heterogeneous domain adaptation. Experiments verified the effectiveness of the KPG-RL model. Xiang Gu 0005, Yucheng Yang 0004, Jian Sun 0009, Zongben Xu |
NeurIPS | 4 |
| 2022 | Learning Generalizable Part-based Feature Representation for 3D Point CloudsabstractDeep networks on 3D point clouds have achieved remarkable success in 3D classification, while they are vulnerable to geometry variations caused by inconsistent data acquisition procedures. This results in a challenging 3D domain generalization (3DDG) problem, that is to generalize a model trained on source domain to an unseen target domain. Based on the observation that local geometric structures are more generalizable than the whole shape, we propose to reduce the geometry shift by a generalizable part-based feature representation and design a novel part-based domain generalization network (PDG) for 3D point cloud classification. Specifically, we build a part-template feature space shared by source and target domains. Shapes from distinct domains are first organized to part-level features and then represented by part-template features. The transformed part-level features, dubbed aligned part-based representations, are then aggregated by a part-based feature aggregation module. To improve the robustness of the part-based representations, we further propose a contrastive learning framework upon part-based shape representation. Experiments and ablation studies on 3DDA and 3DDG benchmarks justify the efficacy of the proposed approach for domain generalization, compared with the previous state-of-the-art methods. Our code will be available on http://github.com/weixmath/PDG. Xiang Gu 0005, Jian Sun 0009 |
NeurIPS | 3 |
| 2022 | Unlabeled data driven cost-sensitive inverse projection sparse representation-based classification with 1/2 regularization
Jian Sun 0009, Zongben Xu |
Sci. China Inf. Sci. | 3 |
| 2022 | Variational HyperAdam: A Meta-Learning Approach to Network TrainingabstractStochastic optimization algorithms have been popular for training deep neural networks. Recently, there emerges a new approach of learning-based optimizer, which has achieved promising performance for training neural networks. However, these black-box learning-based optimizers do not fully take advantage of the experience in human-designed optimizers and heavily rely on learning from meta-training tasks, therefore have limited generalization ability. In this paper, we propose a novel optimizer, dubbed as Variational HyperAdam, which is based on a parametric generalized Adam algorithm, i.e., HyperAdam, in a variational framework. With Variational HyperAdam as optimizer for training neural network, the parameter update vector of the neural network at each training step is considered as random variable, whose approximate posterior distribution given the training data and current network parameter vector is predicted by Variational HyperAdam. The parameter update vector for network training is sampled from this approximate posterior distribution. Specifically, in Variational HyperAdam, we design a learnable generalized Adam algorithm for estimating expectation, paired with a VarBlock for estimating the variance of the approximate posterior distribution of parameter update vector. The Variational HyperAdam is learned in a meta-learning approach with meta-training loss derived by variational inference. Experiments verify that the learned Variational HyperAdam achieved state-of-the-art network training performance for various types of networks on different datasets, such as multilayer perceptron, CNN, LSTM and ResNet. Shipeng Wang 0002, Yan Yang 0007, Jian Sun 0009, Zongben Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Training Networks in Null Space of Feature Covariance for Continual LearningabstractIn the setting of continual learning, a network is trained on a sequence of tasks, and suffers from catastrophic forgetting. To balance plasticity and stability of network in continual learning, in this paper, we propose a novel network training algorithm Adam-NSCL which sequentially optimizes network parameters in the null space of all previous tasks. We first propose two mathematical conditions respectively for achieving network stability and plasticity in continual learning. Based on them, the network training for sequential tasks without forgetting can be simply achieved by projecting the candidate parameter update into the approximate null space of all previous tasks in the network training process, where the candidate parameter update can be generated by Adam. The approximate null space can be derived by applying singular value decomposition to the un-centered covariance matrix of all input features of previous tasks for each linear layer. For efficiency, the uncentered covariance matrix can be incrementally computed after learning each task. We also empirically verify the rationality of the approximate null space at each linear layer. We apply our approach to training networks for continual learning on benchmark datasets of CIFAR-100 and TinyImageNet, and the results suggest that the proposed approach outperforms or matches the state-ot-the-art continual learning approaches. Shipeng Wang 0002, Xiaorong Li, Jian Sun 0009, Zongben Xu |
CVPR | 3 |
| 2021 | Learning Canonical View Representation for 3D Shape Recognition with Arbitrary ViewsabstractIn this paper, we focus on recognizing 3D shapes from arbitrary views, i.e., arbitrary numbers and positions of viewpoints. It is a challenging and realistic setting for view-based 3D shape recognition. We propose a canonical view representation to tackle this challenge. We first transform the original features of arbitrary views to a fixed number of view features, dubbed canonical view representation, by aligning the arbitrary view features to a set of learnable reference view features using optimal transport. In this way, each 3D shape with arbitrary views is represented by a fixed number of canonical view features, which are further aggregated to generate a rich and robust 3D shape representation for shape recognition. We also propose a canonical view feature separation constraint to enforce that the view features in canonical view representation can be embedded into scattered points in a Euclidean space. Experiments on the ModelNet40, ScanObjectNN, and RGBD datasets show that our method achieves competitive results under the fixed viewpoint settings, and significantly outperforms the applicable methods under the arbitrary view setting. Yifei Gong, Fudong Wang 0001, Xing Sun 0001, Jian Sun 0009 |
ICCV | 5 |
| 2021 | A Unified Hyper-GAN Model for Unpaired Multi-contrast MR Image Translation
Heran Yang, Jian Sun 0009, Zongben Xu |
MICCAI (3) | 2 |
| 2021 | Adversarial Reweighting for Partial Domain AdaptationabstractPartial domain adaptation (PDA) has gained much attention due to its practical setting. The current PDA methods usually adapt the feature extractor by aligning the target and reweighted source domain distributions. In this paper, we experimentally find that the feature adaptation by the reweighted distribution alignment in some state-of-the-art PDA methods is not robust to the ``noisy'' weights of source domain data, leading to negative domain transfer on some challenging benchmarks. To tackle the challenge of negative domain transfer, we propose a novel Adversarial Reweighting (AR) approach that adversarially learns the weights of source domain data to align the source and target domain distributions, and the transferable deep recognition network is learned on the reweighted source domain data. Based on this idea, we propose a training algorithm that alternately updates the parameters of the network and optimizes the weights of source domain data. Extensive experiments show that our method achieves state-of-the-art results on the benchmarks of ImageNet-Caltech, Office-Home, VisDA-2017, and DomainNet. Ablation studies also confirm the effectiveness of our approach. Xiang Gu 0005, Yan Yang 0007, Jian Sun 0009, Zongben Xu |
NeurIPS | 4 |
| 2021 | A distribution independence based method for 3D face shape decomposition
Cuican Yu, Huibin Li 0001, Jian Sun 0009, Zongben Xu |
Comput. Vis. Image Underst. | 4 |
| 2021 | Joint Depth and Defocus Estimation From a Single Image Using Physical ConsistencyabstractEstimating depth and defocus maps are two fundamental tasks in computer vision. Recently, many methods explore these two tasks separately with the help of the powerful feature learning ability of deep learning and these methods have achieved impressive progress. However, due to the difficulty in densely labeling depth and defocus on real images, these methods are mostly based on synthetic training dataset, and the performance of learned network degrades significantly on real images. In this paper, we tackle a new task that jointly estimates depth and defocus from a single image. We design a dual network with two subnets respectively for estimating depth and defocus. The network is jointly trained on synthetic dataset with a physical constraint to enforce the physical consistency between depth and defocus. Moreover, we design a simple method to label depth and defocus order on real image dataset, and design two novel metrics to measure accuracies of depth and defocus estimation on real images. Comprehensive experiments demonstrate that joint training for depth and defocus estimation using physical consistency constraint enables these two subnets to guide each other, and effectively improves their depth and defocus estimation performance on real defocused image dataset. Anmei Zhang, Jian Sun 0009 |
IEEE Trans. Image Process. | 2 |
| 2020 | Learning Distribution Independent Latent Representation for 3D Face DisentanglementabstractLearning disentangled 3D face shape representation is beneficial to face attribute transfer, generation and recognition, etc. In this paper, we propose a novel distribution independence-based method to learn to decompose 3D face shapes. Specifically, we design a variational auto-encoder with Graph Convolutional Network (GCN), namely Mesh-Encoder, to model the distributions of identity and expression representations via variational inference. To disentangle facial expression and identity, we eliminate correlation of the two distributions, and enforce them to be independent by adversarial training. Extensive experiments show that the proposed approach can achieve state-of-the-art results in 3D face shape decomposition and expression transfer. Though focusing on disentanglement, our method also achieves the reconstruction accuracies comparable to the state-of-the-art 3D face reconstruction methods. Cuican Yu, Huibin Li 0001, Jian Sun 0009 |
3DV | 4 |
| 2020 | Spherical Space Domain Adaptation With Robust Pseudo-Label LossabstractAdversarial domain adaptation (DA) has been an effective approach for learning domain-invariant features by adversarial training. In this paper, we propose a novel adversarial DA approach completely defined in spherical feature space, in which we define spherical classifier for label prediction and spherical domain discriminator for discriminating domain labels. To utilize pseudo-label robustly, we develop a robust pseudo-label loss in the spherical feature space, which weights the importance of estimated labels of target data by posterior probability of correct labeling, modeled by Gaussian-uniform mixture model in spherical feature space. Extensive experiments show that our method achieves state-of-the-art results, and also confirm effectiveness of spherical classifier, spherical discriminator and spherical robust pseudo-label loss. Xiang Gu 0005, Jian Sun 0009, Zongben Xu |
CVPR | 2 |
| 2020 | View-GCN: View-Based Graph Convolutional Network for 3D Shape AnalysisabstractView-based approach that recognizes 3D shape through its projected 2D images has achieved state-of-the-art results for 3D shape recognition. The major challenge for view-based approach is how to aggregate multi-view features to be a global shape descriptor. In this work, we propose a novel view-based Graph Convolutional Neural Network, dubbed as view-GCN, to recognize 3D shape based on graph representation of multiple views in flexible view configurations. We first construct view-graph with multiple views as graph nodes, then design a graph convolutional neural network over view-graph to hierarchically learn discriminative shape descriptor considering relations of multiple views. The view-GCN is a hierarchical network based on local and non-local graph convolution for feature transform, and selective view-sampling for graph coarsening. Extensive experiments on benchmark datasets show that view-GCN achieves state-of-the-art results for 3D shape classification and retrieval. Ruixuan Yu, Jian Sun 0009 |
CVPR | 3 |
| 2020 | End-to-end Interpretable Learning of Non-blind Image Deblurring
Thomas Eboli, Jian Sun 0009, Jean Ponce |
ECCV (17) | 2 |
| 2020 | Deep Positional and Relational Feature Learning for Rotation-Invariant Point Cloud Analysis
Ruixuan Yu, Federico Tombari, Jian Sun 0009 |
ECCV (10) | 4 |
| 2020 | Model-Driven Deep Attention Network for Ultra-fast Compressive Sensing MRI Guided by Cross-contrast MR Image
Yan Yang 0007, Heran Yang, Jian Sun 0009, Zongben Xu |
MICCAI (2) | 4 |
| 2020 | ADMM-CSNet: A Deep Learning Approach for Image Compressive SensingabstractCompressive sensing (CS) is an effective technique for reconstructing image from a small amount of sampled data. It has been widely applied in medical imaging, remote sensing, image compression, etc. In this paper, we propose two versions of a novel deep learning architecture, dubbed as ADMM-CSNet, by combining the traditional model-based CS method and data-driven deep learning method for image reconstruction from sparsely sampled measurements. We first consider a generalized CS model for image reconstruction with undetermined regularizations in undetermined transform domains, and then two efficient solvers using Alternating Direction Method of Multipliers (ADMM) algorithm for optimizing the model are proposed. We further unroll and generalize the ADMM algorithm to be two deep architectures, in which all parameters of the CS model and the ADMM algorithm are discriminatively learned by end-to-end training. For both applications of fast CS complex-valued MR imaging and CS imaging of real-valued natural images, the proposed ADMM-CSNet achieved favorable reconstruction accuracy in fast computational speed compared with the traditional and the other deep learning methods. Yan Yang 0007, Jian Sun 0009, Huibin Li 0001, Zongben Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Abnormality detection in retinal image by individualized background learning
Benzhi Chen, Lisheng Wang, Xiuying Wang 0001, Jian Sun 0009, David Dagan Feng, Zongben Xu |
Pattern Recognit. | 4 |
| 2020 | Second-Order Spectral Transform Block for 3D Shape Classification and RetrievalabstractIn this paper, we propose a novel network block, dubbed as second-order spectral transform block, for 3D shape retrieval and classification. This network block generalizes the second-order pooling to 3D surface by designing a learnable non-linear transform on the spectrum of the pooled descriptor. The proposed block consists of following two components. First, the second-order average (SO-Avr) and max-pooling (SOMax) operations are designed on 3D surface to aggregate local descriptors, which are shown to be more discriminative than the popular average-pooling or max-pooling. Second, a learnable spectral transform parameterized by mixture of power function is proposed to perform non-linear feature mapping in the space of pooled descriptors, i.e., manifold of symmetric positive definite matrix for SO-Avr, and space of symmetric matrix for SOMax. The proposed block can be plugged into existing network architectures to aggregate local shape descriptors for boosting their performance. We apply it to a shallow network for nonrigid 3D shape analysis and to existing networks for rigid shape analysis, where it improves the first-tier retrieval accuracy by 7.2% on SHREC'14 Real dataset and achieves state-of-the-art classification accuracy on ModelNet40. As an extension, we apply our block to 2D image classification, showing its superiority compared with traditional second-order pooling methods. We also provide theoretical and experimental analysis on stability of the proposed second-order spectral transform block. Ruixuan Yu, Jian Sun 0009, Huibin Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Unsupervised MR-to-CT Synthesis Using Structure-Constrained CycleGANabstractSynthesizing a CT image from an available MR image has recently emerged as a key goal in radiotherapy treatment planning for cancer patients. CycleGANs have achieved promising results on unsupervised MR-to-CT image synthesis; however, because they have no direct constraints between input and synthetic images, cycleGANs do not guarantee structural consistency between these two images. This means that anatomical geometry can be shifted in the synthetic CT images, clearly a highly undesirable outcome in the given application. In this paper, we propose a structure-constrained cycleGAN for unsupervised MR-to-CT synthesis by defining an extra structure-consistency loss based on the modality independent neighborhood descriptor. We also utilize a spectral normalization technique to stabilize the training process and a self-attention module to model the long-range spatial dependencies in the synthetic images. Results on unpaired brain and abdomen MR-to-CT image synthesis show that our method produces better synthetic CT images in both accuracy and visual quality as compared to other unsupervised synthesis methods. We also show that an approximate affine pre-registration for unpaired training data can improve synthesis results. Heran Yang, Jian Sun 0009, Aaron Carass, Can Zhao 0001, Jerry L. Prince, Zongben Xu |
IEEE Trans. Medical Imaging | 2 |
| 2019 | HyperAdam: A Learnable Task-Adaptive Adam for Network TrainingabstractDeep neural networks are traditionally trained using humandesigned stochastic optimization algorithms, such as SGD and Adam. Recently, the approach of learning to optimize network parameters has emerged as a promising research topic. However, these learned black-box optimizers sometimes do not fully utilize the experience in human-designed optimizers, therefore have limitation in generalization ability. In this paper, a new optimizer, dubbed as HyperAdam, is proposed that combines the idea of “learning to optimize” and traditional Adam optimizer. Given a network for training, its parameter update in each iteration generated by HyperAdam is an adaptive combination of multiple updates generated by Adam with varying decay rates . The combination weights and decay rates in HyperAdam are adaptively learned depending on the task. HyperAdam is modeled as a recurrent neural network with AdamCell, WeightCell and StateCell. It is justified to be state-of-the-art for various network training, such as multilayer perceptron, CNN and LSTM. Shipeng Wang 0002, Jian Sun 0009, Zongben Xu |
AAAI | 2 |
| 2019 | A Prior Learning Network for Joint Image and Sensitivity Estimation in Parallel MR Imaging
Nan Meng, Yan Yang 0007, Zongben Xu, Jian Sun 0009 |
MICCAI (4) | 4 |
| 2019 | Neural Diffusion Distance for Image SegmentationabstractDiffusion distance is a spectral method for measuring distance among nodes on graph considering global data structure. In this work, we propose a spec-diff-net for computing diffusion distance on graph based on approximate spectral decomposition. The network is a differentiable deep architecture consisting of feature extraction and diffusion distance modules for computing diffusion distance on image by end-to-end training. We design low resolution kernel matching loss and high resolution segment matching loss to enforce the network's output to be consistent with human-labeled image segments. To compute high-resolution diffusion distance or segmentation mask, we design an up-sampling strategy by feature-attentional interpolation which can be learned when training spec-diff-net. With the learned diffusion distance, we propose a hierarchical image segmentation method outperforming previous segmentation methods. Moreover, a weakly supervised semantic segmentation network is designed using diffusion distance and achieved promising results on PASCAL VOC 2012 segmentation dataset. Jian Sun 0009, Zongben Xu |
NeurIPS | 1 |
| 2019 | A Graph-Based Semisupervised Deep Learning Model for PolSAR Image ClassificationabstractAiming at improving the classification accuracy with limited numbers of labeled pixels in polarimetric synthetic aperture radar (PolSAR) image classification task, this paper presents a graph-based semisupervised deep learning model for PolSAR image classification. It models the PolSAR image as an undirected graph, where the nodes correspond to the labeled and unlabeled pixels, and the weighted edges represent similarities between the pixels. Upon the graph, we design an energy function incorporating a semisupervision term, a convolutional neural network (CNN) term, and a pairwise smoothness term. The employed CNN extracts abstract and data-driven polarimetric features and outputs class label predictions to the graph model. The semisupervision term enforces the category label constraints on the human-labeled pixels. The pairwise smoothness term encourages class label smoothness and the alignment of class label boundaries with the image edges. Starting from an initialized class label map generated based on K-Wishart distribution hypothesis or superpixel segmentation of PauliRGB images, we iteratively and alternately optimize the defined energy function until it converges. We conducted experiments on two real benchmark PolSAR images, and extensive experiments demonstrated that our approach achieved the state-of-the-art results for PolSAR image classification. Haixia Bi, Jian Sun 0009, Zongben Xu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Optimizing a Parameterized Plug-and-Play ADMM for Iterative Low-Dose CT ReconstructionabstractReducing the exposure to X-ray radiation while maintaining a clinically acceptable image quality is desirable in various CT applications. To realize low-dose CT (LdCT) imaging, model-based iterative reconstruction (MBIR) algorithms are widely adopted, but they require proper prior knowledge assumptions in the sinogram and/or image domains and involve tedious manual optimization of multiple parameters. In this paper, we propose a deep learning (DL)-based strategy for MBIR to simultaneously address prior knowledge design and MBIR parameter selection in one optimization framework. Specifically, a parameterized plug-and-play alternating direction method of multipliers (3pADMM) is proposed for the general penalized weighted least-squares model, and then, by adopting the basic idea of DL, the parameterized plug-and-play (3p) prior and the related parameters are optimized simultaneously in a single framework using a large number of training data. The main contribution of this paper is that the 3p prior and the related parameters in the proposed 3pADMM framework can be supervised and optimized simultaneously to achieve robust LdCT reconstruction performance. Experimental results obtained on clinical patient datasets demonstrate that the proposed method can achieve promising gains over existing algorithms for LdCT image reconstruction in terms of noise-induced artifact suppression and edge detail preservation. Ji He 0001, Yan Yang 0007, Dong Zeng, Zhaoying Bian, Hao Zhang 0026, Jian Sun 0009, Zongben Xu, Jianhua Ma 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2018 | Proximal Dehaze-Net: A Prior Learning-Based Deep Network for Single Image Dehazing
Jian Sun 0009 |
ECCV (7) | 2 |
| 2018 | Unsupervised Domain Adaptation with Regularized Optimal Transport for Multimodal 2D+3D Facial Expression RecognitionabstractSince human expressions have strong flexibility and personality, subject-independent facial expression recognition is a typical data bias problem. To address this problem, we propose a novel approach, namely unsupervised domain adaptation with regularized optimal transport for multimodal 2D+3D Facial Expression Recognition (FER). In particular, Wasserstein distance is employed to measure the distribution inconsistency between the training samples (i.e. source domain) and test samples (i.e. target domain). Minimization of this Wasserstein distance is equivalent to finding an optimal transport mapping from training to test samples. Once we find this mapping, original training samples can be transformed into a new space in which the distributions of the mapped training samples and the test samples can be well-aligned. In this case, classifier learned from the transformed training samples can be well generalized to the test samples for expression prediction. In practice, approximate optimal transport can be effectively solved by adding entropy regularization. To fully explore the class label information of training samples, group sparsity regularizer is also used to enforce that the training samples from the same expression class can be mapped to the same group. Experimental results evaluated on the BU-3DFE and Bosphorus databases demonstrate that the proposed approach can achieve superior performance compared with the state-of-the-art methods. Xiaofan Wei, Huibin Li 0001, Jian Sun 0009, Liming Chen 0002 |
FG | 3 |
| 2018 | Surface reconstruction from unorganized points with l0 gradient minimization
Huibin Li 0001, Yibao Li, Ruixuan Yu, Jian Sun 0009, Junseok Kim 0004 |
Comput. Vis. Image Underst. | 4 |
| 2018 | Diverse lesion detection from retinal images by subspace learning over normal samples
Benzhi Chen, Lisheng Wang, Jian Sun 0009, Huai Chen, Yinghua Fu, Shouren Lan, Zongben Xu |
Neurocomputing | 3 |
| 2018 | Neural multi-atlas label fusion: Application to cardiac MR images
Heran Yang, Jian Sun 0009, Huibin Li 0001, Lisheng Wang, Zongben Xu |
Medical Image Anal. | 2 |
| 2018 | A tensor-based nonlocal total variation model for multi-channel image recovery
Wenfei Cao, Jing Yao 0002, Jian Sun 0009 |
Signal Process. | 3 |
| 2018 | BM3D-Net: A Convolutional Neural Network for Transform-Domain Collaborative FilteringabstractDenoising is a fundamental task in image processing with wide applications for enhancing image qualities. BM3D is considered as an effective baseline for image denoising. Although learning-based methods have been dominant in this area recently, the traditional methods are still valuable to inspire new ideas by combining with learning-based approaches. In this letter, we propose a new convolutional neural network inspired by the classical BM3D algorithm, dubbed as BM3D-Net. We unroll the computational pipeline of BM3D algorithm into a convolutional neural network structure, with “extraction” and “aggregation” layers to model block matching stage in BM3D. We apply our network to three denoising tasks: gray-scale image denoising, color image denoising, and depth map denoising. Experiments show that BM3D-Net significantly outperforms the basic BM3D method, and achieves competitive results compared with state of the art on these tasks. Dong Yang 0006, Jian Sun 0009 |
IEEE Signal Process. Lett. | 2 |
| 2017 | Location-sensitive sparse representation of deep normal patterns for expression-robust 3D face recognitionabstractThis paper presents a straight-forward yet efficient, and expression-robust 3D face recognition approach by exploring location sensitive sparse representation of deep normal patterns (DNP). In particular, given raw 3D facial surfaces, we first run 3D face pre-processing pipeline, including nose tip detection, face region cropping, and pose normalization. The 3D coordinates of each normalized 3D facial surface are then projected into 2D plane to generate geometry images, from which three images of facial surface normal components are estimated. Each normal image is then fed into a pre-trained deep face net to generate deep representations of facial surface normals, i.e., deep normal patterns. Considering the importance of different facial locations, we propose a location sensitive sparse representation classifier (LS-SRC) for similarity measure among deep normal patterns associated with different 3D faces. Finally, simple score-level fusion of different normal components are used for the final decision. The proposed approach achieves significantly high performance, and reporting rank-one scores of 98.01%, 97.60%, and 96.13% on the FRGC v2.0, Bosphorus, and BU-3DFE databases when only one sample per subject is used in the gallery. These experimental results reveals that the performance of 3D face recognition would be constantly improved with the aid of training deep models from massive 2D face images, which opens the door for future directions of 3D face recognition. Huibin Li 0001, Jian Sun 0009, Liming Chen 0002 |
IJCB | 2 |
| 2017 | Unsupervised PolSAR Image Classification Using Discriminative ClusteringabstractThis paper presents a novel unsupervised image classification method for polarimetric synthetic aperture radar (PolSAR) data. The proposed method is based on a discriminative clustering framework that explicitly relies on a discriminative supervised classification technique to perform unsupervised clustering. To implement this idea, we design an energy function for unsupervised PolSAR image classification by combining a supervised softmax regression model with a Markov random field smoothness constraint. In this model, both the pixelwise class labels and classifiers are taken as unknown variables to be optimized. Starting from the initialized class labels generated by Cloude-Pottier decomposition and $K$ -Wishart distribution hypothesis, we iteratively optimize the classifiers and class labels by alternately minimizing the energy function with respect to them. Finally, the optimized class labels are taken as the classification result, and the classifiers for different classes are also derived as a side effect. We apply this approach to real PolSAR benchmark data. Extensive experiments justify that our approach can effectively classify the PolSAR image in an unsupervised way and produce higher accuracies than the compared state-of-the-art methods. Haixia Bi, Jian Sun 0009, Zongben Xu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Multimodal 2D+3D Facial Expression Recognition With Deep Fusion Convolutional Neural NetworkabstractThis paper presents a novel and efficient deep fusion convolutional neural network (DF-CNN) for multimodal 2D+3D facial expression recognition (FER). DF-CNN comprises a feature extraction subnet, a feature fusion subnet, and a softmax layer. In particular, each textured three-dimensional (3D) face scan is represented as six types of 2D facial attribute maps (i.e., geometry map, three normal maps, curvature map, and texture map), all of which are jointly fed into DF-CNN for feature learning and fusion learning, resulting in a highly concentrated facial representation (32-dimensional). Expression prediction is performed by two ways: 1) learning linear support vector machine classifiers using the 32-dimensional fused deep features, or 2) directly performing softmax prediction using the six-dimensional expression probability vectors. Different from existing 3D FER methods, DF-CNN combines feature learning and fusion learning into a single end-to-end training framework. To demonstrate the effectiveness of DF-CNN, we conducted comprehensive experiments to compare the performance of DFCNN with handcrafted features, pre-trained deep features, finetuned deep features, and state-of-the-art methods on three 3D face datasets (i.e., BU-3DFE Subset I, BU-3DFE Subset II, and Bosphorus Subset). In all cases, DF-CNN consistently achieved the best results. To the best of our knowledge, this is the first work of introducing deep CNN to 3D FER and deep learning-based featurelevel fusion for multimodal 2D+3D FER. Huibin Li 0001, Jian Sun 0009, Zongben Xu, Liming Chen 0002 |
IEEE Trans. Multim. | 2 |
| 2016 | Deep Fusion Net for Multi-atlas Segmentation: Application to Cardiac MR Images
Heran Yang, Jian Sun 0009, Huibin Li 0001, Lisheng Wang, Zongben Xu |
MICCAI (2) | 2 |
| 2016 | Deep ADMM-Net for Compressive Sensing MRIabstractCompressive Sensing (CS) is an effective approach for fast Magnetic Resonance Imaging (MRI). It aims at reconstructing MR image from a small number of under-sampled data in k-space, and accelerating the data acquisition in MRI. To improve the current MRI system in reconstruction accuracy and computational speed, in this paper, we propose a novel deep architecture, dubbed ADMM-Net. ADMM-Net is defined over a data flow graph, which is derived from the iterative procedures in Alternating Direction Method of Multipliers (ADMM) algorithm for optimizing a CS-based MRI model. In the training phase, all parameters of the net, e.g., image transforms, shrinkage functions, etc., are discriminatively trained end-to-end using L-BFGS algorithm. In the testing phase, it has computational overhead similar to ADMM but uses optimized parameters learned from the training data for CS-based reconstruction task. Experiments on MRI image reconstruction under different sampling ratios in k-space demonstrate that it significantly improves the baseline ADMM algorithm and achieves high reconstruction accuracies with fast computational speed. Yan Yang 0007, Jian Sun 0009, Huibin Li 0001, Zongben Xu |
NIPS | 2 |
| 2016 | Learning Dictionary of Discriminative Part Detectors for Image Categorization and Cosegmentation
Jian Sun 0009, Jean Ponce |
Int. J. Comput. Vis. | 1 |
| 2016 | Total Variation Regularized Tensor RPCA for Background Subtraction From Compressive MeasurementsabstractBackground subtraction has been a fundamental and widely studied task in video analysis, with a wide range of applications in video surveillance, teleconferencing, and 3D modeling. Recently, motivated by compressive imaging, background subtraction from compressive measurements (BSCM) is becoming an active research task in video surveillance. In this paper, we propose a novel tensor-based robust principal component analysis (TenRPCA) approach for BSCM by decomposing video frames into backgrounds with spatial-temporal correlations and foregrounds with spatio-temporal continuity in a tensor framework. In this approach, we use 3D total variation to enhance the spatio-temporal continuity of foregrounds, and Tucker decomposition to model the spatio-temporal correlations of video background. Based on this idea, we design a basic tensor RPCA model over the video frames, dubbed as the holistic TenRPCA model. To characterize the correlations among the groups of similar 3D patches of video background, we further design a patch-group-based tensor RPCA model by joint tensor Tucker decompositions of 3D patch groups for modeling the video background. Efficient algorithms using the alternating direction method of multipliers are developed to solve the proposed models. Extensive experiments on simulated and real-world videos demonstrate the superiority of the proposed approaches over the existing state-of-the-art approaches. Wenfei Cao, Yao Wang 0003, Jian Sun 0009, Deyu Meng, Can Yang 0002, Andrzej Cichocki, Zongben Xu |
IEEE Trans. Image Process. | 3 |
| 2015 | Learning a convolutional neural network for non-uniform motion blur removalabstractIn this paper, we address the problem of estimating and removing non-uniform motion blur from a single blurry image. We propose a deep learning approach to predicting the probabilistic distribution of motion blur at the patch level using a convolutional neural network (CNN). We further extend the candidate set of motion kernels predicted by the CNN using carefully designed image rotations. A Markov random field model is then used to infer a dense non-uniform motion blur field enforcing motion smoothness. Finally, motion blur is removed by a non-uniform deblurring model using patch-level image prior. Experimental evaluations show that our approach can effectively estimate and remove complex non-uniform motion blur that is not handled well by previous approaches. Jian Sun 0009, Wenfei Cao, Zongben Xu, Jean Ponce |
CVPR | 1 |
| 2015 | Color Image Denoising via Discriminatively Learned Iterative ShrinkageabstractIn this paper, we propose a novel model, a discriminatively learned iterative shrinkage (DLIS) model, for color image denoising. The DLIS is a generalization of wavelet shrinkage by iteratively performing shrinkage over patch groups and whole image aggregation. We discriminatively learn the shrinkage functions and basis from the training pairs of noisy/noise-free images, which can adaptively handle different noise characteristics in luminance/chrominance channels, and the unknown structured noise in real-captured color images. Furthermore, to remove the splotchy real color noises, we design a Laplacian pyramid-based denoising framework to progressively recover the clean image from the coarsest scale to the finest scale by the DLIS model learned from the real color noises. Experiments show that our proposed approach can achieve the state-of-the-art denoising results on both synthetic denoising benchmark and real-captured color images. Jian Sun 0009, Jian Sun 0001, Zongben Xu |
IEEE Trans. Image Process. | 1 |
| 2014 | Finding Matches in a Haystack: A Max-Pooling Strategy for Graph Matching in the Presence of OutliersabstractA major challenge in real-world feature matching problems is to tolerate the numerous outliers arising in typical visual tasks. Variations in object appearance, shape, and structure within the same object class make it harder to distinguish inliers from outliers due to clutters. In this paper, we propose a max-pooling approach to graph matching, which is not only resilient to deformations but also remarkably tolerant to outliers. The proposed algorithm evaluates each candidate match using its most promising neighbors, and gradually propagates the corresponding scores to update the neighbors. As final output, it assigns a reliable score to each match together with its supporting neighbors, thus providing contextual information for further verification. We demonstrate the robustness and utility of our method with synthetic and real image experiments. Minsu Cho, Jian Sun 0009, Olivier Duchenne, Jean Ponce |
CVPR | 2 |
| 2013 | Learning to Estimate and Remove Non-uniform Image BlurabstractThis paper addresses the problem of restoring images subjected to unknown and spatially varying blur caused by defocus or linear (say, horizontal) motion. The estimation of the global (non-uniform) image blur is cast as a multi-label energy minimization problem. The energy is the sum of unary terms corresponding to learned local blur estimators, and binary ones corresponding to blur smoothness. Its global minimum is found using Ishikawa's method by exploiting the natural order of discretized blur values for linear motions and defocus. Once the blur has been estimated, the image is restored using a robust (non-uniform) deblurring algorithm based on sparse regularization with global image statistics. The proposed algorithm outputs both a segmentation of the image into uniform-blur layers and an estimate of the corresponding sharp image. We present qualitative results on real images, and use synthetic data to quantitatively compare our approach to the publicly available implementation of Chakrabarti~et al. Florent Couzinie-Devy, Jian Sun 0009, Karteek Alahari, Jean Ponce |
CVPR | 2 |
| 2013 | Learning Discriminative Part Detectors for Image Classification and CosegmentationabstractIn this paper, we address the problem of learning discriminative part detectors from image sets with category labels. We propose a novel latent SVM model regularized by group sparsity to learn these part detectors. Starting from a large set of initial parts, the group sparsity regularizer forces the model to jointly select and optimize a set of discriminative part detectors in a max-margin framework. We propose a stochastic version of a proximal algorithm to solve the corresponding optimization problem. We apply the proposed method to image classification and co segmentation, and quantitative experiments with standard benchmarks show that it matches or improves upon the state of the art. Jian Sun 0009, Jean Ponce |
ICCV | 1 |
| 2013 | Fast image deconvolution using closed-form thresholding formulas of regularization
Wenfei Cao, Jian Sun 0009, Zongben Xu |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Separable Markov Random Field Model and Its Applications in Low Level VisionabstractThis brief proposes a continuously-valued Markov random field (MRF) model with separable filter bank, denoted as MRFSepa, which significantly reduces the computational complexity in the MRF modeling. In this framework, we design a novel gradient-based discriminative learning method to learn the potential functions and separable filter banks. We learn MRFSepa models with 2-D and 3-D separable filter banks for the applications of gray-scale/color image denoising and color image demosaicing. By implementing MRFSepa model on graphics processing unit, we achieve real-time image denoising and fast image demosaicing with high-quality results. Jian Sun 0009, Marshall F. Tappen |
IEEE Trans. Image Process. | 1 |
| 2012 | Unified cardinalized probability hypothesis density filters for extended targets and unresolved targets
Feng Lian, Chongzhao Han, Weifeng Liu 0004, Jing Liu 0011, Jian Sun 0009 |
Signal Process. | 5 |
| 2011 | Learning non-local range Markov Random field for image restorationabstractIn this paper, we design a novel MRF framework which is called Non-Local Range Markov Random Field (NLR-MRF). The local spatial range of clique in traditional MRF is extended to the non-local range which is defined over the local patch and also its similar patches in a non-local window. Then the traditional local spatial filter is extended to the non-local range filter that convolves an image over the non-local ranges of pixels. In this framework, we propose a gradient-based discriminative learning method to learn the potential functions and non-local range filter bank. As the gradients of loss function with respect to model parameters are explicitly computed, efficient gradient-based optimization methods are utilized to train the proposed model. We implement this framework for image denoising and in-painting, the results show that the learned NLR-MRF model significantly outperforms the traditional MRF models and produces state-of-the-art results. Jian Sun 0009, Marshall F. Tappen |
CVPR | 1 |
| 2011 | Gradient Profile Prior and Its Applications in Image Super-Resolution and EnhancementabstractIn this paper, we propose a novel generic image prior-gradient profile prior, which implies the prior knowledge of natural image gradients. In this prior, the image gradients are represented by gradient profiles, which are 1-D profiles of gradient magnitudes perpendicular to image structures. We model the gradient profiles by a parametric gradient profile model. Using this model, the prior knowledge of the gradient profiles are learned from a large collection of natural images, which are called gradient profile prior. Based on this prior, we propose a gradient field transformation to constrain the gradient fields of the high resolution image and the enhanced image when performing single image super-resolution and sharpness enhancement. With this simple but very effective approach, we are able to produce state-of-the-art results. The reconstructed high resolution images or the enhanced images are sharp while have rare ringing or jaggy artifacts. Jian Sun 0009, Jian Sun 0001, Zongben Xu, Harry Shum |
IEEE Trans. Image Process. | 1 |
| 2010 | Context-constrained hallucination for image super-resolutionabstractThis paper proposes a context-constrained hallucination approach for image super-resolution. Through building a training set of high-resolution/low-resolution image segment pairs, the high-resolution pixel is hallucinated from its texturally similar segments which are retrieved from the training set by texture similarity. Given the discrete hallucinated examples, a continuous energy function is designed to enforce the fidelity of high-resolution image to low-resolution input and the constraints imposed by the hallucinated examples and the edge smoothness prior. The reconstructed high-resolution image is sharp with minimal artifacts both along the edges and in the textural regions. Jian Sun 0009, Jiejie Zhu, Marshall F. Tappen |
CVPR | 1 |
| 2010 | Scale selection for anisotropic diffusion filter by Markov random field model
Jian Sun 0009, Zongben Xu |
Pattern Recognit. | 1 |
| 2010 | Image Inpainting by Patch Propagation Using Patch SparsityabstractThis paper introduces a novel examplar-based inpainting algorithm through investigating the sparsity of natural image patches. Two novel concepts of sparsity at the patch level are proposed for modeling the patch priority and patch representation, which are two crucial steps for patch propagation in the examplar-based inpainting approach. First, patch structure sparsity is designed to measure the confidence of a patch located at the image structure (e.g., the edge or corner) by the sparseness of its nonzero similarities to the neighboring patches. The patch with larger structure sparsity will be assigned higher priority for further inpainting. Second, it is assumed that the patch to be filled can be represented by the sparse linear combination of candidate patches under the local patch consistency constraint in a framework of sparse representation. Compared with the traditional examplar-based inpainting approach, structure sparsity enables better discrimination of structure and texture, and the patch sparse representation forces the newly inpainted regions to be sharp and consistent with the surrounding textures. Experiments on synthetic and natural images show the advantages of the proposed approach. Zongben Xu, Jian Sun 0009 |
IEEE Trans. Image Process. | 2 |
| 2008 | Image super-resolution using gradient profile priorabstractIn this paper, we propose an image super-resolution approach using a novel generic image prior - gradient profile prior, which is a parametric prior describing the shape and the sharpness of the image gradients. Using the gradient profile prior learned from a large number of natural images, we can provide a constraint on image gradients when we estimate a hi-resolution image from a low-resolution image. With this simple but very effective prior, we are able to produce state-of-the-art results. The reconstructed hi-resolution image is sharp while has rare ringing or jaggy artifacts. Jian Sun 0009, Zongben Xu, Harry Shum |
CVPR | 1 |
| 2007 | Flash Cut: Foreground Extraction with Flash and No-flash Image PairsabstractIn this paper, we propose a novel approach for foreground layer extraction using flash/no-flash image pairs, which we call flash cut. Flash cut is based on the simple observation that only the foreground is significantly brightened by the flash and the background appearance change is very small, if the background is distant. Changes due to flash, motion, and color information are fused in an MRF framework to produce high quality segmentation results. Flash cut handles some amount of camera shake, and foreground motion, which makes it practical for anyone with a flash-equipped camera to use. We validate our approach on a variety of indoor and outdoor examples. Jian Sun 0009, Jian Sun 0001, Sing Bing Kang, Zongben Xu, Xiaoou Tang, Harry Shum |
CVPR | 1 |
| 2006 | An Edge Preserving Regularization Model for Image Restoration Based on Hopfield Neural Network
Jian Sun 0009, Zongben Xu |
ISNN (2) | 1 |