Jingyuan Xia

dblp:228/7776 · DBLP profile ↗
← Back
32ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0003-4329-0354ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Dynamic Semantic Tokenization for Time Series via Elastic Sampling on Physics-aware Perception
abstract
Despite the remarkable success of semantic token learning in NLP and vision domains, token-level representation mechanisms face fundamental challenges when extended to continuous time series analysis. We identify a core limitation lies in the intrinsic absence of semantically meaningful tokenization boundaries within time-series, which differs substantially from discrete text tokens and presents unique complexities compared to spatially coherent image patches. While existing works mechanically apply fixed-length partitioning, recent evidence from time series foundation models reveals performance ceilings in prediction tasks under such paradigms. This paper introduces a novel tokenization framework known as physics-aware tokenization (PATK), designed to implement adaptive time-frequency tokenization via distribution-sensitive sampling strategies. Key innovations include: 1) A Rate-of-Variation (RoV) distribution is meticulously structured to encompass multi-scale temporal dynamics in the time domain, alongside a Spectral Energy Intensity (SEI) distribution devised to reveal global seasonal patterns within the frequency domain; 2) A physics-aware hidden Markov modeling (PA-HMM) is then established to adaptively breaks down continuous time-series into distinct tokens with elastic lengths, responding to physics-aware probabilities sampled from RoV and SEI distributions. The proposed PATK allows steady integration with both conventional Transformers and advanced large-scale time series models (including LLM-transferred methods and pretrained time series foundation models). Simulations across various datasets demonstrate that PATK excels in classification and forecasting tasks, showing notable adaptability to model long-term dependencies, strengthening resilience against disturbances, and robustness to missing data events.
Huaizhang Liao, Zhixiong Yang 0001, Jingyuan Xia, Yuheng Sun, Yue Zhang 0082, Shengxi Li, Yongxiang Liu
AAAI3
2026 Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution
Chen Wu 0006, Zhuoran Zheng, Jingyuan Xia, Weidong Jiang
ISCAS5
2026 Band-Kernel Stochastic Learning for Unsupervised Blind Hyperspectral Image Super-Resolution
abstract
Hyperspectral image super-resolution (HSI-SR) is fundamentally more difficult than RGB image SR, since its ultrahigh spectral dimensionality. Existing supervised methods rely on labeled training data to obtain data prior, which incurs prohibitive collection costs and limits generalization. Unsupervised methods individually preset the band and kernel with handcrafted priors, whereas this decoupling modeling artificially creates a complexity-performance trade-off in the selected band number. To address these issues, we propose BKX-HMM, a unified statistical framework for blind HSI-SR, which uniformly models the band selection, kernel estimation, and HSI restoration through the state transition of a hidden Markov model (HMM). BKX-HMM redefines the trade-off as a distributional fitting problem: each Markov transition progressively learns optimal parameters of full-band distribution via limited spectral observations. Based on BKX-HMM, we propose BKSR, the first unsupervised blind HSI-SR method, which consists of three synergistic modules: Gibbs sampling-based band selection (GBS), test-time-training kernel estimation (TKE), and robust HSI restoration (RHR). These modules form a closed-loop optimization cycle: i) In GBS, the dynamic ergodicity of Gibbs sampling provides a global spectral view for kernel estimation and HSI restoration while maintaining local spectral computations; ii) In TKE, the GBS-sampled bands guide the kernel estimator update, achieving a learnable sampling-based mechanism, which refines kernel estimation to regularize RHR's diffusion trajectory; iii) In RHR, a spectral hyper-Laplacian prior is integrated into the reverse process of an off-the-shelf diffusion model, which achieves non-i.i.d. noise robust HSI restoration, feedback reweights band and kernel importance for subsequent GBS and TKE iterations. Extensive experiments on both synthetic and real HSI datasets demonstrate our BKSR's superiority over baseline methods across diverse scenarios (e.g., unknown Gaussian/motion kernel, non-i.i.d. noise) while maintaining comparable computational costs to the classic band selection methods.
Zhixiong Yang 0001, Jingyuan Xia, Shengxi Li, Lingyu Zheng, Shuanghui Zhang, Li Liu 0002, Yaowen Fu, Yongxiang Liu
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Ultra-High-Definition Image Restoration via High-Frequency Enhanced Transformer
abstract
Transformer-based architectures exhibit substantial promise in the realm of ultra-high-definition (UHD) image restoration (IR). Nevertheless, they encounter significant challenges in maintaining high-frequency (HF) details, which are crucial for the reconstruction of texture. Conventional methods tackle computational complexity by significantly reducing the resolution (by a factor of 4 to 8). Moreover, the majority of high-frequency components are eliminated due to the inherent characteristics of self-attention mechanisms, as these mechanisms tend to naturally suppress high-frequency elements during non-local feature integration. This paper proposes a dual-branch transformer architecture that synergistically combines native-resolution HF preservation with efficient contextual modeling, named HiFormer. The high-resolution branch utilizes a directionally-sensitive large-kernel decomposition to effectively address anisotropic degradations with fewer parameters and applies depthwise separable convolutions for localized high-frequency (HF) information extraction. Concurrently, the low-resolution branch assimilates these localized HF elements using adaptive channel modulation to offset spectral losses induced by the inherent smoothing effect of self-attention. Comprehensive experiments across numerous UHD image restoration tasks reveal that our approach surpasses current leading methods in both quantitative metrics and qualitative analysis. The code is available at https://github.com/5chen/HiFormer.
Chen Wu 0006, Zhuoran Zheng, Weidong Jiang, Yuning Cui 0001, Jingyuan Xia
IEEE Trans. Circuits Syst. Video Technol.6
2026 Machines Serve Human: A Novel Variable Human-Machine Collaborative Compression Framework
abstract
Human-machine collaborative compression has been receiving increasing research efforts for reducing image/video data, serving as the basis for both human perception and machine intelligence. Existing collaborative methods are dominantly built upon the de facto human-vision compression pipeline, witnessing deficiency on complexity and bit-rates when aggregating the machine-vision compression. Indeed, machine vision solely focuses on the core regions within the image/video, requiring much less information compared with the compressed information for human vision. In this paper, we thus set out the first successful attempt by a novel collaborative compression method based on the machine-vision-oriented compression, instead of human-vision pipeline. In other words, machine vision serves as the basis for human vision within collaborative compression. A plug-and-play variable bit-rate strategy is also developed for machine vision tasks. Then, we propose to progressively aggregate the semantics from the machine-vision compression, whilst seamlessly tailing the diffusion prior to restore high-fidelity details for human vision, thus named as diffusion-prior based feature compression for human and machine visions (Diff-FCHM). Experimental results verify the consistently superior performances of our Diff-FCHM, on both machine-vision and human-vision compression with remarkable margins. The source code is available at https://github.com/bblgbr/Diff-FCHM.
Zifu Zhang, Shengxi Li, Xiancheng Sun, Mai Xu, Zhengyuan Liu, Jingyuan Xia
IEEE Trans. Image Process.6
2025 A Cross-Modal Multi-Attitude Framework for the Generation of Space Target ISAR Images
abstract
Inverse Synthetic Aperture Radar (ISAR) imagery of space targets exhibits superior physical fidelity and satisfactory textural representation of components in ISAR images of targets, even under conditions characterized by sparse input optical samples. This paper introduces an innovative optical-to-radar cross-modal framework for the generation of full-attitude, high-fidelity space target ISAR samples, denominated as AORC. Specifically, the attitude encoding module (AEM) assimilates the prior knowledge of analogous targets across different attitudes through a fine-designed NeRF-based encoder, subsequently deriving the encoded features in the latent space. Subsequently, these comprehensive attitude features are input into the modality transformation module (MTM) to undergo a Brownian-Bridge-based diffusion process, facilitating the transformation between optical and ISAR modalities for each feature from each attitude. Extensive simulations on satellite targets validate the effectiveness of the proposed approach.
Derong Kong, Huaizhang Liao, Jingyuan Xia
ICASSP3
2025 Spherical-Nested Diffusion Model for Panoramic Image Outpainting
abstract
Panoramic image outpainting acts as a pivotal role in immersive content generation, allowing for seamless restoration and completion of panoramic content. Given the fact that the majority of generative outpainting solutions operates on planar images, existing methods for panoramic images address the sphere nature by soft regularisation during the end-to-end learning, which still fails to fully exploit the spherical content. In this paper, we set out the first attempt to impose the sphere nature in the design of diffusion model, such that the panoramic format is intrinsically ensured during the learning procedure, named as spherical-nested diffusion (SpND) model. This is achieved by employing spherical noise in the diffusion process to address the structural prior, together with a newly proposed spherical deformable convolution (SDC) module to intrinsically learn the panoramic knowledge. Upon this, the proposed method is effectively integrated into a pre-trained diffusion model, outperforming existing state-of-the-art methods for panoramic image outpainting. In particular, our SpND method reduces the FID values by more than 50\% against the state-of-the-art PanoDiffusion method. Codes are publicly available at \url{https://github.com/chronos123/SpND}.
Xiancheng Sun, Senmao Ma, Shengxi Li, Mai Xu, Jingyuan Xia, Lai Jiang 0004, Xin Deng 0002
ICML5
2025 Real-time Distributed Force Sensing-Based Position Feedback Control for Fiber-Driven Miniaturized Continuum Robots
abstract
Continuum robots are widely used in the medical scenarios due to their dexterity and flexibility. However, precise end-to-end control of continuum robots remains challenging, limited by the kinematic or kinetostatic accuracy and no enough space for additional sensors configurations. This paper proposes a precise position control method for fiber-driven continuum robots using the reconstructed shape based on distributed force sensing from the same fibers, where the optical fibers serve as both robot actuation and force sensing simultaneously without requiring additional sensors. First, we use single-core optical fibers (SCFs) as the actuation cables of the continuum robot, and each fiber has multiple fiber Bragg grating (FBG) sensors inscribed on it to sense distributed force along the entire cables. Then, the forward kinetostatics model of the fiber-driven continuum robot is established using the known distributed forces as the inputs. Notably, the nonlinear friction between the cables and actuation channels does not require an additional estimation model. Benefiting from this, the shape can be accurately reconstructed after the stiffness calibration of the continuum robot. Finally, a position controller based on real-time feedback from shape is developed to achieve the tip position control of the continuum robot. Experimental results demonstrate that the proposed forward kinetostatics model can achieve the shape reconstruction with the errors of 0.45 mm and 0.57 mm in planar bending and spatial bending states, respectively. By comparison to the traditional constant curvature kinematics-based control method, the proposed methods can achieve the mean absolute error of 0.37 and 0.6 mm in two distinct path tracking tests. The proposed method using distributed forces sensing enables a real-time accurate position feedback control combined with kinetostatic model, instead of modelling the nonlinear friction or adding additional external sensors.
Jingyuan Xia, Zecai Lin, Junling Yang, Guang-Zhong Yang, Anzhu Gao
IROS1
2025 Luminance-Aware Statistical Quantization: Unsupervised Hierarchical Learning for Illumination Enhancement
abstract
Low-light image enhancement (LLIE) faces persistent challenges in balancing reconstruction fidelity with cross-scenario generalization. While existing methods predominantly focus on deterministic pixel-level mappings between paired low/normal-light images, they often neglect the continuous physical process of luminance transitions in real-world environments, leading to performance drop when normal-light references are unavailable. Inspired by empirical analysis of natural luminance dynamics revealing power-law distributed intensity transitions, this paper introduces Luminance-Aware Statistical Quantification (LASQ), a novel framework that reformulates LLIE as a statistical sampling process over hierarchical luminance distributions. Our LASQ re-conceptualizes luminance transition as a power-law distribution in intensity coordinate space that can be approximated by stratified power functions, therefore, replacing deterministic mappings with probabilistic sampling over continuous luminance layers. A diffusion forward process is designed to autonomously discover optimal transition paths between luminance layers, achieving unsupervised distribution emulation without normal-light references. In this way, it considerably improves the performance in practical situations, enabling more adaptable and versatile light restoration. This framework is also readily applicable to cases with normal-light references, where it achieves superior performance on domain-specific datasets alongside better generalization-ability across non-reference datasets. The code is available at: https://github.com/XYLGroup/LASQ.
Derong Kong, Zhixiong Yang 0001, Shengxi Li, Shuaifeng Zhi, Li Liu 0002, Zhen Liu 0004, Jingyuan Xia
NeurIPS7
2025 The SIFT based two-stage STC decoupled learning method for long-tailed SAR target recognition
Jingyuan Xia, Huaizhang Liao, Xu Lan, Weidong Jiang
Neurocomputing2
2025 Despeckling Representation for Data-Efficient SAR Ship Detection
abstract
Deep learning techniques are extensively applied to synthetic aperture radar (SAR) ship detection tasks. Nonetheless, the limited availability of labeled SAR images impedes the neural network’s ability to learn and extract robust object features from SAR ship images. To alleviate the dependence on large datasets, this letter introduces a despeckle-based representation learning approach for SAR ship detection, named despeckling ship detection YOLO (DS-YOLO). The DS-YOLO model integrates a shared feature extractor, a detection head, and a despeckling head, facilitating the concurrent performance of SAR image despeckling and ship detection. The model effectively reduces the potential for neural network overfitting by conducting joint learning of detection and despeckling processes. As a result, DS-YOLO is particularly well-suited for SAR ship detection tasks with constrained training data. Comprehensive experiments indicate that DS-YOLO substantially surpasses the performance of the conventional detection model, particularly in scenarios where labeled data are severely restricted. Source codes are available athttps://github.com/Cthanta/DS-YOLO.
Ruikang Hu, Huangxing Lin, Zhejun Lu, Jingyuan Xia
IEEE Geosci. Remote. Sens. Lett.4
2025 SAIG: Semantic-Aware ISAR Generation via Component-Level Semantic Segmentation
abstract
This paper addresses the challenge of generating high-fidelity Inverse Synthetic Aperture Radar (ISAR) images from optical images, particularly for space targets. We propose a framework for the generation of ISAR images incorporating component refinement, which attains high-fidelity ISAR scattering characteristics through the integration of an advanced generation model predicated on semantic segmentation, designated as Semantic-Aware ISAR Generation (SAIG). SAIG renders ISAR images from optical equivalents by learning mutual semantic segmentation maps. Extensive simulations demonstrate its effectiveness and robustness, outperforming state-of-the-art methods by over 8% across key evaluation metrics.
Huaizhang Liao, Derong Kong, Zhixiong Yang 0001, Jingyuan Xia
IEEE Geosci. Remote. Sens. Lett.5
2025 SAKE: Unsupervised HSI Super-Resolution via Adaptive Kernel Estimation and Reconstruction
abstract
Methods for hyperspectral image (HSI) super-resolution employing deep learning have been pivotal in a range of fields. Despite this, many current approaches face challenges like the scarcity of paired datasets, simplified degradation models, and a lack of image prior information, leading to weak generalization over different datasets and degradation conditions. To overcome these challenges, this paper introduces a single hyperspectral image blind super-resolution algorithm supplemented by an unsupervised blur kernel estimation module. Random kernels from Gaussian distributions form pseudo-label kernels to handle arbitrary degradation kernels. These are processed by a DIP-based network, channel-by-channel, using a batch gradient acceleration algorithm for network parameter updates. This unsupervised, pre-training-free method achieves single HSI super-resolution through alternating iterative optimization. This method necessitates solely the input of the degraded image and requires no extra data, demonstrating a minimal reliance on data. Simulated experiments across datasets and scenarios demonstrate the proposed method’s superior ability to estimate degradation blur kernels, outperforming existing state-of-the-art methods.
Lingyu Zheng, Zhixiong Yang 0001, Jingyuan Xia
IEEE Geosci. Remote. Sens. Lett.4
2025 DeepSN-Net: Deep Semi-Smooth Newton Driven Network for Blind Image Restoration
abstract
The deep unfolding network represents a promising research avenue in image restoration. However, most current deep unfolding methodologies are anchored in first-order optimization algorithms, which suffer from sluggish convergence speed and unsatisfactory learning efficiency. In this paper, to address this issue, we first formulate an improved second-order semi-smooth Newton (ISN) algorithm, transforming the original nonlinear equations into an optimization problem amenable to network implementation. After that, we propose an innovative network architecture based on the ISN algorithm for blind image restoration, namely DeepSN-Net. To the best of our knowledge, DeepSN-Net is the first successful endeavor to design a second-order deep unfolding network for image restoration, which fills the blank of this area. Furthermore, it offers several distinct advantages: 1) DeepSN-Net provides a unified framework to a variety of image restoration tasks in both synthetic and real-world contexts, without imposing constraints on the degradation conditions. 2) The network architecture is meticulously aligned with the ISN algorithm, ensuring that each module possesses robust physical interpretability. 3) The network exhibits high learning efficiency, superior restoration accuracy and good generalization ability across 11 datasets on three typical restoration tasks. The success of DeepSN-Net on image restoration may ignite many subsequent works centered around the second-order optimization algorithms, which is good for the community.
Xin Deng 0002, Lai Jiang 0004, Jingyuan Xia, Mai Xu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Fusion2Void: Unsupervised Multi-Focus Image Fusion Based on Image Inpainting
abstract
Multi-focus image fusion aims to integrate clear segments from different partially focused images, creating an ‘all-in-focus’ composite. Due to the lack of ground-truth for multi-focus image fusion, supervised deep learning methods are deemed inappropriate for this task. In this paper, we present an unsupervised approach for multi-focus image fusion, named Fusion2Void. Fusion2Void ingeniously tackles the challenge of missing ground-truth by framing image inpainting as an auxiliary task. Specifically, Fusion2Void utilizes a fusion network to merge focused regions from multiple source images. Following the fusion process, image patches in the source images are randomly dropped to construct an additional image inpainting task. Subsequently, an image inpainting network uses the fused image as a guide to restore the missing content in the source images. The missing content in the source images includes both focused and defocused regions. Restoring focused image patches is significantly more challenging than restoring their defocused counterparts due to their inclusion of more high-frequency details. If the focused image patches are effectively restored, the repair of the defocused image patches becomes notably easier. Therefore, the image inpainting network implicitly compels the fused image to incorporate all focused content from the source images, as these can be utilized to restore the missing focused regions in the source images perfectly. Based on image inpainting, the fusion network generates ‘all-in-focus’ images in an unsupervised manner. Experiments on several synthetic and real-world datasets highlight Fusion2Void’s state-of-the-art performance relative to other methods.
Huangxing Lin, Yunlong Lin, Jingyuan Xia, Linyu Fan, Yingying Wang 0005, Xinghao Ding
IEEE Trans. Circuits Syst. Video Technol.3
2025 GDROS: A Geometry-Guided Dense Registration Framework for Optical-SAR Images Under Large Geometric Transformations
abstract
Registration of optical and synthetic aperture radar (SAR) remote sensing images serves as a critical foundation for image fusion and visual navigation tasks. This task is particularly challenging because of their modal discrepancy, primarily manifested as severe nonlinear radiometric differences (NRD), geometric distortions, and noise variations. Under large geometric transformations, existing classical template-based and sparse keypoint-based strategies struggle to achieve reliable registration results for optical-SAR image pairs. To address these limitations, we propose GDROS, a geometry-guided dense registration framework leveraging global cross-modal image interactions. First, we extract cross-modal deep features from optical and SAR images through a CNN-Transformer hybrid feature extraction module, upon which a multi-scale 4D correlation volume is constructed and iteratively refined to establish pixel-wise dense correspondences. Subsequently, we implement a least squares regression (LSR) module to geometrically constrain the predicted dense optical flow field. Such geometry guidance mitigates prediction divergence by directly imposing an estimated affine transformation on the final flow predictions. Extensive experiments have been conducted on three representative datasets WHU-Opt-SAR dataset, OS dataset, and UBCv2 dataset with different spatial resolutions, demonstrating robust performance of our proposed method across different imaging resolutions. Qualitative and quantitative results show that GDROS significantly outperforms current state-of-the-art methods in all metrics. Our source code will be released at: https://github.com/Zi-Xuan-Sun/GDROS.
Zixuan Sun, Shuaifeng Zhi, Ruize Li, Jingyuan Xia, Yongxiang Liu, Weidong Jiang
IEEE Trans. Geosci. Remote. Sens.4
2025 Spherical Patch Generative Adversarial Net for Unconditional Panoramic Image Generation
abstract
Recent advancements in virtual reality (VR) and augmented reality (AR) have popularised the emerging panoramic content for the immersive visual experience. The difficulty in acquisition and display of 360° format further highlights the necessity of unconditional panoramic image generation. Existing methods essentially generate planar images mapped from panoramic images, and fail to address the deformation and closed-loop characteristics when inverted back to the panoramic images. Thus leading to the generation of pseudo-panoramic content. This paper aims to directly generate spherical content, in a patch-by-patch style; besides computation friendly, this promises the anywhere continuity on the panoramic image and proper accommodation of panoramic deformation. More specifically, we first propose a novel spherical patch convolution (SPConv) that operates on the local spherical patch, which naturally addresses the deformation of panoramic content. We then propose our spherical patch generative adversarial net (SP-GAN) that consists of spherical local embedding (SLE) and spherical content synthesiser (SCS) modules, which seamlessly incorporate our SPConv so as to generate continuous panoramic patches. To the best of our knowledge, the proposed SP-GAN is the first successful attempt to accommodate the spherical distortion for closed-loop panoramic image generation in a patch-by-patch manner. The experimental results, with human-rated evaluations, have verified the consistently superior performances for unconditional panoramic image generation, from the perspectives of generation quality, computational memory, and generalisation to various resolutions. Codes are publicly available at https://github.com/chronos123/SP-GAN.
Mai Xu, Xiancheng Sun, Shengxi Li, Lai Jiang 0004, Jingyuan Xia, Xin Deng 0002
IEEE Trans. Image Process.5
2024 A Dynamic Kernel Prior Model for Unsupervised Blind Image Super-Resolution
abstract
Deep learning-based methods have achieved significant successes on solving the blind super-resolution (BSR) problem. However, most of them request supervised pretraining on labelled datasets. This paper proposes an unsupervised kernel estimation model, named dynamic kernel prior (DKP), to realize an unsupervised and pretraining-free learning-based algorithm for solving the BSR problem. DKP can adaptively learn dynamic kernel priors to realize real-time kernel estimation, and thereby enables superior HR image restoration performances. This is achieved by a Markov chain Monte Carlo sampling process on random kernel distributions. The learned kernel prior is then assigned to optimize a blur kernel estimation network, which entails a network-based Langevin dynamic optimization strategy. These two techniques ensure the accuracy of the kernel estimation. DKP can be easily used to replace the kernel estimation models in the existing methods, such as Double-DIP and FKP-DIP, or be added to the off-the-shelf image restoration model, such as diffusion model. In this paper, we incorporate our DKP model with DIP and diffusion model, referring to DIP-DKP and Diff-DKP, for validations. Extensive simulations on Gaussian and motion kernel scenarios demonstrate that the proposed DKP model can significantly improve the kernel estimation with comparable runtime and memory usage, leading to state-of-the-art BSR results. The code is available at https://github.com/XYLGroup/DKP.
Zhixiong Yang 0001, Jingyuan Xia, Shengxi Li, Xinghua Huang, Shuanghui Zhang, Zhen Liu 0004, Yaowen Fu, Yongxiang Liu
CVPR2
2024 DURRNET: Deep Unfolded Single Image Reflection Removal Network with Joint Prior
abstract
Single image reflection removal (SIRR) problem can be interpreted as a canonical blind source separation problem and is highly ill-posed. A parameter effective, fast learning and interpretable reflection removal algorithm is essential for many vision analysis applications. In this paper, we propose a novel model-inspired and learning-based SIRR method called Deep Unfolded Reflection Removal Network (DURRNet). It combines the merits of both model-based and learning-based paradigms, leading to a more interpretable and effective deep architecture. To achieve this, we first propose a model-based optimization approach and then obtain DURRNet by unfolding an iterative step into a Unfolded Separation Block (USB) based on proximal gradient descent. Key features of DURR-Net include the use of Invertible Neural Networks to impose the transform-based exclusion prior on the basis of natural image prior, as well as a coarse-to-fine architecture to fine-grain the reflection removal process. Extensive experiments on public datasets demonstrate that DURRNet achieves state-of-the-art results not only visually, quantitatively, but also effectively.
Junjie Huang 0001, Tianrui Liu 0001, Jingyuan Xia, Meng Wang 0001, Pier Luigi Dragotti
ICASSP3
2024 External Interaction Estimation of 6-PSS Parallel Robots with Embodied Mechanical Intelligence
abstract
Traditional interaction perception of parallel robots relies on a six-dimensional force sensor for contact sensing at their distal end. However, the sensor body occupies the space of moving platform and also increases the load on the robot actuations. To enable both minimization and embodied intelligence, this paper proposes an external interaction estimation method with embodied mechanical intelligence by embedding two single-axis force sensors in each leg of 6-PSS parallel robot. The method uses a backward propagation neural network optimized by sparrow search algorithm, and it can simultaneously estimate the external force and its position using information from multiple single-axis force sensors and the encoder of driving motor. The experimental platform is established to collect the data and train the network. The result shows that the force estimation mean error is 2.4% and the position estimation error is 2.9%. A demonstration with a virtual display interface showing the reconstructed parallel robot pose, and the interaction force and its pose using the proposed estimation method, indicates the effectiveness of the proposed interaction method with embodied mechanical intelligence for 6-PSS parallel robot
Jingyuan Xia, Zecai Lin, Xiaojie Ai, Guangjun Yu, Anzhu Gao
IROS1
2024 An Integrated Network for SA-ISAR Image Processing With Adaptive Denoising and Super-Resolution Modules
abstract
This letter focuses on developing an effective and generalizable deep learning approach for inverse synthetic aperture radar (ISAR) image super-resolution (SR). Since the ISAR imaging process is typically carried out under sparse aperture (SA) conditions, imaging results may exhibit striped noise caused by echoes missing, making it challenging to apply conventional SR methods directly. In view of this, we present a blind SR (BSR) method specifically designed for ISAR images with striped noise. The proposed method employs an integrated network that includes an adaptive denoising module and a SR module (AD-SRNet). Experimental results on both synthetic and real ISAR samples demonstrate the superior performance and strong generalization capability of our approach.
Mingyao Chen, Jingyuan Xia, Tianpeng Liu, Li Liu 0002
IEEE Geosci. Remote. Sens. Lett.2
2024 Meta-learning based blind image super-resolution approach to different degradations
Zhixiong Yang 0001, Jingyuan Xia, Shengxi Li, Wende Liu, Shuaifeng Zhi, Shuanghui Zhang, Li Liu 0002, Yaowen Fu, Deniz Gündüz
Neural Networks2
2024 Blind Super-Resolution via Meta-Learning and Markov Chain Monte Carlo Simulation
abstract
Learning based approaches have witnessed great successes in blind single image super-resolution (SISR) tasks, however, handcrafted kernel priors and learning based kernel priors are typically required. In this paper, we propose a meta-learning and Markov Chain Monte Carlo (MCMC) based SISR approach to learn kernel priors from organized randomness. In concrete, a lightweight network is adopted as kernel generator, and is optimized via learning from the MCMC simulation on random Gaussian distributions. This procedure provides an approximation for the rational blur kernel, and introduces a network-level Langevin dynamics into SISR optimization processes, which contributes to preventing bad local optimal solutions for kernel estimation. Meanwhile, a meta-learning based alternating optimization procedure is proposed to optimize the kernel generator and image restorer, respectively. In contrast to the conventional alternating minimization strategy, a meta-learning based framework is applied to learn an adaptive optimization strategy, which is less-greedy and results in better convergence performance. These two procedures are iteratively processed in a plug-and-play fashion, for the first time, realizing a learning-based but plug-and-play blind SISR solution in unsupervised inference. Extensive simulations demonstrate the superior performance and generalization ability of the proposed approach when compared with the Start-of-the-Art solutions on synthesis and real-world datasets.
Jingyuan Xia, Zhixiong Yang 0001, Shengxi Li, Shuanghui Zhang, Yaowen Fu, Deniz Gündüz, Xiang Li 0014
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Meta-Learning Based Domain Prior With Application to Optical-ISAR Image Translation
abstract
This paper focuses on generating Inverse Synthetic Aperture Radar (ISAR) images from optical images, in particular, for orbit space targets. ISAR images are widely applied in space target observation and classification tasks, whereas, limited to the expensive cost of ISAR sample collection, training deep learning-based ISAR image classifiers with insufficient samples and generating ISAR samples from emulation optical images via image translation techniques have attracted increasing attention. Image translation has highlighted significant success and popularity in computer vision, remote sensing and data generation societies. However, most of the existing methods are implemented under the discipline of extracting the explicit pixel-level features and do not perform effectively while entailing translation to domains with specific implicit features, such as ISAR image does. We propose a meta-learning based domain prior to implicit feature modelling and apply it to CycleGAN and UNIT models to realize effective translations between the ISAR and optical domains. Two representative implicit features, ISAR scattering distribution feature from the physical domain and the classification identifying feature from the task domain, are elaborately formulated with explicit modelling in statistic form. A meta-learning based training scheme is introduced to leverage the mutual knowledge of domain priors across different samples, and thus allows few-shot learning capacity with dramatically reduced training samples. Extensive simulations validate that the obtained ISAR images have better visible-authenticity and training-effectiveness than the existing image translation approaches on various synthetic datasets. Source codes are available at.
Huaizhang Liao, Jingyuan Xia, Zhixiong Yang 0001, Fulin Pan, Zhen Liu 0004, Yongxiang Liu
IEEE Trans. Circuits Syst. Video Technol.2
2023 SM-CNN: Separability Measure-Based CNN for SAR Target Recognition
abstract
With the maturity of deep learning algorithm in Synthetic Aperture Radar (SAR) target recognition filed, Convolutional Neural Network (CNN) has become the most effective model. However, the interpretability and the separability of feature maps extracted from convolution layers have not been specially analyzed neither qualitatively nor quantitatively, which makes the traditional model work like a “black box”. To alleviate the problem, a novel model based on separability measure (SM) - CNN is proposed in this letter, which introduces the principle of maximal coding rate reduction to the backbone module. SM-CNN quantitatively analyzes the separability of the feature maps and takes the value as a vital part of the loss function to guide the training process of the model. The calculation process of the separability measure values can be strictly derived mathematically, so it is more interpretable, turning the black box into a “gray box”. Additionally, the proposed model can achieve comparable recognition performance of the backbone networks with reduced computational complexity. Comparative experiments based on MSTAR and OpenSARShip data sets verify the effectiveness and practicability of the method proposed in this letter.
Yifan Zhang 0015, Jingyuan Xia, Xunzhang Gao, Lingyan Xue, Xinyu Zhang 0010, Xiang Li 0014
IEEE Geosci. Remote. Sens. Lett.2
2023 Localized Incomplete Multiple Kernel k-Means With Matrix-Induced Regularization
abstract
Localized incomplete multiple kernel k -means (LI-MKKM) is recently put forward to boost the clustering accuracy via optimally utilizing a quantity of prespecified incomplete base kernel matrices. Despite achieving significant achievement in a variety of applications, we find out that LI-MKKM does not sufficiently consider the diversity and the complementary of the base kernels. This could make the imputation of incomplete kernels less effective, and vice versa degrades on the subsequent clustering. To tackle these problems, an improved LI-MKKM, called LI-MKKM with matrix-induced regularization (LI-MKKM-MR), is proposed by incorporating a matrix-induced regularization term to handle the correlation among base kernels. The incorporated regularization term is beneficial to decrease the probability of simultaneously selecting two similar kernels and increase the probability of selecting two kernels with moderate differences. After that, we establish a three-step iterative algorithm to solve the corresponding optimization objective and analyze its convergence. Moreover, we theoretically show that the local kernel alignment is a special case of its global one with normalizing each base kernel matrices. Based on the above observation, the generalization error bound of the proposed algorithm is derived to theoretically justify its effectiveness. Finally, extensive experiments on several public datasets have been conducted to evaluate the clustering performance of the LI-MKKM-MR. As indicated, the experimental results have demonstrated that our algorithm consistently outperforms the state-of-the-art ones, verifying the superior performance of the proposed algorithm.
Miaomiao Li 0001, Jingyuan Xia, Qing Liao 0001, Xinzhong Zhu, Xinwang Liu 0002
IEEE Trans. Cybern.2
2023 Metalearning-Based Alternating Minimization Algorithm for Nonconvex Optimization
abstract
In this article, we propose a novel solution for nonconvex problems of multiple variables, especially for those typically solved by an alternating minimization (AM) strategy that splits the original optimization problem into a set of subproblems corresponding to each variable and then iteratively optimizes each subproblem using a fixed updating rule. However, due to the intrinsic nonconvexity of the original optimization problem, the optimization can be trapped into a spurious local minimum even when each subproblem can be optimally solved at each iteration. Meanwhile, learning-based approaches, such as deep unfolding algorithms, have gained popularity for nonconvex optimization; however, they are highly limited by the availability of labeled data and insufficient explainability. To tackle these issues, we propose a meta-learning based alternating minimization (MLAM) method that aims to minimize a part of the global losses over iterations instead of carrying minimization on each subproblem, and it tends to learn an adaptive strategy to replace the handcrafted counterpart resulting in advance on superior performance. The proposed MLAM maintains the original algorithmic principle, providing certain interpretability. We evaluate the proposed method on two representative problems, namely, bilinear inverse problem: matrix completion and nonlinear problem: Gaussian mixture models. The experimental results validate the proposed approach outperforms AM-based methods.
Jingyuan Xia, Shengxi Li, Junjie Huang 0001, Zhixiong Yang 0001, Imad Jaimoukha, Deniz Gündüz
IEEE Trans. Neural Networks Learn. Syst.1
2022 Speckle-Variant Attack: Toward Transferable Adversarial Attack to SAR Target Recognition
abstract
Recent advances of deep neural networks (DNNs) highlight the success on synthetic aperture radar automatic target recognition (SAR ATR) with superiority effectiveness and efficiency. However, the DNNs are known to be vulnerable to the adversarial examples, whose performance will be dramatically reduced when the imperceptible perturbation exists. In optical image processing, invisible perturbations are typically embedded in the way of a full-scaled distribution in purely digital setting. Whereas, it is not feasible to achieve this in SAR ATR tasks due to the inaccessibility of SAR system and unique imaging mechanism. In practical, the subtle perturbations could be produced by physical approaches that change the scattering property of the target. Therefore, the adversarial perturbations for SAR ATR should be of good transferability to achieve effective attack on major DNNs classifiers, as well as accessible additive region in SAR images with respect to the realistic target locations. In this letter, we present a novel approach, namely speckle variant attack (SVA). The proposed SVA is composed of two major modules: an iterative gradient based perturbation generator and a target region extractor. The perturbation generator implements a speckle variant transformation that continuously reconstruct the speckle noise pattern during each of the iterations for strong transferability. The target region extractor ensures the feasibility of the additive adversarial perturbations in practical scenarios through restricting the region of the perturbation. Therefore, the proposed SVA is capable of producing adversarial examples that are more transferable and physically feasible. Extensive evaluations on the MSTAR dataset show that the SVA has achieved the superior transferability and competitive time consumption compared with the SOTA transformation-based techniques, including the diverse inputs method and the scale-invariant method.
Bowen Peng, Jie Zhou 0031, Jingyuan Xia, Li Liu 0002
IEEE Geosci. Remote. Sens. Lett.4
2021 Meta-learning Based Beamforming Design for MISO Downlink
abstract
Downlink beamforming is an essential technology for wireless cellular networks; however, the design of beamforming vectors that maximize the weighted sum rate (WSR) is an NP-hard problem and iterative algorithms are typically applied to solve it. The weighted minimum mean square error (WMMSE) algorithm is the most widely used one, which iteratively minimizes the WSR and converges to a local optimal. Motivated by the recent developments in meta-learning techniques to solve non-convex optimization problems, we propose a meta-learning based iterative algorithm for WSR maximization in a MISO downlink channel. A long-short-term-memory (LSTM) network based meta-learning model is built to learn a dynamic optimization strategy to update the variables iteratively. The learned strategy aims to optimize each variable in a less greedy manner compared to WMMSE, which updates variables by computing their first order stationary points at each iteration step. The proposed algorithm outperforms WMMSE significantly in the high signal to noise ratio (SNR) regime and achieves comparable performance when the SNR is low.
Jingyuan Xia, Deniz Gündüz
ISIT1
2021 Efficient and Effective Regularized Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering (IMVC) optimally combines multiple pre-specified incomplete views to improve clustering performance. Among various excellent solutions, the recently proposed multiple kernel k-means with incomplete kernels (MKKM-IK) forms a benchmark, which redefines IMVC as a joint optimization problem where the clustering and kernel matrix imputation tasks are alternately performed until convergence. Though demonstrating promising performance in various applications, we observe that the manner of kernel matrix imputation in MKKM-IK would incur intensive computational and storage complexities, over-complicated optimization and limitedly improved clustering performance. In this paper, we first propose an Efficient and Effective Incomplete Multi-view Clustering (EE-IMVC) algorithm to address these issues. Instead of completing the incomplete kernel matrices, EE-IMVC proposes to impute each incomplete base matrix generated by incomplete views with a learned consensus clustering matrix. Moreover, we further improve this algorithm by incorporating prior knowledge to regularize the learned consensus clustering matrix. Two three-step iterative algorithms are carefully developed to solve the resultant optimization problems with linear computational complexity, and their convergence is theoretically proven. After that, we theoretically study the generalization bound of the proposed algorithms. Furthermore, we conduct comprehensive experiments to study the proposed algorithms in terms of clustering accuracy, evolution of the learned consensus clustering matrix and the convergence. As indicated, our algorithms deliver their effectiveness by significantly and consistently outperforming some state-of-the-art ones.
Xinwang Liu 0002, Miaomiao Li 0001, Chang Tang, Jingyuan Xia, Jian Xiong 0002, Li Liu 0002, Marius Kloft, En Zhu
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Feature Selective Projection with Low-Rank Embedding and Dual Laplacian Regularization
abstract
Feature extraction and feature selection have been regarded as two independent dimensionality reduction methods in most of the existing literature. In this paper, we propose to integrate both approaches into a unified framework and design an unsupervised linear feature selective projection (FSP) for feature extraction with low-rank embedding and dual Laplacian regularization, with the aim to exploit the intrinsic relationship among data and suppress the impact of noise. Specifically, a projection matrix with an l2,1-norm regularization is introduced to project original high dimensional data points into a new subspace with lower dimension, where the l2,1-norm regularization can endow the projection with good interpretability. We deploy a coefficient matrix with low rank constraint to reconstruct the data points and the l2,1-norm is imposed to regularize the data reconstruction errors in the low-dimensional subspace and make FSP robust to noise. Furthermore, a dual graph Laplacian regularization term is imposed on the low dimensional data and data reconstruction matrix for preserving the local manifold geometrical structure of data. Finally, an alternatively iterative algorithm is carefully designed for solving the proposed optimization model. Theoretical convergence and computational complexity analysis of the algorithm are also provided. Comprehensive experiments on various benchmark datasets have been carried out to evaluate the performance of the proposed FSP. As indicated, our algorithm significantly outperforms other state-of-the-art methods for feature extraction.
Chang Tang, Xinwang Liu 0002, Xinzhong Zhu, Jian Xiong 0002, Miaomiao Li 0001, Jingyuan Xia, Xiangke Wang, Lizhe Wang 0001
IEEE Trans. Knowl. Data Eng.6
2019 Multi-view Clustering via Late Fusion Alignment Maximization
abstract
Multi-view clustering (MVC) optimally integrates complementary information from different views to improve clustering performance. Although demonstrating promising performance in many applications, we observe that most of existing methods directly combine multiple views to learn an optimal similarity for clustering. These methods would cause intensive computational complexity and over-complicated optimization. In this paper, we theoretically uncover the connection between existing k-means clustering and the alignment between base partitions and consensus partition. Based on this observation, we propose a simple but effective multi-view algorithm termed {Multi-view Clustering via Late Fusion Alignment Maximization (MVC-LFA)}. In specific, MVC-LFA proposes to maximally align the consensus partition with the weighted base partitions. Such a criterion is beneficial to significantly reduce the computational complexity and simplify the optimization procedure. Furthermore, we design a three-step iterative algorithm to solve the new resultant optimization problem with theoretically guaranteed convergence. Extensive experiments on five multi-view benchmark datasets demonstrate the effectiveness and efficiency of the proposed MVC-LFA.
Siwei Wang 0001, Xinwang Liu 0002, En Zhu, Chang Tang, Jiyuan Liu 0003, Jingtao Hu, Jingyuan Xia, Jianping Yin
IJCAI7