Ni Xu

dblp:91/10752 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3
YearPublicationVenuePosition
2025 Gradient Selection Tuning via Information Bottleneck
abstract
Pre-trained visual models enjoy strong representations, yet suffer from massive parameters to be shifted in downstream practices. Many parameter-efficient fine-tuning methods have been proposed, mostly requiring only 1% additional parameters to achieve comparable results. However, current solutions either consider all feature channels equally or detect saliencies with individual layer, leading to many redundancies reserved. To address current issues, this paper proposes a new parameter fine-tuning method named “Gradient Selection Tuning” (GST), which leverages gradients that are capable of capturing the cascading effects across successive channels. Instead of saliency detection, we turn to compress the redundancies for channel selection, since the computed gradient values enjoy much lower mutual information. With GST facilitated, we further elaborate an Information-Guided Adapter following information bottleneck theory, effectively performing parameter compression yet with task-specific features preserved. Experimental results demonstrate that our method outperforms the baseline methods by adding only 0.075M parameters to ViT-B backbone. On domain generalization, our proposal also enjoys strong performance in low-parameter scenarios.
Xiaoxu Lin, Wei Li 0034, Ni Xu, Honghui Xu 0002, Jianwei Zheng 0001
ECAI4
2025 Object-Level Control for Refined Structure and Appearance in Conditional Image Synthesis
abstract
Recent advances in pretrained diffusion models, particularly the FreeControl, have enabled fine-grained spatial control in text-to-image generation. However, FreeControl still suffers notable limitations in detail generation and appearance synthesis. With deep analysis performed, we reveal that inadequate feature representation in the early generation phases is the main cause of the insufficient structure elaboration. To address this issue, we dig into the temporal evolution of inverse attention features, then extract more expressive structure information as the input to the guidance function, ensuring formation integrity. Moreover, to tackle the surface degradation and ambiguity caused by the dual guidance of structure and appearance, we engineer the Adaptive Instance Normalization (AdaIN) mechanism into a latent space, rather than the typical feature space, during the intermediate generation stage. This improvement not only guarantees the close alignment between the generated image and structural reference but also significantly strengthens the appearance modeling capability and optimizes the texture representation of both foreground and background elements. Extensive experiments demonstrate that our proposal consistently outperforms existing baseline models across multiple metrics, including Self-sim, CLIP, and LPIPS. Both quantitative and qualitative results confirm that our approach achieves superior performance in terms of content consistency, visual quality, and detail preservation.
Ni Xu, Wei Li 0034, Zhouchao Fu, Xiaoxu Lin, Jianwei Zheng 0001
ECAI1
2025 Pipeline-Centered Neighboring Network for Deep Unfolding Pansharpening
abstract
Pansharpening technique is dedicated to enriching the spatial details of low-resolution multispectral images (LRMS) under the guidance of a panchromatic (PAN) image. With the guarantee of promising results, Transformer-based methods have enjoyed a high reputation in this field. However, to reduce computational cost, existing solutions typically divide images into smaller, independent windows, which often weakens inter-window and channel-wise interactions as well as leads to unsmooth edges. To address these issues, we first formulate the pansharpening task as a variational optimization problem, and subsequently solve its data and prior subproblems alternately through an unrolling algorithm. In the prior extractor, we propose a Pipeline-Centered Neighboring Attention (PCNA), which holistically allows all pixels to share the same attention span while fully leveraging channel dependencies, thereby significantly improving the capability to process multispectral images. Moreover, a Multi-Scale Channel-Aware (MSCA) module is designed to capture the edges and structural details. Finally, by sequentially integrating the data and prior modules at each iteration stage, we unroll the iterations into a stage-wise unfolding network. Extensive experiments on three satellite datasets demonstrate the effectiveness and efficiency of our proposal compared to cutting-edge methods.
Yan Li 0083, Qiuju Chen, Chuangjie Fang, Ni Xu, Honghui Xu 0002, Jianwei Zheng 0001
ICASSP4
2025 Spatial-Spectral Fusion Neural Operator
abstract
With the rapid development of deep learning, spatial-spectral fusion (SSF) has emerged as an ideal alternative to traditional, costly hyperspectral image (HSI) acquisition methods. However, current solutions necessitate training and storing multiple models for different scaling factors. Besides, a meticulously designed network architecture to meet desirable performance often suffers from a severe computational burden. To counteract the dilemma, we propose SFNO, a lightweight spatial-spectral fusion neural operator for arbitrary-scale SSF. SFNO leverages approximation theory by embedding features from two degraded functions into a high-dimensional latent space, enabling efficient learning of basis functions. Kernel integration mechanisms are then used to approximate certain priors, followed by dimensionality reduction to generate high-resolution HSIs. Moreover, with the aid of discrete invariance property, we propose a new mechanism of progressive resampling (PR), which allows for the shrinkage of necessary spatial domain without any performance degradation. Extensive experiments on CAVE and Harvard datasets show that SFNO and its variant significantly improve performance, especially in out-of-domain fusion, requiring only 0.098M parameters and 0.966G FLOPs.
Wei Li 0034, Jiawei Jiang 0002, Ni Xu, Yan Li 0083, Jianwei Zheng 0001
ICME3
2025 Semantic-Spatial Attention for Refined Object Placement in Text-to-Image Synthesis
abstract
Solely based on given prompts, text-guided diffusion models have enjoyed a unique capability in generating diverse and creative images. Nevertheless, the conveyance of image information through text presents a series of challenges, particularly in controlling the positioning of objects in synthesized images. Despite attempts of recent efforts in exploring alternative conditions, such as bounding box/mask-image pairs, the requirement of a substantial amount of paired data and time-consuming fine-tuning emerge as new issues. Given the observations that not only prompt-related cross-attention maps reveal the spatial arrangement and centroid positions of the objects, but also out-of-prompt markers enjoy rich semantic information, we thus engineer a weighted optimization loss. Specifically, three spatial sub-losses, namely inner box reinforcement loss, outer box attenuation loss, and centroid loss, are devised and seamlessly integrated into the sampling step of current vanilla diffusion models. Without any annotations of layout data required, the final approach runs in a training-free fashion. Extensive experiments with new performance scores demonstrate that our proposal not only successfully addresses the issue of object positioning but also boosts the capabilities of most current models, such as Stable Diffusion and GLIGEN, in high-quality synthesis and coverage of various concepts. Moreover, the proposed mechanism plays a plug-and-play role.
Jianwei Zheng 0001, Ni Xu, Wei Li 0034, Jiawei Jiang 0002, Xiaoqin Zhang 0002
IEEE Trans. Multim.2
2024 Alias-Free Mamba Neural Operator
abstract
Benefiting from the booming deep learning techniques, neural operators (NO) are considered as an ideal alternative to break the traditions of solving Partial Differential Equations (PDE) with expensive cost. Yet with the remarkable progress, current solutions concern little on the holistic function features--both global and local information-- during the process of solving PDEs. Besides, a meticulously designed kernel integration to meet desirable performance often suffers from a severe computational burden, such as GNO with $O(N(N-1))$, FNO with $O(NlogN)$, and Transformer-based NO with $O(N^2)$. To counteract the dilemma, we propose a mamba neural operator with $O(N)$ computational complexity, namely MambaNO. Functionally, MambaNO achieves a clever balance between global integration, facilitated by state space model of Mamba that scans the entire function, and local integration, engaged with an alias-free architecture. We prove a property of continuous-discrete equivalence to show the capability of MambaNO in approximating operators arising from universal PDEs to desired accuracy. MambaNOs are evaluated on a diverse set of benchmarks with possibly multi-scale solutions and set new state-of-the-art scores, yet with fewer parameters and better efficiency.
Jianwei Zheng 0001, Wei Li 0034, Ni Xu, Xiaoxu Lin, Xiaoqin Zhang 0002
NeurIPS3
2022 Predicting Infection Area of Dengue Fever for Next Week Through Multiple Factors
Cong-Han Zheng, Ping-Yu Hsu 0001, Ming-Shien Cheng, Ni Xu, Yu-Chun Chen
IEA/AIE4
2021 Using Machine Learning to Predict Salaries of Major League Baseball Players
Cheng-Yu Lee, Ping-Yu Hsu 0001, Ming-Shien Cheng, Jun-Der Leu, Ni Xu, Bo-Lun Kan
IEA/AIE (2)5
2014 A 2.5GHz ADPLL with PVT-insensitive ΔΣ dithered time-to-digital conversion by utilizing an ADDLL
abstract
A ΔΣ́ all-digital delay-locked loop (ADDLL) is proposed to realize a PVT-insensitive time-to-digital converter (TDC) with enhanced linearity in an all-digital phase-locked loop (ADPLL). With the proposed TDC, poor timing resolution and nonlinearity problems are mitigated, enabling a low cost, low comparison frequency TDC design without using the advanced CMOS technology. A novel digitally-controlled delay line (DCDL) is proposed to ensure monotonous and linear mapping between a digital control word and a total time delay. A phase error compensator (PEC) is employed to calibrate periodic phase error of the proposed TDC. A 2.5GHz ADPLL is designed in 0.18μm CMOS. Simulation results show that the proposed method effectively reduces fractional spurs caused by the TDC.
Ni Xu, Woogeun Rhee, Zhihua Wang 0001
ISCAS2
2013 A PLL/DLL based CDR with ΔΣ frequency tracking and low algorithmic jitter generation
abstract
A delay- and phase-locked loop (D/PLL) based clock and data recovery (CDR) system enables an independent bandwidth control for jitter transfer and jitter tolerance but requires careful loop design with PVT-sensitive analog building blocks. In this work, an all-digital DLL and a digitally-controlled type-I boosted-gain fractional-N PLL followed by an injection-locked oscillator (ILO) are designed to realize a semidigital CDR system with enhanced frequency tracking capability and low algorithm jitter generation. The proposed CDR designed in 90nm CMOS consumes 26.4mW from a 1.2V supply and occupies the active area of 1.17mm2.
Shuli Geng, Ni Xu, Jun Li 0024, Xueyi Yu, Woogeun Rhee, Zhihua Wang 0001
ISCAS2
2012 A 9.6Gb/s 5+1-lane source synchronous transmitter in 65nm CMOS technology
abstract
This paper describes the design of a low-jitter source-synchronous link transmitter macro for data rates of 9.6 Gb/s. The transmitter macro consists of 5 data channels plus 1 forwarded-clock channel. A low jitter PLL with bandwidth linearization is employed to achieve 0.66ps rms jitter. The power supply induced jitter is minimized by employing a hybrid clock distribution network which is proposed for both jitter and power consideration. To minimize the influence of PVT variation, Successive Approximation Register (SAR) sub block is implemented to accurately set the on chip impedance and the signal amplitude. A CML driver with 4 tap feed forward equalizer is implemented to compensate the channel loss. The transmitter is implemented in 65nm CMOS technology, the active chip area is 3.12 mm2.
Ke Huang 0003, Xuqiang Zheng, Ni Xu, Chun Zhang 0001, Woogeun Rhee, Zhihua Wang 0001
ISCAS4