EDBT 2026 Demo / reviewers in the wild / expert
Wei Li 0034
dblp:64/6025-34
· DBLP profile ↗
23ranked-venue papers
6as first author
11since 2021 · last 2026
0009-0002-8566-6073ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Arbitrary-Scale Fusion Operator for High-Resolution Hyperspectral ImagingabstractFor high-resolution hyperspectral (HrHs) imaging, spatial-spectral fusion offers a promising alternative to expensive equipment. However, retraining multiple models for varied scaling factors is currently unavoidable, costing extra computational resource and human labor. To address this issue, we propose Arbitrary-scale Fusion Operator (AFO), a lightweight solution for HrHs fusion given arbitrary scalings, turning the retraining strategy into “training-free” ones. Specifically, AFO treats low-resolution hyperspectral (LrHs) images and high resolution multispectral (HrMs) images as light-wise degraded functions within the spectrum, which are initially embedded into a high-dimensional space to simulate the original light signals, tapping the potential of enriched prior learning. Then, a flow of kernel integration (KI) is meticulously crafted, followed by a rival step of dimension reduction for HrHs generation. For a well-behaved KI computation, an Attention-Driven Convolution Integration (ADCI) is engineered to restore the broken discretization invariance derived by convolutions, yet with the locally inductive bias preserved. In addition, we propose an Implicit Neural Functional Integration (INFI) to achieve cross domain interaction of spatial degradation functions, followed by the use of Galerkin-type Integration (GI) as a decoder to handle high-frequency information. Finally, the bonded activation functions are improved for the principle of continuous-discrete equivalence. Extensive experiments validate the superiority of our approach over cutting-edge methods. Notably, our proposal holds significantly better generalization on arbitrary scaling factors, yet requires only 0.07M parameters. Honghui Xu 0002, Wei Li 0034, Jiawei Jiang 0002, Zhi Liu 0009, Jianwei Zheng 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Breaking Information Isolation: Accelerating MRI via Inter-sequence Mapping and Progressive MaskingabstractDeep unfolding network (DUN) has shed new light on multi-sequence MRI reconstruction, providing both high interpretability and acceptable performance. However, current approaches still suffer from the plight of information isolation, i.e., learning features of multi-suquences individually and leaving the mask departed from model updating. In this work, we propose a new unfolding solution, namely Information-coupled MRI Acceleration (IMA), to address the isolation issue. Concretely, two specific mechanisms are presented. On the one hand, the latent connections across different sequences are explicitly molded via two auxiliary matrices. While the first matrix is meticulously engineered to assemble the spatial details, the second one hammers at capturing the depth information conditioned on the enriched channels. On the other hand, following a deep analysis on the non-uniform distribution in low- and high-frequency components of the given mask, we elaborate a new unfolding flow using a progressive masking scheme, featuring a dilation-contraction mechanism during forward propagation of successive stages. Massive experiments are conducted under various sampling patterns and acceleration rates, whose results demonstrate that, without any sophisticated architectures, our IMA outperforms the current cutting-edge methods both visually and numerically. Jianwei Zheng 0001, Xiaomin Yao, Guojiang Shen, Wei Li 0034, Jiawei Jiang 0002 |
AAAI | 4 |
| 2025 | LMO: Linear Mamba Operator for MRI ReconstructionabstractInterpretability and consistency have long been crucial factors in MRI reconstruction. While interpretability has been significantly innovated with the emerging deep unfolding networks, current solutions still suffer from inconsistency issues and produce inferior anatomical structures. Especially in out-of-distribution cases, e.g., when the acceleration rate (AR) varies, the generalization performance is often catastrophic. To counteract the dilemma, we propose an innovative Linear Mamba Operator (LMO) to ensure consistency and generalization, while still enjoying desirable interpretability. Theoretically, we argue that mapping between function spaces, rather than between signal instances, provides a solid foundation of high generalization. Technically, LMO achieves a good balance between global integration facilitated by a state space model that scans the whole function domain, and local integration engaged with an appealing property of continuous-discrete equivalence. On that basis, learning holistic features can be guaranteed, tapping the potential of maximizing data consistency. Quantitative and qualitative results demonstrate that LMO significantly outperforms other state-of-the-arts. More importantly, LMO is the unique model that, with AR changed, achieves retraining performance without retraining steps. Codes are available at https://github.com/ZhengJianwei2/LMO. Wei Li 0034, Jiawei Jiang 0002, Kaihao Yu, Jianwei Zheng 0001 |
CVPR | 1 |
| 2025 | Gradient Selection Tuning via Information BottleneckabstractPre-trained visual models enjoy strong representations, yet suffer from massive parameters to be shifted in downstream practices. Many parameter-efficient fine-tuning methods have been proposed, mostly requiring only 1% additional parameters to achieve comparable results. However, current solutions either consider all feature channels equally or detect saliencies with individual layer, leading to many redundancies reserved. To address current issues, this paper proposes a new parameter fine-tuning method named “Gradient Selection Tuning” (GST), which leverages gradients that are capable of capturing the cascading effects across successive channels. Instead of saliency detection, we turn to compress the redundancies for channel selection, since the computed gradient values enjoy much lower mutual information. With GST facilitated, we further elaborate an Information-Guided Adapter following information bottleneck theory, effectively performing parameter compression yet with task-specific features preserved. Experimental results demonstrate that our method outperforms the baseline methods by adding only 0.075M parameters to ViT-B backbone. On domain generalization, our proposal also enjoys strong performance in low-parameter scenarios. Xiaoxu Lin, Wei Li 0034, Ni Xu, Honghui Xu 0002, Jianwei Zheng 0001 |
ECAI | 2 |
| 2025 | Object-Level Control for Refined Structure and Appearance in Conditional Image SynthesisabstractRecent advances in pretrained diffusion models, particularly the FreeControl, have enabled fine-grained spatial control in text-to-image generation. However, FreeControl still suffers notable limitations in detail generation and appearance synthesis. With deep analysis performed, we reveal that inadequate feature representation in the early generation phases is the main cause of the insufficient structure elaboration. To address this issue, we dig into the temporal evolution of inverse attention features, then extract more expressive structure information as the input to the guidance function, ensuring formation integrity. Moreover, to tackle the surface degradation and ambiguity caused by the dual guidance of structure and appearance, we engineer the Adaptive Instance Normalization (AdaIN) mechanism into a latent space, rather than the typical feature space, during the intermediate generation stage. This improvement not only guarantees the close alignment between the generated image and structural reference but also significantly strengthens the appearance modeling capability and optimizes the texture representation of both foreground and background elements. Extensive experiments demonstrate that our proposal consistently outperforms existing baseline models across multiple metrics, including Self-sim, CLIP, and LPIPS. Both quantitative and qualitative results confirm that our approach achieves superior performance in terms of content consistency, visual quality, and detail preservation. Ni Xu, Wei Li 0034, Zhouchao Fu, Xiaoxu Lin, Jianwei Zheng 0001 |
ECAI | 2 |
| 2025 | Spatial-Spectral Fusion Neural OperatorabstractWith the rapid development of deep learning, spatial-spectral fusion (SSF) has emerged as an ideal alternative to traditional, costly hyperspectral image (HSI) acquisition methods. However, current solutions necessitate training and storing multiple models for different scaling factors. Besides, a meticulously designed network architecture to meet desirable performance often suffers from a severe computational burden. To counteract the dilemma, we propose SFNO, a lightweight spatial-spectral fusion neural operator for arbitrary-scale SSF. SFNO leverages approximation theory by embedding features from two degraded functions into a high-dimensional latent space, enabling efficient learning of basis functions. Kernel integration mechanisms are then used to approximate certain priors, followed by dimensionality reduction to generate high-resolution HSIs. Moreover, with the aid of discrete invariance property, we propose a new mechanism of progressive resampling (PR), which allows for the shrinkage of necessary spatial domain without any performance degradation. Extensive experiments on CAVE and Harvard datasets show that SFNO and its variant significantly improve performance, especially in out-of-domain fusion, requiring only 0.098M parameters and 0.966G FLOPs. Wei Li 0034, Jiawei Jiang 0002, Ni Xu, Yan Li 0083, Jianwei Zheng 0001 |
ICME | 1 |
| 2025 | SpecSolver: Solving Spatial-Spectral Fusion via Semantic TransformerabstractBy clustering pixels with locally similar values, superpixel-based approaches have shown great potential in processing hyperspectral images (HSI), thereby reducing the computational burden associated with large spatial dimensions. However, specific for spatial-spectral fusion (SSF), superpixel segmentation is inherently non-differentiable and irreversible; hence it is inapplicable. To address the issues, we propose a semantic transformer-based solver, namely SpecSolver, which is basically inspired by the benefits of superpixel-based approaches, yet with the inner mechanism completely improved. The core idea lies in learning the intrinsic semantic states of HSIs hidden behind discretized pixel representations. Specifically, we propose a new Semantic-Attention to adaptively split the image domain into a series of learnable slices of flexible shapes, where image pixels under similar semantic states will be ascribed to the same slice. By calculating attention to the Semantic-Superpixel tokens encoded from slices, SpecSolver can effectively capture intricate semantic correlations from the vast number of pixels, which also empowers the solver with an endogenous capacity for modeling different magnification scales and allows for efficient computation in linear complexity. On that basis, we elaborate a SpatialNet module, which extracts multiscale local spectral information, and a FreqNet module, which supplements global information, capturing subtle details and variations across different spectra. Experiments on two benchmark SSF datasets verify the state-of-the-art (SOTA) performance of the proposed method, both visually and quantitatively. Also, ablation studies validate the mentioned contributions. Wei Li 0034, Honghui Xu 0002, Jiawei Jiang 0002, Jianwei Zheng 0001 |
ACM Multimedia | 1 |
| 2025 | Arbitrary-scale Fusion Neural OperatorabstractSpatial-spectral fusion offers a promising alternative to expensive equipment in high-resolution hyperspectral (HrHs) imaging. However, training separate models for different scaling factors remains costly. To address this, we propose the Arbitrary-scale Fusion Neural Operator (AFNO), a lightweight solution for HrHs fusion across arbitrary scalings. Instead of entities, AFNO treats low-resolution hyperspectral (LrHs) and high-resolution multispectral (HrMs) images as functions and performs meticulously designed integrations as the mapping operator. The key components include Attention-Driven Convolution Integration (ADCI) to restore discretization invariance disrupted by convolutions, Implicit Neural Functional Integration (INFI) for cross-domain interaction of spatial degradations, and Galerkin-type Integration as a decoder for high-frequency details. Additionally, the bonded activation opeartor are improved for the principle of continuous-discrete equivalence. Extensive experiments validate the superiority of our approach over cutting-edge methods. Notably, AFNO holds significantly better generalization on arbitrary scaling factors, yet requiring only 0.07M parameters. Wei Li 0034, Honghui Xu 0002, Jiawei Jiang 0002, Zhi Liu 0009, Jianwei Zheng 0001 |
ACM Multimedia | 2 |
| 2025 | Solving Partial Differential Equations via Radon Neural OperatorabstractNeural operator is considered a popular data-driven alternative to traditional partial differential equation (PDE) solvers. However, most current solutions, whether fulfilling computations in frequency, Laplacian, and wavelet domains, all deviate far from the intrinsic PDE space. While with meticulous network architecture elaborated, the deviation often leads to biased accuracy. To address the issue, we open a new avenue that pioneers leveraging Radon transform to decompose the input space, finalizing a novel Radon neural operator (RNO) to solve PDEs in infinite-dimensional function space. Distinct from previous solutions, we project the input data into the sinogram domain, shrinking the multi-dimensional transformations to a reduced-dimensional counterpart and fitting compactly with the PDE space. Theoretically, we prove that RNO obeys a property of bilipschitz strongly monotonicity under diffeomorphism, providing deeper insights to guarantee the desired accuracy than typical discrete invariance or continuous-discrete equivalence. Within the sinogram domain, we further evidence that different angles contribute unequally to the overall space, thus engineering a reweighting technique to enable more effective PDE solutions. On that basis, a sinogram-domain convolutional layer is crafted, which operates on a fixed $\theta$-grid that is decoupled from the PDE space, further enjoying a natural guarantee of discrete invariance. Extensive experiments demonstrate that RNO sets new state-of-the-art (SOTA) scores across massive standard benchmarks, with superior generalization performance enjoyed. Code is available at <https://github.com/wenbin-lu/Radon-Neural-Operator>. Wenbin Lu, Junnan Xu, Wei Li 0034, Jianwei Zheng 0001 |
NeurIPS | 4 |
| 2025 | Semantic-Spatial Attention for Refined Object Placement in Text-to-Image SynthesisabstractSolely based on given prompts, text-guided diffusion models have enjoyed a unique capability in generating diverse and creative images. Nevertheless, the conveyance of image information through text presents a series of challenges, particularly in controlling the positioning of objects in synthesized images. Despite attempts of recent efforts in exploring alternative conditions, such as bounding box/mask-image pairs, the requirement of a substantial amount of paired data and time-consuming fine-tuning emerge as new issues. Given the observations that not only prompt-related cross-attention maps reveal the spatial arrangement and centroid positions of the objects, but also out-of-prompt markers enjoy rich semantic information, we thus engineer a weighted optimization loss. Specifically, three spatial sub-losses, namely inner box reinforcement loss, outer box attenuation loss, and centroid loss, are devised and seamlessly integrated into the sampling step of current vanilla diffusion models. Without any annotations of layout data required, the final approach runs in a training-free fashion. Extensive experiments with new performance scores demonstrate that our proposal not only successfully addresses the issue of object positioning but also boosts the capabilities of most current models, such as Stable Diffusion and GLIGEN, in high-quality synthesis and coverage of various concepts. Moreover, the proposed mechanism plays a plug-and-play role. Jianwei Zheng 0001, Ni Xu, Wei Li 0034, Jiawei Jiang 0002, Xiaoqin Zhang 0002 |
IEEE Trans. Multim. | 3 |
| 2024 | Alias-Free Mamba Neural OperatorabstractBenefiting from the booming deep learning techniques, neural operators (NO) are considered as an ideal alternative to break the traditions of solving Partial Differential Equations (PDE) with expensive cost.
Yet with the remarkable progress, current solutions concern little on the holistic function features--both global and local information-- during the process of solving PDEs.
Besides, a meticulously designed kernel integration to meet desirable performance often suffers from a severe computational burden, such as GNO with $O(N(N-1))$, FNO with $O(NlogN)$, and Transformer-based NO with $O(N^2)$.
To counteract the dilemma, we propose a mamba neural operator with $O(N)$ computational complexity, namely MambaNO.
Functionally, MambaNO achieves a clever balance between global integration, facilitated by state space model of Mamba that scans the entire function, and local integration, engaged with an alias-free architecture. We prove a property of continuous-discrete equivalence to show the capability of
MambaNO in approximating operators arising from universal PDEs to desired accuracy. MambaNOs are evaluated on a diverse set of benchmarks with possibly multi-scale solutions and set new state-of-the-art scores, yet with fewer parameters and better efficiency. Jianwei Zheng 0001, Wei Li 0034, Ni Xu, Xiaoxu Lin, Xiaoqin Zhang 0002 |
NeurIPS | 2 |
| 2015 | Multi-Modality Tracker Aggregation: From Generative to Discriminative
Xiaoqin Zhang 0002, Wei Li 0034, Mingyu Fan, Di Wang 0008, Xiuzi Ye |
IJCAI | 2 |
| 2015 | Robust hand tracking via novel multi-cue integration
Xiaoqin Zhang 0002, Wei Li 0034, Xiuzi Ye, Stephen J. Maybank |
Neurocomputing | 2 |
| 2015 | Single and Multiple Object Tracking Using a Multi-Feature Joint Sparse RepresentationabstractIn this paper, we propose a tracking algorithm based on a multi-feature joint sparse representation. The templates for the sparse representation can include pixel values, textures, and edges. In the multi-feature joint optimization, noise or occlusion is dealt with using a set of trivial templates. A sparse weight constraint is introduced to dynamically select the relevant templates from the full set of templates. A variance ratio measure is adopted to adaptively adjust the weights of different features. The multi-feature template set is updated adaptively. We further propose an algorithm for tracking multi-objects with occlusion handling based on the multi-feature joint sparse reconstruction. The observation model based on sparse reconstruction automatically focuses on the visible parts of an occluded object by using the information in the trivial templates. The multi-object tracking is simplified into a joint Bayesian inference. The experimental results show the superiority of our algorithm over several state-of-the-art tracking algorithms. Weiming Hu 0004, Wei Li 0034, Xiaoqin Zhang 0002, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Block covariance based l1 tracker with a subtle template dictionary
Xiaoqin Zhang 0002, Wei Li 0034, Weiming Hu 0004, Haibin Ling, Stephen J. Maybank |
Pattern Recognit. | 2 |
| 2013 | Active Contour-Based Visual Tracking by Integrating Colors, Shapes, and MotionsabstractIn this paper, we present a framework for active contour-based visual tracking using level sets. The main components of our framework include contour-based tracking initialization, color-based contour evolution, adaptive shape-based contour evolution for non-periodic motions, dynamic shape-based contour evolution for periodic motions, and the handling of abrupt motions. For the initialization of contour-based tracking, we develop an optical flow-based algorithm for automatically initializing contours at the first frame. For the color-based contour evolution, Markov random field theory is used to measure correlations between values of neighboring pixels for posterior probability estimation. For adaptive shape-based contour evolution, the global shape information and the local color information are combined to hierarchically evolve the contour, and a flexible shape updating model is constructed. For the dynamic shape-based contour evolution, a shape mode transition matrix is learnt to characterize the temporal correlations of object shapes. For the handling of abrupt motions, particle swarm optimization is adopted to capture the global motion which is applied to the contour in the current frame to produce an initial contour in the next frame. Weiming Hu 0004, Wei Li 0034, Wenhan Luo, Xiaoqin Zhang 0002, Stephen J. Maybank |
IEEE Trans. Image Process. | 3 |
| 2011 | Efficient block-division model for robust multiple object trackingabstractTracking multiple objects under occlusion is one of the most challenging issues in computer vision. Occlusion results in mistaken match when finding the most similar candidate. Adapting to the change of objects is essential for tracking as objects often undergo intrinsic changes, but noise is unavoidably introduced during updating of the object, and this further confuses the tracker. In order to address these problems, a block-division appearance model is introduced to efficiently handle occlusion. In this model, spatial information is introduced to avoid the mistaken match between object and candidate. Based on this model, a selective updating strategy is proposed to incrementally learn the change of the object, avoiding introducing noise when updating. At the same time occlusion is deduced by monitoring the variation of each block. Experimental results in various videos validate the effectiveness of our algorithm in tracking multiple objects under occlusion. Wenhan Luo, Xiaoqin Zhang 0002, Yang Liu 0020, Xi Li 0001, Weiming Hu 0004, Wei Li 0034 |
ICASSP | 6 |
| 2011 | Robust visual tracking via transfer learningabstractIn this paper, we propose a boosting based tracking framework using transfer learning. To deal with complex appearance variations, the proposed tracking framework tries to utilize discriminative information from previous frames to conduct the tracking task in the current frame, and thus transfers some prior knowledge from the previous source data domain to the current target data domain, resulting in a high discriminative tracker for distinguishing the object from the background. The proposed tracking system has been tested on several challenging sequences. Experimental results demonstrate the effectiveness of the proposed tracking framework. Wenhan Luo, Xi Li 0001, Wei Li 0034, Weiming Hu 0004 |
ICIP | 3 |
| 2010 | Horror Image Recognition Based on Emotional Attention
Bing Li 0001, Weiming Hu 0004, Weihua Xiong, Ou Wu 0001, Wei Li 0034 |
ACCV (2) | 5 |
| 2010 | Occlusion Handling with ℓ1-Regularized Sparse Reconstruction
Wei Li 0034, Bing Li 0001, Xiaoqin Zhang 0002, Weiming Hu 0004, Hanzi Wang, Guan Luo |
ACCV (4) | 1 |
| 2010 | Local Outlier Detection Based on Kernel RegressionabstractOutlier detection keeps an important and attractive task of the knowledge discovery in databases. In this paper, a novel approach named Multi-scale Local Kernel Regression is proposed. It transfers the unsupervised learning of outlier detection to the classic non-parameter regression learning. Through preprocessing the original data by the basic local density-based method, it adopts the local kernel regression estimator in the multiple scale neighborhoods to determine outliers. Experiments on several real life data sets demonstrate that this approach is promising in detection performance. Weiming Hu 0004, Wei Li 0034, Zhongfei Zhang, Ou Wu 0001 |
ICPR | 3 |
| 2010 | Discriminative Level Set for Contour TrackingabstractConventional contour tracking algorithms with level set often use generative models to construct the energy function. For tracking through cluttered and noisy background, however, a generative model may not be discriminative enough. In this paper we integrate the discriminative methods into a level set framework when constructing the level set energy function. We train a set of weak classifiers to distinguish the object from the background. Each weak classifier is designed to select the most discriminative feature space and integrated via AdaBoost according to their training errors. We also introduce a novel interaction term to explore the correlation between pixels near the object edge. This term together with the discriminative model both enhance the discriminative power of the level set. The experimental results show that the contour tracked by our approach is more accurate than the conventional algorithms with the generative model. Our algorithm successfully tracks the object contour even in a cluttered environment. Wei Li 0034, Xiaoqin Zhang 0002, Weiming Hu 0004, Haibin Ling |
ICPR | 1 |
| 2009 | Contour tracking with abrupt motionabstractTraditional contour tracking methods can not handle abrupt motion or low frame rate video. This is because the basis of the traditional tracking lies in the assumption that the motion is smooth between consecutive frames. However, the abrupt motion destroys the foundation of this assumption. In this paper, we integrate the stochastic search into the level set evolution to reinstitute the continuity. Our approach can be viewed as a two-layer hierarchical level set-based tracking framework in which Particle Swarm Optimization (PSO) and level set evolution are fused seamlessly. In the first layer, the PSO is adopted to capture the global motion of the object. The coarse contour is obtained by applying the global motion to the contour in the previous frame. For the second layer, the level set evolution based on the coarse contour is carried out to track the local deformation, which results in the actual contour. The promising experimental results for numerous real videos reveal the effectiveness of our approach. Wei Li 0034, Xiaoqin Zhang 0002, Weiming Hu 0004 |
ICIP | 1 |