Weifu Li

dblp:198/9625 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
16since 2021 · last 2025
0000-0002-8444-9782ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Error Analysis Affected by Heavy-Tailed Gradients for Non-Convex Pairwise Stochastic Gradient Descent
abstract
In recent years, there have been a growing number of works studying the generalization properties of stochastic gradient descent (SGD) from the perspective of algorithmic stability. However, few of them devote to simultaneously studying the generalization and optimization for the non-convex setting, especially pairwise SGD with heavy-tailed gradient noise. This paper considers the impact of the heavy-tailed gradient noise obeying sub-Weibull distribution on the stability-based learning guarantees for non-convex pairwise SGD by investigating its generalization and optimization jointly. Specifically, based on two novel pairwise uniform model stability tools, we firstly bound the generalization error of pairwise SGD in the general non-convex setting after bridging the quantitative relationships between stability and generalization error. Then, we further consider the practical heavy-tailed sub-Weibull gradient noise condition to establish a refined generalization bound without the bounded gradient condition. Finally, sharper error bounds for generalization and optimization are built by introducing the gradient dominance condition. Comparing these results reveals that sub-Weibull gradient noise brings some positive dependencies on the heavy-tailed strength for generalization and optimization. Furthermore, we extend our analysis to the corresponding pairwise minibatch SGD and derive the first stability-based near-optimal generalization and optimization bounds which are consistent with many empirical observations.
Hong Chen 0004, Bin Gu 0001, Yingjie Wang 0007, Weifu Li
AAAI6
2025 Pairwise Generalized Importance Weighting for Metric Learning Under Distribution Shift
Richeng Zhou, Weifu Li
ICPADS4
2025 Trajectory-Dependent Generalization Bounds for Pairwise Learning with φ-mixing Samples
abstract
Recently, the mathematical tool from fractal geometry (i.e., fractal dimension) has been employed to investigate optimization trajectory-dependent generalization ability for some pointwise learning models with independent and identically distributed (i.i.d.) observations. This paper goes beyond the limitations of pointwise learning and i.i.d. samples, and establishes generalization bounds for pairwise learning with uniformly strong mixing samples. The derived theoretical results fill the gap of trajectory-dependent generalization analysis for pairwise learning, and can be applied to wide learning paradigms, e.g., metric learning, ranking and gradient learning. Technically, our framework brings concentration estimation with Rademacher complexity and trajectory-dependent fractal dimension together in a coherent way for felicitous learning theory analysis. In addition, the efficient computation of fractal dimension can be guaranteed for random algorithms (e.g., stochastic gradient descent algorithm for deep neural networks) by bridging topological data analysis tools and the trajectory-dependent fractal dimension.
Hong Chen 0004, Weifu Li, Tieliang Gong, Hao Deng 0017, Yulong Wang 0002
IJCAI3
2025 The consistency analysis of gradient learning under independent covariate shift
Chi Xiao 0002, Weifu Li
Neurocomputing4
2025 TSGaussian: Semantic and depth-guided Target-Specific Gaussian Splatting from sparse views
Zehan Bao, Hong Chen 0004, Yaohui Chen 0002, Weifu Li
Image Vis. Comput.6
2025 Generalization Bounds of Deep Neural Networks With τ-Mixing Samples
abstract
Deep neural networks (DNNs) have shown an astonishing ability to unlock the complicated relationships among the inputs and their responses. Along with empirical successes, some approximation analysis of DNNs has also been provided to understand their generalization performance. However, the existing analysis depends heavily on the independently identically distribution (i.i.d.) assumption of observations, which may be too ideal and often violated in real-world applications. To relax the i.i.d. assumption, this article develops the covering number-based concentration estimation to establish generalization bounds of DNNs with $\tau $ -mixing samples, where the dependency between samples is much general including $\alpha $ -mixing process as a special case. By assigning a specific parameter value to the $\tau $ -mixing process, our results are consistent with the existing convergence analysis under the i.i.d. case. Experiments on simulated data validate the theoretical findings.
Yaohui Chen 0002, Weifu Li, Yingjie Wang 0007, Bin Gu 0001, Feng Zheng 0001, Hong Chen 0004
IEEE Trans. Neural Networks Learn. Syst.3
2024 Gradient Learning With the Mode-Induced Loss: Consistency Analysis and Applications
abstract
Variable selection methods aim to select the key covariates related to the response variable for learning problems with high-dimensional data. Typical methods of variable selection are formulated in terms of sparse mean regression with a parametric hypothesis class, such as linear functions or additive functions. Despite rapid progress, the existing methods depend heavily on the chosen parametric function class and are incapable of handling variable selection for problems where the data noise is heavy-tailed or skewed. To circumvent these drawbacks, we propose sparse gradient learning with the mode-induced loss (SGLML) for robust model-free (MF) variable selection. The theoretical analysis is established for SGLML on the upper bound of excess risk and the consistency of variable selection, which guarantees its ability for gradient estimation from the lens of gradient risk and informative variable identification under mild conditions. Experimental analysis on the simulated and real data demonstrates the competitive performance of our method over the previous gradient learning (GL) methods.
Hong Chen 0004, Youcheng Fu, Weifu Li, Yicong Zhou, Feng Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 On the Stability and Generalization of Triplet Learning
abstract
Triplet learning, i.e. learning from triplet data, has attracted much attention in computer vision tasks with an extremely large number of categories, e.g., face recognition and person re-identification. Albeit with rapid progress in designing and applying triplet learning algorithms, there is a lacking study on the theoretical understanding of their generalization performance. To fill this gap, this paper investigates the generalization guarantees of triplet learning by leveraging the stability analysis. Specifically, we establish the first general high-probability generalization bound for the triplet learning algorithm satisfying the uniform stability, and then obtain the excess risk bounds of the order O(log(n)/(√n) ) for both stochastic gradient descent (SGD) and regularized risk minimization (RRM), where 2n is approximately equal to the number of training samples. Moreover, an optimistic generalization bound in expectation as fast as O(1/n) is derived for RRM in a low noise case via the on-average stability analysis. Finally, our results are applied to triplet metric learning to characterize its theoretical underpinning.
Hong Chen 0004, Bin Gu 0001, Weifu Li, Tieliang Gong, Feng Zheng 0001
AAAI5
2023 Stepdown SLOPE for Controlled Feature Selection
abstract
Sorted L-One Penalized Estimation (SLOPE) has shown the nice theoretical property as well as empirical behavior recently on the false discovery rate (FDR) control of high-dimensional feature selection by adaptively imposing the non-increasing sequence of tuning parameters on the sorted L1 penalties. This paper goes beyond the previous concern limited to the FDR control by considering the stepdown-based SLOPE in order to control the probability of k or more false rejections (k-FWER) and the false discovery proportion (FDP). Two new SLOPEs, called k-SLOPE and F-SLOPE, are proposed to realize k-FWER and FDP control respectively, where the stepdown procedure is injected into the SLOPE scheme. For the proposed stepdown SLOPEs, we establish their theoretical guarantees on controlling k-FWER and FDP under the orthogonal design setting, and also provide an intuitive guideline for the choice of regularization parameter sequence in much general setting. Empirical evaluations on simulated data validate the effectiveness of our approaches on controlled feature selection and support our theoretical findings.
Jingxuan Liang, Weifu Li
AAAI4
2023 Stability-Based Generalization Analysis for Mixtures of Pointwise and Pairwise Learning
abstract
Recently, some mixture algorithms of pointwise and pairwise learning (PPL) have been formulated by employing the hybrid error metric of “pointwise loss + pairwise loss” and have shown empirical effectiveness on feature selection, ranking and recommendation tasks. However, to the best of our knowledge, the learning theory foundation of PPL has not been touched in the existing works. In this paper, we try to fill this theoretical gap by investigating the generalization properties of PPL. After extending the definitions of algorithmic stability to the PPL setting, we establish the high-probability generalization bounds for uniformly stable PPL algorithms. Moreover, explicit convergence rates of stochastic gradient descent (SGD) and regularized risk minimization (RRM) for PPL are stated by developing the stability analysis technique of pairwise learning. In addition, the refined generalization bounds of PPL are obtained by replacing uniform stability with on-average stability.
Hong Chen 0004, Bin Gu 0001, Weifu Li
AAAI5
2023 VISN: virus instance segmentation network for TEM images using deep attention transformer
abstract
The identification of viruses from negative staining transmission electron microscopy (TEM) images has mainly depended on experienced experts. Recent advances in artificial intelligence have enabled virus recognition using deep learning techniques. However, most of the existing methods only perform virus classification or semantic segmentation, and few studies have addressed the challenge of virus instance segmentation in TEM images. In this paper, we focus on the instance segmentation of severe acute respiratory syndrome coronavirus type 2 (SARS-CoV-2) and other respiratory viruses and provide experts with more effective information about viruses. We propose an effective virus instance segmentation network based on the You Only Look At CoefficienTs backbone, which integrates the Swin Transformer, dense connections and the coordinate-spatial attention mechanism, to identify SARS-CoV-2, H1N1 influenza virus, respiratory syncytial virus, Herpes simplex virus-1, Human adenovirus type 5 and Vaccinia virus. We also provide a public TEM virus dataset and conduct extensive comparative experiments. Our method achieves a mean average precision score of 83.8 and F1 score of 0.920, outperforming other state-of-the-art instance segmentation algorithms. The proposed automated method provides virologists with an effective approach for recognizing and identifying SARS-CoV-2 and assisting in the diagnosis of viruses. Our dataset and code are accessible at https://github.com/xiaochiHNU/Virus-Instance-Segmentation-Transformer-Network.
Chi Xiao 0002, Shenrong Yang, Minxin Heng, Junyi Su, Jingdong Song, Weifu Li
Briefings Bioinform.8
2022 Error-Based Knockoffs Inference for Controlled Feature Selection
abstract
Recently, the scheme of model-X knockoffs was proposed as a promising solution to address controlled feature selection under high-dimensional finite-sample settings. However, the procedure of model-X knockoffs depends heavily on the coefficient-based feature importance and only concerns the control of false discovery rate (FDR). To further improve its adaptivity and flexibility, in this paper, we propose an error-based knockoff inference method by integrating the knockoff features, the error-based feature importance statistics, and the stepdown procedure together. The proposed inference procedure does not require specifying a regression model and can handle feature selection with theoretical guarantees on controlling false discovery proportion (FDP), FDR, or k-familywise error rate (k-FWER). Empirical evaluations demonstrate the competitive performance of our approach on both simulated and real data.
Xuebin Zhao, Hong Chen 0004, Yingjie Wang 0007, Weifu Li, Tieliang Gong, Yulong Wang 0002, Feng Zheng 0001
AAAI4
2022 Distribution-dependent feature selection for deep neural networks
Xuebin Zhao, Weifu Li, Hong Chen 0004, Yingjie Wang 0007, Vijay John
Appl. Intell.2
2022 Hardware Implementation of Hierarchical Temporal Memory Algorithm
abstract
Hierarchical temporal memory (HTM) is an un-supervised machine learning algorithm that can learn both spatial and temporal information of input. It has been successfully applied to multiple areas. In this paper, we propose a multi-level hierarchical ASIC implementation of HTM, referred to as processor core, to support both spatial and temporal pooling. To improve the unbalanced workload in HTM, the proposed design provides different mapping methods for the spatial and temporal pooling, respectively. In the proposed design, we implement a distributed memory system by assigning one dedicated memory bank to each level of hierarchy to improve the memory bandwidth utilization efficiency. Finally, the hot-spot operations are optimized using a series of customized units. Regarding scalability, we propose a ring-based network consisting of multiple processor cores to support a larger HTM network. To evaluate the performance of our proposed design, we map an HTM network that includes 2,048 columns and 65,536 cells on both the proposed design and NVIDIA Tesla K40c GPU using the KTH database as input. The latency and power of the proposed design is 6.04 ms and 4.1 W using GP 65 nm technology. Compared to the equivalent GPU implementation, the latency and power is improved 12.45× and 57.32×, respectively.
Weifu Li, Paul D. Franzon, Sumon Dey, Joshua Schabel
ACM J. Emerg. Technol. Comput. Syst.1
2021 A Scalable Cluster-based Hierarchical Hardware Accelerator for a Cortically Inspired Algorithm
abstract
This article describes a scalable, configurable and cluster-based hierarchical hardware accelerator through custom hardware architecture for Sparsey, a cortical learning algorithm. Sparsey is inspired by the operation of the human cortex and uses a Sparse Distributed Representation to enable unsupervised learning and inference in the same algorithm. A distributed on-chip memory organization is designed and implemented in custom hardware to improve memory bandwidth and accelerate the memory read/write operations for synaptic weight matrices. Bit-level data are processed from distributed on-chip memory and custom multiply-accumulate hardware is implemented for binary and fixed-point multiply-accumulation operations. The fixed-point arithmetic and fixed-point storage are also adapted in this implementation. At 16 nm, the custom hardware of Sparsey achieved an overall 24.39× speedup, 353.12× energy efficiency per frame, and 1.43× reduction in silicon area against a state-of-the-art GPU.
Sumon Dey, Lee Baker, Joshua Schabel, Weifu Li, Paul D. Franzon
ACM J. Emerg. Technol. Comput. Syst.4
2021 Geolocation Error Estimation and Correction on Long-Term MWRI Data
abstract
Due to the limitation of the satellite attitude measurement accuracy and the system servo control error of the payload scanning mechanism, an optimal use of Micro-Wave Radiation Imager (MWRI) observations requires high geolocation accuracy. In the operational system, the MWRI geolocation accuracy reaches 1 pixel, and there still exists room for improvement. In this article, we improve upon the coastline inflection point method (CIM) and propose to assign the accurate correspondence by employing a nonrigid point set registration method. First, the method identifies a set of latent variables to recognize outliers and then applies nonparametric geometric constraints to the correspondence asa prioridistribution. Second, the maximuma posteriori(MAP) estimation is applied by the expectation–maximization (EM) algorithm to obtain correct inliers. The comparison with other methods demonstrates that the proposed method can provide more accurate estimation of geolocation bias. In addition, the pixel error and changes in spacecraft attitude with the long-term geolocation data in FY-3C MWRI before and after correction were analyzed during the period from April 1 to August 30, 2018. The results have shown that the geolocation errors are reduced from [0.50, 0.60] pixels to [0.20, 0.33] pixels in the along- and cross-track directions after the attitude correction. In addition, the reduction of the standard deviation shows that the geolocation quality of MWRI is improved.
Weifu Li, Jiangtao Peng, Lijun Shen, Hua Han 0001, Peng Zhang 0024, Lei Yang 0035
IEEE Trans. Geosci. Remote. Sens.2
2020 Enhancing Depth Quality of Stereo Vision using Deep Learning-based Prior Information of the Driving Environment
abstract
Generation of high density depth values of the driving environment is indispensable for autonomous driving. Stereo vision is one of the practical and effective methods to generate these depth values. However, the accuracy of the stereo vision is limited by texture-less regions, such as sky and road areas, and repeated patterns in the image. To overcome these problems, we propose to enhance the stereo generated depth by incorporating prior information of the driving environment. Prior information, generated by deep learning-based U-Net model, is utilized in a novel post-processing mathematical framework to refine the stereo generated depth. The proposed mathematical framework is formulated as an optimization problem, which refines the errors due to texture-less regions and repeated patterns. Owing to its mathematical formulation, the post-processing framework is not a black-box and is explainable, and can be readily utilized for depth maps generated by any stereo vision algorithm. The proposed framework is qualitatively validated on the acquired dataset and KITTI dataset. The results obtained show that the proposed framework improves the stereo depth generation accuracy.
Weifu Li, Vijay John, Seiichi Mita
ICPR1
2020 SiamBOMB: A Real-time AI-based System for Home-cage Animal Tracking, Segmentation and Behavioral Analysis
abstract
Biologists often need to handle numerous video-based home-cage animal behavior analysis tasks that require massive workloads. Therefore, we develop an AI-based multi-species tracking and segmentation system, SiamBOMB, for real-time and automatic home-cage animal behavioral analysis. In this system, a background-enhanced Siamese-based network with replaceable modular design ensures the flexibility and generalizability of the system, and a user-friendly interface makes it convenient to use for biologists. This real-time AI system will effectively reduce the burden on biologists.
Xi Chen 0031, Hao Zhai 0003, Danqian Liu, Weifu Li, Chaoyue Ding, Qiwei Xie, Hua Han 0001
IJCAI4
2020 A New Geolocation Error Estimation Method in MWRI Data Aboard FY3 Series Satellites
abstract
Known as input in the numerical weather prediction (NWP) models, microwave radiation imager (MWRI) data have been widely distributed to the user community. Nevertheless, the current operational geolocation accuracy is still on the pixel scale due to the presence of geolocation uncertainty. In this letter, we propose a new method to estimate the geolocation errors in MWRI data. Compared to the traditional coastline inflection method (CIM), the proposed method has two innovations. First, we establish a surface fitting interpolation model by involving more observations to detect the coastline. Second, we employ the iterative closest point (ICP) algorithm to determine the correspondences between the detected coastline and the actual coastline. Simulated experimental results demonstrate that the proposed method can provide a more accurate geolocation error estimation than the CIM. By applying our method, we have processed an MWRI data set from January 1 to February 28 in 2016. The experimental results have shown that the operational FY-3C MWRI geolocation errors are 0.4813 and 0.4909 pixels in the along-track and cross-track directions, respectively, which can be significantly reduced to 0.1299 and 0.1497 pixels after the attitude correction. It means that the geolocation accuracy has an average improvement up to 70%.
Weifu Li, Xinghui Zhao, Jiangtao Peng, Zhicheng Luo, Lijun Shen, Hua Han 0001, Peng Zhang 0024, Lei Yang 0035
IEEE Geosci. Remote. Sens. Lett.1
2019 ℓ0 Sparse Approximation of Coastline Inflection Method on FY-3C MWRI Data
abstract
The microwave radiation imager (MWRI) located onboard the FengYun-3C (FY-3C) satellite provides a considerable amount of critical information for numerical weather predictions. Obtaining accurate geolocation results from the FY-3C MWRI data is of great importance. In this letter, we improve the traditional coastline inflection method (CIM) and propose an$\ell _{0}$sparse approximation model for geolocation error estimation and correction. Specifically, we propose using the jump point of the step function to estimate the true coastline point. This approach can characterize the geolocation errors more accurately than the CIM, which further improves the geolocation accuracy. In the theoretical part, we provide a complete solution to obtain the step function through an iterative blind deconvolution. For a practical use, we demonstrate the effectiveness of the proposed method for geolocation error estimation through quantitative results obtained on the FY-3C MWRI data. The experimental results show that the proposed method can achieve an improvement of up to 33.33% in the standard deviation of geolocation errors (approximately 0.00030) compared to the traditional CIM (approximately 0.00045). Furthermore, we also apply the proposed method to the FY-3C satellite and improve the geolocation accuracy of the MWRI data through geolocation error correction.
Weifu Li, Zhicheng Luo, Chengbao Liu, Lijun Shen, Qiwei Xie, Hua Han 0001, Lei Yang 0035
IEEE Geosci. Remote. Sens. Lett.1
2018 Effective automated pipeline for 3D reconstruction of synapses based on deep learning
abstract
BACKGROUND: The locations and shapes of synapses are important in reconstructing connectomes and analyzing synaptic plasticity. However, current synapse detection and segmentation methods are still not adequate for accurately acquiring the synaptic connectivity, and they cannot effectively alleviate the burden of synapse validation. RESULTS: We propose a fully automated method that relies on deep learning to realize the 3D reconstruction of synapses in electron microscopy (EM) images. The proposed method consists of three main parts: (1) training and employing the faster region convolutional neural networks (R-CNN) algorithm to detect synapses, (2) using the z-continuity of synapses to reduce false positives, and (3) combining the Dijkstra algorithm with the GrabCut algorithm to obtain the segmentation of synaptic clefts. Experimental results were validated by manual tracking, and the effectiveness of our proposed method was demonstrated. The experimental results in anisotropic and isotropic EM volumes demonstrate the effectiveness of our algorithm, and the average precision of our detection (92.8% in anisotropy, 93.5% in isotropy) and segmentation (88.6% in anisotropy, 93.0% in isotropy) suggests that our method achieves state-of-the-art results. CONCLUSIONS: Our fully automated approach contributes to the development of neuroscience, providing neurologists with a rapid approach for obtaining rich synaptic statistics.
Chi Xiao 0002, Weifu Li, Hao Deng 0006, Xi Chen 0031, Qiwei Xie, Hua Han 0001
BMC Bioinform.2
2018 Learning With Coefficient-Based Regularized Regression on Markov Resampling
abstract
Big data research has become a globally hot topic in recent years. One of the core problems in big data learning is how to extract effective information from the huge data. In this paper, we propose a Markov resampling algorithm to draw useful samples for handling coefficient-based regularized regression (CBRR) problem. The proposed Markov resampling algorithm is a selective sampling method, which can automatically select uniformly ergodic Markov chain (u.e.M.c.) samples according to transition probabilities. Based on u.e.M.c. samples, we analyze the theoretical performance of CBRR algorithm and generalize the existing results on independent and identically distributed observations. To be specific, when the kernel is infinitely differentiable, the learning rate depending on the sample size $m$ can be arbitrarily close to $\mathcal {O}(m^{-1})$ under a mild regularity condition on the regression function. The good generalization ability of the proposed method is validated by experiments on simulated and real data sets.
Luoqing Li, Weifu Li, Bin Zou 0002, Yulong Wang 0002, Yuan Yan Tang, Hua Han 0001
IEEE Trans. Neural Networks Learn. Syst.2