Bo Liu 0011

dblp:58/2670-11 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0002-3482-6930ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 NEDA-Net: A neighborhood-enhanced deformable attention network for fMRI brain disease classification
Bo Liu 0011
Expert Syst. Appl.3
2026 Explainable CNN filter pruning based on shapley value approximation in view of game theory
Tongtong Yuan, Bo Liu 0011, Yinan Tang
Expert Syst. Appl.3
2026 Adaptive metric for knowledge distillation by deep Bregman divergence
Tongtong Yuan, Bo Liu 0011, Yinan Tang
Neural Networks3
2026 Spurious Local Minima Provably Exist for Deep CNNs: Theory and Application
abstract
In this article, we prove that a general family of spurious local minima exist in the loss landscape of deep convolutional neural networks (CNNs) with strictly convex loss functions and ReLU activations. Our construction of spurious local minima is general and applies to CNNs with arbitrary architectures. We construct a local minimum $\theta $ at first, and then construct another point $\theta ^{\prime } $ in parameter space with the same empirical risk as $\theta $ . Data samples are split into some groups such that each group behaves differently under the perturbation around $\theta ^{\prime } $ to produce a lower empirical risk. We tackle the challenges caused by convolutional layers in the construction. We show that a differentiation of data samples is always possible somewhere in the feature maps, and despite network parameters being tied in each feature map, our perturbation scheme only affects the output of a single or a few neurons for a group of data samples. We then give an example of nontrivial spurious local minimum in which multiple activation patterns are explicitly constructed. Finally, based on our construction of spurious local minima, we design a deterministic optimization method to escape local minima that is applicable to CNNs, ResNets, MLPs, and transformers. Experimental results on CIFAR-10, CIFAR-100, and ImageNet-1k datasets verify our theoretical findings and show that our optimization method outperforms SGD or Adam in accuracy (by 0.27% on average) consistently on all these architectures and datasets.
Bo Liu 0011, Keyi Fu, Tongtong Yuan, Shen Geng
IEEE Trans. Neural Networks Learn. Syst.1
2025 Layer-Wise Vision Injection With Disentangled Attention for Efficient LVLMs
Xuange Zhang, Dengjie Li, Bo Liu 0011, Zenghao Bao, Baisong Yang, Zhongying Liu, Tongtong Yuan
ICCV3
2025 Temperature Inversion guided Attention for Infrared Small Target Detection
abstract
Targets in infrared images typically lack distinct texture and color features, posing significant challenges for infrared small target detection. Although existing methods have made some progress, they still have limitations in terms of interpretability and are unable to provide sufficient theoretical grounds to explain their effectiveness. Moreover, these methods generally overlook the temperature information in the imaging process, failing to fully exploit the subtle thermal differences in infrared images, which consequently limits detection accuracy. To this end, we propose a novel temperature-aware Transformer network (TATransNet). By introducing temperature information, it allows the model to be more focused on the target area. Specifically, inspired by the temperature inversion technique in remote sensing, we propose a learnable temperature inversion attention module (TIAM), which simulates the inverse process of infrared imaging to perceive the temperature distribution across feature maps of different scales, thereby enhancing the distinguishability between foreground and background. We further integrate TIAM with Swin Transformer to construct a Transformer-CNN hybrid encoder unit, termed temperature-aware Transformer block (TATB). By stacking multiple TATBs, both global context and local temperature features can be captured simultaneously, thereby strengthening the representation of dim and small targets effectively. Experimental results on five benchmark datasets, i.e., Small-ExtIRShip, Small-SSDD, NUAA-SIRST, IRSTD-1K, and IHAST demonstrate that our method not only significantly improves detection accuracy and reduces false alarm rates, but also demonstrates strong generalization capability, enabling effective transfer to other approaches with different types of encoders.
Ting Zhang 0012, Zhaoying Liu, Bo Liu 0011, Changming Sun
IJCNN4
2025 Rethinking cell-based neural architecture search: A theoretical perspective
Bo Liu 0011, Huiwen Zhao, Tongtong Yuan, Ting Zhang 0012, Zhaoying Liu
Neural Networks1
2025 Surveillance Video-and-Language Understanding: From Small to Large Multimodal Models
abstract
Surveillance videos play a crucial role in public security. However, current tasks related to surveillance videos primarily focus on classifying and localizing anomalous events. Despite achieving notable performance, existing methods are restricted to detecting and classifying predefined events and lack satisfactory semantic understanding. To tackle this challenge, we introduce a novel research avenue focused on Video-and-Language Understanding for surveillance (VALU), and construct the first multimodal surveillance video dataset. We manually annotate the real-world surveillance dataset UCF-Crime with fine-grained event content and timing. Our newly annotated dataset, UCA (UCF-Crime Annotation), contains 23,542 sentences, with an average length of 20 words, and its annotated videos are as long as 110.7 hours. Moreover, we evaluate SOTA models on five multimodal tasks using this newly created dataset, establishing new baselines for surveillance VALU, from small to large models. Our experiments reveal that mainstream models, which perform well on previously public datasets, exhibit poor performance on surveillance video, highlighting new challenges in surveillance VALU. In addition to conducting baseline experiments to compare the performance of existing models, we also propose novel methods for multimodal anomaly detection tasks and finetune multimodal large language model models using our dataset. All the experiments highlight the necessity of constructing this multimodal dataset to advance surveillance AI. Upon the experimental results mentioned above, we conduct further in-depth analysis and discussion. The dataset and codes are provided athttps://xuange923.github.io/Surveillance-Video-Understanding.
Tongtong Yuan, Xuange Zhang, Bo Liu 0011, Zhenzhen Jiao
IEEE Trans. Circuits Syst. Video Technol.3
2024 Towards Surveillance Video-and-Language Understanding: New Dataset, Baselines, and Challenges
abstract
Surveillance videos are important for public security. However, current surveillance video tasks mainly focus on classifying and localizing anomalous events. Existing methods are limited to detecting and classifying the predefined events with unsatisfactory semantic understanding, although they have obtained considerable performance. To address this issue, we propose a new research direction of surveillance video-and-language understanding (VALU), and construct the first multimodal surveillance video dataset. We manually annotate the real-world surveillance dataset UCF-Crime with fine-grained event content and timing. Our newly annotated dataset, UCA (UCF-Crime Annotation)11The dataset is provided at https://xuange923.github.io/Surveillance-Video-Understanding., contains 23,542 sentences, with an average length of 20 words, and its annotated videos are as long as 110.7 hours. Furthermore, we benchmark SOTA models for four multimodal tasks on this newly created dataset, which serve as new baselines for surveillance VALU. Through experiments, we find that mainstream models used in previously public datasets perform poorly on surveillance video, demonstrating new challenges in surveillance VALU. We also conducted experiments on multimodal anomaly detection. These results demonstrate that our multimodal surveillance learning can improve the performance of anomaly detection. All the experiments highlight the necessity of constructing this dataset to advance surveillance AI.
Tongtong Yuan, Xuange Zhang, Bo Liu 0011, Zhenzhen Jiao
CVPR4
2024 ARPruning: An automatic channel pruning based on attention map ranking
Tongtong Yuan, Zulin Li, Bo Liu 0011, Yinan Tang
Neural Networks3
2024 Spurious Local Minima are Common for Deep Neural Networks With Piecewise Linear Activations
abstract
In this article, theoretically, it is shown that spurious local minima are common for deep fully connected networks and average-pooling convolutional neural networks (CNNs) with piecewise linear activations and datasets that cannot be fit by linear models. Motivating examples are given to explain why spurious local minima exist: each output neuron of deep fully connected networks and CNNs with piecewise linear activations produces a continuous piecewise linear (CPWL) function, and different pieces of the CPWL output can optimally fit disjoint groups of data samples when minimizing the empirical risk. Fitting data samples with different CPWL functions usually results in different levels of empirical risk, leading to the prevalence of spurious local minima. The results are proved in general settings with arbitrary continuous loss functions and general piecewise linear activations. The main proof technique is to represent a CPWL function as maximization over minimization of linear pieces. Deep networks with piecewise linear activations are then constructed to produce these linear pieces and implement the maximization over minimization operation.
Bo Liu 0011
IEEE Trans. Neural Networks Learn. Syst.1
2023 Infrared Small Target Detection Based on Saliency Guided Multi-Task Learning
abstract
Infrared (IR) small target detection is a challenging task due to the low contrast and low signal-to-noise ratio, generally yields high false alarm rates. To improve the performance of IR small target detection, we propose a saliency guided multi-task leaning model (SGMTLM). The model consists of two parts: feature fusion and saliency detection. The feature fusion module is to integrate shallow information and deep semantic information of small targets. The saliency detection module is used to guide the Feature Pyramid Networks (FPN) to focus on the small target area. It can effectively suppress the non-target information while enhancing the small target information. Finally, experimental results on two datasets Small-ExtIRShip and Small-SSDD demonstrated that, with the help of saliency detection, the proposed method can effectively improve the accuracy of IR small target detection, achieving 95.78% and 98.70% mAP on the two datasets, respectively.
Zhaoying Liu, Junran He, Ting Zhang 0012, Ziqing Han, Bo Liu 0011
ICIP6
2023 Reinforcement learning-enabled efficient data gathering in underground wireless sensor networks
Deng Zhao, Zhangbing Zhou, Shangguang Wang, Bo Liu 0011, Walid Gaaloul
Pers. Ubiquitous Comput.4
2022 Some geometrical and topological properties of DNNs' decision boundaries
Bo Liu 0011, Mengya Shen
Theor. Comput. Sci.1
2021 Effective *-flow schedule for optical circuit switching based data center networks: A comprehensive survey
Yinan Tang, Tongtong Yuan, Bo Liu 0011, Chuangbai Xiao
Comput. Networks3
2021 Optimal function approximation with ReLU neural networks
Bo Liu 0011
Neurocomputing1
2021 Understanding the loss landscape of one-hidden-layer ReLU networks
Bo Liu 0011
Knowl. Based Syst.1
2021 Non-differentiable saddle points and sub-optimal local minima exist for deep ReLU networks
Bo Liu 0011, Zhaoying Liu, Ting Zhang 0012, Tongtong Yuan
Neural Networks1
2020 Joint multi-scale discrimination and region segmentation for person re-ID
Jialiang Huang 0001, Bo Liu 0011
Pattern Recognit. Lett.2
2011 Multiconlitron: A General Piecewise Linear Classifier
abstract
Based on the "convexly separable" concept, we present a solid geometric theory and a new general framework to design piecewise linear classifiers for two arbitrarily complicated nonintersecting classes by using a "multiconlitron," which is a union of multiple conlitrons that comprise a set of hyperplanes or linear functions surrounding a convex region for separating two convexly separable datasets. We propose a new iterative algorithm called the cross distance minimization algorithm (CDMA) to compute hard margin non-kernel support vector machines (SVMs) via the nearest point pair between two convex polytopes. Using CDMA, we derive two new algorithms, i.e., the support conlitron algorithm (SCA) and the support multiconlitron algorithm (SMA) to construct support conlitrons and support multiconlitrons, respectively, which are unique and can separate two classes by a maximum margin as in an SVM. Comparative experiments show that SMA can outperform linear SVM on many of the selected databases and provide similar results to radial basis function SVM on some of them, while SCA performs better than linear SVM on three out of four applicable databases. Other experiments show that SMA and SCA may be further improved to draw more potential in the new research direction of piecewise linear learning.
Bo Liu 0011, Xinwu Yang, Yaozong Fu, Houjun Li
IEEE Trans. Neural Networks2
2008 Boundary Constrained Manifold Unfolding
abstract
A new manifold learning algorithm is proposed in this paper. Our method is motivated by the unit covariance constraint problem of spectral embedding methods, where a unit covariance constraint is imposed to avoid degenerate solutions that map all manifold samples to one point. This constraint distorts the aspect ratio and introduces unwanted correlation between different components of embedding coordinates. Instead, our method uses boundary conditions to pull apart mapped points, and obtains the embedding by solving linear systems under boundary conditions. The mapping of boundary samples is decided by that of a coarse version of manifold, obtained by a graph simplification algorithm designed by us. Comparisons between our method and several other representative manifold learning methods are made, and the results demonstrate the effectiveness of the proposed method.
Bo Liu 0011, Hongbin Zhang 0009, WenAn Chen
ICMLA1
2007 An Energy-Minimizing Mesh Parameterization
abstract
In this paper, we propose a new energy-minimizing mesh parameterization method, which linearly combines two new energies EQand EM. It not only avoids triangles overlap in the parameter domain, but also is invariant under rotation, translation and scale transformations. We first parameterize the original 3D mesh to the parameter plane by using the energy-minimizing parameterization, and get the optimal effect by optimizing the weights wijgradually. Experimental results indicate that this optimized energy-minimizing method has low distortion and good stability.
Li Yong, Bo Liu 0011, Hongbin Zhang 0009
ICME2
2004 Region-of-interest coding of 3D mesh based on wavelet transform
abstract
A scheme for the region of interest (ROI) coding of 3D meshes is proposed for the first time. The ROI is encoded with higher fidelity than the rest region, and the "priority" of ROI relative to the rest region (background, BG) can be specified by encoder or decoder (user). Wavelet transform is used on 3D mesh and zerotrees are adopted to organize the coefficients. The wavelet coefficients of ROI are scaled up and encoded with a modified set partitioning in hierarchical trees (SPIHT) algorithm. In additional, a fast algorithm is proposed for creating the ROI mask. Once the quality of reconstructed ROI becomes high enough, the transmission can be intermitted and much transmission bandwidth and storage space will be saved consequently.
Hongjuan Zheng, Bo Liu 0011, Hongbin Zhang 0009
ICIG2