Yujie Li 0002

dblp:28/7846-2 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-5801-4937ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing vision-and-language transformers through two-stage generative alignment pre-training
Huiming Xie, Shuxue Ding, Yujie Li 0002, Benying Tan
Eng. Appl. Artif. Intell.4
2026 LPOM pretraining as a warm start: Enhancing gradient descent optimization via non-gradient weight initialization
Benying Tan, Jianpeng Wu, Yujie Li 0002, Shuxue Ding
Pattern Recognit.4
2025 Enhancing Large Language Model Fine-Tuning with Sharpness-Aware Minimization Under Split Federated Learning
Benying Tan, Yujie Li 0002, Shuxue Ding, Ahmad Chaddad
ICIC (9)3
2025 An intelligent retrievable object-tracking system with real-time edge inference capability
abstract
Abstract An intelligent retrievable object‐tracking system assists users in quickly and accurately locating lost objects. However, challenges such as real‐time processing on edge devices, low image resolution, and small‐object detection significantly impact the accuracy and efficiency of video‐stream‐based systems, especially in indoor home environments. To overcome these limitations, a novel real‐time intelligent retrievable object‐tracking system is designed. The system incorporates a retrievable object‐tracking algorithm that combines DeepSORT and sliding window techniques to enhance tracking capabilities. Additionally, the YOLOv7‐small‐scale model is proposed for small‐object detection, integrating a specialized detection layer and the convolutional batch normalization LeakyReLU spatial‐depth convolution module to enhance feature capture for small objects. TensorRT and INT8 quantization are used for inference acceleration on edge devices, doubling the frames per second. Experiments on a Jetson Nano (4 GB) using YOLOv7‐small‐scale show an 8.9% improvement in recognition accuracy over YOLOv7‐tiny in video stream processing. This advancement significantly boosts the system's performance in efficiently and accurately locating lost objects in indoor home settings.
Yujie Li 0002, Yifu Wang, Zihang Ma, Xinghe Wang, Benying Tan, Shuxue Ding
IET Image Process.1
2025 Lifted proximal operator machine-based deep nonlinear dictionary learning with multilayer regularization
Benying Tan, Yujie Li 0002, Shuxue Ding
Neurocomputing3
2024 Accelerated Deep Nonlinear Dictionary Learning
Benying Tan, Shuxue Ding, Yujie Li 0002
ACCV (1)5
2024 A flare removal network for night vision perception: Resistant to the interference of complex light
abstract
Abstract The high‐precision visual perception results are easily affected by the lens flare issue when the image sensor is facing to strong light. The existing flare removal methods have poor robustness when confronted with flare interference caused by complex nighttime lighting, which has to preserve natural light source information. A simulated dataset for the removal of night flares is created to solve the problem of collecting complete paired training data, and night flare removal network (NFR‐Net) is proposed to remove the interference caused by various light disturbances at night. The light source extraction module is introduced to retain light source information realistically and effectively in night vision scenes. Extensive experimental results demonstrate that the proposed method is superior to the existing related methods in the various complex night vision scenes. The proposed NFR‐Net can enhance visual perception of nighttime images significantly and improve the performance of night vision tasks.
Yan Liu 0079, Wenting Qi, Yujie Li 0002
IET Image Process.4
2023 Bar transformer: a hierarchical model for learning long-term structure and generating impressive pop music
Huiming Xie, Shuxue Ding, Benying Tan, Yujie Li 0002, Bin Zhao 0007
Appl. Intell.5
2023 A new device-free localization method for RSS data with considering correlations
Ziwei Xia, Benying Tan, Haoli Zhao, Shuxue Ding, Yujie Li 0002
Comput. Commun.5
2023 A DCA-based sparse coding for video summarization with MCP
abstract
Abstract Video summarization offers a summary version that conveys the primary information of a longer video. The main challenges of video summarization are related to keyframe extraction and saliency mapping. Thus, this work proposes a sparse coding model for keyframe extraction and saliency mapping applications. Specifically, the minimax concave penalty (MCP) is utilized as a sparse regularization scheme and the regularized non‐convex MCP problem is solved by decomposing MCP into two convex functions and the convex function's algorithm difference is relied on to solve the resulting sub‐problems. The experimental results demonstrate higher compressed keyframes and saliency maps than current state‐of‐the‐art algorithms. In particular, the model attains a lower summary length of 34% and 19% compared to sparse modeling representation selection (SMRS) and sparse modeling using the determinant sparsity measure (SC‐det), respectively. In addition, the developed scheme has a shorter computation time, requiring 82% and 33% less time than the ITTI and the dense and sparse reconstruction (DSR) methods.
Yujie Li 0002, Zhenni Li, Benying Tan, Shuxue Ding
IET Image Process.1
2021 Gaze prediction for first-person videos based on inverse non-negative sparse coding with determinant sparse measure
abstract
Gaze prediction is a significant approach for processing a large amount of incoming visual information of videos. Recent gaze prediction algorithms often employ sparse models with the assumption that every superpixel in the video frames can be represented as linear combinations of a few salient superpixels. However, they are not actuated enough because of the insufficient knowledge that video signals contain a non-negative request. Hence, we develop a novel gaze prediction based on an inverse sparse coding framework with a determinant sparse measure. By introducing this sparse measure, the solutions are non-negative and sparser than conventional sparse constraints. However, the proposed optimization problem becomes nonconvex, which is difficult to solve. To efficiently address the corresponding nonconvex optimization problem, we propose a novel algorithm based on the difference in convex function programming, which can yield the global solutions. Experimental results indicate the improved accuracy of the proposed approach compared with state-of-the-art algorithms.
Yujie Li 0002, Benying Tan, Shotaro Akaho, Hideki Asoh, Shuxue Ding
J. Vis. Commun. Image Represent.1
2020 A novel dictionary learning method for sparse representation with nonconvex regularizations
Benying Tan, Yujie Li 0002, Haoli Zhao, Xiang Li 0005, Shuxue Ding
Neurocomputing2
2018 A Sparse Coding Framework for Gaze Prediction in Egocentric Video
abstract
To efficiently process and understand a large amount of incoming visual information from first-person perspective (i.e. egocentric vision), predicting human gaze is important. However, even though people continuously gaze in noisy environments, most existing gaze prediction methods mainly use image saliency, which is sensitive to noise in the real-world. To address this issue, we propose a sparse coding-based saliency detection method for gaze prediction. Our model uses a cost function with the 10 norm as a sparse constraint that can control the area of visual saliency in response to the contents of egocentric vision in intuitive and consistent ways. Moreover, we use canonical correlation analysis (CCA) to combine different types of features for reducing noise and the computational complexity. We also utilize the temporal continuity of image frames when defining our saliency. Experiments using a real-world gaze dataset show that our proposed approach outperforms the state-of-the-art algorithms on gaze prediction in egocentric videos.
Yujie Li 0002, Atsunori Kanemura, Hideki Asoh, Taiki Miyanishi, Motoaki Kawanabe
ICASSP1
2018 Manifold optimization-based analysis dictionary learning with an ℓ1∕2-norm regularizer
Zhenni Li, Shuxue Ding, Yujie Li 0002, Zuyuan Yang, Shengli Xie 0001, Wuhui Chen
Neural Networks3
2017 Extracting key frames from first-person videos in the common space of multiple sensors
abstract
Selecting authentic scenes about activities of daily living (ADL) is useful to support our memory of everyday life. Key-frame extraction for first-person vision (FPV) videos is a core technology to realize such memory assistant. However, most existing key-frame extraction methods have mainly focused on stable scenes not related to ADL and only used visual signals of the image sequence even though the activities usually associate with our visual experience. To deal with dynamically changing scenes of FPV about daily activities, integrating motion and visual signals are essential. In this paper, we present a novel key-frame extraction method for ADL, which integrates multi-modal sensor signals to temper noise and detect salient activities. Our proposed method projects motion and visual features to a shared space by a probabilistic canonical correlation analysis and selects key frames there. The experimental results using ADL datasets collected in a house suggest that our key-frame extraction technique running in the shared space improves the precision of extracted key frames and the coverage of the entire video.
Yujie Li 0002, Atsunori Kanemura, Hideki Asoh, Taiki Miyanishi, Motoaki Kawanabe
ICIP1
2017 Key frame extraction from first-person video with multi-sensor integration
abstract
First-person videos (FPVs) in daily living help us to memorize our life experience and information systems to process daily activities. Summarizing FPVs into key frames that represent the entire data would allow us to remember our memory in the past and computers to efficiently process the data. However, most video summarization approaches only use visual information, even though our daily activities consist of multiple modalities such as movements and sounds. FPVs are not as stable as movies or sport scenes since the camera attached to the head shakes frequently, and key frame extraction methods rely only on video frames do not always produce satisfactory results. In this paper, we introduce a novel key frame extraction method for FPVs using multiple wearable sensors. To efficiently integrate multimodal sensor signals, our formulation uses sparse dictionary selection, which minimizes a reconstruction error with a subset (key frames) of the original data. We present experimental results with multimodal datasets captured by wearable sensors in a natural environment. The results suggest multi-sensor information improves the precision of extracted key frames as well as the coverage of an entire video sequence.
Yujie Li 0002, Atsunori Kanemura, Hideki Asoh, Taiki Miyanishi, Motoaki Kawanabe
ICME1
2017 Analysis dictionary learning using block coordinate descent framework with proximal operators
Zhenni Li, Shuxue Ding, Takafumi Hayashi, Yujie Li 0002
Neurocomputing4
2015 A Fast Algorithm for Learning Overcomplete Dictionary for Sparse Representation Based on Proximal Operators
abstract
We present a fast, efficient algorithm for learning an overcomplete dictionary for sparse representation of signals. The whole problem is considered as a minimization of the approximation error function with a coherence penalty for the dictionary atoms and with the sparsity regularization of the coefficient matrix. Because the problem is nonconvex and nonsmooth, this minimization problem cannot be solved efficiently by an ordinary optimization method. We propose a decomposition scheme and an alternating optimization that can turn the problem into a set of minimizations of piecewise quadratic and univariate subproblems, each of which is a single variable vector problem, of either one dictionary atom or one coefficient vector. Although the subproblems are still nonsmooth, remarkably they become much simpler so that we can find a closed-form solution by introducing a proximal operator. This leads to an efficient algorithm for sparse representation. To our knowledge, applying the proximal operator to the problem with an incoherence term and obtaining the optimal dictionary atoms in closed form with a proximal operator technique have not previously been studied. The main advantages of the proposed algorithm are that, as suggested by our analysis and simulation study, it has lower computational complexity and a higher convergence rate than state-of-the-art algorithms. In addition, for real applications, it shows good performance and significant reductions in computational time.
Zhenni Li, Shuxue Ding, Yujie Li 0002
Neural Comput.3