VLDB 2026 Research / reviewers in the wild / expert
Benying Tan
dblp:195/9225
· DBLP profile ↗
20ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0002-9121-8499ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CTMD: A Hybrid Deep Learning Model for Inverse Design of Metasurface
Hedong He, Zihang Ma, Benying Tan, Shuxue Ding |
ICIC | 3 |
| 2026 | NRF-FQ: An Enhanced Numerical Reasoning Framework in Language Models Through FoNE and QLoRA
Benying Tan |
ICIC (22) | 3 |
| 2026 | Enhancing vision-and-language transformers through two-stage generative alignment pre-training
Huiming Xie, Shuxue Ding, Yujie Li 0002, Benying Tan |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Low-rank sparse autoencoders: Unifying efficiency and geometric regularization for large language model interpretability
Jiajia Mu, Benying Tan, Chenchen Luo, Ruibin Bai |
Inf. Sci. | 2 |
| 2026 | DSTG-VS: Dynamic spatio-temporal graph network with global-local feature fusion for video summarization
Haonan Jia, Yu Wang 0064, Benying Tan |
Pattern Recognit. | 6 |
| 2026 | LPOM pretraining as a warm start: Enhancing gradient descent optimization via non-gradient weight initialization
Benying Tan, Jianpeng Wu, Yujie Li 0002, Shuxue Ding |
Pattern Recognit. | 1 |
| 2025 | Device-Free Localization Based on Knowledge Distillation Method
Ruiyang Sun, Zhiyang Cui, Ziwei Xia, Benying Tan |
ICIC (17) | 4 |
| 2025 | Enhancing Large Language Model Fine-Tuning with Sharpness-Aware Minimization Under Split Federated Learning
Benying Tan, Yujie Li 0002, Shuxue Ding, Ahmad Chaddad |
ICIC (9) | 1 |
| 2025 | Global and Local CNN Attention Mechanisms for Video SummarizationabstractThe development of video summarization techniques is crucial for quickly grasping video content. However, existing methods often fall short of comprehending long-range contextual dependencies and capturing critical information effectively so that redundant content is kept. To address these limitations, we propose the Global and Local CNN Attention Mechanisms for Video Summarization (GLA-VS). The proposed GLA-VS model is built upon the deep convolutional neural network architecture of GoogleNet and incorporates the Convolutional Block Attention Module (CBAM) to enhance the ability to identify and focus on key features. Moreover, we introduce a global feature attention module, which enables the model to jointly attend to both local details and global features. Integrating CBAM and a global feature attention module significantly enhances the model’s ability to accurately extract key frames and generate concise video summaries. Experiments conducted on the SumMe and TVSum datasets demonstrate the effectiveness of the GLA-VS model. The model outperforms existing methods in terms of Kendall’s Tau coefficient and Spearman’s rank correlation coefficient. Specifically, the GLA-VS model achieves a 10.57% improvement on the SumMe dataset compared to the state-of-the-art model. The improvements presented in this study provide a new perspective for video summarization, highlighting the potential of leveraging advanced attention mechanisms to boost model performance. Haonan Jia, Benying Tan |
IJCNN | 5 |
| 2025 | EHSF: Enhanced Hybrid Supervision Framework for surface-defect detectionabstractDetecting surface defects in industrial products is crucial for ensuring quality control in manufacturing. Traditional methods face challenges due to the diversity of defect types, the small and ambiguous nature of defects, and the high cost of labeled data. Unsupervised and semi-supervised methods can reduce labeling costs but often fail to meet industrial accuracy requirements. In this paper, we propose an Enhanced Hybrid Supervised Framework (EHSF) designed to improve defect detection accuracy in complex industrial scenarios with fewer labeled samples. The framework incorporates an Adaptive Cross-Scale Feature Enhancement Module (ACFEM) based on Selective Feature Aggregation (SFA), which addresses the limitations of single-level feature representations and significantly enhances the detection of multiscale defects. Additionally, we introduce a novel Dynamic Feature Calibration Network (DFCNet) that synergistically combines global contextual information and local details through dynamic feature calibration and adaptive global semantic enhancement mechanisms. The proposed approach is validated on the DAGM benchmark and three real-world industrial datasets (Kolektor Surface Defect Dataset, Kolektor Surface Defect Dataset2, and Severstal Steel). Experimental results demonstrate that our method outperforms existing techniques by reducing reliance on weakly supervised fine annotations while achieving superior detection accuracy, particularly under weak supervision. Benying Tan, Beibei Ren, Shuxue Ding |
IJCNN | 1 |
| 2025 | An intelligent retrievable object-tracking system with real-time edge inference capabilityabstractAbstract An intelligent retrievable object‐tracking system assists users in quickly and accurately locating lost objects. However, challenges such as real‐time processing on edge devices, low image resolution, and small‐object detection significantly impact the accuracy and efficiency of video‐stream‐based systems, especially in indoor home environments. To overcome these limitations, a novel real‐time intelligent retrievable object‐tracking system is designed. The system incorporates a retrievable object‐tracking algorithm that combines DeepSORT and sliding window techniques to enhance tracking capabilities. Additionally, the YOLOv7‐small‐scale model is proposed for small‐object detection, integrating a specialized detection layer and the convolutional batch normalization LeakyReLU spatial‐depth convolution module to enhance feature capture for small objects. TensorRT and INT8 quantization are used for inference acceleration on edge devices, doubling the frames per second. Experiments on a Jetson Nano (4 GB) using YOLOv7‐small‐scale show an 8.9% improvement in recognition accuracy over YOLOv7‐tiny in video stream processing. This advancement significantly boosts the system's performance in efficiently and accurately locating lost objects in indoor home settings. Yujie Li 0002, Yifu Wang, Zihang Ma, Xinghe Wang, Benying Tan, Shuxue Ding |
IET Image Process. | 5 |
| 2025 | Lifted proximal operator machine-based deep nonlinear dictionary learning with multilayer regularization
Benying Tan, Yujie Li 0002, Shuxue Ding |
Neurocomputing | 2 |
| 2024 | Accelerated Deep Nonlinear Dictionary Learning
Benying Tan, Shuxue Ding, Yujie Li 0002 |
ACCV (1) | 1 |
| 2024 | Vision Transformer with 2D Explicit Position EncodingabstractRecently, the Vision Transformer (ViT) has achieved outstanding performance in various computer vision tasks. Positional encoding is an indispensable component of ViT for handling the inherent structural information of images. However, attaching position encodings manually is a time-consuming process that slows down the training speed of ViT. To address this issue, we propose an explicit approach for positional encoding, distinct from the original ViT’s implicit design. Our new implementation uses a 2D-based explicit positional encoding method that accelerates convergence and improves training efficiency. The proposed approach yields a remarkable improvement, especially in the initial stages of training, where the 2D explicit positional encoding offers improved compatibility with various input lengths and enhanced interpretability. The experimental results on the ImageNet dataset confirm the effectiveness of our proposed 2D explicit positional encoding approach. The proposed explicit 2D coordinate position encoding can achieve a maximum improvement of up to 437%. Zihang Ma, Xinghe Wang, Yifu Wang, Benying Tan |
ICASSP | 5 |
| 2023 | Bar transformer: a hierarchical model for learning long-term structure and generating impressive pop music
Huiming Xie, Shuxue Ding, Benying Tan, Yujie Li 0002, Bin Zhao 0007 |
Appl. Intell. | 4 |
| 2023 | A new device-free localization method for RSS data with considering correlations
Ziwei Xia, Benying Tan, Haoli Zhao, Shuxue Ding, Yujie Li 0002 |
Comput. Commun. | 2 |
| 2023 | A DCA-based sparse coding for video summarization with MCPabstractAbstract Video summarization offers a summary version that conveys the primary information of a longer video. The main challenges of video summarization are related to keyframe extraction and saliency mapping. Thus, this work proposes a sparse coding model for keyframe extraction and saliency mapping applications. Specifically, the minimax concave penalty (MCP) is utilized as a sparse regularization scheme and the regularized non‐convex MCP problem is solved by decomposing MCP into two convex functions and the convex function's algorithm difference is relied on to solve the resulting sub‐problems. The experimental results demonstrate higher compressed keyframes and saliency maps than current state‐of‐the‐art algorithms. In particular, the model attains a lower summary length of 34% and 19% compared to sparse modeling representation selection (SMRS) and sparse modeling using the determinant sparsity measure (SC‐det), respectively. In addition, the developed scheme has a shorter computation time, requiring 82% and 33% less time than the ITTI and the dense and sparse reconstruction (DSR) methods. Yujie Li 0002, Zhenni Li, Benying Tan, Shuxue Ding |
IET Image Process. | 3 |
| 2021 | A fast DC-based dictionary learning algorithm with the SCAD penalty
Zhenni Li, Chao Wan, Benying Tan, Zuyuan Yang, Shengli Xie 0001 |
Neurocomputing | 3 |
| 2021 | Gaze prediction for first-person videos based on inverse non-negative sparse coding with determinant sparse measureabstractGaze prediction is a significant approach for processing a large amount of incoming visual information of videos. Recent gaze prediction algorithms often employ sparse models with the assumption that every superpixel in the video frames can be represented as linear combinations of a few salient superpixels. However, they are not actuated enough because of the insufficient knowledge that video signals contain a non-negative request. Hence, we develop a novel gaze prediction based on an inverse sparse coding framework with a determinant sparse measure. By introducing this sparse measure, the solutions are non-negative and sparser than conventional sparse constraints. However, the proposed optimization problem becomes nonconvex, which is difficult to solve. To efficiently address the corresponding nonconvex optimization problem, we propose a novel algorithm based on the difference in convex function programming, which can yield the global solutions. Experimental results indicate the improved accuracy of the proposed approach compared with state-of-the-art algorithms. Yujie Li 0002, Benying Tan, Shotaro Akaho, Hideki Asoh, Shuxue Ding |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | A novel dictionary learning method for sparse representation with nonconvex regularizations
Benying Tan, Yujie Li 0002, Haoli Zhao, Xiang Li 0005, Shuxue Ding |
Neurocomputing | 1 |