Long Deng

dblp:214/1080 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Graph data management · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › graph processing
multicore graph processing
0.312025
Bubble: Towards Scalable Evolving Graph Processing via Mini-Batch Sorting · SC 2025
Parallel and multicore computing
parallel graph algorithms
0.312025
Bubble: Towards Scalable Evolving Graph Processing via Mini-Batch Sorting · SC 2025

Methods — techniques the papers use, named apart from their topics

mini-batch sorting · 1.7cache optimization · 1.7
YearPublicationVenuePosition
2025 PTWalker: Cache-Efficient Random Walks via Alternating Dual-Subgraph Walker Updating
abstract
Random walks on graphs are essential for various applications such as network analysis, recommendation systems, and graph embedding. However, existing frameworks for random walks face challenges in efficiently managing and updating walkers. These challenges arise from inefficient walker updating due to the subgraph-based iterative updating strategy, which makes walkers update only one step per iteration, and the high management and storage costs associated with the walker array-based management strategy. This paper presents a novel cache-efficient random walk framework called PTWalker, which sets out to have the majority of walkers updated at least two steps per iteration. This is achieved through an alternating dual-subgraph walk updating scheme, which involves dividing the graph into two sets of subgraphs with maximized edge cuts between them. These subgraph sets are then loaded into the cache for walk updating in an alternating manner, prompting most walkers to transition to the other subgraph set and update additional steps. Besides, a thread-level walker pool management strategy is designed to reduce management and storage costs for walkers. Experimental results demonstrate that PTWalker outperforms state-of-the-art random walk systems, achieving a speedup ranging from 0.18 × to 8.65 ×.
Rui Wang 0076, Long Deng, Wenzhe Zhu, Yongkun Li 0001, Yinlong Xu 0001
ICPP4
2025 Bubble: Towards Scalable Evolving Graph Processing via Mini-Batch Sorting
abstract
Evolving graph processing has become a critical component in various applications and is gaining increasing attention. However, existing evolving graph systems suffer from cache contention and workload imbalance between threads, which leads to poor scalability and performance degradation on modern multi-core computers.
Long Deng, Yongkun Li 0001, Yinlong Xu 0001, John C. S. Lui
SC1
2025 Bridging asymmetry between image and video: Cross-modality knowledge transfer based on learning from video
Bingxin Zhou, Jianghao Zhou, Zhongming Chen, Long Deng, Yongxin Ge
Expert Syst. Appl.5
2025 Erratum to: Data-driven soft sensors in blast furnace ironmaking: a survey
Yueyang Luo, Manabu Kano, Long Deng, Chunjie Yang 0001
Frontiers Inf. Technol. Electron. Eng.4
2024 Two-Stream Temporal Feature Aggregation Based on Clustering for Few-Shot Action Recognition
abstract
The metric learning paradigm has achieved notable success in few-shot action recognition; however, it faces unaddressed challenges. Specifically,(1)limited training data could impede the exploration of temporal action relations, and(2)precision would decline from the presence of outliers during the frame-level feature alignment. To address the challenges, we propose a two-stream temporal feature aggregation method based on clustering, incorporating a temporal augmentation module (TAM) and a feature aggregation module (FAM). The TAM adeptly integrates three consecutive grayscale frames into the original RGB frame through weighted summation, thereby addressing the color-related misguidance and enhancing the temporal information extraction. Meanwhile, the FAM employs clustering to aggregate the frame-level features into high semantic sub-actions and replaces the original features with cluster centers to mitigate the adverse impact of outliers on the model performance. Experimental results on benchmark datasets demonstrate the effectiveness of our method in few-shot action recognition. We validate our proposed approach by conducting comprehensive ablation experiments.
Long Deng, Bingxin Zhou, Yongxin Ge
IEEE Signal Process. Lett.1
2023 Data-driven soft sensors in blast furnace ironmaking: a survey
abstract
The blast furnace is a highly energy-intensive, highly polluting, and extremely complex reactor in the ironmaking process. Soft sensors are a key technology for predicting molten iron quality indices reflecting blast furnace energy consumption and operation stability, and play an important role in saving energy, reducing emissions, improving product quality, and producing economic benefits. With the advancement of the Internet of Things, big data, and artificial intelligence, data-driven soft sensors in blast furnace ironmaking processes have attracted increasing attention from researchers, but there has been no systematic review of the data-driven soft sensors in the blast furnace ironmaking process. This review covers the state-of-the-art studies of data-driven soft sensors technologies in the blast furnace ironmaking process. Specifically, we first conduct a comprehensive overview of various data-driven soft sensor modeling methods (multiscale methods, adaptive methods, deep learning, etc.) used in blast furnace ironmaking. Second, the important applications of data-driven soft sensors in blast furnace ironmaking (silicon content, molten iron temperature, gas utilization rate, etc.) are classified. Finally, the potential challenges and future development trends of data-driven soft sensors in blast furnace ironmaking applications are discussed, including digital twin, multi-source data fusion, and carbon peaking and carbon neutrality.
Yueyang Luo, Manabu Kano, Long Deng, Chunjie Yang 0001
Frontiers Inf. Technol. Electron. Eng.4
2018 Image Autoregressive Interpolation Model Using GPU-Parallel Optimization
abstract
With the growth in the consumer electronics industry, it is vital to develop an algorithm for ultrahigh definition products that is more effective and has lower time complexity. Image interpolation, which is based on an autoregressive model, has achieved significant improvements compared with the traditional algorithm with respect to image reconstruction, including a better peak signal-to-noise ratio (PSNR) and improved subjective visual quality of the reconstructed image. However, the time-consuming computation involved has become a bottleneck in those autoregressive algorithms. Because of the high time cost, image autoregressive-based interpolation algorithms are rarely used in industry for actual production. In this study, in order to meet the requirements of real-time reconstruction, we use diverse compute unified device architecture (CUDA) optimization strategies to make full use of the graphics processing unit (GPU) (NVIDIA Tesla K80), including a shared memory and register and multi-GPU optimization. To be more suitable for the GPU-parallel optimization, we modify the training window to obtain a more concise matrix operation. Experimental results show that, while maintaining a high PSNR and subjective visual quality and taking into account the I/O transfer time, our algorithm achieves a high speedup of 147.3 times for a Lena image and 174.8 times for a 720p video, compared to the original single-threaded C CPU code with -O2 compiling optimization.
Jiaji Wu, Long Deng, Gwanggil Jeon
IEEE Trans. Ind. Informatics2