EDBT 2026 Demo / reviewers in the wild / expert
Linchen Yu
dblp:04/7638
· DBLP profile ↗
20ranked-venue papers
9as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 5 since 2021Computer networks · 6 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Achieving Efficient Temporal Graph Transformation on the GPU
Linchen Yu, Jin Zhao 0003, Longlong Lin, Hengshan Yue |
APPT | 1 |
| 2025 | TempGraph: An Efficient Chain-driven Temporal Graph Computing Framework on the GPUabstractTackling temporal path problems in temporal graphs is essential for time-sensitive applications. Although many solutions have been proposed to handle temporal path problems, due to the intrinsic time constraints, these solutions require the vertices of the temporal graph to be sequentially handled along the time-dependent chains (i.e., the temporal dependencies between these vertices) to form the temporal path. This sequential temporal nature poses the challenges of poor parallelism and slow convergence speed, preventing existing solutions from fully leveraging the massive parallelism and high internal bandwidth of GPU to handle temporal path problems. To overcome these challenges, this paper proposes TempGraph, an efficient chain-driven GPU-based temporal graph computing framework. Specifically, it transforms the temporal graph into a set of disjoint time-dependent chains that can elegantly expose the temporal dependency between the vertices while facilitating the fast path exploration along these chains over GPU. Furthermore, TempGraph employs a novel Generate-Activate-Compute execution model to decouple the temporal dependency between different chains through maintaining a set of shortcuts for them, which enables multiple chains to be concurrently handled by massive GPU threads, achieving fast convergence speed and high parallelism on the GPU. Experiments on an A100 GPU show that TempGraph outperforms the state-of-the-art GPU-based solutions by 3.0-16.2×. Besides, TempGraph on an A100 GPU gains 33.9-368.9× speedups compared to the cutting-edge CPU-based system TeGraph on a 128-core CPU machine. Jin Zhao 0003, Qian Wang 0002, Ligang He, Yu Zhang 0027, Sheng Di, Bingsheng He, Hao Qi 0004, Longlong Lin, Linchen Yu, Xiaofei Liao, Hai Jin 0001 |
ASPLOS (3) | 11 |
| 2025 | SDBF: Steep-Decision-Boundary Fingerprinting for Hard-Label Tampering Detection of DNN ModelsabstractCloud-based AI systems offer significant benefits but also introduce vulnerabilities, making deep neural network (DNN) models susceptible to malicious tampering. This tampering may involve harmful behavior injection or resource reduction, compromising model integrity and performance. To detect model tampering, hard-label fingerprinting techniques generate sensitive samples to probe and reveal tampering. Existing fingerprinting methods are mainly based on gradient-defined sensitivity or decision boundary, with the latter showing a manifest superior detection performance. However, all existing fingerprinting methods either suffer from insufficient sensitivity or incur high computational costs.In this paper, we theoretically analyze the black-box co-optimal tampering detection sensitivity of fingerprint samples in the context of decision boundary and gradient-defined sensitivity. Based on this, we further propose Steep-Decision-Boundary Fingerprinting (SDBF), a novel lightweight approach for hard-label tampering detection that inherently and efficiently combines the strengths of existing fingerprinting techniques. SDBF places fingerprint samples near the steep decision boundary, where the outputs of samples are inherently highly sensitive to tampering. We also design a Max Boundary Coverage Strategy (MBCS), which enhances samples’ diversity over the decision boundary. Theoretical analysis and extensive experimental results show that SDBF outperforms existing SOTA hard-label fingerprinting methods in both sensitivity and efficiency. Xiaofan Bai, Shixin Li 0001, Xiaojing Ma 0002, Bin B. Zhu, Dongmei Zhang 0001, Linchen Yu |
CVPR | 6 |
| 2025 | Enhancing Adversarial Transferability with Checkpoints of a Single Model's TrainingabstractAdversarial attacks threaten the integrity of deep neural networks (DNNs), particularly in high-stakes applications. In this paper, we present a novel black-box adversarial attack that leverages the diverse checkpoints generated during a single model’s training trajectory. Unlike conventional ensemble attacks that require multiple surrogate models with diverse architectures, our approach exploits the intrinsic diversity captured over different training stages of a single surrogate model. By decomposing the learned representations into task-intrinsic and task-irrelevant components, we employ an accuracy gap-based selection strategy to identify checkpoints that predominantly capture transferable, task-intrinsic knowledge. Extensive experiments on ImageNet and CIFAR-10 demonstrate that our method consistently outperforms traditional ensemble attacks in terms of transferability, even under resource-constrained and practical settings. This work offers a resource-efficient solution for crafting highly transferable adversarial examples and provides new insights into the dynamics of adversarial vulnerability. Shixin Li 0001, Chaoxiang He, Xiaojing Ma 0002, Bin B. Zhu, Shuo Wang 0012, Hongsheng Hu, Dongmei Zhang 0001, Linchen Yu |
CVPR | 8 |
| 2025 | An Efficient ReRAM-based Accelerator for Asynchronous Iterative Graph ProcessingabstractGraph processing has become a central concern for many real-world applications and is well-known for its low compute-to-communication ratios and poor data locality. By integrating computing logic into memory, resistive random access memory (ReRAM) tackles the demand for high memory bandwidth in graph processing. Despite the years’ research efforts, existing ReRAM-based graph processing approaches still face the challenges of redundant computation overhead . It is because the vertices of many subgraphs are ineffectively and repeatedly processed over the ReRAM crossbars for lots of iterations so as to update their states according to the vertices of other subgraphs regardless of the dependencies among the subgraphs. In this article, we propose ASGraph , a dependency-aware ReRAM-based graph processing accelerator that overcomes the aforementioned performance bottlenecks. Specifically, ASGraph dynamically constructs the subgraph based on the dependencies between vertices’ states and then detects constructed subgraph that owns high value (it is likely that it has accumulated many state propagations from its neighbors and is able to affect more other neighbors) to be preferentially processed. In this way, it makes the vertex states propagate along the dependencies between vertices as much as possible to reduce the redundant computation. Besides, ASGraph employs a hybrid processing scheme to accelerate the state propagations of the tightly connected subgraph, thereby minimizing the redundant computations. Experimental results show that ASGraph achieves 25.5× and 4.8× speedup and 70.8× and 2.2× energy saving on average compared with the state-of-the-art ReRAM-based graph processing accelerators, that is, GraphR and GaaS-X, respectively. Jin Zhao 0003, Yu Zhang 0027, Donghao He, Qikun Li, Weihang Yin, Hao Qi 0004, Xiaofei Liao, Hai Jin 0001, Haikun Liu, Linchen Yu, Zhan Zhang 0003 |
ACM Trans. Archit. Code Optim. | 11 |
| 2024 | An Efficient GCNs Accelerator Using 3D-Stacked Processing-in-Memory ArchitecturesabstractGraph Convolutional Networks (GCNs) hold great promise in facilitating machine learning on graph-structured data. However, the sparsity of graphs often results in a significant number of irregular memory accesses, leading to inefficient data movement for existing GCNs accelerators. With the advancement of 3D stacked technology, the processing-in-memory (PIM) architecture has emerged as a promising solution for graph processing. Nevertheless, existing PIM accelerators are confronted with the challenges of irregular remote access in the aggregation phase of GCNs and dynamic workload variations between phases. In this paper, we present GCNim, a PIM accelerator based on 3D stacked memory, which features two key innovations in terms of the computation model and hardware designs. First, we present a PIM-based hybrid computation model, which employs a remote merging strategy to achieve the outer product in aggregation and the row-wise product in combination. Second, GCNim builds a three-stage aggregation and combination pipeline and integrates unified processing elements (PEs) supporting these three stages at the bank level, achieving load balance among PEs through a lightweight data placement algorithm. Compared with the state-of-the-art software frameworks running on CPUs and GPUs, GCNim achieves an average speedup of 3,736.06× and 76.56×, respectively. Moreover, GCNim outperforms the state-of-the-art GCN hardware accelerators, I-GCN, PEDAL, FlowGNN, and GCIM, with average speedups of 3.35×, 8.97×, 2.24×, and 5.58×, respectively. Ao Hu, Long Zheng 0003, Qinggang Wang, Jingrui Yuan, Haifeng Liu 0003, Linchen Yu, Xiaofei Liao, Hai Jin 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2021 | CDVT: A Cluster-Based Distributed Video Transcoding Scheme for Mobile Stream Services
Wei Ren 0002, Daxi Tu, Linchen Yu, Tianqing Zhu, Yi Ren 0001 |
WASA (1) | 4 |
| 2021 | SVC-based dynamic caching for smart media streaming over the Internet of Things
Linchen Yu |
Future Gener. Comput. Syst. | 1 |
| 2020 | An efficient iterative graph data processing framework based on bulk synchronous parallel modelabstractSummary Graph data processing has been widely applied in a variety of domains such as industry, science, social network, and so on. It therefore has stimulated many efforts devoted to this area. To embrace the fast development trend of big graph data, graph data processing based on Pregel‐like systems has been regarded as one of the most promising ways and has widely attracted the attention of researchers. However, it still remains in its early stage and there still exist many challenges. In Pregel, the superstep synchronization is time consuming as the graph data iteration operation requires multiple synchronizations. Furthermore, the graph data partition strategy adopted by Pregel fails to support load balancing, therefore causing the increase of network I/O overhead as the scale of graph data grows. To address these issues, this paper presents an efficient computational framework for graph data processing based on the bulk synchronous parallel model. The global synchronization control mechanism is improved by determining the start time of the next round of superstep through counting the number of global message files. Furthermore, an improved graph data partition mechanism based on a balanced hash method is proposed to reduce the communication overhead between different partitions of sub‐graph computational tasks. We also re‐design the PageRank algorithm to verify the effectiveness of the proposed framework. Experimental results on different real‐world datasets verify the efficiency of our proposed framework as it outperforms Giraph (an open source Pregel‐like system) by 58%−69%, and achieves 10×−17× performance improvement over Hadoop. Chao Liu 0007, Deze Zeng, Hong Yao, Xuesong Yan 0001, Linchen Yu, Zhangjie Fu 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2020 | CCHybrid: CPU co-scheduling in virtualization environmentabstractSummary Virtualization is very important to build the emerging cloud infrastructure, and a VM (virtual machine) with many kinds of workloads can run on physical machines in cloud environment. The VMM (virtual machine manager) scheduling algorithm asynchronously schedules each VCPU (virtual CPU) of a VM and ensures the CPU time usage of each VM. This proportional share method is widely used, because it simplifies the implementation of VMM CPU scheduling algorithm and can provide near‐perfect performance for most ordinary workloads. However, when a VM runs with parallel workloads, the above method causes performance degradation because of the negative impact of virtualized systems. Therefore, in this paper, we propose an optimized scheduling system, called CCHybrid, for parallel program in the Xen. It uses weight‐based proportion share strategy to ensure the fairness. In order to resolve the impact of virtualization on synchronization, it uses a novel co‐scheduling strategy, which dynamically adjusts the size of co‐scheduling to remit CPU fragmentation and maintains the original asynchronous scheduling policy for non‐parallel applications. In this way, CCHybrid provides CPU resource allocation services for Xen and can decrease the negative impact of virtualized systems, while ensuring the fairness of VMs and the performance of non‐parallel workload. Experimental results show that in the case of multiple VMs, CCHybrid improves the performance of parallel workload from 15% to 50%, and the impact on non‐parallel workload is less than 5%, in comparison with the credit scheduling algorithm of Xen. Linchen Yu |
Concurr. Comput. Pract. Exp. | 1 |
| 2020 | A Hierarchical Encryption and Key Management Scheme for Layered Access Control on H.264/SVC Bitstream in the Internet of ThingsabstractTerminals with diverse technological specifications, heterogeneous network environment, and personalized user requirements raise new challenges to streaming media services. Solutions such as the newly standardized H.264/SVC (scalable video coding; designed to compress original video bitstream into a multilayer video stream according to requirements) have been proposed. With the pervasive application of SVC in applications, such as video on demand, video conferencing, and video surveillance in the Internet of Things (IoT), there has been increased scrutiny on security of H.264/SVC. In this article, we propose a bitstream-oriented layered encryption scheme for SVC bitstream. According to the multilayer bit code structure of SVC, the bitstream is separated and encrypted, respectively, by rearranging the network abstraction layer (NAL) unit of SVC bitstream. This provides hierarchical protection for the multilayer characteristic of SVC. In order to provide sufficient security, as well as achieving improved computational efficiency, we use different cryptographic algorithms for the base layer and enhancement layers according to its requirements. The base layer adopts off-the-shelf high-security encryption algorithms, such as block cipher, to ensure security. Each enhancement layer is encrypted with a different key through the stream cipher with low computational complexity, providing layered control of the video. Furthermore, we propose a hierarchical key management scheme to implement layered access control according to the principle of hierarchical deterministic wallet (H-D wallet). Our scheme can be applied to the user-level distinction in video on demand and video surveillance systems in IoT. The analysis and experiments indicate that the proposed scheme achieves a high-security level, yet incurs reasonably low compression cost and computational complexity. Wei Ren 0002, Linchen Yu, Tianqing Zhu, Kim-Kwang Raymond Choo |
IEEE Internet Things J. | 3 |
| 2018 | Key frame extraction scheme based on sliding window and features
Linchen Yu, Jigang Cao, Mianlong Chen |
Peer-to-Peer Netw. Appl. | 1 |
| 2016 | A Key Frame Selection Algorithm Based on Sliding Window and Image FeaturesabstractNetwork traffic associated with video increases sharply, how to choose the interested information for a number of Internet users is challenging. So, technologies and applications related with video, such as video search, video fast browsing, video index and storage are in great demand. Behind these technologies and applications, a core problem is how to quickly browse massive video data and obtain the main content of the video. To solve this problem, different key frame extraction algorithms have been proposed. Due to the diversity of video content, different video have different characteristics. So the design of general video key frame extraction algorithm to solve the problem is not the reality. The main trend for the problem is to design the key frame extraction algorithm based on the characteristics of the video itself. In this article, we mainly focus on videos with edited boundaries and shot conversions. Aiming at this kind of video, we have designed and implemented video key frame extraction algorithm based on sliding window, the global feature Gist and local feature point detection algorithm SURF. In this algorithm, we use Gist feature to construct the global scene information of frames, and the SURF keypoint detection algorithm to extract local keypoints as local feature for each frame. Then, shot segmentation based on sliding window and shot merging algorithm is applied to dividing the original video into several shots. After that, we select the most representative frames in each video shot as key frames. Finally we evaluate the result of the algorithm from the subjective and objective perspective. Results show that key frames extracted in the algorithm are of high quality and can basically cover the main content of the original video. Jigang Cao, Linchen Yu, Mianlong Chen |
ICPADS | 2 |
| 2013 | CloudWeb: A Cloud Based Webpage Transforming for Mobile DevicesabstractMobile devices have already been widely used to access the Web. However, because most available web pages are designed for desktop PC, it is inconvenient to browse these large web pages on a mobile device with a small screen. In this paper, we propose a new web page transforming scheme to facilitate navigation on a small-form-factor device based on cloud computing. Different with other existing schemes, our approach is based on page segmentation and template-rendering. The semantic structure of a web page is extracted by DOM (Document Object Model) analyzing. Such semantic structure is hierarchical. Each node is corresponded to a block area of one web page and is in charge of reconstructing the whole web page. According to different content styles, the template-rendering scheme generates different display modes. Experimental results show good experience. Linchen Yu, Xiaofei Liao |
MSN | 1 |
| 2013 | Cloud Based Mobile Video Editing SystemabstractMobile video editing systems need efficient streaming transmission and interactive operations. Rather than using a specialized protocol and stream format, the video editing system makes use of a generic mechanism based on chunks in the cloud computing environment. Chunks are in fixed-size and contains a mixture of scalar data and references to other chunks. Chunks allow programmers to expose large, but fine-grained, data structures over the network. The video editing system organizes video clips with simple data types including linked lists and search trees, allowing a client to retrieve. The mobile video editing system supports resource adaptive play-back and "live" streaming of real-time video as well as fast, frame-accurate seeking, bandwidth-efficient high-speed play-back, and compilation of editing decisions from a set of clips. All the editing functions are completed in the cloud center and the editing results are showed on the mobile devices. Evaluations indicate that our system uses less bandwidth than HTTP Live Streaming while providing better support for editing primitives. Linchen Yu, Xiaofei Liao |
MSN | 1 |
| 2012 | A Peer-to-Peer Massive Battle Observing System to Support Game LiveabstractGaming services are attractive for Internet users. How to broadcast live gaming services to large-scale Internet users is still a big problem. Traditional schemes, based on Peer-to-Peer (P2P) live-streaming and based on TV shows, have poor user experiences (good graphic quality with high resolution) and need big bandwidth or TV set support. In order to avoid network congestion, a small number of solutions propose an alternative approach based on game data content instead of video content to provide game live services. At present, all of them who use client/server framework do not emphasize the design of system architecture, so the number of users is constrained and the system is hard to expand. In this paper, we introduce PKTV, a P2P game battle observing system, which is now optimized specially for War craft 3. PKTV system broadcasts game data content to reduce bandwidth overhead, maintains a P2P overlay for each gaming channel and uses a reliable UDP protocol to break the limitation of the TCP connections, decentralizes servers to enhance system's availability and scalability. Furthermore, tests and simulations show that PKTV also have robust functionality and good performance: low-latency, low bandwidth consumption and very considerable bandwidth savings for servers. Linchen Yu, Xiaofei Liao |
APSCC | 1 |
| 2012 | An Efficient Distributed Transactional Memory SystemabstractTransactional memory (TM) is a parallel programming concept which reduces challenges in parallel programming. Existing distributed transactional memory system consumes too much bandwidth and brings high latency. In this work, we present Transactional Memory System for Cluster (Clustm), a generalized and scalable distributed transactional memory system. Our system addresses several open issues posed by this domain, including transactional memory consistency protocol, cache consistency protocol, and the distribution strategy of the metadata of shared data across the cluster. Then, we evaluate our design with several workloads, and the results demonstrate outstanding performance. Xiaofei Liao, Hai Jin 0001, Xuepeng Fan, Xuping Tu, Linchen Yu |
TrustCom | 6 |
| 2012 | Improving Query for P2P SIP VoIPabstractP2PSIP (Peer-to-Peer SIP) is proposed to provide fully distributed multimedia communication systems. However, as the two sides of a coin, P2PSIP has improved the reliability and scalability of the traditional SIP (Session Initiation Protocol) networks, but also reduced the query performance of P2PSIP networks. In P2PSIP networks, the address resolving mechanism of the DHT (Distributed Hash Table) overlay is used to discover the IP address of the corresponding target user. Since there might be multiple overlay hops in resolving procedure, each of which is doing potentially time-intensive operations, it is high likely that the distributed characteristics of DHT overlay might result in an excessive CSD (Call Setup Delay), which is one of important indicators of the query performance of P2PSIP networks. In this paper, to improve the query performance of P2PSIP networks, we propose three approaches, including Bidirectional Chord, Asynchronous processing and Load balancing. By comparing experimental results in terms of RPS (Registrations-Per-Second), CPS (Calls-Per-Second) and CSD under traditional and improved P2PSIP networks, we conclude that the combination of these three approaches can effectively improve the query performance of P2PSIP networks. Linchen Yu |
TrustCom | 1 |
| 2012 | A novel data replication mechanism in P2P VoD system
Xiaofei Liao, Hai Jin 0001, Linchen Yu |
Future Gener. Comput. Syst. | 3 |
| 2011 | Integrated buffering schemes for P2P VoD services
Linchen Yu, Xiaofei Liao, Hai Jin 0001, Wenbin Jiang 0001 |
Peer-to-Peer Netw. Appl. | 1 |