Yingjie Cao

dblp:94/10767 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Security and privacy · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorComputer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 IsolatOS: Detecting Double Fetch Bugs in COTS RTOS by Re-enabling Kernel Isolation
Yingjie Cao, Xiaogang Zhu 0001, Dean Sullivan, Lei Xue 0001, Chenxiong Qian, Minrui Yan, Xiapu Luo
NDSS1
2025 A-Mel: A Resource-Efficient Deep Learning Agent for Precise Early Melanoma Diagnosis
abstract
Early and precise diagnosis of melanoma significantly improves patient prognosis. However, current diagnostic methods are challenged by hair occlusion, diverse lesion morphologies, and limited computational resources in clinical settings. In this paper, we propose A-Mel, a resource-efficient deep learning agent designed for precise early melanoma diagnosis. A-Mel integrates a unified agent framework comprising three sequential modules: hair artifact removal preprocessing, pathological lesion segmentation, and melanoma classification. Specifically, we employ a modified$\mathrm{U}^{2} \text{Net}++$architecture enhanced with depthwise separable convolutions to significantly reduce model parameters and computational load. Furthermore, we introduce a novel Dual-Path Spatial Attention (DPSA) mechanism integrated with Efficient Channel Attention (ECA), achieving comprehensive attention enhancement across multiple scales and dimensions. The network is further augmented with an Atrous Spatial Pyramid Pooling (ASPP) module to strengthen multi-scale feature representation. Experimental evaluations conducted on a test set of approximately 700 dermoscopic images with hair artifacts selected from the ISIC 2019 and ISIC 2020 datasets demonstrate that our A-Mel model achieves a state-of-the-art binary classification accuracy of 99.12 %, outperforming existing approaches through extensive comparative analyses. Our segmentation model is “extremely lightweight” in terms of parameters, volume, and speed, with excellent performance for edge deployment scenarios. The codes for A-Mel are made publicly available at https://github.com/Yaoooyu/A-Mel.
Yaoyu Liu, Yingjie Cao, Denan Liu, Sike Chen, Shaoliang Peng
BIBM3
2025 The Photoacoustic Quality-Enhancement Neural Network Processor with the Scalable and End-to-End Architecture by Improving the Sparsity Level
abstract
Recent advancements have marked significant progress in photoacoustic imaging as an effective method for acquiring deep bio-tissue visuals in modern medical clinical therapy and the efficacy of U-Net and its variants has been established for imaging quality enhancement in this field. Unlike common computer vision datasets such as ImageNet [1] and PASCAL VOC [2], biomedical images exhibit highly structured patterns, low spatial resolution, and single-channel modality, as shown in Fig. 1. Additionally, the U-Net parameters trained for medical super-resolution tasks demonstrate a high sparsity ratio, making them suitable for implementation on edge-computing platforms. Therefore, developing an energy-efficient photoacoustic imaging setup in this area is a natural progression. However, this development is constrained by the current neural network architectures, which are built around a U-Net backbone. The multi-stage feature extractor, skip connection integration across different blocks, and the encoder-decoder backbone design pose significant challenges to cutting-edge computational hardware platforms. In this study, a scalable, sparsity-supported neural network accelerator architecture for bio-tissue imaging quality enhancement is proposed to meet the stringent requirements of latency and energy efficiency, as depicted in Fig. 2. This architecture achieves desired performance improvements by exploring the sparsity possibilities in neural network during the training process and implementing an end-to-end pixel-first hardware design to minimize data movement and support sparsity computation. Compared with the state-of-the-art related works, this optimized architecture has achieved minimum on-chip storage overhead and the fastest frame for the application of photoacoustic imaging quality enhancement. The scalable architecture has also been implemented on a Xilinx XCZU9EG FPGA and attains a performance of PSNR@ 24 dB and a frame rate of 164 fps at a working frequency of 250 MHz.
Zhengyuan Zhang 0002, Caijie Liang, Boyi Dong, Yange Wang, Zhongzhiguang Lu, Xiangjun Yin, Shenglong Zhuo, Yifan Wu 0009, Yingjie Cao, Tianyang Zhou, Jian Qian, Patrick Chiang 0001, Lei Qiu 0002, Yuanjin Zheng
ISCAS12
2024 Revisiting Automotive Attack Surfaces: a Practitioners' Perspective
abstract
As modern vehicles become increasingly complex in terms of both external attack surfaces and internal in-vehicle network (IVN) topology, ensuring their cybersecurity remains a challenge. Existing standards and regulations, such as WP29 R155e and ISO 21434, attempt to establish a baseline for automotive cybersecurity, but their sufficiency in addressing the evolving threats is unclear. To fill in this gap, we first carried out an in-depth interview study with 15 experts in automotive cybersecurity, uncovering the particular challenges encountered during security activities and the limitations of current regulations. We identified 20 key insights from the interview data, ranging from the challenges and gaps in the existing automotive security industry to the limitations and recommendations for current regulations. Notably, we discovered that the quality of threat cases provided by existing regulations is unsatisfactory, and the Threat Analysis and Risk Assessment (TARA) process is often highly inefficient due to the lack of automatic tools. In response to the above limitations, we first built an improved threat database for automotive systems using the collected interview data, which enhanced the existing database both quantitatively and qualitatively. Additionally, we present CarVal, a datalog-based approach designed to infer multi-stage attack paths in IVNs and calculate risk values, thereby making TARA more efficient for automotive systems. By applying CarVal to five real vehicles, we performed extensive security analysis based on the generated attack paths and successfully exploited the corresponding attack chains in the newly gateway-segmented IVN, uncovering new automotive attack surfaces that previous research failed to cover, including the in-vehicle browser, official mobile app, backend server, and in-vehicle malware.
Pengfei Jing, Yingjie Cao, Le Yu 0002, Yuefeng Du 0006, Chenxiong Qian, Xiapu Luo, Sen Nie, Shi Wu
SP3
2024 UAV-Based Emergency Communications: An Iterative Two-Stage Multiagent Soft Actor-Critic Approach for Optimal Association and Dynamic Deployment
abstract
This paper investigates future emergency wireless communication systems based on multiple unmanned vehicles cooperative deployment. A terrestrial carrier vehicle with wireless communication and management capabilities are deployed to release multiple unmanned aerial vehicles (UAVs) which will serve as aerial mobile stations (UAV-BSs) to cover a disaster affected area, forming an emergency Internet of Things (IoT) network. Under the proposed system architecture, we formulate a joint optimization challenge considering the UAV-BSs’ dynamic deployment positions and the association policy between user equipments (UEs) and BSs to maximize the throughput and coverage in dynamic scenarios as a time-varying mixed-integer non-convex sequential programming (MINSP) problem. To solve this problem, we first investigate the impact of decision delay caused by physical networking and computing environment on system performance to illustrate the urgent need for efficient algorithms. Then, a two-stage iterative training algorithm called centralized training multi-agent soft actor-critic with branch-and-cut (CT-MASAC-BAC) is proposed for computing globally optimal solutions. Numerical results show that CT-MASAC-BAC outperforms the heuristic algorithms and other benchmark deep reinforcement learning algorithms in terms of system utility. Furthermore, the experimental results show that the proposed algorithm is scalable with an increasing number of deployed UAV-BSs, contributing to potentially increased performance with more serving UAV-BSs.
Yingjie Cao, Yang Luo 0001, Haifen Yang, Chunbo Luo
IEEE Internet Things J.1
2019 SPFC: An Effective Optimization for Vertex-Centric Graph Processing Systems
abstract
The real-world demands of mining big data and smart data of graph structure have led to an active research of distributed graph processing. Many distributed graph processing systems [19], [22], [23] adopt a vertex-centric programming paradigm. In these systems, messages are passed between vertices to propagate the latest states. The communication efficiency and the high overhead of synchronization are two key considerations of these systems [8], [12]. In this paper, we propose a Slow Passing Fast Consuming (SPFC) approach which can effectively improve the overall performance of vertex-centric graph processing systems. In our approach, the message passing is slow but the consuming is fast. More specifically, at the message sender side, priority is given to those smart messages which contribute more to the algorithm convergence, and at the message receiver side, messages are consumed right after arriving without any delay and intermediate buffer. Besides, by using a two-phase termination check protocol, the global synchronous barrier can be completely eliminated. In addition, based on the slow message passing strategy, further performance improvement can be achieved with some accuracy loss by eliminating those messages which are less useful for algorithm convergence. We implement our approach based on Apache Giraph [1] and evaluate it on a 12-machine cluster. The experimental results show that our method can effectively reduce the amount of message traffic and achieve up to an order of magnitude performance improvement compared with Giraph and GraphLab [3].
Jianxin Li 0002, Yingjie Cao, Yangyang Zhang 0001, Md. Zakirul Alam Bhuiyan, Bo Li 0005
IEEE Trans. Sustain. Comput.2
2018 Parallel Reasoning of Graph Functional Dependencies
abstract
This paper develops techniques for reasoning about graph functional dependencies (GFDs). We study the satisfiability problem, to decide whether a given set of GFDs has a model, and the implication problem, to decide whether a set of GFDs entails another GFD. While these fundamental problems are important in practice, they are coNP-complete and NP-complete, respectively. We establish a small model property for satisfiability, showing that if a set ? of GFDs is satisfiable, then it has a model of a size bounded by the size |Σ| of Σ; similarly we prove a small model property for implication. Based on the properties, we develop algorithms for checking the satisfiability and implication of GFDs. Moreover, we provide parallel algorithms that guarantee to reduce running time when more processors are used, despite the intractability of the problems. We experimentally verify the efficiency and scalability of the algorithms.
Wenfei Fan, Yingjie Cao
ICDE3
2017 PMS: an Effective Approximation Approach for Distributed Large-scale Graph Data Processing and Mining
abstract
Recently, large-scale graph data processing and mining has drawn great attention, and many distributed graph processing systems have been proposed. However, large-scale graph processing remains a challenging problem. Because the computation time in some cases is still unacceptable especially when the time is limited. As illustrated in Table 1, nearly three hours are needed when running Single-Source Shortest Path algorithm on the USA-road dataset using performant open-source distributed graph processing systems.
Yingjie Cao, Yangyang Zhang 0001, Jianxin Li 0002
CIKM1
2013 An FPGA Based PCI-E Root Complex Architecture for Standalone SOPCs
abstract
We present an FPGA (field programmable gate array) based PCI-E (PCI-Express) root complex architecture for SOPCs (System-on-a-Programmable-Chip) in this paper. In our work, the system on the FPGA serves as a PCIE master device rather than a PCIE endpoint, which is usually a common practice as a co-processing device driven by a desktop computer or a server. We use this system to control a PCIE endpoint, which is also an FPGA based endpoint implemented on another FPGA board. This architecture requires only IP cores free of charge. We also provide basic software driver so that specific device driver can be developed on it to control popular PCIE device in the future, i.e. ethernet card or graphic card. The whole architecture has been implemented on Xilinx Virtex-6 FPGAs to indicate that this architecture is a feasible approach to standalone SOPCs, which has better efficiencies than those with additional generic controlling processors.
Yingjie Cao, Yongxin Zhu 0001, Xu Wang 0010, Meikang Qiu
FCCM1
2011 Efficient Pattern Detection for Embedded Optical Bio-sensing System
abstract
To enable pattern detection in optical signals from a novel optical biosensor used in medical embedded system, we propose a set of efficient algorithms and their corresponding implementation on FPGA (field programmable gate array). The optical biosensor is a porous silicon micro cavity membrane, which can generate different optical reflectance spectra for varieties of molecule solutions. In measured reflectance spectra of the membrane, our design is able to detect the shift of the resonant dip which is considered as the pattern to distinguish target molecule solution of different concentration. According to measured results, besides the much higher sensitivity of the novel optical sensor than classic electrochemical methods, our FPGA based implementation of our detection algorithms also shows significant speedup over software implementation on PC. The small chip area cost of FPGA implementation of our detection algorithms further ensures feasibility of ASIC (application specific integrated circuits) to incorporate both sensors and signal processing in near future.
Yingjie Cao, Yongxin Zhu 0001, Guoguang Rong, Meikang Qiu
DASC1