EDBT 2026 Demo / reviewers in the wild / expert
Xiaoguang Guo
dblp:40/8540
· DBLP profile ↗
9ranked-venue papers
0as first author
8since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Weakly Supervised Virtual Immunohistochemistry Staining via Schrödinger Bridge MethodabstractImmunohistochemistry (IHC) staining provides precise localization and qualitative analysis for tumor diagnosis, while its application is often constrained by complexity and high costs. Recently, many researchers have employed virtual staining based on deep learning to translate hematoxylin and eosin (H&E) images into IHC images, presenting a more efficient and cost-effective alternative. However, these methods usually rely on Generative Adversarial Networks (GANs), which are prone to various issues such as mode collapse, thereby limiting the effectiveness. To address this issue, we propose a weakly supervised virtual staining method based on the Schrödinger bridge called StainSB. This method establishes an optimal random process between the source and target distributions, effectively translating H&E images into IHC images. Specifically, we design a regional color state loss to model the pathological similarity between the generated and real IHC images, thereby incorporating pathological information into the generation process. Furthermore, the proposed aggregation strategy enables the generated images to achieve a balance between image quality and pathological consistency. Extensive experiments on two public benchmark datasets show that the proposed StainSB method achieves state-of-the-art performance across multiple metrics. Fanhao Qiu, Xiaoguang Guo, Zhengxia Wang |
BIBM | 3 |
| 2024 | Compressed data direct computing for Chinese dataset on DCU
Yani Liu, Feng Zhang 0007, Zaifeng Pan, Xiaoguang Guo, Yihua Hu 0003, Xiao Zhang 0001, Xiaoyong Du 0001 |
CCF Trans. High Perform. Comput. | 4 |
| 2024 | Enabling Efficient Deep Learning on MCU With Transient Redundancy EliminationabstractDeploying deep neural networks (DNNs) with satisfactory performance in resource-constrained environments is challenging. This is especially true of microcontrollers due to their tight space and computational capabilities. However, there is a growing demand for DNNs on microcontrollers, as executing large DNNs on microcontrollers is critical to reducing energy consumption, increasing performance efficiency, and eliminating privacy concerns. This paper presents a novel and systematic data redundancy elimination method to implement efficient DNNs on microcontrollers through innovations in computation and space optimization. By making the optimization itself a trainable component in the target neural networks, this method maximizes performance benefits while keeping the DNN accuracy stable. Experiments are performed on two microcontroller boards with three popular DNNs, namely CifarNet, ZfNet and SqueezeNet. Experiments show that this solution eliminates more than 96% of computations in DNNs and makes them fit well on microcontrollers, yielding 3.4-5$\times$speedup with little loss of accuracy. Jiesong Liu, Feng Zhang 0007, Jiawei Guan, Hsin-Hsuan Sung, Xiaoguang Guo, Saiqin Long, Xiaoyong Du 0001, Xipeng Shen |
IEEE Trans. Computers | 5 |
| 2024 | Data-Aware Adaptive Compression for Stream ProcessingabstractStream processing has been in widespread use, and one of the most common application scenarios is SQL query on streams. By 2021, the global deployment of IoT endpoints reached 12.3 billion, indicating a surge in data generation. However, the escalating demands for high throughput and low latency in stream processing systems have posed significant challenges due to the increasing data volume and evolving user requirements. We present a compression-based stream processing engine, called CompressStreamDB, which enables adaptive fine-grained stream processing directly on compressed streams, to significantly enhance the performance of existing stream processing solutions. CompressStreamDB utilizes nine diverse compression methods tailored for different stream data types and integrates a cost model to automatically select the most efficient compression schemes. CompressStreamDB provides high throughput with low latency in stream SQL processing by identifying and eliminating redundant data among streams. Our evaluation demonstrates that CompressStreamDB improves average performance by 3.84× and reduces average delay by 68.0% compared to the state-of-the-art stream processing solution for uncompressed streams, along with 68.7% space savings. Besides, our edge trials show an average throughput/price ratio of 9.95× and a throughput/power ratio of 7.32× compared to the cloud design. Yu Zhang 0027, Feng Zhang 0007, Hourun Li, Shuhao Zhang 0001, Xiaoguang Guo, Yuxing Chen 0003, Anqun Pan, Xiaoyong Du 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | G-Learned Index: Enabling Efficient Learned Index on GPUabstractAI and GPU technologies have been widely applied to solve big data problems. The total data volume worldwide reaches 200 zettabytes in 2022. How to efficiently index the required content among massive data becomes serious. Recently, a promising learned index has been proposed to address this challenge: It has extremely high efficiency while retaining marginal space overhead. However, we notice that previous learned indexes have mainly focused on CPU architecture, while ignoring the advantages of GPU. Because traditional indexes like B-Tree, LSM, and bitmap have greatly benefited from GPU acceleration, a combination of a learned index and GPU has great potentials to reach tremendous speedups. In this paper, we propose a GPU-based learned index, called G-Learned Index, to significantly improve the performance of learned index structures. The primary challenges in developing G-Learned Index lie in the use of thousands of GPU cores including minimization of synchronization and branch divergence, data structure design for parallel operations, and usage of memory bandwidth including limited memory transactions and multi-memory hierarchy. To overcome these challenges, a series of novel technologies are developed, including efficient thread organization, succinct data structures, and heterogeneous memory hierarchy utilization. Compared to the state-of-the-art learned index, the proposed G-Learned Index achieves an average of 174× speedup (and 107× of its parallel version). Meanwhile, we attain 2× less query time over the state-of-the-art GPU B-Tree. Our further exploration of range queries shows that G-Learned Index is 17× faster than CPU multi-dimensional learned index. We have made G-Learned Index available athttps://anonymous.4open.science/r/G-Learned-Index-8D89. Jiesong Liu, Feng Zhang 0007, Lv Lu, Chang Qi, Xiaoguang Guo, Dong Deng 0001, Guoliang Li 0001, Huanchen Zhang, Jidong Zhai, Hechen Zhang, Yuxing Chen 0003, Anqun Pan, Xiaoyong Du 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2023 | Space-Efficient TREC for Enabling Deep Learning on MicrocontrollersabstractDeploying deep neural networks (DNNs) for a resource-constrained environment and achieving satisfactory performance is challenging. It is especially so on microcontrollers for their stringent space and computing power. This paper focuses on new ways to make TREC, an optimization recently proposed to enable computation reuse in DNNs, space and time efficient on Microcontrollers. The solution maximizes the performance benefits while keeping the DNN accuracy stable. Experiments show that the solution eliminates over 96% computations in DNNs and makes them fit well into microcontrollers, producing 3.4-5× speedups with only marginal accuracy loss. Jiesong Liu, Feng Zhang 0007, Jiawei Guan, Hsin-Hsuan Sung, Xiaoguang Guo, Xiaoyong Du 0001, Xipeng Shen |
ASPLOS (3) | 5 |
| 2023 | Expanding the Edge: Enabling Efficient Winograd CNN Inference With Deep Reuse on Edge DeviceabstractDeep learning on edge devices is becoming increasingly important, especially with the explosion of IoT devices. For example, the total number of devices connected to IoT reaches 29 billion in 2022. Convolutional neural networks (CNNs), as common deep learning representatives, are among the most popular neural networks in knowledge and data engineering. However, CNN employs a high degree of computing. In comparison to the training phase, the inference process is more frequently done on low-power computing equipments, such as edge devices. The limited computing resource and high computation pressure limit the effective use of CNN algorithms at the edge. Fortunately, a minimal filtering algorithm called Winograd can reduce convolution calculations by minimizing multiplication operations. We find that Winograd convolution can be accelerated further bydeep reusetechnique, which reuses the similar data and computation processes. In this paper, we propose a new inference method, called DREW, which combines deep reuse with Winograd for further accelerating CNNs. DREW handles three difficulties. First, it can detect the similarities from the complex minimal filtering patterns by clustering. Second, it reduces the online clustering cost in a reasonable range. Third, it provides an adjustable method in clustering granularity balancing the performance and accuracy. We perform evaluation on Raspberry PI and NVIDIA Jetson AGX Xavier edge devices, and experiments show that on five popular networks, 1) DREW further accelerates the Winograd convolution by an average of 8.27× speedup. Even for the highly parallel Winograd implementation, DREW still can provide 2.21× speedup. 2) When DREW is applied to end-to-end Winograd CNN inferences, DREW achieves 5.94× the average performance speedup with no ($< $0.4%) accuracy loss. 3) Energy consumption is an important factor for inference in practice. DREW reduces the number of convolution operations to 10% of the original operations, thus achieving up to 60% energy-efficiency benefits than the original Winograd inference. Feng Zhang 0007, Jiawei Guan, Zhen Zheng, Xiaoguang Guo, Xiao Zhang 0001, Xiaoyong Du 0001, Xipeng Shen |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2021 | iMLBench: A Machine Learning Benchmark Suite for CPU-GPU Integrated ArchitecturesabstractUtilizing heterogeneous accelerators, especially GPUs, to accelerate machine learning tasks has shown to be a great success in recent years. GPUs bring huge performance improvements to machine learning and greatly promote the widespread adoption of machine learning. However, the discrete CPU-GPU architecture design with high PCIe transmission overhead decreases the GPU computing benefits in machine learning training tasks. To overcome such limitations, hardware vendors release CPU-GPU integrated architectures with shared unified memory. In this article, we design a benchmark suite for machine learning training on CPU-GPU integrated architectures, called iMLBench, covering a wide range of machine learning applications and kernels. We mainly explore two features on integrated architectures: 1) zero-copy, which means that the PCIe overhead has been eliminated for machine learning tasks and 2) co-running, which means that the CPU and the GPU co-run together to process a single machine learning task. Our experimental results on iMLBench show that the integrated architecture brings an average 7.1× performance improvement over the original implementations. Specifically, the zero-copy design brings 4.65× performance improvement, and co-running brings 1.78× improvement. Moreover, integrated architectures exhibit promising results from both performance-per-dollar and energy perspectives, achieving 6.50× performance-price ratio while 4.06× energy efficiency over discrete GPUs. The benchmark is open-sourced at https://github.com/ChenyangZhang-cs/iMLBench. Chenyang Zhang 0005, Feng Zhang 0007, Xiaoguang Guo, Bingsheng He, Xiao Zhang 0001, Xiaoyong Du 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | A New Approach to Track Multiple Vehicles With the Combination of Robust Detection and Two ClassifiersabstractIt plays an important role to accurately track multiple vehicles in intelligent transportation, especially in intelligent vehicles. Due to complicated traffic environments it is difficult to track multiple vehicles accurately and robustly, especially when there are occlusions among vehicles. To alleviate these problems, a new approach is proposed to track multiple vehicles with the combination of robust detection and two classifiers. An improved ViBe algorithm is proposed for robust and accurate detection of multiple vehicles. It uses the gray-scale spatial information to build dictionary of pixel life length to make ghost shadows and object's residual shadows quickly blended into the samples of the background. The improved algorithm takes good post-processing method to restrain dynamic noise. In this paper, we also design a method using two classifiers to further attack the problem of failure to track vehicles with occlusions and interference. It classifies tracking rectangles with confidence values between two thresholds through combining local binary pattern with support vector machine (SVM) classifier and then using a convolutional neural network (CNN) classifier for the second time to remove the interference areas between vehicles and other moving objects. The two classifiers method has both time efficiency advantage of SVM and high accuracy advantage of CNN. Comparing with several existing methods, the qualitative and quantitative analysis of our experiment results showed that the proposed method not only effectively removed the ghost shadows, and improved the detection accuracy and real-time performance, but also was robust to deal with the occlusion of multiple vehicles in various traffic scenes. Weidong Min, Mengdan Fan, Xiaoguang Guo |
IEEE Trans. Intell. Transp. Syst. | 3 |