Guoqing Bao

dblp:211/4054 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorComputer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 IoT-Based Precision Litchi Tracking and Counting Method Using Gated Metrics
abstract
Accurate and efficient multi-object tracking and counting methods are designed to address the challenges of counting in complex environments.This study presents a novel tracking and counting method called LitchiCount, integrating the multi-object tracking detection model LitchiDet with a counting module to address issues such as missing counts, repeated counts, and the lack of interpretability commonly found in traditional machine learning approaches. The method is designed with the guidance of the visual interpretable method Grad-CAM++, as well as the experimental validation method based on important features. To improve the detection accuracy of small targets under dense occlusion and overlapping, we proposed LitchiDet, which combines a small target detection layer, a decoupled fully connected attention with C3Ghost module (DFC-C3Ghost) and an efficient layer aggregation network block (ELANB). Our counting module improves target tracking accuracy and robustness in dense occlusion scenes while reducing counting errors from scene changes. We propose a Distance-generalized Intersection over Union association metric using a gating mechanism(DG-GM) and an AreaC counting strategy tailored to field intricate scenes. Finally, to enhance IoT deployment, we migrated LitchiCount to the Jetson AGX Xavier platform and optimized the model with TensorRT, significantly improving computational efficiency and real-time performance, particularly in resource-limited IoT environments, meeting real-time and low-power demands. The results demonstrated that our proposed method outperforms state-of-the-art detection models, as well as DeepSort-based counting methods in detection and counting. Importantly, by applying our method to the scenario of detecting and counting litchi from multiple perspectives in a field setting, we achieved low-repetitive and reliable counting, demonstrating the robust performance of this approach in real-world applications.
Jianqiang Lu, Guoqing Bao, Xiaoling Deng, Xiongzhe Han, Yubin Lan, Haiwei Wu
IEEE Internet Things J.2
2024 PresCount: Effective Register Allocation for Bank Conflict Reduction
abstract
Modern processors with large multi-banked register files often rely on hardware solutions to resolve bank conflicts efficiently. However, these hardware-based methods, while flexible, can incur runtime penalties and restrict the exploration of optimized hardware designs. In contrast, compiler-based methods for register bank assignments avoid runtime overhead. However, incorporating bank assignment into the complex register allocation process presents significant challenges, leading existing methods to adopt conservative approaches to avoid potential side effects. This paper introduces the novel register allocation method PresCount, which enhances the coloring strategy for the Register Conflict Graph (RCG) and incorporates a bank pressure tracking mechanism to improve performance. The integrated register bank assigner in PresCount effectively reduces bank conflicts, achieving remarkable reductions of 43.28% and 27.76%, respectively, compared to existing methods on platforms with rich register banks and limited register budgets, as demonstrated by SPECfp and CNN-KERNEL benchmarks. Furthermore, a subgroup splitting technique is introduced to facilitate register allocation under the bank-subgroup register file design, specifically our Domain-Specific Architecture (DSA) for AI computing. This technique demonstrates an impressive 99.85% reduction in bank conflicts for domain-specific kernel functions. By addressing the challenges of bank conflicts in register allocation, the proposed PresCount method showcases significant improvements in performance and efficiency for platforms with different register configurations and domain-specific workloads, allowing for more flexible exploration of optimized hardware designs.
Xiaofeng Guan, Hao Zhou 0009, Guoqing Bao, Handong Li, Jianguo Yao 0002
CGO3
2024 SPHINX: Search Space-Pruning Heterogeneous Task Scheduling for Deep Neural Networks
abstract
Given the tendency of increasingly heterogeneous AI systems and the large workload scale of deep neural networks (DNNs), there is an urgent demand for model scheduling to improve execution performance in heterogeneous computational systems. However, this is very challenging because the task scheduling under the high-dimensional search space is an NP-hard problem. Existing works either schedule under naive search spaces without simplifications or oversimplifies the optimisation, which is hard to strike a balance between efficiency and optimality.
Bowen Yuchi, Heng Shi 0005, Guoqing Bao
ICPP3
2024 UFront: Toward A Unified MLIR Frontend for Deep Learning
abstract
Automatic code generation for ML systems has gained popularity with the advent of compiler techniques like Multi-Level Intermediate Representation (Multi-Level IR, or MLIR). State-of-the-art MLIR frontends, including IREE-TF, Torch-MLIR, and ONNX-MLIR, aim to bridge the gap between ML frameworks and low-level hardware architectures through MLIR's progressive lowering pipeline. However, existing MLIR frontends encounter challenges such as inflexible high-level IR conversion, limited higher-level optimization opportunities, and reduced compatibility and efficiency, leading to software fragmentation and restricting their practical applications within the MLIR ecosystem. To address these challenges, we introduce UFront, a unified MLIR frontend employing a two-stage operator-to-operator compilation workflow. Unlike traditional frontends that compile model source code into binaries step by step with different MLIR transform passes, UFront decouples the process into two distinct stages. It first performs instantaneous model tracing, delegates traced computing nodes as standard Deep Neural Network (DNN) operators and transforms models written in different frameworks into unified high-level IR without relying on MLIR passes, enhancing conversion flexibility. Meanwhile, it performs high-level graph optimizations such as constant folding and operator fusion to produce more efficient high-level IR. In the second stage, UFront directly converts high-level IR into standard TOSA IR using proposed lowering patterns, eliminating transform redundancies and ensuring lower-level compatibility with existing ML compiler backends. This two-stage compilation approach enables consistent end-to-end code generation and optimization of various DNN models written in different formats within a single workflow. Extensive experiments on popular DNN models written in various frameworks demonstrate that UFront exhibits higher compatibility, faster end-to-end compilation, and is capable of producing more efficient binary execution compared to SOTA works.
Guoqing Bao, Heng Shi 0005, Chengyi Cui, Yalin Zhang 0004, Jianguo Yao 0002
ASE1
2022 COVID-MTL: Multitask learning with Shift3D and random-weighted loss for COVID-19 diagnosis and severity assessment
Guoqing Bao, Huai Chen, Tongliang Liu, Guanzhong Gong, Lisheng Wang, Xiuying Wang 0001
Pattern Recognit.1
2021 Identification of lncRNA Signature Associated With Pan-Cancer Prognosis
abstract
Long noncoding RNAs (lncRNAs) have emerged as potential prognostic markers in various human cancers as they participate in many malignant behaviors. However, the value of lncRNAs as prognostic markers among diverse human cancers is still under investigation, and a systematic signature based on these transcripts that related to pan-cancer prognosis has yet to be reported. In this study, we proposed a framework to incorporate statistical power, biological rationale, and machine learning models for pan-cancer prognosis analysis. The framework identified a 5-lncRNA signature (ENSG00000206567, PCAT29, ENSG00000257989, LOC388282, and LINC00339) from TCGA training studies (n = 1,878). The identified lncRNAs are significantly associated (all P ≤ 1.48E-11) with overall survival (OS) of the TCGA cohort (n = 4,231). The signature stratified the cohort into low- and high-risk groups with significantly distinct survival outcomes (median OS of 9.84 years versus 4.37 years, log-rank P = 1.48E-38) and achieved a time-dependent ROC/AUC of 0.66 at 5 years. After routine clinical factors involved, the signature demonstrated better performance for long-term prognostic estimation (AUC of 0.72). Moreover, the signature was further evaluated on two independent external cohorts (TARGET, n = 1,122; CPTAC, n = 391; National Cancer Institute) which yielded similar prognostic values (AUC of 0.60 and 0.75; log-rank P = 8.6E-09 and P = 2.7E-06). An indexing system was developed to map the 5-lncRNA signature to prognoses of pan-cancer patients. In silico functional analysis indicated that the lncRNAs are associated with common biological processes driving human cancers. The five lncRNAs, especially ENSG00000206567, ENSG00000257989 and LOC388282 that never reported before, may serve as viable molecular targets common among diverse cancers.
Guoqing Bao, Xiuying Wang 0001, Jianxiong Ji, Anjing Chen, Beihua Kong, Qifeng Yang, Cunzhong Yuan, Jian Wang 0120
IEEE J. Biomed. Health Informatics1
2020 A Bifocal Classification and Fusion Network for Multimodal Image Analysis in Histopathology
abstract
Recognition of key morphological features in histological slides is crucial for pathological diagnosis and monitoring therapeutic progress. However, the typical routine microscopic workflow is conducted by hand which is time-consuming and has unavoidable intra- and inter-observer variability like all human work. Therefore, we propose a bifocal classification and fusion network for the automated recognition and cross-modality analysis of diagnostic features in whole-slide multimodal images (WSIs). In brief, paired image tiles cropped from digitized tissue sections were fed into a modified dual-path CNN which accepts asymmetric inputs for classification, and then the inference results were converted to feature distribution heatmaps, which permit qualitative as well as quantitative morphological analyses of entire histological sections, even in combination with adjacent sections that have been stained differently. The multimodal heatmaps were aligned using image registration and fused for cross-modality analysis. Our experiments showed that the network achieved high recognition performance (AUCs of 0.985 and 0.988, and accuracies of 94.7% and 96.1% on two WSI modalities, respectively, against expert markings) and outperformed state-of-the-art methods without training on a large cohort or utilizing domain transfer. In addition, the new method involves a self-contained inference and fusion process and thus harbors significant potential for speeding up microscopic analysis workflows.
Guoqing Bao, Manuel B. Graeber, Xiuying Wang 0001
ICARCV1
2020 Depthwise Multiception Convolution for Reducing Network Parameters without Sacrificing Accuracy
abstract
Deep convolutional neural networks have been proven successful in multiple benchmark challenges in recent years. However, the performance improvements are heavily reliant on increasingly complex network architecture and a high number of parameters, which require ever increasing amounts of storage and memory capacity. Depthwise separable convolution (DSConv) can effectively reduce the number of required parameters through decoupling standard convolution into spatial and cross-channel convolution steps. However, the method causes a degradation of accuracy. To address this problem, we present depthwise multiception convolution, termed Multiception, which introduces layer-wise multiscale kernels to learn multiscale representations of all individual input channels simultaneously. We have carried out the experiment on four benchmark datasets, i.e. Cifar-10, Cifar-100, STL-10 and ImageNet32×32, using five popular CNN models, Multiception achieved accuracy promotion in all models and demonstrated higher accuracy performance compared to related works. Meanwhile, Multiception significantly reduces the number of parameters of standard convolution-based models by 32.48 % on average while still preserving accuracy.
Guoqing Bao, Manuel B. Graeber, Xiuying Wang 0001
ICARCV1