Baoquan Liu

dblp:54/784 · DBLP profile ↗
← Back
19ranked-venue papers
13as first author
1since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 12 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Geometric modeling and processing · 38% Rendering · 38% Computer animation and physical simulation · 25%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer animation and physical simulation
interface tracking
0.212016
Multiphase Interface Tracking with Fast Semi-Lagrangian Contouring · IEEE Trans. Vis. Comput. Graph. 2016
Geometric modeling and processing
polygonization
0.212016
Multiphase Interface Tracking with Fast Semi-Lagrangian Contouring · IEEE Trans. Vis. Comput. Graph. 2016
Geometric modeling and processing
surface reconstruction
0.212016
Multiphase Interface Tracking with Fast Semi-Lagrangian Contouring · IEEE Trans. Vis. Comput. Graph. 2016
Rendering › volume rendering
GPU volume rendering
0.212013
Octree Rasterization: Accelerating High-Quality Out-of-Core GPU Volume Rendering · IEEE Trans. Vis. Comput. Graph. 2013
Rendering › rendering optimization › rendering acceleration
out-of-core rendering
0.212013
Octree Rasterization: Accelerating High-Quality Out-of-Core GPU Volume Rendering · IEEE Trans. Vis. Comput. Graph. 2013
Rendering
volume rendering
0.212013
Octree Rasterization: Accelerating High-Quality Out-of-Core GPU Volume Rendering · IEEE Trans. Vis. Comput. Graph. 2013
GPUs and heterogeneous computing › GPU rendering
GPU rasterization
0.012013
Octree Rasterization: Accelerating High-Quality Out-of-Core GPU Volume Rendering · IEEE Trans. Vis. Comput. Graph. 2013

Methods — techniques the papers use, named apart from their topics

ray casting · 0.3octree · 0.3semi-lagrangian contouring · 0.2bisection · 0.2adaptive data structure · 0.2empty-space culling · 0.2empty space culling · 0.2
YearPublicationVenuePosition
2026 Robust Malicious Traffic Representation Under Instance-Dependent Noise via Entropy Selection and Semisupervised Learning
abstract
Traffic labeling methods have inherent limitations, including the potential to mislabel hard samples and introduce noisy labels when annotating real-world malicious traffic. Training on such noisy-labeled samples significantly degrades both the training effectiveness and evaluation reliability of malicious traffic detection models. This study proposes MEn2SLe, a method for robust malicious traffic representation under instance-dependent noise (IDN) based on entropy-guided selection and semi-supervised learning. First, an algorithm for constructing traffic datasets with IDN labels is introduced. Subsequently, leveraging the distributional characteristics of IDN, an information entropy-based sample selection strategy is proposed to identify high-confidence samples by analyzing per-sample training stability. To achieve robust representations, MEn2SLe integrates three complementary training strategies: noise label correction, unsupervised inter-class similarity learning, and ground-truth sample contrastive learning. These strategies collectively mitigate the influence of noisy labels and prevent the model from overfitting to label noise. Extensive experiments were conducted to evaluate the performance of MEn2SLe. The results demonstrate that MEn2SLe surpasses baseline methods in both robust training and data pruning tasks across IDN, CCN, and SYM noise scenarios. Specifically, MEn2SLe achieves classification accuracies of 90% and 82% on the CIC-IDS-2017 and Malicious_TLS datasets, respectively, at a 50% noise level, outperforming the best baseline by 13.6% and 9.87%. To further evaluate generalizability, the impact of noisy labels generated by different labeling methods was also assessed. The results confirm that MEn2SLe effectively resists the adverse effects of noise introduced by diverse annotation approaches, demonstrating strong generalization capability.
Degang Li, Qingjun Yuan, Xi Chen 0045, Baoquan Liu
IEEE Trans. Reliab.4
2018 Improving Utility of GPU in Accelerating Industrial Applications With User-Centered Automatic Code Translation
abstract
Small to medium enterprises (SMEs), particularly those whose business is focused on developing innovative produces, are limited by a major bottleneck in the speed of computation in many applications. The recent developments in GPUs have been the marked increase in their versatility in many computational areas. But due to the lack of specialist GPUprogramming skills, the explosion of GPU power has not been fully utilized in general SME applications by inexperienced users. Also, the existing automatic CPU-to-GPU code translators are mainly designed for research purposes with poor user interface design and are hard to use. Little attentions have been paid to the applicability, usability, and learnability of these tools for normal users. In this paper, we present an online automated CPU-to-GPU source translation system (GPSME) for inexperienced users to utilize the GPU capability in accelerating general SME applications. This system designs and implements a directive programming model with a new kernel generation scheme and memory management hierarchy to optimize its performance. A web service interface is designed for inexperienced users to easily and flexibly invoke the automatic resource translator. Our experiments with nonexpert GPU users in four SMEs reflect that a GPSME system can efficiently accelerate real-world applications with at least 4× and have a better applicability, usability, and learnability than the existing automatic CPU-to-GPU source translators.
Po Yang 0001, Feng Dong 0005, Valeriu Codreanu, David Williams 0002, Jos B. T. M. Roerdink, Baoquan Liu, Amjad Anvari-Moghaddam, Geyong Min
IEEE Trans. Ind. Informatics6
2016 Parallel Marching Blocks: A Practical Isosurfacing Algorithm for Large Data on Many-Core Architectures
abstract
Abstract Interactive isosurface visualisation has been made possible by mapping algorithms to GPU architectures. However, current state‐of‐the‐art isosurfacing algorithms usually consume large amounts of GPU memory owing to the additional acceleration structures they require. As a result, the continued limitations on available GPU memory mean that they are unable to deal with the larger datasets that are now increasingly becoming prevalent. This paper proposes a new parallel isosurface‐extraction algorithm that exploits the blocked organisation of the parallel threads found in modern many‐core platforms to achieve fast isosurface extraction and reduce the associated memory requirements. This is achieved by optimising thread co‐operation within thread‐blocks and reducing redundant computation; ultimately, an indexed triangular mesh can be produced. Experiments have shown that the proposed algorithm is much faster (up to 10×) than state‐of‐the‐art GPU algorithms and has a much smaller memory footprint, enabling it to handle much larger datasets (up to 64×) on the same GPU.
Baoquan Liu, Gordon Clapworthy, Feng Dong 0005, Enhua Wu
Comput. Graph. Forum1
2016 Evaluating automatically parallelized versions of the support vector machine
abstract
Summary The support vector machine (SVM) is a supervised learning algorithm used for recognizing patterns in data. It is a very popular technique in machine learning and has been successfully used in applications such as image classification, protein classification, and handwriting recognition. However, the computational complexity of the kernelized version of the algorithm grows quadratically with the number of training examples. To tackle this high computational complexity, we have developed a directive‐based approach that converts a gradient‐ascent based training algorithm for the CPU to an efficient graphics processing unit (GPU) implementation. We compare our GPU‐based SVM training algorithm to the standard LibSVM CPU implementation, a highly optimized GPU‐LibSVM implementation, as well as to a directive‐based OpenACC implementation. The results on different handwritten digit classification datasets demonstrate an important speed‐up for the current approach when compared to the CPU and OpenACC versions. Furthermore, our solution is almost as fast and sometimes even faster than the highly optimized CUBLAS‐based GPU‐LibSVM implementation, without sacrificing the algorithm's accuracy. Copyright © 2014 John Wiley & Sons, Ltd.
Valeriu Codreanu, Bob Dröge, David Williams 0002, Burhan Yasar, Po Yang 0001, Baoquan Liu, Feng Dong 0005, Olarik Surinta, Lambert Schomaker, Jos B. T. M. Roerdink, Marco A. Wiering
Concurr. Comput. Pract. Exp.6
2016 GSWO: A programming model for GPU-enabled parallelization of sliding window operations in image processing
Po Yang 0001, Gordon Clapworthy, Feng Dong 0005, Valeriu Codreanu, David Williams 0002, Baoquan Liu, Jos B. T. M. Roerdink, Zhikun Deng
Signal Process. Image Commun.6
2016 Multiphase Interface Tracking with Fast Semi-Lagrangian Contouring
abstract
We propose a semi-Lagrangian method for multiphase interface tracking. In contrast to previous methods, our method maintains an explicit polygonal mesh, which is reconstructed from an unsigned distance function and an indicator function, to track the interface of arbitrary number of phases. The surface mesh is reconstructed at each step using an efficient multiphase polygonization procedure with precomputed stencils while the distance and indicator function are updated with an accurate semi-Lagrangian path tracing from the meshes of the last step. Furthermore, we provide an adaptive data structure, multiphase distance tree, to accelerate the updating of both the distance function and the indicator function. In addition, the adaptive structure also enables us to contour the distance tree accurately with simple bisection techniques. The major advantage of our method is that it can easily handle topological changes without ambiguities and preserve both the sharp features and the volume well. We will evaluate its efficiency, accuracy and robustness in the results part with several examples.
Xiaosheng Li, Xiaowei He 0004, Xuehui Liu, Jian J. Zhang 0001, Baoquan Liu, Enhua Wu
IEEE Trans. Vis. Comput. Graph.5
2015 IsoBAS: A binary accelerating structure for fast isosurface rendering on GPUs
Baoquan Liu, Gordon Clapworthy, Feng Dong 0005
Comput. Graph.1
2014 Multiphase surface tracking with explicit contouring
abstract
We introduce a novel framework for tracking multiphase interfaces with explicit contouring technique. In our framework, an unsigned distance function and an additional indicator function are used to represent the multiphase system. Our method maintains the explicit polygonal meshes that define the multiphase interfaces. At each step, distance function and indicator function are updated via semi-Lagrangian path tracing from the meshes of the last step. Interface surfaces are then reconstructed by polygonization procedures with precomputed stencils and further smoothed with a feature-preserving non-manifold smoothing algorithm to stay in good quality. Our method is easy to be implemented and incorporated into multiphase simulation, such as immiscible fluids, crystal grain growth and geometric flows. We demonstrate our method with several level set tests, including advection, propagation, etc., and couple it to some existing fluid simulators. The results show that our approach is stable, flexible, and effective for tracking multiphase interfaces.
Xiaosheng Li, Xiaowei He 0004, Xuehui Liu, Baoquan Liu, Enhua Wu
VRST4
2014 Parallel centerline extraction on the GPU
Baoquan Liu, Alexandru C. Telea, Jos B. T. M. Roerdink, Gordon Clapworthy, David Williams 0002, Po Yang 0001, Feng Dong 0005, Valeriu Codreanu, Alessandro Chiarini
Comput. Graph.1
2013 Octree Rasterization: Accelerating High-Quality Out-of-Core GPU Volume Rendering
abstract
We present a novel approach for GPU-based high-quality volume rendering of large out-of-core volume data. By focusing on the locations and costs of ray traversal, we are able to significantly reduce the rendering time over traditional algorithms. We store a volume in an octree (of bricks); in addition, every brick is further split into regular macrocells. Our solutions move the branch-intensive accelerating structure traversal out of the GPU raycasting loop and introduce an efficient empty-space culling method by rasterizing the proxy geometry of a view-dependent cut of the octree nodes. This rasterization pass can capture all of the bricks that the ray penetrates in a per-pixel list. Since the per-pixel list is captured in a front-to-back order, our raycasting pass needs only to cast rays inside the tighter ray segments. As a result, we achieve two levels of empty space skipping: the brick level and the macrocell level. During evaluation and testing, this technique achieved 2 to 4 times faster rendering speed than a current state-of-the-art algorithm across a variety of data sets.
Baoquan Liu, Gordon Clapworthy, Feng Dong 0005, Edmond C. Prakash
IEEE Trans. Vis. Comput. Graph.1
2011 Accelerating Tumour Growth Simulations on Many-Core Architectures: A Case Study on the Use of GPGPU within VPH
abstract
Simulators of tumour growth can estimate the evolution of tumour volume and the quantity of various categories of cells as functions of time. However, the execution time of each simulation often takes several dozens of minutes (depending upon the dataset resolution), which clearly prevents easy interaction. The modern graphics processing unit (GPU) is not only a powerful graphics engine but also a highly parallel programmable processor featuring peak arithmetic performance and memory bandwidth that substantially outpaces its CPU counterpart. However, despite this, the GPU is little used in the context of the Virtual Physiological Human (VPH). This paper provides a case study to demonstrate the performance advantages that can be gained by using the GPU appropriately in the context of a VPH project in which the study of tumour growth is a central activity. We also analyse the algorithm performance on different modern parallel processing architectures, including multicore CPU and many-core GPU.
Baoquan Liu, Gordon Clapworthy, Feng Dong 0005, Eleni A. Kolokotroni, Georgios S. Stamatakos
IV1
2011 Non-Linear Beam Tracing on a GPU
abstract
Abstract Beam tracing combines the flexibility of ray tracing and the speed of polygon rasterization. However, beam tracing so far only handles linear transformations; thus, it is only applicable to linear effects such as planar mirror reflections but not to non‐linear effects such as curved mirror reflection, refraction, caustics and shadows. In this paper, we introduce non‐linear beam tracing to render these non‐linear effects. Non‐linear beam tracing is highly challenging because commodity graphics hardware supports only linear vertex transformation and triangle rasterization. We overcome this difficulty by designing a non‐linear graphics pipeline and implementing it on top of a commodity GPU. This allows beams to be non‐linear where rays within the same beam do not have to be parallel or intersect at a single point. Using these non‐linear beams, real‐time GPU applications can render secondary rays via polygon streaming similar to how they render primary rays. A major strength of this methodology is that it naturally supports fully dynamic scenes without the need to pre‐store a scene database. Utilizing our approach, non‐linear ray tracing effects can be rendered in real‐time on a commodity GPU under a unified framework.
Baoquan Liu, Li-Yi Wei, Chongyang Ma, Ying-Qing Xu, Baining Guo, Enhua Wu
Comput. Graph. Forum1
2010 Multi-layer Depth Peeling by Single-Pass Rasterisation for Faster Isosurface Raytracing on GPUs
abstract
Abstract Empty‐space skipping is an essential acceleration technique for volume rendering. Image‐order empty‐space skipping is not well suited to GPU implementation, since it must perform checks on, essentially, a per‐sample basis, as in kd‐tree traversal, which can lead to a great deal of divergent branching at runtime, which is very expensive in a modern GPU pipeline. In contrast, object‐order empty‐space skipping is extremely fast on a GPU and has negligible overheads compared with approaches without empty‐space skipping, since it employs the hardware unit for rasterisation. However, previous object‐order algorithms have been able to skip only exterior empty space and not the interior empty space that lies inside or between volume objects. In this paper, we address these issues by proposing a multi‐layer depth‐peeling approach that can obtain all of the depth layers of the tight‐fitting bounding geometry of the isosurface by a single rasterising pass. The maximum count of layers peeled by our approach can be up to thousands, while maintaining 32‐bit float‐point accuracy, which was not possible previously. By raytracing only the valid ray segments between each consecutive pair of depth layers, we can skip both the interior and exterior empty space efficiently. In comparisons with 3 state‐of‐the‐art GPU isosurface rendering algorithms, this technique achieved much faster rendering across a variety of data sets.
Baoquan Liu, Gordon Clapworthy, Feng Dong 0005
Comput. Graph. Forum1
2010 Wavefront raycasting using larger filter kernels for on-the-fly GPU gradient reconstruction
Baoquan Liu, Gordon Clapworthy, Feng Dong 0005
Vis. Comput.1
2009 Multi-layer depth peeling via fragment sort
abstract
We present an accelerated depth peeling algorithm for order-independent transparency rendering on graphics hardware. Unlike traditional depth peeling which only peels one layer of transparent pixels per rendering pass, our algorithm peels multiple layers simultaneously per rendering pass. Our acceleration is achieved via our fragment program which sorts and writes multiple fragment colors and depths via MRT. A notable feature of our algorithm is that it is robust against the unreliable parallel read-after-write behavior in current graphics hardware, guaranteeing correct transparency ordering. For ordinary scenes rendered under RGBA8 color precision, we achieve up to 8x speed-up over conventional depth peeling with current generation graphics hardware. Our algorithm is simple to implement on current GPU without any hardware modification. In addition, it does not require applications to perform any pre-sorting of transparent geometry.
Baoquan Liu, Li-Yi Wei, Ying-Qing Xu, Enhua Wu
CAD/Graphics1
2009 Accelerating Volume Raycasting using Proxy Spheres
abstract
Abstract In this paper, we propose an efficient solution that addresses the performance problems of current single‐pass GPU raycasting algorithms. Our paper provides more control over the rendering process by introducing tighter ray segments for raycasting, while at the same time avoiding the introduction of any new rendering artefacts. We achieve this by dynamically generating, on the GPU, a coarsely fitted proxy geometry, composed of spheres, for the active blocks. The spheres are then rasterised into two z‐buffers by a single rendering pass. The resulting two z‐buffers are used as the first‐hit and last‐hit points for the subsequent raycaster. With this approach, only the valid ray segments between the two z‐buffers need to be sampled during raycasting. This also provides more coherent parallelism on the GPU due to more consistent ray length and avoidance of the overheads and dynamic branching of performing checks on a per‐sample basis during the raycasting pass. Our technique is ideal for dynamic data exploration in which both the transfer function and view parameters need to be changed frequently at runtime. The rendering results of our algorithm are identical to the general cube‐based proxy geometry algorithm, but the performance can be up to 15.7 times faster. Furthermore, the approach can be adopted by any existing raycasting system in a straightforward way.
Baoquan Liu, Gordon Clapworthy, Feng Dong 0005
Comput. Graph. Forum1
2009 Fast Isosurface Rendering on a GPU by Cell Rasterization
abstract
Abstract This paper presents a fast, high‐quality, GPU‐based isosurface rendering pipeline for implicit surfaces defined by a regular volumetric grid. GPUs are designed primarily for use with polygonal primitives, rather than volume primitives, but here we directly treat each volume cell as a single rendering primitive by designing a vertex program and fragment program on a commodity GPU. Compared with previous raycasting methods, ours has a more effective memory footprint (cache locality) and better coherence between multiple parallel SIMD processors. Furthermore, we extend and speed up our approach by introducing a new view‐dependent sorting algorithm to take advantage of the early‐z‐culling feature of the GPU to gain significant performance speed‐up. As another advantage, this sorting algorithm makes multiple transparent isosurfaces rendering available almost for free. Finally, we demonstrate the effectiveness and quality of our techniques in several real‐time rendering scenarios and include analysis and comparisons with previous work.
Baoquan Liu, Gordon Clapworthy, Feng Dong 0005
Comput. Graph. Forum1
2008 Analytic Search Method for Interferometric SAR Image Registration
abstract
We propose an analytic search method for interferometric synthetic aperture radar (InSAR) image registration. An analytic cost function is established, and the subpixel offsets associated with the maximum of the cost function are continuously searched by direction-alternate optimization methods along vertical and horizontal directions. We establish a novel dual-quartic analytic cost function for extraction of subpixel offsets from a corresponding pixel pair established by pixel-level registration. A robust optimization method that is called biiteration algorithm for searching the maximum point of the dual-quartic cost function is developed to solve the subpixel offsets. Since the dual-quartic cost function is quartic if one of two parameters is fixed, one of two substeps including in each iterative step of the efficient biiterative method is to find the solution by the secant method. The performances of the proposed method are shown by using simulated and real data and compared with those of representative existing registration method.
Baoquan Liu, Da-Zheng Feng, Penglang Shui
IEEE Geosci. Remote. Sens. Lett.1
2006 Interactively Rendering Dynamic Caustics on GPU
Baoquan Liu, Enhua Wu, Xuehui Liu
Computer Graphics International1