EDBT 2026 Demo / reviewers in the wild / expert
Yu Mao 0001
dblp:59/5119-1
· DBLP profile ↗
12ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0001-9803-4927ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Parallel-SA: Point Cloud Processing Acceleration via Parallel Set AbstractionabstractPoint-based networks achieve high accuracy by preserving the intrinsic spatial structure of point clouds. The spatial information is effectively extracted by set abstraction, a critical module for feature learning in point-based networks. However, set abstraction introduces a computational bottleneck, and naive parallelization often degrades sampling quality, leading to accuracy loss. To address these challenges, we propose Parallel-SA, a framework that accelerates point-based networks by transforming set abstraction from sequential to parallel processing without sacrificing accuracy. Parallel-SA leverages a multi-scale sampling distribution approximation to preserve sampling quality under parallel execution. In addition, it employs distribution-aware balanced partitioning and adaptive load-balancing refinement to further improve efficiency. Experiments show that Parallel-SA achieves an average 2.38× speedup in set abstraction with minimal accuracy degradation. Dongdong Tang, Weilan Wang, Yu Mao 0001, Nan Guan, Tei-Wei Kuo, Chun Jason Xue |
DATE | 3 |
| 2026 | Hitcher: Efficient GPU-based Vector Search via Cluster-Centric Kernel and Hitch-Ride OrderingabstractSimilarity-based vector search, which retrieves the most similar vectors to a given query vector from a large vector dataset, underlies many applications such as search, recommendation, and Large Language Models (LLMs). Some systems run vector search on GPUs to enjoy GPU's high parallelism, but we observe that they are limited in query throughput and latency. In particular, their query-centric GPU kernel conducts computation independently for each query, failing to reuse data loaded to the GPU shared memory across queries and leading to a low GPU compute utilization. While their batch-based task reordering rearranges computation for queries in a batch to reduce CPU-GPU data transfer, but latency is prolonged since each query needs to wait for its slowest task. To tackle these problems, we propose Hitcher. Specifically, to reuse data across queries and improve GPU utilization, Hitcher implements a cluster-centric GPU kernel to batch computation on the same data for multiple queries. To reduce query latency, Hitcher adopts the hitch-ride ordering, which preserves the arrival order for query processing while batching computation across queries to improve efficiency. Hitcher can also offload computation tasks to the CPU to reduce CPU-GPU data transfer and utilize multiple GPUs. Experimental results show that Hitcher achieves up to 22× lower P99 query latency and 9× higher query throughput when compared with the state-of-the-art GPU-based vector query processing systems. Qihui Zhou, Changji Li, Guanxian Jiang, Chenhao Ma 0001, Xiao Yan 0002, Yu Mao 0001, Ming-Chang Yang, James Cheng |
KDD (1) | 6 |
| 2025 | WISE: A Framework for Gigapixel Whole-Slide-Image Lossless CompressionabstractWhole-Slide Images (WSIs) have revolutionized medical analysis by presenting high-resolution images of the whole tissue slide. Despite avoiding the physical storage of the slides, WSIs require considerable data volume, which makes the storage and maintenance of WSI records costly and unsustainable. To this end, this work presents the first investigation of lossless compression of WSI images. Interestingly, we find that most existing compression methods fail to compress the WSI images effectively. Furthermore, our analysis reveals that the failure of existing compressors is mainly due to information irregularity in WSI images. To resolve this issue, we develop a simple yet effective lossless compressor called WISE, specifically designed for WSI images. WISE employs a hierarchical encoding strategy to extract effective bits, reducing the entropy of the image and then adopting a dictionary-based method to handle the irregular frequency patterns. Through extensive experiments, we show that WISE can effectively compress the gigapixel WSI images to 36 times on average and up to 136 times. Yu Mao 0001, Nan Guan, Chun Jason Xue |
CVPR | 1 |
| 2025 | Easz: An Agile Transformer-based Image Compression Framework for Resource-constrained IoTsabstractNeural image compression, necessary in various machine-to-machine communication scenarios, suffers from its heavy encode-decode structures and inflexibility in switching between different compression levels. Consequently, it raises significant challenges in applying the neural image compression to edge devices that are developed for powerful servers with high computational and storage capacities. We take a step to solve the challenges by proposing a new transformer-based edge-computefree image coding framework called Easz. Easz shifts the computational overhead to the server, and hence avoids the heavy encoding and model switching overhead on the edge. Easz utilizes a patch-erase algorithm to selectively remove image contents using a conditional uniform-based sampler. The erased pixels are reconstructed on the receiver side through a transformer-based framework. To further reduce the computational overhead on the receiver, we then introduce a lightweight transformer-based reconstruction structure to reduce the reconstruction load on the receiver side. Extensive evaluations conducted on a realworld testbed demonstrate multiple advantages of Easz over existing compression approaches, in terms of adaptability to different compression levels, computational efficiency, and image reconstruction quality. Yu Mao 0001, Jingzong Li, Hong Xu 0001, Tei-Wei Kuo, Nan Guan, Chun Jason Xue |
DAC | 1 |
| 2025 | DAWN: Accelerating Point Cloud Object Detection via Object-Aware Partitioning and 3D Similarity-Based FilteringabstractAs a fundamental perception task, 3D point cloud detection has become essential for applications in autonomous driving and robotics. However, point cloud detection faces significant challenges of high computational cost due to complex point processing operations. To address this issue, we propose DAWN, an acceleration framework for point cloud object detection that identifies partial similarities between adjacent frames and reduces computational cost by filtering redundant points. DAWN uses object-aware partitioning that defines boundaries based on previous detection results for localized similarity analysis. Additionally, it applies axis-sorted point selection to refine partitioning for point clouds with non-uniform distribution. An efficient 3D similarity algorithm then filters redundant points to reduce computational load. DAWN enables flexible latencyaccuracy trade-offs by tuning point filtering ratios. Experimental results show that DAWN achieves a $1.59 \times$ average speedup and up to $1.70 \times$ on state-of-the-art detection networks by filtering more than $50 \%$ of points on average, with negligible impact on accuracy. Dongdong Tang, Yu Mao 0001, Weilan Wang, Nan Guan, Tei-Wei Kuo, Chun Jason Xue |
DAC | 2 |
| 2025 | Accelerating point cloud analytics on resource-constrained edge devices
Jingzong Li, Yik Hong Cai, Libin Liu 0001, Yu Mao 0001, Chun Jason Xue, Hong Xu 0001 |
Comput. Networks | 4 |
| 2024 | STEM: Streaming-Based FPGA Acceleration for Large-Scale Compactions in LSM KVabstractLog-Structured-Merge-tree (LSM-tree) has been extensively adopted because of its exceptional write efficiency and high space utilization. Compaction is invoked periodically in LSM-tree based key-value(LSM KV) systems to maintain good system performance. As the size of LSM-KV grows, large-scale compaction is now frequently seen. Compaction throughput significantly degrades with larger inputs, leading to frequent write stalls and decrement in overall write throughput. This paper proposes STEM, a stream-based compaction framework with FPGA to address this issue. A clean-cut algorithm is introduced to enable streaming-based compaction for large-scale data. With a multi-unit pipeline and dynamic pipeline schedule, STEM can handle large-scale compaction tasks efficiently. Based on the experiment result, the compaction throughput of STEM can achieve$27\times$on average and up to$35\times$improvement compared with the current RocksDB compaction,$2.09\times$to$2.27\times$improvement compared with the state-of-the-art FPGA accelerator. Dongdong Tang, Weilan Wang, Yu Mao 0001, Jinghuan Yu, Tei-Wei Kuo, Chun Jason Xue |
ICDE | 3 |
| 2023 | Faster and Stronger Lossless Compression with Optimized Autoregressive FrameworkabstractNeural AutoRegressive (AR) framework has been applied in general-purpose lossless compression recently to improve compression performance. However, this paper found that directly applying the original AR framework causes the duplicated processing problem and the in-batch distribution variation problem, which leads to deteriorated compression performance. The key to address the duplicated processing problem is to disentangle the processing of the history symbol set at the input side. Two new types of neural blocks are first proposed. An individual-block performs separate feature extraction on each history symbol while a mix-block models the correlation between extracted features and estimates the probability. A progressive AR-based compression framework (PAC) is then proposed, which only requires one history symbol from the host at a time rather than the whole history symbol set. In addition, we introduced a trainable matrix multiplication to model the ordered importance, replacing previous hardware-unfriendly Gumble-Softmax sampling. The in-batch distribution variation problem is caused by AR-based compression’s structured batch construction. Based on this observation, a batch-location-aware individual block is proposed to capture the heterogeneous in-batch distributions precisely, improving the performance without efficiency losses. Experimental results show the proposed framework can achieve an average of 130% speed improvement with an average of 3% compression ratio gain across data domains compared to the state-of-the-art. Yu Mao 0001, Jingzong Li, Yufei Cui, Chun Jason Xue |
DAC | 1 |
| 2023 | Moby: Empowering 2D Models for Efficient Point Cloud Analytics on the Edgeabstract3D object detection plays a pivotal role in many applications, most notably autonomous driving and robotics. These applications are commonly deployed on edge devices to promptly interact with the environment, and often require near real-time response. With limited computation power, it is challenging to execute 3D detection on the edge using highly complex neural networks. Common approaches such as offloading to the cloud induce significant latency overheads due to the large amount of point cloud data during transmission. To resolve the tension between wimpy edge devices and compute-intensive inference workloads, we explore the possibility of empowering fast 2D detection to extrapolate 3D bounding boxes. To this end, we present Moby, a novel system that demonstrates the feasibility and potential of our approach. We design a transformation pipeline for Moby that generates 3D bounding boxes efficiently and accurately based on 2D detection results without running 3D detectors. Further, we devise a frame offloading scheduler that decides when to launch the 3D detector judiciously in the cloud to avoid the errors from accumulating. Extensive evaluations on NVIDIA Jetson TX2 with real-world autonomous driving datasets demonstrate that Moby offers up to 91.9% latency improvement with modest accuracy loss over state of the art. Jingzong Li, Yik Hong Cai, Libin Liu 0001, Yu Mao 0001, Chun Jason Xue, Hong Xu 0001 |
ACM Multimedia | 4 |
| 2023 | Variational Nested DropoutabstractNested dropout is a variant of dropout operation that is able to order network parameters or features based on the pre-defined importance during training. It has been explored for: I. Constructing nested nets Cui et al. 2020, Cui et al. 2021: the nested nets are neural networks whose architectures can be adjusted instantly during testing time, e.g., based on computational constraints. The nested dropout implicitly ranks the network parameters, generating a set of sub-networks such that any smaller sub-network forms the basis of a larger one. II. Learning ordered representation Rippel et al. 2014: the nested dropout applied to the latent representation of a generative model (e.g., auto-encoder) ranks the features, enforcing explicit order of the dense representation over dimensions. However, the dropout rate is fixed as a hyper-parameter during the whole training process. For nested nets, when network parameters are removed, the performance decays in a human-specified trajectory rather than in a trajectory learned from data. For generative models, the importance of features is specified as a constant vector, restraining the flexibility of representation learning. To address the problem, we focus on the probabilistic counterpart of the nested dropout. We propose a variational nested dropout (VND) operation that draws samples of multi-dimensional ordered masks at a low cost, providing useful gradients to the parameters of nested dropout. Based on this approach, we design a Bayesian nested neural network that learns the order knowledge of the parameter distributions. We further exploit the VND under different generative models for learning ordered latent distributions. In experiments, we show that the proposed approach outperforms the nested network in terms of accuracy, calibration, and out-of-domain detection in classification tasks. It also outperforms the related generative models on data generation tasks. Yufei Cui, Yu Mao 0001, Ziquan Liu, Qiao Li 0001, Antoni B. Chan, Xue (Steve) Liu, Tei-Wei Kuo, Chun Jason Xue |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Accelerating General-purpose Lossless Compression via Simple and Scalable ParameterizationabstractThe storage of multi-media data can benefit from the advancements in general-purpose lossless compression. The explosive growth of multi-media data volume in data centers demands a higher compression ratio and better compressors' run-time speed. However, recent deep-learning-based compressors with a high compression ratio usually build complicated dependencies on history symbols, leading to a long compression time. This paper investigates the behavior of historical symbols and finds an approximate order of importance. Namely, recent symbols have a substantially larger influence on the probability estimation of the next unknown symbol. This observation guides the designing of an interpretable structure for data compression, rather than learning implicitly from data like Recurrent Neural Network (RNN) and attention. Based on this observation, we disentangle the compression model into order learning and feature learning, which were fused in a large module in previous works. A parameterized ordered mask unit is established to learn the ordered importance of history symbols. A fast Multi-Layer Perceptron (MLP) network is designed for efficient feature learning. The proposed compressor can improve both compression performance and computational efficiency compared with transformer-based or RNN-based compressors. To further enhance computational efficiency, we propose a branch-MLP block to replace the original MLP layer. This block reduces the parameters and the FLOPs of the original MLP to a half, without sacrificing compression performance. Experiments on multi-media data demonstrate that our model improves the compression ratio by 10% on average across data domains while accelerating compression speed by 100% compared with the state-of-the-art. The source code and appendix are released at https://github.com/mynotwo/compressor_via_simple_and_scalable_parameterization.git. Yu Mao 0001, Yufei Cui, Tei-Wei Kuo, Chun Jason Xue |
ACM Multimedia | 1 |
| 2022 | TRACE: A Fast Transformer-based General-Purpose Lossless CompressorabstractDeep-learning-based compressor has received interests recently due to much improved compression ratio. However, modern approaches suffer from long execution time. To ease this problem, this paper targets on cutting down the execution time of deep-learning-based compressors. Building history-dependencies sequentially (e.g., recurrent neural networks) is responsible for long inference latency. Instead, we introduce transformer into deep learning compressors to build history-dependencies in parallel. However, existing transformer is too heavy in computation and incompatible to compression tasks. Yu Mao 0001, Yufei Cui, Tei-Wei Kuo, Chun Jason Xue |
WWW | 1 |