VLDB 2026 Research / reviewers in the wild / expert
Baoyi An 0002
dblp:239/4414-2
· DBLP profile ↗
20ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0001-9471-5223ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Perceptual image compression with textual side information
Shiyu Qin, Bin Chen 0011, Yujun Huang, Baoyi An 0002, Tao Dai 0001, Shutao Xia |
Pattern Recognit. | 4 |
| 2026 | Universal image restoration via task-adaptive diffusion degradation oriented model
Junxi Wu, Sicheng Pan, Naiqi Li, Bin Chen 0011, Baoyi An 0002, Zhi Wang 0001, Yaowei Wang 0001, Shutao Xia |
Pattern Recognit. | 5 |
| 2026 | Resource-Aware Distributed Training Job Placement for GPU Cluster DefragmentationabstractDistributed training (DT) has emerged as a solution to address the growing computational resource demands of training large-scale machine learning models. To meet this need, cloud providers typically build GPU clusters to accommodate DT jobs. For DT job requests, cloud providers need to determine in which GPUs place workers (i.e., job placement). Existing approaches usually place workers on as few idle machines as possible to minimize communication time. However, this scheme will lead to aresource fragmentation problem, which degrades the resource utilization rate of the GPU cluster and increases training costs for cloud providers. In this paper, we propose$\textsf {Titan}$, a novel job placement scheme that mitigates the influence of resource fragmentation by enhancing the utilization of non-idle machines. To further optimize resource allocation, we introduce a dynamic defragmentation algorithm that migrates fragmented jobs to consolidate GPU resources, enabling efficient placement of large-scale training jobs.$\textsf {Titan}$formulates a multi-objective non-linear optimization problem and proves its NP-hardness. To solve this problem,$\textsf {Titan}$presents an effective submodular-based greedy algorithm with a tight approximation ratio ($1-\frac {1}{e}$). We evaluate$\textsf {Titan}$with a large-scale simulation employing real-world job traces and a small-scale testbed consisting of 8 servers with 32 logical GPUs. Experimental results show that$\textsf {Titan}$can achieve near-optimal training throughput while improving the efficiency of the cluster by 74.9% compared to the state-of-the-art solutions. Gongming Zhao, Yichen Dong, Hongli Xu 0001, Baoyi An 0002, Gangyi Luo |
IEEE Trans. Netw. | 6 |
| 2025 | LNeRV: Learnable Hierarchical Encoding Improve Neural Representation Video CodecabstractExisting Implicit Neural Representation (INR) video compression techniques have opened up new avenues in the field of video compression. NeRV maps the temporal coordinates to high-resolution images using neural networks, providing a more flexible and efficient encoding method for video data. However, NeRV implicitly stores all video information in the network, requiring post network compression techniques such as pruning. To integrate explicit compression and implicit representation into an end-to-end framework, this study proposes a novel neural representation-based video compression paradigm called Latent code based Neural representation video compression (LNeRV). Specifically, LNeRV consists of hierarchy feature grids, synthesis network and entropy coding network. With a single-stage training process, LNeRV achieves video compression and dynamically allocates bits according to video complexity, better fitting dynamic videos. We provide a comprehensive compression-to-decompression workflow for our approach. Extensive experimental results verify the effectiveness of our LNeRV. Jiahong Chen, Bin Chen 0011, Baoyi An 0002, Tao Dai 0001, Shutao Xia |
ICASSP | 4 |
| 2025 | Cassic: Towards Content-Adaptive State-Space Models for Learned Image Compression
Shiyu Qin, Jinpeng Wang 0002, Yimin Zhou 0011, Bin Chen 0011, Tianci Luo, Baoyi An 0002, Tao Dai 0001, Shutao Xia, Yaowei Wang 0001 |
ICCV | 6 |
| 2025 | DiffPC: Diffusion-based High Perceptual Fidelity Image Compression with Semantic RefinementabstractReconstructing high-quality images under low bitrates conditions presents a challenge, and previous methods have made this task feasible by leveraging the priors of diffusion models. However, the effective exploration of pre-trained latent diffusion models and semantic information integration in image compression tasks still needs further study. To address this issue, we introduce Diffusion-based High Perceptual Fidelity Image Compression with Semantic Refinement (DiffPC), a two-stage image compression framework based on stable diffusion. DiffPC efficiently encodes low-level image information, enabling the highly realistic reconstruction of the original image by leveraging high-level semantic features and the prior knowledge inherent in diffusion models. Specifically, DiffPC utilizes a multi-feature compressor to represent crucial low-level information with minimal bitrates and employs pre-embedding to acquire more robust hybrid semantics, thereby providing additional context for the decoding end. Furthermore, we have devised a control module tailored for image compression tasks, ensuring structural and textural consistency in reconstruction even at low bitrates and preventing decoding collapses induced by condition leakage. Extensive experiments demonstrate that our method achieves state-of-the-art perceptual fidelity and surpasses previous perceptual image compression methods by a significant margin in statistical fidelity. Yichong Xia, Yimin Zhou 0011, Jinpeng Wang 0002, Baoyi An 0002, Haoqian Wang, Yaowei Wang 0001, Bin Chen 0011 |
ICLR | 4 |
| 2025 | 3D-LMVIC: Learning-based Multi-View Image Compression with 3D Gaussian Geometric PriorsabstractExisting multi-view image compression methods often rely on 2D projection-based similarities between views to estimate disparities. While effective for small disparities, such as those in stereo images, these methods struggle with the more complex disparities encountered in wide-baseline multi-camera systems, commonly found in virtual reality and autonomous driving applications. To address this limitation, we propose 3D-LMVIC, a novel learning-based multi-view image compression framework that leverages 3D Gaussian Splatting to derive geometric priors for accurate disparity estimation. Furthermore, we introduce a depth map compression model to minimize geometric redundancy across views, along with a multi-view sequence ordering strategy based on a defined distance measure between views to enhance correlations between adjacent views. Experimental results demonstrate that 3D-LMVIC achieves superior performance compared to both traditional and learning-based methods. Additionally, it significantly improves disparity estimation accuracy over existing two-view approaches. Yujun Huang, Bin Chen 0011, Niu Lian, Xin Wang 0001, Baoyi An 0002, Tao Dai 0001, Shutao Xia |
ICML | 5 |
| 2025 | EDPC: Accelerating Lossless Compression via Lightweight Probability Models and Decoupled Parallel DataflowabstractThe explosive growth of multi-source multimedia data has significantly increased the demands for transmission and storage, placing substantial pressure on bandwidth and storage infrastructures. While Autoregressive Compression Models (ACMs) have markedly improved compression efficiency through probabilistic prediction, current approaches remain constrained by two critical limitations: suboptimal compression ratios due to insufficient fine-grained feature extraction during probability modeling, and real-time processing bottlenecks caused by high resource consumption and low compression speeds. To address these challenges, we propose Efficient Dual-path Parallel Compression (EDPC), a hierarchically optimized compression framework that synergistically enhances modeling capability and execution efficiency via coordinated dual-path operations. At the modeling level, we introduce the Information Flow Refinement (IFR) metric grounded in mutual information theory, and design a Multi-path Byte Refinement Block (MBRB) to strengthen cross-byte dependency modeling via heterogeneous feature propagation. At the system level, we develop a Latent Transformation Engine (LTE) for compact high-dimensional feature representation and a Decoupled Pipeline Compression Architecture (DPCA) to eliminate encoding-decoding latency through pipelined parallelization. Experimental results demonstrate that EDPC achieves comprehensive improvements over state-of-the-art methods, including a 2.7× faster compression speed, and a 3.2% higher compression ratio. These advancements establish EDPC as an efficient solution for real-time processing of large-scale multimedia data in bandwidth-constrained scenarios. Our code is available at https://github.com/Magie0/EDPC. Zeyi Lu, Yujun Huang, Minxiao Chen, Bin Chen 0011, Baoyi An 0002, Shutao Xia |
ACM Multimedia | 6 |
| 2025 | MB-RACS: Measurement-Bounds-Based Rate-Adaptive Image Compressed Sensing NetworkabstractConventional compressed sensing (CS) algorithms typically apply a uniform sampling rate to different image blocks. A more strategic approach could be to allocate the number of measurements adaptively, based on each image block's complexity. In this paper, we propose a Measurement-Bounds-based Rate-Adaptive Image Compressed Sensing Network (MB-RACS) framework, which aims to adaptively determine the sampling rate for each image block in accordance with traditional measurement bounds theory. Moreover, since in real-world scenarios statistical information about the original image cannot be directly obtained, we suggest a multi-stage rate-adaptive sampling strategy. This strategy sequentially adjusts the sampling ratio allocation based on the information gathered from previous samplings. We formulate the multi-stage rate-adaptive sampling as a convex optimization problem and address it using a combination of Newton's method and binary search techniques. Our experiments demonstrate that the proposed MB-RACS method surpasses current leading methods, with experimental evidence also underscoring the effectiveness of each module within our proposed framework. Yujun Huang, Bin Chen 0011, Naiqi Li, Baoyi An 0002, Shutao Xia, Yaowei Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | An Efficient Implicit Neural Representation Image Codec Based on Mixed Autoregressive Model for Low-Complexity DecodingabstractDisplaying high-quality images on edge devices, such as augmented reality devices, is essential for enhancing the user experience. However, these devices often face power consumption and computing resource limitations, making it challenging to apply many deep learning-based image compression algorithms in this field. Implicit Neural Representation (INR) for image compression is an emerging technology that offers two key benefits compared to cutting-edge autoencoder models: low computational complexity and parameter-free decoding. It also outperforms many traditional and early neural compression methods in terms of quality. In this study, we introduce a new Mixed AutoRegressive Model (MARM) to significantly reduce the decoding time for the current INR codec, along with a new synthesis network to enhance reconstruction quality. MARM includes our proposed AutoRegressive Upsampler (ARU) blocks, which are highly computationally efficient, and ARM from previous work to balance decoding time and reconstruction quality. We also propose enhancing ARU's performance using a checkerboard two-stage decoding strategy. Moreover, the ratio of different modules can be adjusted to maintain a balance between quality and speed. Comprehensive experiments demonstrate that our method significantly improves computational efficiency while preserving image quality. With different parameter settings, our method can achieve over a magnitude acceleration in decoding time without industrial level optimization or achieve state-of-the-art reconstruction quality compared with other INR codecs. To the best of our knowledge, our method is the first INR-based codec comparable with Ballé et al. [1] in both decoding speed and quality while maintaining low complexity. Jiahong Chen, Bin Chen 0011, Zimo Liu, Baoyi An 0002, Shutao Xia, Zhi Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Progressive Learning with Visual Prompt Tuning for Variable-Rate Image CompressionabstractIn this paper, we propose a progressive learning paradigm for transformer-based variable-rate image compression. Our approach covers a wide range of compression rates with the assistance of the Layer-adaptive Prompt Module (LPM). Inspired by visual prompt tuning, we use LPM to extract prompts for input images and hidden features at the encoder side and decoder side, respectively, which are fed as additional information into the swin transformer layer of a pre-trained transformer-based image compression model to affect the allocation of attention region and the bits, which in turn changes the target compression ratio of the model. To ensure the network is more lightweight, we involves the integration of prompt networks with less convolutional layers. Exhaustive experiments show that compared to methods based on multiple models, which are optimized separately for different target rates, the proposed method arrives at the same performance with 80% savings in parameter storage and 90% savings in datasets. Meanwhile, our model outperforms all current variable bitrate image methods in terms of rate-distortion performance and approaches the state-of-the-art fixed bitrate image compression methods trained from scratch. Shiyu Qin, Yimin Zhou 0011, Jin-Peng Wang, Bin Chen 0011, Baoyi An 0002, Tao Dai 0001, Shutao Xia |
ICIP | 5 |
| 2024 | CMCL: Cross-Modal Compressive Learning for Resource-Constrained Intelligent IoT SystemsabstractCompressive Learning (CL) has proven to be highly successful in executing joint signal sampling and inference for intricate vision tasks through resource-limited Internet of Things (IoT) devices. Recent studies have turned their attention towards utilizing the deep neural networks (DNNs) methodology, also known as DeepCL, to enhance performance in unimodal vision tasks. This approach incorporates learnable compressed sensing in a comprehensive, end-to-end manner. Current DeepCL techniques typically employ initial signal reconstruction as the input for subsequent DNNs for inference. However, this practice presents potential risks such as privacy breaches and reduced performance due to information processing inequality. To address these issues, this paper introduces the first cross-modal compressive learning (CMCL) approach that enables image captioning directly on compressed measurements. When compared to previous DeepCL strategies, the proposed CMCL offers significant improvements in computational efficiency and privacy protection. Extensive experiments demonstrate that CMCL performance is nearly on par with leading image captioning methods, showcasing a metric value that is merely 2.75% lower than the uncompressed method when the data is compressed eightfold. Bin Chen 0011, Yujun Huang, Baoyi An 0002, Yaowei Wang 0001, Xuan Wang 0002 |
IEEE Internet Things J. | 4 |
| 2023 | Secure Crowdsensed Data Trading Based on BlockchainabstractCrowdsensed Data Trading (CDT) is a novel data trading paradigm, where each data consumer can publicize its data demand as some crowdsensing tasks, and some mobile users (i.e., data sellers) can compete for these tasks, collect the corresponding data, and sell the results to the consumers. Existing CDT systems generally depend on a data trading broker, which will inevitably cause consumers concerns on the trustworthiness of the systems and truthfulness of the data. To address this problem, we propose a Blockchain-based Crowdsensed Data Trading (BCDT) system, mainly containing a smart contract, called BCDToken. First, we replace the broker with blockchain to guarantee the trustworthiness of data trading. Meanwhile, BCDToken adopts Blockchain-based Reverse Auction (BRA) to assign tasks to data sellers. BRA holds truthfulness and individual rationality, which can ensure the sellers to report costs honestly and prevent sellers to manipulate the auction. Moreover, we implement a Secure Truth Discovery and reliability Rating (STDR) mechanism in BCDToken based on homomorphic cryptography, which can incentivize sellers to upload the truthful data and consumers to rate truthfully the reliabilities of sellers without revealing any privacy of data. Additionally, we also deploy BCDToken to the test network to demonstrate its practicability. Baoyi An 0002, Mingjun Xiao, An Liu 0002, Xiangliang Zhang 0001, Qing Li 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | CMAB-Based Reverse Auction for Unknown Worker Recruitment in Mobile CrowdsensingabstractMobile CrowdSensing (MCS), through which a requester can coordinate a crowd of workers to accomplish some data collection tasks, has been recognized as a promising paradigm for large-scale data acquisition in recent years. Many researches focus on the worker recruitment problem in MCS, but most of them either have the assumption that workers’ qualities are known ahead of time or cannot ensure that workers report costs honestly. In this paper, we propose an incentive mechanism based on Combinatorial Multi-Armed Bandit and reverse Auction, called CMABA, to solve the multiple unknown workers recruitment problem in MCS. Our objective is to determine a recruiting strategy to maximize the total sensing quality under a limited budget, while ensuring truthfulness and individual rationality of sensing workers. We theoretically prove that our CMABA mechanism achieves truthfulness and individual rationality, and then analyze the regret of the mechanism. Based on CMABA, we ulteriorly propose an adaptive incentive mechanism, called ACMABA, to recruit workers via the alternative worker recruitment and quality update, which can achieve a higher total sensing quality and lower regret. Additionally, we also demonstrate significant performances of the CMABA and ACMABA mechanisms through extensive simulations on real-world data traces. Mingjun Xiao, Baoyi An 0002, Jing Wang 0028, Guoju Gao, Sheng Zhang 0001, Jie Wu 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2021 | Crowdsensing Data Trading based on Combinatorial Multi-Armed Bandit and Stackelberg GameabstractCrowdsensing Data Trading (CDT), through which a platform can aggregate some data collected by a group of mobile users with sensing devices (a.k.a., data sellers) and sell the corresponding statistics to data consumers, has been recognized as a promising paradigm for large-scale data trading in recent years. It is critical to select sellers with high sensing qualities and maximize all trading participants' profits simultaneously. However, most existing CDT systems either assume that sellers' sensing qualities are known in advance or cannot realize concurrent profit maximization. In this paper, we propose a data trading mechanism based on Combinatorial Multi-Armed Bandit and three-stage Hierarchical Stackelberg game, called CMAB-HS, to tackle the problem of quality unknown seller selection and incentive strategy design. Our objective is to select a group of sellers to maximize the total sensing quality within time budget, and determine the optimal incentive strategy for each participant to maximize individual profit simultaneously. We theoretically prove that CMAB-HS achieves Stackelberg Equilibrium and a tight bound on regret. Additionally, we demonstrate its significant performances through extensive simulations on real data traces. Baoyi An 0002, Mingjun Xiao, An Liu 0002, Xike Xie, Xiaofang Zhou 0001 |
ICDE | 1 |
| 2021 | Blockchain-Based Double Auction for Edge Cloud Resource Trading with Differential Privacyabstractpresent, edge cloud is becoming more and more important as it provides computing resources required by mobile terminals. The online edge cloud resource trading mechanism has become one of the most important ways for edge cloud resource sharing, attracting much attention. In this paper, we have proposed a Blockchain-based Privacy-preserving Combinatorial Double Auction (BPCDA) mechanism, which combines trustworthiness, truthfulness and privacy protection for the first time. Firstly, BPCDA avoids dependence on third-party brokers by adopting blockchain. Secondly, we draw on the economic theories, and then adopt double auction to ensure the truthfulness of bids while participants in the auction are composed of multiple sellers and multiple buyers. Lastly, we make fell use of the differential privacy mechanism on the blockchain to avoid privacy leakage. Through simulated experiments, we verify the superior performance and practicality of our proposed edge cloud resource trading mechanism. Yin Xu 0004, Baoyi An 0002, Mingjun Xiao |
MASS | 3 |
| 2021 | FedDCS: Federated Learning Framework based on Dynamic Client SelectionabstractFederated Learning, through which a server can coordinate a crowd of clients to accomplish a machine learning task, has been recognized as a promising paradigm for privacy preserving decentralized learning in recent years. Most federated learning researches are on the basis of IID data, and more and more researches focus on the Non-IID data problem. However, they have not considered the process of data acquisition. In fact, in real applications, the training data are typically collected by clients in real scenarios. In this paper, we propose a Federated Learning framework based on Dynamic Client Selection, called FedDCS, to deal with the real scene Non-IID data machine learning problem in federated learning. The objective of FedDCS is to utilize a parameter estimation algorithm to select the optimal clients to join the collaboration and finally acquire a better global machine learning model. Our extensive experiments on several real-world data sets demonstrate the superior performance of FedDCS. Shutong Zou, Mingjun Xiao, Yin Xu 0004, Baoyi An 0002 |
MASS | 4 |
| 2020 | CPchain: A Copyright-Preserving Crowdsourcing Data Trading Framework Based on BlockchainabstractCrowdsourcing data trading is a novel paradigm in which the crowdsourcing technology is adopted to collect big data for trading. At present, existing crowdsourcing data trading systems usually depend on a trusted broker and haven't considered the truthfulness and the quality of data (QoD) simultaneously. Besides, copyright protection is the another issue that has not been properly addressed. To tackle these problems, we propose a Copyright-Preserving crowdsourcing data trading framework based on Blockchain, named CPchain, which mainly includes a smart contract. We design an auction algorithm based on semantic similarity to guarantee the truthfulness and individual rationality while ensuring QoD. Moreover, we combine digital fingerprint technology with blockchain to protect data copyright without a third-party certification authority. Furthermore, we develop a simple prototype of our proposed trading framework on the Ethereum test network. We have carried out a lot of experiments to demonstrate the significant performances of our framework. Dingjie Sheng, Mingjun Xiao, An Liu 0002, Baoyi An 0002, Sheng Zhang 0001 |
ICCCN | 5 |
| 2019 | Truthful Crowdsensed Data Trading Based on Reverse Auction and Blockchain
Baoyi An 0002, Mingjun Xiao, An Liu 0002, Guoju Gao, Hui Zhao 0003 |
DASFAA (1) | 1 |
| 2019 | Reverse-Auction-Based Competitive Order Assignment for Mobile Taxi-Hailing Systems
Hui Zhao 0003, Mingjun Xiao, Jie Wu 0001, An Liu 0002, Baoyi An 0002 |
DASFAA (2) | 5 |