VLDB 2026 Research / reviewers in the wild / expert
Yujun Huang
dblp:252/8134
· DBLP profile ↗
22ranked-venue papers
5as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoE-LC: General-Purpose Lossless Compression for Multi-modal Data via Entropy-Aware Multi-ExpertsabstractThe web-scale surge of multimodal content, including short-video feeds and autonomous sensing streams, has made web-native lossless compression a prerequisite for delivery and storage across browsers and edge–cloud pipelines. However, existing methods often fail to adapt to shifting distributions across different batches and struggle to balance computational resources in the face of large conditional entropy disparities among diverse modalities. To address these limitations, we propose MoE-LC, a new mixture-of-experts framework for multi-modal lossless compression that dynamically accommodates heterogeneous data distributions and varying complexity levels. First, the Batch-Adaptive Experts (BAE) module introduces batch-specific parameters with a residual gating mechanism, ensuring stable modeling under non-stationary distributions. Second, the Entropy-Aware Multi-Expert Selection (MES) strategy adaptively allocates the number of experts according to the data's estimated compression difficulty (entropy), thereby improving resource utilization and computational efficiency. Finally, the Precision-Aware Expert Routing (PER) component applies high-precision computation solely to the most critical experts, significantly reducing overhead without sacrificing compression accuracy. Experimental results across multiple real-world datasets demonstrate that MoE-LC achieves 5.33%--70.89% improvements in compression ratio and 37.25%--1532.41% gains in throughput compared to advanced baselines, offering a scalable solution for real-time, large-scale multi-modal data compression. Our code is available at https://github.com/Magie0/MoE_LC. Zeyi Lu, Yujun Huang, Minxiao Chen, Bin Chen 0011, Shutao Xia |
WWW | 3 |
| 2026 | Perceptual image compression with textual side information
Shiyu Qin, Bin Chen 0011, Yujun Huang, Baoyi An 0002, Tao Dai 0001, Shutao Xia |
Pattern Recognit. | 3 |
| 2025 | 3D-LMVIC: Learning-based Multi-View Image Compression with 3D Gaussian Geometric PriorsabstractExisting multi-view image compression methods often rely on 2D projection-based similarities between views to estimate disparities. While effective for small disparities, such as those in stereo images, these methods struggle with the more complex disparities encountered in wide-baseline multi-camera systems, commonly found in virtual reality and autonomous driving applications. To address this limitation, we propose 3D-LMVIC, a novel learning-based multi-view image compression framework that leverages 3D Gaussian Splatting to derive geometric priors for accurate disparity estimation. Furthermore, we introduce a depth map compression model to minimize geometric redundancy across views, along with a multi-view sequence ordering strategy based on a defined distance measure between views to enhance correlations between adjacent views. Experimental results demonstrate that 3D-LMVIC achieves superior performance compared to both traditional and learning-based methods. Additionally, it significantly improves disparity estimation accuracy over existing two-view approaches. Yujun Huang, Bin Chen 0011, Niu Lian, Xin Wang 0001, Baoyi An 0002, Tao Dai 0001, Shutao Xia |
ICML | 1 |
| 2025 | UITrans: Seamless UI Translation from Android to HarmonyOSabstractSeamless user interface (i.e., UI) translation has emerged as a pivotal technique for modern mobile developers, addressing the challenge of developing separate UI applications for Android and HarmonyOS platforms due to fundamental differences in layout structures and development paradigms.In this paper, we present UITrans, the first automated UI translation tool designed for Android to HarmonyOS.UITrans leverages an LLM-driven multi-agent reflective collaboration framework to convert Android XML layouts into HarmonyOS ArkUI layouts.It not only maps component-level and page-level elements to ArkUI equivalents but also handles project-level challenges, including complex layouts and interaction logic.Our evaluation of six Android applications demonstrates that our UITrans achieves translation success rates of over 90.1%, 89.3%, and 89.2% at the component, page, and project levels, respectively.UITrans is available at https://github.com/OpenSELab/UITransand the demo video can be viewed at https://www.youtube.com/watch?v=iqKOSm CnJG0. Lina Gong, Yujun Huang, Mingqiang Wei |
Internetware | 4 |
| 2025 | TOPP-DWR: Time-Optimal Path Parameterization of Differential-Driven Wheeled Robots Considering Piecewise-Constant Angular Velocity ConstraintsabstractDifferential-driven wheeled robots (DWR) represent the quintessential type of mobile robots and find extensive applications across the robotic field. Most high-performance control approaches for DWR explicitly utilize the linear and angular velocities of the trajectory as control references. However, existing research on time-optimal path parameterization (TOPP) for mobile robots usually neglects the angular velocity and joint velocity constraints, which can result in degraded control performance in practical applications. In this article, a systematic and practical TOPP algorithm named TOPP-DWR is proposed for DWR and other mobile robots. First, the non-uniform B-spline is adopted to represent the initial trajectory in the task space. Second, the piecewise-constant angular velocity, as well as joint velocity, linear velocity, and linear acceleration constraints, are incorporated into the TOPP problem. During the construction of the optimization problem, the aforementioned constraints are uniformly represented as linear velocity constraints. To boost the numerical computational efficiency, we introduce a slack variable to reformulate the problem into second-order-cone programming (SOCP). Subsequently, comparative experiments are conducted to validate the superiority of the proposed method. Quantitative performance indexes show that TOPP-DWR achieves TOPP while adhering to all constraints. Finally, field autonomous navigation experiments are carried out to validate the practicability of TOPP-DWR in real-world applications. Yujun Huang |
IROS | 2 |
| 2025 | EDPC: Accelerating Lossless Compression via Lightweight Probability Models and Decoupled Parallel DataflowabstractThe explosive growth of multi-source multimedia data has significantly increased the demands for transmission and storage, placing substantial pressure on bandwidth and storage infrastructures. While Autoregressive Compression Models (ACMs) have markedly improved compression efficiency through probabilistic prediction, current approaches remain constrained by two critical limitations: suboptimal compression ratios due to insufficient fine-grained feature extraction during probability modeling, and real-time processing bottlenecks caused by high resource consumption and low compression speeds. To address these challenges, we propose Efficient Dual-path Parallel Compression (EDPC), a hierarchically optimized compression framework that synergistically enhances modeling capability and execution efficiency via coordinated dual-path operations. At the modeling level, we introduce the Information Flow Refinement (IFR) metric grounded in mutual information theory, and design a Multi-path Byte Refinement Block (MBRB) to strengthen cross-byte dependency modeling via heterogeneous feature propagation. At the system level, we develop a Latent Transformation Engine (LTE) for compact high-dimensional feature representation and a Decoupled Pipeline Compression Architecture (DPCA) to eliminate encoding-decoding latency through pipelined parallelization. Experimental results demonstrate that EDPC achieves comprehensive improvements over state-of-the-art methods, including a 2.7× faster compression speed, and a 3.2% higher compression ratio. These advancements establish EDPC as an efficient solution for real-time processing of large-scale multimedia data in bandwidth-constrained scenarios. Our code is available at https://github.com/Magie0/EDPC. Zeyi Lu, Yujun Huang, Minxiao Chen, Bin Chen 0011, Baoyi An 0002, Shutao Xia |
ACM Multimedia | 3 |
| 2025 | Generative Adversarial CLIPs for Unsupervised Backlit Image Enhancement
Suoyang Sun, Yujun Huang, Bin Chen 0011 |
WASA (1) | 3 |
| 2025 | MB-RACS: Measurement-Bounds-Based Rate-Adaptive Image Compressed Sensing NetworkabstractConventional compressed sensing (CS) algorithms typically apply a uniform sampling rate to different image blocks. A more strategic approach could be to allocate the number of measurements adaptively, based on each image block's complexity. In this paper, we propose a Measurement-Bounds-based Rate-Adaptive Image Compressed Sensing Network (MB-RACS) framework, which aims to adaptively determine the sampling rate for each image block in accordance with traditional measurement bounds theory. Moreover, since in real-world scenarios statistical information about the original image cannot be directly obtained, we suggest a multi-stage rate-adaptive sampling strategy. This strategy sequentially adjusts the sampling ratio allocation based on the information gathered from previous samplings. We formulate the multi-stage rate-adaptive sampling as a convex optimization problem and address it using a combination of Newton's method and binary search techniques. Our experiments demonstrate that the proposed MB-RACS method surpasses current leading methods, with experimental evidence also underscoring the effectiveness of each module within our proposed framework. Yujun Huang, Bin Chen 0011, Naiqi Li, Baoyi An 0002, Shutao Xia, Yaowei Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | FCA-Net: Accelerating stereo image compression through cascade alignment of side information
Yichong Xia, Yujun Huang, Bin Chen 0011, Genping Wang, Haoqian Wang, Yaowei Wang 0001 |
Pattern Recognit. | 2 |
| 2025 | Image Compression for Resource-Constrained AIoT System With Compressed SensingabstractIn today’s big data era, a key requirement is to implement intelligent semantic analysis (such as image recognition) on data gathered from an extensive array of smart devices in Artificial Intelligence IoT (AIoT) scenarios, all of which is processed at central cloud service providers. Recent advancements in deep-learning-based image compression have fostered semantic compression between machines. However, the deployment of an overparameterized encoder on Internet of Things (IoT) devices remains a challenge due to their restricted computing and storage capabilities. To tackle this issue, we propose a novel approach named compressed sensing (CS)-based asymmetric semantic image compression (CS-ASIC), explicitly designed for resource-constrained AIoT systems. This asymmetric semantic compression scheme intends to surpass the limitations of IoT devices, thereby facilitating efficient semantic compression for machine vision tasks. CS-ASIC notably includes a lightweight front encoder founded on deep image CS techniques, which utilizes rich image priors to learn measurement matrices for sampling. In tandem, a deep iterative decoder is designed cooperatively with the linear encoder offloaded at the server to enhance image reconstruction and semantic analysis across various semantic analysis tasks. Furthermore, we introduce a groundbreaking lossy CS semantic rate-distortion theoretical framework that justifies a compromise in rate for extended semantic distortion. Extensive experimental results underscore the superiority of the proposed CS-ASIC concerning the signal-semantic rate-distortion tradeoff, and its lower encoding complexity over existing codecs in an AIoT simulation environment. Bin Chen 0011, Yujun Huang, Han Qiu 0001, Shutao Xia, Wei Fei, Xuan Wang 0002, Meikang Qiu |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | Expand and Merge: Continual Learning with the Guidance of Fixed Text Embedding SpaceabstractDeep neural networks lack the ability to sequentially learn from new data and adapt to new scenarios. In particular, after learning new data, neural networks will have a significant performance degradation on old knowledge. This phenomenon is known as catastrophic forgetting. To mitigate this issue, we propose expanding and merging additional parameters, encapsulated in a specially designed adapter layer, in the frozen pretrained vision encoder. Along with the newly added parameters, adapter scaling weights in each layer are also introduced to adaptively control the fusion of new and old knowledge. Additionally, the fixed embedding space of a pretrained text encoder is used to guide the continual learning of the vision encoder. Extensive experiments on three datasets demonstrate that the proposed method outperform current state-of-the-art methods. The source code is available at https://github.com/GiantJun/Expand_and_Merge. Yujun Huang, Wentao Zhang 0005 |
IJCNN | 1 |
| 2024 | CMCL: Cross-Modal Compressive Learning for Resource-Constrained Intelligent IoT SystemsabstractCompressive Learning (CL) has proven to be highly successful in executing joint signal sampling and inference for intricate vision tasks through resource-limited Internet of Things (IoT) devices. Recent studies have turned their attention towards utilizing the deep neural networks (DNNs) methodology, also known as DeepCL, to enhance performance in unimodal vision tasks. This approach incorporates learnable compressed sensing in a comprehensive, end-to-end manner. Current DeepCL techniques typically employ initial signal reconstruction as the input for subsequent DNNs for inference. However, this practice presents potential risks such as privacy breaches and reduced performance due to information processing inequality. To address these issues, this paper introduces the first cross-modal compressive learning (CMCL) approach that enables image captioning directly on compressed measurements. When compared to previous DeepCL strategies, the proposed CMCL offers significant improvements in computational efficiency and privacy protection. Extensive experiments demonstrate that CMCL performance is nearly on par with leading image captioning methods, showcasing a metric value that is merely 2.75% lower than the uncompressed method when the data is compressed eightfold. Bin Chen 0011, Yujun Huang, Baoyi An 0002, Yaowei Wang 0001, Xuan Wang 0002 |
IEEE Internet Things J. | 3 |
| 2024 | Multi-scale architectures matter: Examining the adversarial robustness of flow-based lossless compression
Yichong Xia, Bin Chen 0011, Tianshuo Ge, Yujun Huang, Haoqian Wang, Yaowei Wang 0001 |
Pattern Recognit. | 5 |
| 2024 | Continual Learning of Image Classes With Language Guidance From a Vision-Language ModelabstractCurrent deep learning models often catastrophically forget the knowledge of old classes when continually learning new ones. State-of-the-art approaches to continual learning of image classes often require retaining a small subset of old data to partly alleviate the catastrophic forgetting issue, and their performance would be degraded sharply when no old data can be stored due to privacy or safety concerns. In this study, inspired by human learning of visual knowledge with the effective help of language, we propose a novel continual learning framework based on a pre-trained vision-language model (VLM) without retaining any old data. Rich prior knowledge of each new image class is effectively encoded by the frozen text encoder of the VLM, which is then used to guide the learning of new image classes. The output space of the frozen text encoder is unchanged over the whole process of continual learning, through which image representations of different classes become comparable during model inference even when the image classes are learned at different times. Extensive empirical evaluations on multiple image classification datasets under various settings confirm the superior performance of our method over existing ones. The source code is available athttps://github.com/Fatflower/CIL_LG_VLM/. Wentao Zhang 0005, Yujun Huang, Weizhuo Zhang, Tong Zhang 0017, Qicheng Lao, Yue Yu 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Learned Distributed Image Compression with Multi-Scale Patch Matching in Feature DomainabstractBeyond achieving higher compression efficiency over classical image compression codecs, deep image compression is expected to be improved with additional side information, e.g., another image from a different perspective of the same scene. To better utilize the side information under the distributed compression scenario, the existing method only implements patch matching at the image domain to solve the parallax problem caused by the difference in viewing points. However, the patch matching at the image domain is not robust to the variance of scale, shape, and illumination caused by the different viewing angles, and can not make full use of the rich texture information of the side information image. To resolve this issue, we propose Multi-Scale Feature Domain Patch Matching (MSFDPM) to fully utilizes side information at the decoder of the distributed image compression model. Specifically, MSFDPM consists of a side information feature extractor, a multi-scale feature domain patch matching module, and a multi-scale feature fusion network. Furthermore, we reuse inter-patch correlation from the shallow layer to accelerate the patch matching of the deep layer. Finally, we find that our patch matching in a multi-scale feature domain further improves compression rate by about 20% compared with the patch matching method at image domain. Yujun Huang, Bin Chen 0011, Shiyu Qin, Jiawei Li 0006, Yaowei Wang 0001, Tao Dai 0001, Shutao Xia |
AAAI | 1 |
| 2023 | Adaptive ProductAE: SNR-Aware Adaptive Decoding of Neural Product CodesabstractAs a pioneering and deep-learning driven neural channel code, Product Autoencoder (ProductAE) shows significant superiority over Turbo Autoencoder (TurboAE) and other classical codes. However, received noisy codewords with different noise levels have different decoding difficulties for the neural decoder of ProductAE. The existing neural decoder processes all noisy codewords equally without discrimination, and neglects the prior knowledge of signal-to-noise ratio (SNR), thus leading to high decoding complexity. In this paper, we propose a new approach to speed up the decoding process of ProductAE. Intuitively, noisy codewords with different SNR can be recovered by decoders of different complexity, and our proposed novel pipeline, Adaptive ProductAE, an SNR-aware adaptive decoding strategy, adopts deep decoders to different SNR. Specifically, adaptive ProductAE combines an encoding module, an additional classification module, and two independent decoding branches in a unified framework. After receiving noisy codewords, it first judges the decoding difficulty of each signal and assigns them to different branches to get the decoded messages. Furthermore, we introduce a novel training strategy with a gap loss to maintain the classification and decoding performance. It can thus switch to a simpler branch of decoding networks automatically when it comes to signals with a higher SNR, and the overall computational cost can be reduced. Experiments show that our adaptive ProductAE saves up to 30.4% FLOPs for a moderate-length code of parameters (225, 100) and 25.5% FLOPs for a code of parameters (441, 196) in a higher SNR range. Qinshan Zhang, Bin Chen 0011, Yujun Huang, Shutao Xia |
ICC | 3 |
| 2023 | LKBQ: Pushing the Limit of Post-Training Quantization to Extreme 1 bitabstractRecent advances have shown the potential for post-training quantization (PTQ) to reduce excessive hardware resources and quantize deep models to low bits in a short time, compared with Quantization-Aware Training (QAT). However, existing PTQ approaches lose a lot of accuracies when quantizing the model to extremely low bits, e.g., 1 bit. In this work, we propose layer-by-layer self-knowledge distillation binary post-training quantization (LKBQ), the first method capable of quantizing the weights of neural networks to 1 bit in PTQ domain. We show that careful use of layer-by-layer self-distillation within the LKBQ can provide a significant performance boost. Furthermore, our evaluation results show that the initialization of quantized network weights can have a huge impact on the results. Then we propose three methods for weight initialization. Finally, in light of the characteristics of the binarized network, we propose a method named gradient scaling to further improve efficiency. Our experiments show that LKBQ pushes the limit of PTQ to extreme 1-bit for the first time. Bin Chen 0011, Qian-Wei Wang, Yujun Huang, Shutao Xia |
ICIP | 4 |
| 2023 | Adapter Learning in Pretrained Feature Extractor for Continual Learning of Diseases
Wentao Zhang 0005, Yujun Huang, Tong Zhang 0017, Qingsong Zou, Wei-Shi Zheng 0001 |
MICCAI (2) | 2 |
| 2023 | Task-Incremental Medical Image Classification with Task-Specific Batch Normalization
Xuchen Xie, Weizhuo Zhang, Yujun Huang, Wei-Shi Zheng 0001 |
PRCV (13) | 5 |
| 2022 | A Structure-Aware Argument Encoder for Literature Discourse AnalysisabstractExisting research for argument representation learning mainly treats tokens in the sentence equally and ignores the implied structure information of argumentative context. In this paper, we propose to separate tokens into two groups, namely framing tokens and topic ones, to capture structural information of arguments. In addition, we consider high-level structure by incorporating paragraph-level position information. A novel structure-aware argument encoder is proposed for literature discourse analysis. Experimental results on both a self-constructed corpus and a public corpus show the effectiveness of our model. Resources are available at https://github.com/lemuria-wchen/SAE. Yinzi Li, Wei Chen 0088, Zhongyu Wei, Yujun Huang, Chujun Wang, Siyuan Wang 0025, Qi Zhang 0001, Xuanjing Huang 0001, Libo Wu |
COLING | 4 |
| 2022 | Compressive sensing based asymmetric semantic image compression for resource-constrained IoT systemabstractThe widespread application of Internet-of-Things (IoT) and deep learning have made machine-to-machine semantic communication possible. However, it remains challenging to deploy DNN model on IoT devices, due to their limited computing and storage capacity. In this paper, we propose Compressed Sensing based Asymmetric Semantic Image Compression (CS-ASIC) for resource-constrained IoT systems, which consists of a lightweight front encoder and a deep iterative decoder offloaded at the server. We further consider a task-oriented scenario and optimize CS-ASIC for the semantic recognition tasks. The experiment results demonstrate that CS-ASIC achieves considerable data-semantic rate-distortion trade-off, and low encoding complexity over prevailing codecs. Yujun Huang, Bin Chen 0011, Jianghui Zhang, Han Qiu 0001, Shutao Xia |
DAC | 1 |
| 2022 | Improved DC Estimation for JPEG Compression Via Convex RelaxationabstractMass image transmission has undergone an explosion of growth with the development of the internet, DCT-based lossy image compression like JPEG is pervasively conducted to save the transmission bandwidth. Recently, DCT-domain coefficient estimation approaches have been proposed to further improve the compression ratio by discarding DC coefficients at the sender’s end while recovering them at the receiver’s end via DC estimation. However, known DC estimation needs to enumerate all possible DC coefficients. Consequently, they are limited and resource-consuming due to the low delay requirements in real-time transmission. In this paper, we propose an improved DC estimation method via convex relaxation, which achieves state-of-the-art performance in terms of both recovery image quality and time complexity. Extensive experiments across various data sets demonstrate the advantages of our method. Jianghui Zhang, Bin Chen 0011, Yujun Huang, Han Qiu 0001, Zhi Wang 0001, Shutao Xia |
ICIP | 3 |