EDBT 2026 Demo / reviewers in the wild / expert
Yuning Huang
dblp:360/9834
· DBLP profile ↗
18ranked-venue papers
2as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Facade parsing via joint structural priors and phased deep learning
Yuning Huang, Weize Quan, Dong-Ming Yan 0001, Jie Jiang 0017, Yingmei Wei |
Neurocomputing | 2 |
| 2026 | QARV++: An Improved Hierarchical VAE for Learned Image Compressionabstractvalues, stabilizing variable-rate training. Extensive experiments demonstrate that QARV++ achieves superior rate-distortion (R-D) performance among HVAE-based LIC models, exhibiting -12.20% -16.34% -15.23% BD-Rate against VVC Intra mode on the Kodak, Tecnick, and CLIC2020 test datasets, respectively. Our approach also generalizes effectively to existing LICs, delivering substantial improvements. Yuning Huang, Fengqing Zhu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Gain From Give Up: Intuitive Data Augmentation Framework for Image RetrievalabstractModern deep hashing methods rely on data augmentation to fully realize their potential in the face of large-scale retrieval galleries and over-parameterized visual models. However, this work observes that mainstream label-preserving augmentation methods are unreliable for information retrieval because they lead to incomplete alignment between data and labels. This misalignment impairs metric losses in distinguishing original/augmented data during same-class clustering, compromising nearest-neighbor search efficacy. To address these issues, we propose an innovative plug-and-play data augmentation framework tailored for retrieval tasks, based on the concept ofGaining robust features by randomlyGiving up parts of the image (GG). Inspired by the ease with which visual changes induced by discard transformations can be estimated, we design two intuitive augmentation methods along with corresponding semantic shift estimators to measure the semantic changes introduced by each operation. Additionally, we optimize the metric loss based on the semantic retention scores, guiding the metric objective to properly allocate gradients for generated samples. This adjustment mitigates the adverse effects caused by incomplete alignment, optimizing the intra-class distance of both original and augmented data in the Hamming space, while ensuring the relevance and accuracy of the retrieval results. Extensive experiments conducted on six datasets demonstrate the effectiveness and robustness of our proposed framework. Code is available athttps://github.com/wuhulahu/GG. Yurong Qian, Guangqi Yang, Yuning Huang |
IEEE Trans. Multim. | 6 |
| 2025 | Deep Learning-Based Feature Fusion for Emotion Analysis and Suicide Risk Differentiation in Chinese Psychological Support HotlinesabstractMental health is a significant global public health issue, and psychological support hotlines play a crucial role in providing mental health assistance and identifying suicide risks at an early stage. However, the emotional expressions conveyed during these calls remain underexplored in current research. This study introduces a novel method that combines pitch acoustic features with deep learning-based features to analyze and understand emotions expressed during hotline interactions. Using data from China's largest psychological support hotline, which includes 105 subjects, our method achieved an F1-score of 79.13% for negative binary emotion classification. Additionally, the proposed approach was validated on an open dataset for multi-class emotion classification, where it demonstrated better performance compared to the state-of-the-art methods. To explore its clinical relevance, we applied the model to analysis the frequency of negative emotions and the rate of emotional change in the conversation, comparing 46 subjects with suicidal behavior to those without. While the suicidal group exhibited more frequent emotional changes than the non-suicidal group, the difference was not statistically significant. Importantly, our findings suggest that emotional fluctuation intensity and frequency could serve as novel features for psychological assessment scales and suicide risk prediction. The proposed method provides valuable insights into emotional dynamics and has the potential to advance early intervention and improve suicide prevention strategies through integration with clinical tools and assessments. The source code is publicly available at: https://github.com/Sco-field/Speechemotionrecognition/tree/main. Han Wang 0059, Jianqiang Li 0002, Qing Zhao 0005, Zhonglong Chen, Changwei Song, Yuning Huang, Wei Zhai, Yongsheng Tong, Guanghui Fu |
COMPSAC | 7 |
| 2025 | Balanced Rate-Distortion Optimization in Learned Image CompressionabstractLearned image compression (LIC) using deep learning architectures has seen significant advancements, yet standard rate-distortion (R-D) optimization often encounters imbalanced updates due to diverse gradients of the rate and distortion objectives. This imbalance can lead to suboptimal optimization, where one objective dominates, thereby reducing overall compression efficiency. To address this challenge, we reformulate R-D optimization as a multi-objective optimization (MOO) problem and introduce two balanced R-D optimization strategies that adaptively adjust gradient updates to achieve more equitable improvements in both rate and distortion. The first proposed strategy utilizes a coarse-to-fine gradient descent approach along standard RD optimization trajectories, making it particularly suitable for training LIC models from scratch. The second proposed strategy analytically addresses the reformulated optimization as a quadratic programming problem with an equality constraint, which is ideal for fine-tuning existing models. Experimental results demonstrate that both proposed methods enhance the R-D performance of LIC models, achieving around a 2% BD-Rate reduction with acceptable additional training cost, leading to a more balanced and efficient optimization process. Code will be available at https://gitlab.com/viper-purdue/Balanced-RD. Zhihao Duan, Yuning Huang, Fengqing Zhu 0001 |
CVPR | 3 |
| 2025 | Low-Rank Adaptation of Pre-Trained Vision Backbones for Energy-Efficient Image Coding For MachinesabstractImage Coding for Machines (ICM) focuses on optimizing image compression for AI-driven analysis rather than human perception. Existing ICM frameworks often rely on separate codecs for specific tasks, leading to significant storage requirements, training overhead, and computational complexity. To address these challenges, we propose an energy-efficient framework that leverages pre-trained vision backbones to extract robust and versatile latent representations suitable for multiple tasks. We introduce a task-specific low-rank adaptation mechanism, which refines the pre-trained features to be both compressible and tailored to downstream applications. This design minimizes trainable parameters and reduces energy costs for multi-task scenarios. By jointly optimizing task performance and entropy minimization, our method enables efficient adaptation to diverse tasks and datasets without full fine-tuning, achieving high coding efficiency. Extensive experiments demonstrate that our framework significantly outperforms traditional codecs and pre-processors, offering an energy-efficient and effective solution for ICM applications. The code and the supplementary materials will be available at: https://gitlab.com/viper-purdue/efficient-compression. Zhihao Duan, Yuning Huang, Fengqing Zhu 0001 |
ICIP | 3 |
| 2025 | UH-PCC: Unified Octree and Feature Coding for Hierarchical Point Cloud Geometry Compression
Muhammad Asad Lodhi, Jiahao Pang, Junghyun Ahn, Yuning Huang, Dong Tian |
PCS | 4 |
| 2024 | Comparative Analysis of ImageNet Pre-Trained Deep Learning Models and DINOv2 in Medical Imaging ClassificationabstractMedical image analysis frequently encounters data scarcity challenges. Transfer learning has been effective in addressing this issue while conserving computational resources. The recent advent of foundational models like the DINOv2, which uses the vision transformer architecture, has opened new opportunities in the field and gathered significant interest. However, DINOv2's performance on clinical data still needs to be verified. In this paper, we performed a glioma grading task using three clinical modalities of brain MRI data. We compared the performance of various pre-trained deep learning models, including those based on ImageNet and DINOv2, in a transfer learning context. Our focus was on understanding the impact of the freezing mechanism on performance. We also validated our findings on three other types of public datasets: chest radiography, fundus radiography, and dermoscopy. Our findings indicate that in our clinical dataset, DINOv2's performance was not as strong as ImageNet-based pre-trained models, whereas in public datasets, DINOv2 generally outperformed other models, especially when using the frozen mechanism. Similar performance was observed with various sizes of DINOv2 models across different tasks. In summary, DINOv2 is viable for medical image classification tasks, particularly with data resembling natural images. However, its effectiveness may vary with data that significantly differs from natural images such as MRI. In addition, employing smaller versions of the model can be adequate for medical task, offering resource-saving benefits. Our codes are available at https://github.com/GuanghuiFU/medical_dino_eval. Yuning Huang, Jingchen Zou, Lanxi Meng, Xin Yue, Qing Zhao 0005, Jianqiang Li 0002, Changwei Song, Gabriel Jimenez 0001, Shaowu Li, Guanghui Fu |
COMPSAC | 1 |
| 2024 | Progressive Sign Language Video Translation Model for Real-World Complex Background EnvironmentsabstractSign language video translation, which converts sign language information into textual expressions, play a vital role in breaking down the language communication barrier between deaf and healthy people. Existing translation methods are mainly focus on the single and pure background. However, the background in real-world environments is always complex, and these methods are difficult to achieve effective recognition results. To address this issue, we have exploratively constructed a real-world complex background sign language dataset (CBSL), containing sign language videos captured in various authentic environments (e.g., different backgrounds and lighting conditions). Based on this, we propose a progressive sign language translation model to effectively separate sign language users from the background and reduce environmental interference, thus significantly improving the generalization ability. Our proposed method significantly outperforms various comparative methods across all performance metrics on the CBSL dataset. Furthermore, on the publicly available Chinese Sign Language Continuous Recognition dataset(CSL), our method performs comparably to the current state-of-the-art (SOTA). Jingchen Zou, Jianqiang Li 0002, Yuning Huang, Changwei Song, Linna Zhao, Wenxiu Cheng, Chujie Zhu, Suqin Liu |
COMPSAC | 4 |
| 2024 | Theoretical Bound-Guided Hierarchical Vae For Neural Image CodecsabstractRecent studies reveal a significant theoretical link between variational autoencoders (VAEs) and rate-distortion theory, notably in utilizing VAEs to estimate the theoretical upper bound of the information rate-distortion function of images. Such estimated theoretical bounds substantially exceed the performance of existing neural image codecs (NICs). To narrow this gap, we propose a theoretical bound-guided hierarchical VAE (BG-VAE) for NIC. The proposed BG-VAE leverages the theoretical bound to guide the NIC model towards enhanced performance. We implement the BG-VAE using Hierarchical VAEs and demonstrate its effectiveness through extensive experiments. Along with advanced neural network blocks, we provide a versatile, variable-rate NIC that outperforms existing methods when considering both ratedistortion performance and computational complexity. The code is available at $\color{magenta}{\text{BG - VAE}}$. Zhihao Duan, Yuning Huang, Fengqing Zhu 0001 |
ICME | 3 |
| 2024 | MAHFF: An Underwater Image Enhancement Method Based on Multi-scale Attention Hybrid Feature FusionabstractUnderwater images often suffer from image quality degradation due to the influence of refraction and reflection from suspended matter in water, as well as uneven light absorption. This can result in image blurring, loss of detail, fogging, and color distortion, which can significantly impact underwater visual tasks. This paper proposes a method for enhancing the quality of underwater images through multi-scale attention hybrid feature fusion. The input underwater image undergoes multiscale feature extraction using a multiscale attention hybrid feature extraction block to obtain different levels of feature representations. Feature fusion and enhancement are then achieved through a multilevel feature depth coupling module using multiple Resize Blocks to couple the multiscale features in depth. The combination of these two modules effectively solves the problem of degraded picture quality in underwater images. Experimental results demonstrate that the proposed method in this paper improves PSNR by 7.897%, SSIM by 1.724%, and VIF by 2.796%, proving the algorithm’s effectiveness and superiority. Yuning Huang, Yucai Li |
IJCNN | 2 |
| 2024 | Towards Reproducible Learning-Based CompressionabstractA deep learning system typically suffers from a lack of reproducibility that is partially rooted in hardware or software implementation details. The irreproducibility leads to skepticism in deep learning technologies and it can hinder them from being deployed in many applications. In this work, the irreproducibility issue is analyzed where deep learning is employed in compression systems while the encoding and decoding may be run on devices from different manufacturers. The decoding process can even crash due to a single bit difference, e.g., in a learning-based entropy coder. For a given deep learning-based module with limited resources for protection, we first suggest that reproducibility can only be assured when the mismatches are bounded. Then a safeguarding mechanism is proposed to tackle the challenges. The proposed method may be applied for different levels of protection either at the reconstruction level or at a selected decoding level. Furthermore, the overhead introduced for the protection can be scaled down accordingly when the error bound is being suppressed. Experiments demonstrate the effectiveness of the proposed approach for learning-based compression systems, e.g., in image compression and point cloud compression. Jiahao Pang, Muhammad Asad Lodhi, Junghyun Ahn, Yuning Huang, Dong Tian |
MMSP | 4 |
| 2024 | Probing Image Compression for Class-Incremental LearningabstractImage compression emerges as a pivotal tool in the efficient handling and transmission of digital images. Its ability to substantially reduce file size not only facilitates enhanced data storage capacity but also potentially brings advantages to the development of continual machine learning (ML) systems, which learn new knowledge incrementally from sequential data. Continual ML systems often rely on storing representative samples, also known as exemplars, within a limited memory constraint to maintain the performance on previously learned data. These methods are known as memory replay-based algorithms and have proven effective at mitigating the detrimental effects of catastrophic forgetting. Nonetheless, the limited memory buffer size often falls short of adequately representing the entire data distribution. In this paper, we explore the use of image compression as a strategy to enhance the buffer's capacity, thereby increasing exemplar diversity. However, directly using compressed exemplars introduces domain shift during continual ML, marked by a discrepancy between compressed training data and uncompressed testing data. Additionally, it is essential to determine the appropriate compression algorithm and select the most effective rate for continual ML systems to balance the trade-off between exemplar quality and quantity. To this end, we introduce a new framework to incorporate image compression for continual ML including a pre-processing data compression step and an efficient compression rate/algorithm selection method. We conduct extensive experiments on CIFAR-100 and ImageNet datasets and show that our method significantly improves image classification accuracy in continual ML settings. Justin Yang, Zhihao Duan, Andrew Peng, Yuning Huang, Jiangpeng He, Fengqing Zhu 0001 |
PCS | 4 |
| 2024 | MGS-Net: Fusing Global and Local Feature Enhancements for Healthcare Education Management of Myasthenia Gravis Using Speech DataabstractMyasthenia gravis (MG) is a neurological disease that is difficult to diagnose and requires long-term management. The progression of this disease is reflected to some extent in changes in speech, such as hoarseness and articulation disorders. However, it is difficult for general neurologists to grasp the diagnostic patterns of such rare diseases, especially in underdeveloped regions. As an emerging field, speech-based intelligent diagnostic assistance provides a safe, non-invasive, and convenient solution for healthcare education management. To this end, we firstly constructed a novel Chinese speech dataset of myasthenia gravis patients (MGCS). Then we proposed a network named Myasthenia Gravis Speech Net (MGS-Net) for the classification of myasthenia gravis pathological speech, which is mainly composed of two blocks: the Local Feature Enhancement (LFE) block and the Feedforward Dense (FFD) block. The LFE block extracts temporal local features using a sliding window approach, while the FFD block captures the global representation of the data. Compared to existing methods, our pipeline achieves an accuracy of 98.75% and a recall rate of 99.17%. We validated the effectiveness of existing acoustic feature sets in pathological speech classification of MG, which will provide an important tool for health education management of neurological diseases. Jianqiang Li 0002, Jingchen Zou, Yuning Huang, Shujie Ding, Linna Zhao |
SMC | 4 |
| 2024 | Sign Language Recognition and Translation Methods Promote Sign Language Education: A ReviewabstractSign language recognition and translation (SLRT) aims to convert sign language into textual representation, which holds significant importance for the deaf community. Sign language possesses complex and diverse grammatical structures, with each sign language having distinct motion trajectories and gesture variations, making SLRT a complex research domain. In recent years, numerous researchers have proposed differ-ent modeling approaches, achieving significant advancements through the utilization of large language models. In this survey, we systematically review the developmental trajectory of SLRT, encompassing an introduction to key technical approaches at each stage and the latest research progress. Through a comprehensive examination of these methods, valuable insights are provided for future research and practical applications. Lastly, we identify the existing limitations of current methods and propose potential avenues for future research. Jingchen Zou, Jianqiang Li 0002, Yuning Huang, Shujie Ding |
SMC | 4 |
| 2024 | QARV: Quantization-Aware ResNet VAE for Lossy Image CompressionabstractThis paper addresses the problem of lossy image compression, a fundamental problem in image processing and information theory that is involved in many real-world applications. We start by reviewing the framework of variational autoencoders (VAEs), a powerful class of generative probabilistic models that has a deep connection to lossy compression. Based on VAEs, we develop a new scheme for lossy image compression, which we name quantization-aware ResNet VAE (QARV). Our method incorporates a hierarchical VAE architecture integrated with test-time quantization and quantization-aware training, without which efficient entropy coding would not be possible. In addition, we design the neural network architecture of QARV specifically for fast decoding and propose an adaptive normalization operation for variable-rate compression. Extensive experiments are conducted, and results show that QARV achieves variable-rate compression, high-speed decoding, and better rate-distortion performance than existing baseline methods. Zhihao Duan, Ming Lu 0003, Jack Ma, Yuning Huang, Zhan Ma 0001, Fengqing Zhu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Multi-layer collaborative task offloading optimization: balancing competition and cooperation across local edge and cloud resources
Bowen Ling, Xiaoheng Deng, Yuning Huang, Jinsong Gui, Yurong Qian |
J. Supercomput. | 3 |
| 2023 | Efficient Joint Video Denoising and Super-ResolutionabstractDenoising and super-resolution are two important tasks for video enhancement. Despite recent progress for each task, there are very few works that target both tasks simultaneously. In this paper, we propose an efficient noise-robust video super-resolution method that is trained end-to-end for an input video containing observable noises. We investigate current approaches to address this joint denoising and super-resolution task and compare them to our proposed method. Experimental results show that our method achieves competitive reconstruction performance with existing solutions on various datasets while maintaining a low computation cost and a small model size which prove the effectiveness of our joint model design and training. Our code is available at "https://github.com/Eventhyn/EVDSRNet.". Yuning Huang, Qian Lin 0001, Jan P. Allebach, Fengqing Zhu 0001 |
ICIP | 1 |