EDBT 2026 Demo / reviewers in the wild / expert
Ai-Ping Yang
dblp:31/9145 · also Aiping Yang
· DBLP profile ↗
23ranked-venue papers
11as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HG-LMM: Unleashing High-Quality Pixel Grounding Capabilities in Frozen Large Multimodal ModelsabstractLarge Multimodal Models (LMMs) have demonstrated remarkable capabilities in multimodal understanding and conversation. Recently, some researchers have explored fine-tuning LMM for pixel grounding, leading to catastrophic loss of their inherent conversational capabilities. To preserve the conversational capability, some researchers explore to freeze LMM, but employ heavy segmenter SAM for high-quality grounding. In this paper, we propose a novel approach, named HG-LMM, to fully exploit the inherent features of frozen LMM for high-quality pixel grounding. Our HG-LMM introduces two main modules: an LLM-guided instance-aware feature generation (LIFG) and a layer-wise detail and semantic injection (LDSI). The LIFG module employs the output text embeddings of LMM belonging to grounded instances to generate multi-level instance-aware feature maps from the image encoder. Afterwards, we employ the LDSI module to inject more detail and semantic information into these instance-aware feature maps. With these instance-aware feature maps, we employ a simple top-down fusion to predict the segmentation masks of different instances. We perform experiments on various tasks, including referring expression segmentation, panoramic narrative grounding, reasoning segmentation, grounded conversation generation, and visual chain-of-thought reasoning. When using DeepSeekVL-1.3B, our HG-LMM is 12.6% better than F-LMM without SAM in terms of segmentation accuracy on the all set of PNG dataset. Compared to F-LMM with SAM, our HG-LMM achieves comparable segmentation accuracy while being 2.7 times faster. We release our source code and models at https://github.com/WenjieLi2008/HG-LMM. Jiale Cao, Jin Xie 0005, Ai-Ping Yang, Yanwei Pang |
IEEE Trans. Image Process. | 4 |
| 2025 | Self-Convolutional Attention-Based Uncertainty-Aware Network for Single-Image Super-ResolutionabstractCurrent super-resolution (SR) algorithms rely heavily on annotated data and often ignore the uncertainty in image degradation and features, limiting their real-world application. We propose an uncertainty-aware SR network using a self-convolutional attention mechanism. Our approach focuses on an SR reconstruction network enhanced by a cross-scale self-convolutional attention mechanism within the Transformer framework, which leverages local regions at various resolutions as convolution kernels to enhance high-frequency details. We also design a heteroscedastic uncertainty loss function to learn pixel and feature uncertainties, guiding the network to improve textures and edges adaptively. Extensive experiments show that our method achieves superior visual reconstruction on standard real-world datasets. Jinbin Wang, Ai-Ping Yang, Zihao Wei, Qinghua Hu |
ICASSP | 2 |
| 2025 | Uncertainty-Driven Weakly Supervised Dehazing Network: Integrating Dynamic Attention and Multi-Scale Feature FusionabstractHaze significantly affects image clarity, hindering vision-based systems like autonomous driving. However, current dehazing methods rely heavily on labeled data and overlook pixel-level uncertainties, limiting real-world effectiveness. To address this, we propose an uncertainty-driven weakly supervised dehazing network with a regional dynamic attention module and a transformer-based multi-scale feature fusion module. Firstly, the regional dynamic attention module employs a channel-pixel collaborative attention mechanism to adjust weights based on haze distribution, enhancing focus on hazy channels and dense regions. Meanwhile, the multi-scale feature fusion module integrates global semantics with local details through window-based multi-head self-attention and parallel dilated convolutions, effectively capturing long-range dependencies while preserving details. Additionally, we introduce an unsupervised loss function based on heteroscedastic uncertainty to guide training by adaptively correcting contrast, enhancing edges, and improving saturation. Extensive experiments on synthetic and real-world images demonstrate that our network shows superior performance in both objective quality and subjective evaluation. Jinbin Wang, Ai-Ping Yang, Qinghua Hu |
ICME | 2 |
| 2025 | The interpretable deep learning framework and validation for seizure detection in pediatric electroencephalography: An improved accuracy and performance analysis
Qiang Li 0048, Ai-Ping Yang, Ming-Lang Tseng |
Artif. Intell. Medicine | 5 |
| 2025 | Mitigating forgetting in the adaptation of CLIP for few-shot classification
Jiale Cao, Yuanheng Liu, Zhong Ji, Jingren Liu, Ai-Ping Yang, Yanwei Pang |
Comput. Vis. Image Underst. | 5 |
| 2025 | Bi-orientated rectification few-shot segmentation network based on fine-grained prototypes
Ai-Ping Yang, Zijia Sang, Yaran Zhou, Jiale Cao |
Neurocomputing | 1 |
| 2025 | Cross-Scale Atomic Feature Enhanced Network for high-fidelity Single Image Super-Resolution
Ai-Ping Yang, Chenhui Yu, Jinbin Wang, Zihao Wei, Jiale Cao |
Multim. Syst. | 1 |
| 2025 | Few-shot segmentation network based on class-aware prototype fusion
Ai-Ping Yang, Zijia Sang, Yaran Zhou |
Pattern Anal. Appl. | 1 |
| 2025 | Implicit and Explicit Language Guidance for Diffusion-Based Visual PerceptionabstractText-to-image diffusion models have shown powerful ability on conditional image synthesis. With large-scale vision-language pre-training, diffusion models are able to generate high-quality images with rich textures and reasonable structures under different text prompts. However, adapting pre-trained diffusion models for visual perception is an open problem. In this paper, we propose an implicit and explicit language guidance framework for diffusion-based visual perception, named IEDP. Our IEDP comprises an implicit language guidance branch and an explicit language guidance branch. The implicit branch employs a frozen CLIP image encoder to directly generate implicit text embeddings that are fed to the diffusion model without explicit text prompts. The explicit branch uses the ground-truth labels of corresponding images as text prompts to condition feature extraction in diffusion model. During training, we jointly train the diffusion model by sharing the model weights of these two branches. As a result, the implicit and explicit branches can jointly guide feature learning. During inference, we employ only implicit branch for final prediction, which does not require any ground-truth labels. Experiments are performed on two typical perception tasks, including semantic segmentation and depth estimation. Our IEDP achieves promising performance on both tasks. For semantic segmentation, our IEDP has the mIoU$^\text{ss}$score of 55.9% on ADE20K validation set, which outperforms the baseline method VPD by 2.2%. For depth estimation, our IEDP outperforms the baseline method VPD with a relative gain of 11.0%. Hefeng Wang, Jiale Cao, Jin Xie 0005, Ai-Ping Yang, Yanwei Pang |
IEEE Trans. Multim. | 4 |
| 2024 | Multi-query and multi-level enhanced network for semantic segmentation
Jiale Cao, Rao Muhammad Anwer, Jin Xie 0005, Jing Nie 0001, Ai-Ping Yang, Yanwei Pang |
Pattern Recognit. | 6 |
| 2024 | Progressive Semantic Reconstruction Network for Weakly Supervised Referring Expression GroundingabstractWeakly supervised Referring Expression Grounding (REG) aims to localize the target entity in an image based on a given expression, where the mapping between image regions and expressions is unknown during training. It faces two primary challenges. Firstly, conventional methods involve selecting regions to generate reconstructed texts for computing the backpropagation loss between regions and expressions. However, semantic deviations in text reconstruction may result in significant cross-modal bias, leading to substantial losses even in cases of correctly matched regions. Secondly, the absence of region-level ground truth in weakly supervised REG results in a lack of stable and reliable supervision during training. To tackle these challenges, we propose a Progressive Semantic Reconstruction Network (PSRN), which utilizes a two-level matching-reconstruction process based on the key triad and adaptive phrases, respectively. We leverage progressive semantic reconstruction with a three-staged training strategy to mitigate the deviations in the reconstructed texts. Additionally, we introduce a Constrained Interactions operation and an Attention Coordination mechanism to facilitate additional bidirectional supervision between the two matching processes. Experiments on three benchmark datasets of RefCOCO, RefCOCO+ and RefCOCOg demonstrate that the proposed PSRN has the competing results. Our source code will be released athttps://github.com/5jiahe/psrn. Zhong Ji, Jiahe Wu, Ai-Ping Yang, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Multi-feature self-attention super-resolution network
Ai-Ping Yang, Zihao Wei, Jinbin Wang, Jiale Cao, Zhong Ji, Yanwei Pang |
Vis. Comput. | 1 |
| 2023 | Spatial attention-guided deformable fusion network for salient object detection
Ai-Ping Yang, Simeng Cheng, Jiale Cao, Zhong Ji, Yanwei Pang |
Multim. Syst. | 1 |
| 2023 | Zero-reference single underwater image enhancement
Ai-Ping Yang, Chaochen Wang, Jinbin Wang |
Multim. Tools Appl. | 1 |
| 2023 | Visual-quality-driven unsupervised image dehazing
Ai-Ping Yang, Jinbin Wang, Jiale Cao, Zhong Ji, Yanwei Pang |
Neural Networks | 1 |
| 2022 | Saliency detection network with two-stream encoder and interactive decoder
Ai-Ping Yang, Simeng Cheng, Shangyang Song, Jinbin Wang, Zhong Ji, Yanwei Pang, Jiale Cao |
Neurocomputing | 1 |
| 2022 | Non-linear perceptual multi-scale network for single image super-resolution
Ai-Ping Yang, Jinbin Wang, Zhong Ji, Yanwei Pang, Jiale Cao, Zihao Wei |
Neural Networks | 1 |
| 2020 | Lightweight group convolutional network for single image super-resolution
Ai-Ping Yang, Bingwang Yang, Zhong Ji, Yanwei Pang, Ling Shao 0001 |
Inf. Sci. | 1 |
| 2019 | Dual-Path in Dual-Path Network for Single Image DehazingabstractRecently, deep learning-based single image dehazing method has been a popular approach to tackle dehazing. However, the existing dehazing approaches are performed directly on the original hazy image, which easily results in image blurring and noise amplifying. To address this issue, the paper proposes a DPDP-Net (Dual-Path in Dual-Path network) framework by employing a hierarchical dual path network. Specifically, the first-level dual-path network consists of a Dehazing Network and a Denoising Network, where the Dehazing Network is responsible for haze removal in the structural layer, and the Denoising Network deals with noise in the textural layer, respectively. And the second-level dual-path network lies in the Dehazing Network, which has an AL-Net (Atmospheric Light Network) and a TM-Net (Transmission Map Network), respectively. Concretely, the AL-Net aims to train the non-uniform atmospheric light, while the TM-Net aims to train the transmission map that reflects the visibility of the image. The final dehazing image is obtained by nonlinearly fusing the output of the Denoising Network and the Dehazing Network. Extensive experiments demonstrate that our proposed DPDP-Net achieves competitive performance against the state-of-the-art methods on both synthetic and real-world images. Ai-Ping Yang, Zhong Ji, Yanwei Pang, Ling Shao 0001 |
IJCAI | 1 |
| 2019 | Image-attribute reciprocally guided attention network for pedestrian attribute recognition
Zhong Ji, Erlu He, Haoran Wang 0004, Ai-Ping Yang |
Pattern Recognit. Lett. | 4 |
| 2018 | Learning intensity and detail mapping parameters for dehazing
Xuhang Lian, Yanwei Pang, Ai-Ping Yang |
Multim. Tools Appl. | 3 |
| 2014 | Frequency Domain Directional Filtering Based Rain Streaks Removal from a Single Color Image
Changbo Liu, Yanwei Pang, Jian Wang 0087, Ai-Ping Yang |
ICIC (1) | 4 |
| 2009 | All phase biorthogonal transform and its application in JPEG-like image compression
Zheng-Xin Hou, Chengyou Wang, Ai-Ping Yang |
Signal Process. Image Commun. | 3 |