EDBT 2026 Demo / reviewers in the wild / expert
Yingxu Qiao
dblp:279/2348
· DBLP profile ↗
23ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0001-6227-614XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Real-Time Semantic Segmentation Network with Boundary-Focused and Multi-Scale Context FusionabstractABSTRACT Real‐time semantic segmentation is crucial for applications including autonomous driving and augmented reality. While current real‐time semantic segmentation methods achieve a balance between accuracy and speed, an adequate capture of boundary details remains a challenge for many models. Furthermore, as deep learning networks become increasingly complex, certain approaches encounter challenges, including excessive computational overhead and numerous parameters when capturing multi‐scale contextual features. To address these limitations, the boundary‐focused and multi‐scale context fusion network (BFMSNet) is proposed, a lightweight real‐time semantic segmentation model that enhances boundary perception and contextual understanding. A boundary refinement module is designed, which utilizes multi‐level feature fusion and a gating mechanism to precisely capture edge details in complex scenes and achieve pixel‐level boundary alignment and optimization. Furthermore, a hybrid boundary loss is introduced, combining region and boundary supervision signals to effectively guide the network's focus on challenging regions, thereby improving training stability and segmentation accuracy. To reduce model complexity, a lightweight multi‐scale fusion module is implemented based on the multi‐scale frequency‐domain characteristics of wavelet convolution. This module balances context information extraction and computational efficiency, reducing parameters while maintaining feature representation. Experimental results on the Cityscapes and CamVid datasets demonstrate that BFMSNet achieves mIoU of 78.53% and 76.24%, while maintaining real‐time inference speeds of 86.25 FPS and 143.70 FPS, respectively. Preliminary tests indicate that the BFMSNet algorithm effectively balances accuracy and speed requirements. Shan Zhao 0009, Fukai Zhang, Zhanqiang Huo, Yingxu Qiao |
IET Image Process. | 5 |
| 2026 | A bio-inspired photopic-scotopic duplex network for low-light image enhancement
Zihan Cao, Manli Wang, Changsen Zhang, Yannan Shi, Dekui Li, Yingxu Qiao |
Neurocomputing | 6 |
| 2026 | VMamba-LLIE: enhancing low-light images with snr prior-guided and HVI color-assisted triple-branch network
Zhanqiang Huo, Pengyun Shi, Yizhang Meng, Yingxu Qiao, Shan Zhao 0009 |
Multim. Syst. | 4 |
| 2026 | Towards universal object detection: a fine-grained perspective on open-vocabulary object detectionabstractOpen-vocabulary object detection is an innovative computer vision task capable of widely recognizing and localizing various objects in images. Unlike traditional methods, it can handle diverse object categories and is suitable for real-time applications in dynamic environments. Existing methods typically achieve zero-shot detection capabilities by fusing images and text. However, when semantic discrepancies exist between text and images, biased prediction problems arise, diminishing the effectiveness of semantic guidance. To address these issues, we propose a universal open-vocabulary object detection method that leverages foundational models to provide fine-grained semantic guidance for the detection process. We design a multi-level adaptive scene perception algorithm that captures subtle features of target objects in complex scenes, enabling precise separation of background and foreground. Additionally, we introduce the Text-KAN (T-KAN) model, which integrates textual descriptions with image features. By employing learnable activation functions, it resolves dependencies on linear matrices, enhances text interpretability, corrects semantic biases, and achieves precise alignment between images and text at a fine-grained level. We comprehensively evaluate the performance of our proposed method on existing open-vocabulary benchmarks, conducting experiments on the COCO and LVIS datasets. The results demonstrate significant performance gains in detecting novel categories, highlighting the method’s strong generalization capabilities. This work provides valuable insights and references for advancing the field of open-vocabulary object detection. Jing Wang 0093, Yonghua Cao, Zhanqiang Huo, Yingxu Qiao |
Multim. Syst. | 4 |
| 2026 | Hunting for the unknown: Open world object detection from a class-agnostic perspective
Jing Wang 0093, Yonghua Cao, Zhanqiang Huo, Yingxu Qiao |
Neural Networks | 4 |
| 2026 | ACA-ViT: Adaptive class-aware vision transformer with complementary CAMs for weakly-supervised semantic segmentation
Zhanqiang Huo, Qiankun Zhao, Yijiang Wang, Yingxu Qiao |
Pattern Recognit. Lett. | 4 |
| 2026 | AEAFFNet: enhancing real-time semantic segmentation through attention-enhanced adaptive feature fusion
Shan Zhao 0009, Wenjing Fu, Fukai Zhang, Zhanqiang Huo, Yingxu Qiao |
J. Supercomput. | 5 |
| 2026 | SGTNet: real-time semantic segmentation via sparse transformer integration and multi-scale feature fusion
Shan Zhao 0009, Kaiyu Zhou, Fukai Zhang, Zhanqiang Huo, Yingxu Qiao |
Vis. Comput. | 5 |
| 2025 | More Realistic Edges, Textures, and Colors for Image Non-Homogeneous DehazingabstractABSTRACT The existing image dehazing algorithms perform suboptimal in non‐homogeneous and/or dense haze scenarios. The loss of feature information and alteration of color distribution cause images to deviate from real‐world scenes when haze suppresses image details. To address these issues, we design a dual‐branch non‐homogeneous dehazing network integrating discrete wavelet transform (DWT), multi‐scale feature fusion, and color constraints to achieve dehazed images with more realistic edges, textures, and colors. Specifically, we first introduce DWT into a multi‐scale encoder–decoder network structure to capture more details and edge information. Then, a feature supplement and enhancement module (FSEM) combining features from hazy images at different scales and features from the previous stage is devised to enhance the multi‐scale feature capture capability of rich textures in complex scenes. Finally, we propose a pixel‐wise color consistency loss that combines pixel similarity and angular difference to constrain the dehazed images to closely match the color distribution of clear images. Experimental results indicate that the proposed dehazing network outperforms the state‐of‐the‐art non‐homogeneous dehazing methods on relevant public benchmarks and has more realistic edges, textures, and colors. Hairu Guo, Zhanqiang Huo, Shan Zhao 0009, Yingxu Qiao |
IET Image Process. | 5 |
| 2025 | Combining implicit and explicit priors for zero-reference low-light image enhancement and denoising
Jinxia Yu, Fabao Xue, Zhanqiang Huo, Yingxu Qiao |
Multim. Syst. | 4 |
| 2025 | Dark channel map and union training strategy for object detection in foggy scenes
Zhanqiang Huo, Sensen Meng, Yingxu Qiao, Shan Zhao 0009 |
Pattern Recognit. Lett. | 4 |
| 2025 | VMamba-Crowd: Bridging multi-scale features from Visual Mamba for weakly-supervised crowd counting
Zhanqiang Huo, Chunxin Yuan, Kunwei Zhang, Yingxu Qiao, Fen Luo |
Pattern Recognit. Lett. | 4 |
| 2025 | D3-Dehaze: a divide-and-conquer framework for enhanced single image dehazing
Zhanqiang Huo, Xiyan Zhan, Yingxu Qiao, Shan Zhao 0009 |
Vis. Comput. | 3 |
| 2024 | Dual-route synthetic-to-real adaption for single image dehazingabstractAbstract Single image dehazing in real scenarios is still a particularly challenging task due to the large domain discrepancy between synthetic and real hazy images. To address this problem, the dual‐route domain‐aware adaptation framework with the bridging domain is proposed. First, inspired by the effective multistage strategy in low‐level tasks, the bridging domain is constructed with the color‐preserved adaptive histogram equalization to facilitate the adaptation and improve the performance of real hazy images. Second, the totally shared structure is relaxed and the residual dual‐path domain‐aware modules (RDDM) for synthetic and the bridging domains are proposed, which facilitates extracting the domain‐specific haze information with the different parameters. Third, a half‐cyclic constraint is proposed for the unsupervised hazy images to avoid structure distortion during the unsupervised adversarial training process. Finally, for convenience in the inference stage, feature enhancement modules (FEM) are proposed for the original real hazy images to learn the pre‐process operation in the training stage. Extensive qualitative and quantitative experiments demonstrate that the proposed method significantly improves the dehazing performance on synthetic and real hazy images. Yingxu Qiao, Zhanqiang Huo, Sensen Meng |
IET Image Process. | 1 |
| 2024 | Rethinking Zero-DCE for Low-Light Image EnhancementabstractAbstract Zero-Reference Deep Curve Estimation (Zero-DCE) is currently one of the most popular low-light image enhancement methods. Through extensive experimentation, we observe that: (i) the excellent performance of Zero-DCE depends on the training data with multiple exposure levels, (ii) it cannot effectively handle uneven light, extremely low light, or overexposed images in natural environments. Therefore, we propose an improved zero-reference dual-illumination deep curve estimation method for low-light image enhancement named Zero-DiDCE, which can enhance, suppress, or maintain light levels for images. The adaptive light enhancement curve was designed to handle images with different exposure levels. An iterator and amplitude controller are designed to control the curve enhancement intensity by calculating the gap between the input image and the optimal light level. Furthermore, instead of the DCE-Net in Zero-DCE only taking the input image as network input, our DiDCE-Net in Zero-DiDCE takes the input image and the inverted input image simultaneously as network input to ensure that the training set contains samples with multiple exposure levels. A piecewise non-reference loss function is designed to guide the training of DiDCE-Net from the perspective of information loss. Qualitative and quantitative experiments show that our method can handle images with different levels of exposure well and outperforms state-of-the-art methods. In addition, the proposed curve and iterator can be integrated into other methods to improve their enhancement effects. The code is available at https://github.com/Wenhui-Luo/Zero-DiDCE . Aizhong Mi, Wenhui Luo, Yingxu Qiao, Zhanqiang Huo |
Neural Process. Lett. | 3 |
| 2024 | Robust Synthetic-to-Real Ensemble Dehazing Algorithm With the Intermediate DomainabstractLearning-based dehazing methods using synthetic datasets cannot generalize well on real-world hazy images due to the large domain discrepancy. To tackle this issue, we propose a robust synthetic-to-real dehazing framework with the construction of an intermediate domain and ensemble learning strategy. First, by mapping all examples to the intermediate domain, the bidirectional match strategy with adversarial training and the constraint of intermediated results is proposed to suppress the rich domain-specific information, which can facilitate the adaptation and perform image dehazing simultaneously. Furthermore, an ensemble dehazing algorithm based on the intermediate domain is proposed in a semisupervised manner. The reconstruction constraint and the enhanced ground-truths are employed to keep the visual fidelity and remove the dim artifacts of unsupervised dehazing results. Finally, we propose the domain-aware residual groups to deal with the distribution discrepancy between the synthetic and real hazy images. Extensive experiments of various real-world hazy images demonstrate that the proposed method outperforms the state-of-the-art dehazing methods and significantly improves the generalization in the real world. Yingxu Qiao, Hongmin Liu 0001, Zhanqiang Huo |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | Boosting Object Detection in Foggy Scenes via Dark Channel Map and Union Training Strategy
Zhanqiang Huo, Sensen Meng, Yingxu Qiao, Fen Luo |
PRCV (12) | 3 |
| 2023 | PVT-Crowd: Bridging Multi-scale Features from Pyramid Vision Transformer for Weakly-Supervised Crowd Counting
Zhanqiang Huo, Kunwei Zhang, Fen Luo, Yingxu Qiao |
PRCV (9) | 4 |
| 2023 | Domain adaptive crowd counting via dynamic scale aggregation networkabstractAbstract Crowd counting is an important research topic in computer vision. Its goal is to estimate the people's number in an image. Researchers have dramatically improved counting accuracy in recent years by regressing density maps. However, because of the inherent domain shift, the model trained on an expensive manually labelled dataset (source domain) does not perform well on a dataset with scarce labels (target domain). For this issue, a novel dynamic scale aggregation network (DSANet) is proposed to reduce the gaps in style and cross‐domain head scale variations. Specifically, a practical style transfer layer is introduced to reduce the appearance discrepancy between the source and target domains. Then, the translated source and target domain samples are encoded by a generator consisting of the VGG16 network and the dynamic scale aggregation modules (DSA Modules) and produce corresponding density maps. The DSA module can adaptively adjust parameters according to the input features and effectively fuse multi‐scale information to overcome the cross‐domain head scale variations. Next, a discriminator judges the input density map from the source or target domain. Last, domain distributions are aligned through adversarial between the generator and the discriminator. The experiments show that our network outperforms the current state‐of‐the‐art methods and can improve the target domain's performance while maintaining the source domain's performance without significant degradation. Zhanqiang Huo, Yingxu Qiao, Jing Wang 0093, Fen Luo |
IET Comput. Vis. | 3 |
| 2023 | Semantics recalibration and detail enhancement network for real-time semantic segmentationabstractAbstract Real‐time semantic segmentation is a crucial technology in automatic driving scenarios, which needs to meet both high precision and real‐time. The authors observe that learning complex correlations between object categories is vital in the real‐time semantic segmentation task. Moreover, image spatial detail information plays an important role in small object segmentation and preserving edges and textures. A Semantics Recalibration and Detail Enhancement Network for real‐time semantic segmentation based on BiSeNet V2 is proposed. On the one hand, a lightweight Semantics Recalibration module is designed to effectively extract global semantic contextual information, which combines pyramid segmentation and adaptive recalibration operations to learn the correlations between object categories. On the other hand, a Detail Enhancement module takes the feature maps of the shallow layers in the semantics branch as input and refines the feature maps to highlight the detail information. Finally, quantitative and qualitative analyses on Cityscapes and CamVid datasets demonstrate the effectiveness and generalisation of the proposed method. Aizhong Mi, Mingming Gao, Zhanqiang Huo, Yingxu Qiao, Haiyang Jia |
IET Comput. Vis. | 4 |
| 2022 | Efficient photorealistic style transfer with multi-order image statistics
Zhanqiang Huo, Xueli Li, Yingxu Qiao, Panbo Zhou, Jing Wang 0093 |
Appl. Intell. | 3 |
| 2022 | High quality proposal feature generation for crowded pedestrian detection
Jing Wang 0093, Cailing Zhao, Zhanqiang Huo, Yingxu Qiao, Haifeng Sima |
Pattern Recognit. | 4 |
| 2021 | Efficient Style-Corpus Constrained Learning for Photorealistic Style TransferabstractPhotorealistic style transfer is a challenging task, which demands the stylized image remains real. Existing methods are still suffering from unrealistic artifacts and heavy computational cost. In this paper, we propose a novel Style-Corpus Constrained Learning (SCCL) scheme to address these issues. The style-corpus with the style-specific and style-agnostic characteristics simultaneously is proposed to constrain the stylized image with the style consistency among different samples, which improves photorealism of stylization output. By using adversarial distillation learning strategy, a simple fast-to-execute network is trained to substitute previous complex feature transforms models, which reduces the computational cost significantly. Experiments demonstrate that our method produces rich-detailed photorealistic images, with 13 ~ 50 times faster than the state-of-the-art method (WCT2). Yingxu Qiao, Jiabao Cui, Fuxian Huang, Hongmin Liu 0001, Cuizhu Bao, Xi Li 0001 |
IEEE Trans. Image Process. | 1 |