Yujie Dun

dblp:22/2232 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0001-5213-1000ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Human Pose Estimation in Low-Light Condition With Decomposition and Modulation
abstract
Human pose estimation (HPE) is a fundamental problem in computer vision, aiming to locate anatomical keypoints of the human body in a given picture. Benefiting from recent progress in deep learning, dominant HPE methods can achieve more advanced performance. Unfortunately, these methods rely heavily on large-scale, high-quality datasets captured expensively, resulting in limited learning capabilities in data-constrained low-light situations. Existing methods enhance the model's ability for low-light scenarios by performing intermediate feature alignment between low-light image and its well-lit counterpart. However, these methods fall short in fully exploiting explicit semantic feature exploitation that is independent of lighting conditions, resulting in sub-optimal performance. In this paper, we propose a Progressive Decomposition-Modulation network (PDMNet) for human pose estimation in extremely low-light condition. In particular, PDMNet mainly consists of 1) a semantic-specific decomposition module (SDM) for decomposing reflectance component with rich semantic information, and 2) a semantic-specific modulation mechanism (SMM) that enables the reflectance component to modulate the representation learning of human body parts in a tailored manner. Two closely-related components cooperate with each other to achieve more effective content-specific feature learning in low-light conditions. We further equip them progressively into different scales to enhance the feature learning. Experimental results demonstrate the superiority of PDMNet over state-of-the-art models on publicly available datasets. Our code will be released soon.
Chengxu Liu 0001, Yujie Dun, Xueming Qian
IEEE Trans. Multim.3
2025 Extracting Noise and Darkness: Low-Light Image Enhancement via Dual Prior Guidance
abstract
The complex entanglement between darkness and noise hinders the advance of low-light image enhancement. Most existing methods adopted lightening-then-denoising or embedded a special denoising module into enhancement network without specific noise knowledge as supervision to restore low-light images. However, they either fail to remove the amplified noise or blur the detail information. Against above drawbacks, we propose a novel dual prior guidance method for low-light image enhancement that relights darkness and suppresses noise simultaneously. Concretely, the main novelties of our proposed method are three-fold. Firstly, our formulation originates from a statistic observation that darkness can be disentangled into luminance channel, yet noise still exists each channel when low-light images are transformed from RGB space to YCbCr space. It inspires us to design an ingenious method, extracting noise and darkness, termed END, to enhance low-light images. Secondly, we propose a prior extraction network with prior composition module to extract luminance and noise priors from different channels. Thirdly, an image enhancement network deployed with prior guidance module is proposed to progressively lighten the darkness and remove noise. Extensive experiments on multiple benchmarks demonstrate that our proposed method achieves remarkable performance compared to other state-of-the-art low-light image enhancement methods. The source code and trained model can be found inhttps://github.com/WHK-Huake/END.
Huake Wang, Xiaoyang Yan, Xingsong Hou, Kaibing Zhang, Yujie Dun
IEEE Trans. Circuits Syst. Video Technol.5
2024 Division gets better: Learning brightness-aware and detail-sensitive representations for low-light image enhancement
Huake Wang, Xiaoyang Yan, Xingsong Hou, Yujie Dun, Kaibing Zhang
Knowl. Based Syst.5
2024 HF-HRNet: A Simple Hardware Friendly High-Resolution Network
abstract
High-resolution networks have made significant progress in dense prediction tasks such as human pose estimation and semantic segmentation. To better explore this high-resolution mechanism on mobile devices, Lite-HRNet incorporates shuffle operations to reduce computational complexity in the channel dimension, while Dite-HRNet employs dynamic convolution and pooling to capture long-range interactions with low computational complexity in the spatial dimension. The core idea behind both approaches is to efficiently capture information in either the channel or spatial dimension. However, shuffle operations and dynamic operations are not hardware-friendly. As a result, both Lite-HRNet and Dite-HRNet cannot achieve the desired inference speed on specialized devices, including Neural Processing Units (NPUs) and Graphics Processing Units (GPUs). To overcome these limitations, we present a simple Hardware-Friendly Lightweight High-resolution Network (HF-HRNet) based on our proposed Hardware-Friendly Uniform-sized Mug (HUM) block. HUM block mainly consists of the Cascaded Depthwise (CAD) block and Multi-Scale Context Embedding (MCE) block. The CAD block cascades depthwise convolutions to obtain a larger receptive field in the spatial dimension, while the MCE block aggregates multi-scale spatial feature information from different scales and adjusts channel features. Extensive experiments are conducted on human pose estimation (COCO, MPII) and semantic segmentation (Cityscapes), resulting in a better trade-off between inference speed and accuracy on both NPUs and GPUs. It is noteworthy that on the COCO test-dev set, HF-HRNet-30 outperforms Dite-HRNet-30 and Lite-HRNet-30 by 1.9 AP and 2.8 AP, respectively, while running about 13 times faster and 9 times faster on NPUs, respectively. Our code are publicly available for use: https://github.com/zhanghao5201/HF-HRNet.
Hao Zhang 0117, Yujie Dun, Yixuan Pei, Shenqi Lai, Chengxu Liu 0001, Kaipeng Zhang, Xueming Qian
IEEE Trans. Circuits Syst. Video Technol.2
2023 CSDA: Learning Category-Scale Joint Feature for Domain Adaptive Object Detection
abstract
Domain Adaptive Object Detection (DAOD) aims to improve the detection performance of target domains by minimizing the feature distribution between the source and target domain. Recent approaches usually align such distributions in terms of categories through adversarial learning and some progress has been made. However, when objects are non-uniformly distributed at different scales, such category-level alignment causes imbalanced object feature learning, refer as the inconsistency of category alignment at different scales. For better category-level feature alignment, we propose a novel DAOD framework of joint category and scale information, dubbed CSDA, such a design enables effective object learning for different scales. Specifically, our framework is implemented by two closely-related modules: 1) SGFF (Scale-Guided Feature Fusion) fuses the category representations of different domains to learn category-specific features, where the features are aligned by discriminators at three scales. 2) SAFE (Scale-Auxiliary Feature Enhancement) encodes scale coordinates into a group of tokens and enhances the representation of category-specific features at different scales by self-attention. Based on the anchor-based Faster-RCNN and anchor-free FCOS detectors, experiments show that our method achieves state-of-the-art results on three DAOD benchmarks.
Changlong Gao, Chengxu Liu 0001, Yujie Dun, Xueming Qian
ICCV3
2023 Learning Data-Driven Vector-Quantized Degradation Model for Animation Video Super-Resolution
abstract
Existing real-world video super-resolution (VSR) methods focus on designing a general degradation pipeline for open-domain videos while ignoring data intrinsic characteristics which strongly limit their performance when applying to some specific domains (e.g., animation videos). In this paper, we thoroughly explore the characteristics of animation videos and leverage the rich priors in real-world animation data for a more practical animation VSR model. In particular, we propose a multi-scale Vector-Quantized Degradation model for animation video Super-Resolution (VQD-SR) to decompose the local details from global structures and transfer the degradation priors in real-world animation videos to a learned vector-quantized codebook for degradation modeling. A rich-content Real Animation Low-quality (RAL) video dataset is collected for extracting the priors. We further propose a data enhancement strategy for high-resolution (HR) training videos based on our observation that existing HR videos are mostly collected from the Web which contains conspicuous compression artifacts. The proposed strategy is valid to lift the upper bound of animation VSR performance, regardless of the specific VSR model. Experimental results demonstrate the superiority of the proposed VQD-SR over state-of-the-art methods, through extensive quantitative and qualitative evaluations of the latest animation video super-resolution benchmark. The code and pre-trained models can be downloaded at https://github.com/researchmm/VQD-SR.
Zixi Tuo, Huan Yang 0005, Jianlong Fu, Yujie Dun, Xueming Qian
ICCV4
2023 Anomaly detection framework for unmanned vending machines
Zongyang Da, Yujie Dun, Chengxu Liu 0001, Yuanzhi Liang, Xueming Qian
Knowl. Based Syst.2
2023 SCGNet: Shifting and Cascaded Group Network
abstract
Many lightweight networks have been proposed for resource-limited applications, however, they cannot be efficiently applied to neural-network processing units (NPUs) due to the limited operations supported by the NPUs, and few works focus on efficient network design on the NPUs. The basic blocks of networks such as MobileNetV2 and RegNet use smaller convolution kernels with relatively small receptive fields, which are not conducive to capturing large-scale spatial information. To address this weakness, we propose Shifting and Cascaded Group (SCG) block, where we cascade group convolutions with larger kernels to exploit multi-scale information and propose shifting group convolution to communicate channel information between different groups. Besides, we carefully devise our architecture guided by some principles and finally build a very efficient network called Shifting and Cascaded Group Network (SCGNet) on NPUs. To verify the superiority of our method, we conduct extensive experiments on various tasks including image classification, object detection, human pose estimation, person re-identification, and semantic segmentation to comprehensively evaluate the performance. Results on widely used datasets such as ImageNet, PASCAL VOC, COCO, MPII, Market-1501, DukeMTMC-ReID, CUHK03, and Cityscapes demonstrate that the proposed network is a more effective network on the corresponding vision tasks.
Hao Zhang 0117, Shenqi Lai, Yaxiong Wang, Zongyang Da, Yujie Dun, Xueming Qian
IEEE Trans. Circuits Syst. Video Technol.5
2022 Annular-Graph Attention Model for Personalized Sequential Recommendation
abstract
Sequential recommendations aim to predict the user’s next behaviors items based on their successive historical behaviors sequence. It has been widely applied in lots of online services. However, current sequential recommendations use the adjacent behaviors to capture the features of the sequence, ignoring the features among nonadjacent sequential items and the summarized features of the sequence. To address the above problems, in this paper, we propose an annular-graph attention based sequential recommendation (AGSR) model by exploring user’s long-term and short-term preferences for the personalized sequential recommendation. For user’s short-term preferences, AGSR builds an annular-graph on the sequence of user behavior. Then, AGSR proposes an annular-graph attention applying on the sub annular-graph to explore local features and applying annular-graph attention on entire annular-graph to explore the global features and the skip features. For user’s long-term preferences, the latent factor model are introduced in AGSR. The experimental results on two public datasets show that our model outperforms the state-of-the-art methods.
Junmei Hao, Yujie Dun, Guoshuai Zhao 0001, Yuxia Wu, Xueming Qian
IEEE Trans. Multim.2
2021 Image super-resolution based on residually dense distilled attention network
Yujie Dun, Zongyang Da, Xueming Qian
Neurocomputing1
2021 Kernel-attended residual network for single image super-resolution
Yujie Dun, Zongyang Da, Xueming Qian
Knowl. Based Syst.1
2021 Learning Deformable and Attentive Network for image restoration
Xingsong Hou, Yujie Dun, Jie Qin 0004, Li Liu 0004, Xueming Qian, Ling Shao 0001
Knowl. Based Syst.3
2015 A Fine-Resolution Frequency Estimator in the Odd-DFT Domain
abstract
Although many frequency estimation methods are available, few are designed for high-quality speech and audio processing systems, which typically use the modified discrete cosine transform (MDCT) as their analysis filter bank. In this letter, we propose a low complexity frequency estimator that is suitable for MDCT-based systems and that operates in the odd-DFT domain. Taking a complex exponential in noise as the input and deriving the analytical expression of its odd-DFT coefficient, we obtain an interpolated odd-DFT-based frequency estimator. Experiments show that the proposed estimator outperforms all other reported odd-DFT/MDCT domain estimators and has precision that is similar to that of representative DFT domain frequency estimators. The overhead for incorporating this estimator into a speech and audio processing system is small due to the simple odd-DFT to MDCT conversion. The corresponding magnitude and phase estimators are also proposed in this letter.
Yujie Dun, Guizhong Liu
IEEE Signal Process. Lett.1
2008 High Productivity Computing System Based on FPGA and Its Application on Plasma Simulation
abstract
For computational intensive applications, effective and high performance computing capacity is of the most important. Besides the large scale parallel computers or supercomputers, small calculating arrays based on DSPs or FPGAs have become a reasonable selection for relatively lower cost circumstance. In this paper, we have developed a high productivity computing system (HPCS) based on FPGA. In this system, the HPCS computing unit (HCU) board uses two pieces of Altera StratixII 2S60/90/180 and can communicate with an optional host PC via PCI/PCI-x interface. The IEEE-754 compatible double precision floating point arithmetic IP package was realized to meet the computational requirements. Device drivers for the HCU board were written for both Linux and Windows OS. The HPCS has been applied to accelerate plasma simulation code called TRISTAN. When simulating 512 k particles, the result shows that the HPCS performs about 20 times faster than a Pentium-IV 2.8 GHz PC with 1 GB DDR memory.
Yujie Dun, Weixiang Shi, Baogang Miao, Bingo Zhang
HPCC2
2007 Investigation of H.264 intra coding for SAR image
abstract
In this paper we investigate the performance of H.264 Intra Coding for Synthetic Aperture Radar(SAR) image. The results show that H.264 Intra Coding is a high performance coder for the SAR image. However, when the SAR image is despeckled, H.264 Intra coding is not so efficient than the wavelet based image coder, such as JPEG2000 and SPIHT. Then more efficient representation is needed when H.264 Intra Coding is used to code the despeckled SAR image.
Xingsong Hou, Yujie Dun, Rongjing Ji
IGARSS2