Axi Niu

dblp:283/5444 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0001-5238-9917ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021
YearPublicationVenuePosition
2026 History-aware adaptive teacher for cross-domain object detection
Yaoqi Hu, Axi Niu, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001
Expert Syst. Appl.4
2026 Inter-view dual-domain guided stable diffusion for real-world stereo image super-resolution
Yu Zhu 0004, Axi Niu, Jinqiu Sun, Yanning Zhang 0001
Expert Syst. Appl.3
2026 FSCFNet: Lightweight neural networks via multi-dimensional importance-aware optimization
Mengyang Nie, Jinqiu Sun, Hongsong Guoyang, Axi Niu, Yaoqi Hu, Qingsen Yan, Yu Zhu 0004
Neurocomputing4
2025 Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
abstract
We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike prior approaches that rely on classifier-based or classifier-free guidance, MGAudio enables the generative model to guide itself through a dedicated training objective designed for video-conditioned audio generation. The framework integrates three main components: (1) a scalable flow-based Transformer denoiser, (2) a dual-role alignment mechanism where the audio-visual encoder serves both as a conditioning module and as a feature aligner to improve generation quality, and (3) a model-guided objective that enhances cross-modal coherence and audio realism. MGAudio achieves state-of-the-art performance on VGGSound, reducing FAD to 0.40, substantially surpassing the best classifier-free guidance baselines, and consistently outperforms existing methods across FD, IS, and alignment metrics. It also generalizes well to the challenging UnAV-100 benchmark. These results highlight model-guided dual-role alignment as a powerful and scalable paradigm for conditional video-to-audio generation. Code is available at: https://github.com/pantheon5100/mgaudio
Kang Zhang 0008, Trung X. Pham, Suyeon Lee, Axi Niu, Arda Senocak, Joon Son Chung
NeurIPS4
2025 Enhancing the noise robustness of sparse-form patches for image denoising
Liping Qi, Yu Zhu 0004, Wei Sun 0036, Axi Niu, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001
Knowl. Based Syst.5
2025 A multi-scale feature cross-dimensional interaction network for stereo image super-resolution
Yu Zhu 0004, Shengjun Peng, Axi Niu, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001
Multim. Syst.4
2025 Modeling optical imaging pipeline and learning contrastive-based representation for hybrid-corrupted image restoration
Chenyuan Zhao, Yu Zhu 0004, Qingsen Yan, Jinqiu Sun, Axi Niu, Yanning Zhang 0001
Multim. Syst.5
2025 Learning From Multi-Perception Features for Real-Word Image Super-Resolution
abstract
Actual image super-resolution is an extremely challenging task due to complex degradations existing in the image. To solve this problem, two dominant methodologies have emerged: degradation-estimation-based Addressing actual image super-resolution remains a formidable challenge due to the intricate degradations present in images. Two primary methodologies have emerged: degradation-estimation-based and blind-based methods. The former often struggle to accurately estimate degradation, limiting their effectiveness on real low-resolution images. Conversely, blind-based methods rely on a single perceptual perspective, constraining their adaptability to diverse perceptual characteristics. In response to these challenges, we present MPF-Net, a novel super-resolution approach aimed at enhancing real-world image super-resolution tasks by enabling the model to learn multiple perceptual features from input images. Our method features a Multi-Perception Feature Extraction module (MPFE) designed to extract diverse perceptual details, complemented by Cross-Perception Blocks (CPB) facilitating the fusion of this information for efficient super-resolution reconstruction. Additionally, we introduce a contrastive regularization term (CR) to enhance the model’s learning by leveraging newly generated HR and LR images as positive and negative samples. Experimental results on challenging real-world SR datasets demonstrate the superiority of our approach over existing state-of-the-art methods, both qualitatively and quantitatively.
Axi Niu, Kang Zhang 0008, Trung X. Pham, Jinqiu Sun, In-So Kweon, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Multiple Object Tracking Based on Occlusion-Aware Embedding Consistency Learning
abstract
The Joint Detection and Embedding (JDE) framework has achieved remarkable progress for multiple object tracking. Existing methods often employ extracted embeddings to re-establish associations between new detections and previously disrupted tracks. However, the reliability of embeddings diminishes when the region of the occluded object frequently contains adjacent objects or clutters, especially in scenarios with severe occlusion. To alleviate this problem, we propose a novel multiple object tracking method based on visual embedding consistency, mainly including: 1) Occlusion Prediction Module (OPM) and 2) Occlusion-Aware Association Module (OAAM). The OPM predicts occlusion information for each true detection, facilitating the selection of valid samples for consistency learning of the track’s visual embedding. The OAAM leverages occlusion cues and visual embeddings to generate two separate embeddings for each track, guaranteeing consistency in both unoccluded and occluded detections. By integrating these two modules, our method is capable of addressing track interruptions caused by occlusion in online tracking scenarios. Extensive experimental results demonstrate that our approach achieves promising performance levels in both unoccluded and occluded tracking scenarios.
Yaoqi Hu, Axi Niu, Yu Zhu 0004, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001
ICASSP2
2024 Dynamic center point learning for multiple object tracking under Severe occlusions
Yaoqi Hu, Axi Niu, Jinqiu Sun, Yu Zhu 0004, Qingsen Yan, Wei Dong 0010, Marcin Wozniak, Yanning Zhang 0001
Knowl. Based Syst.2
2024 GRAN: ghost residual attention network for single image super resolution
Axi Niu, Yu Zhu 0004, Jinqiu Sun, Qingsen Yan, Yanning Zhang 0001
Multim. Tools Appl.1
2024 KGSR: A kernel guided network for real-world blind super-resolution
Qingsen Yan, Axi Niu, Wei Dong 0010, Marcin Wozniak, Yanning Zhang 0001
Pattern Recognit.2
2023 CDPMSR: Conditional Diffusion Probabilistic Models for Single Image Super-Resolution
abstract
Diffusion probabilistic models (DPM) have been widely adopted in image-to-image translation to generate high-quality images. Prior attempts at applying the DPM to image super-resolution (SR) have shown that iteratively refining a pure Gaussian noise with a conditional image using a U-Net trained on denoising at various-level noises can help obtain a satisfied high-resolution image for the low-resolution one. To further improve the performance and simplify current DPM-based super-resolution methods, we propose a simple but non-trivial DPM-based super-resolution post-process framework, i.e., cDPMSR. After applying a pre-trained SR model on the to-be-test LR image to provide the conditional input, we adapt the standard DPM to conduct conditional image generation and perform super-resolution through a deterministic iterative denoising process. Our method surpasses prior attempts on both qualitative and quantitative results and can generate more photo-realistic counterparts for the low-resolution images with various benchmark datasets including Set5, Set14, Urban100, BSD100, and Manga109. Code will be published after accepted.
Axi Niu, Kang Zhang 0008, Trung X. Pham, Jinqiu Sun, Yu Zhu 0004, In-So Kweon, Yanning Zhang 0001
ICIP1
2022 Dual Temperature Helps Contrastive Learning Without Many Negative Samples: Towards Understanding and Simplifying MoCo
abstract
Contrastive learning (CL) is widely known to require many negative samples, 65536 in MoCo for instance, for which the performance of a dictionary-free framework is often inferior because the negative sample size (NSS) is limited by its mini-batch size (MBS). To decouple the NSS from the MBS, a dynamic dictionary has been adopted in a large volume of CL frameworks, among which arguably the most popular one is MoCo family. In essence, MoCo adopts a momentum-based queue dictionary, for which we perform a fine-grained analysis of its size and consistency. We point out that InfoNCE loss used in MoCo implicitly attract anchors to their corresponding positive sample with various strength of penalties and identify such inter-anchor hardness-awareness property as a major reason for the necessity of a large dictionary. Our findings motivate us to simplify MoCo v2 via the removal of its dictionary as well as momentum. Based on an InfoNCE with the proposed dual temperature, our simplified frameworks, Sim-MoCo and SimCo, outperform MoCo v2 by a visible margin. Moreover, our work bridges the gap between CL and non-CL frameworks, contributing to a more unified under-standing of these two mainstream frameworks in SSL. Code is available at: https://bit.ly/3LkQbaT.
Chaoning Zhang, Kang Zhang 0008, Trung X. Pham, Axi Niu, Zhinan Qiao, Chang Dong Yoo, In-So Kweon
CVPR4
2022 Decoupled Adversarial Contrastive Learning for Self-supervised Adversarial Robustness
Chaoning Zhang, Kang Zhang 0008, Chenshuang Zhang, Axi Niu, Jiu Feng, Chang Dong Yoo, In-So Kweon
ECCV (30)4
2022 MS2Net: Multi-Scale and Multi-Stage Feature Fusion for Blurred Image Super-Resolution
abstract
At present, most mainstream algorithms for single image super-resolution (SISR) assume the image degradation process as an ideal degradation process (e.g. bicubic downscaling), which violates the actual degeneration conditions. In real-world image capturing, objects often move in a dynamic environment, and camera shake also often occurs, which results in serious blurs. Our work focuses on the task of image super-resolution with heavy motion blur, for which we adopt a network with two branches: one branch for image deblurring and the other one for super-resolution. Since the features obtained by the deblurring are rich in details, we apply their features as supplementary information to the super-resolution branch. Based on the adopted dual-branch framework, our major technical novelties lie in two novel modules: Multi-Scale Feature Fusion (MSFF1) module which fuses features of different scale from the deblurring branch to get local and global information, and Multi-Stage Feature Fusion (MSFF2) module which further filters useful information with attention. We evaluate the proposed method under various blur scenarios on the benchmark datasets, demonstrating competitive performance against existing methods.
Axi Niu, Yu Zhu 0004, Chaoning Zhang, Jinqiu Sun, In-So Kweon, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 Non-uniform motion deblurring with blurry component divided guidance
Wei Sun 0036, Qingsen Yan, Axi Niu, Rui Li 0013, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001
Pattern Recognit.4