Miaoyu Li

dblp:304/5218 · DBLP profile ↗
← Back
10ranked-venue papers
8as first author
10since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Two Causally Related Needles in a Video Haystack
abstract
Properly evaluating the ability of Video-Language Models (VLMs) to understand long videos remains a challenge. We propose a long-context video understanding benchmark, Causal2Needles, that assesses two crucial abilities insufficiently addressed by existing benchmarks: (1) extracting information from two separate locations (two needles) in a long video and understanding them jointly, and (2) modeling the world in terms of cause and effect in human behaviors. Causal2Needles evaluates these abilities using noncausal one-needle, causal one-needle, and causal two-needle questions. The most complex question type, causal two-needle questions, require extracting information from both the cause and effect events from a long video and the associated narration text. To prevent textual bias, we introduce two complementary question formats: locating the video clip containing the answer, and verbal description of a visual detail from that video clip. Our experiments reveal that models excelling on existing benchmarks struggle with causal 2-needle questions, and the model performance is negatively correlated with the distance between the two needles. These findings highlight critical limitations in current VLMs.
Miaoyu Li, Qin Chao, Boyang Li 0001
NeurIPS1
2025 Latent Diffusion Enhanced Rectangle Transformer for Hyperspectral Image Restoration
abstract
The restoration of hyperspectral image (HSI) plays a pivotal role in subsequent hyperspectral image applications. Despite the remarkable capabilities of deep learning, current HSI restoration methods face challenges in effectively exploring the spatial non-local self-similarity and spectral low-rank property inherently embedded with HSIs. This paper addresses these challenges by introducing a latent diffusion enhanced rectangle Transformer for HSI restoration, tackling the non-local spatial similarity and HSI-specific latent diffusion low-rank property. In order to effectively capture non-local spatial similarity, we propose the multi-shape spatial rectangle self-attention module in both horizontal and vertical directions, enabling the model to utilize informative spatial regions for HSI restoration. Meanwhile, we propose a spectral latent diffusion enhancement module that generates the image-specific latent dictionary based on the content of HSI for low-rank vector extraction and representation. This module utilizes a diffusion model to generatively obtain representations of global low-rank vectors, thereby aligning more closely with the desired HSI. A series of comprehensive experiments were carried out on four common hyperspectral image restoration tasks, including HSI denoising, HSI super-resolution, HSI reconstruction, and HSI inpainting. The results of these experiments highlight the effectiveness of our proposed method, as demonstrated by improvements in both objective metrics and subjective visual quality.
Miaoyu Li, Ying Fu 0001, Tao Zhang 0042, Ji Liu 0003, Dejing Dou, Chenggang Yan 0001, Yulun Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Supervise-Assisted Self-Supervised Deep-Learning Method for Hyperspectral Image Restoration
abstract
Hyperspectral image (HSI) restoration is a challenging research area, covering a variety of inverse problems. Previous works have shown the great success of deep learning in HSI restoration. However, facing the problem of distribution gaps between training HSIs and target HSI, those data-driven methods falter in delivering satisfactory outcomes for the target HSIs. In addition, the degradation process of HSIs is usually disturbed by noise, which is not well taken into account in existing restoration methods. The existence of noise further exacerbates the dissimilarities within the data, rendering it challenging to attain desirable results without an appropriate learning approach. To track these issues, in this article, we propose a supervise-assisted self-supervised deep-learning method to restore noisy degraded HSIs. Initially, we facilitate the restoration network to acquire a generalized prior through supervised learning from extensive training datasets. Then, the self-supervised learning stage is employed and utilizes the specific prior of the target HSI. Particularly, to restore clean HSIs during the self-supervised learning stage from noisy degraded HSIs, we introduce a noise-adaptive loss function that leverages inner statistics of noisy degraded HSIs for restoration. The proposed noise-adaptive loss consists of Stein's unbiased risk estimator (SURE) and total variation (TV) regularizer and fine-tunes the network with the presence of noise. We demonstrate through experiments on different HSI tasks, including denoising, compressive sensing, super-resolution, and inpainting, that our method outperforms state-of-the-art methods on benchmarks under quantitative metrics and visual quality. The code is available at https://github.com/ying-fu/SSDL-HSI.
Miaoyu Li, Ying Fu 0001, Tao Zhang 0042, Guanghui Wen
IEEE Trans. Neural Networks Learn. Syst.1
2023 Spatial-Spectral Transformer for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising is a crucial preprocessing procedure for the subsequent HSI applications. Unfortunately, though witnessing the development of deep learning in HSI denoising area, existing convolution-based methods face the trade-off between computational efficiency and capability to model non-local characteristics of HSI. In this paper, we propose a Spatial-Spectral Transformer (SST) to alleviate this problem. To fully explore intrinsic similarity characteristics in both spatial dimension and spectral dimension, we conduct non-local spatial self-attention and global spectral self-attention with Transformer architecture. The window-based spatial self-attention focuses on the spatial similarity beyond the neighboring region. While, the spectral self-attention exploits the long-range dependencies between highly correlative bands. Experimental results show that our proposed method outperforms the state-of-the-art HSI denoising methods in quantitative quality and visual results. The code is released at https://github.com/MyuLi/SST.
Miaoyu Li, Ying Fu 0001, Yulun Zhang 0001
AAAI1
2023 Spectral Enhanced Rectangle Transformer for Hyperspectral Image Denoising
abstract
Denoising is a crucial step for hyperspectral image (HSI) applications. Though witnessing the great power of deep learning, existing HSI denoising methods suffer from limitations in capturing the nonlocal self-similarity. Trans-formers have shown potential in capturing longrange de-pendencies, but few attempts have been made with specifically designed Transformer to model the spatial and spec-tral correlation in HSIs. In this paper, we address these issues by proposing a spectral enhanced rectangle Trans-former, driving it to explore the nonlocal spatial similarity and global spectral low-rank property of HSIs. For the former, we exploit the rectangle self-attention horizontally and vertically to capture the nonlocal similarity in the spatial domain. For the latter, we design a spectral enhancement module that is capable of extracting global underlying low-rank property of spatial-spectral cubes to suppress noise, while enabling the interactions among non-overlapping spatial rectangles. Extensive experiments have been conducted on both synthetic noisy HSIs and real noisy HSIs, showing the effectiveness of our proposed method in terms of both objective metric and subjective visual quality. The code is available at https://github.com/MyuLi/SERT.
Miaoyu Li, Ji Liu 0003, Ying Fu 0001, Yulun Zhang 0001, Dejing Dou
CVPR1
2023 Pixel Adaptive Deep Unfolding Transformer for Hyperspectral Image Reconstruction
abstract
Hyperspectral Image (HSI) reconstruction has made gratifying progress with the deep unfolding framework by formulating the problem into a data module and a prior module. Nevertheless, existing methods still face the problem of insufficient matching with HSI data. The issues lie in three aspects: 1) fixed gradient descent step in the data module while the degradation of HSI is agnostic in the pixel-level. 2) inadequate prior module for 3D HSI cube. 3) stage interaction ignoring the differences in features at different stages. To address these issues, in this work, we propose a Pixel Adaptive Deep Unfolding Transformer (PADUT) for HSI reconstruction. In the data module, a pixel adaptive descent step is employed to focus on pixel-level agnostic degradation. In the prior module, we introduce the Non-local Spectral Transformer (NST) to emphasize the 3D characteristics of HSI for recovering. Moreover, inspired by the diverse expression of features in different stages and depths, the stage interaction is improved by the Fast Fourier Transform (FFT). Experimental results on both simulated and real scenes exhibit the superior performance of our method compared to state-of-the-art HSI reconstruction methods. The code is released at: https://github.com/MyuLi/PADUT
Miaoyu Li, Ying Fu 0001, Ji Liu 0003, Yulun Zhang 0001
ICCV1
2023 BEV-DG: Cross-Modal Learning under Bird's-Eye View for Domain Generalization of 3D Semantic Segmentation
abstract
Cross-modal Unsupervised Domain Adaptation aims to exploit the complementarity of 2D-3D data to overcome the lack of annotation in an unknown domain. However, the training of these methods relies on access to target samples, meaning the trained model only works in a specific target domain. In light of this, we propose cross-modal learning under bird’s-eye view for Domain Generalization (DG) of 3D semantic segmentation, called BEV-DG. DG is more challenging because the model cannot access the target domain during training, meaning it needs to rely on cross-modal learning to alleviate the domain gap. Since 3D semantic segmentation requires the classification of each point, existing cross-modal learning is directly conducted point-to-point, which is sensitive to the misalignment in projections between pixels and points. To this end, our approach aims to optimize domain-irrelevant representation modeling with the aid of cross-modal learning under bird’s-eye view. We propose BEV-based Area-to-area Fusion (BAF) to conduct cross-modal learning under bird’s-eye view, which has a higher fault tolerance for point-level misalignment. Furthermore, to model domain-irrelevant representations, we propose BEV-driven Domain Contrastive Learning (BDCL) with the help of cross-modal learning under bird’s-eye view. We design three domain generalization settings based on three 3D datasets, and BEV-DG significantly outperforms state-of-the-art competitors with tremendous margins in all settings.
Miaoyu Li, Yachao Zhang 0001, Xu Ma 0005, Yanyun Qu, Yun Fu 0001
ICCV1
2022 Cross-Domain and Cross-Modal Knowledge Distillation in Domain Adaptation for 3D Semantic Segmentation
abstract
With the emergence of multi-modal datasets where LiDAR and camera are synchronized and calibrated, cross-modal Unsupervised Domain Adaptation (UDA) has attracted increasing attention because it reduces the laborious annotation of target domain samples. To alleviate the distribution gap between source and target domains, existing methods conduct feature alignment by using adversarial learning. However, it is well-known to be highly sensitive to hyperparameters and difficult to train. In this paper, we propose a novel model (Dual-Cross) that integrates Cross-Domain Knowledge Distillation (CDKD) and Cross-Modal Knowledge Distillation (CMKD) to mitigate domain shift. Specifically, we design the multi-modal style transfer to convert source image and point cloud to target style. With these synthetic samples as input, we introduce a target-aware teacher network to learn knowledge of the target domain. Then we present dual-cross knowledge distillation when the student is learning on source domain. CDKD constrains teacher and student predictions under same modality to be consistent. It can transfer target-aware knowledge from the teacher to the student, making the student more adaptive to the target domain. CMKD generates hybrid-modal prediction from the teacher predictions and constrains it to be consistent with both 2D and 3D student predictions. It promotes the information interaction between two modalities to make them complement each other. From the evaluation results on various domain adaptation settings, Dual-Cross significantly outperforms both uni-modal and cross-modal state-of-the-art methods.
Miaoyu Li, Yachao Zhang 0001, Yuan Xie 0006, Zuodong Gao, Cuihua Li, Zhizhong Zhang 0001, Yanyun Qu
ACM Multimedia1
2022 Self-supervised Exclusive Learning for 3D Segmentation with Cross-Modal Unsupervised Domain Adaptation
abstract
2D-3D unsupervised domain adaptation (UDA) tackles the lack of annotations in a new domain by capitalizing the relationship between 2D and 3D data. Existing methods achieve considerable improvements by performing cross-modality alignment in a modality-agnostic way, failing to exploit modality-specific characteristic for modeling complementarity. In this paper, we present self-supervised exclusive learning for cross-modal semantic segmentation under the UDA scenario, which avoids the prohibitive annotation. Specifically, two self-supervised tasks are designed, named "plane-to-spatial'' and "discrete-to-textured''. The former helps the 2D network branch improve the perception of spatial metrics, and the latter supplements structured texture information for the 3D network branch. In this way, modality-specific exclusive information can be effectively learned, and the complementarity of multi-modality is strengthened, resulting in a robust network to different domains. With the help of the self-supervised tasks supervision, we introduce a mixed domain to enhance the perception of the target domain by mixing the patches of the source and target domain samples. Besides, we propose a domain-category adversarial learning with category-wise discriminators by constructing the category prototypes for learning domain-invariant features. We evaluate our method on various multi-modality domain adaptation settings, where our results significantly outperform both uni-modality and multi-modality state-of-the-art competitors.
Yachao Zhang 0001, Miaoyu Li, Yuan Xie 0006, Cuihua Li, Cong Wang 0039, Zhizhong Zhang 0001, Yanyun Qu
ACM Multimedia2
2022 HeadSee: Device-free head gesture recognition with commodity RFID
Miaoyu Li, Baoying Liu, Feng Chen 0002
Peer-to-Peer Netw. Appl.3