Wenyang Luo

dblp:325/1752 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0007-7228-868XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Generative modeling · 47% Language models and text generation · 14% Trustworthy machine learning · 14%
Computer graphics and multimedia
4 papers
Image and video processing · 58% Multimedia analysis and retrieval · 38% Image and video coding · 4%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing
image restoration
1.722025
Reversing Flow for Image Restoration · CVPR 2025
Visual-Instructed Degradation Diffusion for All-in-One Image Restoration · CVPR 2025
Machine learning › Generative modeling › normalizing flow
continuous normalizing flow
0.912025
Reversing Flow for Image Restoration · CVPR 2025
Machine learning › Generative modeling
diffusion model
0.912025
Visual-Instructed Degradation Diffusion for All-in-One Image Restoration · CVPR 2025
Machine learning › Generative modeling
normalizing flow
0.912025
Reversing Flow for Image Restoration · CVPR 2025
Image and video processing › image restoration › multi-task image restoration
all-in-one image restoration
0.912025
Visual-Instructed Degradation Diffusion for All-in-One Image Restoration · CVPR 2025
Multimedia analysis and retrieval › near-duplicate detection
video copy detection
0.912025
FiGVCL: Fine-Grained Benchmark and Method for Video Copy Localization · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Multimedia analysis and retrieval
video retrieval
0.912025
FiGVCL: Fine-Grained Benchmark and Method for Video Copy Localization · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Natural language and speech › Language models and text generation
multilingual language models
0.812024
Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models · ACL (1) 2024
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
neuron analysis
0.812024
Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models · ACL (1) 2024
Computer vision › Video understanding and tracking › video classification
few-shot video classification
0.612022
Long-Short Term Cross-Transformer in Compressed Domain for Few-Shot Video Classification · IJCAI 2022
Machine learning › Deep learning architectures and training
transformer
0.612022
Long-Short Term Cross-Transformer in Compressed Domain for Few-Shot Video Classification · IJCAI 2022
Computer vision › Vision and language › vision-language generation
instruction-guided generation
0.312025
Visual-Instructed Degradation Diffusion for All-in-One Image Restoration · CVPR 2025

Methods — techniques the papers use, named apart from their topics

visual instruction · 1.7velocity field matching · 1.7entropy-preserving flow · 1.7diffusion model · 1.7degradation space denoising · 1.7group of pictures · 1.1cross-attention · 1.1unsupervised training · 0.9local embedding · 0.9continuous normalizing flows · 0.9continuous normalizing flow · 0.9neuron analysis · 0.8motion vector · 0.6
YearPublicationVenuePosition
2025 Visual-Instructed Degradation Diffusion for All-in-One Image Restoration
abstract
Image restoration tasks like deblurring, denoising, and dehazing usually need distinct models for each degradation type, restricting their generalization in real-world scenarios with mixed or unknown degradations. In this work, we propose Defusion, a novel all-in-one image restoration framework that utilizes visual instruction-guided degradation diffusion. Unlike existing methods that rely on task-specific models or ambiguous text-based priors, Defusion constructs explicit visual instructions that align with the visual degradation patterns. These instructions are grounded by applying degradations to standardized visual elements, capturing intrinsic degradation features while agnostic to image semantics. Defusion then uses these visual instructions to guide a diffusion-based model that operates directly in the degradation space, where it reconstructs high-quality images by denoising the degradation effects with enhanced stability and generalizability. Comprehensive experiments demonstrate that Defusion outperforms state-of-the-art methods across diverse image restoration tasks, including complex and real-world degradations.
Wenyang Luo, Haina Qin, Zewen Chen, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004
CVPR1
2025 Reversing Flow for Image Restoration
abstract
Image restoration aims to recover high-quality (HQ) images from degraded low-quality (LQ) ones by reversing the effects of degradation. Existing generative models for image restoration, including diffusion and score-based models, often treat the degradation process as a stochastic transformation, which introduces inefficiency and complexity. In this work, we propose ResFlow, a novel image restoration framework that models the degradation process as a deterministic path using continuous normalizing flows. ResFlow augments the degradation process with an auxiliary process that disambiguates the uncertainty in HQ prediction to enable reversible modeling of the degradation process. ResFlow adopts entropy-preserving flow paths and learns the augmented degradation flow by matching the velocity field. ResFlow significantly improves the performance and speed of image restoration, completing the task in fewer than four sampling steps. Extensive experiments demonstrate that ResFlow achieves state-of-the-art results across various image restoration benchmarks, offering a practical and efficient solution for real-world applications.
Haina Qin, Wenyang Luo, Jingdong Chen, Ming Yang 0007, Bing Li 0001, Weiming Hu 0004
CVPR2
2025 Two-stream transformer tracking with messengers
Miaobo Qiu, Wenyang Luo, Tongfei Liu, Yanqin Jiang, Jiaming Yan, Weiming Hu 0004, Stephen J. Maybank
Image Vis. Comput.2
2025 FiGVCL: Fine-Grained Benchmark and Method for Video Copy Localization
abstract
Content-based video copy localization (VCL) aims to detect and locate copied segments in pairs of videos. VCL requires fine-grained video analysis to robustly identify copied segments that have been edited. Despite recent progress, the prohibitive cost of annotating copied segments and the lack of a fine-grained benchmark hinder the development of effective VCL systems. In this work, we annotate a new real-world dataset, FiGVCL, with challenging scenarios designed to evaluate VCL methods. FiGVCL is carefully annotated to preserve the temporal correspondences observed in copied segments. Moreover, we propose a novel fine-grained VCL benchmark metric based on temporal correspondences to improve discriminability. Finally, we design a simple but effective baseline model that uses fine-grained local embeddings for accurate copied segment localization. We also present an unsupervised training strategy that outperforms previous supervised VCL methods.
Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 iESTA: Instance-Enhanced Spatial-Temporal Alignment for Video Copy Localization
abstract
Video copy Segment Localization (VSL) requires the identification of the temporal segments within a pair of videos that contain copied content. Current methods primarily focus on global temporal modeling, overlooking the complementarity of global semantic and local fine-grained features, which limits their effectiveness. Some related methods attempt to incorporate local spatial information but often disrupt spatial semantic structures, resulting in less accurate matching. To address these issues, we propose the Instance-Enhanced Spatial-Temporal Alignment Framework (iESTA), based on a proper representation granularity that integrates instance-level local features and semantic global features. Specifically, the Instance-relation Graph (IRG) is constructed to capture instance-level features and fine-grained interactions, preserving local information integrity and better representing the video feature space in a proper granularity. An instance-GNN structure is designed to refine these graph representations. For global features, we enhance the representation of semantic information, capturing temporal relationships within videos using a Transformer framework. Additionally, we design a Complementarity-perception Alignment Module (CAM) to effectively process and integrate complementary spatial-temporal information, producing accurate frame-to-frame alignment maps. Our approach also incorporates a differentiable Dynamic Time Warping (DTW) method to utilize latent temporal alignments as weak supervisory signals, improving the accuracy of the matching process. Experimental results indicate that our proposed iESTA outperforms state-of-the-art methods on both the small-scale dataset VCDB and the large-scale dataset VCSL.
Xinmiao Ding, Jinming Lou, Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004
IEEE Trans. Circuits Syst. Video Technol.3
2025 Task-Aware Attentional Dynamic Alignment for Few-Shot Compressed Video Classification
abstract
We present a novel Task-aware Attentional Dynamic Alignment (TADA) framework for visual-based few-shot video classification (FSVC) that addresses two key challenges in this field: efficiency and nuanced spatio-temporal reasoning. Existing methods are often hindered by computationally expensive video decoding processes and neglect the temporal order of videos. In contrast, our method harnesses compressed domain data to extract rich spatio-temporal cues at a fraction of the cost of traditional video processing methods. Specifically, we propose an embedding module to extract informative features from compressed domain data while minimizing computational overheads. Furthermore, to exploit the temporal order of frames, we develop a prototypical ADA module to align and classify videos with an explicit temporal order constraint. Our framework also incorporates a contextual mixer to enrich video embeddings with task-specific context. Extensive experiments on multiple datasets demonstrate that TADA achieves state-of-the-art performance and outperforms existing methods in accuracy and efficiency.
Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank
IEEE Trans. Circuits Syst. Video Technol.1
2024 Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models
abstract
Tianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, Ji-Rong Wen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wenyang Luo, Haoyang Huang, Dongdong Zhang 0001, Xiaolei Wang 0005, Wayne Xin Zhao, Furu Wei, Ji-Rong Wen
ACL (1)2
2022 Long-Short Term Cross-Transformer in Compressed Domain for Few-Shot Video Classification
abstract
Compared with image few-shot learning, most of the existing few-shot video classification methods perform worse on feature matching, because they fail to sufficiently exploit the temporal information and relation. Specifically, frames are usually evenly sampled, which may miss important frames. On the other hand, the heuristic model simply encodes the equally treated frames in sequence, which results in the lack of both long-term and short-term temporal modeling and interaction. To alleviate these limitations, we take advantage of the compressed domain knowledge and propose a long-short term Cross-Transformer (LSTC) for few-shot video classification. For short terms, the motion vector (MV) contains temporal cues and reflects the importance of each frame. For long terms, a video can be natively divided into a sequence of GOPs (Group Of Picture). Using this compressed domain knowledge helps to obtain a more accurate spatial-temporal feature space. Consequently, we design the long-short term selection module, short-term module, and long-term module to comprise the LSTC. Long-short term selection is performed to select informative compressed domain data. Long/short-term modules are utilized to sufficiently exploit the temporal information so that the query and support can be well-matched by cross-attention. Experimental results show the superiority of our method on various datasets.
Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Yanan Miao, Yangxi Li
IJCAI1