VLDB 2026 Research / reviewers in the wild / expert
Wenyang Luo
dblp:325/1752
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0007-7228-868XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 47% Language models and text generation · 14% Trustworthy machine learning · 14% | |
| Computer graphics and multimedia
4 papers |
Image and video processing · 58% Multimedia analysis and retrieval · 38% Image and video coding · 4% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing
image restoration |
1.7 | 2 | 2025 | Reversing Flow for Image Restoration · CVPR 2025 Visual-Instructed Degradation Diffusion for All-in-One Image Restoration · CVPR 2025 |
Machine learning › Generative modeling › normalizing flow
continuous normalizing flow |
0.9 | 1 | 2025 | Reversing Flow for Image Restoration · CVPR 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Visual-Instructed Degradation Diffusion for All-in-One Image Restoration · CVPR 2025 |
Machine learning › Generative modeling
normalizing flow |
0.9 | 1 | 2025 | Reversing Flow for Image Restoration · CVPR 2025 |
Image and video processing › image restoration › multi-task image restoration
all-in-one image restoration |
0.9 | 1 | 2025 | Visual-Instructed Degradation Diffusion for All-in-One Image Restoration · CVPR 2025 |
Multimedia analysis and retrieval › near-duplicate detection
video copy detection |
0.9 | 1 | 2025 | FiGVCL: Fine-Grained Benchmark and Method for Video Copy Localization · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Multimedia analysis and retrieval
video retrieval |
0.9 | 1 | 2025 | FiGVCL: Fine-Grained Benchmark and Method for Video Copy Localization · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Natural language and speech › Language models and text generation
multilingual language models |
0.8 | 1 | 2024 | Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models · ACL (1) 2024 |
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
neuron analysis |
0.8 | 1 | 2024 | Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models · ACL (1) 2024 |
Computer vision › Video understanding and tracking › video classification
few-shot video classification |
0.6 | 1 | 2022 | Long-Short Term Cross-Transformer in Compressed Domain for Few-Shot Video Classification · IJCAI 2022 |
Machine learning › Deep learning architectures and training
transformer |
0.6 | 1 | 2022 | Long-Short Term Cross-Transformer in Compressed Domain for Few-Shot Video Classification · IJCAI 2022 |
Computer vision › Vision and language › vision-language generation
instruction-guided generation |
0.3 | 1 | 2025 | Visual-Instructed Degradation Diffusion for All-in-One Image Restoration · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
visual instruction · 1.7velocity field matching · 1.7entropy-preserving flow · 1.7diffusion model · 1.7degradation space denoising · 1.7group of pictures · 1.1cross-attention · 1.1unsupervised training · 0.9local embedding · 0.9continuous normalizing flows · 0.9continuous normalizing flow · 0.9neuron analysis · 0.8motion vector · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Visual-Instructed Degradation Diffusion for All-in-One Image RestorationabstractImage restoration tasks like deblurring, denoising, and dehazing usually need distinct models for each degradation type, restricting their generalization in real-world scenarios with mixed or unknown degradations. In this work, we propose Defusion, a novel all-in-one image restoration framework that utilizes visual instruction-guided degradation diffusion. Unlike existing methods that rely on task-specific models or ambiguous text-based priors, Defusion constructs explicit visual instructions that align with the visual degradation patterns. These instructions are grounded by applying degradations to standardized visual elements, capturing intrinsic degradation features while agnostic to image semantics. Defusion then uses these visual instructions to guide a diffusion-based model that operates directly in the degradation space, where it reconstructs high-quality images by denoising the degradation effects with enhanced stability and generalizability. Comprehensive experiments demonstrate that Defusion outperforms state-of-the-art methods across diverse image restoration tasks, including complex and real-world degradations. Wenyang Luo, Haina Qin, Zewen Chen, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
CVPR | 1 |
| 2025 | Reversing Flow for Image RestorationabstractImage restoration aims to recover high-quality (HQ) images from degraded low-quality (LQ) ones by reversing the effects of degradation. Existing generative models for image restoration, including diffusion and score-based models, often treat the degradation process as a stochastic transformation, which introduces inefficiency and complexity. In this work, we propose ResFlow, a novel image restoration framework that models the degradation process as a deterministic path using continuous normalizing flows. ResFlow augments the degradation process with an auxiliary process that disambiguates the uncertainty in HQ prediction to enable reversible modeling of the degradation process. ResFlow adopts entropy-preserving flow paths and learns the augmented degradation flow by matching the velocity field. ResFlow significantly improves the performance and speed of image restoration, completing the task in fewer than four sampling steps. Extensive experiments demonstrate that ResFlow achieves state-of-the-art results across various image restoration benchmarks, offering a practical and efficient solution for real-world applications. Haina Qin, Wenyang Luo, Jingdong Chen, Ming Yang 0007, Bing Li 0001, Weiming Hu 0004 |
CVPR | 2 |
| 2025 | Two-stream transformer tracking with messengers
Miaobo Qiu, Wenyang Luo, Tongfei Liu, Yanqin Jiang, Jiaming Yan, Weiming Hu 0004, Stephen J. Maybank |
Image Vis. Comput. | 2 |
| 2025 | FiGVCL: Fine-Grained Benchmark and Method for Video Copy LocalizationabstractContent-based video copy localization (VCL) aims to detect and locate copied segments in pairs of videos. VCL requires fine-grained video analysis to robustly identify copied segments that have been edited. Despite recent progress, the prohibitive cost of annotating copied segments and the lack of a fine-grained benchmark hinder the development of effective VCL systems. In this work, we annotate a new real-world dataset, FiGVCL, with challenging scenarios designed to evaluate VCL methods. FiGVCL is carefully annotated to preserve the temporal correspondences observed in copied segments. Moreover, we propose a novel fine-grained VCL benchmark metric based on temporal correspondences to improve discriminability. Finally, we design a simple but effective baseline model that uses fine-grained local embeddings for accurate copied segment localization. We also present an unsupervised training strategy that outperforms previous supervised VCL methods. Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | iESTA: Instance-Enhanced Spatial-Temporal Alignment for Video Copy LocalizationabstractVideo copy Segment Localization (VSL) requires the identification of the temporal segments within a pair of videos that contain copied content. Current methods primarily focus on global temporal modeling, overlooking the complementarity of global semantic and local fine-grained features, which limits their effectiveness. Some related methods attempt to incorporate local spatial information but often disrupt spatial semantic structures, resulting in less accurate matching. To address these issues, we propose the Instance-Enhanced Spatial-Temporal Alignment Framework (iESTA), based on a proper representation granularity that integrates instance-level local features and semantic global features. Specifically, the Instance-relation Graph (IRG) is constructed to capture instance-level features and fine-grained interactions, preserving local information integrity and better representing the video feature space in a proper granularity. An instance-GNN structure is designed to refine these graph representations. For global features, we enhance the representation of semantic information, capturing temporal relationships within videos using a Transformer framework. Additionally, we design a Complementarity-perception Alignment Module (CAM) to effectively process and integrate complementary spatial-temporal information, producing accurate frame-to-frame alignment maps. Our approach also incorporates a differentiable Dynamic Time Warping (DTW) method to utilize latent temporal alignments as weak supervisory signals, improving the accuracy of the matching process. Experimental results indicate that our proposed iESTA outperforms state-of-the-art methods on both the small-scale dataset VCDB and the large-scale dataset VCSL. Xinmiao Ding, Jinming Lou, Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Task-Aware Attentional Dynamic Alignment for Few-Shot Compressed Video ClassificationabstractWe present a novel Task-aware Attentional Dynamic Alignment (TADA) framework for visual-based few-shot video classification (FSVC) that addresses two key challenges in this field: efficiency and nuanced spatio-temporal reasoning. Existing methods are often hindered by computationally expensive video decoding processes and neglect the temporal order of videos. In contrast, our method harnesses compressed domain data to extract rich spatio-temporal cues at a fraction of the cost of traditional video processing methods. Specifically, we propose an embedding module to extract informative features from compressed domain data while minimizing computational overheads. Furthermore, to exploit the temporal order of frames, we develop a prototypical ADA module to align and classify videos with an explicit temporal order constraint. Our framework also incorporates a contextual mixer to enrich video embeddings with task-specific context. Extensive experiments on multiple datasets demonstrate that TADA achieves state-of-the-art performance and outperforms existing methods in accuracy and efficiency. Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language ModelsabstractTianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, Ji-Rong Wen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Wenyang Luo, Haoyang Huang, Dongdong Zhang 0001, Xiaolei Wang 0005, Wayne Xin Zhao, Furu Wei, Ji-Rong Wen |
ACL (1) | 2 |
| 2022 | Long-Short Term Cross-Transformer in Compressed Domain for Few-Shot Video ClassificationabstractCompared with image few-shot learning, most of the existing few-shot video classification methods perform worse on feature matching, because they fail to sufficiently exploit the temporal information and relation. Specifically, frames are usually evenly sampled, which may miss important frames. On the other hand, the heuristic model simply encodes the equally treated frames in sequence, which results in the lack of both long-term and short-term temporal modeling and interaction. To alleviate these limitations, we take advantage of the compressed domain knowledge and propose a long-short term Cross-Transformer (LSTC) for few-shot video classification. For short terms, the motion vector (MV) contains temporal cues and reflects the importance of each frame. For long terms, a video can be natively divided into a sequence of GOPs (Group Of Picture). Using this compressed domain knowledge helps to obtain a more accurate spatial-temporal feature space. Consequently, we design the long-short term selection module, short-term module, and long-term module to comprise the LSTC. Long-short term selection is performed to select informative compressed domain data. Long/short-term modules are utilized to sufficiently exploit the temporal information so that the query and support can be well-matched by cross-attention. Experimental results show the superiority of our method on various datasets. Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Yanan Miao, Yangxi Li |
IJCAI | 1 |