EDBT 2026 Demo / reviewers in the wild / expert
Jianyi Wang
dblp:39/4327
· DBLP profile ↗
20ranked-venue papers
7as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video RestorationabstractVideo restoration poses non-trivial challenges in maintaining fidelity while recovering temporally consistent details from unknown degradations in the wild. Despite recent advances in diffusion-based restoration, these methods often face limitations in generation capability and sampling efficiency. In this work, we present SeedVR, a diffusion transformer designed to handle real-world video restoration with arbitrary length and resolution. The core design of SeedVR lies in the shifted window attention that facilitates effective restoration on long video sequences. SeedVR further supports variable-sized windows near the boundary of both spatial and temporal dimensions, overcoming the resolution constraints of traditional window attention. Equipped with contemporary practices, including causal video autoencoder, mixed image and video training, and progressive training, SeedVR achieves highly-competitive performance on both synthetic and real-world benchmarks, as well as AI-generated videos. Extensive experiments demonstrate SeedVR’s superiority over existing methods for generic video restoration. Jianyi Wang, Zhijie Lin 0001, Meng Wei 0007, Yang Zhao 0003, Ceyuan Yang, Chen Change Loy |
CVPR | 1 |
| 2025 | STPYOLO: A Lightweight and Efficient Model for Small Target Pest DetectionabstractVision-based pest detection is an important task in smart agriculture. However, existing visual detection algorithms for small target pests often suffer from low detection accuracy, high complexity, and insufficient generalisation. In order to solve these existing algorithms’ deficiencies for small target pest detection. A pest detection model Small Target Pest YOLO (STPYOLO) based on the improved YOLOv10 model is proposed. To enhance the performance of small-target pest detection, this paper introduces a novel feature fusion method called MicroTargetScaleFusion(MTSFusion). This method finely adjusts feature weights, allowing the model to better focus on key areas where small-target pests are located. It also reduces the complexity of the model, making it more lightweight. To compensate for the lack of small-target pest datasets, a self-constructed small-target pest dataset is built, including six pests. Compared to YOLOv10, STPYOLO shows strong performance on two public datasets and the self-built dataset. The model size is reduced by 36.2%, and the number of parameters by 40%. Meanwhile on the self-built dataset mAP50 increased to 72.6%, an improvement of 5.8%. The excellent detection performance is still maintained while lightweighting. Finally, we integrated the STPYOLO model into the designed pest detection and early warning system to realize the real-time pest detection function for farm cotton fields. Jianyi Wang, Zhenhong Jia |
IJCNN | 1 |
| 2025 | Efficient Diffusion Model for Image Restoration by Residual ShiftingabstractWhile diffusion-based image restoration (IR) methods have achieved remarkable success, they are still limited by the low inference speed attributed to the necessity of executing hundreds or even thousands of sampling steps. Existing acceleration sampling techniques, though seeking to expedite the process, inevitably sacrifice performance to some extent, resulting in over-blurry restored outcomes. To address this issue, this study proposes a novel and efficient diffusion model for IR that significantly reduces the required number of diffusion steps. Our method avoids the need for post-acceleration during inference, thereby avoiding the associated performance deterioration. Specifically, our proposed method establishes a Markov chain that facilitates the transitions between the high-quality and low-quality images by shifting their residuals, substantially improving the transition efficiency. A carefully formulated noise schedule is devised to flexibly control the shifting speed and the noise strength during the diffusion process. Extensive experimental evaluations demonstrate that the proposed method achieves superior or comparable performance to current state-of-the-art methods on four classical IR tasks, namely image super-resolution, image inpainting, blind face restoration, and image deblurring, even only with four sampling steps. Zongsheng Yue, Jianyi Wang, Chen Change Loy |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-ResolutionabstractText-based diffusion models have exhibited remarkable success in generation and editing, showing great promise for enhancing visual content with their generative prior. However, applying these models to video super-resolution remains challenging due to the high demands for output fidelity and temporal consistency, which is complicated by the inherent randomness in diffusion models. Our study introduces Upscale-A-Video, a text-guided latent diffusion framework for video upscaling. This framework ensures temporal coherence through two key mechanisms: locally, it integrates temporal layers into U-Net and VAE-Decoder, maintaining consistency within short sequences; globally, without training, a flow-guided recurrent latent propagation module is introduced to enhance overall video stability by propagating and fusing latent across the entire sequences. Thanks to the diffusion paradigm, our model also offers greater flexibility by allowing text prompts to guide texture creation and adjustable noise levels to balance restoration and generation, enabling a trade-off between fidelity and quality. Extensive experiments show that Upscale-A-Video surpasses existing methods in both synthetic and real-world benchmarks, as well as in AI-generated videos, showcasing impressive visual realism and temporal consistency. Shangchen Zhou, Peiqing Yang 0001, Jianyi Wang, Yihang Luo, Chen Change Loy |
CVPR | 3 |
| 2024 | Dual-MambaNet: A Lightweight Dual-Branch Brain Image Segmentation Network Based on Local Attention and Mamba
Dayong Ren, Zhenhong Jia, Jianyi Wang |
ICPR (28) | 5 |
| 2024 | Exploiting Diffusion Prior for Real-World Image Super-Resolution
Jianyi Wang, Zongsheng Yue, Shangchen Zhou, Kelvin C. K. Chan, Chen Change Loy |
Int. J. Comput. Vis. | 1 |
| 2023 | Exploring CLIP for Assessing the Look and Feel of ImagesabstractMeasuring the perception of visual content is a long-standing problem in computer vision. Many mathematical models have been developed to evaluate the look or quality of an image. Despite the effectiveness of such tools in quantifying degradations such as noise and blurriness levels, such quantification is loosely coupled with human language. When it comes to more abstract perception about the feel of visual content, existing methods can only rely on supervised models that are explicitly trained with labeled data collected via laborious user study. In this paper, we go beyond the conventional paradigms by exploring the rich visual language prior encapsulated in Contrastive Language-Image Pre-training (CLIP) models for assessing both the quality perception (look) and abstract perception (feel) of images without explicit task-specific training. In particular, we discuss effective prompt designs and show an effective prompt pairing strategy to harness the prior. We also provide extensive experiments on controlled datasets and Image Quality Assessment (IQA) benchmarks. Our results show that CLIP captures meaningful priors that generalize well to different perceptual assessments. Jianyi Wang, Kelvin C. K. Chan, Chen Change Loy |
AAAI | 1 |
| 2023 | ResShift: Efficient Diffusion Model for Image Super-resolution by Residual ShiftingabstractDiffusion-based image super-resolution (SR) methods are mainly limited by the low inference speed due to the requirements of hundreds or even thousands of sampling steps. Existing acceleration sampling techniques inevitably sacrifice performance to some extent, leading to over-blurry SR results. To address this issue, we propose a novel and efficient diffusion model for SR that significantly reduces the number of diffusion steps, thereby eliminating the need for post-acceleration during inference and its associated performance deterioration. Our method constructs a Markov chain that transfers between the high-resolution image and the low-resolution image by shifting the residual between them, substantially improving the transition efficiency. Additionally, an elaborate noise schedule is developed to flexibly control the shifting speed and the noise strength during the diffusion process. Extensive experiments demonstrate that the proposed method obtains superior or at least comparable performance to current state-of-the-art methods on both synthetic and real-world datasets, \textit{\textbf{even only with 20 sampling steps}}. Our code and model will be made publicly. Zongsheng Yue, Jianyi Wang, Chen Change Loy |
NeurIPS | 2 |
| 2022 | SmoothNet: A Plug-and-Play Network for Refining Human Poses in Videos
Ailing Zeng, Xuan Ju, Jianyi Wang, Qiang Xu 0001 |
ECCV (5) | 5 |
| 2022 | MW-GAN+ for Perceptual Quality Enhancement on Compressed VideoabstractThe great success of deep learning has boosted the fast development of video quality enhancement. However, existing methods mainly focus on enhancing the objective quality of compressed video, and ignore their perceptual quality that plays a key role in determining quality of experience (QoE) of videos. In this paper, we aim at enhancing the perceptual quality of compressed video. Our main observation is that perceptual quality enhancement mostly relies on recovering the high-frequency details with fine textures. Accordingly, we propose a novel generative adversarial network (GAN) based on multi-level wavelet packet transform (WPT), which is called multi-level wavelet-based GAN+ (MW-GAN+), to exploit high-frequency details for enhancing the perceptual quality of compressed video. In MW-GAN+, we first propose a multi-level wavelet pixel-adaptive (MWP) module to extract temporal information across video frames, such that frame similarity can be utilized in recovering high-frequency details. Then, a wavelet reconstruction network, consisting of wavelet-dense residual blocks (WDRB), is developed to recover high-frequency details in a multi-level manner for enhanced frame reconstruction. Finally, we develop a 3D discriminator to encourage temporal coherence with a 3D-CNN based architecture. Experimental results demonstrate the superiority of our method over state-of-the-art methods in enhancing the perceptual quality of compressed video. Our code is available athttps://github.com/IceClear/MW-GAN. Jianyi Wang, Mai Xu, Xin Deng 0002, Liquan Shen, Yuhang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | HiNet: Deep Image Hiding by Invertible NetworkabstractImage hiding aims to hide a secret image into a cover image in an imperceptible way, and then recover the secret image perfectly at the receiver end. Capacity, invisibility and security are three primary challenges in image hiding task. This paper proposes a novel invertible neural network (INN) based framework, HiNet, to simultaneously overcome the three challenges in image hiding. For large capacity, we propose an inverse learning mechanism by simultaneously learning the image concealing and revealing processes. Our method is able to achieve the concealing of a full-size secret image into a cover image with the same size. For high invisibility, instead of pixel domain hiding, we propose to hide the secret information in wavelet domain. Furthermore, we propose a new low-frequency wavelet loss to constrain that secret information is hidden in high-frequency wavelet subbands, which significantly improves the hiding security. Experimental results show that our HiNet significantly outperforms other state-of-the-art image hiding methods, with more than 10 dB PSNR improvement in secret image recovery on ImageNet, COCO and DIV2K datasets. Codes are available at https://github.com/TomTomTommi/HiNet. Junpeng Jing, Xin Deng 0002, Mai Xu, Jianyi Wang, Zhenyu Guan 0002 |
ICCV | 4 |
| 2021 | Attention-Based Deep Reinforcement Learning for Virtual Cinematography of 360$^{\circ}$ VideosabstractVirtual cinematography refers to automatically selecting a natural-looking normal field-of-view (NFOV) from an entire 360$^{\circ}$video. In fact, virtual cinematography can be modeled as a deep reinforcement learning (DRL) problem, in which an agent makes actions related to NFOV selection according to the environment of 360$^{\circ}$video frames. More importantly, we find from our data analysis that the selected NFOVs attract significantly more attention than other regions, i.e., the NFOVs have high saliency. Therefore, in this paper, we propose an attention-based DRL (A-DRL) approach for virtual cinematography in 360$^{\circ}$video. Specifically, we develop a new DRL framework for automatic NFOV selection with the input of both the content, and saliency map of each 360$^{\circ}$frame. Then, we propose a new reward function for the DRL framework in our approach, which considers the saliency values, ground-truth, and smooth transition for NFOV selection. Subsequently, a simplified DenseNet (called Mini-DenseNet) is designed to learn the optimal policy via maximizing the reward. Based on the learned policy, the actions of NFOV can be made in our A-DRL approach for virtual cinematography of 360$^{\circ}$video. Extensive experiments show that our A-DRL approach outperforms other state-of-the-art virtual cinematography methods, over the datasets of Sports-360 video, and Pano2Vid. Jianyi Wang, Mai Xu, Lai Jiang 0004, Yuhang Song 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Arena: A General Evaluation Platform and Building Toolkit for Multi-Agent IntelligenceabstractLearning agents that are not only capable of taking tests, but also innovating is becoming a hot topic in AI. One of the most promising paths towards this vision is multi-agent learning, where agents act as the environment for each other, and improving each agent means proposing new problems for others. However, existing evaluation platforms are either not compatible with multi-agent settings, or limited to a specific game. That is, there is not yet a general evaluation platform for research on multi-agent intelligence. To this end, we introduce Arena, a general evaluation platform for multi-agent intelligence with 35 games of diverse logics and representations. Furthermore, multi-agent intelligence is still at the stage where many problems remain unexplored. Therefore, we provide a building toolkit for researchers to easily invent and build novel multi-agent problems from the provided game set based on a GUI-configurable social tree and five basic multi-agent reward schemes. Finally, we provide Python implementations of five state-of-the-art deep multi-agent reinforcement learning baselines. Along with the baseline implementations, we release a set of 100 best agents/teams that we can train with different training schemes for each game, as the base for evaluating agents with population performance. As such, the research community can perform comparisons under a stable and uniform standard. All the implementations and accompanied tutorials have been open-sourced for the community at https://sites.google.com/view/arena-unity/. Yuhang Song 0001, Andrzej Wojcicki, Thomas Lukasiewicz, Jianyi Wang, Abi Aryan, Zhenghua Xu 0001, Mai Xu, Lianlong Wu |
AAAI | 4 |
| 2020 | Mega-Reward: Achieving Human-Level Play without Extrinsic RewardsabstractIntrinsic rewards were introduced to simulate how human intelligence works; they are usually evaluated by intrinsically-motivated play, i.e., playing games without extrinsic rewards but evaluated with extrinsic rewards. However, none of the existing intrinsic reward approaches can achieve human-level performance under this very challenging setting of intrinsically-motivated play. In this work, we propose a novel megalomania-driven intrinsic reward (called mega-reward), which, to our knowledge, is the first approach that achieves human-level performance in intrinsically-motivated play. Intuitively, mega-reward comes from the observation that infants' intelligence develops when they try to gain more control on entities in an environment; therefore, mega-reward aims to maximize the control capabilities of agents on given entities in a given environment. To formalize mega-reward, a relational transition model is proposed to bridge the gaps between direct and latent control. Experimental studies show that mega-reward (i) can greatly outperform all state-of-the-art intrinsic reward approaches, (ii) generally achieves the same level of performance as Ex-PPO and professional human-level scores, and (iii) has also a superior performance when it is incorporated with extrinsic rewards. Yuhang Song 0001, Jianyi Wang, Thomas Lukasiewicz, Zhenghua Xu 0001, Shangtong Zhang, Andrzej Wojcicki, Mai Xu |
AAAI | 2 |
| 2020 | Multi-level Wavelet-Based Generative Adversarial Network for Perceptual Quality Enhancement of Compressed Video
Jianyi Wang, Xin Deng 0002, Mai Xu, Congyong Chen, Yuhang Song 0001 |
ECCV (14) | 1 |
| 2020 | Drug-target interactions prediction using marginalized denoising model on heterogeneous networksabstractBACKGROUND: Drugs achieve pharmacological functions by acting on target proteins. Identifying interactions between drugs and target proteins is an essential task in old drug repositioning and new drug discovery. To recommend new drug candidates and reposition existing drugs, computational approaches are commonly adopted. Compared with the wet-lab experiments, the computational approaches have lower cost for drug discovery and provides effective guidance in the subsequent experimental verification. How to integrate different types of biological data and handle the sparsity of drug-target interaction data are still great challenges. RESULTS: In this paper, we propose a novel drug-target interactions (DTIs) prediction method incorporating marginalized denoising model on heterogeneous networks with association index kernel matrix and latent global association. The experimental results on benchmark datasets and new compiled datasets indicate that compared to other existing methods, our method achieves higher scores of AUC (area under curve of receiver operating characteristic) and larger values of AUPR (area under precision-recall curve). CONCLUSIONS: The performance improvement in our method depends on the association index kernel matrix and the latent global association. The association index kernel matrix calculates the sharing relationship between drugs and targets. The latent global associations address the false positive issue caused by network link sparsity. Our method can provide a useful approach to recommend new drug candidates and reposition existing drugs. Chunyan Tang, Jianyi Wang |
BMC Bioinform. | 4 |
| 2019 | Diversity-Driven Extensible Hierarchical Reinforcement LearningabstractHierarchical reinforcement learning (HRL) has recently shown promising advances on speeding up learning, improving the exploration, and discovering intertask transferable skills. Most recent works focus on HRL with two levels, i.e., a master policy manipulates subpolicies, which in turn manipulate primitive actions. However, HRL with multiple levels is usually needed in many real-world scenarios, whose ultimate goals are highly abstract, while their actions are very primitive. Therefore, in this paper, we propose a diversitydriven extensible HRL (DEHRL), where an extensible and scalable framework is built and learned levelwise to realize HRL with multiple levels. DEHRL follows a popular assumption: diverse subpolicies are useful, i.e., subpolicies are believed to be more useful if they are more diverse. However, existing implementations of this diversity assumption usually have their own drawbacks, which makes them inapplicable to HRL with multiple levels. Consequently, we further propose a novel diversity-driven solution to achieve this assumption in DEHRL. Experimental studies evaluate DEHRL with nine baselines from four perspectives in two domains; the results show that DEHRL outperforms the state-of-the-art baselines in all four aspects. Yuhang Song 0001, Jianyi Wang, Thomas Lukasiewicz, Zhenghua Xu 0001, Mai Xu |
AAAI | 2 |
| 2019 | Predicting Head Movement in Panoramic Video: A Deep Reinforcement Learning ApproachabstractPanoramic video provides immersive and interactive experience by enabling humans to control the field of view (FoV) through head movement (HM). Thus, HM plays a key role in modeling human attention on panoramic video. This paper establishes a database collecting subjects' HM in panoramic video sequences. From this database, we find that the HM data are highly consistent across subjects. Furthermore, we find that deep reinforcement learning (DRL) can be applied to predict HM positions, via maximizing the reward of imitating human HM scanpaths through the agent's actions. Based on our findings, we propose a DRL-based HM prediction (DHP) approach with offline and online versions, called offline-DHP and online-DHP. In offline-DHP, multiple DRL workflows are run to determine potential HM positions at each panoramic frame. Then, a heat map of the potential HM positions, named the HM map, is generated as the output of offline-DHP. In online-DHP, the next HM position of one subject is estimated given the currently observed HM position, which is achieved by developing a DRL algorithm upon the learned offline-DHP model. Finally, the experiments validate that our approach is effective in both offline and online prediction of HM positions for panoramic video, and that the learned offline-DHP model can improve the performance of online-DHP. Mai Xu, Yuhang Song 0001, Jianyi Wang, Minglang Qiao, Liangyu Huo, Zulin Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | A process-mining-based scenarios generation method for SOA application development
Lihong Jiang, Jianyi Wang, Nazaraf Shah, Hongming Cai 0001, Chengxi Huang, Raymond Farmer |
Serv. Oriented Comput. Appl. | 2 |
| 2006 | Effective Low Frequency Disturbance Feed Forward Compensation Scheme for Mobile HDD ApplicationabstractThe advantage of low cost and large capacity gives strong driving force for hard disk to enter the consumer electronic market. One of the most challenging servo control issue faced in mobile application is to perform normal read/write operation when significant vibration is excited in HDD jogging/walking condition. This paper presents a simple yet effective method to improvement hard disk throughput with an add-on low frequency disturbance feed forward compensator. With the proposed approach, the track following performance is significantly improved in jogging/walking, and the hard disk is capable to meet the read/write throughput requirement with minimum servo controller re-design effort Shuyu Cao, Qiang Bi, Mingzhong Ding, Jianyi Wang, KianKeong Ooi |
ICARCV | 4 |