Jingwen He

dblp:14/2195 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark
abstract
Kai Zou, Ziqi Huang, Yuhao Dong, Shulin Tian, Dian Zheng, Hongbo Liu, Jingwen He, Bin Liu, Yu Qiao, Ziwei Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuhao Dong, Shulin Tian, Dian Zheng, Jingwen He, Bin Liu 0016, Yu Qiao 0001, Ziwei Liu 0002
ACL (1)7
2025 ResMaster: Mastering High-Resolution Image Generation via Structural and Fine-Grained Guidance
abstract
Diffusion models excel at producing high-quality images; however, scaling to higher resolutions, such as 4K, often results in structural distortions, and repetitive patterns. To this end, we introduce ResMaster, a novel, training-free method that empowers resolution-limited diffusion models to generate high-quality images beyond resolution restrictions. Specifically, ResMaster leverages a low-resolution reference image created by a pre-trained diffusion model to provide structural and fine-grained guidance for crafting high-resolution images on a patch-by-patch basis. To ensure a coherent structure, ResMaster meticulously aligns the low-frequency components of high-resolution patches with the low-resolution reference at each denoising step. For fine-grained guidance, tailored image prompts based on the low-resolution reference and enriched textual prompts produced by a vision-language model are incorporated. This approach could significantly mitigate local pattern distortions and improve detail refinement. Extensive experiments validate that ResMaster sets a new benchmark for high-resolution image generation.
Shuwei Shi, Yuechen Zhang, Jingwen He, Biao Gong, Yinqiang Zheng
AAAI4
2025 MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
abstract
The image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware models are appealing to generate diverse motion patterns, yet there lacks a reliable motion estimator for training such models on large-scale video set in the wild. Traditional metrics, e.g., SSIM or optical flow, are hard to generalize to arbitrary videos, while, it is very tough for human annotators to label the abstract motion intensity neither. Furthermore, the motion intensity shall reveal both local object motion and global camera movement, which has not been studied before. This paper addresses the challenge with a new motion estimator, capable of measuring the decoupled motion intensities of objects and cameras in video. We leverage the contrastive learning on randomly paired videos and distinguish the video with greater motion intensity. Such a paradigm is friendly for annotation and easy to scale up to achieve stable performance on motion estimation. We then present a new I2V model, named MotionStone, developed with the decoupled motion estimator. Experimental results demonstrate the stability of the proposed motion estimator and the state-of-the-art performance of MotionStone on I2V generation. These advantages warrant the decoupled motion estimator to serve as a general plug-in enhancer for both data processing and video generation training.
Shuwei Shi, Biao Gong, Zizheng Yang, Yuyuan Li 0001, Jingwen He, Kecheng Zheng, Jingdong Chen, Ming Yang 0007, Yinqiang Zheng
CVPR8
2025 Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation
abstract
Sora unveils the potential of scaling Diffusion Transformer (DiT) for generating photorealistic images and videos at arbitrary resolutions, aspect ratios, and durations, yet it still lacks sufficient implementation details. In this paper, we introduce the Lumina-T2X family -- a series of Flow-based Large Diffusion Transformers (Flag-DiT) equipped with zero-initialized attention, as a simple and scalable generative framework that can be adapted to various modalities, e.g., transforming noise into images, videos, multi-view 3D objects, or audio clips conditioned on text instructions. By tokenizing the latent spatial-temporal space and incorporating learnable placeholders such as |[nextline]| and |[nextframe]| tokens, Lumina-T2X seamlessly unifies the representations of different modalities across various spatial-temporal resolutions. Advanced techniques like RoPE, KQ-Norm, and flow matching enhance the stability, flexibility, and scalability of Flag-DiT, enabling models of Lumina-T2X to scale up to 7 billion parameters and extend the context window to 128K tokens. This is particularly beneficial for creating ultra-high-definition images with our Lumina-T2I model and long 720p videos with our Lumina-T2V model. Remarkably, Lumina-T2I, powered by a 5-billion-parameter Flag-DiT, requires only 35% of the training computational costs of a 600-million-parameter naive DiT (PixArt-alpha), indicating that increasing the number of parameters significantly accelerates convergence of generative models without compromising visual quality. Our further comprehensive analysis underscores Lumina-T2X's preliminary capability in resolution extrapolation, high-resolution editing, generating consistent 3D views, and synthesizing videos with seamless transitions. All code and checkpoints of Lumina-T2X are released at https://github.com/Alpha-VLLM/Lumina-T2X to further foster creativity, transparency, and diversity in the generative AI community.
Peng Gao 0007, Le Zhuo, Ruoyi Du, Longtian Qiu, Rongjie Huang 0001, Shijie Geng, Renrui Zhang, Junlin Xie, Wenqi Shao, Zhengkai Jiang 0001, Tianshuo Yang, Weicai Ye, Tong He 0001, Jingwen He, Junjun He, Yu Qiao 0001, Hongsheng Li 0001
ICLR17
2025 PDNet: Patch-Wise Deformation Network for Cross-Modal Point Cloud Completion
abstract
Point cloud completion aims to reconstruct complete shapes from partial data. Single-modal methods, constrained by limited prior information, see a substantial boost in completion accuracy with additional guidance provided by image modality. However, most existing methods still adopt encoder-decoder architectures, which tends to cause information loss during compression and decompression operations. To address this, we propose the Patch-wise Deformation Network for cross-modal completion. Rather than encode shapes into high-dimensional representations and subsequently decode them, this strategy directly leverages deformations on point cloud in Euclidean space to retain original information adequately. Specifically, we first split input point cloud into similar patches. Then Patch Deformation module decomposes shape attributes from cross-modal inputs into structural and semantic components as guidance for deformations, enabling high-fidelity completion results. Extensive experiments demonstrate our method’s superiority in point cloud completion task.
Jingwen He, Zhenjiang Du, Ning Xie 0003
ICME1
2025 ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
abstract
Recent Vision-Language Models (VLMs) have shown strong performance in general-purpose visual understanding and reasoning, but their ability to comprehend the visual grammar of movie shots remains underexplored and insufficiently evaluated. To bridge this gap, we present \textbf{ShotBench}, a dedicated benchmark for assessing VLMs’ understanding of cinematic language. ShotBench includes 3,049 still images and 500 video clips drawn from more than 200 films, with each sample annotated by trained annotators or curated from professional cinematography resources, resulting in 3,608 high-quality question-answer pairs. We conduct a comprehensive evaluation of over 20 state-of-the-art VLMs across eight core cinematography dimensions. Our analysis reveals clear limitations in fine-grained perception and cinematic reasoning of current VLMs. To improve VLMs capability in cinematography understanding, we construct a large-scale multimodal dataset, named ShotQA, which contains about 70k Question-Answer pairs derived from movie shots. Besides, we propose ShotVL and train this VLM model with a two-stage training strategy, integrating both supervised fine-tuning and Group Relative Policy Optimization (GRPO). Experimental results demonstrate that our model achieves substantial improvements, surpassing all existing strongest open-source and proprietary models evaluated on ShotBench, establishing a new state-of-the-art performance.
Jingwen He, Dian Zheng, Yuhao Dong, Fan Zhang 0045, Yinan He, Weichao Chen 0001, Yu Qiao 0001, Wanli Ouyang, Shengjie Zhao 0001, Ziwei Liu 0002
NeurIPS2
2025 Imagine360: Immersive 360 Video Generation from Perspective Anchor
abstract
$360^\circ$ videos offer a hyper-immersive experience that allows the viewers to explore a dynamic scene from full 360 degrees. To achieve more accessible and personalized content creation in $360^\circ$ video format, we seek to lift standard perspective videos into $360^\circ$ equirectangular videos. To this end, we introduce **Imagine360**, the first perspective-to-$360^\circ$ video generation framework that creates high-quality $360^\circ$ videos with rich and diverse motion patterns from video anchors. Imagine360 learns fine-grained spherical visual and motion patterns from limited $360^\circ$ video data with several key designs. **1)** Firstly we adopt the dual-branch design, including a perspective and a panorama video denoising branch to provide local and global constraints for $360^\circ$ video generation, with motion module and spatial LoRA layers fine-tuned on $360^\circ$ videos. **2)** Additionally, an antipodal mask is devised to capture long-range motion dependencies, enhancing the reversed camera motion between antipodal pixels across hemispheres. **3)** To handle diverse perspective video inputs, we propose rotation-aware designs that adapt to varying video masking due to changing camera poses across frames. **4)** Lastly, we introduce a new 360 video dataset featuring 10K high-quality, trimmed 360 video clips with structured motion to facilitate training. Extensive experiments show Imagine360 achieves superior graphics quality and motion coherence with our curated dataset among state-of-the-art $360^\circ$ video generation methods. We believe Imagine360 holds promise for advancing personalized, immersive $360^\circ$ video creation.
Jing Tan 0002, Shuai Yang 0001, Jingwen He, Yuwei Guo 0002, Ziwei Liu 0002, Dahua Lin
NeurIPS4
2025 Cut2Next: Generating Next Shot via In-Context Tuning
abstract
Effective multi-shot generation demands purposeful, film-like transitions and strict cinematic continuity. Current methods, however, often prioritize basic visual consistency, neglecting crucial editing patterns (e.g., shot/reverse shot, cutaways) that drive narrative flow for compelling storytelling. This yields outputs that may be visually coherent but lack narrative sophistication and true cinematic integrity. To bridge this, we introduce Next Shot Generation (NSG): synthesizing a subsequent, high-quality shot that critically conforms to professional editing patterns while upholding rigorous cinematic continuity. Our framework, Cut2Next, leverages a Diffusion Transformer (DiT). It employs in-context tuning guided by a novel Hierarchical Multi-Prompting strategy. This strategy uses Relational Prompts to define overall context and inter-shot editing styles. Individual Prompts then specify per-shot content and cinematographic attributes. Together, these guide Cut2Next to generate cinematically appropriate next shots. Architectural innovations, Context-Aware Condition Injection (CACI) and Hierarchical Attention Mask (HAM), further integrate these diverse signals without introducing new parameters. We construct RawCuts (large-scale) and CuratedCuts (refined) datasets, both with hierarchical prompts, and introduce CutBench for evaluation. Experiments show Cut2Next excels in visual consistency and text fidelity. Crucially, user studies reveal a strong preference for Cut2Next, particularly for its adherence to intended editing patterns and overall cinematic continuity, validating its ability to generate high-quality, narratively expressive, and cinematically coherent subsequent shots.
Jingwen He, Yu Qiao 0001, Wanli Ouyang, Ziwei Liu 0002
SIGGRAPH Asia1
2025 Towards Efficient SDRTV-to-HDRTV by Learning From Image Formation
abstract
Contemporary display enables video content rendering with high dynamic range (HDR) and wide color gamut (WCG). However, the majority of existing content remains in standard dynamic range (SDR) format. Therefore, the conversion of SDR content to HDRTV standards holds significant value. This paper delineates and analyzes the SDRTV-to-HDRTV conversion by modeling the formation of SDRTV/HDRTV content. The findings reveal that a naive end-to-end supervised training pipeline suffers from severe gamut transition errors. To address this, we propose a new three-step solution called HDRTVNet++, which includes adaptive global color mapping, local enhancement, and highlight refinement. The adaptive global color mapping step utilizes global statistics for image-adaptive color adjustments, followed by a local enhancement network for detail improvement. These two components are integrated as a generator, with GAN-based joint training ensuring highlight consistency. Our method, tailored for ultra-high-definition TV content, offers both effectiveness and computational efficiency in processing 4K resolution images. We also construct HDRTV1K, a dataset comprising HDR videos adhering to the HDR10 standard, featuring 1235 training and 117 testing images at 4K resolution. Furthermore, we employ five metrics to assess SDRTV-to-HDRTV performance. Our results demonstrate state-of-the-art performance both quantitatively and visually. The codes and models are available athttps://github.com/xiaom233/HDRTVNet-plus.
Xiangyu Chen 0006, Zhengwen Zhang, Jimmy S. J. Ren, Yihao Liu 0001, Jingwen He, Yu Qiao 0001, Jiantao Zhou 0001, Chao Dong 0005
IEEE Trans. Multim.6
2024 Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild
abstract
We introduce SUPIR (Scaling-UP Image Restoration), a groundbreaking image restoration method that harnesses generative prior and the power of model scaling up. Lever-aging multi-modal techniques and advanced generative prior, SUPIR marks a significant advance in intelligent and realistic image restoration. As a pivotal catalyst within SUPIR, model scaling dramatically enhances its capabil-ities and demonstrates new potential for image restoration. We collect a dataset comprising 20 million high-resolution, high-quality images for model training, each en-riched with descriptive text annotations. SUPIR provides the capability to restore images guided by textual prompts, broadening its application scope and potential. Moreover, we introduce negative-quality prompts to further improve perceptual quality. We also develop a restoration-guided sampling method to suppress the fidelity issue encountered in generative-based restoration. Experiments demonstrate SUPIR's exceptional restoration effects and its novel capac-ity to manipulate restoration through textual prompts.
Fanghua Yu, Jinjin Gu, Jinfan Hu, Xiangtao Kong, Xintao Wang 0002, Jingwen He, Yu Qiao 0001, Chao Dong 0005
CVPR7
2024 DiffBIR: Toward Blind Image Restoration with Generative Diffusion Prior
Xinqi Lin, Jingwen He, Zhaoyang Lyu, Bo Dai 0002, Fanghua Yu, Yu Qiao 0001, Wanli Ouyang, Chao Dong 0005
ECCV (59)2
2024 AdaptBIR: Adaptive Blind Image Restoration with latent diffusion prior for higher fidelity
abstract
This work aims to help diffusion models get their footing in the low-level vision field, solving the pain point of insufficient fidelity. Specifically, we propose an Adaptive Blind Image Restoration framework with latent diffusion prior — AdaptBIR, which can adaptively distinguish and address various ranges of degradations. First, we quantitatively categorize images through an Image Quality Assessment (IQA) method. Then, a dual-encoder degradation removal module is employed with the guidance of IQA scores to reach better information preservation. Lastly, we utilize a two-phase controller to handle the reconstruction process in an organized manner. Extensive experiments show that applying such an adaptive framework achieves better performance on both fidelity and perceptual metrics. In this way, AdaptBIR represents more than just a novel framework, it paves the way for a broader application of the diffusion model in blind image restoration tasks.
Yingqi Liu, Jingwen He, Yihao Liu 0001, Xinqi Lin, Fanghua Yu, Jinfan Hu, Yu Qiao 0001, Chao Dong 0005
Pattern Recognit.2
2023 DegAE: A New Pretraining Paradigm for Low-Level Vision
abstract
Self-supervised pretraining has achieved remarkable success in high-level vision, but its application in low-level vision remains ambiguous and not well-established. What is the primitive intention of pretraining? What is the core problem of pretraining in low-level vision? In this paper, we aim to answer these essential questions and establish a new pretraining scheme for low-level vision. Specifically, we examine previous pretraining methods in both high-level and low-level vision, and categorize current low-level vision tasks into two groups based on the difficulty of data acqui-sition: low-cost and high-cost tasks. Existing literature has mainly focused on pretraining for low-cost tasks, where the observed performance improvement is often limited. However, we argue that pretraining is more significant for high-cost tasks, where data acquisition is more challenging. To learn a general low-level vision representation that can improve the performance of various tasks, we propose a new pretraining paradigm called degradation autoencoder (De-gAE). DegAE follows the philosophy of designing pretext task for self-supervised pretraining and is elaborately tai-lored to low-level vision. With DegAE pretraining, SwinIR achieves a 6.88dB performance gain on image dehaze task, while Uformer obtains 3.22dB and 0.54dB improvement on dehaze and derain tasks, respectively.
Yihao Liu 0001, Jingwen He, Jinjin Gu, Xiangtao Kong, Yu Qiao 0001, Chao Dong 0005
CVPR2
2023 Structured Sparsity Learning for Efficient Video Super-Resolution
abstract
The high computational costs of video super-resolution (VSR) models hinder their deployment on resource-limited devices, e.g., smartphones and drones. Existing VSR models contain considerable redundant filters, which drag down the inference efficiency. To prune these unimportant filters, we develop a structured pruning scheme called Structured Sparsity Learning (SSL) according to the properties of VSR. In SSL, we design pruning schemes for several key components in VSR models, including residual blocks, recurrent networks, and upsampling networks. Specifically, we develop a Residual Sparsity Connection (RSC) scheme for residual blocks of recurrent networks to liberate pruning restrictions and preserve the restoration information. For upsampling networks, we design a pixel-shuffle pruning scheme to guarantee the accuracy of feature channel-space conversion. In addition, we observe that pruning error would be amplified as the hidden states propagate along with recurrent networks. To alleviate the issue, we design Temporal Finetuning (TF). Extensive experiments show that SSL can significantly outperform recent methods quantitatively and qualitatively. The code is available at https://github.com/Zj-BinXia/SSL.
Bin Xia 0014, Jingwen He, Yulun Zhang 0001, Yapeng Tian, Wenming Yang, Luc Van Gool
CVPR2
2023 Very Lightweight Photo Retouching Network With Conditional Sequential Modulation
abstract
Photo retouching aims at improving the aesthetic visual quality of images that suffer from photographic defects, especially for poor contrast, over/under exposure, and inharmonious saturation. In practice, photo retouching can be accomplished by a series of image processing operations. As most commonly-used retouching operations are pixel-independent, i.e., the manipulation on one pixel is uncorrelated with its neighboring pixels, we can take advantage of this property and design a specialized algorithm for efficient global photo retouching. We analyze these global operations and find that they can be mathematically formulated by a Multi-Layer Perceptron (MLP). Based on this observation, we propose an extremely lightweight framework – Conditional Sequential Retouching Network (CSRNet). Benefiting from the utilization of$1\times 1$convolution, CSRNet only contains less than 37 K trainable parameters, which are orders of magnitude smaller than existing learning-based methods. Experiments show that our method achieves state-of-the-art performance on the benchmark MIT-Adobe FiveK dataset quantitively and qualitatively. In addition to achieve global photo retouching, the proposed framework can be easily extended to learn local enhancement effects. The extended model, namely CSRNet-L, also achieves competitive results in various local enhancement tasks.
Yihao Liu 0001, Jingwen He, Xiangyu Chen 0006, Zhengwen Zhang, Hengyuan Zhao, Chao Dong 0005, Yu Qiao 0001
IEEE Trans. Multim.2
2022 GCFSR: a Generative and Controllable Face Super Resolution Method Without Facial and GAN Priors
abstract
Face image super resolution (face hallucination) usu-ally relies on facial priors to restore realistic details and preserve identity information. Recent advances can achieve impressive results with the help of GAN prior. They ei-ther design complicated modules to modify the fixed GAN prior or adopt complex training strategies to finetune the generator. In this work, we propose a generative and controllable face SR framework, called GCFSR, which can re-construct images with faithful identity information without any additional priors. Generally, GCFSR has an encoder-generator architecture. Two modules called style modu-lation and feature modulation are designed for the multi-factor SR task. The style modulation aims to generate real-istic face details and the feature modulation dynamically fuses the multi-level encoded features and the generated ones conditioned on the upscaling factor. The simple and elegant architecture can be trained from scratch in an end-to-end manner. For small upscaling factors (≤8), GCFSR can produce surprisingly good results with only adversar-ialloss. After adding L1 and perceptual losses, GCFSR can outperform state-of-the-art methods for large upscalingfac-tors (16, 32, 64). During the test phase, we can modulate the generative strength via feature modulation by changing the conditional upscaling factor continuously to achieve various generative effects. Code is available at https://github.com/hejingwenhejingwen/GCFSR
Jingwen He, Wu Shi, Kai Chen 0023, Lean Fu, Chao Dong 0005
CVPR1
2022 Multimodal Fusion-GMM based Gesture Recognition for Smart Home by WiFi Sensing
abstract
With the thriving development of communication systems and the ubiquity of commodity WiFi devices, wireless intelligent sensing has attracted increasingly attention owing to its importance in Human-Computer Interaction (HCI) where hand gesture recognition has been widely studied these years. However, most of gesture recognition approaches have a limited performance owing to the background noises. In this paper, we present a multimodal fusion-Gaussian mixed model (GMM) based gesture recognition scheme by exploiting channel state information (CSI) of WiFi signals. To this end, we firstly design two theoretical underpinnings, including a sensing model and a recognition model. Firstly, the sensing model is established to investigate the impact of gestures on the propagation properties of WiFi signals and quantify the correlation between the CSI dynamics and various gestures. Secondly, the recognition model is used to exploit physical gesture-induced signal changes to infer potential gestures information. Then, we implement the proposed scheme on a set of WiFi devices and evaluate it in both laboratory and corridor environments. The experiment results show that the proposed scheme can achieve average recognition accuracies of 96% and 94% in these two scenarios, respectively. This shows promise for future ubiquitous hands-free gesture-based interaction with mobile devices.
Jianyang Ding, Yong Wang 0013, Hongyan Si, Jiannan Ma, Jingwen He, Shaozhong Fu
VTC Spring5
2022 Interactive Multi-Dimension Modulation for Image Restoration
abstract
Interactive image restoration aims to generate restored images by adjusting a controlling coefficient which determines the restoration level. Previous works are restricted in modulating image with a single coefficient. However, real images always contain multiple types of degradation, which cannot be well determined by one coefficient. To make a step forward, this paper presents a new problem setup, called multi-dimension (MD) modulation, which aims at modulating output effects across multiple degradation types and levels. Compared with the previous single-dimension (SD) modulation, the MD setup to handle multiple degradations adaptively and relief data unbalancing problem in different degradation types. We also propose a deep architecture - CResMD with newly introduced controllable residual connections for multi-dimension modulation. Specifically, we add a controlling variable on the conventional residual connection to allow a weighted summation of input and residual. The values of these weights are generated by another condition network. We further propose a new data sampling strategy based on beta distribution together with a simple loss reweighting approach to balance different degradation types and levels. With corrupted image and degradation information as inputs, the network can output the corresponding restored image. By tweaking the condition vector, users can control the output effects in MD space at test time. Moreover, we also provide an estimation network to predict the condition vector, thus the base network could directly output the restored image without modulation from users. Extensive experiments demonstrate that the proposed CResMD achieves excellent performance on both SD and MD modulation tasks.
Jingwen He, Chao Dong 0005, Yihao Liu 0001, Yu Qiao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Interactive Multi-dimension Modulation with Dynamic Controllable Residual Learning for Image Restoration
Jingwen He, Chao Dong 0005, Yu Qiao 0001
ECCV (20)1
2020 Conditional Sequential Modulation for Efficient Global Image Retouching
Jingwen He, Yihao Liu 0001, Yu Qiao 0001, Chao Dong 0005
ECCV (13)1
2020 Large Margin Nearest Neighbor Classification With Privileged Information for Biometric Applications
abstract
In this paper, a new metric learning algorithm is proposed to improve face verification and person re-identification in RGB images by learning from RGB and Depth (RGB-D) training images. We address this problem by formulating it as a Learning Using Privileged Information problem, in which the additional depth images associated with the RGB training images are not available for the testing process. Based on the large margin nearest neighbor (LMNN) classification framework, we propose an effective metric learning method called large margin nearest neighbor classification with privileged information (LMNN+) by incorporating depth information to improve decision function learning in the training process. Specifically, two distance metrics based on visual features as well as depth features are jointly learned by minimizing the triplet loss in which the within-class difference is minimized, while the between-class difference is maximized. The distances in the depth feature space can be utilized to guide the training process in the visual feature space. In addition, we propose an efficient optimization method which can handle billions of constraints in the optimization problem of LMNN+. The comprehensive experiments on the EUROCOM data set, the CurtinFaces data set as well as the BIWI RGBD-ID data set demonstrate the effectiveness of our algorithm for face verification and person re-identification by leveraging privileged information.
Jingwen He, Dong Xu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 Modulating Image Restoration With Continual Levels via Adaptive Feature Modification Layers
abstract
In image restoration tasks, like denoising and superresolution, continual modulation of restoration levels is of great importance for real-world applications, but has failed most of existing deep learning based image restoration methods. Learning from discrete and fixed restoration levels, deep models cannot be easily generalized to data of continuous and unseen levels. This topic is rarely touched in literature, due to the difficulty of modulating well-trained models with certain hyper-parameters. We make a step forward by proposing a unified CNN framework that consists of little additional parameters than a single-level model yet could handle arbitrary restoration levels between a start and an end level. The additional module, namely AdaFM layer, performs channel-wise feature modification, and can adapt a model to another restoration level with high accuracy. By simply tweaking an interpolation coefficient, the intermediate model – AdaFM-Net could generate smooth and continuous restoration effects without artifacts. Extensive experiments on three image restoration tasks demonstrate the effectiveness of both model training and modulation testing. Besides, we carefully investigate the properties of AdaFM layers, providing a detailed guidance on the usage of the proposed method.
Jingwen He, Chao Dong 0005, Yu Qiao 0001
CVPR1
2009 A commonsense knowledge base supported multi-agent architecture
Jingwen He, Hokyin Lai, Huaiqing Wang
Expert Syst. Appl.1