Yifan Wang 0004

dblp:47/6959-4 · DBLP profile ↗
← Back
31ranked-venue papers
6as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-author · 19 since 2021Artificial intelligence and machine learning · 19 · 4 first-author · 15 since 2021
YearPublicationVenuePosition
2026 Aggregating global-scale pixel-wise forgery cues within a graph
Hengrun Zhao, Yifan Wang 0004, Yunzhi Zhuge, Lijun Wang 0001, Huchuan Lu
Neural Networks2
2025 Mono2Stereo: A Benchmark and Empirical Study for Stereo Conversion
abstract
With the rapid proliferation of 3D devices and the shortage of 3D content, stereo conversion is attracting increasing attention. Recent works introduce pretrained Diffusion Models (DMs) into this task. However, due to the scarcity of large-scale training data and comprehensive benchmarks, the optimal methodologies for employing DMs in stereo conversion and the accurate evaluation of stereo effects remain largely unexplored. In this work, we introduce the Mono2Stereo dataset, providing high-quality training data and benchmark to support in-depth exploration of stereo conversion. With this dataset, we conduct an empirical study that yields two primary findings. 1) The differences between the left and right views are subtle, yet existing metrics consider overall pixels, failing to concentrate on regions critical to stereo effects. 2) Mainstream methods adopt either one-stage left-to-right generation or warp-and-inpaint pipeline, facing challenges of degraded stereo effect and image distortion respectively. Based on these findings, we introduce a new evaluation metric, Stereo Intersection-over-Union, which prioritizes disparity and achieves a high correlation with human judgments on stereo effect. Moreover, we propose a strong baseline model, harmonizing the stereo effect and image quality simultaneously, and notably surpassing current mainstream methods. Our code and data will be open-sourced to promote further research in stereo conversion. Our models are available at mono2stereo-bench.github.io.
Songsong Yu, Zhongang Qi, Zeke Xie, Yifan Wang 0004, Lijun Wang 0001, Ying Shan, Huchuan Lu
CVPR5
2025 GLDesigner: Leveraging Multi-Modal LLMs as Designer for Enhanced Aesthetic Text Glyph Layouts
abstract
Text logo design heavily relies on the creativity and expertise of professional designers, in which arranging element layouts is one of the most important procedures. However, this specific task has received limited attention, often overshadowed by broader layout generation tasks such as document or poster design. In this paper, we propose a Vision-Language Model (VLM)-based framework that generates content-aware text logo layouts by integrating multi-modal inputs with user-defined constraints, enabling more flexible and robust layout generation for real-world applications. We introduce two model techniques that reduce the computational cost for processing multiple glyph images simultaneously, without compromising performance. To support instruction tuning of our model, we construct two extensive text logo datasets that are five times larger than existing public datasets. In addition to geometric annotations (e.g., text masks and character recognition), our datasets include detailed layout descriptions in natural language, enabling the model to reason more effectively in handling complex designs and custom user inputs. Experimental results demonstrate the effectiveness of our proposed framework and datasets, outperforming existing methods on various benchmarks that assess geometric aesthetics and human preferences.
Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Chenyang Li 0007, Jin-Peng Lan, Jun-Yan He, Bin Luo 0008, Yifeng Geng
ACM Multimedia2
2025 From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
abstract
Despite remarkable progress in driving world models, their potential for autonomous systems remains largely untapped: the world models are mostly learned for world simulation and decoupled from trajectory planning. While recent efforts aim to unify world modeling and planning in a single framework, the synergistic facilitation mechanism of world modeling for planning still requires further exploration. In this work, we introduce a new driving paradigm named Policy World Model (PWM), which not only integrates world modeling and trajectory planning within a unified architecture, but is also able to benefit planning using the learned world knowledge through the proposed action-free future state forecasting scheme. Through collaborative state-action prediction, PWM can mimic the human-like anticipatory perception, yielding more reliable planning performance. To facilitate the efficiency of video forecasting, we further introduce a parallel token generation mechanism, equipped with a context-guided tokenizer and an adaptive dynamic focal loss. Despite utilizing only front camera input, our method matches or exceeds state-of-the-art approaches that rely on multi-view and multi-modal inputs. Code will be released at https://github.com/6550Zhao/Policy-World-Model.
Zhida Zhao, Talas Fu, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu
NeurIPS3
2025 Learning pose regression as reliable pixel-level matching for self-supervised depth estimation
Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu
Neurocomputing2
2025 AVS-Mamba: Exploring Temporal and Multi-Modal Mamba for Audio-Visual Segmentation
abstract
The essence of audio-visual segmentation (AVS) lies in locating and delineating sound-emitting objects within a video stream. While Transformer-based methods have shown promise, their handling of long-range dependencies struggles due to quadratic computational costs, presenting a bottleneck in complex scenarios. To overcome this limitation and facilitate complex multi-modal comprehension with linear complexity, we introduce AVS-Mamba, a selective state space model to address the AVS task. Our framework incorporates two key components for video understanding and cross-modal learning: Temporal Mamba Block for sequential video processing and Vision-to-Audio Fusion Block for advanced audio-vision integration. Building on this, we develop the Multi-scale Temporal Encoder, aimed at enhancing the learning of visual features across scales, facilitating the perception of intra- and inter-frame information. To perform multi-modal fusion, we propose the Modality Aggregation Decoder, leveraging the Vision-to-Audio Fusion Block to integrate visual features into audio features across both frame and temporal levels. Further, we adopt the Contextual Integration Pyramid to perform audio-to-vision spatial-temporal context collaboration. Through these innovative contributions, our approach achieves new state-of-the-art results on the AVSBench-object and AVSBench-semantic datasets.
Sitong Gong, Yunzhi Zhuge, Lu Zhang 0053, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu
IEEE Trans. Multim.4
2024 DME: Unveiling the Bias for Better Generalized Monocular Depth Estimation
abstract
This paper aims to design monocular depth estimation models with better generalization abilities. To this end, we have conducted quantitative analysis and discovered two important insights. First, the Simulation Correlation phenomenon, commonly seen in long-tailed classification problems, also exists in monocular depth estimation, indicating that the imbalanced depth distribution in training data may be the cause of limited generalization ability. Second, the imbalanced and long-tail distribution of depth values extends beyond the dataset scale, and also manifests within each individual image, further exacerbating the challenge of monocular depth estimation. Motivated by the above findings, we propose the Distance-aware Multi-Expert (DME) depth estimation model. Unlike prior methods that handle different depth range indiscriminately, DME adopts a divide-and-conquer philosophy where each expert is responsible for depth estimation of regions within a specific depth range. As such, the depth distribution seen by each expert is more uniform and can be more easily predicted. A pixel-level routing module is further designed and learned to stitch the prediction of all experts into the final depth map. Experiments show that DME achieves state-of-the-art performance on both NYU-Depth v2 and KITTI, and also delivers favorable zero-shot generalization capability on unseen datasets.
Songsong Yu, Yifan Wang 0004, Yunzhi Zhuge, Lijun Wang 0001, Huchuan Lu
AAAI2
2024 3D Prompt Learning for RGB-D Tracking
Bocen Li, Yunzhi Zhuge, Lijun Wang 0001, Yifan Wang 0004, Huchuan Lu
ACCV (2)5
2024 Multi-Modal Instruction Tuned LLMs with Fine-Grained Visual Perception
abstract
Multimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs with visual perceiving and grounding capabilities. However, there still remains a gap in providing fine-grained pixel-level perceptions and extending interactions beyond text-specific inputs. In this work, we propose AnyRef, a general MLLM model that can generate pixel-wise object perceptions and natural language descriptions from multi-modality references, such as texts, boxes, images, or audio. This innovation empowers users with greater flexibility to engage with the model beyond textual and regional prompts, without modality-specific designs. Through our proposed refocusing mechanism, the generated grounding output is guided to better focus on the referenced object, implicitly incorporating additional pixel-level supervision. This simple modification utilizes attention scores generated during the inference of LLM, eliminating the need for extra computations while exhibiting performance enhancements in both grounding masks and referring expressions. With only publicly available training data, our model achieves state-of-the-art results across multiple benchmarks, including diverse modality referring segmentation and region-level referring expression generation. Code and models are available at https://github.com/jwh97nn/AnyRef
Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Jun-Yan He, Jin-Peng Lan, Bin Luo 0008, Xuansong Xie
CVPR2
2024 SelM: Selective Mechanism based Audio-Visual Segmentation
abstract
Audio-Visual Segmentation (AVS) aims to segment sound-producing objects in videos according to associated audio cues, where both modalities are affected by noise to different extents, such as the blending of background noises in audio or the presence of distracted objects in video. Most existing methods focus on learning interactions between modalities at high semantic levels but is incapable of filtering low-level noise or achieving fine-grained representational interactions during the early feature extraction phase. Consequently, they struggle with illusion issues, where nonexistent audio cues are erroneously linked to visual objects. In this paper, we present SelM, a novel architecture that leverages selective mechanisms to counteract these illusions. SelM employs State Space model for noise reduction and robust feature selection. By imposing additional bidirectional constraints on audio and visual embeddings, it is able to precisely identify crucial features corresponding to sound-emitting targets. To fill the existing gap in early fusion within AVS, SelM introduces a dual alignment mechanism specifically engineered to facilitate intricate spatio-temporal interactions between audio and visual streams, achieving more fine-grained representations. Moreover, we develop a cross-level decoder for layered reasoning, significantly enhancing segmentation precision by exploring the complex relationships between audio and visual information. SelM achieves state-of-the-art performance in AVS tasks, especially in the challenging Audio-Visual Semantic Segmentation subset. The code can be found at https://github.com/Cyyzpoi/SelM.
Songsong Yu, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu
ACM Multimedia3
2024 LOVD: Large-and-Open Vocabulary Object Detection
abstract
Existing open-vocabulary object detectors require an accurate and compact vocabulary pre-defined during inference. Their performance is largely degraded in real scenarios where the underlying vocabulary may be indeterminate and often exponentially large. To have a more comprehensive understanding of this phenomenon, we propose a new setting called Large-and-Open Vocabulary object Detection, which simulates real scenarios by testing detectors with large vocabularies containing thousands of unseen categories. The vast unseen categories inevitably lead to an increase in category distractors, severely impeding the recognition process and leading to unsatisfactory detection results. To address this challenge, We propose a Large and Open Vocabulary Detector (LOVD) with two core components, termed the Image-to-Region Filtering (IRF) module and Cross-View Verification (CV2) scheme. To relieve the category distractors of the given large vocabularies, IRF performs image-level recognition to build a compact vocabulary relevant to the image scene out of the large input vocabulary, followed by region-level classification upon the compact vocabulary. CV2 further enhances the IRF by conducting image-to-region filtering in both global and local views and produces the final detection categories through a two-branch voting mechanism. Compared to the prior works, our LOVD is more scalable and robust to large input vocabularies, and can be seamlessly integrated with predominant detection methods to improve their open-vocabulary performance. The code can be found at https://github.com/Altria-luo/LOVD.
Shiyu Tang, Zhaofan Luo, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Weibo Su
ACM Multimedia3
2024 MaskMentor: Unlocking the Potential of Masked Self-Teaching for Missing Modality RGB-D Semantic Segmentation
abstract
Existing RGB-D semantic segmentation methods struggle to handle modality missing input, where only RGB images or depth maps are available, leading to degenerated segmentation performance. We tackle this issue using MaskMentor, a new pre-training framework for modality missing segmentation, which advances its counterparts via two novel designs: Masked Modality and Image Modeling (M2IM), and Self-Teaching via Token-Pixel Joint reconstruction (STTP). M2IM simulates modality missing scenarios by combining both modality- and patch-level random masking. Meanwhile, STTP offers an effective self-teaching strategy, where the trained network assumes a dual role, simultaneously acting as both the teacher and the student. The student with modality missing input is supervised by the teacher with complete modality input through both token- and pixel-wise masked modeling, closing the gap between missing and complete input modalities. By integrating M2IM and STTP, MaskMentor significantly improves the generalization ability of the trained model across diverse input conditions and outperforms state-of-the-art methods on two popular benchmarks by a considerable margin. Extensive ablation studies further verify the effectiveness of the above contributions.
Zhida Zhao, Lijun Wang 0001, Yifan Wang 0004, Huchuan Lu
ACM Multimedia4
2024 MOFTrack: Multi-object Formation Tracking in Remote Sensing Videos
Haijiang Sun, Qiaoyuan Liu, Tanlin Li, Lijun Wang 0001, Yifan Wang 0004
PRCV (12)7
2024 CSRNet: Focusing on critical points for depth completion
Bocen Li, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu
Image Vis. Comput.2
2023 Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance Learning
abstract
Depth-aware panoptic segmentation is an emerging topic in computer vision which combines semantic and geometric understanding for more robust scene interpretation. Recent works pursue unified frameworks to tackle this challenge but mostly still treat it as two individual learning tasks, which limits their potential for exploring cross-domain information. We propose a deeply unified framework for depth-aware panoptic segmentation, which performs joint segmentation and depth estimation both in a persegment manner with identical object queries. To narrow the gap between the two tasks, we further design a geometric query enhancement method, which is able to integrate scene geometry into object queries using latent representations. In addition, we propose a bi-directional guidance learning approach to facilitate cross-task feature learning by taking advantage of their mutual relations. Our method sets the new state of the art for depth-aware panoptic segmentation on both Cityscapes-DVPS and SemKITTI-DVPS datasets. Moreover, our guidance learning approach is shown to deliver performance improvement even under incomplete supervision labels. Code and models are available at https://github.com/jwh97nn/DeepDPS.
Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Bin Luo 0008, Jun-Yan He, Jin-Peng Lan, Yifeng Geng, Xuansong Xie
ICCV2
2023 Isomer: Isomerous Transformer for Zero-shot Video Object Segmentation
abstract
Recent leading zero-shot video object segmentation (ZVOS) works devote to integrating appearance and motion information by elaborately designing feature fusion modules and identically applying them in multiple feature stages. Our preliminary experiments show that with the strong long-range dependency modeling capacity of Transformer, simply concatenating the two modality features and feeding them to vanilla Transformers for feature fusion can distinctly benefit the performance but at a cost of heavy computation. Through further empirical analysis, we find that attention dependencies learned in Transformer in different stages exhibit completely different properties: global query-independent dependency in the low-level stages and semantic-specific dependency in the high-level stages. Motivated by the observations, we propose two Transformer variants: i) Context-Sharing Transformer (CST) that learns the global-shared contextual information within image frames with a lightweight computation. ii) Semantic Gathering-Scattering Transformer (SGST) that models the semantic correlation separately for the foreground and background and reduces the computation cost with a soft token merging mechanism. We apply CST and SGST for low-level and high-level feature fusions, respectively, formulating a level-isomerous Transformer framework for ZVOS task. Compared with the baseline that uses vanilla Transformers for multi-stage fusion, ours significantly increase the speed by 13× and achieves new state-of-the-art ZVOS performance. Code is available at https://github.com/DLUT-yyc/Isomer.
Yifan Wang 0004, Lijun Wang 0001, Xiaoqi Zhao 0003, Huchuan Lu, Yu Wang 0108, Weibo Su, Lei Zhang 0006
ICCV2
2023 SACFormer: Unify Depth Estimation and Completion with Prompt
Shiyu Tang, Yifan Wang 0004, Lijun Wang 0001
PRCV (2)3
2023 L2MNet: Enhancing Continual Semantic Segmentation with Mask Matching
Bocen Li, Yifan Wang 0004
PRCV (10)3
2022 You Only Infer Once: Cross-Modal Meta-Transfer for Referring Video Object Segmentation
abstract
We present YOFO (You Only inFer Once), a new paradigm for referring video object segmentation (RVOS) that operates in an one-stage manner. Our key insight is that the language descriptor should serve as target-specific guidance to identify the target object, while a direct feature fusion of image and language can increase feature complexity and thus may be sub-optimal for RVOS. To this end, we propose a meta-transfer module, which is trained in a learning-to-learn fashion and aims to transfer the target-specific information from the language domain to the image domain, while discarding the uncorrelated complex variations of language description. To bridge the gap between the image and language domains, we develop a multi-scale cross-modal feature mining block that aggregates all the essential features required by RVOS from both domains and generates regression labels for the meta-transfer module. The whole system can be trained in an end-to-end manner and shows competitive performance against state-of-the-art two-stage approaches.
Dezhuang Li, Ruoqi Li, Lijun Wang 0001, Yifan Wang 0004, Jinqing Qi, Lu Zhang 0053, Ting Liu 0018, Qingquan Xu, Huchuan Lu
AAAI4
2022 Multi-Source Uncertainty Mining for Deep Unsupervised Saliency Detection
abstract
Deep learning-based image salient object detection (SOD) heavily relies on large-scale training data with pixel-wise labeling. High-quality labels involve intensive labor and are expensive to acquire. In this paper, we propose a novel multi-source uncertainty mining method to facilitate unsupervised deep learning from multiple noisy labels generated by traditional handcrafted SOD methods. We design an Uncertainty Mining Network (UMNet) which consists of multiple Merge-and-Split (MS) modules to recursively analyze the commonality and difference among multiple noisy labels and infer pixel-wise uncertainty map for each label. Meanwhile, we model the noisy labels using Gibbs distribution and propose a weighted uncertainty loss to jointly train the UMNet with the SOD network. As a consequence, our UMNet can adaptively select reliable labels for SOD network learning. Extensive experiments on benchmark datasets demonstrate that our method not only outperforms existing unsupervised methods, but also is on par with fully-supervised state-of-the-art models.
Yifan Wang 0004, Lijun Wang 0001, Ting Liu 0018, Huchuan Lu
CVPR1
2022 Multi-graph convolutional clustering network
abstract
Abstract The relationship between objects can be described from different angles. Although multiple kinds of relationships make the connections between objects complex, they bring in more discriminative information for the clustering tasks. Therefore, how to effectively fuse multiple kinds of relationships becomes a critical problem. In this paper, we propose a novel Multi‐graph Convolutional Clustering Network which deeply explores the feature information of nodes and fuses the multiple kinds of relationships between nodes. Unlike most graph convolutional clustering methods that only exploit the single graph or directly fuse multiple graphs into a unified graph before the graph convolution operation, we firstly build multiple parallelled graph convolution layers for each graph to learn diverse data representations, which fully exploits different statistics information between graphs. Then, a designed multi‐graph attention module fuses above data representations and considers the importance of each graph. Besides, the proposed model completes the transition from single graph to multiple graphs, which reduces the dependence of the quality of the single graph and enhances the robustness to graphs. Experimental results verify that the proposed multi‐graph convolution clustering performs better than the traditional single‐graph convolution clustering.
Boyue Wang, Yifan Wang 0004, Xiaxia He, Yongli Hu
IET Signal Process.2
2022 From Pixels to Semantics: Self-Supervised Video Object Segmentation With Multiperspective Feature Mining
abstract
Existing self-supervised methods pose one-shot video object segmentation (O-VOS) as pixel-level matching to enable segmentation mask propagation across frames. However, the two tasks are not fully equivalent since O-VOS is more reliant on semantic correspondence rather than accurate pixel matching. To remedy this issue, we explore a new self-supervised framework that integrates pixel-level correspondence learning with semantic-level adaptation. The pixel-level correspondence learning is performed through photometric reconstruction of adjacent RGB frames during offline training, while semantic-level adaption operates at test-time by enforcing a bi-directional agreement of the predicted segmentation masks. In addition, we further propose a new network architecture with multi-perspective feature mining mechanism which can not only enhance reliable features but also suppress noisy ones to facilitate more robust image matching. By training the network using the proposed self-supervised framework, we achieve state-of-the-art performance on widely adopted datasets, further closing up the gap between self-supervised learning methods and their fully supervised counterparts.
Ruoqi Li, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Xiaopeng Wei, Qiang Zhang 0008
IEEE Trans. Image Process.2
2021 Can Scale-Consistent Monocular Depth Be Learned in a Self-Supervised Scale-Invariant Manner?
abstract
Geometric constraints are shown to enforce scale consistency and remedy the scale ambiguity issue in self-supervised monocular depth estimation. Meanwhile, scale-invariant losses focus on learning relative depth, leading to accurate relative depth prediction. To combine the best of both worlds, we learn scale-consistent self-supervised depth in a scale-invariant manner. Towards this goal, we present a scale-aware geometric (SAG) loss, which enforces scale consistency through point cloud alignment. Compared to prior arts, SAG loss takes relative scale into consideration during relative motion estimation, enabling more precise alignment and explicit supervision for scale inference. In addition, a novel two-stream architecture for depth estimation is designed, which disentangles scale from depth estimation and allows depth to be learned in a scale-invariant manner. The integration of SAG loss and two-stream network enables more consistent scale inference and more accurate relative depth estimation. Our method achieves state-of-the-art performance under both scale-invariant and scale-dependent evaluation settings.
Lijun Wang 0001, Yifan Wang 0004, Linzhao Wang, Yunlong Zhan, Huchuan Lu
ICCV2
2021 Temporal consistent portrait video segmentation
Yifan Wang 0004, Lijun Wang 0001, Fenghua Yang, Huchuan Lu
Pattern Recognit.1
2021 CSANet for Video Semantic Segmentation With Inter-Frame Mutual Learning
abstract
Video semantic segmentation aims atgenerating temporal consistent segmentation results and is still a very challenging task in the deep learning era. In this work, we improve prior approaches from two aspects. On the network architecture level, we present the cross and self-attention network (CSANet). As opposed to prior methods, CSANet not only propagates temporal features from adjacent frames, but is also designed to aggregate spatial context within the current frame, which is shown to effectively improve the consistency and robustness of the extracted deep features. On the loss function level, we further propose the inter-frame mutual learning strategy which ensures the cross-attention module to focus on semantically correlated context regions, allowing the segmentation results at different frames to be collaboratively improved. By combining the above two novel designs, we show that our proposed method is able to deliver state-of-the-art performance on the Cityscapes and CamVid benchmarks.
Lijun Wang 0001, Yifan Wang 0004
IEEE Signal Process. Lett.3
2020 CLIFFNet for Monocular Depth Estimation with Hierarchical Embedding Loss
Lijun Wang 0001, Jianming Zhang 0001, Yifan Wang 0004, Huchuan Lu, Xiang Ruan
ECCV (5)3
2020 Blind single image super-resolution with a mixture of deep networks
Yifan Wang 0004, Lijun Wang 0001, Hongyu Wang 0001, Peihua Li, Huchuan Lu
Pattern Recognit.1
2019 Resolution-Aware Network for Image Super-Resolution
abstract
In existing deep network-based image super-resolution (SR) methods, each network is only trained for a fixed upscaling factor and can hardly generalize to unseen factors at test time, which is non-scalable in real applications. To mitigate this issue, this paper proposes a resolution-aware network (RAN) for simultaneous SR of multiple factors. The key insight is that SR of multiple factors is essentially different but also shares common operations. To attain stronger generalization across factors, we design an upsampling network (U-Net) consisting of several sub-modules, in which each sub-module implements an intermediate step of the overall image SR and can be shared by SR of different factors. A decision network (D-Net) is further adopted to identify the quality of the input low-resolution image and adaptively select suitable sub-modules to perform SR. U-Net and D-Net together constitute the proposed RAN model, and are jointly trained using a new hierarchical loss function on SR tasks of multiple factors. Experimental evaluations demonstrate that the proposed RAN compares favorably against the state-of-the-art methods and its performance can well generalize across different upscaling factors.
Yifan Wang 0004, Lijun Wang 0001, Hongyu Wang 0001, Peihua Li
IEEE Trans. Circuits Syst. Video Technol.1
2018 Information-Compensated Downsampling for Image Super-Resolution
abstract
A large receptive field of deep networks can better incorporate image context and benefits image super-resolution (SR) in many ways. However, common techniques, like strided pooling and convolutional operations, are not directly applicable to SR due to severe image detail losses. In this letter, we circumvent this issue by proposing a new network architecture, namely the information-compensated (IC) downsampling block. It first uses pooling layers to downsample input feature maps and then immediately upsamples the feature maps back to the original size. To further compensate for information loss, skip connections are added to propagate lost features caused by downsampling to the upsampled output. In addition, pixelwise recurrent units are also applied to the downsampled feature maps to model context coherence. Compared with traditional pooling layers, the IC downsampling blocks cannot only enlarge receptive field and better capture image context, but also preserve image details, which are essential to SR. The final network consists of a stack of IC downsampling blocks and can be trained in an end-to-end manner. Experimental results verify that the proposed method performs favorably against the state-of-the-art approaches.
Yifan Wang 0004, Lijun Wang 0001, Hongyu Wang 0001, Peihua Li
IEEE Signal Process. Lett.1
2017 Learning to Detect Salient Objects with Image-Level Supervision
abstract
Deep Neural Networks (DNNs) have substantially improved the state-of-the-art in salient object detection. However, training DNNs requires costly pixel-level annotations. In this paper, we leverage the observation that image-level tags provide important cues of foreground salient objects, and develop a weakly supervised learning method for saliency detection using image-level tags only. The Foreground Inference Network (FIN) is introduced for this challenging task. In the first stage of our training method, FIN is jointly trained with a fully convolutional network (FCN) for image-level tag prediction. A global smooth pooling layer is proposed, enabling FCN to assign object category tags to corresponding object regions, while FIN is capable of capturing all potential foreground regions with the predicted saliency maps. In the second stage, FIN is fine-tuned with its predicted saliency maps as ground truth. For refinement of ground truth, an iterative Conditional Random Field is developed to enforce spatial label consistency and further boost performance. Our method alleviates annotation efforts and allows the usage of existing large scale training sets with image-level tags. Our model runs at 60 FPS, outperforms unsupervised ones with a large margin, and achieves comparable or even superior performance than fully supervised counterparts.
Lijun Wang 0001, Huchuan Lu, Yifan Wang 0004, Mengyang Feng, Dong Wang 0004, Xiang Ruan
CVPR3
2016 Biologically inspired image enhancement based on Retinex
Yifan Wang 0004, Hongyu Wang 0001, Chuanli Yin
Neurocomputing1