VLDB 2026 Research / reviewers in the wild / expert
Lijun Wang 0001
dblp:96/6702-1
· DBLP profile ↗
53ranked-venue papers
10as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 8 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 9 first-author · 25 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Aggregating global-scale pixel-wise forgery cues within a graph
Hengrun Zhao, Yifan Wang 0004, Yunzhi Zhuge, Lijun Wang 0001, Huchuan Lu |
Neural Networks | 4 |
| 2025 | Mono2Stereo: A Benchmark and Empirical Study for Stereo ConversionabstractWith the rapid proliferation of 3D devices and the shortage of 3D content, stereo conversion is attracting increasing attention. Recent works introduce pretrained Diffusion Models (DMs) into this task. However, due to the scarcity of large-scale training data and comprehensive benchmarks, the optimal methodologies for employing DMs in stereo conversion and the accurate evaluation of stereo effects remain largely unexplored. In this work, we introduce the Mono2Stereo dataset, providing high-quality training data and benchmark to support in-depth exploration of stereo conversion. With this dataset, we conduct an empirical study that yields two primary findings. 1) The differences between the left and right views are subtle, yet existing metrics consider overall pixels, failing to concentrate on regions critical to stereo effects. 2) Mainstream methods adopt either one-stage left-to-right generation or warp-and-inpaint pipeline, facing challenges of degraded stereo effect and image distortion respectively. Based on these findings, we introduce a new evaluation metric, Stereo Intersection-over-Union, which prioritizes disparity and achieves a high correlation with human judgments on stereo effect. Moreover, we propose a strong baseline model, harmonizing the stereo effect and image quality simultaneously, and notably surpassing current mainstream methods. Our code and data will be open-sourced to promote further research in stereo conversion. Our models are available at mono2stereo-bench.github.io. Songsong Yu, Zhongang Qi, Zeke Xie, Yifan Wang 0004, Lijun Wang 0001, Ying Shan, Huchuan Lu |
CVPR | 6 |
| 2025 | GLDesigner: Leveraging Multi-Modal LLMs as Designer for Enhanced Aesthetic Text Glyph LayoutsabstractText logo design heavily relies on the creativity and expertise of professional designers, in which arranging element layouts is one of the most important procedures. However, this specific task has received limited attention, often overshadowed by broader layout generation tasks such as document or poster design. In this paper, we propose a Vision-Language Model (VLM)-based framework that generates content-aware text logo layouts by integrating multi-modal inputs with user-defined constraints, enabling more flexible and robust layout generation for real-world applications. We introduce two model techniques that reduce the computational cost for processing multiple glyph images simultaneously, without compromising performance. To support instruction tuning of our model, we construct two extensive text logo datasets that are five times larger than existing public datasets. In addition to geometric annotations (e.g., text masks and character recognition), our datasets include detailed layout descriptions in natural language, enabling the model to reason more effectively in handling complex designs and custom user inputs. Experimental results demonstrate the effectiveness of our proposed framework and datasets, outperforming existing methods on various benchmarks that assess geometric aesthetics and human preferences. Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Chenyang Li 0007, Jin-Peng Lan, Jun-Yan He, Bin Luo 0008, Yifeng Geng |
ACM Multimedia | 3 |
| 2025 | From Forecasting to Planning: Policy World Model for Collaborative State-Action PredictionabstractDespite remarkable progress in driving world models, their potential for autonomous systems remains largely untapped: the world models are mostly learned for world simulation and decoupled from trajectory planning. While recent efforts aim to unify world modeling and planning in a single framework, the synergistic facilitation mechanism of world modeling for planning still requires further exploration. In this work, we introduce a new driving paradigm named Policy World Model (PWM), which not only integrates world modeling and trajectory planning within a unified architecture, but is also able to benefit planning using the learned world knowledge through the proposed action-free future state forecasting scheme. Through collaborative state-action prediction, PWM can mimic the human-like anticipatory perception, yielding more reliable planning performance. To facilitate the efficiency of video forecasting, we further introduce a parallel token generation mechanism, equipped with a context-guided tokenizer and an adaptive dynamic focal loss. Despite utilizing only front camera input, our method matches or exceeds state-of-the-art approaches that rely on multi-view and multi-modal inputs. Code will be released at https://github.com/6550Zhao/Policy-World-Model. Zhida Zhao, Talas Fu, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu |
NeurIPS | 4 |
| 2025 | Learning pose regression as reliable pixel-level matching for self-supervised depth estimation
Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu |
Neurocomputing | 3 |
| 2025 | AVS-Mamba: Exploring Temporal and Multi-Modal Mamba for Audio-Visual SegmentationabstractThe essence of audio-visual segmentation (AVS) lies in locating and delineating sound-emitting objects within a video stream. While Transformer-based methods have shown promise, their handling of long-range dependencies struggles due to quadratic computational costs, presenting a bottleneck in complex scenarios. To overcome this limitation and facilitate complex multi-modal comprehension with linear complexity, we introduce AVS-Mamba, a selective state space model to address the AVS task. Our framework incorporates two key components for video understanding and cross-modal learning: Temporal Mamba Block for sequential video processing and Vision-to-Audio Fusion Block for advanced audio-vision integration. Building on this, we develop the Multi-scale Temporal Encoder, aimed at enhancing the learning of visual features across scales, facilitating the perception of intra- and inter-frame information. To perform multi-modal fusion, we propose the Modality Aggregation Decoder, leveraging the Vision-to-Audio Fusion Block to integrate visual features into audio features across both frame and temporal levels. Further, we adopt the Contextual Integration Pyramid to perform audio-to-vision spatial-temporal context collaboration. Through these innovative contributions, our approach achieves new state-of-the-art results on the AVSBench-object and AVSBench-semantic datasets. Sitong Gong, Yunzhi Zhuge, Lu Zhang 0053, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu |
IEEE Trans. Multim. | 6 |
| 2024 | DME: Unveiling the Bias for Better Generalized Monocular Depth EstimationabstractThis paper aims to design monocular depth estimation models with better generalization abilities. To this end, we have conducted quantitative analysis and discovered two important insights. First, the Simulation Correlation phenomenon, commonly seen in long-tailed classification problems, also exists in monocular depth estimation, indicating that the imbalanced depth distribution in training data may be the cause of limited generalization ability. Second, the imbalanced and long-tail distribution of depth values extends beyond the dataset scale, and also manifests within each individual image, further exacerbating the challenge of monocular depth estimation. Motivated by the above findings, we propose the Distance-aware Multi-Expert (DME) depth estimation model. Unlike prior methods that handle different depth range indiscriminately, DME adopts a divide-and-conquer philosophy where each expert is responsible for depth estimation of regions within a specific depth range. As such, the depth distribution seen by each expert is more uniform and can be more easily predicted. A pixel-level routing module is further designed and learned to stitch the prediction of all experts into the final depth map. Experiments show that DME achieves state-of-the-art performance on both NYU-Depth v2 and KITTI, and also delivers favorable zero-shot generalization capability on unseen datasets. Songsong Yu, Yifan Wang 0004, Yunzhi Zhuge, Lijun Wang 0001, Huchuan Lu |
AAAI | 4 |
| 2024 | Large Occluded Human Image Completion via Image-Prior CooperatingabstractThe completion of large occluded human body images poses a unique challenge for general image completion methods. The complex shape variations of human bodies make it difficult to establish a consistent understanding of their structures. Furthermore, as human vision is highly sensitive to human bodies, even slight artifacts can significantly compromise image fidelity. To address these challenges, we propose a large occluded human image completion (LOHC) model based on a novel image-prior cooperative completion strategy. Our model leverages human segmentation maps as a prior, and completes the image and prior simultaneously. Compared to the widely adopted prior-then-image completion strategy for object completion, this cooperative completion process fosters more effective interaction between the prior and image information. Our model consists of two stages. The first stage is a transformer-based auto-regressive network that predicts the overall structure of the missing area by generating a coarse completed image at a lower resolution. The second stage is a convolutional network that refines the coarse images. As the coarse result may not always be accurate, we propose a Dynamic Fusion Module (DFM) to selectively fuses the useful features from the coarse image with the original input at spatial and channel levels. Through extensive experiments, we demonstrate our method’s superior performance compared to state-of-the-art methods. Hengrun Zhao, Yu Zeng 0001, Huchuan Lu, Lijun Wang 0001 |
AAAI | 4 |
| 2024 | 3D Prompt Learning for RGB-D Tracking
Bocen Li, Yunzhi Zhuge, Lijun Wang 0001, Yifan Wang 0004, Huchuan Lu |
ACCV (2) | 4 |
| 2024 | PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System SafetyabstractZaibin Zhang, Yongting Zhang, Lijun Li, Hongzhi Gao, Lijun Wang, Huchuan Lu, Feng Zhao, Yu Qiao, Jing Shao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zaibin Zhang, Yongting Zhang, Hongzhi Gao, Yu Qiao 0001, Lijun Wang 0001, Huchuan Lu, Feng Zhao 0004 |
ACL (1) | 7 |
| 2024 | Multi-Modal Instruction Tuned LLMs with Fine-Grained Visual PerceptionabstractMultimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs with visual perceiving and grounding capabilities. However, there still remains a gap in providing fine-grained pixel-level perceptions and extending interactions beyond text-specific inputs. In this work, we propose AnyRef, a general MLLM model that can generate pixel-wise object perceptions and natural language descriptions from multi-modality references, such as texts, boxes, images, or audio. This innovation empowers users with greater flexibility to engage with the model beyond textual and regional prompts, without modality-specific designs. Through our proposed refocusing mechanism, the generated grounding output is guided to better focus on the referenced object, implicitly incorporating additional pixel-level supervision. This simple modification utilizes attention scores generated during the inference of LLM, eliminating the need for extra computations while exhibiting performance enhancements in both grounding masks and referring expressions. With only publicly available training data, our model achieves state-of-the-art results across multiple benchmarks, including diverse modality referring segmentation and region-level referring expression generation. Code and models are available at https://github.com/jwh97nn/AnyRef Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Jun-Yan He, Jin-Peng Lan, Bin Luo 0008, Xuansong Xie |
CVPR | 3 |
| 2024 | SelM: Selective Mechanism based Audio-Visual SegmentationabstractAudio-Visual Segmentation (AVS) aims to segment sound-producing objects in videos according to associated audio cues, where both modalities are affected by noise to different extents, such as the blending of background noises in audio or the presence of distracted objects in video. Most existing methods focus on learning interactions between modalities at high semantic levels but is incapable of filtering low-level noise or achieving fine-grained representational interactions during the early feature extraction phase. Consequently, they struggle with illusion issues, where nonexistent audio cues are erroneously linked to visual objects. In this paper, we present SelM, a novel architecture that leverages selective mechanisms to counteract these illusions. SelM employs State Space model for noise reduction and robust feature selection. By imposing additional bidirectional constraints on audio and visual embeddings, it is able to precisely identify crucial features corresponding to sound-emitting targets. To fill the existing gap in early fusion within AVS, SelM introduces a dual alignment mechanism specifically engineered to facilitate intricate spatio-temporal interactions between audio and visual streams, achieving more fine-grained representations. Moreover, we develop a cross-level decoder for layered reasoning, significantly enhancing segmentation precision by exploring the complex relationships between audio and visual information. SelM achieves state-of-the-art performance in AVS tasks, especially in the challenging Audio-Visual Semantic Segmentation subset. The code can be found at https://github.com/Cyyzpoi/SelM. Songsong Yu, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu |
ACM Multimedia | 4 |
| 2024 | LOVD: Large-and-Open Vocabulary Object DetectionabstractExisting open-vocabulary object detectors require an accurate and compact vocabulary pre-defined during inference. Their performance is largely degraded in real scenarios where the underlying vocabulary may be indeterminate and often exponentially large. To have a more comprehensive understanding of this phenomenon, we propose a new setting called Large-and-Open Vocabulary object Detection, which simulates real scenarios by testing detectors with large vocabularies containing thousands of unseen categories. The vast unseen categories inevitably lead to an increase in category distractors, severely impeding the recognition process and leading to unsatisfactory detection results. To address this challenge, We propose a Large and Open Vocabulary Detector (LOVD) with two core components, termed the Image-to-Region Filtering (IRF) module and Cross-View Verification (CV2) scheme. To relieve the category distractors of the given large vocabularies, IRF performs image-level recognition to build a compact vocabulary relevant to the image scene out of the large input vocabulary, followed by region-level classification upon the compact vocabulary. CV2 further enhances the IRF by conducting image-to-region filtering in both global and local views and produces the final detection categories through a two-branch voting mechanism. Compared to the prior works, our LOVD is more scalable and robust to large input vocabularies, and can be seamlessly integrated with predominant detection methods to improve their open-vocabulary performance. The code can be found at https://github.com/Altria-luo/LOVD. Shiyu Tang, Zhaofan Luo, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Weibo Su |
ACM Multimedia | 4 |
| 2024 | MaskMentor: Unlocking the Potential of Masked Self-Teaching for Missing Modality RGB-D Semantic SegmentationabstractExisting RGB-D semantic segmentation methods struggle to handle modality missing input, where only RGB images or depth maps are available, leading to degenerated segmentation performance. We tackle this issue using MaskMentor, a new pre-training framework for modality missing segmentation, which advances its counterparts via two novel designs: Masked Modality and Image Modeling (M2IM), and Self-Teaching via Token-Pixel Joint reconstruction (STTP). M2IM simulates modality missing scenarios by combining both modality- and patch-level random masking. Meanwhile, STTP offers an effective self-teaching strategy, where the trained network assumes a dual role, simultaneously acting as both the teacher and the student. The student with modality missing input is supervised by the teacher with complete modality input through both token- and pixel-wise masked modeling, closing the gap between missing and complete input modalities. By integrating M2IM and STTP, MaskMentor significantly improves the generalization ability of the trained model across diverse input conditions and outperforms state-of-the-art methods on two popular benchmarks by a considerable margin. Extensive ablation studies further verify the effectiveness of the above contributions. Zhida Zhao, Lijun Wang 0001, Yifan Wang 0004, Huchuan Lu |
ACM Multimedia | 3 |
| 2024 | MOFTrack: Multi-object Formation Tracking in Remote Sensing Videos
Haijiang Sun, Qiaoyuan Liu, Tanlin Li, Lijun Wang 0001, Yifan Wang 0004 |
PRCV (12) | 6 |
| 2024 | CSRNet: Focusing on critical points for depth completion
Bocen Li, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu |
Image Vis. Comput. | 3 |
| 2024 | ITrans: generative image inpainting with transformersabstractAbstract Despite significant improvements, convolutional neural network (CNN) based methods are struggling with handling long-range global image dependencies due to their limited receptive fields, leading to an unsatisfactory inpainting performance under complicated scenarios. To address this issue, we propose the Inpainting Transformer (ITrans) network, which combines the power of both self-attention and convolution operations. The ITrans network augments convolutional encoder–decoder structure with two novel designs, i.e. , the global and local transformers. The global transformer aggregates high-level image context from the encoder in a global perspective, and propagates the encoded global representation to the decoder in a multi-scale manner. Meanwhile, the local transformer is intended to extract low-level image details inside the local neighborhood at a reduced computational overhead. By incorporating the above two transformers, ITrans is capable of both global relationship modeling and local details encoding, which is essential for hallucinating perceptually realistic images. Extensive experiments demonstrate that the proposed ITrans network outperforms favorably against state-of-the-art inpainting methods both quantitatively and qualitatively. Wei Miao 0006, Lijun Wang 0001, Huchuan Lu, Kaining Huang, Xinchu Shi, Bocong Liu |
Multim. Syst. | 2 |
| 2023 | ARKitTrack: A New Diverse Dataset for Tracking Using Mobile RGB-D DataabstractCompared with traditional RGB-only visual tracking, few datasets have been constructed for RGB-D tracking. In this paper, we propose ARKitTrack, a new RGB-D tracking dataset for both static and dynamic scenes captured by consumer-grade LiDAR scanners equipped on Apple's iPhone and iPad. ARKitTrack contains 300 RGB-D sequences, 455 targets, and 229.7K video frames in total. Along with the bounding box annotations and frame-level attributes, we also annotate this dataset with 123.9K pixel-level target masks. Besides, the camera intrinsic and camera pose of each frame are provided for future developments. To demonstrate the potential usefulness of this dataset, we further present a unified baseline for both box-level and pixel-level tracking, which integrates RGB features with bird's-eye-view representations to better explore cross-modality 3D geometry. In-depth empirical analysis has verified that the ARKitTrack dataset can significantly facilitate RGB-D tracking and that the proposed baseline method compares favorably against the state of the arts. The code and dataset is available at https://arkittrack.github.io. Haojie Zhao, Junsong Chen, Lijun Wang 0001, Huchuan Lu |
CVPR | 3 |
| 2023 | Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance LearningabstractDepth-aware panoptic segmentation is an emerging topic in computer vision which combines semantic and geometric understanding for more robust scene interpretation. Recent works pursue unified frameworks to tackle this challenge but mostly still treat it as two individual learning tasks, which limits their potential for exploring cross-domain information. We propose a deeply unified framework for depth-aware panoptic segmentation, which performs joint segmentation and depth estimation both in a persegment manner with identical object queries. To narrow the gap between the two tasks, we further design a geometric query enhancement method, which is able to integrate scene geometry into object queries using latent representations. In addition, we propose a bi-directional guidance learning approach to facilitate cross-task feature learning by taking advantage of their mutual relations. Our method sets the new state of the art for depth-aware panoptic segmentation on both Cityscapes-DVPS and SemKITTI-DVPS datasets. Moreover, our guidance learning approach is shown to deliver performance improvement even under incomplete supervision labels. Code and models are available at https://github.com/jwh97nn/DeepDPS. Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Bin Luo 0008, Jun-Yan He, Jin-Peng Lan, Yifeng Geng, Xuansong Xie |
ICCV | 3 |
| 2023 | Isomer: Isomerous Transformer for Zero-shot Video Object SegmentationabstractRecent leading zero-shot video object segmentation (ZVOS) works devote to integrating appearance and motion information by elaborately designing feature fusion modules and identically applying them in multiple feature stages. Our preliminary experiments show that with the strong long-range dependency modeling capacity of Transformer, simply concatenating the two modality features and feeding them to vanilla Transformers for feature fusion can distinctly benefit the performance but at a cost of heavy computation. Through further empirical analysis, we find that attention dependencies learned in Transformer in different stages exhibit completely different properties: global query-independent dependency in the low-level stages and semantic-specific dependency in the high-level stages. Motivated by the observations, we propose two Transformer variants: i) Context-Sharing Transformer (CST) that learns the global-shared contextual information within image frames with a lightweight computation. ii) Semantic Gathering-Scattering Transformer (SGST) that models the semantic correlation separately for the foreground and background and reduces the computation cost with a soft token merging mechanism. We apply CST and SGST for low-level and high-level feature fusions, respectively, formulating a level-isomerous Transformer framework for ZVOS task. Compared with the baseline that uses vanilla Transformers for multi-stage fusion, ours significantly increase the speed by 13× and achieves new state-of-the-art ZVOS performance. Code is available at https://github.com/DLUT-yyc/Isomer. Yifan Wang 0004, Lijun Wang 0001, Xiaoqi Zhao 0003, Huchuan Lu, Yu Wang 0108, Weibo Su, Lei Zhang 0006 |
ICCV | 3 |
| 2023 | SACFormer: Unify Depth Estimation and Completion with Prompt
Shiyu Tang, Yifan Wang 0004, Lijun Wang 0001 |
PRCV (2) | 4 |
| 2022 | You Only Infer Once: Cross-Modal Meta-Transfer for Referring Video Object SegmentationabstractWe present YOFO (You Only inFer Once), a new paradigm for referring video object segmentation (RVOS) that operates in an one-stage manner. Our key insight is that the language descriptor should serve as target-specific guidance to identify the target object, while a direct feature fusion of image and language can increase feature complexity and thus may be sub-optimal for RVOS. To this end, we propose a meta-transfer module, which is trained in a learning-to-learn fashion and aims to transfer the target-specific information from the language domain to the image domain, while discarding the uncorrelated complex variations of language description. To bridge the gap between the image and language domains, we develop a multi-scale cross-modal feature mining block that aggregates all the essential features required by RVOS from both domains and generates regression labels for the meta-transfer module. The whole system can be trained in an end-to-end manner and shows competitive performance against state-of-the-art two-stage approaches. Dezhuang Li, Ruoqi Li, Lijun Wang 0001, Yifan Wang 0004, Jinqing Qi, Lu Zhang 0053, Ting Liu 0018, Qingquan Xu, Huchuan Lu |
AAAI | 3 |
| 2022 | Multi-Source Uncertainty Mining for Deep Unsupervised Saliency DetectionabstractDeep learning-based image salient object detection (SOD) heavily relies on large-scale training data with pixel-wise labeling. High-quality labels involve intensive labor and are expensive to acquire. In this paper, we propose a novel multi-source uncertainty mining method to facilitate unsupervised deep learning from multiple noisy labels generated by traditional handcrafted SOD methods. We design an Uncertainty Mining Network (UMNet) which consists of multiple Merge-and-Split (MS) modules to recursively analyze the commonality and difference among multiple noisy labels and infer pixel-wise uncertainty map for each label. Meanwhile, we model the noisy labels using Gibbs distribution and propose a weighted uncertainty loss to jointly train the UMNet with the SOD network. As a consequence, our UMNet can adaptively select reliable labels for SOD network learning. Extensive experiments on benchmark datasets demonstrate that our method not only outperforms existing unsupervised methods, but also is on par with fully-supervised state-of-the-art models. Yifan Wang 0004, Lijun Wang 0001, Ting Liu 0018, Huchuan Lu |
CVPR | 3 |
| 2022 | Adaptive Co-teaching for Unsupervised Monocular Depth Estimation
Weisong Ren, Lijun Wang 0001, Yongri Piao, Miao Zhang 0004, Huchuan Lu, Ting Liu 0018 |
ECCV (1) | 2 |
| 2022 | MVSalNet: Multi-view Augmentation for RGB-D Salient Object Detection
Jiayuan Zhou, Lijun Wang 0001, Huchuan Lu, Kaining Huang, Xinchu Shi, Bocong Liu |
ECCV (29) | 2 |
| 2022 | Road extraction from satellite images with iterative cross-task feature enhancement
Weiling Yin, Mingyang Qian, Lijun Wang 0001, Jinqing Qi, Huchuan Lu |
Neurocomputing | 3 |
| 2022 | From Pixels to Semantics: Self-Supervised Video Object Segmentation With Multiperspective Feature MiningabstractExisting self-supervised methods pose one-shot video object segmentation (O-VOS) as pixel-level matching to enable segmentation mask propagation across frames. However, the two tasks are not fully equivalent since O-VOS is more reliant on semantic correspondence rather than accurate pixel matching. To remedy this issue, we explore a new self-supervised framework that integrates pixel-level correspondence learning with semantic-level adaptation. The pixel-level correspondence learning is performed through photometric reconstruction of adjacent RGB frames during offline training, while semantic-level adaption operates at test-time by enforcing a bi-directional agreement of the predicted segmentation masks. In addition, we further propose a new network architecture with multi-perspective feature mining mechanism which can not only enhance reliable features but also suppress noisy ones to facilitate more robust image matching. By training the network using the proposed self-supervised framework, we achieve state-of-the-art performance on widely adopted datasets, further closing up the gap between self-supervised learning methods and their fully supervised counterparts. Ruoqi Li, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Xiaopeng Wei, Qiang Zhang 0008 |
IEEE Trans. Image Process. | 3 |
| 2021 | Video Annotation for Visual Tracking via Selection and RefinementabstractDeep learning based visual trackers entail offline pre-training on large volumes of video datasets with accurate bounding box annotations that are labor-expensive to achieve. We present a new framework to facilitate bounding box annotations for video sequences, which investigates a selection-and-refinement strategy to automatically improve the preliminary annotations generated by tracking algorithms. A temporal assessment network (T-Assess Net) is proposed which is able to capture the temporal coherence of target locations and select reliable tracking results by measuring their quality. Meanwhile, a visual-geometry refinement network (VG-Refine Net) is also designed to further enhance the selected tracking results by considering both target appearance and temporal geometry constraints, allowing inaccurate tracking results to be corrected. The combination of the above two networks provides a principled approach to ensure the quality of automatic video annotation. Experiments on large scale tracking benchmarks demonstrate that our method can deliver highly accurate bounding box annotations and significantly reduce human labor by 94.0%, yielding an effective means to further boost tracking performance with augmented training data. Kenan Dai, Jie Zhao 0014, Lijun Wang 0001, Dong Wang 0004, Huchuan Lu, Xuesheng Qian, Xiaoyun Yang |
ICCV | 3 |
| 2021 | Can Scale-Consistent Monocular Depth Be Learned in a Self-Supervised Scale-Invariant Manner?abstractGeometric constraints are shown to enforce scale consistency and remedy the scale ambiguity issue in self-supervised monocular depth estimation. Meanwhile, scale-invariant losses focus on learning relative depth, leading to accurate relative depth prediction. To combine the best of both worlds, we learn scale-consistent self-supervised depth in a scale-invariant manner. Towards this goal, we present a scale-aware geometric (SAG) loss, which enforces scale consistency through point cloud alignment. Compared to prior arts, SAG loss takes relative scale into consideration during relative motion estimation, enabling more precise alignment and explicit supervision for scale inference. In addition, a novel two-stream architecture for depth estimation is designed, which disentangles scale from depth estimation and allows depth to be learned in a scale-invariant manner. The integration of SAG loss and two-stream network enables more consistent scale inference and more accurate relative depth estimation. Our method achieves state-of-the-art performance under both scale-invariant and scale-dependent evaluation settings. Lijun Wang 0001, Yifan Wang 0004, Linzhao Wang, Yunlong Zhan, Huchuan Lu |
ICCV | 1 |
| 2021 | IPE Transformer for Depth Completion with Input-Aware Positional Embeddings
Bocen Li, Guozhen Li, Haiting Wang, Lijun Wang 0001, Zhenfei Gong, Huchuan Lu |
PRCV (4) | 4 |
| 2021 | Learning Regression and Verification Networks for Robust Long-term Tracking
Lijun Wang 0001, Dong Wang 0004, Jinqing Qi, Huchuan Lu |
Int. J. Comput. Vis. | 2 |
| 2021 | Temporal consistent portrait video segmentation
Yifan Wang 0004, Lijun Wang 0001, Fenghua Yang, Huchuan Lu |
Pattern Recognit. | 3 |
| 2021 | CSANet for Video Semantic Segmentation With Inter-Frame Mutual LearningabstractVideo semantic segmentation aims atgenerating temporal consistent segmentation results and is still a very challenging task in the deep learning era. In this work, we improve prior approaches from two aspects. On the network architecture level, we present the cross and self-attention network (CSANet). As opposed to prior methods, CSANet not only propagates temporal features from adjacent frames, but is also designed to aggregate spatial context within the current frame, which is shown to effectively improve the consistency and robustness of the extracted deep features. On the loss function level, we further propose the inter-frame mutual learning strategy which ensures the cross-attention module to focus on semantically correlated context regions, allowing the segmentation results at different frames to be collaboratively improved. By combining the above two novel designs, we show that our proposed method is able to deliver state-of-the-art performance on the Cityscapes and CamVid benchmarks. Lijun Wang 0001, Yifan Wang 0004 |
IEEE Signal Process. Lett. | 2 |
| 2020 | SDC-Depth: Semantic Divide-and-Conquer Network for Monocular Depth EstimationabstractMonocular depth estimation is an ill-posed problem, and as such critically relies on scene priors and semantics. Due to its complexity, we propose a deep neural network model based on a semantic divide-and-conquer approach. Our model decomposes a scene into semantic segments, such as object instances and background stuff classes, and then predicts a scale and shift invariant depth map for each semantic segment in a canonical space. Semantic segments of the same category share the same depth decoder, so the global depth prediction task is decomposed into a series of category-specific ones, which are simpler to learn and easier to generalize to new scene types. Finally, our model stitches each local depth segment by predicting its scale and shift based on the global context of the image. The model is trained end-to-end using a multi-task loss for panoptic segmentation and depth prediction, and is therefore able to leverage large-scale panoptic segmentation datasets to boost its semantic understanding. We validate the effectiveness of our approach and show state-of-the-art performance on three benchmark datasets. Lijun Wang 0001, Jianming Zhang 0001, Oliver Wang, Zhe Lin 0001, Huchuan Lu |
CVPR | 1 |
| 2020 | CLIFFNet for Monocular Depth Estimation with Hierarchical Embedding Loss
Lijun Wang 0001, Jianming Zhang 0001, Yifan Wang 0004, Huchuan Lu, Xiang Ruan |
ECCV (5) | 1 |
| 2020 | Segmentation based rotated bounding boxes prediction and image synthesizing for object detection of high resolution aerial images
Lijun Wang 0001, Huchuan Lu, You He 0002 |
Neurocomputing | 2 |
| 2020 | Blind single image super-resolution with a mixture of deep networks
Yifan Wang 0004, Lijun Wang 0001, Hongyu Wang 0001, Peihua Li, Huchuan Lu |
Pattern Recognit. | 2 |
| 2019 | Salient Object Detection with Recurrent Fully Convolutional NetworksabstractDeep networks have been proved to encode high-level features with semantic meaning and delivered superior performance in salient object detection. In this paper, we take one step further by developing a new saliency detection method based on recurrent fully convolutional networks (RFCNs). Compared with existing deep network based methods, the proposed network is able to incorpor- ate saliency prior knowledge for more accurate inference. In addition, the recurrent architecture enables our method to automatically learn to refine the saliency map by iteratively correcting its previous errors, yielding more reliable final predictions. To train such a netw- ork with numerous parameters, we propose a pre-training strategy using semantic segmentation data, which simultaneously leverages the strong supervision of segmentation tasks for effective training and enables the network to capture generic representations to chara- cterize category-agnostic objects for saliency detection. Extensive experimental evaluations demonstrate that the proposed method compares favorably against state-of-the-art saliency detection approaches. Additional validations are also performed to study the impact of the recurrent architecture and pre-training strategy on both saliency detection and semantic segmentation, which provides important knowledge for network design and training in the future research. Linzhao Wang, Lijun Wang 0001, Huchuan Lu, Xiang Ruan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Resolution-Aware Network for Image Super-ResolutionabstractIn existing deep network-based image super-resolution (SR) methods, each network is only trained for a fixed upscaling factor and can hardly generalize to unseen factors at test time, which is non-scalable in real applications. To mitigate this issue, this paper proposes a resolution-aware network (RAN) for simultaneous SR of multiple factors. The key insight is that SR of multiple factors is essentially different but also shares common operations. To attain stronger generalization across factors, we design an upsampling network (U-Net) consisting of several sub-modules, in which each sub-module implements an intermediate step of the overall image SR and can be shared by SR of different factors. A decision network (D-Net) is further adopted to identify the quality of the input low-resolution image and adaptively select suitable sub-modules to perform SR. U-Net and D-Net together constitute the proposed RAN model, and are jointly trained using a new hierarchical loss function on SR tasks of multiple factors. Experimental evaluations demonstrate that the proposed RAN compares favorably against the state-of-the-art methods and its performance can well generalize across different upscaling factors. Yifan Wang 0004, Lijun Wang 0001, Hongyu Wang 0001, Peihua Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Structured Siamese Network for Real-Time Visual Tracking
Lijun Wang 0001, Jinqing Qi, Dong Wang 0004, Mengyang Feng, Huchuan Lu |
ECCV (9) | 2 |
| 2018 | Deep visual tracking: Review and experimental comparison
Peixia Li, Dong Wang 0004, Lijun Wang 0001, Huchuan Lu |
Pattern Recognit. | 3 |
| 2018 | Information-Compensated Downsampling for Image Super-ResolutionabstractA large receptive field of deep networks can better incorporate image context and benefits image super-resolution (SR) in many ways. However, common techniques, like strided pooling and convolutional operations, are not directly applicable to SR due to severe image detail losses. In this letter, we circumvent this issue by proposing a new network architecture, namely the information-compensated (IC) downsampling block. It first uses pooling layers to downsample input feature maps and then immediately upsamples the feature maps back to the original size. To further compensate for information loss, skip connections are added to propagate lost features caused by downsampling to the upsampled output. In addition, pixelwise recurrent units are also applied to the downsampled feature maps to model context coherence. Compared with traditional pooling layers, the IC downsampling blocks cannot only enlarge receptive field and better capture image context, but also preserve image details, which are essential to SR. The final network consists of a stack of IC downsampling blocks and can be trained in an end-to-end manner. Experimental results verify that the proposed method performs favorably against the state-of-the-art approaches. Yifan Wang 0004, Lijun Wang 0001, Hongyu Wang 0001, Peihua Li |
IEEE Signal Process. Lett. | 2 |
| 2018 | Constrained Superpixel TrackingabstractIn this paper, we propose a constrained graph labeling algorithm for visual tracking where nodes denote superpixels and edges encode the underlying spatial, temporal, and appearance fitness constraints. First, the spatial smoothness constraint, based on a transductive learning method, is enforced to leverage the latent manifold structure in feature space by investigating unlabeled superpixels in the current frame. Second, the appearance fitness constraint, which measures the probability of a superpixel being contained in the target region, is developed to incrementally induce a long-term appearance model. Third, the temporal smoothness constraint is proposed to construct a short-term appearance model of the target, which handles the drastic appearance change between consecutive frames. All these three constraints are incorporated in the proposed graph labeling algorithm such that induction and transduction, short- and long-term appearance models are combined, respectively. The foreground regions inferred by the proposed graph labeling method are used to guide the tracking process. Tracking results, in turn, facilitate more accurate online update by filtering out potential contaminated training samples. Both quantitative and qualitative evaluations on challenging tracking data sets show that the proposed constrained tracking algorithm performs favorably against the state-of-the-art methods. Lijun Wang 0001, Huchuan Lu, Ming-Hsuan Yang 0001 |
IEEE Trans. Cybern. | 1 |
| 2018 | DeepLens: shallow depth of field from a single imageabstractWe aim to generate high resolution shallow depth-of-field (DoF) images from a single all-in-focus image with controllable focal distance and aperture size. To achieve this, we propose a novel neural network model comprised of a depth prediction module, a lens blur module, and a guided upsampling module. All modules are differentiable and are learned from data. To train our depth prediction module, we collect a dataset of 2462 RGB-D images captured by mobile phones with a dual-lens camera, and use existing segmentation datasets to improve border prediction. We further leverage a synthetic dataset with known depth to supervise the lens blur and guided upsampling modules. The effectiveness of our system and training strategies are verified in the experiments. Our method can generate high-quality shallow DoF images at high resolution, and produces significantly fewer artifacts than the baselines and existing solutions for single image shallow DoF synthesis. Compared with the iPhone portrait mode, which is a state-of-the-art shallow DoF solution based on a dual-lens depth camera, our method generates comparable results, while allowing for greater flexibility to choose focal points and aperture size, and is not limited to one capture setup. Lijun Wang 0001, Xiaohui Shen, Jianming Zhang 0001, Oliver Wang, Zhe Lin 0001, Chih-Yao Hsieh, Sarah Kong, Huchuan Lu |
ACM Trans. Graph. | 1 |
| 2017 | Learning to Detect Salient Objects with Image-Level SupervisionabstractDeep Neural Networks (DNNs) have substantially improved the state-of-the-art in salient object detection. However, training DNNs requires costly pixel-level annotations. In this paper, we leverage the observation that image-level tags provide important cues of foreground salient objects, and develop a weakly supervised learning method for saliency detection using image-level tags only. The Foreground Inference Network (FIN) is introduced for this challenging task. In the first stage of our training method, FIN is jointly trained with a fully convolutional network (FCN) for image-level tag prediction. A global smooth pooling layer is proposed, enabling FCN to assign object category tags to corresponding object regions, while FIN is capable of capturing all potential foreground regions with the predicted saliency maps. In the second stage, FIN is fine-tuned with its predicted saliency maps as ground truth. For refinement of ground truth, an iterative Conditional Random Field is developed to enforce spatial label consistency and further boost performance. Our method alleviates annotation efforts and allows the usage of existing large scale training sets with image-level tags. Our model runs at 60 FPS, outperforms unsupervised ones with a large margin, and achieves comparable or even superior performance than fully supervised counterparts. Lijun Wang 0001, Huchuan Lu, Yifan Wang 0004, Mengyang Feng, Dong Wang 0004, Xiang Ruan |
CVPR | 1 |
| 2016 | STCT: Sequentially Training Convolutional Networks for Visual TrackingabstractDue to the limited amount of training samples, finetuning pre-trained deep models online is prone to overfitting. In this paper, we propose a sequential training method for convolutional neural networks (CNNs) to effectively transfer pre-trained deep features for online applications. We regard a CNN as an ensemble with each channel of the output feature map as an individual base learner. Each base learner is trained using different loss criterions to reduce correlation and avoid over-training. To achieve the best ensemble online, all the base learners are sequentially sampled into the ensemble via important sampling. To further improve the robustness of each base learner, we propose to train the convolutional layers with random binary masks, which serves as a regularization to enforce each base learner to focus on different input features. The proposed online training method is applied to visual tracking problem by transferring deep features trained on massive annotated visual data and is shown to significantly improve tracking performance. Extensive experiments are conducted on two challenging benchmark data set and demonstrate that our tracking algorithm can outperform state-of-the-art methods with a considerable margin. Lijun Wang 0001, Wanli Ouyang, Xiaogang Wang 0001, Huchuan Lu |
CVPR | 1 |
| 2016 | Pattern Mining Saliency
Yuqiu Kong, Lijun Wang 0001, Xiuping Liu, Huchuan Lu, Xiang Ruan |
ECCV (6) | 2 |
| 2016 | Saliency Detection with Recurrent Fully Convolutional Networks
Linzhao Wang, Lijun Wang 0001, Huchuan Lu, Xiang Ruan |
ECCV (4) | 2 |
| 2016 | Visual tracking via shallow and deep collaborative model
Bohan Zhuang, Lijun Wang 0001, Huchuan Lu |
Neurocomputing | 2 |
| 2016 | Visual Tracking via Random Walks on Graph ModelabstractIn this paper, we formulate visual tracking as random walks on graph models with nodes representing superpixels and edges denoting relationships between superpixels. We integrate two novel graphs with the theory of Markov random walks, resulting in two Markov chains. First, an ergodic Markov chain is enforced to globally search for the candidate nodes with similar features to the template nodes. Second, an absorbing Markov chain is utilized to model the temporal coherence between consecutive frames. The final confidence map is generated by a structural model which combines both appearance similarity measurement derived by the random walks and internal spatial layout demonstrated by different target parts. The effectiveness of the proposed Markov chains as well as the structural model is evaluated both qualitatively and quantitatively. Experimental results on challenging sequences show that the proposed tracking algorithm performs favorably against state-of-the-art methods. Xiaoli Li 0011, Zhifeng Han, Lijun Wang 0001, Huchuan Lu |
IEEE Trans. Cybern. | 3 |
| 2015 | Deep networks for saliency detection via local estimation and global searchabstractThis paper presents a saliency detection algorithm by integrating both local estimation and global search. In the local estimation stage, we detect local saliency by using a deep neural network (DNN-L) which learns local patch features to determine the saliency value of each pixel. The estimated local saliency maps are further refined by exploring the high level object concepts. In the global search stage, the local saliency map together with global contrast and geometric information are used as global features to describe a set of object candidate regions. Another deep neural network (DNN-G) is trained to predict the saliency score of each object region based on the global features. The final saliency map is generated by a weighted sum of salient object regions. Our method presents two interesting insights. First, local features learned by a supervised scheme can effectively capture local contrast, texture and shape information for saliency detection. Second, the complex relationship between different global saliency cues can be captured by deep networks and exploited principally rather than heuristically. Quantitative and qualitative experiments on several benchmark data sets demonstrate that our algorithm performs favorably against the state-of-the-art methods. Lijun Wang 0001, Huchuan Lu, Xiang Ruan, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2015 | Visual Tracking with Fully Convolutional NetworksabstractWe propose a new approach for general object tracking with fully convolutional neural network. Instead of treating convolutional neural network (CNN) as a black-box feature extractor, we conduct in-depth study on the properties of CNN features offline pre-trained on massive image data and classification task on ImageNet. The discoveries motivate the design of our tracking system. It is found that convolutional layers in different levels characterize the target from different perspectives. A top layer encodes more semantic features and serves as a category detector, while a lower layer carries more discriminative information and can better separate the target from distracters with similar appearance. Both layers are jointly used with a switch mechanism during tracking. It is also found that for a tracking target, only a subset of neurons are relevant. A feature map selection method is developed to remove noisy and irrelevant feature maps, which can reduce computation redundancy and improve tracking accuracy. Extensive evaluation on the widely used tracking benchmark [36] shows that the proposed tacker outperforms the state-of-the-art significantly. Lijun Wang 0001, Wanli Ouyang, Xiaogang Wang 0001, Huchuan Lu |
ICCV | 1 |
| 2015 | Visual Tracking via Structure Constrained GroupingabstractThis letter introduces a novel two-pass structural grouping algorithm and casts visual tracking as foreground superpixels grouping problem. In the first step, pairwise superpixel grouping is conducted in four orientations. Grouping prototypes containing the prior information of foreground and background are generated to determine whether any pair of neighboring superpixels should be grouped together. In the second step, superpixels selected by the first step are grouped into a single region which serves as the object region. The proposed grouping method has two benefits over the state-of-the-art ones. First, pairwise grouping is independently conducted in four orientations, which exploits the local structure of the foregound/backgroud and facilitates a more robust grouping process. Second, rather than considering the similarity of two neighboring superpixels, the grouping process is performed via accounting for the prior information of the object and the background, which is more suitable for visual tracking. Many experiments on challenging video clips demonstrate that our method achieves good performance than the state-of-the-art trackers in a wide range of tracking scenarios. Lijun Wang 0001, Huchuan Lu, Dong Wang 0004 |
IEEE Signal Process. Lett. | 1 |