EDBT 2026 Demo / reviewers in the wild / expert
Jianbing Shen
dblp:38/5435
· DBLP profile ↗
254ranked-venue papers
26as first author
113since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 166 · 16 first-author · 63 since 2021Artificial intelligence and machine learning · 156 · 9 first-author · 84 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 4 since 2021Security and privacy · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity MasksabstractDespite significant progress in pixel-level medical image analysis, existing medical image segmentation models rarely explore medical segmentation and diagnosis tasks jointly. However, it is crucial for patients that models can provide explainable diagnoses along with medical segmentation results. In this paper, we introduce a medical vision-language task named Medical Diagnosis Segmentation (MDS), which aims to understand clinical queries for medical images and generate the corresponding segmentation masks as well as diagnostic results. To facilitate this task, we first present the Multimodal Multi-disease Medical Diagnosis Segmentation (M3DS) dataset, containing diverse multimodal multi-disease medical images paired with their corresponding segmentation masks and diagnosis chain-of-thought, created via an automated diagnosis chain-of-thought generation pipeline. Moreover, we propose Sim4Seg, a novel framework that improves the performance of diagnosis segmentation by taking advantage of the Region-Aware Vision-Language Similarity to Mask (RVLS2M) module. To improve overall performance, we investigate a test-time scaling strategy for MDS tasks. Experimental results demonstrate that our method outperforms baselines in both segmentation and diagnosis. Lingran Song, Yucheng Zhou 0001, Jianbing Shen |
AAAI | 3 |
| 2026 | Towards High-Fidelity 3D Portrait Generation with Rich Details by Cross-View Prior-Aware DiffusionabstractRecent diffusion-based Single-image 3D portrait generation methods typically employ 2D diffusion models to provide multi-view knowledge, which is then distilled into 3D representations. However, these methods usually struggle to produce high-fidelity 3D models, frequently yielding excessively blurred textures. We attribute this issue to the insufficient consideration of cross-view consistency during the diffusion process, resulting in significant disparities between different views and ultimately leading to blurred 3D representations. In this paper, we address this issue by comprehensively exploiting multi-view priors in both the conditioning and diffusion procedures to produce consistent, detail-rich portraits. From the conditioning standpoint, we propose a Hybrid Priors Diffusion model, which explicitly and implicitly incorporates multi-view priors as conditions to enhance the status consistency of the generated multi-view portraits. From the diffusion perspective, considering the significant impact of the diffusion noise distribution on detailed texture generation, we propose a Multi-View Noise Resampling Strategy integrated within the optimization process leveraging cross-view priors to enhance representation consistency. Extensive experiments show that our method produces 3D portraits with accurate geometry and rich details from a single image. Wencheng Han, Xingping Dong, Jianbing Shen |
AAAI | 4 |
| 2026 | Less Is More: Vision Representation Compression for Efficient Video Generation with Large Language ModelsabstractVideo generation using Large Language Models (LLMs) has shown promising potential, effectively leveraging the extensive LLM infrastructure to provide a unified framework for multimodal understanding and content generation. However, these methods face critical challenges, i.e., token redundancy and inefficiencies arising from long sequences, which constrain their performance and efficiency compared to diffusion-based approaches. In this study, we investigate the impact of token redundancy in LLM-based video generation by information-theoretic analysis and propose Vision Representation Compression (VRC), a novel framework designed to achieve more in both performance and efficiency with less video token representations. VRC introduces learnable representation compressor and decompressor to compress video token representations, enabling autoregressive next-sequence prediction in a compact latent space. Our approach reduces redundancy, shortens token sequences, and improves model's ability to capture underlying video structures. Our experiments demonstrate that VRC reduces token sequence lengths by a factor of 4, achieving more than 9~14x acceleration in inference while maintaining performance comparable to state-of-the-art video generation models. VRC not only accelerates the inference but also significantly reduces memory requirements during both model training and inference. Yucheng Zhou 0001, Jihai Zhang 0002, Guanjie Chen, Jianbing Shen, Yu Cheng 0001 |
AAAI | 4 |
| 2026 | Multimodal Large Language Models for Multi-Subject In-Context Image GenerationabstractRecent advances in text-to-image (T2I) generation have enabled visually coherent image synthesis from descriptions, but generating images containing multiple given subjects remains challenging.As the number of reference identities increases, existing methods often suffer from subject missing and semantic drift.To address this problem, we propose MU-SIC, the first MLLM specifically designed for MUlti-Subject In-Context image generation.To overcome the data scarcity, we introduce an automatic and scalable data generation pipeline that eliminates the need for manual annotation.Furthermore, we enhance the model's understanding of multi-subject semantic relationships through a vision chain-of-thought (CoT) mechanism, guiding step-by-step reasoning from subject images to semantics and generation.To mitigate identity entanglement and manage visual complexity, we develop a novel semantics-driven spatial layout planning method and demonstrate its test-time scalability.By incorporating complex subject images during training, we improve the model's capacity for chained reasoning.In addition, we curate MSIC, a new benchmark tailored for multi-subject in-context generation.Experimental results demonstrate that MUSIC significantly surpasses other methods in both multiand single-subject scenarios. Yucheng Zhou 0001, Dubing Chen, Jianbing Shen |
ACL (1) | 4 |
| 2026 | Compatibility-Aware Dynamic Fine-Tuning for Large Language ModelsabstractSupervised Fine-Tuning (SFT) is the predominant paradigm for aligning large language models (LLMs), yet it suffers from optimization instability and limited generalization.Recent work attributes this issue to pathological gradient scaling and proposes Dynamic Fine-Tuning (DFT) to correct it at the token level.However, DFT assumes all demonstrations are equally suitable learning targets, an assumption violated by the strong heterogeneity of large-scale instruction data, where demonstration-policy mismatch induces highvariance updates at the sample level.We introduce Compatibility-Aware Dynamic Fine-Tuning (CADFT), a principled extension of DFT that controls sample-level optimization variance.CADFT derives a dynamic, policydependent compatibility signal from model likelihoods to modulate supervised updates, suppressing high-variance gradients from incompatible demonstrations.We further propose a delayed, low-frequency compatibilityguided rewriting strategy to transform persistently incompatible demonstrations into learnable targets.We show that CADFT can be interpreted as a variance-controlled estimator that generalizes token-level stabilization in DFT to the sample level.Extensive experiments demonstrate improved stability, generalization, and cold-start reinforcement learning initialization, while remaining fully supervised and independent of explicit reward modeling. Yucheng Zhou 0001, Junwei Sheng, Qianning Wang, Jianbing Shen |
ACL (1) | 4 |
| 2026 | Language Interprets Vision: Adaptive Encoding and Decoding for Referring Image SegmentationabstractReferring image segmentation aims to segment the referent with natural linguistic expressions. Due to the distinct modality properties of the image and language, it is challenging to effectively align token embeddings with visual regions. Different from existing methods of coordinate linguistics for the specific visual region, we propose a novel referring image segmentation paradigm, language interprets vision (LIV), which densely fine-grained aligns the visual and linguistic modalities, and fuse the multi-modal biases effectively. LIV resorts to re-encoding visual features on compositional dimensions of, which interprets vision through linguistic expression and makes cross-modality alignment denser. More specifically, we innovatively consider the adjacency of visual regions on the channel level to promote channel semantic consistency and propagate fine-grained semantics in the whole segmentation procedure. In addition, we also theoretically analyze that LIV effectively enriches the representation space and makes the comprehensive modality-fused biases more generalized, which boosts the precision of mask prediction. Extensive experimental results on three benchmarks validate that our proposed framework significantly outperforms other methods by a remarkable margin. Qi A, Sanyuan Zhao, Xingping Dong, Jianbing Shen |
Comput. Vis. Media | 4 |
| 2026 | MPLDM: Multi-modal prosthetic loosening diagnostic model for total hip arthroplasty
Xiao Chen 0025, Pang Lyu, Wencheng Han, Liyang Yang, Jianbing Shen |
Medical Image Anal. | 7 |
| 2025 | DME-Driver: Integrating Human Decision Logic and 3D Scene Perception in Autonomous DrivingabstractThere are two crucial aspects of reliable autonomous driving systems: the reasoning behind decision-making and the precision of environmental perception. This paper introduces DME-Driver, a new autonomous driving system that enhances performance and robustness by fully leveraging the two crucial aspects. This system comprises two main models. The first, the Decision Maker, is responsible for providing logical driving instructions. The second, the Executor, receives these instructions and generates precise control signals for the vehicles. To ensure explainable and reliable driving decisions, we build the Decision-Maker based on a large vision language model. This model follows the logic employed by experienced human drivers and simulates making decisions in a safe and reasonable manner. On the other hand, the generation of accurate control signals relies on precise and detailed environmental perception, where 3D scene perception models excel. Therefore, a planning-oriented perception model is employed as the Executor. It translates the logical decisions made by the Decision-Maker into accurate control signals for the self-driving cars. To effectively train the proposed system, a new dataset named Human-driver Behavior and Decision-making (HBD) dataset has been collected. This dataset encompasses a diverse range of human driver behaviors and their underlying motivations. By leveraging this dataset, our system achieves high-precision planning accuracy through a logical thinking process. Wencheng Han, Dongqian Guo, Cheng-Zhong Xu 0001, Jianbing Shen |
AAAI | 4 |
| 2025 | Language Prompt for Autonomous DrivingabstractA new trend in the computer vision community is to capture objects of interest following flexible human command represented by a natural language prompt. However, the progress of using language prompts in driving scenarios is stuck in a bottleneck due to the scarcity of paired prompt-instance data. To address this challenge, we propose the first object-centric language prompt set for driving scenes within 3D, multi-view, and multi-frame space, named NuPrompt. It expands nuScenes dataset by constructing a total of 40,147 language descriptions, each referring to an average of 7.4 object tracklets. Based on the object-text pairs from the new benchmark, we formulate a novel prompt-based driving task, \ie, employing a language prompt to predict the described object trajectory across views and frames. Furthermore, we provide a simple end-to-end baseline model based on Transformer, named PromptTrack. Experiments show that our PromptTrack achieves impressive performance on NuPrompt. We hope this work can provide some new insights for the self-driving community. Dongming Wu 0005, Wencheng Han, Yingfei Liu, Tiancai Wang, Cheng-Zhong Xu 0001, Xiangyu Zhang 0005, Jianbing Shen |
AAAI | 7 |
| 2025 | OLiDM: Object-aware LiDAR Diffusion Models for Autonomous DrivingabstractTo enhance autonomous driving, innovative approaches have been proposed to generate simulated LiDAR data. However, these methods often face challenges in producing high-quality and controllable foreground objects. To cater to the needs of object-aware tasks in 3D perception, we introduce OLiDM, a novel framework capable of generating controllable and high-fidelity LiDAR data at both the object and scene levels. OLiDM consists of two pivotal components: the Object-Scene Progressive Generation (OPG) module and the Object Semantic Alignment (OSA) module. OPG adapts to user-specific prompts to generate desired foreground objects, which are subsequently employed as conditions in scene generation, ensuring controllable and diverse output at both the object and scene levels. This also facilitates the association of user-defined object-level annotations with the generated LiDAR scenes. Moreover, OSA aims to rectify the misalignment between foreground objects and background scenes, enhancing the overall quality of the generated objects. The broad efficacy of OLiDM is demonstrated across both unconditional and conditional LiDAR generation tasks, as well as 3D perception tasks. Specifically, on the KITTI-360 dataset, OLiDM surpasses prior state-of-the-art methods such as UltraLiDAR by 11.8 in FPD, producing data that closely mirrors real-world distributions. Additionally, in sparse-to-dense LiDAR completion, OLiDM achieves a significant improvement over LiDARGen, with a 57.47% increase in semantic IoU. Moreover, in 3D object detection, OLiDM enhances the performance of mainstream detectors by 2.4% in mAP and 1.9% in NDS, underscoring its potential in advancing 3D perception models. Tianyi Yan, Junbo Yin, Xianpeng Lang, Ruigang Yang, Cheng-Zhong Xu 0001, Jianbing Shen |
AAAI | 6 |
| 2025 | Improving Medical Large Vision-Language Models with Abnormal-Aware FeedbackabstractExisting Medical Large Vision-Language Models (Med-LVLMs), encapsulating extensive medical knowledge, demonstrate excellent capabilities in understanding medical images. However, there remain challenges in visual localization in medical images, which is crucial for abnormality detection and interpretation. To address these issues, we propose a novel UMed-LVLM designed to unveil medical abnormalities. Specifically, we collect a Medical Abnormalities Unveiling (MAU) dataset and propose a two-stage training method for UMed-LVLM training. To collect MAU dataset, we propose a prompt method utilizing the GPT-4V to generate diagnoses based on identified abnormal areas in medical images. Moreover, the two-stage training method includes Abnormal-Aware Instruction Tuning and Abnormal-Aware Rewarding, comprising Relevance Reward, Abnormal Localization Reward and Vision Relevance Reward. Experimental results demonstrate that our UMed-LVLM significantly outperforms existing Med-LVLMs in identifying and understanding medical abnormalities, achieving a 58% improvement over the baseline. In addition, this work shows that enhancing the abnormality detection capabilities of Med-LVLMs significantly improves their understanding of medical images and generalization capability. Our code and data release at URL. Yucheng Zhou 0001, Lingran Song, Jianbing Shen |
ACL (1) | 3 |
| 2025 | Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy PredictionabstractWe present GDFusion, a temporal fusion method for vision-based 3D semantic occupancy prediction (VisionOcc). GDFusion opens up the underexplored aspects of temporal fusion within the VisionOcc framework, focusing on both temporal cues and fusion strategies. It systematically examines the entire VisionOcc pipeline, identifying three fundamental yet previously overlooked temporal cues: scene-level consistency, motion calibration, and geometric complementation. These cues capture diverse facets of temporal evolution and make distinct contributions across various modules in the VisionOcc framework. To effectively fuse temporal signals across heterogeneous representations, we propose a novel fusion strategy by reinterpreting the formulation of vanilla RNNs. This reinterpretation leverages gradient descent on features to unify the integration of diverse temporal information, seamlessly embedding the proposed temporal cues into the network. Extensive experiments on nuScenes demonstrate that GDFusion significantly outperforms established baselines, achieving 2.2%–4.7% mIoU improvement and reducing memory consumption by 30%–72%. Codes are available at https: //github.com/cdb342/GDFusion. Dubing Chen, Xingping Dong, Xianfei Li, Wenlong Liao, Jianbing Shen |
CVPR | 9 |
| 2025 | LOGICZSL: Exploring Logic-induced Representation for Compositional Zero-shot LearningabstractCompositional zero-shot learning (CZSL) aims to recognize unseen attribute-object compositions by learning the primitive concepts (i.e., attribute and object) from the training set. While recent works achieve impressive results in CZSL by leveraging large vision-language models like CLIP, they ignore the rich semantic relationships between primitive concepts and their compositions. In this work, we propose LogiCzsl, a novel logic-induced learning framework to explicitly model the semantic relationships. Our logic-induced learning framework formulates the relational knowledge constructed from large language models as a set of logic rules, and grounds them onto the training data. Our logic-induced losses are complementary to the widely used CZSL losses, therefore can be employed to inject the semantic information into any existing CZSL methods. Extensive experimental results show that our method brings significant performance improvements across diverse datasets (i.e., CGQA, UT-Zappos50K, MIT-States) with strong CLIP-based methods and settings (i.e., Close World, Open World). Peng Wu 0014, Xiankai Lu, Yongqin Xian, Jianbing Shen, Wenguan Wang |
CVPR | 5 |
| 2025 | DrivingSphere: Building a High-fidelity 4D World for Closed-loop SimulationabstractAutonomous driving evaluation requires simulation environments that closely replicate actual road conditions, including real-world sensory data and responsive feedback loops. However, many existing simulations need to predict waypoints along fixed routes on public datasets or synthetic photorealistic data, i.e., open-loop simulation usually lacks the ability to assess dynamic decision-making. While the recent efforts of closed-loop simulation offer feedback-driven environments, they cannot process visual sensor inputs or produce outputs that differ from real-world data. To address these challenges, we propose DrivingSphere, a realistic and closed-loop simulation framework. Its core idea is to build 4D world representation and generate real-life and controllable driving scenarios. In specific, our framework includes a Dynamic Environment Composition module that constructs a detailed 4D driving world with a format of occupancy equipping with static backgrounds and dynamic objects, and a Visual Scene Synthesis module that transforms this data into high-fidelity, multi-view video outputs, ensuring spatial and temporal consistency. By providing a dynamic and realistic simulation environment, DrivingSphere enables comprehensive testing and validation of autonomous driving algorithms, ultimately advancing the development of more reliable autonomous cars. The benchmark will be publicly released. Tianyi Yan, Dongming Wu 0005, Wencheng Han, Junpeng Jiang, Kun Zhan, Cheng-Zhong Xu 0001, Jianbing Shen |
CVPR | 8 |
| 2025 | Decoupling Fine Detail and Global Geometry for Compressed Depth Map Super-ResolutionabstractRecovering high-quality depth maps from compressed sources has gained significant attention due to the limitations of consumer-grade depth cameras and the bandwidth restrictions during data transmission. However, current methods still suffer from two challenges. First, bitdepth compression produces a uniform depth representation in regions with subtle variations, hindering the recovery of detailed information. Second, densely distributed random noise reduces the accuracy of estimating the global geometric structure of the scene. To address these challenges, we propose a novel framework, termed geometry-decoupled network (GDNet), for compressed depth map super-resolution that decouples the high-quality depth map reconstruction process by handling global and detailed geometric features separately. To be specific, we propose the fine geometry detail encoder (FGDE), which is designed to aggregate fine geometry details in high-resolution low-level image features while simultaneously enriching them with complementary information from low-resolution context-level image features. In addition, we develop the global geometry encoder (GGE) that aims at suppressing noise and extracting global geometric information effectively via constructing compact feature representation in a low-rank space. We conduct experiments on multiple benchmark datasets, demonstrating that our GDNet significantly outperforms current methods in terms of geometric consistency and detail recovery. In the ECCV 2024 AIM Compressed Depth Upsampling Challenge, our solution won the 1st place award. Our codes are available at: https://github.com/Ian0926/GDNet. Wencheng Han, Jianbing Shen |
CVPR | 3 |
| 2025 | ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow Predictions
Dubing Chen, Wencheng Han, Xinjing Cheng, Junbo Yin, Chenzhong Xu, Fahad Shahbaz Khan, Jianbing Shen |
ICCV | 8 |
| 2025 | Semantic Causality-Aware Vision-Based 3D Occupancy PredictionabstractVision-based 3D semantic occupancy prediction is a critical task in 3D vision that integrates volumetric 3D reconstruction with semantic understanding. Existing methods, however, often rely on modular pipelines. These modules are typically optimized independently or use pre-configured inputs, leading to cascading errors. In this paper, we address this limitation by designing a novel causal loss that enables holistic, end-to-end supervision of the modular 2D-to-3D transformation pipeline. Grounded in the principle of 2D-to-3D semantic causality, this loss regulates the gradient flow from 3D voxel representations back to the 2D features. Consequently, it renders the entire pipeline differentiable, unifying the learning process and making previously non-trainable components fully learnable. Building on this principle, we propose the Semantic Causality-Aware 2D-to-3D Transformation, which comprises three components guided by our causal loss: Channel-Grouped Lifting for adaptive semantic mapping, Learnable Camera Offsets for enhanced robustness against camera perturbations, and Normalized Convolution for effective feature propagation. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the Occ3D benchmark, demonstrating significant robustness to camera perturbations and improved 2D-to-3D semantic consistency. Dubing Chen, Yucheng Zhou 0001, Xianfei Li, Wenlong Liao, Jianbing Shen |
ICCV | 8 |
| 2025 | RAGNet: Large-Scale Reasoning-Based Affordance Segmentation Benchmark Towards General GraspingabstractGeneral robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance prediction data, leading to considerable concern about open-world effectiveness. To address this limitation, we build a large-scale grasping-oriented affordance segmentation benchmark with human-like instructions, named RAGNet. It contains 273k images, 180 categories, and 26k reasoning instructions. The images cover diverse embodied data domains, such as wild, robot, ego-centric, and even simulation data. They are carefully annotated with an affordance map, while the difficulty of language instructions is largely increased by removing their category name and only providing functional descriptions. Furthermore, we propose a comprehensive affordance-based grasping framework, named AffordanceNet, which consists of a VLM pre-trained on our massive affordance data and a grasping network that conditions an affordance map to grasp the target. Extensive experiments on affordance segmentation benchmarks and real-robot manipulation tasks show that our model has a powerful open-world generalization ability. Our data and code is available at https://github.com/wudongming97/AffordanceNet. Dongming Wu 0005, Yanping Fu, Saike Huang, Yingfei Liu, Fan Jia 0006, Nian Liu 0002, Tiancai Wang, Rao Muhammad Anwer, Fahad Shahbaz Khan, Jianbing Shen |
ICCV | 11 |
| 2025 | DC-ControlNet: Decoupling Inter- and Intra-Element Conditions in Image Generation with Diffusion ModelsabstractIn this paper, we introduce DC (Decouple)-ControlNet, a highly flexible and precisely controllable framework for multi-condition image generation. The core idea behind DC-ControlNet is to decouple control conditions, transforming global control into a hierarchical system that integrates distinct elements, contents, and layouts. This enables users to mix these individual conditions with greater flexibility, leading to more efficient and accurate image generation control. Previous ControlNet-based models rely solely on global conditions, which affect the entire image and lack the ability of element- or region-specific control. This limitation reduces flexibility and can cause condition misunderstandings in multi-conditional image generation. To address these challenges, we propose both intra-element and Inter-element Controllers in DC-ControlNet. The Intra-Element Controller handles different types of control signals within individual elements, accurately describing the content and layout characteristics of the object. For interactions between elements, we introduce the Inter-Element Controller, which accurately handles multi-element interactions and occlusion based on user-defined relationships. Extensive evaluations show that DC-ControlNet significantly outperforms existing ControlNet models and Layout-to-Image generative models in terms of control flexibility and precision in multi-condition control. Our project website is available at: https://um-lab.github.io/DC-ControlNet/ Wencheng Han, Yucheng Zhou 0001, Jianbing Shen |
ICCV | 4 |
| 2025 | Weak to Strong Generalization for Large Language Models with Multi-capabilitiesabstractAs large language models (LLMs) grow in sophistication, some of their capabilities surpass human abilities, making it essential to ensure their alignment with human values and intentions, i.e., Superalignment. This superalignment challenge is particularly critical for complex tasks, as annotations provided by humans, as weak supervisors, may be overly simplistic, incomplete, or incorrect. Previous work has demonstrated the potential of training a strong model using the weak dataset generated by a weak model as weak supervision. However, these studies have been limited to a single capability. In this work, we conduct extensive experiments to investigate weak to strong generalization for LLMs with multi-capabilities. The experiments reveal that different capabilities tend to remain relatively independent in this generalization, and the effectiveness of weak supervision is significantly impacted by the quality and diversity of the weak datasets. Moreover, the self-bootstrapping of the strong model leads to performance degradation due to its overconfidence and the limited diversity of its generated dataset. To address these issues, we proposed a novel training framework using reward models to select valuable data, thereby providing weak supervision for strong model training. In addition, we propose a two-stage training method on both weak and selected datasets to train the strong model. Experimental results demonstrate our method significantly improves the weak to strong generalization with multi-capabilities. Yucheng Zhou 0001, Jianbing Shen, Yu Cheng 0001 |
ICLR | 2 |
| 2025 | RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video GenerationabstractSynthetic data is crucial for advancing autonomous driving (AD) systems, yet current state-of-the-art video generation models, despite their visual realism, suffer from subtle geometric distortions that limit their utility for downstream perception tasks.
We identify and quantify this critical issue, demonstrating a significant performance gap in 3D object detection when using synthetic versus real data.
To address this, we introduce Reinforcement Learning with Geometric Feedback (RLGF), RLGF uniquely refines video diffusion models by incorporating rewards from specialized latent-space AD perception models.
Its core components include an efficient Latent-Space Windowing Optimization technique for targeted feedback during diffusion, and a Hierarchical Geometric Reward (HGR) system providing multi-level rewards for point-line-plane alignment, and scene occupancy coherence.
To quantify these distortions, we propose GeoScores. Applied to models like DiVE on nuScenes, RLGF substantially reduces geometric errors (e.g., VP error by 21\%, Depth error by 57\%) and dramatically improves 3D object detection mAP by 12.7\%, narrowing the gap to real-data performance. RLGF offers a plug-and-play solution for generating geometrically sound and reliable synthetic videos for AD development. Tianyi Yan, Wencheng Han, Xueyang Zhang, Kun Zhan, Cheng-Zhong Xu 0001, Jianbing Shen |
NeurIPS | 7 |
| 2025 | Alternate Geometric and Semantic Denoising Diffusion for Protein Inverse Folding
Chenglin Wang 0010, Yucheng Zhou 0001, Zijie Zhai, Jianbing Shen, Kai Zhang 0001 |
ECML/PKDD (3) | 5 |
| 2025 | Modality Confusion Learning: A Versatile Framework for Visible-Infrared Re-identification
Sanyuan Zhao, Mang Ye, Ruigang Yang, Jianbing Shen |
Int. J. Comput. Vis. | 5 |
| 2025 | Reconstructing High Quality Raw Video Using Temporal Affinity and Diffusion PriorabstractDue to the rich information and original data distribution, RAW data are widely used in many computer vision applications. However, the use of RAW video remains limited because of the high storage costs associated with data collection. Previous works have attempted to reconstruct RAW frames from sRGB data using small sampled metadata from the original RAW frames. Yet, these algorithms struggle with RAW video reconstruction due to the high computational cost of sampling metadata on cameras. To address these issues, we propose a new RAW video reconstruction pipeline that de-renders high-quality RAW videos from sRGB data using only one initial RAW frame as a reference. Specifically, we introduce three new models to achieve this goal. First, we present the Temporal-Affinity Guided De-rendering Network. This network leverages the temporal affinity between adjacent frames to construct a reference RAW image from previous RAW pixels. The corresponding RAW pixels in the previous frame provide valuable information about the original RAW data distribution, aiding in the precise reconstruction of the current frame. Second, to recover the missing RAW pixels caused by camera and foreground movement, we fully exploit the rich prior information from a pre-trained diffusion model and propose the RAW In-painting Model. This model can accurately fill in hollow regions in a RAW image based on the corresponding sRGB image and the surrounding RAW context. Lastly, we present a lightweight content-aware video clipper that automatically adjusts the clip length used for RAW video reconstruction, thereby balancing storage requirements with reconstruction quality. To better evaluate the performance of the proposed framework across different devices, we introduce the first RAW video reconstruction benchmark that comprises RAW videos from six types of camera devices with challenging scenarios. Experimental results demonstrate that our algorithm can accurately reconstruct RAW videos across all the scenarios. Wencheng Han, Jianbing Shen, David Crandall, Cheng-Zhong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Weakly Supervised Monocular 3D Object Detection by Spatial-Temporal View ConsistencyabstractMonocular 3D object detection plays a crucial role In the field of self-driving cars, estimating the size and location of objects solely based on input images. However, a notable disparity exists between the training and inference of 3D object detectors. This discrepancy arises because during inference, monocular 3D detectors depend solely on images captured by cameras; while during training, these methods require 3D ground truths labeled on point cloud data, which is obtained using specialized devices like LiDAR. This discrepancy creates a break in the data loop, preventing the feedback data from production cars from being utilized to enhance the robustness of the detectors. To address this issue and establish a connection in the data loop, we present a weakly-supervised solution that trains monocular 3D object detectors solely using 2D labels, eliminating the requirement for 3D ground truths. Our approach considers two view consistency: spatial and temporal view consistency, which play a crucial role in regulating the prediction of 3D bounding boxes. Spatial view consistency is achieved by employing projection and multi-view consistency techniques to guide the optimization of the target's location and size. We leverage temporal viewpoint consistency to provide temporal multi-view image pairs, and we further introduce temporal movement consistency to tackle the challenge of dynamic scenes. With only 2D ground truths, our method achieves comparable performance to fully supervised methods. Additionally, our method can be employed as a pre-training method and achieves significant improvement when fine-tuned with a small proportion of fully supervised labels. Wencheng Han, Haibin Ling, Jianbing Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Semantic-Aware Pseudo-Labeling for Unsupervised Meta-LearningabstractIn unsupervised meta-learning, the clustering-based pseudo-labeling approach is an attractive framework, since it is model-agnostic, allowing it to synergize with supervised algorithms to learn from unlabeled data. However, the pseudo-labels suffer from clustering noise and semantic chaos problems, further impacting the effectiveness of meta-learning. In this paper, we analyze and optimize the pseudo-labeling process, including encoding and clustering, aiming to generate semantic-like pseudo-labels to narrow the gap between unsupervised and supervised meta-learning. First, during the encoding, we observe that the embedding space of existing methods lacks clustering-friendly properties, which is the primary reason for clustering noise. To address this issue, we minimize the inter-to-intra-class similarity ratio to generate clustering-friendly embedding features and validate our approach through comprehensive experiments. Then, during the clustering, we find that the semantic quality of pseudo-labels is not adequately controlled, resulting in semantic chaos of pseudo-labels. We propose a semantic-stability index to measure the semantic quality of pseudo-labels quantitatively. Based on this index, we propose the Semantic-aware Pseudo-label Reassignment mechanism to generate semantic-like pseudo-labels for all samples. Our approach is model-agnostic and can easily be integrated into existing supervised methods. To demonstrate its generalization ability, we integrate it into two representative algorithms: MAML and EP. The results on three main few-shot benchmarks clearly show that the proposed method achieves significant improvement compared to state-of-the-art models. Notably, our approach also outperforms the corresponding supervised method in three tasks. Tianran Ouyang, Xingping Dong, Mang Ye, Bo Du 0001, Ling Shao 0001, Jianbing Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | TransBridge: Boost 3D Object Detection by Scene-Level Completion With Transformer Decoderabstract3D object detection is essential in autonomous driving, providing vital information about moving objects and obstacles. Detecting objects in distant regions with only a few LiDAR points is still a challenge, and numerous strategies have been developed to address point cloud sparsity through densification. This paper presents a joint completion and detection framework that improves the detection feature in sparse areas while maintaining costs unchanged. Specifically, we proposeTransBridge, a novel transformer-based up-sampling block that fuses the features from the detection and completion networks. The detection network can benefit from acquiring implicit completion features derived from the completion network. Additionally, we design theDynamic-Static Reconstruction(DSRecon) module to produce dense LiDAR data for the completion network, meeting the requirement for dense point cloud ground truth. Furthermore, we employ the transformer mechanism to establish connections between channels and spatial relations, resulting in a high-resolution feature map used for completion purposes. Extensive experiments on the nuScenes and Waymo datasets demonstrate the effectiveness of the proposed framework. The results show that our framework consistently improves end-to-end 3D object detection, with the mean average precision (mAP) ranging from 0.7 to 1.5 across multiple methods, indicating its generalization ability. For the two-stage detection framework, it also boosts the mAP up to 5.78 points. Qinghao Meng, Chenming Wu, Liangjun Zhang, Jianbing Shen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Towards Better Cephalometric Landmark Detection With Diffusion Data GenerationabstractCephalometric landmark detection is essential for orthodontic diagnostics and treatment planning. Nevertheless, the scarcity of samples in data collection and the extensive effort required for manual annotation have significantly impeded the availability of diverse datasets. This limitation has restricted the effectiveness of deep learning-based detection methods, particularly those based on large-scale vision models. To address these challenges, we have developed an innovative data generation method capable of producing diverse cephalometric X-ray images along with corresponding annotations without human intervention. To achieve this, our approach initiates by constructing new cephalometric landmark annotations using anatomical priors. Then, we employ a diffusion-based generator to create realistic X-ray images that correspond closely with these annotations. To achieve precise control in producing samples with different attributes, we introduce a novel prompt cephalometric X-ray image dataset. This dataset includes real cephalometric X-ray images and detailed medical text prompts describing the images. By leveraging these detailed prompts, our method improves the generation process to control different styles and attributes. Facilitated by the large, diverse generated data, we introduce large-scale vision detection models into the cephalometric landmark detection task to improve accuracy. Experimental results demonstrate that training with the generated data substantially enhances the performance. Compared to methods without using the generated data, our approach improves the Success Detection Rate (SDR) by 6.5%, attaining a notable 82.2%. All code and data are available at: https://um-lab.github.io/cepha-generation/. Dongqian Guo, Wencheng Han, Pang Lyu, Jianbing Shen |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Bilateral Cross-Modality Graph Matching Attention for Feature Fusion in Visual Question AnsweringabstractAnswering semantically complicated questions according to an image is challenging in a visual question answering (VQA) task. Although the image can be well represented by deep learning, the question is always simply embedded and cannot well indicate its meaning. Besides, the visual and textual features have a gap for different modalities, it is difficult to align and utilize the cross-modality information. In this article, we focus on these two problems and propose a graph matching attention (GMA) network. First, it not only builds graph for the image but also constructs graph for the question in terms of both syntactic and embedding information. Next, we explore the intramodality relationships by a dual-stage graph encoder and then present a bilateral cross-modality GMA to infer the relationships between the image and the question. The updated cross-modality features are then sent into the answer prediction module for final answer prediction. Experiments demonstrate that our network achieves the state-of-the-art performance on the GQA dataset and the VQA 2.0 dataset. The ablation studies verify the effectiveness of each module in our GMA network. Jianjian Cao, Xiameng Qin, Sanyuan Zhao, Jianbing Shen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | High-Fidelity and High-Efficiency Talking Portrait Synthesis With Detail-Aware Neural Radiance FieldsabstractIn this paper, we propose a novel rendering framework based on neural radiance fields (NeRF) named HH-NeRF that can generate high-resolution audio-driven talking portrait videos with high fidelity and fast rendering. Specifically, our framework includes a detail-aware NeRF module and an efficient conditional super-resolution module. First, a detail-aware NeRF is proposed to efficiently generate a high-fidelity low-resolution talking head, by using the encoded volume density estimation and audio-eye-aware color calculation. This module can capture natural eye blinks and high-frequency details, and maintain a similar rendering time as previous fast methods. Secondly, we present an efficient conditional super-resolution module on the dynamic scene to directly generate the high-resolution portrait with our low-resolution head. Incorporated with the prior information, such as depth map and audio features, our new proposed efficient conditional super resolution module can adopt a lightweight network to efficiently generate realistic and distinct high-resolution videos. Extensive experiments demonstrate that our method can generate more distinct and fidelity talking portraits on high resolution (900 × 900) videos compared to state-of-the-art methods. Muyu Wang, Sanyuan Zhao, Xingping Dong, Jianbing Shen |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object DetectionabstractVehicle-to-Everything (V2X) collaborative perception has recently gained significant attention due to its capability to enhance scene understanding by integrating information from various agents, e.g., vehicles, and infrastructure. However, current works often treat the information from each agent equally, ignoring the inherent domain gap caused by the utilization of different LiDAR sensors of each agent, thus leading to suboptimal performance. In this paper, we propose DI-V2X, that aims to learn Domain-Invariant representations through a new distillation framework to mitigate the domain discrepancy in the context of V2X 3D object detection. DI-V2X comprises three essential components: a domain-mixing instance augmentation (DMA) module, a progressive domain-invariant distillation (PDD) module, and a domain-adaptive fusion (DAF) module. Specifically, DMA builds a domain-mixing 3D instance bank for the teacher and student models during training, resulting in aligned data representation. Next, PDD encourages the student models from different domains to gradually learn a domain-invariant feature representation towards the teacher, where the overlapping regions between agents are employed as guidance to facilitate the distillation process. Furthermore, DAF closes the domain gap between the students by incorporating calibration-aware domain-adaptive attention. Extensive experiments on the challenging DAIR-V2X and V2XSet benchmark datasets demonstrate DI-V2X achieves remarkable performance, outperforming all the previous V2X models. Code is available at https://github.com/Serenos/DI-V2X. Xiang Li 0001, Junbo Yin, Wei Li 0111, Cheng-Zhong Xu 0001, Ruigang Yang, Jianbing Shen |
AAAI | 6 |
| 2024 | Fine-Grained Distillation for Long Document RetrievalabstractLong document retrieval aims to fetch query-relevant documents from a large-scale collection, where knowledge distillation has become de facto to improve a retriever by mimicking a heterogeneous yet powerful cross-encoder. However, in contrast to passages or sentences, retrieval on long documents suffers from the \textit{scope hypothesis} that a long document may cover multiple topics. This maximizes their structure heterogeneity and poses a granular-mismatch issue, leading to an inferior distillation efficacy. In this work, we propose a new learning framework, fine-grained distillation (FGD), for long-document retrievers. While preserving the conventional dense retrieval paradigm, it first produces global-consistent representations crossing different fine granularity and then applies multi-granular aligned distillation merely during training. In experiments, we evaluate our framework on two long-document retrieval benchmarks, which show state-of-the-art performance. Yucheng Zhou 0001, Tao Shen 0001, Xiubo Geng, Chongyang Tao, Jianbing Shen, Guodong Long, Can Xu 0002, Daxin Jiang |
AAAI | 5 |
| 2024 | IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object DetectionabstractBird's eye view (BEV) representation has emerged as a dominant solution for describing 3D space in autonomous driving scenarios. However, objects in the BEV representation typically exhibit small sizes, and the associated point cloud context is inherently sparse, which leads to great challenges for reliable 3D perception. In this paper, we propose IS-Fusion, an innovative multimodal fusion framework that jointly captures the Instance- and Scene-level contextual information. IS-Fusion essentially differs from existing approaches that only focus on the BEV scene-level fusion by explicitly incorporating instance-level multimodal information, thus facilitating the instance-centric tasks like 3D object detection. It comprises a Hierarchical Scene Fusion (HSF) module and an Instance-Guided Fusion (IGF) module. HSF applies Point-to-Grid and Grid-to-Region transformers to capture the multimodal scene context at different granularities. IGF mines instance candidates, explores their relationships, and aggregates the local multimodal context for each instance. These instances then serve as guidance to enhance the scene feature and yield an instance-aware BEV representation. On the challenging nuScenes benchmark, IS-Fusion outperforms all the published multimodal works to date. Code is available at: https://github.com/yinjunbo/IS-Fusion. Junbo Yin, Jianbing Shen, Runnan Chen, Wei Li 0111, Ruigang Yang, Pascal Frossard, Wenguan Wang |
CVPR | 2 |
| 2024 | Leveraging Frame Affinity for sRGB-to-RAWVideo De-RenderingabstractUnprocessed RAW video has shown distinct advantages over sRGB video in video editing and computer vision tasks. However, capturing RAW video is challenging due to limitations in bandwidth and storage. Various methods have been proposed to address similar issues in single image RAW capture through de-rendering. These methods utilize both the metadata and the sRGB image to perform sRGB-to-RAW de-rendering and recover high-quality single-frame RAW data. However, metadata-based methods always require additional computation for online metadata generation, imposing severe burden on mobile camera device for high frame rate RAW video capture. To address this issue, we propose a framework that utilizes frame affinity to achieve high-quality sRGB-to-RAW video reconstruction. Our approach consists of two main steps. The first step, temporal affinity prior extraction, uses motion information between adjacent frames to obtain a reference RAW image. The second step, spatial feature fusion and mapping, learns a pixel-level mapping function using scene-specific and position-specific features provided by the previous frame. Our method can be easily applied to current mobile camera equipment without complicated adaptations or added burden. To demonstrate the effectiveness of our approach, we introduce the first RAW Video De-rendering Benchmark. In this benchmark, our method outperforms state-of-the-art RAW image reconstruction methods, even without image-level metadata. Wencheng Han, Jianbing Shen, Cheng-Zhong Xu 0001, Wentao Liu 0002 |
CVPR | 4 |
| 2024 | High-Precision Self-supervised Monocular Depth Estimation with Rich-Resource Prior
Wencheng Han, Jianbing Shen |
ECCV (31) | 2 |
| 2024 | RepVF: A Unified Vector Fields Representation for Multi-task 3D Perception
Chunliang Li, Wencheng Han, Jun Yin 0003, Sanyuan Zhao, Jianbing Shen |
ECCV (32) | 5 |
| 2024 | TopoMLP: A Simple yet Strong Pipeline for Driving Topology ReasoningabstractTopology reasoning aims to comprehensively understand road scenes and present drivable routes in autonomous driving. It requires detecting road centerlines (lane) and traffic elements, further reasoning their topology relationship, \textit{i.e.}, lane-lane topology, and lane-traffic topology. In this work, we first present that the topology score relies heavily on detection performance on lane and traffic elements. Therefore, we introduce a powerful 3D lane detector and an improved 2D traffic element detector to extend the upper limit of topology performance. Further, we propose TopoMLP, a simple yet high-performance pipeline for driving topology reasoning. Based on the impressive detection performance, we develop two simple MLP-based heads for topology generation. TopoMLP achieves state-of-the-art performance on OpenLane-V2 dataset, \textit{i.e.}, 41.2\% OLS with ResNet-50 backbone. It is also the 1st solution for 1st OpenLane Topology in Autonomous Driving Challenge. We hope such simple and strong pipeline can provide some new insights to the community. Code is at https://github.com/wudongming97/TopoMLP. Dongming Wu 0005, Fan Jia 0006, Yingfei Liu, Tiancai Wang, Jianbing Shen |
ICLR | 6 |
| 2024 | Prior Metadata-Driven RAW Reconstruction: Eliminating the Need for Per-Image MetadataabstractWhile RAW images are efficient for image editing and perception tasks, their large size can strain camera storage and bandwidth. Reconstruction methods of RAW images from sRGB data typically require additional metadata from the RAW image, which increases camera processing computations. To address this problem, we propose using Prior Meta as a reference to reconstruct the RAW data instead of relying on per-image metadata. Prior metadata is extracted offline from reference RAW images, which are usually part of the training dataset and have similar scenes and light conditions as the target image. With this prior metadata, the camera does not need to provide any extra processing other than the sRGB images, and our model can autonomously find the desired prior information. To achieve this, we design a three-step pipeline. First, we build a pixel searching network that can find the most similar pixels in the reference RAW images as prior information. Then, in the second step, we compress the large-scale reference images to about 0.02% of their original size to reduce the searching cost. Finally, in the last step, we develop a neural network reconstructor to reconstruct the high-fidelity RAW images. Our model achieves comparable, and even better, performance than RAW reconstruction methods based on metadata. Wencheng Han, Wentao Liu 0002, Chen Qian 0006, Cheng-Zhong Xu 0001, Jianbing Shen |
ACM Multimedia | 7 |
| 2024 | A simple but effective vision transformer framework for visible-infrared person re-identification
Yudong Li 0002, Sanyuan Zhao, Jianbing Shen |
Comput. Vis. Image Underst. | 3 |
| 2024 | Asymmetric Convolution: An Efficient and Generalized Method to Fuse Feature Maps in Multiple Vision TasksabstractFusing features from different sources is a critical aspect of many computer vision tasks. Existing approaches can be roughly categorized as parameter-free or learnable operations. However, parameter-free modules are limited in their ability to benefit from offline learning, leading to poor performance in some challenging situations. Learnable fusing methods are often space-consuming and time-consuming, particularly when fusing features with different shapes. To address these shortcomings, we conducted an in-depth analysis of the limitations associated with both fusion methods. Based on our findings, we propose a generalized module named Asymmetric Convolution Module (ACM). This module can learn to encode effective priors during offline training and efficiently fuse feature maps with different shapes in specific tasks. Specifically, we propose a mathematically equivalent method for replacing costly convolutions on concatenated features. This method can be widely applied to fuse feature maps across different shapes. Furthermore, distinguished from parameter-free operations that can only fuse two features of the same type, our ACM is general, flexible, and can fuse multiple features of different types. To demonstrate the generality and efficiency of ACM, we integrate it into several state-of-the-art models on three representative vision tasks. Extensive experimental results on three tasks and several datasets demonstrate that our new module can bring significant improvements and noteworthy efficiency. Wencheng Han, Xingping Dong, David Crandall, Cheng-Zhong Xu 0001, Jianbing Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Multi-threshold deep metric learning for facial expression recognitionabstractFeature representations generated through triplet-based deep metric learning offer significant advantages for facial expression recognition (FER). Each threshold in triplet loss inherently shapes a distinct distribution of inter-class variations, leading to unique representations of expression features. Nonetheless, pinpointing the optimal threshold for triplet loss presents a formidable challenge, as the ideal threshold varies not only across different datasets but also among classes within the same dataset. In this paper, we propose a novel multi-threshold deep metric learning approach that bypasses the complex process of threshold validation and markedly improves the effectiveness in creating expression feature representations. Instead of choosing a single optimal threshold from a valid range, we comprehensively sample thresholds throughout this range, which ensures that the representation characteristics exhibited by the thresholds within this spectrum are fully captured and utilized for enhancing FER. Specifically, we segment the embedding layer of the deep metric learning network into multiple slices, with each slice representing a specific threshold sample. We subsequently train these embedding slices in an end-to-end fashion, applying triplet loss at its associated threshold to each slice, which results in a collection of unique expression features corresponding to each embedding slice. Moreover, we identify the issue that the traditional triplet loss may struggle to converge when employing the widely-used Batch Hard strategy for mining informative triplets, and introduce a novel loss termed dual triplet loss to address it. Extensive evaluations demonstrate the superior performance of the proposed approach on both posed and spontaneous facial expression datasets. Wenwu Yang, Jinyi Yu, Tuo Chen, Zhenguang Liu, Xun Wang 0007, Jianbing Shen |
Pattern Recognit. | 6 |
| 2024 | Uncertainty-Aware Hierarchical Aggregation Network for Medical Image SegmentationabstractMedical image segmentation is an essential process to assist clinics with computer-aided diagnosis and treatment. Recently, a large amount of convolutional neural network (CNN)-based methods have been rapidly developed and achieved remarkable performances in several different medical image segmentation tasks. However, the same type of infected region or lesions often has a diversity of scales, making it a challenging task to achieve accurate medical image segmentation. In this paper, we present a novel Uncertainty-aware Hierarchical Aggregation Network, namely UHA-Net, for medical image segmentation, which can fully make utilization of cross-level and multi-scale features to handle scale variations. Specifically, we propose a hierarchical feature fusion (HFF) module to aggregate high-level features, which is used to produce a global map for the coarse localization of the segmented target. Then, we propose an uncertainty-induced cross-level fusion (UCF) module to fully fuse features from the adjacent levels, which can learn knowledge guidance to capture the contextual information from adjacent resolutions. Further, a scale aggregation module (SAM) is presented to learn multi-scale features by using different convolution kernels, to effectively deal with scale variations. At last, we formulate a unified framework to simultaneously fuse inter-layer convolutional features and learn the discriminability of multi-scale representations from the intra-layer features, leading to accurate segmentation results. We carry out experiments on three different medical image segmentation tasks, and the results demonstrate that our UHA-Net outperforms state-of-the-art segmentation methods. Our implementation code and segmentation maps will be publicly at https://github.com/taozh2017/UHANet. Tao Zhou 0002, Yi Zhou 0007, Geng Chen 0001, Jianbing Shen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Learning 3D Face Reconstruction From the Cycle-Consistency of Dynamic FacesabstractReconstructinga 3D face from a single image is a crucial task in numerous multimedia applications. Face images with ground-truth 3D face shapes are scarce, so unsupervised deep learning methods, which rely primarily on the free supervision signal derived from the visual disparity between the input image and the rendered counterpart of the predicted 3D face, have proven superior for reconstructing 3D faces. However, it is challenging for such techniques to decouple the dynamic 3D face properties such as pose or expression from a single 2D image, especially when similar local visual appearance changes can be caused by both pose and expression motion, resulting in imprecise 3D face reconstruction. In this article, a novel cycle-consistency in dynamic 3D face characteristics is introduced as a free supervisory signal for learning accurate 3D face shapes from unlabeled facial images. The main idea of cycle-consistency is to explicitly inject the head pose or facial expression variation between video frames into a face image, and then to extract and reverse the injected variation in order to reconstruct the face image to its original state. In our model, a CNN network with multiple branches is proposed to disentangle 3D face properties like identity, expression, pose, and texture from 2D facial images, one branch for each 3D face property. During training, our model learns to completely decouple the dynamic 3D face properties (pose and expression) to be useful for performing cycle-consistent face reconstruction. Extensive experiments demonstrate the superiority of our approach. On the challenging AFLW2000-3D, MICC Florence, and NoW datasets, our method outperforms or is on par with the state of the art. Wenwu Yang, Yeqing Zhao, Bailin Yang, Jianbing Shen |
IEEE Trans. Multim. | 4 |
| 2024 | Relational Network via Cascade CRF for Video Language GroundingabstractVideo Language Grounding is one of the most challenging cross-modal video understanding tasks. This task aims to localize a target moment semantically corresponding to a given language query in an untrimmed video. Many existing VLG methods rely on the proposal-based framework, despite the dominant performance achieved, they usually focus on interacting a few internal frames with the query to score segment proposals, trapping in the long-range dependencies when the proposal feature is limited. Meanwhile, adjacent proposals share similar visual semantics, making VLG models hard to align the accurate semantics of video-query contents and degenerating the ranking performance. To remedy the above limitations, we propose VLG-CRF by introducing the conditional random fields (CRFs) to handle the discrete yet indistinguishable proposals. Specifically, VLG-CRF consists of two cascade CRF-based modules. The AttentiveCRFs is developed for multi-modal feature fusion to better integrate temporal and semantic relation between modalities. We also devise a new variant of ConvCRFs to capture the relation of discrete segments and rectify the predicting scores to make relatively high prediction scores clustered in a range. Experiments on three benchmark datasets,i.e., Charades-STA, ActivityNet-Caption, and TACoS, show the superiority of our method and the state-of-the-art performance is achieved. Xiankai Lu, Hao Zhang 0048, Xiushan Nie, Yilong Yin, Jianbing Shen |
IEEE Trans. Multim. | 6 |
| 2024 | A New Framework of Collaborative Learning for Adaptive Metric DistillationabstractThis article presents a new adaptive metric distillation approach that can significantly improve the student networks' backbone features, along with better classification results. Previous knowledge distillation (KD) methods usually focus on transferring the knowledge across the classifier logits or feature structure, ignoring the excessive sample relations in the feature space. We demonstrated that such a design greatly limits performance, especially for the retrieval task. The proposed collaborative adaptive metric distillation (CAMD) has three main advantages: 1) the optimization focuses on optimizing the relationship between key pairs by introducing the hard mining strategy into the distillation framework; 2) it provides an adaptive metric distillation that can explicitly optimize the student feature embeddings by applying the relation in the teacher embeddings as supervision; and 3) it employs a collaborative scheme for effective knowledge aggregation. Extensive experiments demonstrated that our approach sets a new state-of-the-art in both the classification and retrieval tasks, outperforming other cutting-edge distillers under various settings. Mang Ye, Yan Wang 0116, Sanyuan Zhao, Ping Li 0016, Jianbing Shen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Few-Shot Object Detection With Self-Supervising and Cooperative ClassifierabstractFew-shot object detection (FSOD), which detects novel objects with only a few training instances, has recently attracted more attention. Previous works focus on making the most use of label information of objects. Still, they fail to consider the structural and semantic information of the image itself and solve the misclassification between data-abundant base classes and data-scarce novel classes efficiently. In this article, we propose FSOD with Self-Supervising and Cooperative Classifier ( [Formula: see text]) approach to deal with those concerns. Specifically, we analyze the underlying performance degradation of novel classes in FSOD and discover that false-positive samples are the main reason. By looking into these false-positive samples, we further notice that misclassifying novel classes as base classes are the main cause. Thus, we introduce double RoI heads into the existing Fast-RCNN to learn more specific features for novel classes. We also consider using self-supervised learning (SSL) to learn more structural and semantic information. Finally, we propose a cooperative classifier (CC) with the base-novel regularization to maximize the interclass variance between base and novel classes. In the experiment, [Formula: see text] outperforms all the latest baselines in most cases on PASCAL VOC and COCO. Jilin Hu, Jianbing Shen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | LWSIS: LiDAR-Guided Weakly Supervised Instance Segmentation for Autonomous DrivingabstractImage instance segmentation is a fundamental research topic in autonomous driving, which is crucial for scene understanding and road safety. Advanced learning-based approaches often rely on the costly 2D mask annotations for training. In this paper, we present a more artful framework, LiDAR-guided Weakly Supervised Instance Segmentation (LWSIS), which leverages the off-the-shelf 3D data, i.e., Point Cloud, together with the 3D boxes, as natural weak supervisions for training the 2D image instance segmentation models. Our LWSIS not only exploits the complementary information in multimodal data during training but also significantly reduces the annotation cost of the dense 2D masks. In detail, LWSIS consists of two crucial modules, Point Label Assignment (PLA) and Graph-based Consistency Regularization (GCR). The former module aims to automatically assign the 3D point cloud as 2D point-wise labels, while the atter further refines the predictions by enforcing geometry and appearance consistency of the multimodal data. Moreover, we conduct a secondary instance segmentation annotation on the nuScenes, named nuInsSeg, to encourage further research on multimodal perception tasks. Extensive experiments on the nuInsSeg, as well as the large-scale Waymo, show that LWSIS can substantially improve existing weakly supervised segmentation models by only involving 3D data during training. Additionally, LWSIS can also be incorporated into 3D object detectors like PointPainting to boost the 3D detection performance for free. The code and dataset are available at https://github.com/Serenos/LWSIS. Xiang Li 0001, Junbo Yin, Botian Shi, Yikang Li 0002, Ruigang Yang, Jianbing Shen |
AAAI | 6 |
| 2023 | Exposing the Self-Supervised Space-Time Correspondence Learning via Graph KernelsabstractSelf-supervised space-time correspondence learning is emerging as a promising way of leveraging unlabeled video. Currently, most methods adapt contrastive learning with mining negative samples or reconstruction adapted from the image domain, which requires dense affinity across multiple frames or optical flow constraints. Moreover, video correspondence predictive models require mining more inherent properties in videos, such as structural information. In this work, we propose the VideoHiGraph, a space-time correspondence framework based on a learnable graph kernel. Concerning the video as the spatial-temporal graph, the learning objectives of VideoHiGraph are emanated in a self-supervised manner for predicting unobserved hidden graphs via graph kernel manner. We learn a representation of the temporal coherence across frames in which pairwise similarity defines the structured hidden graph, such that a biased random walk graph kernel along the sub-graph can predict long-range correspondence. Then, we learn a refined representation across frames on the node-level via a dense graph kernel. The self-supervision of the model training is formed by the structural and temporal consistency of the graph. VideoHiGraph achieves superior performance and demonstrates its robustness across the benchmark of label propagation tasks involving objects, semantic parts, keypoints, and instances. Our algorithm implementations have been made publicly available at https://github.com/zyqin19/VideoHiGraph. Zheyun Qin, Xiankai Lu, Xiushan Nie, Yilong Yin, Jianbing Shen |
AAAI | 5 |
| 2023 | SSDA3D: Semi-supervised Domain Adaptation for 3D Object Detection from Point CloudabstractLiDAR-based 3D object detection is an indispensable task in advanced autonomous driving systems. Though impressive detection results have been achieved by superior 3D detectors, they suffer from significant performance degeneration when facing unseen domains, such as different LiDAR configurations, different cities, and weather conditions. The mainstream approaches tend to solve these challenges by leveraging unsupervised domain adaptation (UDA) techniques. However, these UDA solutions just yield unsatisfactory 3D detection results when there is a severe domain shift, e.g., from Waymo (64-beam) to nuScenes (32-beam). To address this, we present a novel Semi-Supervised Domain Adaptation method for 3D object detection (SSDA3D), where only a few labeled target data is available, yet can significantly improve the adaptation performance. In particular, our SSDA3D includes an Inter-domain Adaptation stage and an Intra-domain Generalization stage. In the first stage, an Inter-domain Point-CutMix module is presented to efficiently align the point cloud distribution across domains. The Point-CutMix generates mixed samples of an intermediate domain, thus encouraging to learn domain-invariant knowledge. Then, in the second stage, we further enhance the model for better generalization on the unlabeled target set. This is achieved by exploring Intra-domain Point-MixUp in semi-supervised learning, which essentially regularizes the pseudo label distribution. Experiments from Waymo to nuScenes show that, with only 10% labeled target data, our SSDA3D can surpass the fully-supervised oracle model with 100% target label. Our code is available at https://github.com/yinjunbo/SSDA3D. Yan Wang 0116, Junbo Yin, Wei Li 0111, Pascal Frossard, Ruigang Yang, Jianbing Shen |
AAAI | 6 |
| 2023 | Weakly Supervised Monocular 3D Object Detection Using Multi-View Projection and Direction ConsistencyabstractMonocular 3D object detection has become a mainstream approach in automatic driving for its easy application. A prominent advantage is that it does not need Li-DAR point clouds during the inference. However, most current methods still rely on 3D point cloud data for labeling the ground truths used in the training phase. This inconsistency between the training and inference makes it hard to utilize the large-scale feedback data and increases the data collection expenses. To bridge this gap, we propose a new weakly supervised monocular 3D objection detection method, which can train the model with only 2D labels marked on images. To be specific, we explore three types of consistency in this task, i.e. the projection, multi-view and direction consistency, and design a weakly-supervised architecture based on these consistencies. Moreover, we propose a new 2D direction labeling method in this task to guide the model for accurate rotation direction prediction. Experiments show that our weakly-supervised method achieves comparable performance with some fully supervised methods. When used as a pre-training method, our model can significantly outperform the corresponding fully-supervised baseline with only 1/3 3D labels. Wencheng Han, Zhongying Qiu, Cheng-Zhong Xu 0001, Jianbing Shen |
CVPR | 5 |
| 2023 | Referring Multi-Object TrackingabstractExisting referring understanding tasks tend to involve the detection of a single text-referred object. In this paper, we propose a new and general referring understanding task, termed referring multi-object tracking (RMOT). Its core idea is to employ a language expression as a semantic cue to guide the prediction of multi-object tracking. To the best of our knowledge, it is the first work to achieve an arbitrary number of referent object predictions in videos. To push forward RMOT, we construct one benchmark with scalable expressions based on KITTI, named Refer-KITTI. Specifically, it provides 18 videos with 818 expressions, and each expression in a video is annotated with an average of 10.7 objects. Further, we develop a transformer-based architecture TransRMOT to tackle the new task in an online manner, which achieves impressive detection performance and out-performs other counterparts. The Refer-KITTI dataset and the code are released at https://referringmot.github.io. Dongming Wu 0005, Wencheng Han, Tiancai Wang, Xingping Dong, Xiangyu Zhang 0005, Jianbing Shen |
CVPR | 6 |
| 2023 | Self-Supervised Monocular Depth Estimation by Direction-aware Cumulative Convolution NetworkabstractMonocular depth estimation is known as an ill-posed task in which objects in a 2D image usually do not contain sufficient information to predict their depth. Thus, it acts differently from other tasks (e.g., classification and segmentation) in many ways. In this paper, we find that self-supervised monocular depth estimation shows a direction sensitivity and environmental dependency in the feature representation. But the current backbones borrowed from other tasks pay less attention to handling different types of environmental information, limiting the overall depth accuracy. To bridge this gap, we propose a new Direction-aware Cumulative Convolution Network (DaCCN), which improves the depth feature representation in two aspects. First, we propose a direction-aware module, which can learn to adjust the feature extraction in each direction, facilitating the encoding of different types of information. Secondly, we design a new cumulative convolution to improve the efficiency for aggregating important environmental information. Experiments show that our method achieves significant improvements on three widely used benchmarks, KITTI, Cityscapes, and Make3D, setting a new state-of-the-art performance on the popular benchmarks with all three types of self-supervision. https://github.com/wencheng256/DaCCN. Wencheng Han, Junbo Yin, Jianbing Shen |
ICCV | 3 |
| 2023 | OnlineRefer: A Simple Online Baseline for Referring Video Object SegmentationabstractReferring video object segmentation (RVOS) aims at segmenting an object in a video following human instruction. Current state-of-the-art methods fall into an offline pattern, in which each clip independently interacts with text embedding for cross-modal understanding. They usually present that the offline pattern is necessary for RVOS, yet model limited temporal association within each clip. In this work, we break up the previous offline belief and propose a simple yet effective online model using explicit query propagation, named OnlineRefer. Specifically, our approach leverages target cues that gather semantic information and position prior to improve the accuracy and ease of referring predictions for the current frame. Furthermore, we generalize our online model into a semi-online framework to be compatible with video-based backbones. To show the effectiveness of our method, we evaluate it on four benchmarks, i.e., Refer-Youtube-VOS, Refer-DAVIS17, A2D-Sentences, and JHMDB-Sentences. Without bells and whistles, our OnlineRefer with a Swin-L backbone achieves 63.5 J&F and 64.8 J&F on Refer-Youtube-VOS and Refer-DAVIS17, outperforming all other offline methods. Our code is available at https://github.com/wudongming97/OnlineRefer. Dongming Wu 0005, Tiancai Wang, Xiangyu Zhang 0005, Jianbing Shen |
ICCV | 5 |
| 2023 | Clip Fusion with Bi-level Optimization for Human Mesh Reconstruction from Monocular VideosabstractHuman mesh reconstruction (HMR) from monocular video is the key step to many mixed reality and robotic applications. Although existing methods show promising results by capturing frames' temporal information, these methods predict human mesh with the design of implicit temporal learning modules in a sequence to frame manner. To mine more temporal information from the video, we present a bi-level clip inference network for HMR, which leverages both local motion and global context explicitly for dense 3D reconstruction. Specifically, we propose a novel bi-level temporal fusion strategy that takes both neighboring and long-range relations into consideration. In addition, different from traditional frame-wise operation, we investigate an alternative perspective by treating video-based HMR as clip-wise inference. We evaluate the proposed method on multiple datasets (3DPW, Human3.6M, and MPI-INF-3DHP) quantitatively and qualitatively, demonstrating a significant improvement over existing methods (in terms of PA-MPJPE, ACC-Error etc). Furthermore, we extend the proposed method on more challenging Multiple Shots HMR task to demonstrate its generalizability. Some visual demos can be seen https://github.com/bicf0/bicf_demo. Peng Wu 0014, Xiankai Lu, Jianbing Shen, Yilong Yin |
ACM Multimedia | 3 |
| 2023 | Spectrum-irrelevant fine-grained representation for visible-infrared person re-identification
Jiahao Gong, Sanyuan Zhao, Kin-Man Lam 0001, Xin Gao 0001, Jianbing Shen |
Comput. Vis. Image Underst. | 5 |
| 2023 | Full-duplex strategy for video object segmentationabstractPrevious video object segmentation approaches mainly focus on simplex solutions linking appearance and motion, limiting effective feature collaboration between these two cues. In this work, we study a novel and efficient full-duplex strategy network (FSNet) to address this issue, by considering a better mutual restraint scheme linking motion and appearance allowing exploitation of cross-modal features from the fusion and decoding stage. Specifically, we introduce a relational cross-attention module (RCAM) to achieve bidirectional message propagation across embedding sub-spaces. To improve the model’s robustness and update inconsistent features from the spatiotemporal embeddings, we adopt a bidirectional purification module after the RCAM. Extensive experiments on five popular benchmarks show that our FSNet is robust to various challenging scenarios (e.g., motion blur and occlusion), and compares well to leading methods both for video object segmentation and video salient object detection. The project is publicly available at https://github.com/GewelsJI/FSNet . Ge-Peng Ji, Deng-Ping Fan, Keren Fu, Jianbing Shen, Ling Shao 0001 |
Comput. Vis. Media | 5 |
| 2023 | Active Perception for Visual-Language Navigation
Hanqing Wang 0001, Wenguan Wang, Wei Liang 0008, Steven C. H. Hoi, Jianbing Shen, Luc Van Gool |
Int. J. Comput. Vis. | 5 |
| 2023 | Deep understanding of big geo-social data for autonomous vehicles
Shuo Shang, Jianbing Shen, Ji-Rong Wen, Panos Kalnis |
Neural Comput. Appl. | 2 |
| 2023 | Consistency and Diversity Induced Human Motion SegmentationabstractSubspace clustering is a classical technique that has been widely used for human motion segmentation and other related tasks. However, existing segmentation methods often cluster data without guidance from prior knowledge, resulting in unsatisfactory segmentation results. To this end, we propose a novel Consistency and Diversity induced human Motion Segmentation (CDMS) algorithm. Specifically, our model factorizes the source and target data into distinct multi-layer feature spaces, in which transfer subspace learning is conducted on different layers to capture multi-level information. A multi-mutual consistency learning strategy is carried out to reduce the domain gap between the source and target data. In this way, the domain-specific knowledge and domain-invariant properties can be explored simultaneously. Besides, a novel constraint based on the Hilbert Schmidt Independence Criterion (HSIC) is introduced to ensure the diversity of multi-level subspace representations, which enables the complementarity of multi-level representations to be explored to boost the transfer learning performance. Moreover, to preserve the temporal correlations, an enhanced graph regularizer is imposed on the learned representation coefficients and the multi-level representations of the source data. The proposed model can be efficiently solved using the Alternating Direction Method of Multipliers (ADMM) algorithm. Extensive experimental results on public human motion datasets demonstrate the effectiveness of our method against several state-of-the-art approaches. Tao Zhou 0002, Huazhu Fu, Chen Gong 0002, Ling Shao 0001, Fatih Porikli, Haibin Ling, Jianbing Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Adaptive Siamese Tracking With a Compact Latent NetworkabstractIn this article, we provide an intuitive viewing to simplify the Siamese-based trackers by converting the tracking task to a classification. Under this viewing, we perform an in-depth analysis for them through visual simulations and real tracking examples, and find that the failure cases in some challenging situations can be regarded as the issue of missing decisive samples in offline training. Since the samples in the initial (first) frame contain rich sequence-specific information, we can regard them as the decisive samples to represent the whole sequence. To quickly adapt the base model to new scenes, a compact latent network is presented via fully using these decisive samples. Specifically, we present a statistics-based compact latent feature for fast adjustment by efficiently extracting the sequence-specific information. Furthermore, a new diverse sample mining strategy is designed for training to further improve the discrimination ability of the proposed compact latent network. Finally, a conditional updating strategy is proposed to efficiently update the basic models to handle scene variation during the tracking phase. To evaluate the generalization ability and effectiveness and of our method, we apply it to adjust three classical Siamese-based trackers, namely SiamRPN++, SiamFC, and SiamBAN. Extensive experimental results on six recent datasets demonstrate that all three adjusted trackers obtain the superior performance in terms of the accuracy, while having high running speed. Xingping Dong, Jianbing Shen, Fatih Porikli, Jiebo Luo 0001, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Graph Neural Network and Spatiotemporal Transformer Attention for 3D Video Object Detection From Point CloudsabstractPrevious works for LiDAR-based 3D object detection mainly focus on the single-frame paradigm. In this paper, we propose to detect 3D objects by exploiting temporal information in multiple frames, i.e., point cloud videos. We empirically categorize the temporal information into short-term and long-term patterns. To encode the short-term data, we present a Grid Message Passing Network (GMPNet), which considers each grid (i.e., the grouped points) as a node and constructs a k-NN graph with the neighbor grids. To update features for a grid, GMPNet iteratively collects information from its neighbors, thus mining the motion cues in grids from nearby frames. To further aggregate long-term frames, we propose an Attentive Spatiotemporal Transformer GRU (AST-GRU), which contains a Spatial Transformer Attention (STA) module and a Temporal Transformer Attention (TTA) module. STA and TTA enhance the vanilla GRU to focus on small objects and better align moving objects. Our overall framework supports both online and offline video object detection in point clouds. We implement our algorithm based on prevalent anchor-based and anchor-free detectors. Evaluation results on the challenging nuScenes benchmark show superior performance of our method, achieving first on the leaderboard (at the time of paper submission) without any "bells and whistles." Our source code is available at https://github.com/shenjianbing/GMP3D. Junbo Yin, Jianbing Shen, Xin Gao 0001, David Crandall, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Dual-Semantic Consistency Learning for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared person Re-Identification (VI-ReID) conducts comprehensive identity analysis on non-overlapping visible and infrared camera sets for intelligent surveillance systems, which face huge instance variations derived from modality discrepancy. Existing methods employ different kinds of network structure to extract modality-invariant features. Differently, we propose a novel framework, named Dual-Semantic Consistency Learning Network (DSCNet), which attributes modality discrepancy to channel-level semantic inconsistency. DSCNet optimizes channel consistency from two aspects, fine-grained inter-channel semantics, and comprehensive inter-modality semantics. Furthermore, we propose Joint Semantics Metric Learning to simultaneously optimize the distribution of the channel-and-modality feature embeddings. It jointly exploits the correlation between channel-specific and modality-specific semantics in a fine-grained manner. We conduct a series of experiments on the SYSU-MM01 and RegDB datasets, which validates that DSCNet delivers superiority compared with current state-of-the-art methods. On the more challenging SYSU-MM01 dataset, our network can achieve 73.89% Rank-1 accuracy and 69.47% mAP value. Our code is available athttps://github.com/bitreidgroup/DSCNet. Yuhao Kang, Sanyuan Zhao, Jianbing Shen |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | Multi-Granularity Context Network for Efficient Video Semantic SegmentationabstractCurrent video semantic segmentation tasks involve two main challenges: how to take full advantage of multi-frame context information, and how to improve computational efficiency. To tackle the two challenges simultaneously, we present a novel Multi-Granularity Context Network (MGCNet) by aggregating context information at multiple granularities in a more effective and efficient way. Our method first converts image features into semantic prototypes, and then conducts a non-local operation to aggregate the per-frame and short-term contexts jointly. An additional long-term context module is introduced to capture the video-level semantic information during training. By aggregating both local and global semantic information, a strong feature representation is obtained. The proposed pixel-to-prototype non-local operation requires less computational cost than traditional non-local ones, and is video-friendly since it reuses the semantic prototypes of previous frames. Moreover, we propose an uncertainty-aware and structural knowledge distillation strategy to boost the performance of our method. Experiments on Cityscapes and CamVid datasets with multiple backbones demonstrate that the proposed MGCNet outperforms other state-of-the-art methods with high speed and low latency. Zhiyuan Liang, Xiangdong Dai, Xiaogang Jin 0001, Jianbing Shen |
IEEE Trans. Image Process. | 5 |
| 2023 | Reformulating Graph Kernels for Self-Supervised Space-Time Correspondence LearningabstractSelf-supervised space-time correspondence learning utilizing unlabeled videos holds great potential in computer vision. Most existing methods rely on contrastive learning with mining negative samples or adapting reconstruction from the image domain, which requires dense affinity across multiple frames or optical flow constraints. Moreover, video correspondence prediction models need to uncover more inherent properties of the video, such as structural information. In this work, we propose HiGraph+, a sophisticated space-time correspondence framework based on learnable graph kernels. By treating videos as a spatial-temporal graph, the learning objective of HiGraph+ is issued in a self-supervised manner, predicting the unobserved hidden graph via graph kernel methods. First, we learn the structural consistency of sub-graphs in graph-level correspondence learning. Furthermore, we introduce a spatio-temporal hidden graph loss through contrastive learning that facilitates learning temporal coherence across frames of sub-graphs and spatial diversity within the same frame. Therefore, we can predict long-term correspondences and drive the hidden graph to acquire distinct local structural representations. Then, we learn a refined representation across frames on the node-level via a dense graph kernel. The structural and temporal consistency of the graph forms the self-supervision of model training. HiGraph+ achieves excellent performance and demonstrates robustness in benchmark tests involving object, semantic part, keypoint, and instance labeling propagation tasks. Our algorithm implementations have been made publicly available at https://github.com/zyqin19/HiGraph. Zheyun Qin, Xiankai Lu, Dongfang Liu, Xiushan Nie, Yilong Yin, Jianbing Shen, Alexander C. Loui |
IEEE Trans. Image Process. | 6 |
| 2023 | Nested Architecture Search for Point Cloud Semantic SegmentationabstractPoint cloud semantic segmentation (PCSS), for the purpose of labeling a set of points stored in irregular and unordered structures, is an important yet challenging task. It is vital for the task of learning a good representation for each 3D data point, which encodes rich context knowledge and hierarchically structural information. However, despite great success has been achieved by existing PCSS methods, they are limited to make full use of important context information and rich hierarchical features for representation learning. In this paper, we propose to build 'hyperpoint' representations for 3D data point via a nested network architecture, which is able to explicitly exploit multi-scale, pyramidally hierarchical features and construct powerful representations for PCSS. In particular, we introduce a PCSS nested architecture search (PCSS-NAS) algorithm to automatically design the model's side-output branches at different levels as well as its skip-layer structures, enabling the resulting model to best deal with the scale-space problem. Our searched architecture, named Auto-NestedNet, is evaluated on four well-known benchmarks: S3DIS, ScanNet, Semantic3D and Paris-Lille-3D. Experimental results show that the proposed Auto-NestedNet achieves the state-of-the-art performance. Our source code is available at https://github.com/fanyang587/NestedNet. Fan Yang 0054, Xin Li 0079, Jianbing Shen |
IEEE Trans. Image Process. | 3 |
| 2023 | Automatic Schelling Point Detection From MeshesabstractMesh Schelling points explain how humans focus on specific regions of a 3D object. They have a large number of important applications in computer graphics and provide valuable information for perceptual psychology studies. However, detecting mesh Schelling points is time-consuming and expensive since the existing techniques are mostly based on participant observation studies. To overcome these limitations, we propose to employ powerful deep learning techniques to detect mesh Schelling points in an automatic manner, free from participant observation studies. Specifically, we utilize the mesh convolution and pooling operations to extract informative features from mesh objects, and then predict the 3D heat map of Schelling points in an end-to-end manner. In addition, we propose a Deep Schelling Network (DS-Net) to automatically detect the Schelling points, including a multi-scale fusion component and a novel region-specific loss function to improve our network for a better regression of heat maps. To the best of our knowledge, DS-Net is the first deep neural network for detecting Schelling points from 3D meshes. We evaluate DS-Net on a mesh Schelling point dataset obtained from participant observation studies. The experimental results demonstrate that DS-Net is capable of detecting mesh Schelling points effectively and outperforms various state-of-the-art mesh saliency methods and deep learning models, both qualitatively and quantitatively. Geng Chen 0001, Hang Dai, Tao Zhou 0002, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Tree Energy Loss: Towards Sparsely Annotated Semantic SegmentationabstractSparsely annotated semantic segmentation (SASS) aims to train a segmentation network with coarse-grained (i.e., point-, scribble-, and block-wise) supervisions, where only a small proportion of pixels are labeled in each image. In this paper, we propose a novel tree energy loss for SASS by providing semantic guidance for unlabeled pixels. The tree energy loss represents images as minimum spanning trees to model both low-level and high-level pair-wise affini-ties. By sequentially applying these affinities to the net-work prediction, soft pseudo labels for unlabeled pixels are generated in a coarse-to-fine manner, achieving dynamic online self-training. The tree energy loss is effective and easy to be incorporated into existing frameworks by com-bining it with a traditional segmentation loss. Compared with previous SASS methods, our method requires no multi-stage training strategies, alternating optimization proce-dures, additional supervised data, or time-consuming post-processing while outperforming them in all SASS settings. Code is available at https://github.com/megvii-research/TreeEnergyLoss. Zhiyuan Liang, Tiancai Wang, Xiangyu Zhang 0005, Jian Sun 0001, Jianbing Shen |
CVPR | 5 |
| 2022 | Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language NavigationabstractSince the rise of vision-language navigation (VLN), great progress has been made in instruction following - building a follower to navigate environments under the guidance of instructions. However, far less attention has been paid to the inverse task: instruction generation - learning a speaker to generate grounded descriptions for navigation routes. Existing VLN methods train a speaker independently and often treat it as a data augmentation tool to strengthen the follower, while ignoring rich cross-task relations. Here we describe an approach that learns the two tasks simultaneously and exploits their intrinsic correlations to boost the training of each: the follower judges whether the speaker-created instruction explains the original navigation route correctly, and vice versa. Without the need of aligned instruction-path pairs, such cycle-consistent learning scheme is complementary to task-specific training targets defined on labeled data, and can also be applied over unlabeled paths (sampled without paired instructions). Another agent, called creator is added to generate counterfactual environments. It greatly changes current scenes yet leaves novel items - which are vital for the execution of original instructions - unchanged. Thus more informative training scenes are synthesized and the three agents compose a powerful VLN learning system. Extensive experiments on a standard benchmark show that our approach improves the performance of various follower models and produces accurate navigation instructions. Hanqing Wang 0001, Wei Liang 0008, Jianbing Shen, Luc Van Gool, Wenguan Wang |
CVPR | 3 |
| 2022 | Multi-Level Representation Learning with Semantic Alignment for Referring Video Object SegmentationabstractReferring video object segmentation (RVOS) is a challenging language-guided video grounding task, which requires comprehensively understanding the semantic information of both video content and language queries for object prediction. However, existing methods adopt multi-modal fusion at a frame-based spatial granularity. The limitation of visual representation is prone to causing vision-language mismatching and producing poor segmentation results. To address this, we propose a novel multi-level representation learning approach, which explores the inherent structure of the video content to provide a set of discriminative visual embedding, enabling more effective vision-language semantic alignment. Specifically, we embed different visual cues in terms of visual granularity, including multi-frame long-temporal information at video level, intra-frame spatial semantics at frame level, and enhanced object-aware feature prior at object level. With the powerful multi-level visual embedding and carefully-designed dynamic alignment, our model can generate a robust representation for accurate video object segmentation. Extensive experiments on Refer-DAVIS17and Refer-YouTube-VOS demonstrate that our model achieves superior performance both in segmentation accuracy and inference speed. Dongming Wu 0005, Xingping Dong, Ling Shao 0001, Jianbing Shen |
CVPR | 4 |
| 2022 | Learning Disentanglement with Decoupled Labels for Vision-Language Navigation
Wenhao Cheng, Xingping Dong, Salman Khan 0001, Jianbing Shen |
ECCV (36) | 4 |
| 2022 | Rethinking Clustering-Based Pseudo-Labeling for Unsupervised Meta-Learning
Xingping Dong, Jianbing Shen, Ling Shao 0001 |
ECCV (20) | 2 |
| 2022 | BRNet: Exploring Comprehensive Features for Monocular Depth Estimation
Wencheng Han, Junbo Yin, Xiaogang Jin 0001, Xiangdong Dai, Jianbing Shen |
ECCV (38) | 5 |
| 2022 | Semi-supervised 3D Object Detection with Proficient Teachers
Junbo Yin, Dingfu Zhou, Liangjun Zhang, Cheng-Zhong Xu 0001, Jianbing Shen, Wenguan Wang |
ECCV (38) | 6 |
| 2022 | ProposalContrast: Unsupervised Pre-training for LiDAR-Based 3D Object Detection
Junbo Yin, Dingfu Zhou, Liangjun Zhang, Cheng-Zhong Xu 0001, Jianbing Shen, Wenguan Wang |
ECCV (39) | 6 |
| 2022 | Modality Synergy Complement Learning with Cascaded Aggregation for Visible-Infrared Person Re-Identification
Sanyuan Zhao, Yuhao Kang, Jianbing Shen |
ECCV (14) | 4 |
| 2022 | Re-Thinking Co-Salient Object DetectionabstractIn this article, we conduct a comprehensive study on the co-salient object detection (CoSOD) problem for images. CoSOD is an emerging and rapidly growing extension of salient object detection (SOD), which aims to detect the co-occurring salient objects in a group of images. However, existing CoSOD datasets often have a serious data bias, assuming that each group of images contains salient objects of similar visual appearances. This bias can lead to the ideal settings and effectiveness of models trained on existing datasets, being impaired in real-life situations, where similarities are usually semantic or conceptual. To tackle this issue, we first introduce a new benchmark, called CoSOD3k in the wild, which requires a large amount of semantic context, making it more challenging than existing CoSOD datasets. Our CoSOD3k consists of 3,316 high-quality, elaborately selected images divided into 160 groups with hierarchical annotations. The images span a wide range of categories, shapes, object sizes, and backgrounds. Second, we integrate the existing SOD techniques to build a unified, trainable CoSOD framework, which is long overdue in this field. Specifically, we propose a novel CoEG-Net that augments our prior model EGNet with a co-attention projection strategy to enable fast common information learning. CoEG-Net fully leverages previous large-scale SOD datasets and significantly improves the model scalability and stability. Third, we comprehensively summarize 40 cutting-edge algorithms, benchmarking 18 of them over three challenging CoSOD datasets (iCoSeg, CoSal2015, and our CoSOD3k), and reporting more detailed (i.e., group-level) performance analysis. Finally, we discuss the challenges and future works of CoSOD. We hope that our study will give a strong boost to growth in the CoSOD community. The benchmark toolbox and results are available on our project page at https://dpfan.net/CoSOD3K. Deng-Ping Fan, Tengpeng Li, Zheng Lin 0005, Ge-Peng Ji, Dingwen Zhang, Ming-Ming Cheng, Huazhu Fu, Jianbing Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2022 | Siamese Network for RGB-D Salient Object Detection and BeyondabstractExisting RGB-D salient object detection (SOD) models usually treat RGB and depth as independent information and design separate networks for feature extraction from each. Such schemes can easily be constrained by a limited amount of training data or over-reliance on an elaborately designed training process. Inspired by the observation that RGB and depth modalities actually present certain commonality in distinguishing salient objects, a novel joint learning and densely cooperative fusion (JL-DCF) architecture is designed to learn from both RGB and depth inputs through a shared network backbone, known as the Siamese architecture. In this paper, we propose two effective components: joint learning (JL), and densely cooperative fusion (DCF). The JL module provides robust saliency feature learning by exploiting cross-modal commonality via a Siamese network, while the DCF module is introduced for complementary feature discovery. Comprehensive experiments using 5 popular metrics show that the designed framework yields a robust RGB-D saliency detector with good generalization. As a result, JL-DCF significantly advances the SOTAs by an average of ~2.0% (F-measure) across 7 challenging datasets. In addition, we show that JL-DCF is readily applicable to other related multi-modal detection tasks, including RGB-T SOD and video SOD, achieving comparable or better performance. Keren Fu, Deng-Ping Fan, Ge-Peng Ji, Qijun Zhao, Jianbing Shen, Ce Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Deep Object Tracking With Shrinkage LossabstractIn this paper, we address the issue of data imbalance in learning deep models for visual object tracking. Although it is well known that data distribution plays a crucial role in learning and inference models, considerably less attention has been paid to data imbalance in visual tracking. For the deep regression trackers that directly learn a dense mapping from input images of target objects to soft response maps, we identify their performance is limited by the extremely imbalanced pixel-to-pixel differences when computing regression loss. This prevents existing end-to-end learnable deep regression trackers from performing as well as discriminative correlation filters (DCFs) trackers. For the deep classification trackers that draw positive and negative samples to learn discriminative classifiers, there exists heavy class imbalance due to a limited number of positive samples when compared to the number of negative samples. To balance training data, we propose a novel shrinkage loss to penalize the importance of easy training data mostly coming from the background, which facilitates both deep regression and classification trackers to better distinguish target objects from the background. We extensively validate the proposed shrinkage loss function on six benchmark datasets, including the OTB-2013, OTB-2015, UAV-123, VOT-2016, VOT-2018 and LaSOT. Equipped with our shrinkage loss, the proposed one-stage deep regression tracker achieves favorable results against state-of-the-art methods, especially in comparison with DCFs trackers. Meanwhile, our shrinkage loss generalizes well to deep classification trackers. When replacing the original binary cross entropy loss with our shrinkage loss, three representative baseline trackers achieve large performance gains, even setting new state-of-the-art results. Xiankai Lu, Chao Ma 0004, Jianbing Shen, Xiaokang Yang 0001, Ian D. Reid 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Segmenting Objects From Relational Visual DataabstractIn this article, we model a set of pixelwise object segmentation tasks - automatic video segmentation (AVS), image co-segmentation (ICS) and few-shot semantic segmentation (FSS) - in a unified view of segmenting objects from relational visual data. To this end, we propose an attentive graph neural network (AGNN) that addresses these tasks in a holistic fashion, by formulating them as a process of iterative information fusion over data graphs. It builds a fully-connected graph to efficiently represent visual data as nodes and relations between data instances as edges. The underlying relations are described by a differentiable attention mechanism, which thoroughly examines fine-grained semantic similarities between all the possible location pairs in two data instances. Through parametric message passing, AGNN is able to capture knowledge from the relational visual data, enabling more accurate object discovery and segmentation. Experiments show that AGNN can automatically highlight primary foreground objects from video sequences (i.e., automatic video segmentation), and extract common objects from noisy collections of semantically related images (i.e., image co-segmentation). AGNN can even generalize segment new categories with little annotated data (i.e., few-shot semantic segmentation). Taken together, our results demonstrate that AGNN provides a powerful tool that is applicable to a wide range of pixel-wise object pattern understanding tasks with relational visual data. Our algorithm implementations have been made publicly available at https://github.com/carrierlxk/AGNN. Xiankai Lu, Wenguan Wang, Jianbing Shen, David Crandall, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Zero-Shot Video Object Segmentation With Co-Attention Siamese NetworksabstractWe introduce a novel network, called CO-attention siamese network (COSNet), to address the zero-shot video object segmentation task in a holistic fashion. We exploit the inherent correlation among video frames and incorporate a global co-attention mechanism to further improve the state-of-the-art deep learning based solutions that primarily focus on learning discriminative foreground representations over appearance and motion in short-term temporal segments. The co-attention layers in COSNet provide efficient and competent stages for capturing global correlations and scene context by jointly computing and appending co-attention responses into a joint feature space. COSNet is a unified and end-to-end trainable framework where different co-attention variants can be derived for capturing diverse properties of the learned joint feature space. We train COSNet with pairs (or groups) of video frames, and this naturally augments training data and allows increased learning capacity. During the segmentation stage, the co-attention model encodes useful information by processing multiple reference frames together, which is leveraged to infer the frequently reappearing and salient foreground objects better. Our extensive experiments over three large benchmarks demonstrate that COSNet outperforms the current alternatives by a large margin. Our implementations are available at https://github.com/carrierlxk/COSNet. Xiankai Lu, Wenguan Wang, Jianbing Shen, David Crandall, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Towards a Weakly Supervised Framework for 3D Point Cloud Object Detection and AnnotationabstractIt is quite laborious and costly to manually label LiDAR point cloud data for training high-quality 3D object detectors. This work proposes a weakly supervised framework which allows learning 3D detection from a few weakly annotated examples. This is achieved by a two-stage architecture design. Stage-1 learns to generate cylindrical object proposals under inaccurate and inexact supervision, obtained by our proposed BEV center-click annotation strategy, where only the horizontal object centers are click-annotated in bird's view scenes. Stage-2 learns to predict cuboids and confidence scores in a coarse-to-fine, cascade manner, under incomplete supervision, i.e., only a small portion of object cuboids are precisely annotated. With KITTI dataset, using only 500 weakly annotated scenes and 534 precisely labeled vehicle instances, our method achieves 86-97 percent the performance of current top-leading, fully supervised detectors (which require 3,712 exhaustively annotated scenes with 15,654 instances). More importantly, with our elaborately designed network architecture, our trained model can be applied as a 3D object annotator, supporting both automatic and active (human-in-the-loop) working modes. The annotations generated by our model can be used to train 3D object detectors, achieving over 95 percent of their original performance (with manually labeled training data). Our experiments also show our model's potential in boosting performance when given more training data. The above designs make our approach highly practical and open-up opportunities for learning 3D detection at reduced annotation cost. Qinghao Meng, Wenguan Wang, Tianfei Zhou, Jianbing Shen, Yunde Jia, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Distilled Siamese Networks for Visual TrackingabstractIn recent years, Siamese network based trackers have significantly advanced the state-of-the-art in real-time tracking. Despite their success, Siamese trackers tend to suffer from high memory costs, which restrict their applicability to mobile devices with tight memory budgets. To address this issue, we propose a distilled Siamese tracking framework to learn small, fast and accurate trackers (students), which capture critical knowledge from large Siamese trackers (teachers) by a teacher-students knowledge distillation model. This model is intuitively inspired by the one teacher versus multiple students learning method typically employed in schools. In particular, our model contains a single teacher-student distillation module and a student-student knowledge sharing mechanism. The former is designed using a tracking-specific distillation strategy to transfer knowledge from a teacher to students. The latter is utilized for mutual learning between students to enable in-depth knowledge understanding. Extensive empirical evaluations on several popular Siamese trackers demonstrate the generality and effectiveness of our framework. Moreover, the results on five tracking benchmarks show that the proposed distilled trackers achieve compression rates of up to 18× and frame-rates of 265 FPS, while obtaining comparable tracking accuracy compared to base models. Jianbing Shen, Yuanpei Liu, Xingping Dong, Xiankai Lu, Fahad Shahbaz Khan, Steven C. H. Hoi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Salient Object Detection in the Deep Learning Era: An In-Depth SurveyabstractAs an essential problem in computer vision, salient object detection (SOD) has attracted an increasing amount of research attention over the years. Recent advances in SOD are predominantly led by deep learning-based solutions (named deep SOD). To enable in-depth understanding of deep SOD, in this paper, we provide a comprehensive survey covering various aspects, ranging from algorithm taxonomy to unsolved issues. In particular, we first review deep SOD algorithms from different perspectives, including network architecture, level of supervision, learning paradigm, and object-/instance-level detection. Following that, we summarize and analyze existing SOD datasets and evaluation metrics. Then, we benchmark a large group of representative SOD models, and provide detailed analyses of the comparison results. Moreover, we study the performance of SOD algorithms under different attribute settings, which has not been thoroughly explored previously, by constructing a novel SOD dataset with rich attribute annotations covering various salient object types, challenging factors, and scene categories. We further analyze, for the first time in the field, the robustness of SOD models to random input perturbations and adversarial attacks. We also look into the generalization and difficulty of existing SOD datasets. Finally, we discuss several open issues of SOD and outline future research directions. All the saliency prediction maps, our constructed dataset with annotations, and codes for evaluation are publicly available at https://github.com/wenguanwang/SODsurvey. Wenguan Wang, Qiuxia Lai, Huazhu Fu, Jianbing Shen, Haibin Ling, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Hierarchical Human Semantic Parsing With Comprehensive Part-Relation ModelingabstractModeling the human structure is central for human parsing that extracts pixel-wise semantic information from images. We start with analyzing three types of inference processes over the hierarchical structure of human bodies: direct inference (directly predicting human semantic parts using image information), bottom-up inference (assembling knowledge from constituent parts), and top-down inference (leveraging context from parent nodes). We then formulate the problem as a compositional neural information fusion (CNIF) framework, which assembles the information from the three inference processes in a conditional manner, i.e., considering the confidence of the sources. Based on CNIF, we further present a part-relation-aware human parser (PRHP), which precisely describes three kinds of human part relations, i.e., decomposition, composition, and dependency, by three distinct relation networks. Expressive relation information can be captured by imposing the parameters in the relation networks to satisfy specific geometric characteristics of different relations. By assimilating generic message-passing networks with their edge-typed, convolutional counterparts, PRHP performs iterative reasoning over the human body hierarchy. With these efforts, PRHP provides a more general and powerful form of CNIF, and lays the foundation for more sophisticated and flexible human relation patterns of reasoning. Experiments on five datasets demonstrate that our two human parsers outperform the state-of-the-arts in all cases. Wenguan Wang, Tianfei Zhou, Siyuan Qi, Jianbing Shen, Song-Chun Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Deep Learning for Person Re-Identification: A Survey and OutlookabstractPerson re-identification (Re-ID) aims at retrieving a person of interest across multiple non-overlapping cameras. With the advancement of deep neural networks and increasing demand of intelligent video surveillance, it has gained significantly increased interest in the computer vision community. By dissecting the involved components in developing a person Re-ID system, we categorize it into the closed-world and open-world settings. The widely studied closed-world setting is usually applied under various research-oriented assumptions, and has achieved inspiring success using deep learning techniques on a number of datasets. We first conduct a comprehensive overview with in-depth analysis for closed-world person Re-ID from three different perspectives, including deep feature representation learning, deep metric learning and ranking optimization. With the performance saturation under closed-world setting, the research focus for person Re-ID has recently shifted to the open-world setting, facing more challenging issues. This setting is closer to practical applications under specific scenarios. We summarize the open-world Re-ID in terms of five different aspects. By analyzing the advantages of existing methods, we design a powerful AGW baseline, achieving state-of-the-art or at least comparable performance on twelve datasets for four different Re-ID tasks. Meanwhile, we introduce a new evaluation metric (mINP) for person Re-ID, indicating the cost for finding all the correct matches, which provides an additional criteria to evaluate the Re-ID system for real applications. Finally, some important yet under-investigated open issues are discussed. Mang Ye, Jianbing Shen, Gaojie Lin, Tao Xiang 0002, Ling Shao 0001, Steven C. H. Hoi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Augmentation Invariant and Instance Spreading Feature for Softmax EmbeddingabstractDeep embedding learning plays a key role in learning discriminative feature representations, where the visually similar samples are pulled closer and dissimilar samples are pushed away in the low-dimensional embedding space. This paper studies the unsupervised embedding learning problem by learning such a representation without using any category labels. This task faces two primary challenges: mining reliable positive supervision from highly similar fine-grained classes, and generalizing to unseen testing categories. To approximate the positive concentration and negative separation properties in category-wise supervised learning, we introduce a data augmentation invariant and instance spreading feature using the instance-wise supervision. We also design two novel domain-agnostic augmentation strategies to further extend the supervision in feature space, which simulates the large batch training using a small batch size and the augmented features. To learn such a representation, we propose a novel instance-wise softmax embedding, which directly perform the optimization over the augmented instance features with the binary discrmination softmax encoding. It significantly accelerates the learning speed with much higher accuracy than existing methods, under both seen and unseen testing categories. The unsupervised embedding performs well even without pre-trained network over samples from fine-grained categories. We also develop a variant using category-wise supervision, namely category-wise softmax embedding, which achieves competitive performance over the state-of-of-the-arts, without using any auxiliary information or restrict sample mining. Mang Ye, Jianbing Shen, Xu Zhang 0022, Pong C. Yuen, Shih-Fu Chang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Cascaded Parsing of Human-Object Interaction RecognitionabstractThis paper addresses the task of detecting and recognizing human-object interactions (HOI) in images. Considering the intrinsic complexity and structural nature of the task, we introduce a cascaded parsing network (CP-HOI) for a multi-stage, structured HOI understanding. At each cascade stage, an instance detection module progressively refines HOI proposals and feeds them into a structured interaction reasoning module. Each of the two modules is also connected to its predecessor in the previous stage, enabling efficient cross-stage information propagation. The structured interaction reasoning module is built upon a graph parsing neural network (GPNN), which efficiently models potential HOI structures as graphs and mines rich context for comprehensive relation understanding. In particular, GPNN infers a parse graph that i) interprets meaningful HOI structures by a learnable adjacency matrix, and ii) predicts action (edge) labels. Within an end-to-end, message-passing framework, GPNN blends learning and inference, iteratively parsing HOI structures and reasoning HOI representations (i.e., instance and relation features). Further beyond relation detection at a bounding-box level, we make our framework flexible to perform fine-grained pixel-wise relation segmentation; this provides a new glimpse into better relation modeling. A preliminary version of our CP-HOI model reached 1stplace in the ICCV2019 Person in Context Challenge, on both relation detection and segmentation. In addition, our CP-HOI shows promising results on two popular HOI recognition benchmarks,i.e., V-COCO and HICO-DET. Tianfei Zhou, Siyuan Qi, Wenguan Wang, Jianbing Shen, Song-Chun Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Person Re-Identification by Context-Aware Part Attention and Multi-Head Collaborative LearningabstractMost existing works solve the video-based person re-identification (re-ID) problem by computing the representation of each frame independently and finally aggregate the frame-level features. However, these methods often suffer from the challenging factors in videos, such as serious occlusion, background clutter and pose variation. To address these issues, we propose a novel multi-level Context-aware Part Attention (CPA) model to learn discriminative and robust local part features. It is featured in two aspects: 1) the context-aware part attention module improves the robustness by capturing the global relationship among different body parts across different video frames, and 2) the attention module is further extended to multi-level attention mechanism which enhances the discriminability by simultaneously considering low- to high-level features in different convolutional layers. In addition, we propose a novel multi-head collaborative training scheme to improve the performance, which is collaboratively supervised by multiple heads with the same structure but different parameters. It contains two consistency regularization terms, which consider both multi-head and multi-frame consistency to achieve better results. The multi-level CPA model is designed for feature extraction, while the multi-head collaborative training scheme is designed for classifier supervision. They jointly improve our re-ID model from two complementary directions. Extensive experiments demonstrate that the proposed method achieves much better or at least comparable performance compared to the state-of-the-art on four video re-ID datasets. Dongming Wu 0005, Mang Ye, Gaojie Lin, Xin Gao 0001, Jianbing Shen |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | Dynamic Tri-Level Relation Mining With Attentive Graph for Visible Infrared Re-IdentificationabstractMatching the daytime visible and nighttime infrared person images, namely visible infrared person re-identification (VI-ReID), is a challenging cross-modality retrieval problem. Due to the difficulty of data collection and annotation in nighttime surveillance, VI-ReID usually suffers from noise problems, making it challenging to directly learn part discriminative features. In order to improve the discriminability and enhance the robustness against noisy images, this paper proposes a novel dynamic tri-level relation mining (DTRM) framework by simultaneously exploring channel-level, part-level intra-modality, and graph-level cross-modality relation cues. To address the misalignment within the person images, we design an intra-modality weighted-part attention (IWPA) to construct part-aggregated representation. It adaptively integrates the body part relation into the local feature learning with a residual batch normalization (RBN) connection scheme. Besides, a cross-modality graph structured attention (CGSA) is incorporated to improve the global feature learning by utilizing the contextual relation between images from two modalities. This module reduces the negative effects of noisy images. To seamlessly integrate two components, a parameter-free dynamic aggregation strategy is designed in a progressive joint learning manner. To further improve the performance, we additionally design a simple yet effective channel-level learning strategy by exploiting the rich channel information of visible images, which significantly reinforces the performance without modifying the network structure or changing the training process. Extensive experiments on two visible infrared re-identification datasets have verified the effectiveness under various settings. Code is available at:https://github.com/mangye16/DDAG Mang Ye, Cuiqun Chen, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Weakly Supervised Visual Saliency PredictionabstractThe success of current deep saliency models heavily depends on large amounts of annotated human fixation data to fit the highly non-linear mapping between the stimuli and visual saliency. Such fully supervised data-driven approaches are annotation-intensive and often fail to consider the underlying mechanisms of visual attention. In contrast, in this paper, we introduce a model based on various cognitive theories of visual saliency, which learns visual attention patterns in a weakly supervised manner. Our approach incorporates insights from cognitive science as differentiable submodules, resulting in a unified, end-to-end trainable framework. Specifically, our model encapsulates the following important components motivated from biological vision. (a) As scene semantics are closely related to visually attentive regions, our model encodes discriminative spatial information for scene understanding through spatial visual semantics embedding. (b) To model the objectness factors in visual attention deployment, we incorporate object-level semantics embedding and object relation information. (c) Considering the "winner-take-all" mechanism in visual stimuli processing, we model the competition mechanism among objects with softmax based neural attention. (d) Lastly, a conditional center prior is learned to mimic the spatial distribution bias of visual attention. Furthermore, we propose novel loss functions to utilize supervision cues from image-level semantics, saliency prior knowledge, and self-information compression. Experiments show that our method achieves promising results, and even outperforms many of its fully supervised counterparts. Overall, our weakly supervised saliency method makes an essential step towards reducing the annotation budget of current approaches, as well as providing a more comprehensive understanding of the visual attention mechanism. Our code is available at: https://github.com/ashleylqx/WeakFixation.git. Qiuxia Lai, Tianfei Zhou, Salman Khan 0001, Hanqiu Sun, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | Person Foreground Segmentation by Learning Multi-Domain NetworksabstractSeparating the dominant person from the complex background is significant to the human-related research and photo-editing based applications. Existing segmentation algorithms are either too general to separate the person region accurately, or not capable of achieving real-time speed. In this paper, we introduce the multi-domain learning framework into a novel baseline model to construct the Multi-domain TriSeNet Networks for the real-time single person image segmentation. We first divide training data into different subdomains based on the characteristics of single person images, then apply a multi-branch Feature Fusion Module (FFM) to decouple the networks into the domain-independent and the domain-specific layers. To further enhance the accuracy, a self-supervised learning strategy is proposed to dig out domain relations during training. It helps transfer domain-specific knowledge by improving predictive consistency among different FFM branches. Moreover, we create a large-scale single person image segmentation dataset named MSSP20k, which consists of 22,100 pixel-level annotated images in the real world. The MSSP20k dataset is more complex and challenging than existing public ones in terms of scalability and variety. Experiments show that our Multi-domain TriSeNet outperforms state-of-the-art approaches on both public and the newly built datasets with real-time speed. Zhiyuan Liang, Kan Guo, Xiaogang Jin 0001, Jianbing Shen |
IEEE Trans. Image Process. | 5 |
| 2022 | Collaborative Refining for Person Re-Identification With Label NoiseabstractExisting person re-identification (Re-ID) methods usually rely heavily on large-scale thoroughly annotated training data. However, label noise is unavoidable due to inaccurate person detection results or annotation errors in real scenes. It is extremely challenging to learn a robust Re-ID model with label noise since each identity has very limited annotated training samples. To avoid fitting to the noisy labels, we propose to learn a prefatory model using a large learning rate at the early stage with a self-label refining strategy, in which the labels and network are jointly optimized. To further enhance the robustness, we introduce an online co-refining (CORE) framework with dynamic mutual learning, where networks and label predictions are online optimized collaboratively by distilling the knowledge from other peer networks. Moreover, it also reduces the negative impact of noisy labels using a favorable selective consistency strategy. CORE has two primary advantages: it is robust to different noise types and unknown noise ratios; it can be easily trained without much additional effort on the architecture design. Extensive experiments on Re-ID and image classification demonstrate that CORE outperforms its counterparts by a large margin under both practical and simulated noise settings. Notably, it also improves the state-of-the-art unsupervised Re-ID performance under standard settings. Code is available at https://github.com/mangye16/ReID-Label-Noise. Mang Ye, He Li 0054, Bo Du 0001, Jianbing Shen, Ling Shao 0001, Steven C. H. Hoi |
IEEE Trans. Image Process. | 4 |
| 2021 | Video Object Segmentation Using Global and Instance Embedding LearningabstractIn this paper, we propose a feature embedding based video object segmentation (VOS) method which is simple, fast and effective. The current VOS task involves two main challenges: object instance differentiation and cross-frame instance alignment. Most state-of-the-art matching based VOS methods simplify this task into a binary segmentation task and tackle each instance independently. In contrast, we decompose the VOS task into two subtasks: global embedding learning that segments foreground objects of each frame in a pixel-to-pixel manner, and instance feature embedding learning that separates instances. The outputs of these two subtasks are fused to obtain the final instance masks quickly and accurately. Through using the relation among different instances per-frame as well as temporal relation across different frames, the proposed network learns to differentiate multiple instances and associate them properly in one feed-forward manner. Extensive experimental results on the challenging DAVIS[34] and Youtube-VOS [57] datasets show that our method achieves better performances than most counterparts in each case. Wenbin Ge, Xiankai Lu, Jianbing Shen |
CVPR | 3 |
| 2021 | Learning To Fuse Asymmetric Feature Maps in Siamese TrackersabstractRecently, Siamese-based trackers have achieved promising performance in visual tracking. Most recent Siamese-based trackers typically employ a depth-wise cross-correlation (DW-XCorr) to obtain multi-channel correlation information from the two feature maps (target and search region). However, DW-XCorr has several limitations within Siamese-based tracking: it can easily be fooled by distractors, has fewer activated channels and provides weak discrimination of object boundaries. Further, DW-XCorr is a handcrafted parameter-free module and cannot fully benefit from offline learning on large-scale data.We propose a learnable module, called the asymmetric convolution (ACM), which learns to better capture the se-mantic correlation information in offline training on large-scale data. Different from DW-XCorr and its predecessor (XCorr), which regard a single feature map as the convolution kernel, our ACM decomposes the convolution operation on a concatenated feature map into two mathematically equivalent operations, thereby avoiding the need for the feature maps to be of the same size (width and height) during concatenation. Our ACM can incorporate useful prior information, such as bounding-box size, with standard visual features. Furthermore, ACM can easily be integrated into existing Siamese trackers based on DW-XCorr or XCorr. To demonstrate its generalization ability, we integrate ACM into three representative trackers: SiamFC, SiamRPN++ and SiamBAN. Our experiments reveal the benefits of the proposed ACM, which outperforms existing methods on six tracking benchmarks. On the LaSOT test set, our ACM-based tracker obtains a significant improvement of 5.8% in terms of success (AUC), over the baseline. Wencheng Han, Xingping Dong, Fahad Shahbaz Khan, Ling Shao 0001, Jianbing Shen |
CVPR | 5 |
| 2021 | Structured Scene Memory for Vision-Language NavigationabstractRecently, numerous algorithms have been developed to tackle the problem of vision-language navigation (VLN), i.e., entailing an agent to navigate 3D environments through following linguistic instructions. However, current VLN agents simply store their past experiences/observations as latent states in recurrent networks, failing to capture environment layouts and make long-term planning. To address these limitations, we propose a crucial architecture, called Structured Scene Memory (SSM). It is compartmentalized enough to accurately memorize the percepts during navigation. It also serves as a structured scene representation, which captures and disentangles visual and geometric cues in the environment. SSM has a collect-read controller that adaptively collects information for supporting current decision making and mimics iterative algorithms for long-range reasoning. As SSM provides a complete action space, i.e., all the navigable places on the map, a frontier-exploration based navigation decision making strategy is introduced to enable efficient and global planning. Experiment results on two VLN datasets (i.e., R2R and R4R) show that our method achieves state-of-the-art performance on several metrics. Hanqing Wang 0001, Wenguan Wang, Wei Liang 0008, Caiming Xiong, Jianbing Shen |
CVPR | 5 |
| 2021 | Face Forensics in the WildabstractOn existing public benchmarks, face forgery detection techniques have achieved great success. However, when used in multi-person videos, which often contain many people active in the scene with only a small subset having been manipulated, their performance remains far from being satisfactory. To take face forgery detection to a new level, we construct a novel large-scale dataset, called FFIW10K, which comprises 10,000 high-quality forgery videos, with an average of three human faces in each frame. The manipulation procedure is fully automatic, controlled by a domain-adversarial quality assessment network, making our dataset highly scalable with low human cost. In addition, we propose a novel algorithm to tackle the task of multi-person face forgery detection. Supervised by only video-level label, the algorithm explores multiple instance learning and learns to automatically attend to tampered faces. Our algorithm outperforms representative approaches for both forgery classification and localization on FFIW10K, and also shows high generalization ability on existing benchmarks. We hope that our dataset and study will help the community to explore this new field in more depth. Tianfei Zhou, Wenguan Wang, Zhiyuan Liang, Jianbing Shen |
CVPR | 4 |
| 2021 | Cross-Modality Person Re-Identification via Modality Confusion and Center AggregationabstractCross-modality person re-identification is a challenging task due to large cross-modality discrepancy and intramodality variations. Currently, most existing methods focus on learning modality-specific or modality-shareable features by using the identity supervision or modality label. Different from existing methods, this paper presents a novel Modality Confusion Learning Network (MCLNet). Its basic idea is to confuse two modalities, ensuring that the optimization is explicitly concentrated on the modality-irrelevant perspective. Specifically, MCLNet is designed to learn modality-invariant features by simultaneously minimizing inter-modality discrepancy while maximizing cross-modality similarity among instances in a single framework. Furthermore, an identity-aware marginal center aggregation strategy is introduced to extract the centralization features, while keeping diversity with a marginal constraint. Finally, we design a camera-aware learning scheme to enrich the discriminability. Extensive experiments on SYSU-MM01 and RegDB datasets show that MCLNet outperforms the state-of-the-art by a large margin. On the large-scale SYSU-MM01 dataset, our model can achieve 65.40 % and 61.98 % in terms of Rank-1 accuracy and mAP value. Xin Hao, Sanyuan Zhao, Mang Ye, Jianbing Shen |
ICCV | 4 |
| 2021 | Full-Duplex Strategy for Video Object SegmentationabstractAppearance and motion are two important sources of information in video object segmentation (VOS). Previous methods mainly focus on using simplex solutions, lowering the upper bound of feature collaboration among and across these two cues. In this paper, we study a novel framework, termed the FSNet (Full-duplex Strategy Network), which designs a relational cross-attention module (RCAM) to achieve the bidirectional message propagation across embedding subspaces. Furthermore, the bidirectional purification module (BPM) is introduced to update the inconsistent features between the spatial-temporal embeddings, effectively improving the model robustness. By considering the mutual restraint within the full-duplex strategy, our FSNet performs the cross-modal feature-passing (i.e., transmission and receiving) simultaneously before the fusion and decoding stage, making it robust to various challenging scenarios (e.g., motion blur, occlusion) in VOS. Extensive experiments on five popular benchmarks (i.e., DAVIS16, FBMS, MCL, SegTrack-V2, and DAVSOD19) show that our FSNet outperforms other state-of-the-arts for both the VOS and video salient object detection tasks. Ge-Peng Ji, Keren Fu, Deng-Ping Fan, Jianbing Shen, Ling Shao 0001 |
ICCV | 5 |
| 2021 | Robust Shadow Detection by Exploring Effective Shadow ContextsabstractEffective contexts for separating shadows from non-shadow objects can appear in different scales due to different object sizes. This paper introduces a new module, Effective-Context Augmentation (ECA), to utilize these contexts for robust shadow detection with deep structures. Taking regular deep features as global references, ECA enhances the discriminative features from the parallelly computed fine-scale features and, therefore, obtains robust features embedded with effective object contexts by boosting them. We further propose a novel encoder-decoder style of shadow detection method where ECA acts as the main building block of the encoder to extract strong feature representations and the guidance to the classification process of the decoder. Moreover, the networks are optimized with only one loss, which is easy to train and does not have the instability caused by extra losses superimposed on the intermediate features among existing popular studies. Experimental results show that the proposed method can effectively eliminate fake detections. Especially, our method outperforms state-of-the-arts methods and improves over $13.97%$ and $34.67%$ on the challenging SBU and UCF datasets respectively in balance error rate. Xianyong Fang, Xiaohao He, Linbo Wang 0001, Jianbing Shen |
ACM Multimedia | 4 |
| 2021 | RGB-D salient object detection: A surveyabstractSalient object detection, which simulates human visual perception in locating the most significant object(s) in a scene, has been widely applied to various computer vision tasks. Now, the advent of depth sensors means that depth maps can easily be captured; this additional spatial information can boost the performance of salient object detection. Although various RGB-D based salient object detection models with promising performance have been proposed over the past several years, an in-depth understanding of these models and the challenges in this field remains lacking. In this paper, we provide a comprehensive survey of RGB-D based salient object detection models from various perspectives, and review related benchmark datasets in detail. Further, as light fields can also provide depth maps, we review salient object detection models and popular benchmark datasets from this domain too. Moreover, to investigate the ability of existing models to detect salient objects, we have carried out a comprehensive attribute-based evaluation of several representative RGB-D based salient object detection models. Finally, we discuss several challenges and open directions of RGB-D based salient object detection for future research. All collected models, benchmark datasets, datasets constructed for attribute-based evaluation, and related code are publicly available at https://github.com/taozh2017/RGBD-SODsurvey. Tao Zhou 0002, Deng-Ping Fan, Ming-Ming Cheng, Jianbing Shen, Ling Shao 0001 |
Comput. Vis. Media | 4 |
| 2021 | Video person re-identification with global statistic pooling and self-attention distillation
Gaojie Lin, Sanyuan Zhao, Jianbing Shen |
Neurocomputing | 3 |
| 2021 | Deep understanding of big geospatial data for self-driving cars
Shuo Shang, Jianbing Shen, Ji-Rong Wen, Panos Kalnis |
Neurocomputing | 2 |
| 2021 | Dynamical Hyperparameter Optimization via Deep Reinforcement Learning in TrackingabstractHyperparameters are numerical pre-sets whose values are assigned prior to the commencement of a learning process. Selecting appropriate hyperparameters is often critical for achieving satisfactory performance in many vision problems, such as deep learning-based visual object tracking. However, it is often difficult to determine their optimal values, especially if they are specific to each video input. Most hyperparameter optimization algorithms tend to search a generic range and are imposed blindly on all sequences. In this paper, we propose a novel dynamical hyperparameter optimization method that adaptively optimizes hyperparameters for a given sequence using an action-prediction network leveraged on continuous deep Q-learning. Since the observation space for object tracking is significantly more complex than those in traditional control problems, existing continuous deep Q-learning algorithms cannot be directly applied. To overcome this challenge, we introduce an efficient heuristic strategy to handle high dimensional state space, while also accelerating the convergence behavior. The proposed algorithm is applied to improve two representative trackers, a Siamese-based one and a correlation-filter-based one, to evaluate its generalizability. Their superior performances on several popular benchmarks are clearly demonstrated. Our source code is available at https://github.com/shenjianbing/dqltracking. Xingping Dong, Jianbing Shen, Wenguan Wang, Ling Shao 0001, Haibin Ling, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Deeply Supervised Discriminative Learning for Adversarial DefenseabstractDeep neural networks can easily be fooled by an adversary with minuscule perturbations added to an input image. The existing defense techniques suffer greatly under white-box attack settings, where an adversary has full knowledge of the network and can iterate several times to find strong perturbations. We observe that the main reason for the existence of such vulnerabilities is the close proximity of different class samples in the learned feature space of deep models. This allows the model decisions to be completely changed by adding an imperceptible perturbation to the inputs. To counter this, we propose to class-wise disentangle the intermediate feature representations of deep networks, specifically forcing the features for each class to lie inside a convex polytope that is maximally separated from the polytopes of other classes. In this manner, the network is forced to learn distinct and distant decision regions for each class. We observe that this simple constraint on the features greatly enhances the robustness of learned models, even against the strongest white-box attacks, without degrading the classification performance on clean images. We report extensive evaluations in both black-box and white-box attack scenarios and show significant gains in comparison to state-of-the-art defenses. Aamir Mustafa, Salman Khan 0001, Munawar Hayat, Roland Göcke, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Paying Attention to Video Object Pattern Understandingabstract) with dynamic eye-tracking data in the unsupervised video object segmentation (UVOS) setting. For the first time, we quantitatively verified the high consistency of visual attention behavior among human observers, and found strong correlation between human attention and explicit primary object judgments during dynamic, task-driven viewing. Such novel observations provide an in-depth insight of the underlying rationale behind video object pattens. Inspired by these findings, we decouple UVOS into two sub-tasks: UVOS-driven Dynamic Visual Attention Prediction (DVAP) in spatiotemporal domain, and Attention-Guided Object Segmentation (AGOS) in spatial domain. Our UVOS solution enjoys three major advantages: 1) modular training without using expensive video segmentation annotations, instead, using more affordable dynamic fixation data to train the initial video attention module and using existing fixation-segmentation paired static/image data to train the subsequent segmentation module; 2) comprehensive foreground understanding through multi-source learning; and 3) additional interpretability from the biologically-inspired and assessable attention. Experiments on four popular benchmarks show that, even without using expensive video object mask annotations, our model achieves compelling performance compared with state-of-the-arts and enjoys fast processing speed (10 fps on a single GPU). Our collected eye-tracking data and algorithm implementations have been made publicly available at https://github.com/wenguanwang/AGS. Wenguan Wang, Jianbing Shen, Xiankai Lu, Steven C. H. Hoi, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Revisiting Video Saliency Prediction in the Deep Learning EraabstractPredicting where people look in static scenes, a.k.a visual saliency, has received significant research interest recently. However, relatively less effort has been spent in understanding and modeling visual attention over dynamic scenes. This work makes three contributions to video saliency research. First, we introduce a new benchmark, called DHF1K (Dynamic Human Fixation 1K), for predicting fixations during dynamic scene free-viewing, which is a long-time need in this field. DHF1K consists of 1K high-quality elaborately-selected video sequences annotated by 17 observers using an eye tracker device. The videos span a wide range of scenes, motions, object types and backgrounds. Second, we propose a novel video saliency model, called ACLNet (Attentive CNN-LSTM Network), that augments the CNN-LSTM architecture with a supervised attention mechanism to enable fast end-to-end saliency learning. The attention mechanism explicitly encodes static saliency information, thus allowing LSTM to focus on learning a more flexible temporal saliency representation across successive frames. Such a design fully leverages existing large-scale static fixation datasets, avoids overfitting, and significantly improves training efficiency and testing performance. Third, we perform an extensive evaluation of the state-of-the-art saliency models on three datasets : DHF1K, Hollywood-2, and UCF sports. An attribute-based analysis of previous saliency models and cross-dataset generalization are also presented. Experimental results over more than 1.2K testing videos containing 400K frames demonstrate that ACLNet outperforms other contenders and has a fast processing speed (40 fps using a single GPU). Our code and all the results are available at https://github.com/wenguanwang/DHF1K. Wenguan Wang, Jianbing Shen, Jianwen Xie, Ming-Ming Cheng, Haibin Ling, Ali Borji |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Real-time and light-weighted unsupervised video object segmentation network
Zongji Zhao, Sanyuan Zhao, Jianbing Shen |
Pattern Recognit. | 3 |
| 2021 | Visible-Infrared Person Re-Identification via Homogeneous Augmented Tri-Modal LearningabstractMatching person images between the daytime visible modality and night-time infrared modality (VI-ReID) is a challenging cross-modality pedestrian retrieval problem. Existing methods usually learn the multi-modality features in raw image, ignoring the image-level discrepancy. Some methods apply GAN technique to generate the cross-modality images, but it destroys the local structure and introduces unavoidable noise. In this paper, we propose a Homogeneous Augmented Tri-Modal (HAT) learning method for VI-ReID, where an auxiliary grayscale modality is generated from their homogeneous visible images, without additional training process. It preserves the structure information of visible images and approximates the image style of infrared modality. Learning with the grayscale visible images enforces the network to mine structure relations across multiple modalities, making it robust to color variations. Specifically, we solve the tri-modal feature learning from both multi-modal classification and multi-view retrieval perspectives. For multi-modal classification, we learn a multi-modality sharing identity classifier with a parameter-sharing network, trained with a homogeneous and heterogeneous identification loss. For multi-view retrieval, we develop a weighted tri-directional ranking loss to optimize the relative distance across multiple modalities. Incorporated with two invariant regularizers, HAT simultaneously minimizes multiple modality variations. In-depth analysis demonstrates the homogeneous grayscale augmentation significantly outperforms the current state-of-the-art by a large margin. Mang Ye, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Improving Single Shot Object Detection With Feature Scale UnmixingabstractDue to the advantages of real-time detection and improved performance, single-shot detectors have gained great attention recently. To solve the complex scale variations, single-shot detectors make scale-aware predictions based on multiple pyramid layers. Typically, small objects are detected on shallow layers while large objects are detected on deep layers. However, the features in the pyramid are not scale-aware enough, which limits the detection performance. Two common problems in single-shot detectors caused by object scale variations can be observed: (1) false negative problem, i.e., small objects are easily missed due to the weak features; (2) part-false positive problem, i.e., the salient part of a large object is sometimes detected as an object. With this observation, a new Neighbor Erasing and Transferring (NET) mechanism is proposed for feature scale-unmixing to explore scale-aware features in this paper. In NET, a Neighbor Erasing Module (NEM) is designed to erase the salient features of large objects and emphasize the features of small objects in shallow layers. A Neighbor Transferring Module (NTM) is introduced to transfer the erased features and highlight large objects in deep layers. With this mechanism, a single-shot network called NETNet is constructed for scale-aware object detection. In addition, we propose to aggregate nearest neighboring pyramid features to enhance our NET. Experiments on MS COCO dataset and UAVDT dataset demonstrate the effectiveness of our method. NETNet obtains 38.5% AP at a speed of 27 FPS and 32.0% AP at a speed of 55 FPS on MS COCO dataset. As a result, NETNet achieves a better trade-off for real-time and accurate object detection. Yazhao Li, Yanwei Pang, Jiale Cao, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | MSB-FCN: Multi-Scale Bidirectional FCN for Object Skeleton ExtractionabstractThe performance of state-of-the-art object skeleton detection (OSD) methods have been greatly boosted by Convolutional Neural Networks (CNNs). However, the most existing CNN-based OSD methods rely on a 'skip-layer' structure where low-level and high-level features are combined to gather multi-level contextual information. Unfortunately, as shallow features tend to be noisy and lack semantic knowledge, they will cause errors and inaccuracy. Therefore, in order to improve the accuracy of object skeleton detection, we propose a novel network architecture, the Multi-Scale Bidirectional Fully Convolutional Network (MSB-FCN), to better gather and enhance multi-scale high-level contextual information. The advantage is that only deep features are used to construct multi-scale feature representations along with a bidirectional structure for better capturing contextual knowledge. This enables the proposed MSB-FCN to learn semantic-level information from different sub-regions. Moreover, we introduce dense connections into the bidirectional structure to ensure that the learning process at each scale can directly encode information from all other scales. An attention pyramid is also integrated into our MSB-FCN to dynamically control information propagation and reduce unreliable features. Extensive experiments on various benchmarks demonstrate that the proposed MSB-FCN achieves significant improvements over the state-of-the-art algorithms. Fan Yang 0054, Xin Li 0079, Jianbing Shen |
IEEE Trans. Image Process. | 3 |
| 2021 | Modeling and Enhancing Low-Quality Retinal Fundus ImagesabstractRetinal fundus images are widely used for the clinical screening and diagnosis of eye diseases. However, fundus images captured by operators with various levels of experience have a large variation in quality. Low-quality fundus images increase uncertainty in clinical observation and lead to the risk of misdiagnosis. However, due to the special optical beam of fundus imaging and structure of the retina, natural image enhancement methods cannot be utilized directly to address this. In this article, we first analyze the ophthalmoscope imaging system and simulate a reliable degradation of major inferior-quality factors, including uneven illumination, image blurring, and artifacts. Then, based on the degradation model, a clinically oriented fundus enhancement network (cofe-Net) is proposed to suppress global degradation factors, while simultaneously preserving anatomical retinal structures and pathological characteristics for clinical observation and analysis. Experiments on both synthetic and real images demonstrate that our algorithm effectively corrects low-quality fundus images without losing retinal details. Moreover, we also show that the fundus correction method can benefit medical image analysis applications, e.g., retinal vessel segmentation and optic disc/cup detection. Ziyi Shen, Huazhu Fu, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Understanding More About Human and Machine Attention in Deep Neural NetworksabstractHuman visual system can selectively attend to parts of a scene for quick perception, a biological mechanism known asHuman attention. Inspired by this, recent deep learning models encode attention mechanisms to focus on the most task-relevant parts of the input signal for further processing, which is calledMachine/Neural/Artificial attention. Understanding the relation between human and machine attention is important for interpreting and designing neural networks. Many works claim that the attention mechanism offers an extra dimension of interpretability by explaining where the neural networks look. However, recent studies demonstrate that artificial attention maps do not always coincide with common intuition. In view of these conflicting evidence, here we make a systematic study on using artificial attention and human attention in neural network design. With three example computer vision tasks (i.e., salient object segmentation, video action recognition, and fine-grained image classification), diverse representative backbones (i.e., AlexNet, VGGNet, ResNet) and famous architectures (i.e., Two-stream, FCN), corresponding real human gaze data, and systematically conducted large-scale quantitative studies, we quantify the consistency between artificial attention and human visual attention and offer novel insights into existing artificial attention mechanisms by giving preliminary answers to several key questions related to human and artificial attention mechanisms. Overall results demonstrate that human attention can benchmark the meaningful ‘ground-truth’ in attention-driven tasks, where the more the artificial attention is close to human attention, the better the performance; for higher-level vision tasks, it is case-by-case. It would be advisable for attention-driven tasks to explicitly force a better alignment between artificial and human attention to boost the performance; such alignment would also improve the network explainability for higher-level computer vision tasks. Qiuxia Lai, Salman Khan 0001, Yongwei Nie, Hanqiu Sun, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Multim. | 5 |
| 2021 | Capturing Relevant Context for Visual TrackingabstractStudies have shown that contextual information can promote the robustness of trackers. However, trackers based on convolutional neural networks (CNNs) only capture local features, which limits their performance. We propose a novel relevant context block (RCB), which employs graph convolutional networks to capture the relevant context. In particular, it selects the$k$largest contributors as nodes for each query position (unit) that contain meaningful and discriminative contextual information and updates the nodes by aggregating the differences between the query position and its contributors. This operation can be easily incorporated into the existing networks and can be easily end-to-end trained using a standard backpropagation algorithm. To verify the effectiveness of RCB, we apply it to two trackers, SiamFC and GlobalTrack, respectively, and the two improved trackers are referred to as Siam-RCB and GlobalTrack-RCB. Extensive experiments on OTB, VOT, UAV123, LaSOT, TrackingNet, OxUvA, and VOT2018LT show the superiority of our method. For example, our Siam-RCB outperforms SiamFC by a very large margin (up to 11.2% in the success score and 7.8% in the precision score) on the OTB-100 benchmark. Bo Ma 0001, Lianghua Huang, Jianbing Shen |
IEEE Trans. Multim. | 5 |
| 2020 | Multi-Mutual Consistency Induced Transfer Subspace Learning for Human Motion SegmentationabstractHuman motion segmentation based on transfer subspace learning is a rising interest in action-related tasks. Although progress has been made, there are still several issues within the existing methods. First, existing methods transfer knowledge from source data to target tasks by learning domain-invariant features, but they ignore to preserve domain-specific knowledge. Second, the transfer subspace learning is employed in either low-level or high-level feature spaces, but few methods consider fusing multi-level features for subspace learning. To this end, we propose a novel multi-mutual consistency induced transfer subspace learning framework for human motion segmentation. Specifically, our model factorizes the source and target data into distinct multi-layer feature spaces and reduces the distribution gap between them through a multi-mutual consistency learning strategy. In this way, the domain-specific knowledge and domain-invariant properties can be explored simultaneously. Our model also conducts the transfer subspace learning on different layers to capture multi-level structural information. Further, to preserve the temporal correlations, we project the learned representations into a block-like space. The proposed model is efficiently optimized by using the Augmented Lagrange Multiplier (ALM) algorithm. Experimental results on four human motion datasets demonstrate the effectiveness of our method over other state-of-the-art approaches. Tao Zhou 0002, Huazhu Fu, Chen Gong 0002, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
CVPR | 4 |
| 2020 | Camouflaged Object DetectionabstractWe present a comprehensive study on a new task named camouflaged object detection (COD), which aims to identify objects that are “seamlessly” embedded in their surroundings. The high intrinsic similarities between the target object and the background make COD far more challenging than the traditional object detection task. To address this issue, we elaborately collect a novel dataset, called COD10K, which comprises 10,000 images covering camouflaged objects in various natural scenes, over 78 object categories. All the images are densely annotated with category, bounding-box, object-/instance-level, and matting-level labels. This dataset could serve as a catalyst for progressing many vision tasks, e.g., localization, segmentation, and alpha-matting, etc. In addition, we develop a simple but effective framework for COD, termed Search Identification Network (SINet). Without any bells and whistles, SINet outperforms various state-of-the-art object detection baselines on all datasets tested, making it a robust, general framework that can help facilitate future research in COD. Finally, we conduct a large-scale COD study, evaluating 13 cutting-edge models, providing some interesting findings, and showing several potential applications. Our research offers the community an opportunity to explore more in this new field. The code will be available at https://github.com/DengPingFan/SINet/. Deng-Ping Fan, Ge-Peng Ji, Guolei Sun, Ming-Ming Cheng, Jianbing Shen, Ling Shao 0001 |
CVPR | 5 |
| 2020 | Self-Learning With Rectification Strategy for Human ParsingabstractIn this paper, we solve the sample shortage problem in the human parsing task. We begin with the self-learning strategy, which generates pseudo-labels for unlabeled data to retrain the model. However, directly using noisy pseudo-labels will cause error amplification and accumulation. Considering the topology structure of human body, we propose a trainable graph reasoning method that establishes internal structural connections between graph nodes to correct two typical errors in the pseudo-labels, i.e., the global structural error and the local consistency error. For the global error, we first transform category-wise features into a high-level graph model with coarse-grained structural information, and then decouple the high-level graph to reconstruct the category features. The reconstructed features have a stronger ability to represent the topology structure of the human body. Enlarging the receptive field of features can effectively reducing the local error. We first project feature pixels into a local graph model to capture pixel-wise relations in a hierarchical graph manner, then reverse the relation information back to the pixels. With the global structural and local consistency modules, these errors are rectified and confident pseudo-labels are generated for retraining. Extensive experiments on the LIP and the ATR datasets demonstrate the effectiveness of our global and local rectification modules. Our method outperforms other state-of-the-art methods in supervised human parsing tasks. Zhiyuan Liang, Sanyuan Zhao, Jiahao Gong, Jianbing Shen |
CVPR | 5 |
| 2020 | NETNet: Neighbor Erasing and Transferring Network for Better Single Shot Object DetectionabstractDue to the advantages of real-time detection and improved performance, single-shot detectors have gained great attention recently. To solve the complex scale variations, single-shot detectors make scale-aware predictions based on multiple pyramid layers. However, the features in the pyramid are not scale-aware enough, which limits the detection performance. Two common problems in single-shot detectors caused by object scale variations can be observed: (1) small objects are easily missed; (2) the salient part of a large object is sometimes detected as an object. With this observation, we propose a new Neighbor Erasing and Transferring (NET) mechanism to reconfigure the pyramid features and explore scale-aware features. In NET, a Neighbor Erasing Module (NEM) is designed to erase the salient features of large objects and emphasize the features of small objects in shallow layers. A Neighbor Transferring Module (NTM) is introduced to transfer the erased features and highlight large objects in deep layers. With this mechanism, a single-shot network called NETNet is constructed for scale-aware object detection. In addition, we propose to aggregate nearest neighboring pyramid features to enhance our NET. NETNet achieves 38.5% AP at a speed of 27 FPS and 32.0% AP at a speed of 55 FPS on MS COCO dataset. As a result, NETNet achieves a better trade-off for real-time and accurate object detection. Yazhao Li, Yanwei Pang, Jianbing Shen, Jiale Cao, Ling Shao 0001 |
CVPR | 3 |
| 2020 | Learning Video Object Segmentation From Unlabeled VideosabstractWe propose a new method for video object segmentation (VOS) that addresses object pattern learning from unlabeled videos, unlike most existing methods which rely heavily on extensive annotated data. We introduce a unified unsupervised/weakly supervised learning framework, called MuG, that comprehensively captures intrinsic properties of VOS at multiple granularities. Our approach can help advance understanding of visual patterns in VOS and significantly reduce annotation burden. With a carefully-designed architecture and strong representation learning ability, our learned model can be applied to diverse VOS settings, including object-level zero-shot VOS, instance-level zero-shot VOS, and one-shot VOS. Experiments demonstrate promising performance in these settings, as well as the potential of MuG in leveraging unlabeled data to further improve the segmentation accuracy. Xiankai Lu, Wenguan Wang, Jianbing Shen, Yu-Wing Tai, David Crandall, Steven C. H. Hoi |
CVPR | 3 |
| 2020 | Hierarchical Human Parsing With Typed Part-Relation ReasoningabstractHuman parsing is for pixel-wise human semantic understanding. As human bodies are underlying hierarchically structured, how to model human structures is the central theme in this task. Focusing on this, we seek to simultaneously exploit the representational capacity of deep graph networks and the hierarchical human structures. In particular, we provide following two contributions. First, three kinds of part relations, i.e., decomposition, composition, and dependency, are, for the first time, completely and precisely described by three distinct relation networks. This is in stark contrast to previous parsers, which only focus on a portion of the relations and adopt a type-agnostic relation modeling strategy. More expressive relation information can be captured by explicitly imposing the parameters in the relation networks to satisfy the specific characteristics of different relations. Second, previous parsers largely ignore the need for an approximation algorithm over the loopy human hierarchy, while we instead address an iterative reasoning process, by assimilating generic message-passing networks with their edge-typed, convolutional counterparts. With these efforts, our parser lays the foundation for more sophisticated and flexible human relation patterns of reasoning. Comprehensive experiments on five datasets demonstrate that our parser sets a new state-of-the-art on each. Wenguan Wang, Hailong Zhu, Jifeng Dai, Yanwei Pang, Jianbing Shen, Ling Shao 0001 |
CVPR | 5 |
| 2020 | Probabilistic Structural Latent Representation for Unsupervised EmbeddingabstractUnsupervised embedding learning aims at extracting low-dimensional visually meaningful representations from large-scale unlabeled images, which can then be directly used for similarity-based search. This task faces two major challenges: 1) mining positive supervision from highly similar fine-grained classes and 2) generating to unseen testing categories. To tackle these issues, this paper proposes a probabilistic structural latent representation (PSLR), which incorporates an adaptable softmax embedding to approximate the positive concentrated and negative instance separated properties in the graph latent space. It improves the discriminability by enlarging the positive/negative difference without introducing any additional computational cost while maintaining high learning efficiency. To address the limited supervision using data augmentation, a smooth variational reconstruction loss is introduced by modeling the intra-instance variance, which improves the robustness. Extensive experiments demonstrate the superiority of PSLR over state-of-the-art unsupervised methods on both seen and unseen categories with cosine similarity. Code is available at https://github.com/mangye16/PSLR. Mang Ye, Jianbing Shen |
CVPR | 2 |
| 2020 | LiDAR-Based Online 3D Video Object Detection With Graph-Based Message Passing and Spatiotemporal Transformer AttentionabstractExisting LiDAR-based 3D object detectors usually focus on the single-frame detection, while ignoring the spatiotemporal information in consecutive point cloud frames. In this paper, we propose an end-to-end online 3D video object detector that operates on point cloud sequences. The proposed model comprises a spatial feature encoding component and a spatiotemporal feature aggregation component. In the former component, a novel Pillar Message Passing Network (PMPNet) is proposed to encode each discrete point cloud frame. It adaptively collects information for a pillar node from its neighbors by iterative message passing, which effectively enlarges the receptive field of the pillar feature. In the latter component, we propose an Attentive Spatiotemporal Transformer GRU (AST-GRU) to aggregate the spatiotemporal information, which enhances the conventional ConvGRU with an attentive memory gating mechanism. AST-GRU contains a Spatial Transformer Attention (STA) module and a Temporal Transformer Attention (TTA) module, which can emphasize the foreground objects and align the dynamic objects, respectively. Experimental results demonstrate that the proposed 3D video object detector achieves state-of-the-art performance on the large-scale nuScenes benchmark. Junbo Yin, Jianbing Shen, Chenye Guan, Dingfu Zhou, Ruigang Yang |
CVPR | 2 |
| 2020 | A Unified Object Motion and Affinity Model for Online Multi-Object TrackingabstractCurrent popular online multi-object tracking (MOT) solutions apply single object trackers (SOTs) to capture object motions, while often requiring an extra affinity network to associate objects, especially for the occluded ones. This brings extra computational overhead due to repetitive feature extraction for SOT and affinity computation. Meanwhile, the model size of the sophisticated affinity network is usually non-trivial. In this paper, we propose a novel MOT framework that unifies object motion and affinity model into a single network, named UMA, in order to learn a compact feature that is discriminative for both object motion and affinity measure. In particular, UMA integrates single object tracking and metric learning into a unified triplet network by means of multi-task learning. Such design brings advantages of improved computation efficiency, low memory requirement and simplified training procedure. In addition, we equip our model with a task-specific attention module, which is used to boost task-aware feature learning. The proposed UMA can be easily trained end-to-end, and is elegant - requiring only one training stage. Experimental results show that it achieves promising performance on several MOT Challenge benchmarks. Junbo Yin, Wenguan Wang, Qinghao Meng, Ruigang Yang, Jianbing Shen |
CVPR | 5 |
| 2020 | Cascaded Human-Object Interaction RecognitionabstractRapid progress has been witnessed for human-object interaction (HOI) recognition, but most existing models are confined to single-stage reasoning pipelines. Considering the intrinsic complexity of the task, we introduce a cascade architecture for a multi-stage, coarse-to-fine HOI understanding. At each stage, an instance localization network progressively refines HOI proposals and feeds them into an interaction recognition network. Each of the two networks is also connected to its predecessor at the previous stage, enabling cross-stage information propagation. The interaction recognition network has two crucial parts: a relation ranking module for high-quality HOI proposal selection and a triple-stream classifier for relation prediction. With our carefully-designed human-centric relation features, these two modules work collaboratively towards effective interaction understanding. Further beyond relation detection on a bounding-box level, we make ourframework flexible to perform fine-grained pixel-wise relation segmentation; this provides a new glimpse into better relation modeling. Our approach reached the 1st place in the ICCV2019 Person in Context Challenge, on both relation detection and segmentation tasks. It also shows promising results on V-COCO. Tianfei Zhou, Wenguan Wang, Siyuan Qi, Haibin Ling, Jianbing Shen |
CVPR | 5 |
| 2020 | CLNet: A Compact Latent Network for Fast Adjusting Siamese Trackers
Xingping Dong, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
ECCV (20) | 2 |
| 2020 | Video Object Segmentation with Episodic Graph Memory Networks
Xiankai Lu, Wenguan Wang, Martin Danelljan, Tianfei Zhou, Jianbing Shen, Luc Van Gool |
ECCV (3) | 5 |
| 2020 | Weakly Supervised 3D Object Detection from Lidar Point Cloud
Qinghao Meng, Wenguan Wang, Tianfei Zhou, Jianbing Shen, Luc Van Gool, Dengxin Dai |
ECCV (13) | 4 |
| 2020 | Active Visual Information Gathering for Vision-Language Navigation
Hanqing Wang 0001, Wenguan Wang, Tianmin Shu, Wei Liang 0008, Jianbing Shen |
ECCV (22) | 5 |
| 2020 | Dynamic Dual-Attentive Aggregation Learning for Visible-Infrared Person Re-identification
Mang Ye, Jianbing Shen, David Crandall, Ling Shao 0001, Jiebo Luo 0001 |
ECCV (17) | 2 |
| 2020 | PraNet: Parallel Reverse Attention Network for Polyp Segmentation
Deng-Ping Fan, Ge-Peng Ji, Tao Zhou 0002, Geng Chen 0001, Huazhu Fu, Jianbing Shen, Ling Shao 0001 |
MICCAI (6) | 6 |
| 2020 | M2 Net: Multi-modal Multi-channel Network for Overall Survival Time Prediction of Brain Tumor Patients
Tao Zhou 0002, Huazhu Fu, Yu Zhang 0009, Changqing Zhang 0002, Xiankai Lu, Jianbing Shen, Ling Shao 0001 |
MICCAI (2) | 6 |
| 2020 | Efficient Light Deep Network for Street Scene ParsingabstractThe semantic segmentation is a dense pixel label pre-diction task, which takes quite a lot of resources and computation cost in most of the time. In our approach, we pay attention to balance the speed and better performance which outperforms the state of the art in speed and accuracy for real-time performance. We come up with the idea of new efficient deep backbone that can extract more semantic details, reduce the computation cost and be easy to deploy at the same time. We call our new backbone as Cascaded Mobile Network, which is proved to be very useful. Our proposed model achieves 72.1 mIOU on the CityScapes val, and 69.5 on CamVid. We achieve good balance between speed and accuracy. ZheHui Wang, Sanyuan Zhao, Jianbing Shen, Zhengchao Lei |
VCIP | 3 |
| 2020 | Multi-attention deep reinforcement learning and re-ranking for vehicle re-identification
Yu Liu 0074, Jianbing Shen, Haibo He |
Neurocomputing | 2 |
| 2020 | Multiple people tracking with articulation detection and stitching strategy
Yuanpei Liu, Junbo Yin, Dajiang Yu, Sanyuan Zhao, Jianbing Shen |
Neurocomputing | 5 |
| 2020 | Inferring Salient Objects from Human FixationsabstractPrevious research in visual saliency has been focused on two major types of models namely fixation prediction and salient object detection. The relationship between the two, however, has been less explored. In this work, we propose to employ the former model type to identify salient objects. We build a novel Attentive Saliency Network (ASNet)11.Available at: https://github.com/wenguanwang/ASNet. that learns to detect salient objects from fixations. The fixation map, derived at the upper network layers, mimics human visual attention mechanisms and captures a high-level understanding of the scene from a global view. Salient object detection is then viewed as fine-grained object-level saliency segmentation and is progressively optimized with the guidance of the fixation map in a top-down manner. ASNet is based on a hierarchy of convLSTMs that offers an efficient recurrent mechanism to sequentially refine the saliency features over multiple steps. Several loss functions, derived from existing saliency evaluation metrics, are incorporated to further boost the performance. Extensive experiments on several challenging datasets show that our ASNet outperforms existing methods and is capable of generating accurate segmentation maps with the help of the computed fixation prior. Our work offers a deeper insight into the mechanisms of attention and narrows the gap between salient object detection and fixation prediction. Wenguan Wang, Jianbing Shen, Xingping Dong, Ali Borji, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Motion-Aware Rapid Video Saliency DetectionabstractIn this paper, we propose a computationally efficient and consistently accurate spatiotemporal salient object detection method to identify the most noticeable object in a video sequence. Intuitively, the underlying motion in a video is a more stable saliency indicator than the apparent color cues that often contain significant variations and complex structures. Based on this observation, we build an efficient and accurate spatiotemporal saliency detection method that uses motion information as a leverage to locate the most dynamic regions in a video sequence. We first analyze the optical flow field to obtain foreground priors, and then incorporate spatial saliency features such as appearance contrasts and compactness measures, into a multi-cue integration framework to combine various saliency cues and achieve temporal consistency. Rigorous experiments on the challenging SegTrackV1, SegTrackV2, and FBMS datasets demonstrate that our method generates comparable or superior performance to state-of-the-art methods while running almost 100× faster at only 0.08 sec/frame. Promising performance and rapid speed imply that the proposed spatiotemporal saliency method can be easily involved in various vision applications. Wenguan Wang, Ziyi Shen, Jianbing Shen, Ling Shao 0001, Dacheng Tao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Adaptive Nonlocal Random Walks for Image Superpixel SegmentationabstractIn this paper, we propose a novel superpixel segmentation method using an adaptive nonlocal random walk (ANRW) algorithm. There are three main steps in our image superpixel segmentation algorithm. Our method is based on the random walk model, in which the seed points are produced to generate the initial superpixels by a gradient-based method in the first step. In the second step, the ANRW is proposed to get the initial superpixels by adjusting the NRW to obtain a better image and superpixel segmentation. In the last step, these small superpixels are merged to get the final regular and compact superpixels. The experimental results demonstrate that our method achieves a better superpixel performance than the state-of-the-art methods. Our source code will be available at: http://github.com/shenjianbing/ANRW. Jianbing Shen, Junbo Yin, Xingping Dong, Hanqiu Sun, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Text Image Deblurring Using Kernel Sparsity PriorabstractPrevious methods on text image motion deblurring seldom consider the sparse characteristics of the blur kernel. This paper proposes a new text image motion deblurring method by exploiting the sparse properties of both text image itself and kernel. It incorporates the L0-norm for regularizing the blur kernel in the deblurring model, besides the L0sparse priors for the text image and its gradient. Such a L0-norm-based model is efficiently optimized by half-quadratic splitting coupled with the fast conjugate descent method. To further improve the quality of the recovered kernel, a structure-preserving kernel denoising method is also developed to filter out the noisy pixels, yielding a clean kernel curve. Experimental results show the superiority of the proposed method. The source code and results are available at: https://github.com/shenjianbing/text-image-deblur. Xianyong Fang, Jianbing Shen, Christian Jacquemin, Ling Shao 0001 |
IEEE Trans. Cybern. | 3 |
| 2020 | Visual Object Tracking by Hierarchical Attention Siamese NetworkabstractVisual tracking addresses the problem of localizing an arbitrary target in video according to the annotated bounding box. In this article, we present a novel tracking method by introducing the attention mechanism into the Siamese network to increase its matching discrimination. We propose a new way to compute attention weights to improve matching performance by a sub-Siamese network [Attention Net (A-Net)], which locates attentive parts for solving the searching problem. In addition, features in higher layers can preserve more semantic information while features in lower layers preserve more location information. Thus, in order to solve the tracking failure cases by the higher layer features, we fully utilize location and semantic information by multilevel features and propose a new way to fuse multiscale response maps from each layer to obtain a more accurate position estimation of the object. We further propose a hierarchical attention Siamese network by combining the attention weights and multilayer integration for tracking. Our method is implemented with a pretrained network which can outperform most well-trained Siamese trackers even without any fine-tuning and online updating. The comparison results with the state-of-the-art methods on popular tracking benchmarks show that our method achieves better performance. Our source code and results will be available at https://github.com/shenjianbing/HASN. Jianbing Shen, Xingping Dong, Ling Shao 0001 |
IEEE Trans. Cybern. | 1 |
| 2020 | Video Saliency Prediction Using Spatiotemporal Residual Attentive NetworksabstractThis paper proposes a novel residual attentive learning network architecture for predicting dynamic eye-fixation maps. The proposed model emphasizes two essential issues, i.e, effective spatiotemporal feature integration and multi-scale saliency learning. For the first problem, appearance and motion streams are tightly coupled via dense residual cross connections, which integrate appearance information with multi-layer, comprehensive motion features in a residual and dense way. Beyond traditional two-stream models learning appearance and motion features separately, such design allows early, multi-path information exchange between different domains, leading to a unified and powerful spatiotemporal learning architecture. For the second one, we propose a composite attention mechanism that learns multi-scale local attentions and global attention priors end-to-end. It is used for enhancing the fused spatiotemporal features via emphasizing important features in multi-scales. A lightweight convolutional Gated Recurrent Unit (convGRU), which is flexible for small training data situation, is used for long-term temporal characteristics modeling. Extensive experiments over four benchmark datasets clearly demonstrate the advantage of the proposed video saliency model over other competitors and the effectiveness of each component of our network. Our code and all the results will be available at https://github.com/ashleylqx/STRA-Net. Qiuxia Lai, Wenguan Wang, Hanqiu Sun, Jianbing Shen |
IEEE Trans. Image Process. | 4 |
| 2020 | Local Semantic Siamese Networks for Fast TrackingabstractLearning a powerful feature representation is critical for constructing a robust Siamese tracker. However, most existing Siamese trackers learn the global appearance features of the entire object, which usually suffers from drift problems caused by partial occlusion or non-rigid appearance deformation. In this paper, we propose a new Local Semantic Siamese (LSSiam) network to extract more robust features for solving these drift problems, since the local semantic features contain more fine-grained and partial information. We learn the semantic features during offline training by adding a classification branch into the classical Siamese framework. To further enhance the representation of features, we design a generally focal logistic loss to mine the hard negative samples. During the online tracking, we remove the classification branch and propose an efficient template updating strategy to avoid aggressive computing load. Thus, the proposed tracker can run at a high-speed of 100 Frame-per-Second (FPS) far beyond real-time requirement. Extensive experiments on popular benchmarks demonstrate the proposed LSSiam tracker achieves the state-of-the-art performance with a high-speed. Our source code is available at. Zhiyuan Liang, Jianbing Shen |
IEEE Trans. Image Process. | 2 |
| 2020 | Image Super-Resolution as a Defense Against Adversarial AttacksabstractConvolutional Neural Networks have achieved significant success across multiple computer vision tasks. However, they are vulnerable to carefully crafted, human-imperceptible adversarial noise patterns which constrain their deployment in critical security-sensitive systems. This paper proposes a computationally efficient image enhancement approach that provides a strong defense mechanism to effectively mitigate the effect of such adversarial perturbations. We show that deep image restoration networks learn mapping functions that can bring off-the-manifold adversarial samples onto the natural image manifold, thus restoring classification towards correct classes. A distinguishing feature of our approach is that, in addition to providing robustness against attacks, it simultaneously enhances image quality and retains models performance on clean images. Furthermore, the proposed method does not modify the classifier or requires a separate mechanism to detect adversarial images. The effectiveness of the scheme has been demonstrated through extensive experiments, where it has proven a strong defense in gray-box settings. The proposed scheme is simple and has the following advantages: (1) it does not require any model training or parameter optimization, (2) it complements other existing defense mechanisms, (3) it is agnostic to the attacked model and attack type and (4) it provides superior performance across all popular attack algorithms. Our codes are publicly available at https://github.com/aamir-mustafa/super-resolution-adversarial-defense. Aamir Mustafa, Salman Khan 0001, Munawar Hayat, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Cross-Modality Person Re-Identification via Modality-Aware Collaborative Ensemble LearningabstractVisible thermal person re-identification (VT-ReID) is a challenging cross-modality pedestrian retrieval problem due to the large intra-class variations and modality discrepancy across different cameras. Existing VT-ReID methods mainly focus on learning cross-modality sharable feature representations by handling the modality-discrepancy in feature level. However, the modality difference in classifier level has received much less attention, resulting in limited discriminability. In this paper, we propose a novel modality-aware collaborative ensemble (MACE) learning method with middle-level sharable two-stream network (MSTN) for VT-ReID, which handles the modality-discrepancy in both feature level and classifier level. In feature level, MSTN achieves much better performance than existing methods by capturing sharable discriminative middlelevel features in convolutional layers. In classifier level, we introduce both modality-specific and modality-sharable identity classifiers for two modalities to handle the modality discrepancy. To utilize the complementary information among different classifiers, we propose an ensemble learning scheme to incorporate the modality sharable classifier and the modality specific classifiers. In addition, we introduce a collaborative learning strategy, which regularizes modality-specific identity predictions and the ensemble outputs. Extensive experiments on two cross-modality datasets demonstrate that the proposed method outperforms current state-of-the-art by a large margin, achieving rank- 1/mAP accuracy 51.64%/50.11% on the SYSU-MM01 dataset, and 72.37%/69.09% on the RegDB dataset. Mang Ye, Xiangyuan Lan, Qingming Leng, Jianbing Shen |
IEEE Trans. Image Process. | 4 |
| 2020 | MATNet: Motion-Attentive Transition Network for Zero-Shot Video Object SegmentationabstractIn this paper, we present a novel end-to-end learning neural network, i.e., MATNet, for zero-shot video object segmentation (ZVOS). Motivated by the human visual attention behavior, MATNet leverages motion cues as a bottom-up signal to guide the perception of object appearance. To achieve this, an asymmetric attention block, named Motion-Attentive Transition (MAT), is proposed within a two-stream encoder network to firstly identify moving regions and then attend appearance learning to capture the full extent of objects. Putting MATs in different convolutional layers, our encoder becomes deeply interleaved, allowing for close hierarchical interactions between object apperance and motion. Such a biologically-inspired design is proven to be superb to conventional two-stream structures, which treat motion and appearance independently in separate streams and often suffer severe overfitting to object appearance. Moreover, we introduce a bridge network to modulate multi-scale spatiotemporal features into more compact, discriminative and scale-sensitive representations, which are subsequently fed into a boundary-aware decoder network to produce accurate segmentation with crisp boundaries. We perform extensive quantitative and qualitative experiments on four challenging public benchmarks, i.e., DAVIS16, DAVIS17, FBMS and YouTube-Objects. Results show that our method achieves compelling performance against current state-of-the-art ZVOS methods. To further demonstrate the generalization ability of our spatiotemporal learning framework, we extend MATNet to another relevant task: dynamic visual attention prediction (DVAP). The experiments on two popular datasets (i.e., Hollywood-2 and UCF-Sports) further verify the superiority of our model. Our implementations have been made publicly available at https://github.com/tfzhou/MATNet. Tianfei Zhou, Jianwu Li, Shunzhou Wang, Ran Tao 0003, Jianbing Shen |
IEEE Trans. Image Process. | 5 |
| 2020 | Inf-Net: Automatic COVID-19 Lung Infection Segmentation From CT ImagesabstractCoronavirus Disease 2019 (COVID-19) spread globally in early 2020, causing the world to face an existential health crisis. Automated detection of lung infections from computed tomography (CT) images offers a great potential to augment the traditional healthcare strategy for tackling COVID-19. However, segmenting infected regions from CT slices faces several challenges, including high variation in infection characteristics, and low intensity contrast between infections and normal tissues. Further, collecting a large amount of data is impractical within a short time period, inhibiting the training of a deep model. To address these challenges, a novel COVID-19 Lung Infection Segmentation Deep Network (Inf-Net) is proposed to automatically identify infected regions from chest CT slices. In our Inf-Net, a parallel partial decoder is used to aggregate the high-level features and generate a global map. Then, the implicit reverse attention and explicit edge-attention are utilized to model the boundaries and enhance the representations. Moreover, to alleviate the shortage of labeled data, we present a semi-supervised segmentation framework based on a randomly selected propagation strategy, which only requires a few labeled images and leverages primarily unlabeled data. Our semi-supervised framework can improve the learning ability and achieve a higher performance. Extensive experiments on our COVID-SemiSeg and real CT volumes demonstrate that the proposed Inf-Net outperforms most cutting-edge segmentation models and advances the state-of-the-art performance. Deng-Ping Fan, Tao Zhou 0002, Ge-Peng Ji, Yi Zhou 0007, Geng Chen 0001, Huazhu Fu, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2020 | Hi-Net: Hybrid-Fusion Network for Multi-Modal MR Image SynthesisabstractMagnetic resonance imaging (MRI) is a widely used neuroimaging technique that can provide images of different contrasts (i.e., modalities). Fusing this multi-modal data has proven particularly effective for boosting model performance in many tasks. However, due to poor data quality and frequent patient dropout, collecting all modalities for every patient remains a challenge. Medical image synthesis has been proposed as an effective solution, where any missing modalities are synthesized from the existing ones. In this paper, we propose a novel Hybrid-fusion Network (Hi-Net) for multi-modal MR image synthesis, which learns a mapping from multi-modal source images (i.e., existing modalities) to target images (i.e., missing modalities). In our Hi-Net, a modality-specific network is utilized to learn representations for each individual modality, and a fusion network is employed to learn the common latent representation of multi-modal data. Then, a multi-modal synthesis network is designed to densely combine the latent representation with hierarchical features from each modality, acting as a generator to synthesize the target images. Moreover, a layer-wise multi-modal fusion strategy effectively exploits the correlations among multiple modalities, where a Mixed Fusion Block (MFB) is proposed to adaptively weight different fusion strategies. Extensive experiments demonstrate the proposed model outperforms other state-of-the-art medical image synthesis methods. Tao Zhou 0002, Huazhu Fu, Geng Chen 0001, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Reducing Estimation Bias via Triplet-Average Deep Deterministic Policy GradientabstractThe overestimation caused by function approximation is a well-known property in Q-learning algorithms, especially in single-critic models, which leads to poor performance in practical tasks. However, the opposite property, underestimation, which often occurs in Q-learning methods with double critics, has been largely left untouched. In this article, we investigate the underestimation phenomenon in the recent twin delay deep deterministic actor-critic algorithm and theoretically demonstrate its existence. We also observe that this underestimation bias does indeed hurt performance in various experiments. Considering the opposite properties of single-critic and double-critic methods, we propose a novel triplet-average deep deterministic policy gradient algorithm that takes the weighted action value of three target critics to reduce the estimation bias. Given the connection between estimation bias and approximation error, we suggest averaging previous target values to reduce per-update error and further improve performance. Extensive empirical results over various continuous control tasks in OpenAI gym show that our approach outperforms the state-of-the-art methods. Dongming Wu 0005, Xingping Dong, Jianbing Shen, Steven C. H. Hoi |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Shifting More Attention to Video Salient Object DetectionabstractThe last decade has witnessed a growing interest in video salient object detection (VSOD). However, the research community long-term lacked a well-established VSOD dataset representative of real dynamic scenes with high-quality annotations. To address this issue, we elaborately collected a visual-attention-consistent Densely Annotated VSOD (DAVSOD) dataset, which contains 226 videos with 23,938 frames that cover diverse realistic-scenes, objects, instances and motions. With corresponding real human eye-fixation data, we obtain precise ground-truths. This is the first work that explicitly emphasizes the challenge of saliency shift, i.e., the video salient object(s) may dynamically change. To further contribute the community a complete benchmark, we systematically assess 17 representative VSOD algorithms over seven existing VSOD datasets and our DAVSOD with totally ~84K frames (largest-scale). Utilizing three famous metrics, we then present a comprehensive and insightful performance analysis. Furthermore, we propose a baseline model. It is equipped with a saliency shift- aware convLSTM, which can efficiently capture video saliency dynamics through learning human attention-shift behavior. Extensive experiments open up promising future directions for model development and comparison. Deng-Ping Fan, Wenguan Wang, Ming-Ming Cheng, Jianbing Shen |
CVPR | 4 |
| 2019 | Striking the Right Balance With UncertaintyabstractLearning unbiased models on imbalanced datasets is a significant challenge. Rare classes tend to get a concentrated representation in the classification space which hampers the generalization of learned boundaries to new test examples. In this paper, we demonstrate that the Bayesian uncertainty estimates directly correlate with the rarity of classes and the difficulty level of individual samples. Subsequently, we present a novel framework for uncertainty based class imbalance learning that follows two key insights: First, classification boundaries should be extended further away from a more uncertain (rare) class to avoid over-fitting and enhance its generalization. Second, each sample should be modeled as a multi-variate Gaussian distribution with a mean vector and a covariance matrix defined by the sample's uncertainty. The learned boundaries should respect not only the individual samples but also their distribution in the feature space. Our proposed approach efficiently utilizes sample and class uncertainty information to learn robust features and more generalizable classifiers. We systematically study the class imbalance problem and derive a novel loss formulation for max-margin learning based on Bayesian uncertainty measure. The proposed method shows significant performance improvements on six benchmark datasets for face verification, attribute prediction, digit/object classification and skin lesion detection. Salman Khan 0001, Munawar Hayat, Syed Waqas Zamir, Jianbing Shen, Ling Shao 0001 |
CVPR | 4 |
| 2019 | See More, Know More: Unsupervised Video Object Segmentation With Co-Attention Siamese NetworksabstractWe introduce a novel network, called as CO-attention Siamese Network (COSNet), to address the unsupervised video object segmentation task from a holistic view. We emphasize the importance of inherent correlation among video frames and incorporate a global co-attention mechanism to improve further the state-of-the-art deep learning based solutions that primarily focus on learning discriminative foreground representations over appearance and motion in short-term temporal segments. The co-attention layers in our network provide efficient and competent stages for capturing global correlations and scene context by jointly computing and appending co-attention responses into a joint feature space. We train COSNet with pairs of video frames, which naturally augments training data and allows increased learning capacity. During the segmentation stage, the co-attention model encodes useful information by processing multiple reference frames together, which is leveraged to infer the frequently reappearing and salient foreground objects better. We propose a unified and end-to-end trainable framework where different co-attention variants can be derived for mining the rich context within videos. Our extensive experiments over three large benchmarks manifest that COSNet outperforms the current alternatives by a large margin. We will publicly release our implementation and models. Xiankai Lu, Wenguan Wang, Chao Ma 0004, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
CVPR | 4 |
| 2019 | An Iterative and Cooperative Top-Down and Bottom-Up Inference Network for Salient Object DetectionabstractThis paper presents a salient object detection method that integrates both top-down and bottom-up saliency inference in an iterative and cooperative manner. The top-down process is used for coarse-to-fine saliency estimation, where high-level saliency is gradually integrated with finer lower-layer features to obtain a fine-grained result. The bottom-up process infers the high-level, but rough saliency through gradually using upper-layer, semantically-richer features. These two processes are alternatively performed, where the bottom-up process uses the fine-grained saliency obtained from the top-down process to yield enhanced high-level saliency estimate, and the top-down process, in turn, is further benefited from the improved high-level information. The network layers in the bottom-up/top-down processes are equipped with recurrent mechanisms for layer-wise, step-by-step optimization. Thus, saliency information is effectively encouraged to flow in a bottom-up, top-down and intra-layer manner. We show that most other saliency models based on fully convolutional networks (FCNs) are essentially variants of our model. Extensive experiments on several famous benchmarks clearly demonstrate the superior performance, good generalization, and powerful learning ability of our proposed saliency inference framework. Wenguan Wang, Jianbing Shen, Ming-Ming Cheng, Ling Shao 0001 |
CVPR | 2 |
| 2019 | Learning Unsupervised Video Object Segmentation Through Visual AttentionabstractThis paper conducts a systematic study on the role of visual attention in Unsupervised Video Object Segmentation (UVOS) tasks. By elaborately annotating three popular video segmentation datasets (DAVIS, Youtube-Objects and SegTrack V2) with dynamic eye-tracking data in the UVOS setting, for the first time, we quantitatively verified the high consistency of visual attention behavior among human observers, and found strong correlation between human attention and explicit primary object judgements during dynamic, task-driven viewing. Such novel observations provide an in-depth insight into the underlying rationale behind UVOS. Inspired by these findings, we decouple UVOS into two sub-tasks: UVOS-driven Dynamic Visual Attention Prediction (DVAP) in spatiotemporal domain, and Attention-Guided Object Segmentation (AGOS) in spatial domain. Our UVOS solution enjoys three major merits: 1) modular training without using expensive video segmentation annotations, instead, using more affordable dynamic fixation data to train the initial video attention module and using existing fixation-segmentation paired static/image data to train the subsequent segmentation module; 2) comprehensive foreground understanding through multi-source learning; and 3) additional interpretability from the biologically-inspired and assessable attention. Experiments on popular benchmarks show that, even without using expensive video object mask annotations, our model achieves compelling performance in comparison with state-of-the-arts. Wenguan Wang, Hongmei Song, Shuyang Zhao, Jianbing Shen, Sanyuan Zhao, Steven C. H. Hoi, Haibin Ling |
CVPR | 4 |
| 2019 | Salient Object Detection With Pyramid Attention and Salient EdgesabstractThis paper presents a new method for detecting salient objects in images using convolutional neural networks (CNNs). The proposed network, named PAGE-Net, offers two key contributions. The first is the exploitation of an essential pyramid attention structure for salient object detection. This enables the network to concentrate more on salient regions while considering multi-scale saliency information. Such a stacked attention design provides a powerful tool to efficiently improve the representation ability of the corresponding network layer with an enlarged receptive field. The second contribution lies in the emphasis on the importance of salient edges. Salient edge information offers a strong cue to better segment salient objects and refine object boundaries. To this end, our model is equipped with a salient edge detection module, which is learned for precise salient boundary estimation. This encourages better edge-preserving salient object segmentation. Exhaustive experiments confirm that the proposed pyramid attention and salient edges are effective for salient object detection. We show that our deep saliency model outperforms state-of-the-art approaches for several benchmarks with a fast processing speed (25fps on one GPU). Wenguan Wang, Shuyang Zhao, Jianbing Shen, Steven C. H. Hoi, Ali Borji |
CVPR | 3 |
| 2019 | Gaussian Affinity for Max-Margin Class Imbalanced LearningabstractReal-world object classes appear in imbalanced ratios. This poses a significant challenge for classifiers which get biased towards frequent classes. We hypothesize that improving the generalization capability of a classifier should improve learning on imbalanced datasets. Here, we introduce the first hybrid loss function that jointly performs classification and clustering in a single formulation. Our approach is based on an 'affinity measure' in Euclidean space that leads to the following benefits: (1) direct enforcement of maximum margin constraints on classification boundaries, (2) a tractable way to ensure uniformly spaced and equidistant cluster centers, (3) flexibility to learn multiple class prototypes to support diversity and discriminability in feature space. Our extensive experiments demonstrate the significant performance improvements on visual classification and verification tasks on multiple imbalanced datasets. The proposed loss can easily be plugged in any deep architecture as a differentiable block and demonstrates robustness against different levels of data imbalance and corrupted labels. Munawar Hayat, Salman Khan 0001, Syed Waqas Zamir, Jianbing Shen, Ling Shao 0001 |
ICCV | 4 |
| 2019 | Adversarial Defense by Restricting the Hidden Space of Deep Neural NetworksabstractDeep neural networks are vulnerable to adversarial attacks which can fool them by adding minuscule perturbations to the input images. The robustness of existing defenses suffers greatly under white-box attack settings, where an adversary has full knowledge about the network and can iterate several times to find strong perturbations. We observe that the main reason for the existence of such perturbations is the close proximity of different class samples in the learned feature space. This allows model decisions to be totally changed by adding an imperceptible perturbation in the inputs. To counter this, we propose to class-wise disentangle the intermediate feature representations of deep networks. Specifically, we force the features for each class to lie inside a convex polytope that is maximally separated from the polytopes of other classes. In this manner, the network is forced to learn distinct and distant decision regions for each class. We observe that this simple constraint on the features greatly enhances the robustness of learned models, even against the strongest white-box attacks, without degrading the classification performance on clean images. We report extensive evaluations in both black-box and white-box attack scenarios and show significant gains in comparison to state-of-the art defenses. Aamir Mustafa, Salman Khan 0001, Munawar Hayat, Roland Göcke, Jianbing Shen, Ling Shao 0001 |
ICCV | 5 |
| 2019 | Towards Bridging Semantic Gap to Improve Semantic SegmentationabstractAggregating multi-level features is essential for capturing multi-scale context information for precise scene semantic segmentation. However, the improvement by directly fusing shallow features and deep features becomes limited as the semantic gap between them increases. To solve this problem, we explore two strategies for robust feature fusion. One is enhancing shallow features using a semantic enhancement module (SeEM) to alleviate the semantic gap between shallow features and deep features. The other strategy is feature attention, which involves discovering complementary information (i.e., boundary information) from low-level features to enhance high-level features for precise segmentation. By embedding these two strategies, we construct a parallel feature pyramid towards improving multi-level feature fusion. A Semantic Enhanced Network called SeENet is constructed with the parallel pyramid to implement precise segmentation. Experiments on three benchmark datasets demonstrate the effectiveness of our method for robust multi-level feature aggregation. As a result, our SeENet has achieved better performance than other state-of-the-art methods for semantic segmentation. Yanwei Pang, Yazhao Li, Jianbing Shen, Ling Shao 0001 |
ICCV | 3 |
| 2019 | Human-Aware Motion DeblurringabstractThis paper proposes a human-aware deblurring model that disentangles the motion blur between foreground (FG) humans and background (BG). The proposed model is based on a triple-branch encoder-decoder architecture. The first two branches are learned for sharpening FG humans and BG details, respectively; while the third one produces global, harmonious results by comprehensively fusing multi-scale deblurring information from the two domains. The proposed model is further endowed with a supervised, human-aware attention mechanism in an end-to-end fashion. It learns a soft mask that encodes FG human information and explicitly drives the FG/BG decoder-branches to focus on their specific domains. Above designs lead to a fully differentiable motion deblurring network, which can be trained end-to-end. To further benefit the research towards Human-aware Image Deblurring, we introduce a large-scale dataset, named HIDE, which consists of 8,422 blurry and sharp image pairs with 65,784 densely annotated FG human bounding boxes. HIDE is specifically built to span a broad range of scenes, human object sizes, motion patterns, and background complexities. Extensive experiments on public benchmarks and our dataset demonstrate that our model performs favorably against the state-of-the-art motion deblurring methods, especially in capturing semantic details. Ziyi Shen, Wenguan Wang, Xiankai Lu, Jianbing Shen, Haibin Ling, Tingfa Xu, Ling Shao 0001 |
ICCV | 4 |
| 2019 | Zero-Shot Video Object Segmentation via Attentive Graph Neural NetworksabstractThis work proposes a novel attentive graph neural network (AGNN) for zero-shot video object segmentation (ZVOS). The suggested AGNN recasts this task as a process of iterative information fusion over video graphs. Specifically, AGNN builds a fully connected graph to efficiently represent frames as nodes, and relations between arbitrary frame pairs as edges. The underlying pair-wise relations are described by a differentiable attention mechanism. Through parametric message passing, AGNN is able to efficiently capture and mine much richer and higher-order relations between video frames, thus enabling a more complete understanding of video content and more accurate foreground estimation. Experimental results on three video segmentation datasets show that AGNN sets a new state-of-the-art in each case. To further demonstrate the generalizability of our framework, we extend AGNN to an additional task: image object co-segmentation (IOCS). We perform experiments on two famous IOCS datasets and observe again the superiority of our AGNN model. The extensive experiments verify that AGNN is able to learn the underlying semantic/appearance relationships among video frames or related images, and discover the common objects. Wenguan Wang, Xiankai Lu, Jianbing Shen, David Crandall, Ling Shao 0001 |
ICCV | 3 |
| 2019 | Learning Compositional Neural Information Fusion for Human ParsingabstractThis work proposes to combine neural networks with the compositional hierarchy of human bodies for efficient and complete human parsing. We formulate the approach as a neural information fusion framework. Our model assembles the information from three inference processes over the hierarchy: direct inference (directly predicting each part of a human body using image information), bottom-up inference (assembling knowledge from constituent parts), and top-down inference (leveraging context from parent nodes). The bottom-up and top-down inferences explicitly model the compositional and decompositional relations in human bodies, respectively. In addition, the fusion of multi-source information is conditioned on the inputs, i.e., by estimating and considering the confidence of the sources. The whole model is end-to-end differentiable, explicitly modeling information flows and structures. Our approach is extensively evaluated on four popular datasets, outperforming the state-of-the-arts in all cases, with a fast processing speed of 23fps. Our code and results have been released to help ease future research in this direction. Wenguan Wang, Siyuan Qi, Jianbing Shen, Yanwei Pang, Ling Shao 0001 |
ICCV | 4 |
| 2019 | Multi-scale Capsule Attention-Based Salient Object Detection with Multi-crossed Layer ConnectionsabstractWith the popularization of convolutional networks being used for saliency models, saliency detection performance has achieved significant improvement. However, how to integrate accurate and crucial features for modeling saliency is still underexplored. In this paper, we present CapSalNet, which includes a multi-scale Capsule attention module and multi-crossed layer connections for Salient object detection. We first propose a novel capsule attention model, which integrates multi-scale contextual information with dynamic routing. Then, our model adaptively learns to aggregate multi-level features by using multi-crossed skip-layer connections. Finally, the predicted results are efficiently fused to generate the final saliency map in a coarse-to-fine manner. Comprehensive experiments on four benchmark datasets demonstrate that our proposed algorithm outperforms existing state-of-the-art approaches. Sanyuan Zhao, Jianbing Shen, Kin-Man Lam 0001 |
ICME | 3 |
| 2019 | Inter-modality Dependence Induced Data Recovery for MCI Conversion Prediction
Tao Zhou 0002, Kim-Han Thung, Yu Zhang 0009, Huazhu Fu, Jianbing Shen, Dinggang Shen, Ling Shao 0001 |
MICCAI (4) | 5 |
| 2019 | Evaluation of Retinal Image Quality Assessment Networks in Different Color-Spaces
Huazhu Fu, Jianbing Shen, Shanshan Cui, Yanwu Xu 0001, Jiang Liu 0001, Ling Shao 0001 |
MICCAI (1) | 3 |
| 2019 | Extreme Points Derived Confidence Map as a Cue for Class-Agnostic Interactive Segmentation Using Deep Neural Network
Shadab Khan, Ahmed H. Shahin, Javier Villafruela, Jianbing Shen, Ling Shao 0001 |
MICCAI (2) | 4 |
| 2019 | ET-Net: A Generic Edge-aTtention Guidance Network for Medical Image Segmentation
Huazhu Fu, Hang Dai, Jianbing Shen, Yanwei Pang, Ling Shao 0001 |
MICCAI (1) | 4 |
| 2019 | Deep Multi-modal Latent Representation Learning for Automated Dementia Diagnosis
Tao Zhou 0002, Mingxia Liu 0001, Huazhu Fu, Jun Wang 0024, Jianbing Shen, Ling Shao 0001, Dinggang Shen |
MICCAI (4) | 5 |
| 2019 | High-speed video salient object detection with temporal propagation using correlation filter
Sanyuan Zhao, Zhengchao Lei, Jianbing Shen, Yuanyuan Pang |
Neurocomputing | 5 |
| 2019 | A Deep Network Solution for Attention and Aesthetics Aware Photo CroppingabstractWe study the problem of photo cropping, which aims to find a cropping window of an input image to preserve as much as possible its important parts while being aesthetically pleasant. Seeking a deep learning-based solution, we design a neural network that has two branches for attention box prediction (ABP) and aesthetics assessment (AA), respectively. Given the input image, the ABP network predicts an attention bounding box as an initial minimum cropping window, around which a set of cropping candidates are generated with little loss of important information. Then, the AA network is employed to select the final cropping window with the best aesthetic quality among the candidates. The two sub-networks are designed to share the same full-image convolutional feature map, and thus are computationally efficient. By leveraging attention prediction and aesthetics assessment, the cropping model produces high-quality cropping results, even with the limited availability of training data for photo cropping. The experimental results on benchmark datasets clearly validate the effectiveness of the proposed approach. In addition, our approach runs at 5 fps, outperforming most previous solutions. The code and results are available at: https://github.com/shenjianbing/DeepCropping. Wenguan Wang, Jianbing Shen, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Semi-Supervised Video Object Segmentation with Super-TrajectoriesabstractWe introduce a semi-supervised video segmentation approach based on an efficient video representation, called as "super-trajectory". A super-trajectory corresponds to a group of compact point trajectories that exhibit consistent motion patterns, similar appearances, and close spatiotemporal relationships. We generate the compact trajectories using a probabilistic model, which enables handling of occlusions and drifts effectively. To reliably group point trajectories, we adopt a modified version of the density peaks based clustering algorithm that allows capturing rich spatiotemporal relations among trajectories in the clustering process. We incorporate two intuitive mechanisms for segmentation, called as reverse-tracking and object re-occurrence, for robustness and boosting the performance. Building on the proposed video representation, our segmentation method is discriminative enough to accurately propagate the initial annotations in the first frame onto the remaining frames. Our extensive experimental analyses on three challenging benchmarks demonstrate that, given the annotation in the first frame, our method is capable of extracting the target objects from complex backgrounds, and even reidentifying them after prolonged occlusions, producing high-quality video object segments. Wenguan Wang, Jianbing Shen, Fatih Porikli, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | A deep Coarse-to-Fine network for head pose estimation from synthetic data
Wei Liang 0008, Jianbing Shen, Yunde Jia, Lap-Fai Yu |
Pattern Recognit. | 3 |
| 2019 | A stable long-term object tracking method with re-detection strategy
Sanyuan Zhao, Qinghao Meng, Jianbing Shen |
Pattern Recognit. Lett. | 5 |
| 2019 | Better Dense Trajectories by Motion in VideosabstractCurrently, the most widely used point trajectories generation methods estimate the trajectories from the dense optical flow, by using a consistency check strategy to detect the occluded regions. However, these methods will miss some important trajectories, thus resulting in breaking smooth areas without any structure especially around the motion boundaries (MBs). We suggest exploring MBs in video to generate more accurate dense point trajectories. Estimating MBs from the video improves the point trajectory accuracy of the discontinuity or occluded areas. Then, we obtain trajectories by tracking the initial feature points through all frames. The experimental results demonstrate that our method outperforms the state-of-the-art methods on the challenging benchmark. Yu Liu 0074, Jianbing Shen, Wenguan Wang, Hanqiu Sun, Ling Shao 0001 |
IEEE Trans. Cybern. | 2 |
| 2019 | Stereo Video Object Segmentation Using Stereoscopic Foreground TrajectoriesabstractWe present an unsupervised segmentation framework for stereo videos using stereoscopic trajectories. The proposed stereo trajectory shows favorable properties for modeling the long-term motion information through the whole sequence and explicitly capturing the corresponding relationships between two stereo views. The stereo prior is important for inferring the desired object and guarantees the consistent spatial-temporal segmentation, which contributes to an enjoyable stereo experience. We start by deriving stereo trajectories from left and right views simultaneously, which are represented via a graph structure. Then we detect object-like stereo trajectories via the graph structure to efficiently infer the desired object. Finally, an energy optimization function is proposed to produce the stereo segmentation results via leveraging the object information from stereo trajectories. To benefit potential research, we collected a new stereoscopic video benchmark, which consists of a total of 50 stereo video clips and includes many challenges in segmentation. Extensive experimental results demonstrate that our stereo segmentation method achieves higher performance and preserves better stereo structures, compared with prevailing competitors. The source code and results are available at: https://github.com/shenjianbing/StereoSeg. Chang Liu 0071, Wenguan Wang, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Cybern. | 3 |
| 2019 | Multiobject Tracking by Submodular OptimizationabstractIn this paper, we propose a new multiobject visual tracking algorithm by submodular optimization. The proposed algorithm is composed of two main stages. At the first stage, a new selecting strategy of tracklets is proposed to cope with occlusion problem. We generate low-level tracklets using overlap criteria and min-cost flow, respectively, and then integrate them into a candidate tracklets set. In the second stage, we formulate the multiobject tracking problem as the submodular maximization problem subject to related constraints. The submodular function selects the correct tracklets from the candidate set of tracklets to form the object trajectory. Then, we design a connecting process which connects the corresponding trajectories to overcome the occlusion problem. Experimental results demonstrate the effectiveness of our tracking algorithm. Our source code is available at https://github.com/shenjianbing/submodulartrack. Jianbing Shen, Zhiyuan Liang, Jianhong Liu, Hanqiu Sun, Ling Shao 0001, Dacheng Tao |
IEEE Trans. Cybern. | 1 |
| 2019 | Quadruplet Network With One-Shot Learning for Fast Visual Object TrackingabstractIn the same vein of discriminative one-shot learning, Siamese networks allow recognizing an object from a single exemplar with the same class label. However, they do not take advantage of the underlying structure of the data and the relationship among the multitude of samples as they only rely on the pairs of instances for training. In this paper, we propose a new quadruplet deep network to examine the potential connections among the training instances, aiming to achieve a more powerful representation. We design a shared network with four branches that receive a multi-tuple of instances as inputs and are connected by a novel loss function consisting of pair loss and triplet loss. According to the similarity metric, we select the most similar and the most dissimilar instances as the positive and negative inputs of triplet loss from each multi-tuple. We show that this scheme improves the training performance. Furthermore, we introduce a new weight layer to automatically select suitable combination weights, which will avoid the conflict between triplet and pair loss leading to worse performance. We evaluate our quadruplet framework by model-free tracking-by-detection of objects from a single initial exemplar in several visual object tracking benchmarks. Our extensive experimental analysis demonstrates that our tracker achieves superior performance with a real-time processing speed of 78 frames/s. Our source code is available. Xingping Dong, Jianbing Shen, Dongming Wu 0005, Kan Guo, Xiaogang Jin 0001, Fatih Porikli |
IEEE Trans. Image Process. | 2 |
| 2019 | Robust Object Tracking Using Manifold Regularized Convolutional Neural NetworksabstractIn visual tracking, usually only a small number of samples are labeled, and most existing deep learning based trackers ignore abundant unlabeled samples that could provide additional information for deep trackers to boost their tracking performance. An intuitive way to explain unlabeled data is to incorporate manifold regularization into the common classification loss functions, but the high computational cost may prohibit those deep trackers from practical applications. To overcome this issue, we propose a two-stage approach to a deep tracker that takes into account both labeled and unlabeled samples. The annotation of unlabeled samples is propagated from its labeled neighbors first by exploring the manifold space that these samples are assumed to lie in. Then, we refine it by training a deep convolutional neural network using both labeled and unlabeled data in a supervised manner. Online visual tracking is further carried out under the framework of particle filters with the presented manifold regularized deep model being updated every few frames. Experimental results on different tracking datasets demonstrate that our tracker outperforms most existing tracking approaches. The source code and results are available at: https://github.com/shenjianbing/MRCNNTracking. Hongwei Hu, Bo Ma 0001, Jianbing Shen, Hanqiu Sun, Ling Shao 0001, Fatih Porikli |
IEEE Trans. Multim. | 3 |
| 2019 | Editorial: Booming of Neural Networks and Learning SystemsabstractAs you open this January issue of the IEEE Transactions on Neural Networks and Learning Systems (TNNLS), I hope everyone enjoyed a great holiday season and is excited for the new year of 2019. I am very delighted and honored to report several key metrics of IEEE TNNLS to the community. Akira Hirose 0001, Alessio Micheli, Artur S. d'Avila Garcez, Choon Ki Ahn, Gang Pan 0001, Hamid Reza Karimi, Jianbing Shen, José de Jesús Rubio, Lei Zhang 0005, Lingjia Liu 0001, Lorenzo Livi, Nishchal K. Verma, Pedro Antonio Gutiérrez, Qi Tian 0001, Qinglai Wei, Seiichi Ozawa, Stuart Harvey Rubin, Weineng Chen, Xi Li 0001, Xiaofeng Liao 0001, Youmin Zhang 0001, Zhen Ni, Haibo He |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2019 | Submodular Function Optimization for Motion Clustering and Image SegmentationabstractIn this paper, we propose a framework of maximizing quadratic submodular energy with a knapsack constraint approximately, to solve certain computer vision problems. The proposed submodular maximization problem can be viewed as a generalization of the classic 0/1 knapsack problem. Importantly, maximization of our knapsack constrained submodular energy function can be solved via dynamic programing. We further introduce a range-reduction step prior to dynamic programing as a two-stage procedure for more efficient maximization. In order to demonstrate the effectiveness of the proposed energy function and its maximization algorithm, we apply it to two representative computer vision tasks: image segmentation and motion trajectory clustering. Experimental results of image segmentation demonstrate that our method outperforms the classic segmentation algorithms of graph cuts and random walks. Moreover, our framework achieves better performance than state-of-the-art methods on the motion trajectory clustering task. Jianbing Shen, Xingping Dong, Jianteng Peng, Xiaogang Jin 0001, Ling Shao 0001, Fatih Porikli |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Hyperparameter Optimization for Tracking With Continuous Deep Q-LearningabstractHyperparameters are numerical presets whose values are assigned prior to the commencement of the learning process. Selecting appropriate hyperparameters is critical for the accuracy of tracking algorithms, yet it is difficult to determine their optimal values, in particular, adaptive ones for each specific video sequence. Most hyperparameter optimization algorithms depend on searching a generic range and they are imposed blindly on all sequences. Here, we propose a novel hyperparameter optimization method that can find optimal hyperparameters for a given sequence using an action-prediction network leveraged on Continuous Deep Q-Learning. Since the common state-spaces for object tracking tasks are significantly more complex than the ones in traditional control problems, existing Continuous Deep Q-Learning algorithms cannot be directly applied. To overcome this challenge, we introduce an efficient heuristic to accelerate the convergence behavior. We evaluate our method on several tracking benchmarks and demonstrate its superior performance1. Xingping Dong, Jianbing Shen, Wenguan Wang, Yu Liu 0074, Ling Shao 0001, Fatih Porikli |
CVPR | 2 |
| 2018 | Salient Object Detection Driven by Fixation PredictionabstractResearch in visual saliency has been focused on two major types of models namely fixation prediction and salient object detection. The relationship between the two, however, has been less explored. In this paper, we propose to employ the former model type to identify and segment salient objects in scenes. We build a novel neural network called Attentive Saliency Network (ASNet) that learns to detect salient objects from fixation maps. The fixation map, derived at the upper network layers, captures a high-level understanding of the scene. Salient object detection is then viewed as fine-grained object-level saliency segmentation and is progressively optimized with the guidance of the fixation map in a top-down manner. ASNet is based on a hierarchy of convolutional LSTMs (convLSTMs) that offers an efficient recurrent mechanism for sequential refinement of the segmentation map. Several loss functions are introduced for boosting the performance of the ASNet. Extensive experimental evaluation shows that our proposed ASNet is capable of generating accurate segmentation maps with the help of the computed fixation map. Our work offers a deeper insight into the mechanisms of attention and narrows the gap between salient object detection and fixation prediction. Wenguan Wang, Jianbing Shen, Xingping Dong, Ali Borji |
CVPR | 2 |
| 2018 | Revisiting Video Saliency: A Large-Scale Benchmark and a New ModelabstractIn this work, we contribute to video saliency research in two ways. First, we introduce a new benchmark for predicting human eye movements during dynamic scene free-viewing, which is long-time urged in this field. Our dataset, named DHF1K (Dynamic Human Fixation), consists of 1K high-quality, elaborately selected video sequences spanning a large range of scenes, motions, object types and background complexity. Existing video saliency datasets lack variety and generality of common dynamic scenes and fall short in covering challenging situations in unconstrained environments. In contrast, DHF1K makes a significant leap in terms of scalability, diversity and difficulty, and is expected to boost video saliency modeling. Second, we propose a novel video saliency model that augments the CNN-LSTM network architecture with an attention mechanism to enable fast, end-to-end saliency learning. The attention mechanism explicitly encodes static saliency information, thus allowing LSTM to focus on learning more flexible temporal saliency representation across successive frames. Such a design fully leverages existing large-scale static fixation datasets, avoids overfitting, and significantly improves training efficiency and testing performance. We thoroughly examine the performance of our model, with respect to state-of-the-art saliency models, on three large-scale datasets (i.e., DHF1K, Hollywood2, UCF sports). Experimental results over more than 1.2K testing videos containing 400K frames demonstrate that our model outperforms other competitors. Wenguan Wang, Jianbing Shen, Ming-Ming Cheng, Ali Borji |
CVPR | 2 |
| 2018 | Attentive Fashion Grammar Network for Fashion Landmark Detection and Clothing Category ClassificationabstractThis paper proposes a knowledge-guided fashion network to solve the problem of visual fashion analysis, e.g., fashion landmark localization and clothing category classification. The suggested fashion model is leveraged with high-level human knowledge in this domain. We propose two important fashion grammars: (i) dependency grammar capturing kinematics-like relation, and (ii) symmetry grammar accounting for the bilateral symmetry of clothes. We introduce Bidirectional Convolutional Recurrent Neural Networks (BCRNNs) for efficiently approaching message passing over grammar topologies, and producing regularized landmark layouts. For enhancing clothing category classification, our fashion network is encoded with two novel attention mechanisms, i.e., landmark-aware attention and category-driven attention. The former enforces our network to focus on the functional parts of clothes, and learns domain-knowledge centered representations, leading to a supervised attention mechanism. The latter is goal-driven, which directly enhances task-related features and can be learned in an implicit, top-down manner. Experimental results on large-scale fashion datasets demonstrate the superior performance of our fashion grammar network. Wenguan Wang, Yuanlu Xu, Jianbing Shen, Song-Chun Zhu |
CVPR | 3 |
| 2018 | Triplet Loss in Siamese Network for Object Tracking
Xingping Dong, Jianbing Shen |
ECCV (13) | 2 |
| 2018 | Learning Human-Object Interactions by Graph Parsing Neural Networks
Siyuan Qi, Wenguan Wang, Baoxiong Jia, Jianbing Shen, Song-Chun Zhu |
ECCV (9) | 4 |
| 2018 | Pyramid Dilated Deeper ConvLSTM for Video Salient Object Detection
Hongmei Song, Wenguan Wang, Sanyuan Zhao, Jianbing Shen, Kin-Man Lam 0001 |
ECCV (11) | 4 |
| 2018 | Facial landmark detection by semi-supervised deep learning
Jianbing Shen, Tianyuan Du |
Neurocomputing | 3 |
| 2018 | Parallel and efficient approximate nearest patch matching for image editing applications
Hanli Zhao, Heyang Guo, Xiaogang Jin 0001, Jianbing Shen, Xiaoyang Mao, Junru Liu |
Neurocomputing | 4 |
| 2018 | Scene text recognition using residual convolutional recurrent neural network
Zhengchao Lei, Sanyuan Zhao, Hongmei Song, Jianbing Shen |
Mach. Vis. Appl. | 4 |
| 2018 | Saliency-Aware Video Object SegmentationabstractVideo saliency, aiming for estimation of a single dominant object in a sequence, offers strong object-level cues for unsupervised video object segmentation. In this paper, we present a geodesic distance based technique that provides reliable and temporally consistent saliency measurement of superpixels as a prior for pixel-wise labeling. Using undirected intra-frame and inter-frame graphs constructed from spatiotemporal edges or appearance and motion, and a skeleton abstraction step to further enhance saliency estimates, our method formulates the pixel-wise segmentation task as an energy minimization problem on a function that consists of unary terms of global foreground and background models, dynamic location models, and pairwise terms of label smoothness potentials. We perform extensive quantitative and qualitative experiments on benchmark datasets. Our method achieves superior performance in comparison to the current state-of-the-art in terms of accuracy and speed. Wenguan Wang, Jianbing Shen, Ruigang Yang, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Robust Stereoscopic Crosstalk PredictionabstractWe propose a new metric to predict perceived crosstalk using the original images rather than both the original and ghosted images. The proposed metrics are based on color information. First, we extract a disparity map, a color difference map, and a color contrast map from original image pairs. Then, we use those maps to construct two new metrics (Vdispc and Vdlogc). Metric Vdispc considers the effect of the disparity map and the color difference map, while Vdlogc addresses the influence of the color contrast map. The prediction performance is evaluated using various types of stereoscopic crosstalk images. By incorporating Vdispc and Vdlogc, the new metric Vpdlc is proposed to achieve a higher correlation with the perceived subject crosstalk scores. Experimental results show that the new metrics achieve better performance than previous methods, which indicate that color information is one key factor for crosstalk visible prediction. Furthermore, we construct a new data set to evaluate our new metrics. Jianbing Shen, Yan Zhang 0094, Zhiyuan Liang, Chang Liu 0071, Hanqiu Sun, Xiaopeng Hao, Jianhong Liu, Jian Yang 0009, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Video Co-Saliency Guided Co-SegmentationabstractWe introduce the term video co-saliency to denote the task of extracting the common noticeable, or salient, regions from multiple relevant videos. The proposed video co-saliency approach accounts for both inter-video foreground correspondences and intra-video saliency stimuli to emphasize the salient foreground regions of video frames and, at the same time, disregard irrelevant visual information of the background. Compared with image co-saliency, it is more reliable due to the utilization of temporal information of video sequence. Benefiting from the discriminability of video co-saliency, we present a unified framework for segmenting out the common salient regions of relevant videos, guided by video co-saliency prior. Unlike naive video co-segmentation approaches employing simple color differences and local motion features, the presented video co-saliency provides a more powerful indicator for the common salient regions, thus conducting video co-segmentation efficiently. Extensive experiments show that the proposed method successfully infers video co-saliency and extracts the common salient regions, outperforming the state-of-the-art methods. Wenguan Wang, Jianbing Shen, Hanqiu Sun, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Video Saliency Detection Using Object ProposalsabstractIn this paper, we introduce a novel approach to identify salient object regions in videos via object proposals. The core idea is to solve the saliency detection problem by ranking and selecting the salient proposals based on object-level saliency cues. Object proposals offer a more complete and high-level representation, which naturally caters to the needs of salient object detection. As well as introducing this novel solution for video salient object detection, we reorganize various discriminative saliency cues and traditional saliency assumptions on object proposals. With object candidates, a proposal ranking and voting scheme, based on various object-level saliency cues, is designed to screen out nonsalient parts, select salient object regions, and to infer an initial saliency estimate. Then a saliency optimization process that considers temporal consistency and appearance differences between salient and nonsalient regions is used to refine the initial saliency estimates. Our experiments on public datasets (SegTrackV2, Freiburg-Berkeley Motion Segmentation Dataset, and Densely Annotated Video Segmentation) validate the effectiveness, and the proposed method produces significant improvements over state-of-the-art algorithms. Wenguan Wang, Jianbing Shen, Ling Shao 0001, Jian Yang 0009, Dacheng Tao, Yuan Yan Tang |
IEEE Trans. Cybern. | 3 |
| 2018 | Submodular Trajectories for Better Motion Segmentation in VideosabstractWe propose a new trajectory clustering method using submodular optimization for better motion segmentation in videos. A small number of representative trajectories are first selected by submodular maximization automatically. Then all the initial trajectories can be segmented into fragments with the representative trajectories as centers of fragments. At last, fragments are merged into clusters by a two-stage bottom-up clustering method, and each cluster shows the motion of one moving object. The submodular energy function integrates the quality of all trajectories and their correlations. As a result, thousands of initial trajectories are replaced by only dozens of representative trajectories, which will reduce the negative influence of inaccurate initial trajectories on motion segmentation. The representative trajectories will have larger weights while extracting color or texture information of each moving entity at the step of motion segmentation. Experimental results demonstrate that our method can divide trajectories into more accurate clusters. The final motion segmentation results also illustrate that our method outperforms state-of-the-art motion segmentation methods based on trajectory clustering. Jianbing Shen, Jianteng Peng, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Deep Visual Attention PredictionabstractIn this paper, we aim to predict human eye fixation with view-free scenes based on an end-to-end deep learning architecture. Although convolutional neural networks (CNNs) have made substantial improvement on human attention prediction, it is still needed to improve the CNN-based attention models by efficiently leveraging multi-scale features. Our visual attention network is proposed to capture hierarchical saliency information from deep, coarse layers with global saliency information to shallow, fine layers with local saliency response. Our model is based on a skip-layer network structure, which predicts human attention from multiple convolutional layers with various reception fields. Final saliency prediction is achieved via the cooperation of those global and local predictions. Our model is learned in a deep supervision manner, where supervision is directly fed into multi-level layers, instead of previous approaches of providing supervision only at the output layer and propagating this supervision back to earlier layers. Our model thus incorporates multi-level saliency predictions within a single network, which significantly decreases the redundancy of previous approaches of learning multiple network streams with different input scales. Extensive experimental analysis on various challenging benchmark data sets demonstrate our method yields the state-of-the-art performance with competitive inference time. Wenguan Wang, Jianbing Shen |
IEEE Trans. Image Process. | 2 |
| 2018 | Video Salient Object Detection via Fully Convolutional NetworksabstractThis paper proposes a deep learning model to efficiently detect salient regions in videos. It addresses two important issues: 1) deep video saliency model training with the absence of sufficiently large and pixel-wise annotated video data and 2) fast video saliency training and detection. The proposed deep video saliency network consists of two modules, for capturing the spatial and temporal saliency information, respectively. The dynamic saliency model, explicitly incorporating saliency estimates from the static saliency model, directly produces spatiotemporal saliency inference without time-consuming optical flow computation. We further propose a novel data augmentation technique that simulates video training data from existing annotated image data sets, which enables our network to learn diverse saliency information and prevents overfitting with the limited number of training videos. Leveraging our synthetic video data (150K video sequences) and real videos, our deep video saliency model successfully learns both spatial and temporal saliency cues, thus producing accurate spatiotemporal saliency estimate. We advance the state-of-the-art on the densely annotated video segmentation data set (MAE of .06) and the Freiburg-Berkeley Motion Segmentation data set (MAE of .07), and do so with much improved speed (2 fps with all steps).This paper proposes a deep learning model to efficiently detect salient regions in videos. It addresses two important issues: 1) deep video saliency model training with the absence of sufficiently large and pixel-wise annotated video data and 2) fast video saliency training and detection. The proposed deep video saliency network consists of two modules, for capturing the spatial and temporal saliency information, respectively. The dynamic saliency model, explicitly incorporating saliency estimates from the static saliency model, directly produces spatiotemporal saliency inference without time-consuming optical flow computation. We further propose a novel data augmentation technique that simulates video training data from existing annotated image data sets, which enables our network to learn diverse saliency information and prevents overfitting with the limited number of training videos. Leveraging our synthetic video data (150K video sequences) and real videos, our deep video saliency model successfully learns both spatial and temporal saliency cues, thus producing accurate spatiotemporal saliency estimate. We advance the state-of-the-art on the densely annotated video segmentation data set (MAE of .06) and the Freiburg-Berkeley Motion Segmentation data set (MAE of .07), and do so with much improved speed (2 fps with all steps). Wenguan Wang, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Fast Online Tracking With Detection RefinementabstractMost of the existing multiple object tracking (MOT) methods employ the tracking-by-detection framework. Among them, the min-cost network flow optimization techniques become the most popular and standard ones. In these methods, the graph structure models the MOT problem and finds the optimal flow in a connected graph of detections to encode the accurate track trajectories. However, the existing network flow is not suitable for directly online tracking, where the tracking results depend too much on the initial detections. To solve these problems, we present a fast online MOT algorithm by introducing the minimum output sum of squared error filter. The proposed method can adaptively refine the tracking targets according to the proposed rules of correcting the detection mistakes. Furthermore, we introduce an alternative targets hypotheses to reduce the dependence on detections and adaptively refine the object detection boxes. The experimental results on the MOT 2015 benchmark demonstrate that our method achieves comparable or even better results than previous approaches. Jianbing Shen, Dajiang Yu, Leyao Deng, Xingping Dong |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2018 | Manifold Regularized Correlation Object TrackingabstractIn this paper, we propose a manifold regularized correlation tracking method with augmented samples. To make better use of the unlabeled data and the manifold structure of the sample space, a manifold regularization-based correlation filter is introduced, which aims to assign similar labels to neighbor samples. Meanwhile, the regression model is learned by exploiting the block-circulant structure of matrices resulting from the augmented translated samples over multiple base samples cropped from both target and nontarget regions. Thus, the final classifier in our method is trained with positive, negative, and unlabeled base samples, which is a semisupervised learning framework. A block optimization strategy is further introduced to learn a manifold regularization-based correlation filter for efficient online tracking. Experiments on two public tracking data sets demonstrate the superior performance of our tracker compared with the state-of-the-art tracking approaches. Hongwei Hu, Bo Ma 0001, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Robust Object Tracking by Nonlinear LearningabstractWe propose a method that obtains a discriminative visual dictionary and a nonlinear classifier for visual tracking tasks in a sparse coding manner based on the globally linear approximation for a nonlinear learning theory. Traditional discriminative tracking methods based on sparse representation learn a dictionary in an unsupervised way and then train a classifier, which may not generate both descriptive and discriminative models for targets by treating dictionary learning and classifier learning separately. In contrast, the proposed tracking approach can construct a dictionary that fully reflects the intrinsic manifold structure of visual data and introduces more discriminative ability in a unified learning framework. Finally, an iterative optimization approach, which computes the optimal dictionary, the associated sparse coding, and a classifier, is introduced. Experiments on two benchmarks show that our tracker achieves a better performance compared with some popular tracking algorithms. Bo Ma 0001, Hongwei Hu, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Deep Cropping via Attention Box Prediction and Aesthetics AssessmentabstractWe model the photo cropping problem as a cascade of attention box regression and aesthetic quality classification, based on deep learning. A neural network is designed that has two branches for predicting attention bounding box and analyzing aesthetics, respectively. The predicted attention box is treated as an initial crop window where a set of cropping candidates are generated around it, without missing important information. Then, aesthetics assessment is employed to select the final crop as the one with the best aesthetic quality. With our network, cropping candidates share features within full-image convolutional feature maps, thus avoiding repeated feature computation and leading to higher computation efficiency. Via leveraging rich data for attention prediction and aesthetics assessment, the proposed method produces high-quality cropping results, even with the limited availability of training data for photo cropping. The experimental results demonstrate the competitive results and fast processing speed (5 fps with all steps). Wenguan Wang, Jianbing Shen |
ICCV | 2 |
| 2017 | Super-Trajectory for Video SegmentationabstractWe introduce a novel semi-supervised video segmentation approach based on an efficient video representation, called as “super-trajectory”. Each super-trajectory corresponds to a group of compact trajectories that exhibit consistent motion patterns, similar appearance and close spatiotemporal relationships. We generate trajectories using a probabilistic model, which handles occlusions and drifts in a robust and natural way. To reliably group trajectories, we adopt a modified version of the density peaks based clustering algorithm that allows capturing rich spatiotemporal relations among trajectories in the clustering process. The presented video representation is discriminative enough to accurately propagate the initial annotations in the first frame onto the remaining video frames. Extensive experimental analysis on challenging benchmarks demonstrate our method is capable of distinguishing the target objects from complex backgrounds and even reidentifying them after occlusions. Wenguan Wang, Jianbing Shen, Jianwen Xie, Fatih Porikli |
ICCV | 2 |
| 2017 | Diffusion-based saliency detection with optimal seed selection scheme
Sanyuan Zhao, Zhengchao Lei, Meiling Sun, Jianbing Shen |
Neurocomputing | 5 |
| 2017 | A hierarchical visual saliency detection method by combining distinction and background probability maps
Sanyuan Zhao, Jianbing Shen, Fengxia Li |
Multim. Syst. | 2 |
| 2017 | Hierarchical Superpixel-to-Pixel Dense MatchingabstractIn this paper, we propose a novel matching method to establish dense correspondences automatically between two images in a hierarchical superpixel-to-pixel manner. Our method first estimates dense superpixel pairings between the two images in the coarse-grained level to overcome large patch displacements and then utilizes superpixel level pairings to drive the matchings in the pixel level to obtain fine texture details. In order to compensate for the influence of color and illumination variations, we apply a regularization technique to rectify images by a color transfer function. Experimental validation on benchmark data sets demonstrates that our approach achieves better visual quality outperforming the state-of-the-art dense matching algorithms. Xingping Dong, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Single-Image Distance Measurement by a Smart Mobile DeviceabstractExisting distance measurement methods either require multiple images and special photographing poses or only measure the height with a special view configuration. We propose a novel image-based method that can measure various types of distance from single image captured by a smart mobile device. The embedded accelerometer is used to determine the view orientation of the device. Consequently, pixels can be back-projected to the ground, thanks to the efficient calibration method using two known distances. Then the distance in pixel is transformed to a real distance in centimeter with a linear model parameterized by the magnification ratio. Various types of distance specified in the image can be computed accordingly. Experimental results demonstrate the effectiveness of the proposed method. Shang-Wen Chen, Xianyong Fang, Jianbing Shen, Linbo Wang 0001, Ling Shao 0001 |
IEEE Trans. Cybern. | 3 |
| 2017 | Visual Tracking by Sampling in Part SpaceabstractIn this paper, we present a novel part-based visual tracking method from the perspective of probability sampling. Specifically, we represent the target by a part space with two online learned probabilities to capture the structure of the target. The proposal distribution memorizes the historical performance of different parts, and it is used for the first round of part selection. The acceptance probability validates the specific tracking stability of each part in a frame, and it determines whether to accept its vote or to reject it. By doing this, we transform the complex online part selection problem into a probability learning one, which is easier to tackle. The observation model of each part is constructed by an improved supervised descent method and is learned in an incremental manner. Experimental results on two benchmarks demonstrate the competitive performance of our tracker against state-of-the-art methods. Lianghua Huang, Bo Ma 0001, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
IEEE Trans. Image Process. | 3 |
| 2017 | Higher Order Energies for Image SegmentationabstractA novel energy minimization method for general higher order binary energy functions is proposed in this paper. We first relax a discrete higher order function to a continuous one, and use the Taylor expansion to obtain an approximate lower order function, which is optimized by the quadratic pseudo-Boolean optimization or other discrete optimizers. The minimum solution of this lower order function is then used as a new local point, where we expand the original higher order energy function again. Our algorithm does not restrict to any specific form of the higher order binary function or bring in extra auxiliary variables. For concreteness, we show an application of segmentation with the appearance entropy, which is efficiently solved by our method. Experimental results demonstrate that our method outperforms the state-of-the-art methods. Jianbing Shen, Jianteng Peng, Xingping Dong, Ling Shao 0001, Fatih Porikli |
IEEE Trans. Image Process. | 1 |
| 2017 | Selective Video Object CutoutabstractConventional video segmentation approaches rely heavily on appearance models. Such methods often use appearance descriptors that have limited discriminative power under complex scenarios. To improve the segmentation performance, this paper presents a pyramid histogram-based confidence map that incorporates structure information into appearance statistics. It also combines geodesic distance-based dynamic models. Then, it employs an efficient measure of uncertainty propagation using local classifiers to determine the image regions, where the object labels might be ambiguous. The final foreground cutout is obtained by refining on the uncertain regions. Additionally, to reduce manual labeling, our method determines the frames to be labeled by the human operator in a principled manner, which further boosts the segmentation performance and minimizes the labeling effort. Our extensive experimental analyses on two big benchmarks demonstrate that our solution achieves superior performance, favorable computational efficiency, and reduced manual labeling in comparison to the state of the art. Wenguan Wang, Jianbing Shen, Fatih Porikli |
IEEE Trans. Image Process. | 2 |
| 2017 | Occlusion-Aware Real-Time Object TrackingabstractThe online learning methods are popular for visual tracking because of their robust performance for most video sequences. However, the drifting problem caused by noisy updates is still a challenge for most highly adaptive online classifiers. In visual tracking, target object appearance variation, such as deformation and long-term occlusion, easily causes noisy updates. To overcome this problem, a new real-time occlusion-aware visual tracking algorithm is introduced. First, we learn a novel two-stage classifier with circulant structure with kernel, named integrated circulant structure kernels (ICSK). The first stage is applied for transition estimation and the second is used for scale estimation. The circulant structure makes our algorithm realize fast learning and detection. Then, the ICSK is used to detect the target without occlusion and build a classifier pool to save these classifiers with noisy updates. When the target is in heavy occlusion or after long-term occlusion, we redetect it using an optimal classifier selected from the classifier-pool according to an entropy minimization criterion. Extensive experimental results on the full benchmark demonstrate our real-time algorithm achieves better performance than state-of-the-art methods. Xingping Dong, Jianbing Shen, Dajiang Yu, Wenguan Wang, Jianhong Liu |
IEEE Trans. Multim. | 2 |
| 2017 | Stereoscopic Thumbnail Creation via Efficient Stereo Saliency DetectionabstractIn this paper, we propose a framework for automatically producing thumbnails from stereo image pairs. It has two components focusing respectively on stereo saliency detection and stereo thumbnail generation. The first component analyzes stereo saliency through various saliency stimuli, stereoscopic perception and the relevance between two stereo views. The second component uses stereo saliency to guide stereo thumbnail generation. We develop two types of thumbnail generation methods, both changing image size automatically. The first method is called content-persistent cropping (CPC), which aims at cropping stereo images for display devices with different aspect ratios while preserving as much content as possible. The second method is an object-aware cropping method (OAC) for generating the smallest possible thumbnail pair that retains the most important content only and facilitates quick visual exploration of a stereo image database. Quantitative and qualitative experimental evaluations demonstrate promising performance of our thumbnail generation methods in comparison to state-of-the-art algorithms. Wenguan Wang, Jianbing Shen, Yizhou Yu, Kwan-Liu Ma |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2016 | Robust geometric ℓp-norm feature pooling for image classification and action recognition
Teng Li 0001, Bingbing Ni, Jianbing Shen, Meng Wang 0001 |
Image Vis. Comput. | 4 |
| 2016 | Video Supervoxels Using Partially Absorbing Random WalksabstractSupervoxels have been widely used as a preprocessing step to exploit object boundaries to improve the performance of video processing tasks. However, most of the traditional supervoxel algorithms do not perform well in regions with complex textures or weak boundaries. These methods may generate supervoxels with overlapping boundaries. In this paper, we present the novel video supervoxel generation algorithm using partially absorbing random walks to get more accurate supervoxels in these regions. Our spatial-temporal framework is introduced by making full use of the appearance and motion cues, which effectively exploits the temporal consistency in video sequence. Moreover, we build a novel Laplacian optimization structure using two adjacent frames to make our approach more efficient. Experimental results demonstrated that our method achieved better performance than the state-of-the-art supervoxel algorithms. Yuling Liang, Jianbing Shen, Xingping Dong, Hanqiu Sun, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Superpixel Optimization Using Higher Order EnergyabstractA novel superpixel extraction algorithm using a higher order energy optimization framework is proposed in this paper. We first adopt the k-means clustering technique to quickly get an initial superpixel result. Then a higher order energy function is employed to optimize and refine these initial superpixels. We use a more general higher order energy function that includes a first-order data term, a second-order smoothness term, and a higher order term. The presegments are employed to provide the prior information of sufficient edges and segment regions for our higher order energy term. According to the texture measurement in different local regions, our algorithm adaptively computes the proper ratios of different energy terms to obtain a better superpixel performance. The experimental results demonstrate that our method using the higher order energy generates better results with well-aligned boundaries and homogeneous effects than the existing superpixel algorithms. Jianteng Peng, Jianbing Shen, Angela Yao, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Discriminative Tracking Using Tensor PoolingabstractHow to effectively organize local descriptors to build a global representation has a critical impact on the performance of vision tasks. Recently, local sparse representation has been successfully applied to visual tracking, owing to its discriminative nature and robustness against local noise and partial occlusions. Local sparse codes computed with a template actually form a three-order tensor according to their original layout, although most existing pooling operators convert the codes to a vector by concatenating or computing statistics on them. We argue that, compared to pooling vectors, the tensor form could deliver more intrinsic structural information for the target appearance, and can also avoid high dimensionality learning problems suffered in concatenation-based pooling methods. Therefore, in this paper, we propose to represent target templates and candidates directly with sparse coding tensors, and build the appearance model by incrementally learning on these tensors. We propose a discriminative framework to further improve robustness of our method against drifting and environmental noise. Experiments on a recent comprehensive benchmark indicate that our method performs better than state-of-the-art trackers. Bo Ma 0001, Lianghua Huang, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Cybern. | 3 |
| 2016 | High-Order Energies for Stereo SegmentationabstractIn this paper, we propose a novel segmentation approach for stereo images using the high-order energy optimization, which utilizes the disparity maps and statistical information of stereo images to enrich the high-order potential functions. To the best of our knowledge, our approach is the first one to formulate the problem of stereo segmentation as a high-order energy optimization problem, which simultaneously segments the foreground objects in left and right images using the proposed high-order potential function. A new method for designing the penalty function in our high-order term is proposed by the corresponding pixels and their neighboring pixels between left and right images. The relationships of stereo correspondence by disparity maps are further employed to enhance the connections between the left and right stereo images. Experimental results demonstrate that the proposed approach can effectively improve the performance of two kinds of stereo segmentation, including the automatic saliency-aware stereocut and the interactive stereo segmentation with user scribbles. Jianteng Peng, Jianbing Shen, Xuelong Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2016 | Sub-Markov Random Walk for Image SegmentationabstractA novel sub-Markov random walk (subRW) algorithm with label prior is proposed for seeded image segmentation, which can be interpreted as a traditional random walker on a graph with added auxiliary nodes. Under this explanation, we unify the proposed subRW and other popular random walk (RW) algorithms. This unifying view will make it possible for transferring intrinsic findings between different RW algorithms, and offer new ideas for designing novel RW algorithms by adding or changing auxiliary nodes. To verify the second benefit, we design a new subRW algorithm with label prior to solve the segmentation problem of objects with thin and elongated parts. The experimental results on both synthetic and natural images with twigs demonstrate that the proposed subRW method outperforms previous RW algorithms for seeded image segmentation. Xingping Dong, Jianbing Shen, Ling Shao 0001, Luc Van Gool |
IEEE Trans. Image Process. | 2 |
| 2016 | Generalized Pooling for Robust Object TrackingabstractFeature pooling in a majority of sparse coding-based tracking algorithms computes final feature vectors only by low-order statistics or extreme responses of sparse codes. The high-order statistics and the correlations between responses to different dictionary items are neglected. We present a more generalized feature pooling method for visual tracking by utilizing the probabilistic function to model the statistical distribution of sparse codes. Since immediate matching between two distributions usually requires high computational costs, we introduce the Fisher vector to derive a more compact and discriminative representation for sparse codes of the visual target. We encode target patches by local coordinate coding, utilize Gaussian mixture model to compute Fisher vectors, and finally train semi-supervised linear kernel classifiers for visual tracking. In order to handle the drifting problem during the tracking process, these classifiers are updated online with current tracking results. The experimental results on two challenging tracking benchmarks demonstrate that the proposed approach achieves a better performance than the state-of-the-art tracking algorithms. Bo Ma 0001, Hongwei Hu, Jianbing Shen, Yangbiao Liu, Ling Shao 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | Visual Tracking Under Motion BlurabstractMost existing tracking algorithms do not explicitly consider the motion blur contained in video sequences, which degrades their performance in real-world applications where motion blur often occurs. In this paper, we propose to solve the motion blur problem in visual tracking in a unified framework. Specifically, a joint blur state estimation and multi-task reverse sparse learning framework are presented, where the closed-form solution of blur kernel and sparse code matrix is obtained simultaneously. The reverse process considers the blurry candidates as dictionary elements, and sparsely represents blurred templates with the candidates. By utilizing the information contained in the sparse code matrix, an efficient likelihood model is further developed, which quickly excludes irrelevant candidates and narrows the particle scale down. Experimental results on the challenging benchmarks show that our method performs well against the state-of-the-art trackers. Bo Ma 0001, Lianghua Huang, Jianbing Shen, Ling Shao 0001, Ming-Hsuan Yang 0001, Fatih Porikli |
IEEE Trans. Image Process. | 3 |
| 2016 | Real-Time Superpixel Segmentation by DBSCAN Clustering AlgorithmabstractIn this paper, we propose a real-time image superpixel segmentation method with 50 frames/s by using the density-based spatial clustering of applications with noise (DBSCAN) algorithm. In order to decrease the computational costs of superpixel algorithms, we adopt a fast two-step framework. In the first clustering stage, the DBSCAN algorithm with color-similarity and geometric restrictions is used to rapidly cluster the pixels, and then, small clusters are merged into superpixels by their neighborhood through a distance measurement defined by color and spatial features in the second merging stage. A robust and simple distance function is defined for obtaining better superpixels in these two steps. The experimental results demonstrate that our real-time superpixel algorithm (50 frames/s) by the DBSCAN clustering outperforms the state-of-the-art superpixel segmentation methods in terms of both accuracy and efficiency. Jianbing Shen, Xiaopeng Hao, Zhiyuan Liang, Yu Liu 0074, Wenguan Wang, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Correspondence Driven Saliency TransferabstractIn this paper, we show that large annotated data sets have great potential to provide strong priors for saliency estimation rather than merely serving for benchmark evaluations. To this end, we present a novel image saliency detection method called saliency transfer. Given an input image, we first retrieve a support set of best matches from the large database of saliency annotated images. Then, we assign the transitional saliency scores by warping the support set annotations onto the input image according to computed dense correspondences. To incorporate context, we employ two complementary correspondence strategies: a global matching scheme based on scene-level analysis and a local matching scheme based on patch-level inference. We then introduce two refinement measures to further refine the saliency maps and apply the random-walk-with-restart by exploring the global saliency structure to estimate the affinity between foreground and background assignments. Extensive experimental results on four publicly available benchmark data sets demonstrate that the proposed saliency algorithm consistently outperforms the current state-of-the-art methods. Wenguan Wang, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
IEEE Trans. Image Process. | 2 |
| 2016 | Higher-Order Image Co-segmentationabstractA novel interactive image cosegmentation algorithm using likelihood estimation and higher order energy optimization is proposed for extracting common foreground objects from a group of related images. Our approach introduces the higher order clique's, energy into the cosegmentation optimization process successfully. A region-based likelihood estimation procedure is first performed to provide the prior knowledge for our higher order energy function. Then, a new cosegmentation energy function using higher order cliques is developed, which can efficiently cosegment the foreground objects with large appearance variations from a group of images in complex scenes. Both the quantitative and qualitative experimental results on representative datasets demonstrate that the accuracy of our cosegmentation results is much higher than the state-of-the-art cosegmentation methods. Wenguan Wang, Jianbing Shen |
IEEE Trans. Multim. | 2 |
| 2015 | Saliency-aware geodesic video object segmentationabstractWe introduce an unsupervised, geodesic distance based, salient video object segmentation method. Unlike traditional methods, our method incorporates saliency as prior for object via the computation of robust geodesic measurement. We consider two discriminative visual features: spatial edges and temporal motion boundaries as indicators of foreground object locations. We first generate framewise spatiotemporal saliency maps using geodesic distance from these indicators. Building on the observation that foreground areas are surrounded by the regions with high spatiotemporal edge values, geodesic distance provides an initial estimation for foreground and background. Then, high-quality saliency results are produced via the geodesic distances to background regions in the subsequent frames. Through the resulting saliency maps, we build global appearance models for foreground and background. By imposing motion continuity, we establish a dynamic location model for each frame. Finally, the spatiotemporal saliency maps, appearance models and dynamic location models are combined into an energy minimization framework to attain both spatially and temporally coherent object segmentation. Extensive quantitative and qualitative experiments on benchmark video dataset demonstrate the superiority of the proposed method over the state-of-the-art algorithms. Wenguan Wang, Jianbing Shen, Fatih Porikli |
CVPR | 2 |
| 2015 | Linearization to Nonlinear Learning for Visual TrackingabstractDue to unavoidable appearance variations caused by occlusion, deformation, and other factors, classifiers for visual tracking are nonlinear as a necessity. Building on the theory of globally linear approximations to nonlinear functions, we introduce an elegant method that jointly learns a nonlinear classifier and a visual dictionary for tracking objects in a semi-supervised sparse coding fashion. This establishes an obvious distinction from conventional sparse coding based discriminative tracking algorithms that usually maintain two-stage learning strategies, i.e., learning a dictionary in an unsupervised way then followed by training a classifier. However, the treating dictionary learning and classifier training as separate stages may not produce both descriptive and discriminative models for objects. By contrast, our method is capable of constructing a dictionary that not only fully reflects the intrinsic manifold structure of the data, but also possesses discriminative power. This paper presents an optimization method to obtain such an optimal dictionary, associated sparse coding, and a classifier in an iterative process. Our experiments on a benchmark show our tracker attains outstanding performance compared with the state-of-the-art algorithms. Bo Ma 0001, Hongwei Hu, Jianbing Shen, Fatih Porikli |
ICCV | 3 |
| 2015 | Accurate Normal and Reflectance Recovery Using Energy OptimizationabstractIn this paper, we propose a novel energy optimization framework to accurately estimate surface normal and reflectance of an object from an input image sequence. Input images are captured from a fixed viewpoint under varying lighting conditions. In the proposed approach we combine photometric stereo and Retinex constraints into our energy function. To formulate inter-image constraints, shading information is added to the Lambertian model to account for shadows. For intra-image constraints, we moderate the strength of shading smoothness according to shadow mask and normal variations. By minimizing this energy function we are able to recover accurate surface normals and reflectance. Experimental results show that our approach yields more realistic normal map and accurate albedo map than the state-of-the-art uncalibrated photometric stereo algorithms. As for intrinsic image decomposition, results on the real and synthetic scenes show that the proposed approach outperforms previous ones. Tao Luo 0001, Jianbing Shen, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Robust Match Fusion Using OptimizationabstractIn this paper, we present a novel patch-based match and fusion algorithm by taking account of moving scene in a multiple exposure image sequence using optimization. A uniform iterative approach is developed to match and find the corresponding patches in different exposure images, which are then fused in each iteration. Our approach does not need to align the input multiple exposure images before the fusion process. Considering that the pixel values are affected by various exposure time, we design a new patch-based energy function that will be optimized to improve the matching accuracy. An efficient patch-based exposure fusion approach using the random walker algorithm is developed to preserve the moving objects from the input multiple exposure images. To the best of our knowledge, our algorithm is the first patch-based exposure fusion work to preserve the moving objects of dynamic scenes that does not need the registration process of different exposure images. Experimental results of moving scenes demonstrate that our algorithm achieves visually pleasing fusion results without ghosting artifacts, while the results produced by the state-of-the-art exposure fusion and tone mapping algorithms exhibit different levels of ghosting artifacts. Xiameng Qin, Jianbing Shen, Xiaoyang Mao, Xuelong Li 0001, Yunde Jia |
IEEE Trans. Cybern. | 2 |
| 2015 | Interactive Cosegmentation Using Global and Local Energy OptimizationabstractWe propose a novel interactive cosegmentation method using global and local energy optimization. The global energy includes two terms: 1) the global scribbled energy and 2) the interimage energy. The first one utilizes the user scribbles to build the Gaussian mixture model and improve the cosegmentation performance. The second one is a global constraint, which attempts to match the histograms of common objects. To minimize the local energy, we apply the spline regression to learn the smoothness in a local neighborhood. This energy optimization can be converted into a constrained quadratic programming problem. To reduce the computational complexity, we propose an iterative optimization algorithm to decompose this optimization problem into several subproblems. The experimental results show that our method outperforms the state-of-the-art unsupervised cosegmentation and interactive cosegmentation methods on the iCoseg and MSRC benchmark data sets. Xingping Dong, Jianbing Shen, Ling Shao 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Robust Video Object CosegmentationabstractWith ever-increasing volumes of video data, automatic extraction of salient object regions became even more significant for visual analytic solutions. This surge has also opened up opportunities for taking advantage of collective cues encapsulated in multiple videos in a cooperative manner. However, it also brings up major challenges, such as handling of drastic appearance, motion pattern, and pose variations, of foreground objects as well as indiscriminate backgrounds. Here, we present a cosegmentation framework to discover and segment out common object regions across multiple frames and multiple videos in a joint fashion. We incorporate three types of cues, i.e., intraframe saliency, interframe consistency, and across-video similarity into an energy optimization framework that does not make restrictive assumptions on foreground appearance and motion model, and does not require objects to be visible in all frames. We also introduce a spatio-temporal scale-invariant feature transform (SIFT) flow descriptor to integrate across-video correspondence from the conventional SIFT-flow into interframe motion flow from optical flow. This novel spatio-temporal SIFT flow generates reliable estimations of common foregrounds over the entire video data set. Experimental results show that our method outperforms the state-of-the-art on a new extensive data set (ViCoSeg). Wenguan Wang, Jianbing Shen, Xuelong Li 0001, Fatih Porikli |
IEEE Trans. Image Process. | 2 |
| 2015 | Consistent Video Saliency Using Local Gradient Flow Optimization and Global RefinementabstractWe present a novel spatiotemporal saliency detection method to estimate salient regions in videos based on the gradient flow field and energy optimization. The proposed gradient flow field incorporates two distinctive features: 1) intra-frame boundary information and 2) inter-frame motion information together for indicating the salient regions. Based on the effective utilization of both intra-frame and inter-frame information in the gradient flow field, our algorithm is robust enough to estimate the object and background in complex scenes with various motion patterns and appearances. Then, we introduce local as well as global contrast saliency measures using the foreground and background information estimated from the gradient flow field. These enhanced contrast saliency cues uniformly highlight an entire object. We further propose a new energy function to encourage the spatiotemporal consistency of the output saliency maps, which is seldom explored in previous video saliency methods. The experimental results show that the proposed algorithm outperforms state-of-the-art video saliency detection methods. Wenguan Wang, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Video Object Segmentation Via Dense TrajectoriesabstractIn this paper, we propose a novel approach to segment moving object in video by utilizing improved point trajectories . First, point trajectories are densely sampled from video and tracked through optical flow, which provides information of long-term temporal interactions among objects in the video sequence . Second, a novel affinity measurement method considering both global and local information of point trajectories is proposed to cluster trajectories into groups. Finally, we propose a new graph-based segmentation method which adopts both local and global motion information encoded by the tracked dense point trajectories. The proposed approach achieves good performance on trajectory clustering, and it also obtains accurate video object segmentation results on both the Moseg dataset and our new dataset containing more challenging videos. Lin Chen 0021, Jianbing Shen, Wenguan Wang, Bingbing Ni |
IEEE Trans. Multim. | 2 |
| 2015 | Visual Tracking Using Strong Classifier and Structural Local Sparse DescriptorsabstractSparse coding methods have achieved great success in visual tracking, and we present a strong classifier and structural local sparse descriptors for robust visual tracking. Since the summary features considering the sparse codes are sensitive to occlusion and other interfering factors, we extract local sparse descriptors from a fraction of all patches by performing a pooling operation. The collection of local sparse descriptors is combined into a boosting-based strong classifier for robust visual tracking using a discriminative appearance model. Furthermore, a structural reconstruction error based weight computation method is proposed to adjust the classification score of each candidate for more precise tracking results. To handle appearance changes during tracking, we present an occlusion-aware template update scheme. Comprehensive experimental comparisons with the state-of-the-art algorithms demonstrated the better performance of the proposed method. Bo Ma 0001, Jianbing Shen, Yangbiao Liu, Hongwei Hu, Ling Shao 0001, Xuelong Li 0001 |
IEEE Trans. Multim. | 2 |
| 2015 | Structured-Patch Optimization for Dense CorrespondenceabstractThis paper presents a new method to compute the dense correspondences between two images by using the energy optimization and the structured patches. In terms of the property of the sparse feature and the principle that nearest sub-scenes and neighbors are much more similar, we design a new energy optimization to guide the dense matching process and find the reliable correspondences. The sparse features are also employed to design a new structure to describe the patches. Both transformation and deformation with the structured patches are considered and incorporated into an energy optimization framework. Thus, our algorithm can match the objects robustly in complicated scenes. Finally, a local refinement technique is proposed to solve the perturbation of the matched patches. Experimental results demonstrate that our method outperforms the state-of-the-art matching algorithms. Xiameng Qin, Jianbing Shen, Xiaoyang Mao, Xuelong Li 0001, Yunde Jia |
IEEE Trans. Multim. | 2 |
| 2014 | Learning to detect stereo saliencyabstractThis paper develops a novel learning-based method for detecting stereo saliency in stereopair images. The disparity maps computed from stereopair images provide an additional depth cue for stereo saliency detection. To the best of our knowledge, our approach is the first one to simultaneously detect the stereo saliency of both left and right images using support vector machine (SVM). In our work, the disparity maps are used in two aspects. One is to improve the performance of saliency detection for monocular image. The other one is to maintain the consistency between the stereo matching and saliency maps. In order to meet the above requirements, we propose a new combinational saliency feature to train the stereo images with the labeled saliency ground truth, using support vector machine as the classifier. In the test stage, our approach generates the stereo saliency results according to the trained SVM model. Furthermore, a stereopair saliency dataset containing 400 pairs of images is created to perform the challenging experiments. The experimental results have demonstrated that our method achieves better performance than the state-of-the-art algorithms of single-image saliency detection. Jianbing Shen, Xuelong Li 0001 |
ICME | 2 |
| 2014 | A new sparse feature-based patch for dense correspondenceabstractThis paper presents a new method to compute the dense correspondences between two images by using the sparse feature-based patches in an energy optimization framework. Many transformation and deformation cues such as color, scale and rotation should be considered when we finding dense correspondences between images. However, most existing methods only consider part of these transformations, which will introduce the uncorrect correspondence results. In terms of the property of the sparse feature and the principle that nearest sub-scenes and neighbors are much more similar, we design a new energy optimization to guide the dense matching process. Both transformation and deformation are considered in our energy optimization framework since we design the feature-based patches. Thus, our algorithm can match the complicated scenes and objects robustly. At last, a local refinement technique is proposed to solve the perturbation of the matched patches. Experimental results demonstrate that our method outperforms the state-of-the-art algorithms. Xiameng Qin, Jianbing Shen, Xuelong Li 0001, Yunde Jia |
ICME | 2 |
| 2014 | Re-texturing by intrinsic video
Jianbing Shen, Lin Chen 0021, Hanqiu Sun, Xuelong Li 0001 |
Inf. Sci. | 1 |
| 2014 | Interactive Segmentation Using Constrained Laplacian OptimizationabstractWe present a novel interactive image segmentation approach with user scribbles using constrained Laplacian graph optimization. A novel energy framework is developed by adding the smoothing item in the cost function of Laplacian graph energy. To the best of our knowledge, our approach is the first to incorporate the normalized cuts and graph cuts algorithms into a unified energy optimization framework. The proposed approach is further accelerated by running the proposed optimization method on a band region when we segment the large images. Our acceleration strategy enables our approach to efficiently segment the large images, which yields about a 20-80 times speedup. The proposed approach is evaluated on both the publicly available data sets and our own data set with large images. The benefits of the proposed unified framework are also demonstrated both qualitatively and quantitatively. The experimental results show that our segmentation method achieves better performance of both boundary recall and error rate than the existing state-of-the-art approaches. Jianbing Shen, Yunfan Du, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Exposure Fusion Using Boosting Laplacian PyramidabstractThis paper proposes a new exposure fusion approach for producing a high quality image result from multiple exposure images. Based on the local weight and global weight by considering the exposure quality measurement between different exposure images, and the just noticeable distortion-based saliency weight, a novel hybrid exposure weight measurement is developed. This new hybrid weight is guided not only by a single image's exposure level but also by the relative exposure level between different exposure images. The core of the approach is our novel boosting Laplacian pyramid, which is based on the structure of boosting the detail and base signal, respectively, and the boosting process is guided by the proposed exposure weight. Our approach can effectively blend the multiple exposure images for static scenes while preserving both color appearance and texture structure. Our experimental results demonstrate that the proposed approach successfully produces visually pleasing exposure fusion images with better color appearance and more texture details than the existing exposure fusion techniques and tone mapping operators. Jianbing Shen, Ying Zhao 0009, Shuicheng Yan, Xuelong Li 0001 |
IEEE Trans. Cybern. | 1 |
| 2014 | Lazy Random Walks for Superpixel SegmentationabstractWe present a novel image superpixel segmentation approach using the proposed lazy random walk (LRW) algorithm in this paper. Our method begins with initializing the seed positions and runs the LRW algorithm on the input image to obtain the probabilities of each pixel. Then, the boundaries of initial superpixels are obtained according to the probabilities and the commute time. The initial superpixels are iteratively optimized by the new energy function, which is defined on the commute time and the texture measurement. Our LRW algorithm with self-loops has the merits of segmenting the weak boundaries and complicated texture regions very well by the new global probability maps and the commute time strategy. The performance of superpixel is improved by relocating the center positions of superpixels and dividing the large superpixels into small ones with the proposed optimization algorithm. The experimental results have demonstrated that our method achieves better performance than previous superpixel approaches. Jianbing Shen, Yunfan Du, Wenguan Wang, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | Depth-Aware Image Seam CarvingabstractImage seam carving algorithm should preserve important and salient objects as much as possible when changing the image size, while not removing the secondary objects in the scene. However, it is still difficult to determine the important and salient objects that avoid the distortion of these objects after resizing the input image. In this paper, we develop a novel depth-aware single image seam carving approach by taking advantage of the modern depth cameras such as the Kinect sensor, which captures the RGB color image and its corresponding depth map simultaneously. By considering both the depth information and the just noticeable difference (JND) model, we develop an efficient JND-based significant computation approach using the multiscale graph cut based energy optimization. Our method achieves the better seam carving performance by cutting the near objects less seams while removing distant objects more seams. To the best of our knowledge, our algorithm is the first work to use the true depth map captured by Kinect depth camera for single image seam carving. The experimental results demonstrate that the proposed approach produces better seam carving results than previous content-aware seam carving methods. Jianbing Shen, Xuelong Li 0001 |
IEEE Trans. Cybern. | 1 |
| 2013 | Intrinsic Image Decomposition Using Optimization and User ScribblesabstractIn this paper, we present a novel high-quality intrinsic image recovery approach using optimization and user scribbles. Our approach is based on the assumption of color characteristics in a local window in natural images. Our method adopts a premise that neighboring pixels in a local window having similar intensity values should have similar reflectance values. Thus, the intrinsic image decomposition is formulated by minimizing an energy function with the addition of a weighting constraint to the local image properties. In order to improve the intrinsic image decomposition results, we further specify local constraint cues by integrating the user strokes in our energy formulation, including constant-reflectance, constant-illumination, and fixed-illumination brushes. Our experimental results demonstrate that the proposed approach achieves a better recovery result of intrinsic reflectance and illumination components than the previous approaches. Jianbing Shen, Xiaoshan Yang, Xuelong Li 0001, Yunde Jia |
IEEE Trans. Cybern. | 1 |
| 2012 | Interactive image/video retexturing using GPU parallelism
Ping Li 0016, Hanqiu Sun, Jianbing Shen, Yongwei Nie |
Comput. Graph. | 4 |
| 2012 | Image stylization with enhanced structure on GPU
Ping Li 0016, Hanqiu Sun, Bin Sheng 0001, Jianbing Shen |
Sci. China Inf. Sci. | 4 |
| 2012 | Detail-preserving exposure fusion using subband architecture
Jianbing Shen, Ying Zhao 0009, Ying He 0001 |
Vis. Comput. | 1 |
| 2011 | Color-Mood-Aware Clothing Re-texturingabstractIn this paper, we present a novel color-mood-aware technique to re-texture clothing in a photograph. An efficient classification algorithm is developed to classify clothing textures using color mood scheme. To re-texture the clothing, our approach first computes the gradient maps for the cloth region to be replaced and then calculates the texture distortion coordinates on the projected cloth region according to the gradient maps. After the user selects a target clothing texture from the classified clothing texture database, the lighting and shading effects on the original photograph is transferred using the HSV color space. Experimental results show that the proposed approach successfully re-textures the clothes in photographs while preserving the geometry and lighting features. Jianbing Shen, Hanqiu Sun, Xiaoyang Mao, Yanwen Guo 0001, Xiaogang Jin 0001 |
CAD/Graphics | 1 |
| 2011 | Intrinsic images using optimizationabstractIn this paper, we present a novel intrinsic image recovery approach using optimization. Our approach is based on the assumption of in a local window in natural images. Our method adopts a premise that neighboring pixels in a local window of a single image having similar intensity values should have similar reflectance values. Thus the intrinsic image decomposition is formulated by optimizing an energy function with adding a weighting constraint to the local image properties. In order to improve the intrinsic image extraction results, we specify local constrain cues by integrating the user strokes in our energy formulation, including constant-reflectance, constant-illumination and fixed-illumination brushes. Our experimental results demonstrate that our approach achieves a better recovery of intrinsic reflectance and illumination components than by previous approaches. Jianbing Shen, Xiaoshan Yang, Yunde Jia, Xuelong Li 0001 |
CVPR | 1 |
| 2010 | A unified framework for designing textures using energy optimization
Jianbing Shen, Hanqiu Sun, Jiaya Jia, Hanli Zhao, Xiaogang Jin 0001, Shiaofen Fang |
Pattern Recognit. | 1 |
| 2009 | Bilateral filtering using fuzzy-median for image manipulationsabstractThis paper presents a novel bilateral filtering using fuzzy-median for image manipulations such as denoising and tone mapping. Our proposed bilateral filtering consists of the standard bilateral filter and the estimation of the pixel values by the fuzzy median filter. We have applied the proposed fuzzy filtering for image denoising with both the impulse and Gaussian random noise, which achieves better results than the bilateral filtering based denoising approaches, the Perona-Maliks anisotropic diffusion filter, the fuzzy vector median filter and the non-local means filter. Further, we develop the tone mapping algorithm of high dynamic range images incorporating the proposed fuzzy filtering, which does not introduce unpleasant visual halo artifacts. Jianbing Shen, Hanqiu Sun, Hanli Zhao, Xiaogang Jin 0001 |
CAD/Graphics | 1 |
| 2009 | Real-time photo style transferabstractThis paper presents a novel approach for real-time photo style transfer. The automatic image manipulation technique is performed in the oRGB color space, which is a new color model based on the psychologically opponent color theory. We transfer color from an appropriate source image to the target image using a simple statistical analysis. In addition, we match the global luminance histogram to achieve better photographic look. Note that the whole pipeline is highly parallel, enabling a GPU-based real-time implementation. Several experimental results are shown to demonstrate the effectiveness and efficiency of the proposed method. Hanli Zhao, Xiaogang Jin 0001, Jianbing Shen |
CAD/Graphics | 3 |
| 2009 | Ram-based tone mapping for high dynamic range imagesabstractIn this paper we present a novel tone mapping algorithm for high dynamic range (HDR) images using the retinal adaptation model (RAM). The physiological evidence suggests that the RAM is obtained by measuring intensity-response functions to flashes of light presented under varying adaptation conditions, which leads to a theoretic-sound model that can be flexibly adapted for tone reproduction. The multiplicative-subtractive process of the model can provide high quality tone mapping results for rendering the HDR images. The experimental results demonstrate that our RAM-based tone mapping approach is effective to produce pleasing results on HDR images in a wide range of real-world scenarios. Jianbing Shen, Hanqiu Sun, Hanli Zhao, Xiaogang Jin 0001 |
ICME | 1 |
| 2009 | AtelierM++: a fast and accurate marbling system
Hanli Zhao, Xiaogang Jin 0001, Shufang Lu, Xiaoyang Mao, Jianbing Shen |
Multim. Tools Appl. | 5 |
| 2009 | Fast approximation of trilateral filter for tone mapping using a signal processing approach
Jianbing Shen, Shiaofen Fang, Hanli Zhao, Xiaogang Jin 0001, Hanqiu Sun |
Signal Process. | 1 |
| 2009 | Real-time saliency-aware video abstraction
Hanli Zhao, Xiaoyang Mao, Xiaogang Jin 0001, Jianbing Shen, Jieqing Feng |
Vis. Comput. | 4 |
| 2008 | Real-Time Tone Mapping for High-Resolution HDR ImagesabstractHigh dynamic range rendering attempts to take an HDR image and produce a more realistic representation on a limited range computer monitor. Although several tone mapping operators have been proposed in recent years, no evaluation has yet been undertaken to explore which operator is more suitable for hardware implementation. In this paper, we begin with our novel GPU implementations of two state-of-the-art operators in real time. Then several experimental results using eight GPU-based tone mapping operators are presented to evaluate which one is better with regard to running efficiency. Our GPU implementation of the Pattanaik operator can achieve real-time performance even on high-resolution HDR images. In addition, we believe that many real-time applications, including HDR video player and environment mapping with HDR textures in games, will benefit from our novel approach. Hanli Zhao, Xiaogang Jin 0001, Jianbing Shen |
CW | 3 |
| 2008 | Real-time feature-aware video abstraction
Hanli Zhao, Xiaogang Jin 0001, Jianbing Shen, Xiaoyang Mao, Jieqing Feng |
Vis. Comput. | 3 |
| 2007 | Gradient based image completion by solving the Poisson equation
Jianbing Shen, Xiaogang Jin 0001, Chuan Zhou 0008, Charlie C. L. Wang |
Comput. Graph. | 1 |
| 2007 | Deformation-based interactive texture design using energy optimization
Jianbing Shen, Xiaogang Jin 0001, Xiaoyang Mao, Jieqing Feng |
Vis. Comput. | 1 |
| 2007 | High dynamic range image tone mapping and retexturing using fast trilateral filtering
Jianbing Shen, Xiaogang Jin 0001, Hanqiu Sun |
Vis. Comput. | 1 |
| 2006 | Completion-based texture design using deformation
Jianbing Shen, Xiaogang Jin 0001, Xiaoyang Mao, Jieqing Feng |
Vis. Comput. | 1 |