VLDB 2026 Research / reviewers in the wild / expert
Dingwen Zhang
dblp:150/6620
· DBLP profile ↗
171ranked-venue papers
32as first author
124since 2021 · last 2027
0000-0001-8369-8886ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 99 · 23 first-author · 71 since 2021Graphics, computer vision, multimedia, augmented reality and games · 82 · 14 first-author · 49 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 2 first-author · 21 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Mamba-driven sifter for salient object detection
Yi Liu 0038, Dingwen Zhang, Shoukun Xu |
Expert Syst. Appl. | 4 |
| 2027 | Joint color-spatial iterative interaction and metric-based motion filtering for unsupervised polyp segmentation in endoscopic videos
Wenlong Song, Yiwen Jia, Jie Chen 0025, Chenchu Xu, Zhifan Gao, Dingwen Zhang |
Neural Networks | 6 |
| 2026 | FGBVD-KD: Frequency-Guided bias-variance decomposition knowledge distillation for fracture detection
Xiangchun Yu, Dingwen Zhang, Longxiang Teng, Hechang Chen, Huashuai Cai, Miaomiao Liang |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Knowledge distillation-based distributed dynamic 3D Gaussian splatting for large scale scene reconstruction
Sicheng Fei, Xuehao Gao, Jinwen Hu, Xiaolei Hou, Dingwen Zhang |
Expert Syst. Appl. | 7 |
| 2026 | CoSurfGS: 3D Surface Gaussian Splatting with Collaborative Distributed Learning for Large-scale Scene Reconstruction
Yalun Dai, Hao Li 0075, Weicai Ye, Danpeng Chen, Dingwen Zhang, Tong He 0001, Guofeng Zhang 0001, Junwei Han 0001 |
Int. J. Comput. Vis. | 7 |
| 2026 | CLIP-based knowledge projector for image-text matching
Dingwen Zhang, Longfei Han, Huaxiang Zhang 0001, Li Liu 0031, Junwei Han 0001 |
Inf. Process. Manag. | 2 |
| 2026 | VSCode-v2: Dynamic Prompt Learning for General Visual Salient and Camouflaged Object Detection With Two-Stage OptimizationabstractSalient object detection (SOD) and camouflaged object detection (COD) are related but distinct binary mapping tasks, each involving multiple modalities that share commonalities while maintaining unique characteristics. Existing approaches often rely on complex, task-specific architectures, leading to redundancy and limited generalization. Our previous work, VSCode, introduced a generalist model that effectively handles four SOD tasks and two COD tasks. VSCode leveraged VST as its foundation model and incorporated 2D prompts within an encoder-decoder framework to capture domain and task-specific knowledge, utilizing a prompt discrimination loss to optimize the model. Building upon the proven effectiveness of our previous work VSCode, we identify opportunities to further strengthen generalization capabilities through focused modifications in model design and optimization strategy. To unlock this potential, we propose VSCode-v2, an extension that introduces a Mixture of Prompt Experts (MoPE) layer to generate adaptive prompts. We also redesign the training process into a two-stage approach: first learning shared features across tasks, then capturing specific characteristics. To preserve knowledge during this process, we incorporate distillation from our conference version model. Furthermore, we propose a contrastive learning mechanism with data augmentation to strengthen the relationships between prompts and feature representations. VSCode-v2 demonstrates balanced performance improvements across six SOD and COD tasks. Moreover, VSCode-v2 effectively handles various multimodal inputs and exhibits zero-shot generalization capability to novel tasks, such as RGB-D Video SOD. Nian Liu 0002, Xuguang Yang, Dingwen Zhang, Deng-Ping Fan, Fahad Shahbaz Khan, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | High-frequency structure transformer for magnetic resonance image super-resolution
Chaowei Fang, Bolin Fu, De Cheng, Lechao Cheng, Dingwen Zhang |
Pattern Recognit. | 5 |
| 2026 | DecoupleNet: Domain-specific task decoupling network for low-light image enhancement
Peiliang Huang, Xianmin Chen, Xiaoxu Feng, Qiangqiang Wang, Dingwen Zhang, Longfei Han, Junwei Han 0001 |
Pattern Recognit. | 5 |
| 2026 | DecoderTracker: Decoder-only end-to-end method for multiple-object tracking
Pan Liao, Feng Yang 0001, Di Wu 0059, Jinwen Yu, Dingwen Zhang |
Pattern Recognit. | 6 |
| 2026 | Semi-supervised camouflaged fixation prediction via self-evolving pseudo-label learning
Longbin Tang, Chen Xia, Dingwen Zhang |
Pattern Recognit. | 4 |
| 2026 | Learning task-shared and specific knowledge via mixture-of-experts in generative model for continual learning
Weinan Zhao, Yanling Ji, Yan Li 0125, De Cheng, Junwei Han 0001, Dingwen Zhang |
Pattern Recognit. | 6 |
| 2026 | Semantic-based saccadic scanpath prediction for autism spectrum disorder
Wenqi Zhong, Chen Xia, Linzhi Yu, Dingwen Zhang, Kuan Li |
Pattern Recognit. | 5 |
| 2026 | Retinex-RAWMamba: Bridging Demosaicing and Denoising for Low-Light RAW Image EnhancementabstractLow-light image enhancement, particularly in cross-domain tasks such as mapping from the raw domain to the sRGB domain, remains a significant challenge. Many deep learning-based methods have been developed to address this issue and have shown promising results in recent years. However, single-stage methods, which attempt to unify the complex mapping across both domains, leading to limited denoising performance. In contrast, existing two-stage approaches typically overlook the characteristic of demosaicing within the Image Signal Processing (ISP) pipeline, leading to color distortions under varying lighting conditions, especially in low-light scenarios. To address these issues, we propose a novel Mamba-based method customized for low light RAW images, called RAWMamba, to effectively handle raw images with different CFAs. Furthermore, we introduce a Retinex Decomposition Module (RDM) grounded in Retinex prior, which decouples illumination from reflectance to facilitate more effective denoising and automatic non-linear exposure correction, reducing the effect of manual linear illumination enhancement. By bridging demosaicing and denoising, better enhancement for low light RAW images is achieved. Experimental evaluations conducted on public datasets SID and MCR demonstrate that our proposed RAWMamba achieves state-of-the-art performance on cross-domain mapping. The code is available at https://github.com/Cynicarlos/RetinexRawMamba. Xianmin Chen, Longfei Han, Peiliang Huang, Xiaoxu Feng, Dingwen Zhang, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | FastTrackTr: Real-Time Multiobject Tracking With Transformers for Real WorldabstractTransformer-based multiobject tracking (MOT) methods have attracted significant attention from researchers. However, these Transformer-based models often suffer from suboptimal inference speeds due to their architectural complexities or other inherent issues, rendering them difficult to deploy in practical industrial applications. To address this challenge, we revisited the classic joint detection and tracking (JDT) paradigm and analyzed existing models. Drawing inspiration from Detection Transformer’s (DETR) object queries, which naturally encode object appearance features, we constructed a fast and novel JDT-type MOT framework named FastTrackTr by implementing an efficient interframe information transfer mechanism. This framework integrates three key technological innovations: a cross-decoder mechanism that implicitly incorporates historical trajectory information without requiring additional queries or decoders, a historical encoder and decoder pair for refining and utilizing historical feature representations, and a deterministic fixed-shape architecture that enables seamless TensorRT acceleration. Benefiting from these designs, our approach not only reduces the number of queries required for tracking but also avoids introducing excessive network structures, ensuring model simplicity while maintaining high accuracy. Experimental results show that our method achieves real-time tracking while maintaining state-of-the-art accuracy. On an NVIDIA RTX 4090 with an image size of 1333 × 800, it reaches 62.4 higher order tracking accuracy (HOTA) at 86.6 frames per second (FPS) on the DanceTrack dataset, outperforming other advanced Transformer-based methods. Furthermore, it excels on edge devices, such as the NVIDIA Jetson AGX Orin, where its lightweight variant achieves up to 59.2 FPS on 640 × 640 images, meeting real-time requirements for practical applications. Pan Liao, Feng Yang 0001, Di Wu 0059, Jinwen Yu, Xingxin Li, Dingwen Zhang |
IEEE Trans. Ind. Informatics | 6 |
| 2026 | Frequency-Aware B-Line and Pleural Line Analysis in Lung Ultrasound VideosabstractAccurately identifying B-lines and pleural line (P-line) in lung ultrasound (LUS) videos is valuable for evaluating certain lung conditions. However, manual interpretation remains subjective and highly dependenton operator expertise. Existing deep learning methods often suffer from performance degradation due to speckle noise and motion artifacts. Moreover, the limited availability of LUS video data annotated for multiple diagnostic features such as B-lines and the P-line limits model development. Therefore, this paper introduces ILD-LUS, a new clinical LUS database designed based on interstitial lung disease (ILD) analysis by category labeling, comprising 2,149 ultrasound videos (193,410 frames). Also, we construct an external test set based on the public Covid-BLUES dataset for the evaluation of B-lines and P-line recognition in different pulmonary pathologies. Then, we propose a novel video analysis framework that integrates wavelet enhancement with temporal attention modeling. Specifically, we employ a dual-component frequency feature enhancement method using the Discrete Wavelet Transform (DWT), which effectively suppresses noise while preserving important landmarks. Subsequently, an adaptive attention module is introduced to model long-range temporal dependencies and improve dynamic feature representation across consecutive frames. Experimental results show that the proposed method achieves over 94% AUC and 82% ACC for both B-lines and P-line classification on both the ILD-LUS and Covid-BLUES datasets, outperforming existing methods. These findings demonstrate the robustness and generalizability of our approach across different pathological conditions. Overall, the proposed framework shows strong potential for supporting clinical decision-making in LUS analysis. Kaihui Yang, Guangyu Guo 0001, Linxuan Pang, Zhaohui Zheng 0004, Ruyu Liu, Jin Ding, Dingwen Zhang, Junwei Han 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | XLD: A Cross-Lane Dataset for Benchmarking Novel Driving View SynthesisabstractComprehensive testing of autonomous systems through simulation is essential to ensure the safety of autonomous driving vehicles. This requires the generation of safety-critical scenarios that extend beyond the limitations of real-world data collection, as many of these scenarios are rare or rarely encountered on public roads. However, evaluating most existing novel view synthesis (NVS) methods relies on sporadic sampling of image frames from the training data, comparing the rendered images with ground-truth images. Unfortunately, this evaluation protocol falls short of meeting the actual requirements in closed-loop simulations. Specifically, the true application demands the capability to render novel views that extend beyond the original trajectory (such as cross-lane views), which are challenging to capture in the real world. To address this, this paper presents a synthetic dataset for novel driving view synthesis evaluation, which is specifically designed for autonomous driving simulations. This unique dataset includes testing images captured by deviating from the training trajectory by 1–4 meters. It comprises six sequences that cover various times and weather conditions. Each sequence contains 450 training images, 120 testing images, and their corresponding camera poses and intrinsic parameters. Leveraging this novel dataset, we establish the first realistic benchmark for evaluating existing NVS approaches under frontonly and multicamera settings. The experimental findings underscore the significant gap in current approaches, revealing their inadequate ability to fulfill the demanding prerequisites of cross-lane or closed-loop simulation. Our dataset and code are released publicly on the project page: https://3d-aigc.github.io/XLD. Hao Li 0075, Chenming Wu, Chen Zhao 0011, Chunyu Song, Haocheng Feng, Errui Ding, Dingwen Zhang, Jingdong Wang 0001 |
3DV | 9 |
| 2025 | Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose InteractionabstractVideo virtual try-on aims to seamlessly dress a subject in a video with a specific garment. The primary challenge involves preserving the visual authenticity of the garment while dynamically adapting to the pose and physique of the subject. While existing methods have predominantly focused on image-based virtual try-on, extending these techniques directly to videos often results in temporal inconsistencies. Most current video virtual try-on approaches alleviate this challenge by incorporating temporal modules, yet still overlook the critical spatiotemporal pose interactions between human and garment. Effective pose interactions in videos should not only consider spatial alignment between human and garment poses in each frame but also account for the temporal dynamics of human poses throughout the entire video. With such motivation, we propose a new framework, namely Dynamic Pose Interaction Diffusion Models (DPIDM), to leverage diffusion models to delve into dynamic pose interactions for video virtual try-on. Technically, DPIDM introduces a skeleton-based pose adapter to integrate synchronized human and garment poses into the denoising network. A hierarchical attention module is then exquisitely designed to model intra-frame human-garment pose interactions and long-term human pose dynamics across frames through pose-aware spatial and temporal attention mechanisms. Moreover, DPIDM capitalizes on a temporal regularized attention loss between consecutive frames to enhance temporal consistency. Extensive experiments conducted on VITON-HD, VVT and ViViD datasets demonstrate the superiority of our DPIDM against the baseline methods. Notably, DPIDM achieves VFID score of 0.506 on VVT dataset, leading to 60.5% improvement over the state-of-the-art GPD-VVTO approach. Dong Li 0019, Wenqi Zhong, Wei Yu 0004, Yingwei Pan, Dingwen Zhang, Ting Yao 0003, Junwei Han 0001, Tao Mei 0001 |
CVPR | 5 |
| 2025 | CityGS-$\mathcal{X}$: A Scalable Architecture for Efficient and Geometrically Accurate Large-Scale Scene Reconstruction
Hao Li 0069, Zhengyu Zou, Zhihang Zhong, Dingwen Zhang, Junwei Han 0001 |
ICCV | 6 |
| 2025 | Navigating Semantic Drift in Task-Agnostic Class-Incremental LearningabstractClass-incremental learning (CIL) seeks to enable a model to sequentially learn new classes while retaining knowledge of previously learned ones. Balancing flexibility and stability remains a significant challenge, particularly when the task ID is unknown. To address this, our study reveals that the gap in feature distribution between novel and existing tasks is primarily driven by differences in mean and covariance moments. Building on this insight, we propose a novel semantic drift calibration method that incorporates mean shift compensation and covariance calibration. Specifically, we calculate each class's mean by averaging its sample embeddings and estimate task shifts using weighted embedding changes based on their proximity to the previous mean, effectively capturing mean shifts for all learned classes with each new task. We also apply Mahalanobis distance constraint for covariance calibration, aligning class-specific embedding covariances between old and current networks to mitigate the covariance shift. Additionally, we integrate a feature-level self-distillation approach to enhance generalization. Comprehensive experiments on commonly used datasets demonstrate the effectiveness of our approach. The source code is available at https://github.com/fwu11/MACIL.git. Fangwen Wu, Lechao Cheng, Shengeng Tang, Chaowei Fang, Dingwen Zhang, Meng Wang 0001 |
ICML | 6 |
| 2025 | DGTR: Distributed Gaussian Turbo-Reconstruction for Sparse-View Vast ScenesabstractNovel-view synthesis approaches play a critical role in vast scene reconstruction. However, these methods rely heavily on dense image inputs and prolonged training times, making them unsuitable where computational resources are limited. Additionally, few-shot methods often struggle with poor reconstruction quality in vast environments. This paper presents DGTR, a novel distributed framework for efficient Gaussian reconstruction for sparse-view vast scenes. Our approach divides the scene into regions, processed independently by drones with sparse image inputs. Using a feed-forward Gaussian model, we predict high-quality Gaussian primitives, followed by a global alignment algorithm to ensure geometric consistency. Depth priors is incorporated to further enhance training, while a distillation-based model aggregation mechanism enables efficient reconstruction. Our method achieves high-quality large-scale scene reconstruction and novel-view synthesis in significantly reduced training times, outperforming existing approaches in both speed and scalability. We demonstrate the effectiveness of our framework on vast aerial scenes, achieving high-quality results within minutes. Code will released on our project page https://3d-aigc.github.io/DGTR. Hao Li 0075, Haosong Peng, Chenming Wu, Weicai Ye, Yufeng Zhan, Chen Zhao 0011, Dingwen Zhang, Jingdong Wang 0001, Junwei Han 0001 |
ICRA | 8 |
| 2025 | Multi-modal Progressive Fusion for ASD Screening Using Smartphone Video
Wenqi Zhong, Chen Xia, Kuan Li, Dingwen Zhang |
MICCAI (9) | 5 |
| 2025 | STRIDER: Navigation via Instruction-Aligned Structural Decision Space OptimizationabstractThe Zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) task requires agents to navigate previously unseen 3D environments using natural language instructions, without any scene-specific training. A critical challenge in this setting lies in ensuring agents’ actions align with both spatial structure and task intent over long-horizon execution. Existing methods often fail to achieve robust navigation due to a lack of structured decision-making and insufficient integration of feedback from previous actions. To address these challenges, we propose STRIDER (Instruction-Aligned Structural Decision Space Optimization), a novel framework that systematically optimizes the agent’s decision space by integrating spatial layout priors and dynamic task feedback. Our approach introduces two key innovations: 1) a Structured Waypoint Generator that constrains the action space through spatial structure, and 2) a Task-Alignment Regulator that adjusts behavior based on task progress, ensuring semantic alignment throughout navigation. Extensive experiments on the R2R-CE and RxR-CE benchmarks demonstrate that STRIDER significantly outperforms strong SOTA across key metrics; in particular, it improves Success Rate (SR) from 29\% to 35\%, a relative gain of 20.7\%. Such results highlight the importance of spatially constrained decision-making and feedback-guided execution in improving navigation fidelity for zero-shot VLN-CE. Diqi He, Xuehao Gao, Hao Li 0075, Junwei Han 0001, Dingwen Zhang |
NeurIPS | 5 |
| 2025 | Seamless Detection: Unifying Salient Object Detection and Camouflaged Object Detection
Yi Liu 0038, Dingwen Zhang, Shoukun Xu, Jungong Han |
Expert Syst. Appl. | 5 |
| 2025 | Hierarchical candidate recursive network for highlight restoration in endoscopic videos
Chenchu Xu, Jiangnan Wu, Dong Zhang 0009, Longfei Han, Dingwen Zhang, Junwei Han 0001 |
Expert Syst. Appl. | 5 |
| 2025 | Attention correction feature and boundary constraint knowledge distillation for efficient 3D medical image segmentation
Xiangchun Yu, Longxiang Teng, Dingwen Zhang, Hechang Chen |
Expert Syst. Appl. | 3 |
| 2025 | LLaVA-Endo: a large language-and-vision assistant for gastrointestinal endoscopy
Jieru Yao, Xueran Li, Longfei Han, Yiwen Jia, Nian Liu 0002, Dingwen Zhang, Junwei Han 0001 |
Frontiers Comput. Sci. | 7 |
| 2025 | Semantic-Aligned Learning with Collaborative Refinement for Unsupervised VI-ReID
De Cheng, Nannan Wang 0001, Dingwen Zhang, Xinbo Gao 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | Mamba Capsule Routing Towards Part-Whole Relational Camouflaged Object Detection
Dingwen Zhang, Liangbo Cheng, Yi Liu 0038, Xinggang Wang, Junwei Han 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | WeakCLIP: Adapting CLIP for Weakly-Supervised Semantic Segmentation
Lianghui Zhu, Xinggang Wang, Jiapei Feng, Tianheng Cheng, Yingyue Li, Bo Jiang 0011, Dingwen Zhang, Junwei Han 0001 |
Int. J. Comput. Vis. | 7 |
| 2025 | Bias-variance decomposition knowledge distillation for medical image segmentationabstractKnowledge distillation essentially maximizes the mutual information between teacher and student networks. Typically, a variational distribution is introduced to maximize the variational lower bound. However, the heteroscedastic noises derived from this distribution are often unstable, leading to unreliable data-uncertainty modeling. Our research identifies that bias-variance coupling in knowledge distillation causes this instability. We thus propose Bias-variance dEcomposition kNowledge dIstillatioN (BENIN) approach. Initially, we use bias-variance decomposition to decouple these components. Subsequently, we design a lightweight Feature Frequency Expectation Estimation Module (FF-EEM) to estimate the student's prediction expectation, which helps compute bias and variance. Variance learning measures data uncertainty in the teacher's prediction. A balance factor addresses the bias-variance dilemma. Lastly, the bias-variance decomposition distillation loss enables the student to learn valuable knowledge while reducing noise. Experiments on Synapse and Lits17 medical-image-segmentation datasets validate BENIN's effectiveness. FF-EEM also mitigates high-frequency noise from high mask rates, enhancing data-uncertainty estimation and visualization. Our code is available at https://github.com/duanzhongjian/BENIN . Xiangchun Yu, Longxiang Teng, Zhongjian Duan, Dingwen Zhang, Wei Pang 0001, Miaomiao Liang, Liujin Qiu |
Neurocomputing | 4 |
| 2025 | Enhancing 3D multi-organ segmentation via uncertainty guidance and boundary knowledge distillation
Xiangchun Yu, Longjun Ding, Dingwen Zhang |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | EDGE: Edge distillation and gap elimination for heterogeneous networks in 3D medical image segmentation
Xiangchun Yu, Dingwen Zhang, Jianqing Wu 0002 |
Knowl. Based Syst. | 3 |
| 2025 | Ground truth is the best teacher: supervised semantic segmentation inspired by knowledge transfer mechanisms
Xiangchun Yu, Huofa Liu, Dingwen Zhang, Miaomiao Liang, Lingjuan Yu |
Multim. Syst. | 3 |
| 2025 | Hierarchical Region-level Decoupling Knowledge Distillation for semantic segmentation
Xiangchun Yu, Huofa Liu, Dingwen Zhang, Jianqing Wu 0002 |
Multim. Syst. | 3 |
| 2025 | Advanced Discriminative Co-Saliency and Background Mining Transformer for Co-Salient Object DetectionabstractMost existing CoSOD models focus solely on extracting co-saliency cues while neglecting explicit exploration of background regions, potentially leading to difficulties in handling interference from complex background areas. To address this, this paper proposes a Discriminative co-saliency and background Mining Transformer framework (DMT) to explicitly mine both co-saliency and background information and effectively model their discriminability. DMT first learns two types of tokens by disjointly extracting co-saliency and background information from segmentation features, then performs discriminability within the segmentation features guided by these well-learned tokens. In the first phase, we propose economic multi-grained correlation modules for efficient detection information extraction, including Region-to-Region (R2R), Contrast-induced Pixel-to-Token (CtP2T), and Co-saliency Token-to-Token (CoT2T) correlation modules. In the subsequent phase, we introduce Token-Guided Feature Refinement (TGFR) modules to enhance discriminability within the segmentation features. To further enhance the discriminative modeling and practicality of DMT, we first upgrade the original TGFR's intra-image modeling approach to an intra-group one, thus proposing Group TGFR (G-TGFR), which is more suitable for the co-saliency task. Subsequently, we designed a Noise Propagation Suppression (NPS) mechanism to apply our model to a more practical open-world scenario, ultimately presenting our extended version, i.e. DMT+O. Extensive experimental results on both conventional CoSOD and open-world CoSOD benchmark datasets demonstrate the effectiveness of our proposed model. Long Li 0008, Huichao Xie, Nian Liu 0002, Dingwen Zhang, Rao Muhammad Anwer, Hisham Cholakkal, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Unsupervised Pre-Training With Language-Vision Prompts for Low-Data Instance SegmentationabstractIn recent times, following the paradigm of DETR (DEtection TRansformer), query-based end-to-end instance segmentation (QEIS) methods have exhibited superior performance compared to CNN-based models, particularly when trained on large-scale datasets. Nevertheless, the effectiveness of these QEIS methods diminishes significantly when confronted with limited training data. This limitation arises from their reliance on substantial data volumes to effectively train the pivotal queries/kernels that are essential for acquiring localization and shape priors. To address this problem, we propose a novel method for unsupervised pre-training in low-data regimes. Inspired by the recently successful prompting technique, we introduce a new method, Unsupervised Pre-training with Language-Vision Prompts (UPLVP), which improves QEIS models' instance segmentation by bringing language-vision prompts to queries/kernels. Our method consists of three parts: (1) Masks Proposal: Utilizes language-vision models to generate pseudo masks based on unlabeled images. (2) Prompt-Kernel Matching: Converts pseudo masks into prompts and injects the best-matched localization and shape features to their corresponding kernels. (3) Kernel Supervision: Formulates supervision for pre-training at the kernel level to ensure robust learning. With the help of our pre-training method, QEIS models can converge faster and perform better than CNN-based models in low-data regimes. Experimental evaluations conducted on MS COCO, Cityscapes, and CTW1500 datasets indicate that the QEIS models' performance can be significantly improved when pre-trained with our method. Dingwen Zhang, Hao Li 0075, Diqi He, Nian Liu 0002, Lechao Cheng, Jingdong Wang 0001, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | A Learning Paradigm for Selecting Few Discriminative Stimuli in Eye-Tracking ResearchabstractEye-tracking is a reliable method for quantifying visual information processing and holds significant potential for group recognition, such as identifying autism spectrum disorder (ASD). However, eye-tracking research typically faces the heterogeneity of stimuli and is time-consuming due to the large number of observed stimuli. To address these issues, we first mathematically define the stimulus selection problem and introduce the concept of stimulus discrimination ability to reduce the computational complexity of the solution. Then, we construct a scanpath-based recognition model to mine the stimulus discrimination ability. Specifically, we propose cross-subject entropy and cross-subject divergence scores for quantitatively evaluating stimulus discrimination ability, effectively capturing differences in intra-group collective trends and inter-subject consistency within a group. Furthermore, we propose an iterative learning mechanism that employs stimulus-wise attention to focus on discriminative stimuli for discrimination purification. In the experiment, we construct an ASD eye-tracking dataset with diverse stimulus types and conduct extensive tests on three representative models to validate our approach. Remarkably, our method demonstrates superior performance using only 10 selected stimuli compared to models utilizing 220 stimuli. Additionally, we perform experiments on another eye-tracking task, gender prediction, to further validate our method. We believe that our approach is both simple and flexible for integration into existing models, promoting large-scale ASD screening and extending to other eye-tracking research domains. Wenqi Zhong, Chen Xia, Linzhi Yu, Kuan Li, Zhongyu Li 0002, Dingwen Zhang, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Identifying Children With Autism Spectrum Disorder via Transformer-Based Representation Learning From Dynamic Facial CuesabstractRecognizing autism spectrum disorder (ASD) has faced great challenges due to insufficient professional clinicians and complex procedures. Automated data-driven ASD recognition models can reduce the subjectivity and physician dependency of traditional evaluation methods. Facial data, which can encode important perceptual and social behaviors, have emerged in ASD research to explore novel biomarkers for screening, diagnosing, and treating ASD. However, existing research mainly focuses on extracting low-level hand-crafted facial features for analysis and classification. Determining how to learn discriminative deep representations from dynamic facial data for computational model construction remains an unresolved challenge. In this study, we propose an ASD recognition model based on facial videos to fill the lack of temporal correlation learning of facial features. First, we utilize a vision transformer to extract frame-based global facial features. Then, we use a Longformer to establish the correlation of facial features over time. In the experiment, we recruited 146 subjects between 2 and 8 years of age to record their facial videos under a computer-based eye-tracking experiment and 76 subjects to conduct a smartphone-based experiment. Quantitative comparisons have shown the effectiveness and reliability of the proposed model. Furthermore, we have confirmed the correlation between facial and eye-tracking modalities in visual attention. Chen Xia, Hexu Chen, Junwei Han 0001, Dingwen Zhang, Kuan Li |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Achieving Plasticity-Stability Trade-Off in Continual Learning Through Adaptive Orthogonal ProjectionabstractCatastrophic forgetting is the crucial challenge for continual learning. One of the state-of-the-art approaches is the orthogonal projection, which aims to learn each task by updating model parameters in the direction orthogonal to the subspace spanned by the previous task input. Although such strict orthogonal weight constraints ensure no interference with tasks that have been learned to achieve model stability, they greatly sacrifice model plasticity. In this paper, we propose an adaptive balanced orthogonal projection (AdaBOP) method, to search for the optimal network parameter updating direction to address the plasticity-stability dilemma in continual learning. The proposed AdaBOP method can adaptively adjust its tendency towards plasticity-stability trade-off based on the layer-wise feature space correlations of the model between old and new tasks. To further improve the training efficiency, we also implement the AdaBOP method in the uncentered covariance matrix space of the previous tasks, and finally achieve a better stability-plasticity trade-off in continual learning efficiently. Experimental results greatly demonstrate the effectiveness of the proposed method, which achieves superior performances to state-of-the-art continual learning approaches. The code is available athttps://github.com/hyscn/AdaBOP. De Cheng, Yusong Hu, Nannan Wang 0001, Dingwen Zhang, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | TRNet: Two-Tier Recursion Network for Co-Salient Object DetectionabstractCo-salient object detection (CoSOD) is to find the salient and recurring objects from a series of relevant images, where modeling inter-image relationships plays a crucial role. Different from the commonly used direct learning structure that inputs all the intra-image features into some well-designed modules to represent the inter-image relationship, we resort to adopting a recursive structure for inter-image modeling, and propose a two-tier recursion network (TRNet) to achieve CoSOD in this paper. The two-tier recursive structure of the proposed TRNet is embodied in two stages of inter-image extraction and distribution. On the one hand, considering the task adaptability and inter-image correlation, we design an inter-image exploration with recursive reinforcement module to learn the local and global inter-image correspondences, guaranteeing the validity and discriminativeness of the information in the step-by-step propagation. On the other hand, we design a dynamic recursion distribution module to fully exploit the role of inter-image correspondences in a recursive structure, adaptively assigning common attributes to each individual image through an improved semi-dynamic convolution. Experimental results on five prevailing CoSOD benchmarks demonstrate that our TRNet outperforms other competitors in terms of various evaluation metrics. The code and results of our method are available athttps://github.com/rmcong/TRNet_TCSVT2025. Runmin Cong, Ning Yang 0008, Hongyu Liu 0003, Dingwen Zhang, Qingming Huang, Sam Kwong, Wei Zhang 0021 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | SA-MixNet: Structure-Aware Mixup and Invariance Learning for Scribble-Supervised Road Extraction in Remote Sensing ImagesabstractMainstreamed weakly supervised road extractors rely on highly confident pseudo-labels propagated from scribbles, and their performance often degrades gradually as the image scenes tend to vary. We argue that such degradation is due to the poor model’s invariance to scenes with different complexities, whereas existing solutions to this problem are commonly based on crafted priors that cannot be derived from scribbles. To eliminate the reliance on such priors, we propose a novel structure-aware mixup and invariance learning framework (SA-MixNet) for weakly supervised road extraction that improves the model invariance in a data-driven manner. Specifically, we design a structure-aware mixup (SA-Mix) scheme to paste road regions from one image onto another to create an image scene with increased complexity while preserving the road’s structural integrity. Then, an invariance regularization is imposed on the predictions of constructed and origin images to minimize their conflicts, which thus forces the model to behave consistently in various scenes. Moreover, a discriminator-based regularization is designed to enhance connectivity while preserving the structure of roads. Combining these designs, our framework demonstrates superior performance on the DeepGlobe, Wuhan, and Massachusetts datasets, outperforming the state-of-the-art techniques by 1.47%, 2.12%, and 4.09%, respectively, in IoU metrics, and showing its potential as a plug-and-play solution. Our source code is available athttps://github.com/xdu-jjgs. Jie Feng 0003, Junpeng Zhang 0002, Weisheng Dong, Dingwen Zhang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | R2PLoc: A Region-to-Point UAV Visual Geo-Localization Framework Leveraging Hierarchical Semantic RepresentationabstractThe challenges in UAV visual geo-localization primarily stem from discrepancies between satellite maps and aerial images, including scale variations, viewpoint deviations, and spatiotemporal mismatches. Current approaches adopt retrieval-based or keypoint-matching-based localization, and some studies employ a cascaded approach. However, these methods still exhibit limitations in addressing discrepancies. To address these challenges, we propose a region-to-point UAV visual geo-localization framework named R2PLoc. Specifically, we consider UAV visual geo-localization as the process of retrieving corresponding regions from satellite map databases using aerial images while establishing projective relationship between them. First, we employ a shared backbone network for semantic feature extraction to conserve computational resources. Then, the Hierarchical Semantic Aggregation Module (HSAM) is designed to address the feature distribution shifts by fusing multi-scale semantics that combine both global contexts and local structures. Additionally, the Semantic-Enhanced Hierarchical Refinement Matcher (SHRM) is constructed to improve the geometric consistency of keypoint matching by integrating high-level semantic information. Furthermore, the UAV-R2P dataset is constructed for the region-to-point geo-localization task. The qualitative and quantitative experimental results demonstrate that our method outperforms most state-of-the-art methods with similar model size on most available datasets. Ruitao Lu, Yansheng Li 0001, Yunsong Li 0001, Dingwen Zhang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | MROD-YOLO: Multimodal Joint Representation for Small Object Detection in Remote Sensing Imagery via Multiscale Iterative Aggregation
Ruitao Lu, Dingwen Zhang, Weiying Xie, Shuang Su, Zhenyu Zhang 0028 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Weakly Supervised Semantic Segmentation via Alternate Self-Dual TeachingabstractWeakly supervised semantic segmentation (WSSS) is a challenging yet important research field in vision community. In WSSS, the key problem is to generate high-quality pseudo segmentation masks (PSMs). Existing approaches mainly depend on the discriminative object part to generate PSMs, which would inevitably miss object parts or involve surrounding image background, as the learning process is unaware of the full object structure. In fact, both the discriminative object part and the full object structure are critical for deriving of high-quality PSMs. To fully explore these two information cues, we build a novel end-to-end learning framework, alternate self-dual teaching (ASDT), based on a dual-teacher single-student network architecture. The information interaction among different network branches is formulated in the form of knowledge distillation (KD). Unlike the conventional KD, the knowledge of the two teacher models would inevitably be noisy under weak supervision. Inspired by the Pulse Width (PW) modulation, we introduce a PW wave-like selection signal to alleviate the influence of the imperfect knowledge from either teacher model on the KD process. Comprehensive experiments on the PASCAL VOC 2012 and COCO-Stuff 10K demonstrate the effectiveness of the proposed ASDT framework, and new state-of-the-art results are achieved. Dingwen Zhang, Hao Li 0075, Wenyuan Zeng, Chaowei Fang, Lechao Cheng, Ming-Ming Cheng, Junwei Han 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | MHKD: Multi-Step Hybrid Knowledge Distillation for Low-Resolution Whole Slide Images Glomerulus DetectionabstractGlomerulus detection is a critical component of renal histopathology assessment, essential for diagnosing glomerulonephritis. To mitigate the increasing workload on pathologists, AI-assisted diagnostic methods based on high-resolution digital pathology whole slide images have been developed. However, these current AI-assisted approaches are limited to high-resolution whole slide images, necessitating expensive digital scanner equipment, high image storage costs, and significant computational complexity. To address this limitation, this paper pioneers a method for facilitating glomerulus detection in low-resolution human kidney pathology images. Specifically, we propose a novel multi-step hybrid knowledge distillation method. Our method distills both the global features and the semantic information through a hybrid knowledge distillation strategy that integrates offline and online knowledge distillation, where the information from high-resolution pathological images is successively transferred to student model from the global features in the shallow network layers to the semantic information of the back-end through a multi-step training strategy. Experimental results on two datasets show that the proposed method achieves effective detection outcomes for low-resolution kidney pathology images. Compared to other state-of-the-art detection techniques, our method achieves an ${AP}_{0.5:0.95}$ improvement of 23.1% on the private LN dataset and 15.9% on the public HUBMAP dataset. Xiangsen Zhang, Longfei Han, Chenchu Xu, Zhaohui Zheng 0004, Jin Ding, Xianghui Fu, Dingwen Zhang, Junwei Han 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | HV-BEV: Decoupling Horizontal and Vertical Feature Sampling for Multi-View 3D Object DetectionabstractVision-based multi-view perception systems, especially those using bird’s-eye view (BEV) representations, have become increasingly important in autonomous driving. Current state-of-the-art methods typically transform image features from multiple camera views into the BEV space via explicit or implicit depth estimation. However, they often rely on uniform or fixed-height sampling and lack height-aware priors, making them insensitive to the fact that different object categories occupy distinct local height ranges. Furthermore, most approaches treat BEV features as independent across grid locations, overlooking the structured correlations between different parts of an object in 3D space. These limitations hinder accurate spatial reasoning in both vertical and horizontal dimensions. To address these issues, we propose HV-BEV, a novel BEV perception framework that decouples the feature sampling process into Horizontal feature aggregation and Vertical adaptive height-aware reference point sampling. Specifically, for horizontal modeling, we dynamically construct a set of relevant neighboring points on the ground-aligned plane for each 3D reference point, facilitating structured cross-view feature aggregation and promoting consistent representation of large or partially visible objects. For vertical modeling, we introduce an adaptive height-aware module that leverages historical information to guide 3D reference points to focus on the plausible height regions where objects of interest are likely to appear, replacing fixed uniform height sampling. Extensive experiments on the nuScenes dataset demonstrate the effectiveness of our method. Our HV-BEV framework consistently outperforms baselines, achieving 50.5% mAP and 59.8% NDS on the nuScenes test set. Code is available at https://github.com/Uddd821/HV-BEV Di Wu 0059, Feng Yang 0001, Benlian Xu, Pan Liao, Dingwen Zhang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Prompting Vision-Language Model for Nuclei Instance Segmentation and ClassificationabstractNuclei instance segmentation and classification are a fundamental and challenging task in whole slide Imaging (WSI) analysis. Most dense nuclei prediction studies rely heavily on crowd labelled data on high-resolution digital images, leading to a time-consuming and expertise-required paradigm. Recently, Vision-Language Models (VLMs) have been intensively investigated, which learn rich cross-modal correlation from large-scale image-text pairs without tedious annotations. Inspired by this, we build a novel framework, called PromptNu, aiming at infusing abundant nuclei knowledge into the training of the nuclei instance recognition model through vision-language contrastive learning and prompt engineering techniques. Specifically, our approach starts with the creation of multifaceted prompts that integrate comprehensive nuclear knowledge, including visual insights from the GPT-4V model, statistical analyses, and expert insights from the pathology field. Then, we propose a novel prompting methodology that consists of two pivotal vision-language contrastive learning components: the Prompting Nuclei Representation Learning (PNuRL) and the Prompting Nuclei Dense Prediction (PNuDP), which adeptly integrates the expertise embedded in pre-trained VLMs and multifaceted prompts into the feature extraction and prediction process, respectively. Comprehensive experiments on six datasets with extensive WSI scenarios demonstrate the effectiveness of our method for both nuclei instance segmentation and classification tasks. The code is available at https://github.com/NucleiDet/PromptNu. Jieru Yao, Guangyu Guo 0001, Zhaohui Zheng 0004, Longfei Han, Dingwen Zhang, Junwei Han 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Progressive Prompt-Driven Low-Light Image Enhancement With Frequency Aware LearningabstractLow-light Image Enhancement (LLIE) aims to rectify inadequate illumination conditions and achieve superior visual quality in images, which plays a pivotal role in the domain of low-level computer vision. Due to poor illumination in images, many high-frequency details are obscured, which leads to an uneven distribution of low- and high-frequency information. However, most existing LLIE methods do not pay special attention to the restoration of high-frequency detail information and some challenging-to-recover areas in images. To address this issue, we propose a novel progressive prompt-driven LLIE framework with frequency aware learning, through a two-stage coarse-to-fine learning mechanism. Specifically, the proposed method fully utilizes both the specially designed brightness-aware prompt and detail-aware prompt on the prior trained model, to achieve an excellent enhanced image that exhibits more natural brightness and richer detail information. Furthermore, the proposed frequency aware learning objective can adaptively adjust the contribution of individual pixels for image reconstruction based on the statistics of high- and low-frequency features, which enables the network to focus on learning intricate details and other challenging areas in low-light images. Extensive experimental results demonstrate the effectiveness of the proposed method, achieving superior performances to state-of-the-art methods on representative real-world and synthetic datasets. Our source code is available athttps://github.com/MSL502/PPFAL. De Cheng, Yan Li 0125, Nannan Wang 0001, Dingwen Zhang, Xinbo Gao 0001, Jiande Sun 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Learning Video Salient Object Detection Progressively From Unlabeled VideosabstractRecently, deep learning-based video salient object detection (VSOD) has achieved some breakthroughs, but these methods rely on expensive annotated videos with pixel-wise annotations or weak annotations. In this paper, based on the similarities and differences between VSOD and image salient object detection (SOD), we propose a novel VSOD method via a progressive framework that locates and segments salient objects in sequence without utilizing any video annotation. To efficiently use the knowledge learned in the SOD dataset for VSOD efficiently, we introduce dynamic saliency to compensate for the lack of motion information of SOD during the locating process while maintaining the same fine segmenting process. Specifically, we utilize the coarse locating model trained on the image dataset, to identify frames with both static and dynamic saliency. Locating results of these frames are selected as spatiotemporal location labels. Moreover, by tracking salient objects in adjacent frames, the number of spatiotemporal location labels is increased. On the basis of these location labels, a two-stream locating network with an optical flow branch is proposed to capture salient objects in videos. The results with respect to five public benchmarks demonstrate that our method outperforms the state-of-the-art weakly and unsupervised methods. Binwei Xu, Qiuping Jiang, Haoran Liang 0001, Dingwen Zhang, Ronghua Liang, Peng Chen 0008 |
IEEE Trans. Multim. | 4 |
| 2025 | Capsule Networks With Residual Pose RoutingabstractCapsule networks (CapsNets) have been known difficult to develop a deeper architecture, which is desirable for high performance in the deep learning era, due to the complex capsule routing algorithms. In this article, we present a simple yet effective capsule routing algorithm, which is presented by a residual pose routing. Specifically, the higher-layer capsule pose is achieved by an identity mapping on the adjacently lower-layer capsule pose. Such simple residual pose routing has two advantages: 1) reducing the routing computation complexity and 2) avoiding gradient vanishing due to its residual learning framework. On top of that, we explicitly reformulate the capsule layers by building a residual pose block. Stacking multiple such blocks results in a deep residual CapsNets (ResCaps) with a ResNet-like architecture. Results on MNIST, AffNIST, SmallNORB, and CIFAR-10/100 show the effectiveness of ResCaps for image classification. Furthermore, we successfully extend our residual pose routing to large-scale real-world applications, including 3-D object reconstruction and classification, and 2-D saliency dense prediction. The source code has been released on https://github.com/liuyi1989/ResCaps. Yi Liu 0038, De Cheng, Dingwen Zhang, Shoukun Xu, Jungong Han |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | SpFormer: Spatio-Temporal Modeling for Scanpaths with TransformerabstractSaccadic scanpath, a data representation of human visual behavior, has received broad interest in multiple domains. Scanpath is a complex eye-tracking data modality that includes the sequences of fixation positions and fixation duration, coupled with image information. However, previous methods usually face the spatial misalignment problem of fixation features and loss of critical temporal data (including temporal correlation and fixation duration). In this study, we propose a Transformer-based scanpath model, SpFormer, to alleviate these problems. First, we propose a fixation-centric paradigm to extract the aligned spatial fixation features and tokenize the scanpaths. Then, according to the visual working memory mechanism, we design a local meta attention to reduce the semantic redundancy of fixations and guide the model to focus on the meta scanpath. Finally, we progressively integrate the duration information and fuse it with the fixation features to solve the problem of ambiguous location with the Transformer block increasing. We conduct extensive experiments on four databases under three tasks. The SpFormer establishes new state-of-the-art results in distinct settings, verifying its flexibility and versatility in practical applications. The code can be obtained from https://github.com/wenqizhong/SpFormer. Wenqi Zhong, Linzhi Yu, Chen Xia, Junwei Han 0001, Dingwen Zhang |
AAAI | 5 |
| 2024 | A Smooth Conditional Domain Adversarial Training Framework for EEG Motor Imagery DecodingabstractThe brain-computer interface (BCI) based on electroencephalogram (EEG) motor imagery (MI) decoding demonstrates promising application potential. However, the domain shift between training and testing data significantly impacts the model’s decoding efficacy. Domain adaption (DA) has been developed to address this problem recently. Nevertheless, existing DA methods have two limitations. One is that the extracted features are noisy, and the other is that they only align the distribution of features, which leads to limited generalization ability of the model. In this paper, we propose a novel smooth conditional domain adversarial training framework for solving the motor imagery decoding problem under domain shift. The framework uses interactive frequency convolution and channel attention mechanism as feature extractors to obtain effective features, and integrates smooth conditional domain adversarial training with batch spectral penalty to align the joint distribution of features and classes. At the same time, self-iterative training is implemented by generating pseudo-labels and selective outlier removal. Experimental results demonstrate that our proposed framework achieves 80.67% and 86.17% average accuracy in the BCI IV 2a and 2b respectively for cross-session experiments, achieving the best results compared with other methods, proving that the framework can improve the classification ability on the target domain while transferring effective features. Qilong Yuan, Enze Shi, Kui Zhao, Dingwen Zhang, Shu Zhang 0006 |
BIBM | 5 |
| 2024 | GP-NeRF: Generalized Perception NeRF for Context-Aware 3D Scene UnderstandingabstractApplying Neural Radiance Fields (NeRF) to downstream perception tasks for scene understanding and representation is becoming increasingly popular. Most existing methods treat semantic prediction as an additional rendering task, i.e., the “label rendering” task, to build semantic NeRFs. However, by rendering semantic/instance labels per pixel without considering the contextual information of the rendered image, these methods usually suffer from unclear boundary segmentation and abnormal segmentation of pixels within an object. To solve this problem, we propose Generalized Perception NeRF (GP-NeRF), a novel pipeline that makes the widely used segmentation model and NeRF work compatibly under a unified framework, for facilitating context-aware 3D scene perception. To accomplish this goal, we introduce transformers to aggregate radiance as well as semantic embedding fields jointly for novel views and facilitate the joint volumetric rendering of both fields. In addition, we propose two self-distillation mechanisms, i.e., the Semantic Distill Loss and the Depth-Guided Semantic Distill Loss, to enhance the discrimination and quality of the semantic field and the maintenance of geometric consistency. In evaluation, as shown in Fig. 1 we conduct experimental comparisons under two perception tasks (i.e. semantic and instance segmentation) using both synthetic and real-world datasets. Notably, our method outperforms SOTA approaches by 6.94%,11.76%, and 8.47% on generalized semantic segmentation, finetuning semantic segmentation, and instance segmentation, respectively. Project. Hao Li 0075, Dingwen Zhang, Yalun Dai, Nian Liu 0002, Lechao Cheng, Jingfeng Li, Jingdong Wang 0001, Junwei Han 0001 |
CVPR | 2 |
| 2024 | VSCode: General Visual Salient and Camouflaged Object Detection with 2D Prompt LearningabstractSalient object detection (SOD) and camouflaged object detection (COD) are related yet distinct binary mapping tasks. These tasks involve multiple modalities, sharing commonalities and unique cues. Existing research often employs intricate task-specific specialist models, potentially leading to redundancy and suboptimal results. We introduce VS-Code, a generalist model with novel 2D prompt learning, to jointly address four SOD tasks and three COD tasks. We utilize VST as the foundation model and introduce 2D prompts within the encoder-decoder architecture to learn domain and task-specific knowledge on two separate dimensions. A prompt discrimination loss helps disentangle peculiarities to benefit model optimization. VSCode outperforms state-of-the-art methods across six tasks on 26 datasets and exhibits zero-shot generalization to unseen tasks by combining 2D prompts, such as RGB-D COD. Source code has been available at https://github.com/Sssssuperior/VSCode. Nian Liu 0002, Wangbo Zhao, Xuguang Yang, Dingwen Zhang, Deng-Ping Fan, Fahad Shahbaz Khan, Junwei Han 0001 |
CVPR | 5 |
| 2024 | GGRt: Towards Pose-Free Generalizable 3D Gaussian Splatting in Real-Time
Hao Li 0075, Chenming Wu, Dingwen Zhang, Yalun Dai, Chen Zhao 0011, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Junwei Han 0001 |
ECCV (71) | 4 |
| 2024 | CONDA: Condensed Deep Association Learning for Co-salient Object Detection
Long Li 0008, Nian Liu 0002, Dingwen Zhang, Zhongyu Li 0006, Salman Khan 0001, Rao Muhammad Anwer, Hisham Cholakkal, Junwei Han 0001, Fahad Shahbaz Khan |
ECCV (50) | 3 |
| 2024 | Gradient and Brightness Guided Low-Light Enhancement with Attention-Based Self-Paced LearningabstractLow-light image enhancement aims to reconstruct images with insufficient illumination into visually appealing representations with natural brightness. While most existing methods tend to focus on enhancing illumination, they often overlook the restoration of finer details in the enhanced image. Moreover, these methods do not adequately address the varying degradation levels observed in different regions of the image. In this study, we present a gradient and brightness guided low-light image enhancement framework, which can simultaneously augment the detail and illumination during the enhancement process. Our approach involves extracting gradient information from gamma-corrected images, which offers a remarkable advantage in preserving edge details compared to direct extraction from degraded images. To further refine the enhancement process and adaptively adjust the difficulty of samples, thereby boosting learning efficiency, we introduce an attention-based self-paced learning strategy. This strategy assigns different gradient and brightness weights based on the degradation levels within different image regions. Extensive experiments demonstrate the superiority of our proposed method over state-of-the-art approaches. The code is available at https://github.com/MSL502/GBASPL. Yan Li 0125, De Cheng, Dingwen Zhang, Luofeng Zhai, Jiande Sun 0001 |
ICASSP | 4 |
| 2024 | Task-aware Orthogonal Sparse Network for Exploring Shared Knowledge in Continual LearningabstractContinual learning (CL) aims to learn from sequentially arriving tasks without catastrophic forgetting (CF). By partitioning the network into two parts based on the Lottery Ticket Hypothesis—one for holding the knowledge of the old tasks while the other for learning the knowledge of the new task—the recent progress has achieved forget-free CL. Although addressing the CF issue well, such methods would encounter serious under-fitting in long-term CL, in which the learning process will continue for a long time and the number of new tasks involved will be much higher. To solve this problem, this paper partitions the network into three parts—with a new part for exploring the knowledge sharing between the old and new tasks. With the shared knowledge, this part of network can be learnt to simultaneously consolidate the old tasks and fit to the new task. To achieve this goal, we propose a task-aware Orthogonal Sparse Network (OSN), which contains shared knowledge induced network partition and sharpness-aware orthogonal sparse network learning. The former partitions the network to select shared parameters, while the latter guides the exploration of shared knowledge through shared parameters. Qualitative and quantitative analyses, show that the proposed OSN induces minimum to no interference with past tasks, i.e., approximately no forgetting, while greatly improves the model plasticity and capacity, and finally achieves the state-of-the-art performances. Yusong Hu, De Cheng, Dingwen Zhang, Nannan Wang 0001, Tongliang Liu, Xinbo Gao 0001 |
ICML | 3 |
| 2024 | Revisiting the Power of Prompt for Visual TuningabstractVisual prompt tuning (VPT) is a promising solution incorporating learnable prompt tokens to customize pre-trained models for downstream tasks. However, VPT and its variants often encounter challenges like prompt initialization, prompt length, and subpar performance in self-supervised pretraining, hindering successful contextual adaptation. This study commences by exploring the correlation evolvement between prompts and patch tokens during proficient training. Inspired by the observation that the prompt tokens tend to share high mutual information with patch tokens, we propose initializing prompts with downstream token prototypes. The strategic initialization, a stand-in for the previous initialization, substantially improves performance. To refine further, we optimize token construction with a streamlined pipeline that maintains excellent performance with almost no increase in computational expenses compared to VPT. Exhaustive experiments show our proposed approach outperforms existing methods by a remarkable margin. For instance, after MAE pre-training, our method improves accuracy by up to 10%$\sim$30% compared to VPT, and outperforms Full fine-tuning 19 out of 24 cases while using less than 0.4% of learnable parameters. Besides, the experimental results demonstrate the proposed SPT is robust to prompt lengths and scales well with model capacity and training data size. We finally provide an insightful exploration into the amount of target data facilitating the adaptation of pre-trained models to downstream tasks. The code is available at https://github.com/WangYZ1608/Self-Prompt-Tuning. Lechao Cheng, Chaowei Fang, Dingwen Zhang, Manni Duan, Meng Wang 0001 |
ICML | 4 |
| 2024 | ASPS: Augmented Segment Anything Model for Polyp Segmentation
Huiqian Li, Dingwen Zhang, Jieru Yao, Longfei Han, Zhongyu Li 0006, Junwei Han 0001 |
MICCAI (9) | 2 |
| 2024 | M-RRFS: A Memory-Based Robust Region Feature Synthesizer for Zero-Shot Object Detection
Peiliang Huang, Dingwen Zhang, De Cheng, Longfei Han, Pengfei Zhu 0001, Junwei Han 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Deep unsupervised part-whole relational visual saliency
Yi Liu 0038, Dingwen Zhang, Shoukun Xu |
Neurocomputing | 3 |
| 2024 | Contextual Dependency Vision Transformer for spectrogram-based multivariate time series analysis
Jieru Yao, Longfei Han, Kaihui Yang, Guangyu Guo 0001, Nian Liu 0002, Xiankai Huang, Zhaohui Zheng 0004, Dingwen Zhang, Junwei Han 0001 |
Neurocomputing | 8 |
| 2024 | Position-based anchor optimization for point supervised dense nuclei detection
Jieru Yao, Longfei Han, Guangyu Guo 0001, Zhaohui Zheng 0004, Runmin Cong, Xiankai Huang, Jin Ding, Kaihui Yang, Dingwen Zhang, Junwei Han 0001 |
Neural Networks | 9 |
| 2024 | Pixel Distillation: Cost-Flexible Distillation Across Image Sizes and Heterogeneous NetworksabstractPrevious knowledge distillation (KD) methods mostly focus on compressing network architectures, which is not thorough enough in deployment as some costs like transmission bandwidth and imaging equipment are related to the image size. Therefore, we propose Pixel Distillation that extends knowledge distillation into the input level while simultaneously breaking architecture constraints. Such a scheme can achieve flexible cost control for deployment, as it allows the system to adjust both network architecture and image quality according to the overall requirement of resources. Specifically, we first propose an input spatial representation distillation (ISRD) mechanism to transfer spatial knowledge from large images to student's input module, which can facilitate stable knowledge transfer between CNN and ViT. Then, a Teacher-Assistant-Student (TAS) framework is further established to disentangle pixel distillation into the model compression stage and input compression stage, which significantly reduces the overall complexity of pixel distillation and the difficulty of distilling intermediate knowledge. Finally, we adapt pixel distillation to object detection via an aligned feature for preservation (AFP) strategy for TAS, which aligns output dimensions of detectors at each stage by manipulating features and anchors of the assistant. Comprehensive experiments on image classification and object detection demonstrate the effectiveness of our method. Guangyu Guo 0001, Dingwen Zhang, Longfei Han, Nian Liu 0002, Ming-Ming Cheng, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Efficient Statistical Sampling Adaptation for Exemplar-Free Class Incremental LearningabstractDeep learning systems typically suffer from catastrophic forgetting of old knowledge when learning from new data continually. Recently, various class incremental learning (CIL) methods have been proposed to address this issue, and some approaches achieve promising performances by relying on rehearsing the training data of previous tasks. However, storing data from previous tasks would encounter data privacy and memory issues in real-world applications. In this paper, we propose a statistical sampling adaptation method for efficient Exemplar-Free Class-Incremental Learning (EFCIL). Here, instead of preserving the images/features themselves of previous tasks/classes, we store image feature statistics from previous classes to maintain the decision boundary, which is memory-efficient and much semantic-representative. When utilizing the old-class feature statistics, we build a statistical feature adaptation network (SFAN) with a manifold consistency regularization and then train it in a transductive learning paradigm, which can map the outdated statistics onto the current feature space to facilitate a compatible and balanced classifier training subsequently. In this way, the final classifier can be jointly optimized with all the old-class features projected by SFAN and current new-class features, thus alleviating the classification bias problem in EFCIL. Experimental results greatly demonstrate the effectiveness of the proposed method, achieving superior performances than state-of-the-art approaches. Our source code is released inhttps://github.com/yxzhcv/ESSA-EFCIL. De Cheng, Nannan Wang 0001, Guozhang Li, Dingwen Zhang, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Uncertainty Modeling for Gaze EstimationabstractGaze estimation is an important fundamental task in computer vision and medical research. Existing works have explored various effective paradigms and modules for precisely predicting eye gazes. However, the uncertainty for gaze estimation, e.g., input uncertainty and annotation uncertainty, have been neglected in previous research. Existing models use a deterministic function to estimate the gaze, which cannot reflect the actual situation in gaze estimation. To address this issue, we propose a probabilistic framework for gaze estimation by modeling the input uncertainty and annotation uncertainty. We first utilize probabilistic embeddings to model the input uncertainty, representing the input image as a Gaussian distribution in the embedding space. Based on the input uncertainty modeling, we give an instance-wise uncertainty estimation to measure the confidence of prediction results, which is critical in practical applications. Then, we propose a new label distribution learning method, probabilistic annotations, to model the annotation uncertainty, representing the raw hard labels as Gaussian distributions. In addition, we develop an Embedding Distribution Smoothing (EDS) module and a hard example mining method to improve the consistency between embedding distribution and label distribution. We conduct extensive experiments, demonstrating that the proposed approach achieves significant improvements over baseline and state-of-the-art methods on two widely used benchmark datasets, GazeCapture and MPIIFaceGaze, as well as our collected dataset using mobile devices. Wenqi Zhong, Chen Xia, Dingwen Zhang, Junwei Han 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Correction to "A Structure-Aware Relation Network for Thoracic Diseases Detection and Segmentation"abstractIn the above article[1], there are errors on pages 2045 and 2046. Section METHOD.D and Section METHOD.E should be the subsections of Section METHOD.C, i.e., METHOD.C: 1) Relation Graph Construction; 2) Message Passing via Relation Graph; and 3) Mapping Disease Relation to Regions. Jingyu Liu 0004, Shu Zhang 0001, Dingwen Zhang, Yizhou Yu |
IEEE Trans. Medical Imaging | 6 |
| 2024 | Deep Generative Adversarial Reinforcement Learning for Semi-Supervised Segmentation of Low-Contrast and Small Objects in Medical ImagesabstractDeep reinforcement learning (DRL) has demonstrated impressive performance in medical image segmentation, particularly for low-contrast and small medical objects. However, current DRL-based segmentation methods face limitations due to the optimization of error propagation in two separate stages and the need for a significant amount of labeled data. In this paper, we propose a novel deep generative adversarial reinforcement learning (DGARL) approach that, for the first time, enables end-to-end semi-supervised medical image segmentation in the DRL domain. DGARL ingeniously establishes a pipeline that integrates DRL and generative adversarial networks (GANs) to optimize both detection and segmentation tasks holistically while mutually enhancing each other. Specifically, DGARL introduces two innovative components to facilitate this integration in semi-supervised settings. First, a task-joint GAN with two discriminators links the detection results to the GAN's segmentation performance evaluation, allowing simultaneous joint evaluation and feedback. This ensures that DRL and GAN can be directly optimized based on each other's results. Second, a bidirectional exploration DRL integrates backward exploration and forward exploration to ensure the DRL agent explores the correct direction when forward exploration is disabled due to lack of explicit rewards. This mitigates the issue of unlabeled data being unable to provide rewards and rendering DRL unexplorable. Comprehensive experiments on three generalization datasets, comprising a total of 640 patients, demonstrate that our novel DGARL achieves 85.02% Dice and improves at least 1.91% for brain tumors, achieves 73.18% Dice and improves at least 4.28% for liver tumors, and achieves 70.85% Dice and improves at least 2.73% for pancreas compared to the ten most recent advanced methods, our results attest to the superiority of DGARL. Code is available at GitHub. Chenchu Xu, Dong Zhang 0009, Dingwen Zhang, Junwei Han 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Continual All-in-One Adverse Weather Removal With Knowledge Replay on a Unified Network StructureabstractIn real-world applications, image degeneration caused by adverse weather is always complex and changes with different weather conditions from days and seasons. Systems in real-world environments constantly encounter adverse weather conditions that are not previously observed. Therefore, it practically requires adverse weather removal models to continually learn from incrementally collected data reflecting various degeneration types. Existing adverse weather removal approaches, for either single or multiple adverse weathers, are mainly designed for a static learning paradigm, which assumes that the data of all types of degenerations to handle can be finely collected at one time before a single-phase learning process. They thus cannot directly handle the incremental learning requirements. To address this issue, we made the earliest effort to investigate the continual all-in-one adverse weather removal task, in a setting closer to real-world applications. Specifically, we develop a novel continual learning framework with effective knowledge replay (KR) on a unified network structure. Equipped with a principal component projection and an effective knowledge distillation mechanism, the proposed KR techniques are tailored for the all-in-one weather removal task. It considers the characteristics of the image restoration task with multiple degenerations in continual learning, and the knowledge for different degenerations can be shared and accumulated in the unified network structure. Extensive experimental results demonstrate the effectiveness of the proposed method to deal with this challenging task, which performs competitively to existing dedicated or joint training image restoration methods. Our code is available athttps://github.com/xiaojihh/CL_all-in-one. De Cheng, Yanling Ji, Dong Gong, Yan Li 0125, Nannan Wang 0001, Junwei Han 0001, Dingwen Zhang |
IEEE Trans. Multim. | 7 |
| 2024 | Progressive Negative Enhancing Contrastive Learning for Image Dehazing and BeyondabstractImage dehazing is a pivotal preliminary step in the advancement of robust intelligent surveillance system. However, it is an extremely challenging ill-posed problem, as it faces severe information degradation when accurately restoring the clean image from its haze-polluted counterpart. This paper proposes a novel Progressive Negative Enhancing (PNE) contrastive learning mechanism to fully exploit various types of negative information, thereby facilitating the traditional positive-oriented objective function for image dehazing. The proposed method can progressively update the negative samples during model training, to steadily squeeze the restored image towards its desired clean target from various directions. Furthermore, considering the image dehazing task as a many-to-one feature mapping problem, we also make an early effort to enhance the robustness of the dehazing model under variational haze densities. Specifically, a novel density-variational dehazing network is proposed to be optimized under the consistency-regularized framework using the proposed PNE learning mechanism. The consistency regularization ensures consistent output given multi-level degraded hazy images, thereby significantly enhancing the robustness of the model in dealing with various hazy scenarios. Extensive experiments demonstrate that the proposed method exhibits superior performance over existing state-of-the-art methods. It achieves average PSNR boosts of 0.60dB, 0.28dB and 0.82dB on dehazing, deraining and desnowing tasks, respectively. The source code is available athttps://github.com/YanLi-LY/PNE-Net. De Cheng, Yan Li 0125, Dingwen Zhang, Nannan Wang 0001, Jiande Sun 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Separating Noisy Samples From Tail Classes for Long-Tailed Image Classification With Label NoiseabstractMost existing methods that cope with noisy labels usually assume that the classwise data distributions are well balanced. They are difficult to deal with the practical scenarios where training samples have imbalanced distributions, since they are not able to differentiate noisy samples from tail classes' clean samples. This article makes an early effort to tackle the image classification task in which the provided labels are noisy and have a long-tailed distribution. To deal with this problem, we propose a new learning paradigm which can screen out noisy samples by matching between inferences on weak and strong data augmentations. A leave-noise-out regularization (LNOR) is further introduced to eliminate the effect of the recognized noisy samples. Besides, we propose a prediction penalty based on the online classwise confidence levels to avoid the bias toward easy classes which are dominated by head classes. Extensive experiments on five datasets including CIFAR-10, CIFAR-100, MNIST, FashionMNIST, and Clothing1M demonstrate that the proposed method outperforms the existing algorithms for learning with long-tailed distribution and label noise. Chaowei Fang, Lechao Cheng, Yining Mao, Dingwen Zhang, Yixiang Fang, Guanbin Li, Huiyan Qi, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Generalized Weakly Supervised Object LocalizationabstractWith the goal of learning to localize specific object semantics using the low-cost image-level annotation, weakly supervised object localization (WSOL) has been receiving increasing attention in recent years. Although existing literatures have studied a number of major issues in this field, one important yet challenging scenario, where the test object semantics may appear in the training phase (seen categories) or never been observed before (unseen categories), is still beyond the exploration of the existing works. We define this scenario as the generalized WSOL (GWSOL) and make a pioneering effort to study it in this article. By leveraging attribute vectors to associate seen and unseen categories, we involve threefold modeling components, i.e., the class-sensitive modeling, semantic-agnostic modeling, and content-aware modeling, into a unified end-to-end learning framework. Such design enables our model to recognize and localize unconstrained object semantics, learn compact and discriminative features that could represent the potential unseen categories, and customize content-aware attribute weights to avoid localizing on misleading attribute elements. To advance this research direction, we contribute the bounding-box manual annotations to the widely used AwA2 dataset and benchmark the GWSOL methods. Comprehensive experiments demonstrate the effectiveness of our proposed learning framework and each of the considered modeling components. Dingwen Zhang, Guangyu Guo 0001, Wenyuan Zeng, Junwei Han 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Progressive Adapting and Pruning: Domain-Incremental Learning for Saliency PredictionabstractSaliency prediction (SAP) plays a crucial role in simulating the visual perception function of human beings. In practical situations, humans can quickly grasp saliency extraction in new image domains. However, current SAP methods mainly concentrate on training models in single domains, which do not effectively handle diverse content and styles present in real-world images. As a result, it would be of great significance if SAP models could efficiently adjust to new image domains. To this end, this article aims to design SAP models that can imitate the incremental learning ability of human beings on multiple image domains and name domain-incremental saliency prediction (DISAP). To make a tradeoff between preventing the forgetting of historical domains and achieving high performance on new domains, we propose a progressively updated domain incremental encoder. This encoder consists of a domain-sharing branch and a domain-specific branch. The domain-sharing branch includes a feature selection mechanism to preserve crucial parameters after fine-tuning the model on each current domain. The remaining parameters are reserved to absorb knowledge from future domains. Furthermore, to capture the unique characteristics of each domain with relatively low computational overhead, we introduce a lightweight design to construct the domain-specific branch, enabling effective adaptation to new domains. Extensive experiments are conducted on multiple domain-incremental learning settings formed by four saliency prediction datasets, including Salicon, MIT1003, the art subset of CAT2000, and WebSal. The results demonstrate that our method outperforms existing methods significantly. The code is available at https://github.com/KaIi-github/DIL4SAP . Kaihui Yang, Junwei Han 0001, Guangyu Guo 0001, Chaowei Fang, Yingzi Fan, Lechao Cheng, Dingwen Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2023 | Boosting Low-Data Instance Segmentation by Unsupervised Pre-training with Saliency PromptabstractInspired by DETR variants, query-based end-to-end instance segmentation (QEIS) methods have recently outperformed CNN-based models on large-scale datasets. Yet they would lose efficacy when only a small amount of training data is available since it's hard for the crucial queries/kernels to learn localization and shape priors. To this end, this work offers a novel unsupervised pre-training solution for low-data regimes. Inspired by the recent success of the Prompting technique, we introduce a new pre-training method that boosts QEIS models by giving Saliency Prompt for queries/kernels. Our method contains three parts: 1) Saliency Masks Proposal is responsible for generating pseudo masks from unlabeled images based on the saliency mechanism. 2) Prompt-Kernel Matching transfers pseudo masks into prompts and injects the corresponding localization and shape priors to the best-matched kernels. 3) Kernel Supervision is applied to supply supervision at the kernel level for robust learning. From a practical perspective, our pre-training method helps QEIS models achieve a similar convergence speed and comparable performance with CNN-based models in low-data regimes. Experimental results show that our method significantly boosts several QEIS models on three datasets.11Code: https://github.com/lifuguan/saliency.prompt Hao Li 0075, Dingwen Zhang, Nian Liu 0002, Lechao Cheng, Yalun Dai, Xinggang Wang, Junwei Han 0001 |
CVPR | 2 |
| 2023 | Giving Text More Imagination Space for Image-text MatchingabstractImage-text matching is a hot topic in multi-modal analysis. The existing image-text matching algorithms focus on bridging the heterogeneity gap and mapping the feature into a common space under strong alignment assumption. However, these methods have unsatisfactory performance under the weak alignment scenario, which assumes that the text contains more abstract information, and the number of entities in the text is always fewer than objects in image. This is the first time, from our knowledge, to solve the image-text matching problem from the perspective of information difference with weak alignment. In order to both narrow the cross-modal heterogeneity gap and balance the information discrepancy, we proposed an imagination network to enrich the text modality based on pre-trained framework, which is helpful for image-text matching. The imagination network utilizes reinforcement learning to enhance the semantic information for text modality, and an action refinement strategy is designed to constrain the freedom and divergence of imagination. The experiment results show the superiority and generality of the proposed framework based on two pre-trained models, CLIP and BLIP on two most frequently-used datasets MSCOCO and Flickr30K. Longfei Han, Dingwen Zhang, Li Liu 0031, Junwei Han 0001, Huaxiang Zhang 0001 |
ACM Multimedia | 3 |
| 2023 | HCMA '23: 4th International Workshop on Human-Centric Multimedia AnalysisabstractUnderstanding human interactions within diverse media contexts has emerged as a fundamental challenge. The explosive growth of multimedia data not only provides opportunities for human-centirc analysis but also increases the complexity of processing multimodal data. To address this pivotal challenge and explore its multifaceted dimensions, the Fourth International Workshop on Human-Centric Multimedia Analysis is concentrated on the tasks of human-centric analysis with multimedia and multimodal information. By delving into the nuances of human behavior within multimedia, this workshop aims to uncover novel insights, showcase innovative methodologies, and discuss future directions. With a spotlight on cutting-edge research and a focus on real-world applications, the workshop seeks to equip researchers and practitioners with the tools and knowledge to navigate the intricacies of human-centric multimedia analysis. Jingkuan Song, Wu Liu 0005, Xinchen Liu, Dingwen Zhang, Chaowei Fang, Hongyuan Zhu 0002, Wenbing Huang 0001, John R. Smith, Xin Wang 0019 |
ACM Multimedia | 4 |
| 2023 | Equivalent Classification Mapping for Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization is a newly emerging yet widely studied topic in recent years. The existing methods can be categorized into two localization-by-classification pipelines, i.e., the pre-classification pipeline and the post-classification pipeline. The pre-classification pipeline first performs classification on each video snippet, and then, aggregates the snippet-level classification scores to obtain the video-level classification score. In contrast, the post-classification pipeline aggregates the snippet-level features first and then predicts the video-level classification score based on the aggregated feature. Although the classifiers in these two pipelines are used in different ways, the role they play is exactly the same-to classify the given features to identify the corresponding action categories. To this end, an ideal classifier can make both pipelines work. This inspires us to simultaneously learn these two pipelines in a unified framework to obtain an effective classifier. Specifically, in the proposed learning framework, we implement two parallel network streams to model the two localization-by-classification pipelines simultaneously and make the two network streams share the same classifier. This achieves the novel Equivalent Classification Mapping (ECM) mechanism. Moreover, we discover that an ideal classifier may possess two characteristics: 1) the frame-level classification scores obtained from the pre-classification stream and the feature aggregation weights in the post-classification stream should be consistent; and 2) the classification results of these two streams should be identical. Based on these two characteristics, we further introduce a weight-transition module and an equivalent training strategy into the proposed learning framework, which assists to thoroughly mine the equivalence mechanism. Comprehensive experiments are conducted on three benchmarks and ECM achieves accurate action localization results. Tao Zhao 0006, Junwei Han 0001, Le Yang 0008, Dingwen Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Salient Object Detection via Integrity LearningabstractAlthough current salient object detection (SOD) works have achieved significant progress, they are limited when it comes to the integrity of the predicted salient regions. We define the concept of integrity at both a micro and macro level. Specifically, at the micro level, the model should highlight all parts that belong to a certain salient object. Meanwhile, at the macro level, the model needs to discover all salient objects in a given image. To facilitate integrity learning for SOD, we design a novel Integrity Cognition Network (ICON), which explores three important components for learning strong integrity features. 1) Unlike existing models, which focus more on feature discriminability, we introduce a diverse feature aggregation (DFA) component to aggregate features with various receptive fields (i.e., kernel shape and context) and increase feature diversity. Such diversity is the foundation for mining the integral salient objects. 2) Based on the DFA features, we introduce an integrity channel enhancement (ICE) component with the goal of enhancing feature channels that highlight the integral salient objects, while suppressing the other distracting ones. 3) After extracting the enhanced features, the part-whole verification (PWV) method is employed to determine whether the part and whole object features have strong agreement. Such part-whole agreements can further improve the micro-level integrity for each salient object. To demonstrate the effectiveness of our ICON, comprehensive experiments are conducted on seven challenging benchmarks. Our ICON outperforms the baseline methods in terms of a wide range of metrics. Notably, our ICON achieves ∼ 10% relative improvement over the previous best model in terms of average false negative ratio (FNR), on six datasets. Codes and results are available at: https://github.com/mczhuge/ICON. Mingchen Zhuge, Deng-Ping Fan, Nian Liu 0002, Dingwen Zhang, Dong Xu 0001, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Hybrid routing transformer for zero-shot learning
De Cheng, Gerong Wang, Bo Wang 0011, Qiang Zhang 0020, Jungong Han, Dingwen Zhang |
Pattern Recognit. | 6 |
| 2023 | Discriminative and Robust Attribute Alignment for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to learn models that can recognize images of semantically related unseen categories, through transferring attribute-based knowledge learned from training data of seen classes to unseen testing data. As visual attributes play a vital role in ZSL, recent embedding-based methods usually focus on learning a compatibility function between the visual representation and the class semantic attributes. While in this work, in addition to simply learning the region embedding of different semantic attributes to maintain the generalization capability of the learned model, we further consider to improve the discrimination power of the learned visual features themselves by contrastive embedding. It exploits both the class-wise and instance-wise supervision for GZSL, under the attribute guided weakly supervised representation learning framework. To further improve the robustness of the ZSL model, we also propose to train the model under the consistency regularization constraint, through taking full advantages of self-supervised signals of the image under various perturbed augmentation situations, which could make the model robust to some occluded or un-related attribute regions. Extensive experimental results demonstrate the effectiveness of the proposed ZSL method, achieving superior performances to state-of-the-art methods on three widely-used benchmark datasets, namely CUB, SUN, and AWA2. Our source code is released athttps://github.com/KORIYN/CC-ZSL. De Cheng, Gerong Wang, Nannan Wang 0001, Dingwen Zhang, Qiang Zhang 0020, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Reliable Mutual Distillation for Medical Image Segmentation Under Imperfect AnnotationsabstractConvolutional neural networks (CNNs) have made enormous progress in medical image segmentation. The learning of CNNs is dependent on a large amount of training data with fine annotations. The workload of data labeling can be significantly relieved via collecting imperfect annotations which only match the underlying ground truths coarsely. However, label noises which are systematically introduced by the annotation protocols, severely hinders the learning of CNN-based segmentation models. Hence, we devise a novel collaborative learning framework in which two segmentation models cooperate to combat label noises in coarse annotations. First, the complementary knowledge of two models is explored by making one model clean training data for the other model. Secondly, to further alleviate the negative impact of label noises and make sufficient usage of the training data, the specific reliable knowledge of each model is distilled into the other model with augmentation-based consistency constraints. A reliability-aware sample selection strategy is incorporated for guaranteeing the quality of the distilled knowledge. Moreover, we employ joint data and model augmentations to expand the usage of reliable knowledge. Extensive experiments on two benchmarks showcase the superiority of our proposed method against existing methods under annotations with different noise levels. For example, our approach can improve existing methods by nearly 3% DSC on the lung lesion segmentation dataset LIDC-IDRI under annotations with 80% noise ratio. Code is available at: https://github.com/Amber-Believe/ReliableMutualDistillation. Chaowei Fang, Lechao Cheng, Zhifan Gao, Chengwei Pan, Zhaohui Zheng 0004, Dingwen Zhang |
IEEE Trans. Medical Imaging | 8 |
| 2023 | CLRNet: Component-Level Refinement Network for Deep Face ParsingabstractFace parsing aims to assign pixel-wise semantic labels to different facial components (e.g., hair, brows, and lips) in given face images. However, directly predicting pixel-level labels for each facial component over the whole face image would obtain limited accuracy, especially for tiny facial components. To address this problem, some recent works propose to first crop tiny patches from the whole face image and then predict masks for each facial component. However, such cropping-and-segmenting strategy consists of two independent stages, which cannot be jointly optimized. Besides, as one valuable piece of information for parsing the highly structured facial components, context cues are not elaborately explored by the existing works. To address these issues, we propose a component-level refinement network (CLRNet) for precisely segmenting out each facial component. Specifically, we introduce an attention mechanism to bridge the two independent stages together and form an end-to-end trainable pipeline for face parsing. Furthermore, we incorporate the global context information into the refining process for each cropped facial component patch, providing informative cues for accurate parsing. Extensive experiments are carried out on two benchmark datasets, LFW-PL and HELEN. The results demonstrate the superiority of the proposed CLRNet over other state-of-the-art methods, especially for tiny facial components. Peiliang Huang, Junwei Han 0001, Dingwen Zhang, Mingliang Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Deep 3D Vessel Segmentation based on Cross Transformer NetworkabstractThe coronary microvascular disease poses a great threat to human health. Computer-aided analysis/diagnosis systems help physicians intervene in the disease at early stages, where 3D vessel segmentation is a fundamental step. However, there is a lack of carefully annotated dataset to support algorithm development and evaluation. On the other hand, the commonly-used U-Net structures often yield disconnected and inaccurate segmentation results, especially for small vessel structures. In this paper, motivated by the data scarcity, we first construct two large-scale vessel segmentation datasets consisting of 100 and 500 computed tomography (CT) volumes with pixel-level annotations by experienced radiologists. To enhance the U-Net, we further propose the cross transformer network (CTN) for fine-grained vessel segmentation. In CTN, a transformer module is constructed in parallel to a U-Net to learn long-distance dependencies between different anatomical regions; and these dependencies are communicated to the U-Net at multiple stages to endow it with global awareness. Experimental results on the two in-house datasets indicate that this hybrid model alleviates unexpected disconnections by considering topological information across regions. Our codes, together with the trained models are made publicly available at https://github.com/qibaolian/ctn. Chengwei Pan, Baolian Qi, Gangming Zhao, Chaowei Fang, Dingwen Zhang, Jinpeng Li 0002 |
BIBM | 6 |
| 2022 | Incremental Cross-view Mutual Distillation for Self-supervised Medical CT SynthesisabstractDue to the constraints of the imaging device and high cost in operation time, computer tomography (CT) scans are usually acquired with low within-slice resolution. Improving the inter-slice resolution is beneficial to the disease diagnosis for both human experts and computer-aided systems. To this end, this paper builds a novel medical slice synthesis to increase the inter-slice resolution. Considering that the groundtruth intermediate medical slices are always absent in clinical practice, we introduce the incremental cross-view mutual distillation strategy to accomplish this task in the self-supervised learning manner. Specifically, we model this problem from three different views: slice-wise interpolation from axial view and pixel-wise interpolation from coronal and sagittal views. Under this circumstance, the models learned from different views can distill valuable knowledge to guide the learning processes of each other. We can repeat this process to make the models synthesize intermediate slice data with increasing between-slice resolution. To demonstrate the effectiveness of the proposed approach, we conduct comprehensive experiments on a large-scale$CT$dataset. Quantitative and qualitative comparison results show that our method outperforms state-of-the-art algorithms by clear margins. Chaowei Fang, Liang Wang 0001, Dingwen Zhang, Jun Xu 0019, Yixuan Yuan, Junwei Han 0001 |
CVPR | 3 |
| 2022 | Robust Region Feature Synthesizer for Zero-Shot Object DetectionabstractZero-shot object detection aims at incorporating class semantic vectors to realize the detection of (both seen and) unseen classes given an unconstrained test image. In this study, we reveal the core challenges in this research area: how to synthesize robust region features (for unseen objects) that are as intra-class diverse and inter-class separable as the real samples, so that strong unseen object detectors can be trained upon them. To address these challenges, we build a novel zero-shot object detection framework that contains an Intra-class Semantic Diverging component and an Inter-class Structure Preserving component. The former is used to realize the one-to-more mapping to obtain diverse visual features from each class semantic vector, preventing miss-classifying the real unseen objects as image backgrounds. While the latter is used to avoid the synthesized features too scattered to mix up the inter-class and foreground-background relationship. To demonstrate the effectiveness of the proposed approach, comprehensive experiments on PASCAL VOC, COCO, and DIOR datasets are conducted. Notably, our approach achieves the new state-of-the-art performance on PASCAL VOC and COCO and it is the first study to carry out zero-shot object detection in remote sensing imagery. Peiliang Huang, Junwei Han 0001, De Cheng, Dingwen Zhang |
CVPR | 4 |
| 2022 | Colar: Effective and Efficient Online Action Detection by Consulting ExemplarsabstractOnline action detection has attracted increasing research interests in recent years. Current works model historical dependencies and anticipate the future to perceive the action evolution within a video segment and improve the detection accuracy. However, the existing paradigm ignores category-level modeling and does not pay sufficient attention to efficiency. Considering a category, its representative frames exhibit various characteristics. Thus, the category-level modeling can provide complimentary guidance to the temporal dependencies modeling. This paper develops an effective exemplar-consultation mechanism that first measures the similarity between a frame and exemplary frames, and then aggregates exemplary features based on the similarity weights. This is also an efficient mechanism, as both similarity measurement and feature aggregation require limited computations. Based on the exemplar-consultation mechanism, the long-term dependencies can be captured by regarding historical frames as exemplars, while the category-level modeling can be achieved by regarding representative frames from a category as exemplars. Due to the complementarity from the categorylevel modeling, our method employs a lightweight architecture but achieves new high performance on three benchmarks. In addition, using a spatio-temporal network to tackle video frames, our method makes a good trade-off between effectiveness and efficiency. Code is available at https://github.com/VividLe/Online-Action-Detection. Le Yang 0008, Junwei Han 0001, Dingwen Zhang |
CVPR | 3 |
| 2022 | Robust Single Image Dehazing Based on Consistent and Contrast-Assisted ReconstructionabstractSingle image dehazing as a fundamental low-level vision task, is essential for the development of robust intelligent surveillance system. In this paper, we make an early effort to consider dehazing robustness under variational haze density, which is a realistic while under-studied problem in the research filed of singe image dehazing. To properly address this problem, we propose a novel density-variational learning framework to improve the robustness of the image dehzing model assisted by a variety of negative hazy images, to better deal with various complex hazy scenarios. Specifically, the dehazing network is optimized under the consistency-regularized framework with the proposed Contrast-Assisted Reconstruction Loss (CARL). The CARL can fully exploit the negative information to facilitate the traditional positive-orient dehazing objective function, by squeezing the dehazed image to its clean target from different directions. Meanwhile, the consistency regularization keeps consistent outputs given multi-level hazy images, thus improving the model robustness. Extensive experimental results on two synthetic and three real-world datasets demonstrate that our method significantly surpasses the state-of-the-art approaches. De Cheng, Yan Li 0125, Dingwen Zhang, Nannan Wang 0001, Xinbo Gao 0001, Jiande Sun 0001 |
IJCAI | 3 |
| 2022 | Computer-Aided Tuberculosis Diagnosis with Attribute Reasoning Assistance
Chengwei Pan, Gangming Zhao, Junjie Fang, Baolian Qi, Chaowei Fang, Dingwen Zhang, Jinpeng Li 0002, Yizhou Yu |
MICCAI (1) | 7 |
| 2022 | Compound Batch Normalization for Long-tailed Image ClassificationabstractSignificant progress has been made in learning image classification neural networks under long-tail data distribution using robust training algorithms such as data re-sampling, re-weighting, and margin adjustment. Those methods, however, ignore the impact of data imbalance on feature normalization. The dominance of majority classes (head classes) in estimating statistics and affine parameters causes internal covariate shifts within less-frequent categories to be overlooked. To alleviate this challenge, we propose a compound batch normalization method based on a Gaussian mixture. It can model the feature space more comprehensively and reduce the dominance of head classes. In addition, a moving average-based expectation maximization (EM) algorithm is employed to estimate the statistical parameters of multiple Gaussian distributions. However, the EM algorithm is sensitive to initialization and can easily become stuck in local minima where the multiple Gaussian components continue to focus on majority classes. To tackle this issue, we developed a dual-path learning framework that employs class-aware split feature normalization to diversify the estimated Gaussian distributions, allowing the Gaussian components to fit with training samples of less-frequent classes more comprehensively. Extensive experiments on commonly used datasets demonstrated that the proposed method outperforms existing methods on long-tailed image classification. Lechao Cheng, Chaowei Fang, Dingwen Zhang, Guanbin Li, Gang Huang 0004 |
ACM Multimedia | 3 |
| 2022 | Cross-Modality High-Frequency Transformer for MR Image Super-ResolutionabstractImproving the resolution of magnetic resonance (MR) image data is critical to computer-aided diagnosis and brain function analysis. Higher resolution helps to capture more detailed content, but typically induces to lower signal-to-noise ratio and longer scanning time. To this end, MR image super-resolution has become a widely-interested topic in recent times. Existing works establish extensive deep models with the conventional architectures based on convolutional neural networks (CNN). In this work, to further advance this research field, we make an early effort to build a Transformer-based MR image super-resolution framework, with careful designs on exploring valuable domain prior knowledge. Specifically, we consider two-fold domain priors including the high-frequency structure prior and the inter-modality context prior, and establish a novel Transformer architecture, called Cross-modality high-frequency Transformer (Cohf-T), to introduce such priors into super-resolving the low-resolution (LR) MR images. Experiments on two datasets indicate that Cohf-T achieves new state-of-the-art performance. Chaowei Fang, Dingwen Zhang, Liang Wang 0001, Yulun Zhang 0001, Lechao Cheng, Junwei Han 0001 |
ACM Multimedia | 2 |
| 2022 | HCMA'22: 3rd International Workshop on Human-Centric Multimedia AnalysisabstractThe Third International Workshop on Human-Centric Multimedia Analysis concentrates on the tasks of human-centric analysis with multimedia and multimodal information. It involves multiple tasks such as face detection and recognition, human body pattern analysis, person re-identification, human action detection, etc. Today, multiple multimedia sensing technologies and large-scale computing infrastructures are emerging at a rapid velocity a wide variety of big multi-modality data for human-centric analysis, which provides rich knowledge to help tackle these challenges. Researchers have strived to push the limits of human-centric multimedia analysis in a wide variety of applications, such as intelligent surveillance, retailing, fashion design, and services. Therefore, this workshop aims to provide a platform to bridge the gap between the communities of human analysis and multimedia. Dingwen Zhang, Chaowei Fang, Wu Liu 0005, Xinchen Liu, Jingkuan Song, Hongyuan Zhu 0002, Wenbing Huang 0001, John R. Smith |
ACM Multimedia | 1 |
| 2022 | Densely nested top-down flows for salient object detection
Chaowei Fang, Haibin Tian, Dingwen Zhang, Qiang Zhang 0020, Jungong Han, Junwei Han 0001 |
Sci. China Inf. Sci. | 3 |
| 2022 | Onfocus detection: identifying individual-camera eye contact from unconstrained imagesabstractAbstract Onfocus detection aims at identifying whether the focus of the individual captured by a camera is on the camera or not. Based on the behavioral research, the focus of an individual during face-to-camera communication leads to a special type of eye contact, i.e., the individual-camera eye contact, which is a powerful signal in social communication and plays a crucial role in recognizing irregular individual status (e.g., lying or suffering mental disease) and special purposes (e.g., seeking help or attracting fans). Thus, developing effective onfocus detection algorithms is of significance for assisting the criminal investigation, disease discovery, and social behavior analysis. However, the review of the literature shows that very few efforts have been made toward the development of onfocus detector owing to the lack of large-scale public available datasets as well as the challenging nature of this task. To this end, this paper engages in the onfocus detection research by addressing the above two issues. Firstly, we build a large-scale onfocus detection dataset, named as the onfocus detection in the wild (OFDIW). It consists of 20623 images in unconstrained capture conditions (thus called “in the wild”) and contains individuals with diverse emotions, ages, facial characteristics, and rich interactions with surrounding objects and background scenes. On top of that, we propose a novel end-to-end deep model, i.e., the eye-context interaction inferring network (ECIIN), for onfocus detection, which explores eye-context interaction via dynamic capsule routing. Finally, comprehensive experiments are conducted on the proposed OFDIW dataset to benchmark the existing learning models and demonstrate the effectiveness of the proposed ECIIN. Dingwen Zhang, Bo Wang 0011, Gerong Wang, Qiang Zhang 0020, Jungong Han, Zheng You |
Sci. China Inf. Sci. | 1 |
| 2022 | Learning Self-supervised Low-Rank Network for Single-Stage Weakly and Semi-supervised Semantic Segmentation
Junwen Pan, Pengfei Zhu 0001, Kaihua Zhang 0001, Bing Cao 0002, Yu Wang 0106, Dingwen Zhang, Junwei Han 0001, Qinghua Hu |
Int. J. Comput. Vis. | 6 |
| 2022 | Single image dehazing with an independent Detail-Recovery Network
Yan Li 0125, De Cheng, Dingwen Zhang, Nannan Wang 0001, Xinbo Gao 0001, Jiande Sun 0001 |
Knowl. Based Syst. | 3 |
| 2022 | Re-Thinking Co-Salient Object DetectionabstractIn this article, we conduct a comprehensive study on the co-salient object detection (CoSOD) problem for images. CoSOD is an emerging and rapidly growing extension of salient object detection (SOD), which aims to detect the co-occurring salient objects in a group of images. However, existing CoSOD datasets often have a serious data bias, assuming that each group of images contains salient objects of similar visual appearances. This bias can lead to the ideal settings and effectiveness of models trained on existing datasets, being impaired in real-life situations, where similarities are usually semantic or conceptual. To tackle this issue, we first introduce a new benchmark, called CoSOD3k in the wild, which requires a large amount of semantic context, making it more challenging than existing CoSOD datasets. Our CoSOD3k consists of 3,316 high-quality, elaborately selected images divided into 160 groups with hierarchical annotations. The images span a wide range of categories, shapes, object sizes, and backgrounds. Second, we integrate the existing SOD techniques to build a unified, trainable CoSOD framework, which is long overdue in this field. Specifically, we propose a novel CoEG-Net that augments our prior model EGNet with a co-attention projection strategy to enable fast common information learning. CoEG-Net fully leverages previous large-scale SOD datasets and significantly improves the model scalability and stability. Third, we comprehensively summarize 40 cutting-edge algorithms, benchmarking 18 of them over three challenging CoSOD datasets (iCoSeg, CoSal2015, and our CoSOD3k), and reporting more detailed (i.e., group-level) performance analysis. Finally, we discuss the challenges and future works of CoSOD. We hope that our study will give a strong boost to growth in the CoSOD community. The benchmark toolbox and results are available on our project page at https://dpfan.net/CoSOD3K. Deng-Ping Fan, Tengpeng Li, Zheng Lin 0005, Ge-Peng Ji, Dingwen Zhang, Ming-Ming Cheng, Huazhu Fu, Jianbing Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Part-Object Relational Visual SaliencyabstractRecent years have witnessed a big leap in automatic visual saliency detection attributed to advances in deep learning, especially Convolutional Neural Networks (CNNs). However, inferring the saliency of each image part separately, as was adopted by most CNNs methods, inevitably leads to an incomplete segmentation of the salient object. In this paper, we describe how to use the property of part-object relations endowed by the Capsule Network (CapsNet) to solve the problems that fundamentally hinge on relational inference for visual saliency detection. Concretely, we put in place a two-stream strategy, termed Two-Stream Part-Object RelaTional Network (TSPORTNet), to implement CapsNet, aiming to reduce both the network complexity and the possible redundancy during capsule routing. Additionally, taking into account the correlations of capsule types from the preceding training images, a correlation-aware capsule routing algorithm is developed for more accurate capsule assignments at the training stage, which also speeds up the training dramatically. By exploring part-object relationships, TSPORTNet produces a capsule wholeness map, which in turn aids multi-level features in generating the final saliency map. Experimental results on five widely-used benchmarks show that our framework consistently achieves state-of-the-art performance. The code can be found on https://github.com/liuyi1989/TSPORTNet. Yi Liu 0038, Dingwen Zhang, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Background-Click Supervision for Temporal Action LocalizationabstractWeakly supervised temporal action localization aims at learning the instance-level action pattern from the video-level labels, where a significant challenge is action-context confusion. To overcome this challenge, one recent work builds an action-click supervision framework. It requires similar annotation costs but can steadily improve the localization performance when compared to the conventional weakly supervised methods. In this paper, by revealing that the performance bottleneck of the existing approaches mainly comes from the background errors, we find that a stronger action localizer can be trained with labels on the background video frames rather than those on the action frames. To this end, we convert the action-click supervision to the background-click supervision and develop a novel method, called BackTAL. Specifically, BackTAL implements two-fold modeling on the background video frames, i.e., the position modeling and the feature modeling. In position modeling, we not only conduct supervised learning on the annotated video frames but also design a score separation module to enlarge the score differences between the potential action frames and backgrounds. In feature modeling, we propose an affinity module to measure frame-specific similarities among neighboring frames and dynamically attend to informative neighbors when calculating temporal convolution. Extensive experiments on three benchmarks are conducted, which demonstrate the high performance of the established BackTAL and the rationality of the proposed background-click supervision. Le Yang 0008, Junwei Han 0001, Tao Zhao 0006, Dingwen Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Weakly Supervised Object Localization and Detection: A SurveyabstractAs an emerging and challenging problem in the computer vision community, weakly supervised object localization and detection plays an important role for developing new generation computer vision systems and has received significant attention in the past decade. As methods have been proposed, a comprehensive survey of these topics is of great importance. In this work, we review (1) classic models, (2) approaches with feature representations from off-the-shelf deep networks, (3) approaches solely based on deep learning, and (4) publicly available datasets and standard evaluation metrics that are widely used in this field. We also discuss the key challenges in this field, development history of this field, advantages/disadvantages of the methods in each category, the relationships between methods in different categories, applications of the weakly supervised object localization and detection methods, and potential future directions to further promote the development of this research field. Dingwen Zhang, Junwei Han 0001, Gong Cheng 0003, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Weakly Supervised Object Detection Using Proposal- and Semantic-Level RelationshipsabstractIn recent years, weakly supervised object detection has attracted great attention in the computer vision community. Although numerous deep learning-based approaches have been proposed in the past few years, such an ill-posed problem is still challenging and the learning performance is still behind the expectation. In fact, most of the existing approaches only consider the visual appearance of each proposal region but ignore to make use of the helpful context information. To this end, this paper introduces two levels of context into the weakly supervised learning framework. The first one is the proposal-level context, i.e., the relationship of the spatially adjacent proposals. The second one is the semantic-level context, i.e., the relationship of the co-occurring object categories. Therefore, the proposed weakly supervised learning framework contains not only the cognition process on the visual appearance but also the reasoning process on the proposal- and semantic-level relationships, which leads to the novel deep multiple instance reasoning framework. Specifically, built upon a conventional CNN-based network architecture, the proposed framework is equipped with two additional graph convolutional network-based reasoning models to implement object location reasoning and multi-label reasoning within an end-to-end network training procedure. Comprehensive experiments on the widely used PASCAL VOC and MS COCO benchmarks have been implemented, which demonstrate the superior capacity of the proposed approach when compared with other state-of-the-art methods and baseline models. Dingwen Zhang, Wenyuan Zeng, Jieru Yao, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Exploring rich intermediate representations for reconstructing 3D shapes from 2D images
Yang Yang 0009, Junwei Han 0001, Dingwen Zhang, Qi Tian 0001 |
Pattern Recognit. | 3 |
| 2022 | Guest Editorial Introduction to the Special Issue on Advanced Machine Learning Methodologies for Large-Scale Video Object Segmentation and DetectionabstractVideo object segmentation and detection are two important tasks toward intelligent video content understanding. Due to their wide applications in real-world vision tasks, such as video surveillance and automatic driving, they have recently attracted great attention in the computer vision and multimedia processing communities. Although numerous deep learning-based approaches have been proposed to solve these problems, implementing effective and efficient video object segmentation and detection is still very challenging for now, and the principles of solutions to address the problems are still understudied. On the one hand, the features learned by the current deep models are not strong enough to capture the rich spatial and temporal information from the input videos. On the other hand, the annotation information in the video data (especially for unconstrained online videos) is usually insufficient, unspecific, or even absent, thus challenging the current mainstream learning schemes. Dingwen Zhang, Seyed Hamid Rezatofighi, Junwei Han 0001, Nicu Sebe |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Adversarial Prototype Learning for Hyperspectral Image ClassificationabstractIn hyperspectral image (HSI) classification, the training set often contains a very limited number of high-dimensional samples, which can cause overfitting problems, especially in deep learning (DL) frameworks. This situation worsens when a bias exists between the feature distributions of the training and testing sets. In this article, we propose a novel method, referred to as adversarial prototype learning (APL), for learning an accurate HSI classification model in a uniform manner when the training set contains few, high-dimensional, and biased samples. APL consists of a prototype learning module (PLM) and an adversarial alignment module (AAM). The PLM aims to alleviate overfitting by training prototypical classifiers with a simple inductive bias in the initial feature space. The AAM aims to reduce the bias between the feature distributions of the training and testing sets using two adversarial prototypical classifiers learned by the PLM. Iteratively training the PLM and AAM results in alignment of the feature distributions between the training and testing sets while improving the generalization ability of the prototypical classifiers. The theoretical analysis indicates that APL is able to lower the upper error bound when classifying testing samples. We further apply APL in a DL framework to establish the adversarial prototypical network (APNet) architecture. Experimental results on four publicly available HSI datasets demonstrate that the proposed APNet alleviates overfitting, aligns the feature distributions between the training and testing sets, and achieves state-of-the-art performance. Shuai Wang 0059, Bo Du 0001, Dingwen Zhang, Fang Wan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Disentangled Capsule Routing for Fast Part-Object Relational SaliencyabstractRecently, the Part-Object Relational (POR) saliency underpinned by the Capsule Network (CapsNet) has been demonstrated to be an effective modeling mechanism to improve the saliency detection accuracy. However, it is widely known that the current capsule routing operations have huge computational complexity, which seriously limited the usability of the POR saliency models in real-time applications. To this end, this paper takes an early step towards a fast POR saliency inference by proposing a novel disentangled part-object relational network. Concretely, we disentangle horizontal routing and vertical routing from the original omnidirectional capsule routing, thus generating Disentangled Capsule Routing (DCR). This mechanism enjoys two advantages. On one hand, DCR that disentangles orthogonal 1D (i.e., vertical and horizontal) routing greatly reduces parameters and routing complexity, resulting in much faster inference than omnidirectional 2D routing adopted by existing CapsNets. On the other hand, thanks to the light POR cues explored by DCR, we could conveniently integrate the part-object routing process to different feature layers in CNN, rather than just applying it to the small-scaled one as in previous works. This helps to increase saliency inference accuracy. Compared to previous POR saliency detectors, DPORTNet infers visual saliency (5 ∼ 9 ) × faster, and is more accurate. DPORTNet is available under the open-source license at https://github.com/liuyi1989/DCR. Yi Liu 0038, Dingwen Zhang, Nian Liu 0002, Shoukun Xu, Jungong Han |
IEEE Trans. Image Process. | 2 |
| 2022 | Employing Bilinear Fusion and Saliency Prior Information for RGB-D Salient Object DetectionabstractMulti-modal feature fusion and saliency reasoning are two core sub-tasks of RGB-D salient object detection. However, most existing models employ linear fusion strategies (e.g., concatenation) for multi-modal feature fusion and use a simple coarse-to-fine structure for saliency reasoning. Despite their simpleness, they can neither fully capture the cross-modal complementary information nor exploit the multi-level complementary information among the cross-modal features at different levels. To address these issues, a novel RGB-D salient object detection model is presented, where we pay special attention to the aforementioned two sub-tasks. Concretely, a multi-modal feature interaction module is first presented to explore more interactions between the unimodal RGB and depth features. It helps to capture their cross-modal complementary information by jointly using some simple linear fusion strategies and bilinear fusion ones. Then, a saliency prior information guided fusion module is presented to exploit the multi-level complementary information among the fused cross-modal features at different levels. Instead of employing a simple convolutional layer for the final saliency prediction, a saliency refinement and prediction module is designed to better exploit those extracted multi-level cross-modal information for RGB-D saliency detection. Experimental results on several benchmark datasets verify the effectiveness and superiority of the proposed framework over some state-of-the-art methods. Nianchang Huang, Yang Yang 0132, Dingwen Zhang, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Multim. | 3 |
| 2021 | A Comprehensive CT Dataset for Liver Computer Assisted Diagnosis
Qingsen Yan, Bo Wang 0011, Dong Gong, Dingwen Zhang, Yang Yang 0009, Zheng You, Yanning Zhang 0001, Qinfeng Shi |
BMVC | 4 |
| 2021 | ABMDRNet: Adaptive-Weighted Bi-Directional Modality Difference Reduction Network for RGB-T Semantic SegmentationabstractSemantic segmentation models gain robustness against poor lighting conditions by virtue of complementary information from visible (RGB) and thermal images. Despite its importance, most existing RGB-T semantic segmentation models perform primitive fusion strategies, such as concatenation, element-wise summation and weighted summation, to fuse features from different modalities. These strategies, unfortunately, overlook the modality differences due to different imaging mechanisms, so that they suffer from the reduced discriminability of the fused features. To address such an issue, we propose, for the first time, the strategy of bridging-then-fusing, where the innovation lies in a novel Adaptive-weighted Bi-directional Modality Difference Reduction Network (ABMDRNet). Concretely, a Modality Difference Reduction and Fusion (MDRF) subnetwork is designed, which first employs a bi-directional image-to-image translation based method to reduce the modality differences between RGB features and thermal features, and then adaptively selects those discriminative multi-modality features for RGB-T semantic segmentation in a channel-wise weighted fusion way. Furthermore, considering the importance of contextual information in semantic segmentation, a Multi-Scale Spatial Context (MSC) module and a Multi-Scale Channel Context (MCC) module are proposed to exploit the interactions among multi-scale contextual information of cross-modality features together with their long-range dependencies along spatial and channel dimensions, respectively. Comprehensive experiments on MFNet dataset demonstrate that our method achieves new state-of-the-art results. Qiang Zhang 0020, Shenlu Zhao, Yongjiang Luo, Dingwen Zhang, Nianchang Huang, Jungong Han |
CVPR | 4 |
| 2021 | Strengthen Learning Tolerance for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims at learning to localize objects of interest by only using the image-level labels as the supervision. While numerous efforts have been made in this field, recent approaches still suffer from two challenges: one is the part domination issue while the other is the learning robustness issue. Specifically, the former makes the localizer prone to the local discriminative object regions rather than the desired whole object, and the latter makes the localizer over-sensitive to the variations of the input images so that one can hardly obtain localization results robust to the arbitrary visual stimulus. To solve these issues, we propose a novel framework to strengthen the learning tolerance, referred to as SLT-Net, for WSOL. Specifically, we consider two-fold learning tolerance strengthening mechanisms. One is the semantic tolerance strengthening mechanism, which allows the localizer to make mistakes for classifying similar semantics so that it will not concentrate too much on the discriminative local regions. The other is the visual stimuli tolerance strengthening mechanism, which enforces the localizer to be robust to different image transformations so that the prediction quality will not be sensitive to each specific input image. Finally, we implement comprehensive experimental comparisons on two widely-used datasets CUB and ILSVRC2012, which demonstrate the effectiveness of our proposed approach. Guangyu Guo 0001, Junwei Han 0001, Fang Wan 0001, Dingwen Zhang |
CVPR | 4 |
| 2021 | Light Field Saliency Detection with Dual Local Graph Learning and Reciprocative GuidanceabstractThe application of light field data in salient object detection is becoming increasingly popular recently. The difficulty lies in how to effectively fuse the features within the focal stack and how to cooperate them with the feature of the all-focus image. Previous methods usually fuse focal stack features via convolution or ConvLSTM, which are both less effective and ill-posed. In this paper, we model the information fusion within focal stack via graph networks. They introduce powerful context propagation from neighbouring nodes and also avoid ill-posed implementations. On the one hand, we construct local graph connections thus avoiding prohibitive computational costs of traditional graph networks. On the other hand, instead of processing the two kinds of data separately, we build a novel dual graph model to guide the focal stack fusion process using all-focus patterns. To handle the second difficulty, previous methods usually implement one-shot fusion for focal stack and all-focus features, hence lacking a thorough exploration of their supplements. We introduce a reciprocative guidance scheme and enable mutual guidance between these two kinds of information at multiple steps. As such, both kinds of features can be enhanced iteratively, finally benefiting the saliency prediction. Extensive experimental results show that the proposed models are all beneficial and we achieve significantly better results than state-of-the-art methods. Nian Liu 0002, Wangbo Zhao, Dingwen Zhang, Junwei Han 0001, Ling Shao 0001 |
ICCV | 3 |
| 2021 | HUMA'21: 2nd International Workshop on Human-centric Multimedia AnalysisabstractThe Second International Workshop on Human-centric Multimedia Analysis is focused on human-centric analysis using multimedia information. The human-centric multimedia analysis is one of the fundamental and challenging problems of multimedia understanding. It involves various human-centric analysis tasks like face recognition, human pose estimation, person re-identification, human action recognition, person tracking, human-computer interaction, etc. Nowadays, various multimedia sensing devices and large-scale computing infrastructures are generating a wide variety of multi-modality data at a rapid velocity, which supplies rich knowledge to tackle these challenges for human-centric analysis. Researchers and engineers have strived to push the limits of human-centric multimedia analysis in a wide variety of applications, such as smart city, retailing, intelligent manufacturing, and public services. To this end, our workshop aims to provide a platform to promote exchanges and integration for the fields of human analysis and multimedia. Wu Liu 0005, Xinchen Liu, Jingkuan Song, Dingwen Zhang, Wenbing Huang 0001, Junbo Guo, John R. Smith |
ACM Multimedia | 4 |
| 2021 | Disentangling Deep Network for Reconstructing 3D Object Shapes from Single 2D Images
Yang Yang 0009, Junwei Han 0001, Dingwen Zhang, De Cheng |
PRCV (2) | 3 |
| 2021 | SODA: Weakly Supervised Temporal Action Localization Based on Astute Background Response and Self-Distillation Learning
Tao Zhao 0006, Junwei Han 0001, Le Yang 0008, Binglu Wang, Dingwen Zhang |
Int. J. Comput. Vis. | 5 |
| 2021 | Weakly-Supervised Learning of Category-Specific 3D Object ShapesabstractCategory-specific 3D object shape models have greatly boosted the recent advances in object detection, recognition and segmentation. However, even the most advanced approach for learning 3D object shapes still requires heavy manual annotations on large-scale 2D images. Such annotations include object categories, object keypoints, and figure-ground segmentation for the instances in each image. In particular, annotating figure-ground segmentation is unbearably labor-intensive and time-consuming. To address this problem, this paper devotes to learn category-specific 3D shape models under weak supervision, where only object categories and keypoints are required to be manually annotated on the training 2D images. By exploring the underlying relationship between two tasks: object segmentation and category-specific 3D shape reconstruction, we propose a novel weakly-supervised learning framework to jointly address these two tasks and combine them to boost the final performance of the learned 3D shape models. Moreover, learning without using figure-ground segmentation leads to ambiguous solutions. To this end, we develop the confidence weighting schemes in the viewpoint estimation and 3D shape learning procedure. These schemes effectively reduce the confusion caused by the noisy data and thus increase the chances for recovering more reliable 3D object shapes. Comprehensive experiments on the challenging PASCAL VOC benchmark show that our framework achieves comparable performance with the state-of-the-art methods that use expensive manual segmentation-level annotations. In addition, our experiments also demonstrate that our 3D shape models improve object segmentation performance. Junwei Han 0001, Yang Yang 0009, Dingwen Zhang, Dong Huang 0007, Dong Xu 0001, Fernando De la Torre |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Evaluation of Saccadic Scanpath Prediction: Subjective Assessment Database and Recurrent Neural Network Based MetricabstractIn recent years, predicting the saccadic scanpaths of humans has become a new trend in the field of visual attention modeling. Given various saccadic algorithms, determining how to evaluate their ability to model a dynamic saccade has become an important yet understudied issue. To our best knowledge, existing metrics for evaluating saccadic prediction models are often heuristically designed, which may produce results that are inconsistent with human subjective assessment. To this end, we first construct a subjective database by collecting the assessments on 5,000 pairs of scanpaths from ten subjects. Based on this database, we can compare different metrics according to their consistency with human visual perception. In addition, we also propose a data-driven metric to measure scanpath similarity based on the human subjective comparison. To achieve this goal, we employ a long short-term memory (LSTM) network to learn the inference from the relationship of encoded scanpaths to a binary measurement. Experimental results have demonstrated that the LSTM-based metric outperforms other existing metrics. Moreover, we believe the constructed database can be used as a benchmark to inspire more insights for future metric selection. Chen Xia, Junwei Han 0001, Dingwen Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Cross-modality deep feature learning for brain tumor segmentation
Dingwen Zhang, Guohai Huang, Qiang Zhang 0020, Jungong Han, Junwei Han 0001, Yizhou Yu |
Pattern Recognit. | 1 |
| 2021 | Automatic pancreas segmentation based on lightweight DCNN modules and spatial prior propagation
Dingwen Zhang, Qiang Zhang 0020, Jungong Han, Shu Zhang 0001, Junwei Han 0001 |
Pattern Recognit. | 1 |
| 2021 | Revisiting Feature Fusion for RGB-T Salient Object DetectionabstractWhile many RGB-based saliency detection algorithms have recently shown the capability of segmenting salient objects from an image, they still suffer from unsatisfactory performance when dealing with complex scenarios, insufficient illumination or occluded appearances. To overcome this problem, this article studies RGB-T saliency detection, where we take advantage of thermal modality's robustness against illumination and occlusion. To achieve this goal, we revisit feature fusion for mining intrinsic RGB-T saliency patterns and propose a novel deep feature fusion network, which consists of the multi-scale, multi-modality, and multi-level feature fusion modules. Specifically, the multi-scale feature fusion module captures rich contexture features from each modality feature, while the multi-modality and multi-level feature fusion modules integrate complementary features from different modality features and different level of features, respectively. To demonstrate the effectiveness of the proposed approach, we conduct comprehensive experiments on the RGB-T saliency detection benchmark. The experimental results demonstrate that our approach outperforms other state-of-the-art methods and the conventional feature fusion modules by a large margin. Qiang Zhang 0020, Tonglin Xiao, Nianchang Huang, Dingwen Zhang, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | ASIF-Net: Attention Steered Interweave Fusion Network for RGB-D Salient Object DetectionabstractSalient object detection from RGB-D images is an important yet challenging vision task, which aims at detecting the most distinctive objects in a scene by combining color information and depth constraints. Unlike prior fusion manners, we propose an attention steered interweave fusion network (ASIF-Net) to detect salient objects, which progressively integrates cross-modal and cross-level complementarity from the RGB image and corresponding depth map via steering of an attention mechanism. Specifically, the complementary features from RGB-D images are jointly extracted and hierarchically fused in a dense and interweaved manner. Such a manner breaks down the barriers of inconsistency existing in the cross-modal data and also sufficiently captures the complementarity. Meanwhile, an attention mechanism is introduced to locate the potential salient regions in an attention-weighted fashion, which advances in highlighting the salient objects and suppressing the cluttered background regions. Instead of focusing only on pixelwise saliency, we also ensure that the detected salient objects have the objectness characteristics (e.g., complete structure and sharp boundary) by incorporating the adversarial learning that provides a global semantic constraint for RGB-D salient object detection. Quantitative and qualitative experiments demonstrate that the proposed method performs favorably against 17 state-of-the-art saliency detectors on four publicly available RGB-D salient object detection datasets. The code and results of our method are available at https://github.com/Li-Chongyi/ASIF-Net. Chongyi Li, Runmin Cong, Sam Kwong, Junhui Hou, Huazhu Fu, Guopu Zhu, Dingwen Zhang, Qingming Huang |
IEEE Trans. Cybern. | 7 |
| 2021 | Integrating Part-Object Relationship and Contrast for Camouflaged Object DetectionabstractObject detectors that solely rely on image contrast are struggling to detect camouflaged objects in images because of the high similarity between camouflaged objects and their surroundings. To address this issue, in this paper, we investigate the role of the part-object relationship for camouflaged object detection. Specifically, we propose a Part-Object relationship and Contrast Integrated Network (POCINet) covering both search and identification stages, where each stage adopts an appropriate scheme to engage the contrast information and part-object relational knowledge for camouflaged pattern decoding. Besides, we bridge these two stages via a Search-to-Identification Guidance (SIG) module, in which the search result, as well as decoded semantic knowledge, jointly enhances the features encoding ability of the identification stage. Experimental results demonstrate the superiority of our algorithm on three datasets. Notably, our algorithm raises Fβ of the best existing method by approximately 17 points on the CPD1K dataset. The source code will be released soon. Yi Liu 0038, Dingwen Zhang, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Exploring Rich and Efficient Spatial Temporal Interactions for Real-Time Video Salient Object DetectionabstractWe have witnessed a growing interest in video salient object detection (VSOD) techniques in today's computer vision applications. In contrast with temporal information (which is still considered a rather unstable source thus far), the spatial information is more stable and ubiquitous, thus it could influence our vision system more. As a result, the current main-stream VSOD approaches have inferred and obtained their saliency primarily from the spatial perspective, still treating temporal information as subordinate. Although the aforementioned methodology of focusing on the spatial aspect is effective in achieving a numeric performance gain, it still has two critical limitations. First, to ensure the dominance by the spatial information, its temporal counterpart remains inadequately used, though in some complex video scenes, the temporal information may represent the only reliable data source, which is critical to derive the correct VSOD. Second, both spatial and temporal saliency cues are often computed independently in advance and then integrated later on, while the interactions between them are omitted completely, resulting in saliency cues with limited quality. To combat these challenges, this paper advocates a novel spatiotemporal network, where the key innovation is the design of its temporal unit. Compared with other existing competitors (e.g., convLSTM), the proposed temporal unit exhibits an extremely lightweight design that does not degrade its strong ability to sense temporal information. Furthermore, it fully enables the computation of temporal saliency cues that interact with their spatial counterparts, ultimately boosting the overall VSOD performance and realizing its full potential towards mutual performance improvement for each. The proposed method is easy to implement yet still effective, achieving high-quality VSOD at 50 FPS in real-time applications. Chenglizhao Chen, Guotao Wang 0004, Chong Peng 0001, Yuming Fang 0001, Dingwen Zhang, Hong Qin 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | HTD: Heterogeneous Task Decoupling for Two-Stage Object DetectionabstractDecoupling the sibling head has recently shown great potential in relieving the inherent task-misalignment problem in two-stage object detectors. However, existing works design similar structures for the classification and regression, ignoring task-specific characteristics and feature demands. Besides, the shared knowledge that may benefit the two branches is neglected, leading to potential excessive decoupling and semantic inconsistency. To address these two issues, we propose Heterogeneous task decoupling (HTD) framework for object detection, which utilizes a Progressive Graph (PGraph) module and a Border-aware Adaptation (BA) module for task-decoupling. Specifically, we first devise a Semantic Feature Aggregation (SFA) module to aggregate global semantics with image-level supervision, serving as the shared knowledge for the task-decoupled framework. Then, the PGraph module performs progressive graph reasoning, including local spatial aggregation and global semantic interaction, to enhance semantic representations of region proposals for classification. The proposed BA module integrates multi-level features adaptively, focusing on the low-level border activation to obtain representations with spatial and border perception for regression. Finally, we utilize the aggregated knowledge from SFA to keep the instance-level semantic consistency (ISC) of decoupled frameworks. Extensive experiments demonstrate that HTD outperforms existing detection works by a large margin, and achieves single-model 50.4%AP and 33.2% APs on COCO test-dev set using ResNet-101-DCN backbone, which is the best entry among state-of-the-arts under the same configuration. Our code is available at https://github.com/CityU-AIM-Group/HTD. Wuyang Li, Zhen Chen 0013, Baopu Li, Dingwen Zhang, Yixuan Yuan |
IEEE Trans. Image Process. | 4 |
| 2021 | A Structure-Aware Relation Network for Thoracic Diseases Detection and SegmentationabstractInstance level detection and segmentation of thoracic diseases or abnormalities are crucial for automatic diagnosis in chest X-ray images. Leveraging on constant structure and disease relations extracted from domain knowledge, we propose a structure-aware relation network (SAR-Net) extending Mask R-CNN. The SAR-Net consists of three relation modules: 1. the anatomical structure relation module encoding spatial relations between diseases and anatomical parts. 2. the contextual relation module aggregating clues based on query-key pair of disease RoI and lung fields. 3. the disease relation module propagating co-occurrence and causal relations into disease proposals. Towards making a practical system, we also provide ChestX-Det, a chest X-Ray dataset with instance-level annotations (boxes and masks). ChestX-Det is a subset of the public dataset NIH ChestX-ray14. It contains ~3500 images of 13 common disease categories labeled by three board-certified radiologists. We evaluate our SAR-Net on it and another dataset DR-Private. Experimental results show that it can enhance the strong baseline of Mask R-CNN with significant improvements. The ChestX-Det is released at https://github.com/Deepwise-AILab/ChestX-Det-Dataset. Jingyu Liu 0004, Shu Zhang 0001, Dingwen Zhang, Yizhou Yu |
IEEE Trans. Medical Imaging | 6 |
| 2020 | Deep Embedded Complementary and Interactive Information for Multi-View ClassificationabstractMulti-view classification optimally integrates various features from different views to improve classification tasks. Though most of the existing works demonstrate promising performance in various computer vision applications, we observe that they can be further improved by sufficiently utilizing complementary view-specific information, deep interactive information between different views, and the strategy of fusing various views. In this work, we propose a novel multi-view learning framework that seamlessly embeds various view-specific information and deep interactive information and introduces a novel multi-view fusion strategy to make a joint decision during the optimization for classification. Specifically, we utilize different deep neural networks to learn multiple view-specific representations, and model deep interactive information through a shared interactive network using the cross-correlations between attributes of these representations. After that, we adaptively integrate multiple neural networks by flexibly tuning the power exponent of weight, which not only avoids the trivial solution of weight but also provides a new approach to fuse outputs from different deterministic neural networks. Extensive experiments on several public datasets demonstrate the rationality and effectiveness of our method. Jinglin Xu, Wenbin Li 0006, Dingwen Zhang, Junwei Han 0001 |
AAAI | 4 |
| 2020 | Taking a Deeper Look at Co-Salient Object DetectionabstractCo-salient object detection (CoSOD) is a newly emerging and rapidly growing branch of salient object detection (SOD), which aims to detect the co-occurring salient objects in multiple images. However, existing CoSOD datasets often have a serious data bias, which assumes that each group of images contains salient objects of similar visual appearances. This bias results in the ideal settings and the effectiveness of the models, trained on existing datasets, may be impaired in real-life situations, where the similarity is usually semantic or conceptual. To tackle this issue, we first collect a new high-quality dataset, named CoSOD3k, which contains 3,316 images divided into 160 groups with multiple level annotations, i.e., category, bounding box, object, and instance levels. CoSOD3k makes a significant leap in terms of diversity, difficulty and scalability, benefiting related vision tasks. Besides, we comprehensively summarize 34 cutting-edge algorithms, benchmarking 19 of them over four existing CoSOD datasets (MSRC, iCoSeg, Image Pair and CoSal2015) and our CoSOD3k with a total of ~61K images (largest scale), and reporting group-level performance analysis. Finally, we discuss the challenge and future work of CoSOD. Our study would give a strong boost to growth in the CoSOD community. Benchmark toolbox and results are available on our project page. Deng-Ping Fan, Zheng Lin 0005, Ge-Peng Ji, Dingwen Zhang, Huazhu Fu, Ming-Ming Cheng |
CVPR | 4 |
| 2020 | HUMA'20: 1st International Workshop on Human-Centric Multimedia AnalysisabstractThe First International Workshop on Human-Centric MultimediaAnalysis is concentrated on the tasks of human-centric analysis with multimedia and multimodal information. It is one of the fundamental and challenging problems of multimedia understanding. The human-centric multimedia analysis involves multiple tasks such as face detection and recognition, human body pattern analysis, person re-identification, human action detection, person tracking,human-object interaction, and so on. Today, multiple multimedia sensing technologies and large-scale computing infrastructures are producing at a rapid velocity a wide variety of big multi-modality data for human-centric analysis, which provides rich knowledge to help tackle these challenges. Researchers have strived to push the limits of human-centric multimedia analysis in a wide variety of applications, such as intelligent surveillance, retailing, fashion design, and services. Therefore, this workshop aims to provide a platform to bridge the gap between the communities of human analysis and multimedia. Wu Liu 0005, Chuang Gan 0001, Jingkuan Song, Dingwen Zhang, Wenbing Huang 0001, John R. Smith |
ACM Multimedia | 4 |
| 2020 | Few-Cost Salient Object Detection with Adversarial-Paced LearningabstractDetecting and segmenting salient objects from given image scenes has received great attention in recent years. A fundamental challenge in training the existing deep saliency detection models is the requirement of large amounts of annotated data. While gathering large quantities of training data becomes cheap and easy, annotating the data is an expensive process in terms of time, labor and human expertise. To address this problem, this paper proposes to learn the effective salient object detection model based on the manual annotation on a few training images only, thus dramatically alleviating human labor in training models. To this end, we name this new task as the few-cost salient object detection and propose an adversarial-paced learning (APL)-based framework to facilitate the few-cost learning scenario. Essentially, APL is derived from the self-paced learning (SPL) regime but it infers the robust learning pace through the data-driven adversarial learning mechanism rather than the heuristic design of the learning regularizer. Comprehensive experiments on four widely-used benchmark datasets have demonstrated that the proposed approach can effectively approach to the existing supervised deep salient object detection models with only 1k human-annotated training images. Dingwen Zhang, Haibin Tian, Jungong Han |
NeurIPS | 1 |
| 2020 | SPFTN: A Joint Learning Framework for Localizing and Segmenting Objects in Weakly Labeled VideosabstractObject localization and segmentation in weakly labeled videos are two interesting yet challenging tasks. Models built for simultaneous object localization and segmentation have been explored in the conventional fully supervised learning scenario to boost the performance of each task. However, none of the existing works has attempted to jointly learn object localization and segmentation models under weak supervision. To this end, we propose a joint learning framework called Self-Paced Fine-Tuning Network (SPFTN) for localizing and segmenting objects in weakly labelled videos. Learning the deep model jointly for object localization and segmentation under weak supervision is very challenging as the learning process of each single task would face serious ambiguity issue due to the lack of bounding-box or pixel-level supervision. To address this problem, our proposed deep SPFTN model is carefully designed with a novel multi-task self-paced learning objective, which leverages the task-specific prior knowledge and the knowledge that has been already captured to infer the confident training samples for each task. By aggregating the confident knowledge from each single task to mine reliable patterns and learning deep feature representation for both tasks, the proposed learning framework can address the ambiguity issue under weak supervision with simple optimization. Comprehensive experiments on the large-scale YouTube-Objects and DAVIS datasets demonstrate that the proposed approach achieves superior performance when compared with other state-of-the-art methods and the baseline networks/models. Dingwen Zhang, Junwei Han 0001, Le Yang 0008, Dong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Synthesizing Supervision for Learning Deep Saliency Network without Human AnnotationabstractRecently, the research field of salient object detection is undergoing a rapid and remarkable development along with the wide usage of deep neural networks. Being trained with a large number of images annotated with strong pixel-level ground-truth masks, the deep salient object detectors have achieved the state-of-the-art performance. However, it is expensive and time-consuming to provide the pixel-level ground-truth masks for each training image. To address this problem, this paper proposes one of the earliest frameworks to learn deep salient object detectors without requiring any human annotation. The supervisory signals used in our learning framework are generated through a novel supervision synthesis scheme, in which the key insights are "knowledge source transition" and "supervision by fusion". Specifically, in the proposed learning framework, both the external knowledge source and the internal knowledge source are explored dynamically to provide informative cues for synthesizing supervision required in our approach, while a two-stream fusion mechanism is also established to implement the supervision synthesis process. Comprehensive experiments on four benchmark datasets demonstrate that the deep salient object detector trained by our newly proposed learning framework often works well without requiring any human annotated masks, which even approaches to its upper-bound obtained under the fully supervised learning fashion (within only 3 percent performance gap). Besides, we also apply the salient object detector learnt with our annotation-free learning framework to assist the weakly supervised semantic segmentation task, which demonstrates that our approach can also alleviate the heavy supplementary supervision required in the existing weakly supervised semantic segmentation framework. Dingwen Zhang, Junwei Han 0001, Dong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Unsupervised object-level video summarization with online motion auto-encoder
Yujia Zhang 0001, Xiaodan Liang, Dingwen Zhang, Min Tan 0001, Eric P. Xing |
Pattern Recognit. Lett. | 3 |
| 2020 | A structure-aware splitting framework for separating cell clumps in biomedical images
Qiang Zhang 0020, Zaihao Liu, Dingwen Zhang |
Signal Process. | 4 |
| 2020 | Fusion of Multiple Person Re-id Methods With Model and Data-Aware AbilitiesabstractPerson re-identification (person re-id) has attracted rapidly increasing attention in computer vision and pattern recognition research community in recent years. With the goal of providing match ranking results between each query person image and the gallery ones, the person re-id technique has been widely explored and a large number of person re-id methods have been developed. As these algorithms leverage different kinds of prior assumptions, image features, distance matching functions, et al., each of them has its own strengths and weaknesses. Inspired by these facts, this paper proposes a novel person re-id method based on the idea of inferring superior fusion results from a variety of previous base person re-id algorithms using different methodologies or features. To achieve this goal, we propose a novel framework which mainly consists of two steps: 1) a number of existing person re-id methods are implemented, and the ranking results are obtained in the test datasets. and 2) the robust fusion strategy is applied to obtain better re-ranked matching results by simultaneously considering the recognition abilities of various base re-id methods and the difficulties of different gallery person images to be correctly recognized under the generative model of labels, abilities, and difficulties framework. Comprehensive experiments show the effectiveness of our proposed method, and we have received state-of-the-art results on recent popular person re-id datasets. De Cheng, Zhihui Li 0001, Yihong Gong, Dingwen Zhang |
IEEE Trans. Cybern. | 4 |
| 2020 | Revisiting Anchor Mechanisms for Temporal Action LocalizationabstractMost of the current action localization methods follow an anchor-based pipeline: depicting action instances by pre-defined anchors, learning to select the anchors closest to the ground truth, and predicting the confidence of anchors with refinements. Pre-defined anchors set prior about the location and duration for action instances, which facilitates the localization for common action instances but limits the flexibility for tackling action instances with drastic varieties, especially for extremely short or extremely long ones. To address this problem, this paper proposes a novel anchor-free action localization module that assists action localization by temporal points. Specifically, this module represents an action instance as a point with its distances to the starting boundary and ending boundary, alleviating the pre-defined anchor restrictions in terms of action localization and duration. The proposed anchor-free module is capable of predicting the action instances whose duration is either extremely short or extremely long. By combining the proposed anchor-free module with a conventional anchor-based module, we propose a novel action localization framework, called A2Net. The cooperation between anchor-free and anchor-based modules achieves superior performance to the state-of-the-art on THUMOS14 (45.5% vs. 42.8%). Furthermore, comprehensive experiments demonstrate the complementarity between the anchor-free and the anchor-based module, making A2Net simple but effective. Le Yang 0008, Houwen Peng, Dingwen Zhang, Jianlong Fu, Junwei Han 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | RGB-T Salient Object Detection via Fusing Multi-Level CNN FeaturesabstractRGB-induced salient object detection has recently witnessed substantial progress, which is attributed to the superior feature learning capability of deep convolutional neural networks (CNNs). However, such detections suffer from challenging scenarios characterized by cluttered backgrounds, low-light conditions and variations in illumination. Instead of improving RGB based saliency detection, this paper takes advantage of the complementary benefits of RGB and thermal infrared images. Specifically, we propose a novel end-to-end network for multi-modal salient object detection, which turns the challenge of RGB-T saliency detection to a CNN feature fusion problem. To this end, a backbone network (e.g., VGG-16) is first adopted to extract the coarse features from each RGB or thermal infrared image individually, and then several adjacent-depth feature combination (ADFC) modules are designed to extract multi-level refined features for each single-modal input image, considering that features captured at different depths differ in semantic information and visual details. Subsequently, a multi-branch group fusion (MGF) module is employed to capture the cross-modal features by fusing those features from ADFC modules for a RGB-T image pair at each level. Finally, a joint attention guided bi-directional message passing (JABMP) module undertakes the task of saliency prediction via integrating the multi-level fused features from MGF modules. Experimental results on several public RGB-T salient object detection datasets demonstrate the superiorities of our proposed algorithm over the state-of-the-art approaches, especially under challenging conditions, such as poor illumination, complex background and low contrast. Qiang Zhang 0020, Nianchang Huang, Dingwen Zhang, Caifeng Shan, Jungong Han |
IEEE Trans. Image Process. | 4 |
| 2020 | Exploring Task Structure for Brain Tumor Segmentation From Multi-Modality MR ImagesabstractBrain tumor segmentation, which aims at segmenting the whole tumor area, enhancing tumor core area, and tumor core area from each input multi-modality bioimaging data, has received considerable attention from both academia and industry. However, the existing approaches usually treat this problem as a common semantic segmentation task without taking into account the underlying rules in clinical practice. In reality, physicians tend to discover different tumor areas by weighing different modality volume data. Also, they initially segment the most distinct tumor area, and then gradually search around to find the other two. We refer to the first property as the task-modality structure while the second property as the task-task structure, based on which we propose a novel task-structured brain tumor segmentation network (TSBTS net). Specifically, to explore the task-modality structure, we design a modality-aware feature embedding mechanism to infer the important weights of the modality data during network learning. To explore the tasktask structure, we formulate the prediction of the different tumor areas as conditional dependency sub-tasks and encode such dependency in the network stream. Experiments on BraTS benchmarks show that the proposed method achieves superior performance in segmenting the desired brain tumor areas while requiring relatively lower computational costs, compared to other state-of-the-art methods and baseline models. Dingwen Zhang, Guohai Huang, Qiang Zhang 0020, Jungong Han, Junwei Han 0001, Yizhou Wang 0001, Yizhou Yu |
IEEE Trans. Image Process. | 1 |
| 2020 | From Discriminant to Complete: Reinforcement Searching-Agent Learning for Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD) is an interesting yet challenging task in the computer vision community. The core is to discover the image regions that contain the complete object instances under the image-level supervision. Existing works usually solve this problem via a proposal selection strategy, which selects the most discriminative box regions from the weakly labeled training images. However, these regions usually only contain the discriminative object parts rather than the complete object instances. To address this problem, this article proposes to learn a searching-agent to gradually mine desirable object regions under a region searching paradigm, where we formulate the searching process as a Markov decision process and learn the searching-agent under a deep reinforcement learning framework. To learn such a searching-agent under the weak supervision, we extract the pseudo-complete object regions and the corresponding local discriminative object parts and introduce the obtained pseudo-target-part training pairs into the reinforcement learning process of the search-agent. This learning strategy has twofold advantages: 1) it can mimic the searching process to reveal complete object regions from a certain discriminative part of the object under the weak supervision and 2) it will not suffer from the learning difficulty arise from the long-action sequence that happens when searching from the entire image range. Comprehensive experiments on benchmark data sets demonstrate that by integrating the learned searching-agent with the existing WSOD method, we can achieve better performance than the other state-of-the-art and baseline methods. Dingwen Zhang, Junwei Han 0001, Tao Zhao 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Employing Deep Part-Object Relationships for Salient Object DetectionabstractDespite Convolutional Neural Networks (CNNs) based methods have been successful in detecting salient objects, their underlying mechanism that decides the salient intensity of each image part separately cannot avoid inconsistency of parts within the same salient object. This would ultimately result in an incomplete shape of the detected salient object. To solve this problem, we dig into part-object relationships and take the unprecedented attempt to employ these relationships endowed by the Capsule Network (CapsNet) for salient object detection. The entire salient object detection system is built directly on a Two-Stream Part-Object Assignment Network (TSPOANet) consisting of three algorithmic steps. In the first step, the learned deep feature maps of the input image are transformed to a group of primary capsules. In the second step, we feed the primary capsules into two identical streams, within each of which low-level capsules (parts) will be assigned to their familiar high-level capsules (object) via a locally connected routing. In the final step, the two streams are integrated in the form of a fully connected layer, where the relevant parts can be clustered together to form a complete salient object. Experimental results demonstrate the superiority of the proposed salient object detection network over the state-of-the-art methods. Yi Liu 0038, Qiang Zhang 0020, Dingwen Zhang, Jungong Han |
ICCV | 3 |
| 2019 | Leveraging Prior-Knowledge for Weakly Supervised Object Detection Under a Collaborative Self-Paced Curriculum Learning Framework
Dingwen Zhang, Junwei Han 0001, Deyu Meng |
Int. J. Comput. Vis. | 1 |
| 2019 | Dilated temporal relational adversarial network for generic video summarization
Yujia Zhang 0001, Michael Kampffmeyer, Xiaodan Liang, Dingwen Zhang, Min Tan 0001, Eric P. Xing |
Multim. Tools Appl. | 4 |
| 2019 | Learning Object Detectors With Semi-Annotated Weak LabelsabstractFor alleviating the human labor associated with annotating the training data for learning object detectors, recent research has focused on semi-supervised object detection (SSOD) and weakly supervised object detection (WSOD) approaches. In SSOD, instead of annotating all the instances in the whole training set, people only need to annotate the part of the training instances using bounding boxes. In WSOD, people need to annotate the image-level tags on all training images to indicate the object categories contained by the corresponding images since more detailed bounding box annotations are no longer needed. Along this line of research, this paper makes a further step to alleviate the human labor in annotating training data, leading to the problem of object detection with semi-annotated weak labels (ODSAWLs). Instead of labeling image-level tags on all training images, ODSAWL only needs the image-level tags for a small portion of the training images, and then, the object detectors can be learned from a small portion of the weakly-labeled training images and from the remaining unlabeled training images. To address such a challenging problem, this paper proposes a cross model co-training framework that collaborates an object localizer and a tag generator in an alternative optimization procedure. Specifically, during the learning procedure, these two (deep) models can transfer the needed knowledge (including labels and visual patterns) between each other. The whole learning procedure is accomplished in a few stages under the guidance of a progressive learning curriculum. To demonstrate the effectiveness of the proposed approach, we implement the comprehensive experiments on three benchmark datasets, where the obtained experimental results are quite encouraging. Notably, by using only about 15% weakly labeled training images, the proposed approach can effectively approach, or even outperform, the state-of-the-art WSOD methods. Dingwen Zhang, Junwei Han 0001, Guangyu Guo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Multi-Rate Gated Recurrent Convolutional Networks for Video-Based Pedestrian Re-IdentificationabstractMatching pedestrians across multiple camera views has attracted lots of recent research attention due to its apparent importance in surveillance and security applications.While most existing works address this problem in a still-image setting, we consider the more informative and challenging video-based person re-identification problem, where a video of a pedestrian as seen in one camera needs to be matched to a gallery of videos captured by other non-overlapping cameras. We employ a convolutional network to extract the appearance and motion features from raw video sequences, and then feed them into a multi-rate recurrent network to exploit the temporal correlations, and more importantly, to take into account the fact that pedestrians, sometimes even the same pedestrian, move in different speeds across different camera views. The combined network is trained in an end-to-end fashion, and we further propose an initialization strategy via context reconstruction to largely improve the performance. We conduct extensive experiments on the iLIDS-VID and PRID-2011 datasets, and our experimental results confirm the effectiveness and the generalization ability of our model. Zhihui Li 0001, Lina Yao 0001, Feiping Nie 0001, Dingwen Zhang |
AAAI | 4 |
| 2018 | Reinforcement Cutting-Agent Learning for Video Object SegmentationabstractVideo object segmentation is a fundamental yet challenging task in computer vision community. In this paper, we formulate this problem as a Markov Decision Process, where agents are learned to segment object regions under a deep reinforcement learning framework. Essentially, learning agents for segmentation is nontrivial as segmentation is a nearly continuous decision-making process, where the number of the involved agents (pixels or superpixels) and action steps from the seed (super)pixels to the whole object mask might be incredibly huge. To overcome this difficulty, this paper simplifies the learning of segmentation agents to the learning of a cutting-agent, which only has a limited number of action units and can converge in just a few action steps. The basic assumption is that object segmentation mainly relies on the interaction between object regions and their context. Thus, with an optimal object (box) region and context (box) region, we can obtain the desirable segmentation mask through further inference. Based on this assumption, we establish a novel reinforcement cutting-agent learning framework, where the cutting-agent consists of a cutting-policy network and a cutting-execution network. The former learns policies for deciding optimal object-context box pair, while the latter executes the cutting function based on the inferred object-context box pair. With the collaborative interaction between the two networks, our method can achieve the outperforming VOS performance on two public benchmarks, which demonstrates the rationality of our assumption as well as the effectiveness of the proposed learning framework. Junwei Han 0001, Le Yang 0008, Dingwen Zhang, Xiaojun Chang, Xiaodan Liang |
CVPR | 3 |
| 2018 | PoseFlow: A Deep Motion Representation for Understanding Human Behaviors in VideosabstractMotion of the human body is the critical cue for understanding and characterizing human behavior in videos. Most existing approaches explore the motion cue using optical flows. However, optical flow usually contains motion on both the interested human bodies and the undesired background. This "noisy" motion representation makes it very challenging for pose estimation and action recognition in real scenarios. To address this issue, this paper presents a novel deep motion representation, called PoseFlow, which reveals human motion in videos while suppressing background and motion blur, and being robust to occlusion. For learning PoseFlow with mild computational cost, we propose a functionally structured spatial-temporal deep network, PoseFlow Net (PFN), to jointly solve the skeleton localization and matching problems of PoseFlow. Comprehensive experiments show that PFN outperforms the state-of-the-art deep flow estimation models in generating PoseFlow. Moreover, PoseFlow demonstrates its potential on improving two challenging tasks in human video analysis: pose estimation and action recognition. Dingwen Zhang, Guangyu Guo 0001, Dong Huang 0007, Junwei Han 0001 |
CVPR | 1 |
| 2018 | A Unified Metric Learning-Based Framework for Co-Saliency DetectionabstractCo-saliency detection, which focuses on extracting commonly salient objects in a group of relevant images, has been attracting research interest because of its broad applications. In practice, the relevant images in a group may have a wide range of variations, and the salient objects may also have large appearance changes. Such wide variations usually bring about large intra-co-salient objects (intra-COs) diversity and high similarity between COs and background, which makes the co-saliency detection task more difficult. To address these problems, we make the earliest effort to introduce metric learning to co-saliency detection. Specifically, we propose a unified metric learning-based framework to jointly learn discriminative feature representation and co-salient object detector. This is achieved by optimizing a new objective function that explicitly embeds a metric learning regularization term into support vector machine (SVM) training. Here, the metric learning regularization term is used to learn a powerful feature representation that has small intra-COs scatter, but big separation between background and COs and the SVM classifier is used for subsequent co-saliency detection. In the experiments, we comprehensively evaluate the proposed method on two commonly used benchmark data sets. The state-of-the-art results are achieved in comparison with the existing co-saliency detection methods. Junwei Han 0001, Gong Cheng 0003, Dingwen Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Robust Object Co-Segmentation Using Background PriorabstractGiven a set of images that contain objects from a common category, object co-segmentation aims at automatically discovering and segmenting such common objects from each image. During the past few years, object co-segmentation has received great attention in the computer vision community. However, the existing approaches are usually designed with misleading assumptions, unscalable priors, or subjective computational models, which do not have sufficient robustness for dealing with complex and unconstrained real-world image contents. This paper proposes a novel two-stage co-segmentation framework, mainly for addressing the robustness issue. In the proposed framework, we first introduce the concept of union background and use it to improve the robustness for suppressing the image backgrounds contained by the given image groups. Then, we also weaken the requirement for the strong prior knowledge by using the background prior instead. This can improve the robustness when scaling up for the unconstrained image contents. Based on the weak background prior, we propose a novel MR-SGS model, i.e., manifold ranking with the self-learned graph structure, which can infer suitable graph structures in a data-driven manner rather than building the fixed graph structure relying on the subjective design. Such capacity is critical for further improving the robustness in inferring the foreground/background probability of each image pixel. Comprehensive experiments and comparisons with other state-of-the-art approaches can demonstrate the effectiveness of the proposed work. Junwei Han 0001, Rong Quan, Dingwen Zhang, Feiping Nie 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Segmentation in Weakly Labeled Videos via a Semantic Ranking and Optical Warping NetworkabstractWeakly supervised video object segmentation (WSVOS) focuses on generating pixel-level object masks for videos only tagged with class labels, which is an essential yet challenging task. For WSVOS, the algorithm is just aware of rough category information rather than the concrete object size and location cues, besides it lacks reliable annotated exemplars to learn temporal evolution in the investigated videos. Basically, there are three challenging factors which may influence the performance of WSVOS: foreground object discovery in each frame, coarse object semantic consistency within each video, and fine-grained segmentation smoothness within neighbor frames. In this paper, we establish a semantic ranking and optical warping network (SROWN) to simultaneously solve these three challenges in a unified framework. For the first challenge, we apply the still image saliency detection method and discover the foreground object for each frame via a segmentation network. Due to the huge discrepancies between the image saliency and the video object segmentation, we step further and propose two subnetworks to solve the other two challenges. For the second one, we propose an attentive semantic ranking subnetwork to mine video-level tags, which can learn discriminative features for semantic ranking and lead to semantic consistent segmentation masks. For the third one, we propose an optical flow warping subnetwork to constrain fine-grained segmentation smoothness within neighbor frames, which can suppress the large deformation and thus obtain smooth object boundaries for adjacent frames. Experiments on two benchmark datasets, i.e., DAVIS dataset and YouTube-Objects dataset, demonstrate the effectiveness of the proposed approach for segmenting out video objects under weak supervision. Le Yang 0008, Junwei Han 0001, Dingwen Zhang, Nian Liu 0002, Dong Zhang 0009 |
IEEE Trans. Image Process. | 3 |
| 2018 | A Review of Co-Saliency Detection Algorithms: Fundamentals, Applications, and ChallengesabstractCo-saliency detection is a newly emerging and rapidly growing research area in the computer vision community. As a novel branch of visual saliency, co-saliency detection refers to the discovery of common and salient foregrounds from two or more relevant images, and it can be widely used in many computer vision tasks. The existing co-saliency detection algorithms mainly consist of three components: extracting effective features to represent the image regions, exploring the informative cues or factors to characterize co-saliency, and designing effective computational frameworks to formulate co-saliency. Although numerous methods have been developed, the literature is still lacking a deep review and evaluation of co-saliency detection techniques. In this article, we aim at providing a comprehensive review of the fundamentals, challenges, and applications of co-saliency detection. Specifically, we provide an overview of some related computer vision works, review the history of co-saliency detection, summarize and categorize the major algorithms in this research area, discuss some open issues in this area, present the potential applications of co-saliency detection, and finally point out some unsolved challenges and promising future works. We expect this review to be beneficial to both fresh and senior researchers in this field and to give insights to researchers in other related areas regarding the utility of co-saliency detection algorithms. Dingwen Zhang, Huazhu Fu, Junwei Han 0001, Ali Borji, Xuelong Li 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2018 | Unsupervised Salient Object Detection via Inferring From Imperfect Saliency ModelsabstractVisual saliency detection has become an active research direction in recent years. A large number of saliency models, which can automatically locate objects of interest in images, have been developed. As these models take advantage of different kinds of prior assumptions, image features, and computational methodologies, they have their own strengths and weaknesses and may cope with only one or a few types of images well. Inspired by these facts, this paper proposes a novel salient object detection approach with the idea of inferring a superior model from a variety of previous imperfect saliency models via optimally leveraging the complementary information among them. The proposed approach mainly consists of three steps. First, a number of existing unsupervised saliency models are adopted to provide weak/imperfect saliency predictions for each region in the image. Then, a fusion strategy is used to fuse each image region's weak saliency predictions into a strong one by simultaneously considering the performance differences among various weak predictions and various characteristics of different image regions. Finally, a local spatial consistency constraint that ensures high similarity of the saliency labels for neighboring image regions with similar features is proposed to refine the results. Comprehensive experiments on five public benchmark datasets and comparisons with a number of state-of-the-art approaches can demonstrate the effectiveness of the proposed work. Rong Quan, Junwei Han 0001, Dingwen Zhang, Feiping Nie 0001, Xueming Qian, Xuelong Li 0001 |
IEEE Trans. Multim. | 3 |
| 2017 | Learning Category-Specific 3D Shape Models from Weakly Labeled 2D ImagesabstractRecently, researchers have made great processes to build category-specific 3D shape models from 2D images with manual annotations consisting of class labels, keypoints, and ground truth figure-ground segmentations. However, the annotation of figure-ground segmentations is still labor-intensive and time-consuming. To further alleviate the burden of providing such manual annotations, we make the earliest effort to learn category-specific 3D shape models by only using weakly labeled 2D images. By revealing the underlying relationship between the tasks of common object segmentation and category-specific 3D shape reconstruction, we propose a novel framework to jointly solve these two problems along a cluster-level learning curriculum. Comprehensive experiments on the challenging PASCAL VOC benchmark demonstrate that the category-specific 3D shape models trained using our weakly supervised learning framework could, to some extent, approach the performance of the state-of-the-art methods using expensive manual segmentation annotations. In addition, the experiments also demonstrate the effectiveness of using 3D shape models for helping common object segmentation. Dingwen Zhang, Junwei Han 0001, Yang Yang 0009, Dong Huang 0007 |
CVPR | 1 |
| 2017 | SPFTN: A Self-Paced Fine-Tuning Network for Segmenting Objects in Weakly Labelled VideosabstractObject segmentation in weakly labelled videos is an interesting yet challenging task, which aims at learning to perform category-specific video object segmentation by only using video-level tags. Existing works in this research area might still have some limitations, e.g., lack of effective DNN-based learning frameworks, under-exploring the context information, and requiring to leverage the unstable negative video collection, which prevent them from obtaining more promising performance. To this end, we propose a novel self-paced fine-tuning network (SPFTN)-based framework, which could learn to explore the context information within the video frames and capture adequate object semantics without using the negative videos. To perform weakly supervised learning based on the deep neural network, we make the earliest effort to integrate the self-paced learning regime and the deep neural network into a unified and compatible framework, leading to the self-paced fine-tuning network. Comprehensive experiments on the large-scale YouTube-Objects and DAVIS datasets demonstrate that the proposed approach achieves superior performance as compared with other state-of-the-art methods as well as the baseline networks and models. Dingwen Zhang, Le Yang 0008, Deyu Meng, Dong Xu 0001, Junwei Han 0001 |
CVPR | 1 |
| 2017 | Supervision by Fusion: Towards Unsupervised Learning of Deep Salient Object Detector
Dingwen Zhang, Junwei Han 0001 |
ICCV | 1 |
| 2017 | Self-paced Mixture of RegressionsabstractMixture of regressions (MoR) is the well-established and effective approach to model discontinuous and heterogeneous data in regression problems. Existing MoR approaches assume smooth joint distribution for its good anlaytic properties. However, such assumption makes existing MoR very sensitive to intra-component outliers (the noisy training data residing in certain components) and the inter-component imbalance (the different amounts of training data in different components). In this paper, we make the earliest effort on Self-paced Learning (SPL) in MoR, i.e., Self-paced mixture of regressions (SPMoR) model. We propose a novel self-paced regularizer based on the Exclusive LASSO, which improves inter-component balance of training data. As a robust learning regime, SPL pursues confidence sample reasoning. To demonstrate the effectiveness of SPMoR, we conducted experiments on both the sythetic examples and real-world applications to age estimation and glucose estimation. The results show that SPMoR outperforms the state-of-the-arts methods. Longfei Han, Dingwen Zhang, Dong Huang 0007, Xiaojun Chang, Senlin Luo, Junwei Han 0001 |
IJCAI | 2 |
| 2017 | How Unlabeled Web Videos Help Complex Event Detection?abstractThe lack of labeled exemplars is an important factor that makes the task of multimedia event detection (MED) complicated and challenging. Utilizing artificially picked and labeled external sources is an effective way to enhance the performance of MED. However, building these data usually requires professional human annotators, and the procedure is too time-consuming and costly to scale. In this paper, we propose a new robust dictionary learning framework for complex event detection, which is able to handle both labeled and easy-to-get unlabeled web videos by sharing the same dictionary. By employing the lq-norm based loss jointly with the structured sparsity based regularization, our model shows strong robustness against the substantial noisy and outlier videos from open source. We exploit an effective optimization algorithm to solve the proposed highly non-smooth and non-convex problem. Extensive experiment results over standard datasets of TRECVID MEDTest 2013 and TRECVID MEDTest 2014 demonstrate the effectiveness and superiority of the proposed framework on complex event detection. Huan Liu 0012, Minnan Luo, Dingwen Zhang, Xiaojun Chang, Cheng Deng 0002 |
IJCAI | 4 |
| 2017 | Co-Saliency Detection via a Self-Paced Multiple-Instance Learning FrameworkabstractAs an interesting and emerging topic, co-saliency detection aims at simultaneously extracting common salient objects from a group of images. On one hand, traditional co-saliency detection approaches rely heavily on human knowledge for designing hand-crafted metrics to possibly reflect the faithful properties of the co-salient regions. Such strategies, however, always suffer from poor generalization capability to flexibly adapt various scenarios in real applications. On the other hand, most current methods pursue co-saliency detection in unsupervised fashions. This, however, tends to weaken their performance in real complex scenarios because they are lack of robust learning mechanism to make full use of the weak labels of each image. To alleviate these two problems, this paper proposes a new SP-MIL framework for co-saliency detection, which integrates both multiple instance learning (MIL) and self-paced learning (SPL) into a unified learning framework. Specifically, for the first problem, we formulate the co-saliency detection problem as a MIL paradigm to learn the discriminative classifiers to detect the co-saliency object in the "instance-level". The formulated MIL component facilitates our method capable of automatically producing the proper metrics to measure the intra-image contrast and the inter-image consistency for detecting co-saliency in a purely self-learning way. For the second problem, the embedded SPL paradigm is able to alleviate the data ambiguity under the weak supervision of co-saliency detection and guide a robust learning manner in complex scenarios. Experiments on benchmark datasets together with multiple extended computer vision applications demonstrate the superiority of the proposed framework beyond the state-of-the-arts. Dingwen Zhang, Deyu Meng, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Revisiting Co-Saliency Detection: A Novel Approach Based on Two-Stage Multi-View Spectral Rotation Co-clusteringabstractWith the goal of discovering the common and salient objects from the given image group, co-saliency detection has received tremendous research interest in recent years. However, as most of the existing co-saliency detection methods are performed based on the assumption that all the images in the given image group should contain co-salient objects in only one category, they can hardly be applied in practice, particularly for the large-scale image set obtained from the Internet. To address this problem, this paper revisits the co-saliency detection task and advances its development into a new phase, where the problem setting is generalized to allow the image group to contain objects in arbitrary number of categories and the algorithms need to simultaneously detect multi-class co-salient objects from such complex data. To solve this new challenge, we decompose it into two sub-problems, i.e., how to identify subgroups of relevant images and how to discover relevant co-salient objects from each subgroup, and propose a novel co-saliency detection framework to correspondingly address the two sub-problems via two-stage multi-view spectral rotation co-clustering. Comprehensive experiments on two publically available benchmarks demonstrate the effectiveness of the proposed approach. Notably, it can even outperform the state-of-the-art co-saliency detection methods, which are performed based on the image subgroups carefully separated by the human labor. Xiwen Yao, Junwei Han 0001, Dingwen Zhang, Feiping Nie 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | Revealing Event Saliency in Unconstrained Video CollectionabstractRecent progresses in multimedia event detection have enabled us to find videos about a predefined event from a large-scale video collection. Research towards more intrinsic unsupervised video understanding is an interesting but understudied field. Specifically, given a collection of videos sharing a common event of interest, the goal is to discover the salient fragments, i.e., the curt video fragments that can concisely portray the underlying event of interest, from each video. To explore this novel direction, this paper proposes an unsupervised event saliency revealing framework. It first extracts features from multiple modalities to represent each shot in the given video collection. Then, these shots are clustered to build the cluster-level event saliency revealing framework, which explores useful information cues (i.e., the intra-cluster prior, inter-cluster discriminability, and inter-cluster smoothness) by a concise optimization model. Compared with the existing methods, our approach could highlight the intrinsic stimulus of the unseen event within a video in an unsupervised fashion. Thus, it could potentially benefit to a wide range of multimedia tasks like video browsing, understanding, and search. To quantitatively verify the proposed method, we systematically compare the method to a number of baseline methods on the TRECVID benchmarks. Experimental results have demonstrated its effectiveness and efficiency. Dingwen Zhang, Junwei Han 0001, Lu Jiang 0004, Senmao Ye, Xiaojun Chang |
IEEE Trans. Image Process. | 1 |
| 2016 | Object Co-segmentation via Graph Optimized-Flexible Manifold RankingabstractAiming at automatically discovering the common objects contained in a set of relevant images and segmenting them as foreground simultaneously, object co-segmentation has become an active research topic in recent years. Although a number of approaches have been proposed to address this problem, many of them are designed with the misleading assumption, unscalable prior, or low flexibility and thus still suffer from certain limitations, which reduces their capability in the real-world scenarios. To alleviate these limitations, we propose a novel two-stage co-segmentation framework, which introduces the weak background prior to establish a globally close-loop graph to represent the common object and union background separately. Then a novel graph optimized-flexible manifold ranking algorithm is proposed to flexibly optimize the graph connection and node labels to co-segment the common objects. Experiments on three image datasets demonstrate that our method outperforms other state-of-the-art methods. Rong Quan, Junwei Han 0001, Dingwen Zhang, Feiping Nie 0001 |
CVPR | 3 |
| 2016 | Towards Intelligent Visual Understanding under Minimal Supervision
Dingwen Zhang |
IJCAI | 1 |
| 2016 | Bridging Saliency Detection to Weakly Supervised Object Detection Based on Self-Paced Curriculum Learning
Dingwen Zhang, Deyu Meng, Junwei Han 0001 |
IJCAI | 1 |
| 2016 | Detection of Co-salient Objects by Looking Deep and Wide
Dingwen Zhang, Junwei Han 0001, Chao Li 0028, Jingdong Wang 0001, Xuelong Li 0001 |
Int. J. Comput. Vis. | 1 |
| 2016 | Two-Stage Learning to Predict Human Eye Fixations via SDAEsabstractSaliency detection models aiming to quantitatively predict human eye-attended locations in the visual field have been receiving increasing research interest in recent years. Unlike traditional methods that rely on hand-designed features and contrast inference mechanisms, this paper proposes a novel framework to learn saliency detection models from raw image data using deep networks. The proposed framework mainly consists of two learning stages. At the first learning stage, we develop a stacked denoising autoencoder (SDAE) model to learn robust, representative features from raw image data under an unsupervised manner. The second learning stage aims to jointly learn optimal mechanisms to capture the intrinsic mutual patterns as the feature contrast and to integrate them for final saliency prediction. Given the input of pairs of a center patch and its surrounding patches represented by the features learned at the first stage, a SDAE network is trained under the supervision of eye fixation labels, which achieves both contrast inference and contrast integration simultaneously. Experiments on three publically available eye tracking benchmarks and the comparisons with 16 state-of-the-art approaches demonstrate the effectiveness of the proposed framework. Junwei Han 0001, Dingwen Zhang, Shifeng Wen, Lei Guo 0002, Tianming Liu 0001, Xuelong Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2016 | Cosaliency Detection Based on Intrasaliency Prior Transfer and Deep Intersaliency MiningabstractAs an interesting and emerging topic, cosaliency detection aims at simultaneously extracting common salient objects in multiple related images. It differs from the conventional saliency detection paradigm in which saliency detection for each image is determined one by one independently without taking advantage of the homogeneity in the data pool of multiple related images. In this paper, we propose a novel cosaliency detection approach using deep learning models. Two new concepts, called intrasaliency prior transfer and deep intersaliency mining, are introduced and explored in the proposed work. For the intrasaliency prior transfer, we build a stacked denoising autoencoder (SDAE) to learn the saliency prior knowledge from auxiliary annotated data sets and then transfer the learned knowledge to estimate the intrasaliency for each image in cosaliency data sets. For the deep intersaliency mining, we formulate it by using the deep reconstruction residual obtained in the highest hidden layer of a self-trained SDAE. The obtained deep intersaliency can extract more intrinsic and general hidden patterns to discover the homogeneity of cosalient objects in terms of some higher level concepts. Finally, the cosaliency maps are generated by weighted integration of the proposed intrasaliency prior, deep intersaliency, and traditional shallow intersaliency. Comprehensive experiments over diverse publicly available benchmark data sets demonstrate consistent performance gains of the proposed method over the state-of-the-art cosaliency detection methods. Dingwen Zhang, Junwei Han 0001, Jungong Han, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Predicting eye fixations using convolutional neural networksabstractIt is believed that eye movements in free-viewing of natural scenes are directed by both bottom-up visual saliency and top-down visual factors. In this paper, we propose a novel computational framework to simultaneously learn these two types of visual features from raw image data using a multiresolution convolutional neural network (Mr-CNN) for predicting eye fixations. The Mr-CNN is directly trained from image regions centered on fixation and non-fixation locations over multiple resolutions, using raw image pixels as inputs and eye fixation attributes as labels. Diverse top-down visual features can be learned in higher layers. Meanwhile bottom-up visual saliency can also be inferred via combining information over multiple resolutions. Finally, optimal integration of bottom-up and top-down cues can be learned in the last logistic regression layer to predict eye fixations. The proposed approach achieves state-of-the-art results over four publically available benchmark datasets, demonstrating the superiority of our work. Nian Liu 0002, Junwei Han 0001, Dingwen Zhang, Shifeng Wen, Tianming Liu 0001 |
CVPR | 3 |
| 2015 | Co-saliency detection via looking deep and wideabstractWith the goal of effectively identifying common and salient objects in a group of relevant images, co-saliency detection has become essential for many applications such as video foreground extraction, surveillance, image retrieval, and image annotation. In this paper, we propose a unified co-saliency detection framework by introducing two novel insights: 1) looking deep to transfer higher-level representations by using the convolutional neural network with additional adaptive layers could better reflect the properties of the co-salient objects, especially their consistency among the image group; 2) looking wide to take advantage of the visually similar neighbors beyond a certain image group could effectively suppress the influence of the common background regions when formulating the intra-group consistency. In the proposed framework, the wide and deep information are explored for the object proposal windows extracted in each image, and the co-saliency scores are calculated by integrating the intra-image contrast and intra-group consistency via a principled Bayesian formulation. Finally the window-level co-saliency scores are converted to the superpixel-level co-saliency maps through a foreground region agreement strategy. Comprehensive experiments on two benchmark datasets have demonstrated the consistent performance gain of the proposed approach. Dingwen Zhang, Junwei Han 0001, Chao Li 0028, Jingdong Wang 0001 |
CVPR | 1 |
| 2015 | A Self-Paced Multiple-Instance Learning Framework for Co-Saliency DetectionabstractAs an interesting and emerging topic, co-saliency detection aims at simultaneously extracting common salient objects in a group of images. Traditional co-saliency detection approaches rely heavily on human knowledge for designing hand-crafted metrics to explore the intrinsic patterns underlying co-salient objects. Such strategies, however, always suffer from poor generalization capability to flexibly adapt various scenarios in real applications, especially due to their lack of insightful understanding of the biological mechanisms of human visual co-attention. To alleviate this problem, we propose a novel framework for this task, by naturally reformulating it as a multiple-instance learning (MIL) problem and further integrating it into a self-paced learning (SPL) regime. The proposed framework on one hand is capable of fitting insightful metric measurements and discovering common patterns under co-salient regions in a self-learning way by MIL, and on the other hand tends to promise the learning reliability and stability by simulating the human learning process through SPL. Experiments on benchmark datasets have demonstrated the effectiveness of the proposed framework as compared with the state-of-the-arts. Dingwen Zhang, Deyu Meng, Chao Li 0028, Lu Jiang 0004, Qian Zhao 0002, Junwei Han 0001 |
ICCV | 1 |
| 2015 | Weakly Supervised Learning for Target Detection in Remote Sensing ImagesabstractIn this letter, we develop a novel framework of leveraging weakly supervised learning techniques to efficiently detect targets from remote sensing images, which enables us to reduce the tedious manual annotation for collecting training data while maintaining the detection accuracy to large extent. The proposed framework consists of a weakly supervised training procedure to yield the detectors and an effective scheme to detect targets from testing images. Comprehensive evaluations on three benchmarks which have different spatial resolutions and contain different types of targets as well as the comparisons with traditional supervised learning schemes demonstrate the efficiency and effectiveness of the proposed framework. Dingwen Zhang, Junwei Han 0001, Gong Cheng 0003, Zhenbao Liu, Shuhui Bu, Lei Guo 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | Background Prior-Based Salient Object Detection via Deep Reconstruction ResidualabstractDetection of salient objects from images is gaining increasing research interest in recent years as it can substantially facilitate a wide range of content-based multimedia applications. Based on the assumption that foreground salient regions are distinctive within a certain context, most conventional approaches rely on a number of hand-designed features and their distinctiveness is measured using local or global contrast. Although these approaches have been shown to be effective in dealing with simple images, their limited capability may cause difficulties when dealing with more complicated images. This paper proposes a novel framework for saliency detection by first modeling the background and then separating salient objects from the background. We develop stacked denoising autoencoders with deep learning architectures to model the background where latent patterns are explored and more powerful representations of data are learned in an unsupervised and bottom-up manner. Afterward, we formulate the separation of salient objects from the background as a problem of measuring reconstruction residuals of deep autoencoders. Comprehensive evaluations of three benchmark datasets and comparisons with nine state-of-the-art algorithms demonstrate the superiority of this paper. Junwei Han 0001, Dingwen Zhang, Xintao Hu, Lei Guo 0002, Jinchang Ren |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Object Detection in Optical Remote Sensing Images Based on Weakly Supervised Learning and High-Level Feature LearningabstractThe abundant spatial and contextual information provided by the advanced remote sensing technology has facilitated subsequent automatic interpretation of the optical remote sensing images (RSIs). In this paper, a novel and effective geospatial object detection framework is proposed by combining the weakly supervised learning (WSL) and high-level feature learning. First, deep Boltzmann machine is adopted to infer the spatial and structural information encoded in the low-level and middle-level features to effectively describe objects in optical RSIs. Then, a novel WSL approach is presented to object detection where the training sets require only binary labels indicating whether an image contains the target object or not. Based on the learnt high-level features, it jointly integrates saliency, intraclass compactness, and interclass separability in a Bayesian framework to initialize a set of training examples from weakly labeled images and start iterative learning of the object detector. A novel evaluation criterion is also developed to detect model drift and cease the iterative learning. Comprehensive experiments on three optical RSI data sets have demonstrated the efficacy of the proposed approach in benchmarking with several state-of-the-art supervised-learning-based object detection approaches. Junwei Han 0001, Dingwen Zhang, Gong Cheng 0003, Lei Guo 0002, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2014 | Visual attention computation in video of driving environmentabstractWe here study the problem of visual attention computation in video of driving environment via the learning from eye movements. We collect a large-scale database of eye movements from 28 subjects on 30 videos of road scenes, which simulate the driving environment. The analysis on this eye movement database reveals that visual attention in driving environment is directed by high-level cognitive factors such as objects. We then present a new high-level representation called Traffic Object Bank (TOB), which is comprised of many individual road object detectors trained comprehensively in semantic space as well as viewpoint space. TOB provides semantically rich object-level features. Finally, we develop a computational model to predict where drivers look via the mapping from TOB-based representation and to gaze data. Experimental results on our traffic scene video benchmark indicate high accordance with human eye movement and show great promise for further applications. Junwei Han 0001, Liye Sun, Dingwen Zhang, Xintao Hu, Gong Cheng 0003, Lei Guo 0002 |
ICME | 3 |
| 2014 | Saliency detection based on feature learning using Deep Boltzmann MachinesabstractSaliency detection has been a very active research area in recent years. Most traditional methods suffer from the problem that existing visual features are not discriminative or not robust enough to predict salient locations. As a result, the experimental results of these previous methods are still far from satisfactory. In this paper, we propose to utilize a two-layer Deep Boltzmann Machine (DBM) to learn enhanced features from existing contrast-based low-level features, which are more discriminative and reliable. A saliency computation model is then trained to build a mapping from those enhanced features to eye fixation data. The proposed work is amongst the earliest efforts of examining the feasibility of applying deep learning algorithms to saliency detection. Comprehensive evaluations on two publically available benchmark datasets and comparisons with a number of state-of-the-art approaches demonstrate the effectiveness of the proposed work. Shifeng Wen, Junwei Han 0001, Dingwen Zhang, Lei Guo 0002 |
ICME | 3 |