VLDB 2026 Research / reviewers in the wild / expert
Zhihui Wang 0001
dblp:65/2749-1 · also Zhi-Hui Wang 0001
· DBLP profile ↗
105ranked-venue papers
13as first author
70since 2021 · last 2026
0000-0002-5011-9726ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 65 · 5 first-author · 46 since 2021Artificial intelligence and machine learning · 33 · 2 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Computer networks · 4 · 4 since 2021Software engineering, systems software and programming languages · 4 · 4 first-authorSecurity and privacy · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reliable Multi-Prototypical Contrastive Learning for Semi-Supervised Heterogeneous and Multi-Organ Medical Image SegmentationabstractAccurate multi-organ segmentation across heterogeneous medical images is pivotal for real-world surgical navigation. The scarcity of annotation constitutes a well-established consensus in the field, prompting semi-supervised learning to emerge as a prominent solution. However, two critical bottlenecks persist in clinical translation: (1) inter-class feature ambiguity, and (2) high multi-source sample heterogeneity. To tackle these bottlenecks, we propose TP-Net, a semi-supervised framework for multi-organ segmentation in heterogeneous medical images, which innovatively integrates reliable multi-prototype contrastive learning. Specifically, we propose a heterogeneous prototype dynamic evolution mechanism that self-adaptively models fragmented intra-class distributions in multi-source data. Then, to mitigate prototype shift, we first devise an uncertainty-aware cross-domain alignment strategy grounded in the smoothness assumption, which constructs reliable prototypes by propagating reliable pixel prediction distributions from labeled to unlabeled domains. Furthermore, this is synergized with contrastive separation that enforces feature proximity to class-matched prototypes in the embedding space, effectively resolving inter-class ambiguity by minimizing overlap between adjacent organ clusters via prototype repulsion. Experimental results on two public datasets and one in-house dataset prove that the proposed method achieves state-of-the-art performance. Further, clinical validation with our self-developed surgical navigation system demonstrated the clinical viability of the proposed method. The code associated with this work will be made publicly available at https://github.com/JIESHUREN330/TP-Net/tree/main. Xiangjun Yang, Jieshu Ren, Dongpei Liu, Yi Wang 0037, Zhihui Wang 0001, Bin Liu 0040 |
IEEE Trans. Medical Imaging | 8 |
| 2026 | UniAlign: A Universal Cross-Modality Knowledge Alignment Framework for Fine-Grained Action RecognitionabstractThe key to fine-grained video action recognition is identifying subtle differences between action categories. Relying solely on visual features supervised by action labels makes it challenging to characterize robust and discriminative action dynamics from videos. With significant advancements in human pose estimation and the powerful capabilities of Vision-Language Models (VLMs), obtaining reliable and cost-free human pose data and textual semantics has become increasingly feasible, enabling their effective use in fine-grained action recognition. However, the inherent disparities in feature representations across different modalities necessitate a robust alignment strategy to achieve opti mal fusion. To address this, we propose a universal cross-modality knowledge alignment framework, namely UniAlign, to transfer the knowledge from such pre-trained multi-modal models into action recognition models. Specifically, UniAlign introduces two additional branches to extract pose features and textual semantics with the pre-trained pose encoder and VLM. To align the action relevant cues among video features, pose features, and textual semantics, we propose a Cross-Modality Similarity Aggregation module (CMSA) that utilizes the importance of different modal cues while aggregating cross-modal similarities. Additionally, we adopt a fine-tuning mechanism similar to Exponential Moving Average (EMA) to refine the textual semantics, ensuring that the semantic representations encoded by VLMs are preserved while being optimized towards the specific task preferences. Extensive experiments on widely used fine-grained action recognition benchmarks (e.g., FineGym, NTURGB-D, Diving48) and coarse-grained K400 dataset demonstrate the effectiveness of the proposed UniAlign method. Yihan Wang 0011, Baoli Sun, Xinzhu Ma, Zhihui Wang 0001, Zhiyong Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2026 | Toward an Effective Action-Region Tracking Framework for Fine-Grained Video Action RecognitionabstractFine-grained action recognition (FGAR) aims to identify subtle and distinctive differences among fine-grained action categories. However, current recognition methods often capture coarse-grained motion patterns but struggle to identify subtle details in local regions evolving over time. In this work, we introduce the action-region tracking (ART) framework, a novel solution leveraging a query-response mechanism to discover and track the dynamics of distinctive local details, enabling distinguishing similar actions effectively. Specifically, we propose a region-specific semantic activation module that employs discriminative and text-constrained semantics serve as queries to capture the most action-related region responses in each video frame, facilitating interaction among spatial and temporal dimensions with corresponding video features. The captured region responses are then organized into action tracklets, which characterize the region-based action dynamics by linking related responses across different video frames in a coherent sequence. The text-constrained queries are designed to expressly encode nuanced semantic representations derived from the textual descriptions of action labels, as extracted by the language branches within visual language models. To optimize generated action tracklets, we design a multilevel tracklet contrastive constraint among multiple region responses at spatial and temporal levels, which can effectively distinguish individual region responses in each video frame (spatial level) and establish the correlation of similar region responses between adjacent video frames (temporal level). In addition, we implement a task-specific fine-tuning mechanism to refine textual semantics during training. This ensures that the semantic representations encoded by vision language models (VLMs) are not only preserved but also optimized for specific task preferences. Comprehensive experiments on several widely used action recognition benchmarks, i.e., FineGym, Diving48, NTURGB-D, Kinetics, and Something-Something, clearly demonstrate the superiority to previous state-of-the-art baselines. Baoli Sun, Yihan Wang 0011, Xinzhu Ma, Zhihui Wang 0001, Kun Lu 0003, Zhiyong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal TokensabstractRecently, Vision Large Language Models (VLLMs) integrated with vision encoders have shown promising performance in vision understanding. The key of VLLMs is to encode visual content into sequences of visual tokens, enabling VLLMs to simultaneously process both visual and textual content. However, understanding videos, especially long videos, remain a challenge to VLLMs as the number of visual tokens grows rapidly when encoding videos, resulting in the risk of exceeding the context window of VLLMs and introducing heavy computation burden. To restrict the number of visual tokens, existing VLLMs either: (1) uniformly downsample videos into a fixed number of frames or (2) reducing the number of visual tokens encoded from each frame. We argue the former solution neglects the rich temporal cue in videos and the later overlooks the spatial details in each frame. In this work, we present Balanced-VLLM (B-VLLM): a novel VLLM framework that aims to effectively leverage task relevant spatio-temporal cues while restricting the number of visual tokens under the VLLM context window length. At the core of our method, we devise a text-conditioned adaptive frame selection module to identify frames relevant to the visual understanding task. The selected frames are then de-duplicated using a temporal frame token merging technique. The visual tokens of the selected frames are processed through a spatial token sampling module and an optional spatial token merging strategy to achieve precise control over the token count. Experimental results show that B-VLLM is effective in balancing the number of frames and visual tokens in video understanding, yielding superior performance on various video understanding benchmarks. Our code is available at https://github.com/zhuqiangLu/B-VLLM. Zhuqiang Lu, Zhenfei Yin, Mengwei He, Zhihui Wang 0001, Zhiyong Wang 0001, Kun Hu 0008 |
ICCV | 4 |
| 2025 | RobAVA: A Large-Scale Dataset and Baseline Towards Video Based Robotic Arm Action Understanding
Baoli Sun, Xinzhu Ma, Anqi Zou, Chuixuan Fan, Zhihui Wang 0001, Kun Lu 0003, Zhiyong Wang 0001 |
ICCV | 7 |
| 2025 | EchoCardMAE: Video Masked Auto-Encoders Customized for Echocardiography
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Miao Zhang 0004, Yi Wang 0037, Xin Fan 0001, Hongkai Wang 0002, Qingxiong Yue, Xiangjian He, Yen-Wei Chen 0001 |
MICCAI (13) | 4 |
| 2025 | 3DAxisPrompt: Promoting the 3D grounding and reasoning in GPT-4o
Dingning Liu, Cheng Wang 0026, Peng Gao 0007, Renrui Zhang, Xinzhu Ma, Zhihui Wang 0001 |
Neurocomputing | 7 |
| 2025 | Structure-preserving dental plaque segmentation via dynamically complementary information interaction
Rui Xu 0002, Baoli Sun, Tiantian Yan, Zhihui Wang 0001 |
Multim. Syst. | 5 |
| 2025 | Motion-guided semantic alignment for line art animation colorization
Ning Wang 0025, Hairui Yang, Hong Zhang 0011, Zhiyong Wang 0001, Zhihui Wang 0001 |
Pattern Recognit. | 6 |
| 2025 | STRobustNet: Efficient Change Detection via Spatial-Temporal Robust Representations in Remote SensingabstractVision transformers have achieved impressive performance in addressing spatial–temporal inconsistencies of change detection (CD) due to their ability to model long-range dependencies in bitemporal features. However, applying the self-attention operation directly to bitemporal features causes feature confusion and incurs high computational complexity. In this article, we propose a novel CD framework based on spatial–temporal robust representation (STRobustNet). To avoid directly applying self-attention to bitemporal features, we derive a collection of STRobustNets for land-use classes of interest. These representations are used to transform spatial–temporally inconsistent bitemporal features into consistent bitemporal classification representations, enhancing model efficiency and reducing feature confusion. Specifically, to fully perceive and integrate the varied appearances of the same class while minimizing deviation from the current input samples, we design a robust representation generation module (RRGModule), which utilizes the universal spatial–temporal context provided by the entire dataset and the specific spatial–temporal context from the current bitemporal images to generate robust representations for the interested land-use classes, improving robustness to spatial–temporal inconsistencies. Then, these representations are used to activate bitemporal features in different class channels, producing spatial–temporally consistent classification representations. The CD is ultimately performed using these consistent representations, effectively avoiding false detections caused by spatial–temporal inconsistencies. Experimental results demonstrate that STRobustNet achieves performance comparable to top-performing methods and offers the fastest inference speed among transformer-based methods. Code and pretrained models are available athttps://github.com/DLUTTengYH/STRobustNet. Hong Zhang 0011, Yuhang Teng, Zhihui Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Propagating Sparse Depth via Depth Foundation Model for Out-of-Distribution Depth CompletionabstractDepth completion is a pivotal challenge in computer vision, aiming at reconstructing the dense depth map from a sparse one, typically with a paired RGB image. Existing learning-based models rely on carefully prepared but limited data, leading to significant performance degradation in out-of-distribution (OOD) scenarios. Recent foundation models have demonstrated exceptional robustness in monocular depth estimation through large-scale training, and using such models to enhance the robustness of depth completion models is a promising solution. In this work, we propose a novel depth completion framework that leverages depth foundation models to attain remarkable robustness without large-scale training. Specifically, we leverage a depth foundation model to extract environmental cues, including structural and semantic context, from RGB images to guide the propagation of sparse depth information into missing regions. We further design a dual-space propagation approach, without any learnable parameters, to effectively propagate sparse depth in both 3D and 2D spaces to maintain geometric structure and local consistency. To refine the intricate structure, we introduce a learnable correction module to progressively adjust the depth prediction towards the real depth. We train our model on the NYUv2 and KITTI datasets as in-distribution datasets and extensively evaluate the framework on 16 other datasets. Our framework performs remarkably well in the OOD scenarios and outperforms existing state-of-the-art depth completion methods. Our models are released in https://github.com/shenglunch/PSD. Shenglun Chen, Xinzhu Ma, Hong Zhang 0011, Zhihui Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Referring Video Object Segmentation With Cross-Modality Proxy QueriesabstractReferring video object segmentation (RVOS) is an emerging cross-modality task that aims to generate pixel-level maps of the target objects referred by given textual expressions. The main concept involves learning an accurate alignment visual elements and language expressions within a semantic space. Recent approaches address cross-modality alignment through conditional queries, tracking the target object using a queryresponse based mechanism built upon transformer structure. However, they exhibit two limitations: (1) these conditional queries, identifying the same object across different frames through the same query, lack inter-frame dependency and variation modeling, making accurate target tracking challenging amid significant frame-to-frame variations; and (2) they handle the temporal feature of a video and build visual-language interaction sequentially, integrating textual constraints belatedly, which may cause the video features potentially focus on the non-referred objects. Therefore, we propose a novel RVOS architecture called ProxyFormer, which introduces a set of proxy queries to integrate visual and text semantics and facilitate the flow of semantics between them. By progressively updating and propagating proxy queries across multiple stages of video feature encoder, ProxyFormer ensures that the video features are as focused as much possible on the object of interest. This dynamic evolution of the queries across video also enables the proxy queries to establish inter-frame dependencies, enhancing the accuracy and coherence of object tracking throughout the video sequence. To mitigate the high computational costs associated with full spatio-temporal interactions between video and proxy queries, we propose to decouple cross-modality interactions into their temporal and spatial dimensions, respectively. Additionally, we design a Joint Semantic Consistency (JSC) training strategy to align semantic consensus between the proxy queries and the combined videotext pairs. Comprehensive experiments on four widely used RVOS benchmarks, i.e., Ref-Youtube-VOS, Ref-DAVIS17, A2D-Sentences and JHMDB-Sentences, clearly demonstrate the superiority of our ProxyFormer to the state-of-the-art methods Baoli Sun, Xinzhu Ma, Zhihui Wang 0001, Zhiyong Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Real-Time Depth Completion With Multimodal Feature AlignmentabstractAs a key problem in computer vision, depth completion aims to recover dense depth maps from sparse ones [generally derived from light detection and ranging (LiDAR)]. Most methods introduce synchronous RGB images and leverage multimodal fusion to integrate multimodal features from these modalities to describe the complete scene. However, their different natural characteristics lead to inconsistency in features, potentially impacting the effectiveness of multimodal feature fusion. To address this issue, we propose a feature alignment network (FANet) that introduces an alignment scheme to enhance the consistency between multimodal features. This scheme aligns the modality-invariant semantic context, which is invariant to changes in modality and represents the correlation between a pixel and its surroundings. Specifically, we first design an asymmetric context extraction (ACE) module to extract modality-invariant semantic contexts from multimodal features within limited GPU memory, and then pull them closer to improve consistency. Crucially, our alignment scheme is only applied during the training phase, and no additional computation cost is incurred in the inference phase. Moreover, we introduce a simple yet effective refinement module to refine estimated results via residual learning based on intermediate depth maps and sparse depth maps. Extensive experiments on KITTI and VOID datasets demonstrate that our method achieves competitive performance against typical real-time methods. In addition, we embed the proposed alignment scheme and refinement module into other methods to demonstrate their effectiveness. Shenglun Chen, Xinzhu Ma, Hong Zhang 0011, Baoli Sun, Zhihui Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Is a Pure Transformer Effective for Separated and Online Multi-Object Tracking?abstractRecent advances in multi-object tracking (MOT) have demonstrated significant success in short-term association within the separated tracking-by-detection online paradigm. However, long-term tracking remains challenging. While graph-based approaches address this by modeling trajectories as global graphs, these methods are unsuitable for real-time applications due to their non-online nature. In this article, we review the concept of trajectory graphs and propose a novel perspective by representing them as directed acyclic graphs. This representation can be described using frame-ordered object sequences and binary adjacency matrices. We observe that this structure naturally aligns with Transformer attention mechanisms, enabling us to model the association problem using a classic Transformer architecture. Based on this insight, we introduce a concise pure transformer (PuTR) to validate the effectiveness of Transformer in unifying short- and long-term tracking for separated online MOT. Extensive experiments on four diverse datasets (SportsMOT, DanceTrack, MOT17, and MOT20) demonstrate that PuTR effectively establishes a solid baseline compared to existing foundational online methods while exhibiting superior domain adaptation capabilities. Furthermore, the separated nature enables efficient training and inference, making it suitable for practical applications. Implementation code and trained models are available at https://github.com/chongweiliu/PuTR . Chongwei Liu, Zhihui Wang 0001, Rui Xu 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Novelty Detection Based Discriminative Multiple Instance Feature Mining to Classify NSCLC PD-L1 Status on HE-Stained Histopathological Images
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Yi Wang 0037, Hongkai Wang 0002, Dingpin Huang, Fangyi Xu, Yi Gan, Yuan Tu, Hongjie Hu |
MICCAI (4) | 5 |
| 2024 | Fine-grained Video Semantic Distillation for Video-Text Retrieval
Zuyi Pei, Baoli Sun, Zhihui Wang 0001 |
MMAsia | 3 |
| 2024 | Sketch-based 3D Model Retrieval with Cross-Modal Representation
Hairui Yang, Ning Wang 0025, Zhihui Wang 0001, Lei Wang 0005 |
MMAsia | 3 |
| 2024 | FastTrack: A Highly Efficient and Generic GPU-Based Multi-object Tracking Method with Parallel Kalman Filter
Chongwei Liu, Zhihui Wang 0001 |
Int. J. Comput. Vis. | 3 |
| 2024 | Accurate Fine-Grained Object Recognition with Structure-Driven Relation Graph Networks
Shijie Wang 0003, Zhihui Wang 0001, Jianlong Chang, Wanli Ouyang, Qi Tian 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Low-resolution few-shot learning via multi-space knowledge distillation
Xinchen Ye, Baoli Sun, Hairui Yang, Rui Xu 0002, Zhihui Wang 0001 |
Inf. Sci. | 7 |
| 2024 | Consistency-guided pseudo labeling for transductive zero-shot learning
Hairui Yang, Ning Wang 0025, Zhihui Wang 0001, Lei Wang 0005 |
Inf. Sci. | 3 |
| 2024 | Reconciling global and local optimal label assignments for heavily occluded pedestrian detection
Chongwei Liu, Zhihui Wang 0001, Rui Xu 0002 |
Multim. Syst. | 3 |
| 2024 | Application of CLIP for efficient zero-shot learning
Hairui Yang, Ning Wang 0025, Lei Wang 0005, Zhihui Wang 0001 |
Multim. Syst. | 5 |
| 2024 | Content-Aware Rectified Activation for Zero-Shot Fine-Grained Image RetrievalabstractFine-grained image retrieval mainly focuses on learning salient features from the seen subcategories as discriminative embedding while neglecting the problems behind zero-shot settings. We argue that retrieving fine-grained objects from unseen subcategories may rely on more diverse clues, which are easily restrained by the salient features learnt from seen subcategories. To address this issue, we propose a novel Content-aware Rectified Activation model, which enables this model to suppress the activation on salient regions while preserving their discrimination, and spread activation to adjacent non-salient regions, thus mining more diverse discriminative features for retrieving unseen subcategories. Specifically, we construct a content-aware rectified prototype (CARP) by perceiving semantics of salient regions. CARP acts as a channel-wise non-destructive activation upper bound and can be selectively used to suppress salient regions for obtaining the rectified features. Moreover, two regularizations are proposed: 1) a semantic coherency constraint that imposes a restriction on semantic coherency of CARP and salient regions, aiming at propagating the discriminative ability of salient regions to CARP, 2) a feature-navigated constraint to further guide the model to adaptively balance the discrimination power of rectified features and the suppression power of salient features. Experimental results on fine-grained and product retrieval benchmarks demonstrate that our method consistently outperforms the state-of-the-art methods. Shijie Wang 0003, Jianlong Chang, Zhihui Wang 0001, Wanli Ouyang, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | CustomDepth: Customizing point-wise depth categories for depth completion
Shenglun Chen, Xinchen Ye, Hong Zhang 0011, Zhihui Wang 0001 |
Pattern Recognit. Lett. | 5 |
| 2024 | Learning Pixel-Wise Continuous Depth Representation via Clustering for Depth CompletionabstractDepth completion is a long-standing challenge in computer vision, where classification-based methods have made tremendous progress in recent years. However, most existing classification-based methods rely on pre-defined pixel-shared and discrete depth values as depth categories. This representation fails to capture the continuous depth values that conform to the real depth distribution, leading to depth smearing in boundary regions. To address this issue, we revisit depth completion from the clustering perspective and propose a novel clustering-based framework called CluDe which focuses on learning the pixel-wise and continuous depth representation. The key idea of CluDe is to iteratively update the pixel-shared and discrete depth representation to its corresponding pixel-wise and continuous counterpart, driven by the real depth distribution. Specifically, CluDe first utilizes depth value clustering to learn a set of depth centers as the depth representation. While these depth centers are pixel-shared and discrete, they are more in line with the real depth distribution compared to pre-defined depth categories. Then, CluDe estimates offsets for these depth centers, enabling their dynamic adjustment along the depth axis of the depth distribution to generate the pixel-wise and continuous depth representation. Extensive experiments demonstrate that CluDe successfully reduces depth smearing around object boundaries by utilizing pixel-wise and continuous depth representation. Furthermore, CluDe achieves state-of-the-art performance on the VOID datasets and outperforms classification-based methods on the KITTI dataset. Shenglun Chen, Hong Zhang 0011, Xinzhu Ma, Zhihui Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Correlation-Guided Semantic Consistency Network for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) has raised more attention in night-time surveillance applications due to the struggle to capture valid appearance information under poor illumination conditions via visible cameras. Existing works usually separate the modality-specific and modality-irrelevant information in visible and infrared features, or project features of two modalities into a unified embedding feature space directly, which aims to eliminate huge modality discrepancies. However, these methods neglect the intra-modality and inter-modality correlations. We argue that the correlations can implicitly guide the network to discover the modality-irrelevant information, thus more beneficial for eliminating huge modality discrepancies and preserving individual differences. To this end, we propose a novel framework, termed as correlation-guided semantic consistency network (CSC-Net), to explore and exploit the intra-modality and inter-modality correlations. Specifically, CSC-Net consists of a cross-modality semantic alignment (CSA) module, a cross-granularity discrepancy awareness (CDA) module, and a probability consistency constraint (PCC) module. CSA mines the inter-modality correlation by calculating the semantic similarity between modalities to explore modality-irrelevant features, and then transfers the learned features to the backbone network to face the input of only single modality images. To preserve the individual differences, CDA sufficiently utilizes the intra-modality correlation via exploring the multi-granularity discriminative information. Finally, PCC constrains the network at the probability level, cooperating with the CSA which constrains at the feature level, to further alleviate the modality discrepancy. Extensive experiments on two public VI-ReID datasets SYSU-MM01 and RegDB have verified the effectiveness of our approach. Qijie Peng, Shijie Wang 0003, Hong Yu 0005, Zhihui Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Addressing Challenges of Incorporating Appearance Cues Into Heuristic Multi-Object Tracker via a Novel Feature ParadigmabstractIn the field of Multi-Object Tracking (MOT), the incorporation of appearance cues into tracking-by-detection heuristic trackers using re-identification (ReID) features has posed limitations on its advancement. The existing ReID paradigm involves the extraction of coarse-grained object-level feature vectors from cropped objects at a fixed input size using a ReID model, and similarity computation through a simple normalized inner product. However, MOT requires fine-grained features from different object regions and more accurate similarity measurements to identify individuals, especially in the presence of occlusion. To address these limitations, we propose a novel feature paradigm. In this paradigm, we extract the feature map from the entire frame image to preserve object sizes and represent objects using a set of fine-grained features from different object regions. These features are sampled from adaptive patches within the object bounding box on the feature map to effectively capture local appearance cues. We introduce Mutual Ratio Similarity (MRS) to accurately measure the similarity of the most discriminative region between two objects based on the sampled patches, which proves effective in handling occlusion. Moreover, we propose absolute Intersection over Union (AIoU) to consider object sizes in feature cost computation. We integrate our paradigm with advanced motion techniques to develop a heuristic Motion-Feature joint multi-object tracker, MoFe. Within it, we reformulate the track state transition of tracklets to better model their life cycle, and firstly introduce a runtime recorder after MoFe to refine trajectories. Extensive experiments on five benchmarks, i.e., GMOT-40, BDD100k, DanceTrack, MOT17, and MOT20, demonstrate that MoFe achieves state-of-the-art performance in robustness and generalizability without any fine-tuning, and even surpasses the performance of fine-tuned ReID features. Chongwei Liu, Zhihui Wang 0001, Rui Xu 0002 |
IEEE Trans. Image Process. | 3 |
| 2024 | C2ANet: Cross-Scale and Cross-Modality Aggregation Network for Scene Depth Super-ResolutionabstractExisting depth super-resolution (DSR) methods typically utilize an additional high-resolution (HR) color image of the same scene as assistance to recover the low-resolution (LR) depth map. Although these color-guided methods have achieved impressive progress, they easily face with color image under-utilization and mis-utilization issues. In this article, we deeply investigate the above problems and further propose a novel DSR framework to alleviate them. Specifically, we propose a Cross-scale and Cross-modality Aggregation Network(C$^{2}$ANet)to learn abundant and accurate complementarity from color images to help recover the degraded depth map. Our C$^{2}$ANet can simultaneously extract multi-scale representations from color images with parallel network hierarchies, and effectively aggregate cross-scale and cross-modality contexts to boost HR representations in each hierarchy. Then, to appropriately use the guided color image, we further design a Feature Aggregation Module (FAM) to adaptively select and fuse task-relevant features, which consists of (1) afeature alignment blockto learn transformation offsets and align upsampled features with targeted HR features, and (2) afeature fusion blockbased on cross-attention mechanism to maintain strong structural context and suppress texture distraction. Experimental results on synthetic and real-world benchmark datasets demonstrate the superiority of our proposed method in comparison with other state-of-the-art DSR methods. Xinchen Ye, Baoli Sun, Rui Xu 0002, Zhihui Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Transparent Depth Completion Using Segmentation FeaturesabstractEstimating the depth of transparent objects is one of the well-known challenges of RGB-D cameras due to the reflection and refraction effects. Previously, researchers propose to correct the depth of transparent objects by using their estimated segmentation masks because it is possible to recover the internal depth of an object just from its boundary, as illustrated by those depth-from-silhouette methods. However, these algorithms only use segmentation masks. They ignore the internal structure information from the mask segmentation features, which we argue are more useful for transparent depth estimation. In this work, we demonstrate the effectiveness of segmentation features for transparent object depth estimation. We show that it is even possible to recover the depth map just from segmentation features, without any RGB or depth map as input. Based on this observation, we propose DualTransNet which uses segmentation features for transparent depth completion. In our DualTransNet, we feed segmentation features from an extra module to the main network for better depth completion quality. Extensive experiments have shown the superiority of segmentation features as well as the state-of-the-art performance of our network. Boqian Liu, Zhihui Wang 0001, Tianfan Xue |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Digging into Depth and Color Spaces: A Mapping Constraint Network for Depth Super-ResolutionabstractScene depth super-resolution (DSR) poses an inherently ill-posed problem due to the extremely large space of one-to-many mapping functions from a given low-resolution (LR) depth map, which possesses limited depth information, to multiple plausible high-resolution (HR) depth maps. This characteristic renders the task highly challenging, as identifying an optimal solution becomes significantly intricate amidst this multitude of potential mappings. While simplistic constraints have been proposed to address the DSR task, the relationship between LR and HR depth maps and the color image has not been thoroughly investigated. In this paper, we introduce a novel mapping constraint network (MCNet) that incorporates additional constraints derived from both LR depth maps and color images. This integration aims to optimize the space of mapping functions and enhance the performance of DSR. Specifically, alongside the primary DSR network (DSRNet) dedicated to learning LR-to-HR mapping, we have developed an auxiliary degradation network (ADNet) that operates in reverse, generating the LR depth map from the reconstructed HR depth map to obtain depth features in LR space. To enhance the learning process of DSRNet in LR-to-HR mapping, we introduce two mapping constraints in LR space: (1) the cycle-consistent constraint, which offers additional supervision by establishing a closed loop between LR-to-HR and HR-to-LR mappings, and (2) the region-level contrastive constraint, aimed at reinforcing region-specific HR representations by explicitly modeling the consistency between LR and HR spaces. To leverage the color image effectively, we introduce a feature screening module to adaptively fuse color features at different layers, which can simultaneously maintain strong structural context and suppress texture distraction through subspace generation and image projection. Comprehensive experimental results across synthetic and real-world benchmark datasets unequivocally demonstrate the superiority of our proposed method over state-of-the-art DSR methods. Our MCNet achieves an average MAD reduction of 3.7% and 7.5% over state-of-the-art DSR method for ×8 and ×16 cases on Milddleburry dataset, respectively, without incurring additional costs during inference. Baoli Sun, Tiantian Yan, Xinchen Ye, Zhihui Wang 0001, Zhiyong Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Discriminative Segment Focus Network for Fine-grained Video Action RecognitionabstractFine-grained video action recognition aims at identifying minor and discriminative variations among fine categories of actions. While many recent action recognition methods have been proposed to better model spatio-temporal representations, how to model the interactions among discriminative atomic actions to effectively characterize inter-class and intra-class variations has been neglected, which is vital for understanding fine-grained actions. In this work, we devise a Discriminative Segment Focus Network (DSFNet) to mine the discriminability of segment correlations and localize discriminative action-relevant segments for fine-grained video action recognition. Firstly, we propose a hierarchic correlation reasoning (HCR) module which explicitly establishes correlations between different segments at multiple temporal scales and enhances each segment by exploiting the correlations with other segments. Secondly, a discriminative segment focus (DSF) module is devised to localize the most action-relevant segments from the enhanced representations of HCR by enforcing the consistency between the discriminability and the classification confidence of a given segment with a consistency constraint. Finally, these localized segment representations are combined with the global action representation of the whole video for boosting final recognition. Extensive experimental results on two fine-grained action recognition datasets, i.e., FineGym and Diving48, and two action recognition datasets, i.e., Kinetics400 and Something-Something, demonstrate the effectiveness of our approach compared with the state-of-the-art methods. Baoli Sun, Xinchen Ye, Tiantian Yan, Zhihui Wang 0001, Zhiyong Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Fine-Grained Retrieval Prompt TuningabstractFine-grained object retrieval aims to learn discriminative representation to retrieve visually similar objects. However, existing top-performing works usually impose pairwise similarities on the semantic embedding spaces or design a localization sub-network to continually fine-tune the entire model in limited data scenarios, thus resulting in convergence to suboptimal solutions. In this paper, we develop Fine-grained Retrieval Prompt Tuning (FRPT), which steers a frozen pre-trained model to perform the fine-grained retrieval task from the perspectives of sample prompting and feature adaptation. Specifically, FRPT only needs to learn fewer parameters in the prompt and adaptation instead of fine-tuning the entire model, thus solving the issue of convergence to suboptimal solutions caused by fine-tuning the entire model. Technically, a discriminative perturbation prompt (DPP) is introduced and deemed as a sample prompting process, which amplifies and even exaggerates some discriminative elements contributing to category prediction via a content-aware inhomogeneous sampling operation. In this way, DPP can make the fine-grained retrieval task aided by the perturbation prompts close to the solved task during the original pre-training. Thereby, it preserves the generalization and discrimination of representation extracted from input samples. Besides, a category-specific awareness head is proposed and regarded as feature adaptation, which removes the species discrepancies in features extracted by the pre-trained model using category-guided instance normalization. And thus, it makes the optimized features only include the discrepancies among subcategories. Extensive experiments demonstrate that our FRPT with fewer learnable parameters achieves the state-of-the-art performance on three widely-used fine-grained datasets. Shijie Wang 0003, Jianlong Chang, Zhihui Wang 0001, Wanli Ouyang, Qi Tian 0001 |
AAAI | 3 |
| 2023 | Open-Set Fine-Grained Retrieval via Prompting Vision-Language EvaluatorabstractOpen-set fine-grained retrieval is an emerging challenge that requires an extra capability to retrieve unknown subcategories during evaluation. However, current works focus on close-set visual concepts, where all the subcategories are pre-defined, and make it hard to capture discriminative knowledge from unknown subcategories, consequently failing to handle unknown subcategories in open-world scenarios. In this work, we propose a novel Prompting vision-Language Evaluator (PLEor) framework based on the recently introduced contrastive language-image pretraining (CLIP) model, for open-set fine-grained retrieval. PLEor could leverage pre-trained CLIP model to infer the discrepancies encompassing both pre-defined and unknown subcategories, called category-specific discrepancies, and transfer them to the backbone network trained in the close-set scenarios. To make pre-trained CLIP model sensitive to category-specific discrepancies, we design a dual prompt scheme to learn a vision prompt specifying the categoryspecific discrepancies, and turn random vectors with category names in a text prompt into category-specific discrepancy descriptions. Moreover, a vision-language evaluator is proposed to semantically align the vision and text prompts based on CLIP model, and reinforce each other. In addition, we propose an open-set knowledge transfer to transfer the category-specific discrepancies into the backbone network using knowledge distillation mechanism. Quantitative and qualitative experiments show that our PLEor achieves promising performance on open-set fine-grained datasets. Shijie Wang 0003, Jianlong Chang, Zhihui Wang 0001, Wanli Ouyang, Qi Tian 0001 |
CVPR | 4 |
| 2023 | Towards Fair and Comprehensive Comparisons for Image-Based 3D Object DetectionabstractIn this work, we build a modular-designed codebase, formulate strong training recipes, design an error diagnosis toolbox, and discuss current methods for image-based 3D object detection. In particular, different from other highly mature tasks, e.g., 2D object detection, the community of image-based 3D object detection is still evolving, where methods often adopt different training recipes and tricks resulting in unfair evaluations and comparisons. What is worse, these tricks may overwhelm their proposed designs in performance, even leading to wrong conclusions. To address this issue, we build a module-designed code-base and formulate unified training standards for the community. Furthermore, we also design an error diagnosis toolbox to measure the detailed characterization of detection models. Using these tools, we analyze current methods in-depth under varying settings and provide discussions for some open questions, e.g., discrepancies in conclusions on KITTI-3D and nuScenes datasets, which have led to different dominant methods for these datasets. We hope that this work will facilitate future research in image-based 3D object detection. Our codes will be released at https://github.com/OpenGVLab/3dodi. Xinzhu Ma, Yongtao Wang, Yinmin Zhang, Zhiyi Xia, Zhihui Wang 0001, Wanli Ouyang |
ICCV | 6 |
| 2023 | Denser is Better:cost distribution super-resolution network for more accurate sub-pixel disparityabstractThe low-quality cost distribution obtained by simple upsampling leads to disparity maps with many outliers and low sub-pixel accuracy. We propose the Cost Distribution Super-Resolution Network (CDSRNet), which directly extracts high-resolution cost distribution from the low-resolution 4D cost volume.The similarity extraction module of CDSRNet decomposes the task of estimating high-resolution cost distribution into multiple subtasks and completes each subtask by a specific translation block, ensuring high discrimination of the predicted cost distribution. The inter-block aggregation module aggregates information from other subtasks, obtaining a voting volume with global information for correcting errors in the current subtask. Experiment results demonstrate that the proposed method significantly reduces the ratio of outliers and improves the sub-pixel accuracy. Hong Zhang 0011, Shenglun Chen, Zhihui Wang 0001, Wanli Ouyang |
ICME | 3 |
| 2023 | Online Visual SLAM Adaptation against Catastrophic Forgetting with Cycle-Consistent Contrastive LearningabstractVisual SLAM (Simultaneous Localisation and Mapping) aims to simultaneously estimate camera poses and depth maps from navigation videos captured. While recent deep learning based methods have achieved great success on this task, they tend to work well on source domain data and suffer from performance degradation on the unseen data of target domain. Hence, we propose an online adaptation approach to continuously adapt a pre-trained visual SLAM model to changing environments in a self-supervised manner. To preserve pre-learned knowledge against catastrophic forgetting, we perform updating on a novel adapter proposed rather than fine-tuning the whole model for adaptation. The adapter includes a cross-domain feature translation module that translates pre-learned features into translated features suitable for adaptation. Ideally, the translated new features should not only contain pre-learned knowledge but also substantially distinct from pre-learned features since these two features represent different domains. We thus introduce cycle-consistent contrastive learning to maximize the dissimilarity between these two features by enlarging the distance between them in the feature space. Besides, our contrastive learning method exploiting cycle-consistency contraint enables the translated features to be transferred back to the pre-learned ones, which helps the translated features better preserve pre-learned knowledge. Comprehensive experiments on both synthetic and real-world datasets demonstrate superior adaptation performance of our proposed method over several state-of-the-art baselines. Sangni Xu, Hao Xiong 0001, Qiuxia Wu, Zhihui Wang 0001, Zhiyong Wang 0001 |
ICRA | 5 |
| 2023 | Exploring Coarse-to-Fine Action Token Localization and Interaction for Fine-grained Video Action RecognitionabstractVision transformers have achieved impressive performance for video action recognition due to their strong capability of modeling long-range dependencies among spatio-temporal tokens. However, as for fine-grained actions, subtle and discriminative differences mainly exist in the regions of actors, directly utilizing vision transformers without removing irrelevant tokens will compromise recognition performance and lead to high computational costs. In this paper, we propose a coarse-to-fine action token localization and interaction network, namely C2F-ALIN, that dynamically localizes the most informative tokens at a coarse granularity and then partitions these located tokens to a fine granularity for sufficient fine-grained spatio-temporal interaction. Specifically, in the coarse stage, we devise a discriminative token localization module to accurately identify informative tokens and to discard irrelevant tokens, where each localized token corresponds to a large spatial region, thus effectively preserving the continuity of action regions.In the fine stage, we only further partition the localized tokens obtained in the coarse stage into a finer granularity and then characterize fine-grained token interactions in two aspects: (1) first using vanilla transformers to learn compact dependencies among all discriminative tokens; and (2) proposing a global contextual interaction module which enables each fine-grained tokens to communicate with all the spatio-temporal tokens and to embed the global context. As a result, our coarse-to-fine strategy is able to identify more relevant tokens and integrate global context for high recognition accuracy while maintaining high efficiency.Comprehensive experimental results on four widely used action recognition benchmarks, including FineGym, Diving48, Kinetics and Something-Something, clearly demonstrate the advantages of our proposed method in comparison with other state-of-the-art ones. Baoli Sun, Xinchen Ye, Zhihui Wang 0001, Zhiyong Wang 0001 |
ACM Multimedia | 3 |
| 2023 | Learning to Parameterize Visual Attributes for Open-set Fine-grained RetrievalabstractOpen-set fine-grained retrieval is an emerging challenging task that allows to retrieve unknown categories beyond the training set.
The best solution for handling unknown categories is to represent them using a set of visual attributes learnt from known categories, as widely used in zero-shot learning. Though important, attribute modeling usually requires significant manual annotations and thus is labor-intensive. Therefore, it is worth to investigate how to transform retrieval models trained by image-level supervision from category semantic extraction to attribute modeling. To this end, we propose a novel Visual Attribute Parameterization Network (VAPNet) to learn visual attributes from known categories and parameterize them into the retrieval model, without the involvement of any attribute annotations.
In this way, VAPNet could utilize its parameters to parse a set of visual attributes from unknown categories and precisely represent them.
Technically, VAPNet explicitly attains some semantics with rich details via making use of local image patches and distills the visual attributes from these discovered semantics. Additionally, it integrates the online refinement of these visual attributes into the training process to iteratively enhance their quality. Simultaneously, VAPNet treats these attributes as supervisory signals to tune the retrieval models, thereby achieving attribute parameterization. Extensive experiments on open-set fine-grained retrieval datasets validate the superior performance of our VAPNet over existing solutions. Shijie Wang 0003, Jianlong Chang, Zhihui Wang 0001, Wanli Ouyang, Qi Tian 0001 |
NeurIPS | 4 |
| 2023 | Feature enhancement network for stereo matching
Shenglun Chen, Hong Zhang 0011, Baoli Sun, Xinchen Ye, Zhihui Wang 0001 |
Image Vis. Comput. | 6 |
| 2023 | Importance filtered soft label-based deep adaptation network
Wei Wang 0335, Mengzhu Wang, Zhihui Wang 0001 |
Knowl. Based Syst. | 5 |
| 2023 | Face attribute recognition via end-to-end weakly supervised regional location
Geng Sun 0001, Zhihui Wang 0001 |
Multim. Syst. | 4 |
| 2023 | Coloring anime line art videos with transformation region enhancement networkabstractAutomatic colorization of anime line art videos aims to produce color frames given line art frames and reference color images, which is challenging due to various motions and geometric transformations across frame sequences. Existing methods usually utilize the feature maps of reference images directly and treat all the regions in an image equally. However, this may overlook the details of the regions undergoing geometric transformations . To emphasize the regions with significant transformations between the reference and target frames, we propose a Transformation Region Enhancement Network (TRE-Net) to exploit useful reference information and enhance the colorization of key transformation regions with Region Localization Module (RLM) and Feature Enhancement Module (FEM). Specifically, we propose Multi-scale Euclidean Distance Difference (Multi-scale EDD) Maps in RLM which effectively locate geometric transformation regions by contrasting the Euclidean Distance Maps of two line arts and aggregating representations at multiple scales of the network. In addition, FEM is devised to enhance feature learning in the regions with geometric transformation and to ensure proper color alignment. FEM learns locally enhanced features through an attention-gating operation at a low computational cost. With the well-represented key geometric transformation regions, our method exploits the multi-scale reference information well for color alignment, thus produces perceptually pleasing frames. Comprehensive experimental results show that our proposed method is superior to existing methods in terms of the overall quality of colorized anime line art videos. Ning Wang 0020, Muyao Niu, Zhi Dou, Zhihui Wang 0001, Zhiyong Wang 0001, Zhaoyan Ming, Bin Liu 0040 |
Pattern Recognit. | 4 |
| 2023 | Semantic-Guided Information Alignment Network for Fine-Grained Image RecognitionabstractExisting fine-grained image recognition works have attempted to dig into low-level details for emphasizing subtle discrepancies among sub-categories. However, a potential limitation of these methods is that they integrate the low-level details and high-level semantics directly, and neglect their content complementarity and spatial corresponding correlation. To handle this limitation, we propose an end-to-end Semantic-guided Information Alignment Network (SIA-Net) to dynamically pick out the low-level details under the guidance of accurate semantics to make selected details spatially corresponding to high-level semantics and complementary in content. Technically, SIA-Net consists of an Accurate Semantic Calibration (ASC) module for providing accurate semantics and a Discriminative Feature Alignment (DFA) module for aggregating low-level details and high-level semantics using accurate semantics generated by ASC. ASC learns the pixel-level feature shifting caused by convolutional operations, which is utilized for replacing the incorrectly highlighted semantics by shifting discriminative semantics or background features. After obtaining the accurate semantic features, DFA digs into the complementary details and simultaneously makes the selected details spatially corresponding via applying the guidance of accurate semantics to obtain the reassembly features. Finally, the reassembly features, which serve as discriminative cues, are used for more accurate discriminative region localization. Extensive experiments verify that our proposed method yields the best performance under the same settings with the most competitive approaches on CUB-birds, Stanford-Cars, and FGVC Aircraft datasets. Shijie Wang 0003, Zhihui Wang 0001, Jianlong Chang, Wanli Ouyang, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Iterative Class Prototype Calibration for Transductive Zero-Shot LearningabstractZero-shot learning (ZSL) typically suffers from the domain shift issue since the projected feature embedding of unseen samples mismatch with the corresponding class semantic prototypes, making it very challenging to fine-tune an optimal visual-semantic mapping for the unseen domain. Some existing transductive ZSL methods solve this problem by introducing unlabeled samples of the unseen domain, in which the projected features of unseen samples are still not discriminative and tend to be distributed around prototypes of seen classes. Therefore, how to effectively align the projection features of samples in unseen classes with corresponding predefined class prototypes is crucial for promoting the generalization of ZSL models. In this paper, we propose a novel Iterative Class Prototype Calibration (ICPC) framework for transductive ZSL which consists of a pseudo-labeling stage and a model retraining stage to address the above key issue. First, in the labeling stage, we devise a Class Prototype Calibration (CPC) module to calibrate the predefined class prototypes of the unseen domain by estimating the real center of projected feature distribution, which achieves better matching of sample points and class prototypes. Next, in the retraining stage, we devise a Certain Samples Screening (CSS) module to select relatively certain unseen samples with high confidence and align them with predefined class prototypes in the embedding space. A progressive training strategy is adopted to select more certain samples and update the proposed model with augmented training data. Extensive experiments on AwA2, CUB, and SUN datasets demonstrate that the proposed scheme achieves new state-of-the-art in the conventional setting under both standard split (SS) and proposed split (PS). Hairui Yang, Baoli Sun, Baopu Li, Caifei Yang, Zhihui Wang 0001, Jenhui Chen, Lei Wang 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Underwater Depth Estimation via Stereo Adaptation NetworksabstractWith the fast development and wide application of stereo depth estimation, adequate high-quality stereo training data with groundtruth depth information plays an important role, but is not easily acquired in underwater environments. Therefore, satisfactory performance of depth estimation is difficult to achieve in underwater environments. In addition, the domain gap also leads to the failure of directly applying existing models of terrestrial scene to underwater scene. Therefore, this paper proposes a novel underwater depth estimation network which can infer depth maps from real underwater stereo images in an adaptation manner. The proposed learning pipeline mainly contains three different adaptation modules, i.e., style adaptation, semantic adaptation and disparity range adaptation, to progressively adapt a terrestrial depth estimation model to the underwater domain. Specifically, due to the lack of underwater training data, we first propose a depth-aware stereo image translation network to synthesize stylized underwater stereo images from terrestrial dataset, thus benefiting the effective training of depth estimation network. Then, considering the weak generalization to the real underwater data when only trained on the above synthetic data, we present a self-ensembling semantic adaptation for depth estimation network to minimize the semantic domain discrepancy between synthetic and real underwater data. Meanwhile, we design a disparity range adaptation module to address the problem of disparity range miss-match between both data, thus obtaining more accurate depth predictions for large-disparity-span underwater images. Experimental results show that by integrating the proposed adaptation modules into the off-the-shelf depth estimation backbones, our method successfully achieves superior performance of underwater depth estimation compared to other state-of-the-art methods. Xinchen Ye, Yazhi Yuan, Rui Xu 0002, Zhihui Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Region Assisted Sketch ColorizationabstractAutomatic sketch colorization is a challenging task that aims to generate a color image from a sketch, primarily due to its inherently ill-posed nature. While many approaches have shown promising results, two significant challenges remain: limited color patterns and a wide range of artifacts such as color bleeding and semantic inconsistencies among relevant regions. These issues stem from the operation of traditional convolutional structures, which capture structural features in a pixel-wise manner, resulting in inadequate utilization of regional information within the sketch. Therefore, we propose the Region-Assisted Sketch Coloring (RASC) method, which introduces an intermediate representation called the 'Region Map' to explicitly characterize the regional information of the sketch. This Region Map is derived from the input sketch and is effectively formulated by our RASC architecture, enhancing the perception of region-wise features beyond the original pixel-wise features. Specifically, we start by employing the sketch encoder to extract hierarchical feature maps from the input sketches. Subsequently, we introduce a coarse-to-fine decoder comprising a series of Region-based Modulation (RM) blocks. This decoder modulates features that combine the modulation results of its previous block and the sketch features of the corresponding encoder block with our Region Formulation module. Each module explicitly formulates the sketch features in a region-wise manner. This accurately captures both the inner-region local style and inter-region global context dependency, resulting in various color patterns and fewer synthesis artifacts. Our experimental results show that our proposed method surpasses state-of-the-art methods in both synthetic and real sketch datasets. Ning Wang 0025, Muyao Niu, Zhihui Wang 0001, Kun Hu 0008, Bin Liu 0040, Zhiyong Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Semantic Decomposition Network With Contrastive and Structural Constraints for Dental Plaque SegmentationabstractSegmenting dental plaque from images of medical reagent staining provides valuable information for diagnosis and the determination of follow-up treatment plan. However, accurate dental plaque segmentation is a challenging task that requires identifying teeth and dental plaque subjected to semantic-blur regions (i.e., confused boundaries in border regions between teeth and dental plaque) and complex variations of instance shapes, which are not fully addressed by existing methods. Therefore, we propose a semantic decomposition network (SDNet) that introduces two single-task branches to separately address the segmentation of teeth and dental plaque and designs additional constraints to learn category-specific features for each branch, thus facilitating the semantic decomposition and improving the performance of dental plaque segmentation. Specifically, SDNet learns two separate segmentation branches for teeth and dental plaque in a divide-and-conquer manner to decouple the entangled relation between them. Each branch that specifies a category tends to yield accurate segmentation. To help these two branches better focus on category-specific features, two constraint modules are further proposed: 1) contrastive constraint module (CCM) to learn discriminative feature representations by maximizing the distance between different category representations, so as to reduce the negative impact of semantic-blur regions on feature extraction; 2) structural constraint module (SCM) to provide complete structural information for dental plaque of various shapes by the supervision of an boundary-aware geometric constraint. Besides, we construct a large-scale open-source Stained Dental Plaque Segmentation dataset (SDPSeg), which provides high-quality annotations for teeth and dental plaque. Experimental results on SDPSeg datasets show SDNet achieves state-of-the-art performance. Baoli Sun, Xinchen Ye, Zhihui Wang 0001, Xiaolong Luo, Heli Gao |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Rethinking Maximum Mean Discrepancy for Visual Domain AdaptationabstractExisting domain adaptation approaches often try to reduce distribution difference between source and target domains and respect domain-specific discriminative structures by some distribution [e.g., maximum mean discrepancy (MMD)] and discriminative distances (e.g., intra-class and inter-class distances). However, they usually consider these losses together and trade off their relative importance by estimating parameters empirically. It is still under insufficient exploration so far to deeply study their relationships to each other so that we cannot manipulate them correctly and the model's performance degrades. To this end, this article theoretically proves two essential facts: 1) minimizing MMD equals to jointly minimizing their data variance with some implicit weights but, respectively, maximizing the source and target intra-class distances so that feature discriminability degrades and 2) the relationship between intra-class and inter-class distances is as one falls and another rises. Based on this, we propose a novel discriminative MMD with two parallel strategies to correctly restrain the degradation of feature discriminability or the expansion of intra-class distance; specifically: 1) we directly impose a tradeoff parameter on the intra-class distance that is implicit in the MMD according to 1) and 2) we reformulate the inter-class distance with special weights that are analogical to those implicit ones in the MMD and maximizing it can also lead to the intra-class distance falling according to 2). Notably, we do not consider the two strategies in one model due to 2). The experiments on several benchmark datasets not only prove the validity of our revealed theoretical results but also demonstrate that the proposed approach could perform better than some compared state-of-art methods substantially. Our preliminary MATLAB code will be available at https://github.com/WWLoveTransfer/. Wei Wang 0335, Zhengming Ding, Feiping Nie 0001, Junyang Chen 0001, Zhihui Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2022 | Category-Specific Nuance Exploration Network for Fine-Grained Object RetrievalabstractEmploying additional prior knowledge to model local features as a final fine-grained object representation has become a trend for fine-grained object retrieval (FGOR). A potential limitation of these methods is that they only focus on common parts across the dataset (e.g. head, body or even leg) by introducing additional prior knowledge, but the retrieval of a fine-grained object may rely on category-specific nuances that contribute to category prediction. To handle this limitation, we propose an end-to-end Category-specific Nuance Exploration Network (CNENet) that elaborately discovers category-specific nuances that contribute to category prediction, and semantically aligns these nuances grouped by subcategory without any additional prior knowledge, to directly emphasize the discrepancy among subcategories. Specifically, we design a Nuance Modelling Module that adaptively predicts a group of category-specific response (CARE) maps via implicitly digging into category-specific nuances, specifying the locations and scales for category-specific nuances. Upon this, two nuance regularizations are proposed: 1) semantic discrete loss that forces each CARE map to attend to different spatial regions to capture diverse nuances; 2) semantic alignment loss that constructs a consistent semantic correspondence for each CARE map of the same order with the same subcategory via guaranteeing each instance and its transformed counterpart to be spatially aligned. Moreover, we propose a Nuance Expansion Module, which exploits context appearance information of discovered nuances and refines the prediction of current nuance by its similar neighbors, leading to further improvement on nuance consistency and completeness. Extensive experiments validate that our CNENet consistently yields the best performance under the same settings against most competitive approaches on CUB Birds, Stanford Cars, and FGVC Aircraft datasets. Shijie Wang 0003, Zhihui Wang 0001, Wanli Ouyang |
AAAI | 2 |
| 2022 | MonoDistill: Learning Spatial Features for Monocular 3D Object Detection
Zhiyu Chong, Xinzhu Ma, Hong Zhang 0011, Yuxin Yue, Zhihui Wang 0001, Wanli Ouyang |
ICLR | 6 |
| 2022 | Location-Aware Feature Selection Network for Multi-Oriented Scene Text DetectionabstractDirect Regression-based text detection methods have already achieved promising performances with simple network structure. However, there still leaves improvement room in terms of regression accuracy, especially for long and large text instances, which hinders their applicability in some realistic scenarios that need recognition of the detected text. To address this issue, we propose a novel Location Aware feature Selection text detection Network (LASNet) which selects suitable features from different locations to separately predict each component of a bounding box and gets the final bounding box through the combination of them. 1As a result, LASNet predicts the more accurate bounding boxes by effectively making use of features with a learnable feature selection way. The experimental results demonstrate that our LASNet achieves state-of-the-art performance with single model and single-scale testing. Zengyuan Guo, Panliang Fang, Zhihui Wang 0001, Wen Gao 0001 |
ICME | 4 |
| 2022 | Local-Region and Cross-Dataset Contrastive Learning for Retinal Vessel Segmentation
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Yen-Wei Chen 0001 |
MICCAI (2) | 5 |
| 2022 | Low-Dose CT Reconstruction via Dual-Domain Learning and Controllable Modulation
Xinchen Ye, Rui Xu 0002, Zhihui Wang 0001 |
MICCAI (6) | 4 |
| 2022 | Fine-grained Action Recognition with Robust Motion Representation Decoupling and ConcentrationabstractFine-grained action recognition is a challenging task that requires identifying discriminative and subtle motion variations among fine-grained action classes. Existing methods typically focus on spatio-temporal feature extraction and long-temporal modeling to characterize complex spatio-temporal patterns of fine-grained actions. However, the learned spatio-temporal features without explicit motion modeling may emphasize more on visual appearance than on motion, which could compromise the learning of effective motion features required for fine-grained temporal reasoning. Therefore, how to decouple robust motion representations from the spatio-temporal features and further effectively leverage them to enhance the learning of discriminative features still remains less explored, which is crucial for fine-grained action recognition. In this paper, we propose a motion representation decoupling and concentration network (MDCNet) to address these two key issues. First, we devise a motion representation decoupling (MRD) module to disentangle the spatio-temporal representation into appearance and motion features through contrastive learning from video and segment views. Next, in the proposed motion representation concentration (MRC) module, the decoupled motion representations are further leveraged to learn a universal motion prototype shared across all the instances of each action class. Finally, we project the decoupled motion features onto all the motion prototypes through semantic relations to obtain the concentrated action-relevant features for each action class, which can effectively characterize the temporal distinctions of fine-grained actions for improved recognition performance. Comprehensive experimental results on four widely used action recognition benchmarks, i.e., FineGym, Diving48, Kinetics400 and Something-Something, clearly demonstrate the superiority of our proposed method in comparison with other state-of-the-art ones. Baoli Sun, Xinchen Ye, Tiantian Yan, Zhihui Wang 0001, Zhiyong Wang 0001 |
ACM Multimedia | 4 |
| 2022 | From coarse to fine: multi-level feature fusion network for fine-grained image retrieval
Shijie Wang 0003, Zhihui Wang 0001, Ning Wang 0025 |
Multim. Syst. | 2 |
| 2022 | Sequential learning for sketch-based 3D model retrieval
Hairui Yang, Yu Tian 0014, Caifei Yang, Zhihui Wang 0001, Lei Wang 0005 |
Multim. Syst. | 4 |
| 2022 | Discriminative information restoration and extraction for weakly supervised low-resolution fine-grained image recognition
Tiantian Yan, Zhongxuan Luo, Zhihui Wang 0001 |
Pattern Recognit. | 5 |
| 2022 | A New Dataset, Poisson GAN and AquaNet for Underwater Object GrabbingabstractTo boost the object grabbing capability of underwater robots for open-sea farming, we propose a new dataset (UDD) consisting of three categories (seacucumber, seaurchin, and scallop) with 2,227 images. To the best of our knowledge, it is the first 4K HD dataset collected in a real open-sea farm. We also propose a novel Poisson-blending Generative Adversarial Network (Poisson GAN) and an efficient object detection network (AquaNet) to address two common issues within related datasets: the class-imbalance problem and the problem of mass small object, respectively. Specifically, Poisson GAN combines Poisson blending into its generator and employs a new loss called Dual Restriction loss (DR loss), which supervises both implicit space features and image-level features during training to generate more realistic images. By utilizing Poisson GAN, objects of minority class like seacucumber or scallop could be added into an image naturally and annotated automatically, which could increase the loss of minority classes during training detectors to eliminate the class-imbalance problem; AquaNet is a high-efficiency detector to address the problem of detecting mass small objects from cloudy underwater pictures. Within it, we design two efficient components: a depth-wise-convolution-based Multi-scale Contextual Features Fusion (MFF) block and a Multi-scale Blursampling (MBP) module to reduce the parameters of the network to 1.3 million. Both two components could provide multi-scale features of small objects under a short backbone configuration without any loss of accuracy. In addition, we construct a large-scale augmented dataset (AUDD) and a pre-training dataset via Poisson GAN from UDD. Extensive experiments show the effectiveness of the proposed Poisson GAN, AquaNet, UDD, AUDD, and pre-training dataset. Chongwei Liu, Zhihui Wang 0001, Shijie Wang 0003, Yulong Tao, Caifei Yang, Xing Liu 0005, Xin Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Joint Adaptive Dual Graph and Feature Selection for Domain AdaptationabstractDomain adaptation aims to exploit domain-invariant features by aligning the cross-domain distributions in the manifold subspace for applying the classifier trained on the source domain to the target domain. However, two limitations may still deteriorate their performances: (1) the influences of noisy or irrelevant features in the original feature space are ignored, which may unexpectedly hurt the classification of target samples; (2) the graph constructed directly in the original data space cannot accurately capture the inherent local manifold structures of high-dimensional data due to the curse of dimensionality, which may seriously mislead the transferable features learning. In this paper, we propose a novel approach to address these problems, referred to as joint Adaptive Dual Graph and Feature Selection for domain adaptation (ADGFS). Specifically, feature selection can characterize the relative importance of different features through a scaling factor, which enables ADGFS to not only reduce the impacts of noisy or irrelevant features on knowledge transfer but also learn informative domain-invariant features. Meanwhile, ADGFS adaptively optimizes the dual graph by learning the similarity matrices of both instance-level and feature-level graphs in the projected low-dimensional manifold subspace rather than the original high-dimensional space, such that the intrinsic local manifold structures of data can be captured precisely. Moreover, ADGFS simultaneously aligns the marginal and conditional probability distributions in the nonnegative matrix factorization framework to narrow the distribution discrepancies between the two different domains, which can adequately transfer knowledge from the source domain to the target domain. Comprehensive experiments on four benchmark datasets can demonstrate that the effectiveness of the proposed approach in cross-domain image classification. Jing Sun 0012, Zhihui Wang 0001, Wei Wang 0335, Fuming Sun, Zhengming Ding |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Confidence Regularized Label Propagation Based Domain AdaptationabstractIn domain adaptation (DA), label-induced losses generally occupy a dominant position and most previous models regard hard or soft labels as their inputs. However, these two types of labels may mislead the modeling process of label-induced losses since hard label is sensitive to a wrongly-predicted sample while soft label may introduce label noise, thus they may cause negative transfer. To relieve this problem, we propose a novel label learning approach namely confidence regularized label propagation (CRLP) that regularizes the confidence of predicted soft labels with constraints of F-norm or L21-norm. It is validated that maximizing either one of these two constraints equals to minimizing entropy loss. Specially, we illustrate that L21-norm is more suitable for DA than F-norm when the dataset contain a large number of categories. Then, we leverage the regularized soft labels produced by CRLP to reformulate some popular label-induced losses that consider feature transferability and discriminability such as class-wise maximum mean discrepancy, intra-class compactness and inter-class dispersion in a probability manner to present a novel DA method (i.e., CRLP-DA). Comprehensive analysis and experiments on four cross-domain object recognition datasets verify that the proposed CRLP-DA outperforms some state-of-the-art methods, especially 59.5% for Office10+Caltech10 dataset with SURF features. For others to better reproduce, our preliminary Matlab code will be available athttps://github.com/WWLoveTransfer/CRLP-DA/. Wei Wang 0335, Baopu Li, Mengzhu Wang, Feiping Nie 0001, Zhihui Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Discriminative Feature Mining and Enhancement Network for Low-Resolution Fine-Grained Image RecognitionabstractExisting fine-grained image recognition methods are difficult to learn complete discriminative features from low-resolution (LR) data, because the original subtle inter-class distinctions become slimmer with the reduction of the image resolution. Besides, existing methods of LR fine-grained image recognition and general LR image recognition only consider the restoration and extraction of global discriminative features, ignoring unreliable local fine-grained details can be detrimental to final recognition. To address the above problems, we propose a multi-tasking framework, discriminative feature mining and enhancement network (DME-Net), for the LR fine-grained image recognition task, which aims to capture the reliable object descriptions from macro and micro perspectives, respectively. Macroscopically, we train the framework’s ability to recover and extract global discriminative features based on the whole images. Microscopically, we purposefully reinforce the framework’s ability to repair and capture the local discriminative details on the mined informative parts. To precisely excavate the most potential parts, we design an informative part mining (IPM) module, in which we firstly employ a part generation layer to predict several part masks that focus on different discriminative parts under the guidance of discrepancy loss and discriminant loss. Then we introduce a part selection (PS) submodule to further screen out a group of most informative parts from the predicted part masks according to their corresponding scores, which measure the semantic correlation degree of each part to the others. Experimental results on three benchmark datasets and one retail product dataset consistently show that our proposed framework can significantly boost the performance of the baseline model. Besides, extensive ablation studies are conducted, which further prove the effectiveness of each component of our designs. Tiantian Yan, Baoli Sun, Zhihui Wang 0001, Zhongxuan Luo |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | The Farther the Better: Balanced Stereo Matching via Depth-Based Sampling and Adaptive Feature RefinementabstractExisting stereo matching methods achieve satisfactory average accuracy on a whole predicted disparity map under common global metrics, but ignore the fine-grained performance at region level, especially for far regions in the scene, which is more crucial in actual auto-driving scenarios. There are two factors accounting for this problem: 1) Depth resolution. Existing methods use disparity-based sampling to extract matching candidates uniformly according to the disparity range, but leads to sparser sampling density at far regions than that of close regions in terms of depth range, resulting in low depth resolution in far regions. 2) Feature discriminability. Limited image resolution and inferior feature extraction at far regions result in the obtained features with low discrimination, which influences the subsequent matching process between stereo images. To improve the estimation accuracy of far regions and thus achieve a balanced performance at region level, we design a novel two-stage Balanced Stereo Matching Network (BSMNet) to address the above problems. The coarse stage of BSMNet introduces a direct depth-based sampling strategy, which generates matching candidates according to scene depth instead of disparity, thus improving the depth resolution and obtaining initial depth map with more balanced accuracy. Then, a depth refinement stage is proposed to solve the problem of low feature discriminability and further optimizes the initial depth map obtained from the coarse stage. It selects matching candidates and computes their similarity scores from a carefully designed adaptive feature volume guided by a learnable scale map, thus making the final estimation more accurate. Different from the existing methods that construct stereo matching based on disparity prediction, our proposed pipeline is to directly optimize on depth information. Experiments show that our BSMNet can obtain an obvious performance improvement at far regions without discarding that at close regions, so as to largely outperform existing state-of-the-art methods. Hong Zhang 0011, Xinchen Ye, Shenglun Chen, Zhihui Wang 0001, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Dynamic Position-aware Network for Fine-grained Image RecognitionabstractMost weakly supervised fine-grained image recognition (WFGIR) approaches predominantly focus on learning the discriminative details which contain the visual variances and position clues. The position clues can be indirectly learnt by utilizing context information of discriminative visual content. However, this will cause the selected discriminative regions containing some non-discriminative information introduced by the position clues. These analysis motivate us to directly introduce position clues into visual content to only focus on the visual variances, achieving more precise discriminative region localization. Though important, position modelling usually requires significant pixel/region annotations and therefore is labor-intensive. To address this issue, we propose an end-to-end Dynamic Position-aware Network (DP-Net) to directly incorporate the position clues into visual content and dynamically align them without extra annotations, which eliminates the effect of position information for visual variances of subcategories. In particular, the DP-Net consists of: 1) Position Encoding Module, which learns a set of position-aware parts by directly adding the learnable position information into the horizontal/vertical visual content of images; 2) Position-vision Aligning Module, which dynamically aligns both visual content and learnable position information via performing graph convolution on position-aware parts; 3) Position-vision Reorganization Module, which projects the aligned position clues and visual content into the Euclidean space to construct a position-aware feature maps. Finally, the position-aware feature maps are used which is implicitly applied the aligned visual content and position clues for more accurate discriminative regions localization. Extensive experiments verify that DP-Net yields the best performance under the same settings with most competitive approaches, on CUB Bird, Stanford-Cars, and FGVC Aircraft datasets. Shijie Wang 0003, Zhihui Wang 0001, Wanli Ouyang |
AAAI | 3 |
| 2021 | Video Action Recognition with Neural Architecture SearchabstractRecently, deep convolutional neural networks have been widely used in the field of videoaction recognition. Current approaches tend to concentrate on the structure design fordifferent backbone networks, but what kind of network structures can process video botheffectively and quickly still remains to be solved despite the encouraging progress. With thehelp of neural architecture search (NAS), we search for three hyperparameters in the videoprocessing network, which are the number of frames, the number of layers per residual stageand the channel number for all layers. We relax the entire search space into a continuoussearch space, and search for a set of network architectures that balance accuracy andcomputational efficiency by considering accuracy as the primary optimization goal andcomputational complexity as the secondary optimization goal. We conduct experiments onUCF101 and Kinetics400 datasets, validating new state-of-the-art results of the proposedNAS based scheme for video action recognition. Yuanding Zhou, Baopu Li, Zhihui Wang 0001 |
ACML | 3 |
| 2021 | Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-ResolutionabstractExisting color-guided depth super-resolution (DSR) approaches require paired RGB-D data as training samples where the RGB image is used as structural guidance to recover the degraded depth map due to their geometrical similarity. However, the paired data may be limited or expensive to be collected in actual testing environment. Therefore, we explore for the first time to learn the cross-modality knowledge at training stage, where both RGB and depth modalities are available, but test on the target dataset, where only single depth modality exists. Our key idea is to distill the knowledge of scene structural guidance from RGB modality to the single DSR task without changing its network architecture. Specifically, we construct an auxiliary depth estimation (DE) task that takes an RGB image as input to estimate a depth map, and train both DSR task and DE task collaboratively to boost the performance of DSR. Upon this, a cross-task interaction module is proposed to realize bilateral cross-task knowledge transfer. First, we design a cross-task distillation scheme that encourages DSR and DE networks to learn from each other in a teacher-student role-exchanging fashion. Then, we advance a structure prediction (SP) task that provides extra structure regularization to help both DSR and DE networks learn more informative structure representations for depth recovery. Extensive experiments demonstrate that our scheme achieves superior performance in comparison with other DSR methods. Baoli Sun, Xinchen Ye, Baopu Li, Zhihui Wang 0001, Rui Xu 0002 |
CVPR | 5 |
| 2021 | Cost Affinity Learning Network for Stereo MatchingabstractExisting stereo matching methods mainly tend to directly aggregate features output from Convolutional Neural Network to obtain more discriminative cost features, but ignore the affinity of each element in the cost feature which also plays a key role in enhancing the cost feature. In this work, we propose a novel cost affinity learning network(CAL-Net) whose Affinity Enhanced Module(AEM) extracts the affinity of the elements in the cost feature and reconstructs a more discriminative feature. In addition, CAL-Net designs a Disparity Weight Loss(DWL) to guide training. Specifically, AEM takes the advatange of the self-attention mechanism to learn internal affinity between different elements and exploits it to reconstruct the cost feature for emphasizing informative elements. DWL calculates the adaptive weight according to disparity error. As the error decreases, the weight gradually increases and enables the network to gradually transit from the pixel level disparity to sub-pixel level. Experiments demonstrate that CAL-Net boosts the performance, especially in textureless and reflective regions, and achieves better results on Scene Flow and KITTI 2012 benchmarks than some typical related methods. Shenglun Chen, Baopu Li, Wei Wang 0335, Hong Zhang 0011, Zhihui Wang 0001 |
ICASSP | 6 |
| 2021 | Domain adaptation with geometrical preservation and distribution alignment
Jing Sun 0012, Zhihui Wang 0001, Wei Wang 0335, Fuming Sun |
Neurocomputing | 2 |
| 2021 | Sparsely-labeled source assisted domain adaptation
Wei Wang 0335, Shenglun Chen, Yuankai Xiang, Jing Sun 0012, Zhihui Wang 0001, Fuming Sun, Zhengming Ding, Baopu Li |
Pattern Recognit. | 6 |
| 2021 | Dual Color Space Guided Sketch ColorizationabstractAutomatic sketch colorization is a challenging task in both computer graphics and computer vision since all the color, texture, shading generation have to be created based on the abstract sketch. Besides, it is a subjective task in painting process, which needs illustrators to comprehend drawing priori (DP), such as hue variation, saturation contrast and gray contrast and utilize them in the HSV color space which is closer to human visual cognition system. As such, incorporating supplementary supervision in the HSV color space may be beneficial to sketch colorization. However, previous methods improve the colorization quality only in the RGB color space without considering the HSV color space, often causing results with dull color, inappropriate saturation contrast, and artifacts. To address this issue, we propose a novel sketch colorization method, dual color space guided generative adversarial network (DCSGAN), that considers the complementary information contained in both the RGB and HSV color space. Specifically, we incorporate the HSV color space to construct dual color spaces for supervising our method with a color space transformation (CST) network that learns transformation from the RGB to HSV color space. Then, we propose a DP loss that enables the DCSGAN to generate vivid color images with pixel level supervision. Additionally, a novel dual color space adversarial (DCSA) loss is designed to guide the generator at global level to reduce the artifacts to meet audiences' aesthetic expectations. Extensive experiments and ablation studies demonstrate the superiority of the proposed method over previous state-of-the-art (SOTA) methods. Zhi Dou, Ning Wang 0025, Baopu Li, Zhihui Wang 0001, Bin Liu 0040 |
IEEE Trans. Image Process. | 4 |
| 2020 | Weakly Supervised Fine-Grained Image Classification via Guassian Mixture Model Oriented Discriminative LearningabstractExisting weakly supervised fine-grained image recognition (WFGIR) methods usually pick out the discriminative regions from the high-level feature maps directly. We discover that due to the operation of stacking local receptive filed, Convolutional Neural Network causes the discriminative region diffusion in high-level feature maps, which leads to inaccurate discriminative region localization. In this paper, we propose an end-to-end Discriminative Feature-oriented Gaussian Mixture Model (DF-GMM), to address the problem of discriminative region diffusion and find better fine-grained details. Specifically, DF-GMM consists of 1) a low-rank representation mechanism (LRM), which learns a set of low-rank discriminative bases by Gaussian Mixture Model (GMM) to accurately select discriminative details and filter more irrelevant information in high-level semantic feature maps, 2) a low-rank representation reorganization mechanism (LR2M) which resumes the space information of low-rank discriminative bases to reconstruct the low-rank feature maps. By recovering the low-rank discriminative bases into the same embedding space of highlevel feature maps, LR2M alleviates the discriminative region diffusion problem in high-level feature map and discriminative regions can be located more precisely on the new low-rank feature maps. Extensive experiments verify that DF-GMM yields the best performance under the same settings with the most competitive approaches, in CUBBird, Stanford-Cars datasets, and FGVC Aircraft. Zhihui Wang 0001, Shijie Wang 0003, Jianjun Li 0007, Zezhou Li |
CVPR | 1 |
| 2020 | Adaptive Local Neighbors for Transfer Discriminative Feature LearningabstractIn Domain Adaptation (DA), how to reduce the distributional differences across domains and preserve the data structures are two critical issues to obtain domain-invariant features. Existing DA methods either preserve the Local Manifold Structure (LMS) or the Global Discriminative Consistency (GDC), while fail to take those two metrics into account simultaneously. Therefore, the extracted features are either short of discriminative ability or sensitive to the multimodally distributed data. Moreover, the local neighbored relationships among data points are mostly established in original data space, which is unreliable, especially for data with large noises. Therefore, this paper proposes a novel DA approach, i.e., Adaptive Local Neighbors for Transfer Discriminative Feature Learning, to leverage LMS and GDC into a unified transfer feature learning model, where we only focus on the GDC between the local neighbors, so that the extracted features are more discriminative and robust to the multimodally distributed data. Moreover, the data points' local neighbors are revealed adaptively in the learned subspace so that it is insensitive to the data noises. Compared with the state-of-the-art methods, the proposed approach achieves higher performance for different cross-domain image classification tasks, especially 3.0% improved for Office10+Caltech10 dataset. Wei Wang 0335, Zhihui Wang 0001, Zhengming Ding |
ECAI | 2 |
| 2020 | Category-specific Semantic Coherency Learning for Fine-grained Image RecognitionabstractExisting deep learning based weakly supervised fine-grained image recognition (WFGIR) methods usually pick out the discriminative regions from the high-level feature (HLF) maps directly. However, as HLF maps are derived based on spatial aggregation of convolution which is basically a pattern matching process that applies fixed filters, it is ineffective to model visual contents of same semantic but varying posture or perspective. We argue that this will cause the selected discriminative regions of same sub-category are not semantically corresponding and thus degrade the WFGIR performance. To address this issue, we propose an end-to-end Category-specific Semantic Coherency Network (CSC-Net) to semantically align the discriminative regions of the same subcategory. Specifically, CSC-Net consists of: 1) Local-to-Attribute Projecting Module (LPM), which automatically learns a set of latent attributes via collecting the category-specific semantic details while eliminating the varying spatial distributions from the local regions. 2) Latent Attribute Aligning (LAA), which aligns the latent attributes to specific semantic via graph convolution based on their discriminability, to achieve category-specific semantic coherency; 3) Attribute-to-Local Resuming Module (ARM), which resumes the original Euclidean space of latent attributes and construct latent attribute aligned feature maps by a location-embedding graph unpooling operation. Finally, the new feature maps are used which applies the category-specific semantic coherency implicitly for more accurate discriminative regions localization. Extensive experiments verify that CSC-Net yields the best performance under the same settings with most competitive approaches, on CUB Bird, Stanford-Cars, and FGVC Aircraft datasets. Shijie Wang 0003, Zhihui Wang 0001, Wanli Ouyang |
ACM Multimedia | 2 |
| 2020 | Depth Super-Resolution via Deep Controllable Slicing NetworkabstractDue to the imaging limitation of depth sensors, high-resolution (HR) depth maps are often difficult to be acquired directly, thus effective depth super-resolution (DSR) algorithms are needed to generate HR output from its low-resolution (LR) counterpart. Previous methods treat all depth regions equally without considering different extents of degradation at region-level, and regard DSR under different scales as independent tasks without considering the modeling of different scales, which impede further performance improvement and practical use of DSR. To alleviate these problems, we propose a deep controllable slicing network from a novel perspective. Specifically, our model is to learn a set of slicing branches in a divide-and-conquer manner, parameterized by a distance-aware weighting scheme to adaptively aggregate different depths in an ensemble. Each branch that specifies a depth slice (e.g., the region in some depth range) tends to yield accurate depth recovery. Meanwhile, a scale-controllable module that extracts depth features under different scales is proposed and inserted into the front of slicing network, and enables finely-grained control of the depth restoration results of slicing network with a scale hyper-parameter. Extensive experiments on synthetic and real-world benchmark datasets demonstrate that our method achieves superior performance. Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Jing-Yu Yang 0002, Rui Xu 0002, Baopu Li |
ACM Multimedia | 3 |
| 2020 | Multi-attention based cross-domain beauty product image retrieval
Zhihui Wang 0001, Xing Liu 0005, Jiawen Lin, Caifei Yang |
Sci. China Inf. Sci. | 1 |
| 2020 | Geometry and context guided refinement for stereo matchingabstractThe disparity refinement phase of existing end‐to‐end stereo matching networks refines the disparity by learning the mapping from the concatenated coarse disparity and corresponding features to fine disparity. It depends on the scenarios' characteristics, such as the distribution of disparity and semantic categories contained in the domain, which makes the network fail to work on unseen domain. In this paper, we propose a geometry and context guided refinement network (GCGR‐Net) containing a Fine Matching module and an Upsampling module. GCGR‐Net learns to utilize pixels' relationship to get high resolution dense disparity, which is independent of the data's content. The Fine Matching module performs a minimum range search based on the relationship between the possible matching pixel pairs, i.e. the called geometry information, to recover the internal structure of the object. The Upsampling module obtains context information, the relationship between central pixel and the pixels in its neighborhood, to upsample the lower resolution disparity. The final disparity map is obtained step by step through an iterative refinement model. Experiment results show that our method not only has good performance in the training scenarios, but also outperforms previous methods on the unseen domain without fine‐tuning. Hong Zhang 0011, Zhihui Wang 0001, Yuxin Yue, Shenglun Chen |
IET Image Process. | 3 |
| 2020 | DRM-SLAM: Towards dense reconstruction of monocular SLAM with scene depth fusion
Xinchen Ye, Xiang Ji 0005, Baoli Sun, Shenglun Chen, Zhihui Wang 0001 |
Neurocomputing | 5 |
| 2020 | Depth upsampling based on deep edge-aware learning
Zhihui Wang 0001, Xinchen Ye, Baoli Sun, Jing-Yu Yang 0002, Rui Xu 0002 |
Pattern Recognit. | 1 |
| 2020 | Progressive learning for weakly supervised fine-grained classification
Tiantian Yan, Shijie Wang 0003, Zhihui Wang 0001, Zhongxuan Luo |
Signal Process. | 3 |
| 2020 | Deep Joint Depth Estimation and Color Correction From Monocular Underwater Images Based on Unsupervised Adaptation NetworksabstractDegraded visibility and geometrical distortion typically make the underwater vision more intractable than open air vision, which impedes the development of underwater-related machine vision and robotic perception. Therefore, this paper addresses the problem of joint underwater depth estimation and color correction from monocular underwater images, which aims at enjoying the mutual benefits between these two related tasks from a multi-task perspective. Our core ideas lie in our new deep learning architecture. Due to the lack of effective underwater training data, and the weak generalization to the real-world underwater images trained on synthetic data, we consider the problem from a novel perspective of style-level and feature-level adaptation, and propose an unsupervised adaptation network to deal with the joint learning problem. Specifically, a style adaptation network (SAN) is first proposed to learn a style-level transformation to adapt in-air images to the style of underwater domain. Then, we formulate a task network (TN) to jointly estimate the scene depth and correct the color from a single underwater image by learning domain-invariant representations. The whole framework can be trained end-to-end in an adversarial learning manner. Extensive experiments are conducted under air-to-water domain adaptation settings. We show that the proposed method performs favorably against state-of-the-art methods in both depth estimation and color correction tasks. Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Rui Xu 0002, Xin Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | PMBANet: Progressive Multi-Branch Aggregation Network for Scene Depth Super-ResolutionabstractDepth map super-resolution is an ill-posed inverse problem with many challenges. First, depth boundaries are generally hard to reconstruct particularly at large magnification factors. Second, depth regions on fine structures and tiny objects in the scene are destroyed seriously by downsampling degradation. To tackle these difficulties, we propose a progressive multi-branch aggregation network (PMBANet), which consists of stacked MBA blocks to fully address the above problems and progressively recover the degraded depth map. Specifically, each MBA block has multiple parallel branches: 1) The reconstruction branch is proposed based on the designed attention-based error feed-forward/-back modules, which iteratively exploits and compensates the downsampling errors to refine the depth map by imposing the attention mechanism on the module to gradually highlight the informative features at depth boundaries. 2) We formulate a separate guidance branch as prior knowledge to help to recover the depth details, in which the multi-scale branch is to learn a multi-scale representation that pays close attention at objects of different scales, while the color branch regularizes the depth map by using auxiliary color information. Then, a fusion block is introduced to adaptively fuse and select the discriminative features from all the branches. The design methodology of our whole network is well-founded, and extensive experiments on benchmark datasets demonstrate that our method achieves superior performance in comparison with the state-of-the-art methods. Our code and models are available athttps://github.com/Sunbaoli/PMBANet_DSR/. Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Jing-Yu Yang 0002, Rui Xu 0002, Baopu Li |
IEEE Trans. Image Process. | 3 |
| 2019 | Online Single Person Tracking for Unmanned Aerial Vehicles: Benchmark and New BaselineabstractOnline tracking a specific person from a low-altitude unmanned aerial vehicle (UAV) is a very interesting and challenging problem to be solved. However, there exists no large-scale aerial video dataset regarding this online single person tracking (OSPT) task. To promote the study of the OSPT problem in UAV, we first construct a new benchmark dataset including 100 fully annotated aerial videos with nearly 130K frames and 11 challenging factors. Second, we evaluate several state-of-the-art online trackers with real-time performance using our dataset, considering the potential applications in the UAV platform. In addition, with respect to the OSPT problem, we attempt to design a new baseline method with the combination of tracking, detection and re-identification and conduct detailed analysis of different components. This method achieves much better performance than the existing online trackers, which will serve as a new baseline for our benchmark. Zhihui Wang 0001, Dong Wang 0004, Yunwei Qi, Huchuan Lu |
ICASSP | 1 |
| 2019 | Accurate Monocular 3D Object Detection via Color-Embedded 3D Reconstruction for Autonomous DrivingabstractIn this paper, we propose a monocular 3D object detection framework in the domain of autonomous driving. Unlike previous image-based methods which focus on RGB feature extracted from 2D images, our method solves this problem in the reconstructed 3D space in order to exploit 3D contexts explicitly. To this end, we first leverage a stand-alone module to transform the input data from 2D image plane to 3D point clouds space for a better input representation, then we perform the 3D detection using PointNet backbone net to obtain objects' 3D locations, dimensions and orientations. To enhance the discriminative capability of point clouds, we propose a multi-modal feature fusion module to embed the complementary RGB cue into the generated point clouds representation. We argue that it is more effective to infer the 3D bounding boxes from the generated 3D scene space (i.e., X,Y, Z space) compared to the image plane (i.e., R,G,B image plane). Evaluation on the challenging KITTI dataset shows that our approach boosts the performance of state-of-the-art monocular approach by a large margin. Xinzhu Ma, Zhihui Wang 0001, Wanli Ouyang, Xin Fan 0001 |
ICCV | 2 |
| 2019 | Self-Adaption Multi-classifier Fusion Networks for Image RecognitionabstractRecently, many visual recognition related studies have proved that making full use of different levels of features can effectively enhance the representational ability of convolutional neural networks (CNNs). Different from other CNN architecture which are devoted to aggregate features of different scales, we proposed a multi-classifier network (MCN) to make more effective use of these feature maps. Specifically, MCN can directly make full use of features of different levels and fuse intermediate results in a self-adaption way. Note that the auxiliary classifiers not only can optimize the internal features of CNNs directly, but also bring additional gradient which further solves the problem of vanishing-gradient. In addition, MCN is a very flexible architecture and can be combined with existing state-of-the-art networks (ResNet, DenseNet, ResNeXt, etc.) easily. Extensive experiments on three highly competitive benchmark datasets, CIFAR-10, CIFAR-100 and ImageNet, clearly demonstrate superior performance of the proposed MCN over state-of-the-arts. Zengyuan Guo, Xinzhu Ma, Zhihui Wang 0001 |
ICME | 4 |
| 2019 | Accurate And Fast Fine-Grained Image Classification via Discriminative LearningabstractCurrently, most top-performing Weakly supervised Fine-grained Image Classification (WFGIC) schemes tend to pick out discriminative patches. However, those patches usually contain much noise information, which influences the accuracy of the classification. Besides, they rely on a large amount of candidate patches to discover the discriminative ones, thus leading to high computational cost. To address these problems, we propose a novel end-to-end Self-regressive Localization with Discriminative Prior Network (SDN) model, which learns to explore more accurate size of discriminative patches and enables to classify images in real time. Specifically, we design a multi-task discriminative learning network, a self-regressive localization sub-network and a discriminative prior sub-network with the guided loss as well as the consistent loss to simultaneously learn self-regressive coefficients and discriminative prior maps. The self-regressive coefficients can decrease noise information in discriminative patches and the discriminative prior maps through learning discriminative probability values filter thousands of candidate patches to single figure. Extensive experiments demonstrate that the proposed SDN model achieves state-of-the-art both in accuracy and efficiency. Zhihui Wang 0001, Shijie Wang 0003 |
ICME | 1 |
| 2019 | Continuous Scale Adaption for Efficient Box-Based Scene Text DetectionabstractDue to the diversity of text size in scene images, the current box-based methods employ a large amount of fixed-size anchors with different scales to match texts, thus leading to high computational cost. In this paper, we propose to learn the scales of texts and adjust the sizes of anchors accordingly, which can largely reduce the numbers of anchors and therefore significantly reduces the time cost. Moreover, compared to discrete scales used in previous methods, the learned scales are continuous and more reliable. Additionally, we propose Anchor convolution to exploit scaled feature for each anchor by dynamically adjusting the sizes of receptive fields according to the learned scales. Experimental results show that the proposed method significantly improves the computational efficiency of box-based framework(reduce the running time from 0.73s to 0.28s) and enhances its robustness against small texts, while achieving competitive performance with other methods. Bingwang Zhang, Zhihui Wang 0001, Zhongxuan Luo |
ICME | 4 |
| 2019 | Learning to Segment Unseen Category Objects using Gradient Gaussian AttentionabstractExisting semantic segmentation models are trapped in the categories of training sets. Unfortunately, public datasets provide pixel-level annotations only for a small quantity of images and few categories. In this paper, we propose a novel cross-category supervised object segmentation network to explore the similar features among different categories, which can transfer the learned segmentation knowledge from categories with mask annotations to unseen categories that only have bounding boxes. Specifically, we fuse gaussian attention map of an object with guided gradient back-propagation map as an extra input, which gives localizable and discriminative prior cues to obtain precise object segmentation from the bounding box. In addition to computing a segment for each box, we also fuse segments to generate pixel-level labels. Then, without modifying the segmentation training program, the generated labels are still sufficient and achieve about 98.2% of the fully supervised model, in the case of only 50% categories with pixel-level annotations on PASCAL VOC 2012. The proposed method is also effective for interactive segmentation and salient object detection. Zhihui Wang 0001, Xinzhu Ma, Jianjun Li 0007 |
ICME | 2 |
| 2019 | Weakly Supervised Fine-grained Image Classification via Correlation-guided Discriminative LearningabstractWeakly supervised fine-grained image classification (WFGIC) aims at learning to recognize hundreds of subcategories in each basic-level category with only image level labels available. It is extremely challenging and existing methods mainly focus on the discriminative semantic parts or regions localization as the key differences among different subcategories are subtle and local. However, they localize these regions independently while neglecting the fact that regions are mutually correlated and region groups can be more discriminative. Meanwhile, most current work tends to derive features directly from the output of CNN and rarely considers the correlation within the feature vector. To address these issues, we propose an end-to-end Correlation-guided Discriminative Learning (CDL) model to fully mine and exploit the discriminative potentials of correlations for WFGIC globally and locally. From the global perspective, a discriminative region grouping (DRG) sub-network is proposed which first establishes correlation between regions and then enhances each region by weighted aggregating all the correlation from other regions to it. By this means each region's representation encodes the global image-level context and thus is more robust; meanwhile, through learning the correlation between discriminative regions, the network is guided to implicitly discover the discriminative region groups which are more powerful for WFGIC. From the local perspective, a discriminative feature strengthening sub-network (DFS) is proposed to mine and learn the internal spatial correlation among elements of each patch's feature vector, to improve its discriminative power locally by jointly emphasizes informative elements while suppresses the useless ones. Extensive experiments demonstrate the effectiveness of proposed DRG and DFS sub-networks, and show that the CDL model achieves state-of-the-art performance both in accuracy and efficiency. Zhihui Wang 0001, Shijie Wang 0003, Jianjun Li 0007 |
ACM Multimedia | 1 |
| 2019 | Blind Image Deblurring via Adaptive Optimization with Flexible Sparse Structure Control
Risheng Liu, Caisheng Mao, Zhihui Wang 0001 |
J. Comput. Sci. Technol. | 3 |
| 2018 | User-Guided Deep Anime Line Art Colorization with Conditional Adversarial NetworksabstractScribble colors based line art colorization is a challenging computer vision problem since neither greyscale values nor semantic information is presented in line arts, and the lack of authentic illustration-line art training pairs also increases difficulty of model generalization. Recently, several Generative Adversarial Nets (GANs) based methods have achieved great success. They can generate colorized illustrations conditioned on given line art and color hints. However, these methods fail to capture the authentic illustration distributions and are hence perceptually unsatisfying in the sense that they often lack accurate shading. To address these challenges, we propose a novel deep conditional adversarial architecture for scribble based anime line art colorization. Specifically, we integrate the conditional framework with WGAN-GP criteria as well as the perceptual loss to enable us to robustly train a deep network that makes the synthesized images more natural and real. We also introduce a local features network that is independent of synthetic data. With GANs conditioned on features from such network, we notably increase the generalization capability over "in the wild" line arts. Furthermore, we collect two datasets that provide high-quality colorful illustrations and authentic line arts for training and benchmarking. With the proposed model trained on our illustration dataset, we demonstrate that images synthesized by the presented approach are considerably more realistic and precise than alternative approaches. Yuanzheng Ci, Xinzhu Ma, Zhihui Wang 0001, Zhongxuan Luo |
ACM Multimedia | 3 |
| 2018 | Disparity-Based Robust Unstructured Terrain Segmentation
Xinzhu Ma, Zhihui Wang 0001, Zhongxuan Luo |
PRCV (4) | 3 |
| 2018 | Sparse dual graph-regularized NMF for image co-clustering
Jing Sun 0012, Zhihui Wang 0001, Fuming Sun |
Neurocomputing | 2 |
| 2017 | Performance Analysis for Content Distribution in Crowdsourced Content-Centric Mobile Networking
Chengming Li 0004, Xiaojie Wang 0001, Shimin Gong, Zhihui Wang 0001, Qingshan Jiang |
QSHINE | 4 |
| 2016 | 1D Barcode Region Detection Based on the Hough Transform and Support Vector Machine
Zhihui Wang 0001, Ai Chen, Jianjun Li 0007, Zhongxuan Luo |
MMM (2) | 1 |
| 2016 | A magic cube based information hiding scheme of large payload
Qiong Wu 0008, Chin-Chen Chang 0001, Zhihui Wang 0001 |
J. Inf. Secur. Appl. | 5 |
| 2016 | Progressive secret image sharing scheme using meaningful shadowsabstractAbstract This paper proposes a novel secret image sharing scheme, which progressively hides a secret image into multiple different meaningful cover (or host) images by utilizing a magic matrix. The produced shadows are high visual quality, meaningful images that differ from each other. As a result, they are not easy to cause the suspicion by the attackers. Moreover, the secret image in the proposed scheme can be recovered progressively via different numbers of shadows. The more shadows used, the better the quality of the secret image. The experimental results demonstrate the aforementioned advantages of the proposed scheme. Copyright © 2016 John Wiley & Sons, Ltd. Zhihui Wang 0001, Ya-Feng Di, Chin-Chen Chang 0001 |
Secur. Commun. Networks | 1 |
| 2014 | Illumination-based nighttime video contrast enhancement using genetic algorithm
Yunbo Rao, Zhihui Wang 0001, Leiting Chen |
Multim. Tools Appl. | 3 |
| 2013 | Histogram-shifting-imitated reversible data hiding
Zhihui Wang 0001, Chin-Feng Lee, Ching-Yun Chang |
J. Syst. Softw. | 1 |
| 2013 | A high-performance reversible data-hiding scheme for LZW codes
Zhihui Wang 0001, Hai-Rui Yang, Ting-Fang Cheng, Chin-Chen Chang 0001 |
J. Syst. Softw. | 1 |
| 2012 | Reversible Data Hiding Scheme Based on Image InpaintingabstractReversible/lossless image data hiding schemes provide the capability to embed secret information into a cover image where the original carrier can be totally restored after extracting the secret information. This work presents a high performance reversible image data hiding scheme especially in stego image quality control using image inpainting, an efficient image processing skill. Embeddable pixels chosen from a cover image are initialized to a fixed value as preprocessing for inpainting. Subsequently, these initialized pixels are repaired using inpainting technique based on partial differential equations (PDE). These inpainted pixels can be used to carry secret bits and generate a stego image. Experimental results show that the proposed scheme produces low distortion stego images and it also provides satisfactory hiding capacity. Chuan Qin 0001, Zhihui Wang 0001, Chin-Chen Chang 0001, Kuo-Nan Chen |
Fundam. Informaticae | 2 |
| 2012 | Optimizing least-significant-bit substitution using cat swarm optimization strategy
Zhihui Wang 0001, Chin-Chen Chang 0001, Mingchu Li |
Inf. Sci. | 1 |
| 2011 | A Two-Layer Steganography Scheme Using Sudoku for Digital Images
Yi-Hui Chen, Chi-Wei Lan, Zhihui Wang 0001 |
UIC | 3 |
| 2011 | Image Data Hiding Schemes Based on Graph Coloring
Shuai Yue, Zhihui Wang 0001, Ching-Yun Chang, Chin-Chen Chang 0001, Mingchu Li |
UIC | 2 |
| 2010 | An encoding method for both image compression and data lossless information hiding
Zhihui Wang 0001, Chin-Chen Chang 0001, Kuo-Nan Chen, Mingchu Li |
J. Syst. Softw. | 1 |
| 2009 | A reversible information hiding scheme using left-right and up-down chinese character representation
Zhihui Wang 0001, Chin-Chen Chang 0001, Chia-Chen Lin 0001, Mingchu Li |
J. Syst. Softw. | 1 |