EDBT 2026 Demo / reviewers in the wild / expert
Zhuotao Tian
dblp:243/7181
· DBLP profile ↗
43ranked-venue papers
5as first author
41since 2021 · last 2026
0000-0003-2698-6923ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 5 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 3 first-author · 26 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic ManipulationabstractVision-Language-Action (VLA) models have advanced in robotic manipulation, yet practical deployment remains hindered by two key limitations: **1) perceptual redundancy**, where irrelevant visual inputs are processed inefficiently, and **2) superficial instruction-vision alignment**, which hampers semantic grounding of actions. In this paper, we propose **SemanticVLA**, a novel VLA framework that performs Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation. Specifically: **1)** To sparsify redundant perception while preserving semantic alignment, **Semantic-guided Dual Visual Pruner (SD-Pruner)** performs: Instruction-driven Pruner (ID-Pruner) extracts global action cues and local semantic anchors in SigLIP; Spatial-aggregation Pruner (SA-Pruner) compacts geometry-rich features into task-adaptive tokens in DINOv2. **2)** To exploit sparsified features and integrate semantics with spatial geometry, **Semantic-complementary Hierarchical Fuser (SH-Fuser)** fuses dense patches and sparse tokens across SigLIP and DINOv2 for coherent representation. **3)** To enhance the transformation from perception to action, **Semantic-conditioned Action Coupler (SA-Coupler)** replaces the conventional observation-to-DoF approach, yielding more efficient and interpretable behavior modeling for manipulation tasks. Extensive experiments on simulation and real-world tasks show that SemanticVLA sets a new SOTA in both performance and efficiency. SemanticVLA surpasses OpenVLA on LIBERO benchmark by **21.1%** in success rate, while reducing training cost and inference latency by **3.0×** and **2.7×**. Renshan Zhang, Rui Shao 0001, Zhijian Fang, Kaiwen Zhou 0001, Zhuotao Tian, Liqiang Nie |
AAAI | 6 |
| 2026 | Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning FrameworkabstractChenyuan Zhang, Qiguang Chen, Xie Chen, Zhuotao Tian, Bowen Xing, Meishan Zhang, Libo Qin, Baotian Hu, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qiguang Chen, Xie Chen 0001, Zhuotao Tian, Meishan Zhang, Libo Qin 0001, Baotian Hu, Min Zhang 0005 |
ACL (1) | 4 |
| 2026 | Failure Detection in Image Segmentation Under Conditions of Semantic and Covariate ShiftsabstractWhether deep neural networks can provide reliable confidence is of great significance, especially in risk-sensitive scenarios. This work explores the impact of covariate and semantic shifts on segmentation tasks, an area which has not been extensively studied. Covariate shift refers to changes in the data distribution without alterations in the label space, while semantic shift involves changes in both data distribution and label space. We find that model-unknown distributional shifts in test data can transform an overconfidence problem into a situation of making random predictions with arbitrary confidence. The paper proposes a novel approach for effective failure detection that combines holistic image-level analysis and detailed pixel-level information. This approach involves the use of a Gray Level Co-occurrence Matrix (GLCM) to analyze the prediction randomness between adjacent pixels and a Magnitude-Direction Confidence Score Function (MD-CSF) for determining pixel acceptance or rejection. Furthermore, we introduce a new benchmark dataset, the Robot Inspection dataset for Semantic and Covariate shift in Segmentation (RISKS, the homophone of RISCS), to fill the need for datasets capable of evaluating the simultaneous impact of semantic and covariate shifts. Experimental results demonstrate that our method successfully detects image-level failures in segmentation, with MD-CSF outperforming other pluggable CSFs. The code and RISKS dataset will be available at https://github.com/liuyijungoon/MD-CSF. Yijun Liu 0012, Zhuotao Tian, Hang Zhao 0019, Zipeng Zhu, Jingyong Su |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | RefDetector: A Simple Yet Effective Matching-based Method for Referring Expression ComprehensionabstractDespite the rapid and substantial advancements in object detection, it continues to face limitations imposed by pre-defined category sets. Current methods for visual grounding primarily focus on how to better leverage the visual backbone to generate text-tailored visual features, which may require adjusting the parameters of the entire model. Besides, some early methods, \ie, matching-based method, build upon and extend the functionality of existing object detectors by enabling them to localize an object based on free-form linguistic expressions, which have good application potential. However, the untapped potential of the matching-based approach has not been fully realized due to inadequate exploration. In this paper, we first analyze the limitations that exist in the current matching-based method (\ie, mismatch problem and complicated fusion mechanisms), and then present a simple yet effective matching-based method, namely RefDetector. To tackle the above issues, we devise a simple heuristic rule to generate proposals with improved referent recall. Additionally, we introduce a straightforward vision-language interaction module that eliminates the need for intricate manually-designed mechanisms. Moreover, we have explored the visual grounding based on the modern detector DETR, and achieved significant performance improvement. Extensive experiments on three REC benchmark datasets, \ie, RefCOCO, RefCOCO+, and RefCOCOg validate the effectiveness of the proposed method. Zhuotao Tian, Sanping Zhou, Le Wang 0003 |
AAAI | 2 |
| 2025 | DeCLIP: Decoupled Learning for Open-Vocabulary Dense PerceptionabstractDense visual prediction tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbounded. While Vision-Language Models (VLMs) like CLIP have shown promise in open-vocabulary tasks, their direct application to dense prediction often leads to suboptimal performance due to limitations in local feature representation. In this work, we present our observation that CLIP’s image tokens struggle to effectively aggregate information from spatially or semantically related regions, resulting in features that lack local discriminability and spatial consistency. To address this issue, we propose DeCLIP, a novel framework that enhances CLIP by decoupling the self-attention module to obtain "content" and "context" features respectively. The "content" features are aligned with image crop representations to improve local discriminability, while "context" features learn to retain the spatial correlations under the guidance of vision foundation models, such as DINO. Extensive experiments demonstrate that DeCLIP significantly outperforms existing methods across multiple open-vocabulary dense prediction tasks, including object detection and semantic segmentation. Code is available at https://github.com/xiaomoguhz/DeCLIP. Bin Chen 0022, Yulin Li 0003, Bin Kang, Yichi Chen 0002, Zhuotao Tian |
CVPR | 6 |
| 2025 | VisionZip: Longer is Better but Not Necessary in Vision Language ModelsabstractRecent advancements in vision-language models have enhanced performance by increasing the length of visual tokens, making them much longer than text tokens and significantly raising computational costs. However, we observe that the visual tokens generated by popular vision encoders, such as CLIP and SigLIP, contain significant redundancy. To address this, we introduce VisionZip, a simple yet effective method that selects a set of informative tokens for input to the language model, reducing visual token redundancy and improving efficiency while maintaining model performance. The proposed VisionZip can be widely applied to image and video understanding tasks and is well-suited for multi-turn dialogues in real-world scenarios, where previous methods tend to underperform. Experimental results show that VisionZip outperforms the previous state-of-the-art method by at least 5% performance gains across nearly all settings. Moreover, our method significantly enhances model inference speed, improving the prefilling time by 8× and enabling the LLaVA-Next 13B model to infer faster than the LLaVA-Next 7B model while achieving better results. Furthermore, we analyze the causes of this redundancy and encourage the community to focus on extracting better visual features rather than merely increasing token length. Our code is available at https://github.com/dvlab-research/VisionZip. Senqiao Yang, Yukang Chen, Zhuotao Tian, Chengyao Wang, Jingyao Li 0001, Bei Yu 0001, Jiaya Jia |
CVPR | 3 |
| 2025 | C2AD: Dual Consistency Learning for Zero-Shot Anomaly DetectionabstractZero-shot anomaly detection (ZSAD) is dedicated to detecting anomalies without having any seen normal or abnormal samples for the target set. Existing approaches utilize the pre-trained CLIP to assess normality/abnormality by exploiting the similarity between images and text with the frozen visual encoder. However, the frozen CLIP visual encoder impedes performance improvements. Additionally, their representations of anomalies are sensitive to contextual variations, leading to poor localization of unseen abnormalities. Therefore, this paper introduces the Dual Consistency Learning for Zero-Shot Anomaly Detection (C2AD), comprising two components: semantic and contextual consistency. Semantic consistency enhances generalization by maintaining correlational semantic consistency, while contextual consistency encourages representations to be robust to contextual changes. C2AD improves the model training without adding extra computational overhead during inference. Comprehensive experiments demonstrate that C2AD can boost the performance of ZSAD in anomaly detection and localization, achieving state-of-the-art results. Ruilong Xing, Zhuotao Tian, Yijun Liu 0012, Senqiao Yang, Jingyong Su |
ICASSP | 3 |
| 2025 | Edit360: 2D Image Edits to 3D Assets From Any AngleabstractRecent advances in diffusion models have significantly improved image generation and editing, but extending these capabilities to 3D assets remains challenging, especially for fine-grained edits that require multi-view consistency. Existing methods typically restrict editing to predetermined viewing angles, severely limiting their flexibility and practical applications. We introduce Edit360, a tuning-free framework that extends 2D modifications to multi-view consistent 3D editing. Built upon video diffusion models, Edit360 enables user-specific editing from arbitrary viewpoints while ensuring structural coherence across all views. The framework selects anchor views for 2D modifications and propagates edits across the entire 360-degree range. To achieve this, Edit360 introduces a novel Anchor-View Editing Propagation mechanism, which effectively aligns and merges multi-view information within the latent and attention spaces of diffusion models. The resulting edited multi-view sequences facilitate the reconstruction of high-quality 3D assets, enabling customizable 3D content creation. Junchao Huang, Xinting Hu, Shaoshuai Shi, Zhuotao Tian, Li Jiang 0009 |
ICCV | 4 |
| 2025 | Enhancing Spatial Reasoning in Multimodal Large Language Models Through Reasoning-Based SegmentationabstractRecent advances in point cloud perception have demonstrated remarkable progress in scene understanding through vision-language alignment leveraging large language models (LLMs). However, existing methods may still encounter challenges in handling complex instructions that require accurate spatial reasoning, even if the 3D point cloud data provides detailed spatial cues such as size and position for identifying the targets. To tackle this issue, we propose Relevant Reasoning Segmentation (R$^2$S), a reasoning-based segmentation framework. The framework emulates human cognitive processes by decomposing spatial reasoning into two sequential stages: first identifying relevant elements, then processing instructions guided by their associated visual priors. Furthermore, acknowledging the inadequacy of existing datasets in complex reasoning tasks, we introduce 3D ReasonSeg, a reasoning-based segmentation dataset comprising 25,185 training samples and 3,966 validation samples with precise annotations. Both quantitative and qualitative experiments demonstrate that the R$^2$S and 3D ReasonSeg effectively endow 3D point cloud perception with stronger spatial reasoning capabilities, and we hope that they can serve as a new baseline and benchmark for future work. Zhenhua Ning, Zhuotao Tian, Shaoshuai Shi, Guangming Lu 0002, Daojing He, Wenjie Pei, Li Jiang 0009 |
ICCV | 2 |
| 2025 | Mitigating Object Hallucinations via Sentence-Level Early Intervention
Shangpin Peng, Songiao Yang, Li Jiang 0009, Zhuotao Tian |
ICCV | 4 |
| 2025 | CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image RetrievalabstractExisting Visual Language Models (VLMs) suffer structural limitations where a few low contribution tokens may excessively capture global semantics, dominating the information aggregation process and suppressing the discriminative features in text-driven image retrieval tasks. To address this, we introduce \textbf{CalibCLIP}, a training-free method designed to calibrate the suppressive effect of dominant tokens. Specifically, in the visual space, we propose the Contrastive Visual Enhancer (CVE), which decouples visual features into target and low information regions. Subsequently, it identifies dominant tokens and dynamically suppresses their representations.In the textual space, we introduce the Discriminative Concept Calibrator (DCC), which aims to differentiate between general and discriminative concepts within the text query. By mitigating the challenges posed by generic concepts and improving the representations of discriminative concepts, DCC strengthens the differentiation among similar samples. Finally, extensive experiments demonstrate consistent improvements across seven benchmarks spanning three image retrieval tasks, underscoring the effectiveness of CalibCLIP. Code is available at: https://github.com/kangbin98/CalibCLIP Bin Kang, Bin Chen 0022, Yulin Li 0003, Junzhi Zhao, Junle Wang, Zhuotao Tian |
ACM Multimedia | 7 |
| 2025 | Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe PriorabstractRecent advances in Video Large Language Models (VLLMs) have achieved remarkable video understanding capabilities, yet face critical efficiency bottlenecks due to quadratic computational growth with lengthy visual token sequences of long videos. While existing keyframe sampling methods can improve temporal modeling efficiency, additional computational cost is introduced before feature encoding, and the binary frame selection paradigm is found suboptimal. Therefore, in this work, we propose **Dy**namic **To**ken compression via LLM-guided **K**eyframe prior (**DyToK**), a training-free paradigm that enables dynamic token compression by harnessing VLLMs' inherent attention mechanisms. Our analysis reveals that VLLM attention layers naturally encoding query-conditioned keyframe priors, by which DyToK dynamically adjusts per-frame token retention ratios, prioritizing semantically rich frames while suppressing redundancies. Extensive experiments demonstrate that DyToK achieves state-of-the-art efficiency-accuracy tradeoffs. DyToK shows plug-and-play compatibility with existing compression methods, such as VisionZip and FastV, attaining 2.5x faster inference while preserving accuracy across multiple VLLMs, such as LLaVA-OneVision and Qwen2.5-VL. Code and models will be made publicly available. Yulin Li 0003, Haokun Gui, Ziyang Fan, Bin Kang, Bin Chen 0022, Zhuotao Tian |
NeurIPS | 7 |
| 2025 | Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMsabstractLarge Language Models (LLMs) have emerged as powerful tools for diverse applications. However, their uniform token processing paradigm introduces critical vulnerabilities in instruction handling, particularly when exposed to adversarial scenarios. In this work, we identify and propose a novel class of vulnerabilities, termed Tool-Completion Attack (TCA), which exploits function-calling mechanisms to subvert model behavior. To evaluate LLM robustness against such threats, we introduce the Tool-Completion benchmark, a comprehensive security assessment framework, which reveals that even state-of-the-art models remain susceptible to TCA, with surprisingly high attack success rates. To address these vulnerabilities, we introduce Context-Aware Hierarchical Learning (CAHL), a sophisticated mechanism that dynamically equilibrates semantic comprehension with role-specific instruction constraints. CAHL leverages the contextual correlations between different instruction segments to establish a robust, context-aware instruction hierarchy. Extensive experiments demonstrate that CAHL significantly enhances LLM robustness against both conventional attacks and the proposed TCA, exhibiting strong generalization capabilities in zero-shot evaluations while still preserving model performance on generic tasks. Our code is available at https://github.com/S2AILab/CAHL. Tengyun Ma, Daojing He, Shihao Peng, Yu Li 0007, Shaohui Liu, Zhuotao Tian |
NeurIPS | 7 |
| 2025 | Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial RepresentationsabstractHumans learn abstract concepts through multisensory synergy, and once formed, such representations can often be recalled from a single modality. Inspired by this principle, we introduce Concerto, a minimalist simulation of human concept learning for spatial cognition, combining 3D intra-modal self-distillation with 2D-3D cross-modal joint embedding. Despite its simplicity, Concerto learns more coherent and informative spatial features, as demonstrated by zero-shot visualizations. It outperforms both standalone SOTA 2D and 3D self-supervised models by 14.2\% and 4.8\%, respectively, as well as their feature concatenation, in linear probing for 3D scene perception. With full fine-tuning, Concerto sets new SOTA results across multiple scene understanding benchmarks (e.g., 80.7\% mIoU on ScanNet). We further present a variant of Concerto tailored for video-lifted point cloud spatial understanding, and a translator that linearly projects Concerto representations into CLIP’s language space, enabling open-world perception.
These results highlight that Concerto emerges spatial representations with superior fine-grained geometric and semantic consistency. Yujia Zhang 0003, Xiaoyang Wu 0002, Yixing Lao, Chengyao Wang, Zhuotao Tian, Naiyan Wang, Hengshuang Zhao |
NeurIPS | 5 |
| 2025 | BFRA: A Bi-Level Feature Relation Alignment Method for Cross-Domain Few-Shot LearningabstractWhile existing Few-Shot Learning (FSL) techniques demonstrate strong performance on uniform datasets, they encounter domain shift challenges when presented with domainagnostic queries in real-world scenarios. So we investigate it in Cross-Domain Few-Shot Learning (CD-FSL) and propose to learn more universal feature representations to enhance generalization on unseen domains. Toward this issue, we pinpoint two issues in current multi-model fusion approaches: 1) the entanglement of domain and class information, and 2) feature overlap across distinct domains. To address these challenges, we introduce a Bi-level Feature Relation Alignment method, BFRA, which facilitates the acquisition of more versatile features by decoupling domain-class relationships and aligning feature relations. Through the segregation of domain and class feature learning, we devise a smoothing layer prior for domain feature alignment to mitigate inter-domain discrepancies. This approach enables our model to acquire domain-consistent features, diminishing interference in subsequent class feature alignment procedures. During the class feature alignment, we notice that class feature representations from various in-domain models may intersect, leading to a diminished distinction between classes. To address this, we adopt a topological perspective to train our target model, by aligning feature relations instead of features between our target model and multiple in-domain models. The integration of these components results in the establishment of a bi-level feature relation alignment framework aimed at acquiring more universal features. Furthermore, we partially fine-tune the plug-in layer-wise affine adapter on domain-agnostic queries to expedite adaptation without impacting the known domains. Experiments of 21 datasets on meta-dataset and BSCD-FSL benchmark demonstrate the effectiveness of our method. The code are made publicly available at https://github.com/leaves162/BFRA. Tong Shao, Zhuotao Tian, Jingyong Su |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Layer-Wise Mutual Information Meta-Learning Network for Few-Shot SegmentationabstractThe goal of few-shot segmentation (FSS) is to segment unlabeled images belonging to previously unseen classes using only a limited number of labeled images. The main objective is to transfer label information effectively from support images to query images. In this study, we introduce a novel meta-learning framework called layer-wise mutual information (LayerMI), which enhances the propagation of label information by maximizing the mutual information (MI) between support and query features at each layer. Our approach involves the utilization of a LayerMI Block based on information-theoretic co-clustering. This block performs online co-clustering on the joint probability distribution obtained from each layer, generating a target-specific attention map. The LayerMI Block can be seamlessly integrated into the meta-learning framework and applied to all convolutional neural network (CNN) layers without altering the training objectives. Notably, the LayerMI Block not only maximizes MI between support and query features but also facilitates internal clustering within the image. Extensive experiments demonstrate that LayerMI significantly enhances the performance of baseline and achieves competitive performance compared to state-of-the-art methods on three challenging benchmarks: PASCAL- $5^{i}$ , COCO- $20^{i}$ , and FSS-1000. Xiaoliu Luo, Zhao Duan, Anyong Qin, Zhuotao Tian, Ting Xie 0004, Taiping Zhang, Yuan Yan Tang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | LISA: Reasoning Segmentation via Large Language ModelabstractAlthough perception systems have made remarkable ad-vancements in recent years, they still rely on explicit human instruction or pre-defined categories to identify the target objects before executing visual recognition tasks. Such systems cannot actively reason and comprehend implicit user intention. In this work, we propose a new segmentation task - reasoning segmentation. The task is designed to output a segmentation mask given a complex and implicit query text. Furthermore, we establish a benchmark comprising over one thousand image-instruction-mask data samples, incorporating intricate reasoning and world knowledge for evaluation purposes. Finally, we present LISA: large Language Instructed Segmentation Assistant, which inherits the language generation capabilities of multimodal Large Language Models (LLMs) while also possessing the ability to produce segmentation masks. We expand the original vocabulary with atoken and propose the embedding-as-mask paradigm to unlock the segmentation capability. Remarkably, LISA can handle cases involving complex rea-soning and world knowledge. Also, it demonstrates robust zero-shot capability when trained exclusively on reasoning-free datasets. In addition, fine-tuning the model with merely 239 reasoning segmentation data samples results in further performance enhancement. Both quantitative and qualitative experiments show our method effectively unlocks new reasoning segmentation capabilities for multimodal LLMs. Code, models, and data are available at github.com/dvlab-research/LISA. Zhuotao Tian, Yukang Chen, Yuhui Yuan, Shu Liu 0005, Jiaya Jia |
CVPR | 2 |
| 2024 | OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic SegmentationabstractThe booming of 3D recognition in the 2020s began with the introduction of point cloud transformers. They quickly overwhelmed sparse CNNs and became state-of-the-art models, especially in 3D semantic segmentation. However, sparse CNNs are still valuable networks, due to their efficiency treasure, and ease of application. In this work, we reexamine the design distinctions and test the limits of what a sparse CNN can achieve. We discover that the key credit to the performance difference is adaptivity. Specifically, we propose two key components, i.e., adaptive receptive fields (spatially) and adaptive relation, to bridge the gap. This exploration led to the creation of Omni-Adaptive 3D CNNs (OA-CNNs), a family of networks that integrates a lightweight module to greatly enhance the adaptivity of sparse CNNs at minimal computational cost. Without any self-attention modules, OA-CNNs favorably surpass point transformers in terms of accuracy in both indoor and outdoor scenes, with much less latency and memory cost. Notably, it achieves 76.1%, 78.9%, and 70.6% mIoU on ScanNet v2, nuScenes, and SemanticKITTI validation benchmarks respectively, while maintaining at most 5× better speed than transformer counterparts. This revelation highlights the potential of pure sparse CNNs to outperform transformer-related networks. Our code is built upon Pointcept [9], which is available at here11https://github.com/Pointcept/Pointcept. Bohao Peng, Xiaoyang Wu 0002, Li Jiang 0009, Yukang Chen, Hengshuang Zhao, Zhuotao Tian, Jiaya Jia |
CVPR | 6 |
| 2024 | GroupContrast: Semantic-Aware Self-Supervised Representation Learning for 3D UnderstandingabstractSelf-supervised 3D representation learning aims to learn effective representations from large-scale unlabeled point clouds. Most existing approaches adopt point discrimination as the pretext task, which assigns matched points in two distinct views as positive pairs and unmatched points as negative pairs. However, this approach often results in semantically identical points having dissimilar representations, leading to a high number of false negatives and introducing a “semantic conflict” problem. To address this issue, we propose Group Contrast, a novel approach that combines segment grouping and semantic-aware contrastive learning. Segment grouping partitions points into semantically meaningful regions, which enhances semantic coherence and provides semantic guidance for the subsequent contrastive representation learning. Semantic-aware contrastive learning augments the semantic information extracted from segment grouping and helps to alleviate the issue of “semantic conflict”. We conducted extensive experiments on multiple 3D scene understanding tasks. The results demonstrate that GroupContrast learns semantically meaningful representations and achieves promising transfer learning performance. Chengyao Wang, Li Jiang 0009, Xiaoyang Wu 0002, Zhuotao Tian, Bohao Peng, Hengshuang Zhao, Jiaya Jia |
CVPR | 4 |
| 2024 | SaCo Loss: Sample-Wise Affinity Consistency for Vision-Language Pre-TrainingabstractVision-language pre-training (VLP) aims to learn joint representations of vision and language modalities. The contrastive paradigm is currently dominant in this field. However, we observe a notable misalignment phenomenon, that is, the affinity between samples has an obvious disparity across different modalities, namely “Affinity Inconsistency Problem”. Our intuition is that, for a well-aligned model, two images that look similar to each other should have the same level of similarity as their corresponding texts that describe them. In this paper, we first investigate the reason of this inconsistency problem. We discover that the lack of consideration for sample-wise affinity consistency across modalities in existing training objectives is the central cause. To address this problem, we propose a novel loss function, named Sample-wise affinity Consistency (SaCo) loss, which is designed to enhance such consistency by minimizing the distance between image embedding similarity and text embedding similarity for any two samples. Our SaCo loss can be easily incorporated into existing vision-language models as an additional loss due to its complementarity for most training objectives. In addition, considering that pre-training from scratch is computationally expensive, we also provide a more efficient way to continuously pre-train on a converged model by integrating our loss. Experimentally, the model trained with our SaCo loss significantly outperforms the baseline on a variety of vision and language tasks. Sitong Wu, Haoru Tan, Zhuotao Tian, Yukang Chen, Xiaojuan Qi 0001, Jiaya Jia |
CVPR | 3 |
| 2024 | Towards Large-Scale 3D Representation Learning with Multi-Dataset Point Prompt TrainingabstractThe rapid advancement of deep learning models is often attributed to their ability to leverage massive training data. In contrast, such privilege has not yet fully benefited 3D deep learning, mainly due to the limited availability of large-scale 3D datasets. Merging multiple available data sources and letting them collaboratively train a single model is a potential solution. However, due to the large domain gap between 3D point cloud datasets, such mixed supervision could adversely affect the model's performance and lead to degenerated performance (i.e., negative transfer) compared to single-dataset training. In view of this challenge, we introduce Point Prompt Training (PPT), a novel framework for multi-dataset synergistic learning in the context of 3D representation learning that supports multiple pre-training paradigms. Based on this framework, we propose Prompt-driven Normalization, which adapts the model to different datasets with domain-specific prompts and Language-guided Categorical Alignment that decently unifies the multiple-dataset label spaces by leveraging the relationship between label text. Extensive experiments verify that PPT can overcome the negative transfer associated with synergistic learning and produce generalizable representations. Notably, it achieves state-of-the-art performance on each dataset using a single weight-shared model with supervised multi-dataset training. Moreover, when served as a pre-training framework, it outperforms other pre-training approaches regarding representation quality and attains remarkable state-of-the-art performance across over ten diverse downstream tasks spanning both indoor and outdoor 3D scenarios. Xiaoyang Wu 0002, Zhuotao Tian, Xin Wen 0004, Bohao Peng, Xihui Liu, Kaicheng Yu, Hengshuang Zhao |
CVPR | 2 |
| 2024 | Unified Language-Driven Zero-Shot Domain AdaptationabstractThis paper introduces Unified Language-driven Zero-shot Domain Adaptation (ULDA), a novel task setting that enables a single model to adapt to diverse target domains without explicit domain-ID knowledge. We identify the constraints in the existing language-driven zero-shot domain adaptation task, particularly the requirement for domain IDs and domain-specific models, which may restrict flexibility and scalability. To overcome these issues, we propose a new framework for ULDA, consisting of Hierarchical Context Alignment (HCA), Domain Consistent Representation Learning (DCRL), and Text-Driven Rectifier (TDR). These components work synergistically to align simulated features with target text across multiple visual levels, retain semantic correlations between different regional representations, and rectify biases between simulated and real target visual features, respectively. Our extensive empirical evaluations demonstrate that this framework achieves competitive performance in both settings, surpassing even the model that requires domain-ID, showcasing its superiority and generalization ability. The proposed method is not only effective but also maintains practicality and efficiency, as it does not introduce additional computational costs during inference. The code is available on the project website11Senqiaoyang.com/project/ULDA. Senqiao Yang, Zhuotao Tian, Li Jiang 0009, Jiaya Jia |
CVPR | 2 |
| 2024 | Explore the Potential of CLIP for Training-Free Open Vocabulary Semantic Segmentation
Tong Shao, Zhuotao Tian, Hang Zhao 0019, Jingyong Su |
ECCV (86) | 2 |
| 2024 | Mind the Interference: Retaining Pre-trained Knowledge in Parameter Efficient Continual Learning of Vision-Language Models
Longxiang Tang, Zhuotao Tian, Kai Li 0012, Chunming He, Hantao Zhou, Hengshuang Zhao, Xiu Li 0001, Jiaya Jia |
ECCV (36) | 2 |
| 2024 | Scalable Language Model with Generalized Continual LearningabstractContinual learning has gained increasing importance as it facilitates the acquisition and refinement of scalable knowledge and skills in language models. However, existing methods typically encounter strict limitations and challenges in real-world scenarios, such as reliance on experience replay, optimization constraints, and inference task-ID. In this study, we introduce the Scalable Language Model (SLM) to overcome these limitations within a more challenging and generalized setting, representing a significant advancement toward practical applications for continual learning. Specifically, we propose the Joint Adaptive Re-Parameterization (JARe), integrated with Dynamic Task-related Knowledge Retrieval (DTKR), to enable adaptive adjustment of language models based on specific downstream tasks. This approach leverages the task distribution within the vector space, aiming to achieve a smooth and effortless continual learning process. Our method demonstrates state-of-the-art performance on diverse backbones and benchmarks, achieving effective continual learning in both full-set and few-shot scenarios with minimal forgetting. Moreover, while prior research primarily focused on a single task type such as classification, our study goes beyond, with the large language model, i.e., LLaMA-2, to explore the effects across diverse domains and task types, such that a single language model can be decently scaled to broader applications. The code and models will be released to the public. Bohao Peng, Zhuotao Tian, Shu Liu 0005, Ming-Chang Yang, Jiaya Jia |
ICLR | 2 |
| 2024 | Decoupled Kullback-Leibler Divergence LossabstractIn this paper, we delve deeper into the Kullback–Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss that consists of 1) a weighted Mean Square Error ($\mathbf{w}$MSE) loss and 2) a Cross-Entropy loss incorporating soft labels.
Thanks to the decomposed formulation of DKL loss, we have identified two areas for improvement.
Firstly, we address the limitation of KL/DKL in scenarios like knowledge distillation by breaking its asymmetric optimization property. This modification ensures that the $\mathbf{w}$MSE component is always effective during training, providing extra constructive cues.
Secondly, we introduce class-wise global information into KL/DKL to mitigate bias from individual samples.
With these two enhancements, we derive the Improved Kullback–Leibler (IKL) Divergence loss and evaluate its effectiveness by conducting experiments on CIFAR-10/100 and ImageNet datasets, focusing on adversarial training, and knowledge distillation tasks. The proposed approach achieves new state-of-the-art adversarial robustness on the public leaderboard --- \textit{RobustBench} and competitive performance on knowledge distillation, demonstrating the substantial practical merits. Our code is available at https://github.com/jiequancui/DKL. Jiequan Cui, Zhuotao Tian, Zhisheng Zhong, Xiaojuan Qi 0001, Bei Yu 0001, Hanwang Zhang |
NeurIPS | 2 |
| 2024 | Typicalness-Aware Learning for Failure DetectionabstractDeep neural networks (DNNs) often suffer from the overconfidence issue, where incorrect predictions are made with high confidence scores, hindering the applications in critical systems. In this paper, we propose a novel approach called Typicalness-Aware Learning (TAL) to address this issue and improve failure detection performance.
We observe that, with the cross-entropy loss, model predictions are optimized to align with the corresponding labels via increasing logit magnitude or refining logit direction. However, regarding atypical samples, the image content and their labels may exhibit disparities. This discrepancy can lead to overfitting on atypical samples, ultimately resulting in the overconfidence issue that we aim to address.
To address this issue, we have devised a metric that quantifies the typicalness of each sample, enabling the dynamic adjustment of the logit magnitude during the training process. By allowing relatively atypical samples to be adequately fitted while preserving reliable logit direction, the problem of overconfidence can be mitigated. TAL has been extensively evaluated on benchmark datasets, and the results demonstrate its superiority over existing failure detection methods. Specifically, TAL achieves a more than 5\% improvement on CIFAR100 in terms of the Area Under the Risk-Coverage Curve (AURC) compared to the state-of-the-art. Code is available at https://github.com/liuyijungoon/TAL. Yijun Liu 0012, Jiequan Cui, Zhuotao Tian, Senqiao Yang, Qingdong He, Jingyong Su |
NeurIPS | 3 |
| 2024 | Referencing Where to Focus: Improving Visual Grounding with Referential QueryabstractVisual Grounding aims to localize the referring object in an image given a natural language expression. Recent advancements in DETR-based visual grounding methods have attracted considerable attention, as they directly predict the coordinates of the target object without relying on additional efforts, such as pre-generated proposal candidates or pre-defined anchor boxes. However, existing research primarily focuses on designing stronger multi-modal decoder, which typically generates learnable queries by random initialization or by using linguistic embeddings. This vanilla query generation approach inevitably increases the learning difficulty for the model, as it does not involve any target-related information at the beginning of decoding. Furthermore, they only use the deepest image feature during the query learning process, overlooking the importance of features from other levels. To address these issues, we propose a novel approach, called RefFormer. It consists of the query adaption module that can be seamlessly integrated into CLIP and generate the referential query to provide the prior context for decoder, along with a task-specific decoder. By incorporating the referential query into the decoder, we can effectively mitigate the learning difficulty of the decoder, and accurately concentrate on the target object. Additionally, our proposed query adaption module can also act as an adapter, preserving the rich knowledge within CLIP without the need to tune the parameters of the backbone network. Extensive experiments demonstrate the effectiveness and efficiency of our proposed method, outperforming state-of-the-art approaches on five visual grounding benchmarks. Zhuotao Tian, Qingpei Guo, Sanping Zhou, Ming Yang 0007, Le Wang 0003 |
NeurIPS | 2 |
| 2024 | Generalized Parametric Contrastive LearningabstractIn this paper, we propose the Generalized Parametric Contrastive Learning (GPaCo/PaCo) which works well on both imbalanced and balanced data. Based on theoretical analysis, we observe supervised contrastive loss tends to bias on high-frequency classes and thus increases the difficulty of imbalanced learning. We introduce a set of parametric class-wise learnable centers to rebalance from an optimization perspective. Further, we analyze our GPaCo/PaCo loss under a balanced setting. Our analysis demonstrates that GPaCo/PaCo can adaptively enhance the intensity of pushing samples of the same class close as more samples are pulled together with their corresponding centers and benefit hard example learning. Experiments on long-tailed benchmarks manifest the new state-of-the-art for long-tailed recognition. On full ImageNet, models from CNNs to vision transformers trained with GPaCo loss show better generalization performance and stronger robustness compared with MAE models. Moreover, GPaCo can be applied to semantic segmentation task and obvious improvements are observed on 4 most popular benchmarks. Jiequan Cui, Zhisheng Zhong, Zhuotao Tian, Shu Liu 0005, Bei Yu 0001, Jiaya Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | PFENet++: Boosting Few-Shot Semantic Segmentation With the Noise-Filtered Context-Aware Prior MaskabstractIn this work, we revisit the prior mask guidance proposed in “Prior Guided Feature Enrichment Network for Few-Shot Segmentation”. The prior mask serves as an indicator that highlights the region of interests of unseen categories, and it is effective in achieving better performance on different frameworks of recent studies. However, the current method directly takes the maximum element-to-element correspondence between the query and support features to indicate the probability of belonging to the target class, thus the broader contextual information is seldom exploited during the prior mask generation. To address this issue, first, we propose the Context-aware Prior Mask (CAPM) that leverages additional nearby semantic cues for better locating the objects in query images. Second, since the maximum correlation value is vulnerable to noisy features, we take one step further by incorporating a lightweight Noise Suppression Module (NSM) to screen out the unnecessary responses, yielding high-quality masks for providing the prior knowledge. Both two contributions are experimentally shown to have substantial practical merit, and the new model named PFENet++ significantly outperforms the baseline PFENet as well as all other competitors on three challenging benchmarks PASCAL-5$^{i}$, COCO-20$^{i}$and FSS-1000. The new state-of-the-art performance is achieved without compromising the efficiency, manifesting the potential for being a new strong baseline in few-shot semantic segmentation. Xiaoliu Luo, Zhuotao Tian, Taiping Zhang, Bei Yu 0001, Yuan Yan Tang, Jiaya Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Learning Context-Aware Classifier for Semantic SegmentationabstractSemantic segmentation is still a challenging task for parsing diverse contexts in different scenes, thus the fixed classifier might not be able to well address varying feature distributions during testing. Different from the mainstream literature where the efficacy of strong backbones and effective decoder heads has been well studied, in this paper, additional contextual hints are instead exploited via learning a context-aware classifier whose content is data-conditioned, decently adapting to different latent distributions. Since only the classifier is dynamically altered, our method is model-agnostic and can be easily applied to generic segmentation models. Notably, with only negligible additional parameters and +2\% inference time, decent performance gain has been achieved on both small and large models with challenging benchmarks, manifesting substantial practical merits brought by our simple yet effective method. The implementation is available at https://github.com/tianzhuotao/CAC. Zhuotao Tian, Jiequan Cui, Li Jiang 0009, Xiaojuan Qi 0001, Shu Liu 0005, Jiaya Jia |
AAAI | 1 |
| 2023 | Hierarchical Dense Correlation Distillation for Few-Shot SegmentationabstractFew-shot semantic segmentation (FSS) aims to form class-agnostic models segmenting unseen classes with only a handful of annotations. Previous methods limited to the semantic feature and prototype representation suffer from coarse segmentation granularity and train-set overfitting. In this work, we design Hierarchically Decoupled Matching Network (HDMNet) mining pixel-level support correlation based on the transformer architecture. The self-attention modules are used to assist in establishing hierarchical dense features, as a means to accomplish the cascade matching between query and support features. Moreover, we propose a matching module to reduce train-set overfitting and introduce correlation distillation leveraging semantic correspondence from coarse resolution to boost fine-grained segmentation. Our method performs decently in experiments. We achieve 50.0% mIoU on COCO-20idataset one-shot setting and 56.0% on five-shot segmentation, respectively. The code is available on the project website11https://github.com/Pbihao/HDMNet. Bohao Peng, Zhuotao Tian, Xiaoyang Wu 0002, Chengyao Wang, Shu Liu 0005, Jingyong Su, Jiaya Jia |
CVPR | 2 |
| 2023 | Boosting Few-shot 3D Point Cloud Segmentation via Query-Guided EnhancementabstractAlthough extensive research has been conducted on 3D point cloud segmentation, effectively adapting generic models to novel categories remains a formidable challenge. This paper proposes a novel approach to improve point cloud few-shot segmentation (PC-FSS) models. Unlike existing PC-FSS methods that directly utilize categorical information from support prototypes to recognize novel classes in query samples, our method identifies two critical aspects that substantially enhance model performance by reducing contextual gaps between support prototypes and query features. Specifically, we (1) adapt support background prototypes to match query context while removing extraneous cues that may obscure foreground and background in query samples, and (2) holistically rectify support prototypes under the guidance of query features to emulate the latter having no semantic gap to the query targets. Our proposed designs are agnostic to the feature extractor, rendering them readily applicable to any prototype-based methods. The experimental results on S3DIS and ScanNet demonstrate notable practical benefits, as our approach achieves significant improvements while still maintaining high efficiency. The code for our approach is available at https://github.com/AaronNZH/Boosting-Few-shot-3D-Point-Cloud-Segmentation-via-Query-Guided-Enhancement Zhenhua Ning, Zhuotao Tian, Guangming Lu 0002, Wenjie Pei |
ACM Multimedia | 2 |
| 2023 | ResLT: Residual Learning for Long-Tailed RecognitionabstractDeep learning algorithms face great challenges with long-tailed data distribution which, however, is quite a common case in real-world scenarios. Previous methods tackle the problem from either the aspect of input space (re-sampling classes with different frequencies) or loss space (re-weighting classes with different weights), suffering from heavy over-fitting to tail classes or hard optimization during training. To alleviate these issues, we propose a more fundamental perspective for long-tailed recognition, i.e., from the aspect of parameter space, and aims to preserve specific capacity for classes with low frequencies. From this perspective, the trivial solution utilizes different branches for the head, medium, tail classes respectively, and then sums their outputs as the final results is not feasible. Instead, we design the effective residual fusion mechanism - with one main branch optimized to recognize images from all classes, another two residual branches are gradually fused and optimized to enhance images from medium+tail classes and tail classes respectively. Then the branches are aggregated into final results by additive shortcuts. We test our method on several benchmarks, i.e., long-tailed version of CIFAR-10, CIFAR-100, Places, ImageNet, and iNaturalist 2018. Experimental results manifest the effectiveness of our method. Our code is available at https://github.com/jiequancui/ResLT. Jiequan Cui, Shu Liu 0005, Zhuotao Tian, Zhisheng Zhong, Jiaya Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Adaptive Perspective Distillation for Semantic SegmentationabstractStrong semantic segmentation models require large backbones to achieve promising performance, making it hard to adapt to real applications where effective real-time algorithms are needed. Knowledge distillation tackles this issue by letting the smaller model (student) produce similar pixel-wise predictions to that of a larger model (teacher). However, the classifier, which can be deemed as the perspective by which models perceive the encoded features for yielding observations (i.e., predictions), is shared by all training samples, fitting a universal feature distribution. Since good generalization to the entire distribution may bring the inferior specification to individual samples with a certain capacity, the shared universal perspective often overlooks details existing in each sample, causing degradation of knowledge distillation. In this paper, we propose Adaptive Perspective Distillation (APD) that creates an adaptive local perspective for each individual training sample. It extracts detailed contextual information from each training sample specifically, mining more details from the teacher and thus achieving better knowledge distillation results on the student. APD has no structural constraints to both teacher and student models, thus generalizing well to different semantic segmentation models. Extensive experiments on Cityscapes, ADE20K, and PASCAL-Context manifest the effectiveness of our proposed APD. Besides, APD can yield favorable performance gain to the models in both object detection and instance segmentation without bells and whistles. Zhuotao Tian, Pengguang Chen, Li Jiang 0009, Shu Liu 0005, Hengshuang Zhao, Bei Yu 0001, Ming-Chang Yang, Jiaya Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Generalized Few-shot Semantic SegmentationabstractTraining semantic segmentation models requires a large amount of finely annotated data, making it hard to quickly adapt to novel classes not satisfying this condition. Few- Shot Segmentation (FS-Seg) tackles this problem with many constraints. In this paper, we introduce a new benchmark, called Generalized Few-Shot Semantic Segmentation (GFS- Seg), to analyze the generalization ability of simultaneously segmenting the novel categories with very few examples and the base categories with sufficient examples. It is the first study showing that previous representative state-of-the-art FS-Seg methods fall short in GFS-Seg and the performance discrepancy mainly comes from the constrained setting of FS-Seg. To make GFS-Seg tractable, we set up a GFS-Seg baseline that achieves decent performance without structural change on the original model. Then, since context is essential for semantic segmentation, we propose the Context-Aware Prototype Learning (CAPL) that significantly improves performance by 1) leveraging the co-occurrence prior knowledge from support samples, and 2) dynamically enriching contextual information to the classifier, conditioned on the content of each query image. Both two contributions are experimentally manifested for their substantial practical merit. Extensive experiments on Pascal-Voc and COCO also show that CAPL generalizes well to FS-Seg by achieving competitive performance. Code is available at https://github.com/dvlab-research/GFS-Seg. Zhuotao Tian, Li Jiang 0009, Shu Liu 0005, Michelle Shu, Hengshuang Zhao, Jiaya Jia |
CVPR | 1 |
| 2022 | DecoupleNet: Decoupled Network for Domain Adaptive Semantic Segmentation
Zhuotao Tian, Xiaogang Xu 0002, Ying-Cong Chen, Shu Liu 0005, Hengshuang Zhao, Liwei Wang 0009, Jiaya Jia |
ECCV (33) | 2 |
| 2022 | Spatial Pruned Sparse Convolution for Efficient 3D Object Detectionabstract3D scenes are dominated by a large number of background points, which is redundant for the detection task that mainly needs to focus on foreground objects. In this paper, we analyze major components of existing sparse 3D CNNs and find that 3D CNNs ignores the redundancy of data and further amplifies it in the down-sampling process, which brings a huge amount of extra and unnecessary computational overhead. Inspired by this, we propose a new convolution operator named spatial pruned sparse convolution (SPS-Conv), which includes two variants, spatial pruned submanifold sparse convolution (SPSS-Conv) and spatial pruned regular sparse convolution (SPRS-Conv), both of which are based on the idea of dynamically determine crucial areas for performing computations to reduce redundancy. We empirically find that magnitude of features can serve as an important cues to determine crucial areas which get rid of the heavy computations of learning-based methods. The proposed modules can easily be incorporated into existing sparse 3D CNNs without extra architectural modifications. Extensive experiments on the KITTI and nuScenes datasets demonstrate that our method can achieve more than 50% reduction in GFLOPs without compromising the performance. Yukang Chen, Xiaoqing Ye, Zhuotao Tian, Xiao Tan 0001, Xiaojuan Qi 0001 |
NeurIPS | 4 |
| 2022 | Prior Guided Feature Enrichment Network for Few-Shot SegmentationabstractState-of-the-art semantic segmentation methods require sufficient labeled data to achieve good results and hardly work on unseen classes without fine-tuning. Few-shot segmentation is thus proposed to tackle this problem by learning a model that quickly adapts to new classes with a few labeled support samples. Theses frameworks still face the challenge of generalization ability reduction on unseen classes due to inappropriate use of high-level semantic information of training classes and spatial inconsistency between query and support targets. To alleviate these issues, we propose the Prior Guided Feature Enrichment Network (PFENet). It consists of novel designs of (1) a training-free prior mask generation method that not only retains generalization power but also improves model performance and (2) Feature Enrichment Module (FEM) that overcomes spatial inconsistency by adaptively enriching query features with support features and prior masks. Extensive experiments on PASCAL-5$^i$and COCO prove that the proposed prior generation method and FEM both improve the baseline method significantly. Our PFENet also outperforms state-of-the-art methods by a large margin without efficiency loss. It is surprising that our model even generalizes to cases without labeled support samples. Zhuotao Tian, Hengshuang Zhao, Michelle Shu, Ruiyu Li, Jiaya Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Semi-Supervised Semantic Segmentation With Directional Context-Aware ConsistencyabstractSemantic segmentation has made tremendous progress in recent years. However, satisfying performance highly depends on a large number of pixel-level annotations. Therefore, in this paper, we focus on the semi-supervised segmentation problem where only a small set of labeled data is provided with a much larger collection of totally unlabeled images. Nevertheless, due to the limited annotations, models may overly rely on the contexts available in the training data, which causes poor generalization to the scenes un-seen before. A preferred high-level representation should capture the contextual information while not losing self-awareness. Therefore, we propose to maintain the context-aware consistency between features of the same identity but with different contexts, making the representations robust to the varying environments. Moreover, we present the Directional Contrastive Loss (DC Loss) to accomplish the consistency in a pixel-to-pixel manner, only requiring the feature with lower quality to be aligned towards its counterpart. In addition, to avoid the false-negative samples and filter the uncertain positive samples, we put forward two sampling strategies. Extensive experiments show that our simple yet effective method surpasses current state-of-the-art methods by a large margin and also generalizes well with extra image-level annotations. Zhuotao Tian, Li Jiang 0009, Shu Liu 0005, Hengshuang Zhao, Liwei Wang 0009, Jiaya Jia |
CVPR | 2 |
| 2021 | Guided Point Contrastive Learning for Semi-supervised Point Cloud Semantic SegmentationabstractRapid progress in 3D semantic segmentation is inseparable from the advances of deep network models, which highly rely on large-scale annotated data for training. To address the high cost and challenges of 3D point-level labeling, we present a method for semi-supervised point cloud semantic segmentation to adopt unlabeled point clouds in training to boost the model performance. Inspired by the recent contrastive loss in self-supervised tasks, we propose the guided point contrastive loss to enhance the feature representation and model generalization ability in semi-supervised setting. Semantic predictions on unlabeled point clouds serve as pseudo-label guidance in our loss to avoid negative pairs in the same category. Also, we design the confidence guidance to ensure high-quality feature learning. Besides, a category-balanced sampling strategy is proposed to collect positive and negative samples to mitigate the class imbalance problem. Extensive experiments on three datasets (ScanNet V2, S3DIS, and SemanticKITTI) show the effectiveness of our semi-supervised method to improve the prediction quality with unlabeled data. Li Jiang 0009, Shaoshuai Shi, Zhuotao Tian, Shu Liu 0005, Chi-Wing Fu, Jiaya Jia |
ICCV | 3 |
| 2019 | Homomorphic Latent Space Interpolation for Unpaired Image-To-Image TranslationabstractGenerative adversarial networks have achieved great success in unpaired image-to-image translation. Cycle consistency allows modeling the relationship between two distinct domains without paired data. In this paper, we propose an alternative framework, as an extension of latent space interpolation, to consider the intermediate region between two domains during translation. It is based on the fact that in a flat and smooth latent space, there exist many paths that connect two sample points. Properly selecting paths makes it possible to change only certain image attributes, which is useful for generating intermediate images between the two domains. We also show that this framework can be applied to multi-domain and multi-modal translation. Extensive experiments manifest its generality and applicability to various tasks. Ying-Cong Chen, Xiaogang Xu 0002, Zhuotao Tian, Jiaya Jia |
CVPR | 3 |
| 2019 | Learning Shape-Aware Embedding for Scene Text DetectionabstractWe address the problem of detecting scene text in arbitrary shapes, which is a challenging task due to the high variety and complexity of the scene. Specifically, we treat text detection as instance segmentation and propose a segmentation-based framework, which extracts each text instance as an independent connected component. To distinguish different text instances, our method maps pixels onto an embedding space where pixels belonging to the same text are encouraged to appear closer to each other and vise versa. In addition, we introduce a Shape-Aware Loss to make training adaptively accommodate various aspect ratios of text instances and the tiny gaps among them, and a new post-processing pipeline to yield precise bounding box predictions. Experimental results on three challenging datasets (ICDAR15, MSRA-TD500 and CTW1500) demonstrate the effectiveness of our work. Zhuotao Tian, Michelle Shu, Pengyuan Lv, Ruiyu Li, Chao Zhou 0001, Xiaoyong Shen, Jiaya Jia |
CVPR | 1 |