VLDB 2026 Research / reviewers in the wild / expert
Hongtao Xie 0001
dblp:25/588-1
· DBLP profile ↗
195ranked-venue papers
12as first author
122since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 154 · 11 first-author · 96 since 2021Artificial intelligence and machine learning · 71 · 1 first-author · 53 since 2021Databases, data management, data science and information retrieval · 14 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 4 since 2021Systems, architecture and hardware · 3 · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document UnderstandingabstractMulti-modal Retrieval-Augmented Generation (RAG) has become a critical method for empowering LLMs by leveraging candidate visual documents. However, current methods consider the entire document as the basic retrieval unit, introducing substantial irrelevant visual content in two ways: 1) Relevant documents often contain large regions unrelated to the query, diluting the focus on salient information; 2) Retrieving multiple documents to increase recall further introduces redundant and irrelevant documents. These redundant contexts distract the model's attention and further degrade the performance. To address this challenge, we propose RegionRAG, a novel framework that shifts the retrieval paradigm from the document level to the region level. During training, we design a hybrid supervision strategy from both labeled data and unlabeled data to pinpoint relevant patches. During inference, we propose a dynamic pipeline that intelligently groups salient patches into complete semantic regions. By delegating the task of identifying relevant regions to the retriever, RegionRAG enables the generator to focus solely on concise, query-relevant visual content, improving both efficiency and accuracy. Experiments on six benchmarks demonstrate that RegionRAG achieves state-of-the-art performance. It improves retrieval accuracy by 10.02% in R@1 on average, and boosts question answering accuracy by 3.56% while using only 71.42% visual tokens compared with prior methods. Yinglu Li, Zhiying Lu, Chuanbin Liu 0001, Hongtao Xie 0001 |
AAAI | 6 |
| 2026 | SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding CapabilityabstractMultimodal Large Language Models (MLLMs) have shown remarkable progress in temporal or spatial localization tasks, but struggle with joint spatio-temporal video grounding (STVG). We identify two key bottlenecks hindering this capability: (1) the sheer number of visual tokens makes long-range and fine-grained visual modeling challenging; (2) generating a long sequence of bounding boxes in text makes it hard to accurately align each box with its specific video frame. Distinct from prior efforts that rely on attaching complex modules, we argue for a more elegant paradigm that unlocks the inherent potential of MLLMs and leverages their strengths. To this end, we propose \textbf{\textit{SpaceVLLM}}, a MLLM equipped with spatio-temporal video grounding capabilities. Specifically, we propose Spatio-Temporal Aware Queries, interleaved with video frames, to guide the MLLM in capturing both static appearance and dynamic motion features. We further present a lightweight Query-Guided Space Head that maps queries to precise spatial coordinates, bypassing the need for direct textual coordinate generation and enabling the MLLM to focus on video understanding. To further facilitate research in this area, we propose an automated data synthesis pipeline to construct \textbf{V-STG} dataset, comprising 110K STVG instances. Extensive experiments show that \textit{SpaceVLLM} achieves the state-of-the-art performance on STVG benchmarks and maintains strong performance on various video understanding tasks, validating our approach's effectiveness. Jiankang Wang, Jiannan Ge, Hongtao Xie 0001, Yongdong Zhang 0001 |
AAAI | 6 |
| 2026 | LRANet++: Low-Rank Approximation Network for Accurate and Efficient Text SpottingabstractEnd-to-end text spotting aims to jointly optimize text detection and recognition within a unified framework. Despite significant progress, designing an accurate and efficient end-to-end text spotter for arbitrary-shaped text remains challenging. We identify the primary bottleneck as the lack of a reliable and efficient text detection method. To address this, we propose a novel parameterized text shape representation based on low-rank approximation for precise detection and a triple assignment detection head for fast inference. Specifically, unlike current data-irrelevant shape representation methods, we exploit shape correlations among labeled text boundaries to construct a robust low-rank subspace. By minimizing an $\ell _{1}$ℓ1-norm objective, we extract orthogonal vectors that capture the intrinsic text shape from noisy annotations, enabling precise reconstruction via the linear combination of only a few basis vectors. Next, the triple assignment scheme decouples training complexity from inference speed. It utilizes a deep sparse branch to guide an ultra-lightweight inference branch, while a dense branch provides rich parallel supervision. Building upon these advancements, we integrate the enhanced detection module with a lightweight recognition branch to form an end-to-end text spotting framework, termed LRANet++, capable of accurately and efficiently spotting arbitrary-shaped text. Extensive experiments on challenging benchmarks demonstrate the superiority of LRANet++ compared to state-of-the-art methods. Zhineng Chen, Yongkun Du, Zuxuan Wu, Hongtao Xie 0001, Yu-Gang Jiang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | HiFi3D: Improving Text-to-3D With High-Fidelity Multi-View DiffusionabstractRecent advances in score distillation sampling (SDS) have revolutionized the field of text-to-3D generation, enabling the distillation of prior knowledge from diffusion models for 3D generation. Although exhibiting impressive texture quality, these methods often suffer from geometric inconsistencies (“Janus” issue), as the prior 2D diffusion model inherently lacks 3D awareness. Recent work fine-tunes the 2D diffusion model on 3D data to obtain a multi-view diffusion model as the SDS prior, which addresses the Janus issue but is at the cost of sacrificing texture quality, as available 3D training data always have unrealistic texture. Thus, a natural question arises — Is therean ideal prior diffusion model for 3D generationthat simultaneously has 3D awareness and high texture fidelity? In response, we present HiFi3D, a tuning-free method to establish a new hybrid diffusion model that can generate consistent multi-view images with photorealism appearances. We accomplish this by novelly marring a 3D multi-view diffusion model with a 2D image diffusion model through our unique designs. We find that such a high-fidelity multi-view diffusion model harbors an innate agency to serve as a strong prior for SDS optimization. Additionally, we introduce a depth-guided multi-view attention strategy to further improve the 3D consistency across views during optimization. Extensive experiments demonstrate that our HiFi3D outperforms previous state-of-the-art methods in faithfully generating 3D content with realistic textual details and consistent geometry. Runxin Liu, Yang Chen 0048, Yingwei Pan, Hongtao Xie 0001, Yongdong Zhang 0001, Ting Yao 0003, Tao Mei 0001 |
IEEE Trans. Multim. | 4 |
| 2026 | SwapController: Toward Improving Identity and Attribute Control for Diffusion-Based Face SwappingabstractFace swapping efforts strive to achieve high-fidelity and well-controlled generation effects. Owing to the remarkable generative capabilities, diffusion models deliver promising high-fidelity solutions. However, their intrinsic stochastic properties complicate the accurate modeling of facial representations, introducing new challenges for identity and attribute consistency of the generated faces. In this paper, we introduce a novel diffusion-based face-swapping framework, named SwapController, which achieves high-fidelity generation via careful facial identity and attribute modeling. Specifically, our facial modeling mainly involves facial structure and facial texture. For structure modeling, 3D facial priors are leveraged to provide explicit structure supervision, enabling accurate head structure control. On this basis, two novel components are proposed to deeply mine facial textural representations from identity and attribute aspects. To enhance identity control, multi-grained source identity embeddings are obtained from various functional encoders to convey critical global identity and fine-grained identity details. To improve attribute modeling, identity-shifted attribute embeddings are derived by applying identity modulation to the most salient textural attribute features of the target face. Moreover, in line with the diffusion denoising characteristics, a timestep-aware identity optimization objective is introduced to optimize identity consistency guidelines and overall fidelity. Extensive experiments demonstrate the effectiveness of our SwapController in generating identity-consistent portrait images while faithfully preserving target attributes, which obtains a 98.32 ID Retrieval, exceeding the SOTA DiffSFSR by 7.32 $\uparrow$↑. Lingyun Yu 0002, Quanwei Yang, Runxin Liu, Yongdong Zhang 0001, Hongtao Xie 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | IDseq: Decoupled and Sequentially Detecting and Grounding Multi-Modal Media ManipulationabstractDetecting and grounding multi-modal media manipulation aims to categorize the type and localize the region of manipulation for image-text pairs in both two modalities. Existing methods have not sufficiently explored the intrinsic properties of the manipulated images, which contain both forgery and content features, leading to inefficient utilization. To address this problem, we propose an Image-Driven Decoupled Sequential Framework (IDseq), designed to decouple image features and rationally integrate them to accomplish different sub-tasks effectively. Specifically, IDseq employs two specially designed disentangled losses to guide the disentangled learning of forgery and content features. To efficiently leverage these features, we propose a Decoupled Image Manipulation Decoder (DIMD) that processes image tasks within a decoupled schema. We mitigate their exclusive competition by separating the image tasks into forgery-relevant and content-relevant components and training them without gradient interaction. Additionally, we utilize content features enhanced by the proposed Manipulation Indicator Generator (MIG) for the text tasks, which provide the maximal visual information as a reference while eliminating interference from unverified image data. Extensive experiments show the superiority of our IDseq, where it notably outperforms SOTA methods on the fine-grained classification by 3.8% in mAP and the forgery face grounding by 8.7% in IoUmean, even 1.3% in F1 on the most challenging manipulated text grounding. Runxin Liu, Jiaming Li 0017, Lingyun Yu 0002, Hongtao Xie 0001 |
AAAI | 5 |
| 2025 | PosterMaker: Towards High-Quality Product Poster Generation with Accurate Text RenderingabstractProduct posters, which integrate subject, scene, and text, are crucial promotional tools for attracting customers. Creating such posters using modern image generation methods is valuable, while the main challenge lies in accurately rendering text, especially for complex writing systems like Chinese, which contains over 10,000 individual characters. In this work, we identify the key to precise text rendering as constructing a character-discriminative visual feature as a control signal. Based on this insight, we propose a robust character-wise representation as control and we develop TextRenderNet, which achieves a high text rendering accuracy of over 90%. Another challenge in poster generation is maintaining the fidelity of user-specific products. We address this by introducing SceneGenNet, an inpainting-based model, and propose subject fidelity feedback learning to further enhance fidelity. Based on TextRenderNet and SceneGenNet, we present PosterMaker, an end-to-end generation framework. To optimize PosterMaker efficiently, we implement a two-stage training strategy that decouples text rendering and background generation learning. Experimental results show that PosterMaker outperforms existing baselines by a remarkable margin, which demonstrates its effectiveness. Yifan Gao 0011, Zihang Lin, Chuanbin Liu 0001, Tiezheng Ge, Bo Zheng 0007, Hongtao Xie 0001 |
CVPR | 7 |
| 2025 | Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language ModelsabstractRecent Multi-modal Large Language Models (MLLMs) have been challenged by the computational overhead resulting from massive video frames, often alleviated through compression strategies. However, the visual content is not equally contributed to user instructions, existing strategies (e.g., average pool) inevitably lead to the loss of potentially useful information. To tackle this, we propose the Hybridlevel Instruction Injection Strategy for Conditional Token Compression in MLLMs (HICom), utilizing the instruction as a condition to guide the compression from both local and global levels. This encourages the compression to retain the maximum amount of user-focused information while reducing visual tokens to minimize computational burden. Specifically, the instruction condition is injected into the grouped visual tokens at the local level and the learnable tokens at the global level, and we conduct the attention mechanism to complete the conditional compression. From the hybrid-level compression, the instruction-relevant visual parts are highlighted while the temporal-spatial structure is also preserved for easier understanding of LLMs. To further unleash the potential of HICom, we introduce a new conditional pre-training stage with our proposed dataset HICom-248K. Experiments show that our HICom can obtain distinguished video understanding ability with fewer tokens, increasing the performance by 2.43% average on three multiple-choice QA benchmarks and saving 78.8% tokens compared with the SOTA method. The code is available at https://github.com/lntzm/HICom. Chen-Wei Xie, Pandeng Li, Longxiang Tang, Chuanbin Liu 0001, Hongtao Xie 0001 |
CVPR | 8 |
| 2025 | Mask^2DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video GenerationabstractSora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video generation, which offers broader applications, remains relatively underexplored. To bridge this gap, we propose Mask2DiT, a novel approach that establishes fine-grained, one-to-one alignment between video segments and their corresponding text annotations. Specifically, we introduce a symmetric binary mask at each attention layer within the DiT architecture, ensuring that each text annotation applies exclusively to its respective video segment while preserving temporal coherence across visual tokens. This attention mechanism enables precise segment-level textual-to-visual alignment, allowing the DiT architecture to effectively handle video generation tasks with a fixed number of scenes. To further equip the DiT architecture with the ability to generate additional scenes based on existing ones, we incorporate a segment-level conditional mask, which conditions each newly generated segment on the preceding video segments, thereby enabling auto-regressive scene extension. Both qualitative and quantitative experiments confirm that Mask2DiT excels in maintaining visual consistency across segments while ensuring semantic alignment between each segment and its corresponding text description. Our project page is https://tianhao-qi.github.io/Mask2DiTProject/. Tianhao Qi, Jianlong Yuan, Wanquan Feng, Shancheng Fang, Jiawei Liu 0001, SiYu Zhou 0002, Hongtao Xie 0001, Yongdong Zhang 0001 |
CVPR | 8 |
| 2025 | SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled SynthesisabstractDue to the limited scale of multimodal table understanding (MTU) data, model performance is constrained. A straightforward approach is to use multimodal large language models to obtain more samples, but this may cause hallucinations, generate incorrect sample pairs, and cost significantly. To address the above issues, we design a simple yet effective synthesis framework that consists of two independent steps: table image rendering and table question and answer (Q&A) pairs generation. We use table codes (HTML, LaTeX, Markdown) to synthesize images and generate Q&A pairs with large language model (LLM). This approach leverages LLMs high concurrency and low cost to boost annotation efficiency and reduce expenses. By inputting code instead of images, LLMs can directly access the content and structure of the table, reducing hallucinations in table understanding and improving the accuracy of generated Q&A pairs. Finally, we synthesize a large-scale MTU dataset, SynTab, containing 636K images and 1.8M samples costing within $200 in US dollars. We further introduce a generalist tabular multimodal model, SynTab-LLaVA. This model not only effectively extracts local textual content within the table but also enables global modeling of relationships between cells. SynTab-LLaVA achieves SOTA performance on 21 out of 24 in-domain and out-of-domain benchmarks, demonstrating the effectiveness and generalization of our method. The Code is available at SynTab-LLaVA. Bangbang Zhou, Zuan Gao, Boqiang Zhang, Zhineng Chen, Hongtao Xie 0001 |
CVPR | 7 |
| 2025 | CPL: Curriculum Pseudo Labeling for Weakly Supervised Temporal Forgery LocalizationabstractIn forgery detection, temporal forgery localization offers a more nuanced perspective than binary detection by providing more precise temporal boundaries of manipulations. However, its need for frame-wise annotations limits real-world practicality. Therefore, we present the task of Weakly Supervised Temporal Forgery Localization (WS-TFL), which aims to localize forgery segments in videos given only video-level labels for training. To tackle the absence of frame-wise annotations in weakly supervised settings, pseudo-label learning presents a viable solution for WS-TFL. However, pseudo labels often suffer from inaccuracy (i.e. contain noise) due to lack of supervision. Inspired by the ability of curriculum learning in handling noisy data, this work proposes Curriculum Pseudo Labeling (CPL), a simple yet effective strategy to address noise in pseudo labels. Specifically, pseudo-label learning leverages proposals generated from a base model as coarse pseudo labels to retrain the model for better performance. To refine pseudo labels and reduce noise, we design a learning curriculum that ranks them by quality, where we analyze the multiple-instance learning essence of WS-TFL and further introduce forgery prior knowledge. Experiments show that CPL consistently improves performance across different baselines, with a 35.24% and 27.63% increase in average precision to the default baseline on Lav-DF and TVIL datasets, respectively. Dijia Zhang, Mingqi Fang, Zhiying Lu, Hongtao Xie 0001 |
ICASSP | 4 |
| 2025 | SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionabstractConnectionist temporal classification (CTC)-based scene text recognition (STR) methods, e.g., SVTR, are widely employed in OCR applications, mainly due to their simple architecture, which only contains a visual model and a CTC-aligned linear classifier, and therefore fast inference. However, they generally exhibit worse accuracy than encoder-decoder-based methods (EDTRs) due to struggling with text irregularity and linguistic missing. To address these challenges, we propose SVTRv2, a CTC model endowed with the ability to handle text irregularities and model linguistic context. First, a multi-size resizing strategy is proposed to resize text instances to appropriate predefined sizes, effectively avoiding severe text distortion. Meanwhile, we introduce a feature rearrangement module to ensure that visual features accommodate the requirement of CTC, thus alleviating the alignment puzzle. Second, we propose a semantic guidance module. It integrates linguistic context into the visual features, allowing CTC model to leverage language information for accuracy improvement. This module can be omitted at the inference stage and would not increase the time cost. We extensively evaluate SVTRv2 in both standard and recent challenging benchmarks, where SVTRv2 is fairly compared to popular STR models across multiple scenarios, including different types of text irregularity, languages, long text, and whether employing pretraining. SVTRv2 surpasses most EDTRs across the scenarios in terms of accuracy and inference speed. Code: https://github.com/Topdu/OpenOCR. Yongkun Du, Zhineng Chen, Hongtao Xie 0001, Caiyan Jia, Yu-Gang Jiang 0001 |
ICCV | 3 |
| 2025 | Forensic-MoE: Exploring Comprehensive Synthetic Image Detection Traces With Mixture of Experts
Mingqi Fang, Ziguang Li, Lingyun Yu 0002, Quanwei Yang, Hongtao Xie 0001, Yongdong Zhang 0001 |
ICCV | 5 |
| 2025 | CLIP-Adapted Region-to-Text Learning for Generative Open-Vocabulary Semantic Segmentation
Jiannan Ge, Lingxi Xie, Hongtao Xie 0001, Pandeng Li, Sun-Ao Liu, Xiaopeng Zhang 0008, Qi Tian 0001, Yongdong Zhang 0001 |
ICCV | 3 |
| 2025 | Igd: Instructional Graphic Design With Multimodal Layer Generatio
Yadong Qu, Hongtao Xie 0001, Yongdong Zhang 0001, Shancheng Fang, Yuxin Wang 0002, Zhineng Chen |
ICCV | 2 |
| 2025 | Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design
Gaowen Liu, Hongtao Xie 0001, Sijia Liu 0001 |
ICCV | 4 |
| 2025 | GestureHYDRA: Semantic Co-Speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
Quanwei Yang, Luying Huang, Kaisiyuan Wang, Jiazhi Guan, Shengyi He, Fengguo Li, Hang Zhou 0009, Lingyun Yu 0002, Haocheng Feng, Hongtao Xie 0001 |
ICCV | 11 |
| 2025 | IterMeme: Expert-Guided Multimodal LLM for Interactive Meme Creation with Layout-Aware GenerationabstractMeme creation is a creative process that blends images and text. However, existing methods lack critical components, failing to support intent-driven caption-layout generation and personalized generation, making it difficult to generate high-quality memes. To address this limitation, we propose IterMeme, an end-to-end interactive meme creation framework that utilizes a unified Multimodal Large Language Model (MLLM) to facilitate seamless collaboration among multiple components. To overcome the absence of a caption-layout generation component, we develop a robust layout representation method and construct a large-scale image-caption-layout dataset, MemeCap, which enhances the model’s ability to comprehend emotions and coordinate caption-layout generation effectively. To address the lack of a personalization component, we introduce a parameter-shared dual-LLM architecture that decouples the intricate representations of reference images and text. Furthermore, we incorporate the expert-guided M³OE for fine-grained identity properties (IP) feature extraction and cross-modal fusion. By dynamically injecting features into every layer of the model, we enable adaptive refinement of both visual and semantic information. Experimental results demonstrate that IterMeme significantly advances the field of meme creation by delivering consistently high-quality outcomes. The code, model, and dataset will be open-sourced to the community. Yaqi Cai, Shancheng Fang, Yadong Qu, Meng Shao, Hongtao Xie 0001 |
IJCAI | 6 |
| 2025 | Proactive Deepfake Detection via Self-Verifiable Semantic WatermarkingabstractMalicious Deepfakes pose serious security risks by producing highly realistic forged faces. While numerous countermeasures have been developed to train binary Deepfake classifiers, their limited generalization capacity restricts practical deployment. To proactively defend against Deepfakes, we propose SVS-WM, a Self-Verifiable Semantic Watermarking strategy. The core idea behind SVS-WM is to embed pairs of correlated watermarks within facial semantics, leveraging the inherent fragility of these features, i.e., any semantic modification will disrupt the watermark correlation, thereby enabling robust Deepfake detection. SVS-WM employs a facial semantic disentanglement and reconstruction network, allowing semi-fragile watermarks to be embedded concurrently across multiple semantic levels, including identity and multi-levels of attributes. Specifically, pairs of pseudo-random noise watermarks are adaptively injected into facial attribute and identity features. During propagation stage, the protected image may encounter identity or facial attributes manipulations, we then detect Deepfakes by verifying the correlation result between the decoded attribute watermark and the extracted identity vector. This unique cross-verification mechanism enables authentication without requiring original reference watermark, thereby realizing blind Deepfake detection. Extensive experiments validate the effectiveness of our approach, achieving an average detection accuracy of 98.19% across diverse Deepfake manipulations. Peiqi Jiang, Bohan Lei, Lingyun Yu 0002, Zhineng Chen, Hongtao Xie 0001, Yongdong Zhang 0001 |
ACM Multimedia | 6 |
| 2025 | CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and ThoroughnessabstractVisual captioning benchmarks have become outdated with the emergence of modern multimodal large language models (MLLMs), as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent benchmarks attempt to address this by focusing on keyword extraction or object-centric evaluation, they remain limited to vague-view or object-view analyses and incomplete visual element coverage. In this paper, we introduce CAPability, a comprehensive multi-view benchmark for evaluating visual captioning across 12 dimensions spanning six critical views. We curate nearly 11K human-annotated images and videos with visual element annotations to evaluate the generated captions. CAPability stably assesses both the correctness and thoroughness of captions with \textit{precision} and \textit{hit} metrics. By converting annotations to QA pairs, we further introduce a heuristic metric, \textit{know but cannot tell} ($K\bar{T}$), indicating a significant performance gap between QA and caption capabilities. Our work provides a holistic analysis of MLLMs' captioning abilities, as we identify their strengths and weaknesses across various dimensions, guiding future research to enhance specific aspects of their capabilities. Chen-Wei Xie, Feiwu Yu, Jixuan Chen, Pandeng Li, Boqiang Zhang, Nianzu Yang, Yinglu Li, Zuan Gao, Hongtao Xie 0001 |
NeurIPS | 12 |
| 2025 | GRIP: A Graph-Based Reasoning Instruction ProducerabstractLarge-scale, high-quality data is essential for advancing the reasoning capabilities of large language models (LLMs). As publicly available Internet data becomes increasingly scarce, synthetic data has emerged as a crucial research direction. However, existing data synthesis methods often suffer from limited scalability, insufficient sample diversity, and a tendency to overfit to seed data, which constrains their practical utility. In this paper, we present \textit{\textbf{GRIP}}, a \textbf{G}raph-based \textbf{R}easoning \textbf{I}nstruction \textbf{P}roducer that efficiently synthesizes high-quality and diverse reasoning instructions. \textit{GRIP} constructs a knowledge graph by extracting high-level concepts from seed data, and uniquely leverages both explicit and implicit relationships within the graph to drive large-scale and diverse instruction data synthesis, while employing open-source multi-model supervision to ensure data quality. We apply \textit{GRIP} to the critical and challenging domain of mathematical reasoning. Starting from a seed set of 7.5K math reasoning samples, we construct \textbf{GRIP-MATH}, a dataset containing 2.1 million synthesized question-answer pairs. Compared to similar synthetic data methods, \textit{GRIP} achieves greater scalability and diversity while also significantly reducing costs. On mathematical reasoning benchmarks, models trained with GRIP-MATH demonstrate substantial improvements over their base models and significantly outperform previous data synthesis methods. Jiankang Wang, Yuxin Wang 0002, Mengting Xing, Shancheng Fang, Hongtao Xie 0001 |
NeurIPS | 7 |
| 2025 | THGS: Lifelike Talking Human Avatar Synthesis From Monocular Video Via 3D Gaussian SplattingabstractAbstract Despite the remarkable progress in 3D talking head generation, directly generating 3D talking human avatars still suffers from rigid facial expressions, distorted hand textures and out‐of‐sync lip movements. In this paper, we extend speaker‐specific talking head generation task to talking human avatar synthesis and propose a novel pipeline, THGS, that animates lifelike Talking Human avatars using 3D Gaussian Splatting (3DGS). Given speech audio, expression and body poses as input, THGS effectively overcomes the limitations of 3DGS human re‐construction methods in capturing expressive dynamics, such as mouth movements, facial expressions and hand gestures, from a short monocular video. Firstly, we introduce a simple yet effective Learnable Expression Blendshapes (LEB) for facial dynamics re‐construction, where subtle facial dynamics can be generated by linearly combining the static head model and expression blendshapes. Secondly, a Spatial Audio Attention Module (SAAM) is proposed for lip‐synced mouth movement animation, building connections between speech audio and mouth Gaussian movements. Thirdly, we employ a body pose, expression and skinning weights joint optimization strategy to optimize these parameters on the fly, which aligns hand movements and expressions better with video input. Experimental results demonstrate that THGS can achieve high‐fidelity 3D talking human avatar animation at 150+ fps on a web‐based rendering system, improving the requirements of real‐time applications. Our project page is at https://sora158.github.io/THGS.github.io/ . Lingyun Yu 0002, Quanwei Yang, Aihua Zheng, Hongtao Xie 0001 |
Comput. Graph. Forum | 5 |
| 2025 | DHVT: Dynamic Hybrid Vision Transformer for Small Dataset RecognitionabstractThe performance gap between Vision Transformers (ViTs) and Convolutional Neural Networks (CNNs) persists due to the lack of inductive bias, notably when training from scratch with limited datasets. This paper identifies two crucial shortcomings in ViTs: spatial relevance and diverse channel representation. Thus, ViTs struggle to grasp fine-grained spatial features and robust channel representation due to insufficient data. We propose the Dynamic Hybrid Vision Transformer (DHVT) to address these challenges. Regarding the spatial aspect, DHVT introduces convolution in the feature embedding phase and feature projection modules to enhance spatial relevance. Regarding the channel aspect, the dynamic aggregation mechanism and a groundbreaking design "head token" facilitate the recalibration and harmonization of disparate channel representations. Moreover, we investigate the choices of the network meta-structure and adopt the optimal multi-stage hybrid structure without the conventional class token. The methods are then modified with a novel dimensional variable residual connection mechanism to leverage the potential of the structure sufficiently. This updated variant, called DHVT2, offers a more computationally efficient solution for vision-related tasks. DHVT and DHVT2 achieve state-of-the-art image recognition results, effectively bridging the performance gap between CNNs and ViTs. The downstream experiments further demonstrate their strong generalization capacities. Zhiying Lu, Chuanbin Liu 0001, Xiaojun Chang, Yongdong Zhang 0001, Hongtao Xie 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | A Detail-Aware Transformer to Generalizable Face Forgery DetectionabstractGeneralisable face forgery detectors strive to detect forgeries generated by unseen manipulations. Recently advanced detection methods have managed to capture subtle blending traces, but their neglect of the diversity of blending traces in different regions leads to limited generalization. Towards this, transformer with global receptive fields and dynamic weight mechanism is a promising solution, but vanilla transformer is weak at capturing subtle blending traces. In this paper, we propose a novel Detail-Aware Transformer (DAT) able to focus on both diverse and subtle blending traces caused by inconsistencies in the low-level image details. The intrinsic multi-head self-attention mechanism of the transformer allows our DAT to adaptively capture diverse blending traces in different regions. Furthermore, we improve the transformer’s capability of capturing subtle blending traces by two inference overhead-free measures,$i.e$., self-supervised pre-training based on patch augmentation and region-level contrastive learning. Specifically, the self-supervised pre-training encourages the model to focus on the inconsistencies in low-level image details through a patch number prediction task. The region-level contrastive learning employs a contrastive loss on representations of regions with different low-level details to further improve the transformer’s ability to handle subtle blending traces. Extensive experiments show that our method substantially improves the generalization performance and outperforms the state-of-the-art methods on CDF, DFDC, DFDCP, FFIW, and WildDeepfake datasets. Jiaming Li 0017, Lingyun Yu 0002, Runxin Liu, Hongtao Xie 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | ReferSAM: Unleashing Segment Anything Model for Referring Image SegmentationabstractThe Segment Anything Model (SAM) has demonstrated remarkable capability as a general segmentation model given visual prompts such as points or boxes. While SAM is conceptually compatible with text prompts, it merely employs linguistic features from vision-language models as prompt embeddings and lacks fine-grained cross-modal interaction. This deficiency limits its application in referring image segmentation (RIS), where the targets are specified by free-form natural language expressions. In this paper, we introduce ReferSAM, a novel SAM-based framework that enhances cross-modal interaction and reformulates prompt encoding, thereby unleashing SAM’s segmentation capability for RIS. Specifically, ReferSAM incorporates the Vision-Language Interactor (VLI) to integrate linguistic features with visual features during the image encoding stage of SAM. This interactor introduces fine-grained alignment between linguistic features and multi-scale visual representations without altering the architecture of pre-trained models. Additionally, we present the Vision-Language Prompter (VLP) to generate dense and sparse prompt embeddings by aggregating the aligned linguistic and visual features. Consequently, the generated embeddings sufficiently prompt SAM’s mask decoder to provide precise segmentation results. Extensive experiments on five public benchmarks demonstrate that ReferSAM achieves state-of-the-art performance on both classic and generalized RIS tasks. The code and models are available at https://github.com/lsa1997/ReferSAM. Sun'ao Liu, Hongtao Xie 0001, Jiannan Ge, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Distilling Multi-Level Semantic Cues Across Multi-Modalities for Face Forgery DetectionabstractExisting face forgery detection methods attempt to identify low-level forgery artifacts (e.g., blending boundary, flickering) in spatial-temporal domains or high-level semantic inconsistencies (e.g., abnormal lip movements) between visual-auditory modalities for generalized face forgery detection. However, they still suffer from significant performance degradation when dealing with out-of-domain artifacts, as they only consider single semantic mode inconsistencies, but ignore the complementarity of forgery traces at different levels and different modalities. In this paper, we propose a novel Multi-modal Multi-level Semantic Cues Distillation Detection framework that adopts the teacher-student protocol to focus on both spatial-temporal artifacts and visual-auditory incoherence to capture multi-level semantic cues. Specifically, our framework primarily comprises the Spatial-Temporal Pattern Learning module and the Visual-Auditory Consistency Modeling module. The Spatial-Temporal Pattern Learning module employs a mask-reconstruction strategy, in which the student network learns diverse spatial-temporal patterns from a pixel-wise teacher network to capture low-level forgery artifacts. The Visual-Auditory Consistency Modeling module is designed to enhance the student network’s ability to identify high-level semantic irregularities, with a visual-auditory consistency modeling expert serving as a guide. Furthermore, a novel Real-Similarity loss is proposed to enhance the proximity of real faces in feature space without explicitly penalizing the distance from manipulated faces, which prevents the overfitting in particular manipulation methods and improves the generalization capability. Extensive experiments show that our method substantially improves the generalization and robustness performance. Particularly, our approach outperforms the SOTA detector by 1.4% in generalization performance on DFDC with large domain gaps, and by 2.0% in the robustness evaluation on the FF++ dataset under various extreme settings. Our code is available athttps://github.com/TianXie834/M2SD. Lingyun Yu 0002, Chuanbin Liu 0001, Guoqing Jin, Zhiguo Ding 0006, Hongtao Xie 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Denoised and Dynamic Alignment Enhancement for Zero-Shot LearningabstractZero-shot learning (ZSL) focuses on recognizing unseen categories by aligning visual features with semantic information. Recent advancements have shown that aligning each attribute with its corresponding visual region significantly improves zero-shot learning performance. However, the crude semantic proxies used in these methods fail to capture the varied appearances of each attribute, and are also easily confused by the presence of semantically redundant backgrounds, leading to suboptimal alignment. To combat these issues, we introduce a novel Alignment-Enhanced Network (AENet), designed to denoise the visual features and dynamically perceive semantic information, thus enhancing visual-semantic alignment. Our approach comprises two key innovations. (1) A visual denoising encoder, employing a class-agnostic mask to filter out semantically redundant visual information, thus producing refined visual features adaptable to unseen classes. (2) A dynamic semantic generator that crafts content-aware semantic proxies adaptively, steered by visual features, enabling AENet to discriminate fine-grained variations in visual contents. Additionally, we integrate a cross-fusion module to ensure comprehensive interaction between the denoised visual features and the generated dynamic semantic proxies, further facilitating visual-semantic alignment. Through extensive experiments across three datasets, the proposed method demonstrates that it narrows down the visual-semantic gap and sets a new benchmark in this setting. Jiannan Ge, Pandeng Li, Lingxi Xie, Yongdong Zhang 0001, Qi Tian 0001, Hongtao Xie 0001 |
IEEE Trans. Image Process. | 7 |
| 2025 | Exploiting Pre-Trained Language Models for Black-Box Attack against Knowledge Graph EmbeddingsabstractDespite the emerging research on adversarial attacks against knowledge graph embedding (KGE) models, most of them focus on white-box attack settings. However, white-box attacks are difficult to apply in practice compared to black-box attacks since they require access to model parameters that are unlikely to be provided. In this article, we propose a novel black-box attack method that only requires access to knowledge graph data, making it more realistic in real-world attack scenarios. Specifically, we utilize pre-trained language models (PLMs) to encode text features of the knowledge graphs, an aspect neglected by previous research. We then employ these encoded text features to identify the most influential triples for constructing corrupted triples for the attack. To improve the transferability of the attack, we further propose to fine-tune the PLM model by enriching triple embeddings with structure information. Extensive experiments conducted on two knowledge graph datasets illustrate the effectiveness of our proposed method. Guangqian Yang, Lei Zhang 0119, Yi Liu 0148, Hongtao Xie 0001, Zhendong Mao 0001 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2025 | High Fidelity Face Swapping via Facial Texture and Structure Consistency Mining
Lingyun Yu 0002, Quanwei Yang, Meng Shao, Hongtao Xie 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Masked Text Pre-Training for Scene Text Detection
Hongtao Xie 0001, Keran Wang, Bangbang Zhou, Yuxin Wang 0002, Weigang Qi, Yadong Qu, Zuan Gao, Dongming Zhang 0004 |
IEEE Trans. Multim. | 1 |
| 2025 | Leveraging Concise Concepts With Probabilistic Modeling for Interpretable Visual RecognitionabstractInterpretable visual recognition is essential for decision-making in high-stakes situations. Recent advancements have automated the construction of interpretable models by leveraging Visual Language Models (VLMs) and Large Language Models (LLMs) with Concept Bottleneck Models (CBMs), which process a bottleneck layer associated with human-understandable concepts. However, existing methods suffer from two main problems: a) the collected concepts from LLMs could be redundant with task-irrelevant descriptions, resulting in an inferior concept space with potential mismatch. b) VLMs directly map the global deterministic image embeddings with fine-grained concepts results in an ambiguous process with imprecise mapping results. To address the above two issues, we propose a novel solution for CBMs with Concise Concept and Probabilistic Modeling (CCPM) that can achieve superior classification performance via high-quality concepts and precise mapping strategy. Fisrt, we leverage in-context examples as category-related clues to guide LLM concept generation process. To mitigate redundancy in the concept space, we propose a Relation-Aware Selection (RAS) module to obtain a concise concept set that is discriminative and relevant based on image-concept and inter-concept relationships. Second, for precise mapping, we employ a Probabilistic Distribution Adapter (PDA) that estimates the inherent ambiguity of the image embeddings of pre-trained VLMs to capture the complex relationships with concepts. Extensive experiments indicate that our model achieves state-of-the-art results with a 5.48% improvement in classification accuracy on eight mainstream recognition benchmarks as well as reliable explainability through interpretable analysis. Chuanbin Liu 0001, Yifan Gao 0011, Zhiying Lu, Hongtao Xie 0001, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment RetrievalabstractVideo Moment Retrieval (VMR) aims to retrieve temporal segments in untrimmed videos corresponding to a given language query by constructing cross-modal alignment strategies. However, these existing strategies are often sub-optimal since they ignore the modality imbalance problem, i.e., the semantic richness inherent in videos far exceeds that of a given limited-length sentence. Therefore, in pursuit of better alignment, a natural idea is enhancing the video modality to filter out query-irrelevant semantics, and enhancing the text modality to capture more segment-relevant knowledge. In this paper, we introduce Modal-Enhanced Semantic Modeling (MESM), a novel framework for more balanced alignment through enhancing features at two levels. First, we enhance the video modality at the frame-word level through word reconstruction. This strategy emphasizes the portions associated with query words in frame-level features while suppressing irrelevant parts. Therefore, the enhanced video contains less redundant semantics and is more balanced with the textual modality. Second, we enhance the textual modality at the segment-sentence level by learning complementary knowledge from context sentences and ground-truth segments. With the knowledge added to the query, the textual modality thus maintains more meaningful semantics and is more balanced with the video modality. By implementing two levels of MESM, the semantic information from both modalities is more balanced to align, thereby bridging the modality gap. Experiments on three widely used benchmarks, including the out-of-distribution settings, show that the proposed framework achieves a new start-of-the-art performance with notable generalization ability (e.g., 4.42% and 7.69% average gains of [email protected] on Charades-STA and Charades-CG). The code will be available at https://github.com/lntzm/MESM. Hongtao Xie 0001, Pandeng Li, Jiannan Ge, Sun'ao Liu, Guoqing Jin |
AAAI | 3 |
| 2024 | DEADiff: An Efficient Stylization Diffusion Model with Disentangled RepresentationsabstractThe diffusion-based text-to-image model harbors im-mense potential in transferring reference style. However, current encoder-based approaches significantly impair the text controllability of text-to-image models while transfer-ring styles. In this paper, we introduce DEADiff to address this issue using the following two strategies: 1) a mecha-nism to decouple the style and semantics of reference images. The decoupled feature representations are first extracted by Q-Formers which are instructed by different text descriptions. Then they are injected into mutually exclusive subsets of cross-attention layers for better disentanglement. 2) A non-reconstructive learning method. The Q-Formers are trained using paired images rather than the identical target, in which the reference image and the ground-truth image are with the same style or semantics. We show that DEADiff attains the best visual stylization results and optimal balance between the text controllability inherent in the text-to-image model and style similarity to the reference image, as demonstrated both quantitatively and qualitatively. Our project page is https://tianhao-qi.github.io/DEADiff‘/. Tianhao Qi, Shancheng Fang, Yanze Wu, Hongtao Xie 0001, Jiawei Liu 0001, Lang Chen, Yongdong Zhang 0001 |
CVPR | 4 |
| 2024 | DiffAM: Diffusion-Based Adversarial Makeup Transfer for Facial Privacy ProtectionabstractWith the rapid development of face recognition (FR) systems, the privacy of face images on social media is facing severe challenges due to the abuse of unauthorized FR systems. Some studies utilize adversarial attack techniques to defend against malicious FR systems by generating adversarial examples. However, the generated adversarial examples, i.e., the protected face images, tend to suffer from sub-par visual quality and low transferability. In this paper, we propose a novel face protection approach, dubbed DiffAM, which leverages the powerful generative ability of diffusion models to generate high-quality protected face images with adversarial makeup transferred from reference images. To be specific, we first introduce a makeup removal module to generate non-makeup images utilizing a fine-tuned diffusion model with guidance of textual prompts in CLIP space. As the inverse process of makeup transfer, makeup removal can make it easier to establish the deterministic relationship between makeup domain and non-makeup domain regardless of elaborate text prompts. Then, with this relationship, a CLIP-based makeup loss along with an ensemble attack strategy is introduced to jointly guide the direction of adversarial makeup domain, achieving the generation of protected face images with natural-looking makeup and high black-box transferability. Extensive experiments demonstrate that DiffAM achieves higher visual quality and attack success rates with a gain of 12.98% under black-box setting compared with the state of the arts. The code will be available at https://github.com/HansSunYIDiffAM. Lingyun Yu 0002, Hongtao Xie 0001, Jiaming Li 0017, Yongdong Zhang 0001 |
CVPR | 3 |
| 2024 | OTE: Exploring Accurate Scene Text Recognition Using One TokenabstractIn this paper, we propose a novel framework to fully exploit the potential of a single vector for scene text recognition (STR). Different from previous sequence-to-sequence methods that rely on a sequence of visual tokens to rep-resent scene text images, we prove that just one token is enough to characterize the entire text image and achieve ac-curate text recognition. Based on this insight, we introduce a new paradigm for STR, called One Token rEcognizer (OTE). Specifically, we implement an image-to-vector en-coder to extract the fine-grained global semantics, elimi-nating the need for sequential features. Furthermore, an elegant yet potent vector-to-sequence decoder is designed to adaptively diffuse global semantics to corresponding character locations, enabling both autoregressive and non-autoregressive decoding schemes. by executing decoding within a high-level representational space, our vector-to-sequence (V2S) approach avoids the alignment issues between visual tokens and character embeddings prevalent in traditional sequence-to-sequence methods. Remarkably, due to introducing character-wise fine-grained information, such global tokens also boost the performance of scene text retrieval tasks. Extensive experiments on synthetic and real datasets demonstrate the effectiveness of our method by achieving new state-of-the-art results on various public STR benchmarks. Our code is available at h t t$P$https://github.com/Xu-Jianjun/OTE. Yuxin Wang 0002, Hongtao Xie 0001, Yongdong Zhang 0001 |
CVPR | 3 |
| 2024 | Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and EditingabstractScene text images contain not only style information (font, background) but also content information (character, texture). Different scene text tasks need different information, but previous representation learning methods use tightly coupled features for all tasks, resulting in sub-optimal performance. We propose a Disentangled Representation Learning framework (DARLING) aimed at disentangling these two types of features for improved adaptability in better addressing various downstream tasks (choose what you really need). Specifically, we synthesize a dataset of image pairs with identical style but different content. Based on the dataset, we decouple the two types of features by the supervision design. Clearly, we directly split the visual representation into style and content features, the content features are supervised by a text recognition loss, while an alignment loss aligns the style features in the image pairs. Then, style features are employed in reconstructing the counterpart image via an image decoder with a prompt that indicates the counterpart's content. Such an operation effectively decouples the features based on their distinctive properties. To the best of our knowledge, this is the first time in the field of scene text that disentangles the inherent properties of the text images. Our method achieves state-of-the-art performance in Scene Text Recognition, Removal, and Editing. Boqiang Zhang, Hongtao Xie 0001, Zuan Gao |
CVPR | 2 |
| 2024 | AlignZeg: Mitigating Objective Misalignment for Zero-Shot Semantic Segmentation
Jiannan Ge, Lingxi Xie, Hongtao Xie 0001, Pandeng Li, Xiaopeng Zhang 0008, Yongdong Zhang 0001, Qi Tian 0001 |
ECCV (43) | 3 |
| 2024 | Leveraging Text Localization for Scene Text Removal via Text-Aware Masked Image Modeling
Zixiao Wang 0002, Hongtao Xie 0001, Yuxin Wang 0002, Yadong Qu, Fengjun Guo, Pengwei Liu |
ECCV (66) | 2 |
| 2024 | Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition
Zuan Gao, Yuxin Wang 0002, Yadong Qu, Boqiang Zhang, Zixiao Wang 0002, Hongtao Xie 0001 |
IJCAI | 7 |
| 2024 | Focus on the Whole Character: Discriminative Character Modeling for Scene Text Recognition
Bangbang Zhou, Yadong Qu, Zixiao Wang 0002, Boqiang Zhang, Hongtao Xie 0001 |
IJCAI | 6 |
| 2024 | LATextSpotter: Empowering Transformer Decoder with Length Perception AbilityabstractScene text spotting aims to integrate scene text detection and recognition into a unified framework. The existing transformer-based methods lack fine-grained positional information and linguistic information, limiting the convergence and performance of the model. In this paper, we propose a Length-Awear Text Spotter (LATextSpotter) to alleviate this problem by explicitly introducing two types of prior knowledge. First, the location of each character is initialized by coarsely locating the text instance and predicting the length, which provides effective guidance for the subsequent position-sensitive decoder. It is worth noting that the model requires only word-level supervision to achieve decent performance in the absence of expensive character-level annotations. Second, we design a mask prediction strategy based on the length information that masks character information at the feature level, and guides the model to predict the missing part. It empowers the decoder with language modeling capability without introducing extra modules. Additionally, considering the coordination between each module, a multi-stage training strategy is proposed to optimize the convergence process. Quantitative experiments demonstrate that LATextSpotter achieves the optimal end-to-end performance on arbitrary-shaped benchmarks by 76.6% and competitive spotting performance on multi-oriented datasets. Yadong Qu, Hongtao Xie 0001, Yongdong Zhang 0001 |
ISCAS | 3 |
| 2024 | Control-Talker: A Rapid-Customization Talking Head Generation Method for Multi-Condition Control and High-Texture EnhancementabstractIn recent years, the field of talking head generation has made significant strides. However, the need for substantial computational resources for model training, coupled with a scarcity of high-quality video data, poses challenges for the rapid customization of model to specific individual. Additionally, existing models usually only support single-modal control, lacking the ability to generate vivid facial expressions and controllable head poses based on multiple conditions such as audio, video, etc. These limitations restricts the models' widespread application. In this paper, we introduce a two-stage method called Control-Talker to achieve rapid customization of identity in talking head model and high-quality generation based on multimodal conditions. Specifically, we divide the training process into two stages: prior learning stage and identity rapid-customization stage. 1) In the prior learning stage, we leverage a diffusion-based model pre-trained on the high-quality image dataset to acquire a robust controllable facial prior. Meanwhile, we innovatively propose a high-frequency ControlNet structure to enhance the fidelity of the synthesized results. This structure adeptly extracts a high-frequency feature map from the source image, serving as a facial texture prior, thereby excellently preserving facial texture of the source image. 2) In the identity rapid-customization stage, the identity is fixed by fine-tuning the U-Net part of the diffusion model on merely several images of a specific individual. The entire fine-tuning process for identity customization can be completed within approximately ten minutes, thereby significantly reducing training costs. Further, we propose a unified driving method for both audio and video, enabling the model to precisely control expressions, poses, and lighting under multi conditions. Extensive experiments and visual results demonstrate that our method outperforms other state-of-the-art models. Additionally, our model demonstrates reduced training costs and lower data requirements. Yiding Li, Lingyun Yu 0002, Li Wang 0154, Hongtao Xie 0001 |
ACM Multimedia | 4 |
| 2024 | Boosting Semi-Supervised Scene Text Recognition via Viewing and SummarizingabstractExisting scene text recognition (STR) methods struggle to recognize challenging texts, especially for artistic and severely distorted characters. The limitation lies in the insufficient exploration of character morphologies, including the monotonousness of widely used synthetic training data and the sensitivity of the model to character morphologies. To address these issues, inspired by the human learning process of viewing and summarizing, we facilitate the contrastive learning-based STR framework in a self-motivated manner by leveraging synthetic and real unlabeled data without any human cost. In the viewing process, to compensate for the simplicity of synthetic data and enrich character morphology diversity, we propose an Online Generation Strategy to generate background-free samples with diverse character styles. By excluding background noise distractions, the model is encouraged to focus on character morphology and generalize the ability to recognize complex samples when trained with only simple synthetic data. To boost the summarizing process, we theoretically demonstrate the derivation error in the previous character contrastive loss, which mistakenly causes the sparsity in the intra-class distribution and exacerbates ambiguity on challenging samples. Therefore, a new Character Unidirectional Alignment Loss is proposed to correct this error and unify the representation of the same characters in all samples by aligning the character features in the student model with the reference features in the teacher model. Extensive experiment results show that our method achieves SOTA performance (94.7\% and 70.9\% average accuracy on common benchmarks and Union14M-Benchmark). Code will be available. Yadong Qu, Bangbang Zhou, Hongtao Xie 0001, Yongdong Zhang 0001 |
NeurIPS | 5 |
| 2024 | ShowMaker: Creating High-Fidelity 2D Human Video via Fine-Grained Diffusion ModelingabstractAlthough significant progress has been made in human video generation, most previous studies focus on either human facial animation or full-body animation, which cannot be directly applied to produce realistic conversational human videos with frequent hand gestures and various facial movements simultaneously.
To address these limitations, we propose a 2D human video generation framework, named ShowMaker, capable of generating high-fidelity half-body conversational videos via fine-grained diffusion modeling.
We leverage dual-stream diffusion models as the backbone of our framework and carefully design two novel components for crucial local regions (i.e., hands and face) that can be easily integrated into our backbone.
Specifically, to handle the challenging hand generation caused by sparse motion guidance, we propose a novel Key Point-based Fine-grained Hand Modeling module by amplifying positional information from raw hand key points and constructing a corresponding key point-based codebook.
Moreover, to restore richer facial details in generated results, we introduce a Face Recapture module, which extracts facial texture features and global identity features from the aligned human face and integrates them into the diffusion process for face enhancement.
Extensive quantitative and qualitative experiments demonstrate the superior visual quality and temporal consistency of our method. Quanwei Yang, Jiazhi Guan, Kaisiyuan Wang, Lingyun Yu 0002, Wenqing Chu, Hang Zhou 0009, ZhiQiang Feng, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Hongtao Xie 0001 |
NeurIPS | 11 |
| 2024 | How Control Information Influences Multilingual Text Image Generation and Editing?abstractVisual text generation has significantly advanced through diffusion models aimed at producing images with readable and realistic text. Recent works primarily use a ControlNet-based framework, employing standard font text images to control diffusion models. Recognizing the critical role of control information in generating high-quality text, we investigate its influence from three perspectives: input encoding, role at different stages, and output features. Our findings reveal that: 1) Input control information has unique characteristics compared to conventional inputs like Canny edges and depth maps. 2) Control information plays distinct roles at different stages of the denoising process. 3) Output control features significantly differ from the base and skip features of the U-Net decoder in the frequency domain. Based on these insights, we propose TextGen, a novel framework designed to enhance generation quality by optimizing control information. We improve input and output features using Fourier analysis to emphasize relevant information and reduce noise. Additionally, we employ a two-stage generation framework to align the different roles of control information at different stages. Furthermore, we introduce an effective and lightweight dataset for training. Our method achieves state-of-the-art performance in both Chinese and English text generation. The code and dataset are available at https://github.com/CyrilSterling/TextGen. Boqiang Zhang, Zuan Gao, Yadong Qu, Hongtao Xie 0001 |
NeurIPS | 4 |
| 2024 | TALK-Act: Enhance Textural-Awareness for 2D Speaking Avatar Reenactment with Diffusion Model
Jiazhi Guan, Quanwei Yang, Kaisiyuan Wang, Hang Zhou 0009, Shengyi He, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Hongtao Xie 0001, Youjian Zhao, Ziwei Liu 0002 |
SIGGRAPH Asia | 10 |
| 2024 | Symmetrical Siamese Network for pose-guided person synthesis
Quanwei Yang, Lingyun Yu 0002, Yun Song, Meng Shao, Guoqing Jin, Hongtao Xie 0001 |
Comput. Vis. Image Underst. | 7 |
| 2024 | CDistNet: Perceiving Multi-domain Character Distance for Robust Text Recognition
Tianlun Zheng, Zhineng Chen, Shancheng Fang, Hongtao Xie 0001, Yu-Gang Jiang 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | Generalizable Speech Spoofing Detection Against Silence Trimming With Data Augmentation and Multi-Task Meta-LearningabstractA major difficulty in speech spoofing detection lies in improving the generalization ability to detect unknown forgery methods. However, most previous methods do not consider the interference of silence information on the generalization performance of speech spoofing detection. Notably, we experimentally observe that the generalization performance of existing methods drops sharply when silence segments are trimmed. This indicates that previous works have two problems: a) they do not remove the interference of silence and over-rely on silence information, and b) they lack the ability to uncover general forgery traces in utterance segments. To solve the above two problems, we propose a novel Silence-Agnostic Speech Spoofing Detection (SASSD) framework. To be specific, unlike previous methods trained on speech samples with silence information, we completely remove the leading and trailing silence segments from all speech samples to eliminate the interference of silence and focus on utterance information. Meanwhile, to uncover general forgery traces in utterance segments and improve the generalization ability, we view speech spoofing detection as a domain generalization problem and employ meta-learning to simulate the actual domain shift scenarios, which can reduce overfitting to specific forgery methods. In addition, to improve the domain generalization of meta-learning, a novel data augmentation method named ShuffleMix is proposed. Unlike previous methods that only consider inter-speech patterns, our method additionally introduces an intra-speech augmentation technique, which performs enhancements within a single speech and across multiple speech to generate more diverse forged samples. Extensive experiments show that our method achieves SOTA on the ASVspoof 2019LA dataset. In particular, our method achieves 0.231% EER and 2.529% EER on the original dataset with silence information and the silence-trimmed dataset, respectively. Li Wang 0154, Lingyun Yu 0002, Yongdong Zhang 0001, Hongtao Xie 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | STIDNet: Identity-Aware Face Forgery Detection With Spatiotemporal Knowledge DistillationabstractThe impressive development of facial manipulation techniques has raised severe public concerns. Identity-aware methods, especially suitable for protecting celebrities, are seen as one of promising face forgery detection approaches with additional reference video. However, without in-depth observation of fake video’s characteristics, most existing identity-aware algorithms are just naive imitation of face verification model and fail to exploit discriminative information. In this article, we argue that it is necessary to take both spatial and temporal perspectives into consideration for adequate inconsistency clues and propose a novel forgery detector named SpatioTemporal IDentity network (STIDNet). To effectively capture heterogeneous spatiotemporal information in a unified formulation, our STIDNet is following a knowledge distillation architecture that the student identity extractor receives supervision from a spatial information encoder (SIE) and a temporal information encoder (TIE) through multiteacher training. Specifically, a regional sensitive identity modelling paradigm is proposed in SIE by introducing facial blending augmentation but with uniform identity label, thus encourage model to focus on spatial discriminative region like outer face. Meanwhile, considering the strong temporal correlation between audio and talking face video, our TIE is devised in a cross-modal pattern that the audio information is introduced to supervise model exploiting temporal personalized movements. Benefit from knowledge transfer from SIE and TIE, STIDNet is able to capture individual’s essential spatiotemporal identity attributes and sensitive to even subtle identity deviation caused by manipulation. Extensive experiments indicate the superiority of our STIDNet compared with previous works. Moreover, we also demonstrate STIDNet is more suitable for real-world implementation in terms of model complexity and reference set size. Mingqi Fang, Lingyun Yu 0002, Hongtao Xie 0001, Qingfeng Tan, Zhiyuan Tan 0001, Amir Hussain 0001, Zezheng Wang 0002, Zhihong Tian 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | DCFP: Distribution Calibrated Filter Pruning for Lightweight and Accurate Long-Tail Semantic SegmentationabstractRecently, semantic segmentation has made promising progress, but the high cost of processing still limits its application. With focusing on removing the parameters of the networks, filter pruning using the importance criterion is a straightforward and effective technique to obtain the lightweight sub-network. However, we argue that the long-tail distribution in segmentation datasets poses two significant problems which are ignored in existing pruning algorithms: 1) The importance criterion is dominated by head classes which contain numerous positive samples, where the knowledge of tail classes is easily degenerated. 2) The degenerated knowledge of tail classes is hard to recover as their samples are also insufficient during fine-tuning. To address these issues, we propose a Distribution Calibrated Filter Pruning (DCFP) framework for segmentation. Firstly, a gradient-based Equalization Importance Criterion (EIC) is designed to generate a class-balanced pruning procedure. It avoids the bias on head classes by discarding the imbalanced positive gradients. Secondly, we introduce a Geometric-Semantic Re-balanced Loss (GSRL) to emphasize the learning on tail classes during fine-tuning. The GSRL consists of two cooperative components to calibrate the imbalanced optimization on geometric and semantic domains dynamically. Compared with previous methods, DCFP explores a novel distribution-aware pruning framework to obtain lightweight architectures with accurate results. Extensive experiments proved that DCFP achieves impressive performance on four popular segmentation benchmarks. Zixiao Wang 0002, Hongtao Xie 0001, Yuxin Wang 0002, Guoqing Jin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Cascade Semantic Prompt Alignment Network for Image CaptioningabstractImage captioning (IC) takes an image as input and generates open-form descriptions in the domain of natural language. IC requires the detection of objects, modeling of relations between them, an assessment of the semantics of the scene and representing the extracted knowledge in a language space. Previous detector-based models suffer from limited semantic perception capability due to predefined object detection classes and semantic inconsistency between visual region features and numeric labels of the detector. Inspired by the fact that text prompts in pre-trained multi-modal models contain specific linguistic knowledge rather than discrete labels, and excel at an open-form semantic understanding of visual inputs and their representation in the domain of natural language. We aim to distill and leverage the transferable language knowledge from the pre-trained RegionCLIP model to remedy the detector for generating rich image captioning. In this paper, we propose a novel Cascade Semantic Prompt Alignment Network (CSA-Net) to produce an aligned fine-grained regional semantic-visual space where rich and consistent textual semantic details are automatically incorporated to region features. Specifically, we first align the object semantic prompt and region features to produce semantic grounded object features. Then, we employ these object features and relation semantic prompt to predict the relations between objects. Finally, these enhanced object and relation features are fed into the language decoder, generating rich descriptions. Extensive experiments conducted on the MSCOCO dataset show that our method achieves a new state-of-the-art performance with 145.2% (single model) and 147.0% (ensemble of 4 models) CIDEr scores on the ‘Karpathy’ split, 141.6% (c5) and 144.1% (c40) CIDEr scores on the official online test server. Significantly, CSA-Net outperforms in generating captions with higher quality and diversity, achieving a RefCLIP-S score of 83.2. Moreover, we expand the testbeds to other challenging captioning benchmarks, i.e., nocaps datasets, CSA-Net demonstrates superior zero-shot capability. Source codes released at https://github.com/CrossmodalGroup/CSA-Net. Lei Zhang 0119, Kun Zhang 0040, Bo Hu 0036, Hongtao Xie 0001, Zhendong Mao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Exploring Bi-Level Inconsistency via Blended Images for Generalizable Face Forgery DetectionabstractThe challenge of generalization in face forgery detection has become increasingly prominent as manipulation techniques continue to evolve. Although recent image blending-based methods have demonstrated remarkable potential, they often encounter a significant performance drop when applied to datasets exhibiting significant domain gaps. This limitation stems from the exclusive reliance of prior methods on blending unaltered faces with various augmentations to produce common artifacts, which ignores the inherent characteristics of the forged regions. To fully exploit the potential of image blending-based methods for generalizable Deepfake detection, we propose a novel image synthesis framework called Bi-Level Inconsistency Generator (Bi-LIG) to introduce bi-level inconsistency in the synthesized images. Specifically, Bi-LIG generates synthetic images by blending source and target images from both pristine and forged image sets, introducing a) Extrinsic-Inconsistency between real and pseudo-forged regions, and b) Inherent-Inconsistency between real and manipulated areas. In this way, Bi-LIG creates a diverse synthesized image set and establishes a generalizable training domain. Furthermore, we propose a novel face forgery detection network named Token Consistency Constrained Vision Transformer, in which two modules are developed based on patch consistency learning. Firstly, a Patch Token Contrast module is employed to learn the bi-level patch inconsistencies. Secondly, a Progressive Patch Token Assemble module is adopted to aggregate local patch relations and enhance the inconsistency representations. Experimental results demonstrate the effectiveness and superiority of our method on both in-dataset and cross-dataset evaluations. Notably, our approach outperforms state-of-the-art methods by 5.09% and 10.15% on cross-dataset evaluations in DFDCp and DFDC, respectively. Peiqi Jiang, Hongtao Xie 0001, Lingyun Yu 0002, Guoqing Jin, Yongdong Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | IEIRNet: Inconsistency Exploiting Based Identity Rectification for Face Forgery DetectionabstractFace forgery detection has attracted much attention due to the ever-increasing social concerns caused by facial manipulation techniques. Recently, identity-based detection methods have made considerable progress, which is especially suitable in the celebrity protection scenario. However, they still suffer from two main limitations: (a) generic identity extractor is not specifically designed for forgery detection, leading to nonnegligibleIdentity Representation Biasto forged images. (b) existing methods only analyze the identity representation of each image individually, but ignores the query-reference interaction for inconsistency exploiting. To address these issues, a novelInconsistency Exploiting based Identity Rectification Network(IEIRNet) is proposed in this paper. Firstly, for the identity bias rectification, the IEIRNet follows an effective two-branches structure. Besides theGeneric Identity Extractor(GIE) branch, an essentialBias Diminishing Module(BDM) branch is proposed to eliminate the identity bias through a novelAttention-based Bias Rectification(ABR) component, accordingly acquiring the ultimate discriminative identity representation. Secondly, for query-reference inconsistency exploiting, anInconsistency Exploiting Module(IEM) is applied in IEIRNet to comprehensively exploit the inconsistency clues from both spatial and channel perspectives. In the spatial aspect, an innovative region-aware kernel is derived to activate the local region inconsistency with deep spatial interaction. Afterward in the channel aspect, a coattention mechanism is utilized to model the channel interaction meticulously, and accordingly highlight the channel-wise inconsistency with adaptive weight assignment and channel-wise dropout. Our IEIRNet has shown effectiveness and superiority in various generalization and robustness experiments. Mingqi Fang, Lingyun Yu 0002, Yun Song, Yongdong Zhang 0001, Hongtao Xie 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Towards Discriminative Feature Generation for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) aims to recognize both seen and unseen categories by establishing visual and semantic relations. Recently, generation-based methods that focus on synthesizing fictitious visual features from corresponding attributes have gained significant attention. However, these generated features often lack discriminative capabilities due to inadequate training of the generative model. To address this issue, we propose a novel Discriminative Enhanced Network (DENet) to harness the potential of the generative model by adapting the training features and imposing constraints on the generated features. Our approach incorporates three pivotal modules: (1) Before the generative network training, we implement a Pre-Tuning Module (PTM) to eliminate irrelevant background noise in the raw features extracted from a fixed CNN backbone. Therefore, PTM can provide tuned training features without redundant noise for generative model. (2) During the generative network training, we propose an Asymmetry Cross-authenticity Contrastive (AC2) loss to group visual features of the same category while repel features from different categories by optimizing a large number of sample pairs. Additionally, we incorporate intra-class and relation-specific inter-class boundaries within the AC2 loss to enrich sample diversity and preserve valid semantic information. (3) Also within the generative network training, a Dual-semantic Alignment Module (DAM) is designed to align visual features with both attributes and label embeddings, enabling the model to learn attribute-related information and discriminative extended semantics. Experiments on four standard benchmarks demonstrate that our approach learns more discriminative features and surpasses the existing methods. Jiannan Ge, Hongtao Xie 0001, Pandeng Li, Lingxi Xie, Shaobo Min, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Balanced Classification: A Unified Framework for Long-Tailed Object DetectionabstractConventional detectors suffer from performance degradation when dealing with long-tailed data due to a classification bias towards the majority head categories. In this article, we contend that the learning bias originates from two factors: 1) the unequal competition arising from the imbalanced distribution of foreground categories, and 2) the lack of sample diversity in tail categories. To tackle these issues, we introduce a unified framework calledBAlancedCLassification (BACL), which enables adaptive rectification of inequalities caused by disparities in category distribution and dynamic intensification of sample diversities in a synchronized manner. Specifically, a novel foreground classification balance loss (FCBL) is developed to ameliorate the domination of head categories and shift attention to difficult-to-differentiate categories by introducing pairwise class-aware margins and auto-adjusted weight terms, respectively. This loss prevents the over-suppression of tail categories in the context of unequal competition. Moreover, we propose a dynamic feature hallucination module (FHM), which enhances the representation of tail categories in the feature space by synthesizing hallucinated samples to introduce additional data variances. In this divide-and-conquer approach, BACL sets a new state-of-the-art on the challenging LVIS benchmark with a decoupled training pipeline, surpassing vanilla Faster R-CNN with ResNet-50-FPN by 5.8% AP and 16.1% AP for overall and tail categories. Extensive experiments demonstrate that BACL consistently achieves performance improvements across various datasets with different backbones and architectures. Tianhao Qi, Hongtao Xie 0001, Pandeng Li, Jiannan Ge, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Semantic-Enhanced Proxy-Guided Hashing for Long-Tailed Image RetrievalabstractHashing has been studied extensively for large-scale image retrieval due to its efficient computation and storage. Deep hashing methods typically train models with category-balanced data and suffer from a serious performance deterioration when dealing with long-tailed training samples. Recently, several long-tailed hashing methods focus on this newly emerging field for practical purpose. However, existing methods still face challenges that fixed category centers with limited semantic information cannot effectively improve the discriminative ability of tail-category hash codes. To tackle the issue, we propose a novel method called Semantic-enhanced Proxy-guided Hashing in this paper. We leverage two sets of learnable category proxies in the feature space and the Hamming space respectively, which can describe category semantics by getting updated continuously along with the whole model via back-propagation. Based on this, we introduce the Mahalanobis distance metric to characterize relationships accurately and enhance the semantic representation of both proxies and samples concurrently, improving the hash learning process. Moreover, we capture the multilateral correlations between proxies and samples in the feature space and extend a hypergraph neural network to transfer semantic knowledge from proxies to samples in the Hamming space. Extensive experiments show that our method achieves the state-of-the-art performance and surpasses existing methods by 1.47%–7.56% MAP on long-tailed benchmarks, demonstrating the superiority of learnable category proxies and the effectiveness of our proposed learning algorithm for long-tailed hashing. Hongtao Xie 0001, Lei Zhang 0119, Pandeng Li, Dongming Zhang 0004, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Exploring Stroke-Level Modifications for Scene Text EditingabstractScene text editing (STE) aims to replace text with the desired one while preserving background and styles of the original text. However, due to the complicated background textures and various text styles, existing methods fall short in generating clear and legible edited text images. In this study, we attribute the poor editing performance to two problems: 1) Implicit decoupling structure. Previous methods of editing the whole image have to learn different translation rules of background and text regions simultaneously. 2) Domain gap. Due to the lack of edited real scene text images, the network can only be well trained on synthetic pairs and performs poorly on real-world images. To handle the above problems, we propose a novel network by MOdifying Scene Text image at strokE Level (MOSTEL). Firstly, we generate stroke guidance maps to explicitly indicate regions to be edited. Different from the implicit one by directly modifying all the pixels at image level, such explicit instructions filter out the distractions from background and guide the network to focus on editing rules of text regions. Secondly, we propose a Semi-supervised Hybrid Learning to train the network with both labeled synthetic images and unpaired real scene text images. Thus, the STE model is adapted to real-world datasets distributions. Moreover, two new datasets (Tamper-Syn2k and Tamper-Scene) are proposed to fill the blank of public evaluation datasets. Extensive experiments demonstrate that our MOSTEL outperforms previous methods both qualitatively and quantitatively. Datasets and code will be available at https://github.com/qqqyd/MOSTEL. Yadong Qu, Qingfeng Tan, Hongtao Xie 0001, Yuxin Wang 0002, Yongdong Zhang 0001 |
AAAI | 3 |
| 2023 | Learning Orthogonal Prototypes for Generalized Few-Shot Semantic SegmentationabstractGeneralized few-shot semantic segmentation (GFSS) distinguishes pixels of base and novel classes from the background simultaneously, conditioning on sufficient data of base classes and a few examples from novel class. A typical GFSS approach has two training phases: base class learning and novel class updating. Nevertheless, such a stand-alone updating process often compromises the well-learnt features and results in performance drop on base classes. In this paper, we propose a new idea of leveraging Projection onto Orthogonal Prototypes (POP), which updates features to identify novel classes without compromising base classes. POP builds a set of orthogonal prototypes, each of which represents a semantic class, and makes the prediction for each class separately based on the features projected onto its prototype. Technically, POP first learns prototypes on base data, and then extends the prototype set to novel classes. The orthogonal constraint of POP encourages the orthogonality between the learnt prototypes and thus mitigates the influence on base class features when generalizing to novel prototypes. Moreover, we capitalize on the residual of feature projection as the background representation to dynamically fit semantic shifting (i.e., background no longer includes the pixels of novel classes in updating phase). Extensive experiments on two benchmarks demonstrate that our POP achieves superior performances on novel classes without sacrificing much accuracy on base classes. Notably, POP outperforms the state-of-the-art fine-tuning by 3.93% overall mIoU on PASCAL-5iin 5-shot scenario. Sun'ao Liu, Zhaofan Qiu, Hongtao Xie 0001, Yongdong Zhang 0001, Ting Yao 0003 |
CVPR | 4 |
| 2023 | Progressive Spatio-Temporal Prototype Matching for Text-Video RetrievalabstractThe performance of text-video retrieval has been significantly improved by vision-language cross-modal learning schemes. The typical solution is to directly align the global video-level and sentence-level features during learning, which would ignore the intrinsic video-text relations, i.e., a text description only corresponds to a spatio-temporal part of videos. Hence, the matching process should consider both fine-grained spatial content and various temporal semantic events. To this end, we propose a text-video learning framework with progressive spatio-temporal prototype matching. Specifically, the matching process is decomposed into two complementary phases: object-phrase prototype matching and event-sentence prototype matching. In the object-phrase prototype matching phase, the spatial prototype generation mechanism predicts key patches or words, which are aggregated into object or phrase prototypes. Importantly, optimizing the local alignment between object-phrase prototypes helps the model perceive spatial details. In the event-sentence prototype matching phase, we design a temporal prototype generation mechanism to associate intra-frame objects and interact inter-frame temporal relations. Such progressively generated event prototypes can reveal semantic diversity in videos for dynamic matching. Validated by comprehensive experiments, our method consistently outperforms the state-of-the-art methods on four video retrieval benchmark.1 Pandeng Li, Chen-Wei Xie, Hongtao Xie 0001, Jiannan Ge, Deli Zhao, Yongdong Zhang 0001 |
ICCV | 4 |
| 2023 | Difference-Aware Iterative Reasoning Network for Key Relation DetectionabstractScene graph serves as a crucial visual representation of an image, with salient objects providing richer semantics for detecting key relations. However, most methods use a one-step reasoning manner for key relation detection, which may not utilize potential clues effectively. Humans usually review and revise to achieve the final answer, and semantics of relations offer further linguistic clues. Therefore, we propose the Difference-aware Iterative Reasoning Network (DIRNet) to predict key relations in a multi-step manner. Our model estimates visual saliency, encodes contexts globally with message passing, and then refines predictions iteratively by considering the difference in predicted relation semantics and contextual information across iterations. Extensive experiments show that our model outperforms state-of-the-art methods in key relation prediction on the VG-KR benchmark, and achieves competitive results in common relation prediction on VG, demonstrating its generalization and superiority. Weidong Chen 0013, Bo Hu 0036, Hongtao Xie 0001, Zhendong Mao 0001 |
ICME | 4 |
| 2023 | Linguistic More: Taking a Further Step toward Efficient and Accurate Scene Text RecognitionabstractVision model have gained increasing attention due to their simplicity and efficiency in Scene Text Recognition (STR) task. However, due to lacking the perception of linguistic knowledge and information, recent vision models suffer from two problems: (1) the pure vision-based query results in attention drift, which usually causes poor recognition and is summarized as linguistic insensitive drift (LID) problem in this paper. (2) the visual feature is suboptimal for the recognition in some vision-missing cases (e.g. occlusion, etc.). To address these issues, we propose a Linguistic Perception Vision model (LPV), which explores the linguistic capability of vision model for accurate text recognition. To alleviate the LID problem, we introduce a Cascade Position Attention (CPA) mechanism that obtains high-quality and accurate attention maps through step-wise optimization and linguistic information mining. Furthermore, a Global Linguistic Reconstruction Module (GLRM) is proposed to improve the representation of visual features by perceiving the linguistic information in the visual space, which gradually converts visual features into semantically rich ones during the cascade process. Different from previous methods, our method obtains SOTA results while keeping low complexity (92.4% accuracy with only 8.11M parameters). Code is available at https://github.com/CyrilSterling/LPV. Boqiang Zhang, Hongtao Xie 0001, Yuxin Wang 0002, Yongdong Zhang 0001 |
IJCAI | 2 |
| 2023 | TPS++: Attention-Enhanced Thin-Plate Spline for Scene Text RecognitionabstractText irregularities pose significant challenges to scene text recognizers. Thin-Plate Spline (TPS)-based rectification is widely regarded as an effective means to deal with them. Currently, the calculation of TPS transformation parameters purely depends on the quality of regressed text borders. It ignores the text content and often leads to unsatisfactory rectified results for severely distorted text. In this work, we introduce TPS++, an attention-enhanced TPS transformation that incorporates the attention mechanism to text rectification for the first time. TPS++ formulates the parameter calculation as a joint process of foreground control point regression and content-based attention score estimation, which is computed by a dedicated designed gated-attention block. TPS++ builds a more flexible content-aware rectifier, generating a natural text correction that is easier to read by the subsequent recognizer. Moreover, TPS++ shares the feature backbone with the recognizer in part and implements the rectification at feature-level rather than image-level, incurring only a small overhead in terms of parameters and inference time. Experiments on public benchmarks show that TPS++ consistently improves the recognition and achieves state-of-the-art accuracy. Meanwhile, it generalizes well on different backbones and recognizers. Code is at https://github.com/simplify23/TPS_PP. Tianlun Zheng, Zhineng Chen, Jinfeng Bai, Hongtao Xie 0001, Yu-Gang Jiang 0001 |
IJCAI | 4 |
| 2023 | RAIRNet: Region-Aware Identity Rectification for Face Forgery DetectionabstractThe malicious usage of facial manipulation techniques boosts the desire of face forgery detection research. Recently, identity-based approaches have attracted much attention due to the effective observation of identity inconsistency. However, there are still several nonnegligible problems: (1) generic identity extractor is totally trained on real images, leading to enormous identity representation bias during processing forged content; (2) the identity information of forged image is hybrid and presents regional distribution, while the single global identity feature is hard to reflect this local identity inconsistency. To solve the above problems, in this paper a novel Region-Aware Identity Rectification Network (RAIRNet) is proposed to effectively rectify the identity bias and adaptively exploit the inconsistency local region. Firstly, for the identity bias problem, our RAIRNet is devised in a two-branch architecture, which consists of a Generic Identity Extractor (GIE) branch and a Bias Diminishing Module (BDM) branch. The BDM branch is designed to rectify the bias introduced by GIE branch through a prototype-based training schema. This two-branch architecture effectively promotes model to adapt to forged content while maintaining the focus on identity space. Secondly, for local identity inconsistency exploiting, a novel Meta Identity Filter Generator (MIFG) is devised in a meta-learning way to generate the region-aware filter based on identity prior. This region-aware filter can adaptively exploit the local inconsistency clues and activate the discriminative local region. Moreover, to balance the local-global information and highlight the forensic clues, an Adaptive Weight Assignment Mechanism (AWAM) is proposed to assign adaptive importance weight to two branches. Extensive experiments on various datasets show the superiority of our RAIRNet. In particular, on the challenging DFDCp dataset, our approach outperforms previous binary-based and identity-based methods by 10.3% and 5.5% respectively. Mingqi Fang, Lingyun Yu 0002, Hongtao Xie 0001, Junqiang Wu, Zezheng Wang 0002, Yongdong Zhang 0001 |
ACM Multimedia | 3 |
| 2023 | TextPainter: Multimodal Text Image Generation with Visual-harmony and Text-comprehension for Poster DesignabstractText design is one of the most critical procedures in poster design, as it relies heavily on the creativity and expertise of humans to design text images considering the visual harmony and text-semantic. This study introduces TextPainter, a novel multimodal approach that leverages contextual visual information and corresponding text semantics to generate text images. Specifically, TextPainter takes the global-local background image as a hint of style and guides the text image generation with visual harmony. Furthermore, we leverage the language model and introduce a text comprehension module to achieve both sentence-level and word-level style variations. Besides, we construct the PosterT80K dataset, consisting of about 80K posters annotated with sentence-level bounding boxes and text contents. We hope this dataset will pave the way for further research on multimodal text image generation. Extensive quantitative and qualitative experiments demonstrate that TextPainter can generate visually-and-semantically-harmonious text images for posters. Yifan Gao 0011, Jinpeng Lin, Chuanbin Liu 0001, Hongtao Xie 0001, Tiezheng Ge, Yuning Jiang 0001 |
ACM Multimedia | 5 |
| 2023 | Dual Dynamic Proxy Hashing Network for Long-tailed Image RetrievalabstractDeep hashing has been extensively explored for image retrieval due to fast computation and efficient storage. Since conventional deep hashing methods are not suitable for the common scenario in real life that data exhibits a long-tailed distribution, several long-tailed hashing methods have been proposed recently. However, existing long-tail hashing methods seek to utilize fixed class centroids and cannot fully develop the discriminative ability of hash codes for tail-class samples. Specifically, fixed class centroids cannot characterize authentic semantics of tail classes or provide effective semantic information for hash codes learning under the long-tailed setting. To this end, we propose a novel Dual Dynamic Proxy Hashing Network (DDPHN) with two sets of learnable dynamic proxies, i.e. hash proxies and feature proxies, to improve the discrimination of hash codes for tail-class samples. Compared with fixed class centroids, learnable proxies can be optimized constantly via the proxy learning loss and depict accurate class semantics despite the scarcity of tail-class samples. Apart from low-dimensional binary hash proxies, we introduce high-dimensional continuous feature proxies that can describe semantic relationships more precisely, contributing to hash codes learning as well. To further leverage semantic information carried by proxies, we build a hypergraph by exploring neighborhood relationships in the feature space and then introduce a hypergraph neural network to transfer knowledge from proxies to samples in the Hamming space. Extensive experiments show the superiority of our learnable dynamic proxies and demonstrate that our method outperforms numerous deep hashing models and recent state-of-the-art long-tailed hashing methods. Hongtao Xie 0001, Lei Zhang 0119, Pandeng Li, Dongming Zhang 0004, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2023 | High Fidelity Face Swapping via Semantics Disentanglement and Structure EnhancementabstractIn this paper, we propose a novel Semantics and Structure-aware face Swapping framework (S2Swap) that exploits semantics disentanglement and structure enhancement for high fidelity face generation. Different from previous methods that either 1) suffer from degraded generation fidelity due to insufficient identity-attributes disentanglement or 2) neglect the importance of structure information for identity consistency, our approach can achieve local facial semantics disentanglement beyond global identity while boosting identity consistency through structure enhancement. Specifically, to achieve identity-attributes disentanglement, our S2Swap is designed from global-local perspectives. Firstly, an Oriented Identity Transfer module is proposed to globally disentangle target identity and attributes under global identity semantics prior. Such global disentanglement enables source identity transfer to the individual target identity. Secondly, a Local Semantics Disentanglement module is devised to disentangle local identity and identity-irrelevant facial semantics, providing local semantic compensation for the global counterpart. Moreover, to boost identity consistency, a Structure-Aware Head Modeling module is introduced to provide the desired face structure enhancement through an intuitive face sketch. Finally, considering the identity-attributes trade-off, we adaptively integrate semantics and structure information in a self-learning manner. Extensive experiments qualitatively and quantitatively show that our method outperforms SOTA face swapping methods in terms of both identity transfer and attribute preservation. Lingyun Yu 0002, Hongtao Xie 0001, Chuanbin Liu 0001, Zhiguo Ding 0006, Quanwei Yang, Yongdong Zhang 0001 |
ACM Multimedia | 3 |
| 2023 | CARIS: Context-Aware Referring Image SegmentationabstractReferring image segmentation aims to segment the target object described by a natural-language utterance. Recent approaches typically distinguish pixels by aligning pixel-wise visual features with linguistic features extracted from the referring description. Nevertheless, such a free-form description only specifies certain discriminative attributes of the target object or its relations to a limited number of objects, which fails to represent the rich visual context adequately. The stand-alone linguistic features are therefore unable to align with all visual concepts, resulting in inaccurate segmentation. In this paper, we propose to address this issue by incorporating rich visual context into linguistic features for sufficient vision-language alignment. Specifically, we present Context-Aware Referring Image Segmentation (CARIS), a novel architecture that enhances the contextual awareness of linguistic features via sequential vision-language attention and learnable prompts. Technically, CARIS develops a context-aware mask decoder with sequential bidirectional cross-modal attention to integrate the linguistic features with visual context, which are then aligned with pixel-wise visual features. Furthermore, two groups of learnable prompts are employed to delve into additional contextual information from the input image and facilitate the alignment with non-target pixels, respectively. Extensive experiments demonstrate that CARIS achieves new state-of-the-art performances on three public benchmarks. Code is available at https://github.com/lsa1997/CARIS. Sun'ao Liu, Zhaofan Qiu, Hongtao Xie 0001, Yongdong Zhang 0001, Ting Yao 0003 |
ACM Multimedia | 4 |
| 2023 | Symmetrical Linguistic Feature Distillation with CLIP for Scene Text RecognitionabstractIn this paper, we explore the potential of the Contrastive Language-Image Pretraining (CLIP) model in scene text recognition (STR), and establish a novel Symmetrical Linguistic Feature Distillation framework (named CLIP-OCR) to leverage both visual and linguistic knowledge in CLIP. Different from previous CLIP-based methods mainly considering feature generalization on visual encoding, we propose a symmetrical distillation strategy (SDS) that further captures the linguistic knowledge in the CLIP text encoder. By cascading the CLIP image encoder with the reversed CLIP text encoder, a symmetrical structure is built with an image-to-text feature flow that covers not only visual but also linguistic information for distillation. Benefiting from the natural alignment in CLIP, such guidance flow provides a progressive optimization objective from vision to language, which can supervise the STR feature forwarding process layer-by-layer. Besides, a new Linguistic Consistency Loss (LCL) is proposed to enhance the linguistic capability by considering second-order statistics during the optimization. Overall, CLIP-OCR is the first to design a smooth transition between image and text for the STR task. Extensive experiments demonstrate the effectiveness of CLIP-OCR with 93.8% average accuracy on six popular STR benchmarks. Code will be available at https://github.com/wzx99/CLIPOCR. Zixiao Wang 0002, Hongtao Xie 0001, Yuxin Wang 0002, Boqiang Zhang, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2023 | Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text DetectionabstractScene text detection has made great progress recently with the wide use of pre-training. Nonetheless, existing scene text detection methods still suffer from two problems: 1) Limited annotated real data reduces the feature robustness. 2) Detectors perform poorly on text lacking of visual information. In this paper, we explore the potential of the CLIP model, and propose a novel self-supervised Masked Text Modeling (MTM) pre-training method for scene text detection, which can be trained with unlabeled data and improve the linguistic reasoning ability for text occlusion. Different from previous randomly pixel-level masking methods, MTM performs a targeted text-aware masking process under an unsupervised manner. Specifically, MTM consists of text perception and masked text modeling. In the text perception step, benefiting from the text-friendliness of CLIP, a Text Perception Module is proposed to attend to text area by computing the similarity between the text and image tokens from CLIP model. In the masked text modeling step, a Text-aware Masking Strategy is designed to mask the text area, and the Masked Text Modeling Module is used to reconstruct the masked texts. MTM obtains the ability to reason the linguistic information of masked texts with the reconstruction. This robust feature extraction learned by MTM ensures a more discriminative representation for the text lacking of visual information. Moreover, a new text dataset named OcclusionText is proposed to evaluate the robustness for text occlusion of detection methods. Extensive experiments on public benchmarks demonstrate that our MTM can boost the performance of existing text detectors. Keran Wang, Hongtao Xie 0001, Yuxin Wang 0002, Dongming Zhang 0004, Yadong Qu, Zuan Gao, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2023 | Frequency-based Zero-Shot Learning with Phase AugmentationabstractZero-Shot Learning (ZSL) aims to recognize images from seen and unseen classes by aligning visual and semantic knowledge (e.g., attribute descriptions). However, the fine-grained attributes in the RGB domain can be easily affected by background noise (e.g., the grey bird tail blending with the ground), making it difficult to effectively distinguish them. Analyzing the features in the frequency domain assists in better distinguishing the attributes since their patterns remain consistent across different images, unlike noise which may be more variable. Nevertheless, existing ZSL methods typically learn visual features directly from the RGB domain, which can impede the recognition of certain attributes. To overcome this limitation, we propose a novel ZSL method named Frequency-based Phase Augmentation (FPA) network, which learns an effective representation of the attributes in the frequency domain. Specifically, we introduce a Hybrid Phase Augmentation (HPA) module to transform visual features into the frequency domain and augment the phase component for better retention of semantic information of the attributes. The use of phase-augmented features enables FPA to capture more semantic knowledge that can be challenging to distinguish in the RGB domain, suppress noise, and highlight significant attributes. Our extensive experiments show that FPA achieves state-of-the-art performance across four standard datasets. Wanting Yin, Hongtao Xie 0001, Lei Zhang 0119, Jiannan Ge, Pandeng Li, Chuanbin Liu 0001, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2023 | MomentDiff: Generative Video Moment Retrieval from Random to RealabstractVideo moment retrieval pursues an efficient and generalized solution to identify the specific temporal segments within an untrimmed video that correspond to a given language description.
To achieve this goal, we provide a generative diffusion-based framework called MomentDiff, which simulates a typical human retrieval process from random browsing to gradual localization.
Specifically, we first diffuse the real span to random noise, and learn to denoise the random noise to the original span with the guidance of similarity between text and video.
This allows the model to learn a mapping from arbitrary random locations to real moments, enabling the ability to locate segments from random initialization.
Once trained, MomentDiff could sample random temporal segments as initial guesses and iteratively refine them to generate an accurate temporal boundary.
Different from discriminative works (e.g., based on learnable proposals or queries), MomentDiff with random initialized spans could resist the temporal location biases from datasets.
To evaluate the influence of the temporal location biases, we propose two ``anti-bias'' datasets with location distribution shifts, named Charades-STA-Len and Charades-STA-Mom.
The experimental results demonstrate that our efficient framework consistently outperforms state-of-the-art methods on three public benchmarks, and exhibits better generalization and robustness on the proposed anti-bias datasets.
The code, model, and anti-bias evaluation datasets will be released publicly. Pandeng Li, Chen-Wei Xie, Hongtao Xie 0001, Lei Zhang 0119, Deli Zhao, Yongdong Zhang 0001 |
NeurIPS | 3 |
| 2023 | ABINet++: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text SpottingabstractScene text spotting is of great importance to the computer vision community due to its wide variety of applications. Recent methods attempt to introduce linguistic knowledge for challenging recognition rather than pure visual classification. However, how to effectively model the linguistic rules in end-to-end deep networks remains a research challenge. In this paper, we argue that the limited capacity of language models comes from 1) implicit language modeling; 2) unidirectional feature representation; and 3) language model with noise input. Correspondingly, we propose an autonomous, bidirectional and iterative ABINet++ for scene text spotting. First, the autonomous suggests enforcing explicitly language modeling by decoupling the recognizer into vision model and language model and blocking gradient flow between both models. Second, a novel bidirectional cloze network (BCN) as the language model is proposed based on bidirectional feature representation. Third, we propose an execution manner of iterative correction for the language model which can effectively alleviate the impact of noise input. Additionally, based on an ensemble of the iterative predictions, a self-training method is developed which can learn from unlabeled images effectively. Finally, to polish ABINet++ in long text recognition, we propose to aggregate horizontal features by embedding Transformer units inside a U-Net, and design a position and content attention module which integrates character order and content to attend to character features precisely. ABINet++ achieves state-of-the-art performance on both scene text recognition and scene text spotting benchmarks, which consistently demonstrates the superiority of our method in various environments especially on low-quality images. Besides, extensive experiments including in English and Chinese also prove that, a text spotter that incorporates our language modeling method can significantly improve its performance both in accuracy and speed compared with commonly used attention-based recognizers. Code is available at https://github.com/FangShancheng/ABINet-PP. Shancheng Fang, Zhendong Mao 0001, Hongtao Xie 0001, Yuxin Wang 0002, Chenggang Yan 0001, Yongdong Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Neighborhood-Adaptive Multi-Cluster Ranking for Deep Metric LearningabstractDeep metric learning methods generally concentrate on designing distance-based losses to learn sample embeddings, which tacitly presuppose the neighborhood structure around each sample (e.g., hypersphere for Euclidean distance). However, this supposition is overly optimistic: 1) visual data is often located on low-dimensional manifolds curved in high-dimensional space, and all regions of the manifold may hardly share the same local structures in the input space; 2) it is unlikely that the local structure in the output embedding space is as homogeneous as assumed due to the non-linearity of neural networks. Hence, simply characterizing sample embeddings while ignoring the respective neighborhood structures leads to limitations. To address this problem, this paper presents a Neighborhood-Adaptive Multi-cluster Ranking (NAMR) framework by leveraging the heterogeneity of local structures. Specifically, considering that indexing algorithms are usually required in large-scale retrieval, NAMR characterizes an image from two kinds of embeddings (i.e., sample embedding and structure embedding). The sample embedding can be trained using any distance-based loss, while the structure embedding representing the neighborhood structure can be jointly learned with the sample embedding in a self-supervised multi-cluster ranking manner. In this way, existing indexing algorithms can seamlessly support large-scale retrieval employing NAMR embeddings without any modifications. We evaluate the proposed model on five standard benchmarks, consistently and explicitly improving four baselines (especially the simplest triplet loss) and achieving state-of-the-art performance. Pandeng Li, Hongtao Xie 0001, Jiannan Ge, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Prototypical Matching Networks for Video Object SegmentationabstractSemi-supervised video object segmentation is the task of segmenting the target in sequential frames given the ground truth mask in the first frame. The modern approaches usually utilize such a mask as pixel-level supervision and typically exploit pixel-to-pixel matching between the reference frame and current frame. However, the matching at pixel level, which overlooks the high-level information beyond local areas, often suffers from confusion caused by similar local appearances. In this paper, we present Prototypical Matching Networks (PMNet) - a novel architecture that integrates prototypes into matching-based video objection segmentation frameworks as high-level supervision. Specifically, PMNet first divides the foreground and background areas into several parts according to the similarity to the global prototypes. The part-level prototypes and instance-level prototypes are generated by encapsulating the semantic information of identical parts and identical instances, respectively. To model the correlation between prototypes, the prototype representations are propagated to each other by reasoning on a graph structure. Then, PMNet stores both the pixel-level features and prototypes in the memory bank as the target cues. Three affinities, i.e., pixel-to-pixel affinity, prototype-to-pixel affinity, and prototype-to-prototype affinity, are derived to measure the similarity between the query frame and the features in the memory bank. The features aggregated from the memory bank using these affinities provide powerful discrimination from both the pixel-level and prototype-level perspectives. Extensive experiments conducted on four benchmarks demonstrate superior results than the state-of-the-art video object segmentation techniques. Fanchao Lin, Zhaofan Qiu, Chuanbin Liu 0001, Ting Yao 0003, Hongtao Xie 0001, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | What is the Real Need for Scene Text Removal? Exploring the Background Integrity and Erasure Exhaustivity PropertiesabstractAs a crucial application in privacy protection, scene text removal (STR) has received amounts of attention in recent years. However, existing approaches coarsely erasing texts from images ignore two important properties: the background texture integrity (BI) and the text erasure exhaustivity (EE). These two properties directly determine the erasure performance, and how to maintain them in a single network is the core problem for STR task. In this paper, we attribute the lack of BI and EE properties to the implicit erasure guidance and imbalanced multi-stage erasure respectively. To improve these two properties, we propose a new ProgrEssively Region-based scene Text eraser (PERT). There are three key contributions in our study. First, a novel explicit erasure guidance is proposed to enhance the BI property. Different from implicit erasure guidance modifying all the pixels in the entire image, our explicit one accurately performs stroke-level modification with only bounding-box level annotations. Second, a new balanced multi-stage erasure is constructed to improve the EE property. By balancing the learning difficulty and network structure among progressive stages, each stage takes an equal step towards the text-erased image to ensure the erasure exhaustivity. Third, we propose two new evaluation metrics called BI-metric and EE-metric, which make up the shortcomings of current evaluation tools in analyzing BI and EE properties. Compared with previous methods, PERT outperforms them by a large margin in both BI-metric ( ↑ 6.13 %) and EE-metric ( ↑ 1.9 %), obtaining SOTA results with high speed (71 FPS) and at least 25% lower parameter complexity. Code will be available at https://github.com/wangyuxin87/PERT. Yuxin Wang 0002, Hongtao Xie 0001, Zixiao Wang 0002, Yadong Qu, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Discriminative Feature Mining Based on Frequency Information and Metric Learning for Face Forgery DetectionabstractFace forgery detection has received considerable attention due to security concerns about abnormal faces generated by face forgery technology. While recent researches have made prominent progress, they still suffer from two limitations: a) the learned features supervised by softmax loss are insufficiently discriminative, since the softmax loss fails to explicitly boost inter-class separability and intra-class compactness; b) hand-crafted features are unable to effectively mine forgery patterns from frequency domain. To address the two problems, this paper proposes a novel frequency-aware discriminative feature learning framework. Specifically, we design an innovative single-center loss which compresses mere intra-class variations of natural faces while encouraging inter-class differences between natural and manipulated faces in the embedding space. Supervised by such a loss, more discriminative features can be learned with less optimization difficulty. As for frequency-related features, a frequency feature adaptively generated module is developed to capture frequency clues in a data-driven manner. Besides, to better fuse the features of both RGB domain and frequency domain, this paper devises a fusion module based on positional correlation of features. The effectiveness and superiority of our framework have been proved by extensive experiments and our approach achieves state-of-the-art performance in both in-dataset and cross-dataset evaluation. Jiaming Li 0017, Hongtao Xie 0001, Lingyun Yu 0002, Xingyu Gao 0001, Yongdong Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Learning Cross-Channel Representations for Semantic SegmentationabstractSemantic segmentation is a fundamental problem in multimedia which requires delicate per-pixel predictions of object categories. Recently, many researchers strive to refine the pixel-wise feature withspatial-contextual information. However, many of them still neglect the invisible hand of cross-channelinformation which provides inherent semantics to facilitate the segmentation performance. On the one hand, in the feature extraction stage, enhancing informative channels and suppressing trivial ones contribute to the acquisition of valuable semantic features, and thus improving the segmentation accuracy. On the other hand, in the prediction stage, we can predict the complete objects more clearly by finding the connections and complements between different channels, which can also contribute to the pixel prediction. And based on this idea, we propose a novel Channel-Adaptive Network for semantic segmentation, which is capable of enhancing the features from the perspective of channels in both feature extraction stage and prediction stage. Specifically, we propose two modules: (i) the Comprehensive Information Channel Attention (CiCA) module that addresses the shortcomings of existing channel attention by learning both low and high frequency components within each channel for emphasizing the informative channels; (ii) the Inter-Channel Relationship Reasoning (iCRR) module which is applied on the top of the feature extractor to adaptively enhance the interdependent channels by mining the complementary associations between them. Besides, our Channel-Adaptive Network is highly flexible, with a plug-and-play design. Extensive experiments have demonstrated that our method achieves the state-of-the-art segmentation performance on three challenging datasets, including Cityscapes (82.1%), ADE20K (46.51%) and PASCAL Context (55.0%). Lingfeng Ma, Hongtao Xie 0001, Chuanbin Liu 0001, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | ADNet: Rethinking the Shrunk Polygon-Based Approach in Scene Text DetectionabstractTo localize text regions and separate close instances, the shrunk polygon is widely used in recent scene text detection methods. However, there exist two problems: 1) Existing methods fail to consider the aspect ratio sensitive problem when reconstructing the text instance from shrunk polygon. 2) Texts with extreme aspect ratios will lead to the fracture of shrunk polygons. To handle these two problems, in this paper, we propose a novel Adaptive Dilation Network (ADNet) to focus on the reconstruction process from shrunk polygon, which aims to provide a tight and complete text representation. Firstly, instead of using a fixed dilation factor, ADNet uses an aspect ratio-wise dilation factor to reconstruct the text region from shrunk polygon for each text instance. Such an instance-wise dilation factor considers the scale correlation between the original and shrunk polygon, and thus can guide an adaptive text region reconstruction for texts with large aspect ratio variance. Secondly, to deal with the fracture of detection results, a new Efficient Spatial Relationship Module (ESRM) is devised to capture long-range dependencies with low computation cost. ESRM uses a novel Weighted Pooling to reduce the resolution of feature maps without much information loss. Compared with the existing methods, ADNet further explores the potential of shrunk polygon-based approaches and obtains excellent detection results at an impressive speed. Extensive experiments on several datasets (Total-Text, CTW1500, MSRA-TD500 and ICDAR2015) verify the superiority of our method. Yadong Qu, Hongtao Xie 0001, Shancheng Fang, Yuxin Wang 0002, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Learning Pixel Affinity Pyramid for Arbitrary-Shaped Text DetectionabstractArbitrary-shaped text detection in natural images is a challenging task due to the complexity of the background and the diversity of text properties. The difficulty lies in two aspects: accurate separation of adjacent texts and sufficient text feature representation. To handle these problems, we consider text detection as instance segmentation and propose a novel text detection framework, which jointly learns semantic segmentation and a pixel affinity pyramid in a unified fully convolutional network. Specifically, the pixel affinity pyramid is proposed to encode multi-scale instance affiliation relationships of pixels, which is not only robust to varying shapes of text but also provides an accurate boundary description for separating closely located texts. In the inference phase, a simple but effective post-processing is presented to reconstruct text instances from the semantic segmentation results under the guidance of the learned pixel affinity pyramid, achieving good accuracy and efficiency. Furthermore, to enhance the representation of text features in the neural network, two modules — the Region Enhancement Module (REM) and Attentional Fusion Module (AFM) — are proposed. The REM models the semantic correlations of regional features to enhance the features from the text area, which effectively suppresses false-positive detection. The AFM adaptively fuses multi-scale textual information through an attention mechanism to obtain abundant text semantic features, which benefits multi-sized text detection. Extensive ablation experiments are conducted demonstrating the effectiveness of the REM and AFM. Evaluation results on standard benchmarks, including Total-Text, ICDAR2015, SCUT-CTW1500, and MSRA-TD500, show that our method surpasses most existing text detectors and achieves state-of-the-art performance, denoting its superior capability in detecting arbitrary-shaped texts. Zilong Fu, Hongtao Xie 0001, Shancheng Fang, Yuxin Wang 0002, Mengting Xing, Yongdong Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Constructing Spatio-Temporal Graphs for Face Forgery DetectionabstractRecently, advanced development of facial manipulation techniques threatens web information security, thus, face forgery detection attracts a lot of attention. It is clear that both spatial and temporal information of facial videos contains the crucial manipulation traces, which are inevitably created during the generation process. However, most existing face forgery detectors only focus on the spatial artifacts or the temporal incoherence, and they are struggling to learn a significant and general kind of representations for manipulated facial videos. In this work, we propose to construct spatial-temporal graphs for fake videos to capture the spatial inconsistency and the temporal incoherence at the same time. To model the spatial-temporal relationship among the graph nodes, a novel forgery detector named Spatio-Temporal Graph Network (STGN) is proposed, which contains two kinds of graph-convolution-based units, the Spatial Relation Graph Unit (SRGU) and the Temporal Attention Graph Unit (TAGU). To exploit spatial information, the SRGU models the inconsistency between each pair of patches in the same frame, instead of focusing on the low-level local spatial artifacts which are vulnerable to samples created by unseen manipulation methods. And, the TAGU is proposed to model the long-distance temporal relation among the patches at the same spatial position in different frames with a graph attention mechanism based on the inter-node similarity. With the SRGU and the TAGU, our STGN can combine the discriminative power of spatial inconsistency and the generalization capacity of temporal incoherence for face forgery detection. Our STGN achieves state-of-the-art performances on several popular forgery detection datasets. Extensive experiments demonstrate both the superiority of our STGN on intra manipulation evaluation and the effectiveness for new sorts of face forgery videos on cross manipulation evaluation. Zhihua Shang, Hongtao Xie 0001, Lingyun Yu 0002, Zhengjun Zha, Yongdong Zhang 0001 |
ACM Trans. Web | 2 |
| 2023 | Multi-task hourglass network for online automatic diagnosis of developmental dysplasia of the hip
Hongtao Xie 0001, Qingfeng Tan, Chuanbin Liu 0001, Zhendong Mao 0001, Yongdong Zhang 0001 |
World Wide Web (WWW) | 2 |
| 2022 | Neighborhood-Adaptive Structure Augmented Metric LearningabstractMost metric learning techniques typically focus on sample embedding learning, while implicitly assume a homogeneous local neighborhood around each sample, based on the metrics used in training ( e.g., hypersphere for Euclidean distance or unit hyperspherical crown for cosine distance). As real-world data often lies on a low-dimensional manifold curved in a high-dimensional space, it is unlikely that everywhere of the manifold shares the same local structures in the input space. Besides, considering the non-linearity of neural networks, the local structure in the output embedding space may not be homogeneous as assumed. Therefore, representing each sample simply with its embedding while ignoring its individual neighborhood structure would have limitations in Embedding-Based Retrieval (EBR). By exploiting the heterogeneity of local structures in the embedding space, we propose a Neighborhood-Adaptive Structure Augmented metric learning framework (NASA), where the neighborhood structure is realized as a structure embedding, and learned along with the sample embedding in a self-supervised manner. In this way, without any modifications, most indexing techniques can be used to support large-scale EBR with NASA embeddings. Experiments on six standard benchmarks with two kinds of embeddings, i.e., binary embeddings and real-valued embeddings, show that our method significantly improves and outperforms the state-of-the-art methods. Pandeng Li, Yan Li 0068, Hongtao Xie 0001, Lei Zhang 0119 |
AAAI | 3 |
| 2022 | Partial Class Activation Attention for Semantic SegmentationabstractCurrent attention-based methods for semantic segmentation mainly model pixel relation through pairwise affinity and coarse segmentation. For the first time, this paper explores modeling pixel relation via Class Activation Map (CAM). Beyond the previous CAM generated from image-level classification, we present Partial CAM, which sub-divides the task into region-level prediction and achieves better localization performance. In order to eliminate the intra-class inconsistency caused by the variances of local context, we further propose Partial Class Activation Attention (PCAA) that simultaneously utilizes local and global class-level representations for attention calculation. Once obtained the partial CAM, PCAA collects local class centers and computes pixel-to-class relation locally. Applying local-specific representations ensures reliable results under different local contexts. To guarantee global consistency, we gather global representations from all local class centers and conduct feature aggregation. Experimental results confirm that Partial CAM outperforms the previous two strategies as pixel relation. Notably, our method achieves state-of-the-art performance on several challenging benchmarks including Cityscapes, Pascal Context, and ADE20K. Code is available at https://github.com/lsa1997/PCAA. Sun'ao Liu, Hongtao Xie 0001, Yongdong Zhang 0001, Qi Tian 0001 |
CVPR | 2 |
| 2022 | Dual-Stream Knowledge-Preserving Hashing for Unsupervised Video Retrieval
Pandeng Li, Hongtao Xie 0001, Jiannan Ge, Lei Zhang 0119, Shaobo Min, Yongdong Zhang 0001 |
ECCV (14) | 2 |
| 2022 | Detecting Tampered Scene Text in the Wild
Yuxin Wang 0002, Hongtao Xie 0001, Mengting Xing, Jing Wang 0221, Shenggao Zhu, Yongdong Zhang 0001 |
ECCV (28) | 2 |
| 2022 | Weakly Supervised Pediatric Bone Age Assessment Using Ultrasonic Images via Automatic Anatomical RoI DetectionabstractBone age assessment (BAA) is vital in pediatric clinical diagnosis. Existing deep learning methods predict bone age based on Regions of Interest (RoIs) detection or segmentation of hand radiograph, which requires expensive annotations. Limitations of radiographic technique on imaging and cost hinder their clinical application as well. Compared to X-ray images, ultrasonic images are rather clean, cheap and flexible, but the deep learning research on ultrasonic BAA is still a white space. For this purpose, we propose a weakly supervised interpretable framework entitled USB-Net, utilizing ultrasonic pelvis images and only image-level age annotations. USB-Net consists of automatic anatomical RoI detection stage and age assessment stage. In the detection stage, USB-Net locates the discriminative anatomical RoIs of pelvis through attention heatmap without any extra RoI supervision. In the assessment stage, the cropped anatomical RoI patch is fed as fine-grained input to estimate age. In addition, we provide the first ultrasonic BAA dataset composed of 1644 ultrasonic hip joint images with image-level labels of age and gender. The experimental results verify that our model keeps consistent attention with human knowledge and achieves 16.24 days mean absolute error (MAE) on USBAA dataset. Yunyan Yan, Chuanbin Liu 0001, Hongtao Xie 0001, Zhendong Mao 0001 |
ICMR | 3 |
| 2022 | Geometry Aligned Variational Transformer for Image-conditioned Layout GenerationabstractLayout generation is a novel task in computer vision, which combines the challenges in both object localization and aesthetic appraisal, widely used in advertisements, posters and slides design. An accurate and pleasant layout should consider both the intra-domain relationship within layout elements and the inter-domain relationship between layout elements and image. However, most previous methods simply focus on image-content-agnostic layout generation, without leveraging the complex visual information from the image. To this end, we explore a novel paradigm entitled image-conditioned layout generation, which aims to add text overlays to an image in a semantically coherent manner. Specifically, we propose an Image-Conditioned Variational Transformer (ICVT) that autoregressively generates various layouts in an image. First, self-attention mechanism is adopted to model the contextual relationship within layout elements, while cross-attention mechanism is used to fuse the visual information of conditional images. Subsequently, we take them as building blocks of conditional variational autoencoder (CVAE), which demonstrates appealing diversity. Second, in order to alleviate the gap between layout elements domain and visual domain, we design a Geometry Alignment module, in which the geometric information of the image is aligned with the layout representation. In addition, we construct a large-scale advertisement poster layout designing dataset with delicate layout and saliency map annotations. Experimental results show that our model can adaptively generate layouts in the non-intrusive area of the image, resulting in a harmonious layout design. Yunning Cao, Chuanbin Liu 0001, Hongtao Xie 0001, Tiezheng Ge, Yuning Jiang 0001 |
ACM Multimedia | 5 |
| 2022 | Dual Part Discovery Network for Zero-Shot LearningabstractZero-Shot Learning (ZSL) aims to recognize unseen classes by transferring knowledge from seen classes. Recent methods focus on learning a common semantic space to align visual and attribute information. However, they always over-relied on provided attributes and ignored the category discriminative information that contributes to accurate unseen class recognition, resulting in weak transferability. To this end, we propose a novel Dual Part Discovery Network (DPDN) that considers both attribute and category discriminative information by discovering attribute-guided parts and category-guided parts simultaneously to improve knowledge transfer. Specifically, for attribute-guided parts discovery, DPDN can localize the regions with specific attribute information and significantly bridge the gap between visual and semantic information guided by the given attributes. For category-guided parts discovery, the local parts are explored to discover other important regions that bring latent crucial details ignored by attributes, with the guidance of adaptive category prototypes. To better mine the transferable knowledge, we impose class correlations constraints to regularize the category prototypes. Finally, attribute- and category-guided parts complement each other and provide adequate discriminative subtle information for more accurate unseen class recognition. Extensive experimental results demonstrate that DPDN can discover discriminative parts and outperform state-of-the-art methods on three standard benchmarks. Jiannan Ge, Hongtao Xie 0001, Shaobo Min, Pandeng Li, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2022 | Wavelet-enhanced Weakly Supervised Local Feature Learning for Face Forgery DetectionabstractFace forgery detection is getting increasing attention due to the security threats caused by forged faces. Recently, local patch-based approaches have achieved sound achievements due to effective attention to local details. However, there are still unignorable problems: a) local feature learning requires patch-level labels to circumvent label noise, which is not practical in real-world scenarios; b) the commonly used DCT (FFT) transform loses all spatial information, which brings difficulty in handling local details. To compensate for such limitations, a novel wavelet-enhanced weakly supervised local feature learning framework is proposed in this paper. Specifically, to supervise the learning of local features with only image-level labels, two modules are devised based on the idea of multi-instance learning: local relation constraint module (LRCM) and category knowledge-guided local feature aggregation module (CKLFA). LRCM constrains the maximum distance between local features of forged face images greater than that of real face images. CKLFA adaptively aggregates local features based on their correlation to global embedding containing global category information. Combining these two modules, the network is encouraged to learn discriminative local features supervised only by image-level labels. Besides, a multi-level wavelet-powered feature enhancement module is developed to promote the network mining local forgery artifacts from spatio-frequency domain, which is beneficial to learning discriminative local features. Extensive experiments show that our approach outperforms previous state-of-the-art methods when only image-level labels are available and achieves comparable or even better performance than counterparts using patch-level labels. Jiaming Li 0017, Hongtao Xie 0001, Lingyun Yu 0002, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2022 | Proxy Probing Decoder for Weakly Supervised Object Localization: A Baseline InvestigationabstractWeakly supervised object localization (WSOL) aims to localize the object with only image category labels. Existing methods generally fine-tune the models with manually selected training epochs and subjective loss functions to mitigate the partial activation problem of the classification-based model. However, such fine-tuning scheme would cause the model to degrade, e.g. affect the classification performance and generalization capabilities of the pre-trained model. In this paper, we propose a novel method named Proxy Probing Decoder (PPD) to meet these challenges, which utilizes the segmentation property of self-attention map in the self-supervised vision transformer and breaks through model fine-tuning with a novel proxy probing decoder. Specifically, we utilize the self-supervised vision transformer to capture long-range dependencies and avoid partial activation. Then we simply adopt a proxy consisting of a series of decoding layers to transform the feature representations into the heatmap of the objects' foreground and conduct localization. The backbone parameters are frozen during training while the proxy is used to decode the feature and localize the object. In this way, the vision transformer model can maintain the feature representation capabilities and only the proxy is required for adapting to the task. Without bells and whistles, our framework achieves 55.0% Top-1 Loc on the ILSVRC2012 dataset and 78.8% Top-1 Loc on the CUB-200-2011 dataset, which surpasses state-of-the-art by a large margin and provides a simple baseline. Codes and models will be available on Github. Hongtao Xie 0001, Chuanbin Liu 0001, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2022 | Boat in the Sky: Background Decoupling and Object-aware Pooling for Weakly Supervised Semantic SegmentationabstractPrevious image-level weakly-supervised semantic segmentation methods based on Class Activation Map (CAM) have two limitations: 1) focusing on partial discriminative foreground regions and 2) containing undesirable background. The above issues are attributed to the spurious correlations between the object and background (semantic ambiguity) and the insufficient spatial perception ability of the classification network (spatial ambiguity). In this work, we propose a novel self-supervised framework to mitigate the semantic and spatial ambiguity from the perspectives of background bias and object perception. First, a background decoupling mechanism (BDM) is proposed to handle the semantic ambiguity by regularizing the consistency of predicted CAMs from the samples with identical foregrounds but different backgrounds. Thus, a decoupled relationship is constructed to reduce the dependence between the object instance and the scene information. Second, a global object-aware pooling (GOP) is introduced to alleviate spatial ambiguity. The GOP utilizes a learnable object-aware map to dynamically aggregate spatial information and further improve the performance of CAMs. Extensive experiments demonstrate the effectiveness of our method by achieving new state-of-the-art results on both the Pascal VOC 2012 and MS COCO 2014 datasets. Hongtao Xie 0001, Yuxin Wang 0002, Sun'ao Liu, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2022 | REMOT: A Region-to-Whole Framework for Realistic Human Motion TransferabstractHuman Video Motion Transfer (HVMT) aims to, given an image of a source person, generate his/her video that imitates the motion of the driving person. Existing methods for HVMT mainly exploit Generative Adversarial Networks (GANs) to perform the warping operation based on the flow estimated from the source person image and each driving video frame. However, these methods always generate obvious artifacts due to the dramatic differences in poses, scales, and shifts between the source person and the driving person. To overcome these challenges, this paper presents a novel REgion-to-whole human MOtion Transfer (REMOT) framework based on GANs. To generate realistic motions, the REMOT adopts a progressive generation paradigm: it first generates each body part in the driving pose without flow-based warping, then composites all parts into a complete person of the driving motion. Moreover, to preserve the natural global appearance, we design a Global Alignment Module to align the scale and position of the source person with those of the driving person based on their layouts. Furthermore, we propose a Texture Alignment Module to keep each part of the person aligned according to the similarity of the texture. Finally, through extensive quantitative and qualitative experiments, our REMOT achieves state-of-the-art results on two public benchmarks. Quanwei Yang, Xinchen Liu, Wu Liu 0005, Hongtao Xie 0001, Xiaoyan Gu 0001, Lingyun Yu 0002, Yongdong Zhang 0001 |
ACM Multimedia | 4 |
| 2022 | Bridging the Gap Between Vision Transformers and Convolutional Neural Networks on Small DatasetsabstractThere still remains an extreme performance gap between Vision Transformers (ViTs) and Convolutional Neural Networks (CNNs) when training from scratch on small datasets, which is concluded to the lack of inductive bias. In this paper, we further consider this problem and point out two weaknesses of ViTs in inductive biases, that is, the spatial relevance and diverse channel representation. First, on spatial aspect, objects are locally compact and relevant, thus fine-grained feature needs to be extracted from a token and its neighbors. While the lack of data hinders ViTs to attend the spatial relevance. Second, on channel aspect, representation exhibits diversity on different channels. But the scarce data can not enable ViTs to learn strong enough representation for accurate recognition. To this end, we propose Dynamic Hybrid Vision Transformer (DHVT) as the solution to enhance the two inductive biases. On spatial aspect, we adopt a hybrid structure, in which convolution is integrated into patch embedding and multi-layer perceptron module, forcing the model to capture the token features as well as their neighboring features. On channel aspect, we introduce a dynamic feature aggregation module in MLP and a brand new "head token" design in multi-head self-attention module to help re-calibrate channel representation and make different channel group representation interacts with each other. The fusion of weak channel representation forms a strong enough representation for classification. With this design, we successfully eliminate the performance gap between CNNs and ViTs, and our DHVT achieves a series of state-of-the-art performance with a lightweight model, 85.68% on CIFAR-100 with 22.8M parameters, 82.3% on ImageNet-1K with 24.0M parameters. Code is available at https://github.com/ArieSeirack/DHVT. Zhiying Lu, Hongtao Xie 0001, Chuanbin Liu 0001, Yongdong Zhang 0001 |
NeurIPS | 2 |
| 2022 | Attention-guided transformation-invariant attack for black-box adversarial examplesabstractWith the development of media convergence, information acquisition is no longer limited to traditional media, such as newspapers and televisions, but more from digital media on the Internet, where media contents should be under supervision by platforms. At present, the media content analysis technology of Internet platforms relies on deep neural networks (DNNs). However, DNNs show vulnerability to adversarial examples, which results in security risks. Therefore, it is necessary to adequately study the internal mechanism of adversarial examples to build more effective supervision models. When coming to practical applications, supervision models are mostly faced with black-box attacks, where cross-model transferability of adversarial examples has attracted increasing attention. In this paper, to improve the transferability of adversarial examples, we propose an attention-guided transformation-invariant adversarial attack method, which incorporates an attention mechanism to disrupt the most distinctive features and simultaneously ensures adversarial attack invariance under different transformations. Specifically, we dynamically weight the latent features according to an attention mechanism and disrupt them accordingly. Meanwhile, considering the lack of semantics in low-level features, high-level semantics are introduced as spatial guidance to make low-level feature perturbations concentrate on the most discriminative regions. Moreover, since the attention heatmaps may vary significantly across different models, a transformation-invariant aggregated attack strategy is proposed to alleviate overfitting to the proxy model attention. Comprehensive experimental results show that the proposed method can significantly improve the transferability of adversarial examples. Lingyun Yu 0002, Hongtao Xie 0001, Bo Wu 0018, Yongdong Zhang 0001 |
Int. J. Intell. Syst. | 4 |
| 2022 | Semi-Supervised Text Detection With Accurate Pseudo-LabelsabstractRecent scene text detection methods have made great progress. However, existing methods rely heavily on extensive labeled data, which is very time-consuming and expensive. In this letter, we propose a novel semi-supervised text detection method to alleviate the dependence of text detectors on labeled data by generating accurate pseudo-labels and performing effective data augmentations. Specifically, a dual-threshold pseudo-label generation algorithm is designed to divide the prediction results into background regions, text regions, and uncertain regions. The definition of uncertain regions obviously improves the accuracy of pseudo-labels. To obtain accurate pseudo-labels for text at various scales, we first design a scale-aware loss function to adaptively adjust the loss weight of different scale texts. Then, a multi-scale feature extraction module is proposed to extract multi-scale text features and adaptively weight these features according to the scale of the text. Moreover, effective data augmentations are explored to use unlabeled data to improve the robustness of the model to various texts. Experiments show that our method achieves state-of-the-art performance on several datasets(e.g., outperforms existing methods by 2.0% on TD500). Yu Zhou 0016, Hongtao Xie 0001, Shancheng Fang, Yongdong Zhang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Self-Supervised Synthesis Ranking for Deep Metric LearningabstractThe core purpose of deep metric learning is to construct an embedding space, where objects belonging to the same class are gathered together and the ones from different classes are pushed apart. Most existing approaches typically insist to inter-class characteristics,e.g., class-level information or instance-level similarity, to obtain semantic relevance of data points and get a large margin between different classes in the embedding space. However, the intra-class characteristics,e.g., local manifold structure or relative relationship within the same class, are usually overlooked in the learning process. Hence the output embeddings have limitation in retrieving a good ranking result if existing multiple positive samples. And the local data structure of embedding space cannot be fully exploited since lack of relative ranking information. As a result, the model is prone to overfitting on a train set and get low generalization on the test set (unseen classes) when losing sight of intra-class variance. This paper presents a novel self-supervised synthesis ranking auxiliary framework, which captures intra-class characteristics as well as inter-class characteristics for better metric learning. Our method designs a synthetic samples generation of polar coordinates to generate measurable intra-class variance with different strength and diversity in the latent space, which can simulate the various local structure change of intra-class in the initial data domain. And then formulates a self-supervised learning procedure to fully exploit this property and preserve it in the embedding space. As a result, the learned embedding space not only keeps inter-class discrimination but also owns subtle intra-class diversity, leading to better global and local embedding structures. Extensive experiments on five benchmarks show that our method significantly improves and outperforms the state-of-the-art methods on the performances of both retrieval and ranking by 2%-4% (personal use of this material is permitted. However, permission to use this material for any other purposes must be obtained from the IEEE by sending an email to [email protected]). Zheren Fu, Zhendong Mao 0001, Chenggang Yan 0001, Anan Liu, Hongtao Xie 0001, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Bilateral Temporal Re-Aggregation for Weakly-Supervised Video Object SegmentationabstractWeakly-supervised video object segmentation is an emerging video task to track and segment the target given a simple bounding box label, which requires the method to fully catch and utilize the target information. Most existing approaches only rely on the guidance of a single frame and ignore the interaction between different frames when gathering information, making them hard to achieve reliable target representation. In this paper, we propose to capture the temporal dependencies and gather information from multiple frames through bilateral temporal re-aggregation. We explore three schemes to build the aggregation: 1) a two-stage re-aggregation mechanism is applied to provide target prior to the current frame, which obtains more valid feature matching and information aggregation; 2) a query-memory bilateral aggregation module is proposed to aggregate features from an unlimited amount of past frames and enable the mutual perception between different frames to validate the gathered information; 3) we guide the learning of aggregation modules through a novel cross-task representation distillation, transferring the knowledge from a semi-supervised model to our weakly-supervised model without increasing the inference latency. These schemes collaboratively build an efficient and competent aggregation process, thus we can fully exploit the video context to make the inference. Experimental results on four benchmarks show that our method achieves superior performance than previous methods and still maintains the efficiency ($e.g$., overall scores of 70.4% and 72.5% on the YouTube-VOS and DAVIS 2017 validation sets, respectively). Fanchao Lin, Hongtao Xie 0001, Chuanbin Liu 0001, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Deep Fourier Ranking Quantization for Semi-Supervised Image RetrievalabstractTo reduce the extreme label dependence of supervised product quantization methods, the semi-supervised paradigm usually employs massive unlabeled data to assist in regularizing deep networks, thereby improving model performance. However, the existing method focuses on the overall distribution consistency between unlabeled data and class prototypes, while ignoring subtle individual variances between unlabeled instances. Therefore, the local neighborhood structure is not fully explored, which will cause the model to easily overfit in the training set. In this paper, we introduce a new Fourier perspective to alleviate this issue by exploring the semantic relations between unlabeled instances in a self-supervised manner. Specifically, based on Fourier Transform, we first design a Phase Mixing (PM) strategy, which can manipulate the mixing area and values of the phase component between two images to control the proportion of semantic information. In this way, we can construct multi-level similarity neighbors naturally for unlabeled data. Then, a ranking quantization loss is formulated to perceive multi-level semantic variances in neighbor instances, which improves the robustness and generalization of the model. Extensive experiments in three different semi-supervised settings show that our method outperforms existing state-of-the-art methods by averaged 3.95% improvement on four datasets. Pandeng Li, Hongtao Xie 0001, Shaobo Min, Jiannan Ge, Xun Chen 0001, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | PETR: Rethinking the Capability of Transformer-Based Language Model in Scene Text RecognitionabstractThe exploration of linguistic information promotes the development of scene text recognition task. Benefiting from the significance in parallel reasoning and global relationship capture, transformer-based language model (TLM) has achieved dominant performance recently. As a decoupled structure from the recognition process, we argue that TLM's capability is limited by the input low-quality visual prediction. To be specific: 1) The visual prediction with low character-wise accuracy increases the correction burden of TLM. 2) The inconsistent word length between visual prediction and original image provides a wrong language modeling guidance in TLM. In this paper, we propose a Progressive scEne Text Recognizer (PETR) to improve the capability of transformer-based language model by handling above two problems. Firstly, a Destruction Learning Module (DLM) is proposed to consider the linguistic information in the visual context. DLM introduces the recognition of destructed images with disordered patches in the training stage. Through guiding the vision model to restore patch orders and make word-level prediction on the destructed images, visual prediction with high character-wise accuracy is obtained by exploring inner relationship between the local visual patches. Secondly, a new Language Rectification Module (LRM) is proposed to optimize the word length for language guidance rectification. Through progressively implementing LRM in different language modeling steps, a novel progressive rectification network is constructed to handle some extremely challenging cases (e.g. distortion, occlusion, etc.). By utilizing DLM and LRM, PETR enhances the capability of transformer-based language model from a more general aspect, that is, focusing on the reduction of correction burden and rectification of language modeling guidance. Compared with parallel transformer-based methods, PETR obtains 1.0% and 0.8% improvement on regular and irregular datasets respectively while introducing only 1.7M additional parameters. The extensive experiments on both English and Chinese benchmarks demonstrate that PETR achieves the state-of-the-art results. Yuxin Wang 0002, Hongtao Xie 0001, Shancheng Fang, Mengting Xing, Jing Wang 0221, Shenggao Zhu, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Dynamic-Aware Federated Learning for Face Forgery Video DetectionabstractThe spread of face forgery videos is a serious threat to information credibility, calling for effective detection algorithms to identify them. Most existing methods have assumed a shared or centralized training set. However, in practice, data may be distributed on devices of different enterprises that cannot be centralized to share due to security and privacy restrictions. In this article, we propose a Federated Learning face forgery detection framework to train a global model collaboratively while keeping data on local devices. In order to make the detection model more robust, we propose a novel Inconsistency-Capture module (ICM) to capture the dynamic inconsistencies between adjacent frames of face forgery videos. The ICM contains two parallel branches. The first branch takes the whole face of adjacent frames as input to calculate a global inconsistency representation. The second branch focuses only on the inter-frame variation of critical regions to capture the local inconsistency. To the best of our knowledge, this is the first work to apply federated learning to face forgery video detection, which is trained with decentralized data. Extensive experiments show that the proposed framework achieves competitive performance compared with existing methods that are trained with centralized data, with higher-level security and privacy guarantee. Ziheng Hu, Hongtao Xie 0001, Lingyun Yu 0002, Xingyu Gao 0001, Zhihua Shang, Yongdong Zhang 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2022 | Online Residual Quantization Via Streaming Data Correlation PreservingabstractRecently, the online retrieval task has been receiving widespread attention, which is closely related to many real-world applications. However, existing online retrieval methods based on hashing suffer from two main problems: a) the models tend to be biased towards the current streaming data due to unavailable history streaming data; b) when new streaming data comes in and the hashing functions have been updated, all history binary codes should be recomputed, which takes much computation burden. To address the above two issues, we propose a novel Online Residual Quantization (ORQ) method that can achieve efficient streaming data quantization via the small-scale residual quantization codebooks. For the first problem, we design a residual quantization module by learning multiple residual codebooks to quantize the float streaming data, which effectively reduces the quantization error and enables the binary codes to be easily reconstructed back to original float data. Then, with the reconstructed history data, a balanced affinity matrix is developed to model the semantic relationship,e.g.,similarity and difference, between the history and current data distributions, which can prevent the model from being biased towards the current data distribution. For the second problem, when inputting current streaming data, only the residual codebooks should be updated, instead of the whole history binary codes in hashing-based methods, which significantly reduces the computation burden. Comprehensive experiments on six benchmarks demonstrate that ORQ yields significant improvements (i.e.,1.2%$\sim$4.9% in average mAP) compared to the state-of-the-art methods. Pandeng Li, Hongtao Xie 0001, Shaobo Min, Zhengjun Zha, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Boundary-Aware Arbitrary-Shaped Scene Text Detector With Learnable Embedding NetworkabstractBenefiting from the popularity of deep learning theory, scene text detection algorithms have developed rapidly in recent years. Methods representing text region by text segmentation map are proved to capture arbitrary-shaped text in a more flexible and accurate way. However, such segmentation-based methods are prone to be disturbed by the text-like background patterns (like the fence, grass, etc.), which generally suffer from imprecise boundary detail problem. In this paper, LEMNet is proposed to handle the imprecise boundary problem by guiding the generation of text boundary based on a priori constraint. In the training stage, Boundary Segmentation Branch is firstly constructed to predict coarse boundary mask for each text instance. Then, through mapping pixels into an embedding space, the proposed Pixel Embedding Branch makes the embedding representation of boundary points learn to be more similar, meanwhile enlarging the characteristic distance between background points and boundary points. During inference, noise in the coarse boundary segmentation map can be effectively suppressed by a Noisy Point Suppression Algorithm among pixel embedding vectors. In this way, LEMNet can generate a more precise boundary description of text regions. To further enhance the distinguishability of boundary features, we propose a Context Enhancement Module to capture feature interactions in different representation subspaces, in which features are parallelly performed attention and concatenated to generate enhanced features. Extensive experiments are conducted over four challenging datasets, which demonstrate the effectiveness of LEMNet. Specifically, LEMNet achieves F-measure of 85.2%, 87.6% and 85.2% on CTW1500, Total-Text and MSRA-TD500 respectively, which is the latest SOTA. Mengting Xing, Hongtao Xie 0001, Qingfeng Tan, Shancheng Fang, Yuxin Wang 0002, Zhengjun Zha, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Multimodal Learning for Temporally Coherent Talking Face Generation With Articulator SynergyabstractTalking face generation is a demanding task to synthesize a high quality video with accurate lip synchronization and rhythmic head motion. Any subtle artifacts could be sensitively captured by humans and lead to poor visual quality. Existing methods tend to employ a conditional generation solution, which introduces facial landmarks to bridge the input information and output videos. However, these methods always suffer from unrealistic facial animations, because 1) they only take single-mode input, but ignore the complementarity of multimodal inputs for lip-sync improvement; 2) they only explore lip movements, but ignore the articulator synergy between lips and jaw; 3) they generate each video frame in a temporal-independent way, but ignore the temporal continuity among the entire video. To address these limitations, in this paper, we present a novel method to generate realistic and temporally coherent talking heads by considering multimodal inputs, articulator synergy, inter-frame consistency and intra-frame consistency. Firstly, for landmark prediction, a novel Multiple Synergy Network (MSN) is proposed to improve the accuracy of landmark prediction by incorporating multimodal inputs (i.e., audio and text inputs). Besides, instead of merely considering lip landmarks, we also explore the jaw movements to ensure articulator synergy among lips and jaw. Secondly, for realistic video generation, a Video Consistency Network (VCN) is proposed conditioned on the predicted landmarks. In VCN, the optical flow is adopted to model the temporal continuity between frames to ensure inter-frame consistency. Meanwhile, a mouth generation branch is proposed to enhance mouth texture and the corresponding mouth mask is employed to ensure intra-frame consistency between the mouth area and the others. Extensive experiments demonstrate that our approach exhibits excellent superiority on lip-sync and can generate photo-realistic facial animations. Project is available at http://imcc.ustc.edu.cn/project/tfgen/. Lingyun Yu 0002, Hongtao Xie 0001, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Semantic-guided Reinforced Region Embedding for Generalized Zero-Shot LearningabstractGeneralized zero-shot Learning (GZSL) aims to recognize images from either seen or unseen domain, mainly by learning a joint embedding space to associate image features with the corresponding category descriptions. Recent methods have proved that localizing important object regions can effectively bridge the semantic-visual gap. However, these are all based on one-off visual localizers, lacking of interpretability and flexibility. In this paper, we propose a novel Semantic-guided Reinforced Region Embedding (SR2E) network that can localize important objects in the long-term interests to construct semantic-visual embedding space. SR2E consists of Reinforced Region Module (R2M) and Semantic Alignment Module (SAM). First, without the annotated bounding box as supervision, R2M encodes the semantic category guidance into the reward and punishment criteria to teach the localizer serialized region searching. Besides, R2M explores different action spaces during the serialized searching path to avoid local optimal localization, which thereby generates discriminative visual features with less redundancy. Second, SAM preserves the semantic relationship into visual features via semantic-visual alignment and designs a domain detector to alleviate the domain confusion. Experiments on four public benchmarks demonstrate that the proposed SR2E is an effective GZSL method with reinforced embedding space, which obtains averaged 6.1% improvements. Jiannan Ge, Hongtao Xie 0001, Shaobo Min, Yongdong Zhang 0001 |
AAAI | 2 |
| 2021 | Query-Memory Re-Aggregation for Weakly-supervised Video Object SegmentationabstractWeakly-supervised video object segmentation (WVOS) is an emerging video task that can track and segment the target given a simple bounding box label. However, existing WVOS methods are still unsatisfied in either speed or accuracy, since they only use the exemplar frame to guide the prediction while they neglect the reference from other frames. To solve the problem, we propose a novel Re-Aggregation based framework, which uses feature matching to efficiently find the target and capture the temporal dependencies from multiple frames to guide the segmentation. Based on a two-stage structure, our framework builds an information-symmetric matching process to achieve robust aggregation. In each stage, we design a Query-Memory Aggregation (QMA) module to gather features from the past frames and make bidirectional aggregation to adaptively weight the aggregated features, which relieves the latent misguidance in unidirectional aggregation. To further exploit the information from different aggregation stages, we propose a novel coarse-fine constraint by using the Cascaded Refinement Module (CRM) to combine the predictions from different stages and further boosts the performance. Experimental results on three benchmarks show that our method achieves the state-of-the-art performance in WVOS (e.g., an overall score of 84.7% on the DAVIS 2016 validation set). Fanchao Lin, Hongtao Xie 0001, Yan Li 0068, Yongdong Zhang 0001 |
AAAI | 2 |
| 2021 | Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text RecognitionabstractLinguistic knowledge is of great benefit to scene text recognition. However, how to effectively model linguistic rules in end-to-end deep networks remains a research challenge. In this paper, we argue that the limited capacity of language models comes from: 1) implicitly language modeling; 2) unidirectional feature representation; and 3) language model with noise input. Correspondingly, we propose an autonomous, bidirectional and iterative ABINet for scene text recognition. Firstly, the autonomous suggests to block gradient flow between vision and language models to enforce explicitly language modeling. Secondly, a novel bidirectional cloze network (BCN) as the language model is proposed based on bidirectional feature representation. Thirdly, we propose an execution manner of iterative correction for language model which can effectively alleviate the impact of noise input. Additionally, based on the ensemble of iterative predictions, we propose a self-training method which can learn from unlabeled images effectively. Extensive experiments indicate that ABINet has superiority on low-quality images and achieves state-of-the-art results on several mainstream benchmarks. Besides, the ABINet trained with ensemble self-training shows promising improvement in realizing human-level recognition. Code is available at https://github.com/FangShancheng/ABINet. Shancheng Fang, Hongtao Xie 0001, Yuxin Wang 0002, Zhendong Mao 0001, Yongdong Zhang 0001 |
CVPR | 2 |
| 2021 | Frequency-Aware Discriminative Feature Learning Supervised by Single-Center Loss for Face Forgery DetectionabstractFace forgery detection is raising ever-increasing interest in computer vision since facial manipulation technologies cause serious worries. Though recent works have reached sound achievements, there are still unignorable problems: a) learned features supervised by softmax loss are separable but not discriminative enough, since softmax loss does not explicitly encourage intra-class compactness and inter-class separability; and b) fixed filter banks and hand-crafted features are insufficient to capture forgery patterns of frequency from diverse inputs. To compensate for such limitations, a novel frequency-aware discriminative feature learning framework is proposed in this paper. Specifically, we design a novel single-center loss (SCL) that only compresses intra-class variations of natural faces while boosting inter-class differences in the embedding space. In such a case, the network can learn more discriminative features with less optimization difficulty. Besides, an adaptive frequency feature generation module is developed to mine frequency clues in a completely data-driven fashion. With the above two modules, the whole framework can learn more discriminative features in an end-to-end manner. Extensive experiments demonstrate the effectiveness and superiority of our framework on three versions of the FF++ dataset. Jiaming Li 0017, Hongtao Xie 0001, Zhongyuan Wang 0006, Yongdong Zhang 0001 |
CVPR | 2 |
| 2021 | From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkabstractIn this paper, we abandon the dominant complex language model and rethink the linguistic learning process in the scene text recognition. Different from previous methods considering the visual and linguistic information in two separate structures, we propose a Visual Language Modeling Network (VisionLAN), which views the visual and linguistic information as a union by directly enduing the vision model with language capability. Specially, we introduce the text recognition of character-wise occluded feature maps in the training stage. Such operation guides the vision model to use not only the visual texture of characters, but also the linguistic information in visual context for recognition when the visual cues are confused (e.g. occlusion, noise, etc.). As the linguistic information is acquired along with visual features without the need of extra language model, Vision-LAN significantly improves the speed by 39% and adaptively considers the linguistic information to enhance the visual features for accurate recognition. Furthermore, an Occlusion Scene Text (OST) dataset is proposed to evaluate the performance on the case of missing character-wise visual cues. The state of-the-art results on several benchmarks prove our effectiveness. Code and dataset are available at https://github.com/wangyuxin87/VisionLAN. Yuxin Wang 0002, Hongtao Xie 0001, Shancheng Fang, Jing Wang 0221, Shenggao Zhu, Yongdong Zhang 0001 |
ICCV | 2 |
| 2021 | Global Characteristic Guided Landmark Detection for Genu Valgus and Varus Diagnosis
Lingfeng Ma, Chuanbin Liu 0001, Hongtao Xie 0001 |
ICIG (2) | 5 |
| 2021 | Dynamic Inconsistency-aware DeepFake Video DetectionabstractThe spread of DeepFake videos causes a serious threat to information security, calling for effective detection methods to distinguish them. However, the performance of recent frame-based detection methods become limited due to their ignorance of the inter-frame inconsistency of fake videos. In this paper, we propose a novel Dynamic Inconsistency-aware Network to handle the inconsistent problem, which uses a Cross-Reference module (CRM) to capture both the global and local inter-frame inconsistencies. The CRM contains two parallel branches. The first branch takes faces from adjacent frames as input, and calculates a structure similarity map for a global inconsistency representation. The second branch only focuses on the inter-frame variation of independent critical regions, which captures the local inconsistency. To the best of our knowledge, this is the first work to totally use the inter-frame inconsistency information from the global and local perspectives. Compared with existing methods, our model provides a more accurate and robust detection on FaceForensics++, DFDC-preview and Celeb-DFv2 datasets. Ziheng Hu, Hongtao Xie 0001, Yuxin Wang 0002, Zhongyuan Wang 0006, Yongdong Zhang 0001 |
IJCAI | 2 |
| 2021 | Look Back Again: Dual Parallel Attention Network for Accurate and Robust Scene Text RecognitionabstractNowadays, it is a trend that using a parallel-decoupled encoder-decoder (PDED) framework in scene text recognition for its flexibility and efficiency. However, due to the inconsistent information content between queries and keys in the parallel positional attention module (PPAM) used in this kind of framework(queries: position information, keys: context and position information), visual misalignment tends to appear when confronting hard samples(e.g., blurred texts, irregular texts, or low-quality images). To tackle this issue, in this paper, we propose a dual parallel attention network (DPAN), in which a newly designed parallel context attention module (PCAM) is cascaded with the original PPAM, using linguistic contextual information to compensate for the information inconsistency between queries and keys. Specifically, in PCAM, we take the visual features from PPAM as inputs and present a bidirectional language model to enhance them with linguistic contexts to produce queries. In this way, we make the information content of the queries and keys consistent in PCAM, which helps to generate more precise visual glimpses to improve the entire PDED framework's accuracy and robustness. Experimental results verify the effectiveness of the proposed PCAM, showing the necessity of keeping the information consistency between queries and keys in the attention mechanism. On six benchmarks, including regular text and irregular text, the performance of DPAN surpasses the existing leading methods by large margins, achieving new state-of-the-art performance. The code is available on \urlhttps://github.com/Jackandrome/DPAN. Zilong Fu, Hongtao Xie 0001, Guoqing Jin, Junbo Guo |
ICMR | 2 |
| 2021 | End-to-end Boundary Exploration for Weakly-supervised Semantic SegmentationabstractIt is full of challenges for weakly supervised semantic segmentation (WSSS) acquiring the pixel-level object location with only image-level annotations. Especially, the single-stage methods learn image- and pixel-level labels simultaneously to avoid complicated multi-stage computations and sophisticated training procedures. In this paper, we argue that using a single model to accomplish image- and pixel-level classification will fall into the balance of multi-target and consequently weakens the recognition capability. Because the image-level task tends to learn position-independent features, but the pixel-level task tends to be position-sensitive. Hence, we propose an effective encoder-decoder framework to explore object boundaries and solve the above dilemma. The encoder and decoder learn position-independent and position-sensitive features independently during the end-to-end training. In addition, a global soft pooling is suggested to suppress background pixels' activation for the encoder training and further improve the class activation map (CAM) performance. The edge annotations for the decoder training are synthesized by the high confidence CAMs, which do not requires extra supervision. The extensive experiments on the Pascal VOC12 dataset demonstrate that our method achieves state-of-the-art compared to the end-to-end approaches. It gets 63.6% and 65.7% mIoU scores on val and test sets respectively. Shancheng Fang, Hongtao Xie 0001, Zhengjun Zha, Yue Hu 0002, Jianlong Tan |
ACM Multimedia | 3 |
| 2021 | Cluster and Scatter: A Multi-grained Active Semi-supervised Learning Framework for Scalable Person Re-identificationabstractActive learning has recently attracted increasing attention in the task of person re-identification, due to its unique scalability that not only maximally reduces the annotation cost but also retains the satisfying performance. Although some preliminary active learning methods have been explored in scalable person re-identification task, they have the following two problems: 1) the inefficiency in the selection process of image pairs due to the huge search space, and 2) the ineffectiveness caused by ignoring the impact of unlabeled data in model training. Considering that, we propose a Multi-grained Active Semi-Supervised learning framework, named MASS, to address the scalable person re-identification problem existing in the practical scenarios. Specifically, we firstly design a cluster-scatter procedure to alleviate the inefficiency problem, which consists of two components: cluster step and scatter step. The cluster step shrinks the search space into individual small clusters by a coarse-grained clustering method, and the subsequent scatter step further mines the hard distinguished image pairs from unlabelled set to purify the learned clusters by a novel centrality-based adaptive purification strategy. Afterward, we introduce a customized purification loss for the purified clustering, which utilizes the complementary information in both labeled and unlabeled data to optimize the model for solving the ineffectiveness problem. The cluster-scatter procedure and the model optimization are performed in an iterative fashion to achieve the promising performance while greatly reducing the annotation cost. Extensive experimental results have demonstrated that MASS can even achieve a competitive performance with fully supervised methods in the case of extremely less annotation requirements. Bingyu Hu, Zhengjun Zha, Jiawei Liu 0001, Xierong Zhu, Hongtao Xie 0001 |
ACM Multimedia | 5 |
| 2021 | TDI TextSpotter: Taking Data Imbalance into Account in Scene Text SpottingabstractRecent scene text spotters that integrate text detection module and recognition module have made significant progress. However, existing methods encounter two problems. 1). The data imbalance issue between text detection module and text recognition module limits the performance of text spotters. 2). The default left-to-right reading direction leads to errors in unconventional text spotting. In this paper, we propose a novel scene text spotter TDI to solve these problems. Firstly, in order to solve the data imbalance problem, a sample generation algorithm is proposed to generate plenty of samples online for training the text recognition module by using character features and character labels. Secondly, a weakly supervised character generation algorithm is designed to generate character-level labels from word-level labels for the sample generation algorithm and the training of the text detection module. Finally, in order to spot arbitrarily arranged text correctly, a direction perception module is proposed to perceive the reading direction of text instance. Experiments on several benchmarks show that these designs can significantly improve the performance of text spotter. Specifically, our method outperforms state-of-the-art methods on three public datasets in both text detection and end-to-end text recognition, which fully proves the effectiveness and robustness of our method. Yu Zhou 0016, Hongtao Xie 0001, Shancheng Fang, Jing Wang 0221, Zhengjun Zha, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2021 | Hierarchical multi-view context modelling for 3D object classification and retrieval
Anan Liu, Heyu Zhou, Weizhi Nie, Zhenguang Liu, Wu Liu 0005, Hongtao Xie 0001, Zhendong Mao 0001, Xuanya Li, Dan Song 0006 |
Inf. Sci. | 6 |
| 2021 | PRRNet: Pixel-Region relation network for face forgery detection
Zhihua Shang, Hongtao Xie 0001, Zhengjun Zha, Lingyun Yu 0002, Yan Li 0068, Yongdong Zhang 0001 |
Pattern Recognit. | 2 |
| 2021 | Self-Supervised Attention Mechanism for Pediatric Bone Age Assessment With Efficient Weak AnnotationabstractPediatric bone age assessment (BAA) is a common clinical practice to investigate endocrinology, genetic and growth disorders of children. Different specific bone parts are extracted as anatomical Regions of Interest (RoIs) during this task, since their morphological characters have important biological identification in skeletal maturity. Following this clinical prior knowledge, recently developed deep learning methods address BAA with an RoI-based attention mechanism, which segments or detects the discriminative RoIs for meticulous analysis. Great strides have been made, however, these methods strictly require large and precise RoIs annotations, which limits the real-world clinical value. To overcome the severe requirements on RoIs annotations, in this paper, we propose a novel self-supervised learning mechanism to effectively discover the informative RoIs without the need of extra knowledge and precise annotation-only image-level weak annotation is all we take. Our model, termed PEAR-Net for Part Extracting and Age Recognition Network, consists of one Part Extracting (PE) agent for discriminative RoIs discovering and one Age Recognition (AR) agent for age assessment. Without precise supervision, the PE agent is designed to discover and extract RoIs fully automatically. Then the proposed RoIs are fed into AR agent for feature learning and age recognition. Furthermore, we utilize the self-consistency of RoIs to optimize PE agent to understand the part relation and select the most useful RoIs. With this self-supervised design, the PE agent and AR agent can reinforce each other mutually. To the best of our knowledge, this is the first end-to-end bone age assessment method which can discover RoIs automatically with only image-level annotation. We conduct extensive experiments on the public RSNA 2017 dataset and achieve state-of-the-art performance with MAE 3.99 months. Project is available at http://imcc.ustc.edu.cn/project/ssambaa/. Chuanbin Liu 0001, Hongtao Xie 0001, Yongdong Zhang 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2021 | Hip Landmark Detection With Dependency Mining in Ultrasound ImageabstractDevelopmental dysplasia of the hip (DDH) is a common and serious disease in infants. Hip landmark detection plays a critical role in diagnosing the development of neonatal hip in the ultrasound image. However, the local confusion and the regional weakening make this task challenging. To solve these challenges, we explore the stable hip structure and the distinguishable local features to provide dependencies for hip landmark detection. In this paper, we propose a novel architecture named Dependency Mining ResNet (DM-ResNet), which investigates end-to-end dependency mining for more accurate and much faster hip landmark detection. First of all, we convert the landmark detection to the heatmap estimation by ResNet to build a strong baseline architecture for fast and accurate detection. Secondly, a dependency mining module is explored to mine the dependencies and leverage both the local and global information to decline the local confusion and strengthen the weakening region. Thirdly, we propose a simple but effective local voting algorithm (LVA) that seeks trade-off between long-range and short-range dependencies in the hip ultrasound image. Besides, a dataset with 2000 annotated hip ultrasound images is constructed in our work. It is the first public hip ultrasound dataset for open research. Experimental results show that our method achieves excellent precision in hip landmark detection (average point error of 0.719mm and successful detection rate within 1mm of 79.9%). Hongtao Xie 0001, Chuanbin Liu 0001, Xun Chen 0001, Yongdong Zhang 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2021 | A Mutually Attentive Co-Training Framework for Semi-Supervised RecognitionabstractSelf-training plays an important role in practical recognition applications where sufficient clean labels are unavailable. Existing methods focus on generating reliable pseudo labels to retrain a model, while ignoring the importance of improving model reliability to those inevitably mislabeled data. In this paper, we propose a novel Mutually Attentive Co-training Framework (MACF) that can effectively alleviate the negative impacts of incorrect labels on model retraining by exploring deep model disagreements. Specifically, MACF trains two symmetrical sub-networks that have the same input and are connected by several attention modules at different layers. Each attention module analyzes the inferred features from two sub-networks for the same input and feedback attention maps for them to indicate noisy gradients. This is realized by exploring the back-propagation process of incorrect labels at different layers to design attention modules. By multi-layer interception, the noisy gradients caused by incorrect labels can be effectively reduced for both sub-networks, leading to robust training to potential incorrect labels. In addition, a hierarchical distillation strategy is developed to improve the pseudo labels by aggregating the predictions from multi-models and data transformations. The experiments on six general benchmarks, including classification and biomedical segmentation, demonstrate that MACF is much robust to noisy labels than previous methods. Shaobo Min, Xuejin Chen, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | Domain-Oriented Semantic Embedding for Zero-Shot LearningabstractZero-Shot Learning (ZSL) targets to recognize images from new classes. Existing methods focus on learning a projection function to associate the visual features and category descriptions in the seen domain, which is directly transferred to the unseen domain. However, due to the inherent domain shift, a single shared projection cannot fully capture the domain difference and similarity, thereby making the unseen samples tend to be recognized as seen categories. In this paper, we propose a novel Domain-Oriented Semantic Embedding (DOSE) network that learns specific projections for different domains to better capture the domain characteristics for unbiased ZSL. Besides a domain-shared projection, DOSE learns two auxiliary domain-specific sub-projections to model the semantic-visual association in respective seen and unseen domains. Specifically, the domain-specific projections are learned in a cycle consistency way to capture domain characteristics, and a domain division constraint is developed to penalize the margin between two domain embeddings. Furthermore, to boost semantic-visual association, a semantic-visual dual attention module is designed to automatically remove trivial information in both visual and semantic embeddings under a co-guidance learning manner. Experiments on four public benchmarks prove that the proposed DOSE is robust to the domain shift problem in ZSL and obtains an averaged 5.6% improvement in terms of harmonic mean. Shaobo Min, Hantao Yao, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | R-Net: A Relationship Network for Efficient and Accurate Scene Text DetectionabstractThis paper introduces a novel bi-directional con-volutional framework to cope with the large-variance scale problem in scene text detection. Due to the lack of scale normalization in recent CNN-based methods, text instances with large-variance scale are activated inconsistently in feature maps, which makes it hard for CNN-based methods to accurately locate multi-size text instances. Thus, we propose the relationship network (R-Net) that maps multi-scale convolutional features to a scale-invariant space to obtain consistent activation of multi-size text instances. Firstly, we implement an FPN-like backbone with a Spatial Relationship Module (SPM) to extract multi-scale features with powerful spatial semantics. Then, a Scale Relationship Module (SRM) constructed on feature pyramid propagates contextual scale information in sequential features through a bi-directional convolutional operation. SRM supplements the multi-scale information in different feature maps to obtain consistent activation of multi-size text instances. Compared with previous approaches, R-Net effectively handles the large-variance scale problem without complicated post processing and complex hand-crafted hyperparameter setting. Extensive experiments conducted on several benchmarks verify that our R-Net obtains state-of-the-art performance on both accuracy and efficiency. More specifically, R-Net achieves an F-measure of 85.6% at 21.4 frames/s and an F-measure of 81.7% at 11.8 frames/s for ICDAR 2015 and MSRA-TD500 datasets respectively, which is the latest SOTA. The code is available on https://github.com/wangyuxin87/R-Net. Yuxin Wang 0002, Hongtao Xie 0001, Zhengjun Zha, Youliang Tian, Zilong Fu, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Filtration and Distillation: Enhancing Region Attention for Fine-Grained Visual CategorizationabstractDelicate attention of the discriminative regions plays a critical role in Fine-Grained Visual Categorization (FGVC). Unfortunately, most of the existing attention models perform poorly in FGVC, due to the pivotal limitations in discriminative regions proposing and region-based feature learning. 1) The discriminative regions are predominantly located based on the filter responses over the images, which can not be directly optimized with a performance metric. 2) Existing methods train the region-based feature extractor as a one-hot classification task individually, while neglecting the knowledge from the entire object. To address the above issues, in this paper, we propose a novel “Filtration and Distillation Learning” (FDL) model to enhance the region attention of discriminate parts for FGVC. Firstly, a Filtration Learning (FL) method is put forward for discriminative part regions proposing based on the matchability between proposing and predicting. Specifically, we utilize the proposing-predicting matchability as the performance metric of Region Proposal Network (RPN), thus enable a direct optimization of RPN to filtrate most discriminative regions. Go in detail, the object-based feature learning and region-based feature learning are formulated as “teacher” and “student”, which can furnish better supervision for region-based feature learning. Accordingly, our FDL can enhance the region attention effectively, and the overall framework can be trained end-to-end without neither object nor parts annotations. Extensive experiments verify that FDL yields state-of-the-art performance under the same backbone with the most competitive approaches on several FGVC tasks. Chuanbin Liu 0001, Hongtao Xie 0001, Zhengjun Zha, Lingfeng Ma, Lingyun Yu 0002, Yongdong Zhang 0001 |
AAAI | 2 |
| 2020 | CircleNet for Hip Landmark DetectionabstractLandmark detection plays a critical role in diagnosis of Developmental Dysplasia of the Hip (DDH). Heatmap and anchor-based object detection techniques could obtain reasonable results. However, they have limitations in both robustness and precision given the complexities and inhomogeneity of hip X-ray images. In this paper, we propose a much simpler and more efficient framework called CircleNet to improve the accuracy of landmark detection by predicting landmark and corresponding radius. Using the CircleNet, we not only constrain the relationship between landmarks but also integrate landmark detection and object detection into an end-to-end framework. In order to capture the effective information of the long-range dependency of landmarks in the DDH image, here we propose a new context modeling framework, named the Local Non-Local (LNL) block. The LNL block has the benefits of both non-local block and lightweight computation. We construct a professional DDH dataset for the first time and evaluate our CircleNet on it. The dataset has the largest number of DDH X-ray images in the world to our knowledge. Our results show that the CircleNet can achieve the state-of-the-art results for landmark detection on the dataset with a large margin of 1.8 average pixels compared to current methods. The dataset and source code will be publicly available. Hongtao Xie 0001, Chuanbin Liu 0001, Zhengjun Zha, Jun Sun 0016, Yongdong Zhang 0001 |
AAAI | 2 |
| 2020 | Curriculum Learning for Natural Language UnderstandingabstractWith the great success of pre-trained language models, the pretrain-finetune paradigm now becomes the undoubtedly dominant solution for natural language understanding (NLU) tasks.At the fine-tune stage, target task data is usually introduced in a completely random order and treated equally.However, examples in NLU tasks can vary greatly in difficulty, and similar to human learning procedure, language models can benefit from an easy-to-difficult curriculum.Based on this idea, we propose our Curriculum Learning approach.By reviewing the trainset in a crossed way, we are able to distinguish easy examples from difficult ones, and arrange a curriculum for language models.Without any manual model architecture design or use of external data, our Curriculum Learning approach obtains significant and universal performance improvements on a wide range of NLU tasks. Benfeng Xu, Licheng Zhang 0002, Zhendong Mao 0001, Quan Wang 0002, Hongtao Xie 0001, Yongdong Zhang 0001 |
ACL | 5 |
| 2020 | Graph Structured Network for Image-Text MatchingabstractImage-text matching has received growing interest since it bridges vision and language. The key challenge lies in how to learn correspondence between image and text. Existing works learn coarse correspondence based on object co-occurrence statistics, while failing to learn fine-grained phrase correspondence. In this paper, we present a novel Graph Structured Matching Network (GSMN) to learn fine-grained correspondence. The GSMN explicitly models object, relation and attribute as a structured phrase, which not only allows to learn correspondence of object, relation and attribute separately, but also benefits to learn fine-grained correspondence of structured phrase. This is achieved by node-level matching and structure-level matching. The node-level matching associates each node with its relevant nodes from another modality, where the node can be object, relation or attribute. The associated nodes then jointly infer fine-grained correspondence by fusing neighborhood associations at structure-level matching. Comprehensive experiments show that GSMN outperforms state-of-the-art methods on benchmarks, with relative Recall@1 improvements of nearly 7% and 2% on Flickr30K and MSCOCO, respectively. Code will be released at: https://github.com/CrossmodalGroup/GSMN. Zhendong Mao 0001, Tianzhu Zhang 0001, Hongtao Xie 0001, Bin Wang 0004, Yongdong Zhang 0001 |
CVPR | 4 |
| 2020 | Domain-Aware Visual Bias Eliminating for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning aims to recognize images from seen and unseen domains. Recent methods focus on learning a unified semantic-aligned visual representation to transfer knowledge between two domains, while ignoring the effect of semantic-free visual representation in alleviating the biased recognition problem. In this paper, we propose a novel Domain-aware Visual Bias Eliminating (DVBE) network that constructs two complementary visual representations, i.e., semantic-free and semantic-aligned, to treat seen and unseen domains separately. Specifically, we explore cross-attentive second-order visual statistics to compact the semantic-free representation, and design an adaptive margin Softmax to maximize inter-class divergences. Thus, the semantic-free representation becomes discriminative enough to not only predict seen class accurately but also filter out unseen images, i.e., domain detection, based on the predicted class entropy. For unseen images, we automatically search an optimal semantic-visual alignment architecture, rather than manual designs, to predict unseen classes. With accurate domain detection, the biased recognition problem towards the seen domain is significantly reduced. Experiments on five benchmarks for classification and segmentation show that DVBE outperforms existing methods by averaged 5.7% improvement. Shaobo Min, Hantao Yao, Hongtao Xie 0001, Chaoqun Wang 0011, Zhengjun Zha, Yongdong Zhang 0001 |
CVPR | 3 |
| 2020 | ContourNet: Taking a Further Step Toward Accurate Arbitrary-Shaped Scene Text DetectionabstractScene text detection has witnessed rapid development in recent years. However, there still exists two main challenges: 1) many methods suffer from false positives in their text representations; 2) the large scale variance of scene texts makes it hard for network to learn samples. In this paper, we propose the ContourNet, which effectively handles these two problems taking a further step toward accurate arbitrary-shaped text detection. At first, a scale-insensitive Adaptive Region Proposal Network (Adaptive-RPN) is proposed to generate text proposals by only focusing on the Intersection over Union (IoU) values between predicted and ground-truth bounding boxes. Then a novel Local Orthogonal Texture-aware Module (LOTM) models the local texture information of proposal features in two orthogonal directions and represents text region with a set of contour points. Considering that the strong unidirectional or weakly orthogonal activation is usually caused by the monotonous texture characteristic of false-positive patterns (e.g. streaks.), our method effectively suppresses these false positives by only outputting predictions with high response value in both orthogonal directions. This gives more accurate description of text regions. Extensive experiments on three challenging datasets (Total-Text, CTW1500 and ICDAR2015) verify that our method achieves the state-of-the-art performance. Code is available at https://github.com/wangyuxin87/ContourNet. Yuxin Wang 0002, Hongtao Xie 0001, Zhengjun Zha, Mengting Xing, Zilong Fu, Yongdong Zhang 0001 |
CVPR | 2 |
| 2020 | Real-World Automatic Makeup via Identity Preservation Makeup NetabstractThis paper focuses on the real-world automatic makeup problem. Given one non-makeup target image and one reference image, the automatic makeup is to generate one face image, which maintains the original identity with the makeup style in the reference image. In the real-world scenario, face makeup task demands a robust system against the environmental variants. The two main challenges in real-world face makeup could be summarized as follow: first, the background in real-world images is complicated. The previous methods are prone to change the style of background as well; second, the foreground faces are also easy to be affected. For instance, the ``heavy'' makeup may lose the discriminative information of the original identity. To address these two challenges, we introduce a new makeup model, called Identity Preservation Makeup Net (IPM-Net), which preserves not only the background but the critical patterns of the original identity. Specifically, we disentangle the face images to two different information codes, i.e., identity content code and makeup style code. When inference, we only need to change the makeup style code to generate various makeup images of the target person. In the experiment, we show the proposed method achieves not only better accuracy in both realism (FID) and diversity (LPIPS) in the test set, but also works well on the real-world images collected from the Internet. Zhikun Huang, Zhedong Zheng, Chenggang Yan 0001, Hongtao Xie 0001, Yaoqi Sun, Jiyong Zhang 0001 |
IJCAI | 4 |
| 2020 | Learning Rich Attention for Pediatric Bone Age Assessment
Chuanbin Liu 0001, Hongtao Xie 0001, Yunyan Yan, Zhendong Mao 0001, Yongdong Zhang 0001 |
MICCAI (1) | 2 |
| 2020 | Multi-Features Fusion and Decomposition for Age-Invariant Face RecognitionabstractAlthough the General Face Recognition (GFR) research achieves great success, Age-Invariant Face Recognition (AIFR) is still a challenging problem since facial appearance changing over time brings significant intra-class variations. The existing discriminative methods for the AIFR task mostly focus on decomposing the facial feature from a sigle image into age-related feature and age-independent feature for recognition, which suffer from the loss of facial identity information. To address this issue, in this work we propose a novel Multi-Features Fusion and Decomposition (MFFD) framework to learn more discriminative feature representations and alleviate the intra-class variations for AIFR. Specifically, we first sample multiple face images of different ages with the same identity as a face time series. Next, we combine feature decomposition with fusion based on the face time series to ensure that the final age-independent features effectively represent the identity information of the face and have stronger robustness against aging. Moreover, we also present two feature fusion methods and several different training strategies to explore the impact on the model. Extensive experiments on several cross-age datasets (CACD, CACD-VS) demonstrate the effectiveness of our proposed method. Besides, our method also shows comparable generalization performance on the well-known LFW dataset. Lixuan Meng, Chenggang Yan 0001, Jian Yin 0003, Wu Liu 0005, Hongtao Xie 0001, Liang Li 0003 |
ACM Multimedia | 6 |
| 2020 | March on Data Imperfections: Domain Division and Domain Generalization for Semantic SegmentationabstractSignificant progress has been made in semantic segmentation by deep neural networks, most of which concentrate on discriminative representation learning. However, model performances suffer from deterioration when the training process is optimized without awareness of data imperfections (e.g., data imbalance and label noise). In contrast to previous works, we present a novel model-agnostic training optimization algorithm which has two prominent components: Domain Division and Domain Generalization. Rather than sampling all pixels uniformly, an uncertainty-based Domain Division method is proposed to deal with data imbalance, which dynamically decomposes the pixels into meta-train and meta-test domains according to whether they lie near the classification boundary. The meta-train domain corresponds to highly-uncertain but more informative pixels and determines the current main update direction. Furthermore, to alleviate the degradation caused by label noise, we propose a Domain Generalization technique with a meta-optimization objective which ensures that update on the meta-train domain should generalize to the meta-test domain. Comprehensive experimental results on three public benchmarks across multi-modalities show that the proposed optimization algorithm is superior to other segmentation optimization methods and significantly outperforms conventional methods without introducing additional model parameters. Hongtao Xie 0001, Zhengjun Zha, Sun'ao Liu, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2020 | CRNet: A Center-aware Representation for Detecting Text of Arbitrary ShapesabstractExisting scene text detection methods achieve state-of-the-art performance by designing elaborate anchors or complex post-processing. Nonetheless, most methods still face the dilemma of detecting adjacent texts as one instance and long text with large character spacing as multiple fragments. To tackle these problems, we propose an anchor-free scene text detector leveraging Center-aware Representation to achieve accurate arbitrary-shaped scene text detection namely CRNet. Firstly, we propose a center-aware location algorithm to explicitly learn center regions and center points of text instances, which is able to separate adjacent text instances effectively. Then, a multi-scale context extraction module capable of extracting local context, long-range dependencies and global context adaptively is designed to effectively perceive long text with large character spacing. Finally, a low-level features enhancement block is introduced to enhance the geometric information of text. Extensive experiments conducted on several benchmarks including SCUT-CTW1500, Total-Text, ICDAR2015, ICDAR2017 MLT, and MSRA-TD500 demonstrate the effectiveness of our method. Specifically, without any anchor and complicated post-processing, our CRNet achieves 84.2% and 85.1% on CTW1500 and MSRA-TD500 in F-measure, outperforming all state-of-the-art anchor-based and anchor-free methods. Yu Zhou 0016, Hongtao Xie 0001, Shancheng Fang, Yan Li 0068, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2020 | Law Is Order: Protecting Multimedia Network Transmission by Game Theory and Mechanism Design
Chuanbin Liu 0001, Youliang Tian, Hongtao Xie 0001 |
MMM (2) | 3 |
| 2020 | Improving Brain Tumor Segmentation with Dilated Pseudo-3D Convolution and Multi-direction Fusion
Sun'ao Liu, Hongtao Xie 0001 |
MMM (1) | 4 |
| 2020 | Hierarchical Granularity Transfer LearningabstractIn the real world, object categories usually have a hierarchical granularity tree. Nowadays, most researchers focus on recognizing categories in a specific granularity, \emph{e.g.,} basic-level or sub(ordinate)-level. Compared with basic-level categories, the sub-level categories provide more valuable information, but its training annotations are harder to acquire. Therefore, an attractive problem is how to transfer the knowledge learned from basic-level annotations to sub-level recognition. In this paper, we introduce a new task, named Hierarchical Granularity Transfer Learning (HGTL), to recognize sub-level categories with basic-level annotations and semantic descriptions for hierarchical categories. Different from other recognition tasks, HGTL has a serious granularity gap,~\emph{i.e.,} the two granularities share an image space but have different category domains, which impede the knowledge transfer. To this end, we propose a novel Bi-granularity Semantic Preserving Network (BigSPN) to bridge the granularity gap for robust knowledge transfer. Explicitly, BigSPN constructs specific visual encoders for different granularities, which are aligned with a shared semantic interpreter via a novel subordinate entropy loss. Experiments on three benchmarks with hierarchical granularities show that BigSPN is an effective framework for Hierarchical Granularity Transfer Learning. Shaobo Min, Hongtao Xie 0001, Hantao Yao, Xuran Deng, Zhengjun Zha, Yongdong Zhang 0001 |
NeurIPS | 2 |
| 2020 | Global context and boundary structure-guided network for cross-modal organ segmentation
Hongtao Xie 0001, Yongdong Zhang 0001 |
Inf. Process. Manag. | 2 |
| 2020 | Multi-Objective Matrix Normalization for Fine-Grained Visual RecognitionabstractBilinear pooling achieves great success in fine-grained visual recognition (FGVC). Recent methods have shown that the matrix power normalization can stabilize the second-order information in bilinear features, but some problems, e.g., redundant information and over-fitting, remain to be resolved. In this paper, we propose an efficient Multi-Objective Matrix Normalization (MOMN) method that can simultaneously normalize a bilinear representation in terms of square-root, low-rank, and sparsity. These three regularizers can not only stabilize the second-order information, but also compact the bilinear features and promote model generalization. In MOMN, a core challenge is how to jointly optimize three non-smooth regularizers of different convex properties. To this end, MOMN first formulates them into an augmented Lagrange formula with approximated regularizer constraints. Then, auxiliary variables are introduced to relax different constraints, which allow each regularizer to be solved alternately. Finally, several updating strategies based on gradient descent are designed to obtain consistent convergence and efficient implementation. Consequently, MOMN is implemented with only matrix multiplication, which is well-compatible with GPU acceleration, and the normalized bilinear features are stabilized and discriminative. Experiments on five public benchmarks for FGVC demonstrate that the proposed MOMN is superior to existing normalization-based methods in terms of both accuracy and efficiency. The code is available: https://github.com/mboboGO/MOMN. Shaobo Min, Hantao Yao, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Mining Spatial-Temporal Similarity for Visual TrackingabstractCorrelation filter (CF) is a critical technique to improve accuracy and speed in the field of visual object tracking. Despite being studied extensively, most existing CF methods suffer from failing to make the most of the inherent spatial-temporal prior of videos. To address this limitation, as consecutive frames are eminently resemble in most videos, we investigate a novel scheme to predict targets' future state by exploiting previous observations. Specifically, in this paper, we propose a prediction based CF tracking framework by learning the spatial-temporal similarity of consecutive frames for sample managing, template regularization, and training response pre-weighting. We model the learning problem theoretically as a novel objective and provide effective optimization algorithms to solve the learning task. In addition, we implement two CF trackers with different features. Extensive experiments are conducted on three popular benchmarks to validate our scheme. The encouraging results demonstrate that the proposed scheme can significantly boost the accuracy of CF tracking, and the two trackers achieve competitive performances against state-of-the-art trackers. We finally present a comprehensive analysis on the efficacy of our proposed method and the efficiency of our trackers to facilitate real-world visual tracking applications. Yu Zhang 0102, Xingyu Gao 0001, Zhenyu Chen 0003, Huicai Zhong, Hongtao Xie 0001, Chenggang Yan 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Misshapen Pelvis Landmark Detection With Local-Global Feature Learning for Diagnosing Developmental Dysplasia of the HipabstractDevelopmental dysplasia of the hip (DDH) is one of the most common orthopedic disorders in infants and young children. Accurately detecting and identifying the misshapen anatomical landmarks plays a crucial role in the diagnosis of DDH. However, the diversity during the calcification and the deformity due to the dislocation lead it a difficult task to detect the misshapen pelvis landmarks for both human expert and computer. Generally, the anatomical landmarks exhibit stable morphological features in part regions and rigid structural features in long ranges, which can be strong identification for the landmarks. In this paper, we investigate the local morphological features and global structural features for the misshapen landmark detection with a novel Pyramid Non-local UNet (PN-UNet). Firstly, we mine the local morphological features with a series of convolutional neural network (CNN) stacks, and convert the detection of a landmark to the segmentation of the landmark's local neighborhood by UNet. Secondly, a non-local module is employed to capture the global structural features with high-level structural knowledge. With the end-to-end and accurate detection of pelvis landmarks, we realize a fully automatic and highly reliable diagnosis of DDH. In addition, a dataset with 10,000 pelvis X-ray images is constructed in our work. It is the first public dataset for diagnosing DDH and has been already released for open research. To the best of our knowledge, this is the first attempt to apply deep learning method in the diagnosis of DDH. Experimental results show that our approach achieves an excellent precision in landmark detection (average point to point error of 0.9286mm) and illness diagnosis over human experts. Project is available at http://imcc.ustc.edu.cn/project/ddh/. Chuanbin Liu 0001, Hongtao Xie 0001, Zhendong Mao 0001, Jun Sun 0016, Yongdong Zhang 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Bidirectional Attention-Recognition Model for Fine-Grained Object ClassificationabstractFine-grained object classification (FGOC) is a challenging research topic in multimedia computing with machine learning, which faces two pivotal conundrums: focusing attention on the discriminate part regions, and then processing recognition with the part-based features. Existing approaches generally adopt a unidirectional two-step structure, that first locate the discriminate parts and then recognize the part-based features. However, they neglect the truth that part localization and feature recognition can be reinforced in a bidirectional process. In this paper, we propose a novel bidirectional attention-recognition model (BARM) to actualize the bidirectional reinforcement for FGOC. The proposed BARM consists of one attention agent for discriminate part regions proposing and one recognition agent for feature extraction and recognition. Meanwhile, a feedback flow is creatively established to optimize the attention agent directly by recognition agent. Therefore, in BARM the attention agent and the recognition agent can reinforce each other in a bidirectional way and the overall framework can be trained end-to-end without neither object nor parts annotations. Moreover, a novel Multiple Random Erasing data augmentation is proposed, and it exhibits impressive pertinency and superiority for FGOC. Conducted on several extensive FGOC benchmarks, BARM outperforms the present state-of-the-art methods in classification accuracy. Furthermore, BARM exhibits a clear interpretability and keeps consistent with the human perception in visualization experiments. Chuanbin Liu 0001, Hongtao Xie 0001, Zhengjun Zha, Lingyun Yu 0002, Zhineng Chen, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Robust Deep Co-Saliency Detection With Group Semantic and Pyramid AttentionabstractHigh-level semantic knowledge in addition to low-level visual cues is essentially crucial for co-saliency detection. This article proposes a novel end-to-end deep learning approach for robust co-saliency detection by simultaneously learning high-level groupwise semantic representation as well as deep visual features of a given image group. The interimage interaction at the semantic level and the complementarity between the group semantics and visual features are exploited to boost the inferring capability of co-salient regions. Specifically, the proposed approach consists of a co-category learning branch and a co-saliency detection branch. While the former is proposed to learn a groupwise semantic vector using co-category association of an image group as supervision, the latter is to infer precise co-salient maps based on the ensemble of group-semantic knowledge and deep visual cues. The group-semantic vector is used to augment visual features at multiple scales and acts as a top-down semantic guidance for boosting the bottom-up inference of co-saliency. Moreover, we develop a pyramidal attention (PA) module that endows the network with the capability of concentrating on important image patches and suppressing distractions. The co-category learning and co-saliency detection branches are jointly optimized in a multitask learning manner, further improving the robustness of the approach. We construct a new large-scale co-saliency data set COCO-SEG to facilitate research of the co-saliency detection. Extensive experimental results on COCO-SEG and a widely used benchmark Cosal2015 have demonstrated the superiority of the proposed approach compared with state-of-the-art methods. Zhengjun Zha, Chong Wang 0023, Dong Liu 0002, Hongtao Xie 0001, Yongdong Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Robust Deep Co-Saliency Detection with Group SemanticabstractHigh-level semantic knowledge in addition to low-level visual cues is essentially crucial for co-saliency detection. This paper proposes a novel end-to-end deep learning approach for robust co-saliency detection by simultaneously learning highlevel group-wise semantic representation as well as deep visual features of a given image group. The inter-image interaction at semantic-level as well as the complementarity between group semantics and visual features are exploited to boost the inferring of co-salient regions. Specifically, the proposed approach consists of a co-category learning branch and a co-saliency detection branch. While the former is proposed to learn group-wise semantic vector using co-category association of an image group as supervision, the latter is to infer precise co-salient maps based on the ensemble of group semantic knowledge and deep visual cues. The group semantic vector is broadcasted to each spatial location of multi-scale visual feature maps and is used as a top-down semantic guidance for boosting the bottom-up inferring of co-saliency. The co-category learning and co-saliency detection branches are jointly optimized in a multi-task learning manner, further improving the robustness of the approach. Moreover, we construct a new large-scale co-saliency dataset COCO-SEG to facilitate research of co-saliency detection. Extensive experimental results on COCO-SEG and a widely used benchmark Cosal2015 have demonstrated the superiority of the proposed approach as compared to the state-of-the-art methods. Chong Wang 0023, Zhengjun Zha, Dong Liu 0002, Hongtao Xie 0001 |
AAAI | 4 |
| 2019 | Accurate Segmentation of Synaptic Cleft with Contour Growing Concatenated with a ConvnetabstractSynaptic cleft is an important area for neuroscientists to analyze the macromolecular complexes related to neurotransmitter transmission. However, the large amount of noise and low signal-to-noise ratio in raw electron micrographs make it challenging to extract this region automatically. In this paper, we propose a simple but effective framework to automatically extract accurate boundaries of synaptic cleft regions. Our approach concatenates a novel contour growing algorithm to a fully convolutional network (FCN), so that it takes both advantages of large receptive field of FCNs and fine-level localization of contour evolution. The contour growing algorithm is based on the flexible evolving tension and synchronous growing controlling to localize the opening contour of clef region. With consideration of both global localization and local segmentation, our approach is more robust to noisy electron micrographs and outperforms all existing single-model FCNs on accurate segmentation of synaptic clefts. Shaobo Min, Xuejin Chen, Hongtao Xie 0001, Zhengjun Zha, Guoqiang Bi, Feng Wu 0001, Yongdong Zhang 0001 |
ICIP | 3 |
| 2019 | Semantic-Embedding and Shape-Aware U-Net for Ultrasound Eyeball SegmentationabstractSegmentation of eyeball region from ultrasound images is a new research direction for the diagnosis of ophthalmic diseases. Despite the advantages of convenience and cheapness, ultrasound images bring more noise and fuzzy contour compared with other medical images. Existing methods fail to give a segmentation with reasonable eyeball shape, especially when the contour is ambiguous. In this paper, we propose a novel framework based on convolutional neural network, named semantic-embedding and shape-aware U-Net, to deal with the segmentation in blurred images. A signed distance field is used as label instead of the traditional binary mask label to add shape prior in network. The applying of semantic embedding modules fuses semantic information between different stages of the network. Experimental results show that our method improves the ability to segment image with blurred edges and outperforms existing methods in the accuracy of segmentation. Fanchao Lin, Chuanbin Liu 0001, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001 |
ICME | 3 |
| 2019 | MLTS: A Multi-Language Scene Text SpotterabstractScene text detection and recognition are popular research topics in computer vision due to its various applications such as autonomous driving, blind assistance and text translation. However, many methods currently can only detect or recog-nize the text of one language. In scene text images, we can often see text in multi-language appearing on the same image. However, there is no valid model for multi-language text spotting. In this paper, an end-to-end method for multi-language scene text detection, recognition and script identification is proposed. The method, called MLTS, is an abbreviation of a Multi-Language Scene Text Spotter. By designing a special backbone for text and combining two different kinds of attention. MLTS achieves state-of-the-art performance for both joint localization and script identification in natural images and in cropped word script identification, the precision, recall and F-measure are 0.7145, 0.6583 and 0.6852 respectively, while the corresponding values of the best existing methods are 0.5759, 0.6207, 0.5974 respectively. Additionally, our MLTS achieves comparable performance on ICDAR2013 and ICDAR2015, which proves the effectiveness of the model. Yu Zhou 0016, Shancheng Fang, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001 |
ICME | 3 |
| 2019 | Semi-supervised User Profiling with Heterogeneous Graph Attention NetworksabstractAiming to represent user characteristics and personal interests, the task of user profiling is playing an increasingly important role for many real-world applications, e.g., e-commerce and social networks platforms. By exploiting the data like texts and user behaviors, most existing solutions address user profiling as a classification task, where each user is formulated as an individual data instance. Nevertheless, a user's profile is not only reflected from her/his affiliated data, but also can be inferred from other users, e.g., the users that have similar co-purchase behaviors in e-commerce, the friends in social networks, etc. In this paper, we approach user profiling in a semi-supervised manner, developing a generic solution based on heterogeneous graph learning. On the graph, nodes represent the entities of interest (e.g., users, items, attributes of items, etc.), and edges represent the interactions between entities. Our heterogeneous graph attention networks (HGAT) method learns the representation for each entity by accounting for the graph structure, and exploits the attention mechanism to discriminate the importance of each neighbor entity. Through such a learning scheme, HGAT can leverage both unsupervised information and limited labels of users to build the predictor. Extensive experiments on a real-world e-commerce dataset verify the effectiveness and rationality of our HGAT for user profiling. Weijian Chen 0001, Yulong Gu, Zhaochun Ren, Xiangnan He 0001, Hongtao Xie 0001, Dawei Yin 0001, Yongdong Zhang 0001 |
IJCAI | 5 |
| 2019 | Learning to Draw Text in Natural Images with Conditional Adversarial NetworksabstractIn this work, we propose an entirely learning-based method to automatically synthesize text sequence in natural images leveraging conditional adversarial networks. As vanilla GANs are clumsy to capture structural text patterns, directly employing GANs for text image synthesis typically results in illegible images. Therefore, we design a two-stage architecture to generate repeated characters in images. Firstly, a character generator attempts to synthesize local character appearance independently, so that the legible characters in sequence can be obtained. To achieve style consistency of characters, we propose a novel style loss based on variance-minimization. Secondly, we design a pixel-manipulation word generator constrained by self-regularization, which learns to convert local characters to plausible word image. Experiments on SVHN dataset and ICDAR, IIIT5K datasets demonstrate our method is able to synthesize visually appealing text images. Besides, we also show the high-quality images synthesized by our method can be used to boost the performance of a scene text recognition algorithm. Shancheng Fang, Hongtao Xie 0001, Jianlong Tan, Yongdong Zhang 0001 |
IJCAI | 2 |
| 2019 | DSRN: A Deep Scale Relationship Network for Scene Text DetectionabstractNowadays, scene text detection has become increasingly important and popular. However, the large variance of text scale remains the main challenge and limits the detection performance in most previous methods. To address this problem, we propose an end-to-end architecture called Deep Scale Relationship Network (DSRN) to map multi-scale convolution features onto a scale invariant space to obtain uniform activation of multi-size text instances. Firstly, we develop a Scale-transfer module to transfer the multi-scale feature maps to a unified dimension. Due to the heterogeneity of features, simply concatenating feature maps with multi-scale information would limit the detection performance. Thus we propose a Scale Relationship module to aggregate the multi-scale information through bi-directional convolution operations. Finally, to further reduce the miss-detected instances, a novel Recall Loss is proposed to force the network to concern more about miss-detected text instances by up-weighting poor-classified examples. Compared with previous approaches, DSRN efficiently handles the large-variance scale problem without complex hand-crafted hyperparameter settings (e.g. scale of default boxes) and complicated post processing. On standard datasets including ICDAR2015 and MSRA-TD500, the proposed algorithm achieves the state-of-art performance with impressive speed (8.8 FPS on ICDAR2015 and 13.3 FPS on MSRA-TD500). Yuxin Wang 0002, Hongtao Xie 0001, Zilong Fu, Yongdong Zhang 0001 |
IJCAI | 2 |
| 2019 | Extract Bone Parts Without Human Prior: End-to-end Convolutional Neural Network for Pediatric Bone Age Assessment
Chuanbin Liu 0001, Hongtao Xie 0001, Zhengjun Zha, Fanchao Lin, Yongdong Zhang 0001 |
MICCAI (6) | 2 |
| 2019 | Misshapen Pelvis Landmark Detection by Spatial Local Correlation Mining for Diagnosing Developmental Dysplasia of the Hip
Chuanbin Liu 0001, Hongtao Xie 0001, Jun Sun 0016, Yongdong Zhang 0001 |
MICCAI (6) | 2 |
| 2019 | Deep Cascaded Attention Network for Multi-task Brain Tumor Segmentation
Hongtao Xie 0001, Chuandong Cheng, Chaoshi Niu, Yongdong Zhang 0001 |
MICCAI (3) | 2 |
| 2019 | ACE-Net: Biomedical Image Segmentation with Augmented Contracting and Expansive Paths
Yanhao Zhu, Zhineng Chen, Hongtao Xie 0001, Wenming Guo, Yongdong Zhang 0001 |
MICCAI (1) | 4 |
| 2019 | Domain-Specific Embedding Network for Zero-Shot RecognitionabstractZero-Shot Learning (ZSL) seeks to recognize a sample from either seen or unseen domain by projecting the image data and semantic labels into a joint embedding space. However, most existing methods directly adapt a well-trained projection from one domain to another, thereby ignoring the serious bias problem caused by domain differences. To address this issue, we propose a novel Domain-Specific Embedding Network (DSEN) that can apply specific projections to different domains for unbiased embedding, as well as several domain constraints. In contrast to previous methods, the DSEN decomposes the domain-shared projection function into one domain-invariant and two domain-specific sub-functions to explore the similarities and differences between two domains. To prevent the two specific projections from breaking the semantic relationship, a semantic reconstruction constraint is proposed by applying the same decoder function to them in a cycle consistency way. Furthermore, a domain division constraint is developed to directly penalize the margin between real and pseudo image features in respective seen and unseen domains, which can enlarge the inter-domain difference of visual features. Extensive experiments on four public benchmarks demonstrate the effectiveness of DSEN with an average of $9.2%$ improvement in terms of harmonic mean. The code is available in \urlhttps://github.com/mboboGO/DSEN-for-GZSL. Shaobo Min, Hantao Yao, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001 |
ACM Multimedia | 3 |
| 2019 | Question-Aware Tube-Switch Network for Video Question AnsweringabstractVideo Question & Answering (VideoQA), a task to answer questions in videos, involves rich spatio-temporal content (e.g., appearance and motion) and requires multi-hop reasoning process. However, existing methods usually deal with appearance and motion separately and fail to synchronize the attentions on appearance and motion features, neglecting two key properties of video QA: (1) appearance and motion features are usually concomitant and complementary to each other at time slice level. Some questions rely on joint representations of both kinds of features at some point in the video; (2) appearance and motion have different importance in multi-step reasoning. In this paper, we propose a novel Question- Aware Tube-Switch Network (TSN) for video question answering which contains (1) a Mix module to synchronously combine the appearance and motion representation at time slice level, achieving fine-grained temporal alignment and correspondence between appearance and motion at every time slice and (2) a Switch mod- ule to adaptively choose appearance or motion tube as primary at each reasoning step, guiding the multi-hop reasoning process. To end-to-end train TSN, we utilize the Gumbel-Softmax strategy to account for the discrete tube-switch process. Extensive experimental results on two benchmarks: MSVD-QA and MSRVTT-QA, have demonstrated that the proposed TSN consistently outperforms state-of-the-art on all metrics. Tianhao Yang, Zhengjun Zha, Hongtao Xie 0001, Meng Wang 0001, Hanwang Zhang |
ACM Multimedia | 3 |
| 2019 | Adaptive Bilinear Pooling for Fine-grained Representation LearningabstractFine-grained representation learning targets to generate discriminative description for fine-grained visual objects. Recently, the bilinear feature interaction has been proved effective in generating powerful high-order representation with spatially invariant information. However, the existing methods apply a fixed feature interaction strategy to all samples, which ignore the image and region heterogeneity in a dataset. To this end, we propose a generalized feature interaction method, named Adaptive Bilinear Pooling (ABP), which can adaptively infer a suitable pooling strategy for a given sample based on image content. Specifically, ABP consists of two learning strategies: p-order learning (P-net) and spatial attention learning (S-net). The p-order learning predicts an optimal exponential coefficient rather than a fixed order number to extract moderate visual information from an image. The spatial attention learning aims to infer a weighted score that measures the importance of each local region, which can compact the image representations. To make ABP compatible with kernelized bilinear feature interaction, a crossed two-branch structure is utilized to combine the P-net and S-net. This structure can facilitate complementary information exchange between two different visual branches. The experiments on three widely used benchmarks, including fine-grained object classification and action recognition, demonstrate the effectiveness of the proposed method. Shaobo Min, Hongtao Xie 0001, Youliang Tian, Hantao Yao, Yongdong Zhang 0001 |
MMAsia | 2 |
| 2019 | WaveCSN: Cascade Segmentation Network for Hip Landmark DetectionabstractLandmark detection in hip X-ray images plays a critical role in diagnosis of Developmental Dysplasia of the Hip (DDH) and surgeries of Total Hip Arthroplasty (THA). Regression and heatmap techniques of convolution network could obtain reasonable results. However, they have limitations in either robustness or precision given the complexities and intensity inhomogeneities of hip X-ray images. In this paper, we propose a Wave-like Cascade Segmentation Network (WaveCSN) to improve the accuracy of landmark detection by transforming landmark detection into area segmentation. The WaveCSN consists of three basic sub-networks and each sub-network is composed of a U-net module, an indicate module and a max-MSER module. The U-net undertakes the task to generate masks, and the indicate module is trained to distinguish the masks and ground truth. The U-net and indicate module are trained in turns, in which process the generated masks are supervised to be more and more alike to the ground truth. The max-MSER module ensures landmarks can be extracted from the generated masks precisely. We present two professional datasets (DDH and THA) for the first time and evaluate the WaveCSN on them. Our results prove that the WaveCSN can improve 2.66 and 4.11 pixels at least on these two datasets compared to other methods, and achieves the state-of-the-art for landmark detection in hip X-ray images. Hongtao Xie 0001, Fanchao Lin, Jun Sun 0016, Yongdong Zhang 0001 |
MMAsia | 2 |
| 2019 | Adaptive Alignment Network for Person Re-identification
Xierong Zhu, Jiawei Liu 0001, Hongtao Xie 0001, Zhengjun Zha |
MMM (2) | 3 |
| 2019 | Name-face association with web facial image supervision
Zhineng Chen, Wei Zhang 0031, Hongtao Xie 0001, Xiaoyan Gu 0001 |
Multim. Syst. | 4 |
| 2019 | Supervised deep hashing for image content security
Yanping Ma, Dongbao Yang, Hongtao Xie 0001, Jian Yin 0003 |
Multim. Tools Appl. | 3 |
| 2019 | Automated pulmonary nodule detection in CT images using deep convolutional neural networks
Hongtao Xie 0001, Dongbao Yang, Nannan Sun, Zhineng Chen, Yongdong Zhang 0001 |
Pattern Recognit. | 1 |
| 2019 | Double-Bit Quantization and Index Hashing for Nearest Neighbor SearchabstractAs binary code is storage efficient and fast to compute, it has become a trend to compact real-valued data to binary codes for the nearest neighbors (NN) search in a large-scale database. However, the use of binary code for the NN search leads to low retrieval accuracy. To increase the discriminability of the binary codes of existing hash functions, in this paper, we propose a framework of double-bit quantization and index hashing for an effective NN search. The main contributions of our framework are: first, a novel double-bit quantization (DBQ) is designed to assign more bits to each dimension for higher retrieval accuracy; second, a double-bit index hashing (DBIH) is presented to efficiently index binary codes generated by DBQ; and third, a weighted distance measurement for DBQ binary codes is put forward to re-rank the search results from DBIH. The empirical results on three benchmark databases demonstrate the superiority of our framework over existing approaches in terms of both retrieval accuracy and query efficiency. Specifically, we observe an absolute improvement on precision of 10%-25% in most cases and the query speed increases over 30 times compared to traditional binary embedding methods and linear scan, respectively. Hongtao Xie 0001, Zhendong Mao 0001, Yongdong Zhang 0001, Chenggang Yan 0001, Zhineng Chen |
IEEE Trans. Multim. | 1 |
| 2019 | Convolutional Attention Networks for Scene Text RecognitionabstractIn this article, we present Convoluitional Attention Networks (CAN) for unconstrained scene text recognition. Recent dominant approaches for scene text recognition are mainly based on Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN), where the CNN encodes images and the RNN generates character sequences. Our CAN is different from these methods; our CAN is completely built on CNN and includes an attention mechanism. The distinctive characteristics of our method include (i) CAN follows encoder-decoder architecture, in which the encoder is a deep two-dimensional CNN and the decoder is a one-dimensional CNN; (ii) the attention mechanism is applied in every convolutional layer of the decoder, and we propose a novel spatial attention method using average pooling; and (iii) position embeddings are equipped in both a spatial encoder and a sequence decoder to give our networks a sense of location. We conduct experiments on standard datasets for scene text recognition, including Street View Text , IIIT5K, and ICDAR datasets. The experimental results validate the effectiveness of different components and show that our convolutional-based method achieves state-of-the-art or competitive performance over prior works, even without the use of RNN. Hongtao Xie 0001, Shancheng Fang, Zhengjun Zha, Yating Yang, Yan Li 0068, Yongdong Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Deep Convolutional Nets for Pulmonary Nodule Detection and Classification
Nannan Sun, Dongbao Yang, Shancheng Fang, Hongtao Xie 0001 |
KSEM (2) | 4 |
| 2018 | Attention and Language Ensemble for Scene Text Recognition with Convolutional Sequence ModelingabstractRecent dominant approaches for scene text recognition are mainly based on convolutional neural network (CNN) and recurrent neural network (RNN), where the CNN processes images and the RNN generates character sequences. Different from these methods, we propose an attention-based architecture1 which is completely based on CNNs. The distinctive characteristics of our method include: (1) the method follows encoder-decoder architecture, in which the encoder is a two-dimensional residual CNN and the decoder is a deep one-dimensional CNN. (2) An attention module that captures visual cues, and a language module that models linguistic rules are designed equally in the decoder. Therefore the attention and language can be viewed as an ensemble to boost predictions jointly. (3) Instead of using a single loss from language aspect, multiple losses from attention and language are accumulated for training the networks in an end-to-end way. We conduct experiments on standard datasets for scene text recognition, including Street View Text, IIIT5K and ICDAR datasets. The experimental results show our CNN-based method has achieved state-of-the-art performance on several benchmark datasets, even without the use of RNN. Shancheng Fang, Hongtao Xie 0001, Zhengjun Zha, Nannan Sun, Jianlong Tan, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2018 | CA3Net: Contextual-Attentional Attribute-Appearance Network for Person Re-IdentificationabstractPerson re-identification aims to identify the same pedestrian across non-overlapping camera views. Deep learning techniques have been applied for person re-identification recently, towards learning representation of pedestrian appearance. This paper presents a novel Contextual-Attentional Attribute-Appearance Network ($\rm CA^3Net$) for person re-identification. The $\rm CA^3Net$ simultaneously exploits the complementarity between semantic attributes and visual appearance, the semantic context among attributes, visual attention on attributes as well as spatial dependencies among body parts, leading to discriminative and robust pedestrian representation. Specifically, an attribute network within $\rm CA^3Net$ is designed with an Attention-LSTM module. It concentrates the network on latent image regions related to each attribute as well as exploits the semantic context among attributes by a LSTM module. An appearance network is developed to learn appearance features from the full body, horizontal and vertical body parts of pedestrians with spatial dependencies among body parts. The $\rm CA^3Net$ jointly learns the attribute and appearance features in a multi-task learning manner, generating comprehensive representation of pedestrians. Extensive experiments on two challenging benchmarks, i.e., Market-1501 and DukeMTMC-reID datasets, have demonstrated the effectiveness of the proposed approach. Jiawei Liu 0001, Zhengjun Zha, Hongtao Xie 0001, Zhiwei Xiong, Yongdong Zhang 0001 |
ACM Multimedia | 3 |
| 2018 | Uyghur Text Localization with Fast Component Detection
Hongtao Xie 0001, Yue Hu 0002, Chenggang Yan 0001 |
MMM (1) | 2 |
| 2018 | Potential of Attention Mechanism for Classification of Optical Coherence Tomography ImagesabstractDeep neural network (DNN) can extract high-dimensional feature of images for computer vision tasks including Optical Coherence Tomography (OCT) images classification. However, OCT images are usually processed by DNN just like natural images, thus the performance of DNN is not satisfactory. We present an end-to-end DNN targeting OCT images classification. Considering the characteristic of OCT images, we introduce attention mechanism into classifier to extract more specific feature of OCT images. Our network demonstrates its capacity to enhance the features that represent the disease region. Our method achieves the state-of-the-art performance with average accuracy of 99.5% and F1-score of 0.995 on the OCT images dataset. Zhihua Shang, Zilong Fu, Chuanbin Liu 0001, Hongtao Xie 0001, Yongdong Zhang 0001 |
VCIP | 4 |
| 2018 | Effective Uyghur Language Text Detection in Complex Background Images for Traffic Prompt IdentificationabstractText detection in complex background images is a challenging task for intelligent vehicles. Actually, almost all the widely-used systems focus on commonly used languages while for some minority languages, such as the Uyghur language, text detection is paid less attention. In this paper, we propose an effective Uyghur language text detection system in complex background images. First, a new channel-enhanced maximally stable extremal regions (MSERs) algorithm is put forward to detect component candidates. Second, a two-layer filtering mechanism is designed to remove most non-character regions. Third, the remaining component regions are connected into short chains, and the short chains are extended by a novel extension algorithm to connect the missed MSERs. Finally, a two-layer chain elimination filter is proposed to prune the non-text chains. To evaluate the system, we build a new data set by various Uyghur texts with complex backgrounds. Extensive experimental comparisons show that our system is obviously effective for Uyghur language text detection in complex background images. The F-measure is 85%, which is much better than the state-of-the-art performance of 75.5%. Chenggang Yan 0001, Hongtao Xie 0001, Jian Yin 0003, Yongdong Zhang 0001, Qionghai Dai |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Supervised Hash Coding With Deep Neural Network for Environment Perception of Intelligent VehiclesabstractImage content analysis is an important surround perception modality of intelligent vehicles. In order to efficiently recognize the on-road environment based on image content analysis from the large-scale scene database, relevant images retrieval becomes one of the fundamental problems. To improve the efficiency of calculating similarities between images, hashing techniques have received increasing attentions. For most existing hash methods, the suboptimal binary codes are generated, as the hand-crafted feature representation is not optimally compatible with the binary codes. In this paper, a one-stage supervised deep hashing framework (SDHP) is proposed to learn high-quality binary codes. A deep convolutional neural network is implemented, and we enforce the learned codes to meet the following criterions: 1) similar images should be encoded into similar binary codes, and vice versa; 2) the quantization loss from Euclidean space to Hamming space should be minimized; and 3) the learned codes should be evenly distributed. The method is further extended into SDHP+ to improve the discriminative power of binary codes. Extensive experimental comparisons with state-of-the-art hashing algorithms are conducted on CIFAR-10 and NUS-WIDE, the MAP of SDHP reaches to 87.67% and 77.48% with 48 b, respectively, and the MAP of SDHP+ reaches to 91.16%, 81.08% with 12 b, 48 b on CIFAR-10 and NUS-WIDE, respectively. It illustrates that the proposed method can obviously improve the search accuracy. Chenggang Yan 0001, Hongtao Xie 0001, Dongbao Yang, Jian Yin 0003, Yongdong Zhang 0001, Qionghai Dai |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | A Fast Uyghur Text Detector for Complex Background ImagesabstractUyghur text localization in images with complex backgrounds is a challenging yet important task for many applications. Generally, Uyghur characters in images consist of strokes with uniform features, and they are distinct from backgrounds in color, intensity, and texture. Based on these differences, we propose a FASTroke keypoint extractor, which is fast and stroke-specific. Compared with the commonly used MSER detector, FASTroke produces less than twice the amount of components and recognizes at least 10% more characters. While the characters in a line usually have uniform features such as size, color, and stroke width, a component similarity based clustering is presented without component-level classification. It incurs no extra errors by incorporating a component-level classifier while the computing cost is drastically reduced. The experiments show that the proposed method can achieve the best performance on the UICBI-500 benchmark dataset. Chenggang Yan 0001, Hongtao Xie 0001, Zhengjun Zha, Xinhong Hao, Yongdong Zhang 0001, Qionghai Dai |
IEEE Trans. Multim. | 2 |
| 2017 | Double-bit quantization and weighting for nearest neighbor searchabstractBinary embedding is an effective way for nearest neighbor (NN) search as binary code is storage efficient and fast to compute. It tries to convert real-value signatures into binary codes while preserving similarity of the original data. However, it greatly decreases the discriminability of original signatures due to the huge loss of information. In this paper, we propose a novel method double-bit quantization and weighting (DBQW) to solve the problem by mapping each dimension to double-bit binary code and assigning different weights according to their spatial relationship. The proposed method is applicable to a wide variety of embedding techniques, such as SH, PCA-ITQ and PCA-RR. Experimental comparisons on two datasets show that DBQW for NN search can achieve remarkable improvements in query accuracy compared to original binary embedding methods. Hongtao Xie 0001, Zhendong Mao 0001, Chuan Zhou 0001 |
ICASSP | 2 |
| 2017 | CPMF: A collective pairwise matrix factorization model for upcoming event recommendationabstractDue to the rapid growth of event-based social networks (EBSNs), event recommendation which helps users find their preferred events has become a popular topic. Different from movies or books in conventional recommendation problem, events usually have recommendation lifetimes and almost all the events to be recommended are upcoming, which brings a severe cold start problem. To achieve better event recommendation performance, we formulates multiple interactions among users, events, groups and locations into an unified framework and propose a collective pairwise matrix factorization (CPMF) model to estimate users' pairwise preferences on events, groups and locations. We further develop an efficient stochastic gradient descent algorithm for the model learning. We conduct experiments on real-world Meetup datasets and the experimental results demonstrate that our CPMF model can outperform the state-of-the-art methods. Chun-Yi Liu 0003, Chuan Zhou 0001, Jia Wu 0001, Hongtao Xie 0001, Yue Hu 0002, Li Guo 0001 |
IJCNN | 4 |
| 2017 | Uyghur Language Text Detection in Complex Background Images Using Enhanced MSERs
Hongtao Xie 0001, Chuan Zhou 0001, Zhendong Mao 0001 |
MMM (1) | 2 |
| 2017 | RICS-DFA: a space and time-efficient signature matching algorithm with Reduced Input Character SetabstractSummary Regular expression matching as a core component of deep packet inspection is widely used in various kinds of modern network intrusion detection system, traffic classification system, network monitoring system, and so on. In these systems, regular expressions are typically converted to a deterministic finite automaton (DFA), which takes O(1) to scan each input character. However, DFA generally consumes a large amount of memory. This paper proposes a novel, space‐efficient and time‐efficient DFA presentation, called reduced input character set DFA (RICS‐DFA). A character escaping and replacing scheme is first introduced to decrease the size of DFA's character set and then to reduce DFA's space requirement with a series of optimization techniques. Based on transition rewriting, a RICS‐DFA constructing algorithm with time complexity of O(n) is presented in this paper. For real rule‐sets, RICS‐DFA reduces the memory consumption by 68–92%, compared with the original DFA. Finally, this paper designs a scalable RICS‐DFA matching engine on field‐programmable gate array platform in which the reduced state transition matrix is mapped to on‐chip memories. The throughput of executing deep packet inspection for real rule‐sets can achieve 7–50.5 Gbps. Copyright © 2016 John Wiley & Sons, Ltd. Qiu Tang, Lei Jiang 0003, Qiong Dai, Majing Su, Hongtao Xie 0001, Binxing Fang |
Concurr. Comput. Pract. Exp. | 5 |
| 2017 | Detecting Uyghur text in complex background images with convolutional neural network
Shancheng Fang, Hongtao Xie 0001, Zhineng Chen, Shiai Zhu, Xiaoyan Gu 0001, Xingyu Gao 0001 |
Multim. Tools Appl. | 2 |
| 2017 | Residual domain dictionary learning for compressed sensing video recovery
Yun Song, Gaobo Yang, Hongtao Xie 0001, Dengyong Zhang, Xingming Sun |
Multim. Tools Appl. | 3 |
| 2017 | Robust and parallel Uyghur text localization in complex background images
Yun Song, Hongtao Xie 0001, Zhineng Chen, Xingyu Gao 0001 |
Mach. Vis. Appl. | 3 |
| 2017 | Triple-Bit Quantization with Asymmetric Distance for Image Content Security
Hongtao Xie 0001, Chenggang Yan 0001 |
Mach. Vis. Appl. | 2 |
| 2015 | Data-oriented multi-index hashingabstractMulti-index hashing (MIH) is the state-of-the-art method for indexing binary codes, as it divides long codes into substrings and builds multiple hash tables. However, MIH is based on the dataset codes uniform distribution assumption, and will lose efficiency in dealing with non-uniformly distributed codes. Besides, there are lots of results sharing the same Hamming distance to a query, which makes the distance measure ambiguous. In this paper, we propose a data-oriented multi-index hashing method. We first compute the covariance matrix of bits and learn adaptive projection vector for each binary substring. Instead of using substrings as direct indices into hash tables, we project them with corresponding projection vectors to generate new indices. With adaptive projection, the indices in each hash table are near uniformly distributed. Then with covariance matrix, we propose a ranking method for the binary codes. By assigning different bit-level weights to different bits, the returned binary codes are ranked at a finer-grained binary code level. Experiments conducted on reference large scale datasets show that compared to MIH the time performance of our method can be improved by 36.9%-87.4%, and the search accuracy can be improved by 22.2%. Qingyun Liu 0001, Hongtao Xie 0001, Li Guo 0001 |
ICME | 2 |
| 2015 | Hierarchical Encoding of Binary Descriptors for Image MatchingabstractBinary descriptors are increasingly popular such as BRIEF, ORB, and BRISK. Typically, binary descriptors are computed by comparing pairs of image pixel intensities over a sampling pattern. To improve matching performance, lots of progresses have been made on the selection of pixel pairs, yet the discriminative power of pixel pairs is not fully studied. Zhendong Mao 0001, Lingling Tong, Hongtao Xie 0001, Qi Tian 0001 |
ICMR | 3 |
| 2015 | Fast approximate matching of binary codes with distinctive bits
Chenggang Yan 0001, Hongtao Xie 0001, Yanping Ma, Qiong Dai |
Frontiers Comput. Sci. | 2 |
| 2015 | Corrigendum to "Fast and scalable lock methods for video coding on many-core architecture" [J. Visual Communication and Image Representation 25(7) (2014) 1758-1762]
Weizhi Xu 0001, Hui Yu 0010, Dianjie Lu, Fenglong Song, Xiaochun Ye, Songwei Pei, Dongrui Fan, Hongtao Xie 0001 |
J. Vis. Commun. Image Represent. | 9 |
| 2015 | Corrigendum to "Fast and scalable lock methods for video coding on many-core architecture" [J. Visual Communication and Image Representation 25 (7) (2014) 1758-1762]
Weizhi Xu 0001, Hui Yu 0010, Dianjie Lu, Fenglong Song, Xiaochun Ye, Songwei Pei, Dongrui Fan, Hongtao Xie 0001 |
J. Vis. Commun. Image Represent. | 9 |
| 2014 | Video to Article Hyperlinking by Multiple Tag Property Exploration
Zhineng Chen, Bailan Feng, Hongtao Xie 0001, Rong Zheng 0005, Bo Xu 0002 |
MMM (1) | 3 |
| 2014 | Fusing audio vocabulary with visual features for pornographic video detection
Hongtao Xie 0001, Sheng Tang |
Future Gener. Comput. Syst. | 3 |
| 2014 | Fast and scalable lock methods for video coding on many-core architecture
Weizhi Xu 0001, Hui Yu 0010, Dianjie Lu, Fenglong Song, Xiaochun Ye, Songwei Pei, Dongrui Fan, Hongtao Xie 0001 |
J. Vis. Commun. Image Represent. | 9 |
| 2014 | Extracting salient region for pornographic image detection
Chenggang Yan 0001, Hongtao Xie 0001, Zhuhua Liao, Jian Yin 0003 |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | Contextual Query Expansion for Image RetrievalabstractIn this paper, we study the problem of image retrieval by introducing contextual query expansion to address the shortcomings of bag-of-words based frameworks: semantic gap of visual word quantization, and the efficiency and storage loss due to query expansion. Our method is built on common visual patterns (CVPs), which are the distinctive visual structures between two images and have rich contextual information. With CVPs, two contextual query expansions on visual word-level and image-level are explored, respectively. For visual word-level expansion, we find contextual synonymous visual words (CSVWs) and expand a word in the query image with its CSVWs to boost retrieval accuracy. CSVWs are the words that appear in the same CVPs and have same contextual meaning, i.e. similar spatial layout and geometric transformations. For image-level expansion, the database images that have the same CVPs are organized by linked list and the images that have the same CVPs as the query image, but not included in the results are automatically expanded. The main computation of these two expansions is carried out offline, and they can be integrated into the inverted file and efficiently applied to all images in the dataset. Experiments conducted on three reference datasets and a dataset of one million images demonstrate the effectiveness and efficiency of our method. Hongtao Xie 0001, Yongdong Zhang 0001, Jianlong Tan, Li Guo 0001, Jintao Li 0001 |
IEEE Trans. Multim. | 1 |
| 2013 | Robust common visual pattern discovery using graph matching
Hongtao Xie 0001, Yongdong Zhang 0001, Ke Gao 0012, Sheng Tang, Kefu Xu, Li Guo 0001, Jintao Li 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2011 | Local geometric consistency constraint for image retrievalabstractIn state-of-the-art image retrieval systems, an image is represented by bag-of-features (BOF). As BOF representation discards geometric relationships among local features, exploiting geometric constraints as post-processing procedure has been shown to greatly improve retrieval precision. However, full geometric constraints are computationally expensive and weak geometric constraints have limited range of applications. To efficiently handle common transformations and deformations, we present a novel local geometric consistency constraint (LGC) method. It utilizes the local similarity characteristic of deformations, and measures the pairwise geometric similarity of matches between two sets of local features. Besides, we propose a new method to accurately calculate the transformation matrix between two matched features, with the information provided by their local neighbors. Experiments performed on famous datasets show the excellent performance of our method. Hongtao Xie 0001, Ke Gao 0012, Yongdong Zhang 0001, Jintao Li 0001 |
ICIP | 1 |
| 2011 | Pairwise weak geometric consistency for large scale image searchabstractState-of-the-art image search systems mostly build on bag-of-features (BOF) representation. As BOF ignores geometric relationships among local features, geometric consistency constraints have been proposed to improve search precision. However, exploiting full geometric constraints are too computational expensive. Weak geometric constraints have strong assumptions and can only deal with uniform transformations. To handle view point changes and nonrigid deformations, in this paper we present a novel pairwise weak geometric consistency constraint (P-WGC) method. It utilizes the local similarity characteristic of deformations, and measures the pairwise geometric similarity of matches between two sets of local features. Experiments performed on four famous datasets and a dataset of one million of images show a significant improvement due to P-WGC as well as its efficiency. Further improvement of search accuracy is obtained when it is combined with full geometric verification. Hongtao Xie 0001, Ke Gao 0012, Yongdong Zhang 0001, Jintao Li 0001 |
ICMR | 1 |
| 2011 | Common visual pattern discovery via graph matchingabstractDiscovering common visual patterns (CVPs) between two images is a challenging problem, due to the significant photometric and geometric transformations, and the high computational cost. In this paper, we formulate CVPs discovery as a graph matching problem, depending on pairwise geometric compatibility between feature correspondences. To efficiently find all CVPs, we propose two algorithms--Preliminary Initialization Optimization (PIO) and Post Agglomerative Combining (PAC). PIO reduces the search space of CVPs discovery based on the internal homogeneity of CVPs, while PAC refines the discovery result in an agglomerative way. Experiments on object recognition and near-duplicate image re-trieval validate the effectiveness and efficiency of our method. Hongtao Xie 0001, Ke Gao 0012, Yongdong Zhang 0001, Jintao Li 0001, Huamin Ren |
ACM Multimedia | 1 |
| 2011 | Efficient Feature Detection and Effective Post-Verification for Large Scale Near-Duplicate Image SearchabstractState-of-the-art near-duplicate image search systems mostly build on the bag-of-local features (BOF) representation. While favorable for simplicity and scalability, these systems have three shortcomings: 1) high time complexity of the local feature detection; 2) discriminability reduction of local descriptors due to BOF quantization; and 3) neglect of the geometric relationships among local features after BOF representation. To overcome these shortcomings, we propose a novel framework by using graphics processing units (GPU). The main contributions of our method are: 1) a new fast local feature detector coined Harris-Hessian (H-H) is designed according to the characteristics of GPU to accelerate the local feature detection; 2) the spatial information around each local feature is incorporated to improve its discriminability, supplying semi-local spatial coherent verification (LSC); and 3) a new pairwise weak geometric consistency constraint (P-WGC) algorithm is proposed to refine the search result. Additionally, part of the system is implemented on GPU to improve efficiency. Experiments conducted on reference datasets and a dataset of one million images demonstrate the effectiveness and efficiency of H-H, LSC, and P-WGC. Hongtao Xie 0001, Ke Gao 0012, Yongdong Zhang 0001, Sheng Tang, Jintao Li 0001 |
IEEE Trans. Multim. | 1 |
| 2010 | GPU-based fast scale invariant interest point detectorabstractTo take full advantage of the powerful computing capability of graphics processing units (GPU) to speed up local feature detection, we present a novel GPU-based scale invariant interest point detector, coined Harris-Hessian(H-H). H-H detects Harris points in low scale and refines their location and scale in higher scale-space with the determinant of Hessian matrix. Compared to the existing methods, H-H significantly reduces the pixel-level computation complexity and has better parallelism. The experiment results show that with the assistance of GPU, H-H achieves up to a 10-20x speedup than CPU-based method. It only takes 6.3ms to detect a 640 × 480 image with high detection accuracy, meeting the need of real-time detection. Hongtao Xie 0001, Ke Gao 0012, Yongdong Zhang 0001, Jintao Li 0001 |
ICASSP | 1 |