Chuanbin Liu 0001

dblp:239/7365-1 · DBLP profile ↗
← Back
32ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0002-2840-6235ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 16 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
abstract
Multi-modal Retrieval-Augmented Generation (RAG) has become a critical method for empowering LLMs by leveraging candidate visual documents. However, current methods consider the entire document as the basic retrieval unit, introducing substantial irrelevant visual content in two ways: 1) Relevant documents often contain large regions unrelated to the query, diluting the focus on salient information; 2) Retrieving multiple documents to increase recall further introduces redundant and irrelevant documents. These redundant contexts distract the model's attention and further degrade the performance. To address this challenge, we propose RegionRAG, a novel framework that shifts the retrieval paradigm from the document level to the region level. During training, we design a hybrid supervision strategy from both labeled data and unlabeled data to pinpoint relevant patches. During inference, we propose a dynamic pipeline that intelligently groups salient patches into complete semantic regions. By delegating the task of identifying relevant regions to the retriever, RegionRAG enables the generator to focus solely on concise, query-relevant visual content, improving both efficiency and accuracy. Experiments on six benchmarks demonstrate that RegionRAG achieves state-of-the-art performance. It improves retrieval accuracy by 10.02% in R@1 on average, and boosts question answering accuracy by 3.56% while using only 71.42% visual tokens compared with prior methods.
Yinglu Li, Zhiying Lu, Chuanbin Liu 0001, Hongtao Xie 0001
AAAI5
2025 PosterMaker: Towards High-Quality Product Poster Generation with Accurate Text Rendering
abstract
Product posters, which integrate subject, scene, and text, are crucial promotional tools for attracting customers. Creating such posters using modern image generation methods is valuable, while the main challenge lies in accurately rendering text, especially for complex writing systems like Chinese, which contains over 10,000 individual characters. In this work, we identify the key to precise text rendering as constructing a character-discriminative visual feature as a control signal. Based on this insight, we propose a robust character-wise representation as control and we develop TextRenderNet, which achieves a high text rendering accuracy of over 90%. Another challenge in poster generation is maintaining the fidelity of user-specific products. We address this by introducing SceneGenNet, an inpainting-based model, and propose subject fidelity feedback learning to further enhance fidelity. Based on TextRenderNet and SceneGenNet, we present PosterMaker, an end-to-end generation framework. To optimize PosterMaker efficiently, we implement a two-stage training strategy that decouples text rendering and background generation learning. Experimental results show that PosterMaker outperforms existing baselines by a remarkable margin, which demonstrates its effectiveness.
Yifan Gao 0011, Zihang Lin, Chuanbin Liu 0001, Tiezheng Ge, Bo Zheng 0007, Hongtao Xie 0001
CVPR3
2025 Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models
abstract
Recent Multi-modal Large Language Models (MLLMs) have been challenged by the computational overhead resulting from massive video frames, often alleviated through compression strategies. However, the visual content is not equally contributed to user instructions, existing strategies (e.g., average pool) inevitably lead to the loss of potentially useful information. To tackle this, we propose the Hybridlevel Instruction Injection Strategy for Conditional Token Compression in MLLMs (HICom), utilizing the instruction as a condition to guide the compression from both local and global levels. This encourages the compression to retain the maximum amount of user-focused information while reducing visual tokens to minimize computational burden. Specifically, the instruction condition is injected into the grouped visual tokens at the local level and the learnable tokens at the global level, and we conduct the attention mechanism to complete the conditional compression. From the hybrid-level compression, the instruction-relevant visual parts are highlighted while the temporal-spatial structure is also preserved for easier understanding of LLMs. To further unleash the potential of HICom, we introduce a new conditional pre-training stage with our proposed dataset HICom-248K. Experiments show that our HICom can obtain distinguished video understanding ability with fewer tokens, increasing the performance by 2.43% average on three multiple-choice QA benchmarks and saving 78.8% tokens compared with the SOTA method. The code is available at https://github.com/lntzm/HICom.
Chen-Wei Xie, Pandeng Li, Longxiang Tang, Chuanbin Liu 0001, Hongtao Xie 0001
CVPR7
2025 TalkingAvatar: Learning 3D talking human avatar via NeRF
Lingyun Yu 0002, Chuanbin Liu 0001, Wu Liu 0005, Quanwei Yang, Meng Shao
Neurocomputing3
2025 DHVT: Dynamic Hybrid Vision Transformer for Small Dataset Recognition
abstract
The performance gap between Vision Transformers (ViTs) and Convolutional Neural Networks (CNNs) persists due to the lack of inductive bias, notably when training from scratch with limited datasets. This paper identifies two crucial shortcomings in ViTs: spatial relevance and diverse channel representation. Thus, ViTs struggle to grasp fine-grained spatial features and robust channel representation due to insufficient data. We propose the Dynamic Hybrid Vision Transformer (DHVT) to address these challenges. Regarding the spatial aspect, DHVT introduces convolution in the feature embedding phase and feature projection modules to enhance spatial relevance. Regarding the channel aspect, the dynamic aggregation mechanism and a groundbreaking design "head token" facilitate the recalibration and harmonization of disparate channel representations. Moreover, we investigate the choices of the network meta-structure and adopt the optimal multi-stage hybrid structure without the conventional class token. The methods are then modified with a novel dimensional variable residual connection mechanism to leverage the potential of the structure sufficiently. This updated variant, called DHVT2, offers a more computationally efficient solution for vision-related tasks. DHVT and DHVT2 achieve state-of-the-art image recognition results, effectively bridging the performance gap between CNNs and ViTs. The downstream experiments further demonstrate their strong generalization capacities.
Zhiying Lu, Chuanbin Liu 0001, Xiaojun Chang, Yongdong Zhang 0001, Hongtao Xie 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Distilling Multi-Level Semantic Cues Across Multi-Modalities for Face Forgery Detection
abstract
Existing face forgery detection methods attempt to identify low-level forgery artifacts (e.g., blending boundary, flickering) in spatial-temporal domains or high-level semantic inconsistencies (e.g., abnormal lip movements) between visual-auditory modalities for generalized face forgery detection. However, they still suffer from significant performance degradation when dealing with out-of-domain artifacts, as they only consider single semantic mode inconsistencies, but ignore the complementarity of forgery traces at different levels and different modalities. In this paper, we propose a novel Multi-modal Multi-level Semantic Cues Distillation Detection framework that adopts the teacher-student protocol to focus on both spatial-temporal artifacts and visual-auditory incoherence to capture multi-level semantic cues. Specifically, our framework primarily comprises the Spatial-Temporal Pattern Learning module and the Visual-Auditory Consistency Modeling module. The Spatial-Temporal Pattern Learning module employs a mask-reconstruction strategy, in which the student network learns diverse spatial-temporal patterns from a pixel-wise teacher network to capture low-level forgery artifacts. The Visual-Auditory Consistency Modeling module is designed to enhance the student network’s ability to identify high-level semantic irregularities, with a visual-auditory consistency modeling expert serving as a guide. Furthermore, a novel Real-Similarity loss is proposed to enhance the proximity of real faces in feature space without explicitly penalizing the distance from manipulated faces, which prevents the overfitting in particular manipulation methods and improves the generalization capability. Extensive experiments show that our method substantially improves the generalization and robustness performance. Particularly, our approach outperforms the SOTA detector by 1.4% in generalization performance on DFDC with large domain gaps, and by 2.0% in the robustness evaluation on the FF++ dataset under various extreme settings. Our code is available athttps://github.com/TianXie834/M2SD.
Lingyun Yu 0002, Chuanbin Liu 0001, Guoqing Jin, Zhiguo Ding 0006, Hongtao Xie 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Leveraging Concise Concepts With Probabilistic Modeling for Interpretable Visual Recognition
abstract
Interpretable visual recognition is essential for decision-making in high-stakes situations. Recent advancements have automated the construction of interpretable models by leveraging Visual Language Models (VLMs) and Large Language Models (LLMs) with Concept Bottleneck Models (CBMs), which process a bottleneck layer associated with human-understandable concepts. However, existing methods suffer from two main problems: a) the collected concepts from LLMs could be redundant with task-irrelevant descriptions, resulting in an inferior concept space with potential mismatch. b) VLMs directly map the global deterministic image embeddings with fine-grained concepts results in an ambiguous process with imprecise mapping results. To address the above two issues, we propose a novel solution for CBMs with Concise Concept and Probabilistic Modeling (CCPM) that can achieve superior classification performance via high-quality concepts and precise mapping strategy. Fisrt, we leverage in-context examples as category-related clues to guide LLM concept generation process. To mitigate redundancy in the concept space, we propose a Relation-Aware Selection (RAS) module to obtain a concise concept set that is discriminative and relevant based on image-concept and inter-concept relationships. Second, for precise mapping, we employ a Probabilistic Distribution Adapter (PDA) that estimates the inherent ambiguity of the image embeddings of pre-trained VLMs to capture the complex relationships with concepts. Extensive experiments indicate that our model achieves state-of-the-art results with a 5.48% improvement in classification accuracy on eight mainstream recognition benchmarks as well as reliable explainability through interpretable analysis.
Chuanbin Liu 0001, Yifan Gao 0011, Zhiying Lu, Hongtao Xie 0001, Yongdong Zhang 0001
IEEE Trans. Multim.2
2023 TextPainter: Multimodal Text Image Generation with Visual-harmony and Text-comprehension for Poster Design
abstract
Text design is one of the most critical procedures in poster design, as it relies heavily on the creativity and expertise of humans to design text images considering the visual harmony and text-semantic. This study introduces TextPainter, a novel multimodal approach that leverages contextual visual information and corresponding text semantics to generate text images. Specifically, TextPainter takes the global-local background image as a hint of style and guides the text image generation with visual harmony. Furthermore, we leverage the language model and introduce a text comprehension module to achieve both sentence-level and word-level style variations. Besides, we construct the PosterT80K dataset, consisting of about 80K posters annotated with sentence-level bounding boxes and text contents. We hope this dataset will pave the way for further research on multimodal text image generation. Extensive quantitative and qualitative experiments demonstrate that TextPainter can generate visually-and-semantically-harmonious text images for posters.
Yifan Gao 0011, Jinpeng Lin, Chuanbin Liu 0001, Hongtao Xie 0001, Tiezheng Ge, Yuning Jiang 0001
ACM Multimedia4
2023 High Fidelity Face Swapping via Semantics Disentanglement and Structure Enhancement
abstract
In this paper, we propose a novel Semantics and Structure-aware face Swapping framework (S2Swap) that exploits semantics disentanglement and structure enhancement for high fidelity face generation. Different from previous methods that either 1) suffer from degraded generation fidelity due to insufficient identity-attributes disentanglement or 2) neglect the importance of structure information for identity consistency, our approach can achieve local facial semantics disentanglement beyond global identity while boosting identity consistency through structure enhancement. Specifically, to achieve identity-attributes disentanglement, our S2Swap is designed from global-local perspectives. Firstly, an Oriented Identity Transfer module is proposed to globally disentangle target identity and attributes under global identity semantics prior. Such global disentanglement enables source identity transfer to the individual target identity. Secondly, a Local Semantics Disentanglement module is devised to disentangle local identity and identity-irrelevant facial semantics, providing local semantic compensation for the global counterpart. Moreover, to boost identity consistency, a Structure-Aware Head Modeling module is introduced to provide the desired face structure enhancement through an intuitive face sketch. Finally, considering the identity-attributes trade-off, we adaptively integrate semantics and structure information in a self-learning manner. Extensive experiments qualitatively and quantitatively show that our method outperforms SOTA face swapping methods in terms of both identity transfer and attribute preservation.
Lingyun Yu 0002, Hongtao Xie 0001, Chuanbin Liu 0001, Zhiguo Ding 0006, Quanwei Yang, Yongdong Zhang 0001
ACM Multimedia4
2023 Frequency-based Zero-Shot Learning with Phase Augmentation
abstract
Zero-Shot Learning (ZSL) aims to recognize images from seen and unseen classes by aligning visual and semantic knowledge (e.g., attribute descriptions). However, the fine-grained attributes in the RGB domain can be easily affected by background noise (e.g., the grey bird tail blending with the ground), making it difficult to effectively distinguish them. Analyzing the features in the frequency domain assists in better distinguishing the attributes since their patterns remain consistent across different images, unlike noise which may be more variable. Nevertheless, existing ZSL methods typically learn visual features directly from the RGB domain, which can impede the recognition of certain attributes. To overcome this limitation, we propose a novel ZSL method named Frequency-based Phase Augmentation (FPA) network, which learns an effective representation of the attributes in the frequency domain. Specifically, we introduce a Hybrid Phase Augmentation (HPA) module to transform visual features into the frequency domain and augment the phase component for better retention of semantic information of the attributes. The use of phase-augmented features enables FPA to capture more semantic knowledge that can be challenging to distinguish in the RGB domain, suppress noise, and highlight significant attributes. Our extensive experiments show that FPA achieves state-of-the-art performance across four standard datasets.
Wanting Yin, Hongtao Xie 0001, Lei Zhang 0119, Jiannan Ge, Pandeng Li, Chuanbin Liu 0001, Yongdong Zhang 0001
ACM Multimedia6
2023 Recombining Vision Transformer Architecture for Fine-Grained Visual Categorization
Xuran Deng, Chuanbin Liu 0001, Zhiying Lu
MMM (2)2
2023 Prototypical Matching Networks for Video Object Segmentation
abstract
Semi-supervised video object segmentation is the task of segmenting the target in sequential frames given the ground truth mask in the first frame. The modern approaches usually utilize such a mask as pixel-level supervision and typically exploit pixel-to-pixel matching between the reference frame and current frame. However, the matching at pixel level, which overlooks the high-level information beyond local areas, often suffers from confusion caused by similar local appearances. In this paper, we present Prototypical Matching Networks (PMNet) - a novel architecture that integrates prototypes into matching-based video objection segmentation frameworks as high-level supervision. Specifically, PMNet first divides the foreground and background areas into several parts according to the similarity to the global prototypes. The part-level prototypes and instance-level prototypes are generated by encapsulating the semantic information of identical parts and identical instances, respectively. To model the correlation between prototypes, the prototype representations are propagated to each other by reasoning on a graph structure. Then, PMNet stores both the pixel-level features and prototypes in the memory bank as the target cues. Three affinities, i.e., pixel-to-pixel affinity, prototype-to-pixel affinity, and prototype-to-prototype affinity, are derived to measure the similarity between the query frame and the features in the memory bank. The features aggregated from the memory bank using these affinities provide powerful discrimination from both the pixel-level and prototype-level perspectives. Extensive experiments conducted on four benchmarks demonstrate superior results than the state-of-the-art video object segmentation techniques.
Fanchao Lin, Zhaofan Qiu, Chuanbin Liu 0001, Ting Yao 0003, Hongtao Xie 0001, Yongdong Zhang 0001
IEEE Trans. Image Process.3
2023 Learning Cross-Channel Representations for Semantic Segmentation
abstract
Semantic segmentation is a fundamental problem in multimedia which requires delicate per-pixel predictions of object categories. Recently, many researchers strive to refine the pixel-wise feature withspatial-contextual information. However, many of them still neglect the invisible hand of cross-channelinformation which provides inherent semantics to facilitate the segmentation performance. On the one hand, in the feature extraction stage, enhancing informative channels and suppressing trivial ones contribute to the acquisition of valuable semantic features, and thus improving the segmentation accuracy. On the other hand, in the prediction stage, we can predict the complete objects more clearly by finding the connections and complements between different channels, which can also contribute to the pixel prediction. And based on this idea, we propose a novel Channel-Adaptive Network for semantic segmentation, which is capable of enhancing the features from the perspective of channels in both feature extraction stage and prediction stage. Specifically, we propose two modules: (i) the Comprehensive Information Channel Attention (CiCA) module that addresses the shortcomings of existing channel attention by learning both low and high frequency components within each channel for emphasizing the informative channels; (ii) the Inter-Channel Relationship Reasoning (iCRR) module which is applied on the top of the feature extractor to adaptively enhance the interdependent channels by mining the complementary associations between them. Besides, our Channel-Adaptive Network is highly flexible, with a plug-and-play design. Extensive experiments have demonstrated that our method achieves the state-of-the-art segmentation performance on three challenging datasets, including Cityscapes (82.1%), ADE20K (46.51%) and PASCAL Context (55.0%).
Lingfeng Ma, Hongtao Xie 0001, Chuanbin Liu 0001, Yongdong Zhang 0001
IEEE Trans. Multim.3
2023 Multi-task hourglass network for online automatic diagnosis of developmental dysplasia of the hip
Hongtao Xie 0001, Qingfeng Tan, Chuanbin Liu 0001, Zhendong Mao 0001, Yongdong Zhang 0001
World Wide Web (WWW)5
2022 Weakly Supervised Pediatric Bone Age Assessment Using Ultrasonic Images via Automatic Anatomical RoI Detection
abstract
Bone age assessment (BAA) is vital in pediatric clinical diagnosis. Existing deep learning methods predict bone age based on Regions of Interest (RoIs) detection or segmentation of hand radiograph, which requires expensive annotations. Limitations of radiographic technique on imaging and cost hinder their clinical application as well. Compared to X-ray images, ultrasonic images are rather clean, cheap and flexible, but the deep learning research on ultrasonic BAA is still a white space. For this purpose, we propose a weakly supervised interpretable framework entitled USB-Net, utilizing ultrasonic pelvis images and only image-level age annotations. USB-Net consists of automatic anatomical RoI detection stage and age assessment stage. In the detection stage, USB-Net locates the discriminative anatomical RoIs of pelvis through attention heatmap without any extra RoI supervision. In the assessment stage, the cropped anatomical RoI patch is fed as fine-grained input to estimate age. In addition, we provide the first ultrasonic BAA dataset composed of 1644 ultrasonic hip joint images with image-level labels of age and gender. The experimental results verify that our model keeps consistent attention with human knowledge and achieves 16.24 days mean absolute error (MAE) on USBAA dataset.
Yunyan Yan, Chuanbin Liu 0001, Hongtao Xie 0001, Zhendong Mao 0001
ICMR2
2022 Geometry Aligned Variational Transformer for Image-conditioned Layout Generation
abstract
Layout generation is a novel task in computer vision, which combines the challenges in both object localization and aesthetic appraisal, widely used in advertisements, posters and slides design. An accurate and pleasant layout should consider both the intra-domain relationship within layout elements and the inter-domain relationship between layout elements and image. However, most previous methods simply focus on image-content-agnostic layout generation, without leveraging the complex visual information from the image. To this end, we explore a novel paradigm entitled image-conditioned layout generation, which aims to add text overlays to an image in a semantically coherent manner. Specifically, we propose an Image-Conditioned Variational Transformer (ICVT) that autoregressively generates various layouts in an image. First, self-attention mechanism is adopted to model the contextual relationship within layout elements, while cross-attention mechanism is used to fuse the visual information of conditional images. Subsequently, we take them as building blocks of conditional variational autoencoder (CVAE), which demonstrates appealing diversity. Second, in order to alleviate the gap between layout elements domain and visual domain, we design a Geometry Alignment module, in which the geometric information of the image is aligned with the layout representation. In addition, we construct a large-scale advertisement poster layout designing dataset with delicate layout and saliency map annotations. Experimental results show that our model can adaptively generate layouts in the non-intrusive area of the image, resulting in a harmonious layout design.
Yunning Cao, Chuanbin Liu 0001, Hongtao Xie 0001, Tiezheng Ge, Yuning Jiang 0001
ACM Multimedia4
2022 Proxy Probing Decoder for Weakly Supervised Object Localization: A Baseline Investigation
abstract
Weakly supervised object localization (WSOL) aims to localize the object with only image category labels. Existing methods generally fine-tune the models with manually selected training epochs and subjective loss functions to mitigate the partial activation problem of the classification-based model. However, such fine-tuning scheme would cause the model to degrade, e.g. affect the classification performance and generalization capabilities of the pre-trained model. In this paper, we propose a novel method named Proxy Probing Decoder (PPD) to meet these challenges, which utilizes the segmentation property of self-attention map in the self-supervised vision transformer and breaks through model fine-tuning with a novel proxy probing decoder. Specifically, we utilize the self-supervised vision transformer to capture long-range dependencies and avoid partial activation. Then we simply adopt a proxy consisting of a series of decoding layers to transform the feature representations into the heatmap of the objects' foreground and conduct localization. The backbone parameters are frozen during training while the proxy is used to decode the feature and localize the object. In this way, the vision transformer model can maintain the feature representation capabilities and only the proxy is required for adapting to the task. Without bells and whistles, our framework achieves 55.0% Top-1 Loc on the ILSVRC2012 dataset and 78.8% Top-1 Loc on the CUB-200-2011 dataset, which surpasses state-of-the-art by a large margin and provides a simple baseline. Codes and models will be available on Github.
Hongtao Xie 0001, Chuanbin Liu 0001, Yongdong Zhang 0001
ACM Multimedia3
2022 Bridging the Gap Between Vision Transformers and Convolutional Neural Networks on Small Datasets
abstract
There still remains an extreme performance gap between Vision Transformers (ViTs) and Convolutional Neural Networks (CNNs) when training from scratch on small datasets, which is concluded to the lack of inductive bias. In this paper, we further consider this problem and point out two weaknesses of ViTs in inductive biases, that is, the spatial relevance and diverse channel representation. First, on spatial aspect, objects are locally compact and relevant, thus fine-grained feature needs to be extracted from a token and its neighbors. While the lack of data hinders ViTs to attend the spatial relevance. Second, on channel aspect, representation exhibits diversity on different channels. But the scarce data can not enable ViTs to learn strong enough representation for accurate recognition. To this end, we propose Dynamic Hybrid Vision Transformer (DHVT) as the solution to enhance the two inductive biases. On spatial aspect, we adopt a hybrid structure, in which convolution is integrated into patch embedding and multi-layer perceptron module, forcing the model to capture the token features as well as their neighboring features. On channel aspect, we introduce a dynamic feature aggregation module in MLP and a brand new "head token" design in multi-head self-attention module to help re-calibrate channel representation and make different channel group representation interacts with each other. The fusion of weak channel representation forms a strong enough representation for classification. With this design, we successfully eliminate the performance gap between CNNs and ViTs, and our DHVT achieves a series of state-of-the-art performance with a lightweight model, 85.68% on CIFAR-100 with 22.8M parameters, 82.3% on ImageNet-1K with 24.0M parameters. Code is available at https://github.com/ArieSeirack/DHVT.
Zhiying Lu, Hongtao Xie 0001, Chuanbin Liu 0001, Yongdong Zhang 0001
NeurIPS3
2022 Bilateral Temporal Re-Aggregation for Weakly-Supervised Video Object Segmentation
abstract
Weakly-supervised video object segmentation is an emerging video task to track and segment the target given a simple bounding box label, which requires the method to fully catch and utilize the target information. Most existing approaches only rely on the guidance of a single frame and ignore the interaction between different frames when gathering information, making them hard to achieve reliable target representation. In this paper, we propose to capture the temporal dependencies and gather information from multiple frames through bilateral temporal re-aggregation. We explore three schemes to build the aggregation: 1) a two-stage re-aggregation mechanism is applied to provide target prior to the current frame, which obtains more valid feature matching and information aggregation; 2) a query-memory bilateral aggregation module is proposed to aggregate features from an unlimited amount of past frames and enable the mutual perception between different frames to validate the gathered information; 3) we guide the learning of aggregation modules through a novel cross-task representation distillation, transferring the knowledge from a semi-supervised model to our weakly-supervised model without increasing the inference latency. These schemes collaboratively build an efficient and competent aggregation process, thus we can fully exploit the video context to make the inference. Experimental results on four benchmarks show that our method achieves superior performance than previous methods and still maintains the efficiency ($e.g$., overall scores of 70.4% and 72.5% on the YouTube-VOS and DAVIS 2017 validation sets, respectively).
Fanchao Lin, Hongtao Xie 0001, Chuanbin Liu 0001, Yongdong Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 Global Characteristic Guided Landmark Detection for Genu Valgus and Varus Diagnosis
Lingfeng Ma, Chuanbin Liu 0001, Hongtao Xie 0001
ICIG (2)2
2021 Self-Supervised Attention Mechanism for Pediatric Bone Age Assessment With Efficient Weak Annotation
abstract
Pediatric bone age assessment (BAA) is a common clinical practice to investigate endocrinology, genetic and growth disorders of children. Different specific bone parts are extracted as anatomical Regions of Interest (RoIs) during this task, since their morphological characters have important biological identification in skeletal maturity. Following this clinical prior knowledge, recently developed deep learning methods address BAA with an RoI-based attention mechanism, which segments or detects the discriminative RoIs for meticulous analysis. Great strides have been made, however, these methods strictly require large and precise RoIs annotations, which limits the real-world clinical value. To overcome the severe requirements on RoIs annotations, in this paper, we propose a novel self-supervised learning mechanism to effectively discover the informative RoIs without the need of extra knowledge and precise annotation-only image-level weak annotation is all we take. Our model, termed PEAR-Net for Part Extracting and Age Recognition Network, consists of one Part Extracting (PE) agent for discriminative RoIs discovering and one Age Recognition (AR) agent for age assessment. Without precise supervision, the PE agent is designed to discover and extract RoIs fully automatically. Then the proposed RoIs are fed into AR agent for feature learning and age recognition. Furthermore, we utilize the self-consistency of RoIs to optimize PE agent to understand the part relation and select the most useful RoIs. With this self-supervised design, the PE agent and AR agent can reinforce each other mutually. To the best of our knowledge, this is the first end-to-end bone age assessment method which can discover RoIs automatically with only image-level annotation. We conduct extensive experiments on the public RSNA 2017 dataset and achieve state-of-the-art performance with MAE 3.99 months. Project is available at http://imcc.ustc.edu.cn/project/ssambaa/.
Chuanbin Liu 0001, Hongtao Xie 0001, Yongdong Zhang 0001
IEEE Trans. Medical Imaging1
2021 Hip Landmark Detection With Dependency Mining in Ultrasound Image
abstract
Developmental dysplasia of the hip (DDH) is a common and serious disease in infants. Hip landmark detection plays a critical role in diagnosing the development of neonatal hip in the ultrasound image. However, the local confusion and the regional weakening make this task challenging. To solve these challenges, we explore the stable hip structure and the distinguishable local features to provide dependencies for hip landmark detection. In this paper, we propose a novel architecture named Dependency Mining ResNet (DM-ResNet), which investigates end-to-end dependency mining for more accurate and much faster hip landmark detection. First of all, we convert the landmark detection to the heatmap estimation by ResNet to build a strong baseline architecture for fast and accurate detection. Secondly, a dependency mining module is explored to mine the dependencies and leverage both the local and global information to decline the local confusion and strengthen the weakening region. Thirdly, we propose a simple but effective local voting algorithm (LVA) that seeks trade-off between long-range and short-range dependencies in the hip ultrasound image. Besides, a dataset with 2000 annotated hip ultrasound images is constructed in our work. It is the first public hip ultrasound dataset for open research. Experimental results show that our method achieves excellent precision in hip landmark detection (average point error of 0.719mm and successful detection rate within 1mm of 79.9%).
Hongtao Xie 0001, Chuanbin Liu 0001, Xun Chen 0001, Yongdong Zhang 0001
IEEE Trans. Medical Imaging3
2020 Filtration and Distillation: Enhancing Region Attention for Fine-Grained Visual Categorization
abstract
Delicate attention of the discriminative regions plays a critical role in Fine-Grained Visual Categorization (FGVC). Unfortunately, most of the existing attention models perform poorly in FGVC, due to the pivotal limitations in discriminative regions proposing and region-based feature learning. 1) The discriminative regions are predominantly located based on the filter responses over the images, which can not be directly optimized with a performance metric. 2) Existing methods train the region-based feature extractor as a one-hot classification task individually, while neglecting the knowledge from the entire object. To address the above issues, in this paper, we propose a novel “Filtration and Distillation Learning” (FDL) model to enhance the region attention of discriminate parts for FGVC. Firstly, a Filtration Learning (FL) method is put forward for discriminative part regions proposing based on the matchability between proposing and predicting. Specifically, we utilize the proposing-predicting matchability as the performance metric of Region Proposal Network (RPN), thus enable a direct optimization of RPN to filtrate most discriminative regions. Go in detail, the object-based feature learning and region-based feature learning are formulated as “teacher” and “student”, which can furnish better supervision for region-based feature learning. Accordingly, our FDL can enhance the region attention effectively, and the overall framework can be trained end-to-end without neither object nor parts annotations. Extensive experiments verify that FDL yields state-of-the-art performance under the same backbone with the most competitive approaches on several FGVC tasks.
Chuanbin Liu 0001, Hongtao Xie 0001, Zhengjun Zha, Lingfeng Ma, Lingyun Yu 0002, Yongdong Zhang 0001
AAAI1
2020 CircleNet for Hip Landmark Detection
abstract
Landmark detection plays a critical role in diagnosis of Developmental Dysplasia of the Hip (DDH). Heatmap and anchor-based object detection techniques could obtain reasonable results. However, they have limitations in both robustness and precision given the complexities and inhomogeneity of hip X-ray images. In this paper, we propose a much simpler and more efficient framework called CircleNet to improve the accuracy of landmark detection by predicting landmark and corresponding radius. Using the CircleNet, we not only constrain the relationship between landmarks but also integrate landmark detection and object detection into an end-to-end framework. In order to capture the effective information of the long-range dependency of landmarks in the DDH image, here we propose a new context modeling framework, named the Local Non-Local (LNL) block. The LNL block has the benefits of both non-local block and lightweight computation. We construct a professional DDH dataset for the first time and evaluate our CircleNet on it. The dataset has the largest number of DDH X-ray images in the world to our knowledge. Our results show that the CircleNet can achieve the state-of-the-art results for landmark detection on the dataset with a large margin of 1.8 average pixels compared to current methods. The dataset and source code will be publicly available.
Hongtao Xie 0001, Chuanbin Liu 0001, Zhengjun Zha, Jun Sun 0016, Yongdong Zhang 0001
AAAI3
2020 Learning Rich Attention for Pediatric Bone Age Assessment
Chuanbin Liu 0001, Hongtao Xie 0001, Yunyan Yan, Zhendong Mao 0001, Yongdong Zhang 0001
MICCAI (1)1
2020 Law Is Order: Protecting Multimedia Network Transmission by Game Theory and Mechanism Design
Chuanbin Liu 0001, Youliang Tian, Hongtao Xie 0001
MMM (2)1
2020 Misshapen Pelvis Landmark Detection With Local-Global Feature Learning for Diagnosing Developmental Dysplasia of the Hip
abstract
Developmental dysplasia of the hip (DDH) is one of the most common orthopedic disorders in infants and young children. Accurately detecting and identifying the misshapen anatomical landmarks plays a crucial role in the diagnosis of DDH. However, the diversity during the calcification and the deformity due to the dislocation lead it a difficult task to detect the misshapen pelvis landmarks for both human expert and computer. Generally, the anatomical landmarks exhibit stable morphological features in part regions and rigid structural features in long ranges, which can be strong identification for the landmarks. In this paper, we investigate the local morphological features and global structural features for the misshapen landmark detection with a novel Pyramid Non-local UNet (PN-UNet). Firstly, we mine the local morphological features with a series of convolutional neural network (CNN) stacks, and convert the detection of a landmark to the segmentation of the landmark's local neighborhood by UNet. Secondly, a non-local module is employed to capture the global structural features with high-level structural knowledge. With the end-to-end and accurate detection of pelvis landmarks, we realize a fully automatic and highly reliable diagnosis of DDH. In addition, a dataset with 10,000 pelvis X-ray images is constructed in our work. It is the first public dataset for diagnosing DDH and has been already released for open research. To the best of our knowledge, this is the first attempt to apply deep learning method in the diagnosis of DDH. Experimental results show that our approach achieves an excellent precision in landmark detection (average point to point error of 0.9286mm) and illness diagnosis over human experts. Project is available at http://imcc.ustc.edu.cn/project/ddh/.
Chuanbin Liu 0001, Hongtao Xie 0001, Zhendong Mao 0001, Jun Sun 0016, Yongdong Zhang 0001
IEEE Trans. Medical Imaging1
2020 Bidirectional Attention-Recognition Model for Fine-Grained Object Classification
abstract
Fine-grained object classification (FGOC) is a challenging research topic in multimedia computing with machine learning, which faces two pivotal conundrums: focusing attention on the discriminate part regions, and then processing recognition with the part-based features. Existing approaches generally adopt a unidirectional two-step structure, that first locate the discriminate parts and then recognize the part-based features. However, they neglect the truth that part localization and feature recognition can be reinforced in a bidirectional process. In this paper, we propose a novel bidirectional attention-recognition model (BARM) to actualize the bidirectional reinforcement for FGOC. The proposed BARM consists of one attention agent for discriminate part regions proposing and one recognition agent for feature extraction and recognition. Meanwhile, a feedback flow is creatively established to optimize the attention agent directly by recognition agent. Therefore, in BARM the attention agent and the recognition agent can reinforce each other in a bidirectional way and the overall framework can be trained end-to-end without neither object nor parts annotations. Moreover, a novel Multiple Random Erasing data augmentation is proposed, and it exhibits impressive pertinency and superiority for FGOC. Conducted on several extensive FGOC benchmarks, BARM outperforms the present state-of-the-art methods in classification accuracy. Furthermore, BARM exhibits a clear interpretability and keeps consistent with the human perception in visualization experiments.
Chuanbin Liu 0001, Hongtao Xie 0001, Zhengjun Zha, Lingyun Yu 0002, Zhineng Chen, Yongdong Zhang 0001
IEEE Trans. Multim.1
2019 Semantic-Embedding and Shape-Aware U-Net for Ultrasound Eyeball Segmentation
abstract
Segmentation of eyeball region from ultrasound images is a new research direction for the diagnosis of ophthalmic diseases. Despite the advantages of convenience and cheapness, ultrasound images bring more noise and fuzzy contour compared with other medical images. Existing methods fail to give a segmentation with reasonable eyeball shape, especially when the contour is ambiguous. In this paper, we propose a novel framework based on convolutional neural network, named semantic-embedding and shape-aware U-Net, to deal with the segmentation in blurred images. A signed distance field is used as label instead of the traditional binary mask label to add shape prior in network. The applying of semantic embedding modules fuses semantic information between different stages of the network. Experimental results show that our method improves the ability to segment image with blurred edges and outperforms existing methods in the accuracy of segmentation.
Fanchao Lin, Chuanbin Liu 0001, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001
ICME2
2019 Extract Bone Parts Without Human Prior: End-to-end Convolutional Neural Network for Pediatric Bone Age Assessment
Chuanbin Liu 0001, Hongtao Xie 0001, Zhengjun Zha, Fanchao Lin, Yongdong Zhang 0001
MICCAI (6)1
2019 Misshapen Pelvis Landmark Detection by Spatial Local Correlation Mining for Diagnosing Developmental Dysplasia of the Hip
Chuanbin Liu 0001, Hongtao Xie 0001, Jun Sun 0016, Yongdong Zhang 0001
MICCAI (6)1
2018 Potential of Attention Mechanism for Classification of Optical Coherence Tomography Images
abstract
Deep neural network (DNN) can extract high-dimensional feature of images for computer vision tasks including Optical Coherence Tomography (OCT) images classification. However, OCT images are usually processed by DNN just like natural images, thus the performance of DNN is not satisfactory. We present an end-to-end DNN targeting OCT images classification. Considering the characteristic of OCT images, we introduce attention mechanism into classifier to extract more specific feature of OCT images. Our network demonstrates its capacity to enhance the features that represent the disease region. Our method achieves the state-of-the-art performance with average accuracy of 99.5% and F1-score of 0.995 on the OCT images dataset.
Zhihua Shang, Zilong Fu, Chuanbin Liu 0001, Hongtao Xie 0001, Yongdong Zhang 0001
VCIP3