EDBT 2026 Demo / reviewers in the wild / expert
Bo Liu 0005
dblp:58/2670-5
· DBLP profile ↗
71ranked-venue papers
10as first author
40since 2021 · last 2026
0000-0002-8678-2271ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 5 first-author · 28 since 2021Artificial intelligence and machine learning · 37 · 5 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparse-Scale Transformer with Bidirectional Awareness for Time Series ForecastingabstractTime series forecasting (TSF) plays a crucial role in many real-world applications, such as weather prediction and economic planning. While Transformer-based models have shown strong capabilities in modeling long-range dependencies, effectively capturing the multi-scale temporal dynamics inherent in time series remains a major challenge. Existing methods often adopt time-windows of varying sizes, which may introduce noisy or irrelevant representations when mismatched with the underlying temporal patterns, potentially leading to overfitting. In this paper, we propose Sparse-Scale Transformer (SSformer) with Bidirectional Awareness for Time Series Forecasting to enhance the multi-scale modeling for time series. Specifically, we propose a novel Sparse-Scale Convolution (SSC) block that imposes sparsity on scales to obtain the informative representations by evaluating the intra-scale segment similarity of time series, and utilizes scale-specific convolutions to extract local patterns. Furthermore, we design a Bidirectional-Scale Interaction (BSI) block to explicitly model scale correlations in both coarse-to-fine and fine-to-coarse directions. Finally, scale predictions are ensembled to fully exploit the complementary forecasting capabilities across scales. Extensive experiments on various real-world datasets demonstrate that SSformer achieves state-of-the-art performance with superior efficiency. Ying Liu 0096, Bo Liu 0005, Sheng Huang 0001, Wenbo Hu 0001, Meng Wang 0001, Richang Hong |
AAAI | 2 |
| 2026 | DICE: Discrete Inversion Enabling Controllable Editing for Masked Generative ModelsabstractRecent advances in discrete diffusion models have demonstrated strong performance in image generation and masked language modeling, yet they remain limited in their capacity for controlled content editing. We propose DICE (Discrete Inversion for Controllable Editing), a novel framework that pioneers precise inversion capabilities for discrete diffusion models, including both masked generative and multinomial diffusion variants. Our key innovation lies in capturing noise sequences and masking patterns during reverse diffusion process, enabling both accurate reconstruction and flexible editing without relying on predefined masks or attention-based manipulations. Through comprehensive experiments across image and text modalities using models such as Paella, VQ-Diffusion, RoBERTa and LLaDA, we demonstrate that DICE successfully maintains high fidelity to the original data while significantly expanding editing capabilities. These results establish new possibilities for fine-grained content manipulation in discrete spaces. Xiaoxiao He, Quan Dao, Ligong Han, Song Wen 0001, Minhao Bai, Di Liu 0003, Han Zhang 0010, Felix Juefei-Xu, Chaowei Tan, Bo Liu 0005, Martin Renqiang Min, Kang Li 0004, Faez Ahmed, Akash Srivastava, Hongdong Li, Junzhou Huang, Dimitris N. Metaxas |
WACV | 10 |
| 2026 | Large Sign Language Models: Toward 3D American Sign Language TranslationabstractWe present Large Sign Language Models (LSLM), a novel framework for translating 3D American Sign Language (ASL) by leveraging Large Language Models (LLMs) as the backbone, which can benefit hearing-impaired individuals’ virtual communication. Unlike existing sign language recognition methods that rely on 2D video, our approach directly utilizes 3D sign language data to capture rich spatial, gestural, and depth information in 3D scenes. This enables more accurate and resilient translation, enhancing digital communication accessibility for the hearing-impaired community. Beyond the task of ASL translation, our work explores the integration of complex, embodied multimodal languages into the processing capabilities of LLMs, moving beyond purely text-based inputs to broaden their understanding of human communication. We investigate both direct translation from 3D gesture features to text and an instruction-guided setting where translations can be modulated by external prompts, offering greater flexibility. This work provides a foundational step toward inclusive, multimodal intelligent systems capable of understanding diverse forms of language. Xiaoxiao He, Di Liu 0003, Zhaoyang Xia, Chaowei Tan, Vivian Li, Bo Liu 0005, Dimitris N. Metaxas, Mubbasir Kapadia |
WACV | 8 |
| 2026 | Multiple Instance Learning Framework with Masked Hard Instance Mining for Gigapixel Histopathology Image Analysis
Sheng Huang 0001, Fengtao Zhou, Bo Liu 0005, Qingshan Liu 0001 |
Int. J. Comput. Vis. | 5 |
| 2026 | HDG-CLIP: Hierarchical Dual-Granularity Vision-Semantic Alignment for Open-Vocabulary Multi-Label Image ClassificationabstractOpen-Vocabulary Multi-Label Image Classification (OV-MLIC) is an emerging task in computer vision aimed at recognizing unseen categories in real-world scenarios, leveraging Vision and Language Pre-training (VLP) models like CLIP. However, existing methods overlook the impact of category coupling and scale variation on cross-category knowledge transfer, thereby restricting performance on unseen categories. To address this issue, we propose a novel OV-MLIC method called Hierarchical Dual-Granularity Alignment-CLIP (HDG-CLIP), which emphasizes the complementary characteristics of different modalities and introduces a sample-category matching mechanism. Specifically, to address the category coupling issue, we construct semantic category prototypes to enhance cross-category knowledge transfer. Through the interaction between visual embeddings and category prototypes, we decouple category-specific information from mixed visual features and leverage the visual context of samples to learn category-level visual features. For mitigating the scale variation issue, we build a sample-category dual-granularity matching mechanism based on the difference in capture capability of different modalities across scales, thereby improving the object localization accuracy from a multi-dimensional perspective. Extensive experimental results show that HDG-CLIP exhibits state-of-art performance over existing methods on both the NUS-WIDE and the Open-Images datasets. Our code is available at https://github.com/wakihy/HDG-CLIP. Beiyan Liu, Sheng Huang 0001, Bo Liu 0005, Xiaobin Huang, Nankun Mu, Richang Hong, Meng Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Scan-Invariant Mamba With Differentiated Sequence Contrastive Learning in Computational PathologyabstractMultiple instance learning (MIL) is a commonly used paradigm for histopathological analysis due to the ultra-high resolution and coarse-grained labels of Whole Slide Images (WSIs). Recent studies apply Mamba architecture to WSI classification by modeling MIL as long-sequence tasks, but a key discrepancy remains: Mamba's output is sensitive to scanning modes, whereas MIL requires scan-invariant predictions. To address this problem, we propose Scan-invariant Mamba with Differentiated Sequence Contrastive Learning (SMDC-MIL), a novel Mamba-based MIL approach enabling bag-level feature learning independent of input modes. Our method mitigates scanning-mode impacts and adapts Mamba to learn the bag discrimination features that are independent of the input mode via two innovations: 1) a differentiated sequence generation mechanism that employs instance rearrangement, augmentation, and masking to simulate real-world scanning variations by maximizing differences in sequence order, length, and composition from the same WSI; and 2) a differentiated sequence contrastive learning architecture that enforces consistent bag-level representations and predictions across diverse sequences using the same Mamba model, guiding it to prioritize scan-invariant discriminative features. Experimental results on 4 computational pathology tasks and 10 datasets demonstrate that our SMDC-MIL achieves state-of-the-art performance compared to other methods. The corresponding code is available at https://github.com/LianYueZ/SMDCMIL.git. Sheng Huang 0001, Xin Zhang 0131, Bo Liu 0005, Fengtao Zhou, Kang Li 0004, Hao Chen 0011, Meng Wang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Hier-pFedMe: Hierarchical Personalized Federated Learning with Moreau EnvelopesabstractMost existing Personalized Federated Learning (PFL) approaches rely on frequent client-to-cloud communication to ensure convergence, making them vulnerable to bandwidth constraints and network latency. To address this deficiency, we propose Hier-pFedMe, a novel client-edge-cloud tri-level PFL framework formulated as hierarchical optimization with Moreau envelopes. Unlike traditional client-cloud bi-level architectures, Hier-pFedMe introduces an additional intermediate edge-server level to coordinate client training, alleviating the communication burden on the central cloud server and improving efficiency through parallel edge server operations. Moreover, the hierarchical design enhances privacy by restricting client updates to small groups via edge servers, limiting direct access to client updates by the central server. Experiments on CIFAR-10, CIFAR-100 and Tiny-ImageNet datasets demonstrate that Hier-pFedMe improves personalized learning performance while reducing communication overhead on heterogeneous data. Fanfan Ji, Bo Liu 0005, Xiao-Tong Yuan |
ICME | 3 |
| 2025 | Enhanced prototype network with gated point recyclable feature mining for few-shot 3D point cloud classification
Sheng Huang 0001, Luwen Huangfu, Ma Rui, Bo Liu 0005 |
Knowl. Based Syst. | 5 |
| 2025 | Dual-View Alignment Learning With Hierarchical-Prompt for Class-Imbalance Multi-Label Image ClassificationabstractReal-world datasets often exhibit class imbalance across multiple categories, manifesting as long-tailed distributions and few-shot scenarios. This is especially challenging in Class-Imbalanced Multi-Label Image Classification (CI-MLIC) tasks, where data imbalance and multi-object recognition present significant obstacles. To address these challenges, we propose a novel method termed Dual-View Alignment Learning with Hierarchical Prompt (HP-DVAL), which leverages multi-modal knowledge from vision-language pretrained (VLP) models to mitigate the class-imbalance problem in multi-label settings. Specifically, HP-DVAL employs dual-view alignment learning to transfer the powerful feature representation capabilities from VLP models by extracting complementary features for accurate image-text alignment. To better adapt VLP models for CI-MLIC tasks, we introduce a hierarchical prompt-tuning strategy that utilizes global and local prompts to learn task-specific and context-related prior knowledge. Additionally, we design a semantic consistency loss during prompt tuning to prevent learned prompts from deviating from general knowledge embedded in VLP models. The effectiveness of our approach is validated on two CI-MLIC benchmarks: MS-COCO and VOC2007. Extensive experimental results demonstrate the superiority of our method over SOTA approaches, achieving mAP improvements of 10.0% and 5.2% on the long-tailed multi-label image classification task, and 6.8% and 2.9% on the multi-label few-shot image classification task. Sheng Huang 0001, Jiexuan Yan, Beiyan Liu, Bo Liu 0005, Richang Hong |
IEEE Trans. Image Process. | 4 |
| 2025 | Learning Self-Corrective Network via Adaptive Self-Labeling and Dynamic NMS for High-Performance Long-Term TrackingabstractThis article presents a self-corrective network-based long-term tracker (SCLT) including a self-modulated tracking reliability evaluator (STRE) and a self-adjusting proposal postprocessor (SPPP). The targets in the long-term sequences often suffer from severe appearance variations. Existing long-term trackers often online update their models to adapt the variations, but the inaccurate tracking results introduce cumulative error into the updated model that may cause severe drift issue. To this end, a robust long-term tracker should have the self-corrective capability that can judge whether the tracking result is reliable or not, and then it is able to recapture the target when severe drift happens caused by serious challenges (e.g., full occlusion and out-of-view). To address the first issue, the STRE designs an effective tracking reliability classifier that is built on a modulation subnetwork. The classifier is trained using the samples with pseudo labels generated by an adaptive self-labeling strategy. The adaptive self-labeling can automatically label the hard negative samples that are often neglected in existing trackers according to the statistical characteristics of target state, and the network modulation mechanism can guide the backbone network to learn more discriminative features without extra training data. To address the second issue, after the STRE has been triggered, the SPPP follows it with a dynamic NMS to recapture the target in time and accurately. In addition, the STRE and the SPPP demonstrate good transportability ability, and their performance is improved when combined with multiple baselines. Compared to the commonly used greedy NMS, the proposed dynamic NMS leverages an adaptive strategy to effectively handle the different conditions of in view and out of view, thereby being able to select the most probable object box that is essential to accurately online update the basic tracker. Extensive evaluations on four large-scale and challenging benchmark datasets including VOT2021LT, OxUvALT, TLP, and LaSOT demonstrate superiority of the proposed SCLT to a variety of state-of-the-art long-term trackers in terms of all measures. Source codes and demos can be found at https://github.com/TJUT-CV/SCLT. Wanli Xue, Kaihua Zhang 0001, Bo Liu 0005, Chengwei Zhang 0001, Jingen Liu, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Generalizable Fourier Augmentation for Unsupervised Video Object SegmentationabstractThe performance of existing unsupervised video object segmentation methods typically suffers from severe performance degradation on test videos when tested in out-of-distribution scenarios. The primary reason is that the test data in real- world may not follow the independent and identically distribution (i.i.d.) assumption, leading to domain shift. In this paper, we propose a generalizable fourier augmentation method during training to improve the generalization ability of the model. To achieve this, we perform Fast Fourier Transform (FFT) over the intermediate spatial domain features in each layer to yield corresponding frequency representations, including amplitude components (encoding scene-aware styles such as texture, color, contrast of the scene) and phase components (encoding rich semantics). We produce a variety of style features via Gaussian sampling to augment the training data, thereby improving the generalization capability of the model. To further improve the cross-domain generalization performance of the model, we design a phase feature update strategy via exponential moving average using phase features from past frames in an online update manner, which could help the model to learn cross-domain-invariant features. Extensive experiments show that our proposed method achieves the state-of-the-art performance on popular benchmarks. Huihui Song 0003, Tiankang Su, Yuhui Zheng, Kaihua Zhang 0001, Bo Liu 0005, Dong Liu 0002 |
AAAI | 5 |
| 2024 | Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyabstractMultiple instance learning (MIL) is the most widely used framework in computational pathology, encompassing sub-typing, diagnosis, prognosis, and more. However, the ex-isting MIL paradigm typically requires an offline instance feature extractor, such as a pre-trained ResNet or a foun-dation model. This approach lacks the capability for feature fine-tuning within the specific downstream tasks, limiting its adaptability and performance. To address this issue, we propose a Re-embedded Regional Transformer (R2T) for re-embedding the instance features online, which captures fine-grained local features and establishes connections across different regions. Unlike existing works that focus on pre-training powerful feature extractor or designing sophisticated instance aggregator, R2T is tailored to re-embed instance features online. It serves as a portable module that can seamlessly integrate into mainstream MIL models. Extensive experimental results on common computational pathology tasks validate that: 1) feature re-embedding improves the performance of MIL models based on ResNet-50 features to the level of foundation model features, and further enhances the performance of foundation model features; 2) the R2T can introduce more signifi-cant performance improvements to various MIL models; 3) R2T-MIL, as an R2T-enhanced AB-MIL, outperforms other latest methods by a large margin. The code is available at: https://github.com/DearCaat/RRT-MIL. Fengtao Zhou, Sheng Huang 0001, Yi Zhang 0113, Bo Liu 0005 |
CVPR | 6 |
| 2024 | SAM-MIL: A Spatial Contextual Aware Multiple Instance Learning Approach for Whole Slide Image ClassificationabstractMultiple Instance Learning (MIL) represents the predominant framework in Whole Slide Image (WSI) classification, covering aspects such as sub-typing, diagnosis, and beyond. Current MIL models predominantly rely on instance-level features derived from pretrained models such as ResNet. These models segment each WSI into independent patches and extract features from these local patches, leading to a significant loss of global spatial context and restricting the model's focus to merely local features. To address this issue, we propose a novel MIL framework, named SAM-MIL, that emphasizes spatial contextual awareness and explicitly incorporates spatial context by extracting comprehensive, image-level information. The Segment Anything Model (SAM) represents a pioneering visual segmentation foundational model that can capture segmentation features without the need for additional fine-tuning, rendering it an outstanding tool for extracting spatial context directly from raw WSIs. Our approach includes the design of group feature extraction based on spatial context and a SAM-Guided Group Masking strategy to mitigate class imbalance issues. We implement a dynamic mask ratio for different segmentation categories and supplement these with representative group features of categories. Moreover, SAM-MIL divides instances to generate additional pseudo-bags, thereby augmenting the training set, and introduces consistency of spatial context across pseudo-bags to further enhance the model's performance. Experimental results on the CAMELYON-16 and TCGA Lung Cancer datasets demonstrate that our proposed SAM-MIL model outperforms existing mainstream methods in WSIs classification. Our open-source implementation code is is available at https://github.com/FangHeng/SAM-MIL. Sheng Huang 0001, Luwen Huangfu, Bo Liu 0005 |
ACM Multimedia | 5 |
| 2024 | Category-Prompt Refined Feature Learning for Long-Tailed Multi-Label Image ClassificationabstractReal-world data consistently exhibits a long-tailed distribution, often spanning multiple categories. This complexity underscores the challenge of content comprehension, particularly in scenarios requiring Long-Tailed Multi-Label image Classification (LTMLC). In such contexts, imbalanced data distribution and multi-object recognition pose significant hurdles. To address this issue, we propose a novel and effective approach for LTMLC, termed Category-Prompt Refined Feature Learning (CPRFL), utilizing semantic correlations between different categories and decoupling category-specific visual representations for each category. Specifically, CPRFL initializes category-prompts from the pretrained CLIP's embeddings and decouples category-specific visual representations through interaction with visual features, thereby facilitating the establishment of semantic correlations between the head and tail classes. To mitigate the visual-semantic domain bias, we design a progressive Dual-Path Back-Propagation mechanism to refine the prompts by progressively incorporating context-related visual information into prompts. Simultaneously, the refinement process facilitates the progressive purification of the category-specific visual representations under the guidance of the refined prompts. Furthermore, taking into account the negative-positive sample imbalance, we adopt the Asymmetric Loss as our optimization objective to suppress negative samples across all classes and potentially enhance the head-to-tail recognition performance. We validate the effectiveness of our method on two LTMLC benchmarks and extensive experiments demonstrate the superiority of our work over baselines.The code is available at https://github.com/jiexuanyan/CPRFL. Jiexuan Yan, Sheng Huang 0001, Nankun Mu, Luwen Huangfu, Bo Liu 0005 |
ACM Multimedia | 5 |
| 2024 | Dual temporal memory network with high-order spatio-temporal graph learning for video object segmentation
Jiaqing Fan, Shenglong Hu, Kaihua Zhang 0001, Bo Liu 0005 |
Image Vis. Comput. | 5 |
| 2024 | Text kernel expansion for real-time scene text detection
Sheng Huang 0001, Bo Liu 0005 |
Pattern Anal. Appl. | 4 |
| 2024 | Gloss Prior Guided Visual Feature Learning for Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) is to recognize the glosses in a sign language video. Enhancing the generalization ability of CSLR's visual feature extractor is a worthy area of investigation. In this paper, we model glosses as priors that help to learn more generalizable visual features. Specifically, the signer-invariant gloss feature is extracted by a pre-trained gloss BERT model. Then we design a gloss prior guidance network (GPGN). It contains a novel parallel densely-connected temporal feature extraction (PDC-TFE) module for multi-resolution visual feature extraction. The PDC-TFE captures the complex temporal patterns of the glosses. The pre-trained gloss feature guides the visual feature learning through a cross-modality matching loss. We propose to formulate the cross-modality feature matching into a regularized optimal transport problem, it can be efficiently solved by a variant of the Sinkhorn algorithm. The GPGN parameters are learned by optimizing a weighted sum of the cross-modality matching loss and CTC loss. The experiment results on German and Chinese sign language benchmarks demonstrate that the proposed GPGN achieves competitive performance. The ablation study verifies the effectiveness of several critical components of the GPGN. Furthermore, the proposed pre-trained gloss BERT model and cross-modality matching can be seamlessly integrated into other RGB-cue-based CSLR methods as plug-and-play formulations to enhance the generalization ability of the visual feature extractor. Leming Guo, Wanli Xue, Bo Liu 0005, Kaihua Zhang 0001, Tiantian Yuan, Dimitris N. Metaxas |
IEEE Trans. Image Process. | 3 |
| 2023 | Distilling Cross-Temporal Contexts for Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) aims to recognize glosses in a sign language video. State-of-the-art methods typically have two modules, a spatial perception module and a temporal aggregation module, which are jointly learned end-to-end. Existing results in [9, 20, 25, 36] have indicated that, as the frontal component of the over-all model, the spatial perception module used for spatial feature extraction tends to be insufficiently trained. In this paper, we first conduct empirical studies and show that a shallow temporal aggregation module allows more thor-ough training of the spatial perception module. However, a shallow temporal aggregation module cannot well capture both local and global temporal context information in sign language. To address this dilemma, we propose a cross-temporal context aggregation (CTCA) model. Specifically, we build a dual-path network that contains two branches for perceptions of local temporal context and global temporal context. We further design a cross-context knowledge distil-lation learning objective to aggregate the two types of con-text and the linguistic prior. The knowledge distillation en-ables the resultant one-branch temporal aggregation mod-ule to perceive local-global temporal and semantic context. This shallow temporal perception module structure facili-tates spatial perception module learning. Extensive exper-iments on challenging CSLR benchmarks demonstrate that our method outperforms all state-of-the-art methods. Leming Guo, Wanli Xue, Qing Guo 0005, Bo Liu 0005, Kaihua Zhang 0001, Tiantian Yuan, Shengyong Chen |
CVPR | 4 |
| 2023 | Co-Salient Object Detection with Uncertainty-Aware Group Exchange-MaskingabstractThe traditional definition of co-salient object detection (CoSOD) task is to segment the common salient objects in a group of relevant images. Existing CoSOD models by-default adopt the group consensus assumption. This brings about model robustness defect under the condition of irrelevant images in the testing image group, which hinders the use of CoSOD models in real-world applications. To address this issue, this paper presents a group exchange-masking (GEM) strategy for robust CoSOD model learning. With two group of image containing different types of salient object as input, the GEM first selects a set of images from each group by the proposed learning based strategy, then these images are exchanged. The proposed feature extraction module considers both the uncertainty caused by the irrelevant images and group consensus in the remaining relevant images. We design a latent variable generator branch which is made of conditional variational autoencoder to generate uncertainly-based global stochastic features. A CoSOD transformer branch is devised to capture the correlation-based local features that contain the group consistency information. At last, the output of two branches are concatenated and fed into a transformer-based decoder, producing robust co-saliency prediction. Extensive evaluations on co-saliency detection with and without irrelevant images demonstrate the superiority of our method over a variety of state-of-the-art methods. Huihui Song 0003, Bo Liu 0005, Kaihua Zhang 0001, Dong Liu 0002 |
CVPR | 3 |
| 2023 | Unsupervised Video Object Segmentation with Online Adversarial Self-TuningabstractThe existing unsupervised video object segmentation methods depend heavily on the segmentation model trained offline on a labeled training video set, and cannot well generalize to the test videos from a different domain with possible distribution shifts. We propose to perform online fine-tuning on the pre-trained segmentation model to adapt to any ad-hoc videos at the test time. To achieve this, we design an offline semi-supervised adversarial training process, which leverages the unlabeled video frames to improve the model generalizability while aligning the features of the labeled video frames with the features of the unlabeled video frames. With the trained segmentation model, we further conduct an online self-supervised adversarial finetuning, in which a teacher model and a student model are first initialized with the pre-trained segmentation model weights, and the pseudo label produced by the teacher model is used to supervise the student model in an adversarial learning framework. Through online finetuning, the student model is progressively updated according to the emerging patterns in each test video, which significantly reduces the test-time domain gap. We integrate our offline training and online fine-tuning in a unified framework for unsupervised video object segmentation and dub our method Online Adversarial Self-Tuning (OAST). The experiments show that our method outperforms the state-of-the-arts with significant gains on the popular video object segmentation datasets. Tiankang Su, Huihui Song 0003, Dong Liu 0002, Bo Liu 0005, Qingshan Liu 0001 |
ICCV | 4 |
| 2023 | Multiple Instance Learning Framework with Masked Hard Instance Mining for Whole Slide Image ClassificationabstractThe whole slide image (WSI) classification is often formulated as a multiple instance learning (MIL) problem. Since the positive tissue is only a small fraction of the gigapixel WSI, existing MIL methods intuitively focus on identifying salient instances via attention mechanisms. However, this leads to a bias towards easy-to-classify instances while neglecting hard-to-classify instances. Some literature has revealed that hard examples are beneficial for modeling a discriminative boundary accurately. By applying such an idea at the instance level, we elaborate a novel MIL framework with masked hard instance mining (MHIM-MIL), which uses a Siamese structure (Teacher-Student) with a consistency constraint to explore the potential hard instances. With several instance masking strategies based on attention scores, MHIM-MIL employs a momentum teacher to implicitly mine hard instances for training the student model, which can be any attention-based MIL model. This counter-intuitive strategy essentially enables the student to learn a better discriminating boundary. Moreover, the student is used to update the teacher with an exponential moving average (EMA), which in turn identifies new hard instances for subsequent training iterations and stabilizes the optimization. Experimental results on the CAMELYON-16 and TCGA Lung Cancer datasets demonstrate that MHIM-MIL outperforms other latest methods in terms of performance and training cost. The code is available at: https://github.com/DearCaat/MHIM-MIL. Sheng Huang 0001, Xiaoxian Zhang, Fengtao Zhou, Yi Zhang 0113, Bo Liu 0005 |
ICCV | 6 |
| 2023 | DMCVR: Morphology-Guided Diffusion Model for 3D Cardiac Volume Reconstruction
Xiaoxiao He, Chaowei Tan, Ligong Han, Bo Liu 0005, Leon Axel, Kang Li 0004, Dimitris N. Metaxas |
MICCAI (7) | 4 |
| 2023 | Temporally Efficient Gabor Transformer for Unsupervised Video Object SegmentationabstractSpatial-temporal structural details of targets in video (e.g. varying edges, textures over time) are essential to accurate Unsupervised Video Object Segmentation (UVOS). The vanilla multi-head self-attention in the Transformer-based UVOS methods usually concentrates on learning the general low-frequency information (e.g. illumination, color), while neglecting the high-frequency texture details, leading to unsatisfying segmentation results. To address this issue, this paper presents a Temporally efficient Gabor Transformer (TGFormer) for UVOS. The TGFormer jointly models the spatial dependencies and temporal coherence intra- and inter-frames, which can fully capture the rich structural details for accurate UVOS. Concretely, we first propose an effective learnable Gabor filtering Transformer to mine the structural texture details of the object for accurate UVOS. Then, to adaptively store the redundant neighboring historical information, we present an efficient dynamic neighboring frame selection module to automatically choose the useful temporal information, which simultaneously relieves the blurry frame and reduces the computation burden. Finally, we make the UVOS model be a fully Transformer architecture, meanwhile aggregating the information from space, Gabor and time domains, yielding a strong representation with rich structure details. Extensive experiments on five mainstream UVOS benchmarks (DAVIS2016, FBMS, DAVSOD, ViSal, and MCL) demonstrate the superiority of the presented solution to sate-of-the-art methods. Jiaqing Fan, Tiankang Su, Kaihua Zhang 0001, Bo Liu 0005, Qingshan Liu 0001 |
ACM Multimedia | 4 |
| 2023 | Click-Conversion Multi-Task Model with Position Bias Mitigation for Sponsored Search in eCommerceabstractPosition bias, the phenomenon whereby users tend to focus on higher-ranked items of the search result list regardless of the actual relevance to queries, is prevailing in many ranking systems. Position bias in training data biases the ranking model, leading to increasingly unfair item rankings, click-through-rate (CTR), and conversion rate (CVR) predictions. To jointly mitigate position bias in both item CTR and CVR prediction, we propose two position-bias-free CTR and CVR prediction models: Position-Aware Click-Conversion (PACC) and PACC via Position Embedding (PACC-PE). PACC is built upon probability decomposition and models position information as a probability. PACC-PE utilizes neural networks to model product-specific position information as embedding. Experiments on the E-commerce sponsored product search dataset show that our proposed models have better ranking effectiveness and can greatly alleviate position bias in both CTR and CVR prediction. Yibo Wang 0001, Yanbing Xue, Bo Liu 0005, Musen Wen, Wenting Zhao 0006, Stephen D. Guo, Philip S. Yu |
SIGIR | 3 |
| 2023 | Intellectual property protection for deep semantic segmentation models
Hongjia Ruan, Huihui Song 0003, Bo Liu 0005, Yong Cheng 0002, Qingshan Liu 0001 |
Frontiers Comput. Sci. | 3 |
| 2023 | Deep Object Co-Segmentation and Co-Saliency Detection via High-Order Spatial-Semantic Network ModulationabstractObject co-segmentation (CSG) is to segment the common objects of the same category in multiple relevant images while the co-saliency detection (CSD) aims to discover the salient and common foreground objects in a group of images. To process both tasks simultaneously, this paper presents an adaptive spatially and high-order semantically modulated deep network framework. A backbone network is first adopted to extract multi-resolution image features. With the multi-resolution features of the relevant images as input, we design an adaptive spatial modulator to learn a spatial representation that can highlight the co-object regions for each image. The adaptive spatial modulator fully captures the rich correlations of all image feature descriptors via unsupervised clustering and a graph aggregation strategy. The learned representation can well localize the common foreground object while effectively suppressing the background signals. For the high-order semantic modulator, we model it as a supervised image classification task. We propose a hierarchical high-order pooling module to learn the rich semantic features for classification use. The outputs of the two modulators manipulate the multi-resolution features by a shift-and-scale operation so that the features focus on segmenting common object regions. The proposed model is trained end-to-end without any intricate post-processing. Extensive experiments on three CSG benchmark datasets (MSRC, i-Coseg, and PASCAL-VOC) and three CSD datasets (Cosal2015, CoCA, and CoSOD3k) demonstrate the superior accuracy of the proposed method compared to state-of-the-art methods on both tasks. Kaihua Zhang 0001, Mingliang Dong, Bo Liu 0005, Dong Liu 0002, Qingshan Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | Boosting Multi-Label Image Classification with Complementary Parallel Self-DistillationabstractMulti-Label Image Classification (MLIC) appro-aches usually exploit label correlations to achieve good performance. However, emphasizing correlation like co-occurrence may overlook discriminative features and lead to model overfitting. In this study, we propose a generic framework named Parallel Self-Distillation (PSD) for boosting MLIC models. PSD decomposes the original MLIC task into several simpler MLIC sub-tasks via two elaborated complementary task decomposition strategies named Co-occurrence Graph Partition (CGP) and Dis-occurrence Graph Partition (DGP). Then, the MLIC models of fewer categories are trained with these sub-tasks in parallel for respectively learning the joint patterns and the category-specific patterns of labels. Finally, knowledge distillation is leveraged to learn a compact global ensemble of full categories with these learned patterns for reconciling the label correlation exploitation and model overfitting. Extensive results on MS-COCO and NUS-WIDE datasets demonstrate that our framework can be easily plugged into many MLIC approaches and improve performances of recent state-of-the-art approaches. The source code is released at https://github.com/Robbie-Xu/CPSD. Jiazhi Xu, Sheng Huang 0001, Fengtao Zhou, Luwen Huangfu, Daniel Dajun Zeng, Bo Liu 0005 |
IJCAI | 6 |
| 2022 | Multi-label out-of-distribution detection via exploiting sparsity and co-occurrence of labels
Lei Wang 0062, Sheng Huang 0001, Luwen Huangfu, Bo Liu 0005, Xiaohong Zhang 0002 |
Image Vis. Comput. | 4 |
| 2022 | Learning interlaced sparse Sinkhorn matching network for video super-resolution
Huihui Song 0003, Yutong Jin, Yong Cheng 0002, Bo Liu 0005, Dong Liu 0002, Qingshan Liu 0001 |
Pattern Recognit. | 4 |
| 2022 | Semi-Supervised Video Object Segmentation via Learning Object-Aware Global-Local CorrespondenceabstractIn semi-supervised video object segmentation (VOS) task, temporal coherent object-level cues play a key role yet are hard to accurately model. To this end, this paper presents an object-aware global-local correspondence architecture, which enables to extract the inter-frame temporal coherent object-level features for accurate VOS. Specifically, we first generate a set of object masks by the ground-truth segmentation, and then we squeeze the current frame representation inside the object masks into a set of global object embeddings. Second, we compute the similarity between each embedding and the feature map, producing an object-aware weight for each pixel. The object-aware feature at each pixel is then constructed by summing the object embeddings weighted by their corresponding object-aware weights, which is able to capture rich object category information. Third, to establish the accurate correspondences between the inter-frame temporal coherent cues, we further design a novel global-local correspondence module to refine the temporal feature representations. Finally, we augment the object-aware features with the global-local aligned information to produce a strong spatio-temporal representation, which is essential to a more reliable pixel-wise segmentation prediction. Extensive evaluations are conducted on three popular VOS benchmarks containing Youtube-VOS, Davis2017 and Davis2016, demonstrating that the proposed method achieves favourable performance compared to the state-of-the-arts. Jiaqing Fan, Bo Liu 0005, Kaihua Zhang 0001, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Multi-Label Image Classification via Category Prototype Compositional LearningabstractReal-world images are often compositions of multiple objects with different categories, scales, poses and locations. Adding nonexistent objects to an image (composing) or removing existent objects from an image (decomposing) leads to higher discrepancy in appearance, which reveals an important but long-neglected compositional nature of multi-label images. In light of this observation, we propose a novel end-to-end compositional learning framework named Category Prototype Compositional Learning (CPCL) to model such compositional nature for multi-label image classification. In CPCL, each image is represented by a collection of category-related features used to eliminate the negative effects from location information. Then, a compositional learning module is introduced to compose and decompose the category-related features with their corresponding category prototypes, which are derived from the semantic representations of categories. If the image has the given object, the output after composing should be closer to the original input than the output after decomposing. Contrarily, if the image does not have the given object, the output after decomposing should be closer to the original input than the output after composing. We introduce the Transformed Appearance Distance (TAD) to measure the appearance change between the composed and decomposed features relative to the category-related features with respect to each category. Finally, multi-label image classification is accomplished by performing a TAD-based metric learning. Experimental results on three multi-label image classification benchmarks,i.e., NUS-WIDE, MS-COCO and VOC 2007, validate the effectiveness and superiority of our work in comparison with the state-of-the-arts. The source codes of our model have been released onhttps://github.com/ZFT-CQU/CPCL. Fengtao Zhou, Sheng Huang 0001, Bo Liu 0005, Dan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Image Co-Saliency Detection and Instance Co-Segmentation Using Attention Graph Clustering Based Graph Convolutional NetworkabstractCo-Saliency Detection (CSD) is to explore the concurrent patterns and salient objects from a group of relevant images, while Instance Co-Segmentation (ICS) aims to identify and segment out all of these co-salient instances, generating corresponding mask for each instance. To simultaneously tackle these two tasks, we present a novel adaptive graph convolutional network with attention graph clustering (GCAGC) for CSD and ICS, termed as GCAGC-CSD and GCAGC-ICS, respectively. The GCAGC-CSD contains three key model designs: first, we develop a graph convolutional network architecture to extract multi-scale representations to characterize the intra- and inter-image consistency. Second, we propose an attention graph clustering algorithm to distinguish the salient foreground objects from common areas in an unsupervised manner. Third, we present a unified framework with encoder-decoder structure to jointly train and optimize the graph convolutional network, attention graph cluster, and CSD decoder in an end-to-end fashion. Afterwards, we design a salient instance segmentation network for GCAGC-ICS, and combine the outputs of GCAGC-CSD and the instance segmentation branch to obtain instance-aware co-segmentation masks. The proposed GCAGC-CSD and GCAGC-ICS are extensively evaluated on four CSD benchmark datasets (iCoseg, Cosal2015, COCO-SEG and CoSOD3k) and five ICS benchmark datasets (CoSOD3k, COCO-NONVOC, COCO-VOC, VOC12 and SOC), and achieve superior performance over state-of-the-arts on both tasks. Tengpeng Li, Kaihua Zhang 0001, Shiwen Shen, Bo Liu 0005, Qingshan Liu 0001, Zhu Li 0001 |
IEEE Trans. Multim. | 4 |
| 2021 | LRSC: Learning Representations for Subspace ClusteringabstractDeep learning based subspace clustering methods have attracted increasing attention in recent years, where a basic theme is to non-linearly map data into a latent space, and then uncover subspace structures based upon the data self-expressiveness property. However, almost all existing deep subspace clustering methods only rely on target domain data, and always resort to shallow neural networks for modeling data, leaving huge room to design more effective representation learning mechanisms tailored for subspace clustering. In this paper, we propose a novel subspace clustering framework through learning precise sample representations. In contrast to previous approaches, the proposed method aims to leverage external data through constructing lots of relevant tasks to guide the training of the encoder, motivated by the idea of meta-learning. Considering limited layer structures of current deep subspace clustering models, we intend to distill knowledge from a deeper network trained on the external data, and transfer it into the shallower model. To reach the above two goals, we propose a new loss function to realize them in a joint framework. Moreover, we propose to construct a new pretext task for self-supervised training of the model, such that the representation ability of the model can be further improved. Extensive experiments are performed on four publicly available datasets, and experimental results clearly demonstrate the efficacy of our method, compared to state-of-the-art methods. Bo Liu 0005, Ye Yuan 0001, Guoren Wang |
AAAI | 3 |
| 2021 | DeepACG: Co-Saliency Detection via Semantic-Aware Contrast Gromov-Wasserstein DistanceabstractThe objective of co-saliency detection is to segment the co-occurring salient objects in a group of images. To address this task, we introduce a new deep network architecture via semantic-aware contrast Gromov-Wasserstein distance (DeepACG). We first adopt the Gromov-Wasserstein (GW) distance to build dense 4D correlation volumes for all pairs of image pixels within the image group. These dense correlation volumes enable the network to accurately discover the structured pair-wise pixel similarities among the common salient objects. Second, we develop a semantic-aware co-attention module (SCAM) to enhance the foreground co-saliency through predicted categorical information. Specifically, SCAM recognizes the semantic class of the foreground co-objects, and this information is then modulated to the deep representations to localize the related pixels. Third, we design a contrast edge-enhanced module (EEM) to capture richer contexts and preserve fine-grained spatial information. We validate the effectiveness of our model using three largest and most challenging benchmark datasets (Cosal2015, CoCA, and CoSOD3k). Extensive experiments have demonstrated the substantial practical merit of each module. Compared with the existing works, DeepACG shows significant improvements and achieves state-of-the-art performance. Kaihua Zhang 0001, Mingliang Dong, Bo Liu 0005, Xiao-Tong Yuan, Qingshan Liu 0001 |
CVPR | 3 |
| 2021 | Deep Transport Network for Unsupervised Video Object SegmentationabstractThe popular unsupervised video object segmentation methods fuse the RGB frame and optical flow via a two-stream network. However, they cannot handle the distracting noises in each input modality, which may vastly deteriorate the model performance. We propose to establish the correspondence between the input modalities while suppressing the distracting signals via optimal structural matching. Given a video frame, we extract the dense local features from the RGB image and optical flow, and treat them as two complex structured representations. The Wasserstein distance is then employed to compute the global optimal flows to transport the features in one modality to the other, where the magnitude of each flow measures the extent of the alignment between two local features. To plug the structural matching into a two-stream network for end-to-end training, we factorize the input cost matrix into small spatial blocks and design a differentiable long-short Sinkhorn module consisting of a long-distant Sinkhorn layer and a short-distant Sinkhorn layer. We integrate the module into a dedicated two-stream network and dub our model TransportNet. Our experiments show that aligning motion-appearance yields the state-of-the-art results on the popular video object segmentation datasets. Kaihua Zhang 0001, Zicheng Zhao, Dong Liu 0002, Qingshan Liu 0001, Bo Liu 0005 |
ICCV | 5 |
| 2021 | Oriented Object Detection in Aerial Images with Box Boundary-Aware VectorsabstractOriented object detection in aerial images is a challenging task as the objects in aerial images are displayed in arbitrary directions and are usually densely packed. Cur-rent oriented object detection methods mainly rely on two-stage anchor-based detectors. However, the anchor-based detectors typically suffer from a severe imbalance issue be-tween the positive and negative anchor boxes. To address this issue, in this work we extend the horizontal keypoint-based object detector to the oriented object detection task. In particular, we first detect the center keypoints of the objects, based on which we then regress the box boundary-aware vectors (BBAVectors) to capture the oriented bounding boxes. The box boundary-aware vectors are distributed in the four quadrants of a Cartesian coordinate system for all arbitrarily oriented objects. To relieve the difficulty of learning the vectors in the corner cases, we further classify the oriented bounding boxes into horizontal and rotational bounding boxes. In the experiment, we show that learning the box boundary-aware vectors is superior to directly predicting the width, height, and angle of an oriented bounding box, as adopted in the baseline method. Besides, the proposed method competes favorably with state-of-the-art methods. Code is available at https://github.com/yijingru/BBAVectors-Oriented-Object-Detection. Jingru Yi, Pengxiang Wu, Bo Liu 0005, Qiaoying Huang, Dimitris N. Metaxas |
WACV | 3 |
| 2021 | A lightweight multi-scale aggregated model for detecting aerial images captured by UAVs
Zhaokun Li, Xueliang Liu, Ye Zhao 0001, Bo Liu 0005, Zhen Huang 0006, Richang Hong |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Video saliency prediction using enhanced spatiotemporal alignment network
Huihui Song 0003, Kaihua Zhang 0001, Bo Liu 0005, Qingshan Liu 0001 |
Pattern Recognit. | 4 |
| 2021 | Multi-Stage Feature Fusion Network for Video Super-ResolutionabstractVideo super-resolution (VSR) is to restore a photo-realistic high-resolution (HR) frame from both its corresponding low-resolution (LR) frame (reference frame) and multiple neighboring frames (supporting frames). An important step in VSR is to fuse the feature of the reference frame with the features of the supporting frames. The major issue with existing VSR methods is that the fusion is conducted in a one-stage manner, and the fused feature may deviate greatly from the visual information in the original LR reference frame. In this paper, we propose an end-to-end Multi-Stage Feature Fusion Network that fuses the temporally aligned features of the supporting frames and the spatial feature of the original reference frame at different stages of a feed-forward neural network architecture. In our network, the Temporal Alignment Branch is designed as an inter-frame temporal alignment module used to mitigate the misalignment between the supporting frames and the reference frame. Specifically, we apply the multi-scale dilated deformable convolution as the basic operation to generate temporally aligned features of the supporting frames. Afterwards, the Modulative Feature Fusion Branch, the other branch of our network accepts the temporally aligned feature map as a conditional input and modulates the feature of the reference frame at different stages of the branch backbone. This enables the feature of the reference frame to be referenced at each stage of the feature fusion process, leading to an enhanced feature from LR to HR. Experimental results on several benchmark datasets demonstrate that our proposed method can achieve state-of-the-art performance on VSR task. Huihui Song 0003, Dong Liu 0002, Bo Liu 0005, Qingshan Liu 0001, Dimitris N. Metaxas |
IEEE Trans. Image Process. | 4 |
| 2021 | Object-Guided Instance Segmentation With Auxiliary Feature Refinement for Biological ImagesabstractInstance segmentation is of great importance for many biological applications, such as study of neural cell interactions, plant phenotyping, and quantitatively measuring how cells react to drug treatment. In this paper, we propose a novel box-based instance segmentation method. Box-based instance segmentation methods capture objects via bounding boxes and then perform individual segmentation within each bounding box region. However, existing methods can hardly differentiate the target from its neighboring objects within the same bounding box region due to their similar textures and low-contrast boundaries. To deal with this problem, in this paper, we propose an object-guided instance segmentation method. Our method first detects the center points of the objects, from which the bounding box parameters are then predicted. To perform segmentation, an object-guided coarse-to-fine segmentation branch is built along with the detection branch. The segmentation branch reuses the object features as guidance to separate target object from the neighboring ones within the same bounding box region. To further improve the segmentation quality, we design an auxiliary feature refinement module that densely samples and refines point-wise features in the boundary regions. Experimental results on three biological image datasets demonstrate the advantages of our method. The code will be available at https://github.com/yijingru/ObjGuided-Instance-Segmentation. Jingru Yi, Pengxiang Wu, Bo Liu 0005, Qiaoying Huang, Lianyi Han, Wei Fan 0001, Daniel J. Hoeppner, Dimitris N. Metaxas |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Robust Conditional GAN from Uncertainty-Aware Pairwise ComparisonsabstractConditional generative adversarial networks have shown exceptional generation performance over the past few years. However, they require large numbers of annotations. To address this problem, we propose a novel generative adversarial network utilizing weak supervision in the form of pairwise comparisons (PC-GAN) for image attribute editing. In the light of Bayesian uncertainty estimation and noise-tolerant adversarial training, PC-GAN can estimate attribute rating efficiently and demonstrate robust performance in noise resistance. Through extensive experiments, we show both qualitatively and quantitatively that PC-GAN performs comparably with fully-supervised methods and outperforms unsupervised baselines. Code and Supplementary can be found on the project website*. Ligong Han, Ruijiang Gao, Mun Kim, Bo Liu 0005, Dimitris N. Metaxas |
AAAI | 5 |
| 2020 | Object-Guided Instance Segmentation for Biological ImagesabstractInstance segmentation of biological images is essential for studying object behaviors and properties. The challenges, such as clustering, occlusion, and adhesion problems of the objects, make instance segmentation a non-trivial task. Current box-free instance segmentation methods typically rely on local pixel-level information. Due to a lack of global object view, these methods are prone to over- or under-segmentation. On the contrary, the box-based instance segmentation methods incorporate object detection into the segmentation, performing better in identifying the individual instances. In this paper, we propose a new box-based instance segmentation method. Mainly, we locate the object bounding boxes from their center points. The object features are subsequently reused in the segmentation branch as a guide to separate the clustered instances within an RoI patch. Along with the instance normalization, the model is able to recover the target object distribution and suppress the distribution of neighboring attached objects. Consequently, the proposed model performs excellently in segmenting the clustered objects while retaining the target object details. The proposed method achieves state-of-the-art performances on three biological datasets: cell nuclei, plant phenotyping dataset, and neural cells. Jingru Yi, Pengxiang Wu, Bo Liu 0005, Daniel J. Hoeppner, Dimitris N. Metaxas, Lianyi Han, Wei Fan 0001 |
AAAI | 4 |
| 2020 | Deep Object Co-Segmentation via Spatial-Semantic Network ModulationabstractObject co-segmentation is to segment the shared objects in multiple relevant images, which has numerous applications in computer vision. This paper presents a spatial and semantic modulated deep network framework for object co-segmentation. A backbone network is adopted to extract multi-resolution image features. With the multi-resolution features of the relevant images as input, we design a spatial modulator to learn a mask for each image. The spatial modulator captures the correlations of image feature descriptors via unsupervised learning. The learned mask can roughly localize the shared foreground object while suppressing the background. For the semantic modulator, we model it as a supervised image classification task. We propose a hierarchical second-order pooling module to transform the image features for classification use. The outputs of the two modulators manipulate the multi-resolution features by a shift-and-scale operation so that the features focus on segmenting co-object regions. The proposed model is trained end-to-end without any intricate post-processing. Extensive experiments on four image co-segmentation benchmark datasets demonstrate the superior accuracy of the proposed method compared to state-of-the-art methods. The codes are available at http://kaihuazhang.net/. Kaihua Zhang 0001, Bo Liu 0005, Qingshan Liu 0001 |
AAAI | 3 |
| 2020 | Generative Adversarial Attributed Network Anomaly DetectionabstractAnomaly detection is a useful technique in many applications such as network security and fraud detection. Due to the insufficiency of anomaly samples as training data, it is usually formulated as an unsupervised model learning problem. In recent years there is a surge of adopting graph data structure in numerous applications. Detecting anomaly in an attributed network is more challenging than the sample based task because of the sample information representations in the form of graph nodes and edges. In this paper, we propose a generative adversarial attributed network (GAAN) anomaly detection framework. The fake graph nodes are generated by a generator module with Gaussian noise as input. An encoder module is employed to map both real and fake graph nodes into a latent space. To encode the graph structure information into the node latent representation, we compute the sample covariance matrix for real nodes and fake nodes respectively. A discriminator is trained to recognize whether two connected nodes are from the real or fake graph. With the learned encoder module output, an anomaly evaluation measurement considering the sample reconstruction error and real-sample identification confidence is employed to make prediction. We conduct extensive experiments on benchmark datasets and compare with state-of-the-art attributed graph anomaly detection methods. The superior AUC score demonstrates the effectiveness of the proposed method. Zhenxing Chen, Bo Liu 0005, Peng Dai 0001, Liefeng Bo |
CIKM | 2 |
| 2020 | Adaptive Graph Convolutional Network With Attention Graph Clustering for Co-Saliency DetectionabstractCo-saliency detection aims to discover the common and salient foregrounds from a group of relevant images. For this task, we present a novel adaptive graph convolutional network with attention graph clustering (GCAGC). Three major contributions have been made, and are experimentally shown to have substantial practical merits. First, we propose a graph convolutional network design to extract information cues to characterize the intra- and inter-image correspondence. Second, we develop an attention graph clustering algorithm to discriminate the common objects from all the salient foreground objects in an unsupervised fashion. Third, we present a unified framework with encoder-decoder structure to jointly train and optimize the graph convolutional network, attention graph cluster, and co-saliency detection decoder in an end-to-end manner. We evaluate our proposed GCAGC method on three co-saliency detection benchmark datasets (iCoseg, Cosal2015 and COCO-SEG). Our GCAGC method obtains significant improvements over the state-of-the-arts on most of them. Kaihua Zhang 0001, Tengpeng Li, Shiwen Shen, Bo Liu 0005, Qingshan Liu 0001 |
CVPR | 4 |
| 2020 | Meta-learning with Network Pruning
Hongduan Tian, Bo Liu 0005, Xiao-Tong Yuan, Qingshan Liu 0001 |
ECCV (19) | 2 |
| 2020 | Dual Temporal Memory Network for Efficient Video Object SegmentationabstractVideo Object Segmentation (VOS) is typically formulated in a semi-supervised setting. Given the ground-truth segmentation mask on the first frame, the task of VOS is to track and segment the single or multiple objects of interests in the rest frames of the video at the pixel level. One of the fundamental challenges in VOS is how to make the most use of the temporal information to boost the performance. We present an end-to-end network which stores short- and long-term video sequence information preceding the current frame as the temporal memories to address the temporal modeling in VOS. Our network consists of two temporal sub-networks including a short-term memory sub-network and a long-term memory sub-network. The short-term memory sub-network models the fine-grained spatial-temporal interactions between local regions across neighboring frames in video via a graph-based learning framework, which can well preserve the visual consistency of local regions over time. The long-term memory sub-network models the long-range evolution of object via a Simplified-Gated Recurrent Unit (S-GRU), making the segmentation be robust against occlusions and drift errors. In our experiments, we show that our proposed method achieves a favorable and competitive performance on three frequently-used VOS datasets, including DAVIS 2016, DAVIS 2017 and Youtube-VOS in terms of both speed and accuracy. Kaihua Zhang 0001, Dong Liu 0002, Bo Liu 0005, Qingshan Liu 0001, Zhu Li 0001 |
ACM Multimedia | 4 |
| 2020 | Hierarchical Representations with Discriminative Meta-filters in Dual Path Network for Tracking
Ning Wang 0020, Yuncong Yao, Wankou Yang, Kaihua Zhang 0001, Bo Liu 0005 |
PRCV (2) | 6 |
| 2020 | Dual Iterative Hard ThresholdingabstractIterative Hard Thresholding (IHT) is a popular class of first-order greedy selection methods for loss minimization under cardinality constraint. The existing IHT-style algorithms, however, are proposed for minimizing the primal formulation. It is still an open issue to explore duality theory and algorithms for such a non-convex and NP-hard combinatorial optimization problem. To address this issue, we develop in this article a novel duality theory for $\ell_2$-regularized empirical risk minimization under cardinality constraint, along with an IHT-style algorithm for dual optimization. Our sparse duality theory establishes a set of sufficient and/or necessary conditions under which the original non-convex problem can be equivalently or approximately solved in a concave dual formulation. In view of this theory, we propose the Dual IHT (DIHT) algorithm as a super-gradient ascent method to solve the non-smooth dual problem with provable guarantees on primal-dual gap convergence and sparsity recovery. Numerical results confirm our theoretical predictions and demonstrate the superiority of DIHT to the state-of-the-art primal IHT-style algorithms in model estimation accuracy and computational efficiency. Xiao-Tong Yuan, Bo Liu 0005, Lezi Wang, Qingshan Liu 0001, Dimitris N. Metaxas |
J. Mach. Learn. Res. | 2 |
| 2020 | Dual-Path Attention Network for Compressed Sensing Image ReconstructionabstractAlthough deep neural network methods achieved much success in compressed sensing image reconstruction in recent years, they still have some issues, especially in preserving texture details. In this paper, we propose a new dual-path attention network for compressed sensing image reconstruction, which is composed of a structure path, a texture path and a texture attention module. Motivated by the classical paradigm of image structure-texture decomposition, the structure path aims to reconstruct the dominant structure component of the original image, and the texture path targets at recovering the remaining texture details. To better bridge the information between two paths, the texture attention module is designed to deliver the useful structure information to the texture path and predict the texture region, thereby facilitating the recovery of texture details. Two paths are optimized with a unified loss function. In the testing phase, given the measurement vector of a new image, it can be well reconstructed by carrying out the well trained dual-path attention network and integrating the outputs of the structure path and the texture path. Experimental results on the SET5, SET11 and BSD68 testing datasets demonstrate that the proposed method achieves comparable or better results compared with some state-of-the-art deep learning based methods and conventional iterative optimization based methods in terms of reconstruction quality and robustness to noise. Yubao Sun, Qingshan Liu 0001, Bo Liu 0005, Guodong Guo |
IEEE Trans. Image Process. | 4 |
| 2019 | Distributed Inexact Newton-type Pursuit for Non-convex Sparse LearningabstractIn this paper, we present a sample distributed greedy pursuit method for non-convex sparse learning under cardinality constraint. Given the training samples uniformly randomly partitioned across multiple machines, the proposed method alternates between local inexact sparse minimization of a Newton-type approximation and centralized global results aggregation. Theoretical analysis shows that for a general class of convex functions with Lipschitze continues Hessian, the method converges linearly with contraction factor scaling inversely to the local data size; whilst the communication complexity required to reach desirable statistical accuracy scales logarithmically with respect to the number of machines for some popular statistical learning models. For nonconvex objective functions, up to a local estimation error, our method can be shown to converge to a local stationary sparse solution with sub-linear communication complexity. Numerical results demonstrate the efficiency and accuracy of our method when applied to large-scale sparse learning tasks including deep neural nets pruning Bo Liu 0005, Xiao-Tong Yuan, Lezi Wang, Qingshan Liu 0001, Junzhou Huang, Dimitris N. Metaxas |
AISTATS | 1 |
| 2019 | Co-Saliency Detection via Mask-Guided Fully Convolutional Networks With Multi-Scale Label SmoothingabstractIn image co-saliency detection problem, one critical issue is how to model the concurrent pattern of the co-salient parts, which appears both within each image and across all the relevant images. In this paper, we propose a hierarchical image co-saliency detection framework as a coarse to fine strategy to capture this pattern. We first propose a mask-guided fully convolutional network structure to generate the initial co-saliency detection result. The mask is used for background removal and it is learned from the high-level feature response maps of the pre-trained VGG-net output. We next propose a multi-scale label smoothing model to further refine the detection result. The proposed model jointly optimizes the label smoothness of pixels and superpixels. Experiment results on three popular image co-saliency detection benchmark datasets including iCoseg, MSRC and Cosal2015 demonstrate the remarkable performance compared with the state-of-the-art methods. Kaihua Zhang 0001, Tengpeng Li, Bo Liu 0005, Qingshan Liu 0001 |
CVPR | 3 |
| 2019 | Sharpen Focus: Learning With Attention Separability and ConsistencyabstractRecent developments in gradient-based attention modeling have seen attention maps emerge as a powerful tool for interpreting convolutional neural networks. Despite good localization for an individual class of interest, these techniques produce attention maps with substantially overlapping responses among different classes, leading to the problem of visual confusion and the need for discriminative attention. In this paper, we address this problem by means of a new framework that makes class-discriminative attention a principled part of the learning process. Our key innovations include new learning objectives for attention separability and cross-layer consistency, which result in improved attention discriminability and reduced visual confusion. Extensive experiments on image classification benchmarks show the effectiveness of our approach in terms of improved classification accuracy, including CIFAR-100 (+3.33%), Caltech-256 (+1.64%), ImageNet (+0.92%), CUB-200-2011 (+4.8%) and PASCAL VOC2012 (+5.73%). Lezi Wang, Ziyan Wu 0001, Srikrishna Karanam, Kuan-Chuan Peng, Rajat Vikram Singh, Bo Liu 0005, Dimitris N. Metaxas |
ICCV | 6 |
| 2019 | Multi-scale Cell Instance Segmentation with Keypoint Graph Based Bounding Boxes
Jingru Yi, Pengxiang Wu, Qiaoying Huang, Bo Liu 0005, Daniel J. Hoeppner, Dimitris N. Metaxas |
MICCAI (1) | 5 |
| 2017 | Dual Iterative Hard Thresholding: From Non-convex Sparse Minimization to Non-smooth Concave MaximizationabstractIterative Hard Thresholding (IHT) is a class of projected gradient descent methods for optimizing sparsity-constrained minimization models, with the best known efficiency and scalability in practice. As far as we know, the existing IHT-style methods are designed for sparse minimization in primal form. It remains open to explore duality theory and algorithms in such a non-convex and NP-hard setting. In this article, we bridge the gap by establishing a duality theory for sparsity-constrained minimization with $\ell_2$-regularized objective and proposing an IHT-style algorithm for dual maximization. Our sparse duality theory provides a set of sufficient and necessary conditions under which the original NP-hard/non-convex problem can be equivalently solved in a dual space. The proposed dual IHT algorithm is a super-gradient method for maximizing the non-smooth dual objective. An interesting finding is that the sparse recovery performance of dual IHT is invariant to the Restricted Isometry Property (RIP), which is required by all the existing primal IHT without sparsity relaxation. Moreover, a stochastic variant of dual IHT is proposed for large-scale stochastic optimization. Numerical results demonstrate that dual IHT algorithms can achieve more accurate model estimation given small number of training data and have higher computational efficiency than the state-of-the-art primal IHT-style algorithms. Bo Liu 0005, Xiao-Tong Yuan, Lezi Wang, Qingshan Liu 0001, Dimitris N. Metaxas |
ICML | 1 |
| 2017 | Parallel Sparse Subspace Clustering via Joint Sample and Parameter Blockwise PartitionabstractSparse subspace clustering (SSC) is a classical method to cluster data with specific subspace structure for each group. It has many desirable theoretical properties and has been shown to be effective in various applications. However, under the condition of a large-scale dataset, learning the sparse sample affinity graph is computationally expensive. To tackle the computation time cost challenge, we develop a memory-efficient parallel framework for computing SSC via an alternating direction method of multiplier (ADMM) algorithm. The proposed framework partitions the data matrix into column blocks and then decomposes the original problem into parallel multivariate Lasso regression subproblems and samplewise operations. The proposed method allows us to allocate multiple cores/machines for the processing of individual column blocks. We propose a stochastic optimization algorithm to minimize the objective function. Experimental results on real-world datasets demonstrate that the proposed blockwise ADMM framework is substantially more efficient than its matrix counterpart used by SSC, without sacrificing performance in applications. Moreover, our approach is directly applicable to parallel neighborhood selection for Gaussian graphical models structure estimation. Bo Liu 0005, Xiao-Tong Yuan, Yang Yu 0010, Qingshan Liu 0001, Dimitris N. Metaxas |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2016 | Decentralized Robust Subspace ClusteringabstractWe consider the problem of subspace clustering using the SSC (Sparse Subspace Clustering) approach, which has several desirable theoretical properties and has been shown to be effective in various computer vision applications.We develop a large scale distributed framework for the computation of SSC via an alternating direction method of multiplier (ADMM) algorithm. The proposed framework solves SSC in column blocks and only involves parallel multivariate Lasso regression subproblems and sample-wise operations. This appealing property allows us to allocate multiple cores/machines for the processing of individual column blocks.We evaluate our algorithm on a shared-memory architecture. Experimental results on real-world datasets confirm that the proposed block-wise ADMM framework is substantially more efficient than its matrix counterpart used by SSC,without sacrificing accuracy. Moreover, our approach is directly applicable to decentralized neighborhood selection for Gaussian graphical models structure estimation. Bo Liu 0005, Xiao-Tong Yuan, Yang Yu 0010, Qingshan Liu 0001, Dimitris N. Metaxas |
AAAI | 1 |
| 2016 | Efficient k-Support-Norm Regularized Minimization via Fully Corrective Frank-Wolfe Method
Bo Liu 0005, Xiao-Tong Yuan, Shaoting Zhang 0001, Qingshan Liu 0001, Dimitris N. Metaxas |
IJCAI | 1 |
| 2014 | 3D Face Tracking and Multi-Scale, Spatio-temporal Analysis of Linguistically Significant Facial Expressions and Head Positions in ASL
Bo Liu 0005, Jingjing Liu 0001, Xiang Yu 0002, Dimitris N. Metaxas, Carol Neidle |
LREC | 1 |
| 2014 | Non-manual grammatical marker recognition based on multi-scale, spatio-temporal analysis of head pose and facial expressions
Jingjing Liu 0001, Bo Liu 0005, Shaoting Zhang 0001, Fei Yang 0001, Peng Yang 0001, Dimitris N. Metaxas, Carol Neidle |
Image Vis. Comput. | 2 |
| 2012 | Learning active facial patches for expression analysisabstractIn this paper, we present a new idea to analyze facial expression by exploring some common and specific information among different expressions. Inspired by the observation that only a few facial parts are active in expression disclosure (e.g., around mouth, eye), we try to discover the common and specific patches which are important to discriminate all the expressions and only a particular expression, respectively. A two-stage multi-task sparse learning (MTSL) framework is proposed to efficiently locate those discriminative patches. In the first stage MTSL, expression recognition tasks, each of which aims to find dominant patches for each expression, are combined to located common patches. Second, two related tasks, facial expression recognition and face verification tasks, are coupled to learn specific facial patches for individual expression. Extensive experiments validate the existence and significance of common and specific patches. Utilizing these learned patches, we achieve superior performances on expression recognition compared to the state-of-the-arts. Lin Zhong 0002, Qingshan Liu 0001, Peng Yang 0001, Bo Liu 0005, Junzhou Huang, Dimitris N. Metaxas |
CVPR | 4 |
| 2012 | Recognition of Nonmanual Markers in American Sign Language (ASL) Using Non-Parametric Adaptive 2D-3D Face Tracking
Dimitris N. Metaxas, Bo Liu 0005, Fei Yang 0001, Peng Yang 0001, Nicholas Michael, Carol Neidle |
LREC | 2 |
| 2010 | In-sequence video duplicate detection with fast point-to-line matchingabstractA computational geometry approach is developed to detect video duplicate with mild transformations. We model the video sequence as a trajectory after scaling and projection. Through interpolation and equal curve length sampling, part of the frame points is selected. A simplified video representation is the line segment set connecting the left neighboring points. For a given query, match distortion is calculated by projecting the query frame points to the line segment set guided by the frame temporal relationship. Experiments demonstrate the effectiveness of the proposed approach. Bo Liu 0005, Zhu Li 0001, Meng Wang 0001, Aggelos K. Katsaggelos |
ICIP | 1 |
| 2010 | Efficient video duplicate detection via compact curve matchingabstractEfficient video duplicate detection is of great practical significance in many applications. In this paper, we propose a computational geometry approach for video duplicate detection. In the proposed scheme, video clips are modeled as curves in a scaled appearance space. Curves are simplified through spline modeling and distortion threshold controlled point selection. Duplicate detection is therefore achieved by a fast line segment matching algorithm. Experiments demonstrate the effectiveness of the proposed method. Bo Liu 0005, Zhu Li 0001, Meng Wang 0001 |
ICME | 1 |
| 2010 | Metric learning with feature decomposition for image categorization
Meng Wang 0001, Bo Liu 0005, Jinhui Tang 0001, Xian-Sheng Hua 0001 |
Neurocomputing | 2 |
| 2010 | Accessible image search for colorblindnessabstractThis article introduces an intelligent system that accommodates colorblind users in image search. Color plays an important role in the human perception and recognition of images. However, there are about 8% of men and 0.8% of women suffering from colorblindness. We show that the existing image search techniques cannot provide satisfactory results for these users since many images will not be well perceived by them due to the loss of color information. To deal with this difficulty, we introduce a system named Accessible Image Search (AIS) to accommodate these users. Different from the general image search scheme that aims at returning more relevant results, AIS further takes into account the colorblind accessibilities of the returned results, that is, the image qualities in the eyes of colorblind users. The system contains three components: accessibility assessment, accessibility improvement, and color indication. The accessibility assessment component measures the accessibility scores of images, and consequently different reranking methods can be performed to prioritize images with high accessibilities. In the accessibility improvement component, we propose an efficient recoloring algorithm to modify the colors of the images such that they can be better perceived by colorblind users. Color indication aims to indicate the name of the interesting color in an image. We evaluate the introduced system with more than 60 queries and 20 anonymous colorblind users, and the empirical results demonstrate its effectiveness and usefulness. Meng Wang 0001, Bo Liu 0005, Xian-Sheng Hua 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2010 | In-Image Accessibility IndicationabstractThere are about 8% of men and 0.8% of women suffering from colorblindness. Due to the loss of certain color information, regions or objects in several images cannot be recognized by these viewers and this may degrade their perception and understanding of the images. This paper introduces an in-image accessibility indication scheme, which aims to automatically point out regions in which the content can hardly be recognized by colorblind viewers in a manually designed image. The proposed method first establishes a set of points around which the patches are not prominent enough for colorblind viewers due to the loss of color information. The inaccessible regions are then detected based on these points via a regularization framework. This scheme can be applied to check the accessibility of designed images, and consequently it can be used to help designers improve the images, such as modifying the colors of several objects or components. To our best knowledge, this is the first work that attempts to detect regions with accessibility problems in images for colorblindness. Experiments are conducted on 1994 poster images and empirical results have demonstrated the effectiveness of our approach. Meng Wang 0001, Yelong Sheng, Bo Liu 0005, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 3 |
| 2010 | Joint Learning of Labels and Distance MetricabstractMachine learning algorithms frequently suffer from the insufficiency of training data and the usage of inappropriate distance metric. In this paper, we propose a joint learning of labels and distance metric (JLLDM) approach, which is able to simultaneously address the two difficulties. In comparison with the existing semi-supervised learning and distance metric learning methods that focus only on label prediction or distance metric construction, the JLLDM algorithm optimizes the labels of unlabeled samples and a Mahalanobis distance metric in a unified scheme. The advantage of JLLDM is multifold: 1) the problem of training data insufficiency can be tackled; 2) a good distance metric can be constructed with only very few training samples; and 3) no radius parameter is needed since the algorithm automatically determines the scale of the metric. Extensive experiments are conducted to compare the JLLDM approach with different semi-supervised learning and distance metric learning methods, and empirical results demonstrate its effectiveness. Bo Liu 0005, Meng Wang 0001, Richang Hong, Zhengjun Zha, Xian-Sheng Hua 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2009 | Efficient image and video re-coloring for colorblindnessabstractThere are about 8% of men and 0.8% of women suffering from colorblindness. These viewers have difficulty in discriminating several colors, and thus many colorful images and videos that have high qualities for normal viewers may not be readily perceptible for them. In this paper, we propose an efficient re-coloring approach which can modify the colors of images and videos to enhance their perceptibility for colorblind users. Given an image, the re-coloring is accomplished by two color rotation steps in CIELAB color space. We first perform a local color rotation such that information of b*axis can be enhanced, and then adopt a global color rotation to refine the results. We will show that this method is simple yet effective, and it is able to outperform the traditional re-coloring algorithms in both performance and computational efficiency. We apply the method to process video frames and adopt several strategies to further reduce computational cost to realize real-time video re-coloring. We also explore the structure information of video data to avoid the color inconsistency problem. Specifically, we enforce the color mapping function to be identical in each shot and vary smoothly in adjacent shots. We conduct experiments on real-world images and videos with diverse content, and empirical results demonstrate the effectiveness of the proposed methods. Bo Liu 0005, Meng Wang 0001, Linjun Yang, Xiuqing Wu, Xian-Sheng Hua 0001 |
ICME | 1 |
| 2009 | Accessible image searchabstractThere are about 8% of men and 0.8% of women suffering from colorblindness. We show that the existing image search techniques cannot provide satisfactory results for these users, since many images will not be well perceived by them due to the loss of color information. In this paper, we introduce a scheme named Accessible Image Search (AIS) to accommodate these users. Different from the general image search scheme that aims at returning more relevant results, AIS further takes into account the colorblind accessibilities of the returned results, i.e., the image qualities in the eyes of colorblind users. The scheme includes two components: accessibility assessment and accessibility improvement. For accessibility assessment, we introduce an analysisbased method and a learning-based method. Based on the measured accessibility scores, different reranking methods can be performed to prioritize the images with high accessibilities. In accessibility improvement component, we propose an efficient recoloring algorithm to modify the colors of the images such that they can be better perceived by colorblind users. We also propose the Accessibility Average Precision (AAP) for AIS as a complementary performance evaluation measure to the conventional relevance-based evaluation methods. Experimental results with more than 60,000 images and 20 anonymous colorblind users demonstrate the effectiveness and usefulness of the proposed scheme. Meng Wang 0001, Bo Liu 0005, Xian-Sheng Hua 0001 |
ACM Multimedia | 2 |
| 2009 | Accommodating colorblind users in image searchabstractThere are about 8% of men and 0.8% of women suffering from colorblindness. Due to certain loss of color information, the existing image search techniques may not provide satisfactory results for these users. In this demonstration, we show an image search system that can accommodate colorblind users. It can help these special users find and enjoy what they want by providing multiple services for them, including search results reranking, image recoloring and color indication. Meng Wang 0001, Bo Liu 0005, Linjun Yang, Xian-Sheng Hua 0001 |
SIGIR | 2 |