Qi Bi

dblp:59/27 · DBLP profile ↗
← Back
72ranked-venue papers
32as first author
52since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 18 first-author · 29 since 2021Artificial intelligence and machine learning · 30 · 15 first-author · 29 since 2021Computer networks · 17 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 SAM3-I: Segment Anything with Instructions
abstract
Jingjing Li, Yue Feng, Yuchen Guo, Jincai Huang, Wei Ji, Qi Bi, Yongri Piao, Miao Zhang, Xiaoqi Zhao, Qiang Chen, Shihao Zou, Huchuan Lu, Li Cheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jincai Huang 0003, Wei Ji 0011, Qi Bi, Yongri Piao, Miao Zhang 0004, Xiaoqi Zhao 0003, Qiang Chen 0007, Shihao Zou, Huchuan Lu, Li Cheng 0001
ACL (1)6
2026 Coverage Probability Density: A Spatially-Resolved Performance Analysis for Non-Homogeneous Maritime Networks
abstract
The performance analysis of maritime communication networks is frequently constrained by the widespread yet unrealistic assumption of a uniform spatial distribution of vessels. This paper presents a unified stochastic geometry framework that integrates high-fidelity physical channel models with nonhomogeneous vessel topologies to provide a more accurate analytical foundation. We first systematically demonstrate that the macroscopic geographical distribution of vessels is a dominant factor governing overall network performance, with significant performance variations observed across different deployment scenarios. Subsequently, to analyze the spatial origins of this performance, we introduce the coverage probability density function (Cpdf) as a novel metric. The Cpdf enables a finegrained spatial decomposition of the total coverage probability, thereby identifying key geographical areas that contribute most to network performance and quantifying the precise spatial impact of physical phenomena, such as multipath fading nulls. Our results affirm the primacy of realistic spatial modeling and provide a new analytical tool for the design and optimization of next-generation maritime networks.
Wen-Yu Dong, Shaoshi Yang, Weiliang Xie, Junwei Hou, Rui-Si Han, Qi Bi, Sheng Chen 0001
ICC8
2026 Federated Learning With Doubly Adaptive Quantization in Unreliable Wireless Networks: Convergence Analysis and Low-Latency Design
abstract
Federated learning (FL) over wireless networks has become a key enabler for privacy-preserving distributed artificial intelligence (AI). However, high learning latency remains a critical bottleneck due to the presence of stragglers, limited wireless resources, and frequent model uploads. While model quantization can mitigate this issue by reducing communication overhead, its effectiveness is sensitive to device heterogeneity and time-varying channel conditions. To address this issue, we proposeFedDamQu, a communication-efficient FL framework with doubly-adaptive model quantization, which dynamically adjusts quantization bit-widths across devices and communication rounds to balance latency and accuracy. Our objective is to maximize the model performance under learning latency constraints. The main contributions are summarized as follows. 1) Convergence Analysis under Unreliable Channels: We derive a novel convergence error upper bound forFedDamQu, which explicitly quantifies the impact of device selection, unreliable transmission, and quantization error on the global model performance, under both fixed and dynamic quantization gain settings. 2) Joint Optimization Framework: Based on the knowledge from the proposed theoretical bound, we formulate a joint mixed integer nonlinear programming (MINLP) problem that integrates device selection, quantization bit-width configuration, and bandwidth allocation to minimize the convergence error under latency constraints. 3) Efficient Solution Design: The MINLP problem is decomposed into three subproblems, where closed-form solutions for quantization bit-width configuration and bandwidth allocation subproblems are derived, and a lightweight yet effective iterative algorithm is developed to obtain a suboptimal solution for the device selection subproblem. Extensive experiment results validate the theoretical analysis and demonstrate thatFedDamQuconsistently outperforms existing methods in terms of convergence rate and model accuracy, while significantly reducing the overall learning latency.
Jingsheng Tan, Shaoshi Yang, Hou-Yu Zhai, Zhiyong Feng 0001, Qi Bi
IEEE Internet Things J.6
2026 Harmonized medical federated learning via redundancy-aware client consistency
Jingjun Yi, Yuexiang Li, Qi Bi, Wei Ji 0011, Huimin Huang 0002, Yawen Huang, Yefeng Zheng 0001, Feiyue Huang
Pattern Recognit.4
2026 Revisiting Fine-Grained Image Analysis by Semantic-Part Alignment
abstract
Fine-grained image analysis is widely recognized as highly challenging, since distinguishing individual differences within a certain category, species, or type often depends on tiny, subtle patterns. However, learning fine-grained semantic categories from these subtle part patterns is inherently fragile, as they can easily be overwhelmed by the dominant patterns resting in the coarse-category information. Therefore, how to enhance the relation between the fine-grained semantics and these subtle patterns is the key. To push this frontier, a novel semantic-part alignment (SPA) learning scheme is proposed in this paper. Its general idea is to firstly measure the relevance of each part to the fine-grained semantics, and then regularize the fine-grained visual representation learning. Specifically, it consists of three key components, namely, joint semantic-part modeling, semantic-part set modeling, and optimal semantic-part transport. The joint semantic-part modeling associates each part in an image with the fine-grained semantics in a latent space. Then, the optimal semantic-part transport component is devised to enhance the relation between fine-grained semantic embeddings and the discriminative part embeddings. Notably, the proposed SPA is plug-in-and-play, easy-to-implement, and insensitive to the latent embedding dimension and loss weight. Experiments show the proposed method can substantially boost performance on multiple fine-grained image analysis tasks across various baselines.
Qi Bi, Jingjun Yi, Haolan Zhan, Wei Ji 0011, Gui-Song Xia
IEEE Trans. Image Process.1
2026 Forwarding or Learning? A Flexible Low-Latency Low-Energy-Consumption Wireless Federated Learning Architecture With UE-to-Network Relay
abstract
Wireless federated learning (FL) is an emerging artificial intelligence (AI) technique capable of leveraging the data and computing capacity of networked wireless devices while ensuring their individual data privacy and security. However, in geographical areas with poor wireless signal coverage, implementing FL is challenging. Additionally, intensive computation and communication put significant strain on resource-limited wireless devices. To address these issues, firstly, we propose a user equipment (UE)-to-network relay aided FL (UNR-FL) architecture that facilitates a low-cost and flexible implementation of wireless FL, without densifying network equipment deployment. Secondly, we propose an adaptive network control scheme that jointly optimizes device scheduling, network topology construction, and multi-type resource allocation to achieve low latency and low energy consumption. The second contribution is threefold. 1) For solving the device scheduling problem, we propose a voting-based strategy to identify the most suitable wireless UEs as relays. 2) Regarding the network topology optimization problem, we derive the optimal solutions under certain conditions, and propose a tabu search based meta-heuristic algorithm to find feasible solutions under the other conditions. 3) For solving the multi-type resource allocation problem, we analyze its mathematical structure and propose an iterative algorithm that has significantly lower computational complexity than the traditional method. This algorithm is capable of jointly optimizing the usage of transmission time resource, computing capacity, and transmit power. Extensive experimental results demonstrate that the proposed UNR-FL architecture and the adaptive network control scheme are capable of substantially reducing the learning latency and the total energy consumption.
Jingsheng Tan, Shaoshi Yang, Hou-Yu Zhai, Ping Zhang 0003, Qi Bi
IEEE Trans. Wirel. Commun.6
2025 DGFamba: Learning Flow Factorized State Space for Visual Domain Generalization
abstract
Domain generalization aims to learn a representation from the source domain, which can be generalized to arbitrary unseen target domains. A fundamental challenge for visual domain generalization is the domain gap caused by the dramatic style variation whereas the image content is stable. The realm of selective state space, exemplified by VMamba, demonstrates its global receptive field in representing the content. However, the way exploiting the domain-invariant property for selective state space is rarely explored. In this paper, we propose a novel Flow Factorized State Space model, dubbed as DGFamba, for visual domain generalization. To maintain domain consistency, we innovatively map the style-augmented and the original state embeddings by flow factorization. In this latent flow space, each state embedding from a certain style is specified by a latent probability path. By aligning these probability paths in the latent space, the state embeddings are able to represent the same content distribution regardless of the style differences. Extensive experiments conducted on various visual domain generalization settings show its state-of-the-art performance.
Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li
AAAI1
2025 Learning Fine-grained Domain Generalization via Hyperbolic State Space Hallucination
abstract
Fine-grained domain generalization (FGDG) aims to learn a fine-grained representation that can be well generalized to unseen target domains when only trained on the source domain data. Compared with generic domain generalization, FGDG is particularly challenging in that the fine-grained category can be only discerned by some subtle and tiny patterns. Such patterns are particularly fragile under the cross-domain style shifts caused by illumination, color and etc. To push this frontier, this paper presents a novel Hyperbolic State Space Hallucination (HSSH) method. It consists of two key components, namely, state space hallucination (SSH) and hyperbolic manifold consistency (HMC). SSH enriches the style diversity for the state embeddings by firstly extrapolating and then hallucinating the source images. Then, the pre- and post- style hallucinate state embeddings are projected into the hyperbolic manifold. The hyperbolic state space models the high-order statistics, and allows a better discernment of the fine-grained patterns. Finally, the hyperbolic distance is minimized, so that the impact of style variation on fine-grained patterns can be eliminated. Experiments on three FGDG benchmarks demonstrate its state-of-the-art performance.
Qi Bi, Jingjun Yi, Haolan Zhan, Wei Ji 0011, Gui-Song Xia
AAAI1
2025 MaCP: Minimal yet Mighty Adaptation via Hierarchical Cosine Projection
abstract
We present a new adaptation method MaCP, Minimal yet Mighty adaptive Cosine Projection, that achieves exceptional performance while requiring minimal parameters and memory for fine-tuning large foundation models.Its general idea is to exploit the superior energy compaction and decorrelation properties of cosine projection to improve both model efficiency and accuracy.Specifically, it projects the weight change from the low-rank adaptation into the discrete cosine space.Then, the weight change is partitioned over different levels of the discrete cosine spectrum, and each partition's most critical frequency components are selected.Extensive experiments demonstrate the effectiveness of MaCP across a wide range of single-modality tasks, including natural language understanding, natural language generation, text summarization, as well as multimodality tasks such as image classification and video understanding.MaCP consistently delivers superior accuracy, significantly reduced computational complexity, and lower memory requirements compared to existing alternatives.
Yixian Shen, Qi Bi, Jia-Hong Huang, Hongyi Zhu 0004, Andy D. Pimentel, Anuj Pathania
ACL (1)2
2025 NightAdapter: Learning a Frequency Adapter for Generalizable Night-time Scene Segmentation
abstract
Night-time scene segmentation is a critical yet challenging task in the real-world applications, primarily due to the complicated lighting conditions. However, existing methods lack sufficient generalization ability to unseen nighttime scenes with varying illumination. In light of this issue, we focus on investigating generalizable paradigms for night-time scene segmentation and propose an efficient fine-tuning scheme, dubbed NightAdapter, alleviating the domain gap across various scenes. Interestingly, different properties embedded in the day-time and night-time features can be characterized by the bands after discrete sine transform, which can be categorized into illumination-sensitive/-insensitive bands. Hence, our NightAdapter is powered by two appealing designs: (1) Illumination-Insensitive Band Adaptation that provides a foundation for understanding the prior, enhancing the robustness to illumination shifts; (2) Illumination-Sensitive Band Adaptation that fine-tunes the randomized frequency bands, mitigating the domain gap between the day-time and various night-time scenes. As a consequence, illumination-insensitive enhancement improves the domain invariance, while illumination-sensitive diminution strengthens the domain shift between different scenes. NightAdapter yields significant improvements over the state-of-the-art methods under various day-to-night, night-to-night, and in-domain night segmentation experiments. Source code is available at https://github.com/BiQiWHU/NightAdapter.
Qi Bi, Jingjun Yi, Huimin Huang 0002, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001
CVPR1
2025 Integrated Sensing and Communication Systems based on Millimeter wave: Frame Structure Design
Xue Ding 0001, Yuanhao Cui, Weiliang Xie, Qi Bi
GLOBECOM6
2025 Communication-Efficient Federated Learning with Doubly-Adaptive Model Quantization in Unreliable Wireless Networks
abstract
To address the latency bottleneck in wireless federated learning (FL) systems, we propose FedDamQu, a communication-efficient framework that adaptively adjusts model quantization bit-width across both devices and communication rounds to cope with device heterogeneity and dynamic wireless conditions. Our objective is to maximize global model performance under strict latency constraints. The key contributions are threefold. First, we derive a novel convergence error upper bound that explicitly characterizes the effects of device selection, unreliable transmission, and quantization error. Second, we formulate a joint mixed-integer nonlinear programming (MINLP) problem that integrates device selection, quantization bit-width configuration, and bandwidth allocation to minimize the convergence error under latency constraints. 3) Third, we decompose the MINLP into three tractable subproblems, obtain closed-form solutions for quantization bit-width configuration and bandwidth allocation, and develop a lightweight iterative algorithm for device selection. Extensive experiments demonstrate that FedDamQu consistently outperforms existing methods in terms of convergence speed and model accuracy, while significantly reducing the overall learning latency.
Jingsheng Tan, Shaoshi Yang, Hou-Yu Zhai, Zhiyong Feng 0001, Qi Bi
GLOBECOM6
2025 A Flexible Low-Latency Low-Energy-Consumption Wireless Federated Learning Architecture with UE-to-Network Relay
abstract
This paper addresses the difficulty of implementing federated learning (FL) in geographical areas with poor wireless signal coverage, and alleviates the high burden imposed by intensive computation and communication on resource-limited wireless devices. Firstly, we propose a user equipment (UE)-tonetwork relay aided FL (UNR-FL) architecture that facilitates a low-cost and flexible implementation of wireless FL, without densifying network equipment deployment. Secondly, we propose an adaptive network control scheme that jointly optimizes resource allocation, network topology construction, and device scheduling, to achieve low latency and low energy consumption. The second contribution is threefold. 1) For allocating resources, we propose a linear-complexity algorithm which is capable of jointly optimizing the transmission time resource and the computing power. 2) For constructing network topology, we derive the optimal closed-form solution under certain conditions, and propose a tabu search based meta-heuristic algorithm to find feasible solutions under the other conditions. 3) For scheduling devices, we propose a voting-based device scheduling algorithm that is near-optimal. Extensive experimental results demonstrate that the proposed UNR-FL architecture and the adaptive network control scheme are capable of substantially reducing the learning latency and the total energy consumption.
Jingsheng Tan, Shaoshi Yang, Hou-Yu Zhai, Ping Zhang 0003, Qi Bi
ICC6
2025 AdaDCP: Learning an Adapter with Discrete Cosine Prior for Clear-to-Adverse Domain Generalization
Qi Bi, Yixian Shen, Jingjun Yi, Gui-Song Xia
ICCV1
2025 A Simple Yet Mighty Hartley Diffusion Versatilist for Generalizable Dense Vision Tasks
Qi Bi, Jingjun Yi, Huimin Huang 0002, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001
ICCV1
2025 D-CAM: Learning Generalizable Weakly-Supervised Medical Image Segmentation from Domain-Invariant CAM
Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Huimin Huang 0002, Yuexiang Li, Shaoxin Li 0001, Xian Wu 0001, Yefeng Zheng 0001, Feiyue Huang
MICCAI (5)2
2025 Single Domain Generalization for Multimodal Cross-Cancer Prognosis via Dirac Rebalancer and Distribution Entanglement
abstract
Deep learning has shown remarkable performance in integrating multimodal data for survival prediction. However, existing multimodal methods mainly focus on single cancer types and overlook the challenge of generalization across cancers. In this work, we are the first to reveal that multimodal prognosis models often generalize worse than unimodal ones in cross-cancer scenarios, despite the critical need for such robustness in clinical practice. To address this, we propose a new task: Cross-Cancer Single Domain Generalization for Multimodal Prognosis, which evaluates whether models trained on a single cancer type can generalize to unseen cancers. We identify two key challenges: degraded features from weaker modalities and ineffective multimodal integration. To tackle these, we introduce two plug-and-play modules: Sparse Dirac Information Rebalancer (SDIR) and Cancer-aware Distribution Entanglement (CADE). SDIR mitigates the dominance of strong features by applying Bernoulli-based sparsification and Dirac-inspired stabilization to enhance weaker modality signals. CADE, designed to synthesize the target domain distribution, fuses local morphological cues and global gene expression in latent space. Experiments on a four-cancer-type benchmark demonstrate superior generalization, laying the foundation for practical, robust cross-cancer multimodal prognosis. Code is available at here.
Jiaxuan Jiang 0001, Jiashuai Liu 0001, Zhong Wang 0006, Qi Bi, Yefeng Zheng 0001
ACM Multimedia6
2025 AtlantisGS: Underwater Sparse-View Scene Reconstruction via Gaussian Splatting
Jingjun Yi, Qi Bi, Hao Zheng 0008, Huimin Huang 0002, Haolan Zhan, Yixian Shen, Wei Ji 0011, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001
ACM Multimedia2
2025 SSH: Sparse Spectrum Adaptation via Discrete Hartley Transformation
abstract
Yixian Shen, Qi Bi, Jia-hong Huang, Hongyi Zhu, Andy D. Pimentel, Anuj Pathania. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yixian Shen, Qi Bi, Jia-Hong Huang, Hongyi Zhu 0004, Andy D. Pimentel, Anuj Pathania
NAACL (Long Papers)2
2025 Degradation-Aware Dynamic Schrödinger Bridge for Unpaired Image Restoration
abstract
Image restoration is a fundamental task in computer vision and machine learning, which learns a mapping between the clear images and the degraded images under various conditions (e.g., blur, low-light, haze). Yet, most existing image restoration methods are highly restricted by the requirement of degraded and clear image pairs, which limits the generalization and feasibility to enormous real-world scenarios without paired images. To address this bottleneck, we propose a Degradation-aware Dynamic Schr\"{o}dinger Bridge (DDSB) for unpaired image restoration. Its general idea is to learn a Schr\"{o}dinger Bridge between clear and degraded image distribution, while at the same time emphasizing the physical degradation priors to reduce the accumulation of errors during the restoration process. A Degradation-aware Optimal Transport (DOT) learning scheme is accordingly devised. Training a degradation model to learn the inverse restoration process is particularly challenging, as it must be applicable across different stages of the iterative restoration process. A Dynamic Transport with Consistency (DTC) learning objective is further proposed to reduce the loss of image details in the early iterations and therefore refine the degradation model. Extensive experiments on multiple image degradation tasks show its state-of-the-art performance over the prior arts.
Jingjun Yi, Qi Bi, Hao Zheng 0008, Huimin Huang 0002, Yixian Shen, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001
NeurIPS2
2025 Learning a Cross-Modal Schrödinger Bridge for Visual Domain Generalization
abstract
Domain generalization aims to train models that perform robustly on unseen target domains without access to target data. The realm of vision-language foundation model has opened a new venue owing to its inherent out-of-distribution generalization capability. However, the static alignment to class-level textual anchors remains insufficient to handle the dramatic distribution discrepancy from diverse domain-specific visual features. In this work, we propose a novel cross-domain Schrödinger Bridge (SB) method, namely SBGen, to handle this challenge, which explicitly formulates the stochastic semantic evolution, to gain better generalization to unseen domains. Technically, the proposed \texttt{SBGen} consists of three key components: (1) \emph{text-guided domain-aware feature selection} to isolate semantically aligned image tokens; (2) \emph{stochastic cross-domain evolution} to simulate the SB dynamics via a learnable time-conditioned drift; and (3) \emph{stochastic domain-agnostic interpolation} to construct semantically grounded feature trajectories. Empirically, \texttt{SBGen} achieves state-of-the-art performance on domain generalization in both classification and segmentation. This work highlights the importance of modeling domain shifts as structured stochastic processes grounded in semantic alignment.
Hao Zheng 0008, Jingjun Yi, Qi Bi, Huimin Huang 0002, Haolan Zhan, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001
NeurIPS3
2025 Intra-hour photovoltaic power point-interval prediction using a dual-view deep neural network
Zhi-Ru Chen, Yulong Bai 0001, Hao-yu Qin, Qi Bi
Expert Syst. Appl.5
2025 Learning Generalized Medical Image Representation by Decoupled Feature Queries
abstract
Medical images are usually collected from multiple clinical centers with various types of scanners. When confronted with such significant cross-domain distribution discrepancy, a deep network tends to capture similar patterns by multiple channels, while different cross-domain patterns are also allowed to rest in the same channel. Such channel redundancy limits the expressive capability of a representation, resulting in less preferable generalization ability. To address this fundamental yet challenging issue, we propose a novel decoupled feature as query (DFQ) framework for domain generalized medical image representation learning. Its general idea is to leverage the channel-wise decoupled deep features as queries. Particularly, a deep instance whitening transform with restricted isometry is proposed, which enforces each channel orthogonal to the rest channels after decoupling. Besides, the long-range dependency between decoupled deep and shallow features is implicitly constrained to minimize channel redundancy throughout training. Extensive experiments show its state-of-the-art performance on three medical domain generalization tasks with four modalities.
Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 GAD: Domain generalized diabetic retinopathy grading by grade-aware de-stylization
Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Wei Ji 0011, Yuexiang Li, Yefeng Zheng 0001
Pattern Recognit.1
2025 Cross-Level Multi-Instance Distillation for Self-Supervised Fine-Grained Visual Categorization
abstract
High-quality annotation of fine-grained visual categories demands great expert knowledge, which is taxing and time consuming. Alternatively, learning fine-grained visual representation from enormous unlabeled images (e.g., species, brands) by self-supervised learning becomes a feasible solution. However, recent investigations find that existing self-supervised learning methods are less qualified to represent fine-grained categories. The bottleneck lies in that the pre-trained class-agnostic representation is built from every patch-wise embedding, while fine-grained categories are only determined by several key patches of an image. In this paper, we propose a Cross-level Multi-instance Distillation (CMD) framework to tackle this challenge. Our key idea is to consider the importance of each image patch in determining the fine-grained representation by multiple instance learning. To comprehensively learn the relation between informative patches and fine-grained semantics, the multi-instance knowledge distillation is implemented on both the region/image crop pairs from the teacher and student net, and the region-image crops inside the teacher / student net, which we term as intra-level multi-instance distillation and inter-level multi-instance distillation. Extensive experiments on several commonly used datasets, including CUB-200-2011, Stanford Cars and FGVC Aircraft, demonstrate that the proposed method outperforms the contemporary methods by up to 10.14% and existing state-of-the-art self-supervised learning approaches by up to 19.78% on both top-1 accuracy and Rank-1 retrieval metric. Source code is available at https://github.com/BiQiWHU/CMD.
Qi Bi, Wei Ji 0011, Jingjun Yi, Haolan Zhan, Gui-Song Xia
IEEE Trans. Image Process.1
2025 Universal Fine-Grained Visual Categorization by Concept Guided Learning
abstract
Existing fine-grained visual categorization (FGVC) methods assume that the fine-grained semantics rest in the informative parts of an image. This assumption works well on favorable front-view object-centric images, but can face great challenges in many real-world scenarios, such as scene-centric images (e.g., street view) and adverse viewpoint (e.g., object reidentification, remote sensing). In such scenarios, the mis-/over-feature activation is likely to confuse the part selection and degrade the fine-grained representation. In this paper, we are motivated to design a universal FGVC framework for real-world scenarios. More precisely, we propose a concept guided learning (CGL), which models concepts of a certain fine-grained category as a combination of inherited concepts from its subordinate coarse-grained category and discriminative concepts from its own. The discriminative concepts is utilized to guide the fine-grained representation learning. Specifically, three key steps are designed, namely, concept mining, concept fusion, and concept constraint. On the other hand, to bridge the FGVC dataset gap under scene-centric and adverse viewpoint scenarios, a Fine-grained Land-cover Categorization Dataset (FGLCD) with 59,994 fine-grained samples is proposed. Extensive experiments show the proposed CGL: 1) has a competitive performance on conventional FGVC; 2) achieves state-of-the-art performance on fine-grained aerial scenes & scene-centric street scenes; 3) good generalization on object re-identification and fine-grained aerial object detection. The dataset and source code will be available at https://github.com/BiQiWHU/CGL.
Qi Bi, Beichen Zhou, Wei Ji 0011, Gui-Song Xia
IEEE Trans. Image Process.1
2024 Learning Generalized Segmentation for Foggy-Scenes by Bi-directional Wavelet Guidance
abstract
Learning scene semantics that can be well generalized to foggy conditions is important for safety-crucial applications such as autonomous driving. Existing methods need both annotated clear images and foggy images to train a curriculum domain adaptation model. Unfortunately, these methods can only generalize to the target foggy domain that has seen in the training stage, but the foggy domains vary a lot in both urban-scene styles and fog styles. In this paper, we propose to learn scene segmentation well generalized to foggy-scenes under the domain generalization setting, which does not involve any foggy images in the training stage and can generalize to any arbitrary unseen foggy scenes. We argue that an ideal segmentation model that can be well generalized to foggy-scenes need to simultaneously enhance the content, de-correlate the urban-scene style and de-correlate the fog style. As the content (e.g., scene semantic) rests more in low-frequency features while the style of urban-scene and fog rests more in high-frequency features, we propose a novel bi-directional wavelet guidance (BWG) mechanism to realize the above three objectives in a divide-and-conquer manner. With the aid of Haar wavelet transformation, the low frequency component is concentrated on the content enhancement self-attention, while the high frequency component is shifted to the style and fog self-attention for de-correlation purpose. It is integrated into existing mask-level Transformer segmentation pipelines in a learnable fashion. Large-scale experiments are conducted on four foggy-scene segmentation datasets under a variety of interesting settings. The proposed method significantly outperforms existing directly-supervised, curriculum domain adaptation and domain generalization segmentation methods. Source code is available at https://github.com/BiQiWHU/BWG.
Qi Bi, Shaodi You, Theo Gevers
AAAI1
2024 Learning Content-Enhanced Mask Transformer for Domain Generalized Urban-Scene Segmentation
abstract
Domain-generalized urban-scene semantic segmentation (USSS) aims to learn generalized semantic predictions across diverse urban-scene styles. Unlike generic domain gap challenges, USSS is unique in that the semantic categories are often similar in different urban scenes, while the styles can vary significantly due to changes in urban landscapes, weather conditions, lighting, and other factors. Existing approaches typically rely on convolutional neural networks (CNNs) to learn the content of urban scenes. In this paper, we propose a Content-enhanced Mask TransFormer (CMFormer) for domain-generalized USSS. The main idea is to enhance the focus of the fundamental component, the mask attention mechanism, in Transformer segmentation models on content information. We have observed through empirical analysis that a mask representation effectively captures pixel segments, albeit with reduced robustness to style variations. Conversely, its lower-resolution counterpart exhibits greater ability to accommodate style variations, while being less proficient in representing pixel segments. To harness the synergistic attributes of these two approaches, we introduce a novel content-enhanced mask attention mechanism. It learns mask queries from both the image feature and its down-sampled counterpart, aiming to simultaneously encapsulate the content and address stylistic variations. These features are fused into a Transformer decoder and integrated into a multi-resolution content-enhanced mask attention learning scheme. Extensive experiments conducted on various domain-generalized urban-scene segmentation datasets demonstrate that the proposed CMFormer significantly outperforms existing CNN-based methods by up to 14.0% mIoU and the contemporary HGFormer by up to 1.7% mIoU. The source code is publicly available at https://github.com/BiQiWHU/CMFormer.
Qi Bi, Shaodi You, Theo Gevers
AAAI1
2024 Learning Generalized Medical Image Segmentation from Decoupled Feature Queries
abstract
Domain generalized medical image segmentation requires models to learn from multiple source domains and generalize well to arbitrary unseen target domain. Such a task is both technically challenging and clinically practical, due to the domain shift problem (i.e., images are collected from different hospitals and scanners). Existing methods focused on either learning shape-invariant representation or reaching consensus among the source domains. An ideal generalized representation is supposed to show similar pattern responses within the same channel for cross-domain images. However, to deal with the significant distribution discrepancy, the network tends to capture similar patterns by multiple channels, while different cross-domain patterns are also allowed to rest in the same channel. To address this issue, we propose to leverage channel-wise decoupled deep features as queries. With the aid of cross-attention mechanism, the long-range dependency between deep and shallow features can be fully mined via self-attention and then guides the learning of generalized representation. Besides, a relaxed deep whitening transformation is proposed to learn channel-wise decoupled features in a feasible way. The proposed decoupled fea- ture query (DFQ) scheme can be seamlessly integrate into the Transformer segmentation model in an end-to-end manner. Extensive experiments show its state-of-the-art performance, notably outperforming the runner-up by 1.31% and 1.98% with DSC metric on generalized fundus and prostate benchmarks, respectively. Source code is available at https://github.com/BiQiWHU/DFQ.
Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001
AAAI1
2024 Constraint Latent Space Matters: An Anti-anomalous Waveform Transformation Solution from Photoplethysmography to Arterial Blood Pressure
abstract
Arterial blood pressure (ABP) holds substantial promise for proactive cardiovascular health management. Notwithstanding its potential, the invasive nature of ABP measurements confines their utility primarily to clinical environments, limiting their applicability for continuous monitoring beyond medical facilities. The conversion of photoplethysmography (PPG) signals into ABP equivalents has garnered significant attention due to its potential in revolutionizing cardiovascular disease management. Recent strides in PPG-to-ABP prediction encompass the integration of generative and discriminative models. Despite these advances, the efficacy of these models is curtailed by the latent space shift predicament, stemming from alterations in PPG data distribution across disparate hardware and individuals, potentially leading to distorted ABP waveforms. To tackle this problem, we present an innovative solution named the Latent Space Constraint Transformer (LSCT), leveraging a quantized codebook to yield robust latent spaces by employing multiple discretizing bases. To facilitate improved reconstruction, the Correlation-boosted Attention Module (CAM) is introduced to systematically query pertinent bases on a global scale. Furthermore, to enhance expressive capacity, we propose the Multi-Spectrum Enhancement Knowledge (MSEK), which fosters local information flow within the channels of latent code and provides additional embedding for reconstruction. Through comprehensive experimentation on both publicly available datasets and a private downstream task dataset, the proposed approach demonstrates noteworthy performance enhancements compared to existing methods. Extensive ablation studies further substantiate the effectiveness of each introduced module.
Cheng Bian, Xiaoyu Li 0007, Qi Bi, Guangpu Zhu, Jiegeng Lyu, Weile Zhang, Yelei Li, Zijing Zeng
AAAI3
2024 Self-Supervised Cross-Level Consistency Learning For Fundus Image Classification
abstract
The rapid development of intelligent systems for eye disease diagnosis decreases the risk of people suffering from vision impairment. However, the superior discrimination ability of existing retinal disease diagnosis methods heavily relies on the large-scale high-quality annotations. In this work, we adapt the self-supervised technique for fundus image classification with the merits of bypassing the over-dependence of labeled data. Unlike most current self-supervised approaches, which only learn global pre-text representations from view-level, our method further incorporates the region-level representations into the learning process, since the pathological changes in fundus images are usually subtle and scattered. Specifically, we propose a novel self-supervised cross-level consistency learning scheme (S2C2L), which leverages both view-level and region-level representations of a vision Transformer to improve the robustness of extracted self-supervised representation. A diagnosis perception module (DPM) is constructed to enhance the activation of local pathological regions from both region and view levels, and a cross-level consistency loss is dedicated to align the representations from both levels. Extensive experiments on iChallenge-AMD, LAG and APTOS2019 datasets validate the state-of-the-art performance of our method for three common eye diseases.
Qi Bi, Hao Zheng 0008, Xu Sun 0006, Jingjun Yi, Wentian Zhang, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001
ICASSP1
2024 Boosting Fine-Grained Oriented Object Detection via Text Features
Beichen Zhou, Qi Bi, Jian Ding 0001, Gui-Song Xia
ICPR (16)2
2024 Hallucinated Style Distillation for Single Domain Generalization in Medical Image Segmentation
Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Shaoxin Li 0001, Yuexiang Li, Yefeng Zheng 0001, Feiyue Huang
MICCAI (10)2
2024 Learning Spectral-Decomposited Tokens for Domain Generalized Semantic Segmentation
abstract
The rapid development of Vision Foundation Model (VFM) brings inherent out-domain generalization for a variety of down-stream tasks. Among them, domain generalized semantic segmentation (DGSS) holds unique challenges as the cross-domain images share common pixel-wise content information but vary greatly in terms of the style. In this paper, we present a novel Spectral-dEcomposed Token (SET) learning framework to advance the frontier. Delving into further than existing fine-tuning token & frozen backbone paradigm, the proposed SET especially focuses on the way learning style-invariant features from these learnable tokens. Particularly, the frozen VFM features are first decomposed into the phase and amplitude components in the frequency space, which mainly contain the information of content and style, respectively, and then separately processed by learnable tokens for task-specific information extraction. Particularly, the frozen VFM features are first decomposed into the phase and amplitude components in the frequency space, which mainly contain the information of content and style, respectively, and then separately processed by learnable tokens for task-specific information extraction.After the decomposition, style variation primarily impacts the token-based feature enhancement within the amplitude branch. To address this issue, we further develop an attention optimization method to bridge the gap between style-affected representation and static tokens during inference. Extensive cross-domain experiments show its state-of-the-art performance.
Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001
ACM Multimedia2
2024 Samba: Severity-aware Recurrent Modeling for Cross-domain Medical Image Grading
abstract
Disease grading is a crucial task in medical image analysis. Due to the continuous progression of diseases, i.e., the variability within the same level and the similarity between adjacent stages, accurate grading is highly challenging. Furthermore, in real-world scenarios, models trained on limited source domain datasets should also be capable of handling data from unseen target domains. Due to the cross-domain variants, the feature distribution between source and unseen target domains can be dramatically different, leading to a substantial decrease in model performance. To address these challenges in cross-domain disease grading, we propose a Severity-aware Recurrent Modeling (Samba) method in this paper. As the core objective of most staging tasks is to identify the most severe lesions, which may only occupy a small portion of the image, we propose to encode image patches in a sequential and recurrent manner. Specifically, a state space model is tailored to store and transport the severity information by hidden states. Moreover, to mitigate the impact of cross-domain variants, an Expectation-Maximization (EM) based state recalibration mechanism is designed to map the patch embeddings into a more compact space. We model the feature distributions of different lesions through the Gaussian Mixture Model (GMM) and reconstruct the intermediate features based on learnable severity bases. Extensive experiments show the proposed Samba outperforms the VMamba baseline by an average accuracy of 23.5\%, 5.6\% and 4.1\% on the cross-domain grading of fatigue fracture, breast cancer and diabetic retinopathy, respectively. Source code is available at \url{https://github.com/BiQiWHU/Samba}.
Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Haolan Zhan, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001
NeurIPS1
2024 Learning Frequency-Adapted Vision Foundation Model for Domain Generalized Semantic Segmentation
abstract
The emerging vision foundation model (VFM) has inherited the ability to generalize to unseen images. Nevertheless, the key challenge of domain-generalized semantic segmentation (DGSS) lies in the domain gap attributed to the cross-domain styles, i.e., the variance of urban landscape and environment dependencies. Hence, maintaining the style-invariant property with varying domain styles becomes the key bottleneck in harnessing VFM for DGSS. The frequency space after Haar wavelet transformation provides a feasible way to decouple the style information from the domain-invariant content, since the content and style information are retained in the low- and high- frequency components of the space, respectively. To this end, we propose a novel Frequency-Adapted (FADA) learning scheme to advance the frontier. Its overall idea is to separately tackle the content and style information by frequency tokens throughout the learning process. Particularly, the proposed FADA consists of two branches, i.e., low- and high- frequency branches. The former one is able to stabilize the scene content, while the latter one learns the scene styles and eliminates its impact to DGSS. Experiments conducted on various DGSS settings show the state-of-the-art performance of our FADA and its versatility to a variety of VFMs. Source code is available at \url{https://github.com/BiQiWHU/FADA}.
Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Wei Ji 0011, Yuexiang Li, Yefeng Zheng 0001
NeurIPS1
2023 Blockchain-based distributed operation and incentive solution for P-RAN
Xiaoou Liu, Qi Bi
Comput. Commun.3
2023 Learning rotation equivalent scene representation from instance-level semantics: A novel top-down perspective
abstract
This paper focuses on rotation variant scene recognition. Different from existing rotation invariant recognition approaches which learn from either rotated images or rotated convolutional filters in a bottom-up manner, a new top-down perspective by learning is explored from instance-level semantic representation. The goal is to eliminate the convolutional feature differences in bottom-up feature propagation caused by the rotation sensitive nature of convolution operation. Our rotation equivalent convolutional neural network (RE-CNN) scheme consists of three components. Firstly, our key instance selection module highlights the instances strongly related to the scene scheme regardless of their orientation. Secondly, our key instance aggregation module builds a scene representation invariant to the position change of each instance caused by rotation. Finally, our semantic fusion module allows the framework to be organized as a whole and implements rotation regularization. Notably, our RE-CNN scheme can be adapted to existing CNNs in a plug-in-and-play manner. Extensive experiments on rotation variant scene recognition benchmarks from four domains demonstrate the state-of-the-art performance and generalization capability of the proposed RE-CNN.
Qi Bi, Shaodi You, Wei Ji 0011, Theo Gevers
Comput. Vis. Image Underst.1
2023 MIL-ViT: A multiple instance vision transformer for fundus image classification
Qi Bi, Xu Sun 0006, Kai Ma 0002, Cheng Bian, Munan Ning, Nanjun He, Yawen Huang, Yuexiang Li, Hanruo Liu, Yefeng Zheng 0001
J. Vis. Commun. Image Represent.1
2023 Interactive Learning of Intrinsic and Extrinsic Properties for All-Day Semantic Segmentation
abstract
Scene appearance changes drastically throughout the day. Existing semantic segmentation methods mainly focus on well-lit daytime scenarios and are not well designed to cope with such great appearance changes. Naively using domain adaption does not solve this problem because it usually learns a fixed mapping between the source and target domain and thus have limited generalization capability on all-day scenarios (i. e., from dawn to night). In this paper, in contrast to existing methods, we tackle this challenge from the perspective of image formulation itself, where the image appearance is determined by both intrinsic (e. g., semantic category, structure) and extrinsic (e. g., lighting) properties. To this end, we propose a novel intrinsic-extrinsic interactive learning strategy. The key idea is to interact between intrinsic and extrinsic representations during the learning process under spatial-wise guidance. In this way, the intrinsic representation becomes more stable and, at the same time, the extrinsic representation gets better at depicting the changes. Consequently, the refined image representation is more robust to generate pixel-wise predictions for all-day scenarios. To achieve this, we propose an All-in-One Segmentation Network (AO-SegNet) in an end-to-end manner. Large scale experiments are conducted on three real datasets (Mapillary, BDD100K and ACDC) and our proposed synthetic All-day CityScapes dataset. The proposed AO-SegNet shows a significant performance gain against the state-of-the-art under a variety of CNN and ViT backbones on all the datasets.
Qi Bi, Shaodi You, Theo Gevers
IEEE Trans. Image Process.1
2022 Label-Efficient Hybrid-Supervised Learning for Medical Image Segmentation
abstract
Due to the lack of expertise for medical image annotation, the investigation of label-efficient methodology for medical image segmentation becomes a heated topic. Recent progresses focus on the efficient utilization of weak annotations together with few strongly-annotated labels so as to achieve comparable segmentation performance in many unprofessional scenarios. However, these approaches only concentrate on the supervision inconsistency between strongly- and weakly-annotated instances but ignore the instance inconsistency inside the weakly-annotated instances, which inevitably leads to performance degradation. To address this problem, we propose a novel label-efficient hybrid-supervised framework, which considers each weakly-annotated instance individually and learns its weight guided by the gradient direction of the strongly-annotated instances, so that the high-quality prior in the strongly-annotated instances is better exploited and the weakly-annotated instances are depicted more precisely. Specially, our designed dynamic instance indicator (DII) realizes the above objectives, and is adapted to our dynamic co-regularization (DCR) framework further to alleviate the erroneous accumulation from distortions of weak annotations. Extensive experiments on two hybrid-supervised medical segmentation datasets demonstrate that with only 10% strong labels, the proposed framework can leverage the weak labels efficiently and achieve competitive performance against the 100% strong-label supervised scenario.
Junwen Pan, Qi Bi, Yanzhan Yang, Pengfei Zhu 0001, Cheng Bian
AAAI2
2022 Promoting Saliency From Depth: Deep Unsupervised RGB-D Saliency Detection
Wei Ji 0011, Qi Bi, Chuan Guo 0002, Jie Liu 0044, Li Cheng 0001
ICLR3
2022 All Grains, One Scheme (AGOS): Learning Multigrain Instance Representation for Aerial Scene Classification
abstract
Aerial scene classification remains challenging as: 1) the size of key objects in determining the scene scheme varies greatly; 2) many objects irrelevant to the scene scheme are often flooded in the image. Hence, how to effectively perceive the region of interests (RoIs) from a variety of sizes and build more discriminative representation from such complicated object distribution is vital to understand an aerial scene. In this paper, we propose a novelall grains, one scheme(AGOS) framework to tackle these challenges.To the best of our knowledge, it is the first work to extend the classic multiple instance learning into multi-grain formulation. Specially, it consists of a multi-grain perception module (MGP), a multi-branch multi-instance representation module (MBMIR) and a self-aligned semantic fusion (SSF) module. Firstly, our MGP preserves the differential dilated convolutional features from the backbone, which magnifies the discriminative information from multi-grains. Then, our MBMIR highlights the key instances in the multi-grain representation under the MIL formulation. Finally, our SSF allows our framework to learn the same scene scheme from multi-grain instance representations and fuses them, so that the entire framework is optimized as a whole. Notably, our AGOS is flexible and can be easily adapted to existing CNNs in a plug-and-play manner. Extensive experiments on UCM, AID and NWPU benchmarks demonstrate that our AGOS achieves a comparable performance against the state-of-the-art methods.
Qi Bi, Beichen Zhou, Kun Qin, Qinghao Ye, Gui-Song Xia
IEEE Trans. Geosci. Remote. Sens.1
2021 Calibrated RGB-D Salient Object Detection
abstract
Complex backgrounds and similar appearances between objects and their surroundings are generally recognized as challenging scenarios in Salient Object Detection (SOD). This naturally leads to the incorporation of depth information in addition to the conventional RGB image as input, known as RGB-D SOD or depth-aware SOD. Meanwhile, this emerging line of research has been considerably hindered by the noise and ambiguity that prevail in raw depth images. To address the aforementioned issues, we propose a Depth Calibration and Fusion (DCF) framework that contains two novel components: 1) a learning strategy to calibrate the latent bias in the original depth maps towards boosting the SOD performance; 2) a simple yet effective cross reference module to fuse features from both RGB and depth modalities. Extensive empirical experiments demonstrate that the proposed approach achieves superior performance against 27 state-of-the-art methods. Moreover, our depth calibration strategy alone can work as a preprocessing step; empirically it results in noticeable improvements when being applied to existing cutting-edge RGB-D SOD models. Source code is available at https://github.com/jiwei0921/DCF.
Wei Ji 0011, Miao Zhang 0004, Yongri Piao, Shunyu Yao 0004, Qi Bi, Kai Ma 0002, Yefeng Zheng 0001, Huchuan Lu, Li Cheng 0001
CVPR7
2021 Learning Calibrated Medical Image Segmentation via Multi-Rater Agreement Modeling
abstract
In medical image analysis, it is typical to collect multiple annotations, each from a different clinical expert or rater, in the expectation that possible diagnostic errors could be mitigated. Meanwhile, from the computer vision practitioner viewpoint, it has been a common practice to adopt the ground-truth labels obtained via either the majority-vote or simply one annotation from a preferred rater. This process, however, tends to overlook the rich information of agreement or disagreement ingrained in the raw multi-rater annotations. To address this issue, we propose to explicitly model the multi-rater (dis-)agreement, dubbed MRNet, which has two main contributions. First, an expertise-aware inferring module or EIM is devised to embed the expertise level of individual raters as prior knowledge, to form high-level semantic features. Second, our approach is capable of reconstructing multi-rater gradings from coarse predictions, with the multi-rater (dis-)agreement cues being further exploited to improve the segmentation performance. To our knowledge, our work is the first in producing calibrated predictions under different expertise levels for medical image segmentation. Extensive empirical experiments are conducted across five medical segmentation tasks of diverse imaging modalities. In these experiments, superior performance of our MRNet is observed comparing to the state-of-the-arts, indicating the effectiveness and applicability of our MRNet toward a wide range of medical segmentation tasks. Source code is publicly available.
Wei Ji 0011, Kai Ma 0002, Cheng Bian, Qi Bi, Hanruo Liu, Li Cheng 0001, Yefeng Zheng 0001
CVPR6
2021 Differential Convolution Feature Guided Deep Multi-Scale Multiple Instance Learning for Aerial Scene Classification
abstract
Aerial image classification is challenging for current deep learning models due to the varied geo-spatial object scales and the complicated scene spatial arrangement. Thus, it is necessary to stress the key local feature response from a variety of scales so as to represent discriminative convolutional features. In this paper, we propose a deep multi-scale multiple instance learning (DMSMIL) framework to tackle the above challenges. Firstly, we develop a differential multi-scale dilated convolution feature extractor to exploit the different patterns from different scales. Then, the deep features of each scale are fed into a multiple instance learning module to generate a bag-level probability prediction. Lastly, probability predictions from all the MIL branches are fused to generate the final semantic prediction. Extensive experiments on three widely-utilized aerial scene classification benchmarks demonstrate that our proposed DMSMIL outperforms the state-of-the-art approaches by a large margin.
Beichen Zhou, Jingjun Yi, Qi Bi
ICASSP3
2021 Temporal Cue Guided Video Highlight Detection with Low-Rank Audio-Visual Fusion
abstract
Video highlight detection plays an increasingly important role in social media content filtering, however, it remains highly challenging to develop automated video highlight detection methods because of the lack of temporal annotations (i.e., where the highlight moments are in long videos) for supervised learning. In this paper, we propose a novel weakly supervised method that can learn to detect highlights by mining video characteristics with video level annotations (topic tags) only. Particularly, we exploit audio-visual features to enhance video representation and take temporal cues into account for improving detection performance. Our contributions are threefold: 1) we propose an audio-visual tensor fusion mechanism that efficiently models the complex association between two modalities while reducing the gap of the heterogeneity between the two modalities; 2) we introduce a novel hierarchical temporal context encoder to embed local temporal clues in between neighboring segments; 3) finally, we alleviate the gradient vanishing problem theoretically during model optimization with attention-gated instance aggregation. Extensive experiments on two benchmark datasets (YouTube Highlights and TVSum) have demonstrated our method outperforms other state-of-the-art methods with remarkable improvements.
Qinghao Ye, Xiyue Shen, Yuan Gao 0017, Qi Bi, Ping Li 0006, Guang Yang 0006
ICCV5
2021 Local-Global Dual Perception Based Deep Multiple Instance Learning for Retinal Disease Classification
Qi Bi, Wei Ji 0011, Cheng Bian, Lijun Gong, Hanruo Liu, Kai Ma 0002, Yefeng Zheng 0001
MICCAI (8)1
2021 MIL-VT: Multiple Instance Learning Enhanced Vision Transformer for Fundus Image Classification
Kai Ma 0002, Qi Bi, Cheng Bian, Munan Ning, Nanjun He, Yuexiang Li, Hanruo Liu, Yefeng Zheng 0001
MICCAI (8)3
2021 Joint Semantic Mining for Weakly Supervised RGB-D Salient Object Detection
abstract
Training saliency detection models with weak supervisions, e.g., image-level tags or captions, is appealing as it removes the costly demand of per-pixel annotations. Despite the rapid progress of RGB-D saliency detection in fully-supervised setting, it however remains an unexplored territory when only weak supervision signals are available. This paper is set to tackle the problem of weakly-supervised RGB-D salient object detection. The key insight in this effort is the idea of maintaining per-pixel pseudo-labels with iterative refinements by reconciling the multimodal input signals in our joint semantic mining (JSM). Considering the large variations in the raw depth map and the lack of explicit pixel-level supervisions, we propose spatial semantic modeling (SSM) to capture saliency-specific depth cues from the raw depth and produce depth-refined pseudo-labels. Moreover, tags and captions are incorporated via a fill-in-the-blank training in our textual semantic modeling (TSM) to estimate the confidences of competing pseudo-labels. At test time, our model involves only a light-weight sub-network of the training pipeline, i.e., it requires only an RGB image as input, thus allowing efficient inference. Extensive evaluations demonstrate the effectiveness of our approach under the weakly-supervised setting. Importantly, our method could also be adapted to work in both fully-supervised and unsupervised paradigms. In each of these scenarios, superior performance has been attained by our approach with comparing to the state-of-the-art dedicated methods. As a by-product, a CapS dataset is constructed by augmenting existing benchmark training set with additional image tags and captions.
Wei Ji 0011, Qi Bi, Miao Zhang 0004, Yongri Piao, Huchuan Lu, Li Cheng 0001
NeurIPS3
2021 Multi-scale stacking attention pooling for remote sensing scene classification
Qi Bi, Han Zhang 0052, Kun Qin
Neurocomputing1
2021 Local Semantic Enhanced ConvNet for Aerial Scene Recognition
abstract
Aerial scene recognition is challenging due to the complicated object distribution and spatial arrangement in a large-scale aerial image. Recent studies attempt to explore the local semantic representation capability of deep learning models, but how to exactly perceive the key local regions remains to be handled. In this paper, we present a local semantic enhanced ConvNet (LSE-Net) for aerial scene recognition, which mimics the human visual perception of key local regions in aerial scenes, in the hope of building a discriminative local semantic representation. Our LSE-Net consists of a context enhanced convolutional feature extractor, a local semantic perception module and a classification layer. Firstly, we design a multi-scale dilated convolution operators to fuse multi-level and multi-scale convolutional features in a trainable manner in order to fully receive the local feature responses in an aerial scene. Then, these features are fed into our two-branch local semantic perception module. In this module, we design a context-aware class peak response (CACPR) measurement to precisely depict the visual impulse of key local regions and the corresponding context information. Also, a spatial attention weight matrix is extracted to describe the importance of each key local region for the aerial scene. Finally, the refined class confidence maps are fed into the classification layer. Exhaustive experiments on three aerial scene classification benchmarks indicate that our LSE-Net achieves the state-of-the-art performance, which validates the effectiveness of our local semantic perception module and CACPR measurement.
Qi Bi, Kun Qin, Han Zhang 0052, Gui-Song Xia
IEEE Trans. Image Process.1
2020 RADC-Net: A residual attention based convolution network for aerial scene classification
Qi Bi, Kun Qin, Han Zhang 0052, Zhili Li
Neurocomputing1
2020 APDC-Net: Attention Pooling-Based Convolutional Network for Aerial Scene Classification
abstract
Deep learning methods have boosted the performance of a series of visual tasks. However, the aerial image scene classification remains challenging. The object distribution and spatial arrangement in aerial scenes are often more complicated than in natural image scenes. Possible solutions include highlighting local semantics relevant to the scene label and preserving more discriminative features. To tackle this challenge, in this letter, we propose an attention pooling-based dense connected convolutional network (APDC-Net) for aerial scene classification. First, it uses a simplified dense connection structure as the backbone to preserve features from different levels. Then, we propose a trainable pooling to down-sample the feature maps and to enhance the local semantic representation capability. Finally, we introduce a multi-level supervision strategy, so that features from different levels are all allowed to supervise the training process directly. Exhaustive experiments on three aerial scene classification benchmarks demonstrate that our proposed APDC-Net outperforms other state-of-the-art methods with much fewer parameters and validate the effectiveness of our attention-based pooling and multi-level supervision strategy.
Qi Bi, Kun Qin, Han Zhang 0052, Jiafen Xie, Zhili Li
IEEE Geosci. Remote. Sens. Lett.1
2020 Deep Multiple Instance Convolutional Neural Networks for Learning Robust Scene Representations
abstract
The accuracy and efficiency of scene classification have immensely improved with the extensive application of deep convolutional neural networks (CNNs). However, standard CNNs classify images mostly based on the global features from the last fully connected layer, which may cause the negligence of discriminative local information and the sensitivity to various spatial transformations. In this article, we consider the problem of scene classification from the perspective of multiple instance learning (MIL) and propose an end-to-end multiple instance CNN (MI-CNN) for learning more robust scene representations. In MI-CNN, a scene is represented as a bag of local patches (instances). An instance-level classifier is trained to obtain the label of each patch in an MIL fashion, which makes the classifier more sensitive to the discriminative local patches. The patch labels are then aggregated into an image label by an MIL pooling layer, which is invariant to the order of local patches and helps construct more robust representations. We present extensive experiments on UC Merced Land use (UCM), Aerial Image data set (AID), and NWPU-RESISC (NWPU) data sets. Experimental results show that the proposed method achieves 1.17%, 1.70%, and 3.61% accuracy improvements with 90% parameter reduction compared with the standard CNNs.
Zhili Li, Jiafen Xie, Qi Bi, Kun Qin
IEEE Trans. Geosci. Remote. Sens.4
2020 A Multiple-Instance Densely-Connected ConvNet for Aerial Scene Classification
abstract
In contrast with nature scenes, aerial scenes are often composed of many objects crowdedly distributed on the surface in bird's view, the description of which usually demands more discriminative features as well as local semantics. However, when applied to scene classification, most of the existing convolution neural networks (ConvNets) tend to depict global semantics of images, and the loss of low- and mid-level features can hardly be avoided, especially when the model goes deeper. To tackle these challenges, in this paper, we propose a multiple-instance densely-connected ConvNet (MIDC-Net) for aerial scene classification. It regards aerial scene classification as a multiple-instance learning problem so that local semantics can be further investigated. Our classification model consists of an instance-level classifier, a multiple instance pooling and followed by a bag-level classification layer. In the instance-level classifier, we propose a simplified dense connection structure to effectively preserve features from different levels. The extracted convolution features are further converted into instance feature vectors. Then, we propose a trainable attention-based multiple instance pooling. It highlights the local semantics relevant to the scene label and outputs the bag-level probability directly. Finally, with our bag-level classification layer, this multiple instance learning framework is under the direct supervision of bag labels. Experiments on three widely-utilized aerial scene benchmarks demonstrate that our proposed method outperforms many state-of-the-art methods by a large margin with much fewer parameters.
Qi Bi, Kun Qin, Zhili Li, Han Zhang 0052, Gui-Song Xia
IEEE Trans. Image Process.1
2019 Multiple Instance Dense Connected Convolution Neural Network for Aerial Image Scene Classification
abstract
With the development of deep learning, many state-of-the-art natural image scene classification methods have demonstrated impressive performance. While the current convolution neural network tends to extract global features and global semantic information in a scene, the geo-spatial objects can be located at anywhere in an aerial image scene and their spatial arrangement tends to be more complicated. One possible solution is to preserve more local semantic information and enhance feature propagation. In this paper, an end to end multiple instance dense connected convolution neural network (MIDCCNN) is proposed for aerial image scene classification. First, a 23 layer dense connected convolution neural network (DCCNN) is built and served as a backbone to extract convolution features. It is capable of preserving middle and low level convolution features. Then, an attention based multiple instance pooling is proposed to highlight the local semantics in an aerial image scene. Finally, we minimize the loss between the bag-level predictions and the ground truth labels so that the whole framework can be trained directly. Experiments on three aerial image datasets demonstrate that our proposed methods can outperform current baselines by a large margin.
Qi Bi, Kun Qin, Zhili Li, Han Zhang 0052
ICIP1
2019 Joint precoding and scheduling algorithm for massive MIMO in FDD multi-cell network
Bin Han 0006, Zheng Jiang 0005, Peng Chen 0028, Fengyi Yang, Qi Bi
Wirel. Networks6
2013 Performance of LTE-Advanced macro-pico heterogeneous networks
abstract
Heterogenous network is currently under intensive discussion within 3GPP as an effective solution to improve network coverage and capacity for long term evolution (LTE)Advanced system, which consists of a mix of high-power macro base stations and a diverse set of low-power nodes, i.e., pico, femto and relay nodes. In this paper, we will particularly focus on the mixed deployment of macro and pico cells. Co-channel deployment is typically assumed for heterogenous network to make the best use of spectrum resource; therefore, a critical issue for heterogenous network is co-channel interference between macro and pico cells, which, if not properly addressed, may significantly degrade the performance of heterogenous network. Furthermore, other deployment options, e.g., the number and the positions of pi co nodes will also affect system performance greatly. In this paper, we will carry out extensive simulation campaigns to evaluate the downlink and uplink performance of co-channel macro-pico heterogeneous network. Results show that huge performance benefit can be obtained if heterogenous network is well designed and parameterized.
Xuetian Zhu, Zheng Jiang 0005, Fengyi Yang, Qi Bi
WCNC6
2006 Delay Performance of Enhanced Access Channel in 1xEV-DO Revision A Systems
abstract
In CDMA 1timesEV-DO Revision A system, an enhanced access channel (EAC) on reverse link with multiple transmission data rates is introduced with improved access performance for the general usage of connection setup. Data Over Signaling (DOS) protocol, also introduced in 1timesEV-DO Rev. A allows an access terminal to send and receive short user data messages over the EAC on reverse link and control channel on forward link DOS is useful for delay sensitive applications such as the call setup messages for push-to-talk and VoIP services. In this paper, we investigate the performance of EAC using both analytical and simulation models. In particular, the delay performance of DOS using the EAC with different data rates and message sizes is analyzed and compared. It is shown that the analytical and simulation results match very well. Furthermore, we propose a method for selecting access channel and traffic channel for transmitting data traffic. The proposed method takes better advantage of DOS for reduced latency and also improves the RF coverage.
Sigen Ye, Qinqing Zhang, Qi Bi
VTC Fall3
2006 An analysis of VoIP service using 1×EV-DO revision a system
abstract
While voice-over-Internet protocol (VoIP) on wireline network is maturing, VoIP on wireless mobile network is still in its infancy. This disparity is due to the fact that the wireline bandwidth is abundant and can be traded off for delay performance and overhead, whereas bandwidth in wireless mobile network is still a scarce resource. With the deployment of 1$times$EV-DO revision 0 (DOr0) worldwide, the spectrum efficiency has been significantly improved. However, DOr0 still lacks of features essential for VoIP. For this reason, 1$times$EV-DO revision A (DOrA) has been standardized in the 3GPP2 with many improvements favorable for VoIP implementation. In this paper, we identify challenges and explore the feasibility of implementing VoIP using DOrA. We develop both analytical and simulation models to evaluate the VoIP capacity and delay performance over the air interface.
Qi Bi, Pi-Chun Chen, Qinqing Zhang
IEEE J. Sel. Areas Commun.1
2006 Quality-of-service provisioning and efficient resource utilization in CDMA cellular communications
abstract
One of the major challenges in supporting multimedia services over Internet protocol (IP)-based code-division multiple-access (CDMA) wireless networks is the quality-of-service (QoS) provisioning with efficient resource utilization. Compared with the circuit-switched voice service in the second-generation CDMA systems (i.e., IS-95), heterogeneous multimedia applications in future IP-based CDMA networks require more complex QoS provisioning and more sophisticated management of the scarce radio resources. This paper provides an overview of the CDMA-related QoS provisioning techniques in the avenues of packet scheduling, power allocation, and network coordination, summarizes state-of-the-art research results, and identifies further research issues.
Hai Jiang 0001, Weihua Zhuang, Xuemin Shen, Qi Bi
IEEE J. Sel. Areas Commun.4
2005 Special issue: emerging multiple access technologies
Hsiao-Hwa Chen, Daoben Li, Qi Bi
Wirel. Commun. Mob. Comput.3
2003 The future evolution of wireless mobile communications
abstract
Abstract At the start of the 21st century, wireless mobile communications has witnessed an unprecedented growth fueled by information explosion and technology revolution. After much fierce technology competitions and uncertainties, two third‐generation (3G) wideband standards have emerged to dominate the wireless mobile communications for years to come. They are the CDMA2000 [ 1 ] and the universal mobile terrestrial system (UMTS) [ 2 ] standards, both of which are based on code‐division multiple‐access (CDMA) technology. In this paper, we shall indicate how current 2G digital wireless systems might evolve into these 3G systems. We shall examine the likely innovative techniques to be adopted for further evolution of the two systems and the possible capacity gains obtainable by utilizing these techniques. Copyright © 2003 John Wiley & Sons, Ltd.
Qi Bi, James P. Seymour
Wirel. Commun. Mob. Comput.1
2003 Ultra broadband wireless communications
Hsiao-Hwa Chen, Daoben Li, Qi Bi
Wirel. Commun. Mob. Comput.3
2003 Ultra-wideband wireless communications
abstract
Abstract Ultra‐wideband (UWB) communication techniques have attracted a great interest in both academia and industry in the past few years for applications in short‐range wireless mobile systems. This is due to the potential advantages of UWB transmissions such as low power, high rate, immunity to multipath propagation, less complex transceiver hardware, and low interference. However, tremendous R&D efforts are required to face various technical challenges in developing UWB wireless systems, including UWB channel characterization, transceiver design, coexistence and interworking with other narrowband wireless systems, design of the link and network layers to benefit from UWB transmission characteristics. This paper is to provide an overview of UWB communications, summarize the previous research results, and identify further research issues that need to be tackled. The emphasis is placed on the commercial wireless communications. Copyright © 2003 John Wiley & Sons, Ltd.
Weihua Zhuang, Xuemin Shen, Qi Bi
Wirel. Commun. Mob. Comput.3
2002 Forward and reverse link capacity for 1xEV-DO: third generation wireless high-speed data systems
abstract
This paper presents the performance results of a 1xEV-DO system. With different fading channels, data rate control lengths and overhead channel power levels, the system performance will vary. While optimizing the forward-link throughput becomes critical for asymmetric traffic demand, it is important to ensure that the reverse-link throughput is sufficient to support the data traffic ratio of 4:1 to 6:1 (forward/reverse). In this paper, the forward-link throughput is simulated. The final throughput is based on field measured channels for different morphologies. On the reverse link capacity, the throughput is analyzed based on an analytical formula that considers pilot, DRC and data channels.
Ching-Yao Huang, Qi Bi, Asif D. Gandhi, Ronald R. Brown, Dongzhe Cui
GLOBECOM2
2002 Performance analysis of 3G-1X EVDO high data rate system
abstract
The 3G-1X EVDO (Evolution Data Only) standard has been recently finalized by the 3GPP2 (3G Partnership Project 2) and published by the Telecommunications Industry Association as IS-856 (Interim Standard). This system allows data rates of up to 2.4 Mbit/s in 1.25 MHz bandwidth compatible with the frequency plan, of second and third generation systems based on IS-95 and IS-2000. Expected to be launched throughout the world in the near future, this system is capable of substantially enhancing end-user wireless data experience by providing high speed "always on" connectivity in a wide area mobile environment. In this paper we investigate various aspects of 3G-1X EVDO system performance through simulations and analysis. Link level simulator is used to evaluate rate assignment in different fading environments. Long-range prediction of fast fading is used to model operation of the 3G-1X EVDO mobile terminal. System level performance with multiple users in the sector is simulated using realistic assumptions of RF environment and mobility mix. Scheduler operation is included into system level simulation. We further propose a methodology for semianalytical evaluation of packet data sector throughput under assumption of bursty traffic. In this approach, an additional criterion is introduced to establish acceptable operating point for the system. By placing a restriction on a minimum tolerable user throughput (maximum allowed delay), we demonstrate a methodology for evaluating sector throughput in a consistent manner. This methodology ties together minimum performance guarantees for individual users and applications with maximum achievable system level performance and capacity. In this paper we present results demonstrating how the 3G-1X EVDO system capacity depends upon data traffic statistics and user requirements. The notion of system fairness, which is closely related to individual user performance guarantees, is also discussed with respect to 3G-1X EVDO performance.
Qi Bi, Stan Vitebsky
WCNC1
1992 Performance analysis of a CDMA cellular system in the multipath fading environment
abstract
One advantage of CDMA cellular systems is that they allow mobiles to transmit at any time instant without mutual synchronization. The cost of this flexibility is the tolerance of interference from other users in the same cell. In cellular applications, however, the effects of interference are further compounded by the multipath fading environment, because each interferer can generate several interfering rays. To reduce the effects of these interferences, M-ary signaling can be used together with the diversity antenna combining. A CDMA cellular system is analyzed that utilizes these techniques.>
Qi Bi
PIMRC1
1992 Fast full search equivalent encoding algorithms for image compression using vector quantization
abstract
Three fast search routines to be used in the encoding phase of vector quantization (VQ) image compression systems are presented. These routines, which are based on geometric considerations, provide the same results as an exhaustive (or full) search. Examples show that the proposed algorithms need only 3-20% of the number of mathematical operations required by a full search and fewer than 50% of the operations required by recently proposed alternatives.
C.-M. Huang, Qi Bi, Gardiner S. Stiles, R. W. Harris
IEEE Trans. Image Process.2
1992 Obtaining the recovery function from the interval histogram and the poststimulus time histogram
abstract
Timing patterns from auditory-nerve discharges are known to depend not only on the acoustic stimuli and the encoding process of the peripheral auditory system, but also on the refractory effects due to the recovery process of the neural cells. Thus, to study clues of encoding schemes of the peripheral auditory system, one has to separate refractory effects from the timing patterns. This can be achieved by defining a function, called the recovery function, that describes the refractory effects. The author demonstrates that this recovery function is readily obtainable from two histograms, the interval histogram and the poststimulus time histogram. Once the recovery function is known, refractory effects on the timing patterns can be easily removed.>
Qi Bi
IEEE Trans. Syst. Man Cybern.1
1984 A neural-counting model based on physiological characteristics of the peripheral auditory system. V. Application to loudness estimation and intensity discrimination
abstract
For pt.IV see ibid., vol.SMC-13, no.5, p.964-72 (1983). The psychophysical properties of a multiple-channel neural-counting model are investigated. Each channel represents a peripheral afferent fiber (or a group of such fibers) and consists of a cascade of signal-processing transformations, each of which has a physiological correlate in the auditory system. The acoustic signal is passed by a mathematical construct (which may be a pure tone or Gaussian noise) through a series of transformations. Spontaneous neural activity is independently incorporated into each channel by means of an additive refractoriness-modified Poisson process. A union process at a more distal center in the nervous system is generated by a parallel collection of such channels with a density (in frequency) determined by the cochlear mapping function. The statistics of the union count (in a fixed time) are then processed at a decision center in a manner that depends on the psychophysical paradigm under consideration. This random count number is assumed to contain all of the information for the examples considered. The model has been used to calculate psychophysical functions for pure-tone loudness estimation, pure-tone and variable-bandwidth noise intensity discrimination, and variable-bandwidth noise loudness summation. The theoretical results are in good agreement with human psychophysical data.
Gerard Lachs, R. Al-Shaikh, Qi Bi, Rosalie A. Saia, Malvin Carl Teich
IEEE Trans. Syst. Man Cybern.3