EDBT 2026 Demo / reviewers in the wild / expert
Ruiheng Zhang 0001
dblp:50/10861-1
· DBLP profile ↗
37ranked-venue papers
10as first author
33since 2021 · last 2026
0000-0002-5460-7196ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 16 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Superpixel-based Visual Feature Enhancement for Compositional Zero-Shot Learning
Wenlong Du, Xianglin Bao, Ruiheng Zhang 0001 |
Inf. Process. Manag. | 5 |
| 2026 | It Takes Two: Multi-Frequency Perception With Complementary Fusion Network for Complex Scene SegmentationabstractComplex scene segmentation aims to segment objects with intricate details or those concealed within the background. Despite significant advancements, a persistent challenge remains: accurately identifying object edges in backgrounds with high inherent similarity and complex structures. To address this, we identify the prevalent spectral bias in image segmentation, where networks preferentially learn low-frequency information, as a key impediment to recognizing and learning object edges, which are rich in high-frequency details. To mitigate this bias, we propose MCNet, a segmentation framework designed to promote balanced frequency learning. MCNet comprises two primary components: multi-frequency perception (MP), which independently captures high-frequency details and low-frequency structural components of objects, and complementary fusion (CF), which intelligently fuses these distinct frequency features through learnable, adaptive mechanisms. Crucially, MCNet employs a novel frequency-aware consistency adversarial loss to explicitly guide the learning across different frequency bands. MCNet effectively integrates MP and CF, enhancing the detection of high-frequency details and low-frequency structures, thereby alleviating challenges posed by spectral bias. We evaluate the proposed method on complex scene segmentation tasks, including camouflaged object detection and dichotomous image segmentation. Through extensive comparisons with 31 existing methods across 8 benchmark datasets, we demonstrate the superiority of the proposed method. Jin Zhang 0021, Ruiheng Zhang 0001, Zhe Cao 0001, Lixin Xu 0001, Xi Chen 0090, Min Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | RADCI: A Synchronized Radar-RGBT Object Detecting-Tracking Dataset And A BenchmarkabstractHigh-quality perception is crucial in autonomous driving and monitoring systems, where millimeter-wave radar and infrared cameras play important roles due to their robustness and reliability under harsh conditions. Both technologies can serve as low-cost supplements to optical image detection, improving overall system robustness. However, there is currently a lack of widely applicable feature-level fusion methods and multimodal datasets to effectively integrate visible light with these two heterogeneous data types for multiple tasks. In this work, we collect a new multimodal dataset, RADCI8, which synchronizes data from a camera, an infrared camera, and a radar for target detection and tracking. The dataset includes 2D image annotations, radar RAD tensor data with distance, angle, and Doppler information, as well as target ID annotations in both data formats. In addition, to address the incomplete use of radar data in previous fusion algorithms, we propose a detection method that fuses image and radar features using feature concatenation and an attention mechanism. Our proposed algorithm achieves 51.5% AP with an IOU of 50:95 on 2D bounding box prediction, significantly improving average detection accuracy over vision-based methods and maintaining robustness even when a single sensor degrades. Ruiheng Zhang 0001, Zhe Cao 0001, Biwen Yang, Jin Zhang 0021, Guanyu Liu |
ICASSP | 2 |
| 2025 | IRGPT: Understanding Real-World Infrared Image with Bi-Cross-Modal Curriculum on Large-Scale BenchmarkabstractReal-world infrared imagery presents unique challenges for vision-language models due to the scarcity of aligned text data and domain-specific characteristics. Although existing methods have advanced the field, their reliance on synthetic infrared images generated through style transfer from visible images, which limits their ability to capture the unique characteristics of the infrared modality. To address this, we propose IRGPT, the first multi-modal large language model for real-world infrared images, built upon a large-scale InfraRed-Text Dataset (IR-TD) comprising over 260K authentic image-text pairs. The proposed IR-TD dataset contains real infrared images paired with meticulously handcrafted texts, where the initial drafts originated from two complementary processes: (1) LLM-generated descriptions of visible images, and (2) rule-based descriptions of annotations. Furthermore, we introduce a bi-cross-modal curriculum transfer learning strategy that systematically transfers knowledge from visible to infrared domains by considering the difficulty scores of both infrared-visible and infrared-text. Evaluated on a benchmark of 9 tasks (e.g., recognition, grounding), IRGPT achieves state-of-the-art performance even compared with larger-scale models. Zhe Cao 0001, Jin Zhang 0021, Ruiheng Zhang 0001 |
ICCV | 3 |
| 2025 | MSD-HENet: Multi-Scale Detail-Preserving Holistic Enhancement Network for Infrared ImagesabstractInfrared images, widely utilized in various applications, often suffer from noise, contrast degradation, and detail loss. Existing image enhancement (IE) methods, predominantly designed for the RGB domain, often fail to perform effectively in the infrared domain. This suboptimal performance arises from their inadequate consideration of the intricate interplay between noise and contrast during multi-stage processing, which ultimately results in the loss of fine details. To address these challenges, this paper introduces the Multi-Scale Detail-Preserving Holistic Enhancement Network (MSD-HENet), a framework that leverages a novel interaction mechanism to achieve robust detail preservation while simultaneously denoising and contrast improvement. Specifically, a Detail Information Extractor (DIE) is proposed to effectively extract detail information through multi-scale differential convolution channels during the Deep Denoiser (D2) process, enabling the Contrast Improver (CI) to perform contrast enhancement without losing details, significantly enhancing overall image quality. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art approaches in terms of PSNR, SSIM, and visual effect. Yijing Zhao, Guanyu Liu, Ruiheng Zhang 0001 |
ICME | 5 |
| 2025 | MMCSBench: A Fine-Grained Benchmark for Large Vision-Language Models in Camouflage ScenesabstractCurrent camouflaged object detection methods predominantly follow discriminative segmentation paradigms and heavily rely on predefined categories present in the training data, limiting their generalization to unseen or emerging camouflage objects. This limitation is further compounded by the labor-intensive and time-consuming nature of collecting camouflage imagery. Although Large Vision-Language Models (LVLMs) show potential to improve such issues with their powerful generative capabilities, their understanding of camouflage scenes is still insufficient. To bridge this gap, we introduce MMCSBench, the first comprehensive multimodal benchmark designed to evaluate and advance LVLM capabilities in camouflage scenes. MMCSBench comprises 22,537 images and 76,843 corresponding image-text pairs across five fine-grained camouflage tasks. Additionally, we propose a new task, Camouflage Efficacy Assessment (CEA), aimed at quantitatively evaluating the camouflage effectiveness of objects in images and enabling automated collection of camouflage images from large-scale databases. Extensive experiments on 26 LVLMs reveal significant shortcomings in models' ability to perceive and interpret camouflage scenes. These findings highlight the fundamental differences between natural and camouflaged visual inputs, offering insights for future research in advancing LVLM capabilities within this challenging domain. Jing Zhang 0037, Ruiheng Zhang 0001, Zhe Cao 0001, Kaizheng Chen |
NeurIPS | 2 |
| 2025 | A novel compositional zero-shot learning approach based on hierarchical multi-scale feature fusion
Wenlong Du, Xianglin Bao, Ruiheng Zhang 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Breaking the alignment barrier: A spatiotemporal alignment-free RGBT tracking approach
Meibo Lv, Daming Zhou, Lingyu Si, Ruiheng Zhang 0001 |
Neurocomputing | 5 |
| 2025 | Two-stage sand-dust image enhancement method based on concentration scaling and domain adaptation
Yuting Yu, Zhidong Yang, Ruiheng Zhang 0001, Bosheng Ding, Lixin Xu 0001, He Zhao 0002 |
Knowl. Based Syst. | 3 |
| 2025 | A Novel Deep Generative Model via Semantic-Based Knowledge Distillation for Zero-Shot LearningabstractZero-Shot Learning (ZSL) aims to identify unseen target classes that lack training data. Most existing methods address the ZSL problem by generating samples of unseen classes based on the training data of seen classes and the semantic representations of unseen classes. However, due to the inherent limitations of ZSL, the generated unseen samples tend to be biased towards the data of seen classes, resulting in a label shift problem in the model's projection domain. To address these issues, we propose a novel generation-based ZSL approach that incorporates semantic-based constraints and knowledge distillation. Specifically, the semantic regularization and preservation constraints are designed to improve the distribution and discriminability of the generated unseen data, respectively. Furthermore, the semantic-based knowledge distillation strategy is introduced to enhance the generative model's feature encoding ability, thereby improving the quality of the generated unseen data. Extensive experiments on two standard ZSL benchmark datasets demonstrate that the proposed model achieves superior performance on both traditional and generalized ZSL tasks. Xianglin Bao, Ruiheng Zhang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Visible-Infrared Person Re-Identification With Real-World Label NoiseabstractIn recent years, growing needs for advanced security and traffic management have significantly heightened the prominence of the visible-infrared person re-identification community (VI-ReID), garnering considerable attention. A critical challenge in VI-ReID is the performance degradation attributable to label noise, an issue that becomes even more pronounced in cross-modal scenarios due to an increased likelihood of data confusion. While previous methods have achieved notable successes, they often overlook the complexities of instance-dependent and real-world noise, creating a disconnect from the practical applications of person re-identification. To bridge this gap, our research analyzes the primary sources of label noise in real-world settings, which include a) instantiated identities, b) blurry infrared images, and c) annotators’ errors. In response to these challenges, we develop a Robust Hybrid Loss function (RHL) that enables targeted recognition and retrieval optimization through a more fine-grained division of the noisy dataset. The proposed method categorises data into three sets: clean, obviously noisy, and indistinguishably noisy, with bespoke loss calculations for each category. The identification loss is structured to address the varied nature of these sets specifically. For the retrieval sub-task, we utilize an enhanced triplet loss, adept at handling noisy correspondences. Furthermore, to empirically validate our method, we have re-annotated a real-world dataset, SYSU-Real. Our experiments on SYSU-MM01 and RegDB, conducted under various noise ratios of random and instance-dependent label noise, demonstrate the generalized robustness and effectiveness of our proposed approach. Ruiheng Zhang 0001, Zhe Cao 0001, Yan Huang 0023, Shuo Yang 0006, Lixin Xu 0001, Min Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Dif-CDFusion: A Diffusion-Based Common-Differential Network for Infrared and Visible Image FusionabstractInfrared and visible image fusion aims to enhance scene representation by integrating complementary sensor data. However, existing methods fail to reconcile spectral fidelity with structural consistency. For one thing, grayscale fusion approaches preserve structural details by discarding color information, inherently sacrificing spectral fidelity. For another, color fusion techniques maintain spectral authenticity but compromise details and structural consistency due to the misaligned chromatic information. To bridge the gap, we present the Dif-CDFusion, which resolves the conflict between spectral fidelity and the preservation of structural details through diffusion-based feature extraction and common-differential alternating feature fusion. By individually constructing a denoising diffusion process in latent space to model multi-channel spectral distributions, our approach extracts diffusion features that preserve color integrity while capturing complete spectral information for texture retention. Subsequently, we design a common-differential alternate fusion module to alternately integrate differential and common mode components within diffusion features, enhancing both structual details and thermal target salience. Extensive experiments demonstrate that our Dif-CDFusion achieves state-of-the-art performance both quantitatively and qualitatively. The code and datasets are publicly available at https://github.com/ChickenEating/Dif-CDFusion. Guanyu Liu, Ruiheng Zhang 0001, Lixin Xu 0001, Qi Zhang 0004, Daming Zhou |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Detail-Aware Network for Infrared Image EnhancementabstractInfrared (IR) images inherently face the dual challenges of noise contamination and reduced contrast. However, existing image enhancement methods often overlook the intrinsic correlations between these factors—noise and low contrast—during multistage enhancement processes. Consequently, this oversight leads to a significant reduction in the fidelity of intricate details in IR images. In this article, we present a synergistic IR image enhancement network that simultaneously achieves denoising, contrast improvement, and detail preservation (DCDNet), which breaks down the overall enhancement process into more manageable steps. DCDNet is comprised of a detail awareness unit (DAU), a deep denoising prior (DDP), and a contrast improvement module (CIM). To maintain the details in the IR image, DAU is developed to extract the original detail feature information in DDP and integrate them into the CIM during contrast improvement to improve the final result. The detail information is derived from the encoder of the DDP, which focuses on denoising. The preserved detail features are subsequently incorporated into the decoder of the CIM, which is dedicated to enhancing contrast. Experimental results validate that our proposed approach surpasses other state-of-the-art methods for enhancing IR images in terms of the peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), visual information fidelity (VIF), and performance in downstream tasks. The code and dataset are publicly available athttps://github.com/ChickenEating/IR-Enhancement. Ruiheng Zhang 0001, Guanyu Liu, Qi Zhang 0004, Xiankai Lu, Renwei Dian, Yang Yang 0074, Lixin Xu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | A Benchmark and Frequency Compression Method for Infrared Few-Shot Object DetectionabstractInfrared few-shot object detection (IFSOD) aims to detect infrared objects with limited labeled examples. Current infrared datasets, however, suffer from limited diversity in object types and classes, hindering robust evaluation of model generalization on novel classes. To systematically assess dataset quality, we propose metrics for class diversity, instance variability, and object density. By integrating three widely used infrared datasets, we construct the first dataset specifically tailored for IFSOD, increasing instance density to 4.8 (a 1.1 improvement) and expanding the number of classes to 18 (a 5-class increase) compared to the source datasets. Furthermore, frequency analysis of spatial features reveals that sparse annotations introduce spectral bias in the frequency domain. Directly transforming spatial features to the frequency domain, however, mixes background noise with object features, causing spectral leakage and impairing the learning of discriminative features for novel classes. To address these issues, we propose the frequency compression few-shot detection (FC-fsd) method, which incorporates a frequency compression (FC) module. The FC module leverages Discrete Cosine Transform (DCT) within localized windows to reduce spectral leakage and enhance feature clarity. With minimal additional computational overhead, FC-fsd significantly outperforms state-of-the-art methods, achieving nAP50 scores of 28.57 (+13.37) and 35.63 (+2.59) in 1-shot and 2-shot settings, respectively. Our dataset is published athttps://github.com/RuihengZhang/IFSOD-dataset. Ruiheng Zhang 0001, Biwen Yang, Lixin Xu 0001, Yan Huang 0023, Qi Zhang 0070, Zhizhuo Jiang, Yu Liu 0005 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | FasterSal: Robust and Real-Time Single-Stream Architecture for RGB-D Salient Object DetectionabstractRGB-D Salient Object Detection (SOD) aims to segment the most prominent areas and objects in a given pair of RGB and depth images. Most current models adopt a dual-stream structure to extract information from both RGB and depth images. However, this leads to an exponential increase in the number of parameters and computations in the model. Moreover, the discrepancy between RGB pretrained and the 3D geometric relationships in depth maps present a challenge for the encoder in capturing spatial structural details. These issues impact the model's accuracy in locating salient objects and distinguishing edge details. To address these, we propose a novel early feature fusion network, named FasterSal, which enables more efficient RGB-D SOD. FasterSal uses a single stream structure to receive RGB images and depth maps, extracting features based on the 3D geometric relationships in the depth map while fully leveraging the pretrained RGB encoder. This approach effectively avoids the inconsistencies between depth modality and the RGB pretrained encoder. It also significantly reduces the number of network parameters while maintaining efficient feature encoding capabilities. To achieve finer edge learning, the detail-aware loss and texture enhancement module are introduced. These modules are designed to extract latent details in high-frequency component features and to enhance the edge learning capability of the model using distance information. Experimental results on several benchmark datasets confirm the effectiveness and superiority of our method over the state-of-the-art approaches, achieving a good balance between performance and speed with only 3.4 million parameters and a CPU operating speed of 63 FPS. Jing Zhang 0037, Ruiheng Zhang 0001, Lixin Xu 0001, Xiankai Lu, Yushu Yu, Min Xu 0001, He Zhao 0002 |
IEEE Trans. Multim. | 2 |
| 2025 | Cognition-Driven Structural Prior for Instance-Dependent Label Transition Matrix EstimationabstractThe label transition matrix has emerged as a widely accepted method for mitigating label noise in machine learning. In recent years, numerous studies have centered on leveraging deep neural networks to estimate the label transition matrix for individual instances within the context of instance-dependent noise. However, these methods suffer from low search efficiency due to the large space of feasible solutions. Behind this drawback, we have explored that the real murderer lies in the invalid class transitions, that is, the actual transition probability between certain classes is zero but is estimated to have a certain value. To mask the invalid class transitions, we introduced a human-cognition-assisted method with structural information from human cognition. Specifically, we introduce a structured transition matrix network (STMN) designed with an adversarial learning process to balance instance features and prior information from human cognition. The proposed method offers two advantages: 1) better estimation effectiveness is obtained by sparing the transition matrix and 2) better estimation accuracy is obtained with the assistance of human cognition. By exploiting these two advantages, our method parametrically estimates a sparse label transition matrix, effectively converting noisy labels into true labels. The efficiency and superiority of our proposed method are substantiated through comprehensive comparisons with state-of-the-art methods on three synthetic datasets and a real-world dataset. Our code will be available at https://github.com/WheatCao/STMN-Pytorch. Ruiheng Zhang 0001, Zhe Cao 0001, Shuo Yang 0006, Lingyu Si, Lixin Xu 0001, Fuchun Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | ALMA: Adjustable Location and Multi-Angle Attention for Fine-Grained Visual ClassificationabstractFine-grained visual classification (FGVC) is a challenging but realistic problem that recognizes objects from common categories with subtle differences. Most previous work focused on identifying more regional features while neglecting the fact that these regions still contain a large amount of secondary information. To alleviate the interference of the secondary information, in this paper, we propose a novel Adjustable Location and Multi-angle Attention (ALMA) network to solve the FGVC problem. ALMA consists of two branches, i.e. the adjustable location module and the multi-angle attention module. Specifically, in the adjustable localization module, we first locate the interested area of the object and obtain the adjusted cropped area by adjusting the interested area through the background masking. Then, the adjusted regions will be gathered to locate objects with better prediction performance. Furthermore, we design the multi-angle attention module to gradually maximize the difference between the original attention map and the randomly selected attention map. Consequently, the model can focus on the main information which represents the entire object. To evaluate the effectiveness of the proposed model, we conduct extensive experiments on three public fine-grained benchmark datasets. Experimental results demonstrate that the proposed ALMA model has significant superiority over other FGVC methods. Boyu Ding, Xianglin Bao, Ruiheng Zhang 0001 |
CSCWD | 5 |
| 2024 | Learning Camouflaged Object Detection from Noisy Pseudo Label
Jin Zhang 0021, Ruiheng Zhang 0001, Yanjiao Shi, Zhe Cao 0001, Nian Liu 0002, Fahad Shahbaz Khan |
ECCV (1) | 2 |
| 2024 | Mind the Boundary: Coreset Selection via Reconstructing the Decision BoundaryabstractExisting paradigms of pushing the state of the art require exponentially more training data in many fields. Coreset selection seeks to mitigate this growing demand by identifying the most efficient subset of training data. In this paper, we delve into geometry-based coreset methods and preliminarily link the geometry of data distribution with models’ generalization capability in theoretics. Leveraging these theoretical insights, we propose a novel coreset construction method by selecting training samples to reconstruct the decision boundary of a deep neural network learned on the full dataset. Extensive experiments across various popular benchmarks demonstrate the superiority of our method over multiple competitors. For the first time, our method achieves a 50% data pruning rate on the ImageNet-1K dataset while sacrificing less than 1% in accuracy. Additionally, we showcase and analyze the remarkable cross-architecture transferability of the coresets derived from our approach. Shuo Yang 0006, Zhe Cao 0001, Ruiheng Zhang 0001, Ping Luo 0002, Shengping Zhang, Liqiang Nie |
ICML | 4 |
| 2024 | Tight Fusion of Odometry and Kinematic Constraints for Multiple Aerial Vehicles in Physical InterconnectionabstractIntegrated aerial Platforms (IAPs), comprising multiple aircrafts, are typically fully actuated and hold significant potential for aerial manipulation tasks. Differing from a multiple aerial swarm, the aircrafts within the IAP are interconnected, presenting promising opportunities for enhancing localization. Incorporating the physical constraints of these multiple aircrafts to improve the accuracy and reliability of integrated aircraft positioning and navigation systems is a challenging yet highly significant problem. In this paper, we introduce a distributed multi-aircraft visual-inertial-range odometry system that analyzes the position, velocity, and attitude constraints within the IAP. Leveraging constraint relationships in the IAP, we propose corresponding methods that tightly fuse visual-inertial-range odometry and kinematic constraints to optimize odometry accuracy. Our system’s performance is validated using a collected dataset, resulting in a notable 28.7% reduction in drift compared to the baseline. Yingjun Fan, Chuanbeibei Shi, Ganghua Lai, Ruiheng Zhang 0001, Yushu Yu, Fuchun Sun 0001, Yiqun Dong |
ICRA | 4 |
| 2024 | ABNT: Attention Binary Navigation Tree for Fine-Grained Visual ClassificationabstractFine-grained visual categorization (FGVC) presents a notable challenge owing to the high intra-class variance and minimal inter-class variability. In multi-stage FGVC tasks, the initial attention region’s influence on subsequent stages proves profoundly significant. Prior approaches tended to excessively concentrate on discriminatory regions, which overlooked the crucial aspect of effectively focusing on objects in the initial phase. In this paper, we propose the Attention Binary Navigation Tree (ABNT) model to augment the discernment capabilities of the initial leaf node when distinguishing between object and background information. This strategy makes the model direct the attention to the object and can provide more effective guidance for subsequent stages. Moreover, multiple branch routing modules are integrated via decision trees to rationally distribute the contribution of each leaf node. Subsequently, predictions from the leaf nodes are aggregated to obtain the final decision. Extensive experiments on two fine-grained benchmark datasets are conducted to validate the effectiveness of the proposed model. Experimental results demonstrate the marked superiority of the proposed ABNT model over other state-of-the-art FGVC methods. Boyu Ding, Xianglin Bao, Ruiheng Zhang 0001 |
IJCNN | 5 |
| 2024 | Concentrating Estimation Attention: Human Prior Constrained Methods for Robust Classification
Zhe Cao 0001, Shuo Yang 0006, Hongbin Pei, Yan Huang 0023, Yushu Yu, Ruiheng Zhang 0001 |
PRCV (15) | 8 |
| 2024 | Completing Saliency from Details
Jin Zhang 0021, Lingxiang Wu, Renwei Dian, Yiheng Yao, Shihao Huang, Yang Yang 0074, Ruiheng Zhang 0001 |
PRCV (13) | 8 |
| 2024 | Structural Transformer with Region Strip Attention for Video Object Segmentation
Qingfeng Guan 0002, Hao Fang 0010, Chenchen Han, Zhicheng Wang 0017, Ruiheng Zhang 0001, Xiankai Lu |
Neurocomputing | 5 |
| 2024 | Wireless Localization and Formation Control With Asynchronous AgentsabstractThe formation control of multi-agent systems has increasingly drawn attention for fulfilling numerous emerging applications and services. To achieve high-accuracy formation, the location awareness of all agents becomes an essential requirement. In this paper, we address the problem of network localization and formation control in a cooperative system with asynchronous agents. In particular, we formulate the joint localization and synchronization of agents as a statistical inference problem. The underlying probabilistic model is represented by a factor graph from which a message-passing algorithm is designed that computes approximations of the marginals of unknown variables, i.e. agents’ locations and clock offsets. Due to the Euclidean-norm operator involved in their computation no parametric closed-form expressions of the messages exist. As a compromise, implemented message-passing methods therefore resort to approximations of these messages. Conventional methods rely either on a first-order Taylor expansion of the norm operation or on non-parametric representations, e.g. by means particle filters (PFs), to compute such approximations. However, the former approach suffers from poor performance while the latter one experiences high complexity. The proposed message-passing algorithm in this paper is parametric. Specifically, it passes Gaussian messages that can be essentially obtained by suitably augmenting the factor graph and applying on it a hybrid method for combining belief propagation and variational message passing. Subsequently, the agents can exploit the estimated locations for determining the control policy. Two types of control policy are designed based on the optimization of a generalized cost function. We show that the proposed scheme enjoys a reduced complexity for multi-agent localization while achieving the desired formation with excellent accuracy. Weijie Yuan 0001, Zhaohui Yang 0001, Liangming Chen, Ruiheng Zhang 0001, Yiheng Yao, Yuanhao Cui, Hong Zhang 0013, Derrick Wing Kwan Ng |
IEEE J. Sel. Areas Commun. | 4 |
| 2024 | Differential Feature Awareness Network Within Antagonistic Learning for Infrared-Visible Object DetectionabstractThe combination of infrared and visible videos aims to gather more comprehensive feature information from multiple sources and reach superior results on various practical tasks, such as detection and segmentation, over that of a single modality. However, most existing dual-modality object detection algorithms ignore the modal differences and fail to consider the correlation between feature extraction and fusion, which leads to incomplete extraction and inadequate fusion of dual-modality features. Hence, there raises an issue of how to preserve each unique modal feature and fully utilize the complementary infrared and visible information. Facing the above challenges, we propose a novel Differential Feature Awareness Network (DFANet) within antagonistic learning for infrared and visible object detection. The proposed model consists of an Antagonistic Feature Extraction with Divergence (AFED) module used to extract the differential infrared and visible features with unique information, and an Attention-based Differential Feature Fusion (ADFF) module used to fully fuse the extracted differential features. We conduct performance comparisons with existing state-of-the-art models on two benchmark datasets to represent the robustness and superiority of DFANet, and numerous ablation experiments to illustrate its effectiveness. Ruiheng Zhang 0001, Qi Zhang 0004, Jin Zhang 0021, Lixin Xu 0001, Baomin Zhang, Binglu Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | U2D2Net: Unsupervised Unified Image Dehazing and Denoising Network for Single Hazy Image EnhancementabstractHazy images captured under ill-posed scenarios with scattering medium (i.e. haze, fog, or smoke) are contaminated in visibility. Inevitably, these images are further degraded by noises owing to real-world imaging. Most existing hazy image enhancement methods perform image dehazing and denoising stage by stage, with the undesirable result that the estimation error of the former stage has to be propagated and amplified in the latter stage, e.g., noise amplification after dehazing. To address this inconsistent degradation, we present an Unsupervised Unified Image Dehazing and Denoising Network, U2D2Net, to remove the haze and suppress the noise simultaneously for a single hazy image. U2D2Net is mainly comprised of an unsupervised dehazing module, an unsupervised denoising module, and a region-similarity fusion strategy. Specifically, we propose an unsupervised transmission-aware dehazing module to restore visibility and suppress depth-dependent noise propagation in the dehazing module. Besides, we design an unsupervised network with a Mean/Max Sub-Sampler in the denoising module. To exploit the correlation and complementary between the previous outputs, a region-similarity fusion strategy is developed to compute the final qualified result. Extensive experiments on both synthetic and real-world datasets illustrate that U2D2Net outperforms other state-of-the-art dehazing and denoising methods in terms of PSNR, SSIM, and subjective visual effects. Bosheng Ding, Ruiheng Zhang 0001, Lixin Xu 0001, Guanyu Liu, Shuo Yang 0006, Qi Zhang 0004 |
IEEE Trans. Multim. | 2 |
| 2024 | Part-Aware Correlation Networks for Few-Shot LearningabstractFew-shot learning brings the machine close to human thinking which enables fast learning with limited samples. Recent work considers local features to achieve contextual semantic complementation, while they are merely coarsened feature observations that can only extract insignificant label correlations. On the contrary, partial properties of few-shot examples significantly draw the implicit feature observations that can reveal the underlying label correlation of rare label classification. To fully explore the correlation between labels and partial features, this paper proposes a Part-Aware Correlation Network (PACNet) based on Partial Representation (PR) and Semantic Covariance Matrix (SCM). Specifically, we develop a partial representing module of an object that eliminates object-independent information and allows the model to focus on more distinctive parts. Furthermore, a semantic covariance measure function is redefined as a way to learn the semantic relationships of partial representations and to compute the partial similarity between the query sample and the support set. Experiments on three benchmark datasets consistently show that the proposed method outperforms the state-of-the-art counterparts,e.g., on the PartImageNet dataset, the performance gains of up to 12% and 5.9% are observed for the 5-way 1-shot and 5-way 5-shot settings, respectively. Ruiheng Zhang 0001, Jinyu Tan, Zhe Cao 0001, Lixin Xu 0001, Lingyu Si, Fuchun Sun 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Few-Shot Infrared Image Classification with Partial Concept Feature
Jinyu Tan, Ruiheng Zhang 0001, Qi Zhang 0004, Zhe Cao 0001, Lixin Xu 0001 |
PRCV (4) | 2 |
| 2023 | An end-to-end deep generative approach with meta-learning optimization for zero-shot object classification
Xianglin Bao, Ruiheng Zhang 0001, Guifu Lu |
Inf. Process. Manag. | 4 |
| 2022 | Objects in Semantic Topology
Shuo Yang 0006, Peize Sun, Yi Jiang 0009, Xiaobo Xia, Ruiheng Zhang 0001, Zehuan Yuan, Changhu Wang, Ping Luo 0002, Min Xu 0001 |
ICLR | 5 |
| 2022 | Graph-based few-shot learning with transformed feature propagation and optimal class allocation
Ruiheng Zhang 0001, Shuo Yang 0006, Qi Zhang 0004, Lixin Xu 0001, Yang He 0002, Fan Zhang 0007 |
Neurocomputing | 1 |
| 2022 | Deep-IRTarget: An Automatic Target Detector in Infrared Imagery Using Dual-Domain Feature Extraction and AllocationabstractRecently, convolutional neural networks (CNNs) have brought impressive improvements for object detection. However, detecting targets in infrared images still remains challenging, because the poor texture information, low resolution and high noise levels of the thermal imagery restrict the feature extraction ability of CNNs. In order to deal with these difficulties in the feature extraction, we propose a novel backbone network named Deep-IRTarget, composing of a frequency feature extractor, a spatial feature extractor and a dual-domain feature resource allocation model. Hypercomplex Infrared Fourier Transform is developed to calculate the infrared intensity saliency by designing hypercomplex representations in the frequency domain, while a convolutional neural network is invoked to extract feature maps in the spatial domain. Features from the frequency domain and spatial domain are stacked to construct Dual-domain features. To efficiently integrate and recalibrate them, we propose a Resource Allocation model for Features (RAF). The well-designed channel attention block and position attention block are used in RAF to respectively extract interdependent relationships among channel and position dimensions, and capture channel-wise and position-wise contextual information. Extensive experiments are conducted on three challenging infrared imagery databases. We achieve 10.14%, 9.1% and 8.05% improvement on mAP scores, compared to the current state of the art method on MWIR, BITIR and WCIR respectively. Ruiheng Zhang 0001, Lixin Xu 0001, Zhengyu Yu, Chengpo Mu, Min Xu 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Infrared Target Detection Using Intensity Saliency And Self-AttentionabstractInfrared target detection is essential for many computer vision tasks. Generally, the IR images present common infrared characteristics, such as poor texture information, low resolution, and high noise. However, these characteristics are ignored in the existing detection methods, making them fail in real-world scenarios. In this paper, we take infrared intensity into account and propose a novel backbone network named Deep-IRTarget. We first extract infrared intensity saliency by a convolution with a Gaussian kernel filtering the images in the frequency domain. We then propose the triple self-attention network to further extract spatial domain image saliency by selectively emphasize interdependent semantic features in each channel. Jointly exploiting infrared characteristics in the frequency domain and the overall semantic interdependencies in the spatial domain, the proposed Deep-IRTarget outperforms existing methods in real-world Infrared target detection tasks. Experimental results on two infrared imagery datasets demonstrate the superiorly of our model. Ruiheng Zhang 0001, Min Xu 0001, Yaxin Shi, Jian Fan, Chengpo Mu, Lixin Xu 0001 |
ICIP | 1 |
| 2020 | Multi-camera Sports Players 3D Localization with Identification ReasoningabstractMulti-camera sports players 3D localization is always a challenging task due to heavy occlusions in crowded sports scenes. Traditional methods can only provide players' locations without identifying information. Existing methods of localization may cause ambiguous detection and unsatisfactory precision and recall, especially when heavy occlusions occur. To solve this problem, we propose a generic localization method by providing distinguishable results that have the probabilities of locations being occupied by players with unique ID labels. We design the algorithms with a multi-dimensional Bayesian model to create a Probabilistic and Identified Occupancy Map (PIOM). By using this model, we jointly apply deep learning-based object segmentation and identification to obtain sports players' probable positions and their likely identification labels. This approach not only provides players 3D locations but also gives their ID information that is distinguishable from others. Experimental results demonstrate that our method outperforms the previous localization approaches with reliable and distinguishable outcomes. Yukun Yang 0003, Ruiheng Zhang 0001, Wanneng Wu, Min Xu 0001 |
ICPR | 2 |
| 2020 | Multi-camera multi-player tracking with deep player identification in sports video
Ruiheng Zhang 0001, Lingxiang Wu, Yukun Yang 0003, Wanneng Wu, Yueqiang Chen, Min Xu 0001 |
Pattern Recognit. | 1 |
| 2019 | Learning Image-Specific Attributes by Hyperbolic Neighborhood Graph PropagationabstractAs a kind of semantic representation of visual object descriptions, attributes are widely used in various computer vision tasks. In most of existing attribute-based research, class-specific attributes (CSA), which are class-level annotations, are usually adopted due to its low annotation cost for each class instead of each individual image. However, class-specific attributes are usually noisy because of annotation errors and diversity of individual images. Therefore, it is desirable to obtain image-specific attributes (ISA), which are image-level annotations, from the original class-specific attributes. In this paper, we propose to learn image-specific attributes by graph-based attribute propagation. Considering the intrinsic property of hyperbolic geometry that its distance expands exponentially, hyperbolic neighborhood graph (HNG) is constructed to characterize the relationship between samples. Based on HNG, we define neighborhood consistency for each sample to identify inconsistent samples. Subsequently, inconsistent samples are refined based on their neighbors in HNG. Extensive experiments on five benchmark datasets demonstrate the significant superiority of the learned image-specific attributes over the original class-specific attributes in the zero-shot object classification task. Ivor W. Tsang, Xiaofeng Cao 0002, Ruiheng Zhang 0001, Chuancai Liu |
IJCAI | 4 |