VLDB 2026 Research / reviewers in the wild / expert
Mao Ye 0001
dblp:36/2301-1
· DBLP profile ↗
155ranked-venue papers
13as first author
97since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 99 · 13 first-author · 59 since 2021Graphics, computer vision, multimedia, augmented reality and games · 61 · 42 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Databases, data management, data science and information retrieval · 7 · 4 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domain-Auxiliary Infrared Moving Small Target Detection by Learning to Overlook Domain DiscrepancyabstractCurrently, almost all traditional infrared small target detection methods work on the assumption that training and test sets always belong to the same domain, and training samples are sufficient. However, in real applications, a new detection task could often have no sufficient training samples from a special domain. In this situation, adopting the auxiliary data from big-sample domains is usually believed to be one of the most potential solutions. However, exceeding expectations, it is found that simply adding auxiliary samples cannot often be always effective, even causing performance decline, due to existing infrared domain shift. To overcome this unexpected problem, we propose the first infrared moving small target detection framework with domain-auxiliary supports by Learning to Overlook Domain Discrepancy (Loddis). This framework consists of three primary processing stages: correlation weakening, domain confusing, and target consistency contrastive learning. Breaking through traditional learning paradigm, through auxiliary data, it enables the model to focus more on targets themselves, and less on image backgrounds, minimizing the sensitivity to domain discrepancy. The extensive experiments on 6 different-domain datasets show the effectiveness and superiority of the proposed Loddis framework for infrared small target detection. Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001 |
AAAI | 5 |
| 2026 | Cross-domain Joint Learning with Prototype-guided Mixture-of-Experts for Infrared Moving Small Target DetectionabstractInfrared small target detection often faces significant domain gaps across datasets due to varying sensors and scene distributions. Currently, most existing methods are typically based on single-domain learning (i.e., training and test are on the same dataset), requiring training separate detectors when considering different datasets. However, they overlook the valuable public knowledge across domains and limit the applicability in multiple infrared scenarios. To break through single-domain learning, implementing only one universal detector simultaneously on multiple datasets, as the first exploration, we propose a cross-domain joint learning task framework with prototype-guided Mixture-of-Experts (CoMoE). Specifically, it designs a hyperspherical prototype learning to adaptively maintain both domain-specific prototypes and global prototypes, enhancing cross-domain feature representation. Meanwhile, a domain-aware Mixture-of-Experts with Top-K routing strategy is proposed to select the optimal domain experts. Moreover, to enhance cross-domain feature alignment, we design an adaptive cross-domain feature modulation with noise-guided contrastive learning. The extensive experiments on a newly constructed benchmark comprising three datasets verify the superiority of our CoMoE, even under limited data settings. It could often surpass general joint learning methods, and state-of-the-art (SOTA) single-domain ones. Luping Ji, Jianghong Huang, Sicheng Zhu, Mao Ye 0001 |
AAAI | 5 |
| 2026 | Hierarchical Frequency-Guided Alignment Transformer for Compressed Video Quality EnhancementabstractDuring the video encoding process, the original spatial domain signal is first transformed into the frequency domain, followed by quantization and compression. As a result, the quality degradation in compressed videos primarily stems from distortions in the frequency domain information. However, existing video enhancement methods typically directly fuse information from adjacent frames in the spatial domain, making it difficult for models to effectively compensate for frequency domain distortions, which leads to suboptimal detail restoration. To address this issue, we propose a Hierarchical Frequency-Guided Alignment Transformer. Additionally, by analyzing the characteristics of the frequency domain, we find that different frequency bands exhibit both correlations and a certain degree of independence. Based on this, we introduce a Frequency-Aware Transformer module that employs a combination of independent and mixed processing to optimize information exchange across different frequency domains, effectively mitigating cross-interference from irrelevant information. Experimental results demonstrate that, compared to existing methods, our approach achieves state-of-the-art performance in objective metrics (PSNR/SSIM), perceptual quality (LPIPS), and subjective visual effects, while reducing model complexity. Liuhan Peng, Shuai Li 0005, Yanbo Gao, Mao Ye 0001, Chong Lv |
AAAI | 4 |
| 2026 | Multi-exposure high dynamic range reconstruction by incorporating imaging knowledge
Mao Ye 0001, Dengyan Luo, Yan Gan |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Progressive category prototype optimization for black-box domain adaptation
Lihua Zhou, Song Tang 0001, Yan Gan, Mao Ye 0001 |
Neurocomputing | 5 |
| 2026 | Language-guided cross-modal collaborative learning for unsupervised domain adaptation
Ai Peng, Zhouli Shen, Mao Ye 0001, Jianwei Zhang 0001 |
Neurocomputing | 5 |
| 2026 | FAST: Foreground-aware active self-training for domain adaptive object detection
Hongmin Deng, Hailin Wang 0002, Zhekai Du, Guisong Liu, Jingjing Li 0001, Mao Ye 0001 |
Neural Networks | 7 |
| 2026 | BeltCrack: The first sequential-image industrial conveyor belt crack detection dataset and its baseline with triple-domain feature learning
Jianghong Huang, Luping Ji, Mao Ye 0001 |
Pattern Recognit. | 4 |
| 2026 | Deformable Feature Alignment and Refinement for moving infrared small target detection
Dengyan Luo, Yanping Xiang, Luping Ji, Shuai Li 0005, Mao Ye 0001 |
Pattern Recognit. | 6 |
| 2026 | Source-Free Domain Adaptive Object Detection with semantics compensation
Song Tang 0001, Jiuzheng Yang, Mao Ye 0001, Yan Gan, Xiatian Zhu |
Pattern Recognit. | 3 |
| 2026 | Adaptive Feature Enhancement for SAR-Based Ship DetectionabstractSynthetic Aperture Radar (SAR) enables reliable maritime surveillance, yet ship detection remains challenging due to limited training data and complex ocean clutter. To address these issues, we propose a concise yet effective two-fold enhancement framework for robust SAR ship detection. First, a target-driven data enhancement strategy is introduced by extracting ship targets and synthesizing them with pure ocean backgrounds to reduce the impact of ship wakes and complex land backgrounds on ship data generation. Second, an adaptive feature enhancement mechanism based on a Mixture of Experts (MoE) is designed, where a gating network dynamically selects specialized expert sub-networks to refine multi-scale feature representations. This adaptive routing enables spatially aware feature enhancement and improves discriminability under complex maritime conditions. Extensive experiments on benchmark SAR datasets demonstrate that the proposed framework consistently outperforms state-of-the-art approaches, achieving superior detection accuracy and robustness. These results validate the effectiveness of combining data-level enhancement with expert-driven feature refinement for SAR ship detection. Mao Ye 0001 |
IEEE Signal Process. Lett. | 2 |
| 2026 | A Noise Constrained Diffusion (NC-Diffusion) Framework for High-Fidelity Image CompressionabstractWith the great success of diffusion models in image generation, diffusion-based image compression is attracting increasing interests. However, due to the random noise introduced in the diffusion learning, they usually produce reconstructions with deviation from the original images, leading to suboptimal compression results. To address this problem, in this paper, we propose a Noise Constrained Diffusion (NC-Diffusion) framework for high fidelity image compression. Unlike existing diffusion-based compression methods that add random Gaussian noise and direct the noise into the image space, the proposed NC-Diffusion formulates the quantization noise originally added in the learned image compression as the noise in the forward process of diffusion. Then a noise constrained diffusion process is constructed from the ground-truth image to the initial compression result generated with quantization noise. The NC-Diffusion overcomes the problem of noise mismatch between compression and diffusion, significantly improving the inference efficiency. In addition, an adaptive frequency-domain filtering module is developed to enhance the skip connections in the U-Net based diffusion architecture, in order to enhance high-frequency details. Moreover, a zero-shot sample-guided enhancement method is designed to further improve the fidelity of the image. Experiments on multiple benchmark datasets demonstrate that our method can achieve the best performance compared with existing methods. Yanbo Gao, Shuai Li 0005, Hui Yuan 0001, Mao Ye 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | From Point to Flow: Enhancing Unsupervised Domain Adaptation With Flow ClassificationabstractUnsupervised domain adaptation aims to transfer knowledge from a labeled source domain to an unlabeled target domain. Existing methods, whether based on distribution matching or self-supervised learning, often focus solely on classifying individual source samples, potentially overlooking discriminative information. To address this limitation, we propose FlowUDA, a novel plugin method that enhances existing UDA frameworks by constructing semantically invariant flows from individual source samples to corresponding target samples, forming cross-domain trajectories. By leveraging a diffusion network guided by ordinary differential equations, FlowUDA ensures these flows preserve the topological structure of the source domain, maintaining their distinguishability. Our method then classifies these flows by sampling points along them and transferring labels from source samples, effectively capturing spatial relationships between domains. In essence, FlowUDA transforms the traditional point-based classification on individual source samples into flow-based classification on flows, allowing the model to learn richer, more discriminative features that bridge the gap between source and target domains. Extensive experiments on standard benchmarks demonstrate that integrating FlowUDA into existing UDA methods leads to notable performance gains, highlighting its effectiveness in addressing domain shift challenges. Lihua Zhou, Mao Ye 0001, Nianxin Li, Song Tang 0001, Xu-Qian Fan, Lei Deng 0001, Zhen Lei 0001, Xiatian Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | CWRNN-INVR: A Coupled WarpRNN Based Implicit Neural Video RepresentationabstractImplicit Neural Video Representation (INVR) has emerged as a novel approach for video representation and compression, using learnable grids and neural networks. Existing methods focus on developing new grid structures efficient for latent representation and neural network architectures with large representation capability, lacking the study on their roles in video representation. In this paper, the difference between INVR based on neural network and INVR based on grid is first investigated from the perspective of video information composition to specify their own advantages, i.e., neural network for general structure while grid for specific detail. Accordingly, an INVR based on mixed neural network and residual grid framework is proposed, where the neural network is used to represent the regular and structured information and the residual grid is used to represent the remaining irregular information in a video. A Coupled WarpRNN-based multi-scale motion representation and compensation module is specifically designed to explicitly represent the regular and structured information, thus terming our method as CWRNN-INVR. For the irregular information, a mixed residual grid is learned where the irregular appearance and motion information are represented together. The mixed residual grid can be combined with the coupled WarpRNN in a way that allows for network reuse. Experiments show that our method achieves the best reconstruction results compared with the existing methods, with an average PSNR of 33.73 dB on the UVG dataset under the 3M model and outperforms existing INVR methods in other downstream tasks. The code can be found athttps://github.com/yiyang-sdu/CWRNN-INVR.git. Yanbo Gao, Shuai Li 0005, Jinglin Zhang 0001, Hui Yuan 0001, Mao Ye 0001, Xingyu Gao 0001 |
IEEE Trans. Multim. | 7 |
| 2026 | Long-Short Match for Lost Control in UAV Multi-Object TrackingabstractMulti-Object Tracking (MOT) in Unmanned Aerial Vehicles (UAV) aims to continuously and stably detect and track objects in videos captured by UAVs. In existing MOT tracking-by-detection schemes, the tracker with a fixed step size is always employed, and a fixed length of past tracking information is input to the tracker to guide position prediction. However, the limited prediction range of a single-scale tracker leads to frequent tracking losses, and limited historical information also reduces tracking accuracy. To address these limitations, we propose a novel Long-Short Match (LSMTrack) tracking method. The key idea is to use long and short trackers and maintain a long-term motion state to improve tracking performance, thus reducing the likelihood of entering the lost status. To this end, a new Mamba-based tracker and a long-short match strategy are proposed. For long and short trackers, the same architecture is used based on Mamba. Unlike the previous Mamba-based approach, the proposed tracker maintains a long-term state while updating the state and making position predictions in each time step, so we call it a step Mamba tracker. Meanwhile, we devise a long-short match strategy at the inference stage to integrate long and short trackers, and design a lost control operation which updates the long-term states using historical state values. In this way, the matching probability and the inference efficiency are guaranteed. Experimental results on two UAV MOT datasets confirm the state-of-the-art performance. Specifically, the best results are achieved in terms of two popular MOTA and IDF1 tracking evaluation metrics. Zi-Zhuang Zou, Mao Ye 0001, Luping Ji, Lihua Zhou, Song Tang 0001, Yan Gan, Shuai Li 0005 |
IEEE Trans. Multim. | 2 |
| 2025 | Motion Prior Knowledge Learning with Homogeneous Language Descriptions for Moving Infrared Small Target DetectionabstractDifferent from traditional object detection, pure vision is not enough to infrared small target detection, due to small target size and weak background contrast. For promoting detection performance, more target representations are needed. Currently, motion representations have been proved to be one of the most potential feature kinds for infrared small target detection. Existing methods have an obvious weakness, that besides vision features, they could only capture coarse motion representations from temporal domain. With vision features, fine motion representations could be more effective to enhance detection performance. To overcome this weakness, inspired by prevalent vision-language models, we propose the first vision-language framework with motion prior knowledge learning (MoPKL). Breaking through traditional pure-vision modality, it utilizes homogeneous language descriptions, formatted for moving targets, to directionally guide vision channel learning motion prior knowledge. With the facilitation of motion-vision alignment and motion-relation mining, the motion of infrared small targets is further refined by graph attention, to generate more fine motion representations. The extensive experiments on datasets ITSDT-15K and IRDST show that our framework is effective. It could often obviously outperform other methods. Shengjia Chen, Luping Ji, Mao Ye 0001 |
AAAI | 5 |
| 2025 | Queryable Prototype Multiple Instance Learning with Vision-Language Models for Incremental Whole Slide Image ClassificationabstractWhole Slide Image (WSI) classification has very significant applications in clinical pathology, e.g., tumor identification and cancer diagnosis. Currently, most research attention is focused on Multiple Instance Learning (MIL) using static datasets. One of the most obvious weaknesses of these methods is that they cannot efficiently preserve and utilize previously learned knowledge. With any new data arriving, classification models are required to be re-trained on both previous and current new data. To overcome this shortcoming and break through traditional vision modality, this paper proposes the first Vision-Language-based framework with Queryable Prototype Multiple Instance Learning (QPMIL-VL) specially designed for incremental WSI classification. This framework mainly consists of two information processing branches: one is for generating bag-level features by prototype-guided aggregation of instance features, while the other is for enhancing class features through a combination of class ensemble, tunable vector and class similarity loss. The experiments on four public WSI datasets demonstrate that our QPMIL-VL framework is effective for incremental WSI classification and often significantly outperforms other compared methods, achieving state-of-the-art (SOTA) performance. Jiaxiang Gou, Luping Ji, Pei Liu 0008, Mao Ye 0001 |
AAAI | 4 |
| 2025 | Self-Prompting Analogical Reasoning for UAV Object DetectionabstractUnmanned Aerial Vehicle Object Detection (UAVOD) presents unique challenges due to varying altitudes, dynamic backgrounds, and the small size of objects. Traditional detection methods often struggle with these challenges, as they typically rely on visual feature only and fail to extract the semantic relations between the objects. To address these limitations, we propose a novel approach named Self-Prompting Analogical Reasoning (SPAR). Our method utilizes the vision-language model (CLIP) to generate context-aware prompts based on image feature, providing rich semantic information that guides analogical reasoning. SPAR includes two main modules: self-prompting and analogical reasoning. Self-prompting module based on learnable description and CLIP-text encoder generates context-aware prompt by combining specific image feature; then an objectness prompt score map is produced by computing the similarity between pixel-level features and context-aware prompt. With this score map, multi-scale image features are enhanced and pixel-level features are chosen for graph construction. While for analogical reasoning module, graph nodes consists of category-level prompt nodes and pixel-level image feature nodes. Analogical inference is based graph convolution. Under the guidance of category-level nodes, different-scale object features have been enhanced, which helps achieve more accurate detection of challenging objects. Extensive experiments illustrate that SPAR outperforms traditional methods, offering a more robust and accurate solution for UAVOD. Nianxin Li, Mao Ye 0001, Lihua Zhou, Song Tang 0001, Yan Gan, Zizhuo Liang, Xiatian Zhu |
AAAI | 2 |
| 2025 | Pseudo Visible Feature Fine-Grained Fusion for Thermal Object DetectionabstractThermal object detection is a critical task in various fields, such as surveillance and autonomous driving. Current state-of-the-art (SOTA) models always leverage a prior Thermal-To-Visible (T2V) translation model to obtain visible spectrum information, followed by a cross-modality aggregation module to fuse information from both modalities. However, this fusion approach does not fully exploit the complementary visible spectrum information beneficial for thermal detection. To address this issue, we propose a novel cross-modal fusion method called Pseudo Visible Feature Fine-Grained Fusion (PFGF). Specifically, a graph is constructed with nodes generated from multi-level thermal features and pseudo-visual latent features produced by the T2V model. Each level of features corresponds to a subgraph. An Inter-Mamba block is proposed to perform cross-modality fusion between nodes at the lowest level; while a Cascade Knowledge Integration (CKI) strategy is used to fuse low-level fused information to high-level subgraphs in a cascade manner. After several iterations of graph node updating, each subgraph outputs an aggregated feature to the detection head respectively. Unlike previous cross-modal fusion methods, our approach explicitly models high-level relationships between cross-modal data, effectively fusing different granularity information. Experimental results demonstrate that our method achieves SOTA detection performance. Code is available at https://github.com/liting1018/PFGF. Mao Ye 0001, Tianwen Wu, Nianxin Li, Shuaifeng Li, Song Tang 0001, Luping Ji |
CVPR | 2 |
| 2025 | Domain Adaptive Diabetic Retinopathy Grading with Model Absence and Flowing DataabstractDomain shift (the difference between source and target domains) poses a significant challenge in clinical applications, e.g., Diabetic Retinopathy (DR) grading. Despite considering certain clinical requirements, like source data privacy, conventional transfer methods are predominantly model-centered and often struggle to prevent model-targeted attacks. In this paper, we address a challenging Online Model-aGnostic Domain Adaptation (OMG-DA) setting, driven by the demands of clinical environments. This setting is characterized by the absence of the model and the flow of target data. To tackle the new challenge, we propose a novel approach, Generative Unadversarial ExampleS (GUES), which enables adaptation from a data-centric perspective. Specifically, we first theoretically reformulate conventional perturbation optimization in a generative way—learning a perturbation generation function with a latent input variable. During model instantiation, we leverage a Variational AutoEncoder to express this function. The encoder with the reparameterization trick predicts the latent input, whilst the decoder is responsible for the generation. Furthermore, the saliency map is selected as pseudo-perturbation labels. Because it not only captures potential lesions but also theoretically provides an upper bound on the function input, enabling the identification of the latent variable. Extensive experiments on DR benchmarks with both frozen pre-trained models and trainable models demonstrate the superiority of GUES, showing robustness even with small batch size. The source code and data are available at https://github.com/tntek/GUES. Wenxin Su, Song Tang 0001, Xiaojing Yi, Mao Ye 0001, Chunxiao Zu, Xiatian Zhu |
CVPR | 5 |
| 2025 | Bayesian Test-Time Adaptation for Vision-Language ModelsabstractTest-time adaptation with pre-trained vision-language models, such as CLIP, aims to adapt the model to new, potentially out-of-distribution test data. Existing methods calculate the similarity between visual embedding and learnable class embeddings, which are initialized by text embeddings, for zero-shot image classification. In this work, we first analyze this process based on Bayes theorem, and observe that the core factors influencing the final prediction are the likelihood and the prior. However, existing methods essentially focus on adapting class embeddings to adapt likelihood, but they often ignore the importance of prior. To address this gap, we propose a novel approach, Bayesian Class Adaptation (BCA), which in addition to continuously updating class embeddings to adapt likelihood, also uses the posterior of incoming samples to continuously update the prior for each class embedding. This dual updating mechanism allows the model to better adapt to distribution shifts and achieve higher prediction accuracy. Our method not only surpasses existing approaches in terms of performance metrics but also maintains superior inference rates and memory usage, making it highly efficient and practical for real-world applications. Lihua Zhou, Mao Ye 0001, Shuaifeng Li, Nianxin Li, Xiatian Zhu, Lei Deng 0001, Hongbin Liu 0001, Zhen Lei 0001 |
CVPR | 2 |
| 2025 | Proxy Denoising for Source-Free Domain AdaptationabstractSource-Free Domain Adaptation (SFDA) aims to adapt a pre-trained source model to an unlabeled target domain with no access to the source data. Inspired by the success of large Vision-Language (ViL) models in many applications, the latest research has validated ViL's benefit for SFDA by using their predictions as pseudo supervision. However, we observe that ViL's supervision could be noisy and inaccurate at an unknown rate, potentially introducing additional negative effects during adaption. To address this thus-far ignored challenge, we introduce a novel Proxy Denoising (__ProDe__) approach. The key idea is to leverage the ViL model as a proxy to facilitate the adaptation process towards the latent domain-invariant space. Concretely, we design a proxy denoising mechanism to correct ViL's predictions. This is grounded on a proxy confidence theory that models the dynamic effect of proxy's divergence against the domain-invariant space during adaptation. To capitalize the corrected proxy, we further derive a mutual knowledge distilling regularization. Extensive experiments show that ProDe significantly outperforms the current state-of-the-art alternatives under both conventional closed-set setting and the more challenging open-set, partial-set, generalized SFDA, multi-target, multi-source, and test-time settings. Our code and data are available at https://github.com/tntek/source-free-domain-adaptation. Song Tang 0001, Wenxin Su, Yan Gan, Mao Ye 0001, Jianwei Zhang 0001, Xiatian Zhu |
ICLR | 4 |
| 2025 | Interpretable Vision-Language Survival Analysis with Ordinal Inductive Bias for Computational PathologyabstractHistopathology Whole-Slide Images (WSIs) provide an important tool to assess cancer prognosis in computational pathology (CPATH). While existing survival analysis (SA) approaches have made exciting progress, they are generally limited to adopting highly-expressive network architectures and only coarse-grained patient-level labels to learn visual prognostic representations from gigapixel WSIs. Such learning paradigm suffers from critical performance bottlenecks, when facing present scarce training data and standard multi-instance learning (MIL) framework in CPATH. To overcome it, this paper, for the first time, proposes a new Vision-Language-based SA (**VLSA**) paradigm. Concretely, (1) VLSA is driven by pathology VL foundation models. It no longer relies on high-capability networks and shows the advantage of *data efficiency*. (2) In vision-end, VLSA encodes textual prognostic prior and then employs it as *auxiliary signals* to guide the aggregating of visual prognostic features at instance level, thereby compensating for the weak supervision in MIL. Moreover, given the characteristics of SA, we propose i) *ordinal survival prompt learning* to transform continuous survival labels into textual prompts; and ii) *ordinal incidence function* as prediction target to make SA compatible with VL-based prediction. Notably, VLSA's predictions can be interpreted intuitively by our Shapley values-based method. The extensive experiments on five datasets confirm the effectiveness of our scheme. Our VLSA could pave a new way for SA in CPATH by offering weakly-supervised MIL an effective means to learn valuable prognostic clues from gigapixel WSIs. Our source code is available at https://github.com/liupei101/VLSA. Pei Liu 0008, Luping Ji, Jiaxiang Gou, Bo Fu 0007, Mao Ye 0001 |
ICLR | 5 |
| 2025 | High Dynamic Range Novel View Synthesis with Single ExposureabstractHigh Dynamic Range Novel View Synthesis (HDR-NVS) aims to establish a 3D scene HDR model from Low Dynamic Range (LDR) imagery. Typically, multiple-exposure LDR images are employed to capture a wider range of brightness levels in a scene, as a single LDR image cannot represent both the brightest and darkest regions simultaneously. While effective, this multiple-exposure HDR-NVS approach has significant limitations, including susceptibility to motion artifacts (e.g., ghosting and blurring), high capture and storage costs. To overcome these challenges, we introduce, for the first time, the single-exposure HDR-NVS problem, where only single exposure LDR images are available during training. We further introduce a novel approach, Mono-HDR-3D, featuring two dedicated modules formulated by the LDR image formation principles, one for converting LDR colors to HDR counterparts, and the other for transforming HDR images to LDR format so that unsupervised learning is enabled in a closed loop. Designed as a meta-algorithm, our approach can be seamlessly integrated with existing NVS models. Extensive experiments show that Mono-HDR-3D significantly outperforms previous methods. Source code is released at https://github.com/prinasi/Mono-HDR-3D. Minxian Li, Mingwu Ren, Mao Ye 0001, Xiatian Zhu |
ICML | 5 |
| 2025 | SAM-Guided Semantic Knowledge Fusion for Visible-Infrared Object DetectionabstractVisible-infrared object detection has gained significant attention because of its applications in autonomous driving, video surveillance, and related fields. The effective fusion of multimodal information is fundamental to its success. The existing approaches concentrate on improving the pixel-level fusion mechanisms; detection performance has reached a plateau. We propose a new framework for SAM-guided semantic knowledge fusion ( SemFusion ). The core idea is to leverage semantic priors from large models while incorporating a lightweight cross-modal fusion strategy. Specifically, our method comprises two stages. In the first stage, the Flow-Guided RGB Feature Alignment (FGRA) module establishes object-aware correspondences between multimodalities based on SAM-generated masks. This ensures semantic-level feature matching by deformable convolution alignment. In the second stage, the Semantic Knowledge Distillation (SKD) strategy facilitates the transfer of large-model knowledge to the detection model through SAM feature, offset, and mask level distillations. For the detector model, three blocks are designed to augment any off-the-shelf detector. They are deformable cross-modal alignment, spatio-channel preliminary fusion, and mask-guided feature refinement. By alignment with SAM masks, semantic alignment and fusion can be achieved, breaking the pixel-level fusion barrier. Extensive experiments demonstrate that our method, as a plugin, exhibits superior performance on the DroneVehicle, VEDAI, and LLVIP datasets. Code is available at https://github.com/liting1018/SemFusion. Shuaifeng Li, Xiaolin Qin, Maoyuan Zhao, Luping Ji, Mao Ye 0001 |
ACM Multimedia | 7 |
| 2025 | Multimodal Causal Reasoning for UAV Object DetectionabstractUnmanned Aerial Vehicle (UAV) object detection faces significant challenges due to complex environmental conditions and different imaging conditions. These factors introduce significant changes in scale and appearance, particularly for small objects that occupy limited pixels and exhibit limited information, complicating detection tasks. To address these challenges, we propose a Multimodel Causal Reasoning framework based on YOLO backbone for UAV Object Detection (MCR-UOD). The key idea is to use the backdoor adjustment to discover the condition-invariant object representation for easy detection. Specifically, the YOLO backbone is first adjusted to incorporate the pre-trained vision-language model. The original category labels are replaced with semantic text prompts, and the detection head is replaced with text-image contrastive learning. Based on this backbone, our method consists of two parts. The first part, named language guided region exploration, discovers the regions with high probability of object existence using text embeddings based on vision-language model such as CLIP. Another part is the backdoor adjustment casual reasoning module, which constructs a confounder dictionary tailored to different imaging conditions to capture global image semantics and derives a prior probability distribution of shooting conditions. During causal inference, we use the confounder dictionary and the prior to intervene on local instance features, disentangling condition variations, and obtaining condition-invariant representations. Experimental results on several public datasets confirm the state-of-the-art performance of our approach. The code, data and models will be released upon publication of this paper. Nianxin Li, Mao Ye 0001, Lihua Zhou, Shuaifeng Li, Song Tang 0001, Luping Ji, Ce Zhu |
NeurIPS | 2 |
| 2025 | ODE-based generative modeling: Learning from a single natural image
Jian Yue, Yan Gan, Lihua Zhou, Shuaifeng Li, Mao Ye 0001 |
Expert Syst. Appl. | 6 |
| 2025 | Multi-level semantic-assisted prototype learning for Few-Shot Action Recognition
Mao Ye 0001, Jianwei Zhang 0001 |
Neurocomputing | 4 |
| 2025 | Boosting few-shot action recognition via time-enhanced multimodal adaptation learning
Ai Peng, Zhouli Shen, Feng Zhang 0052, Mao Ye 0001, Jianwei Zhang 0001 |
Neurocomputing | 6 |
| 2025 | Few-shot medical image segmentation with high-fidelity prototypesabstractFew-shot Semantic Segmentation (FSS) aims to adapt a pretrained model to new classes with as few as a single labeled training sample per class. Despite the prototype based approaches have achieved substantial success, existing models are limited to the imaging scenarios with considerably distinct objects and not highly complex background, e.g., natural images. This makes such models suboptimal for medical imaging with both conditions invalid. To address this problem, we propose a novel D etail S elf-refined P rototype Net work ( DSPNet ) to construct high-fidelity prototypes representing the object foreground and the background more comprehensively. Specifically, to construct global semantics while maintaining the captured detail semantics, we learn the foreground prototypes by modeling the multimodal structures with clustering and then fusing each in a channel-wise manner. Considering that the background often has no apparent semantic relation in the spatial dimensions, we integrate channel-specific structural information under sparse channel-aware regulation. Extensive experiments on three challenging medical image benchmarks show the superiority of DSPNet over previous state-of-the-art methods. The code and data are available at https://github.com/tntek/DSPNet . • A novel prototypical FSS approach DSPNet that enhances prototypes’ self-representation. • A class prototype self-refining method FSPA integrating the cluster prototypes. • A background prototype self-refining method BCMA coding channel-specific structure. Song Tang 0001, Shaxu Yan, Xiaozhi Qi, Mao Ye 0001, Jianwei Zhang 0001, Xiatian Zhu |
Medical Image Anal. | 5 |
| 2025 | SAM-Net: Semantic-assisted multimodal network for action recognition in RGB-D videos
Jinpeng Mi, Mao Ye 0001, Qingdu Li, Jianwei Zhang 0001 |
Pattern Recognit. | 4 |
| 2025 | Adaptive Surveillance Video Compression With Background HyperpriorabstractNeural surveillance video compression methods have demonstrated significant improvements over traditional video compression techniques. In current surveillance video compression frameworks, the first frame in a Group of Pictures (GOP) is usually compressed fully as an I frame, and the subsequent P frames are compressed by referencing this I frame at Low Delay P (LDP) encoding mode. However, this compression approach overlooks the utilization of background information, which limits its adaptability to different scenarios. In this paper, we propose a novel Adaptive Surveillance Video Compression framework based on background hyperprior, dubbed as ASVC. This background hyperprior is related with side information to assist in coding both the temporal and spatial domains. Our method mainly consists of two components. First, the background information from a GOP is extracted, modeled as hyperprior and is compressed by exiting methods. Then these hyperprior is used as side information to compress both I frames and P frames. ASVC effectively captures the temporal dependencies in the latent representations of surveillance videos by leveraging background hyperprior for auxiliary video encoding. The experimental results demonstrate that applying ASVC to traditional and learning based methods significantly improves performance. Song Tang 0001, Mao Ye 0001 |
IEEE Signal Process. Lett. | 3 |
| 2025 | LESEP: Boosting Adversarial Transferability via Latent Encoding and Semantic Embedding PerturbationsabstractTransferability and imperceptibility of adversarial examples are pivotal for assessing the efficacy of black-box attacks. While diffusion models have been employed to generate adversarial examples, leveraging their advanced image generation capability to enhance transferability and imperceptibility, these methods typically focus only on perturbing the image or latent space. They often ignore the critical role of semantic information in the denoising process, thereby impeding the improvement of the transferability of adversarial examples. Furthermore, the modification of high-level semantics inevitably introduces image blurring. This degradation in visual quality makes the adversarial examples more susceptible to detection. To overcome the above limitations, we are the first to utilize image latent encoding and semantic embedding perturbations to enhance the performance of adversarial attacks. Then, the LESEP method is proposed. In the LESEP framework, we first apply image latent encoding attack to achieve deception of the target model. Second, the semantic embedding attack enhances the transferability of adversarial examples. Additionally, we utilize the image restoration technique to guarantee the high imperceptibility of the crafted adversarial examples. Through comprehensive experiments on diverse datasets, different network architectures and defense methods, we have demonstrated that the LESEP method achieves outstanding transferability and imperceptibility while displaying strong robustness. Yan Gan, Chengqian Wu, Deqiang Ouyang, Song Tang 0001, Mao Ye 0001, Tao Xiang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Language-Driven Motion Prior Knowledge Learning for Moving Infrared Small Target DetectionabstractDifferent from traditional object detection, pure vision is often not enough to infrared small target detection (ISTD), due to the small target size and weak background contrast. For promoting performance, more target representations are often needed. Currently, motion representations have been proved to be one of the most potential feature patterns for infrared small targets. Besides vision features, existing methods have an obvious weakness that they could only capture coarse motion representations from the temporal domain. By vision features, fine motion representations could often be more effective to enhance detection performance. To overcome this weakness, and inspired by prevalent vision-language models (VLMs), the first vision-language framework with motion prior knowledge learning (MoPKL) was proposed in our previous work. To further extend it, we repropose an improved version, i.e., iMoPKL. Breaking through traditional pure-vision modality, it utilizes the homogeneous language descriptions, specially formatted for moving targets, to directionally guide vision channels to learn the motion prior knowledge of targets. In detail, it learns the distribution of target motion reconstruction corresponding to the language description as a type of prior knowledge. With the facilitation of language-driven motion alignment, the motion of infrared small targets could be further refined by motion-relation learning, to generate more fine motion representations. The extensive experiments on ITSDT-15K, DAUB-R, and IRDST-H show that our improvement version is effective. It could often obviously outperform the other methods, including our original MoPKL. Our source codes are available athttps://github.com/UESTC-nnLab/MoPKL Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001, Yongsheng Sang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Weakly Supervised Contrastive Learning With Quantity Prompts for Moving Infrared Small Target Detection
Luping Ji, Shengjia Chen, Sicheng Zhu, Jianghong Huang, Mao Ye 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Semi-Supervised Multiview Prototype Learning With Motion Reconstruction for Moving Infrared Small Target DetectionabstractMoving infrared small target detection is critical for various applications, e.g., remote sensing and military. Due to tiny target size and limited labeled data, accurately detecting targets is highly challenging. Currently, existing methods primarily focus on fully-supervised learning, which relies heavily on numerous annotated frames for training. However, annotating a large number of frames for each video is often expensive, time-consuming, and redundant, especially for low-quality infrared images. To break through traditional fully-supervised framework, we propose a new semi-supervised multi-view prototype (S2MVP) learning scheme that incorporates motion reconstruction. In our scheme, we design a bi-temporal motion perceptor based on bidirectional ConvGRU cells to effectively model the motion paradigms of targets by perceiving both forward and backward. Additionally, to explore the potential of unlabeled data, it generates the multi-view feature prototypes of targets as soft labels to guide feature learning by calculating cosine similarity. Imitating human visual system, it retains only the feature prototypes of recent frames. Moreover, it eliminates noisy pseudo-labels to enhance the quality of pseudo-labels through anomaly-driven pseudo-label filtering. Furthermore, we develop a target-aware motion reconstruction loss to provide additional supervision and prevent the loss of target details. To our best knowledge, the proposed S2MVP is the first work to utilize large-scale unlabeled video frames to detect moving infrared small targets. Although 10% labeled training samples are used, the experiments on three public benchmarks (DAUB, ITSDT-15K and IRDST) verify the superiority of our scheme compared to other methods. Source codes are available at https://github.com/UESTC-nnLab/S2MVP. Luping Ji, Jianghong Huang, Shengjia Chen, Sicheng Zhu, Mao Ye 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Approximately Invertible Neural Network for Learned Image CompressionabstractLearned image compression has attracted considerable interests in recent years. An analysis transform and a synthesis transform, which can be regarded as coupled transforms, are used to encode an image to latent feature and decode the feature after quantization to reconstruct the image. Inspired by the success of invertible neural networks in generative modeling, invertible modules can be used to construct the coupled analysis and synthesis transforms. Considering the noise introduced in the feature quantization invalidates the invertible process, this paper proposes an Approximately Invertible Neural Network (A-INN) framework for learned image compression. It formulates the rate-distortion optimization in lossy image compression when using INN with quantization, which differentiates from using INN for generative modelling. Generally speaking, A-INN can be used as the theoretical foundation for any INN based lossy compression method. Based on this formulation, A-INN with a progressive denoising module (PDM) is developed to effectively reduce the quantization noise in the decoding. Moreover, a Cascaded Feature Recovery Module (CFRM) is designed to learn high-dimensional feature recovery from low-dimensional ones to further reduce the noise in feature channel compression. In addition, a Frequency-enhanced Decomposition and Synthesis Module (FDSM) is developed by explicitly enhancing the high-frequency components in an image to address the loss of high-frequency information inherent in neural network based image compression, thereby enhancing the reconstructed image quality. Extensive experiments demonstrate that the proposed A-INN framework achieves better or comparable compression efficiency than the conventional image compression approach and state-of-the-art learned image compression methods. Yanbo Gao, Shuai Li 0005, Chong Lv, Hui Yuan 0001, Mao Ye 0001 |
IEEE Trans. Image Process. | 8 |
| 2025 | MICPL: Motion-Inspired Cross-Pattern Learning for Small-Object Detection in Satellite VideosabstractFor small-object detection, vision patterns can only provide limited support to feature learning. Most prior schemes mainly depend on a single vision pattern to learn object features, seldom considering more latent motion patterns. In the real world, humans often efficiently perceive small objects through multipattern signals. Inspired by this observation, this article attempts to address small-object detection from a new prospective of latent pattern learning. To fulfill this purpose, it regards a real-world moving object as the spatiotemporal sequences of a static object to capture latent motion patterns. In view of this, we propose a motion-inspired cross-pattern learning (MICPL) scheme to capture the motion patterns for moving small-object scenarios. This scheme mainly consists of two crucial parts: motion pattern mining (MPM) and motion-vision adaption. The former is designed to effectively mine the motion pattern from time-dependent representation space. The latter is devised to correlate between motion patterns and vision semantics. In the meanwhile, we explore their cross-pattern interactions to guide MICPL to capture motion patterns effectively. Comparison experiments verify that, cooperated by motion pattern, even a simple detector could often refresh state-of-the-art (SOTA) results on moving small-object detection. Moreover, the experiments on two small-object-related tasks further prove the adaptivity and advantages of our cross-pattern feature learning scheme. Our source codes are available at https://github.com/UESTC-nnLab/MICPL. Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Generative Adversarial Networks with Learnable Auxiliary Module for Image SynthesisabstractTraining generative adversarial networks (GANs) for noise-to-image synthesis is a challenge task, primarily due to the instability of GANs’ training process. One of the key issues is the generator’s sensitivity to input data, which can cause sudden fluctuations in the generator’s loss value with certain inputs. This sensitivity suggests an inadequate ability to resist disturbances in the generator, causing the discriminator’s loss value to oscillate and negatively impacting the discriminator. Then, the negative feedback of discriminator is also not conducive to updating generator’s parameters, leading to suboptimal image generation quality. In response to this challenge, we present an innovative GANs model equipped with a learnable auxiliary module that processes auxiliary noise. The core objective of this module is to enhance the stability of both the generator and discriminator throughout the training process. To achieve this target, we incorporate a learnable auxiliary penalty and an augmented discriminator, designed to control the generator and reinforce the discriminator’s stability, respectively. We further apply our method to the Hinge and LSGANs loss functions, illustrating its efficacy in reducing the instability of both the generator and the discriminator. The tests we conducted on LSUN, CelebA, Market-1501, and Creative Senz3D datasets serve as proof of our method’s ability to improve the training stability and overall performance of the baseline methods. Yan Gan, Chenxue Yang, Mao Ye 0001, Renjie Huang, Deqiang Ouyang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Source-Free Domain Adaptation with Frozen Multimodal Foundation ModelabstractSource-Free Domain Adaptation (SFDA) aims to adapt a source model for a target domain, with only access to unlabeled target training data and the source model pretrained on a supervised source domain. Relying on pseudo labeling and/or auxiliary supervision, conventional methods are inevitably error-prone. To mitigate this limitation, in this work we for the first time explore the potentials of off-the-shelf vision-language (ViL) multimodal models (e.g., CLIP) with rich whilst heterogeneous knowledge. We find that directly applying the ViL model to the target domain in a zero-shot fashion is unsatisfactory, as it is not specialized for this particular task but largely generic. To make it task specific, we propose a novel Distilling multImodal Foundation mOdel (DIFO) approach. Specifically, DIFO alternates between two steps during adaptation: (i) Customizing the ViL model by maximizing the mutual information with the target model in a prompt learning manner, (ii) Distilling the knowledge of this customized ViL model to the target model. For more fine-grained and reliable distillation, we further introduce two effective regularization terms, namely most-likely category encouragement and predictive consistency. Extensive experiments show that DIFO significantly outperforms the state-of-the-art alternatives. Code is here. Song Tang 0001, Wenxin Su, Mao Ye 0001, Xiatian Zhu |
CVPR | 3 |
| 2024 | Adversarial Experts Model for Black-box Domain AdaptationabstractBlack-box domain adaptation treats the source domain model as a black box. During the transfer process, the only available information about the target domain is the noisy labels output by the black-box model. This poses significant challenges for domain adaptation. Conventional approaches typically tackle the black-box noisy label problem from two aspects: self-knowledge distillation and pseudo-label denoising, both achieving limited performance due to limited knowledge information. To mitigate this issue, we explore the potential of off-the-shelf vision-language (ViL) multimodal models with rich semantic information for black-box domain adaptation by introducing an Adversarial Experts Model (AEM). Specifically, our target domain model is designed as one feature extractor and two classifiers, trained over two stages: In the knowledge transferring stage, with a shared feature extractor, the black-box source model and the ViL model act as two distinct experts for joint knowledge contribution, guiding the learning of one classifier each. While contributing their respective knowledge, the experts are also updated due to their own limitation and bias. In the adversarial alignment stage, to further distill expert knowledge to the target domain model, adversarial learning is conducted between the feature extractor and the two classifiers. A new consistency-max loss function is proposed to measure two classifier consistency and further improve classifier prediction certainty. Extensive experiments on multiple datasets demonstrate the effectiveness of our approach. Code is available at https://github.com/singinger/AEM. Siying Xiao, Mao Ye 0001, Qichen He, Shuaifeng Li, Song Tang 0001, Xiatian Zhu |
ACM Multimedia | 2 |
| 2024 | Cloud Object Detector Adaptation by Integrating Different Source KnowledgeabstractWe propose to explore an interesting and promising problem, Cloud Object Detector Adaptation (CODA), where the target domain leverages detections provided by a large cloud model to build a target detector. Despite with powerful generalization capability, the cloud model still cannot achieve error-free detection in a specific target domain. In this work, we present a novel Cloud Object detector adaptation method by Integrating different source kNowledge (COIN). The key idea is to incorporate a public vision-language model (CLIP) to distill positive knowledge while refining negative knowledge for adaptation by self-promotion gradient direction alignment. To that end, knowledge dissemination, separation, and distillation are carried out successively. Knowledge dissemination combines knowledge from cloud detector and CLIP model to initialize a target detector and a CLIP detector in target domain. By matching CLIP detector with the cloud detector, knowledge separation categorizes detections into three parts: consistent, inconsistent and private detections such that divide-and-conquer strategy can be used for knowledge distillation. Consistent and private detections are directly used to train target detector; while inconsistent detections are fused based on a consistent knowledge generation network, which is trained by aligning the gradient direction of inconsistent detections to that of consistent detections, because it provides a direction toward an optimal target detector. Experiment results demonstrate that the proposed COIN method achieves the state-of-the-art performance. Shuaifeng Li, Mao Ye 0001, Lihua Zhou, Nianxin Li, Siying Xiao, Song Tang 0001, Xiatian Zhu |
NeurIPS | 2 |
| 2024 | Shooting condition insensitive unmanned aerial vehicle object detection
Jinzong Cui, Mao Ye 0001, Xiatian Zhu, Song Tang 0001 |
Expert Syst. Appl. | 3 |
| 2024 | Disentanglement then reconstruction: Unsupervised domain adaptation by twice distribution alignments
Lihua Zhou, Mao Ye 0001, Xinpeng Li 0005, Ce Zhu, Yiguang Liu, Xue Li 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Source-Free Domain Adaptation via Target Prediction Distribution SearchingabstractAbstract Existing Source-Free Domain Adaptation (SFDA) methods typically adopt the feature distribution alignment paradigm via mining auxiliary information (eg., pseudo-labelling, source domain data generation). However, they are largely limited due to that the auxiliary information is usually error-prone whilst lacking effective error-mitigation mechanisms. To overcome this fundamental limitation, in this paper we propose a novel Target Prediction Distribution Searching (TPDS) paradigm. Theoretically, we prove that in case of sufficient small distribution shift, the domain transfer error could be well bounded. To satisfy this condition, we introduce a flow of proxy distributions that facilitates the bridging of typically large distribution shift from the source domain to the target domain. This results in a progressive searching on the geodesic path where adjacent proxy distributions are regularized to have small shift so that the overall errors can be minimized. To account for the sequential correlation between proxy distributions, we develop a new pairwise alignment with category consistency algorithm for minimizing the adaptation errors. Specifically, a manifold geometry guided cross-distribution neighbour search is designed to detect the data pairs supporting the Wasserstein distance based shift measurement. Mutual information maximization is then adopted over these pairs for shift regularization. Extensive experiments on five challenging SFDA benchmarks show that our TPDS achieves new state-of-the-art performance. The code and datasets are available at https://github.com/tntek/TPDS . Song Tang 0001, An Chang, Fabian Zhang, Xiatian Zhu, Mao Ye 0001, Changshui Zhang |
Int. J. Comput. Vis. | 5 |
| 2024 | Temporal cues enhanced multimodal learning for action recognition in RGB-D videos
Zhiyuan Ma 0001, Jinpeng Mi, Yan Gan, Mao Ye 0001, Jianwei Zhang 0001 |
Neurocomputing | 7 |
| 2024 | SPGAN: Siamese projection Generative Adversarial Networks
Yan Gan, Tao Xiang 0001, Deqiang Ouyang, Mingliang Zhou 0001, Mao Ye 0001 |
Knowl. Based Syst. | 5 |
| 2024 | Compressed-SDR to HDR Video ReconstructionabstractThe new generation of organic light emitting diode display is designed to enable the high dynamic range (HDR), going beyond the standard dynamic range (SDR) supported by the traditional display devices. However, a large quantity of videos are still of SDR format. Further, most pre-existing videos are compressed at varying degrees for minimizing the storage and traffic flow demands. To enable movie-going experience on new generation devices, converting the compressed SDR videos to the HDR format (i.e., compressed-SDR to HDR conversion) is in great demands. The key challenge with this new problem is how to solve the intrinsic many-to-many mapping issue. However, without constraining the solution space or simply imitating the inverse camera imaging pipeline in stages, existing SDR-to-HDR methods can not formulate the HDR video generation process explicitly. Besides, they ignore the fact that videos are often compressed. To address these challenges, in this work we propose a novel imaging knowledge-inspired parallel networks (termed as KPNet) for compressed-SDR to HDR (CSDR-to-HDR) video reconstruction. KPNet has two key designs: Knowledge-Inspired Block (KIB) and Information Fusion Module (IFM). Concretely, mathematically formulated using some priors with compressed videos, our conversion from a CSDR-to-HDR video reconstruction is conceptually divided into four synergistic parts: reducing compression artifacts, recovering missing details, adjusting imaging parameters, and reducing image noise. We approximate this process by a compact KIB. To capture richer details, we learn HDR representations with a set of KIBs connected in parallel and fused with the IFM. Extensive evaluations show that our KPNet achieves superior performance over the state-of-the-art methods. Mao Ye 0001, Xiatian Zhu, Shuai Li 0005, Xue Li 0001, Ce Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Source-free domain adaptation with Class Prototype Discovery
Lihua Zhou, Nianxin Li, Mao Ye 0001, Xiatian Zhu, Song Tang 0001 |
Pattern Recognit. | 3 |
| 2024 | Toward Dense Moving Infrared Small Target Detection: New Datasets and BaselineabstractAs an important research branch of infrared small target detection, dense target detection (e.g., drone swarm detection) has always been a topic worth exploring. Currently, existing datasets cover only one or several (sparse) targets, with almost no dataset available for the research on dense small target detection. To advance this kind of search, for the first time, we synthesize two special dense moving target datasets (DMIST-60 and DMIST-100) on DAUB data. They both contain far more than 50 infrared small targets per frame. In the meantime, for evaluating our new datasets and flourishing detection methodology research, we propose a linking-aware sliced network (LASNet) as the baseline of our datasets. It mainly consists of visual feature extraction, motion feature extraction and motion-affinity fusion. The comprehensive experiments on our synthesized datasets confirm: i) both datasets are practical and effective for dense moving infrared small target detection and ii) proposed LASNet could always obviously outperform other compared methods in both sparse and dense target scenarios. Our new datasets and source codes are currently available athttps://github.com/UESTC-nnLab/DMIST. Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001, Haohao Ren, Yongsheng Sang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | SSTNet: Sliced Spatio-Temporal Network With Cross-Slice ConvLSTM for Moving Infrared Dim-Small Target DetectionabstractInfrared dim-small target detection, as an important branch of object detection, has been attracting research attention in recent decades. Its challenges mainly lie in the small target sizes and dim contrast to background images. Recent research schemes on it mainly focus on improving the feature representation of spatio-temporal domains only in single-slice temporal scope. More cross-slice motion, i.e., past and future, is seldom considered to enhance target features. To use cross-slice motion context, this article proposes a sliced spatio-temporal network (SSTNet) with cross-slice enhancement for moving infrared dim-small target detection. In our scheme, a new cross-slice ConvLSTM node is designed to capture spatio-temporal motion features from both inner slice and inter-slices. Moreover, to improve infrared small target motion feature learning, we extend conventional loss function by adopting a new motion-coordination loss (MCL) term. On these, we propose a motion-coupling neck to assist feature extractor in facilitating the capturing and utilization of motion features from multiframes. To our best knowledge, our work is the first one to explore the cross-slice spatio-temporal motion modeling for infrared dim-small targets. Experiments verify that our SSTNet could refresh most state-of-the-art metrics on two public benchmarks (DAUB and IRDST). Our source codes are available athttps://github.com/UESTC-nnLab/SSTNet. Shengjia Chen, Luping Ji, Jiewen Zhu, Mao Ye 0001, Xiaoyong Yao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Triple-Domain Feature Learning With Frequency-Aware Memory Enhancement for Moving Infrared Small Target DetectionabstractAs a subfield of object detection, moving infrared small target detection (ISTD) presents significant challenges due to tiny target sizes and low contrast against backgrounds. Currently existing methods primarily rely on the features extracted only from spatiotemporal domain. Frequency domain has hardly been concerned yet, although it has been widely applied in image processing. To extend feature source domains and enhance feature representation, we propose a new triple-domain strategy (Tridos) with the frequency-aware memory enhancement on spatiotemporal domain for ISTD. In this scheme, it effectively detaches and enhances frequency features by a local-global frequency-aware module (LGFM) with Fourier transform (FT). Inspired by human visual system (HVS), our memory enhancement is designed to capture the spatial relationships of infrared targets among video frames. Furthermore, it encodes temporal dynamics motion features via differential learning and residual enhancing. In addition, we further design a residual compensation to reconcile possible cross-domain feature mismatches. To our best knowledge, proposed Tridos is the first work to explore infrared target feature learning comprehensively in spatiotemporal-frequency domains. The extensive experiments on three datasets (i.e., DAUB, ITSDT-15K, and IRDST) validate that our triple-domain infrared feature learning scheme could often be obviously superior to state-of-the-art (SOTA) ones. Source codes are available athttps://github.com/UESTC-nnLab/Tridos. Luping Ji, Shengjia Chen, Sicheng Zhu, Mao Ye 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Illumination Distribution-Aware Thermal Pedestrian DetectionabstractPedestrian detection is an important task in computer vision, which is also an important part of intelligent transportation systems. For privacy protection, thermal images are widely used in pedestrian detection problems. However, thermal pedestrian detection is challenging due to the significant effect of temperature variation on the illumination of images and that fine-grained illumination annotations are difficult to be acquired. The existing methods have attempted to exploit coarse-grained day/night labels, which however even hampers the model performance. In this work, we introduce a novel idea of regressing conditional thermal-visible feature distribution, dubbed as Illumination Distribution-Aware adaptation (IDA). The key idea is to predict the conditional visible feature distribution given a thermal image, subject to their pre-computed joint distribution. Specifically, we first estimate the thermal-visible feature joint distribution by constructing feature co-occurrence matrices, offering a conditional probability distribution for any given thermal image. With this pairing information, we then form a conditional probability distribution regression task for model optimization. Critically, as a model agnostic strategy, this allows the visible feature knowledge to be transferred to the thermal counterpart implicitly for learning more discriminating feature representation. Experiment results show that our method outperforms the prior art methods, which use extra illumination annotations. Besides, as a plug-in, our method can averagely reduce about 2% MR on KAIST dataset, and improve about 1% mAP on FLIR-aligned and Autonomous Vehicles datasets without extra calculation for test. Code is available athttps://github.com/HaMeow-lst1/IDA. Mao Ye 0001, Luping Ji, Song Tang 0001, Yan Gan, Xiatian Zhu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Progressive Source-Aware Transformer for Generalized Source-Free Domain AdaptationabstractSource-free domain adaptation (SFDA) tends to forget the source domain, suffering from limitations in real-world scenarios. Recently, generalized source-free domain adaptation (GSFDA) problem naturally emerges, aiming for good performance on both target and source domains. The existing methods attempt to retain model parameters associated with the source domain to prevent such forgetting. However, this strategy is not conducive to improving cross-domain performance on the target domain, prioritizing mitigating forgetting on the source domain. This article introduces a Progressive Source-Aware Transformer approach for GSFDA, dubbed PSAT-GDA. Our core idea is to enforce the domain adaptation process to remember the source domain by imposing source guidance, offering a target domain-centric anti-forgetting mechanism. Specifically, for each epoch, a Transformer-based deep network is adapted to do domain alignment like the traditional SFDA method, because the transformer working on the image patch sequence helps to reduce image noise caused by domain shift. Meanwhile, another Transformer is designed to generate source guidance supervising domain alignment. By augmenting target sample and mining the source information from the historical models before current epoch, source injected feature group is constructed. Based on the Transformer mechanism, the attention block can select useful source information for each target sample. From it, we devise neighbour-based and augmentation-based regularizations to shape the source guidance. Experiments on three challenging datasets show that our method can achieve evident cross-domain improvement on the target domains. Also, it can mitigate forgetting on all domains after adapting to single or multiple target domains. Song Tang 0001, Yuji Shi, Mao Ye 0001, Changshui Zhang, Jianwei Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Homeomorphism Alignment for Unsupervised Domain AdaptationabstractExisting unsupervised domain adaptation (UDA) methods rely on aligning the features from the source and target domains explicitly or implicitly in a common space (i.e., the domain invariant space). Explicit distribution matching ignores the discriminability of learned features, while the implicit counterpart such as self-supervised learning suffers from pseudo-label noises. With distribution alignment, it is challenging to acquire a common space which maintains fully the discriminative structure of both domains. In this work, we propose a novel HomeomorphisM Alignment (HMA) approach characterized by aligning the source and target data in two separate spaces. Specifically, an invertible neural network based homeomorphism is constructed. Distribution matching is then used as a sewing up tool for connecting this homeomorphism mapping between the source and target feature spaces. Theoretically, we show that this mapping can preserve the data topological structure (e.g., the cluster/group structure). This property allows for more discriminative model adaptation by leveraging both the original and transformed features of source data in a supervised manner, and those of target domain in an unsupervised manner (e.g., prediction consistency). Extensive experiments demonstrate that our method can achieve the state-of-the-art results. Code is released at https://github.com/buerzlh/HMA. Lihua Zhou, Mao Ye 0001, Xiatian Zhu, Siying Xiao, Xuqian Fan, Ferrante Neri |
ICCV | 2 |
| 2023 | Bit Allocation using OptimizationabstractIn this paper, we consider the problem of bit allocation in Neural Video Compression (NVC). First, we reveal a fundamental relationship between bit allocation in NVC and Semi-Amortized Variational Inference (SAVI). Specifically, we show that SAVI with GoP (Group-of-Picture)-level likelihood is equivalent to pixel-level bit allocation with precise rate & quality dependency model. Based on this equivalence, we establish a new paradigm of bit allocation using SAVI. Different from previous bit allocation methods, our approach requires no empirical model and is thus optimal. Moreover, as the original SAVI using gradient ascent only applies to single-level latent, we extend the SAVI to multi-level such as NVC by recursively applying back-propagating through gradient ascent. Finally, we propose a tractable approximation for practical implementation. Our method can be applied to scenarios where performance outweights encoding speed, and serves as an empirical bound on the R-D performance of bit allocation. Experimental results show that current state-of-the-art bit allocation algorithms still have a room of $\approx 0.5$ dB PSNR to improve compared with ours. Code is available at https://github.com/tongdaxu/Bit-Allocation-Using-Optimization. Tongda Xu, Han Gao 0012, Chenjian Gao, Dailan He, Jinyong Pi, Jixiang Luo, Mao Ye 0001, Hongwei Qin, Yan Wang 0080, Ya-Qin Zhang |
ICML | 9 |
| 2023 | Unsupervised Feature Selection Using Both Similar and Dissimilar Structures
Tao Xiang 0006, Mao Ye 0001 |
ICONIP (7) | 5 |
| 2023 | Independent Feature Decomposition and Instance Alignment for Unsupervised Domain AdaptationabstractExisting Unsupervised Domain Adaptation (UDA) methods typically attempt to perform knowledge transfer in a domain-invariant space explicitly or implicitly. In practice, however, the obtained features is often mixed with domain-specific information which causes performance degradation. To overcome this fundamental limitation, this article presents a novel independent feature decomposition and instance alignment method (IndUDA in short). Specifically, based on an invertible flow, we project the base features into a decomposed latent space with domain-invariant and domain-specific dimensions. To drive semantic decomposition independently, we then swap the domain-invariant part across source and target domain samples with the same category and require their inverted features are consistent in class-level with the original features. By treating domain-specific information as noise, we replace it by Gaussian noise and further regularize source model training by instance alignment, i.e., requiring the base features close to the corresponding reconstructed features, respectively. Extensive experiment results demonstrate that our method achieves state-of-the-art performance on popular UDA benchmarks. The appendix and code are available at https://github.com/ayombeach/IndUDA. Qichen He, Siying Xiao, Mao Ye 0001, Xiatian Zhu, Ferrante Neri, Dongde Hou |
IJCAI | 3 |
| 2023 | Single-image HDR reconstruction by dual learning the camera imaging process
Lei She, Mao Ye 0001, Shuai Li 0005, Ce Zhu |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Cross-domain video action recognition via adaptive gradual learning
Zhenwei Bao, Jinpeng Mi, Yan Gan, Mao Ye 0001, Jianwei Zhang 0001 |
Neurocomputing | 5 |
| 2023 | Generative adversarial networks with adaptive learning strategy for noise-to-image synthesis
Yan Gan, Tao Xiang 0001, Hangcheng Liu, Mao Ye 0001, Mingliang Zhou 0001 |
Neural Comput. Appl. | 4 |
| 2023 | Pre-encoding based temporal dependent rate-distortion optimization for HEVC
Hongwei Guo 0001, Ce Zhu, Mao Ye 0001, Lei Luo 0003, Xu Yang 0030 |
Signal Process. Image Commun. | 3 |
| 2023 | Multi-Frame Compressed Video Quality Enhancement by Spatio-Temporal Information BalanceabstractIn recent years, the performance of multi-frame quality enhancement algorithms for compressed videos has been greatly improved compared with single-frame based algorithms. However, the existing methods mainly focus on mining the temporal information of multiple frames. The large number of reference frames reduces the exploration of spatial information, although the existing single-frame based algorithms for enhancement, denoising, and super-resolution demonstrate the significance of the spatial information. To address this problem, we propose a plug-and-play module called Spatio-temporal Information Balance (STIB) to adaptively balance the spatial and temporal information. In our method, we use a feature extractor to exploit richer spatial information, and use a refinement module to refine the aligned temporal information, to be more conducive to the fusion of spatio-temporal information. Finally, we use the deformable convolution based re-alignment module to do alignment and fusion in feature space for balancing the spatio-temporal information. Experiments show that our module can significantly improve the performance of the existing multi-frame based enhancement algorithms. Zeyang Wang, Mao Ye 0001, Shuai Li 0005, Xue Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2023 | Adaptive Mutual Learning for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation aims to transfer knowledge from labeled source domain to unlabeled target domain. The semi-supervised method based on mean-teacher framework is one of the main stream approaches. By enforcing consistency constraints, it is hopeful that the teacher network will distill useful source domain knowledge to the student network. However, in practice negative transfer often emerges because the performance of the teacher network is not guaranteed to be always better than the student network. To address this limitation, a novel Adaptive Mutual Learning (AML) strategy is proposed in this paper. Specifically, given a target sample, the network with worse prediction will be optimized by pushing its prediction close to the better prediction. This is in the spirit of traditional knowledge distillation. On the other hand, the network with better prediction is further refined by requiring its prediction to stay away from the worse prediction. This can be regarded conceptually as reverse knowledge distillation. In this way, two networks learn from each other according to their respective performance. At inference phase, the averaged output of these two networks can be taken as the final prediction. Experimental results demonstrate that our AML achieves competitive results. Lihua Zhou, Siying Xiao, Mao Ye 0001, Xiatian Zhu, Shuaifeng Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Spatio-Temporal Detail Information Retrieval for Compressed Video Quality EnhancementabstractThe past few years have witnessed the great success of multi-frame quality enhancement for compressed video. Although the existing methods based on deformable alignment have achieved the state-of-the-art performance, they do not pay enough attention to the recovery of detail information. In this work, we propose a Spatio-Temporal Detail Retrieval (STDR) method to promote the recovery of detail information. To alleviate the problem of inaccurate deformable offsets caused by the fixed receptive field, motivated by multi-task learning, we design a plug-and-play Multi-path Deformable Alignment (MDA) module to generate more accurate offsets by integrating the alignment features of different receptive fields, so that the temporal detail information can be better recovered. For the spatial detail information restoration, several residual dense blocks with channel attention layer are utilized in the reconstruction module to explore valuable high-frequency spatial information from the fused multi-path alignment features. Meanwhile, a complementary loss function based on the Pearson correlation coefficient is developed to ameliorate the over-smoothing shortcoming caused by pixel-wise mean square or absolute value loss. Experimental results demonstrate that the proposed STDR network achieves superior performance compared with the state-of-the-art methods in both quantitative and qualitative evaluations. Dengyan Luo, Mao Ye 0001, Shuai Li 0005, Ce Zhu, Xue Li 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Source-Free Object Detection by Learning to Overlook Domain StyleabstractSource-free object detection (SFOD) needs to adapt a detector pre-trained on a labeled source domain to a tar-get domain, with only unlabeled training data from the tar-get domain. Existing SFOD methods typically adopt the pseudo labeling paradigm with model adaption alternating between predicting pseudo labels and fine-tuning the model. This approach suffers from both unsatisfactory accuracy of pseudo labels due to the presence of domain shift and lim-ited use of target domain training data. In this work, we present a novel Learning to Overlook Domain Style (LODS) method with such limitations solved in a principled man-ner. Our idea is to reduce the domain shift effect by en-forcing the model to overlook the target domain style, such that model adaptation is simplified and becomes easier to carry on. To that end, we enhance the style of each tar-get domain image and leverage the style degree difference between the original image and the enhanced image as a self-supervised signal for model adaptation. By treating the enhanced image as an auxiliary view, we exploit a student- teacher architecture for learning to overlook the style de-gree difference against the original image, also character-ized with a novel style enhancement algorithm and graph alignment constraint. Extensive experiments demonstrate that our LODS yields new state-of-the-art performance on four benchmarks. Shuaifeng Li, Mao Ye 0001, Xiatian Zhu, Lihua Zhou |
CVPR | 2 |
| 2022 | A Multiscale Gradient-Backpropagation Optimization Framework for Deformable Convolution Based Compressed Video EnhancementabstractDeep learning based compressed video quality enhancement has raised lots of interest recently. To explore the information over multiple frames, deformable convolution has been used for temporal alignment. However, in the existing methods, the deformable convolution is used in a relatively naïve way, without differing the characteristics of offset and features, and their behavior in gradient backpropagation. In this paper, a multiscale gradient-backpropagation optimization framework is proposed for the deformable convolution based compressed video quality enhancement. By analyzing the gradient backpropagation mechanism of deformable convolution, a multi-scale deformable convolution alignment structure is developed to facilitate the gradient backpropagation at all scales. Moreover, a progressive offset prediction module is developed, which decouples the offset prediction from the feature up-sampling, thus reducing the noise flow over scales. Experimental results show that the proposed method achieves the state-of-the-art performance, with 25.6% BD-rate saving compared to the HEVC reference software (HM). Yanbo Gao, Menghu Jia, Shuai Li 0005, Mao Ye 0001, Frédéric Dufaux |
ICASSP | 5 |
| 2022 | Content Adaptive Compressed Screen Content Video Quality EnhancementabstractIn recent years, with the rise of various online learning plat-forms and game live broadcasting industry, screen content video is explosively increasing. There is an urgent demand to reduce the inevitable compression artifacts produced by traditional lossy compression method. However, there do not exist any research on compressed Screen Content Video (SCV) quality enhancement. Since SCV frames always con-sist of two main types of contents with different characteris-tics, i.e., text and graphic, a Content Adaptive model based on Two branches (CAT) is proposed in this paper. For en-hancing graphic, we utilize temporal information by motion compensation, while for enhancing text, we explore spatial relevance in horizontal and vertical directions to recover sharp edges. Furthermore, a content adaptive block is used to select content related features for the collaborative enhancement of both contents. We build a large SCV dataset compressed by H.266/VVC Test Model (VTM 12.1). On this dataset, experi-mental results demonstrate that the proposed method achieves the state-of-the-art performance on 10 different kinds of SCV test sequences. Mao Ye 0001, Yanbo Gao, Shuai Li 0005, Xue Li 0001 |
ICME | 2 |
| 2022 | KUNet: Imaging Knowledge-Inspired Single HDR Image ReconstructionabstractRecently, with the rise of high dynamic range (HDR) display devices, there is a great demand to transfer traditional low dynamic range (LDR) images into HDR versions. The key to success is how to solve the many-to-many mapping problem. However, the existing approaches either do not consider constraining solution space or just simply imitate the inverse camera imaging pipeline in stages, without directly formulating the HDR image generation process. In this work, we address this problem by integrating LDR-to-HDR imaging knowledge into an UNet architecture, dubbed as Knowledge-inspired UNet (KUNet). The conversion from LDR-to-HDR image is mathematically formulated, and can be conceptually divided into recovering missing details, adjusting imaging parameters and reducing imaging noise. Accordingly, we develop a basic knowledge-inspired block (KIB) including three subnetworks corresponding to the three procedures in this HDR imaging process. The KIB blocks are cascaded in the similar way to the UNet to construct HDR image with rich global information. In addition, we also propose a knowledge inspired jump-connect structure to fit a dynamic range gap between HDR and LDR images. Experimental results demonstrate that the proposed KUNet achieves superior performance compared with the state-of-the-art methods. The code, dataset and appendix materials are available at https://github.com/wanghu178/KUNet.git. Mao Ye 0001, Xiatian Zhu, Shuai Li 0005, Ce Zhu, Xue Li 0001 |
IJCAI | 2 |
| 2022 | Recurrent Deformable Fusion for Compressed Video Artifact ReductionabstractThe compressed video inevitably appears in compression artifacts, which seriously affect the Quality of Experience. The state-of-the-art methods employ deformable alignment to gather similar information from multiple neighborhood frames to enhance target frame quality. However, they always align multiple frames to the target frame simultaneously, which brings repetitive and useless information because of multiple and imperfect alignments. In this paper, we propose a recurrent deformable fusion method which considers the alignment quality distortion caused by time distance from the target frame. Specifically, a Deformable Alignment (DA) module aligns each pair of the target frame and an adjacent frame following the time line. At the same time, a Recurrent Fusion (RF) module integrates the current aligned feature with the previous fused feature. After that, the fused features are concatenated along the time line. Then, a Multi-Scale Attention Reconstruction (MSAR) module is proposed to gather useful information from the fused features. Compared with the previous multi-frame alignment approach, our method can avoid obtaining a lot of repetitive and useless information. Experiment results confirm that our method achieves state-of-the-art performance on the standard test sequences. Liuhan Peng, Askar Hamdulla, Mao Ye 0001, Shuai Li 0005, Hongwei Guo 0001 |
ISCAS | 3 |
| 2022 | Structure-Preserving Motion Estimation for Learned Video CompressionabstractFollowing the conventional hybrid video coding framework, existing learned video compression methods rely on the decoded previous frame as the reference for motion estimation considering that it is available to the decoder. Diving into its essential advantage of strong representation capability with CNNs, however, we find this strategy is suboptimal due to two reasons: (1) Motion estimation based on the decoded (often distorted) frame would damage both the spatial structure of motion information inferred and the corresponding residual for each frame, making it difficult to be spatially encoded on the whole image basis using CNNs; (2) Typically, it would break the consistent nature across frames since the estimated motion information is no longer consistent with the movement in the original video due to the distortion in the decoded video, lowering the overall temporal coding efficiency. To overcome these problems, a novel asymmetric Structure-Preserving Motion Estimation (SPME) method is proposed, with the aim to fully explore the ignored original previous frame at the encoder side while complying with the decoded previous frame at the decoder side. Concretely, SPME estimates superior spatially structure-preserving and temporally consistent motion field by aggregating the motion prediction of both the original and the decoded reference frames w.r.t the current frame. Critically, our method can be universally applied to the existing feature prediction based video compression methods. Extensive experiments on several standard test datasets show that our SPME can significantly enhance the state-of-the-art methods. Han Gao 0012, Jinzhong Cui, Mao Ye 0001, Shuai Li 0005, Xiatian Zhu |
ACM Multimedia | 3 |
| 2022 | Geometric Warping Error Aware CNN for DIBR Oriented View SynthesisabstractDepth Image based Rendering (DIBR) oriented view synthesis is an important virtual view generation technique. It warps the reference view images to the target viewpoint based on their depth maps, without requiring many available viewpoints. However, in the 3D warping process, pixels are warped to fractional pixel locations and then rounded (or interpolated) to integer pixels, resulting in geometric warping error and reducing the image quality. This resembles, to some extent, the image super-resolution problem, but with unfixed fractional pixel locations. To address this problem, we propose a geometric warping error aware CNN (GWEA) framework to enhance the DIBR oriented view synthesis. First, a deformable convolution based geometric warping error aware alignment (GWEA-DCA) module is developed, by taking advantage of the geometric warping error preserved in the DIBR module. The offset learned in the deformable convolution can account for the geometric warping error to facilitate the mapping from the fractional pixels to integer pixels. Moreover, in view that the pixels in the warped images are of different qualities due to the different strengths of warping errors, an attention enhanced view blending (GWEA-AttVB) module is further developed to adaptively fuse the pixels from different warped images. Finally, a partial convolution based hole filling and refinement module fills the remaining holes and improves the quality of the overall image. Experiments show that our model can synthesize higher-quality images than the existing methods, and ablation study is also conducted, validating the effectiveness of each proposed module. Shuai Li 0005, Yanbo Gao, Mao Ye 0001 |
ACM Multimedia | 5 |
| 2022 | Class Discriminative Adversarial Learning for Unsupervised Domain AdaptationabstractAs a state-of-the-art family of Unsupervised Domain Adaptation (UDA), bi-classifier adversarial learning methods are formulated in an adversarial (minimax) learning framework with a single feature extractor and two classifiers. Model training alternates between two steps: (I) constraining the learning of the two classifiers to maximize the prediction discrepancy of unlabeled target domain data, and (II) constraining the learning of the feature extractor to minimize this discrepancy. Despite being an elegant formulation, this approach has a fundamental limitation: Maximizing and minimizing the classifier discrepancy is not class discriminative for the target domain, finally leading to a suboptimal adapted model. To solve this problem, we propose a novel Class Discriminative Adversarial Learning (CDAL) method characterized by discovering class discrimination knowledge and leveraging this knowledge to discriminatively regulate the classifier discrepancy constraints on-the-fly. This is realized by introducing an evaluation criterion for judging each classifier's capability and each target domain sample's feature reorientation via objective loss reformulation. Extensive experiments on three standard benchmarks show that our CDAL method yields new state-of-the-art performance. Our code is made available at https://github.com/buerzlh/CDAL. Lihua Zhou, Mao Ye 0001, Xiatian Zhu, Shuaifeng Li, Yiguang Liu |
ACM Multimedia | 2 |
| 2022 | MetaTeacher: Coordinating Multi-Model Domain Adaptation for Medical Image ClassificationabstractIn medical image analysis, we often need to build an image recognition system for a target scenario with the access to small labeled data and abundant unlabeled data, as well as multiple related models pretrained on different source scenarios. This presents the combined challenges of multi-source-free domain adaptation and semi-supervised learning simultaneously. However, both problems are typically studied independently in the literature, and how to effectively combine existing methods is non-trivial in design. In this work, we introduce a novel MetaTeacher framework with three key components: (1) A learnable coordinating scheme for adaptive domain adaptation of individual source models, (2) A mutual feedback mechanism between the target model and source models for more coherent learning, and (3) A semi-supervised bilevel optimization algorithm for consistently organizing the adaption of source models and the learning of target model. It aims to leverage the knowledge of source models adaptively whilst maximize their complementary benefits collectively to counter the challenge of limited supervision. Extensive experiments on five chest x-ray image datasets show that our method outperforms clearly all the state-of-the-art alternatives. The code is available at https://github.com/wongzbb/metateacher. Zhenbin Wang, Mao Ye 0001, Xiatian Zhu, Liuhan Peng, Yingying Zhu 0003 |
NeurIPS | 2 |
| 2022 | Attention-to-Embedding Framework for Multi-instance Learning
Mei Yang 0002, Mao Ye 0001, Fan Min 0001 |
PAKDD (2) | 3 |
| 2022 | Quality enhancement of compressed screen content video by cross-frame information fusion
Jiawang Huang, Jinzhong Cui, Mao Ye 0001, Shuai Li 0005 |
Neurocomputing | 3 |
| 2022 | Learning missing instances in latent space for incomplete multi-view clustering
Zhiqi Yu, Mao Ye 0001, Siying Xiao |
Knowl. Based Syst. | 2 |
| 2022 | An explicit self-attention-based multimodality CNN in-loop filter for versatile video coding
Menghu Jia, Yanbo Gao, Shuai Li 0005, Jian Yue, Mao Ye 0001 |
Multim. Tools Appl. | 5 |
| 2022 | End-to-end video compression for surveillance and conference videos
Shenhao Wang, Han Gao 0012, Mao Ye 0001, Shuai Li 0005 |
Multim. Tools Appl. | 4 |
| 2022 | Domain adaptation based on source category prototypes
Lihua Zhou, Mao Ye 0001, Siying Xiao |
Neural Comput. Appl. | 2 |
| 2022 | Semantic consistency learning on manifold for source data-free unsupervised domain adaptation
Song Tang 0001, Yan Zou, Jianzhi Lyu, Mao Ye 0001, Shouming Zhong, Jianwei Zhang 0001 |
Neural Networks | 6 |
| 2022 | Source data-free domain adaptation for a faster R-CNN
Mao Ye 0001, Yan Gan, Yiguang Liu |
Pattern Recognit. | 2 |
| 2022 | Self-Alignment for Black-Box Domain Adaptation of Image ClassificationabstractRecently, black-box domain adaptation attracts a lot of attention, which is a new concept to realize domain adaptation with only a cloud API service instead of the source data or well-trained source model, reflecting the focus on development of cloud services and concerns about data security. However, the existing black-box domain adaptation methods always only use high-confidence samples which limits their performance. We propose a self-alignment approach based on statistic moment matching to realize black-box domain adaptation. We construct a model for target domain in the initial stage of our work. Then, we put target data into source model API to obtain the pseudo-labels and divide the target data into high-confidence and low-confidence parts according to their pseudo-labels confidence. By matching the data distributions between these two parts and self-supervised learning on high-confidence part, the performance on both parts samples can be boosted respectively. Information maximization is also applied to the target data to further improve their classification performance. Experiment results confirm that our method achieves state-of-the-art performance. Lihua Zhou, Mao Ye 0001, Xue Li 0001 |
IEEE Signal Process. Lett. | 3 |
| 2022 | Coarse-to-Fine Spatio-Temporal Information Fusion for Compressed Video Quality EnhancementabstractWith the successful application of deformable convolution in aligning different video frames, it has also been used in video compression artifact reduction. The existing methods based on deformable convolution only apply 2D convolutional layers to generate the features for predicting alignment offsets, which is inaccurate due to limited receptive field. In this letter, we propose a new end-to-end network called Coarse-to-Fine Spatio-Temporal Information Fusion (CF-STIF) for compressed video quality enhancement by predicting better offsets with a larger receptive field. Specifically, several 3D convolutional layers are first to roughly fuse the spatio-temporal information in the video sequence, and then a Multi-level Residual Fusion Module (MLRF) is developed to generate global and local fused fine features from different levels for predicting deformable offsets. Thanks to the inherent advantages of 3D convolution and multi-scale strategy, the receptive field is greatly increased in both spatial and temporal dimensions, so that information from neighboring frames can be efficiently aggregated. In the end, the enhanced frame is derived by the proposed reconstruction module (REModule). Both qualitative and quantitative experimental results show that the proposed CF-STIF performs better than the state-of-the-art approaches. Dengyan Luo, Mao Ye 0001, Shuai Li 0005, Xue Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Prototype-Based Multisource Domain AdaptationabstractUnsupervised domain adaptation aims to transfer knowledge from labeled source domain to unlabeled target domain. Recently, multisource domain adaptation (MDA) has begun to attract attention. Its performance should go beyond simply mixing all source domains together for knowledge transfer. In this article, we propose a novel prototype-based method for MDA. Specifically, for solving the problem that the target domain has no label, we use the prototype to transfer the semantic category information from source domains to target domain. First, a feature extraction network is applied to both source and target domains to obtain the extracted features from which the domain-invariant features and domain-specific features will be disentangled. Then, based on these two kinds of features, the named inherent class prototypes and domain prototypes are estimated, respectively. Then a prototype mapping to the extracted feature space is learned in the feature reconstruction process. Thus, the class prototypes for all source and target domains can be constructed in the extracted feature space based on the previous domain prototypes and inherent class prototypes. By forcing the extracted features are close to the corresponding class prototypes for all domains, the feature extraction network is progressively adjusted. In the end, the inherent class prototypes are used as a classifier in the target domain. Our contribution is that through the inherent class prototypes and domain prototypes, the semantic category information from source domains is transformed into the target domain by constructing the corresponding class prototypes. In our method, all source and target domains are aligned twice at the feature level for better domain-invariant features and more closer features to the class prototypes, respectively. Several experiments on public data sets also prove the effectiveness of our method. Lihua Zhou, Mao Ye 0001, Ce Zhu, Luping Ji |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Teacher-Supervised Generative Adversarial NetworksabstractAlthough generative adversarial networks (GANs) show impressive effects on image generation, existing GANs suffer an unstable training process, and thus result in poor image quality sometimes. To solve this problem, we first introduce a supervision mechanism into GANs and propose a teacher-supervised GAN (GAN-T) model. Specifically, we design a teacher supervision mechanism to inspect whether the features of generated images are as real as those of real images. If not, we add action into the generator. The action takes the encoding of the real image as prior knowledge to guide the generation of samples. We then apply our proposed method to existing GANs to show its compatibility with them. Finally, we conduct extensive experiments on the tasks of noise-to-image generation and image translation, and experimental results show that our proposed method can significantly stabilize the training process of the generator and improve the quality of generated images. Yan Gan, Tao Xiang 0001, Hangcheng Liu, Mao Ye 0001 |
ICME | 4 |
| 2021 | Source-Style Transferred Mean Teacher for Source-data Free Object DetectionabstractUnsupervised cross-domain object detection transfers a detection model trained on a source domain to the target domain that has a different data distribution from the source domain. Conventional domain adaptation detection protocols need source domain data during adaptation. However, due to some reasons such as data security, privacy and storage, we cannot access the source data in many practical applications. In this paper, we focus on source-data free domain adaptive object detection, which uses the pre-trained source model instead of the source data for cross-domain adaptation. Due to the lack of source data, we cannot directly align domain distribution between domains. To challenge this, we propose the Source style transferred Mean Teacher (SMT) for source-data free Object Detection. The batch normalization layers in the pre-trained model contain the style information and the data distribution of the non-observed source data. Thus we use the batch normalization information from the pre-trained source model to transfer the target domain feature to the source-like style feature to make full use of the knowledge from the pre-trained source model. Meanwhile, we use the consistent regularization of the Mean Teacher to further distill the knowledge from the source domain to the target domain. Furthermore, we found that by adding perturbations associated with the target domain distribution, the model can increase the robustness of domain-specific information, thus making the learned model generalized to the target domain. Experiments on multiple domain adaptation object detection benchmarks verify that our method is able to achieve state-of-the-art performance. Mao Ye 0001, Shuaifeng Li, Xue Li 0001 |
MMAsia | 2 |
| 2021 | Alignment-Free Video Compression Artifact ReductionabstractThe past few years have witnessed the great success of using multi-frame information to enhance the quality of compressed video. Most existing methods do frame or feature alignments to bring similar information from neighborhood frames much closer to enhance frame quality. However, inaccurate motion estimation will bring new artifacts. In this paper, we propose a new approach without alignment, which takes each non-Peak Quality Frame (non-PQF) and its two adjacent Peak Quality Frames (PQFs) as input. A Pre-processing module based on multi-scale feature extraction strategy is used to broaden receptive field of the network. Then, an enhancement module uses a two-stream feature extraction architecture to combine deep architecture and attention mechanism to further gather similar information. In this module, the lost high-frequency and similar information can be further retrieved from the adjacent PQFs. The proposed network is trained in an end-to-end manner. Compared with the alignment based methods, the competitive results can be obtained. A large number of qualitative and quantitative experimental results demonstrate the robustness and effectiveness of the proposed method. Dengyan Luo, Mao Ye 0001, Shengjie Chen, Xue Li 0001 |
VCIP | 2 |
| 2021 | Local-Global Attentive Adaptation for Object Detection
Jingjing Li 0001, Xingpeng Li, Zhekai Du, Mao Ye 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2021 | A novel hybrid augmented loss discriminator for text-to-image synthesisabstractFor the text-to-image synthesis task, most discriminators in existing generative adversarial networks based methods tend to fall into a local suboptimal state too early in the training process, resulting in the poor quality of generated images. To address the above problems, a hybrid augmented loss discriminator is designed. In this designed discriminator, to reduce the sensitivity of the discriminator classification recognition, make it pay attention to the semantic and structural changes, we add the loss value of the fake sample to the loss value of the real sample. Moreover, to indirectly guide the generator to generate samples, the loss value of the real sample is added to the fake sample. The loss value mixed with real and fake samples actually augments signal transmission. It perturbs parameter update of the discriminator during optimization and prevents the discriminator from falling into the local suboptimal state prematurely. Whereafter, we apply the proposed discriminator to two kinds of text-to-image synthesis tasks. Experimental results show that the proposed method can help the baseline models to improve performance. Yan Gan, Mao Ye 0001, Shangming Yang, Tao Xiang 0001 |
Int. J. Intell. Syst. | 2 |
| 2021 | Source data-free domain adaptation of object detector through domain-specific perturbationabstractThe current unsupervised cross-domain detection methods need source domain data to retrain the detection model in target domain. However, the source domain data may be unavailable due to privacy, decentralization, or computation resource restrictions. A natural idea is to optimize the parameters of the source domain model by self-supervised learning based on pseudo labels. We propose another approach from the viewpoint of noise perturbation without pseudo-labeling. It can be assumed that the source and target domains are actually derived from a domain invariant space through domain-specific perturbations, respectively. A super target domain can be constructed by augmenting more target domain perturbations to the target domain images. The optimal direction of the target domain to the domain invariant space can be approximated as the alignment direction from the super target domain to the target domain. Based on this idea, we propose a novel method called SOAP (SOurce data-free domain Adaptation through domain Perturbation) which can remove domain perturbation from the target domain. The image-level, instance-level, and category consistency regularizations based on Mean Teacher structure are proposed to learn the correct alignment direction. Specifically, the category consistency can also further improve the classification accuracy. Extensive experiments on multiple domain adaptation scenarios demonstrate that SOAP achieves better performance surpassing the baseline (Faster R-CNN) and multiple state-of-the-art domain adaptation methods which need to access source domain data. Mao Ye 0001, Yan Gan, Xue Li 0001, Yingying Zhu 0003 |
Int. J. Intell. Syst. | 2 |
| 2021 | Rain streaks removal from single image based on texture constraint of background scene
Shuangli Du, Yiguang Liu, Mao Ye 0001, Minghua Zhao |
Neurocomputing | 3 |
| 2021 | A new financial data forecasting model using genetic algorithm and long short-term memory network
Yelin Gao, Yan Gan, Mao Ye 0001 |
Neurocomputing | 4 |
| 2021 | Unsupervised feature selection via multi-step markov probability relationship
Yan Min, Mao Ye 0001, Yulin Jian, Ce Zhu, Shangming Yang |
Neurocomputing | 2 |
| 2021 | Domain adaptation of object detector using scissor-like networks
Mao Ye 0001, Yan Gan, Dongde Hou |
Neurocomputing | 2 |
| 2021 | Learning various length dependence by dual recurrent neural networks
Chenpeng Zhang, Shuai Li 0005, Mao Ye 0001, Ce Zhu, Xue Li 0001 |
Neurocomputing | 3 |
| 2021 | Bidirectional generative transductive zero-shot learning
Xinpeng Li 0005, Mao Ye 0001, Xue Li 0001, Qiang Dou, Qiao Lv |
Neural Comput. Appl. | 3 |
| 2020 | Distribution-Aware Coordinate Representation for Human Pose EstimationabstractWhile being the de facto standard coordinate representation for human pose estimation, heatmap has not been investigated in-depth. This work fills this gap. For the first time, we find that the process of decoding the predicted heatmaps into the final joint coordinates in the original image space is surprisingly significant for the performance. We further probe the design limitations of the standard coordinate decoding method, and propose a more principled distributionaware decoding method. Also, we improve the standard coordinate encoding process (i.e. transforming ground-truth coordinates to heatmaps) by generating unbiased/accurate heatmaps. Taking the two together, we formulate a novel Distribution-Aware coordinate Representation of Keypoints (DARK) method. Serving as a model-agnostic plug-in, DARK brings about significant performance boost to existing human pose estimation models. Extensive experiments show that DARK yields the best results on two common benchmarks, MPII and COCO. Besides, DARK achieves the 2nd place entry in the ICCV 2019 COCO Keypoints Challenge. The code is available online. Feng Zhang 0052, Xiatian Zhu, Hanbin Dai, Mao Ye 0001, Ce Zhu |
CVPR | 4 |
| 2020 | Sentence guided object color change by adversarial learning
Yan Gan, Kedi Liu, Mao Ye 0001 |
Neurocomputing | 3 |
| 2020 | Improving Action Recognition Using Sequence Prediction LearningabstractSkeleton-based action recognition distinguishes human actions using the trajectories of skeleton joints, which can be a good representation of human behaviors. Conventional methods usually construct classifiers with hand-crafted or the learned features to recognize human actions. Different from constructing a direct action classifier for action recognition task, this paper attempts to identify human actions based on the development trends of behavior sequences. Specifically, we first utilize the memory neural network to construct action predictors for each kind of activity. These action predictors can then output the action trends at the next time step. According to the predictions of these action predictors at each time step and the removal rule, the poor predictors can be eliminated step by step, and the IDentity(ID) number of the last predictor left is considered as the label of the action sequence to be categorized. We compare the proposed action recognition algorithm using sequence prediction learning with other methods on two publicly available datasets. Our experimental results consistently demonstrate the feasibility and effectiveness of the suggested method. It also proves the importance of prediction learning for action recognition. Mao Ye 0001, Jianwei Zhang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2020 | Knowledge based domain adaptation for semantic segmentation
Mao Ye 0001, Yan Gan, Wencong Zhang |
Knowl. Based Syst. | 2 |
| 2020 | Recurrent matching networks of spatial alignment learning for person re-identification
Mao Ye 0001, Jiuxia Guo |
Multim. Tools Appl. | 4 |
| 2020 | Generative adversarial networks with denoising penalty and sample augmentation
Yan Gan, Kedi Liu, Mao Ye 0001 |
Neural Comput. Appl. | 3 |
| 2020 | Attention guided neural network models for occluded pedestrian detection
Tengtao Zou, Shangming Yang, Mao Ye 0001 |
Pattern Recognit. Lett. | 4 |
| 2019 | Fast Human Pose EstimationabstractExisting human pose estimation approaches often only consider how to improve the model generalisation performance, but putting aside the significant efficiency problem. This leads to the development of heavy models with poor scalability and cost-effectiveness in practical use. In this work, we investigate the under-studied but practically critical pose model efficiency problem. To this end, we present a new Fast Pose Distillation (FPD) model learning strategy. Specifically, the FPD trains a lightweight pose neural network architecture capable of executing rapidly with low computational cost. It is achieved by effectively transferring the pose structure knowledge of a strong teacher network. Extensive evaluations demonstrate the advantages of our FPD method over a broad range of state-of-the-art pose estimation approaches in terms of model cost-effectiveness on two standard benchmark datasets, MPII Human Pose and Leeds Sports Pose. Feng Zhang 0052, Xiatian Zhu, Mao Ye 0001 |
CVPR | 3 |
| 2019 | Generative adversarial networks with augmentation and penalty
Yan Gan, Kedi Liu, Mao Ye 0001 |
Neurocomputing | 3 |
| 2019 | Adaptive pedestrian detection by predicting classifier
Song Tang 0001, Mao Ye 0001, Pei Xu 0009, Xudong Li 0001 |
Neural Comput. Appl. | 2 |
| 2019 | Adaptive Deep Convolutional Neural Networks for Scene-Specific Object DetectionabstractA deep convolutional neural network (CNN) becomes a widely used tool for object detection. Many previous works have achieved excellent performance on object detection benchmarks. However, these works present generic detectors whose performance will drop rapidly when they are applied to a surveillance scene. In this paper, we propose an efficient method to construct a scene-specific regression model based on a generic CNN-based classifier. Our regression model is an adaptive deep CNN (ADCNN), which can predict object locations in the surveillance scene. First, we transfer the generic CNN-based classifier to the surveillance scene by selecting useful kernels. Second, we learn the context information of the surveillance scene in our regression model for accurate location prediction. Our main contributions are: 1) a transfer learning method that selects useful kernels in the generic CNN-based classifier; 2) a special architecture that can effectively learn the local and global context information in the surveillance scene; and 3) a new objective function to effectively train parameters in ADCNN. Compared with some state-of-the-art models, ADCNN achieves the best performance on three surveillance data sets for pedestrian detection and one surveillance data set for vehicle detection. Xudong Li 0001, Mao Ye 0001, Yiguang Liu, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Unpaired cross domain image translation with augmented auxiliary domain information
Yan Gan, Junxin Gong, Mao Ye 0001, Kedi Liu |
Neurocomputing | 3 |
| 2018 | Road segmentation for all-day outdoor robot navigation
Haiqiang Chen, Mao Ye 0001, Xi Cai |
Neurocomputing | 4 |
| 2018 | Single image deraining via decorrelating the rain streaks and background scene in gradient domain
Shuangli Du, Yiguang Liu, Mao Ye 0001, Jian Guo Liu 0005 |
Pattern Recognit. | 3 |
| 2017 | Memory-based pedestrian detection through sequence learningabstractHuman recognize an object through eyes scanning in a certain order. We think that the proper order is helpful for capturing useful characteristics, which makes our recognition process rapidly and accurately. Therefore, we propose a memory-based sequence learning model to simulate the human recognition process. Firstly, we divide the image without overlapping to generate the sequence. Then, a convolutional neural network is used for feature extraction. Next, the sequence is re-sorted by order of importance. Finally, a long short-term memory successively receives the sequence to memorize the sequential patterns and predict the sequence label. In addition, we propose a joint learning method to make our model efficiently learn both of the sequence order and the sequence patterns. Our model is applied in the region-based detection framework for pedestrian detection. Compared with the state-of-the-art methods on two pedestrian datasets, our method achieves the comparable performance in term of accuracy and speed. Xudong Li 0001, Mao Ye 0001, Yiguang Liu, Ce Zhu |
ICME | 2 |
| 2017 | Age invariant face recognition and retrieval by coupled auto-encoder networks
Chenfei Xu, Qihe Liu, Mao Ye 0001 |
Neurocomputing | 3 |
| 2017 | Accurate object detection using memory-based models in surveillance scenes
Xudong Li 0001, Mao Ye 0001, Yiguang Liu, Feng Zhang 0052, Song Tang 0001 |
Pattern Recognit. | 2 |
| 2017 | A Bayesian Approach to Camouflaged Moving Object DetectionabstractMoving object detection is about foreground and background separation based on motion detection. Detecting moving objects from similarly colored background (known as camouflage problem) has been a long-standing open question in this field. Discriminative modeling (DM), which focuses on enhancing the performance to distinguish foreground from background with discriminative features and well-designed classifiers, has been widely used for moving object detection. However, DM may tend to fail when encountering the camouflage problem, as the class separability in camouflaged areas is generally poor. In this paper, we propose a new strategy, camouflage modeling (CM), to identify camouflaged foreground pixels. In view of the fact that camouflage involves both foreground and background, we need to model both the background and the foreground, and compare them in a well-designed way in camouflage detection. Specifically, we develop a global model for the background, and an integration of global and local models for the foreground, respectively. Based on both background and foreground models, we introduce a factor to measure the degree of camouflage, and further identify truly camouflaged areas. In view of the fact that a moving object is usually composed of both camouflaged and noncamouflaged areas, CM and DM are fused in a Bayesian framework to perform complete object detection. Experiments are conducted on testing sequences to demonstrate the effectiveness of the proposed algorithm. Xiang Zhang 0006, Ce Zhu, Yipeng Liu 0001, Mao Ye 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | Geometry Guided Multi-Scale Depth Map Fusion via Graph OptimizationabstractIn depth discontinuous and untextured regions, depth maps created by multiple view stereopsis are with heavy noises, but existing depth map fusion methods cannot handle it explicitly. To tackle the problem, two novel strategies are proposed: 1) a more discriminative fusion method, which is based on geometry consistency, measuring the consistency, and stability of surface geometry computed on both partial and global surfaces, different from traditional methods only using visibility consistency; 2) a graph optimization method which fuses pyramids of depth maps as mutual complementary information is available in different scales, and differs from existing multi-scale fusion methods. The method considers both sampling scale of a point and relations among points, and is proven to be solvable by graph cuts. Experimental results verify the superior performance of the proposed method to the traditional visibility consistency-based methods, and the proposed method is also compared favorably with a number of state-of-the-art methods. Moreover, the proposed method achieves the highest completeness among all the methods compared. Pengfei Wu 0002, Yiguang Liu, Mao Ye 0001, Yunan Zheng |
IEEE Trans. Image Process. | 3 |
| 2017 | Fast and Adaptive 3D Reconstruction With Extensively High CompletenessabstractThe seed-and-expand scheme is appropriate for multiple view stereo, since it can build dense point clouds adaptively by avoiding unnecessary computation. However, due to the irregularity of the algorithm, it is not suitable for parallel computing on general public utilities (GPU). This paper is the first attempt to implement the irregular seed-and-expand method on GPU for multiple view stereo problems. Meanwhile, a hierarchical parallel computing architecture is also proposed to maximize the usage of both CPU and GPU. The adaptivity of the seed-and-expand scheme is pushed further by processing a pixel several rounds while, in order to maintain regularity for GPU implementation, every seed has exactly the same behavior in a single round of optimization. The high adaptivity also improves the robustness of the proposed method, thus aggressive matching score and a view selection method can be used to improve the reconstruction completeness extensively, without smearing out local details and lowering the accuracy. Compared with the state of the art, the proposed method achieves higher accuracy and completeness on standard datasets. The proposed method is also very fast. It is maximally five times faster than other methods running on a CPU and is on par with the regular depth map-based methods on GPU, which are naturally suitable for GPU acceleration. Pengfei Wu 0002, Yiguang Liu, Mao Ye 0001, Shuangli Du |
IEEE Trans. Multim. | 3 |
| 2016 | Memory-based Gait Recognition
Mao Ye 0001, Xudong Li 0001, Feng Zhang 0052 |
BMVC | 2 |
| 2016 | Memory-based object detection in surveillance scenesabstractObject detection is a significant step of intelligent video surveillance. The existing methods achieve the goals by technically designing or learning special features and detection models. Conversely, we propose a method to simulate the mechanism of memory and prediction in our brain. Firstly, a fix-sized window is slid on a static image to generate sequences. Then, a convolutional neural network extracts the sequence features. Finally, a long short-term memory receives these sequence features in proper order to memorize and recognize the sequential patterns. Our contributions are 1) a memory-based classification model in which both of feature learning and sequence learning are integrated subtly, and 2) a memory-based prediction model which is specially designed to predict the potential object locations in the surveillance scene. Compared with the state-of-the-art methods, our method obtains the best performance on three surveillance datasets. Our method may give some new insights on object detection researches. Xudong Li 0001, Mao Ye 0001, Feng Zhang 0052, Song Tang 0001 |
ICME | 2 |
| 2016 | Action recognition by learning temporal slowness invariant features
Lishen Pei, Mao Ye 0001, Xuezhuan Zhao, Yumin Dou, Jiao Bao |
Vis. Comput. | 2 |
| 2015 | Fast crowd density estimation with convolutional neural networks
Pei Xu 0009, Xudong Li 0001, Qihe Liu, Mao Ye 0001, Ce Zhu |
Eng. Appl. Artif. Intell. | 5 |
| 2015 | Gas Recognition under Sensor Drift by Using Deep LearningabstractMachine olfaction is an intelligent system that combines a cross-sensitivity chemical sensor array and an effective pattern recognition algorithm for the detection, identification, or quantification of various odors. Data collected by the sensor array are the multivariate time series signals with a complex structure, and these signals become more difficult to analyze due to sensor drift. In this work, we focus on improving the classification performance under sensor drift by using the deep learning method, which is popular nowadays. Compared with other methods, our method can effectively tackle sensor drift by automatically extracting features, thus not only removing the complexity of designing the hand-made features but also making it pervasive for a variety of application in machine olfaction. Our experimental results show that the deep learning method can learn the features that are more robust to drift than the original input and achieves high classification accuracy. Qihe Liu, Xiaonan Hu, Mao Ye 0001, Xianqiong Cheng |
Int. J. Intell. Syst. | 3 |
| 2015 | Fast multi-class action recognition by querying inverted index tables
Lishen Pei, Mao Ye 0001, Pei Xu 0009 |
Multim. Tools Appl. | 2 |
| 2015 | Learning to pool high-level features for face representation
Renjie Huang, Mao Ye 0001, Pei Xu 0009, Yumin Dou |
Vis. Comput. | 2 |
| 2014 | Motion detection via a couple of auto-encoder networksabstractMotion detection is a basis step for video processing. Previous works of motion detection based on deep learning need clean foreground or background images which always do not exist in practice. To address this challenge, a novel and practical method is proposed based on auto-encoder neural networks. First, the approximate background images are obtained via an auto-encoder network (called Reconstruction Network) from video frames. Then, a background model is learned based on these images by using another auto-encoder network (called Background Network). To be more resilient, our background model can be updated on-line to absorb more training samples. Our main contributions are 1) the architecture of the couple of auto-encoder networks which can model the background very efficiently; 2) the online learning algorithm in which a method of searching the minimizing effect parameters is adopted to accelerate the training of the Reconstruction Network. Our approach improves the motion detection performance on three data sets. Pei Xu 0009, Mao Ye 0001, Qihe Liu, Xudong Li 0001, Lishen Pei |
ICME | 2 |
| 2014 | Dynamic Background Learning through Deep Auto-encoder NetworksabstractBackground learning is a pre-processing of motion detection which is a basis step of video analysis. For the static background, many previous works have already achieved good performance. However, the results on learning dynamic background are still much to be improved. To address this challenge, in this paper, a novel and practical method is proposed based on deep auto-encoder networks. Firstly, dynamic background images are extracted through a deep auto-encoder network (called Background Extraction Network) from video frames containing motion objects. Then, a dynamic background model is learned by another deep auto-encoder network (called Background Learning Network) using the extracted background images as the input. To be more flexible, our background model can be updated on-line to absorb more training samples. Our main contributions are 1) a cascade of two deep auto-encoder networks which can deal with the separation of dynamic background and foregrounds very efficiently; 2) a method of online learning is adopted to accelerate the training of Background Extraction Network. Compared with previous algorithms, our approach obtains the best performance over six benchmark data sets. Especially, the experiments show that our algorithm can handle large variation background very well. Pei Xu 0009, Mao Ye 0001, Xue Li 0001, Qihe Liu, Yi Yang 0001 |
ACM Multimedia | 2 |
| 2014 | An extension to Rough c-means clustering based on decision-theoretic Rough Sets model
Mao Ye 0001 |
Int. J. Approx. Reason. | 2 |
| 2014 | One example based action detection in hough space
Lishen Pei, Mao Ye 0001, Pei Xu 0009, Xuezhuan Zhao, Guanjun Guo |
Multim. Tools Appl. | 2 |
| 2014 | Convergence Analysis of Graph Regularized Non-Negative Matrix FactorizationabstractGraph regularized non-negative matrix factorization (NMF) algorithms can be applied to information retrieval, image processing, and pattern recognition. However, challenge that still remains is to prove the convergence of this class of learning algorithms since the geometrical structure of the data space is considered. This paper presents the convergence properties of the graph regularized NMF learning algorithms. In the analysis, we focus on the study of Euclidian distance based algorithms. The structures of the fixed points are presented. The non-divergence of the learning algorithms is analyzed by constructing invariant sets for update rules. Based on Lyapunov indirect method, the stability of the algorithms is discussed in detail. The analysis shows that this class of NMF algorithms can converge to their fixed points under some given conditions. In the simulations, theoretical results presented in the paper are confirmed. For different initializations and data sets, variations of cost functions and decomposition data in the learning are presented to show the convergence features of the discussed NMF update rules, and the convergence speed of the algorithms is also investigated. Shangming Yang, Zhang Yi 0001, Mao Ye 0001, Xiaofei He 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2014 | Structured Streaming Skeleton - A New Feature for Online Human Gesture RecognitionabstractOnline human gesture recognition has a wide range of applications in computer vision, especially in human-computer interaction applications. The recent introduction of cost-effective depth cameras brings a new trend of research on body-movement gesture recognition. However, there are two major challenges: (i) how to continuously detect gestures from unsegmented streams, and (ii) how to differentiate different styles of the same gesture from other types of gestures. In this article, we solve these two problems with a new effective and efficient feature extraction method—Structured Streaming Skeleton (SSS)—which uses a dynamic matching approach to construct a feature vector for each frame. Our comprehensive experiments on MSRC-12 Kinect Gesture, Huawei/3DLife-2013, and MSR-Action3D datasets have demonstrated superior performances than the state-of-the-art approaches. We also demonstrate model selection based on the proposed SSS feature, where the classifier of squared loss regression with l 2,1 norm regularization is a recommended classifier for best performance. Xin Zhao 0013, Xue Li 0001, Chaoyi Pang, Quan Z. Sheng, Sen Wang 0001, Mao Ye 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2014 | Discriminative Hough context model for object detection
Mao Ye 0001 |
Vis. Comput. | 2 |
| 2013 | Multi-class action recognition based on inverted index of action statesabstractA fast inverted index based algorithm is introduced for multi-class action recognition. At first, we compute the shape-motion features of the automatically localized actor. Secondly, a binary state tree is built by hierarchically clustering of the extracted features, and the action states are the cluster centers. Then videos are represented as sequences of states by searching the state binary tree. With the labeled state sequences, we create the inverted index tables. During testing, the state and the state transition scores are computed by querying the inverted index tables. With the learned weight, we compute an action recognition score vector. The recognized action class is the index of the maximum score element. Our key contribution is that we propose a fast inverted index based multi-class action recognition approach. Experiments on several challenging data sets confirm the performance of this approach. Lishen Pei, Mao Ye 0001, Pei Xu 0009, Xuezhuan Zhao |
ICIP | 2 |
| 2013 | Shadow compensation and illumination normalization of face image
Mao Ye 0001, Shangming Yang |
Mach. Vis. Appl. | 2 |
| 2013 | Global Minima Analysis of Lee and Seung's NMF Algorithms
Shangming Yang, Mao Ye 0001 |
Neural Process. Lett. | 2 |
| 2012 | Abnormal crowd behavior detection using high-frequency and spatio-temporal features
Mao Ye 0001, Xue Li 0001, Fengjuan Zhao |
Mach. Vis. Appl. | 2 |
| 2011 | Extracting Specific Signal from Post-nonlinear Mixture Based on Maximum Negentropy
Dongxiao Ren, Mao Ye 0001, Yuanxiang Zhu |
ISNN (2) | 2 |
| 2010 | Improved exponential stability criteria for discrete-time neural networks with time-varying delay
Shu Lü, Shouming Zhong, Mao Ye 0001 |
Neurocomputing | 4 |
| 2009 | Projected outlier detection in high-dimensional mixed-attributes data set
Mao Ye 0001, Xue Li 0001, Maria E. Orlowska |
Expert Syst. Appl. | 1 |
| 2009 | Global asymptotic stability analysis of nonlinear differential equations in hybrid bidirectional associative memory neural networks with distributed time-varying delays
Yongqiang Zhou, Shouming Zhong, Mao Ye 0001 |
Neurocomputing | 3 |
| 2008 | A few online algorithms for extracting minor generalized eigenvectorsabstractRecently, a few adaptive algorithms for generalized eigen-decomposition have been proposed, which are very useful in many applications such as digital mobile communications, Blind signal separation, etc. These algorithms are all focusing on extracting principal generalized eigenvectors. However, in many practical applications such as dimension reduction and signal processing, extracting the minor generalized eigenvectors adaptively are needed. Because of little literatures in the community, we discuss several approaches that lead to a few novel algorithms for extracting minor generalized eigenvectors. First, we derive an adaptive algorithms by using a single-layer linear forward neural network from the viewpoint of linear discriminant analysis(LDA). And the algorithm to extract multiple minor generalized eigenvectors are also proposed by using orthogonality property. Second, by using gradient ascent approach of some objective functions, we can derive more algorithms and explain the first algorithm. Then, we extend these algorithms to minor generalized eigenvector problem. Theoretical analysis shows that these algorithms are stable and convergent to the minor generalized eigenvectors. Simulations have been conducted for illustration of the efficiency and effectiveness of our algorithms. Mao Ye 0001, Yongguo Liu, Qihe Liu |
IJCNN | 1 |
| 2008 | A tabu search approach for the minimum sum-of-squares clustering problem
Yongguo Liu, Zhang Yi 0001, Mao Ye 0001, Kefei Chen |
Inf. Sci. | 4 |
| 2007 | Blind Separation of Positive Signals by Using Genetic Algorithm
Mao Ye 0001, Zengan Gao, Xue Li 0001 |
ISNN (3) | 1 |
| 2007 | An Efficient Measure of Signal Temporal Predictability for Blind Source Separation
Mao Ye 0001, Xue Li 0001 |
Neural Process. Lett. | 1 |
| 2006 | Monotonic Convergence of a Nonnegative ICA Algorithm on Stiefel Manifold
Mao Ye 0001, Xuqian Fan, Qihe Liu |
ICONIP (1) | 1 |
| 2006 | Convergence Analysis of a Discrete-Time Single-Unit Gradient ICA Algorithm
Mao Ye 0001, Xue Li 0001, Chengfu Yang, Zengan Gao |
ISNN (1) | 1 |
| 2006 | Global convergence analysis of a discrete time nonnegative ICA algorithmabstractWhen the independent sources are known to be nonnegative and well-grounded, which means that they have a nonzero pdf in the region of zero, Oja and Plumbley have proposed a "Nonnegative principal component analysis (PCA)" algorithm to separate these positive sources. Generally, it is very difficult to prove the convergence of a discrete-time independent component analysis (ICA) learning algorithm. However, by using the skew-symmetry property of this discrete-time "Nonnegative PCA" algorithm, if the learning rate satisfies suitable condition, the global convergence of this discrete-time algorithm can be proven. Simulation results are employed to further illustrate the advantages of this theory. Mao Ye 0001 |
IEEE Trans. Neural Networks | 1 |
| 2006 | A Class of Self-Stabilizing MCA Learning AlgorithmsabstractIn this letter, we propose a class of self-stabilizing learning algorithms for minor component analysis (MCA), which includes a few well-known MCA learning algorithms. Self-stabilizing means that the sign of the weight vector length change is independent of the presented input vector. For these algorithms, rigorous global convergence proof is given and the convergence rate is also discussed. By combining the positive properties of these algorithms, a new learning algorithm is proposed which can improve the performance. Simulations are employed to confirm our theoretical results. Mao Ye 0001, Xuqian Fan, Xue Li 0001 |
IEEE Trans. Neural Networks | 1 |
| 2005 | Face Recognition Using Fisher Non-negative Matrix Factorization with Sparseness Constraints
Xiaorong Pu, Zhang Yi 0001, Ziming Zheng, Mao Ye 0001 |
ISNN (2) | 5 |
| 2005 | Robust Beamforming by a Globally Convergent MCA Neural Network
Mao Ye 0001 |
ISNN (1) | 1 |
| 2005 | Global convergence analysis of a self-stabilizing MCA learning algorithm
Mao Ye 0001 |
Neurocomputing | 1 |
| 2005 | A globally convergent learning algorithm for PCA neural networks
Mao Ye 0001, Zhang Yi 0001, Jiancheng Lv 0001 |
Neural Comput. Appl. | 1 |
| 2005 | Complete Convergence of Competitive Neural Networks with Different Time Scales
Mao Ye 0001, Zhang Yi 0001 |
Neural Process. Lett. | 1 |
| 2005 | Convergence analysis of a deterministic discrete time system of Oja's PCA learning algorithmabstractThe convergence of Oja's principal component analysis (PCA) learning algorithms is a difficult topic for direct study and analysis. Traditionally, the convergence of these algorithms is indirectly analyzed via certain deterministic continuous time (DCT) systems. Such a method will require the learning rate to converge to zero, which is not a reasonable requirement to impose in many practical applications. Recently, deterministic discrete time (DDT) systems have been proposed instead to indirectly interpret the dynamics of the learning algorithms. Unlike DCT systems, DDT systems allow learning rates to be constant (which can be a nonzero). This paper will provide some important results relating to the convergence of a DDT system of Oja's PCA learning algorithm. It has the following contributions: 1) A number of invariant sets are obtained, based on which we can show that any trajectory starting from a point in the invariant set will remain in the set forever. Thus, the nondivergence of the trajectories is guaranteed. 2) The convergence of the DDT system is analyzed rigorously. It is proven, in the paper, that almost all trajectories of the system starting from points in an invariant set will converge exponentially to the unit eigenvector associated with the largest eigenvalue of the correlation matrix. In addition, exponential convergence rate are obtained, providing useful guidelines for the selection of fast convergence learning rate. 3) Since the trajectories may diverge, the careful choice of initial vectors is an important issue. This paper suggests to use the domain of unit hyper sphere as initial vectors to guarantee convergence. 4) Simulation results will be furnished to illustrate the theoretical results achieved. Zhang Yi 0001, Mao Ye 0001, Jiancheng Lv 0001, Kok Kiong Tan |
IEEE Trans. Neural Networks | 2 |
| 2004 | Convergence Analysis for Oja+ MCA Learning Algorithm
Jiancheng Lv 0001, Mao Ye 0001, Zhang Yi 0001 |
ISNN (1) | 2 |
| 2004 | On the Discrete Time Dynamics of the MCA Neural Networks
Mao Ye 0001, Zhang Yi 0001 |
ISNN (1) | 1 |