EDBT 2026 Demo / reviewers in the wild / expert
Aidong Men
dblp:48/7458
· DBLP profile ↗
78ranked-venue papers
1as first author
30since 2021 · last 2026
0000-0001-6168-9276ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 14 since 2021Artificial intelligence and machine learning · 15 · 13 since 2021Computer networks · 8 · 4 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Zero-shot diverse audio captioning with diffusion models
Yonggang Zhu, Yiming Zhang 0025, Li Xiao 0005, Wenwu Wang 0001, Aidong Men |
Knowl. Based Syst. | 5 |
| 2025 | Filter or Compensate: Towards Invariant Representation from Distribution Shift for Anomaly DetectionabstractRecent Anomaly Detection (AD) methods have achieved great success with In-Distribution (ID) data. However, real-world data often exhibits distribution shift, causing huge performance decay on traditional AD methods. From this perspective, few previous work has explored AD with distribution shift, and the distribution-invariant normality learning has been proposed based on the Reverse Distillation (RD) framework. However, we observe the misalignment issue between the teacher and the student network that causes detection failure, thereby propose FiCo, Filter or Compensate, to address the distribution shift issue in AD. FiCo firstly compensates the distribution-specific information to reduce the misalignment between the teacher and student network via the Distribution-Specific Compensation (DiSCo) module, and secondly filters all abnormal information to capture distribution-invariant normality with the Distribution-Invariant Filter (DiIFi) module. Extensive experiments on three different AD benchmarks demonstrate the effectiveness of FiCo, which outperforms all existing state-of-the-art (SOTA) methods, and even achieves better results on the ID scenario compared with RD-based methods. Zining Chen, Xingshuang Luo, Weiqiu Wang, Zhicheng Zhao 0001, Aidong Men |
AAAI | 6 |
| 2025 | Radar2ECG: Multi-Scale Bottleneck Fusion and Cross-modal Semantic Distillation for Conditional Electrocardiogram Generation from Radar Heart SoundabstractThe field of conditional Electrocardiogram(ECG) generation focuses on generating specified ECGs under given conditions for medical purposes. Existing methods are typically based on conditions of simple inputs like text or lead types. However, they struggle to handle the complexity of radar heart sound signals due to the lack of effective feature extraction, which hinders capturing the intricate waveform correlations between radar heart sounds and ECGs. Considering that radar-detected heart sound signals are contactless, the application is of essential value in a real-world deployment like sleep scenarios. Moreover, no prior approaches have addressed this specific task. To tackle this challenge, we propose a novel multi-scale feature fusion network framework, Radar2ECG. This model leverages pre-trained autoencoders for heart sound and ECG signals, aligning and integrating multi-layer features through a bottleneck structure to enhance receptive fields and reduce redundant features, thereby capturing the correlations between heart sounds and ECGs. Finally, we employ knowledge distillation to transfer knowledge from the ECG decoder to the heart sound decoder. We present three anomaly type datasets and extensive experiments conducted on both normal and abnormal datasets demonstrate that our method outperforms existing models in both accuracy and robustness. The multi-scale feature fusion significantly improves performance, showcasing strong potential in ECG generation and heart sound anomaly detection tasks. Jinye Li, Aidong Men, Yang Liu 0105, Pengda Han, Qingchao Chen |
ICASSP | 2 |
| 2025 | Camera-Invariant Meta-Learning Network for Single-Camera-Training Person ReidentificationabstractSingle-camera-training person reidentification (SCT re-ID) aims to train a reidentification (re-ID) model using single-camera-training (SCT) datasets where each person appears in only one camera. The main challenge of SCT re-ID is to learn camera-invariant feature representations without cross-camera same-person (CCSP) data as supervision. Previous methods address it by assuming that the most similar person should be found in another camera. However, this assumption is not guaranteed to be correct. In this article, we propose a novel solution: the camera-invariant meta-learning network (CIMN) for SCT re-ID. CIMN operates under the premise that camera-invariant feature representations should remain robust despite changes in camera settings. To achieve this, we partition the training data into a meta-train set and a meta-test set based on camera IDs. We then conduct a cross-camera simulation (CCS) using a meta-learning strategy, aiming to enforce the feature representations learned from the meta-train set to be robust when applied to the meta-test set. We further introduce three specific loss functions to leverage potential identity relations between the meta-train set and the meta-test set. Through the CCS and the introduced loss functions, CIMN can extract feature representations that are both camera-invariant and identity-discriminative even in the absence of CCSP data. Our experimental results demonstrate that CIMN can extract feature representations that are both camera-invariant and identity-discriminative, even in the absence of CCSP data. our method achieves comparable performance with and without the use of CCSP data, and outperforms state-of-the-art methods on three SCT re-ID benchmarks. Jiangbo Pei, Zhuqing Jiang, Aidong Men, Haiying Wang 0005, Haiyong Luo, Shiping Wen 0001 |
IEEE Internet Things J. | 3 |
| 2025 | Understand and Detect: Multi-step zero-shot detection with image-level specific prompt
Miaotian Guo, Kewei Wu, Zhuqing Jiang, Haiying Wang 0005, Aidong Men |
Knowl. Based Syst. | 5 |
| 2025 | Selection, Ensemble, and Adaptation: Advancing Multi-Source-Free Domain Adaptation via Architecture ZooabstractConventional Multi-Source Free Domain Adaptation (MSFDA) assumes that each source domain provides a single source model, and all source models adopt a uniform architecture. This paper introduces Zoo-MSFDA, a more general setting that allows each source domain to offer a zoo of multiple source models with different architectures. While it enriches the source knowledge, Zoo-MSFDA risks being dominated by suboptimal/harmful models. To address this issue, we theoretically analyze the model selection problem in Zoo-MSFDA, and introduce two principles: transferability principle and diversity principle. Recognizing the challenge of measuring transferability, we subsequently propose a novel Source-Free Unsupervised Transferability Estimation (SUTE). It enables assessing and comparing transferability across multiple source models with different architectures under domain shift, without requiring target labels and source data. Based on above, we introduce a Selection, Ensemble, and Adaptation (SEA) framework to address Zoo-MSFDA, which consists of: 1) source models selection based on the proposed principles and SUTE; 2) ensemble construction based on SUTE-estimated transferability; 3) target-domain adaptation of the ensemble model. Evaluations demonstrate that our SEA framework, with the introduced Zoo-MSFDA setting, significantly improves adaptation performance in 2D image classification tasks. Additionally, our SUTE achieves state-of-the-art performance in transferability estimation. Jiangbo Pei, Aidong Men, Yang Liu 0105, Xiahai Zhuang, Qingchao Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Incorporating Pre-Training Data Matters in Unsupervised Domain AdaptationabstractIn deep learning, initializing models with pre-trained weights has become the de facto practice for various downstream tasks. Many unsupervised domain adaptation (UDA) methods typically adopt a backbone pre-trained on ImageNet, and focus on reducing the source-target domain discrepancy. However, the impact of pre-training on adaptation received little attention. In this study, we delve into UDA from the novel perspective of pre-training. We first demonstrate the impact of pre-training by analyzing the dynamic distribution discrepancies between pre-training data domain and the source/ target domain during adaptation. Then, we reveal that the target error also stems from the pre-training in the following two factors: 1) empirically, target error arises from the gradually degenerative pre-trained knowledge during adaptation; 2) theoretically, the error bound depends on difference between the gradient of loss function, i.e., on the target domain and pre-training data domain. To address these two issues, we redefine UDA as a three-domain problem, i.e., source domain, target domain, and pre-training data domain; then we propose a novel framework, named TriDA. We maintain the pre-trained knowledge and improve the error bound by incorporating pre-training data into adaptation for both vanilla UDA and source-free UDA scenarios. For efficiency, we introduce a selection strategy for pre-training data, and offer a solution with synthesized images when pre-training data is unavailable during adaptation. Notably, TriDA is effective even with a small amount of pre-training or synthesized images, and seamlessly complements the two scenario UDA methods, demonstrating state-of-the-art performance across multiple benchmarks. We hope our work provides new insights for better understanding and application of domain adaptation. Yinsong Xu 0002, Aidong Men, Yang Liu 0105, Xiahai Zhuang, Qingchao Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | PracticalDG: Perturbation Distillation on Vision-Language Models for Hybrid Domain GeneralizationabstractDomain Generalization (DG) aims to resolve distribution shifts between source and target domains, and current DG methods are default to the setting that data from source and target domains share identical categories. Nevertheless, there exists unseen classes from target domains in practical scenarios. To address this issue, Open Set Domain Generalization (OSDG) has emerged and several methods have been exclusively proposed. However, most existing methods adopt complex architectures with slight improvement compared with DG methods. Recently, vision-language models (VLMs) have been introduced in DG following the fine-tuning paradigm, but consume huge training overhead with large vision models. Therefore, in this paper, we innovate to transfer knowledge from VLMs to lightweight vision models and improve the robustness by introducing Perturbation Distillation (PD) from three perspectives, including Score, Class and Instance (SCI), named SCI-PD. Moreover, previous methods are oriented by the benchmarks with identical and fixed splits, ignoring the divergence between source domains. These methods are revealed to suffer from sharp performance decay with our proposed new benchmark Hybrid Domain Generalization (HDG) and a novel metric H2-CV, which construct various splits to comprehensively assess the robustness of algorithms. Extensive experiments demonstrate that our method outperforms state-of-the-art algorithms on multiple datasets, especially improving the robustness when confronting data scarcity. Zining Chen, Weiqiu Wang, Zhicheng Zhao 0001, Aidong Men, Hongying Meng |
CVPR | 5 |
| 2024 | Selective Cross-Correlation Consistency Loss for Out-of-Distribution GeneralizationabstractDeep learning methods usually succeed in independent and identically distributed (IID) data distribution, but suffer from sharp performance decay in real-world out-of-distribution (OOD) data. OOD generalization emerges to alleviate the large distribution shift between source and target domains. Recently, domain-invariant learning has boosted the research on OOD generalization, but most methods indulge complex architectures and training strategies. Hence, we propose a simple yet effective Selective Cross-Correlation Consistency (SC3) loss to align the cross-correlation matrix of features from identical categories. Specifically, we design the Semantic-Oriented Selection (SOS) algorithm in SC3loss to eliminate negative effects on spurious channels. Extensive experiments demonstrate that the SC3loss achieves superior performance on multiple OOD scenarios, including domain generalization (DG) and single domain gener-alization (SDG) tasks. Also, our loss consumes negligible computational resource which conforms to real-world applications. Source code is available at https://github.com/znchen666/SC3. Zining Chen, Weiqiu Wang, Zhicheng Zhao 0001, Aidong Men |
ICME | 5 |
| 2024 | Poisson Ordinal Network for Gleason Group Estimation Using Bi-Parametric MRI
Yinsong Xu 0002, Ziyi Shen, Iani J. M. B. Gayo, Natasha Thorley, Shonit Punwani, Aidong Men, Dean C. Barratt, Qingchao Chen, Yipeng Hu |
MICCAI (5) | 7 |
| 2024 | IoT-V2E: An Uncertainty-Aware Cross-Modal Hashing Retrieval Between Infrared-Videos and EEGs for Automated Sleep State AnalysisabstractEstimating and monitoring the sleep states at home using ubiquitous infrared (IR) visual camera sensors is an essential healthcare problem. Currently, the common challenge of using IoT sensors to predict sleep stages is the “semantic gap” between the IoT sensory signals and the medical signals, where fewer correlations between IoT sensory signals and the sleep stage labels are observed. To bridge this gap, we propose a novel systematic and methodological IoT design (IoT-V2E) to retrieve the most similar electroencephalogram signal representations in a database given an IR visual query for sleep-related analysis. Specifically, we make the following specific contributions: 1) we collect a cross-modal retrieval data set, including the IR sensory signals and the synchronized Polysomnography signals with sleep stage ground-truth annotations; 2) we propose a novel uncertainty-aware hashing retrieval method, presenting superior performances, sufficient interpretability, and high memory efficiency; 3) our method achieves the state-of-the-art sleep stage retrieval results and provides the uncertainty for each query in the inference; and 4) most importantly, our system is evaluated to be able to assist the physicians not only in diagnosing sleep-related diseases but also finding the subjects with the most similar sleep patterns. Our project is available athttps://github.com/SPIresearch/IoT-V2E. Aidong Men, Yang Liu 0105, Ziming Yao, Shaoxing Zhang, Qingchao Chen |
IEEE Internet Things J. | 2 |
| 2024 | Pedestrian Navigation Activity Recognition Based on Segmentation TransformerabstractIn the context of the Internet of Things, utilizing the inherent inertial sensors in smartphones for human activity recognition (HAR) has garnered considerable attention owing to its wide-ranging applications. However, prevailing HAR approaches primarily treat activity identification as a single-label classification task, focusing solely on discerning pedestrian motion modes or device usage modes, while disregarding their interrelatedness. Additionally, HAR methods employing sliding windows encounter challenges associated with the multiclass window problem, wherein certain sample labels differ from the label assigned to the window. This paper aims to address these issues. This paper presents a novel approach for simultaneously recognizing pedestrian motion and device usage modes by utilizing the segmentation transformer. The proposed joint recognition framework effectively annotates sensor data at each timestamp and achieves dense prediction of time-series data through the encoding and decoding of the annotated data. To optimize the utilization of information extracted from each Transformer layer, a global up-sampling decoder based on the pyramid attention module is introduced, enabling dense decoding of features obtained from each Transformer layer. We performed experiments on two publicly available datasets to comprehensively assess the effectiveness of the proposed methodology. The results demonstrate that our approach achieves an accuracy of 99.79% and a weighted F-score of 99.77%, surpassing the performance of existing state-of-the-art methods. Furthermore, we constructed heterogeneous datasets to validate the robustness of our method. The extensive experimental findings indicate that the joint recognition framework effectively uncovers the inherent correlations between pedestrian motion and device usage modes, leading to enhanced accuracy in recognition and addressing the challenges posed by the multiclass window problem. Qu Wang, Jiahui Ning, Zhuqing Jiang, Liangliang Guo, Haiyong Luo, Haiying Wang 0005, Aidong Men, Xiaofei Cheng |
IEEE Internet Things J. | 8 |
| 2024 | Evidential Multi-Source-Free Unsupervised Domain AdaptationabstractMulti-Source-Free Unsupervised Domain Adaptation (MSFUDA) requires aggregating knowledge from multiple source models and adapting it to the target domain. Two challenges remain: 1) suboptimal coarse-grained (domain-level) aggregation of multiple source models, and 2) risky semantics propagation based on local structures. In this article, we propose an evidential learning method for MSFUDA, where we formulate two uncertainties, i.e. Evidential Prediction Uncertainty (EPU) and Evidential Adjacency-Consistent Uncertainty (EAU), respectively for addressing the two challenges. The former, EPU, captures the uncertainty of a sample fitted to a source model, which can suggest the preferences of target samples for different source models. Based on this, we develop an EPU-Based Multi-Source Aggregation module to achieve fine-grained, instance-level source knowledge aggregation. The latter, EAU, provides a robust measure of consistency among adjacent samples in the target domain. Utilizing this, we develop an EAU-Guided Local Structure Mining module to ensure the trustworthy propagation of semantics. The two modules are integrated into the Evidential Aggregation and Adaptation Framework (EAAF), and we demonstrated that this framework achieves state-of-the-art performances on three MSFUDA benchmarks. Jiangbo Pei, Aidong Men, Yang Liu 0105, Xiahai Zhuang, Qingchao Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Instance Paradigm Contrastive Learning for Domain GeneralizationabstractDomain Generalization (DG) aims to develop models that can learn from data in source domains and generalize to unseen target domains. Recently, some domain generalization algorithms have emerged, but most of them were designed with complex modules. Among all the prior methods under DG settings, contrastive learning has become a promising solution for simplicity and efficiency. However, existing contrastive learning neglects distribution shifts that causes severe domain confusions. In this paper, we propose an instance paradigm contrastive learning framework, introducing contrast between original features and novel paradigms to alleviate domain-specific distractions. And then we explore hard-pair information, an essential factor in contrastive learning, based on domain label and feature similarity. Moreover, to produce domain-invariant instance paradigms, we generate multiple views of the original images and design a novel channel-wise attention mechanism to dynamically combine features from all the views. Furthermore, a test-time feature integration module is designed to mimic the paradigms during the training process to improve generalization ability. Extensive experiments show that our method achieves state-of-the-art performance. The proposed algorithm can also serve as a plug-and-play module which improves performance of existing methods with a relatively large margin. Zining Chen, Weiqiu Wang, Zhicheng Zhao 0001, Aidong Men |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | EviPrompt: A Training-Free Evidential Prompt Generation Method for Adapting Segment Anything Model in Medical ImagesabstractMedical image segmentation is a critical task in clinical applications. Recently, the Segment Anything Model (SAM) has demonstrated potential for natural image segmentation. However, the requirement for expert labour to provide prompts, and the domain gap between natural and medical images pose significant obstacles in adapting SAM to medical images. To overcome these challenges, this paper introduces a novel prompt generation method named EviPrompt. The proposed method requires only a single reference image-annotation pair, making it a training-free solution that significantly reduces the need for extensive labelling and computational resources. First, prompts are automatically generated based on the similarity between features of the reference and target images, and evidential learning is introduced to improve reliability. Then, to mitigate the impact of the domain gap, committee voting and inference-guided in-context learning are employed, generating prompts primarily based on human prior knowledge and reducing reliance on extracted semantic information. EviPrompt represents an efficient and robust approach to medical image segmentation. We evaluate it across a broad range of tasks and modalities, confirming its efficacy. The source code is available at https://github.com/SPIresearch/EviPrompt. Yinsong Xu 0002, Jiaqi Tang 0012, Aidong Men, Qingchao Chen |
IEEE Trans. Image Process. | 3 |
| 2024 | Cluster-Instance Normalization: A Statistical Relation-Aware Normalization for Generalizable Person Re-IdentificationabstractPerson re-identification (ReID) has achieved great improvement under supervised settings, but suffers from considerable degradation when large distribution shifts between training and testing sets exist. Domain generalization (DG ReID) emerges to promote the generalization ability of models, overcoming the distribution shifts issue between source domains and unseen target domains. Among most prior methods in DG ReID, instance normalization (IN) serves as a promising solution for removing domain-specific information, however, it damages the discriminative ability simultaneously. In this article, we propose a new normalization method called Cluster-Instance Normalization (CINorm) to extract information from clusters for information compensation. The relations between samples in a batch can be mined to establish evolving clusters with aggregated samples during the forward training process. In this way, high intra-cluster congregation can eliminate the impacts of outliers to avoid overfitting, and high inter-cluster variances can synthesize diverse novel statistics to compensate discriminative information. Therefore, a Relation-Aware Normalization (RANorm) with a Dynamic ReCalibration (DRC) module is designed to integrate normalized features between evolving clusters and instances efficiently. Furthermore, a novel Group-based Triplet (G-Triplet) loss is proposed to divide a batch into multiple groups with greater compactness for hard-pair mining. Extensive experiments show that our method outperforms state-of-the-art algorithms on multiple DG benchmarks by a large margin. The proposed method can also achieve superior performance on image classification tasks under DG settings without using domain labels. Zining Chen, Weiqiu Wang, Zhicheng Zhao 0001, Aidong Men |
IEEE Trans. Multim. | 5 |
| 2023 | Multiple Prompt Fusion for Zero-Shot Lesion Detection Using Vision-Language Models
Miaotian Guo, Huahui Yi, Ziyuan Qin 0001, Haiying Wang 0005, Aidong Men, Qicheng Lao |
MICCAI (5) | 5 |
| 2023 | Uncertainty-Induced Transferability Representation for Source-Free Unsupervised Domain AdaptationabstractSource-free unsupervised domain adaptation (SFUDA) aims to learn a target domain model using unlabeled target data and the knowledge of a well-trained source domain model. Most previous SFUDA works focus on inferring semantics of target data based on the source knowledge. Without measuring the transferability of the source knowledge, these methods insufficiently exploit the source knowledge, and fail to identify the reliability of the inferred target semantics. However, existing transferability measurements require either source data or target labels, which are infeasible in SFUDA. To this end, firstly, we propose a novel Uncertainty-induced Transferability Representation (UTR), which leverages uncertainty as the tool to analyse the channel-wise transferability of the source encoder in the absence of the source data and target labels. The domain-level UTR unravels how transferable the encoder channels are to the target domain and the instance-level UTR characterizes the reliability of the inferred target semantics. Secondly, based on the UTR, we propose a novel Calibrated Adaption Framework (CAF) for SFUDA, including i) the source knowledge calibration module that guides the target model to learn the transferable source knowledge and discard the non-transferable one, and ii) the target semantics calibration module that calibrates the unreliable semantics. With the help of the calibrated source knowledge and the target semantics, the model adapts to the target domain safely and ultimately better. We verified the effectiveness of our method using experimental results and demonstrated that the proposed method achieves state-of-the-art performances on the three SFUDA benchmarks. Code is available at https://github.com/SPIresearch/UTR. Jiangbo Pei, Zhuqing Jiang, Aidong Men, Yang Liu 0105, Qingchao Chen |
IEEE Trans. Image Process. | 3 |
| 2022 | An Efficient Method for Model Pruning Using Knowledge Distillation with Few SamplesabstractDeep neural network compression methods can produce small-scale networks and utilizes fine-tuning to get back the dropped accuracy. Despite their remarkable performance, the fine-tuning procedure is limited to the requirement of a huge training dataset, which is a time-consuming progress. To address the issue, few-sample knowledge distillation (FSKD) has been proposed for data efficiency. However, FSKD needs to add additional convolution layers for compressed networks during training, which increases the complexity of network structure. In this paper, we present Progressive Feature Distribution Distillation (PFDD) without modifying network structures, which surpasses FSKD. Concretely, it is based on a progressive training strategy that is efficient for matching feature distributions between compressed network and original network. Thus, we can notably exploit both external information from samples and internal information from network, where using a small proportion of training dataset can yield quite considerable results. Experiments on various datasets and architectures demonstrate that our distillation approach is remarkably efficient and effective in improving compressed networks’ performance while only few samples have been applied. ZhaoJing Zhou, Zhuqing Jiang, Aidong Men, Haiying Wang 0005 |
ICASSP | 4 |
| 2022 | Mixed In Time And Modality: Curse Or Blessingƒ Cross-Instance Data Augmentation for Weakly Supervised Multimodal Temporal FusionabstractIn multimodal video event localization, we usually leverage feature fusion across different axes, such as the modality and temporal axes, for better context. To reduce the costs of detailed annotations, recent solutions explore weakly supervised settings. However, we observe that when feature fusion meets weakly supervised localization, problems can occur. It may cause "feature cross-interference", which produces a smearing effect on the localization result and can’t be effectively supervised with conventional multiple instance learning loss. We verify it quantitatively on the audio-visual video parsing (AVVP) task, and propose a cross-instance data-augmentation framework, which can preserve the benefits of feature fusion while providing explicit feedbacks for feature cross-interference. We show that our method can enhance performance of existing models on two weakly supervised audio-visual localization tasks, i.e. AVVP and AVE. Yonggang Zhu, Zhuqing Jiang, Aidong Men, Haiying Wang 0005, Qingchao Chen |
ICASSP | 4 |
| 2022 | Delving into the Continuous Domain AdaptationabstractExisting domain adaptation methods assume that domain discrepancies are caused by a few discrete attributes and variations, e.g., art, real, painting, quickdraw, etc. We argue that this is not realistic as it is implausible to define the real-world datasets using a few discrete attributes. Therefore, we propose to investigate a new problem namely the Continuous Domain Adaptation (CDA) through the lens where infinite domains are formed by continuously varying attributes. Leveraging knowledge of two labeled source domains and several observed unlabeled target domains data, the objective of CDA is to learn a generalized model for whole data distribution with the continuous attribute. Besides the contributions of formulating a new problem, we also propose a novel approach as a strong CDA baseline. To be specific, firstly we propose a novel alternating training strategy to reduce discrepancies among multiple domains meanwhile generalize to unseen target domains. Secondly, we propose a continuity constraint when estimating the cross-domain divergence measurement. Finally, to decouple the discrepancy from the mini-batch size, we design a domain-specific queue to maintain the global view of the source domain that further boosts the adaptation performances. Our method is proven to achieve the state-of-the-art in CDA problem using extensive experiments. The code is available at https://github.com/SPIresearch/CDA. Yinsong Xu 0002, Zhuqing Jiang, Aidong Men, Yang Liu 0105, Qingchao Chen |
ACM Multimedia | 3 |
| 2022 | Taylor saves for later: Disentanglement for video prediction using Taylor representation
Zhuqing Jiang, Shiping Wen 0001, Aidong Men, Haiying Wang 0005 |
Neurocomputing | 5 |
| 2022 | Toward a perceptive pretraining framework for Audio-Visual Video Parsing
Jianning Wu, Zhuqing Jiang, Qingchao Chen, Shiping Wen 0001, Aidong Men, Haiying Wang 0005 |
Inf. Sci. | 5 |
| 2022 | Shedding light on images: Multi-level image brightness enhancement guided by arbitrary referencesabstractThe non-linearity between human perception and image brightness levels results in different definitions of NORMAL-light. Thus, most existing low-light image enhancement methods which produce one-to-one mapping can not meet the aesthetic demand. Other pioneers enhance low-light images guided by a given value. However, the inherent problem of non-linearity will cause poor usability. To this end, we propose a user-friendly neural network for multi-level low-light image enhancement. Inspired by style transfer, our method decomposes an image into content component feature and luminance component feature in the latent space. Then we enhance the image brightness to different levels by concatenating the content components from low-light images and the luminance components from reference images. The network meets various user requirements by selecting different brightness references. Moreover, information except for brightness is preserved to alleviate color distortion. Extensive experiments demonstrate the superiority of our network against existing methods. Zhuqing Jiang, Aidong Men, Haiying Wang 0005 |
Pattern Recognit. | 5 |
| 2021 | Integration-and-Diffusion Network for Low-Light Image EnhancementabstractImages captured under extreme low-light conditions often suffer from low Signal-to-Noise Ratio(SNR) caused by low photon count, making low-light image enhancement challenging. Deep learning-based methods have recently yielded impressive progress by reconstructing extreme low-light images from raw sensor data. Despite their promising results, they still fail at recovering detailed textures and corresponding colors. To address these issues, we propose an Information Integration-and-Diffusion (InD) module to reconstruct excellent details from extreme low-light raw images. Precisely, a pixel-intensive global information matrix is calculated by separately integrating spatial-wise and channel-wise information and then diffusing them to each other by a matrix multiplication operation. In addition to this, we propose a Bottleneck Guided Channel Attention (BGCA) module to achieve unified channel information through low-light image enhancement networks for better color recovery. Extensive experimental results show that the networks equipped with our proposed modules outperform state-of-the-art approaches both quantitatively and visually. Pengliang Tang, Guodong Ju, Liangheng Shen, Aidong Men |
ICIP | 5 |
| 2021 | Image Brightness Adjustment with Unpaired Training
Chaojian Liu, Aidong Men |
ICONIP (4) | 3 |
| 2021 | Multi-DIP: A General Framework for Unsupervised Multi-degraded Image Restoration
Qiansong Wang, Haiying Wang 0005, Aidong Men, Zhuqing Jiang |
ICONIP (4) | 4 |
| 2021 | A switched view of Retinex: Deep self-regularized low-light image enhancementabstractSelf-regularized low-light image enhancement does not require any normal-light image in training, thereby freeing from the chains of paired or unpaired training data that are time-consuming to obtain. However, existing methods suffer color deviation and fail to generalize to various lighting conditions. This paper presents a novel self-regularized method based on Retinex, which, inspired by HSV, preserves all colors (Hue, Saturation) and only integrates Retinex theory into brightness (Value). Besides, we design a novel random brightness disturbance approach to generate another abnormal brightness of the same scene. It is combined with the original form of brightness to estimate the same reflectance, which is achieved by a CNN. The reflectance, which is assumed irrelevant to any illumination according to the Retinex theory, is treated as the enhanced brightness. Our method is efficient as a low-light image is decoupled into two subspaces, i.e., color and brightness, for better preservation and enhancement. Extensive experiments demonstrate that our method outperforms multiple state-of-the-art algorithms qualitatively and quantitatively and adapts to more lighting conditions. Our code is available at https://github.com/Github-LHT/A-Switched-View-of-Retinex-Deep-Self-Regularized-Low-Light-Image-Enhancement. Zhuqing Jiang, Liangjie Liu, Aidong Men, Haiying Wang 0005 |
Neurocomputing | 4 |
| 2021 | Pedestrian Dead Reckoning Based on Walking Pattern Recognition and Online Magnetic Fingerprint Trajectory CalibrationabstractWith the explosive development of pervasive computing and the Internet of Things (IoT), indoor positioning and navigation have attracted immense attention over recent years. Pedestrian dead reckoning (PDR) is a potential autonomous localization technology that obtains position estimation employing built-in sensors. However, most existing PDR methods assume that the smartphone is held horizontally and points to the walking direction. To solve reckoning errors caused by inconsistency of headings between walking heading and pointing of smartphone, we design an accurate and robust PDR method based on walking patterns, which is identified by multihead convolutional neural networks. In addition to adaptively adjust the threshold of step detection and select the most suitable step length model according to the results of walking pattern recognition, a novel heading estimation approach independent of device orientation is proposed. To mitigate accumulative errors, we proposed an online trajectory calibration method based on forward and backward magnetic fingerprint trajectory matching. We conduct extensive and well-designed experiments in typical scenarios, and the experimental results indicate that the 75th percentile localization accuracy of the three scenarios is 1.06, 1.08, and 1.22 m, respectively, using the commercial smartphone embedded sensor without any dedicated infrastructures or training data. Despite the intricate pedestrian locomotion, the proposed PDR method has great potential in pedestrian positioning. Qu Wang, Haiyong Luo, Aidong Men, Fang Zhao 0003, Ming Xia 0009, Changhai Ou |
IEEE Internet Things J. | 4 |
| 2021 | Multi-view feature fusion for person re-identificationabstractPerson re-identification (ReID) suffers from camera view variants. Existing works, which typically learn a feature for each image, share a limitation that the learned features are single-view: each feature only contains information in one camera view. Thus, view bias occurs when matching pedestrians across camera views. In this paper, we seek to mitigate the view bias by generating multi-view features (fusion of features from a fixed number of cameras). To this end, we define the complementary-view features (complementary features to generate multi-view features with single-view features) and perform in-depth analysis. Based on this insight, we alleviate the view bias in testing and training, respectively. In testing, we present Multi-view Message Passing (MVMP), which generates multi-view features by aggregating single-view features from the neighborhood. In training, we propose Multi-view Feature Fusion Network (MFFN), which involves the single-view feature extractor and the complementary-view feature aggregator. MFFN makes the network sensitive to view-specific cues by adding constraints on multi-view features rather than single-view features. In addition, MVMP and MFFN have two key advantages: (1) They are parameter-free. (2) They can be applied to any Convolutional Neural Networks (CNNs) readily without extra supervision. Extensive experiments are conducted to validate the superiority of our method for person ReID over state-of-the-art methods on four benchmark datasets (Market-1501, DukeMTMC-reID, CUHK03, and MSMT17). The code is available at https://github.com/Yinsongxu/MVMP_MFFN. Yinsong Xu 0002, Zhuqing Jiang, Aidong Men, Haiying Wang 0005, Haiyong Luo |
Knowl. Based Syst. | 3 |
| 2020 | Attention-Enhanced And More Balanced R-CNN For Object DetectionabstractAttention mechanisms have been widely used in deep neural convolution networks and different fields, such as object detection and instance segmentation. Many attention mechanisms will cost too much calculation, so in this paper, we incorporate a kind of light attention mechanism, the attended residual module, into our object detection backbone to get an accuracy-efficiency trade-off. Besides, to solve the imbalance problem in region sample level, we use the cascade region proposal network (RPN) module to gain anchors of higher quality resulting in higher average recall (AR). Furthermore, we replace the non-local attention module in feature fusion level with the criss-cross attention module to reduce computation and improve performance. With them all, our method significantly improves the detection performance and achieves 43.6 AP in COCO test-dev. Ruohong Mei, Haiying Wang 0005, Aidong Men |
ICIP | 3 |
| 2020 | Split to Be Slim: An Overlooked Redundancy in Vanilla ConvolutionabstractMany effective solutions have been proposed to reduce the redundancy of models for inference acceleration. Nevertheless, common approaches mostly focus on eliminating less important filters or constructing efficient operations, while ignoring the pattern redundancy in feature maps. We reveal that many feature maps within a layer share similar but not identical patterns. However, it is difficult to identify if features with similar patterns are redundant or contain essential details. Therefore, instead of directly removing uncertain redundant features, we propose a split based convolutional operation, namely SPConv, to tolerate features with similar patterns but require less computation. Specifically, we split input feature maps into the representative part and the uncertain redundant part, where intrinsic information is extracted from the representative part through relatively heavy computation while tiny hidden details in the uncertain redundant part are processed with some light-weight operation. To recalibrate and fuse these two groups of processed features, we propose a parameters-free feature fusion module. Moreover, our SPConv is formulated to replace the vanilla convolution in a plug-and-play way. Without any bells and whistles, experimental results on benchmarks demonstrate SPConv-equipped networks consistently outperform state-of-the-art baselines in both accuracy and inference time on GPU, with FLOPs and parameters dropped sharply. Qiulin Zhang, Zhuqing Jiang, Qishuo Lu, Zhengxin Zeng, Shanghua Gao, Aidong Men |
IJCAI | 7 |
| 2020 | Personalized Stride-Length Estimation Based on Active Online LearningabstractThe ability to accurately estimate a user's stride length plays a great important role in various applications. For a new target pedestrian or device, their heterogeneity dramatically reduces the performance of the current stride-length estimation (SLE) methods. To address the issue of heterogeneity, in this article, we propose an SLE method based on a long short-term memory (LSTM) network and denoising autoencoders (DAEs). The LSTM network is used to mine temporal dependencies and extract significant eigenvectors from the corrupted inertial sensor observations. Then, DAEs are adopted to automatically eliminate the inherent noise in eigenvectors and obtain denoised eigenvectors. Finally, a regression module maps the denoised eigenvectors to the resulting stride length. To mitigate the heterogeneity, we propose an unperceived model updating framework based on active online learning to establish a personalized model for a given target pedestrian or device. The proposed framework utilizes a magnetism-aided map-matching approach to automatically generate personalized training data and utilizes online learning technologies to evolve the stride-length model. The extensive experimental results demonstrate that the proposed method outperforms other state-of-the-art algorithms and achieves a promising accuracy with a stride-length error rate of 4.59% at a confidence level of 80%. Qu Wang, Haiyong Luo, Langlang Ye, Aidong Men, Fang Zhao 0003, Yan Huang 0035, Changhai Ou |
IEEE Internet Things J. | 4 |
| 2019 | Recursive Multi-Stage Upscaling Network with Discriminative Fusion for Super-ResolutionabstractSince convolutional neural networks have fundamentally changed how computers learn features, many super-resolution (SR) methods focus on extracting informative features to improve performance by using more layers or more innovative skip-connections. However, using more layers in feature extraction while keeping the upscaling module unchanged will exacerbate structural imbalances. Besides, blindly fusing different stage features by concatenation may make them interfere with each other. To address these issues, we proposed a recursive multi-stage upscaling network (RMUN) with discriminative fusion module (DFM). Specifically, we construct multiple upscaling paths to produce various high-resolution features in the forward propagation and deliver error loss in the back propagation. Furthermore, we fuse and re-weight those features by DFM to avoid mutual interference and boost reconstruction quality. Experiments show that RMUN is superior to the state-of-the-art methods, especially for large scale SR tasks. Zhuqing Jiang, Guodong Ju, Liangheng Shen, Aidong Men |
ICME | 5 |
| 2019 | Attentional Part-based Network for Person Re-identificationabstractPart-based network is an effective method to improve performance in person re-identification (re-ID). Most existing methods assume the availability of well-aligned person bounding box images as model input. However, automatic detection in some datasets causes misalignment which negatively affects the performance. In this work, we propose an Attentional Part-based CNN (AP-CNN) model which combines learning partial features and attention selection. First, we partition feature map into several horizontal stripes. Second, we use attention selection in each stripe to align the pedestrian images. Inside, we introduce a free-parameter attention model with skip-layer connection which maximizes the complementary information of different levels without increasing the complexity of network. Results on four datasets validate the competitiveness of AP-CNN over the state-of-the-art achieving Rank-1 accuracy of 94.4% on Market-1501, 87.3% on DukeMTMC-ReID, 73.7% on CUHK03-labeled and 72.6% on CUHK03-detected. Yinsong Xu 0002, Zhuqing Jiang, Aidong Men, Jiangbo Pei, Guodong Ju, Bo Yang 0007 |
VCIP | 3 |
| 2019 | Pyramid Real Image Denoising NetworkabstractWhile deep Convolutional Neural Networks (CNNs) have shown extraordinary capability of modelling specific noise and denoising, they still perform poorly on real-world noisy images. The main reason is that the real-world noise is more sophisticated and diverse. To tackle the issue of blind denoising, in this paper, we propose a novel pyramid real image denoising network (PRIDNet), which contains three stages. First, the noise estimation stage uses channel attention mechanism to recalibrate the channel importance of input noise. Second, at the multi-scale denoising stage, pyramid pooling is utilized to extract multi-scale features. Third, the stage of feature fusion adopts a kernel selecting operation to adaptively fuse multi-scale features. Experiments on two datasets of real noisy photographs demonstrate that our approach can achieve competitive performance in comparison with state-of-the-art denoisers in terms of both quantitative measure and visual perception quality. Yiyun Zhao, Zhuqing Jiang, Aidong Men, Guodong Ju |
VCIP | 3 |
| 2019 | Object detection using convolutional networks with adaptively adjusting receptive field of convolutional filterabstractThe receptive field size of a convolutional filter in a deep convolutional network is a crucial issue for object detection task, as the output must response to a suitable size of area in the image to capture proper information. Receptive field size of convolutional filter is fixed due to the inherently fixed geometric structure in its building module. However, objects of interest vary significantly in size within the images for object detection. Different locations of images correspond to objects with different scales, and high level convolutional layers encode semantic features over spatial positions, thus adaptive determination of receptive field size of convolutional filter is desirable for object detection. The authors propose a new module to adaptively determine the receptive field size of convolutional filter, named adaptive convolution. It is based on the idea of dilating the convolutional filter with multiple dilation values and choosing the maximum activation as output, without adding any other parameters. The plain counterparts in existing convolutional neural networks can be easily replaced by adaptive convolution, giving rise to adaptive convolutional networks. Adequate experiments have proven the effectiveness of authors’ method. Qishuo Lu, Zhuqing Jiang, Aidong Men, Pengliang Tang |
IET Comput. Vis. | 3 |
| 2018 | Deep Network with Spatial and Channel Attention for Person Re-identificationabstractMost existing person re-identification (Re-ID) methods assume pedestrian images are well-aligned within tightly surrounded bounding boxes or require additional annotation information to calibrate misaligned images. In this work, we propose a novel deep network to address the misalignment problem in person re-identification task without requiring additional annotation. Spatial attention selection mechanism is introduced in our network to align the pedestrian images. Moreover, we present a channel attention selection mechanism to integrate the global image feature maps and regional feature maps more effectively by explicitly modelling interdependencies between channels and recalibrates feature response in each channel. Extensive experiments and comparative evaluations demonstrate the effectiveness of our approach and the superiority of this novel network for person re-identification over a wide variety of state-of-the-art methods on two datasets including Market-1501 and CUHK03(both detected and labeled sets). Tiansheng Guo, Dongfei Wang, Zhuqing Jiang, Aidong Men |
VCIP | 4 |
| 2017 | Real-time object detection by a multi-feature fully convolutional networkabstractPrior work on object detection depends on region proposals to guide the search for object instances. Generally, several thousand proposals must be processed, thus hurting the detection efficiency. In this paper, we propose a new model free from region proposals for object detection which treats detection task as a regression problem. To improve small-size object detection and localization, we employ the deep hierarchical features extracted from convolutional neural networks (CNNs). The hierarchical architecture combines appearance information from a shallow layer with semantic information from a deep layer. Our approach can predict bounding boxes and class probabilities simultaneously from a full input image. We transfer a classification network called Darknet into fully convolutional network and fine-tune it for the detection task. Experiments on PASCAL VOC dataset demonstrate that our approach outperforms other detection models. Yajing Guo, Zhuqing Jiang, Aidong Men |
ICIP | 4 |
| 2017 | Coupled analysis-synthesis dictionary learning for person re-identificationabstractIn this paper, we propose a novel coupled dictionary learning method, namely coupled analysis-synthesis dictionary learning, to improve the performance of person re-identification in the non-overlapping fields of different camera views. Most of the existing coupled dictionary learning methods train a coupled synthesis dictionary directly on the original feature spaces, which limits the representation ability of the dictionary. To handle the diversities of different original spaces, We first employ local Fisher discriminant analysis (LFDA) to learn a common feature space for close relationship of the same people in different views. In order to enhance the representation power of the coupled synthesis dictionary, we then learn a coupled analysis dictionary by transforming the common feature space into the coupled feature space. Experimental results on two publicly available VIPeR and CUHK01 datasets have validated the effectiveness of the proposed method. Lingchuan Sun, Zhuqing Jiang, Aidong Men |
ICIP | 4 |
| 2017 | Improve object detection via a multi-feature and multi-task CNN modelabstractCurrent state-of-the-art object detection methods have made dramatic performance improvements in recent few years. However, there are still several challenges. In particular, it still struggles for precise localization of small-sized objects, mainly due to coarse resolutions of feature maps and excessive surroundings such as ground and water. To address the issues, we propose an object detection system based on standard Fast R-CNN object detection branch and DeepLap semantic segmentation branch: (1) multi-feature aggregates hierarchical features for more finer feature maps to detect objects at multiple scales. (2) multi-task uses semantic segmentation for more contextual information to assist object detection via a cross structure between the two tasks. (3) a novel overlap loss function is used for bounding box regression that adjusts region proposals to improve localization. The fusion network improves results over the Fast R-CNN baseline detector by 2.8% mAP and by 4.8% mAP for small objects based on PASCAL VOC datasets. Yingxin Lou, Guangtao Fu, Zhuqing Jiang, Aidong Men |
VCIP | 4 |
| 2015 | No-reference image quality assessment based on phase congruency and spectral entropiesabstractWe develop an efficient general-purpose blind/no-reference image quality assessment (IQA) algorithm that utilizes curvelet domain features of phase congruency values and local spectral entropy values in distorted images. A 2-stage framework of distortion classification followed by quality assessment is used for mapping feature vectors to prediction scores. We utilize a support vector machine (SVM) to train an image distortion and quality prediction model. The resulting algorithm which we name Phase Congruency and Spectral Entropy based Quality (PCSEQ) index is capable of assessing the quality of distorted images across multiple distortion categories. We explain the advantages of phase congruency features and spectral entropy features. We also thoroughly test the algorithm on the LIVE IQA databse and find that PCSEQ correlates well with human judgments of quality. It is superior to the full-reference (FR) IQA algorithm SSIM and several top-performance no-reference (NR) IQA methods such as DIIVINE and SSEQ. We also tested PCSEQ on the TID2008 database to ascertain whether it has performance that is database independent. Maozheng Zhao, Qin Tu, Yongyu Chang, Bo Yang 0007, Aidong Men |
PCS | 6 |
| 2015 | Visual saliency detection based on mutual information in compressed domainabstractSaliency detection of video sequences has attracted great attention in recent years. In this paper, we propose a saliency detection framework in the compressed domain. Spatial and temporal saliency map are derived by calculating mutual information of feature distribution between center window and surrounding window respectively. Motion vector is used to get the temporal saliency map, and features extracted from the discrete cosine transformation coefficients including luminance, color and texture are used to obtain the spatial saliency map. Then we combine the spatial and temporal saliency maps to obtain a spatiotemporal saliency map. Finally, a convex-hull-based center bias is added to optimize the saliency maps. Experimental results show that the proposed method outperforms the existing state-of-the-art saliency detection methods. Qin Tu, Aidong Men |
VCIP | 6 |
| 2015 | Gradient magnitude similarity for tone-mapped image quality assessmentabstractRecently, high dynamic range (HDR) image which can accurately reflect the real scene prompts increasing interest. For visualizing the HDR images on standard display devices, it is needed to convert it to low dynamic range (LDR) images. This process is named as tone-mapping. Due to the reduction of the dynamic range, evaluating the tone-mapped image becomes important. Gradient magnitude measure has been used in many state-of-the-art image quality assessment methods. In our research we find that the gradient magnitude measure is also effective for the quality assessment of tone-mapped images. The more different between the gradient magnitude maps of the HDR image and the corresponding LDR image, the worse quality of the LDR image. This paper using the gradient magnitude feature develops a gradient magnitude based method for tone-mapped image quality assessment. The similarity of gradient magnitude maps between the HDR image and the corresponding LDR image is computed. A tone-mapped image should look natural. The naturalness measure used in tone-mapped image quality index (TMQI) is as the supplement measure. In the experiment on the subject-rated tone-mapped image database provided by the authors of TMQI, our proposed method gets better performance than state-of-the-art tone-mapped image assessment methods. Qin Tu, Maozheng Zhao, Aidong Men, Bo Yang 0007 |
VCIP | 5 |
| 2015 | Video saliency detection incorporating temporal information in compressed domain
Qin Tu, Aidong Men, Zhuqing Jiang |
Signal Process. Image Commun. | 2 |
| 2014 | An adaptive CU mode decision mechanism based on Bayesian decision theory for H.265/HEVCabstractH.265/HEVC aims to provide significant improvement of compression performance compared with previous coding standards at cost of significant computation complexity. In this paper, an adaptive CU mode decision mechanism (ACMD) based on Bayesian decision theory is proposed to accelerate mode selection procedure. Specifically, the homogeneous determination is firstly utilized to filter out non-split LCU and then the feature space related to CU mode decision is introduced, which is divided into two regions according to Bayesian risk. Bayesian classifier is employed in low-risk region while in high-risk region, the mode decision is made by rate distortion cost. Experimental results demonstrate that the proposed algorithm provides averagely 34.28% encoding time reduction while maintaining the same level of perceptual visual quality, compared with HEVC test mode (HM 10.0) encoder with low-delay configurations. Qin Tu, Jingfeng Feng, Aidong Men |
ICME | 4 |
| 2014 | Temporarily static object detection in surveillance video using double foregrounds and superpixelsabstractIn surveillance video applications, temporarily static regions indicate move-then-stop objects, such as the abandoned/removed objects, parked vehicles. This paper presents an approach using both pixel-level and region-level analysis together. In the pixel-level foreground extraction process, two improved Gaussian Mixture Model(GMM) are adopted to obtain the binary foreground mask. A residence map is also set to measure the lasting time for an abandoned/removed object. In the region-level analysis, we apply a superpixel-based method, which could refine the foreground extraction result and further generate the output of exact object regions instead of only bounding boxes in other related works. Experimental results show the proposed method could detect the temporarily static objects effectively and accurately. Aidong Men, Yuanyuan Cui, Bo Yang 0007 |
VCIP | 2 |
| 2014 | An Adaptive CU Depth Selection Mechanism Based on Visual Sensitivity for HEVC Inter CodingabstractHigh Efficiency Video Coding (HEVC) is the latest video coding standard. It aims to provide significant improvement of compression performance with considerable increment of computational complexity, compared with all exiting video coding standards. In this paper, we propose an adaptive CU depth selection mechanism (ACDSM) for inter mode decision. To avoid unnecessary CU partition by rate distortion optimization with Lagrange multiplier, an adaptive CU depth is derived from spatial homogeneity based on visual sensitivity, intensity gradient filter and the temporal stationarity characteristics of video objects, in ACDSM mechanism. Experimental results demonstrate that the proposed algorithm provides averagely 32.93% encoding time reduction while maintaining the same level of perceptual visual quality, compared with HEVC test mode (HM 10.0) encoder with low-delay configurations. Qin Tu, Aidong Men, Bo Yang 0007 |
VTC Spring | 3 |
| 2014 | A Frame-Level HEVC Rate Control Algorithm for Videos with Complex Scene over Wireless NetworkabstractIn this paper, we propose a novel frame-level rate control algorithm for videos with complex scene in High Efficiency Video Coding(HEVC). To achieve a constant bitrate output and reduce the mismatch ratio of bitrate for time varying wireless links, the relationship of content complexity, QP and bitrate is utilized to adjust bit allocation for each GOP and frame via linear extrapolation. In addition, hierarchical frame structure is also considered to optimize QP selection. Experimental results show that compared with the latest rate control algorithm in JCTVC-K0103, our proposed algorithm can significantly improve the bitrate accurate at comparable coding performance in terms of constant quality and average PSNR, especially for videos with complex scene. Qin Tu, Aidong Men |
VTC Spring | 3 |
| 2014 | A Novel Routing Algorithm Design of Time Evolving Graph Based on Pairing Heap for MEO Satellite NetworkabstractAs a tradeoff of GEO and LEO, MEO satellite system has more acceptable service performance and it is more appropriate to provide global mobile communications. A MEO satellite system model communicating according to time slots is constructed in the paper. Moreover, in order to improve comprehensive performance of the network, a novel routing algorithm applying Time Evolving Graph based on Pairing Heap is proposed. The Time Evolving Graph is employed to analysis the dynamic topology of the network and the Pairing Heap is applied in the Dijkstra algorithm to reduce the time complexity. By contrast, Fibonacci Heap is also used to optimize Dijkstra algorithm. Finally simulation results show that routing algorithm applying Time Evolving Graph based on Pairing Heap can perform better and reduce the time complexity obviously, and at the same time, Pairing Heap works better than Fibonacci Heap when the number of nodes grows bigger. Yupeng Wang 0002, Zhuqing Jiang, Chengkai Huang, Aidong Men, Bo Yang 0007, Kaifeng Qi |
VTC Fall | 6 |
| 2013 | A novel temporal error concealment framework in H.264/AVCabstractIn this paper, we propose a novel framework for temporal error concealment in H.264/AVC. The proposed framework combines the properties of Skip Mode and Intra Mode in H.264/AVC Inter frame and robust motion inpainting in the computer vision community. In our framework, modes of the correctly received surrounding blocks of the missing macroblocks are checked first. If Intra Mode or Skip Mode blocks exist, depending on the boundary matching results, motion compensation based temporal error concealment method is used to conceal the missing macroblocks. Otherwise, improved robust motion inpainting based error concealment method is used to conceal the missing macroblocks. Experiments on several videos show visually pleasing results and low computation cost. Aidong Men, Bo Yang 0007 |
ICME | 4 |
| 2013 | A joint reconstruction algorithm for multi-view compressed imagingabstractAs compressed sensing can capture signal at sub-Nyquist rate, it is suitable to apply multi-view compressed imaging framework in vision sensor networks. The image views in such networks are correlated with each other, and therefore the performance of independent view reconstruction can be further improved by joint reconstruction. In this paper, we propose a joint reconstruction algorithm, where disparity estimation and disparity compensation are used to exploit the correlation between views. The target optimization problem is divided into two sub-problems and they are solved alternately by proximal-gradient method. We show by experiments that, for a given sub-rate, the proposed joint reconstruction scheme outperforms the independent reconstruction in terms of image quality. Kan Chang, Tuanfa Qin, Aidong Men |
ISCAS | 4 |
| 2013 | Efficient rate-distortion optimization for HEVC using SSIM and motion homogeneityabstractRate-distortion optimization is widely used in modern video codecs to make various encoder decisions in order to optimize the trade-off between bit-rate and quality. The distortion models used in HEVC are mean squared error (MSE) and sum of absolute difference (SAD), both of which are not always reflective of perceptual quality. In this paper, we show that SSIM, which has been found to be a good indicator of image visual quality, can be used as the distortion metric in the RDO framework in a simple yet effective manner by modifying the Lagrange multiplier used in RDO with the difference of motion homogeneity. Experiments show that the proposed scheme can achieve better rate-SSIM performance and provide higher subjective quality when compared with HEVC test model 10 (HM10.0) anchor. Fan Su, Qin Tu, Aidong Men |
PCS | 5 |
| 2013 | An SSIM-motivated LCU-level rate control algorithm for HEVCabstractThe rate control model used in the HEVC reference software HM10.0 is a R-λ model. In this model the distortion is measured by Mean Squared Error (MSE) and Sum of Absolute Difference(SAD). However, MSE and SAD are not correlated well with perceptual image quality. In this paper, we propose an improved LCU-level rate control algorithm based on the structural similarity (SSIM) index which was proved to be a better indicator of perceived image quality than MSE. What's more, we use the the structural similarity index to decide the weight of LCU-level bit allocation in the R-λ model and incorporate it to the rate-distortion optimization. Experiments show that the proposed scheme can achieve significant gain in terms of rate-SSIM performance and better subjective quality when compared with HEVC. Huiling Zhao, Aidong Men |
PCS | 5 |
| 2013 | A novel R-Q model based rate control scheme in HEVCabstractHigh Efficiency Video Coding (HEVC) standard, which has been published as ITU-T H.265-ISO/IEC 23008-2, is the latest video coding standard of the ITU-T and the ISO/IEC. The main goal of the HEVC standardization is to improve the compression performance significantly, about 50% bit rate reduction for equal perceptual video quality, compared to the H.264/AVC standard. For any practically applied video coding standard, rate control is always an integral part. This paper proposed a novel rate-quantization (R-Q) model based rate control scheme to further reduce the bitrate error. The experimental results show that the proposed algorithm has better performance compared with the existing algorithm. The bitrate error of the proposed algorithm is much lower than the existing algorithms, while the Y-PSNR loss is less. Xiaochuan Liang, Yinhe Zhou, Binji Luo, Aidong Men |
VCIP | 5 |
| 2013 | A novel motion compensated prediction framework using weighted AMVP prediction for HEVCabstractIn this paper, we propose a novel motion compensated prediction (MCP) framework which combines the properties of motion vector restriction and weighted advanced motion vector prediction (AMVP) to achieve higher prediction accuracy for High Efficiency Video Coding (HEVC). In our framework, motion vectors of the prediction units (PUs) surrounding the current PU are checked first by a motion field model, and the geometric relationship between motion vectors of the current PU and its neighboring coded PUs is analyzed. Then whether to use the weighted AMVP prediction method is determined by the motion vector restriction criterion. True motion vectors therefore can be obtained. Experimental results show that the proposed framework achieves BD-PSNR increments ranging from 0.03dB to 0.22dB and the BD-rate saving is up to 6.4%. Guangtao Fu, Aidong Men, Binji Luo, Huiling Zhao |
VCIP | 3 |
| 2013 | Radio-Frequency Tomographic Tracking of a Time-Varying Number of Targets with Wireless Sensor NetworksabstractRadio-frequency (RF) tomographic tracking is an emerging technology which tracks moving targets by analyzing changes of received signal strength (RSS) in wireless links. This paper presents and evaluates a novel RF tomographic tracking system that is capable of tracking a time-varying number of targets in wireless sensor networks (WSNs). The system incorporates two major contributions: a RSS histogram based observation model and a multi-target filtering algorithm based on multi-Bernoulli approximation. In addition, the sequential Monte Carlo method is applied to implement the multi-target filter. To evaluate the tracking system, an experiment involving 3 targets is performed within an indoor area of 50 square meters. Experimental results demonstrate that the proposed tracking system achieves high performance in accuracy and efficiency. Chonghua Liu, Weiping Shu, Aidong Men |
VTC Spring | 4 |
| 2012 | Intra prediction with enhanced inpainting method and vector predictor for HEVCabstractAs the successor to H.264/AVC, High Efficiency Video Coding (HEVC) will provide 50% reduction on compression data compared to H.264/AVC. In this paper, we propose an intra prediction method based on inpainting algorithms and vector predictor for HEVC. Our method utilizes a combination of two important inpainting algorithms: Laplace partial differential equation (PDE) and total variation (TV) model. Experiment results show that, compared to HEVC Test Model (HM) 2.0, our proposal achieves an average of 1.65% bitrate reduction. Xingli Qi, Aidong Men, Bo Yang 0007 |
ICASSP | 4 |
| 2012 | Depth map compression via edge-based inpaintingabstractThis work presents a novel intra-frame coding scheme for depth maps that makes use of the characteristics of depth images: large smooth areas separated by sharp edges. The proposed method is block-based, thus can be easily integrated into H.264/AVC framework. By using the extracted edge information, each block is divided into several regions and each region is handled independently. Most regions are predicted via image inpainting based on the Laplace equation. The others that could not be inpainted are predicted by their mean values or a default value. Both the extracted edge information and the mean values are compressed under the rate-distortion (RD) optimization principle. Compared with H.264/AVC, the results show that the proposed scheme achieves both PSNR gain in depth maps and visual quality improvement in rendered views. Jinhong Di, Aidong Men |
PCS | 5 |
| 2012 | Scrolling text processing for motion compensated frame interpolationabstractThis paper proposes a low complexity scrolling text processing method for motion compensated frame interpolation. An update strategy is given to determine whether to add update values to the candidate motion vectors, thus reducing interruption brought by the incorrect update values whereas not affecting the convergence speed. Without complex detection of the scrolling text, a geometric-based motion vector constraint method in motion estimation process can effectively keep the motion consistence of the scrolling text. Experimental results show that our interpolation results have much better visual quality than other methods in the presence of scrolling text. Besides, the proposed scheme is robust for video sequences which contain whether slow or fast motion. Test patterns without scrolling text also benefit from the suggested method. Aidong Men, Bo Yang 0007, Jinhong Di |
PCS | 2 |
| 2012 | Feedback-free distributed video coding using parallelized designabstractDistributed video coding (DVC) has a great development during the past decade for its simpler encoder and higher robustness against the channel noises. However, typical DVC schemes, such as DISCOVER, are almost based on the feedback channel between the encoder and the decoder, and the decoding process is very complex, which causes difficulty in feedback-free or tight delay constraint video application. In order to resolve the problems mentioned above, a feedback-free distributed video codec with parallelized design is presented using general-purpose computing on graphics processing units (GPU). A rate control algorithm based on offline correlation noise model (OCNM) is proposed to improve the performance of the DVC system. Then the parallel decoding methods are applied, including parallel low density parity check code accumulate (LDPCA) code decoding algorithm and simplified 3DRS design. Compared to the system without using parallel technology, our parallel codec is shown to be 11~25 times faster by using GPU. Aidong Men, Jinhong Di, Bo Yang 0007 |
PCS | 2 |
| 2012 | A Novel Motion Tracking System with Sparse Radio-Frequency Sensor NetworkabstractThis paper presents a novel device-free motion tracking (DFMT) system using an RF sensor network with very few nodes. The system includes a newly designed measurement model and corresponding tracking algorithm. Based on the uniform theory of diffraction (UTD), we prove that the received signal strength (RSS) measurements can be separated into two parts, the long-term one and the short-term one, reflecting the shadowing and scattering effects of the target. Then a model considering both the two effects is proposed to obtain positions of target using RSS measurements, which expands the sensing range of each wireless link. Particle filtering algorithm is applied to perform tracking. Experiment results illustrates that the system works well with very few nodes in different settings. Aidong Men |
VTC Fall | 1 |
| 2012 | RSS-Based Node Localization in the Existence of Moving ObstructionsabstractIn the context of wireless sensor networks, a node's location must be known for its data to be meaningful in many cases. Received signal strength (RSS)-based localization has been widely used because of low complexity and easy deployment. This paper proposes a novel method to localize nodes in the presence of randomly moving obstructions. We introduce background learning to reduce interferences caused by moving obstructions such as people or other objects. Based on our experimental results, each link of data is modeled as a mixture of Gaussians (MoG) and its parameters are updated by background learning. In this way, we can reduce the interferences of moving obstructions from obtained RSS measurements. Then we use least-square (LS) cooperative localization algorithm to implement node localization and the experimental results show good performance. Bo Yang 0007, Aidong Men, Qingchao Chen |
VTC Fall | 4 |
| 2012 | A novel fusion method in distributed multi-view video coding over wireless video sensor networkabstractTo meet the special requirements of resource-limited video sensors in wireless video sensor network (WVSN), low-complexity video encoding technique is highly desired. In distributed multi-view video coding (DMVC) system, multi-view video sources are encoded separately and decoded dependently, so the burden of huge computation is shifted from the encoder side to the decoder side. The generation of the side information (SI) is an important part in the design of a DMVC system as it directly relates to the system's performance. In this paper, a new fusion method combining the histogram matching and the minimum sum of the absolute differences (SAD) criterion is proposed to generate the final SI. The simulation results show that the proposed method generates more qualified SI and improves the peak signal to noise ratio (PSNR) performance at an average 1.6dB gain for reconstructed frame when compared to the traditional fusion methods in DMVC system. Manman Fan, Jinhong Di, Aidong Men |
WCNC | 4 |
| 2012 | Through-wall tracking with radio tomography networks using foreground detectionabstractThis paper presents a novel method for tracking a moving person or object through walls using wireless networks. The method takes advantage of the motion-induced variation of received signal strength (RSS) measurements in a radio tomography network. Based on real measurements of a deployed network, we show that the RSS distribution on a wireless link can be modeled as a mixture of Gaussians. An online learning algorithm is then proposed to update the model and detect whether the link is affected by the motion. Using spatial locations of the affected links, we apply the sequential Monte Carlo (SMC) methods to track the coordinates of a moving target. Experimental results show that the proposed method achieves high tracking accuracy in time-varying environment without the need for offline training. Aidong Men |
WCNC | 2 |
| 2011 | Real-time affine invariant patch matching using DCT descriptor and affine space quantizationabstractA novel framework for real-time image patch matching is proposed in this paper. It is composed of two parts: (a) a one-way descriptor based on Discrete Cosine Transform (DCT), (b) an optimal affine parameter quantization for descriptor dimensionality reduction. With these improvements, the patch matching is much faster than the state-of-art method and the training stage is fast enough to be performed online without the loss of accuracy or robustness. Experiments demonstrate the effectiveness of proposed framework in matching patches, while keeping affine invariant. Moreover, it works well in object detection and pose estimation applications. Fuguo Zhu, Aidong Men |
ICIP | 4 |
| 2011 | Sequential Monte Carlo for simultaneous passive device-free tracking and sensor localization using received signal strength measurements
Xi Chen 0042, Andrea Edelstein, Yunpeng Li 0001, Mark Coates, Michael G. Rabbat, Aidong Men |
IPSN | 6 |
| 2011 | Block-level adaptive optimization for inter-layer texture up-sampling in H.264/SVCabstractH.264 Scalable Video Coding (SVC) extension has spatial scalability which is able to provide various resolution sequences for a single encoded bit-stream. In order to reduce redundancies between different layers, for spatial scalable intra-coded frames, co-located reconstructed 8×8 sub-macroblock in base layer (BL) is up-sampled to predict the marcoblock (MB) in enhancement layer (EL). Unfortunately, simple 1-D poly-phase up-sampling filter used in current SVC isn't cable of achieving ideal result, which limits the performance of inter-layer intra prediction (ILIP). This paper proposes an adaptive optimization method for inter-layer texture up-sampling by applying wiener filter and controlling it at block level. Working as an additional part of ILIP, the proposed method can greatly reduce the prediction error between the original EL signals and the up-sampled BL signals. Experimental results show that, the proposed method achieves bit rate reduction up to 14.25% and PSNR increment up to 0.97 dB when compared with the traditional method in current SVC. Kan Chang, Tuanfa Qin, Wenhao Zhang 0001, Aidong Men |
MMSP | 4 |
| 2011 | An Improved Distributed Video Coding Scheme for Wireless Video Sensor NetworkabstractDistributed video coding (DVC) is considered as one promising video coding scheme for Wireless Video Sensor Network (WVSN) due to its high compression efficiency and error resilience functionalities, as well as the low encoding complexity. This paper presents an improved DVC scheme based on the source classification in wavelet domain. In this scheme, the new side information (SI) refinement method is adopted. Initial SI is achieved by motion-compensated weighted interpolation. Then we use a motion compensated refinement of the partially decoded Wyner-Ziv (WZ) frame to update SI. A better reconstruction of the WZ frame is eventually obtained with the refined SI. The results indicate that the proposed scheme can improve the whole performance of DVC codec compared to state-of-the-art DVC and be deployed over a real visual sensor platform. Jinhong Di, Aidong Men, Bo Yang 0007 |
VTC Fall | 2 |
| 2011 | An Improved Wyner-Ziv Video Coding for Sensor NetworkabstractWyner-Ziv video coding is a new compression paradigm based on two key Information Theory results: the Slepian-Wolf and Wyner-Ziv theorems. It shifts the complexity to the decoder, resulting in a low- complexity encoder suitable for mobile video communications and visual sensor networks. This paper presents an improved Wyner-Ziv video coding scheme for sensor network. An improved key frame encoding method based on the correlation noise model (CNM) is proposed, and then a 3DRS-assisted motion estimation algorithm and AOBMC technique are used to improve the rate-distortion performance of the codec. The results show that our coding scheme can achieve 2-4 dB gain compared to state-of-the-art TDWZ codec and be deployed over a real visual sensor platform. Aidong Men, Kan Chang, Jinhong Di |
VTC Spring | 2 |
| 2011 | An Unscented Kalman Filter for ICI Cancellation in High-Mobility OFDM SystemabstractOFDM system suffers from inter-carrier interference due to frequency offset produced by the movement of terminals. Several schemes have been proposed to mitigate this type of interference. In this paper, an unscented Kalman filter (UKF) based methodology is addressed to estimate the carrier frequency offset (CFO).We have compared the BER performance of UKF with other schemes and also analyzed the convergence as well as accuracy behavior between UKF and EKF. The simulation result shows that comparing to conventional non-iterative methods, UKF and EKF have higher accuracy and efficiency. Furthermore, UKF surpasses EKF in convergence rate and consistency. Bo Yang 0007, Aidong Men |
VTC Spring | 4 |
| 2010 | An improved Wyner-Ziv video coding with feedback channelabstractThis paper presents an improved feedback-assisted low complexity WZVC scheme. The performance of this scheme is improved by two enhancements: an improved mode-based key frame encoding and a 3DRS-assisted (three-dimensional recursive search assisted) motion estimation algorithm for WZ encoding. Experimental results show that our coding scheme can achieve significant gain compared to state-of-the-art TDWZ codec while still low encoding complexity. Aidong Men, Bo Yang 0007, Manman Fan, Kan Chang |
PCS | 2 |
| 2009 | GOP-Level Transmission Distortion Modeling for Video Streaming over Mobile NetworksabstractA major challenge in video coding and transmission over mobile networks is that the wireless channel is error-prone and the channel resources are limited. In this work, we analyze the picture distortion caused by channel errors and the distortion propagation behavior in its subsequent frames along the motion prediction path. We propose a linear fitting approach algorithm to achieve a low complexity GOP-level transmission distortion modeling. It is a predictive modeling which allows the encoder to predict the transmission distortion before the whole GOP is compressed and transmitted. The simulation results demonstrate that the proposed modeling has low computational complexity and high accuracy. It can be used in allocating the limited channel resources optimally for mobile video applications. Aidong Men, Kan Chang, Ziyi Quan |
IAS | 2 |
| 2009 | Adaptive Inter-layer Intra Prediction in Scalable Video CodingabstractIn the scalable video coding (SVC) standard, interlayer intra prediction (ILIP) is one of the most fundamental coding tools used to reduce bit rate. The prediction block of macroblock (MB) in enhancement layer (EL) is obtained by upsampling the co-located base layer (BL) block. In this paper, we propose a new ILIP method by introducing the adaptive signal processing techniques. Adaptive Wiener filters, which are calculated for each frame independently, are used to generate a prediction signal with minimum error energy. On average, a coding gain of 3.6% reduction of bit rate is obtained for QCIF-CIF scenario. Up to 9% bit rate reduction is achieved for CIF-4CIF scenario. Wenhao Zhang 0001, Aidong Men, Pinhua Chen |
ISCAS | 2 |
| 2009 | Adaptive optimizing filter for inter-layer intra prediction in SVCabstractSeveral inter-layer prediction techniques are adopted in the H.264 Scalable Video Coding Extension to increase the compression efficiency. For spatial scalable intra-coded frames, the prediction of macroblock in the enhancement layer is obtained by upsampling the co-located reconstructed block in the base layer. This inter-layer intra prediction can remove the redundancies between the different layers; however, its performance is limited by the non-ideal upsampling filter and the coding losses of base layer. In order to improve the performance of the inter-layer intra prediction, we propose an optimizing filtering method to enhance the prediction signal in this paper. Adaptive Wiener filters, which are calculated for each slice independently, are used to generate a prediction signal with minimum error energy in a statisitcal way. After this optimizing, the coding efficiency can be increased progressively. Experiments show that up to 6.45% and 15.45% bit rate reduction is achieved for QCIF-CIF scenario and CIF-4CIF scenario, respectively. Wenhao Zhang 0001, Aidong Men, Kan Chang |
ACM Multimedia | 2 |
| 2009 | An optimal bit allocation framework for error resilient Scalable Video CodingabstractIn this work, we propose an optimal joint bit allocation framework in order to provide unequal error protection (UEP) for scalable video coding. In the proposed scheme, video sequence is coded based on scalable video coding (SVC) extension of the H.264/AVC standard. We generate different streams according to changing quantization parameters. Each stream divides into different sub-streams in terms of time and quality scalability. These sub-streams with different significance levels are channel-coded using low density parity check (LDPC) codes with different coding rates respectively. Motivated by optimally allocating the bits to achieve resilience against channel induced corruptions, we propose an algorithm that optimally allocates the channel coding rates using a dynamic programming approach. We employ a probability model for decoding failure using LDPC codes. The results achieved clearly establish that the proposed framework can optimally allocate LDPC codes with varying coding rates to different sub-streams with the limited bandwidth. Saad B. Qaisar, Hayder Radha, Aidong Men |
PCS | 4 |
| 2005 | Perceptually Optimized Error-Resilient H.264 Video Streaming System over the Best-Effort InternetabstractStreaming video over Internet suffers varied quality degradation from packet delay, loss and spatiotemporal error propagation. This paper first proposes a perceptually optimized INTRA refreshing (POIR) H.264 video coding scheme, which stops error propagation quickly. It was inspired by the high-level recognition feature of the human visual system (HVS), and takes into consideration the error concealment (EC) function of decoder in an effort to optimize the end-toend distortion. Further, a spatio-temporal EC method at decoder based on refined boundary match criterion is introduced. Simulation illustrates that the POIR scheme:(1) maintains a best rate-distortion tradeoff point by selecting the number and placement of INTRA updated macroblock; (2) outperforms random intra refresh (RIR) by several dBs in terms of PSNR depending on sequences, especially at low coding bit rate and at high packet loss rate; (3) delivers a good tradeoff between performance and computational load, which is critical for real-time situations and portable devices. Liao Ning, Quan Zi-Yi, Aidong Men |
PDCAT | 3 |
| 2003 | A novel scheme of coding and modulation for digital television terrestrial broadcastingabstractIn this paper high-coding-rate block turbo codes (BTC) are adopted in digital television terrestrial broadcasting (DTTB) to replace the conventional applications of serially concatenated codes. A new multi-resolution 64QAM constellation is designed for hierarchical modulation to offer three tiers of services with dedicated protection in one DTTB channel. In the receiver, the simulation result on fixed-point iterative turbo decoding algorithm based on MAX-LOG-MAP criterion is aided to develop the ASIC core. Simulation in terms of bit-error-rate (BER) on additive white Gaussian noise channel shows that the required E/sub b//N/sub O/ High priority stream is only 3.9 dB at the threshold of visibility, which makes it more competent for mobile and indoor reception or other error-sensitive services. At the same time, the medium and low priority streams can support fixed outdoor reception. Zhixing Yang, Changyong Pan, Aidong Men |
PIMRC | 4 |