EDBT 2026 Demo / reviewers in the wild / expert
Bin Liu 0016
dblp:35/837-16
· DBLP profile ↗
133ranked-venue papers
6as first author
70since 2021 · last 2026
0000-0002-3977-8800ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 92 · 3 first-author · 53 since 2021Artificial intelligence and machine learning · 33 · 23 since 2021Computer networks · 14 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 since 2021Security and privacy · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and UnderstandingabstractIn this work, we present EchoGen, a unified framework for layout-to-image generation and image grounding, capable of generating images with both accurate layout and high fidelity to the text description.(e.g., spatial relationship), and grounding the image robustly at the same time. We believe that image grounding possesses strong text and layout understanding abilities, which can compensate for the corresponding limitations in layout-to-image generation. At the same time, images generated from layouts exhibit high diversity in content, thereby enhancing the robustness of image grounding. Jointly training both tasks within a unified model can promote performance improvements for each. However, we identify that this joint training paradigm encounters several optimization challenges and results in restricted performance. To address these issues, we propose progressive training strategies. First, the Parallel Multi-Task Pre-training (PMTP) stage equips the model with basic abilities for both tasks, leveraging shared tokens to accelerate training. Next, the Dual Joint Optimization (DJO) stage exploits task duality to sequentially integrate the two tasks, enabling unified optimization. Finally, the Cycle RL stage eliminates reliance on visual supervision by using consistency constraints as rewards, significantly enhancing the model’s unified capabilities via the GRPO strategy. Extensive experiments demonstrate state-of-the-art results on both layout-to-image generation and image grounding benchmarks, and reveal clear synergistic gains from optimizing the two tasks together. Dian Zheng, Jianxiong Gao, Bin Liu 0016 |
AAAI | 6 |
| 2026 | Uni-MMMU: A Massive Multi-discipline Multimodal Unified BenchmarkabstractKai Zou, Ziqi Huang, Yuhao Dong, Shulin Tian, Dian Zheng, Hongbo Liu, Jingwen He, Bin Liu, Yu Qiao, Ziwei Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuhao Dong, Shulin Tian, Dian Zheng, Jingwen He, Bin Liu 0016, Yu Qiao 0001, Ziwei Liu 0002 |
ACL (1) | 8 |
| 2026 | Advancing Aesthetic Image Generation via Composition Transfer
Bin Liu 0016, Nenghai Yu |
Int. J. Comput. Vis. | 3 |
| 2026 | DiffLoc+: Toward Robust Wi-Fi Hidden Camera Localization Based on Electromagnetic DiffractionabstractThe proliferation of hidden WiFi cameras has raised serious privacy concerns, making their accurate detection and localization essential for the secure development of future intelligent wireless networks. However, existing solutions often require substantial user involvement, large movement spaces, predefined system parameters, or pre-collected training data, limiting their practicality and scalability. In this paper, we present DiffLoc+, a novel and low-cost system that localizes hidden WiFi cameras by harnessing the fundamental physical principle of electromagnetic diffraction. When an obstacle crosses the line-of-sight path between a transmitter and a receiver, it causes a distinctive signal attenuation pattern. We theoretically analyze the feasibility of exploiting this phenomenon for localization and identify two key conditions for building an unbiased diffraction-based model: symmetry and observability. To satisfy these conditions, DiffLoc+ introduces a controllable diffraction generation mechanism that precisely rotates a small metal plate around a WiFi receiver (e.g. a Raspberry Pi), producing a stable and predictable diffraction “shadowing” effect. We then construct an unbiased localization model that maps this effect to the azimuth of the camera. To ensure the robustness of the theoretical model in real-world applications, DiffLoc+ further introduces two robustness-enhancing mechanisms: (1) an attenuation-region difference-driven subcarrier selection method, which filters subcarriers that reliably reflect the diffraction attenuation pattern by quantifying the signal contrast between diffraction- and reflection-dominated regions; and (2) an uncertainty evaluation framework that integrates result consistency and diffraction signal quality to eliminate unreliable estimates. Implemented entirely with commodity off-the-shelf (COTS) hardware, DiffLoc+ achieves an average angular error of 11.92° across six diverse indoor environments and eleven commercial camera models, demonstrating its effectiveness and robustness. Huan Yan 0004, Jian Liu 0055, Xiang Zhang 0011, Zhi Liu 0002, Bin Liu 0016, Meng Li 0006, Ming Gao 0023, Fusang Zhang |
IEEE J. Sel. Areas Commun. | 5 |
| 2026 | Dual-level modality debiasing learning for unsupervised visible-infrared person re-identification
Yan Lu 0001, Bin Liu 0016, Guojun Yin, Mang Ye |
Pattern Recognit. | 3 |
| 2025 | Rethinking Masked Data Reconstruction Pretraining for Strong 3D Action Representation LearningabstractIn 3D human action recognition, limited supervised data makes it challenging to fully tap into the modeling potential of powerful networks such as transformers. As a result, researchers have been actively investigating effective self-supervised pre-training strategies. For example, MAMP shows that instead of following the prevalent masked joint reconstruction, explicit masked motion reconstruction is key to the success of learning effective feature representation for 3D action recognition. However, we find that if we make a simple and effective change to the reconstructed target of masked joint reconstruction, masked joint reconstruction can achieve the same results as masked motion reconstruction. The devil is in the special characteristic of 3D skeleton data and the normalization process of training targets. We need to dig for all effective information of targets during normalization. Besides, considering that mask data reconstruction focuses more on learning local relations in input data for fulfilling the reconstruction task, instead of modeling the relation among samples, we further employ contrastive learning to learn more discriminative 3D action representations. We show that contrastive learning can consistently boost the performance of model pre-trained by masked joint prediction under various settings, especially in the semi-supervised setting that has a very limited number of labeled samples. Extensive experiments on NTU-60, NTU-120, and PKU-MMD datasets show that the proposed pre-training strategy achieves state-of-the-art results without bells and whistles. Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
AAAI | 3 |
| 2025 | Training-free Open-Vocabulary Semantic Segmentation via Diverse Prototype Construction and Sub-region MatchingabstractOpen-vocabulary semantic segmentation (OVSS) aims to segment images of arbitrary categories specified by class labels. While previous approaches relied on extensive image-text pairs or dense semantic annotations, recent training-free methods attempted to overcome these limitations by constructing semantic prototypes in the construction stage and image-to-image matching (i.e., prototype matching) during testing. However, these methods often struggle to effectively capture the visual characteristics of categories and fail to utilize local features during prototype matching. To deal with these problems, we propose a novel training-free framework for OVSS that constructs diverse prototypes and performs fine-grained sub-region matching. Specifically, our method leverages Large Language Models (LLMs) to guide support image generation by descriptions of different attributes of categories and employs coarse-fine clustering to obtain diverse and robust part-level prototypes in the construction stage. During testing, we propose a sub-region matching method, which assigns part-level prototypes to sub-regions utilizing optimal transport, to fully utilize local image features among part-level prototypes. Extensive experiments demonstrate the effectiveness of our method and show that our method achieves state-of-the-art performance, outperforming previous methods across five datasets. Xuanpu Zhao, Dianmo Sheng, Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
AAAI | 7 |
| 2025 | UNICL-SAM: Uncertainty-Driven In-Context Segmentation with Part Prototype DiscoveryabstractRecent advancements in in-context segmentation generalists have demonstrated significant success in performing various image segmentation tasks using a limited number of labeled example images. However, real-world applications present challenges due to the variability of support examples, which often exhibit quality issues resulting from various sources and inaccurate labeling. How to extract more robust representations from these examples has always been one of the goals of in-context visual learning. In response, we propose UNICL-SAM, to better model the example distribution and extract robust representations to help in-context segmentation. We incorporate an uncertainty probabilistic module to quantify each example’s reliability during both the training and testing phases. Utilizing this uncertainty estimation, we introduce an uncertainty-guided graph augmentation and feature refinement strategy, aimed at mitigating the impact of high-uncertainty regions to enhance the learning of robust representations. Subsequently, we construct prototypes for each example by aggregating part information, thereby creating reliable in-context instruction that effectively represents fine-grained local semantics. This approach serves as a valuable complement to traditional global pooling features. Experimental results demonstrate the effectiveness of the proposed framework, underscoring its potential for real-world applications. Dianmo Sheng, Dongdong Chen 0001, Zhentao Tan, Qiankun Liu 0001, Qi Chu 0001, Bin Liu 0016, Wenbin Tu, Shengwei Xu, Nenghai Yu |
CVPR | 7 |
| 2025 | Source-Free Domain Adaptation via Perceptual Semantic Decoupling for WiFi Gesture RecognitionabstractGeneralizable WiFi gesture recognition has gained increasing attention for its contactless operation, ubiquitous infrastructure and enhanced robustness. Among existing methods, source-free domain adaptation (SFDA) stands out by preserving privacy and reducing computational demands without relying on source data. Current methods typically process low-level WiFi signals and their high-level semantic representations from a unified perspective, making temporal semantic learning highly susceptible to low-level signal noise and lacking consistent semantic guidance for cross domain alignment, thereby limiting the effectiveness. In this paper, we propose ViFi, a novel SFDA framework specifically designed for cross-domain WiFi gesture recognition. Unlike prior work, ViFi introduces a viewpoint-hierarchical strategy that explicitly processes cross-domain sensing from two perspectives: the perceptual (signal-level) and the semantic (gesture-level). This separation mitigates the impact of signal noise on high-level semantics while preventing semantic space drift during domain alignment. ViFi operates in two key stages. First, it anchors the perceptual encoder and employs masked signal semantic reconstruction to learn robust high-level temporal semantics. Then, it freezes the semantic encoder and aligns the perceptual encoder across domains, again leveraging masked reconstruction to ensure alignment under a unified and meaningful semantic space. We evaluate ViFi on a public dataset, and experimental results show that our viewpoint-hierarchical method achieves over 15% improvement compared to the baseline and significantly outperforms state-of-the-art approaches. Yelin Wei, Xiang Zhang 0011, Bin Liu 0016, Songming Jia, Jinyang Huang, Zhi Liu 0002, Huan Yan 0004 |
GLOBECOM | 3 |
| 2025 | Training an Anti-KD Model that Cannot Teach Students via Similarity DisruptionabstractKnowledge Distillation (KD) aims to enhance the performance of student models by transferring knowledge from teacher models. While reaping the benefits of KD, the intellectual property risks associated with it cannot be ignored. Even if models are released without training data or provided as a service, potential adversaries can still clone the target model using KD. To mitigate the risks, some researchers propose training the anti-KD model that cannot teach student models. However, we find existing methods cannot defend against representation-based KD. To address the knowledge leakage from representations, we introduce Similarity Disruption (SD). SD increases the distance between the representation similarity matrices of our anti-KD model and the normal model, thereby reducing the effective information in the representation space. Extensive experiments demonstrate the proposed method can effectively defend against representation-based KD. Qi Chu 0001, Bin Liu 0016, Quanchen Zou, Deyue Zhang, Nenghai Yu |
ICASSP | 4 |
| 2025 | CMGait: Enhancing Cross-Modality Gait Recognition between LiDAR and RGB through Contrastive Identity-consistent Feature AggregationabstractCombination usage of LiDAR and RGB cameras for gait recognition can achieve cross space recognition and privacy protection. In addition, the widespread application of LiDAR cameras with 3D geometry information and the large amount of RGB gaits has led to the demand for cross-modality gait recognition on LiDAR and RGB modalities. To address the challenge of cross-modality recognition, we proposed a novel cross-modality gait recognition paradigm called CMGait. The key innovations include a novel projection method for transforming LiDAR point clouds into depth maps, feature alignment modules, Transformer-based identity encoders, and an embedding distance fusion method with similarity matrices based contrastive learning. Experimental results showcase state-of-the-art performance with Rank-1 accuracy of 62.8% and 66.1% for different directions in cross-modality gait recognition. Ablation experiments validate the effectiveness of the proposed methods, highlighting advancements in feature alignment and modality fusion techniques. Yubo Wang 0011, Bin Liu 0016, Jixiang Niu, Qi Chu 0001, Nenghai Yu |
ICASSP | 2 |
| 2025 | FE-CLIP: Frequency Enhanced CLIP Model for Zero-Shot Anomaly Detection and Segmentation
Qi Chu 0001, Bin Liu 0016, Wei Zhou 0021, Nenghai Yu |
ICCV | 3 |
| 2025 | Exploiting Feature Gating and Injection For Multi-modal Manipulation Detection and Grounding
Jiazhen Wang, Bin Liu 0016, Changtao Miao, Qi Chu 0001, Nenghai Yu |
ICIG (3) | 2 |
| 2025 | Exploring Generalized Features For LLM-Generated Text Detection
Jiazhen Wang, Bin Liu 0016, Changtao Miao, Qi Chu 0001, Quanchen Zou, Deyue Zhang, Nenghai Yu |
ICIG (3) | 2 |
| 2025 | Multimodal Consistency-Driven Deepfake Detection
Bin Liu 0016, Qi Chu 0001, Nenghai Yu |
ICIG (2) | 2 |
| 2025 | Remote Sensing Target Detector with Multi Scale Attention MechanismabstractMost of the existing rotation detection models focus on solving problems such as feature misalignment and boundary discontinuity, but ignore the use of contextual information in remote sensing images. However, context information plays a vital role in the accurate detection of small instance targets. Especially when the receptive field of the model is limited, it is very easy to cause deviations in the detection results. Therefore, we propose a Remote Sensing Target Detector with Multi-scale Attention Mechanism Network(MAMNet) that gradually enhances the features in the region of interest by utilizing the spatial, local, and global information of the image input features, thereby supplementing the attention mechanism at multiple scales, better identifying targets, and improving the detection performance of the model. Our model was experimented on DOTAv1.0, DOTAv1.5, HRSC2016 and has achieved the state of the art on DOTAv1.0. Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
ICIP | 4 |
| 2025 | A Watermark Updating Framework for Multi-stage Image Content DistributionabstractDeep image watermarking embeds identification data into images to facilitate source tracking. However, existing schemes are primarily designed for single-stage transmission scenarios, and in practical multi-stage distribution requirements, current methods degrade image quality and reduce watermark extraction accuracy. In this paper, we introduces WaterUp, a deep watermark updating framework. WaterUp automatically updates watermark information as the image is transmitted, preserving image quality while accurately recording the transmission path for traceability. The core of WaterUp is a flow-based encoder-decoder (FED), which utilizes a forward and backward network to enable efficient watermark updating with minimal computational and storage demands. Experimental results show that WaterUp outperforms state-of-the-art methods, maintaining high visual quality with a PSNR exceeding 38 dB across multiple transmissions. Bin Liu 0016, Jie Zhang 0073, Xiang Zhang 0011, Zehua Ma, Nenghai Yu |
ICME | 2 |
| 2025 | Adversarial Examples Detection Based on Adversarial Attack SensitivityabstractDeep neural networks have found widespread application in critical fields but remain vulnerable to adversarial attacks. Existing detection methods aim to achieve defense without modifying the model, but they generally struggle with generalization to unseen attacks. To address this limitation, we investigate the underlying principles of max-loss and min-distance adversarial attacks and uncover a strong positive correlation between perturbation magnitude, prediction confidence, and the distance to the decision boundary. Building on this insight, we introduce Adversarial Detection via Adversarial Sensitivity (ADAS), a novel approach that detects adversarial attacks by analyzing the sensitivity of a model's predictions to perturbation magnitude. ADAS estimates the distance to the decision boundary through sensitivity analysis by simulating adversarial attacks on input samples, identifying anomalies indicative of adversarial manipulation. Extensive experiments demonstrate the robustness and generalizability of ADAS across diverse and previously unseen adversarial attack scenarios, establishing its efficacy as a versatile and reliable detection framework. Cong Ming 0001, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
ICME | 6 |
| 2025 | Towards Anytime Retrieval: A Benchmark for Anytime Person Re-IdentificationabstractIn real applications, person re-identification (ReID) expects to retrieve the target person at any time, including both daytime and nighttime, ranging from short-term to long-term. However, existing ReID tasks and datasets cannot meet this requirement, as they are constrained by available time and only provide training and evaluation for specific scenarios. Therefore, we investigate a new task called Anytime Person Re-identification (AT-ReID), which aims to achieve effective retrieval in multiple scenarios based on variations in time. To address the AT-ReID problem, we collect the first large-scale dataset, AT-USTC, which contains 135k images of individuals wearing multiple clothes captured by RGB and IR cameras. Our data collection spans over an entire year and 270 volunteers were photographed on average 29.1 times across different dates or scenes, 4-15 times more than current datasets, providing conditions for follow-up investigations in AT-ReID. Further, to tackle the new challenge of multi-scenario retrieval, we propose a unified model named Uni-AT, which comprises a multi-scenario ReID (MS-ReID) framework for scenario-specific features learning, a Mixture-of-Attribute-Experts (MoAE) module to alleviate inter-scenario interference, and a Hierarchical Dynamic Weighting (HDW) strategy to ensure balanced training across all scenarios. Extensive experiments show that our model leads to satisfactory results and exhibits excellent generalization to all scenarios. Xulin Li, Yan Lu 0001, Bin Liu 0016, Qinhong Yang, Qi Chu 0001, Mang Ye, Nenghai Yu |
IJCAI | 3 |
| 2025 | CamLopa: A Hidden Wireless Camera Localization Framework via Signal Propagation Path AnalysisabstractHidden wireless cameras pose significant privacy threats, necessitating effective detection and localization methods. However, existing localization solutions often require impractical activity spaces, expensive specialized devices, or pre-collected training data, limiting their practical deployment. To address these limitations, we introduce CamLopa, a training-free wireless camera localization framework that operates with minimal activity space constraints using low-cost, commercial-off-the-shelf (COTS) devices. CamLopa can achieve detection and localization in just 45 seconds of user activities with a Raspberry Pi board. During this short period, it analyzes the causal relationship between wireless traffic and user movement to detect the presence of a hidden camera. Upon detection, CamLopa utilizes a novel azimuth localization model based on wireless signal propagation path analysis for localization. This model leverages the time ratio of user paths crossing the First Fresnel Zone (FFZ) to determine the camera's azimuth angle. Subsequently, CamLopa refines the localization by identifying the camera's quadrant. We evaluate CamLopa across various devices and environments, demonstrating its effectiveness with a 95.37% detection accuracy for snooping cameras and an average localization error of 17.23°, under the significantly reduced activity space requirements and without the need for training. Our code and demo are available at https://github.com/CamLoPA/CamLoPA-Code. Xiang Zhang 0011, Jie Zhang 0073, Zehua Ma, Jinyang Huang, Meng Li 0006, Huan Yan 0004, Peng Zhao 0024, Zijian Zhang 0001, Bin Liu 0016, Qing Guo 0005, Tianwei Zhang 0004, Nenghai Yu |
SP | 9 |
| 2025 | DiffLoc: WiFi Hidden Camera Localization Based on Electromagnetic Diffraction
Xiang Zhang 0011, Jie Zhang 0073, Huan Yan 0004, Jinyang Huang, Zehua Ma, Bin Liu 0016, Meng Li 0006, Kejiang Chen, Qing Guo 0005, Tianwei Zhang 0004, Zhi Liu 0002 |
USENIX Security Symposium | 6 |
| 2025 | Context-Aware Weakly Supervised Image Manipulation Localization With SAM RefinementabstractMalicious image manipulation poses societal risks, increasing the importance of effective image manipulation detection methods. Recent approaches in image manipulation detection have largely been driven by fully supervised approaches, which require labor-intensive pixel-level annotations. Thus, it is essential to explore weakly supervised image manipulation localization methods that only require image-level binary labels for training. However, existing weakly supervised image manipulation methods overlook the importance of edge information for accurate localization, leading to suboptimal localization performance. To address this, we propose a Context-Aware Boundary Localization (CABL) module to aggregate boundary features and learn context-inconsistency for localizing manipulated areas. Furthermore, by leveraging Class Activation Mapping (CAM) and Segment Anything Model (SAM), we introduce the CAM-Guided SAM Refinement (CGSR) module to generate more accurate manipulation localization maps. By integrating two modules, we present a novel weakly supervised framework based on a dual-branch Transformer-CNN architecture. Our method achieves outstanding localization performance across multiple datasets. Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
IEEE Signal Process. Lett. | 4 |
| 2025 | ReSup: Reliable Label Noise Suppression for Facial Expression RecognitionabstractBecause of the ambiguous and subjective property of the facial expression, the label noise is widely existing in the FER dataset. For this problem, in the training phase, current methods often directly predict whether the label is noised or not, aiming to reduce the contribution of the noised data. However, we argue that this kind of method suffers from the low reliability of such noise data decision operation. It makes that some mistakenly abounded clean data are not utilized sufficiently and some mistakenly kept noised data disturbing the model learning. In this paper, we propose a more reliable noise-label suppression method called ReSup. First, instead of directly predicting noised or not, ReSup makes the noise data decision by modeling the distribution of noise and clean labels simultaneously according to the disagreement between the prediction and the target. Specifically, to achieve optimal distribution modeling, ReSup models the similarity distribution of all samples. To further enhance the reliability of our noise decision results, ReSup uses two networks to jointly achieve noise suppression. Specifically, ReSup utilize the property that two networks are less likely to make the same mistakes, making two networks swap decisions and tending to trust decisions with high agreement. Extensive experiments on popular datasets shows the effectiveness of ReSup. Xiang Zhang 0011, Yan Lu 0001, Huan Yan 0005, Jinyang Huang, Yu Gu 0003, Yusheng Ji, Zhi Liu 0002, Bin Liu 0016 |
IEEE Trans. Affect. Comput. | 8 |
| 2025 | Bootstrapping Audio-Visual Video Segmentation by Strengthening Audio CuesabstractHow to effectively interact audio with vision has garnered considerable interest within the multi-modality research field. Recently, a novel audio-visual video segmentation (AVS) task has been proposed, aiming to segment the sounding objects in video frames under the guidance of audio cues. However, most existing AVS methods are hindered by a modality imbalance where the visual features tend to dominate those of the audio modality, due to a unidirectional and insufficient integration of audio cues. This imbalance skews the feature representation towards the visual aspect, impeding the learning of joint audio-visual representations and potentially causing segmentation inaccuracies. To address this issue, we propose AVSAC. Our approach features a Bidirectional Audio-Visual Decoder (BAVD) with integrated bidirectional bridges, enhancing audio cues and fostering continuous interplay between audio and visual modalities. This bidirectional interaction narrows the modality imbalance, facilitating more effective learning of integrated audio-visual representations. Additionally, we present a strategy for audio-visual frame-wise synchrony as fine-grained guidance of BAVD. This strategy enhances the share of auditory components in visual features, contributing to a more balanced audio-visual representation learning. Extensive experiments show that our method has state-of-the-art performance on several AVS public benchmarks. Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu, Le Lu 0001, Jieping Ye |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | WiOpen: A Robust Wi-Fi-Based Open-Set Gesture Recognition FrameworkabstractRecent years have witnessed a growing interest in Wi-Fi-based gesture recognition. However, existing works have predominantly focused on closed-set paradigms, where all testing gestures are predefined during training. This poses a significant challenge in real-world applications, as unseen gestures might be misclassified as known class during testing. To address this issue, we propose WiOpen, a robust Wi-Fi-based open-set gesture recognition (OSGR) framework. Implementing OSGR requires addressing challenges caused by the unique uncertainty in Wi-Fi sensing. This uncertainty, resulting from noise and domains, leads to widely scattered and irregular data distributions in collected Wi-Fi sensing data. Consequently, data ambiguity between classes and challenges in defining appropriate decision boundaries to identify unknowns arise. To tackle these challenges, WiOpen adopts a twofold approach to eliminate uncertainty and define precise decision boundaries. Initially, it addresses uncertainty induced by noise during data preprocessing by utilizing the channel state information (CSI) ratio. Next, it designs the OSGR network based on an uncertainty quantification method. Throughout the learning process, this network effectively mitigates uncertainty stemming from domains. Ultimately, the network leverages relationships among samples' neighbors to dynamically define open-set decision boundaries, successfully realizing OSGR. Comprehensive experiments on publicly accessible datasets confirm WiOpen's effectiveness. Xiang Zhang 0011, Jinyang Huang, Huan Yan 0004, Yuanhao Feng, Peng Zhao 0024, Guohang Zhuang, Zhi Liu 0002, Bin Liu 0016 |
IEEE Trans. Hum. Mach. Syst. | 8 |
| 2025 | Multi-spectral Class Center Network for Face Manipulation LocalizationabstractAs Deepfake content proliferates online, advancing face manipulation forensics has become crucial. To combat this emerging threat, previous methods mainly focus on studying how to distinguish authentic and manipulated face images. Although impressive, image-level classification lacks explainability and is limited to specific application scenarios, spurring recent research on pixel-level prediction for face manipulation forensics. However, existing forgery localization methods suffer from exploring frequency-based forgery traces in the localization network. In this paper, we observe that multi-frequency spectrum information is effective for identifying tampered regions. To this end, a novel Multi-spectral Class Center Network (MSCCNet) is proposed for face manipulation localization. Specifically, we design a Multi-spectral Class Center (MSCC) module to learn more generalizable and multi-frequency features. Based on the features of different frequency bands, the MSCC module collects multi-spectral class centers and computes pixel-to-class relations. Applying multi-spectral class-level representations suppresses the semantic information of the visual concepts which is insensitive to manipulated regions of forgery images. Furthermore, we propose a Multi-level Features Aggregation (MFA) module to employ more low-level forgery artifacts and structural textures. Meanwhile, we conduct a comprehensive localization benchmark based on pixel-level FF++ and Dolos datasets. Experimental results quantitatively and qualitatively demonstrate the effectiveness and superiority of the proposed MSCCNet. We expect this work to inspire more studies on pixel-level face manipulation localization. The codes are available. Changtao Miao, Qi Chu 0001, Zhentao Tan, Zhenchao Jin, Wanyi Zhuang, Bin Liu 0016, Honggang Hu, Nenghai Yu |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2024 | TCI-Former: Thermal Conduction-Inspired Transformer for Infrared Small Target DetectionabstractInfrared small target detection (ISTD) is critical to national security and has been extensively applied in military areas. ISTD aims to segment small target pixels from background. Most ISTD networks focus on designing feature extraction blocks or feature fusion modules, but rarely describe the ISTD process from the feature map evolution perspective. In the ISTD process, the network attention gradually shifts towards target areas. We abstract this process as the directional movement of feature map pixels to target areas through convolution, pooling and interactions with surrounding pixels, which can be analogous to the movement of thermal particles constrained by surrounding variables and particles. In light of this analogy, we propose Thermal Conduction-Inspired Transformer (TCI-Former) based on the theoretical principles of thermal conduction. According to thermal conduction differential equation in heat dynamics, we derive the pixel movement differential equation (PMDE) in the image domain and further develop two modules: Thermal Conduction-Inspired Attention (TCIA) and Thermal Conduction Boundary Module (TCBM). TCIA incorporates finite difference method with PMDE to reach a numerical approximation so that target body features can be extracted. To further remove errors in boundary areas, TCBM is designed and supervised by boundary masks to refine target body features with fine boundary details. Experiments on IRSTD-1k and NUAA-SIRST demonstrate the superiority of our method. Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
AAAI | 5 |
| 2024 | MotionGPT: Finetuned LLMs Are General-Purpose Motion GeneratorsabstractGenerating realistic human motion from given action descriptions has experienced significant advancements because of the emerging requirement of digital humans. While recent works have achieved impressive results in generating motion directly from textual action descriptions, they often support only a single modality of the control signal, which limits their application in the real digital human industry. This paper presents a Motion General-Purpose generaTor (MotionGPT) that can use multimodal control signals, e.g., text and single-frame poses, for generating consecutive human motions by treating multimodal signals as special input tokens in large language models (LLMs). Specifically, we first quantize multimodal control signals into discrete codes and then formulate them in a unified prompt instruction to ask the LLMs to generate the motion answer. Our MotionGPT demonstrates a unified human motion generation model with multimodal control signals by tuning a mere 0.4% of LLM parameters. To the best of our knowledge, MotionGPT is the first method to generate human motion by multimodal control signals, which we hope can shed light on this new direction. Visit our webpage at https://qiqiapink.github.io/MotionGPT/. Bin Liu 0016, Shixiang Tang, Yan Lu 0001, Lu Chen 0001, Lei Bai 0001, Qi Chu 0001, Nenghai Yu, Wanli Ouyang |
AAAI | 3 |
| 2024 | Unifying Multi-Modal Uncertainty Modeling and Semantic Alignment for Text-to-Image Person Re-identificationabstractText-to-Image person re-identification (TI-ReID) aims to retrieve the images of target identity according to the given textual description. The existing methods in TI-ReID focus on aligning the visual and textual modalities through contrastive feature alignment or reconstructive masked language modeling (MLM). However, these methods parameterize the image/text instances as deterministic embeddings and do not explicitly consider the inherent uncertainty in pedestrian images and their textual descriptions, leading to limited image-text relationship expression and semantic alignment. To address the above problem, in this paper, we propose a novel method that unifies multi-modal uncertainty modeling and semantic alignment for TI-ReID. Specifically, we model the image and textual feature vectors of pedestrian as Gaussian distributions, where the multi-granularity uncertainty of the distribution is estimated by incorporating batch-level and identity-level feature variances for each modality. The multi-modal uncertainty modeling acts as a feature augmentation and provides richer image-text semantic relationship. Then we present a bi-directional cross-modal circle loss to more effectively align the probabilistic features between image and text in a self-paced manner. To further promote more comprehensive image-text semantic alignment, we design a task that complements the masked language modeling, focusing on the cross-modality semantic recovery of global masked token after cross-modal interaction. Extensive experiments conducted on three TI-ReID datasets highlight the effectiveness and superiority of our method over state-of-the-arts. Bin Liu 0016, Yan Lu 0001, Qi Chu 0001, Nenghai Yu |
AAAI | 2 |
| 2024 | Towards More Unified In-Context Visual UnderstandingabstractThe rapid advancement of large language models (LLMs) has accelerated the emergence of in-context learning (ICL) as a cutting-edge approach in the natural language processing domain. Recently, ICL has been employed in visual understanding tasks, such as semantic segmentation and image captioning, yielding promising results. However, existing visual ICL framework can not enable producing content across multiple modalities, whicd limits their potential usage scenarios. To address this issue, we present a new ICLframeworkfor visual understanding with multi-modal output enabled. First, we quantize and embed both text and visual prompt into a unified representational space, structured as interleaved in-context sequences. Then a decoder-only sparse transformer architecture is employed to perform generative modeling on them, facilitating in-context learning. Thanks to this design, the model is capable of handling in-context vision understanding tasks with multimodal output in a unified pipeline. Experimental re-sults demonstrate that our model achieves competitive performance compared with specialized models and previous ICL baselines. Overall, our research takes a further step toward unified multimodal in-context learning. Dianmo Sheng, Dongdong Chen 0001, Zhentao Tan, Qiankun Liu 0001, Qi Chu 0001, Jianmin Bao, Bin Liu 0016, Shengwei Xu, Nenghai Yu |
CVPR | 8 |
| 2024 | Exploiting Modality-Specific Features for Multi-Modal Manipulation Detection and GroundingabstractAI-synthesized text and images have gained significant attention, particularly due to the widespread dissemination of multi-modal manipulations on the internet, which has resulted in numerous negative impacts on society. Existing methods for multi-modal manipulation detection and grounding primarily focus on fusing vision-language features to make predictions, while overlooking the importance of modality-specific features, leading to sub-optimal results. In this paper, we construct a simple and novel transformer-based framework for multi-modal manipulation detection and grounding tasks. Our framework simultaneously explores modality-specific features while preserving the capability for multi-modal alignment. To achieve this, we introduce visual/language pre-trained encoders and dual-branch cross-attention (DCA) to extract and fuse modality-unique features. Furthermore, we design decoupled fine-grained classifiers (DFC) to enhance modality-specific feature mining and mitigate modality competition. Moreover, we propose an implicit manipulation query (IMQ) that adaptively aggregates global contextual cues within each modality using learnable queries, thereby improving the discovery of forged details. Extensive experiments on the DGM4dataset demonstrate the superior performance of our proposed model compared to state-of-the-art approaches. Jiazhen Wang, Bin Liu 0016, Changtao Miao, Wanyi Zhuang, Qi Chu 0001, Nenghai Yu |
ICASSP | 2 |
| 2024 | DSIS: A Novel (K, N) Threshold Deniable Secret Image Sharing Scheme with Lossless RecoveryabstractSecret image sharing (SIS) schemes have undergone significant development. However, to the best of our knowledge, none of the existing schemes has considered the deniable property during secret sharing. This presents a problem when we need to share secret images through an untrusted and supervised channel, where we may be coerced to reveal the secret to adversary. Here we propose a deniable SIS (DSIS) scheme. Before sharing the secret image, we manipulate the secret area of the image to create a forged image that possesses deniability. Then we employ SIS to distribute the forged image, while generating an auxiliary matrix derived from secret key. This matrix governs rules for sharing secret area. In the event of coercion to reveal the secret, we have the capability to present the adversary with the forged image instead, thereby retaining control over the disclosure of the secret area at our discretion. In DSIS, we can obtain a secret image with the small-sized secret key losslessly, while we can recover another visually-meaningful fake image to safeguard the secrecy and protect ourselves when facing coercion. Zikai Xu, Bin Liu 0016, Weihai Li, Nenghai Yu |
ICASSP | 2 |
| 2024 | SE-SIS: Shadow-Embeddable Lossless Secret Image Sharing for Greyscale ImagesabstractSecret image sharing (SIS) has made significant progress in research and has found wide applications. However, we note that shadows of traditional SIS contain a large amount of redundancy. A novel Shadow-Embeddable Secret Image Sharing scheme (SE-SIS) leveraging the redundancy in the shadows is proposed in this paper. SE-SIS utilizes the random values in Lagrange polynomials of traditional secret image sharing (SIS) scheme, and modifies a shadow to embed another secret image with a secret key. Then other shadows are modified simultaneously according to the properties of Lagrange polynomials to ensure the accurate recovery of the previously shared image. It is worth noting that embedding process does not impact the recovery of the shared image, and the embedded shadow is indistinguishable from the others. SE-SIS modifies noise-liked shadows into other noise-liked ones without affecting the recovery process, thereby achieving a high embedding rate. Meanwhile, SE-SIS realizes lossless recovery for both the shared image and the secret image. Experimental results indicate SE-SIS constructs randomized shadows and exhibits excellent performance in terms of Peak Signal to Noise Ratio (PSNR) and embedding rate. Zikai Xu, Bin Liu 0016, Weihai Li, Nenghai Yu |
ICASSP | 2 |
| 2024 | Hidden WiFi Camera Localization via Signal Propagation Path AnalysisabstractHidden WiFi cameras pose significant privacy threats, necessitating effective localization methods. In this work, we introduce CamLoPA, a system designed for the detection and localization of WiFi cameras. CamLoPA achieves this in just 45 seconds of user walking. It begins by analyzing the causal relationship between WiFi traffic and user movement to identify the presence of a snooping camera. Upon detection, CamLoPA utilizes a novel azimuth location model based on WiFi signal propagation path analysis to localize the hidden camera. Comprehensive evaluations demonstrate that CamLoPA can accurately and swiftly detect and localize snooping WiFi cameras with minimal constraints. Xiang Zhang 0011, Zehua Ma, Jinyang Huang, Huan Yan 0004, Meng Li 0006, Zhi Liu 0002, Bin Liu 0016 |
MobiCom | 7 |
| 2024 | Detect Text Forgery with Non-forged Image Features: A Framework for Detection and Grounding of Image-Text Manipulation
Changtao Miao, Qi Chu 0001, Dianmo Sheng, Jiazhen Wang, Bin Liu 0016, Nenghai Yu |
PRCV (11) | 7 |
| 2024 | Feature Preservation and Shape Cues Assist Infrared Small Target DetectionabstractInfrared small target detection (ISTD) aims to segment small target pixels from infrared images and has extensive applications in many fields. Despite multiple progress, challenges remain as present methods still easily suffer from missed detection. Also, present methods are not sensitive enough to irregular target shapes. We argue that the main reason is that some informative small target features get lost during the aggressive downsampling in the encoder without effective recovery. In this article, we propose a new network with a dual-branch encoder-decoder structure for ISTD to address the two challenges. Specifically, to better preserve small target body features for more accurate target locations, we propose to maintain a relatively high resolution of feature maps in one encoder branch. For the other encoder branch, we gradually enlarge feature channels while shrinking resolutions and devise Perona-Malik diffusion (PMD) blocks to preserve shape cues inspired by the shape-preserving effect of PMD in denoising. The encoded high-resolution target body features and high-channel shape cues actually complement each other, so we design channel-resolution interact modules (CRIMs) to combine them. In the decoder, we propose orthogonal central difference fusion (OCDF) that relies on mining contrast differences to further refine shape-aware ISTD quality. Experiments on NUAA-SIRST and IRSTD-1k prove the superiority of our method. Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | MiM-ISTD: Mamba-in-Mamba for Efficient Infrared Small-Target DetectionabstractRecently, infrared small-target detection (ISTD) has made significant progress, thanks to the development of basic models. Specifically, the models combining CNNs with Transformers can successfully extract both local and global features. However, the disadvantage of the Transformer is also inherited, that is, the quadratic computational complexity to sequence length. Inspired by the recent basic model with linear complexity for long-distance modeling, Mamba, we explore the potential of this state-space model (SSM) for ISTD tasks in terms of effectiveness and efficiency in the article. However, directly applying Mamba achieves suboptimal performances due to the insufficient harnessing of local features, which are imperative for detecting small targets. Instead, we tailor a nested structure, Mamba-in-Mamba (MiM-ISTD), for efficient ISTD. It consists of Outer and Inner Mamba blocks to adeptly capture both global and local features. Specifically, we treat the local patches as “visual sentences” and use the Outer Mamba to explore the global information. We then decompose each visual sentence into subpatches as “visual words” and use the Inner Mamba to further explore the local information among words in the visual sentence with negligible computational costs. By aggregating the visual word and visual sentence features, our MiM-ISTD can effectively explore both global and local information. Experiments on NUAA-SIRST and IRSTD-1k show the superior accuracy and efficiency of our method. Specifically, MiM-ISTD is$8\times $faster than the SOTA method and reduces GPU memory usage by 62.2% when testing on$2048 \times 2048$images, overcoming the computation and memory constraints on high-resolution infrared images. Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu, Jieping Ye |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | PhyFinAtt: An Undetectable Attack Framework Against PHY Layer Fingerprint-Based WiFi AuthenticationabstractWiFi connection has been suffering from MAC forgery attacks due to the loose authentication mechanism between access points (APs) and clients. To address this problem, the physical (PHY) layer information-based fingerprint has been adopted for safe WiFi authentication. Since such a fingerprint is constant and unique for each specific network interface card (NIC), it can effectively prevent MAC forgery attacks. However, the PHY layer information-based fingerprint is still vulnerable to malicious attacks as it is extracted from Channel State Information (CSI), and its stability can be affected by the wireless environment. In this paper, we propose a novel undetectable attack framework, called PhyFinAtt, base on which the attacker can undermine the stability of the PHY layer-based authentication fingerprints through human movement and further attack the WiFi authentication protocols. Specifically, we first demonstrate that human movement at a designated location can affect the PHY fingerprint. We then illustrate the impact of human movement on the PHY fingerprint and the relationship between the movement and the channel quality to ensure that the PHY fingerprint is destroyed by the movement in an undetected way without affecting normal communication. Extensive experiments in real-world scenarios show that our proposed attack can effectively disrupt the stability of the PHY fingerprints and significantly degrade the performance of the authentication protocols based on such fingerprints. To the best of our knowledge, this is the first study on effective attacks against the PHY information-based WiFi authentication protocols. Furthermore, we also present a practical defense mechanism without involving any additional equipment to mitigate attacks similar to PhyFinAtt. Jinyang Huang, Bin Liu 0016, Chenglin Miao, Xiang Zhang 0011, Jianchun Liu, Lu Su 0001, Zhi Liu 0002, Yu Gu 0003 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Joint Identity-Aware Mixstyle and Graph-Enhanced Prototype for Clothes-Changing Person Re-IdentificationabstractIn recent years, considerable progress has been witnessed in the person re-identification (Re-ID). However, in a more realistic long-term scenario, the appearance shift arising from the clothes-changing inevitably deteriorates the conventional methods that heavily depend on the clothing color. Although the current clothes-changing person Re-ID methods introduce external human knowledge (i.e, contour, mask) and sophisticated feature decoupling strategy to alleviate the clothing shift, they still face the risk of overfitting to clothing due to the limited clothing diversity of training set. To more efficiently and effectively promote the clothes-irrelevant feature learning, we present a novel joint Identity-aware Mixstyle and Graph-enhanced Prototype method for clothes-changing person Re-ID. Specifically, by treating the cloth-changing as fine-grained domain/style shift, the identity-aware mixstyle (IMS) is proposed from the perspective of domain generalization, which mixes the instance-level feature statistics of samples within each identity to synthesize novel and diverse clothing styles, while retaining the correspondence between synthesized samples and latent label space. By incorporating the IMS module, the more diverse styles can be exploited to train a clothing-shift robust model. To further reduce the feature discrepancy caused by clothing variations, the graph-enhanced prototype constraint (GEP) module is proposed to explore the graph similarity structure of style-augmented samples across memory bank to build informative and robust prototypes, which serve as powerful exemplars for better clothing-irrelevant metric learning. The two modules are integrated into a joint learning framework and benefit each other. The extensive experiments conducted on clothes-changing person Re-ID datasets validate the superiority and effectiveness of our method. In addition, our method also shows good universality and corruption robustness on other Re-ID tasks. Bin Liu 0016, Yan Lu 0001, Qi Chu 0001, Nenghai Yu, Chang Wen Chen |
IEEE Trans. Multim. | 2 |
| 2023 | BAUENet: Boundary-Aware Uncertainty Enhanced Network for Infrared Small Target DetectionabstractInfrared small target detection (ISTD) is indispensable in remote sensing and military surveillance. Existing ISTD methods can discover regularly-shaped and clear objects well, but tend to overlook the tough-to-detect ones, such as targets with irregular shapes or blurry boundaries, causing inaccurate segmentation and missed detection. Considering that boundary areas assemble rich uncertainty information, we propose the Boundary-Aware Uncertainty Enhanced Network (BAUENet), where Uncertainty Enhanced Context Refinement (UECR) and Adaptive Feature Fusion Modules (AFFM) are devised to address this problem. Specifically, UECR extracts spatial contexts and refines them with uncertain area maps derived from backbone intermediate outputs, so as to distinguish boundary areas from other regions. AFFM adaptively aggregates cross-level features via balancing low-level details and high-level semantics for finer boundary preservation in both channel and spatial dimensions during up-sampling feature fusion. Experiments on several public datasets demonstrate the effectiveness of the proposed method, especially for irregular shape and blurry boundary cases. Qi Chu 0001, Zhentao Tan, Bin Liu 0016, Nenghai Yu |
ICASSP | 4 |
| 2023 | Dual-Feature Enhancement for Weakly Supervised Temporal Action LocalizationabstractWeakly-supervised Temporal Action Localization (WTAL) aims at localizing actions in untrimmed videos with only video-level labels. Most existing methods embrace a "localization by classification" paradigm and adopt a model that pre-trained with recognition task for feature extraction. The gap between recognition and localization tasks leads to inferior performance. Some recent works attempt to utilize feature enhancement to obtain better feature for localization and boost the performance to some extent. However, they are limited to intra-video information exploiting, while ignoring meaningful inter-video information in the dataset. In this paper, we propose a novel Dual-Feature Enhancement (DFE) method for WTAL, which can utilize both intra-and inter-video information. For intra-video, a local feature enhancement module is designed to promote the feature interaction along the temporal dimension within each video. For inter-video information, a global memory module is firstly designed to learn the representations for different categories across different videos. Then, a global feature enhancement module is used to enhance the video features with the help of those global representations in the memory. Besides, to reduce the extra computational cost caused by global enhancement module in the inference stage, a distillation loss is applied to enforce the local branch to learn the information from global branch, so the global enhancement module could be removed during inference. The proposed method achieves state-of-the-art performance on popular benchmarks. Qiankun Liu 0001, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
ICASSP | 4 |
| 2023 | Dual-Uncertainty Guided Curriculum Learning and Part-Aware Feature Refinement for Domain Adaptive Person Re-IdentificationabstractUnsupervised Domain Adaptative person re-identification (UDA ReID) aims to transfer the knowledge of pre-trained model from labeled source domain to unlabeled target domain. Although the current clustering-based methods have achieved promising success, they neglect the tolerance of the model to cope with different-level noise, which may cause the model to memorize some incorrect patterns caused by label noise and overfit on them rapidly in the early stages. In this paper, we introduce a novel Dual Uncertainty guided Curriculum Learning (DUCL) method to tackle the above problems. Specifically, the reliability-based curriculum allocation is proposed to enforce the sample adaptation in an easy-to-hard manner, which is further assisted by a novel dual-uncertainty re-weighting strategy to alleviate the influence of label noise. In addition, we design Part-aware Feature Refinement (PAFR) to enhance the discrimination of model and thereby acquiring more reliable pseudo-labels. Specifically, the part-aware attention maps are exploited in the PAFR to integrate fine-grained semantics into holistic representation. Extensive experiments have validated the superiority of the proposed method. Zhangping Liu, Bin Liu 0016, Qi Chu 0001, Nenghai Yu |
ICASSP | 2 |
| 2023 | Evopose: A Recursive Transformer for 3D Human Pose Estimation with Kinematic Structure PriorsabstractTransformer is popular in recent 3D human pose estimation, which utilizes long-term modeling to lift 2D keypoints into the 3D space. However, current transformer-based methods do not fully exploit the prior knowledge of the human skeleton provided by the kinematic structure. In this paper, we propose a novel transformer-based model EvoPose to introduce the human body prior knowledge for 3D human pose estimation effectively. Specifically, a Structural Priors Representation (SPR) module represents human priors as structural features carrying rich body patterns, e.g. joint relationships. The structural features are interacted with 2D pose sequences and help the model to achieve more informative spatiotemporal features. Moreover, a Recursive Refinement (RR) module is applied to refine the 3D pose outputs by utilizing estimated results and further injects human priors simultaneously. Extensive experiments demonstrate the effectiveness of EvoPose which achieves a new state of the art on two most popular benchmarks, Human3.6M and MPI-INF-3DHP. Yan Lu 0001, Bin Liu 0016, Qi Chu 0001, Nenghai Yu |
ICASSP | 3 |
| 2023 | Enhancing Adversarial Transferability from the Perspective of Input Loss Landscape
Yinhu Xu, Qi Chu 0001, Zixiang Luo, Bin Liu 0016, Nenghai Yu |
ICIG (1) | 5 |
| 2023 | Revisiting TENT for Test-Time Adaption Semantic Segmentation and Classification Head Adjustment
Xuanpu Zhao, Qi Chu 0001, Changtao Miao, Bin Liu 0016, Nenghai Yu |
ICIG (3) | 4 |
| 2023 | ABMNet: Coupling Transformer with CNN Based on Adams-Bashforth-Moulton Method for Infrared Small Target DetectionabstractInfrared small target detection (ISTD) aims at segmenting the small targets from infrared images, which has wide applications in military surveillance. Present methods are mainly based on CNN and focus on modelling locality while ignoring global dependencies, which are indispensable because the local areas similar to small targets always spread over most of the background, causing heavy target ambiguity. Recently, RKformer [1] has combined local features with global dependencies and further introduced Runge-Kutta method, a one-step Ordinary Differential Equation (ODE) solver, to ISTD and performed well. However, the method simply fuses features from original transformer and residual blocks by naive concatenation, causing insufficient feature interaction. Also, it inevitably brings effective information loss, which greatly impairs ambiguous target features. To address above problems and target ambiguity, we introduce Adams-Bashforth-Moulton method and propose ABMNet, which has (1) multi-step memory and self-rectification mechanisms, guaranteeing more sufficient information usage and more accurate detection, (2) and achieves more sufficient interaction of both local and global information. Experiments on MDFA and IRSTD-1k demonstrate the superiority of our method. Qi Chu 0001, Zhentao Tan, Bin Liu 0016, Nenghai Yu |
ICME | 4 |
| 2023 | Fluid Dynamics-Inspired Network for Infrared Small Target DetectionabstractMost infrared small target detection (ISTD) networks focus on building effective neural blocks or feature fusion modules but none describes the ISTD process from the image evolution perspective. The directional evolution of image pixels influenced by convolution, pooling and surrounding pixels is analogous to the movement of fluid elements constrained by surrounding variables ang particles. Inspired by this, we explore a novel research routine by abstracting the movement of pixels in the ISTD process as the flow of fluid in fluid dynamics (FD). Specifically, a new Fluid Dynamics-Inspired Network (FDI-Net) is devised for ISTD. Based on Taylor Central Difference (TCD) method, the TCD feature extraction block is designed, where convolution and Transformer structures are combined for local and global information. The pixel motion equation during the ISTD process is derived from the Navier–Stokes (N-S) equation, constructing a N-S Refinement Module that refines extracted features with edge details. Thus, the TCD feature extraction block determines the primary movement direction of pixels during detection, while the N-S Refinement Module corrects some skewed directions of the pixel stream to supplement the edge details. Experiments on IRSTD-1k and SIRST demonstrate that our method achieves SOTA performance in terms of evaluation metrics. Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
IJCAI | 3 |
| 2023 | Semantic Probability Distribution Modeling for Diverse Semantic Image SynthesisabstractSemantic image synthesis, translating semantic layouts to photo-realistic images, is a one-to-many mapping problem. Though impressive progress has been recently made, diverse semantic synthesis that can efficiently produce semantic-level or even instance-level multimodal results, still remains a challenge. In this article, we propose a novel diverse semantic image synthesis framework from the perspective of semantic class distributions, which naturally supports diverse generation at both semantics and instance level. We achieve this by modeling class-level conditional modulation parameters as continuous probability distributions instead of discrete values, and sampling per-instance modulation parameters through instance-adaptive stochastic sampling that is consistent across the network. Moreover, we propose prior noise remapping, through linear perturbation parameters encoded from paired references, to facilitate supervised training and exemplar-based instance style control at test time. To further extend the user interaction function of the proposed method, we also introduce sketches into the network. In addition, specially designed generator modules, Progressive Growing Module and Multi-Scale Refinement Module, can be used as a general module to improve the performance of complex scene generation. Extensive experiments on multiple datasets show that our method can achieve superior diversity and comparable quality compared to state-of-the-art methods. Codes are available at https://github.com/tzt101/INADE.git. Zhentao Tan, Qi Chu 0001, Menglei Chai, Dongdong Chen 0001, Jing Liao 0001, Qiankun Liu 0001, Bin Liu 0016, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | PhaseAnti: An Anti-Interference WiFi-Based Activity Recognition System Using Interference-Independent Phase ComponentabstractDriven by a wide range of essential applications, significant achievements have recently been made to explore WiFi-based Human Activity Recognition (HAR) techniques that utilize the information collected by commercial off-the-shelf (COTS) WiFi infrastructures to infer human activities without the need for the subject to carry any devices. Although existing WiFi-based HAR systems achieve satisfactory performance in some instances, they are faced with a severe challenge that the impacts of ubiquitous Co-channel Interference (CCI) on WiFi signals are inevitable. This downgrades the performance of these HAR systems significantly. To address this challenge, we propose PhaseAnti in this paper, a novel WiFi-based HAR system to exploit the CCI-independent phase component, Nonlinear Phase Error Variation (NLPEV), of WiFi Channel State Information (CSI) to cope with the negative effects of CCI. The stability of NLPEV data and the sensibility of this component to motions are rigorously analyzed. Furthermore, validated by extensive properly designed experiments, this phase component across subcarriers is invariant under various CCI scenarios while sufficiently distinct for different motions. Therefore, the NLPEV data can be used and processed effectively to perform HAR in CCI scenarios. Extensive experiments with various daily activities in different indoor rooms demonstrate the superior effectiveness and generalizability of the proposed PhaseAnti system under various CCI scenarios. Specifically, PhaseAnti achieves a$ 96.5\%$recognition accuracy rate (RAR) on average in different CCI scenarios, which can improve up to a$ 16.7\%$RAR compared with the amplitude component in the presence of CCI. Furthermore, the recognition speed is 10.3 × faster than the state-of-the-art solution. Jinyang Huang, Bin Liu 0016, Chenglin Miao, Yan Lu 0001, Qijia Zheng, Yu Wu 0020, Jiancun Liu, Lu Su 0001, Chang Wen Chen |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | AutoMA: Towards Automatic Model Augmentation for Transferable Adversarial AttacksabstractRecent adversarial attack works attempt to improve the transferability by applying various differentiable transformations on input images. Considering the differentiable transformations and the original model together as a new model, these methods can be regarded as model augmentation that effectively derives an ensemble of models from the single original model. Despite their impressive performance, the model augmentation policies used in these methods are manually designed by experimental attempts, leaving the design of model augmentation policy an open question. In this paper, we propose an Automatic Model Augmentation (AutoMA) approach to find a strong model augmentation policy for transferable adversarial attacks. Specifically, we design a discrete search space that contains various diffierentiable transformations with different parameters and adopt reinforcement learning to search for the strong augmentation policy. The sampled augmentation policies together with the rewards they obtain during the searching process reveal several valuable observations for designing more powerful attacks using model augmentation policy:1) Augmentation transformations on color space are less effective; 2) The transformation type diversity matters; and 3) Using small distortion for geometric transformations while larger distortion for intensity transformations.Extensive experiments show that the augmentation policy found by AutoMA achieves superior performance than existing manually designed policies in a wide range of cases. Qi Chu 0001, Feng Zhu 0006, Rui Zhao 0001, Bin Liu 0016, Nenghai Yu |
IEEE Trans. Multim. | 5 |
| 2022 | Affinity-Aware Relation Network for Oriented Object Detection in Aerial Images
Tingting Fang, Bin Liu 0016, Qi Chu 0001, Nenghai Yu |
ACCV (5) | 2 |
| 2022 | Counterfactual Intervention Feature Transfer for Visible-Infrared Person Re-identification
Xulin Li, Yan Lu 0001, Bin Liu 0016, Guojun Yin, Qi Chu 0001, Jinyang Huang, Feng Zhu 0006, Rui Zhao 0001, Nenghai Yu |
ECCV (26) | 3 |
| 2022 | Towards Intrinsic Common Discriminative Features Learning for Face Forgery Detection Using Adversarial LearningabstractExisting face forgery detection methods usually treat face forgery detection as a binary classification problem and adopt deep convolution neural networks to learn discriminative features. The ideal discriminative features should be only related to the real/fake labels of facial images. However, we observe that the features learned by vanilla classification networks are correlated to unnecessary properties, such as forgery methods and facial identities. Such phenomenon would limit forgery detection performance especially for the generalization ability. Motivated by this, we propose a novel method which utilizes adversarial learning to eliminate the negative effect of different forgery methods and facial identities, which helps classification network to learn intrinsic common discriminative features for face forgery detection. To leverage data lacking ground truth label of facial identities, we design a special identity discriminator based on similarity information derived from off-the-shelf face recognition model. Extensive experiments demonstrate the effectiveness of the proposed method under both intra-dataset and cross-dataset evaluation settings. Wanyi Zhuang, Qi Chu 0001, Changtao Miao, Bin Liu 0016, Nenghai Yu |
ICME | 5 |
| 2022 | Cloth-Aware Center Cluster Loss for Cloth-Changing Person Re-identification
Xulin Li, Bin Liu 0016, Yan Lu 0001, Qi Chu 0001, Nenghai Yu |
PRCV (1) | 2 |
| 2022 | Multi-view Geometry Distillation for Cloth-Changing Person ReID
Hanlei Yu, Bin Liu 0016, Yan Lu 0001, Qi Chu 0001, Nenghai Yu |
PRCV (1) | 2 |
| 2022 | Online multi-object tracking with unsupervised re-identification learning and occlusion estimation
Qiankun Liu 0001, Dongdong Chen 0001, Qi Chu 0001, Lu Yuan 0001, Bin Liu 0016, Lei Zhang 0001, Nenghai Yu |
Neurocomputing | 5 |
| 2021 | Joint Color-irrelevant Consistency Learning and Identity-aware Modality Adaptation for Visible-infrared Cross Modality Person Re-identificationabstractVisible-infrared cross modality person re-identification (VI-ReID) is a core but challenging technology in the 24-hours intelligent surveillance system. How to eliminate the large modality gap lies in the heart of VI-ReID. Conventional methods mainly focus on directly aligning the heterogeneous modalities into the same space. However, due to the unbalanced color information between the visible and infrared images, the features of visible images tend to overfit the clothing color information, which would be harmful to the modality alignment. Besides, these methods mainly align the heterogeneous feature distributions in dataset-level while ignoring the valuable identity information, which may cause the feature misalignment of some identities and weaken the discrimination of features. To tackle above problems, we propose a novel approach for VI-ReID. It learns the color-irrelevant features through the color-irrelevant consistency learning (CICL) and aligns the identity-level feature distributions by the identity-aware modality adaptation (IAMA). The CICL and IAMA are integrated into a joint learning framework and can promote each other. Extensive experiments on two popular datasets SYSU-MM01 and RegDB demonstrate the superiority and effectiveness of our approach against the state-of-the-art methods. Bin Liu 0016, Qi Chu 0001, Yan Lu 0001, Nenghai Yu |
AAAI | 2 |
| 2021 | Diverse Semantic Image Synthesis via Probability Distribution ModelingabstractSemantic image synthesis, translating semantic layouts to photo-realistic images, is a one-to-many mapping problem. Though impressive progress has been recently made, diverse semantic synthesis that can efficiently produce semantic-level multimodal results, still remains a challenge. In this paper, we propose a novel diverse semantic image synthesis framework from the perspective of semantic class distributions, which naturally supports diverse generation at semantic or even instance level. We achieve this by modeling class-level conditional modulation parameters as continuous probability distributions instead of discrete values, and sampling per-instance modulation parameters through instance-adaptive stochastic sampling that is consistent across the network. Moreover, we propose prior noise remapping, through linear perturbation parameters encoded from paired references, to facilitate supervised training and exemplar-based instance style control at test time. Extensive experiments on multiple datasets show that our method can achieve superior diversity and comparable quality compared to state-of-the-art methods. Code will be available at https://github.com/tzt101/INADE.git Zhentao Tan, Menglei Chai, Dongdong Chen 0001, Jing Liao 0001, Qi Chu 0001, Bin Liu 0016, Gang Hua 0001, Nenghai Yu |
CVPR | 6 |
| 2021 | ISNet: Integrate Image-Level and Semantic-Level Context for Semantic SegmentationabstractCo-occurrent visual pattern makes aggregating contextual information a common paradigm to enhance the pixel representation for semantic image segmentation. The existing approaches focus on modeling the context from the perspective of the whole image, i.e., aggregating the image-level contextual information. Despite impressive, these methods weaken the significance of the pixel representations of the same category, i.e., the semantic-level contextual information. To address this, this paper proposes to augment the pixel representations by aggregating the image-level and semantic-level contextual information, respectively. First, an image-level context module is designed to capture the contextual information for each pixel in the whole image. Second, we aggregate the representations of the same category for each pixel where the category regions are learned under the supervision of the ground-truth segmentation. Third, we compute the similarities between each pixel representation and the image-level contextual information, the semantic-level contextual information, respectively. At last, a pixel representation is augmented by weighted aggregating both the image-level contextual information and the semantic-level contextual information with the similarities as the weights. Integrating the image-level and semantic-level context allows this paper to report state-of-the-art accuracy on four benchmarks, i.e., ADE20K, LIP, COCOStuff and Cityscapes1. Zhenchao Jin, Bin Liu 0016, Qi Chu 0001, Nenghai Yu |
ICCV | 2 |
| 2021 | Improve Unsupervised Pretraining for Few-label TransferabstractUnsupervised pretraining has achieved great success and many recent works have shown unsupervised pretraining can achieve comparable or even slightly better transfer performance than supervised pretraining on downstream target datasets. But in this paper, we find this conclusion may not hold when the target dataset has very few labeled samples for finetuning, i.e., few-label transfer. We analyze the possible reason from the clustering perspective: 1) The clustering quality of target samples is of great importance to few-label transfer; 2) Though contrastive learning is essential to learn how to cluster, its clustering quality is still inferior to supervised pretraining due to lack of label supervision. Based on the analysis, we interestingly discover that only involving some unlabeled target domain into the unsupervised pretraining can improve the clustering quality, subsequently reducing the transfer performance gap with supervised pretraining. This finding also motivates us to propose a new progressive few-label transfer algorithm for real applications, which aims to maximize the transfer performance under a limited annotation budget. To support our analysis and proposed method, we conduct extensive experiments on nine different target datasets. Experimental results show our proposed method can significantly boost the few-label transfer performance of unsupervised pretraining. Suichan Li, Dongdong Chen 0001, Yinpeng Chen, Lu Yuan 0001, Lei Zhang 0001, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
ICCV | 7 |
| 2021 | Talking Face Video Generation with Editable Expression
Luchuan Song, Bin Liu 0016, Nenghai Yu |
ICIG (3) | 2 |
| 2021 | Towards More Powerful Multi-column Convolutional Network for Crowd Counting
Jiabin Zhang, Qi Chu 0001, Weihai Li, Bin Liu 0016, Weiming Zhang 0001, Nenghai Yu |
ICIG (1) | 4 |
| 2021 | Deepfake Video Detection Using 3D-Attentional Inception Convolutional Neural NetworkabstractThe current spike of deepfake techniques has received considerable attention due to security concerns. To mitigate the potential risks brought by deepfake techniques, many detection methods have been proposed. However, most existing works merely leverage spatial information from separate frames and ignore valuable inter-frame temporal information. In this paper, we propose a deepfake detection scheme that uses 3D-attentional inception network. The proposed model encompasses both spatial and temporal information simultaneously with the 3D kernels. Furthermore, the channel and spatial-temporal attention modules are applied to improve detection capabilities. Comprehensive experiments demonstrate that our scheme outperforms state-of-the-art methods. Changlei Lu, Bin Liu 0016, Wenbo Zhou 0004, Qi Chu 0001, Nenghai Yu |
ICIP | 2 |
| 2021 | Fsft-Net: Face Transfer Video Generation With Few-Shot ViewsabstractTo transfer head pose and expression with few photographs is a novel yet challenging task in deepfake generation. Despite impressive results have been achieved in related works, there are still two limitations in the existing methods: 1) most of the methods are based on computer graphics, which take a lot of computing resources, while lacking of generalization for different identity, 2) few-shot based methods cannot handle the few-shot style transfer video generation. To address these distortion problems, we propose a novel deep learning framework, named as Few-Shot Face Transfer Networks(FSFT-Net) which works for the face transfer video generation. The proposed FSFT-Net driven by arbitrary portrait video involves a cascaded-based style generator to synthesize stable video with few free-view images. In addition, the frame and video discriminators are adopted for optimization of the proposed generator. The FSFT-Net performs long-term adversarial training on large-scale video datasets. Extensive experiments demonstrate that our FSFT-Net outperforms state-of-the-art methods both quantitatively and qualitatively results. Luchuan Song, Guojun Yin, Bin Liu 0016, Nenghai Yu |
ICIP | 3 |
| 2021 | Content-Independent Online Handwriting Verification Based on Multi-Modal FusionabstractUser identity authentication is essencial for ensuring information security. With the widespread use of electronic devices, online handwriting verification becomes more important in identity authentication based on biometrics and widely used in financial, commercial, and forensic fields. In this paper, we propose a multi-path feature fusion network for multi-modal fusion of static and dynamic handwriting obtained by electronic devices to intensify the handwriting verification. Since traditional handwritten signature verification, of which the handwritten content just the writer’s name, is vulnerable to skilled forgery attacks, we propose a content-independent handwriting verification scheme to solve this problem. We also build a handwriting dataset with approximately 5400 samples of 30 individuals’ handwriting, which contributes to extracting content-independent handwriting style features. We test our method on widely used BiosecurID dataset and our dataset. The experimental results demonstrate the feasibility of the proposed method. Bin Liu 0016, Yan Lu 0001, Qi Chu 0001, Zhenchao Jin, Nenghai Yu |
ICME | 2 |
| 2021 | Efficient Open-Set Adversarial Attacks on Deep Face RecognitionabstractDifferent from close-set classification task, deep face recognition models are often used in open-set scenarios, where the models need to handle arbitrary faces. Open-set adversarial attacks can identify the vulnerability of deep face recognition models. Compared to time-consuming iterative gradient-based methods, generator-based methods can produce adversarial examples with only one forward pass, which greatly improves attack efficiency. However, existing generator-based attack methods need to train an individual model for each target identity and can only generate a fixed perturbation pattern regardless of different attack intensity constraints, which is impractical and sub-optimal for open-set adversarial attacks. In this paper, we propose an efficient generator-based Single Model ARbitrary Target (SMART) approach for open-set adversarial attacks against deep face recognition models. Given an arbitrary source-target face image pair, SMART first generates an additive perturbation and then adds it to the source image to obtain the final adversarial face image. After the training with various source-target pairs randomly sampled on large scale face images, SMART could effectively learn inherent perturbation patterns for arbitrary source-target face images pairs. Besides, we also propose a novel Constraint-aware Adversarial Decoder (CAD) module, which makes SMART the first generator-based method that could produce adaptive adversarial patterns according to different constraints on attack intensity. Extensive experimental results in various settings demonstrate the effectiveness of the proposed method. Qi Chu 0001, Feng Zhu 0006, Rui Zhao 0001, Bin Liu 0016, Nenghai Yu |
ICME | 5 |
| 2021 | I Know Your Keyboard Input: A Robust Keystroke Eavesdropper Based-on Acoustic SignalsabstractRecently, smart devices equipped with microphones have become increasingly popular in people's lives. However, when users type on a keyboard near devices with microphones, the acoustic signals generated by different keystrokes may leak the user's privacy. This paper proposes a robust side-channel attack scheme to infer keystrokes on the surrounding keyboard, leveraging the smart devices' microphones. To address the challenge of non-cooperative attacking environments, we propose an efficient scheme to estimate the relative position between the microphones and the keyboard, and extract two robust features from the acoustic signals to alleviate the impact of various victims and keyboards. As a result, we can realize the side-channel attack through acoustic signals, regardless of the exact location of microphones, the victims, and the type of keyboards. We implement the proposed scheme on the commercial smartphone and conduct extensive experiments to evaluate its performance. Experimental results show that the proposed scheme could achieve good performance in predicting keyboard input under various conditions. Overall, we can correctly identify 91.2% of keystrokes with 10-fold cross-validation. When predicting keystrokes from unknown victims, the attack can obtain a Top-5 accuracy of 91.52%. Furthermore, the Top-5 accuracy of predicting keystrokes can reach 72.25% when the victims and keyboards are both unknown. When predicting meaningful contents, we can obtain a Top-5 accuracy of 96.67% for the words entered by the victim. Jia-Xuan Bai, Bin Liu 0016, Luchuan Song |
ACM Multimedia | 2 |
| 2021 | TACR-Net: Editing on Deep Video and Voice PortraitsabstractUtilizing an arbitrary speech clip to edit the mouth of the portrait in the target video is a novel yet challenging task. Despite impressive results have been achieved, there are still three limitations in the existing methods: 1) since the acoustic features are not completely decoupled from person identity, there is no global speech to facial features (i.e., landmarks, expression blendshape) mapping method. 2) the audio-driven talking face sequences generated by simple cascade structure usually lack of temporal consistency and spatial correlation, which leads to defects in the consistency of changes in details. 3) the operation of forgery is always at the video level, without considering the forgery of the voice, especially the synchronization of the converted voice and the mouth. To address these distortion problems, we propose a novel deep learning framework, named Temporal-Refinement Autoregressive-Cascade Rendering Network (TACR-Net) for audio-driven dynamic talking face editing. The proposed TACR-Net encodes facial expression blendshape based on the given acoustic features without separately training for special video. Then TACR-Net also involves a novel autoregressive cascade structure generator for video re-rendering. Finally, we transform the in-the-wild speech to the target portrait and obtain a photo-realistic and audio-realistic video. Luchuan Song, Bin Liu 0016, Guojun Yin, Xiaoyi Dong, Yufei Zhang 0006, Jia-Xuan Bai |
ACM Multimedia | 2 |
| 2021 | WiLay: A Two-Layer Human Localization and Activity Recognition System Using WiFiabstractHuman activity monitoring (HAM) in the home environment has become increasingly important due to its broad applications including elder care, and well-being management. Recently, some state-of-the-art WiFi-based HAM systems have been proposed due to its properties of non-intrusive and privacy-friendly. However, their key drawback lies in ignoring the crucial impact of human position on HAM. To solve this problem, we present a two-layer WiFi-based HAM system (WiLay), which combines human activity recognition (HAR) with indoor human location (IHL) to provide more integrated information for HAM. Specifically, in the first layer, WiLay adopts the high-frequency energy (HFE) feature of WiFi signals to detect human moving. Then, in the second layer, different processing methods are employed for processing different types of motions accordingly. When the subject activities are static (SAs, the activity without position change), e.g., standing and sitting, WiLay locates the subject before recognizing the specific motion. On the contrary, when the activities are the moving activities (MAs), to reduce the loss of motion information, WiLay employs a comprehensive classifier generated by all different subcarrier classifiers voting, to recognize these MAs accurately. Extensive experimental results show that WiLay has high accuracy with a 99.9% SA/MA detection accuracy rate in the first layer, and a 99.7% location accuracy rate with 98.1% recognition performance for SAs and 90.2% recognition performance for MAs in the second layer. Jinyang Huang, Bin Liu 0016, Hongxin Jin, Nenghai Yu |
VTC Spring | 2 |
| 2021 | Real Time Video Object Segmentation in Compressed DomainabstractMany of the recent methods for semi-supervised video object segmentation are still far from being applicable for real time applications due to their slow inference speed. Therefore, we explore a propagation based segmentation method in compressed domain to accelerate inference speed in this paper. In particular, we only extract the features of I-frames by traditional deep convolutional neural network and produce the features of P-frames through information flow propagation. In the process of feature propagation, we propose two effective components to enhance the representation ability of simply warped features in terms of appearance and location. Specifically, we propose a residual supplement module to supplement appearance information which is lost in direct warping and a spatial attention module that can mine extra spatial saliency to provide the location information of the specified object. Besides, we propose a metric based decoder module which consists of a feature match module and a multi-level refinement module to transform information from semantic representation to shape segmentation mask. Extensive experiments on several video datasets demonstrate that the proposed method can achieve comparable accuracy while much faster inference speed when compared to the state-of-the-art algorithms. Zhentao Tan, Bin Liu 0016, Qi Chu 0001, Hangshi Zhong, Weihai Li, Nenghai Yu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | DASOT: A Unified Framework Integrating Data Association and Single Object Tracking for Online Multi-Object TrackingabstractIn this paper, we propose an online multi-object tracking (MOT) approach that integrates data association and single object tracking (SOT) with a unified convolutional network (ConvNet), named DASOTNet. The intuition behind integrating data association and SOT is that they can complement each other. Following Siamese network architecture, DASOTNet consists of the shared feature ConvNet, the data association branch and the SOT branch. Data association is treated as a special re-identification task and solved by learning discriminative features for different targets in the data association branch. To handle the problem that the computational cost of SOT grows intolerably as the number of tracked objects increases, we propose an efficient two-stage tracking method in the SOT branch, which utilizes the merits of correlation features and can simultaneously track all the existing targets within one forward propagation. With feature sharing and the interaction between them, data association branch and the SOT branch learn to better complement each other. Using a multi-task objective, the whole network can be trained end-to-end. Compared with state-of-the-art online MOT methods, our method is much faster while maintaining a comparable performance. Qi Chu 0001, Wanli Ouyang, Bin Liu 0016, Feng Zhu 0006, Nenghai Yu |
AAAI | 3 |
| 2020 | Density-Aware Graph for Deep Semi-Supervised Visual RecognitionabstractSemi-supervised learning (SSL) has been extensively studied to improve the generalization ability of deep neural networks for visual recognition. To involve the unlabelled data, most existing SSL methods are based on common density-based cluster assumption: samples lying in the same high-density region are likely to belong to the same class, including the methods performing consistency regularization or generating pseudo-labels for the unlabelled images. Despite their impressive performance, we argue three limitations exist: 1) Though the density information is demonstrated to be an important clue, they all use it in an implicit way and have not exploited it in depth. 2) For feature learning, they often learn the feature embedding based on the single data sample and ignore the neighborhood information. 3) For label-propagation based pseudo-label generation, it is often done offline and difficult to be end-to-end trained with feature learning. Motivated by these limitations, this paper proposes to solve the SSL problem by building a novel density-aware graph, based on which the neighborhood information can be easily leveraged and the feature learning and label propagation can also be trained in an end-to-end way. Specifically, we first propose a new Density-aware Neighborhood Aggregation(DNA) module to learn more discriminative features by incorporating the neighborhood information in a density-aware manner. Then a novel Density-ascending Path based Label Propagation(DPLP) module is proposed to generate the pseudo-labels for unlabeled samples more efficiently according to the feature distribution characterized by density. Finally, the DNA module and DPLP module evolve and improve each other end-to-end. Extensive experiments demonstrate the effectiveness of the newly proposed density-aware graph based SSL framework and our approach can outperform current state-of-the-art methods by a large margin. Suichan Li, Bin Liu 0016, Dongdong Chen 0001, Qi Chu 0001, Lu Yuan 0001, Nenghai Yu |
CVPR | 2 |
| 2020 | Cross-Modality Person Re-Identification With Shared-Specific Feature TransferabstractCross-modality person re-identification (cm-ReID) is a challenging but key technology for intelligent video analysis. Existing works mainly focus on learning modality-shared representation by embedding different modalities into a same feature space, lowering the upper bound of feature distinctiveness. In this paper, we tackle the above limitation by proposing a novel cross-modality shared-specific feature transfer algorithm (termed cm-SSFT) to explore the potential of both the modality-shared information and the modality-specific characteristics to boost the reidentification performance. We model the affinities of different modality samples according to the shared features and then transfer both shared and specific features among and across modalities. We also propose a complementary feature learning strategy including modality adaption, project adversarial learning and reconstruction enhancement to learn discriminative and complementary shared and specific features of each modality, respectively. The entire cmSSFTalgorithm can be trained in an end-to-end manner. We conducted comprehensive experiments to validate the superiority ofthe overall algorithm and the effectiveness ofeach component. The proposed algorithm significantly outperforms state-of-the-arts by 22.5% and 19.3% mAP on the two mainstream benchmark datasets SYSU-MM01 and RegDB, respectively. Yan Lu 0001, Bin Liu 0016, Tianzhu Zhang 0001, Baopu Li, Qi Chu 0001, Nenghai Yu |
CVPR | 3 |
| 2020 | Spatial-Temporal Feature Aggregation Network For Video Object DetectionabstractVideo object detection is a challenging problem in computer vision. In this paper, we propose a novel spatial-temporal feature aggregation network to deal with this issue. Specifically, we present a novel instance-level feature aggregation module as complementary to traditional pixel-level feature aggregation, in which we build a new movement estimation module to learn instance movements across frames. Then the Graph Convolutional Networks (GCNs) is applied to obtain temporal relation among instances over frames to implement instance-level feature aggregation. At last, we combine pixel-level and instance-level features by learnable soft weights to make use of their complementary information. Our framework is simple to implement and enables end-to-end training, which achieves state-of-art performance on the ImageNet VID dataset by extensive experiments. Weihai Li, Chi Fei, Bin Liu 0016, Nenghai Yu |
ICASSP | 4 |
| 2020 | GSM: Graph Similarity Model for Multi-Object TrackingabstractThe popular tracking-by-detection paradigm for multi-object tracking (MOT) focuses on solving data association problem, of which a robust similarity model lies in the heart. Most previous works make effort to improve feature representation for individual object while leaving the relations among objects less explored, which may be problematic in some complex scenarios. In this paper, we focus on leveraging the relations among objects to improve robustness of the similarity model. To this end, we propose a novel graph representation that takes both the feature of individual object and the relations among objects into consideration. Besides, a graph matching module is specially designed for the proposed graph representation to alleviate the impact of unreliable relations. With the help of the graph representation and the graph matching module, the proposed graph similarity model, named GSM, is more robust to the occlusion and the targets sharing similar appearance. We conduct extensive experiments on challenging MOT benchmarks and the experimental results demonstrate the effectiveness of the proposed method. Qiankun Liu 0001, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
IJCAI | 3 |
| 2020 | Towards Anti-interference WiFi-based Activity Recognition System Using Interference-Independent Phase ComponentabstractHuman activity recognition (HAR) has become increasingly essential due to its potential to support a broad array of applications, e.g., elder care, and VR games. Recently, some pioneer WiFi-based HAR systems have been proposed due to its privacy-friendly and device-free characteristics. However, their crucial limitation lies in ignoring the inevitable impact of co-channel interference (CCI), which degrades the performance of these HAR systems significantly. To address this challenge, we propose PhaseAnti, a novel HAR system to exploit the CCI- independent phase component, NLPEV (Nonlinear Phase Error Variation), of Channel State Information (CSI) to cope with the impact of CCI. We provide a rigorous analysis of NLPEV data with respect to its stability and otherness. Validated by our experiments, this phase component across subcarriers is invariant to various CCI scenarios, while different for distinct motions. Based on the analysis, we use NLPEV data to perform HAR in CCI scenarios. Extensive experiments demonstrate that PhaseAnti can reliably recognize activity in various CCI scenarios. Specifically, PhaseAnti achieves a 95% recognition accuracy rate (RAR) on average, which improves up to 16% RAR in the presence of CCI. Moreover, the recognition speed is 9× faster than the state-of-the-art solution. Jinyang Huang, Bin Liu 0016, Yu Wu 0020, Chi Zhang 0001, Nenghai Yu |
INFOCOM | 2 |
| 2020 | Efficient and privacy-preserving authentication scheme for wireless body area networks
Mengxia Shuai, Bin Liu 0016, Nenghai Yu, Ling Xiong, Changhui Wang |
J. Inf. Secur. Appl. | 2 |
| 2020 | SAFNet: A Semi-Anchor-Free Network With Enhanced Feature Pyramid for Object DetectionabstractIn recent years, the field of object detection has made significant progress. The success of most of the state-of-the-art object detectors is derived from the use of feature pyramid and the carefully designed anchor boxes. However, the current methods of constructing feature pyramid usually blindly integrate multi-scale representations on each feature hierarchy. Furthermore, these detectors also suffer from some drawbacks brought by the hand-designed anchors. To mitigate the adverse effects caused thereby, we introduce a one-stage object detector, named as the semi-anchor-free network with enhanced feature pyramid (SAFNet). Specifically, to better construct feature pyramid, we propose a novel enhanced feature pyramid generation paradigm, which mainly consists of two modules, i.e., adaptive feature fusion module (AFFM) and self-enhanced module (SEM). The paradigm adaptively integrates multi-scale representations in a non-linear method meanwhile suppress the redundant semantic information for each pyramid level, such that a clean and enhanced feature pyramid could be obtained. In addition, an adaptive anchor generator (AAG) is designed to yield fewer but more suitable anchor boxes for each input image. Benefiting from the enhanced feature pyramid, AAG is capable of generating more accurate anchor boxes by introducing few priors. Thus, AAG has the ability to alleviate the drawbacks caused by the preset anchor hyper-parameters and helps to decrease the computation cost. Extensive experiments demonstrate the effectiveness of our approach. Profited from the proposed modules, SAFNet significantly boosts the detection performance, i.e., achieving 2 points and 2.1 points higher Average Precision (AP) than RetinaNet (our baseline) on PASCAL VOC and MS COCO respectively. Codes will be publicly available soon. Zhenchao Jin, Bin Liu 0016, Qi Chu 0001, Nenghai Yu |
IEEE Trans. Image Process. | 2 |
| 2019 | A Channel Hopping Strategy Based on the Human Trajectory Similarity for WBANsabstractWireless Body Area Network (WBAN) has permeated in various fields, such as e-health, entertainment, sports and so on. However, Inter-Wban interference makes it rather difficult to ensure the reliability of data transmission. While a WBAN comes into the interference range of others, the collision is inevitable if they are working on the same channel. In this paper, we design a channel hopping strategy based on the human movement trajectory, which allows two WBANs with high meeting probability to hop to different channels, thus decreases the probability of interference. Furthermore, we design a more practical metric, which reflects the asynchronous nature of WBANs' channel hopping status, to measure the probability of two WBANs hopping to the same channel. In addition, the proposed strategy is energy efficiency and low latency without channel sensing or negotiating with other WBANs. Simulation results demonstrate the performance of the proposed strategy. Xiaoyu Zhang 0002, Bin Liu 0016 |
BSN | 2 |
| 2019 | WristPress: Hand Gesture Classification with two-array Wrist-Mounted pressure sensorsabstractThis paper presents a hand gesture recognition system WristPress based on only the pressure sensors, which can reflect the different pressure changes of different hand gestures. Two arrays of force sensitive resistors (FSRs) are arranged around the wrist to capture the pressure fluctuation with the subtle muscle and tendon movements of different gestures, which can help to identify similar gestures for achieving more functions. For distinguishing more gestures with similar muscle and tendon movements, the temporal features and the spatial features of pressures are selected and designed to characterize the relation of every tiny pressure changes corresponding to the muscle and tendon movements at different positions around the wrist. In the WristPress system, 24 kinds of one-gestures, which cover not only the finger movements but also rotations around the wrist and forearm, are classified with an overall 10-fold cross validation classification accuracy of 97.40%. In addition, the WristPress prototype is non-obtrusive with a small size, and is well suited to existing wearable device forms, such as smart watches and a bracelet that are already mounted on the wrist. Our study shows that the temporal features and the spatial features of these pressures can reflect the the correlation between different pressure sensors can improve the accuracy of the hand gesture classification, and the kNN classifier has the best classification accuracy performance 97.40% with a low time complexity. Yufei Zhang 0006, Bin Liu 0016, Jinyang Huang |
BSN | 2 |
| 2019 | Semantics Disentangling for Text-To-Image GenerationabstractSynthesizing photo-realistic images from text descriptions is a challenging problem. Previous studies have shown remarkable progresses on visual quality of the generated images. In this paper, we consider semantics from the input text descriptions in helping render photo-realistic images. However, diverse linguistic expressions pose challenges in extracting consistent semantics even they depict the same thing. To this end, we propose a novel photo-realistic text-to-image generation model that implicitly disentangles semantics to both fulfill the high-level semantic consistency and low-level semantic diversity. To be specific, we design (1) a Siamese mechanism in the discriminator to learn consistent high-level semantics, and (2) a visual-semantic embedding strategy by semantic-conditioned batch normalization to find diverse low-level semantics. Extensive experiments and ablation studies on CUB and MS-COCO datasets demonstrate the superiority of the proposed method in comparison to state-of-the-art methods. Guojun Yin, Bin Liu 0016, Lu Sheng, Nenghai Yu, Xiaogang Wang 0001 |
CVPR | 2 |
| 2019 | Context and Attribute Grounded Dense CaptioningabstractDense captioning aims at simultaneously localizing semantic regions and describing these regions-of-interest (ROIs) with short phrases or sentences in natural language. Previous studies have shown remarkable progresses, but they are often vulnerable to the aperture problem that a caption generated by the features inside one ROI lacks contextual coherence with its surrounding context in the input image. In this work, we investigate contextual reasoning based on multi-scale message propagations from the neighboring contents to the target ROIs. To this end, we design a novel end-to-end context and attribute grounded dense captioning framework consisting of 1) a contextual visual mining module and 2) a multi-level attribute grounded description generation module. Knowing that captions often co-occur with the linguistic attributes (such as who, what and where), we also incorporate an auxiliary supervision from hierarchical linguistic attributes to augment the distinctiveness of the learned captions. Extensive experiments and ablation studies on Visual Genome dataset demonstrate the superiority of the proposed model in comparison to state-of-the-art methods. Guojun Yin, Lu Sheng, Bin Liu 0016, Nenghai Yu, Xiaogang Wang 0001 |
CVPR | 3 |
| 2019 | Memory-Based Neighbourhood Embedding for Visual RecognitionabstractLearning discriminative image feature embeddings is of great importance to visual recognition. To achieve better feature embeddings, most current methods focus on designing different network structures or loss functions, and the estimated feature embeddings are usually only related to the input images. In this paper, we propose Memory-based Neighbourhood Embedding (MNE) to enhance a general CNN feature by considering its neighbourhood. The method aims to solve two critical problems, i.e., how to acquire more relevant neighbours in the network training and how to aggregate the neighbourhood information for a more discriminative embedding. We first augment an episodic memory module into the network, which can provide more relevant neighbours for both training and testing. Then the neighbours are organized in a tree graph with the target instance as the root node. The neighbourhood information is gradually aggregated to the root node in a bottom-up manner, and aggregation weights are supervised by the class relationships between the nodes. We apply MNE on image search and few shot learning tasks. Extensive ablation studies demonstrate the effectiveness of each component, and our method significantly outperforms the state-of-the-art approaches. Suichan Li, Dapeng Chen, Bin Liu 0016, Nenghai Yu, Rui Zhao 0001 |
ICCV | 3 |
| 2019 | Enhanced Video Segmentation with Object Tracking
Zheran Hong, Zhentao Tan, Qiankun Liu 0001, Bin Liu 0016, Nenghai Yu |
ICIG (1) | 5 |
| 2019 | Learning Cross Camera Invariant Features with CCSC Loss for Person Re-identification
Bin Liu 0016, Weihai Li, Nenghai Yu |
ICIG (1) | 2 |
| 2019 | Dhff: Robust Multi-Scale Person Search by Dynamic Hierarchical Feature FusionabstractPerson Search plays the role of the ultimate destination of person re-identification (re-ID) in real applications. It has many challenges that person re-ID doesn't need to handle, such as mis-detections, false alarms and multi-scale matching. In contrast to previous works, we show that a strong multi-scale person matching system can result in a good person search performance with a common deep object detector (e.g. Faster-RCNN). In this work, we provide a robust person search method called Dynamic Hierarchical Feature Fusion (DHFF) which is based on multi-level feature fusion to tackle with multi-scale matching. In addition, A Multi-Metric loss is proposed to train the model effectively and stably with numerous identities. We evaluate our method on two large person search benchmark data sets: CUHK-SYSU and PRW. Experiments show that the proposed algorithm outperforms other state-of-the-art person search methods. Yan Lu 0001, Zheran Hong, Bin Liu 0016, Weihai Li, Nenghai Yu |
ICIP | 3 |
| 2019 | Cascaded Residual Density Network for Crowd CountingabstractCrowd counting is a challenging task due to the issues such as scale variation and perspective variation in real crowd scenes. In this paper, we propose a novel Cascaded Residual Density Network (CRDNet) in a coarse-to-fine approach to generate the high-quality density map for crowd counting more accurately. (1) We estimate the residual density maps by multi-scale pyramidal features through cascaded residual density modules. It can improve the quality of density map layer by layer effectively. (2) A novel additional local count loss is presented to refine the accuracy of crowd counting, which reduces the errors of pixel-wise Euclidean loss by restricting the number of people in the local crowd areas. Experiments on two public benchmark datasets show that the proposed method achieves effective improvement compared with the state-of-the-art methods. Bin Liu 0016, Luchuan Song, Weihai Li, Nenghai Yu |
ICIP | 2 |
| 2019 | Real Time Compressed Video Object SegmentationabstractVideo object segmentation is a challenging task with wide variety of applications. Although recent CNN based methods have achieved great performance, they are far from being applicable for real time applications. In this paper, we propose a propagation based video object segmentation method in compressed domain to accelerate inference speed. We only extract features from I-frames by the traditional deep segmentation network. And the features of P-frames are propagated from I-frames. Apart from feature warping, we propose two effective modules in the process of feature propagation to ensure the representation ability of propagated features in terms of appearance and location. Residual supplement module is used to supplement appearance information lost in warping, and spatial attention module mines accurate spatial saliency prior to highlight the specified object. Compared with recent state-of-the-art algorithms, the proposed method achieves comparable accuracy while much faster inference speed. Zhengtao Tan, Bin Liu 0016, Weihai Li, Nenghai Yu |
ICME | 2 |
| 2019 | Tracking Assisted Faster Video Object DetectionabstractRecent approaches have achieved great success on still image object detection. Despite the high accuracy, directly applying image object detectors for video object detection is rather slow. Inspired from the fact that object tracking is much more efficient than object detection, we propose to combine object detection and tracking for fast video object detection. Computational expensive detection network is applied on sparsely arranged key frames, while proposals of non-key frames are obtained through tracking and regression of previous frame's proposals. Assisted with an adaptive key-frame arrangement module, our method can adaptively decide whether to track or to detect based on tracking quality. Extensive experiments show that the proposed method can significantly boost detection speed with a rather small drop in detection accuracy. Wenfei Yang, Bin Liu 0016, Weihai Li, Nenghai Yu |
ICME | 2 |
| 2019 | PPML: Metric Learning with Prior Probability for Video Object SegmentationabstractVideo object segmentation plays an important role in computer vision and has attracted much attention. Although many recent works have removed the fine-tuning process in pursuit of fast inference speed, while achieving high segmentation accuracy, they are still far from being real-time. In this paper, we regard this task as a feature matching problem and propose a prior probability based metric learning (PPML) method for faster inference speed and higher segmentation accuracy. The proposed method consists of two ingredients: a novel template space updating strategy that improves the efficiency of segmentation by avoiding the explosion of data in template space, and a novel feature matching method which applies more potential probability information through integrating the prior of the first frame and the predicted score of previous frames. Experimental results on DAVIS datasets demonstrate that the proposed method reaches the state-of-the-art competitive performance and is more efficient in time consumption. Hangshi Zhong, Zhentao Tan, Bin Liu 0016, Weihai Li, Nenghai Yu |
VCIP | 3 |
| 2019 | Using multi-label classification to improve object detection
Bin Liu 0016, Qi Chu 0001, Nenghai Yu |
Neurocomputing | 2 |
| 2019 | Object and patch based anomaly detection and localization in crowded scenes
Weihai Li, Bin Liu 0016, Nenghai Yu |
Multim. Tools Appl. | 3 |
| 2019 | Lightweight and Secure Three-Factor Authentication Scheme for Remote Patient Monitoring Using On-Body Wireless NetworksabstractOn-body wireless networks (oBWNs) play a crucial role in improving the ubiquitous healthcare services. Using oBWNs, the vital physiological information of the patient can be gathered from the wearable sensor nodes and accessed by the authorized user like the health professional or the doctor. Since the open nature of wireless communication and the sensitivity of physiological information, secure communication has always been the vital issue in oBWNs-based systems. In recent years, several authentication schemes have been proposed for remote patient monitoring. However, most of these schemes are so susceptible to security threats and not suitable for practical use. Specifically, all these schemes using lightweight cryptographic primitives fail to provide forward secrecy and suffer from the desynchronization attack. To overcome the historical security problems, in this paper, we present a lightweight and secure three-factor authentication scheme for remote patient monitoring using oBWNs. The proposed scheme adopts one-time hash chain technique to ensure forward secrecy, and the pseudonym identity method is employed to provide user anonymity and resist against desynchronization attack. The formal and informal security analyses demonstrate that the proposed scheme not only overcomes the security weaknesses in previous schemes but also provides more excellent security and functional features. The comparisons with six state-of-the-art schemes indicate that the proposed scheme is practical with acceptable computational and communication efficiency. Mengxia Shuai, Bin Liu 0016, Nenghai Yu, Ling Xiong |
Secur. Commun. Networks | 2 |
| 2018 | WristMouse: Wearable mouse controller based on pressure sensorsabstractIn this paper, we present a wearable mouse controller based on pressure sensors, which can recognize the hand gestures in real-time and translate them into the movements of the mouse. Only four force sensitive resistors ( FSRs) are placed around the wrist to capture the wrist pressure distribution of different hand gestures. A novel real-time recognition framework is proposed to accurately and quickly identify the starting and ending position of the gesture. In the proposed framework, a pressure-parameter adaptive updating strategy is designed for improving the robust of the system to cope with different hand positions, different users and re-wear scenarios. To ensure the real-time of the system, we design some simple features and classifiers to recognize hand gestures for decreasing the time complexity. In addition, the WristMouse prototype has a small volume, a low energy consumption and is easy to integrate with the normal wearable devices like smart watch and smart wristband. Our study shows that the proposed system has a good performance in the real-time hand gesture recognition with a high F-score of 92.55%, and the average time delay per second is less than 50ms for processing data. Yufei Zhang 0006, Bin Liu 0016 |
BSN | 2 |
| 2018 | Zoom-Net: Mining Deep Feature Interactions for Visual Relationship Recognition
Guojun Yin, Lu Sheng, Bin Liu 0016, Nenghai Yu, Xiaogang Wang 0001, Chen Change Loy |
ECCV (3) | 3 |
| 2018 | Object-Oriented Anomaly Detection in Surveillance VideosabstractDetecting and localizing anomalies in surveillance videos is an ongoing challenge. Most existing methods are patch or trajectory-based, which lack semantic understanding of scenes and may split targets into pieces. To handle this problem, this paper proposes a novel and effective algorithm by incorporating deep object detection and tracking with full utilization of spatial and temporal information. We propose a new dynamic image by fusing both appearance and motion information and feed it into object detection network, which can detect and classify objects precisely even in dim and crowd scenes. Based on the detected objects, we develop an effective and scale-insensitive feature, named histogram variance of optical flow angle (HVOFA), together with motion energy to find abnormal motion patterns. In order to further discover missing anomalies and reduce false detected ones, we conduct a post-processing step with abnormal object tracking. The proposed algorithm outperforms state-of-the-art methods on standard benchmarks. Weihai Li, Bin Liu 0016, Qiankun Liu 0001, Nenghai Yu |
ICASSP | 3 |
| 2018 | Pyramid Sub-Region Sensitive Network for Object DetectionabstractIn prevalent two-stage object detectors, ROI pooling or position sensitive ROI pooling (PS ROI pooling) is usually used to extract features of proposal. But ROI pooling or PS ROI pooling ignores the local or global information of proposal respectively. It motivates us to design a kind of pooling method which can capture both global information and local information of proposal. In this paper, we propose pyramid sub-region sensitive network (PSSNet) for object detection which uses pyramid sub-region sensitive ROI pooling (PSS ROI pooling) to extract features of proposal. The PSS ROI pooling can capture the global and coarse-to-fine local information of proposal. Then, we explore different weighting strategies to utilize the PSS ROI features using self-adapting learning factors. Our PSSNet achieves the state-of-art result on PASCAL VOC 2007, PASCAL VOC 2012 datasets and competitive result on MS COCO dataset. Bin Liu 0016, Weihai Li, Nenghai Yu |
ICIP | 2 |
| 2018 | Flow Guided Siamese Network for Visual TrackingabstractHow to effectively utilize the temporal information in video has been an important problem in visual tracking. In this paper, we try to address this problem from two aspects. At first, we use optical flow to take advantage of the inter-frame information, when we obtain the position of the last frame, we predict the approximate location in present frame by calculating the optical flow and generate samples around it. Secondly, we use the tracked patches as reference to identify the target better. We designed a Siamese network which take image pairs consisted of exemplars and samples as inputs, the objective function is also modified by adding a priori probability. Further, we conducted experiments on the OTB benchmark and achieve competitive result both on accuracy and robustness, which demonstrate the effectiveness of our proposed algorithm. Guokun Wang, Bin Liu 0016, Weihai Li, Nenghai Yu |
ICIP | 2 |
| 2018 | Robust Anomaly Detection via Fusion of Appearance and Motion FeaturesabstractAnomaly detection in crowded scenes is an important issue in computer vision. In this paper, we propose a novel framework which takes both appearance and motion characteristics into consideration to detect anomalies. A new foreground object localization method is put forward at first to extract object proposals. For motion representation, we present a novel local motion based descriptor named as Spatially Localized Multi-scale Histogram of Optical Flow (SL-MHOF) to capture the local motion statistics for each object proposal. For appearance representation, we apply convolutional neural networks (CNNs) because of their high visual discriminative capacities. These two features are then fed into Gaussian Mixture Model (GMM) Classifiers respectively to generate anomaly scores, which are fused with a softmax function to produce the final anomaly detection results. Experiments on UCSD datasets indicate the effectiveness of our proposed approach, which achieves state-of-the-art performance. Weihai Li, Chi Fei, Bin Liu 0016, Nenghai Yu |
VCIP | 4 |
| 2017 | Performance analysis for ZigBee under WiFi interference in smart homeabstractSmart home not only makes people's lives more convenient, but also saves energy and daily expenses for households. Both WiFi and ZigBee are widely deployed in the smart home and operated in the same 2.4GHz ISM band, which results in the coexistence interference. Since the transmitting power of WiFi devices is much higher than that of ZigBee devices, ZigBee is more susceptible to coexisting interference. In this paper, we propose an analytical model to evaluate the performance of ZigBee under WiFi interference in the practical smart home scenario. By considering both channel access behavior and path loss behavior of the WiFi and ZigBee devices, the proposed model is built by involving both the Markov chain model and the indoor path loss model, and the ZigBee performance under WiFi interference is theoretically derived. The simulation results demonstrate the effectiveness of the proposed model for evaluating the ZigBee performance under WiFi interference. Bin Liu 0016, Chang Wen Chen |
ICC | 2 |
| 2017 | Online Multi-object Tracking Using CNN-Based Single Object Tracker with Spatial-Temporal Attention MechanismabstractIn this paper, we propose a CNN-based framework for online MOT. This framework utilizes the merits of single object trackers in adapting appearance models and searching for target in the next frame. Simply applying single object tracker for MOT will encounter the problem in computational efficiency and drifted results caused by occlusion. Our framework achieves computational efficiency by sharing features and using ROI-Pooling to obtain individual features for each target. Some online learned target-specific CNN layers are used for adapting the appearance model for each target. In the framework, we introduce spatial-temporal attention mechanism (STAM) to handle the drift caused by occlusion and interaction among targets. The visibility map of the target is learned and used for inferring the spatial attention map. The spatial attention map is then applied to weight the features. Besides, the occlusion status can be estimated from the visibility map, which controls the online updating process via weighted loss on training samples with different occlusion statuses in different frames. It can be considered as temporal attention mechanism. The proposed algorithm achieves 34.3% and 46.0% in MOTA on challenging MOT15 and MOT16 benchmark dataset respectively. Qi Chu 0001, Wanli Ouyang, Hongsheng Li 0001, Xiaogang Wang 0001, Bin Liu 0016, Nenghai Yu |
ICCV | 5 |
| 2017 | TCCF: Tracking Based on Convolutional Neural Network and Correlation Filters
Qiankun Liu 0001, Bin Liu 0016, Nenghai Yu |
ICIG (1) | 2 |
| 2017 | PPEDNet: Pyramid Pooling Encoder-Decoder Network for Real-Time Semantic Segmentation
Zhentao Tan, Bin Liu 0016, Nenghai Yu |
ICIG (1) | 2 |
| 2017 | Deep Scale Feature for Visual Tracking
Wenyi Tang, Bin Liu 0016, Nenghai Yu |
ICIG (1) | 2 |
| 2017 | Key-Region Representation Learning for Anomaly Detection
Wenfei Yang, Bin Liu 0016, Nenghai Yu |
ICIG (1) | 2 |
| 2017 | A Structural Coupled-Layer Tracking Method Based on Correlation Filters
Bin Liu 0016, Chang Wen Chen |
MMM (1) | 2 |
| 2017 | Medium Access Control for Wireless Body Area Networks with QoS Provisioning and Energy Efficient DesignabstractWith the promising applications in e-Health and entertainment services, wireless body area network (WBAN) has attracted significant interest. One critical challenge for WBAN is to track and maintain the quality of service (QoS), e.g., delivery probability and latency, under the dynamic environment dictated by human mobility. Another important issue is to ensure the energy efficiency within such a resource-constrained network. In this paper, a new medium access control (MAC) protocol is proposed to tackle these two important challenges. We adopt a TDMA-based protocol and dynamically adjust the transmission order and transmission duration of the nodes based on channel status and application context of WBAN. The slot allocation is optimized by minimizing energy consumption of the nodes, subject to the delivery probability and throughput constraints. Moreover, we design a new synchronization scheme to reduce the synchronization overhead. Through developing an analytical model, we analyze how the protocol can adapt to different latency requirements in the healthcare monitoring service. Simulations results show that the proposed protocol outperforms CA-MAC and IEEE 802.15.6 MAC in terms of QoS and energy efficiency under extensive conditions. It also demonstrates more effective performance in highly heterogeneous WBAN. Bin Liu 0016, Zhisheng Yan, Chang Wen Chen |
IEEE Trans. Mob. Comput. | 1 |
| 2016 | An energy-efficient and QoS-effective resource allocation scheme in WBANsabstractWireless Body Area Networks (WBANs) represent one of the most promising networks to provide health applications for improving the quality of life, such as ubiquitous e-Health services and real-time health monitoring. The resource allocation of an energy-constrained, heterogeneous WBAN is a critical issue that should consider both energy efficiency and Quality of Service (QoS) requirements with the dynamic link characteristics, especially when the limited resource cannot satisfy the expected QoS requirements. In this paper, we propose an Energy-efficient and QoS-effective resource allocation that considers a mix-cost parameter characterizing both energy cost and QoS cost between attainable QoS support and QoS requirements. Based on the mix-cost parameter, we first formulate the resource allocation problem as a mixed integer nonlinear programming (MINP) for optimizing the transmission power, the transmission rate and allocated time slots for each sensor to minimize total mix-cost of the system. Then we propose a sub-optimal greedy resource allocation algorithm, which has a much lower complexity compared to exhaustive search. Simulation results demonstrate the advantage of the mix-cost parameter to evaluate energy efficiency and attainable QoS support, as well as verifying the effectiveness of the proposed resource allocation algorithm. Bin Liu 0016, Chang Wen Chen |
BSN | 2 |
| 2016 | Buffer-aware and QoS-effective resource allocation scheme in WBANsabstractWireless Body Area Network (WBAN) represents one of the most promising networks to provide health applications for improving the quality of life, such as ubiquitous e-Health services and real-time health monitoring. The resource allocation of an energy-constrained, heterogeneous WBAN is a critical issue that should consider both energy efficiency and Quality of Service (QoS) requirements with the dynamic link characteristics, especially when the limited resource cannot satisfy the expected QoS requirements. In this paper, a buffer aware Energy-efficient and QoS-effective resource allocation scheme is proposed in which the sensor queue buffer states, constraints of QoS metrics and the characteristics of dynamic links are considered. Specifically, a buffer aware sensor evaluation method is designed to dynamically evaluate the sensor state with considering the sensor buffer states for improving the system performance. We then formulate the resource allocation problem for optimizing the transmission power, the transmission rate and the allocated time slots for each sensor to minimize the sum mix-cost, which is defined to characterize the energy cost and QoS cost between attainable QoS support and QoS requirements. Simulation results demonstrate the effectiveness of the buffer aware sensor evaluation method and the proposed energy-efficient and QoS-effective resource allocation scheme. Bin Liu 0016, Chang Wen Chen |
HealthCom | 2 |
| 2016 | A two-layer and multi-strategy framework for human activity recognition using smartphoneabstractHuman Activity Recognition (HAR) is widely used in many applications and HAR using smartphone only has been proved to be effective, flexible and unobtrusive for activity recognition. In this paper, a two-layer and multi-strategy HAR framework is proposed to overcome the major challenge of HAR using smartphone only, i.e., the variation in orientation and position of the device. In the first layer, the activities are classified into different groups with high accuracy and for each group in the second layer, the appropriate strategy is designed according to the characteristics of the group to improve the recognition performance. For static activity group, the transitional activities are introduced to help classifying the activities indirectly. For dynamic activity group sensitive to the position variation of the smartphone, a position-assisted strategy is proposed to alleviate the influence of position variation. The simulation results demonstrate the effectiveness of the proposed two-layer multi-strategy HAR framework. Bin Liu 0016, Chang Wen Chen |
ICC | 2 |
| 2015 | Energy-Efficient Resource Allocation with QoS Support in Wireless Body Area NetworksabstractWireless Body Area Network (WBAN) has become a promising type of networks to provide applications such as real-time health monitoring and ubiquitous e-Health services. One challenge in the design of WBAN is that energy efficiency needs to be ensured to increase the network lifetime in such a resourceconstrained network. Another critical challenge for WBAN is that quality of service (QoS) requirements, including packet loss rate (PLR), throughput and delay, should be guaranteed even under the highly dynamic environment due to changing of body postures. In this paper, we design a unified framework of energy efficient resource allocation scheme for WBAN, in which both constraints of QoS metrics and the characteristics of dynamic links are considered. A transmission rate allocation policy (TRAP) is proposed to carefully adjust the transmission rate at each sensor such that more strict PLR requirement could be achieved even when the link quality is very poor. A QoS optimization problem is then formulated to optimize the transmission power and allocated time slots for each sensor, which minimizes energy consumption subject to the QoS constraints. Numerical results demonstrate the effectiveness of the proposed transmission rate allocation policy and the resource allocation scheme. Bin Liu 0016, Chang Wen Chen |
GLOBECOM | 2 |
| 2015 | QoS-Driven Power Control for Inter-WBAN Interference MitigationabstractWireless Body Area Networks (WBANs) are usually designed for pervasive healthcare applications. Since the primary traffic in WBAN is vital physiological signals, guaranteeing the Quality of Service (QoS) is crucial while designing WBAN. However, QoS of WBAN will be degraded in strong inter-WBAN interference environment such as hospitals and senior communities, where WBANs are densely deployed. In this paper, by focusing on a more practical WBAN model, we propose a non- cooperative power control game to mitigate inter- WBAN interference, in which the cost function is well designed by considering both QoS requirement and energy constraint. The existence of at least one Nash equilibrium (NE) point for the game is proved and a sufficient condition for the uniqueness of the NE is derived. To guarantee non-cooperative among WBANs, an interference segmentation estimate (ISE) algorithm is proposed to obtain an approximation of the NE point. Simulation results demonstrate the effectiveness of the proposed ISE algorithm. Xiaosong Zhao, Bin Liu 0016, Chang Wen Chen |
GLOBECOM | 2 |
| 2015 | Sequentially ordered backoff: Towards implicit resource reservation for wireless LANsabstractIn this paper, we present SOBO, a novel hybrid MAC protocol using sequentially ordered backoff in wireless LANs. SOBO eliminates packet collisions and wasted idle backoff slots by introducing implicit resource reservation into 802.11 DCF. In SOBO, the AP divides time into repeating cycles by beacon frames. Exploiting the implicit information of successful transmission order in every cycle, sequentially ordered backoff in a distributed manner during reservation period is achieved without extra control packets. In addition, we propose a novel scheme to estimate the number of contention stations, and design an adaptive contention window algorithm. We also analyze the robustness of SOBO against message losses in realistic networks with channel errors. The performance of SOBO is verified via extensive simulations with different scenarios. Our simulation results show that SOBO achieves a significant increase in network throughput compared to the legacy 802.11 DCF. Bing Feng, Chi Zhang 0001, Bin Liu 0016, Yuguang Fang |
ICC | 3 |
| 2015 | A Robust Occlusion Judgment Scheme for Target Tracking Under the Framework of Particle Filter
Kejia Liu, Bin Liu 0016, Nenghai Yu |
ICIG (1) | 2 |
| 2015 | A hierarchical anti-occlusion tracking algorithm based on DMPF and ORBabstractAn important issue in video target tracking is to deal with occlusion problem. In this paper, a hierarchical anti-occlusion tracking algorithm based on Dual Mode Particle Filter (DMPF) and Oriented FAST and Rotated BRIEF (ORB) is proposed to improve the location accuracy under different occlusion conditions in video target tracking. In the first layer, DMPF is used to track target and preliminarily locate its position under various occlusion status. By using corner matching in the second layer, the target is precisely located according to ORB similarity and corner coordinate, thus a more accurate position of target is obtained. A modularized occlusion judgment scheme is also presented to switch tracking mode timely and accurately in DMPF and a Quantified Matrix based on Weighted RGB color space (WQM) is introduced for both color feature creation and corner detection to save operation time. Simulation results show that the proposed algorithm could provide high tracking accuracy in a real-time manner for different occlusion status. Kejia Liu, Bin Liu 0016, Chang Wen Chen |
ICIP | 2 |
| 2014 | JSM-2 Based Joint ECG Compression Exploiting Temporal and Structural DependencyabstractIn Wireless Body Area Networks (WBAN), the electrocardiogram (ECG) signal is an important class of bio-signals which needs to be transmitted and stored for diseasesdiagnostics. Due to the resource limitation in WBAN, the large amount of ECG signals need to be compressed before transmission and reconstructed with high accuracy. In this paper, we propose a novel CS-based ECG compression scheme, which considers both structural dependency and temporal dependency among ECG signals. The received ECG heartbeats are first classified into different classes and the statistical support information (SSI) is then established for each class. By using the corresponding SSI, a more accurate partially known support (PKS) will be obtained and the joint reconstruction performance of ECG signals could be improved consequently. Simulation results show that the proposed ECG compression scheme outperforms existing schemes, especially when the dimension of sampled measurements is low. Jinguo Luo, Bin Liu 0016, Chang Wen Chen |
BSN | 2 |
| 2014 | Bayesian game based power control scheme for inter-WBAN interference mitigationabstractWireless body area network (WBAN) is an emerging technology that provides socialized health monitoring service. However, the quality of service can be severely degraded by concomitant inter-WBAN interference in some specific environments where multiple WBANs are densely deployed, e.g., hospitals and senior citizen communities. In this work, we propose a Bayesian game based power control scheme to mitigate the impact of inter-WBAN interference. By modeling WBANs as players and active links as types of players in the Bayesian game model, the proposed power control scheme tries to maximize each player's expected payoff involving both throughput and energy efficiency. We prove the existence of Bayesian equilibrium (BE) for the proposed power control game and also derive a practical sufficient condition for the uniqueness of BE. A harmonic mean based algorithm is then proposed to obtain an approximation of BE point without the need to pass message among WBANs, which satisfies the non-cooperative manner for inter-WBAN interference mitigation. Simulation results show that the proposed algorithm can converge to the BE point effectively. Bin Liu 0016, Chang Wen Chen |
GLOBECOM | 2 |
| 2014 | Admission Control for Wireless Adaptive HTTP Streaming: An Evidence Theory Based ApproachabstractIn this research, we propose an evidence theory based admission control scheme for wireless cellular adaptive HTTP streaming systems. This novel scheme allows us to effectively address the uncertainty and inaccuracy in QoE management and network estimation, and seamlessly grant or deny the access requests. Specifically, based on recent work of QoE continuum model and QoE continuum driven adaptation algorithm, we utilize Dempster-Shafer evidence theory to assign proper degree of belief to admission, rejection and an uncertainty decision for each user's evidence. We then can strategically combine the weighted evidence of multiple users and make the final decision. The evaluation results show that the proposed scheme can provide satisfactory QoE for both existing and new users while still achieving comparable bandwidth efficiency. Zhisheng Yan, Chang Wen Chen, Bin Liu 0016 |
ACM Multimedia | 3 |
| 2013 | JSM-2 based ECG compression with statistical support predictionabstractThis paper addresses the problem of developing an efficient compression scheme with high quality and low computational complexity for ECG signal compression. Taking into account the joint sparsity existing in ECG data and the temporal dependencies in ECG signal sequence, a novel scheme for JSM-2 based ECG compression is developed to exploit these characteristics. We first predict support information in sparse domain from the previous ECG data for the current recovery process. Then a modified Simultaneous Orthogonal Matching Pursuit Algorithm (SOMP) algorithm is proposed to incorporate the idea of support information establishment for JSM-2 based ECG compression. Simulation results show that the proposed JSM-2 based ECG compression scheme with statistical support prediction outperforms existing schemes with enhanced performance and low computational complexity. Sucheng Yu, Bin Liu 0016, Chi Zhang 0001, Chang Wen Chen |
Healthcom | 2 |
| 2013 | MOS-Based Channel Allocation Schemes for Mixed Services over Cognitive Radio NetworksabstractIn cognitive radio (CR) networks, secondary users (SUs) may have various applications such as multimedia delivery and file download, resulting in different bandwidth requirements. In this paper, we propose a channel allocation scheme for mixed services, especially video streaming, based on mean opinion score (MOS) maximization. MOS is an effective metric of Quality of Experience (QoE) that directly measures the satisfaction of the end users. The cognitive radio network base station (CRNBS) collects all the SUs' application information and allocates available channel resource to the SUs with the overall user perceived MOS maximized and fairness among SUs ensured. The simulation results confirm that the proposed MOS-based channel allocation scheme outperforms the conventional good put-based scheme in terms of overall user satisfaction. Bin Liu 0016, Yuanzhi Yao, Nenghai Yu, Chang Wen Chen |
ICIG | 2 |
| 2013 | Moving Object Detection for Moving Cameras on Superpixel LevelabstractIn this paper, we will present a novel algorithm for detecting moving objects from relatively dynamic background in particular video sequences. To achieve our goal, we introduced a new conception called "feature super pixel". As is well known, by tracking feature points we can easily tell whether a feature point belongs to moving object or not. However, it's more difficult to obtain an optimal pixel-wise foreground and background labeling. Considering the assumption that pixels in the same super pixel belong to the same object, this paper designed a method to describe and match super pixels, which has been proved to be efficient in experiments. Qiyu Liao, Nenghai Yu, Bin Liu 0016 |
ICIG | 3 |
| 2013 | High Precision Image Rotation Angle Estimation with Periodicity of Pixel VarianceabstractThe credibility of digital images is decreasing with the fast development of easy-to-use image manipulation tools such as Photoshop and Picasso. Although good forged images are deceptive to human eyes, the manipulation process often leaves some statistical traces in the forged images. Very often, the forged images are rotated partly or as a whole. If we can identify image rotation or even determine image rotation angle, the authenticity of images could be verified. In this paper, we propose an image rotation angle estimation algorithm based on properties of 2-D DFT of pixel variance. This algorithm is a blind one because it works in absence of any prior information such as digital watermark. We show the efficacy of this approach in our experiment. Ruohan Qian, Weihai Li, Nenghai Yu, Bin Liu 0016 |
ICIG | 4 |
| 2013 | Prediction-based dynamic relay transmission scheme for Wireless Body Area NetworksabstractTo support long-term pervasive healthcare services, communications in Wireless Body Area Networks (WBANs) need to be both reliable and energy-efficient. As a cooperative transmission method, relay transmission scheme works effectively in resisting shadowing effect and improving reliability in WBANs. However, the extra energy consumption introduced by relay transmission is very high, which can shorten the lifetime of the whole network. In this paper, temporal and spatial correlation models for on-body channels are first presented to better characterize the slow fading effect of on-body channels. Then a prediction-based dynamic relay transmission (PDRT) scheme that makes full use of the correlation characteristics of on-body channels is proposed. In the PDRT scheme, “when to relay” and “who to relay” are decided in an optimal way based on the last known channel states. Moreover, neither extra signaling procedure nor dedicated channel sensing period is needed. Simulation results show that the PDRT scheme achieves significant performance improvement in energy efficiency, as well as ensuring the transmission reliability. Bin Liu 0016, Zhisheng Yan, Chi Zhang 0001, Chang Wen Chen |
PIMRC | 2 |
| 2012 | JSM-2 based joint ECG compressed sensing with partially known support establishmentabstractCompressed sensing (CS) is a technique that enables sparse signal reconstruction from much fewer samples. In this paper, we propose ECG compressed sensing methods based on distributed compressed sensing to exploit the joint sparsity for both single- and multi-lead ECG signals. We apply JSM-2 (joint sparse model type 2) for jointly sparse ECG signals and formulate how to establish a partially known support based on this type of sparse model. Through careful analysis of joint partially known support, two-step ECG signal reconstruction schemes for single-lead and multi-lead ECG signals are developed. Simulation results show that the proposed schemes based on partially known support establishment outperforms existing schemes with enhanced performance measured by percentage root mean square difference (PRD). Bin Liu 0016, Chang Wen Chen |
Healthcom | 2 |
| 2012 | QoS-driven scheduling approach using optimal slot allocation for Wireless Body Area NetworksabstractWireless Body Area Network (WBAN) is a promising type of networks that mainly targets at applications in ubiquitous communication and e-Health services. Different from other types of networks, one important challenge for WBAN is that its quality of service (QoS) requirement, in terms of delivery probability and data rate, will be time varying since human body is a highly dynamic physical environment. Another significant challenge for WBAN is that energy efficiency needs to be guaranteed in such a resource-limited network. In this paper, a QoS-driven scheduling approach is proposed to address these challenges. We model the WBAN channel as a Markov model as suggested by the emerging IEEE 802.15.6 BAN standard and propose a threshold-based scheme to adjust the transmission order of nodes. The number of slots for each node is optimally assigned according to the QoS requirement while minimizing the energy consumption of nodes. The results from extensive simulations show that the proposed approach can provide high QoS and energy efficiency under different network conditions, especially in highly heterogeneous ones in WBAN. Zhisheng Yan, Bin Liu 0016, Chang Wen Chen |
Healthcom | 2 |
| 2012 | Block-based variable density compressed image samplingabstractCompressed sampling (CS) is a technique that enables signal reconstruction at sub-Nyquist sampling rate. A key problem in CS is how to design the sampling scheme. In this paper, we propose a novel sampling method for compressed image sampling, which exploits a priori information and uses a block-based strategy to improve image reconstruction. Our block-based sampling scheme assigns more samples to blocks with more high-frequency contents while making sure that important coefficients of each block are sampled. Simulation results show that our proposed method outperforms existing methods on both reconstruction quality and running time. Bin Liu 0016, Zixiang Xiong, Gonzalo R. Arce, Javier Garcia-Frías, Wenwu Zhu 0001, Zhisheng Yan |
ICIP | 2 |
| 2012 | Block-based compressed sampling with non-linear coding for image transmissionabstractWe propose a novel block-based image transmission system, which exploits the a prior information existing in the DCT domain of images and combines both linear and non-linear coding schemes accommodated to a block-based DCT domain compressed sampling method. An image is firstly divided into blocks and each block is separately sampled in DCT domain. Different coding schemes are used to transmit the samples based on their properties. With block-based strategy, each image block can be processed and transmitted separately, which reduces a lot of latency. Besides, an efficient system optimization algorithm is proposed by jointly optimizing the power allocation scheme and the transmission parameters to search for the maximum peak signal-to-noise ratio (PSNR) of the reconstructed image. Simulation results show that the proposed system provides a good performance with less latency. Bin Liu 0016, Zixiang Xiong, Gonzalo R. Arce, Javier Garcia-Frías |
MMSP | 1 |
| 2011 | A context aware MAC protocol for medical Wireless Body Area NetworkabstractDuring long-term medical monitoring in Wireless Body Area Networks (WBAN), network requirements (i.e. traffic loads and latency) of various data sources may be different at different time. High traffic loads may lead to data overload and unacceptable latency, which makes potential danger of patients undiagnosed. It is important that real-time transmission of life-critical data can be always guaranteed. To address this problem, a context-aware MAC protocol is presented in this paper. According to analysis of collected life parameters, the protocol can switch between normal state and emergency state. As a result, data rate and duty cycle of sensor nodes are dynamically changed to meet the requirement of latency and traffic loads in a contexta-ware way. To save the power consumption, a TDMA-based MAC frame structure is used. Moreover, a novel optional synchronization scheme is proposed to decrease the overhead caused by traditional TDMA synchronization scheme. Simulation results show significant improvements of our design on latency and power consumption. Zhisheng Yan, Bin Liu 0016 |
IWCMC | 2 |
| 2010 | SVD based linear filtering in DCT domainabstractEfficient linear filtering in DCT domain is important in the area of processing and manipulation of image and video streams compressed in DCT-based method. In this paper, we proposed a novel method for linear filtering in DCT domain, regardless of filter type. We decompose any filter by SVD into weighted separable sub-filters which are well studied. Then we do fast linear filtering using these separable subfilters in DCT domain, and combine their results. To our best knowledge, it is the first method capable to do linear filtering with any type of filters directly in DCT domain. The scheme is demonstrated and discussed by doing Gabor filtering in DCT domain. Experiment results show that convolution result using the proposed solution is the same as that in spatial domain. Furthermore, our scheme is well suitable for distributed computing, which will improve computing speed greatly. Liansheng Zhuang, Rui Zhao 0001, Nenghai Yu, Bin Liu 0016 |
ICIP | 4 |
| 2007 | Sensor fusion enhancement via optimized stochastic resonance at local sensorsabstractThis paper considers the decentralized fusion problem involving local sensor detection as well as the fusion of decisions transmitted over non-ideal transmission channels in a wireless sensor network. Prime emphasis is given to the enhancement of several fusion rules using a recently developed stochastic resonance methodology applied at the local sensors. Further, it is shown that the optimal form of the stochastic resonance probability mass density for the decentralized sensor fusion problem retains the same form as that previously developed for the single sensor case. Bin Liu 0016, Satish G. Iyengar, Hao Chen 0001, James H. Michels, Pramod K. Varshney |
FUSION | 1 |
| 2006 | Decentralized Detection in Wireless Sensor Networks with Channel Fading StatisticsabstractExisting channel aware signal processing design for decentralized detection in wireless sensor networks typically assumes the clairvoyant case, i.e., global information regarding the transmission channels is known at the design stage. In this paper, we consider the distributed detection problem where only the channel fading statistics, instead of the instant channel state information (CSI), is available to the designer. We investigate the design of local decision rules for the following two cases: 1. Fusion center has the instant CSI; 2. Fusion center does not have the instant CSI. We show that, for both cases, the optimal local decision rules that minimize the error probability at the fusion center amount to a likelihood-ratio test (LRT), as in the previous work with known CSI. The proposed approach enables distributed design for a decentralized detection problem. Bin Liu 0016, Biao Chen 0001 |
ICASSP (4) | 1 |
| 2006 | Channel-Optimized Quantizers for Decentralized Detection in Sensor NetworksabstractMotivated by the delay and resource constraints omnipresent in most wireless sensor network applications, we design channel-optimized scalar quantizers for a canonical decentralized detection system. Aimed at minimizing the error probability of the fusion center output, we first establish the optimality of monotone likelihood ratio partition of the observation space for the local quantizer design. We then devise an iterative algorithm to construct distributed quantizers that are person-by-person optimal. The channel-optimized approach is shown to offer better performance compared with various alternatives. It also exhibits inherent adaptivity in resource (bit) allocation in response to varying channel conditions. Bin Liu 0016, Biao Chen 0001 |
IEEE Trans. Inf. Theory | 1 |
| 2005 | Exploiting the finite-alphabet property for cooperative relaysabstractWe consider, in this paper, the design of a cooperative relay strategy by exploiting the finite-alphabet property of the source. Assuming a single source-sink pair with L relay nodes all communicating in orthogonal channels, we derive necessary conditions for optimal relay signaling that minimizes the error probability at the sink node. The derived conditions allow us to construct an iterative algorithm to find the distributed relay signaling that is at least locally optimal. As a byproduct, one can show that the so-called decode-and-forward (DF) relay scheme does not satisfy the necessary condition hence is not optimal in its error probability performance. Indeed, numerical examples show that the proposed scheme provides substantial performance improvement over both DF and the amplify-and-forward approach. Bin Liu 0016, Biao Chen 0001, Rick S. Blum |
ICASSP (3) | 1 |