Bin Chen 0006

dblp:22/5523-6 · DBLP profile ↗
← Back
38ranked-venue papers
13as first author
29since 2021 · last 2026
0000-0003-3503-2291ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 8 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AIPO: Adaptive Information Guided Token-Level Reinforcement Learning for Large Language Model Reasoning
abstract
Reinforcement Learning with Verifiable Rewards (RLVR) improves the reasoning capability of Large Language Models (LLMs). Current RLVR trains LLMs on all generated tokens, rather than exploring which tokens actually contribute to reasoning. We propose AIPO(Adaptive–Information Policy Optimization), which focuses updates on those decisive tokens discovered on the fly. AIPO estimates each hidden state’s mutual information to score tokens. Policy gradients are then computed only on these critical tokens, using an advantage that blends information gain and verifiable correctness. To improve the efficiency of mutual-information estimation, AIPO adopts a Random–Fourier approximation of the Hilbert–Schmidt Independence Criterion. Across five math and science benchmarks, AIPO yields up to +20% accuracy over strong RLVR baselines while updating merely 10% of tokens, demonstrating superior efficiency and effectiveness. Our findings highlight the importance of information–driven token selection for efficient and effective reinforcement learning of LLM reasoning.
Bin Chen 0006, Hongfei Ye, Wenxi Liu, Yu Zhang 0296, Furui Liu
ACL (1)1
2026 ParaSuite: Boosting LLM Reasoning via Paradox Resolution
abstract
Logical reasoning is a key capability of large language models, yet current benchmarks focus almost entirely on tasks that just check basic logical consistency and overlook the reflective reasoning required for paradox detection and resolution.To fill the gap, we present ParaSuite, the first pipeline dedicated to paradox research that automates data synthesis, evaluation, and training.We introduce PARADOX, a synthetic, high-quality data spanning two difficulty tiers and three academic domains, accompanied by specialized evaluation metrics and solving algorithms.We propose ParadoxBreaker-7B, trained with Mutual-Information Guided Fine-Tuning and reinforcement learning step verify paradox reward(PAPO).Experiments demonstrate significant improvements in both paradoxical and general STEM reasoning.
Bin Chen 0006, Yu Zhang 0296, Hongfei Ye, Wenxi Liu, Hongyang Chen 0001
ACL (1)1
2026 Assessing Color Vision Test in Large Vision-language Models
abstract
With the widespread adoption of large vision-language models, the capacity for color vision in these models is crucial. However, the color vision abilities of large visual-language models have not yet been thoroughly explored. To address this gap, we define a color vision testing task for large vision-language models and construct a dataset that covers multiple categories of test questions and tasks of varying difficulty levels. Furthermore, we analyze the types of errors made by large vision-language models and propose a chain-of-thought prompting strategy to enhance their performance in color vision tests.
Hongfei Ye, Bin Chen 0006, Wenxi Liu, Yu Zhang 0296, Zhao Li 0007, Dandan Ni, Hongyang Chen 0001
ICMR2
2026 ITGO: A general framework for text-guided image outpainting
Bin Chen 0006, Yuanbo Zhou, Xinlin Zhang, Yuanbin Chen, Qinquan Gao, Wenxi Liu, Tong Tong 0001
Expert Syst. Appl.1
2026 Multi-Agent Image Restoration
Gehui Li, Bin Chen 0006, Jian Zhang 0018
Int. J. Comput. Vis.3
2026 Covert Performance of Bidirectional Air-Water Cross-Boundary Optical Wireless Communication Systems
abstract
Optical wireless communication (OWC) has been recognized as a promising solution for establishing high-speed, low-latency data connectivity between air and water. During the propagation of optical beams across dynamic water surfaces, random scattering and refraction can result in information leakage issues, significantly increasing the probability of an adversary listener (Willie) detecting the signal transmission. In this paper, we investigate the covert performance of bidirectional air-water cross-boundary OWC systems under the detection of arbitrary underwater/aerial Willie for the first time. To characterize the detection accuracy at Willie and communication quality between a legitimate transmitter (Alice) and a legitimate receiver (Bob), the covert performance metrics including the covertness outage probability and covert throughput are computed based on the theoretical derivations of channel gain under dynamic wave, surface refraction, and particle absorption and scattering. The numerical simulation reveals the variation patterns of covert performance metrics with respect to multiple environmental and system parameters, which shows that wave intensity and transmitted beam’s divergence angle have opposite effects on detection performance of Willie locating at different positions, and both the optical wavelength beyond blue-green band and strong ambient light can increase Willie’s detection uncertainty. Furthermore, it can be seen that the upper bound on covertness outage probability reaches 1 and the achievable covert throughput approaches 0 in some cases, while covert threat regions exist in both atmospheric and underwater areas within several meters, which indicates that covert transmission strategy and system parameters are required to be further designed and optimized to enhance communication security.
Qingqing Hu, Bin Chen 0006, Chen Gong 0001, Murat Uysal
IEEE J. Sel. Areas Commun.2
2026 A Weakly Hybrid Decoding Algorithm for Staircase Codes via Multi-Level Bit Marking
abstract
Reducing the complexity of soft-decision (SD) decoding algorithm or improving the performance of hard-decision (HD) decoding algorithm becomes an emerging trend on forward error corrections to ensure reliable optical transmission with high performance and low cost. In this paper, a hybrid decoding algorithm is proposed for staircase codes (SCCs), which performs HD decoding with the help of channel soft information via multi-level bit-marking (MLBM). The HD and SD in the proposed MLBM decoding algorithm isweakly hybridthat the soft information is only used for one-shot bit marking and does not involve in the decoding directly, thus providing good performance-complexity tradeoff. The marked bits indicate four reliability levels that are used to prevent miscorrections by checking conflicts with highly reliable bits and to enhance the error-correcting capability of the HD decoder by flipping unreliable bits. Particularly, three priority levels are set for bit flipping by combining with the decoding statuses of the associated component codes. The decoding results will, in turn, help to update the marked bits among the four reliability levels. Numerical results show up to 1.22 dB gains with steeper waterfall region and lower error floor when compared to conventional HD decoding for SCCs, at the cost of approximately a two-fold increase in the memory storage.
Bin Chen 0006, Minghua Cao, Zhongyi Guo 0001
IEEE Trans. Commun.3
2026 GAN4RM: A CWGAN-Based Framework for Radio Maps Generation in Real Cellular Networks
abstract
With the evolution of mobile networks towards Artificial Intelligence as a Service (AIaaS), generative radio maps not only need to reflect the signal strength distribution in specific areas, but also possess the capability of proactive prediction. However, due to the rapid updates in urban infrastructure and the network iterations, crafting radio maps in complex urban environments represents a substantial challenge. In this paper, a multi-output framework for generating radio maps in real multi-building scenarios is proposed, based on Reference Signal Received Power (RSRP) and Reference Signal Received Quality (RSRQ) extracted from actual urban and suburban Measurement Reports (MRs). Specifically, An image encoding method integrating environmental features and base station system information is designed, while considering the sector antenna characteristics in actual communication environments. Then, a multi-output Conditional Wasserstein Generative Adversarial Network (CWGAN) is constructed for image conversion, and the radio maps are generated by learning the mapping from environmental & system information to RSRP & RSRQ radio maps, on the basis of image encoding that incorporates the physical laws of radio propagation. By calculating the priority of communication link gains at receiving points, it provides generative networks with reliable theoretical basis and conditional information, for serving cells and first neighboring cells. Experimental results show that the root mean square errors (RMSE) of the proposed method for RSRP / RSRQ of serving and neighboring cells are 1.7821 / 2.2251 and 0.8108 / 1.5121, which demonstrates the proposed method outperforms the baseline results. Simultaneously radio maps generation endows the cellular network with a certain “prophetic" capability, significantly enhancing the live service experience.
Lei Zhang 0139, Wanting Su, Jiawangnan Lu, Bin Chen 0006
IEEE Trans. Netw. Serv. Manag.5
2025 Adversarial Diffusion Compression for Real-World Image Super-Resolution
abstract
Real-world image super-resolution (Real-ISR) aims to reconstruct high-resolution images from low-resolution inputs degraded by complex, unknown processes. While many Stable Diffusion (SD)-based Real-ISR methods have achieved remarkable success, their slow, multi-step inference hinders practical deployment. Recent SD-based one-step networks like OSEDiff and S3Diff alleviate this issue but still incur high computational costs due to their reliance on large pretrained SD models. This paper proposes a novel Real-ISR method, AdcSR, by distilling the one-step diffusion network OSEDiff into a streamlined diffusion-GAN model under our Adversarial Diffusion Compression (ADC) framework. We meticulously examine the modules of OSEDiff, categorizing them into two types: (1) Removable (VAE encoder, prompt extractor, text encoder, etc.) and (2) Prunable (denoising UNet and VAE decoder). Since direct removal and pruning can degrade the model’s generation capability, we pretrain our pruned VAE decoder to restore its ability to decode images and employ adversarial distillation to compensate for performance loss. This ADC-based diffusion-GAN hybrid design effectively reduces complexity by 73% in inference time, 78% in computation, and 74% in parameters, while preserving the model’s generation capability. Experiments manifest that our proposed AdcSR achieves competitive recovery quality on both synthetic and real-world datasets, offering up to 9.3× speedup over previous one-step diffusion-based methods. Code and models are available at https://github.com/Guaishou74851/AdcSR.
Bin Chen 0006, Gehui Li, Rongyuan Wu, Jie Chen 0001, Jian Zhang 0018, Lei Zhang 0001
CVPR1
2025 OSMamba: Omnidirectional Spectral Mamba with Dual-Domain Prior Generator for Exposure Correction
abstract
Exposure correction is a fundamental problem in computer vision and image processing. Recently, frequency domainbased methods have achieved impressive improvement, yet they still struggle with complex real-world scenarios under extreme exposure conditions. This is due to the local convolutional receptive fields failing to model long-range dependencies in the spectrum, and the non-generative learning paradigm being inadequate for retrieving lost details from severely degraded regions. In this paper, we propose Omnidirectional Spectral Mamba (OSMamba), a novel exposure correction network that incorporates the advantages of state space models and generative diffusion models to address these limitations. Specifically, OSMamba introduces an omnidirectional spectral scanning mechanism that adapts Mamba to the frequency domain to capture comprehensive long-range dependencies in both the amplitude and phase spectra of deep image features, hence enhancing illumination correction and structure recovery. Furthermore, we develop a dual-domain prior generator that learns from well-exposed images to generate a degradation-free diffusion prior containing correct information about severely under- and over-exposed regions for better detail restoration. Extensive experiments on multiple-exposure and mixed-exposure datasets demonstrate that the proposed OSMamba achieves state-of-the-art performance both quantitatively and qualitatively. Our code and models can be found at https://github.com/cvsym/OSMamba.
Gehui Li, Bin Chen 0006, Chen Zhao 0002, Lei Zhang 0001, Jian Zhang 0018
CVPR2
2025 OmniGuard: Hybrid Manipulation Localization via Augmented Versatile Deep Image Watermarking
abstract
With the rapid growth of generative AI and its widespread application in image editing, new risks have emerged regarding the authenticity and integrity of digital content. Existing versatile watermarking approaches suffer from tradeoffs between tamper localization precision and visual quality. Constrained by the limited flexibility of previous framework, their localized watermark must remain fixed across all images. Under AIGC-editing, their copyright extraction accuracy is also unsatisfactory. To address these challenges, we propose OmniGuard, a novel augmented versatile watermarking approach that integrates proactive embedding with passive, blind extraction for robust copyright protection and tamper localization. OmniGuard employs a hybrid forensic framework that enables flexible localization watermark selection and introduces a degradation-aware tamper extraction network for precise localization under challenging conditions. Additionally, a lightweight AIGC-editing simulation layer is designed to enhance robustness across global and local editing. Extensive experiments show that OmniGuard achieves superior fidelity, robustness, and flexibility. Compared to the recent state-of-the-art approach EditGuard, our method outperforms it by 4.25dB in PSNR of the container image, 20.7% in F1-Score under noisy conditions, and 14.8% in average bit accuracy.
Xuanyu Zhang 0003, Zecheng Tang, Zhipei Xu, Runyi Li, Youmin Xu, Bin Chen 0006, Jian Zhang 0018
CVPR6
2025 Text-guided Multimodal Fusion for the Multimodal Emotion and Intent Joint Understanding
abstract
Emotion and Intent Joint Understanding in Multi-modal Conversation is a challenging task in the field of affective computing, aiming to decode the semantic information manifested in the multimodal conversational while simultaneously inferring the emotions and intents of the utterance. To address this challenge, we propose the Text-guided Multimodal Emotion-Intent Joint Recognition method. By leveraging the text modality to guide the fusion process, it effectively reduces the noise introduced by other modalities. To strengthen the text modality’s guiding role, we use large language models (LLMs) for multi-turn targeted data augmentation and oversampling strategies to address data imbalance. Our approach achieved first place in Track 1 (English) of the ICASSP 2025 MEIJU Challenge, demonstrating its effectiveness in practical applications.
Yu Zhang 0133, Bin Chen 0006, Hongfei Ye, Zijian Gao, Tianjiao Wan, Long Lan, Kele Xu
ICASSP2
2025 Self-supervised Scalable Deep Compressed Sensing
Bin Chen 0006, Xuanyu Zhang 0003, Yongbing Zhang 0002, Jian Zhang 0018
Int. J. Comput. Vis.1
2025 On Shaping Gain of Multidimensional Constellations in Linear and Nonlinear Optical Fiber Channel
abstract
Utilizing the multi-dimensional (MD) space for constellation shaping has been proven to be an effective approach for achieving shaping gains. Despite there exists a variety of MD modulation formats tailored for specific optical transmission scenarios, there remains a notable absence of a dependable comparison method for efficiently and promptly re-evaluating their performance in arbitrary transmission systems. In this paper, we introduce an analytical nonlinear interference (NLI) power model-based shaping gain estimation method to enable a fast performance evaluation of various MD modulation formats in coherent dual-polarization (DP) optical transmission system. In order to extend the applicability of this method to a broader set of modulation formats, we extend the established NLI model to take the 4D joint distribution into account and thus able to analyze the complex interactions of non-iid signaling in DP systems. With the help of the NLI model, we conduct a comprehensive analysis of the state-of-the-art modulation formats and investigate their actual shaping gains in two types of optical fiber communication scenarios (multi-span and single-span). The numerical simulation shows that for arbitrary modulation formats, the NLI power and relative shaping gains in terms of signal-to-noise ratio can be more accurately estimated by capturing the statistics of MD symbols. Furthermore, the proposed method further validates the effectiveness of the reported NLI-tolerant modulation format in the literature, which reveals that the linear shaping gains and modulation-dependent NLI should be jointly considered for nonlinearity mitigation.
Bin Chen 0006, Jingxin Deng, Shen Li 0006, Gabriele Liga
IEEE J. Sel. Areas Commun.1
2025 Practical Compact Deep Compressed Sensing
abstract
Recent years have witnessed the success of deep networks in compressed sensing (CS), which allows for a significant reduction in sampling cost and has gained growing attention since its inception. In this paper, we propose a new practical and compact network dubbed PCNet for general image CS. Specifically, in PCNet, a novel collaborative sampling operator is designed, which consists of a deep conditional filtering step and a dual-branch fast sampling step. The former learns an implicit representation of a linear transformation matrix into a few convolutions and first performs adaptive local filtering on the input image, while the latter then uses a discrete cosine transform and a scrambled block-diagonal Gaussian matrix to generate under-sampled measurements. Our PCNet is equipped with an enhanced proximal gradient descent algorithm-unrolled network for reconstruction. It offers flexibility, interpretability, and strong recovery performance for arbitrary sampling rates once trained. Additionally, we provide a deployment-oriented extraction scheme for single-pixel CS imaging systems, which allows for the convenient conversion of any linear sampling operator to its matrix form to be loaded onto hardware like digital micro-mirror devices. Extensive experiments on natural image CS, quantized CS, and self-supervised CS demonstrate the superior reconstruction accuracy and generalization ability of PCNet compared to existing state-of-the-art methods, particularly for high-resolution images. Code is available at https://github.com/Guaishou74851/PCNet.
Bin Chen 0006, Jian Zhang 0018
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Invertible Diffusion Models for Compressed Sensing
abstract
While deep neural networks (NNs) significantly advance image compressed sensing (CS) by improving reconstruction quality, the necessity of training current CS NNs from scratch constrains their effectiveness and hampers rapid deployment. Although recent methods utilize pre-trained diffusion models for image reconstruction, they struggle with slow inference and restricted adaptability to CS. To tackle these challenges, this paper proposes Invertible Diffusion Models (IDM), a novel efficient, end-to-end diffusion-based CS method. IDM repurposes a large-scale diffusion sampling process as a reconstruction model, and fine-tunes it end-to-end to recover original images directly from CS measurements, moving beyond the traditional paradigm of one-step noise estimation learning. To enable such memory-intensive end-to-end fine-tuning, we propose a novel two-level invertible design to transform both 1) multi-step sampling process and 2) noise estimation U-Net in each step into invertible networks. As a result, most intermediate features are cleared during training to reduce up to 93.8% GPU memory. In addition, we develop a set of lightweight modules to inject measurements into noise estimator to further facilitate reconstruction. Experiments demonstrate that IDM outperforms existing state-of-the-art CS networks by up to 2.64 dB in PSNR. Compared to the recent diffusion-based approach DDNM, our IDM achieves up to 10.09 dB PSNR gain and 14.54 times faster inference.
Bin Chen 0006, Zhenyu Zhang 0005, Chen Zhao 0002, Jiwen Yu, Shijie Zhao 0001, Jie Chen 0001, Jian Zhang 0018
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Flow Adjustment and Scale-Free Reliable Topology Control for Underwater Acoustic Sensor Networks
abstract
To improve the reliability and efficiency of underwater acoustic sensor networks (UASNs) under limited node energy and network attacks, a joint scheme of enhancing robustness and extending network lifetime is proposed in this paper. Scale-free networks exhibit strong robustness against random failures, and because of the lower number of link connections, nodes consume energy more slowly. We first introduce a new scale-free topology evolution model that adjusts the initial number of nodes according to the characteristics of UASNs. This model incorporates factors such as flow load, energy consumption, and distance into the preferential attachment mechanism, balancing the network load and enhancing its resistance to attacks. Further, based on this topology, we propose a network flow adjustment algorithm that features the joint selection of paths and corresponding power levels. Energy consumption is balanced among nodes in proportion to their residual energy, rather than by minimizing the absolute consumed power. Numerical simulations show that different topologies significantly affect network lifetime, which is the longest when the number of links added is 2. The network lifetime of the proposed scheme surpasses the state-of-the-art schemes, such as ETFLA, Initial-BA, and VODA, by up to 37.18%, 46.23%, and 128.80%, respectively.
Na Xia, Bin Chen 0006, Yutao Yin, Lei Chen 0081, Sizhou Wei, Ke Zhang 0034
IEEE Trans. Netw. Serv. Manag.3
2024 ResVR: Joint Rescaling and Viewport Rendering of Omnidirectional Images
abstract
With the advent of virtual reality technology, omnidirectional image (ODI) rescaling techniques are increasingly embraced to reduce transmitted and stored file sizes while preserving high image quality. Despite this progress, current ODI rescaling methods predominantly focus on enhancing the quality of images in equirectangular projection (ERP) format, which overlooks the fact that the content viewed on head-mounted displays (HMDs) is actually a rendered viewport instead of an ERP image. In this work, we emphasize that focusing solely on ERP quality results in inferior viewport visual experiences for users. Thus, we propose ResVR, which is the first comprehensive framework for the joint Rescaling and Viewport Rendering of ODIs. ResVR allows obtaining LR ERP images for transmission while rendering high-quality viewports for users to watch on HMDs. In our ResVR, a novel discrete pixel sampling strategy is developed to tackle the complex mapping between the viewport and ERP, enabling end-to-end training of the ResVR pipeline. Furthermore, a spherical pixel shape representation technique is innovatively derived from spherical differentiation to significantly improve the visual quality of rendered viewports. Extensive experiments demonstrate that our ResVR outperforms existing methods in viewport rendering tasks across different fields of view, resolutions, and view directions while keeping a low transmission overhead. Code is available at https://github.com/lwq20020127/ResVR.
Shijie Zhao 0001, Bin Chen 0006, Xinhua Cheng, Li Zhang 0006, Jian Zhang 0018
ACM Multimedia3
2024 D3C2-Net: Dual-Domain Deep Convolutional Coding Network for Compressive Sensing
abstract
By mapping iterative optimization algorithms into neural networks (NNs), deep unfolding networks (DUNs) exhibit well-defined and interpretable structures and achieve remarkable success in the field of compressive sensing (CS). However, most existing DUNs solely rely on the image-domain unfolding, which restricts the information transmission capacity and reconstruction flexibility, leading to their loss of image details and unsatisfactory performance. To overcome these limitations, this paper develops a dual-domain optimization framework that combines the priors of (1) image- and (2) convolutional-coding-domains and offers generality to CS and other inverse imaging tasks. By converting this optimization framework into deep NN structures, we present a Dual-Domain Deep Convolutional Coding Network (D3C2-Net), which enjoys the ability to efficiently transmit high-capacity self-adaptive convolutional features across all its unfolded stages. Our theoretical analyses and experiments on simulated and real captured data, covering 2D and 3D natural, medical, and scientific signals, demonstrate the effectiveness, practicality, superior performance, and generalization ability of our method over other competing approaches and its significant potential in achieving a balance among accuracy, complexity, and interpretability. Code is available at https://github.com/lwq20020127/D3C2-Net.
Bin Chen 0006, Shijie Zhao 0001, Bowen Du 0002, Yongbing Zhang 0002, Jian Zhang 0018
IEEE Trans. Circuits Syst. Video Technol.2
2024 Progressive Content-Aware Coded Hyperspectral Snapshot Compressive Imaging
abstract
Hyperspectral imaging plays a pivotal role across diverse applications, like remote sensing, medicine, and cytology. The utilization of 2D sensors to acquire 3D hyperspectral images (HSIs) via a coded aperture snapshot spectral imaging (CASSI) system has proven successful, owing to its hardware-friendly implementation and fast sampling speed. Nevertheless, for less spectrally sparse scenes, the use of a single snapshot and unreasonable coded aperture design limits the efficacy of CASSI systems and renders HSI reconstruction more ill-posed, leading to compromised spatial and spectral fidelity. This paper proposes a novel Progressive Content-Aware CASSI (PCA-CASSI) framework, which progressively captures HSIs using multiple optimized content-aware coded apertures and fuses all snapshot measurements for reconstruction. By unlocking the full potential of CASSI systems and elevating their performance ceilings, this framework offers researchers new avenues for improving imaging quality. Furthermore, we develop the RndHRNet, a Range-Null space Decomposition (RND)-inspired deep unfolding network with multiple iterative phases for HSI recovery. Each unfolded recovery phase efficiently exploits the physical information within the coded apertures via explicit RND and adaptively explores the spatial-spectral correlation by dual transformer blocks. Through comprehensive experiments, our approach demonstrates superior performance compared to existing state-of-the-art methods in both the multiple- and single-shot compressive HSI imaging tasks with substantial improvements. Code is available athttps://github.com/xuanyuzhang21/PCA-CASSI.
Xuanyu Zhang 0003, Bin Chen 0006, Wenzhen Zou, Yongbing Zhang 0002, Ruiqin Xiong, Jian Zhang 0018
IEEE Trans. Circuits Syst. Video Technol.2
2024 A Location-Independent Human Activity Recognition Method Based on CSI: System, Architecture, Implementation
abstract
In the application of human activity recognition (HAR) based on channel state information (CSI), due to the high dynamic characteristics of wireless channel to different environments, the features of human activity samples in different locations are different. In addition, the existing CSI-based HAR approaches limit the extraction of activity features to the Euclidean space and ignores the rich relational information between samples, categories and locations, which result in insufficient generalization performance for location-independent HAR. To address this challenge, this paper proposes a CSI-based location-independent HAR system CSI-MTGN. The system represents the classification task under each training sample collection location (TSCL) as a task, which is composed of three interactive parts: sample hidden representation, activity features extraction based on hierarchical graph neural network (HGNN) and information exchange based on multi-task learning. The proposed system improves the sample hidden representation, which is benefit for activity feature extraction and classification. The HGNN is designed to express various relationship information between samples, categories and locations in the form of graph structure, and the classification task under each TSCL is constructed through data augmentation, so as to improve the knowledge understanding and inference capabilities of the recognition model. The multi-task learning is used to achieve implicit data augmentation by sharing parameters among tasks through soft parameter sharing, and improves the generalization performance of the system. To validate the performance of the proposed system, experiments were conducted in a hall and a conference room, where samples of 10 categories of activities under 7 TSCLs were used for training the system, and the HAR accuracy rates at any locations were 94.1% and 93.3%, respectively.
Yong Zhang 0044, Andong Cheng, Bin Chen 0006, Yujie Wang 0002
IEEE Trans. Mob. Comput.3
2023 Deep Physics-Guided Unrolling Generalization for Compressed Sensing
Bin Chen 0006, Jiechong Song, Jingfen Xie, Jian Zhang 0018
Int. J. Comput. Vis.1
2023 Deep Memory-Augmented Proximal Unrolling Network for Compressive Sensing
Jiechong Song, Bin Chen 0006, Jian Zhang 0018
Int. J. Comput. Vis.2
2023 MPGSE-D-LinkNet: Multiple-Parameters-Guided Squeeze-and-Excitation Integrated D-LinkNet for Road Extraction in Remote Sensing Imagery
abstract
In road extraction task, traditional Squeeze-and-Excitation (SE) module only calculates the mean value of each channel to represent the salient features of the roads, but it easily causes false detection due to the interference such as water, roofs, and so on. This letter specifically proposes a Multiple-Parameters-Guided Squeeze-and-Excitation (MPGSE) module for road extraction by incorporating two key parameters of the variance, and the coefficient of variation into the SE module. Further, MPGSE module adaptively adjusts the weights of different features to suppress the redundant information while enhancing the informative features, which makes the roads more separable from other disturbances. MPGSE greatly increases the between-class distance and decrease the within-class distance, thus enhancing the separation capability of the road features compared with other interference. In addition, MPGSE module is integrated into D-LinkNet to optimally fuse features, thus further improving the completeness of road feature representation. Undoubtedly, MPGSE-D-LinkNet can achieve better road extraction performance than other methods. The superiority of MPGSE-D-LinkNet is verified on the RoadNet benchmark dataset (RNBD) and Massachusetts road dataset.
Jiaqiu Ai, Shaofan Hou, Bin Chen 0006
IEEE Geosci. Remote. Sens. Lett.4
2023 Dynamic Path-Controllable Deep Unfolding Network for Compressive Sensing
abstract
Deep unfolding network (DUN) that unfolds the optimization algorithm into a deep neural network has achieved great success in compressive sensing (CS) due to its good interpretability and high performance. Each stage in DUN corresponds to one iteration in optimization. At the test time, all the sampling images generally need to be processed by all stages, which comes at a price of computation burden and is also unnecessary for the images whose contents are easier to restore. In this paper, we focus on CS reconstruction and propose a novel Dynamic Path-Controllable Deep Unfolding Network (DPC-DUN). DPC-DUN with our designed path-controllable selector can dynamically select a rapid and appropriate route for each image and is slimmable by regulating different performance-complexity tradeoffs. Extensive experiments show that our DPC-DUN is highly flexible and can provide excellent performance and dynamic adjustment to get a suitable tradeoff, thus addressing the main requirements to become appealing in practice. Codes are available at https://github.com/songjiechong/DPC-DUN.
Jiechong Song, Bin Chen 0006, Jian Zhang 0018
IEEE Trans. Image Process.2
2023 IMF2O2: A Fully Connected Sensor Deployment Algorithm for Underwater Sensor Networks
abstract
To address the problems of node deployment schemes in existing underwater sensor networks that lack consideration of network connectivity and high deployment costs, this article constructs an optimization model that maximizes network coverage and minimizes deployment costs while ensuring full connectivity. For the NP-hard property of this optimization model, an improved moth flame optimization node deployment algorithm based on fuzzy operators (IMF 2 O 2 ) is proposed. First, comprehensively considering the two performance metrics of network coverage and network connectivity, a multi-objective selection mechanism based on fuzzy operators is proposed to improve network coverage while ensuring full connectivity. Second, a fixed number of nodes are used to monitor the target event points, transforming the node deployment of sensors into an optimal problem and proposing an improved moth flame optimization algorithm to solve this problem. Finally, the two metrics of coverage and deployment cost are measured and the fuzzy operator is used to select the optimal number of nodes to be deployed. Numerical results showed that the proposed algorithm improved network coverage rate by 10%, 22%, and 25%, and improved network connectivity rate by 12%, 20%, and 8% as compared to PSSD, RAWS, and VODA, respectively, while ensuring full connectivity.
Na Xia, Bin Chen 0006, Huazheng Du, Chaonong Xu, Rong Zheng 0001
ACM Trans. Sens. Networks3
2022 Content-Aware Scalable Deep Compressed Sensing
abstract
To more efficiently address image compressed sensing (CS) problems, we present a novel content-aware scalable network dubbed CASNet which collectively achieves adaptive sampling rate allocation, fine granular scalability and high-quality reconstruction. We first adopt a data-driven saliency detector to evaluate the importance of different image regions and propose a saliency-based block ratio aggregation (BRA) strategy for sampling rate allocation. A unified learnable generating matrix is then developed to produce sampling matrix of any CS ratio with an ordered structure. Being equipped with the optimization-inspired recovery subnet guided by saliency information and a multi-block training scheme preventing blocking artifacts, CASNet jointly reconstructs the image blocks sampled at various sampling rates with one single model. To accelerate training convergence and improve network robustness, we propose an SVD-based initialization scheme and a random transformation enhancement (RTE) strategy, which are extensible without introducing extra parameters. All the CASNet components can be combined and learned end-to-end. We further provide a four-stage implementation for evaluation and practical deployments. Experiments demonstrate that CASNet outperforms other CS networks by a large margin, validating the collaboration and mutual supports among its components and strategies. Codes are available at https://github.com/Guaishou74851/CASNet.
Bin Chen 0006, Jian Zhang 0018
IEEE Trans. Image Process.1
2021 Memory-Augmented Deep Unfolding Network for Compressive Sensing
abstract
Mapping a truncated optimization method into a deep neural network, deep unfolding network (DUN) has attracted growing attention in compressive sensing (CS) due to its good interpretability and high performance. Each stage in DUNs corresponds to one iteration in optimization. By understanding DUNs from the perspective of the human brain's memory processing, we find there exists two issues in existing DUNs. One is the information between every two adjacent stages, which can be regarded as short-term memory, is usually lost seriously. The other is no explicit mechanism to ensure that the previous stages affect the current stage, which means memory is easily forgotten. To solve these issues, in this paper, a novel DUN with persistent memory for CS is proposed, dubbed Memory-Augmented Deep Unfolding Network (MADUN). We design a memory-augmented proximal mapping module (MAPMM) by combining two types of memory augmentation mechanisms, namely High-throughput Short-term Memory (HSM) and Cross-stage Long-term Memory (CLM). HSM is exploited to allow DUNs to transmit multi-channel short-term memory, which greatly reduces information loss between adjacent stages. CLM is utilized to develop the dependency of deep information across cascading stages, which greatly enhances network representation capability. Extensive CS experiments on natural and MR images show that with the strong ability to maintain and balance information our MADUN outperforms existing state-of-the-art methods by a large margin. The source code is available at https://github.com/jianzhangcs/MADUN/.
Jiechong Song, Bin Chen 0006, Jian Zhang 0018
ACM Multimedia2
2021 COAST: COntrollable Arbitrary-Sampling NeTwork for Compressive Sensing
abstract
Recent deep network-based compressive sensing (CS) methods have achieved great success. However, most of them regard different sampling matrices as different independent tasks and need to train a specific model for each target sampling matrix. Such practices give rise to inefficiency in computing and suffer from poor generalization ability. In this paper, we propose a novel COntrollable Arbitrary-Sampling neTwork, dubbed COAST, to solve CS problems of arbitrary-sampling matrices (including unseen sampling matrices) with one single model. Under the optimization-inspired deep unfolding framework, our COAST exhibits good interpretability. In COAST, a random projection augmentation (RPA) strategy is proposed to promote the training diversity in the sampling space to enable arbitrary sampling, and a controllable proximal mapping module (CPMM) and a plug-and-play deblocking (PnP-D) strategy are further developed to dynamically modulate the network features and effectively eliminate the blocking artifacts, respectively. Extensive experiments on widely used benchmark datasets demonstrate that our proposed COAST is not only able to handle arbitrary sampling matrices with one single model but also to achieve state-of-the-art performance with fast speed.
Di You, Jian Zhang 0018, Jingfen Xie, Bin Chen 0006, Siwei Ma 0001
IEEE Trans. Image Process.4
2019 Fingerprint template protection using minutia-pair spectral representations
abstract
Storage of biometric data requires some form of template protection in order to preserve the privacy of people enrolled in a biometric database. One approach is to use a Helper Data System. Here it is necessary to transform the raw biometric measurement into a fixed-length representation. In this paper, we extend the spectral function approach of Stanko and Škorić (IEEE Workshop on Information Forensics and Security (WIFS), 2017) which provides such a fixed-length representation for fingerprints. First, we introduce a new spectral function that captures different information from the minutia orientations. It is complementary to the original spectral function, and we use both of them to extract information from a fingerprint image. Second, we construct a helper data system consisting of zero-leakage quantisation followed by the Code Offset Method. We show empirical data on matching performance and entropy content. On the negative side, transforming a list of minutiae to the spectral representation degrades the matching performance significantly. On the positive side, adding privacy protection to the spectral representation can be done with little loss of performance.
Taras Stanko, Bin Chen 0006, Boris Skoric
EURASIP J. Inf. Secur.2
2019 Complexity-adjustable SC decoding of polar codes for energy consumption reduction
abstract
This study proposes an enhanced list‐aided successive cancellation stack (ELSCS) decoding algorithm with adjustable decoding complexity. Also, a logarithmic likelihood ratio‐threshold based path extension scheme is designed to further reduce the memory consumption of stack decoding. Numerical simulation results show that without affecting the error correction performance, the proposed ELSCS decoding algorithm provides a flexible trade‐off between time complexity and computational complexity, while reducing storage space up to 70%. Based on the fact that most mobile devices operate in environments with stringent energy budget to support diverse applications, the proposed scheme is a promising candidate for meeting requirements of different applications while maintaining a low computational complexity and computing resource utilisation.
Bin Chen 0006, Luis F. Abanto-Leon, Zizheng Cao, Antonius M. J. Koonen
IET Commun.2
2019 Secret Key Generation Over Biased Physical Unclonable Functions With Polar Codes
abstract
Internet-of-Things (IoT) devices are usually small, low cost, and have limited resources, which makes them vulnerable to physical and cloning attacks. To secure IoT devices, physical unclonable functions (PUFs) are relatively new security primitives used for device authentication and device-specific secret-key generation. In this paper, we focus on designing a robust construction to derive secret keys from static randomaccess memory (SRAM)-PUFs, which enjoy the uniqueness and randomness properties stemming from the manufacturing variations of SRAM memory cells. We make use of a polar code construction. Based on the fact that SRAM memory can often be found in today's IoT devices, and since polar codes have been selected as error-correction technique in the fifth generation standard, this makes the proposed scheme a promising candidate for reducing the extra cost and securing resource-constrained IoT devices. In this paper, we propose a novel construction method to eliminate the effect of noise and bias in SRAM-PUFs. We shall prove that the secrecy leakage of the helper data about the secretkey can be made negligible due to polarization and proper code construction design. Results show that the proposed scheme provides a significant improvement of the reliability (achieve a failure probability below 10-6) and of the realizable secret-key rate, which is also evaluated by the theoretical analysis. In addition, the proposed scheme provides the possibility to tradeoff complexity, secrecy, and reliability with the same code construction for different IoT applications.
Bin Chen 0006, Frans M. J. Willems
IEEE Internet Things J.1
2019 Improved Decoding of Staircase Codes: The Soft-Aided Bit-Marking (SABM) Algorithm
abstract
Staircase codes (SCCs) are typically decoded using iterative bounded-distance decoding (BDD) and hard decisions. In this paper, a novel decoding algorithm is proposed, which partially uses soft information from the channel. The proposed algorithm is based on marking certain number of highly reliable and highly unreliable bits. These marked bits are used to improve the miscorrection-detection capability of the SCC decoder and the error-correcting capability of BDD. For SCCs with 2-error-correcting Bose-Chaudhuri-Hocquenghem component codes, our algorithm improves upon standard SCC decoding by up to 0.30 dB at a bit-error rate (BER) of 10-7. The proposed algorithm is shown to achieve almost half of the gain achievable by a genie decoder with this structure. The increased complexity caused by bit marking and additional calls to the component BDD decoder is discussed as well. Our algorithm is also extended (with minor modifications) to product codes. The simulation results show that in this case, the algorithm offers gains of up to 0.5 dB at a BER of 10-7.
Bin Chen 0006, Gabriele Liga, Xiong Deng, Zizheng Cao, Jianqiang Li 0003, Kun Xu 0008, Alex Alvarado
IEEE Trans. Commun.2
2018 Mitigating LED Nonlinearity to Enhance Visible Light Communications
abstract
This paper addresses the nonlinear memory effects in the response of typical illumination light emitting diodes (LEDs), in order to enhance the performance of visible light communication (VLC) systems. These LEDs have a limited bandwidth of only several MHz. To reflect the physical mechanisms in the quantum well, we describe the LED transient response by a nonlinear dynamic differential equation. Three different mechanisms of the nonlinearity are relevant in the double hetero-structure LEDs, which result in dynamic nonlinearities, that is, a mixture of nonlinearities and memory effects. Hitherto, generic pre-distorter and non-linear equalizers have been studied for the LEDs. Yet this paper shows that recombination rates of photon generation can be translated into an equivalent discrete-time circuit that can be inverted. This allows us to develop a new pre-distorter with a simpler and more efficient structure than previously studied and overly generic approaches. The novel pre-distorter along with a parameter estimation can effectively overcome LED nonlinearity for high-speed VLC with amplitude-based single carrier modulations, including ON-OFF keying and pulse amplitude modulation-4 systems, and with the multi-carrier orthogonal frequency-division multiplexing. We report experimentally obtained eye-diagrams, first to justify our choice for the LED model on which our nonlinear pre-distorter have been based, and second to verify the effectiveness in enhancing the VLC link performance to the extent predicted by our model.
Xiong Deng, Shokoufeh Mardanikorani, Yan Wu 0001, Kumar Arulandu, Bin Chen 0006, Amir M. Khalid, Jean-Paul Linnartz
IEEE Trans. Commun.5
2017 A Robust SRAM-PUF Key Generation Scheme Based on Polar Codes
abstract
Physical unclonable functions (PUFs) are relatively new security primitives used for device authentication and device-specific secret key generation. In this paper we focus on SRAM- PUFs. The SRAM-PUFs enjoy uniqueness and randomness properties stemming from the intrinsic randomness of SRAM memory cells, which is a result of manufacturing variations. This randomness can be translated into the cryptographic keys thus avoiding the need to store and manage the device cryptographic keys. Therefore these properties, combined with the fact that SRAM memory can be often found in today's IoT devices, make SRAM-PUFs a promising candidate for securing and authentication of the resource-constrained IoT devices. PUF observations are always affected by noise and environmental changes. Therefore secret- generation schemes with helper data are used to guarantee reliable regeneration of the PUF-based secret keys. Error correction codes (ECCs) are an essential part of these schemes. In this work, we propose a practical error correction construction for PUF-based secret generation that are based on polar codes. The resulting scheme can generate 128-bit keys using 1024 SRAM-PUF bits and 896 helper data bits and achieve a failure probability of 10^{-9} or lower for a practical SRAM-PUFs setting with bit error probability of 15%. The method is based on successive cancellation combined with list decoding and hash-based checking that makes use of the hash that is already available at the decoder. In addition, an adaptive list decoder for polar codes is investigated. This decoder increases the list size only if needed.
Bin Chen 0006, Tanya Ignatenko, Frans M. J. Willems, Roel Maes, Erik van der Sluis, Georgios N. Selimis
GLOBECOM1
2016 Spatially-coupled LDPC coding in cooperative wireless networks
abstract
This paper proposes a new technique of spatially-coupled low-density parity-check (SC-LDPC) code-based soft information relaying scheme for a two-way relay system. We introduce an optimized SC-LDPC codes in relay channels. A more precise model is proposed to characterize the soft noise on the soft symbols, using a pre-calculated look-up table at the destination. This requires less signalling overhead compared to existing soft noise modelling techniques. We also introduce a variance correction factor to provide a rectification to the equivalent total noise variance at the destination. Finally, we modify the LLR former at the destination which is tailored to the proposed soft information relaying technique. Simulation results demonstrate that the proposed relay protocol yields an improved BER performance compared to competitive schemes proposed in the literature.
Dushantha N. K. Jayakody, Vitaly Skachek, Bin Chen 0006
WCNC3
2015 Low-Density Lattice Coded Relaying With Joint Iterative Decoding
abstract
Low-density lattice codes (LDLCs) are known for their high decoding efficiency and near-capacity performance on point-to-point Gaussian channels. In this paper, we present a distributed LDLC-based cooperative relaying scheme for the multiple-access relay channel (MARC). The relay node decodes LDLC-coded packets from two sources and forwards a network-coded combination to the destination. At the destination, a joint iterative decoding structure is designed to exploit the diversity gain as well as coding gain. For the LDLC-based network coding operation at the relay, we consider two alternative methods which offer a tradeoff between implementation complexity and performance, called superposition LDLC (S-LDLC) and modulo-addition LDLC (MA-LDLC). Soft symbol relaying is considered as an alternative to hard decision relaying which is capable of reducing the effect of error propagation at the relay. Simulation results show that the proposed scheme can provide greater diversity gain and up to 6.2 dB coding gain when compared with noncooperative LDLC coding and uncoded network-coded transmission. The proposed scheme also achieves 2.5 dB gain over network-turbo-coded cooperation, for the same code rate and overall transmitted power. Also, soft symbol relaying is shown to provide approximately 2 dB gain over hard decision relaying when the source-relay link suffers from deep fading.
Bin Chen 0006, Dushantha N. K. Jayakody, Mark F. Flanagan
IEEE Trans. Commun.1
2014 A multilevel soft quantize-and-forward scheme for multiple access relay systems
abstract
This paper proposes the novel technique of multilevel threshold based soft quantization (MLT-SQ) for a multiple access relay system (MARS). The scheme is suitable for systems using binary phase-shift keying (BPSK) and network coding at the relay. In the proposed MLT-SQ protocol, the relay evaluates the reliabilities, expressed as log-likelihood ratios (LLRs), of the received signals from the two sources. It then computes the LLRs of the network-coded packet and quantizes these using a set of optimized multilevel thresholds, forwarding the resulting “quantized soft symbols” to the destination. We provide the derivation for the bit error rate (BER) at the destination, based on which we optimize the multilevel thresholds to minimize the BER. Compared to competing schemes, the performance of our system is superior in terms of BER when the same amount of channel state information (CSI) is exploited.
Dushantha N. K. Jayakody, Jun Li 0004, Bin Chen 0006, Mark F. Flanagan
PIMRC3