Biao Dong

dblp:127/2996 · DBLP profile ↗
← Back
14ranked-venue papers
9as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Inference-Optimal ISAC via Task-Oriented Feature Transmission and Power Allocation
abstract
This work is concerned with the coordination gain in integrated sensing and communication (ISAC) systems under a compress-and-estimate (CE) framework, wherein inference performance is leveraged as the key metric. To enable tractable transceiver design and resource optimization, we characterize inference performance via an error probability bound as a monotonic function of the discriminant gain (DG). This raises the natural question of whether maximizing DG, rather than minimizing mean squared error (MSE), can yield better inference performance. Closed-form solutions for DG-optimal and MSE-optimal transceiver designs are derived, revealing water-filling-type structures and explicit sensing and communication (S\&C) tradeoff. Numerical experiments confirm that DG-optimal design achieves more power-efficient transmission, especially in the low signal-to-noise ratio (SNR) regime, by selectively allocating power to informative features and thus saving transmit power for sensing.
Biao Dong, Bin Cao 0003, Qinyu Zhang 0001
ICC1
2026 Robust Deep Joint Source-Channel Coding Enabled Distributed Image Transmission With Imperfect Channel State Information
abstract
This work is concerned with robust distributed multi-view image transmission over a severe fading channel with imperfect channel state information (CSI), wherein the sources are slightly correlated. In contrast to point-to-point deep joint source-channel coding (DJSCC), the distributed setting introduces the key challenge of exploiting inter-source correlations without direct communication, especially under imperfect CSI. To tackle this problem, we leverage the complementarity and consistency characteristics among the distributed, yet correlated sources, and propose an robust distributed DJSCC, namely RDJSCC. In RDJSCC, we design a novel cross-view information extraction (CVIE) mechanism to capture more nuanced cross-view patterns and dependencies. In addition, a complementarity-consistency fusion (CCF) mechanism is utilized to fuse the complementarity and consistency from multi-view information in a symmetric and compact manner. Theoretical analysis and simulation results show that our proposed RDJSCC can effectively leverage the advantages of correlated sources even under severe fading conditions, leading to an improved reconstruction performance.
Biao Dong, Bin Cao 0003, Guan Gui 0001, Qinyu Zhang 0001
IEEE Trans. Wirel. Commun.1
2025 M3Diff:Semantic Mask-Guided 3D Medical Image Synthesis via Mamba-U-Net Hybrid for Data Augmentation
Haixia Pan, Biao Dong, Xirui Wu, Zhaohui Tian, Yunfei Lei, Huolong Ye
ICIC (25)2
2025 NuPreX: Times Series Forecasting for the Nuclear Steam Supply System
Huobin Tan, Biao Dong
ICIC (7)4
2025 T4Di: A Hybrid TTT-Transformer Backbone for Scalable and Efficient Diffusion Model
Xirui Wu, Haixia Pan, Biao Dong, Huolong Ye
ICIC (18)4
2025 Talking Head Generation via Viewpoint and Lighting Simulation Based on Global Representation
abstract
NeRF-based talking head generation has made great progress, but existing methods still lack in achieving high-quality detail fidelity, mainly manifested in detail loss and intermittent blur. We attribute this to the limitations of the training video data in terms of viewpoint and lighting, which leads to the inability to fully model the global depth and brightness information of spatial points. Specifically, a fixed viewpoint may fail to provide sufficient depth information for high-frequency details, leading to inaccurate volume density estimation and the loss of details such as hair. Furthermore, constant lighting often fails to adapt to the drastic brightness changes of continuous video frames, resulting in color accumulation errors and blurring artifacts. To address these issues, we propose a novel talking head generation method that combines layered viewpoint simulation (LVS) and continuous lighting simulation (CLS). LVS simulates multiple viewpoints through the multi-scale features of the video frame to construct the global depth representation, which can improve the accuracy of volume density estimation and enhance detail description. CLS simulates multiple lighting through brightness changes of continuous video frames to construct the global brightness representation, thereby alleviating color accumulation errors and eliminating blur. Extensive experiments demonstrate that our method significantly improves the detail quality compared to the state-of-the-art methods.
Biao Dong, Lei Zhang 0021
ACM Multimedia1
2025 Enhanced Temporal Representation and Spatial Alignment for High-Fidelity Talking Video Generation
Biao Dong
Vis. Comput.1
2025 Integrating models of real aboveground scene and underground geological structures at an open pit mine
abstract
Background As information technology has advanced and been popularized, open pit mining has rapidly developed toward integration and digitization. The three-dimensional reconstruction technology has been successfully applied to geological reconstruction and modeling of surface scenes in open pit mines. However, an integrated modeling method for surface and underground mine sites has not been reported. Methods In this study, we propose an integrated modeling method for open pit mines that fuses a real scene on the surface with an underground geological model. Based on oblique photography, a real-scene model was established on the surface. Based on the surface-stitching method proposed, the upper and lower surfaces and sides of the model were constructed in stages to construct a complete underground three-dimensional geological model, and the aboveground and underground models were registered together to build an integrated open pit mine model. Results The oblique photography method used reconstructed a surface model of an open pit mine using a real scene. The surface-stitching algorithm proposed was compared with the ball-pivoting and Poisson algorithms, and the integrity of the reconstructed model was markedly superior to that of the other two reconstruction methods. In addition, the surface-stitching algorithm was applied to the reconstruction of different formation models and showed good stability and reconstruction efficiency. Finally, the aboveground and underground models were accurately fitted after registration to form an integrated model. Conclusions The proposed method can efficiently establish an integrated open pit model. Based on the integrated model, an open pit auxiliary planning system was designed and realized. It supports the functions of mining planning and output calculation, assists users in mining planning and operation management, and improves production efficiency and management levels.
Biao Dong, Wenjun Tan, Weichao Chang, Baoting Li, Yanliang Guo, Quanxing Hu, Guangwei Liu
Virtual Real. Intell. Hardw.1
2024 RDJSCC: Robust Deep Joint Source-Channel Coding Enabled Distributed Image Transmision over Severe Fading Channel
abstract
In this paper, we investigate the effects of severe channel fading in the scenario of distributed deep learning-based joint source-channel coding (DJSCC) for image transmission without perfect channel state information (CSI). To tackle the challenges posed by imperfect CSI, we propose a robust DJSCC (RDJSCC) scheme that operates at three levels: modulation, encoding, and decoding, respectively. Firstly, at the modulation level, we adopt orthogonal frequency division multiplexing (OFDM) modulation for exploring the tradeoff between reconstruction performance and peak-to-average power ratio (PAPR). Secondly, at the encoding level, two parameter-efficient operators are introduced to combat channel fading with low encoding complexity. Finally, at the decoding level, we divide the decoding process into two stages, i.e., denoising and recovery, aiming to maximize the correlation between the encoded representations. Theoretic analysis and simulation results show that our proposed RDJSCC can effectively alleviate the effects of severe fading with imperfect CSI, leading to an improved reconstruction performance while maintaining low PAPR and encoding complexity.
Biao Dong, Wenkai Tian, Bin Cao 0003, Yu Wang 0078
GLOBECOM1
2024 Joint ROI Guidance and Spatial Analysis for Task-Aware Distributed Deep Joint Source-Channel Coding
abstract
In this paper, we investigate the system performance of deep joint source-channel coding (JSCC) for task-oriented transmission in the Wyner-Ziv scenario, i.e., a distributed coding scenario, aiming to improve the image reconstruction performance and task accuracy. Unlike existing deep JSCC based methods, we introduce regions of interest (ROI), which facilitates the effective utilization of side information for enhancing task performance. Meanwhile, we incorporate a spatial analysis mechanism to fuse the side information. By integrating these two mechanisms, we propose a novel distributed deep JSCC scheme that further leverages task relevance within the side information. Simulation results show that our proposed scheme outperforms the benchmark in terms of image reconstruction performance and task accuracy. The code is available on the project website1.
Wenkai Tian, Biao Dong, Bin Cao 0003
GLOBECOM2
2024 Spatially and Temporally Optimized Audio-Driven Talking Face Generation
abstract
Abstract Audio‐driven talking face generation is essentially a cross‐modal mapping from audio to video frames. The main challenge lies in the intricate one‐to‐many mapping, which affects lip sync accuracy. And the loss of facial details during image reconstruction often results in visual artifacts in the generated video. To overcome these challenges, this paper proposes to enhance the quality of generated talking faces with a new spatio‐temporal consistency. Specifically, the temporal consistency is achieved through consecutive frames of the each phoneme, which form temporal modules that exhibit similar lip appearance changes. This allows for adaptive adjustment in the lip movement for accurate sync. The spatial consistency pertains to the uniform distribution of textures within local regions, which form spatial modules and regulate the texture distribution in the generator. This yields fine details in the reconstructed facial images. Extensive experiments show that our method can generate more natural talking faces than previous state‐of‐the‐art methods in both accurate lip sync and realistic facial details.
Biao Dong, Bo-Yao Ma
Comput. Graph. Forum1
2024 Study on defect detection of metal castings based on supervised enhancement and attention distillation
Haixia Pan, Xingyun Wei, Biao Dong, Jiahua Lan
Mach. Vis. Appl.5
2022 Decentralized Automatic Modulation Classification Method Based on Lightweight Neural Network
abstract
Due to the computing capability and memory limitations, it is difficult to apply the traditional deep learning (DL) models to the edge devices (EDs) for realizing automatic modulation classification (AMC). In this paper, a lightweight neural network for decentralized learning-based automatic modulation classification (DecentAMC) method is proposed. Specifically, group convolutional neural network (GCNN) is designed by replacing the standard convolution layer with the group convolution layer, replacing the flatten layer with the global average pooling (GAP) layer and removing part of fully connected layers. DecentAMC method is achieved by the cooperation in which multiple EDs update and upload the model weight to a central device (CD) for model aggregation to avoid the data privacy disclosure. Experimental results show that the proposed GCNN-based DecentAMC method can improve training efficiency to about 4 times and 57 times than that of GCNN-based centralized AMC (CentAMC) and CNN-based DecentAMC respectively. GCNN-based DecentAMC method can effectively reduce the communication cost and save storage of EDs when compared with CNN-based DecentAMC. Meanwhile, the time complexity and the space complexity of GCNN is significantly decreased when compared with CNN and SCNN, which is suitable to be deployed in EDs.
Biao Dong, Guozhen Xu, Xue Fu, Guan Gui 0001, Haris Gacanin, Fumiyuki Adachi
PIMRC1
2022 A Lightweight Decentralized-Learning-Based Automatic Modulation Classification Method for Resource-Constrained Edge Devices
abstract
Due to the computing capability and memory limitations, it is difficult to apply the traditional deep learning (DL) models to the edge devices (EDs) for realizing lightweight automatic modulation classification (AMC). Recently, many works attempt to use different ways to realize lightweight AMC methods for EDs. However, the lightweight seems to be a contradiction with the classification performance in these lightweight networks. In this article, we propose an efficient lightweight decentralized-learning-based AMC (DecentAMC) method using spatiotemporal hybrid deep neural network based on multichannels and multifunction blocks (MCMBNN). Specifically, the lightweight network is designed from the perspectives of comprehensive consideration of lightweight and classification performance, which is composed of three parts to extract different features for realizing high classification performance and they are phase estimator and transformer (PET) block, spatial feature extraction block and temporal feature extraction & Softmax block. In addition, we use a multichannel input to extract complementary features of different channels for a better classification performance. The proposed DecentAMC method is an efficient training method, which is achieved by the cooperation in which multiple EDs update and upload the model weight to a central device (CD) for model aggregation to avoid the data privacy disclosure and reduce the computing power and storage pressure of CD. Experimental results show that the proposed MCMBNN can obtain an improved classification accuracy while reducing model complexity with the contributions of three blocks. Moreover, the proposed DecentAMC method can be deployed on EDs efficiently. Thus, the method has the advantages of avoiding data leakage on EDs and relieving the computing pressure of CD with relatively lower communication overhead. The simulation code and datasets are shared on GitHub.
Biao Dong, Guan Gui 0001, Xue Fu, Bamidele Adebisi, Haris Gacanin, Hikmet Sari
IEEE Internet Things J.1