Bingjie Xia

dblp:27/8058 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-2711-6228ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PMENet: Pre-assessment Modality Enhancement Network for RGB-T salient object detection
Hongbo Bi, Bingjie Xia, Weihan Sun, Yina Zhou
Image Vis. Comput.2
2025 Comprehensive RISC- V Floating-Point Verification: Efficient Coverage Models and Constraint-Based Test Generation
abstract
The increasing complexity of processor architectures necessitates more rigorous functional verification. Floating-point operations, in particular, present significant challenges due to their extensive range of computational cases that require verification. This paper proposes a comprehensive approach for generating floating-point instruction sequences to enhance the verification of RISC-V. We introduce a constraint-based method for floating-point test generation and design efficient coverage models as input constraints for this process. The resulting representative floating-point tests are integrated with RISC-V instruction sequence generation through a memory-bound register update method. Experimental results demonstrate that our approach improves the functional coverage of RISC-V floating-point instruction sequences from 93.32% to 98.34%, while simultaneously reducing the number of required instructions by 66.67% compared to the Google RISCV-DV generator. Additionally, our method achieves more comprehensive coverage of floating-point types in instruction write-back data compared to RISCV-DV. Using the proposed approach, we successfully detect representative floating-point-related faults injected into the RISC-V processor CV32E40P, thereby demonstrating its effectiveness.
Tianyao Lu, Anlin Liu, Bingjie Xia, Peng Liu 0016
DATE3
2025 Efficient Multiple-Precision Floating-Point Multiply-Add Architecture for Deep Learning Applications
abstract
To fully exploit the potential of parallel computing while minimizing hardware costs, floating-point multiply-add (FMA) units that support multiple precisions are widely used in deep learning. However, achieving a balance between accuracy, performance, and hardware overhead remains a significant challenge. This paper presents a unified multiple-precision floating-point FMA architecture that supports four precision formats: single-precision (SP), Bfloat16 (BF16), TensorFloat32+ (TF32+), and INT8. The proposed architecture allows for the parallel execution of nine BF16 FMA operations, two TF32+ FMA operations, one SP FMA operation, or nine INT8 multiply-add (MA) operations. Through careful data format selection and an optimized architectural design, the architecture achieves multiplier utilization rates of 100%, 88.9%, 100%, and 100% for the four precision modes, respectively, with all multipliers operating at full bit width. Compared to state-of-the-art multiple-precision FMA designs, this architecture delivers over nine times the BF16 throughput while increasing the area by only 49%. The flexible data format configuration makes the proposed architecture suitable for a wide range of deep learning applications.
Songtai Liang, Bingjie Xia, Wen Wang 0015, Peng Liu 0016
ISCAS2
2024 Mantissa-Aware Floating-Point Eight-Term Fused Dot Product Unit
abstract
Floating-point dot product is widely used in various applications. A conventional discrete construction of multipliers and adders often leads to accumulated errors and lower speed. This paper presents a floating-point eight-term fused dot product unit with mantissa-aware hardware design to attack these problems. In the proposed design, multiplication and addition of numbers are operated in a fused manner with exception controller. A pre-shift and post-shift combined scheme for mantissa alignment is utilized to eliminate the latency between significand multiplication and mantissa alignment. A one-path mantissa compress and addition structure is employed that can effectively reduce the footprint. An impact of internal mantissa datapath width on calculation accuracy is analyzed. Compared to the discrete method, the proposed design delivers a significant reduction up to 45.7%, 26.9%, and 33.0%, in terms of latency, area, and power, respectively.
Wen Wang 0015, Bingjie Xia, Peng Liu 0016
ISCAS2
2021 Inversion Based on a Detached Dual-Channel Domain Method for StyleGAN2 Embedding
abstract
A style-based generative adversarial network (StyleGAN2) yields remarkable results in image-to-latent embedding. This work proposes a Detached Dual-channel Domain Encoder as an effective and robust method to embed an image to a latent code, i.e., GAN inversion. It infers a latent code from two aspects: a) a detached dual-channel design to support faithful image reconstruction; and b) a local skip connection that allows conveying pieces of information with image details. We further introduce a hierarchical progressive training strategy that allows the proposed encoder to separately capture different semantic features. The qualitative and quantitative experimental results show that the well-trained encoder can embed an image into a latent code in StyleGAN2 latent space with less time than its peers while preserving facial identity and image details well.
MengChu Zhou, Bingjie Xia, Xiwang Guo 0001, Liang Qi 0001
IEEE Signal Process. Lett.3