Yuxian Jiang

dblp:134/1105 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Evolutionary Guided Decoding: Iterative Value Refinement for LLMs
abstract
Zhenhua Liu, Lijun Li, Ruizhe Chen, Yuxian Jiang, Tong Zhu, Zhaochen Su, Wenliang Chen, Jing Shao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ruizhe Chen, Yuxian Jiang, Tong Zhu 0002, Zhaochen Su, Wenliang Chen
ACL (1)4
2026 Multi-structure segmentation in CBCT volumes: The ToothFairy2 challenge
abstract
Cone-beam computed tomography (CBCT) is widely used for dento-maxillofacial diagnostics and treatment planning, and comprehensive multi-structure segmentation remains time-consuming, limiting large-scale, reproducible research. In this article, we present ToothFairy2, a MICCAI 2024 challenge on multi-structure segmentation in maxillofacial CBCT. The accompanying dataset comprises 530 CBCT volumes (480 public training, 50 hidden test) with expert 3D annotations of 42 classes, including maxilla, mandible, crowns, bridges, implants, inferior alveolar canals, maxillary sinuses, pharynx, and teeth labeled according to the International Tooth Numbering System (FDI). 26 international teams participated in ToothFairy2, and their methods were run and evaluated for voxel-wise multi-class segmentation using a standardized protocol. This report extends the evaluation of teeth to also investigate the current capabilities of tooth detection and FDI numbering. Furthermore, ranking stability was analyzed to assess the robustness of the final challenge outcome. Overall, challenge participants achieved consistently high performance for large, high-contrast structures such as jawbones, pharynx, and most teeth, while maxillary sinuses, dental restorations, and fine structures remain challenging due to class imbalance and metal artifacts. Analysis of tooth-related metrics further revealed that assigning correct FDI numbers was more challenging than delineating individual teeth. By releasing CBCT data, 3D annotations, baseline models, and evaluation code, ToothFairy2 establishes a long-term benchmark to drive the development of automated methods for robust, clinically meaningful multi-structure segmentation in maxillofacial CBCT.
Federico Bolelli, Luca Lumetti, Niels van Nistelrooij, Shankeeth Vinayahalingam, Mattia Di Bartolomeo, Kevin Marchesini, Arrigo Pellacani, Ettore Candeloro, Gabriele Rosati, Tong Xi 0001, Fabian Isensee, Yannick Kirchhoff, Lars Krämer, Maximilian Rokuss, Constantin Ulrich, Klaus H. Maier-Hein, Yuxian Jiang, Yusheng Liu 0001, Lisheng Wang, Haoshen Wang, Zhiming Cui 0001, Zhaohong Pan, Xiaokun Liang, Ender Konukoglu, Marek Wodzinski, Henning Müller, Haipeng Mai, Xiaobing Dang, Shrajan Bhandary, Radu Grosu, Stefaan Bergé, Alexandre Anesi, Costantino Grana
Medical Image Anal.17
2026 MetaAccel: A High-Performance and Agile Accelerator Design Framework With Multi Clock Domain Optimization for Complex CNN
abstract
Edge computing for artificial intelligence (AI) has become a new focus today. At the edge, the growing complexity and diversity of AI models has made FPGA, which has shorter development cycles, a good choice. Traditional AI accelerators on FPGA are mainly based on the Compute Engine (CE) architecture, suffering from low resource utilization and suboptimal speed. In contrast, the pipeline architecture achieves higher performance through its algorithm-structure-aware feature and fully on-chip data flow. However, customized designs and large bandwidth demands bring new challenges to its development agility and memory utilization, while high-performance acceleration for complex neural networks is still hard. In this article, we proposed MetaAccel, a novel fully pipelined accelerator. It uses two clock domains to manage data scheduling and calculation, significantly improving computing resource efficiency and on-chip memory utilization. Besides that, we built a hyperparameter-driven resource estimation model that can match the most appropriate design solutions for specific network structures. Based on this architecture, processing method for networks with complex branch structures and various operations is given, which makes MetaAccel suitable for Convolutional Neural Networks (CNNs) in different fields, such as image classification, object detection, and image segmentation. For typical networks, MetaAccel can achieve a throughput of more than 0.7TOPS and a DSP efficiency of up to 2.0GOPS/DSP and outperforms previous FPGA work in other metrics, showing its advantages in complex CNNs’ acceleration.
Yuxian Jiang, Zhihan Zhang 0004, Qunkang Meng, Hao Wang 0046, Qijun Huang, Sheng Chang 0003
IEEE Trans. Circuits Syst. I Regul. Pap.2
2026 Morphology Prior Enhanced Teeth Segmentation for High-Resolution Oral Scans
abstract
Deep learning methods have been proposed for tooth segmentation on high-resolution intra-oral scans (IOS) that plays a crucial role in clinical dental practice. However, they generally segment teeth in a low-resolution data with a fixed receptive field and generate final segmentation by up-sampling interpolation, and neglect teeth's morphology priors: their similar dental arch structures and significantly different curvatures in different parts of each tooth. They thus lack adaptability to different parts of each tooth, and show less accurate segmentation of boundary points between teeth and gums due to the up-sampling computation. Further, cluttered poses of IOS limit their generalization and usability of teeth location and geometric information. To address these limitations, a morphology prior enhanced teeth segmentation framework is proposed in this paper. Firstly, a robust preprocessing is introduced to align poses of different IOS by computing their dental arch orientations, thereby improving segmentation generalization and usability of IOS geometric information. Secondly, a decomposition-merging strategy is designed to avoid the up-sampling limitation, which decomposes an IOS into multiple low-resolution data and merges their segmentation outcomes into a high-resolution result. Thirdly, an innovative module integrating semantic and geometric features is proposed to adaptively select deformable receptive fields. It geometrically samples within a variable probability space to construct receptive fields with varied graph relationships for different points, facilitating adaptive segmentation of different parts of each tooth. Experimental results on 6238 IOS from four centers demonstrate that our method significantly outperforms 11 state-of-the-art methods, achieving a 6.93% enhancement for cross-center testing.
Yuxian Jiang, Xiuying Wang 0001, Tao Yang 0037, Changkai Ji, Lanshan He, Yusheng Liu 0001, Junyu Shi, Huayan Guo, Lisheng Wang
IEEE J. Biomed. Health Informatics1
2025 A High-Intensity Solution of Hardware Accelerator for Sparse and Redundant Computations in Semantic Segmentation Models
abstract
The rapid development of artificial intelligence (AI) has met people’s personalized needs. However, with the increase of data capacities and computing requirements, the imbalance between large-scale data transmission and limited network bandwidth has become increasingly prominent. To improve the speed of embedded system, real-time intelligent computing is gradually moving from the cloud to the edge. Traditional FPGA-based AI accelerators mainly utilize PE architecture, but the low computing throughput and resource utilization make it difficult to meet the power requirement of edge AI application scenarios such as image segmentation. In recent years, AI accelerators based on streaming architecture have become a trend, and it is necessary to customize high-performance streaming accelerators for specific segmentation algorithms. In this paper, we design a high-intensity pixel-level fully pipelined accelerator with customized strategies to eliminate the sparse and redundant computations in specific algorithms of semantic segmentation, which significantly improve the accelerator’s computing throughput and hardware resources utilization. On Xilinx FPGA, our acceleration of two typical semantic segmentation networks-ESPNet and DeepLabV3, achieves optimized throughputs of 171.3 GOPS and 1324.8 GOPS, and computing efficiency of 9.26 and 9.01, respectively. It provides the possibility of hardware deployment in real-time application with high computing intensity.
Yuxian Jiang, Zhihan Zhang 0004, Hao Wang 0046, Sheng Chang 0003
IEEE Trans. Computers3
2025 PEDSA: High-Throughput Pipeline-Based FPGA Accelerator for Convolutional Encoder-Decoder Segmentation Networks
abstract
In the era of artificial intelligence (AI), rapidly growing data and computing demands stimulate a shift toward more intelligent processing at the edge in the Internet of Things (IoT). AI application scenarios, such as image segmentation, pose new challenges to the computing capability of edge hardware, which cannot be solved by traditional AI accelerators with the traditional processing element (PE) architecture. Recently, the streaming architecture has received more attention due to its higher performance. To improve the throughput of edge platforms for segmentation tasks, customizing streaming accelerators for segmentation models is now necessary. Based on these motivations, we proposed pipelined encoder-decoder segmentation model accelerator (PEDSA), a fully pipelined streaming accelerator for convolutional encoder-decoder segmentation networks. PEDSA maps all the layers in the network into a pixel-level pipeline. All operations, especially upsampling (unpooling, deconvolution, etc.) which involves complex data rearrangement, are integrated into a regular, concise, and fully on-chip data flow. On Xilinx field-programmable gate arrays (FPGAs), our acceleration of SegNet-Basic and U-Net reached performances of 2676.47 and 7646.29 GOPS, respectively, outperforming previous accelerators for this kind of network. This work provides new ideas for the deployment of segmentation algorithms at the edge.
Yuxian Jiang, Zhihan Zhang 0004, Hao Wang 0046, Sheng Chang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 Individual Graph Representation Learning for Pediatric Tooth Segmentation From Dental CBCT
abstract
Pediatric teeth exhibit significant changes in type and spatial distribution across different age groups. This variation makes pediatric teeth segmentation from cone-beam computed tomography (CBCT) more challenging than that in adult teeth. Existing methods mainly focus on adult teeth segmentation, which however cannot be adapted to spatial distribution of pediatric teeth with individual changes (SDPTIC) in different children, resulting in limited accuracy for segmenting pediatric teeth. Therefore, we introduce a novel topology structure-guided graph convolutional network (TSG-GCN) to generate dynamic graph representation of SDPTIC for improved pediatric teeth segmentation. Specifically, this network combines a 3D GCN-based decoder for teeth segmentation and a 2D decoder for dynamic adjacency matrix learning (DAML) to capture SDPTIC information for individual graph representation. 3D teeth labels are transformed into specially-designed 2D projection labels, which is accomplished by first decoupling 3D teeth labels into class-wise volumes for different teeth via one-hot encoding and then projecting them to generate instance-wise 2D projections. With such 2D labels, DAML can be trained to adaptively describe SDPTIC from CBCT with dynamic adjacency matrix, which is then incorporated into GCN for improving segmentation. To ensure inter-task consistency at the adjacency matrix level between the two decoders, a novel loss function is designed. It can address the issue with inconsistent prediction and unstable TSG-GCN convergence due to two heterogeneous decoders. The TSG-GCN approach is finally validated with both public and multi-center datasets. Experimental results demonstrate its effectiveness for pediatric teeth segmentation, with significant improvement over seven state-of-the-art methods.
Yusheng Liu 0001, Xiyi Wu, Tao Yang 0037, Yuchen Pei, Huayan Guo, Yuxian Jiang, Zhien Feng, Yu-Ping Wang 0002, Lisheng Wang
IEEE Trans. Medical Imaging7
2024 ClarityDiffuseNet: Enhancing fundus image quality under black shadows with diffusion model-based research
Jiadi Dong, Tianwei Qian, Yuxian Jiang, Lei Bi 0001, Jinman Kim, Lisheng Wang
Pattern Recognit. Lett.3
2012 An improved algorithm to determine the relative attitude during rolling phase of spacecraft rendezvous and docking
abstract
Relative attitude is computed with improved TRIAD algorithm when relative position is far and relative attitude is large during rolling phase of spacecraft rendezvous and docking. On the precondition that the distances between identification points on target spacecraft and the centre of camera lens have been worked out with existing algorithms, orthogonal basis is constructed in each vector space. Rotation matrix can be calculated by the transformation of orthogonal basis and relative attitude can be computed further. The orthogonal basis reflects the relative attitudes between two coordinate systems exactly, avoiding the shortage that traditional TRIAD algorithm can not reflect the rotation along body axis and will induce the large computing errors even wrong results. Improved TRIAD can be used to calculate large attitude. Simulations show that the algorithm is more reliable, more accurate and is not confined to the size of attitude. It is superior to traditional TRIAD algorithm remarkably.
Yuxian Jiang
INDIN2