Jiaming Xie

dblp:230/3669 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0002-5671-5753ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Integrating Eye Tracking With Grouped Fusion Networks for Semantic Segmentation on Mammogram Images
abstract
Medical image segmentation has seen great progress in recent years, largely due to the development of deep neural networks. However, unlike in computer vision, high-quality clinical data is relatively scarce, and the annotation process is often a burden for clinicians. As a result, the scarcity of medical data limits the performance of existing medical image segmentation models. In this paper, we propose a novel framework that integrates eye tracking information from experienced radiologists during the screening process to improve the performance of deep neural networks with limited data. Our approach, a grouped hierarchical network, guides the network to learn from its faults by using gaze information as weak supervision. We demonstrate the effectiveness of our framework on mammogram images, particularly for handling segmentation classes with large scale differences. We evaluate the impact of gaze information on medical image segmentation tasks and show that our method achieves better segmentation performance compared to state-of-the-art models. A robustness study is conducted to investigate the influence of distraction or inaccuracies in gaze collection. We also develop a convenient system for collecting gaze data without interrupting the normal clinical workflow. Our work offers novel insights into the potential benefits of integrating gaze information into medical image segmentation tasks.
Jiaming Xie, Zhiming Cui 0001, Chong Ma 0004, Wenping Wang 0001, Dinggang Shen
IEEE Trans. Medical Imaging1
2025 Tooth Motion Monitoring in Orthodontic Treatment by Mobile Device-Based Multi-View Stereo
abstract
Nowadays, orthodontics has become an important part of modern personal life to assist one in improving mastication and raising self-esteem. However, the quality of orthodontic treatment still heavily relies on the empirical evaluation of experienced doctors, which lacks quantitative assessment and requires patients to visit clinics frequently for in-person examination. To resolve the aforementioned problem, we propose a novel and practical mobile device-based framework for precisely measuring tooth movement in treatment, so as to simplify and strengthen the traditional tooth monitoring process. To this end, we formulate the tooth movement monitoring task as a multi-view multi-object pose estimation problem via different views that capture multiple texture-less and severely occluded objects (i.e. teeth). Specifically, we exploit a pre-scanned 3D tooth model and a sparse set of multi-view tooth images as inputs for our proposed tooth monitoring framework. After extracting tooth contours and localizing the initial camera pose of each view from the initial configuration, we propose a joint pose estimation scheme to precisely estimate the 3D pose of each individual tooth, so as to infer their relative offsets during treatment. Furthermore, we introduce the metric of Relative Pose Bias to evaluate the individual tooth pose accuracy in a small scale. We demonstrate that our approach is capable of reaching high accuracy and efficiency as practical orthodontic treatment monitoring requires.
Jiaming Xie, Congyi Zhang 0001, Guangshun Wei, Peng Wang 0099, Guodong Wei, Wenxi Liu, Min Gu 0003, Ping Luo 0002, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.1
2023 No Reference Image Quality Assessment Via Quality Difference Learning
abstract
For human beings, there is a natural preference for judging the relative quality rather than directly predicting the quality score of an image. Based on this view, we propose an image quality difference learning network (IQDLNet) for evaluating image quality in a no-reference manner. Specifically, the proposed IQDLNet consists of a quality difference-aware network (QDAN) and a quality assessment network (QAN). The QDAN aims to predict the score difference between two randomly matched images and the QAN aims to predict the quality score of these two images. To further enhance the mutual understanding of image semantics, a semantic interaction module (SIM) is proposed with a dual regressor set up to carry out competitive learning in combination with the quality difference-aware feature. Experimental results on five IQA datasets demonstrate the superior performance of the proposed method over eight state-of-the-arts.
Jiaming Xie, Yu Luo 0004, Jie Ling 0002, Guanghui Yue 0001
ICME1
2023 Mammo-Net: Integrating Gaze Supervision and Interactive Information in Multi-view Mammogram Classification
Changkai Ji, Changde Du, Sheng Wang 0014, Chong Ma 0004, Jiaming Xie, Huiguang He, Dinggang Shen
MICCAI (7)6
2022 An Efficient Hardware Design for Accelerating Sparse CNNs With NAS-Based Models
abstract
Deep convolutional neural networks (CNNs) have achieved remarkable performance at the cost of huge computation. As the CNN models become more complex and deeper, compressing CNNs to sparse by pruning the redundant connection in the networks has emerged as an attractive approach to reduce the amount of computation and memory requirement. On the other hand, FPGAs have been demonstrated to be an effective hardware platform to accelerate CNN inference. However, most existing FPGA accelerators focus on dense CNN models, which are inefficient when executing sparse models as most of the arithmetic operations involve addition and multiplication with zero operands. In this work, we propose an accelerator with software–hardware co-design for sparse CNNs on FPGAs. To efficiently deal with the irregular connections in the sparse convolutional layers, we propose a weight-oriented dataflow that exploits element–matrix multiplication as the key operation. Each weight is processed individually, which yields low decoding overhead. Then, we design an FPGA accelerator that features a tile look-up table (TLUT) and a channel multiplexer (CMUX). The TLUT is designed to match the index between sparse weights and input pixels. Using TLUT, the runtime decoding overhead is mitigated by using an efficient indexing operation. Moreover, we propose a weight layout to enable efficient on-chip memory access without conflicts. To cooperate with the weight layout, a CMUX is inserted to locate the address. Finally, we build a neural architecture search (NAS) engine that leverages the reconfigurability of FPGAs to generate an efficient CNN model and choose the optimal hardware design parameters. The experiments demonstrate that our accelerator can achieve 223.4-309.0 GOP/s for the modern CNNs on Xilinx ZCU102, which provides a$2.4\times $–$12.9\times $speedup over previous dense CNN accelerators on FPGAs. Our FPGA-aware NAS approach shows$2\times $speedup over MobileNetV2 with 1.5% accuracy loss.
Yun Liang 0001, Liqiang Lu, Yicheng Jin, Jiaming Xie, Ruirui Huang, Jiansong Zhang 0001, Wei Lin 0016
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 FCNNLib: A Flexible Convolution Algorithm Library for Deep Learning on FPGAs
abstract
Convolution features huge complexity and demands high computation capability. Among hardware platforms, field programmable gate array (FPGA) emerges as a promising solution for its substantial available parallelism and energy efficiency. Besides, convolution can be implemented with different algorithms, including conventional, general matrix–matrix multiplication (GEMM), Winograd, and fast Fourier transformation (FFT) algorithms, which are diverse in arithmetic complexity, resource requirement, etc. Different convolutional neural network (CNN) models have different topologies and structures, favoring different convolution algorithms. In response, software libraries such as cuDNN provide a variety of computational primitives to support these algorithms. However, supporting such libraries on FPGAs is challenging. First, multiple algorithms can share the FPGA resources spatially as well as temporally, introducing either reconfiguration overhead or resource underutilization. Second, FPGA implementation remains a significant challenge for library developers. It typically requires significant specialized hardware knowledge. In this article, we proposeFCNNLib, an efficient and scalable convolution algorithm library on FPGAs. To coordinate multiple convolution algorithms on FPGAs, we develop three schedulings: 1) spatial; 2) temporal; and 3) hybrid, which exhibit different tradeoffs in latency and throughput. We explore these schedulings by balancing the reconfiguration overhead, resource utilization, and optimization objectives of the CNNs. Then, we provide efficient and tunable algorithm templates that allow performance tuning through performance and resource models. To arm the users,FCNNLibexposes a set of interfaces to support high-level application designs. We demonstrate the usability ofFCNNLibwith state-of-the-art CNNs.FCNNLibachieves up to$44.6\times $and$1.76\times $energy efficiency in various scenarios compared with software libraries for CPUs and GPUs, respectively.
Yun Liang 0001, Qingcheng Xiao, Liqiang Lu, Jiaming Xie
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2021 Secure synchronization of stochastic complex networks subject to deception attack with nonidentical nodes and internal disturbance
Jianwen Feng, Jiaming Xie, Jingyi Wang 0001, Yi Zhao 0002
Inf. Sci.2
2021 OMNI: A Framework for Integrating Hardware and Software Optimizations for Sparse CNNs
abstract
Convolution neural networks (CNNs) as one of today's main flavor of deep learning techniques dominate in various image recognition tasks. As the model size of modern CNNs continues to grow, neural network compression techniques have been proposed to prune the redundant neurons and synapses. However, prior techniques disconnect the software neural networks compression and hardware acceleration, which fail to balance multiple design parameters, including sparsity, performance, hardware area cost, and efficiency. More concretely, prior unstructured pruning techniques achieve high sparsity at the expense of extra performance overhead, while prior structured pruning techniques relying on strict sparse patterns lead to low sparsity and extra hardware cost. In this article, we propose OMNI, a framework for accelerating sparse CNNs on hardware accelerators. The innovation of OMNI stems from that it uses hardware amenable on-chip memory partition patterns to seamlessly engage the software CNN model compression and hardware CNN acceleration. To accelerate the compute-intensive convolution kernel, a promising hardware optimization approach is memory partition, which divides the original weight kernels into several groups so that the different hardware processing elements can simultaneously access the weight. We exploit the memory partition patterns including block, cyclic, or hybrid as a means of CNN compression patterns. Our software CNN model compression balances the sparsity across different groups and our hardware accelerator employs hardware parallelization coordinately with the sparse patterns, leading to a desirable compromise between sparsity and performance. We further develop performance models to help the designers to quickly identify the pattern factors subject to an area constraint. Last, we evaluate our design on application specific integrated circuit (ASIC) and field-programmable gate array (FPGA) platform. Experiments demonstrate that OMNI achieves 3.4×- 6.2× speedup for the modern CNNs, over a comparably ideal dense CNN accelerator. OMNI shows 114.7× energy efficiency improvement compared with GPU platform. OMNI is also evaluated on Xilinx ZC706 and ZCU102 FPGA platforms, achieving 41.5 GOP/s and 125.3 GOP/s, respectively.
Yun Liang 0001, Liqiang Lu, Jiaming Xie
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 FCNNLib: An Efficient and Flexible Convolution Algorithm Library on FPGAs
abstract
Convolutions can be implemented with different algorithms, which are diverse in arithmetic complexity, resource requirement, etc. Multiple algorithms can share the FPGA resources spatially as well as temporally, introducing either reconfiguration overhead or resource underutilization. In this paper, we propose an efficient library FCNNLib to coordinate multiple convolution algorithms on FPGAs. We develop three scheduling techniques: spatial, temporal, and hybrid, which exhibit different trade-offs in latency and throughput. We also expose a set of interfaces to arm the users. Experiments using modern CNNs demonstrate FCNNLib achieves up to 1.315X latency improvement compared with dedicated accelerators and 1.755X energy efficiency improvement compared with cuDNN.
Qingcheng Xiao, Liqiang Lu, Jiaming Xie, Yun Liang 0001
DAC3
2019 SPART: Optimizing CNNs by Utilizing Both Sparsity of Weights and Feature Maps
Jiaming Xie, Yun Liang 0001
APPT1
2019 An Efficient Hardware Accelerator for Sparse Convolutional Neural Networks on FPGAs
abstract
Deep convolutional neural networks (CNN) have achieved remarkable performance with the cost of huge computation. As the CNN model becomes more complex and deeper, compressing CNN to sparse by pruning the redundant connection in networks has emerged as an attractive approach to reduce the amount of computation and memory requirement. In recent years, FPGAs have been demonstrated to be an effective hardware platform to accelerate CNN inference. However, most existing FPGA architectures focus on dense CNN models. The architecture designed for dense CNN models are inefficient when executing sparse models as most of the arithmetic operations involve addition and multiplication with zero operands. On the other hand, recent sparse FPGA accelerators only focus on FC layers. In this work, we aim to develop an FPGA accelerator for sparse CNNs. To efficiently deal with the irregular connection in the sparse convolutional layer, we propose a weight-oriented dataflow that processes each weight individually. Then we design an FPGA architecture which can handle input-weight connection and weight-output connection efficiently. For input-weight connection, we design a tile look-up table to eliminate the runtime indexing match of compressed weights. Moreover, we develop a weight layout to enable high on-chip memory access. To cooperate with the weight layout, a channel multiplexer is inserted to locate the address which can ensure no data access conflict. Experiments demonstrate that our accelerator can achieve 223.4-309.0 GOP/s for the modern CNNs on Xilinx ZCU102, which provides a 3.6x-12.9x speedup over previous dense CNN FPGA accelerators.
Liqiang Lu, Jiaming Xie, Ruirui Huang, Jiansong Zhang 0001, Wei Lin 0016, Yun Liang 0001
FCCM2