VLDB 2026 Research / reviewers in the wild / expert
Wenyu Sun
dblp:92/6021
· DBLP profile ↗
25ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic-aware Sparsity and Adaptive Octree Partition for Efficient Deployment of Point Cloud Models on Edge Devices
Weigui Li, Le Qiu, Weichen Gao, Wenyu Sun, Yongpan Liu |
ISCAS | 4 |
| 2025 | MPICC: Multiple-Precision Inter-Combined MAC Unit with Stochastic Rounding for Ultra-Low-Precision TrainingabstractRecent studies have proved the feasibility of ultra-low-precision (≤ 8-bit) training. However, most of the existing operational circuits support only a few higher precisions (FP16, FP32, etc.) for multiplication, and the bit width for accumulation cannot be reduced, which results in low area and power efficiencies. In this paper, we propose MPICC, a multiple-precision Multiply-Accumulate (MAC) unit designed for ultra-low-precision training. It supports inter-combined computations across 18 different precision combinations, including LOG4 (radix-4 FP4), FP4, FP6, FP8, and INT4. It also reduces the accumulation precision from FP16 to FP12 through an optimized Stochastic Rounding (SR) strategy to further save logic resources. Moreover, a low-cost emulating controller, which time-division multiplexed the low-precision MAC unit, is also designed to accomplish high-precision computations for critical DNN layers. Compared with the existing multiple-precision computing units, the area/energy efficiencies of this design are improved by 1.17×/1.19× at FP8, and 4.69×/3.64× at FP4, respectively. The SR strategy further reduces the area/power consumption by 15.6%/14.9% of the floating-point accumulator. Leran Huang, Yongpan Liu, Xinyuan Lin, Chenhan Wei, Wenyu Sun, Zengwei Wang, Boran Cao, Xiaoxia Fu |
ASP-DAC | 5 |
| 2025 | Point2skh: End-to-end Parametric Primitive Inference from Point Clouds with Improved Denoising TransformerabstractRecovering the CAD command sequence from the point cloud is an essential component in CAD reverse engineering. In this paper, we strive to solve this problem from both the perspectives of artificial intelligence and the procedures of procedural CAD models. We propose a CAD reconstruction method based on an end-to-end point-to-sketch network (Point2Skh) that can produce the CAD modeling sequence from the input geometrical point cloud by recovering the inverse sketch-and-extrude process. The point cloud is first segmented into point sets corresponding to the same extrusion. The modeling sequence can then be recovered by combining the network prediction of each point set. The proposed Point2Skh can detect and infer command vectors of sketch curves (line, arc, and circle) and the extrusion operation from the input point cloud of a single extrusion. By directly representing the sketch with its curves and inferring the command parameters, accurate sketch reconstruction is produced, which further leads to precise CAD reconstruction with sharp edges. The produced CAD modeling sequence is human-interpretable and can be readily edited by importing it into CAD tools. Experiments show that the Chamfer Distance (CD) between the predicted results and the ground truth is 0.312, and the primitive type and parameter accuracy are 93.87% and 83.24%, respectively. • We propose a CAD reconstruction method based on an point-to-sketch network (Point2Skh). • The CAD modeling sequence can be directly predicted from the input point cloud and imported into CAD tools. • The Point2Skh can directly infer the CAD command type and parameters from the extrusion point cloud. Cheng Wang 0026, Wenyu Sun, Xinzhu Ma |
Comput. Aided Des. | 2 |
| 2025 | Improving Transformer Inference Through Optimized Nonlinear Operations With Quantization-Approximation-Based StrategyabstractTransformers have recently shown significant performance across various tasks, such as natural language processing (NLP) and computer vision (CV). However, the performance comes at the cost of large memory and computation overhead. Existing researches primarily focus on accelerating matrix multiplication (MatMul) through techniques like quantization and pruning, notably increasing the proportion of nonlinear operations in inference runtime. Meanwhile, previous approaches designed for nonlinear operations struggle with inefficient implementation as they are incapable of achieving both computation and memory efficiency. Additionally, these methods often require retraining or fine-tuning leading to substantial costs and inconveniences. To overcome these problems, we propose efficient implementation of nonlinear operations with quantization-approximation-based strategy. Through an in-depth analysis of the dataflow and data distribution of nonlinear operations, we design distinct quantization and approximation strategies tailored for different operations. Specifically, log2 quantization and power-of-two factor quantization have been employed in Softmax and LayerNorm, complemented by logarithmic function and low-precision statistic calculation as approximation strategies. Furthermore, the proposed efficient GeLU implementation integrates a nonuniform lookup procedure alongside low-bit-width quantization. Experimental results demonstrate negligible accuracy drops without the need for retraining or fine-tuning. By implementing the hardware design, it achieves$3.14\times - 6.34\times $energy-efficiency and$3.01\times - 10.1\times $area-efficiency improvements compared to state-of-the-art application-specific-integrated-circuit (ASIC) designs. In system-level evaluation, substantial speedup and reductions in energy consumption of 15% to 35% are achieved for end-to-end inference across both GPU and ASIC accelerator platforms. Wenxun Wang, Wenyu Sun, Yongpan Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | A Multichiplet Computing-in-Memory Architecture Exploration Framework Based on Various CIM DevicesabstractComputing-in-memory (CIM) architectures based on various devices, such as resistive random access memory, SRAM, DRAM, etc., have demonstrated promising energy efficiency. Single-device-based CIM chips show different advantages on performance, power, or area metrics under different workload/operators sizes and application requirements. Some nonidealities, such as the write endurance of some nonvolatile devices, also influence the design choices. Motivated by the emerging 2.5-D/3-D chiplet integration, this work aims to combine the advantages of CIM/storage chips based on different devices, and proposes a design exploration framework to combine the advantages of CIM chips based on these devices in a 3-D-stack architecture. This work proposes: 1) an evaluation method for the power, performance, and area metrics of the multichiplet CIM architecture; 2) an abstraction for the single-device-based CIM chiplets and artificial intelligence algorithm operators; and 3) a mapping and optimization strategy to explore the 2.5-D/3-D CIM chiplet set. The effectiveness of the mapping strategy is verified with a small-scale brute-force search. The proposed design exploration framework can help to find a better-multichiplet CIM architecture. Under a simple design case, the proposed 3-D CIM architecture shows$4.68\times $–$53.32\times $energy efficiency compared with the single CIM chip baselines. The abstracted chiplet library is open-source available in the open-sourcehttps://github.com/dai0dai/3D_CIM_Chiplet_Architecture_Exploration. Zhuoyu Dai, Feibin Xiang, Xiangqu Fu, Yifan He 0003, Wenyu Sun, Yongpan Liu, Guanhua Yang, Feng Zhang 0014, Jinshan Yue, Ling Li 0013 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Semantic Guided Fine-Grained Point Cloud Quantization Framework for 3D Object DetectionabstractUnlike the grid-paced RGB images, network compression, i.e.pruning and quantization, for the irregular and sparse 3D point cloud face more challenges. Traditional quantization ignores the unbalanced semantic distribution in 3D point cloud. In this work, we propose a semantic-guided adaptive quantization framework for 3D point cloud. Different from traditional quantization methods that adopt a static and uniform quantization scheme, our proposed framework can adaptively locate the semantic-rich foreground points in the feature maps to allocate a higher bitwidth for these "important" points. Since the foreground points are in a low proportion in the sparse 3D point cloud, such adaptive quantization can achieve higher accuracy than uniform compression under a similar compression rate. Furthermore, we adopt a block-wise fine-grained compression scheme in the proposed framework to fit the larger dynamic range in the point cloud. Moreover, a 3D point cloud based software and hardware co-evaluation process is proposed to evaluate the effectiveness of the proposed adaptive quantization in actual hardware devices. Based on the nuScenes dataset, we achieve 12.52% precision improvement under average 2-bit quantization. Compared with 8-bit quantization, we can achieve 3.11× energy efficiency based on co-evaluation results. Xiaoyu Feng, Zongkai Zhang, Wenyu Sun, Yongpan Liu |
ASP-DAC | 4 |
| 2023 | Razor SNN: Efficient Spiking Neural Network with Temporal Embeddings
Yuan Zhang 0020, Jian Cao 0002, Wenyu Sun, Yuan Wang 0001 |
ICANN (5) | 4 |
| 2023 | SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer InferenceabstractTransformers have shown remarkable performance in both natural language processing (NLP) and computer vision (CV) tasks. However, their real-time inference speed and efficiency are limited due to the inefficiency in Softmax and Layer Normalization (LayerNorm). Previous works based on function approximation suffer from inefficient implementation as they place emphasis on computation while disregarding memory overhead concerns. Moreover, such methods rely on retraining to compensate for approximation error which can be costly and inconvenient. In this paper, we present SOLE, a hardware-software co-design for Softmax and LayerNorm which is composed of E2Softmax and AILayerNorm. E2Softmax utilizes log2 quantization of exponent function and log-based division to approximate Softmax while AILayerNorm adopts low-precision statistic calculation. Compared with state-of-the-art designs, we achieve both low-precision calculation and low bit-width storage on Softmax and LayerNorm. Experiments show that SOLE maintains inference accuracy without retraining while offering orders of magnitude speedup and energy savings over GPU, achieving 3.04×, 3.86× energy-efficiency improvements and 2.82×, 3.32× area-efficiency improvements over prior state-of-the-art custom hardware for Softmax and LayerNorm, respectively. Wenxun Wang, Shuchang Zhou 0001, Wenyu Sun, Peiqin Sun, Yongpan Liu |
ICCAD | 3 |
| 2023 | Simultaneously Training and Compressing Vision-and-Language Pre-Training ModelabstractModel compression is an essential step for large-scale pre-training models toward practical application and deployment on the edge device. However, when conventional compression methods following ‘pre-training then compressing’ two-phase pipeline are applied to Vision-and-Language Pre-training (VLP) models, it will lead to a high calculation and memory overhead. In this work, we break the two-phase pipeline and propose an efficient and effective one-phase VLP model compression mechanism, namedREDUCER, which stands for ‘simultaneously training and compREssing’ VLP model via progressive moDUle replaCing and nEtworkRewiring. Specifically, REDUCER consists of three insightful designs. Firstly, we design a one-phase compression framework to train and compress the VLP model simultaneously to avoid the extra calculation and memory cost caused by an isolated model compression phase in the conventional two-phase pipeline. Secondly, we propose an adaptive progressive module replacing mechanism to compress the model depth free from explicit knowledge distillation losses, relieving the multi-task optimization problems. Thirdly, we integrate pruning techniques into VLP model compression to simultaneously compress the model in width and depth. Overall, we obtain a lightweight VLP model with only one pre-training phase, and it is the first one-phase compression method for VLP models. Extensive experiments have been conducted on representative VLP models,i.e., ClipBERT and VICTOR, and the experimental results show a superior trade-off between performance and efficiency. Qiaosong Qi, Aixi Zhang, Yue Liao, Wenyu Sun, Si Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | Sparsity-Aware Non-Volatile Computing-In-Memory Macro with Analog Switch Array and Low-Resolution Current-Mode ADCabstractNon-volatile computing-in-memory (nvCIM) is a novel architecture used for deep neural networks (DNNs) because it can reduce the movement of data between computing units and memory units. As sparsity has made great progress in DNNs, the existing nvCIM architecture is only optimized for structured sparsity but little for unstructured sparsity. To solve this problem, the sparsity-aware nvCIM macro is proposed to improve the computing performance and network classification accuracy, and to support both structured and unstructured sparsity. First, the analog switch array is used to take advantage of the structured sparsity and to improve the computing parallelism. Second, the low-resolution current-mode analog-to-digital converter (CMADC) is designed to optimize the unstructured sparsity. Experimental results show that the peak equivalent energy efficiency of the proposed nvCIM macro is 9.1 TOPS/W (A8W8, 8-bit activations and 8-bit weights) with only 0.51% accuracy loss, and 584.9 TOPS/W (A1W1), which is 4.8 -$7.5\times$compared to the state-of-the-art nvCIM macros. Yifan He 0003, Jinshan Yue, Wenyu Sun, Huazhong Yang, Yongpan Liu |
ASP-DAC | 4 |
| 2022 | Dynamic CNN Accelerator Supporting Efficient Filter Generator with Kernel Enhancement and Online Channel PruningabstractDeep neural network achieves exciting performance in several tasks with heavy storing and computing costs. Previous works adopt pruning-based methods to slim deep network. For traditional pruning, either the convolution kernel or the network inference is static, which cannot fully compress the model parameter and restrains their performance. In this paper, we propose an online pruning algorithm to support dynamic kernel generation and dynamic network inference at the same time. Two novel techniques including the filter generator and the importance-level based channel pruning are proposed. Moreover, we validate the success of the proposed method by the implementation on Ultra96-v2 FPGA. Compared with state-of-art static or dynamic pruning methods, our method can reduce the top-5 accuracy drop by nearly 50% for ResNet model on ImageNet at similar compressing level. It can also achieve better accuracy while up to 50% fewer weights are reduced to be saved on chip. Wenyu Sun, Wenxun Wang, Yongpan Liu |
ASP-DAC | 2 |
| 2022 | Toward Low-Bit Neural Network Training Accelerator by Dynamic Group AccumulationabstractLow-bit quantization is a big challenge for neural network training. Conventional training hardware adopts FP32 to accumulate the partial-sum result, which seriously degrades energy efficiency. In this paper, a technology called dynamic group accumulation (DGA) is proposed to reduce the accumulation error. First, we model the proposed group accumulation method and give the optimal DGA algorithm. Second, we design a training architecture and implement a hardware-efficient DGA unit. Third, we make a comprehensive analysis of the DGA algorithm and training architecture. The proposed method is evaluated on CIFAR and ImageNet datasets, and results show that DGA can reduce accumulation bit-width by 6 bits while achieving the same precision as the static group method. With the FP12 DGA, the CNN algorithm only loses 0.11% accuracy in ImageNet training, and our architecture saves 32% of power consumption compared to the FP32 baseline. Yixiong Yang, Ruoyang Liu, Wenyu Sun, Jinshan Yue, Huazhong Yang, Yongpan Liu |
ASP-DAC | 3 |
| 2022 | C-RRAM: A Fully Input Parallel Charge-Domain RRAM-based Computing-in-Memory Design with High Tolerance for RRAM VariationsabstractPrevious RRAM-based computing-in-memory works mainly focus on the current-domain approach. However, the performance and accuracy of current computation are limited by large read currents of RRAM cells and their variations. This work presents a novel RRAM-based charge-domain design, C-RRAM, to resolve these limitations. A 3TlRlC cell is proposed to execute MAC operations by capacitor discharging. The resistance variations can be tolerated with reasonable discharging time, which is accelerated by a positive feedback loop. Also, the output from each cell is accumulated by charge sharing instead of summing currents to eliminate the static current path in readout circuits. In this way, robust and efficient RRAM-based CIM operation is enabled with fully input parallelism. A $512\times 514$ RRAM array is implemented to evaluate the benefits of the proposed charge-domain approach. The experiment results show that C-RRAM can suppress the output variations by $41\times$ and incur negligible accuracy loss for ResNet-18 on Cifar10 dataset. Compared to previous ITIR current-domain RRAM designs, it achieves $1.2\times$ energy efficiency and $127\times$ area efficiency due to improved parallelism. Yifan He 0003, Jinshan Yue, Wenyu Sun, Lu Zhang 0074, Yongpan Liu |
ISCAS | 4 |
| 2022 | Efficient Neural Networks with Spatial Wise Sparsity Using Unified Importance MapabstractExploiting neural network sparsity is one of the most important directions to accelerate CNN executions. Plenty of techniques are proposed to exploit neural network sparsity, where spatial-wise pruning is quite effective for input image. However, previous spatial-wise pruning methods need nontrivial hardware overhead for dynamic execution, due to layer-by-layer binary sampling and online scheduling. This paper proposes a structured configured, spatial-wise pruning technique. Numerous computation will be saved by skipping unimportant region. By using a unified importance map, the computing graph could be compiled in advance to make it more hardware friendly. Additionally, due to multi-level measurement of importance for each region, our method can have a better performance on various tasks. On image classification task, the method can have around 50% fewer top-1 accuracy drop than previous spatialwise pruning methods at similar sparse level. On super resolution and image deraining task, the method can bring $5 \times$ to $19 \times$ acceleration while causing neglectable effect on reconstruction quality. Hardware implementation is also included. Wenyu Sun, Wenxun Wang, Zhuqing Yuan, Yongpan Liu |
ISCAS | 2 |
| 2021 | Part Uncertainty Estimation Convolutional Neural Network For Person Re-IdentificationabstractDue to the large amount of noisy data in person re-identification (ReID) task, the ReID models are usually affected by the data uncertainty. Therefore, the deep uncertainty estimation method is important for improving the model robustness and matching accuracy. To this end, we propose a part-based uncertainty convolutional neural network (PUCNN), which introduces the part-based uncertainty estimation into the baseline model. On the one hand, PUCNN improves the model robustness to noisy data by distributilizing the feature embedding and constraining the part-based uncertainty. On the other hand, PUCNN improves the cumulative matching characteristics (CMC) performance of the model by filtering out low-quality training samples according to the estimated uncertainty score. The experiments on both non-video datasets, the noised Market-1501 and DukeMTMC, and video datasets, PRID2011, iLiDS-VID and MARS, demonstrate that our proposed method achieves encouraging and promising performance. Wenyu Sun, Jiyang Xie 0001, Jiayan Qiu, Zhanyu Ma |
ICIP | 1 |
| 2020 | High-Quality Single-Model Deep Video Compression with Frame-Conv3D and Multi-frame Differential Modulation
Wenyu Sun, Weigui Li, Zhuqing Yuan, Huazhong Yang, Yongpan Liu |
ECCV (30) | 1 |
| 2020 | A derivative-free algorithm for spherically constrained optimization
Min Xi, Wenyu Sun, Yannan Chen |
J. Glob. Optim. | 2 |
| 2019 | AERIS: area/energy-efficient 1T2R ReRAM based processing-in-memory neural network system-on-a-chipabstractReRAM-based processing-in-memory (PIM) architecture is a promising solution for deep neural networks (NN), due to its high energy efficiency and small footprint. However, traditional PIM architecture has to use a separate crossbar array to store either positive or negative (P/N) weights, which limits both energy efficiency and area efficiency. Even worse, imbalance running time of different layers and idle ADCs/DACs even lower down the whole system efficiency. This paper proposes AERIS, an Area/Energy-efficient 1T2R ReRAM based processing-In-memory NN System-on-a-chip to enhance both energy and area efficiency. We propose an area-efficient 1T2R ReRAM structure to represent both P/N weights in a single array, and a reference current cancelling scheme (RCS) is also presented for better accuracy. Moreover, a layer-balance scheduling strategy, as well as the power gating technique for interface circuits, such as ADCs/DACs, is adopted for higher energy efficiency. Experiment results show that compared with state-of-the-art ReRAM-based architectures, AERIS achieves 8.5x/1.3x peak energy/area efficiency improvements in total, due to layer-balance scheduling for different layers, power gating of interface circuits, and 1T2R ReRAM circuits. Furthermore, we demonstrate that the proposed RCS compensates the non-ideal factors of ReRAM and improves NN accuracy by 5.2% in the XNOR net on CIFAR-10 dataset. Jinshan Yue, Yongpan Liu, Fang Su, Shuangchen Li, Zhibo Wang 0004, Wenyu Sun, Xueqing Li 0002, Huazhong Yang |
ASP-DAC | 7 |
| 2019 | On semi-definiteness and minimal H-eigenvalue of a symmetric space tensor using nonnegative polynomial optimization techniques
Liqun Qi 0001, Wenyu Sun |
Signal Process. Image Commun. | 3 |
| 2019 | Design Methodology for TFT-Based Pseudo-CMOS Logic Array With Multilayer Interconnection Architecture and Optimization AlgorithmsabstractThin-film transistor (TFT) circuits are important for flexible electronics which are promising in the area of wearable devices and Internet of Things. However, most flexible TFT technologies only have unipolar devices and the process variation and defective rate are relatively high, which impose challenges to TFT circuit design. In this paper, we propose a novel logic array design based on pseudo-CMOS logic to address the problems of unipolar TFT circuit design. A multilayer interconnection architecture is presented to improve the routability of circuit and the area efficiency. Cell mapping and wire routing algorithms, which aim to map the logic gates of circuit to logic array and then route the interconnection wires, are devised to improve the performance of circuit in consideration of parameter variations of TFT and meanwhile enhance the routability. The experimental results show that the proposed logic array along with design methodologies can reduce more than 80% area compared with transistor level scheme and help to improve performance significantly. Qinghang Zhao, Wenyu Sun, Jiaqing Zhao, Jian Zhao 0004, Hailong Yao 0002, Tsung-Yi Ho, Huazhong Yang, Yongpan Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | Mechanical strain and temperature aware design methodology for thin-film transistor based pseudo-CMOS logic arrayabstractThin-film transistor (TFT) circuits are facing the challenges of unipolar device, process variation, and yield problems, which can be addressed by pseudo-CMOS logic array with multi-layer interconnect. However, existing design methodology does not take mechanical strain and temperature into consideration which may seriously affect the carrier mobility of TFT and thus the performance of whole logic array circuits. This paper presents a novel cell mapping algorithm including intrarow mapping step and inter-row mapping step for flexible logic array to mitigate the mobility influence. Experimental results indicate that there is more than 40% performance improvement in critical path delay at best case with the proposed algorithm. Wenyu Sun, Qinghang Zhao, Fei Qiao, Tsung-Yi Ho, Huazhong Yang, Yongpan Liu |
ASP-DAC | 1 |
| 2017 | Design Methodology for Thin-Film Transistor Based Pseudo-CMOS Logic Array with Multi-Layer Interconnect ArchitectureabstractThin-film transistor (TFT) circuits are important for flexible electronics which are promising in the area of wearable devices. However, most TFT technologies only have unipolar devices and the process variation and defective rate are relatively high, which impose challenges to TFT circuit design. In this paper, we propose a novel logic array based on pseudo-CMOS logic to address the problem of unipolar TFT circuit design. A multi-layer interconnect architecture and wire routing methodology are presented to improve the routability and meanwhile the area efficiency. The experimental results show that the proposed logic array reduces more than 80% area compared with transistor level scheme. Qinghang Zhao, Yongpan Liu, Wenyu Sun, Jiaqing Zhao, Hailong Yao 0002, Huazhong Yang |
DAC | 3 |
| 2017 | An 8b 0.8kS/s configurable VCO-based ADC using oxide TFTs with Inkjet printing interconnectionabstractFlexible electronic is a promising technology for flexible and large-area sensing IoT applications, where ADC is a fundamental component This paper proposes a configurable and flexible VCO-based ADC, implemented with Oxide Thin-Film Transistors(TFT) technology. A VCO with four connecting modes is designed to configure the VCO-based ADC working under different power and resolutions. An Inkjet printing interconnection technology is introduced to enable the configurability of ADC, even after all TFT transistors are fabricated. It allows the ADC to be customized for different applications and avoids fabrication failure of devices. Experimental results show that the proposed ADC achieves a 0.8kS/s sampling rate. Its power consumption ranges from 541 to 866uW with ENOB from 3 to 6b. Wenyu Sun, Qinghang Zhao, Fei Qiao, Yongpan Liu, Huazhong Yang |
ISCAS | 1 |
| 2016 | HW/SW co-design of nonvolatile IO system in energy harvesting sensor nodes for optimal data acquisitionabstractEnergy harvesting has been widely investigated as a promising alternative for future wearable sensors or internet-of-things. However, power and performance overhead is induced when IO operations are interrupted by power failures because non-preemptive characteristic of IO operations causes expensive re-executions. Furthermore, the state-of-art IO devices need long and power hungry initializing process, which makes IO operations inefficient in transient powered systems. This paper proposed a HW/SW co-design approach for nonvolatile IO system to maximize data acquisition. A ferroelectric flip-flop based nonvolatile IO architecture is adopted to reduce IO initialization overhead by 3-4 orders of magnitude. Based on the nonvolatile IO interface, we further formulate the optimal data acquisition as an INLP problem and a risk-aware online scheduler is presented to solve the problem efficiently. Experimental results show that the proposed HW/SW co-design architecture improves data acquisition by 2-5 times compared with conventional HW/SW architecture. Yongpan Liu, Chun Jason Xue, Zhangyuan Wang, Wenyu Sun, Jiwu Shu, Huazhong Yang |
DAC | 7 |
| 2013 | Positive Semidefinite Generalized Diffusion Tensor Imaging via Quadratic Semidefinite ProgrammingabstractThe positive definiteness of a diffusion tensor is important in magnetic resonance imaging because it reflects the phenomenon of water molecular diffusion in complicated biological tissue environments. To preserve this property, we represent it as an explicit positive semidefinite (PSD) matrix constraint and some linear matrix equalities. The objective function is the regularized linear least squares fitting for the log-linearized Stejskal--Tanner equation. The regularization term is the heuristic nuclear norm of the PSD matrix, since we expect it to be of low rank. In this way, we establish a convex quadratic semidefinite programming (SDP) model, whose global solution exists. The optimal solution could be solved by three efficient methods. While there are two state-of-the-art solvers---SDPT3 and QSDP---for the primal problem, we design a new augmented Lagrangian based alternating direction method (ADM) for the dual problem. Sensitivity analyses on the coefficients of the optimal diffusion tensor and the optimal objective function value with respect to noise-corrupted signals are presented. Experiments on synthetic data with multiple fibers show that the new method is robust to the Rician noise and outperforms several existing methods. Furthermore, when the fiber orientation distribution function is considered, the new method is competitive with the Q-ball imaging. Using the human brain data, we illustrate that the new method could capture the crossing of three nervous fiber bundles. Additionally, the new method generates positive definite generalized diffusion tensors in all voxels, while the unconstrained least squares fitting fails. Finally, we confirm that the ADM solver is more efficient than SDPT3 and QSDP for this special problem. Yannan Chen, Yuhong Dai, Deren Han, Wenyu Sun |
SIAM J. Imaging Sci. | 4 |