EDBT 2026 Demo / reviewers in the wild / expert
Minghua Tang
dblp:35/1725
· DBLP profile ↗
24ranked-venue papers
15as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 11 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A multi-modal framework for generating and evaluating diverse jokes using large language models
Minghua Tang |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Are LLMs qualified evaluators and generators of one-liner jokes
Minghua Tang |
Neural Networks | 1 |
| 2026 | NoiseGuard: A Comprehensive Framework With Noise Modeling, Noise-Aware Training, and Noise Compensation for In-Memory Computing SoC
Guangyao Wang, Yizhe Chen, Yuexi Lv, Yuannuo Feng, Jenny Ma, Saiya Wang, Guilin Zhao, Yong Pei, Minghua Tang, Wang Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 12 |
| 2025 | CIM-BLAS: Computing-in-Memory Accelerator for BLASabstractBasic Linear Algebra Subprograms (BLAS) is a foundational software library for linear algebra kernels, which is widely used in scientific and engineering computing. Existing BLAS accelerations mainly rely on CPUs and GPUs. Many operations in BLAS are data intensive, so they are constrained by the limited memory bandwidth of CPUs and GPUs. The computing-in-memory (CIM) technology can effectively alleviate the memory wall bottleneck and is particularly suitable for accelerating BLAS. In this paper, we propose the first CIM accelerator for BLAS, CIM-BLAS, based on non-volatile memory. CIM-BLAS includes a unified floating-point pipeline to support high-precision arithmetics. High efficiency of the accelerator is achieved by developing configurable data flows to support various BLAS functions. Compared with GPU implementations, CIMBLAS demonstrates several orders of magnitude performance and energy efficiency improvements for executing level-1 and level-2 BLAS functions, and can achieve an energy efficiency improvement of 2.6-24.1 $\times$ for executing level-3 BLAS functions. The improvement increases with the size of the matrix, indicating excellent scalability of CIM-BLAS. Application-level evaluations also demonstrate the potential of CIM for accelerating BLAS. Rui Liu 0045, Zerun Li, Xiaoyu Zhang 0009, Xiaoming Chen 0003, Yinhe Han 0001, Minghua Tang |
DAC | 6 |
| 2025 | Key-concept thinking prompting for improved reasoning in large language models
Minghua Tang, Chen Bian, Xueling Zhong |
Neurocomputing | 1 |
| 2024 | MemSort: In-Memory Sorting ArchitectureabstractSorting is one of the most fundamental operations in computer programming and used in countless algorithms. The performance of traditional von Neumann computers running sorting is limited by the bandwidth between memories and processors. Computing-in-memory (CiM) is a promising technology which has the potential to solve the “memory wall” bottleneck. CiM is suitable for data-intensive applications, and it is ideal for accelerating large-scale data sorting. In this paper, we propose a novel in-memory sorting accelerator, named MemSort, based on a proposed in-memory comparison array design based on emerging non-volatile devices. MemSort supports three sort operations including counting sort, merging sort, and the combination of counting sort and merging sort. We build a performance model for the combination sort which enables flexible allocation of resources under given constraints to meet the requirements of various applications for sorting. The evaluation results show that MemSort shows significant performance improvement and energy efficiency at both the system level and application level when processing large-scale data sorting. Compared with the CPU implementation, MemSort achieves energy savings of 19.69-72.75x and speedups of 24.48-38.58 x with the same power constraint. MemSort's throughput is at least 4.86 x higher than that of the recent FPGA-based sorting accelerator FANS. MemSort exhibits more than 11 x throughput and 4.03 x area efficiency, compared with the recent CiM - based sorting accelerator, RIME. Rui Liu 0045, Xiaoyu Zhang 0009, Xinyu Wang 0040, Feng Min, Zhejian Luo, Xiaoming Chen 0003, Yinhe Han 0001, Minghua Tang |
ICCD | 8 |
| 2024 | Labelrepair: Sequence Labelling for Compilation Errors RepairabstractManual fixing of compilation errors could be a tedious and time-consuming task for novice programmers, and even for experienced ones. In recent years, an increasing number of automated repair techniques have been proposed to guide novice programmers and improve the efficiency of software development. Among them, learning-based automated repair techniques have achieved promising results in terms of repair accuracy. However, existing approaches neglect the time efficiency of patch generation, and often treat the compilation errors repair as a neural machine translation task. The end-to-end repair model decoding cannot be parallelized during the inference stage and suffers from redundant decoding search space. Furthermore, the large search space brought by the model poses a potential risk of semantic tampering. To this end, we propose Labelrepair, a novel repair technique that treats the repair of compilation errors as a sequence labelling task. Labelrepair discards the decoding model and converts the search for patches to a search for the error mapping actions between broken code and patch code pairs. In this way, the search space for patch tokens is plum-meted from the entire vocabulary to the size of edit action labels, and the logical semantics of the original program are preserved to some extent. The time complexity of inference is reduced from$O(n)$to$O$(1) owing to the parallel generation of edit action labels. Through a comprehensive evaluation of Labelrepair on two datasets, we demonstrate that Labelrepair is able to generate patches instantly (0.71ms on average), which is 28 times faster than existing end-to-end repair models. Compared with existing edit-based repair models, Labelrepair achieves the state-of-art repair accuracy. Deheng Yang, Yan Lei 0005, Huan Xie 0002, Minghua Tang, Maojin Li |
SANER | 5 |
| 2024 | IDSSI: Image Deturbulence with Semantic and Spatial-Temporal Information
Xiangqing Liu, Shaoan Yan, Yongguang Xiao, Jianbin Xie, Minghua Tang |
Pattern Recognit. | 8 |
| 2024 | Correction to: The position-based compression techniques for DNN model
Minghua Tang, Enrico Russo 0002, Maurizio Palesi |
J. Supercomput. | 1 |
| 2023 | FeCrypto: Instruction Set Architecture for Cryptographic Algorithms Based on FeFET-Based In-Memory ComputingabstractRecently, computing-in-memory (CiM) becomes a promising technology for alleviating the memory wall bottleneck. CiM is suitable for data-intensive applications, especially, cryptographic algorithms. Most current cryptographic accelerators are specific to a single function. It is expensive to accelerate different cryptographic algorithms with different accelerators. In this work, we first introduce a CiM architecture FeMIC that supports multioperand CiM operations, by exploring advantages of state-of-the-art ferroelectric field-effect transistors. Based on that, we propose a novel instruction set together with an accelerator architecture named FeCrypto which supports the acceleration of various cryptographic algorithms. Evaluation results show that FeCrypto has better performance and energy efficiency than software implementations. The energy-delay product (EDP) of FeCrypto is$118.4\times $and$1.93\times $lower than that of the dedicated AES accelerator AIM that is built based on phase-change memories (PCMs) and magnetic random-access memories (MRAMs), respectively. EDP is reduced by$44.7\times $compared with PCM-based EIM, a recent AES accelerator. Compared with MRAM-based EIM, the EDP overhead of FeCrypto for supporting multiple functions is 23.2%. Rui Liu 0045, Xiaoyu Zhang 0009, Zhiwen Xie, Xinyu Wang 0040, Zerun Li, Xiaoming Chen 0003, Yinhe Han 0001, Minghua Tang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2023 | EAF-WGAN: Enhanced Alignment Fusion-Wasserstein Generative Adversarial Network for Turbulent Image RestorationabstractBecause of optical distortion induced by atmospheric turbulence and the limitations of optical devices, acquired images of space objects are blurred and degraded. This effect results in anamorphosis in the output of two-dimensional imaging systems. In this work, we present a novel Enhanced Alignment Fusion-Wasserstein Generative Adversarial Network, called EAF-WGAN, for turbulent image restoration. This model is characterized by the innovative use of two modules, including the Align Module (AM) and the Feature Fusion Module (FFM) in the generator, especially in the process of feature fusion, in which 3D convolution is used. Through 3D convolution, the temporal and spatial information of the input image is obtained. Therefore, the formation mode and intensity of turbulence are not considered and the image can be reconstructed. The ability of EAF-WGAN is proved by algorithmically simulated data, physically simulated data, and real-world data. Xiangqing Liu, Zhenyang Zhao, Shaoan Yan, Jianbin Xie, Minghua Tang |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2023 | The position-based compression techniques for DNN model
Minghua Tang, Enrico Russo 0002, Maurizio Palesi |
J. Supercomput. | 1 |
| 2022 | FeMIC: Multi-Operands in-Memory Computing Based on FeFETsabstractThe “memory wall” bottleneck caused by the performance gap between processors and memories is getting worse. Computing-in-memory (CiM), a promising technology to alleviate the “memory wall” bottleneck, has recently attracted much attention. Conventional CiM architectures based on emerging nonvolatile devices have a major drawback that they need${N\,-\,1}$clock cycles to complete a CiM operation with${N}$operands, as they are natively designed for processing two operands. In this work, we propose FeMIC, a new CiM architecture based on ferroelectric field-effect transistors (FeFETs), which natively supports the computation of multiple operands. For a CiM operation with${N}$operands, FeMIC only needs$\left\lfloor {N/2} \right\rfloor$clock cycles. The simulation results based on a calibrated FeFET model reveal that FeMIC can significantly reduce the energy consumption when processing multi-operand CiM operations, compared with state-of-the-arts that use conventional CiM mechanisms. Rui Liu 0045, Xiaoyu Zhang 0009, Xiaoming Chen 0003, Yinhe Han 0001, Minghua Tang |
ASP-DAC | 5 |
| 2018 | The Suboptimal Routing Algorithm for 2D Mesh NetworkabstractDue to the huge routing algorithm search space for 2D mesh based Network-on-Chip (NoC), Divide-Conquer method is presented to effectively explore the search space. When using Divide-Conquer method, a large number of routing algorithms will be created. In order to get the final results in an acceptable time, a precise metric is needed to measure routing performance and discard the poor performance routings. In this paper, we propose a new routing performance metric, namely, network pressure. Network pressure has the following three advantages: (1) it could measure the whole network congestion state; (2) network pressure of a network and that of its partial component is highly related, under the same routing; (3) it is closely related with routing performance. Based on network pressure and Divide-Conquer method, high performance routing could be achieved. The obtained routing is called suboptimal routing due to the following two reasons: (1) there is only a little gap between its performance and that of the fully adaptive routings under both transpose1 and transpose2 traffics. (2) the search space of routing algorithms is systematically and widely exploited. Minghua Tang, Maurizio Palesi |
IEEE Trans. Computers | 1 |
| 2017 | The Repetitive Turn Model for Adaptive RoutingabstractFor 2D mesh based Network-on-Chip (NoC), the prohibited turns of routing algorithms should be repetitively distributed in order for the routing algorithms to be implemented by logic-based circuit. In this paper, we aim to exploit the designing space for logicbased routing algorithms, and propose new logic-based routing algorithms that outperform the state-of-the-art counterparts. Toward this direction, we firstly construct all routing algorithms for 5 x 5 2D mesh topology. Then we select those routing algorithms which have repetitive prohibited turns across both the network rows and columns. In addition, we chose those routing algorithms that have smaller routing pressures than Odd-Even routing algorithm. Then the routing algorithms for 2D mesh topology ranging from 6 x 6 to 15 x 15 are respectively constructed according to the prohibited turns distribution of the selected routing algorithms. Two routing algorithms that have smaller routing pressures than Odd-Even routing algorithm are obtained for all considered networks. The obtained logic-based routing algorithms are called as Repetitive Turn Model (RTM). Simulation results show that RTM could achieve up to 51% performance improvement as compared to Odd-Even routing algorithm. Minghua Tang, Xiaola Lin, Maurizio Palesi |
IEEE Trans. Computers | 1 |
| 2016 | Local Congestion Avoidance in Network-on-ChipabstractNetwork-on-Chip (NoC) has been made the communication infrastructure for many-core architecture. NoC are subject to congestion, which is claimed to be avoided by many researchers. However, there is no completely understanding of congestion in literature, which hinders its solution. Toward this direction, we firstly carry out study on congestion in this paper. We find that congestion usually occurs at a portion of nodes in a local network region. Moreover, local congestion will significantly decrease system performance and mostly impact some particular communication pairs. Then we attempt to solve local congestion by addressing different local region size, based on Divide-Conquer approach and routing pressure. It avoids congestion in every local region by keeping routing pressure of every local region minimum. Using different local region size will create different routings. Our study shows that the local region size is closely related with the routing performance. When local region size is 5 × 5 the optimal routing performance of large size network could be achieved. Minghua Tang, Xiaola Lin, Maurizio Palesi |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | An Offline Method for Designing Adaptive Routing Based on Pressure ModelabstractAs a scalable substitute of on-chip bus, network-on-chip (NoC) is proposed as the communication infrastructure in modern multi/many-core system-on-chip (SoC). Efficient communication in NoC is critical to the overall SoC performance. Although local congestion has an important impact on communication delay, it is barely taken into account when designing routing algorithms. In this paper, we propose an offline methodology of designing routing algorithm based on channel pressure model to address the local congestion issue. Specifically, the proposed methodology uses divide-conquer with the aim of generating high performance routing algorithms, which are able to balance the load over the network with a consequent reduction of local congestion. By using the proposed methodology, the obtained routing could achieve up to 37% performance improvement (in terms of average communication delay) as compared to the well-known odd-even routing algorithm for 15 × 15 network. Minghua Tang, Xiaola Lin, Maurizio Palesi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2015 | Routing Pressure: A Channel-Related and Traffic-Aware Metric of Routing AlgorithmabstractHow to precisely measure performance of routing algorithm is an important issue when studying routing algorithm of network-on-chip (NoC). The degree of adaptiveness is the most widely used metric in the literature. However, our study shows that the degree of adaptiveness cannot precisely measure performance of routing algorithm. It cannot account for why routing algorithm with high degree of adaptiveness may have poor performance. Simulation has to be carried out to evaluate performance of routing algorithm. In this paper, we propose a new metric of routing pressure for measuring performance of routing algorithm. It has higher precision of measuring routing algorithm performance than the degree of adaptiveness. Performance of routing algorithm can be evaluated through routing pressure without simulation. It can explain why congestion takes place in network. In addition, where and when congestion takes place can be pointed out without simulation. Minghua Tang, Xiaola Lin, Maurizio Palesi |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Reconfigurable All-Band RF CMOS Transceiver for GPS/GLONASS/Galileo/Beidou With Digitally Assisted CalibrationabstractThis paper presents a fully integrated reconfigurable all-band RF transceiver for GPS/GLONASS/ Galileo/Beidou in 55-nm CMOS. The transceiver incorporates three low-IF receivers (RXs) and one direct up-conversion BPSK transmitter (TX), which can be configured to receive any two global navigation satellite system (GNSS) signals or switched to process the Chinese Beidou(I) signals. A switching module is integrated to provide the connectivity between different RF front-end and IF channel (IFC), which will effectively simplify the design complexity of the IFC and save power consumption, while the GNSS signals are received. A flexible frequency plan with two frequency synthesizers is utilized to satisfy different local oscillator requirements of the transceiver. An optimized automatic frequency calibration scheme using an error compensation logic enables fast and high-precision calibration process for optimum phase-locked loop operation. Several digitally assisted calibration modules are integrated to ensure that the chip performance only shows a weak process, voltage and temperature (PVT) dependence. While drawing about 21.5-30.2 mA per RX channel from a 1.2-V supply, the RXs achieve an image rejection ratio more than 49 dB after I/Q mismatch calibration, an automatic gain control range of 88 dB, and an input-referred 1 dB compression point of better than -25 dBm with a minimum noise figure of about 2 dB. The output power of the TX is about 5 dBm with about 6% error vector magnitude (EVM) and 30-mA current from a 1.2-V supply. The whole transceiver consumes a die area of 2.8 × 3 mm2. Songting Li, Jiancheng Li, Xiaochen Gu, Hongyi Wang 0003, Jianfei Wu, Minghua Tang |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2010 | QLT: A flow control strategy for wormhole network-on-chipabstractThe fact that the on-chip network will suddenly get saturated indicates that it is necessary to devise an effective flow control strategy. Through tracing the status of the network buffer space we found that the payload of the on-chip network can not get beyond an upper bound to avoid vicious congestion. Specifically, quarter of the total network buffer space is such a threshold, which is termed Quarter Load Threshold (QLT). Based on this fact we present the Quarter Load Threshold (QLT) flow control strategy. The performance of the proposed strategy is evaluated by the open source simulator Noxim [2]. Simulation results show that the on-chip network runs smoothly and no serious congestion is encountered any more. Minghua Tang, Xiaola Lin |
AICCSA | 1 |
| 2010 | Network-on-Chip Routing Algorithms by Breaking Cycles
Minghua Tang, Xiaola Lin |
ICA3PP (1) | 1 |
| 2010 | Quarter Load Threshold (QLT) flow control for wormhole switching in mesh-based Network-on-Chip
Minghua Tang, Xiaola Lin |
J. Syst. Archit. | 1 |
| 2009 | An Advanced NoP Selection Strategy for Odd-Even Routing Algorithm in Network-on-Chip
Minghua Tang, Xiaola Lin |
ICA3PP | 1 |
| 2008 | A Novel Scheme to Balance the Cache Sharing in High Performance Computing SystemabstractCaching is an important technique to improve the performance of computer systems, especially for high performance computing systems. In traditional n-way set associative cache scheme, all sets have the same number of blocks. The blocks in a set cannot be used by other sets, resulting in inefficient cache utility for most applications with non-uniform cache access patterns. In this paper, we propose a novel cache sharing method, called the Way-Level Share Cache structure (WLSC), to enhance the caching performance by balancing the cache sharing. In addition to retaining the major properties of the traditional method, our method allows the blocks in some sets to be used by other sets and it can be done dynamically in the running processes. It can adapt to different applications. Simulation indicates that the proposed method can greatly increase the hit rate and thus improve the caching performance in the systems. Minghua Tang, Xiaola Lin |
HPCC | 1 |