VLDB 2026 Research / reviewers in the wild / expert
Qiang Li 0021
dblp:72/872-21
· DBLP profile ↗
37ranked-venue papers
0as first author
19since 2021 · last 2026
0000-0001-9503-995XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 13 since 2021Computer networks · 7 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Hierarchical End-to-End Learning Framework for MIMO Load Modulation Under Stochastic Channels and Malicious JammingabstractWireless channels are inherently vulnerable to malicious jamming. Recently, end-to-end (E2E) learning has emerged as a promising paradigm for intelligent physical-layer design, particularly in multiple-input multiple-output (MIMO) load modulation array (LMA) systems, due to its channel-adaptive capability. By treating jamming as part of the channel state, we investigate robust E2E communication under adversarial conditions. However, two major challenges hinder practical deployment: (1) the high training and communication overhead caused by frequent retraining under dynamic channels and jamming, and (2) the unrealistic reliance on perfect channel state information (CSI) and feedback in resource-constrained systems. To address these issues, we propose Channel-Adaptive Hierarchical End-to-End learning (CAHE), a framework that reformulates E2E learning under stochastic channels as a dynamic mapping from channel realizations to optimal network parameters. CAHE consists of a channel-adaptive generator layer that dynamically produces encoder and decoder parameters, and an E2E learning layer that deploys these parameters for adaptive transceiver design. To reduce CSI feedback, we introduce a compact channel embedding mechanism for efficient representation. For jamming scenarios, we further propose a differential block-averaged channel estimation method to enable accurate anti-jamming channel estimation. In addition, we design a jamming-resilient network that embeds convolutional neural network (CNN)-learned jamming features into the decoder generator, enhancing perception of the communication state and improving robustness to random and dynamic jamming. Extensive analysis demonstrates the interpretability and feasibility of CAHE, and simulations confirm its significant performance gains over existing schemes. Ercong Yu, Qiang Li 0021, Hongyang Chen 0001 |
IEEE Trans. Commun. | 2 |
| 2026 | RIS-Based QAM Modulation via Composite PSK Spatial Superposition and Grouping Optimization
Ziyun Yue, Ercong Yu, Qiang Li 0021, Hongyang Chen 0001, Zhu Han 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2025 | End-to-End Learning for MIMO Load Modulation in Stochastic Channels: A Hierarchical ApproachabstractEnd-to-end (E2E) learning has emerged as a key technology for intelligent physical-layer communications, demonstrating significant potential in multiple-input multiple-output (MIMO) load-modulation array (LMA) systems. However, practical deployment faces two major challenges: high online training overhead due to retraining on new channel realizations and impractical reliance on perfect channel state information (CSI) and feedback in resource-constrained systems. To address these issues, we propose Channel-Adaptive Hierarchical End-to-End Learning (CAHE), a novel framework that reformulates E2E learning under stochastic channels as a dynamic mapping problem between channel realizations and optimal E2E network parameters. CAHE comprises a channel-adaptive generator layer, which dynamically generates encoder and decoder parameters tailored to each channel, and an E2E learning layer, which deploys these parameters for adaptive encoding and decoding. Moreover, to reduce channel feedback overhead, we introduce an efficient channel embedding mechanism, achieving up to 75% lossless compression of channel data. Furthermore, we develop a pilot-assisted alternative scheme to address scenarios where perfect CSI is unavailable. Simulation results demonstrate that CAHE significantly outperforms existing methods, achieving up to 9 dB gain at a bit error rate of 3×10−5, while the pilot-based scheme exhibits robust performance in imperfect CSI scenarios. Ercong Yu, Qiang Li 0021, Hongyang Chen 0001 |
GLOBECOM | 2 |
| 2025 | RIS-Based Composite Phase Shift Keying: An Efficient Rectangular QAM ApproachabstractThe reconfigurable intelligent surface (RIS)-based carrier modulation technology, which operates without conventional radio-frequency chains, offers hardware simplicity and cost efficiency. In particular, most current RIS-based amplitude modulation schemes rely on the ON/OFF state switching or amplifier gain control of RIS elements, and ambiguity-free high-order phase modulation requires correspondingly highresolution phase shifts. However, practical hardware limitations often impose a low-resolution phase shift and near-constant reflecting amplitude of RIS elements, complicating high-order modulation. To address this, we propose an efficient composite phase shift keying (CPSK) scheme to generate virtual highorder rectangular quadrature amplitude modulation (RQAM) signals at the receiver, referred to as CPSK-RQAM. By decomposing high-order RQAM signals into multiple low-order quadrature phase shift keying signal components, with each component transmitted by a distinct RIS group, CPSK-RQAM eliminates the need for additional amplitude control and ensures robustness to phase-shift quantization errors. Unlike the similar schemes based on conventional I/Q decomposition and equal grouping, CPSK-RQAM achieves a significantly larger minimum Euclidean distance by employing a tailored grouping criterion. Furthermore, we derive the closed-form expressions for the approximate symbol error probability (SEP) of CPSK-RQAM under maximum likelihood detection. Simulation results validate the theoretical analysis and demonstrate the superiority of CPSKRQAM over the state-of-the-art schemes. Ziyun Yue, Erbo Jizi, Ercong Yu, Qiang Li 0021, Hongyang Chen 0001, Zhu Han 0001 |
ICC | 4 |
| 2025 | Graph-Based Joint Client Clustering and Resource Allocation for Wireless Distributed Learning: A New Hierarchical Federated Learning Framework With Non-IID DataabstractHierarchical federated learning (HFL) is a key technology enabling distributed learning with reduced communication overhead. However, practical HFL systems encounter two major challenges: limited resources and data heterogeneity. In particular, limited resources can result in intolerable system latency, while heterogeneous data across clients can significantly degrade model accuracy and convergence rates. To address these issues and fully leverage the potential of HFL, we propose a novel framework called graph-based joint client and resource orchestration. This framework addresses the challenges of practical networks through joint client clustering and resource allocation. First, we propose a learning process where edge servers employ hypernetworks to achieve edge aggregation. This method can generate personalized client models and extract data distributions without directly exposing data distributions. Then, to characterize the joint effects of limited resources and data heterogeneity, we propose a graph-based modeling method and formulate a joint optimization problem that aims to balance data distributions and minimize latency. Subsequently, we propose a graph neural network-based algorithm to tackle the formulated problem with low-complexity optimization. Numerical results demonstrate significant benefits over existing algorithms in terms of convergence latency, model accuracy, scalability, and adaptability to new distributions. Ercong Yu, Shanyun Liu, Qiang Li 0021, Hongyang Chen 0001, H. Vincent Poor, Shlomo Shamai |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | A Compact Sub-nW/kHz Relaxation Oscillator Using a Negative-Offset Comparator With Chopping and Piecewise Charge-Acceleration in 28-nm CMOSabstractThis work presents a compact and power-efficient kHz-range relaxation (RC) oscillator with robust performance against temperature and voltage variations. By deliberately introducing a negative-offset voltage into the comparator, an offset cancellation scheme leveraging chopping and piecewise charge-acceleration facilitates a low temperature coefficient. A low-power comparator with a tail resistor and a low oscillation amplitude improves the energy efficiency. The die area is compact by introducing leakage-based temperature compensation that eliminates bulky resistors and complex calibration. Prototyped in a 28-nm CMOS process and measured at 28.5 kHz, our oscillator occupies 0.0046 mm$^{2}$and dissipates 27.6 nW at a 0.8-V supply. The energy efficiency is 0.97 nW/kHz, and the temperature coefficient is 33.3 ppm/$^{\circ}$C over$-$40 to 85$^{\circ}$C with 1-point calibration. The corresponding FoM of 164.9 dB compares favorably with the recent arts. The start-up time is rapid, and the period settling time is within one cycle of$\sim$5.7$\mu$s. The Allan deviation is$\le $40 ppm for measurement intervals of$>$0.5 s. Yueduo Liu, Rongxin Bao, Jiahui Lin, Jun Yin 0001, Qiang Li 0021, Pui-In Mak, Shiheng Yang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | Deep Learning Assisted Multiuser MIMO Load Modulated Systems for Enhanced Downlink mmWave CommunicationsabstractThis paper is focused on multiuser load modulation arrays (MU-LMAs) which are attractive due to their low system complexity and reduced cost for millimeter wave (mmWave) multi-input multi-output (MIMO) systems. The existing precoding algorithm for downlink MU-LMA relies on a sub-array structured (SAS) transmitter which may suffer from decreased degrees of freedom and complex system configuration. Furthermore, a conventional LMA codebook with codewords uniformly distributed on a hypersphere may not be channel-adaptive and may lead to increased signal detection complexity. In this paper, we conceive an MU-LMA system employing a full-array structured (FAS) transmitter and propose two algorithms accordingly. The proposed FAS-based system addresses the SAS structural problems and can support larger numbers of users. For LMA-imposed constant-power downlink precoding, we propose an FAS-based normalized block diagonalization (FAS-NBD) algorithm. However, the forced normalization may result in performance degradation. This degradation, together with the aforementioned codebook design problems, is difficult to solve analytically. This motivates us to propose a Deep Learning-enhanced (FAS-DL-NBD) algorithm for adaptive codebook design and codebook-independent decoding. It is shown that the proposed algorithms are robust to imperfect knowledge of channel state information and yield excellent error performance. Moreover, the FAS-DL-NBD algorithm enables signal detection with low complexity as the number of bits per codeword increases. Ercong Yu, Jinle Zhu, Qiang Li 0021, Zi Long Liu 0001, Hongyang Chen 0001, Shlomo Shamai, H. Vincent Poor |
IEEE Trans. Wirel. Commun. | 3 |
| 2023 | A 128-GS/s Timing-Robust Sampling Architecture Exploiting Analog FFTabstractTiming errors are the fundamental limitations of high-speed sampling systems. This paper proposes a high-speed sampling architecture with significantly improved timing-error robustness. During the analog calculation of the fast Fourier transform (FFT), the signal and noise behave differently with the superposition and correlation operations, resulting in improved sampling performance. In particular, the impact of timing skews, jitters and bandwidth mismatch is relaxed significantly. Exploiting a 64-channel radix-4 AFFT sampler, the proposed sampling system achieves 128GS/s with 7.93b effective number of bits (ENOB) at 30fs rms jitter. Compared with state-of-the-art time-interleaving systems, >10dB linearity enhancement has been observed with the same level of timing errors. Xingchen Chao, Qiang Li 0021 |
ISCAS | 2 |
| 2023 | A Low-Noise and Settling-Enhanced Switched-Capacitor Amplifier With Correlated Level Shifting and Bandwidth SwitchingabstractIn this article, a novel technique for reducing settling error and noise in switched-capacitor amplifiers is presented. The technique combines Correlated Level Shifting (CLS) with Bandwidth Switching to achieve improved performance. Closed-loop amplification with CLS can provide true rail-to-rail performance and reduce settling error from finite opamp gain and limited opamp output swing. The purpose of CLS is to remove the signal from the output of the opamp by sampling the estimated signal in the first phase of CLS. Bandwidth switching is used to reduce the final output noise of amplification by changing the noise bandwidth during settling. The common feature of CLS and Bandwidth switching is multi-phase operation. Therefore, the proposed closed-loop amplification operation can integrate CLS with bandwidth switching without extra clock phase overhead. From simulation results, the proposed technique improves SFDR from 51 dB to 88 dB and reduces the output noise from 583.3 uV to 450.8 uV for 80 mV input voltage. Feng Tai, Ziqin Nie, Yinhao Wang, Xingchen Chao, Qiang Li 0021 |
ISCAS | 5 |
| 2023 | A 0.0043-mm2 0.085-μW/MHz Relaxation Oscillator Using Charge-Prestored Asymmetric Swings R-RC NetworkabstractIn this brief, a charge-prestored 21.2-MHz relaxation oscillator is proposed for ultralow-power applications. It occupies only 0.0043 mm2in 0.18-$\mu \text{m}$CMOS by resistor reusing and is reference-free. The simulated temperature coefficient (TC) of the output frequency is 15.2 ppm/° from −30 °C to 125 °C. By generating an asymmetric capacitor charging swing, our charge-prestored technique reduces significantly the power consumed by the swing-boostingRCnetwork during the charging phase. Also, the R-RCstructure further improves the energy efficiency. The total power consumption of the oscillator core is$1.806 \mu \text{W}$at 0.8 V, corresponding to an energy efficiency of$0.085 \mu \text{W}$/MHz that compares favorably with the state of the art. Shiheng Yang, Yueduo Liu, Rongxin Bao, Jiahui Lin, Zehao Zhang, Yong Chen 0005, Jun Yin 0001, Pui-In Mak, Qiang Li 0021 |
IEEE Trans. Very Large Scale Integr. Syst. | 11 |
| 2022 | Bitwise ELD Compensation in Δ∑ ModulatorsabstractExcess loop delay (ELD) in high speed continuoustime (CT) Delta-Sigma-Modulators (DSMs) imposes design challenges and limits the use of high resolution, e.g. successiveapproximation-register (SAR) based internal quantizers, as usual compensation techniques like the use of a direct path around the quantizer come with increased swings and reduced maximum stable amplitude (MSA). In this paper, two bitwise ELD compensation approaches applicable to cascade-of-integrators with distributed feedback (CIFB) loop filters are proposed, which alleviate this problem for SAR and other multi-step quantizers by sequentially feeding bits into the feedback loop when they are available (MSB first, LSB last). Loop-filter equivalence for such bitwise ELD compensation is analytically derived. System-level simulations using Matlab & Simulink for exemplary 4-bit 2nd, 3rd and 4th-order modulators show 40% reduction of the last integrator output swing compared to the conventional direct path compensation. This potentially allows to avoid the swing, quantizer scaling and noise trade-offs due to ELD in state of the art designs. Michael Pietzko, Jonathan Ungethüm, John G. Kauffman, Qiang Li 0021, Maurits Ortmanns |
ISCAS | 4 |
| 2022 | A 100dB-TCMRR 8-Channel Bio-Potential Front-End with Multi-Channel Common-Mode ReplicationabstractFor a multichannel bio-potential signal acquisition front-end, it's total CMRR (TCMRR) is restricted by these three factors: the inherent CMRR (ICMRR) of the amplifier used in the acquisition front-end, the impedance mismatch of electrodes, and the systematic mismatch due to the shared reference electrode. This paper presents a 8-channel bio-potential front-end with multi-channel common-mode replication(MC-CMR). The frontend has been implemented in a standard CMOS 0.18-$\mu$m process, and the measurement results show that the amplifier has an adjustable gain of 46 to 64 dB, the input reference noise within 1Hz to 1k Hz is 1.62$\mu \mathrm{V}_{\mathrm{r}\mathrm{m}\mathrm{s}}$, TCMRR of the front-end is 100dB (at 50Hz). A single channel consumes 1.86$\mu$A at 1. 8V supply voltage, and the NEF is 2.7. Borui Tan, Sanfeng Zhang 0001, Qiang Li 0021 |
ISCAS | 5 |
| 2022 | Maximizing the Inter-Stage Gain in CT 0-X MASH Delta-Sigma-ModulatorsabstractThe performance of a 0-X MASH DSM is primarily defined by the DSM noise-shaping and the inter-stage gain (ISG). Due to system specification like modulator order and OSR, the DSM noise shaping is limited. All further performance in the 0-X MASH modulator is achieved by setting a large ISG, which amplifies the residue signal before it is processed by the fine stage DSM. This paper presents an analysis of the factors influencing the ISG and proposes measures, which can be used to maximize it. Most importantly, it is shown, that the CIFB DSM architecture is preferred, since the wideband residue overloads the CIFF architecture with its peaking OOB STF much earlier, thus CIFB results in more ISG and less internal swings. The presented analysis can be used as a guideline for achieving optimum performance in a 0-X MASH DSM and yielding the maximum SQNR. Jonathan Ungethüm, Michael Pietzko, John G. Kauffman, Qiang Li 0021, Maurits Ortmanns |
ISCAS | 4 |
| 2022 | A 13-bit 1-MS/s SAR ADC With Rotation-Based Mismatch Error CancellationabstractThis paper presents a mismatch error cancellation technique for high-resolution successive approximation register (SAR) analog-to-digital converter (ADC). The proposed technique that combines residue voltage oversampling and capacitors rotation significantly diminishes the impact of capacitor mismatch without calibration. A VCO-based comparator is adopted to achieve good noise performance with high energy efficiency. A 13-bit 1-MS/s SAR ADC is designed in a 180-nm CMOS technology to prove this technique. The post-layout simulated SAR ADC consumes $154.45 \mu \mathrm{W}$, achieves SNDR of 75.25 dB and SFDR of 90.34 dB at Nyquist input, resulting in a Schreier Figure of merit (FoM) of 170.35 dB. Maurits Ortmanns, Qiang Li 0021 |
ISCAS | 5 |
| 2022 | Accurate Performance Evaluation of Jitter-Power FOM for Multiplying Delay-Locked LoopabstractAn accurate performance evaluation of jitter-power figure-of-merit (FOM) for multiplying delay-locked loop (MDLL) is presented. For a typical MDLL employing a single-ended multiplying-delay ring voltage-controlled oscillator (MDVCO), it can be tuned via the capacitive loads to uphold a constant normalized phase noise (PN) across different frequencies for better jitter-power performance. Linear approximation and z-domain PN model are utilized to simplify the analysis with excellent agreement between the time-domain simulation and z-domain approximation. The influences of the asymmetric waveform, flicker noise corner, frequency error and reference noise are also discussed and then ignored based on the reasonable approximation. Under a given process, reference clock frequency and supply voltage, the ideal FOM can be derived theoretically. Based on these insights, the predicted FOM can be the benchmark metric in the design of MDLL for the possible best jitter-power performance. Yueduo Liu, Rongxin Bao, Shiheng Yang, Jun Yin 0001, Pui-In Mak, Qiang Li 0021 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2021 | A Ku-Band SiGe Phased-Array Transceiver with 6-Bit Phase and Attenuation ControlabstractA Ku-band phased-array transceiver using 0.13- μm SiGe BiCMOS technology is proposed in this paper for satellite communication applications. With a 6-bit inductive compensated attenuator and a 6-bit phase shifter integrated in the common-leg of the receive (RX) and transmit (TX) paths, the transceiver realizes 6-bit phase and attenuation control of the RF signal. The transceiver realizes larger than 20 dB gain, with RMS gain/phase error less than 1 dB/10° within 16-18 GHz. The whole transceiver occupies 2.7 × 2.5 mm2chip area. Bojun Hu, Sanfeng Zhang 0001, Qiang Li 0021 |
ISCAS | 4 |
| 2021 | A Sampling Speed Enhancement Technique for Near-Threshold SAR ADCsabstractThis paper presents a sampling speed-enhancement technique for SAR ADCs under near-threshold supply voltages. The proposed level-shifted boosting circuit generates sharp falling edges for the sampling clock, which is found a key factor limiting the sample speed under ultra-low voltages. A 0.35V 8b 12MS/s SAR ADC is designed in a 65nm CMOS technology to prove this technique. The post-layout simulated SAR ADC consumes only 6.71/iW and achieves SNDR of 48.8dB at Nyquist input, resulting in a figure-of-merit (FoM) of 2.47 fj/convertion-step. Simulation results show the proposed speed-enhancement technique improve the sampling rate of SAR ADC significantly under near-threshold supply voltages. Bojun Hu, Sanfeng Zhang 0001, Xiangxin Pan, Zhaoming Ding, Qiang Li 0021 |
ISCAS | 6 |
| 2021 | A 40nm 1Mb 35.6 TOPS/W MLC NOR-Flash Based Computation-in-Memory Structure for Machine LearningabstractComputation-in-memory (CIM) is a feasible method to overcome "Von-Neumann bottleneck" with high throughput and energy efficiency. In this paper, we proposed a 1Mb Multi-Level (MLC) NOR Flash based CIM (MLFlash- CIM) structure with 40nm technology node. A multi-bit readout circuit was proposed to realize adaptive quantization, which comprises a current interface circuit, a multi-level analog shift amplifier (AS-Amp) and an 8-bit SAR-ADC. When applied to a modified VGG-16 Network with 16 layers, the proposed MLFlash-CIM can achieve 92.73% inference accuracy under CIFAR-10 dataset. This CIM structure also achieved a peak throughput of 3.277 TOPS and an energy efficiency of 35.6 TOPS/W with 4-bit multiplication and accumulation (MAC) operations. Sitao Zeng, Zhiguo Zhu, Zhaolong Qin, Chen Wang 0130, Jingjing Li 0001, Sanfeng Zhang 0001, Yajuan He, Chunmeng Dou, Xin Si, Meng-Fan Chang, Qiang Li 0021 |
ISCAS | 12 |
| 2021 | ECDR$^{2}$2: Error Corrector and Detector Relocation Router for Network-on-ChipabstractNetwork-on-chip (NoC) is commonly used in modern many-core systems due to their high bandwidth and flexibility. As the manufacturing process keeps scaling, the reliability challenge in NoCs increases as well. The error correction code (ECC) is widely adopted in error correction NoCs to improve the data correctness. At the same time, extra stages are introduced in the router pipeline to improve the error correction capability. As a result, conventional error correction routers suffer from high network latency. Motivated by this limitation, i.e., we remove the extra pipeline stages delicately introduced for error correction. We propose an error correction router, called error corrector and detector relocation router (ECDR2), whose architecture optimizes the pipeline flow of the router. As a result, it can achieve both low latency and high error correction. Experimental results show that, compared with the baseline design, ECDR2obtains 13.67 and 39.4 percent less average latency under the uniform traffic pattern and Dedup benchmark, respectively, in an 8 × 8 mesh NoC. The circuit area of ECR is also 7.9 percent less than that of the baseline design under 45-nm technology. Letian Huang, Chikun Yuan, Junshi Wang, Masoumeh Ebrahimi, Qiang Li 0021 |
IEEE Trans. Computers | 6 |
| 2019 | A 12-bit 30MS/s SAR ADC with VCO-Based Comparator and Split-and-Recombination Redundancy for Bypass LogicabstractThis paper presents a 12-bit asynchronous successive approximation register (SAR) analog-to-digital converter (ADC) implemented in 40nm CMOS technology. A VCO-based comparator is employed to adjust the noise level adaptively and its oscillation number is harnessed to bypass unnecessary cycles for saving energy. A 1-bit split-and-recombination redundancy and a digital error correction for bypass logic are proposed to address the settling issues and refine the bypass window size. The sampling speed of the ADC reaches up to 30MS/s, which is the highest among SAR ADCs with single time-domain comparator. The ADC achieves an SFDR of 85.35 dB and 11.12-bit ENOB with Nyquist input consuming 0.38mW at a 1.1V supply, resulting in a figure of merit (FoM) of 5.69 fJ/conversion-step. Zhaoming Ding, Qiang Li 0021 |
ISCAS | 4 |
| 2019 | A Loss-Compensated 5-Bit Ka-Band Digital Phase Shifter with Low RMS Phase/Gain Error Over Wide Temperature RangesabstractThis paper presents the design and demonstration of a 5-bit Ka-band phase shifter with amplifiers to compensate for the passive loss using a 0.13-μm SiGe BiCMOS technology. Adopting the switched low-pass (LP)/high-pass (HP) networks in the phase shift stages, the phase shifter exhibits measured RMS phase and gain errors less than 4° and 0.6 dB, respectively, within the frequency range of 30-40 GHz under different temperature conditions (from -40 °C to 140 °C). To the authors' knowledge, our proposed phase shifter achieves the best RMS phase/gain error performance at Ka-band among the published works. Ruoman Yang, Qiang Li 0021 |
ISCAS | 6 |
| 2019 | Online Path-Based Test Method for Network-on-ChipabstractA considerable amount of routers and links remains idle after each mapping application onto the Network-on-Chip based many-core systems. Online path-based test method is a kind of self-test for these idle components. In this paper, a path-based fabric for NoC is firstly proposed. A path serves as the basic component, covering one link and its associated control logic in the routers. One possibility is to apply fault detection on the idle paths, while the other paths continue to operate normally. Moreover, this paper details the hardware implementation, targeting the stuck-at and bridging faults. It suggests a good trade-off between fault coverage, hardware overhead and test time. Experimental results show that the approach achieves 93% of the stuck-at faults in control unit and cover 100% of the stuck-at and bridging faults on the global link within 256 clock cycles. Junkai Zhan, Letian Huang, Junshi Wang, Masoumeh Ebrahimi, Qiang Li 0021 |
ISCAS | 5 |
| 2019 | Efficient Design-for-Test Approach for Networks-on-ChipabstractTo achieve high reliability in on-chip networks, it is necessary to test the network continuously with Built-in Self-Tests (BIST) so that the faults can be detected quickly and the number of affected packets can be minimized. However, BIST causes significant performance loss due to data dependencies. We propose EsyTest, a comprehensive test strategy with minimized influence on system performance. EsyTest tests the data path and the control path separately. The data path test starts periodically, but the actual test performs in the free time slots to avoid deactivating the router for testing. A reconfigurable router architecture and an adaptive fault-tolerant routing algorithm are proposed to guarantee the access to the processing core when the associated router is under test. During the whole test procedure of the network, all processing cores are accessible, and thus the system performance is maintained during the test. At the same time, EsyTest provides a full test coverage for the NoC and a better hardware compatibility comparing with the existing test strategies. Under the PARSEC benchmark and different test frequencies, the execution time increases less than 5 percent at the cost of 9.9 percent more area and 4.6 percent more power in comparison with the execution where no test procedure is applied. Junshi Wang, Masoumeh Ebrahimi, Letian Huang, Qiang Li 0021, Guangjun Li, Axel Jantsch |
IEEE Trans. Computers | 5 |
| 2018 | A lifetime-aware mapping algorithm to extend MTTF of Networks-on-ChipabstractFast aging of components has become one of the major concerns in Systems-on-Chip with further scaling of the submicron technology. This problem accelerates when combined with improper working conditions such as unbalanced components' utilization. Considering the mapping algorithms in the Networks-on-Chip domain, some routers/links might be frequently selected for mapping while others are underutilized. Consequently, the highly utilized components may age faster than others which results in disconnecting the related cores from the network. To address this issue, we propose a mapping algorithm, called lifetime-aware neighborhood allocation (LaNA), that takes the aging of components into account when mapping applications. The proposed method is able to balance the wear-out of NoC components, and thus extending the service time of NoC. We model the lifetime as a resource consumed over time and accordingly define the lifetime budget metric. LaNA selects a suitable node for mapping which has the maximum lifetime budget. Experimental results show that the lifetime-aware mapping algorithm could improve the minimal MTTF of NoC around 72.2%, 58.3%, 46.6% and 48.2% as compared to NN, CoNA, WeNA and CASqA, respectively. Letian Huang, Masoumeh Ebrahimi, Junshi Wang, Shuyan Jiang, Qiang Li 0021 |
ASP-DAC | 7 |
| 2018 | Optimizing dynamic mapping techniques for on-line NoC testabstractWith the aggressive scaling of submicron technology, intermittent faults are becoming one of the limiting factors in achieving a high reliability in Network-on-Chip (NoC). Increasing test frequency is necessary to detect intermittent faults, which in turn interrupts the execution of applications. On the other hand, the main goal of traditional mapping algorithms is to allocate applications to the NoC platform, ignoring about the test requirement. In this paper, we propose a novel testing-aware mapping algorithm (TAMA) for NoC, targeting intermittent faults on the paths between crossbars. In this approach, the idle links are identified and the components between two crossbars are tested when the application is mapped to the platform. The components can be tested if there is enough time from when the application leaves the platform and a new application enters it. The mapping algorithm is tuned to give a higher priority to the tested paths in the next application mapping. This leaves enough time to test the links and the belonging components that have not been tested in the expected time. Experiment results show that the proposed testing-aware mapping algorithm leads to a significant improvement over FF, NN, CoNA, and WeNA. Shuyan Jiang, Junshi Wang, Masoumeh Ebrahimi, Letian Huang, Qiang Li 0021 |
ASP-DAC | 7 |
| 2018 | Delta-Measurement Low-Power SAR ADC Architecture with Adaptive Threshold-First SwitchingabstractA delta-measurement successive approximation register (SAR) analog-to-digital converter (ADC) architecture with adaptive threshold-first switching scheme is designed to achieve low power for biomedical signals. Delta-measurement architecture achieves the delta voltage from the current sample to the last one with few the least weighted capacitors switched when the input signal changes slowly. Adaptive threshold-first switching scheme is utilized to decrease the number of successive approximation (SA) cycles. Analysis demonstrates the proposed delta-measurement architecture will not introduce nonlinearity. Simulation results show that this work can decrease the number of SA cycles, and can save more than 90% switching energy than the most effective LSB-first switching scheme when dealing with slowly-changing signal. Zhaoming Ding, Qiang Li 0021 |
ISCAS | 3 |
| 2018 | An Energy-Efficient Approximate DCT for Wireless Capsule Endoscopy ApplicationabstractWireless capsule endoscopy is widely used as an efficient way to receive images of gastrointestinal tract for medical diagnostics. Due to its energy compaction property for correlated image pixels, DCT) is highly preferred in image compression for this real time application. To accommodate the tight power budget and small size condition in the wireless capsule endoscopy application, an energy efficient DCT based on CMOS implementation is proposed in this paper. Three approximating methods are combined and applied in a multiplier-less DCT architecture. The simulation results indicate that the proposed DCT reduces 53.85% energy and 25.32% area with a PSNR penalty of 1.9dB when compared with the integer DCT. Compared with CORDIC-Loeffler DCT, it achieves 19.38% reduction on energy dissipation and 14.4dB more PSNR at the cost of only 1.39% area overhead. Ziji Zhang 0001, Yiduan Qian, Qiang Li 0021, Yajuan He |
ISCAS | 4 |
| 2018 | Micro-Architecture Design for Low Overhead Fault Tolerant Network-on-ChipabstractAggressive technology scaling results in reliability decrease of Network-on-Chips (NoCs). Error Correction Codes (ECC) is commonly used to correct error data. It is necessary to balance the reliability of transmissions and the overhead introduced by encoders and decoders. This work utilizes a mechanism reusing decoders in Network Interfaces (NIs), which is named Send-Back ECC. This paper proposes the detailed hardware implementation of Send-Back NoC after hardware overhead optimizing. The design details of the routers and NIs are described. Simulation results prove that the latency of Send-Back ECC is lower than H2H ECC when bit error rate is lower than 0.0002 and always lower than E2E ECC. The hardware overhead of Send-Back ECC is 10.6% lower than H2H ECC, while the energy consumption is also less than both E2E ECC and H2H ECC. Chikun Yuan, Letian Huang, Junshi Wang, Qiang Li 0021 |
ISCAS | 4 |
| 2018 | Optimal Slope Ranking: An Approximate Computing Approach for Circuit PruningabstractIn recent years, approximate computing has been widely used in low power digital circuits design. As it incurs error in computation, there are always tradeoffs between hardware performances and computational accuracy. Thus, the circuit implementation can vary from applications to applications and even in the same application, which makes difficult to evaluate the design efficiency in terms of delay, power and computational accuracy. In this paper, approximate efficiency (AE) is proposed as a new design metric dedicated in energy-efficient approximate computation. An automatic pruning approach is presented based on quickly estimated AE, which leads to an efficient approximate circuit design and a significant reduction on energy dissipation. A 32-bit adder is demonstrated using proposed method and compared with other approximation methods, including direct truncation and significance-activity pruning method. For legitimate comparisons, same computational accuracy is obtained for all imprecise adders. The simulation results show that assessed on mean absolute error, the proposed adder outperforms the others. With given error threshold, it can reach 31.15% energy-delay-product reduction compared with the most competitive contender. Ziji Zhang 0001, Yajuan He, Xilin Yi, Qiang Li 0021, Bo Zhang 0027 |
ISCAS | 5 |
| 2017 | A low latency fault tolerant transmission mechanism for Network-on-ChipabstractReliability of Network-on-Chip has become a critical problem because of the aggressive technology scaling. A variety of transmission mechanism to tolerant the bit errors has been proposed to achieve the best trade-off between performance and overhead. In this work, a transmission mechanism for NoC based on a novel combination of error detection, error correction, and retransmission is proposed. The light-weight error detectors are integrated into the input ports of routers to check the correctness of head flits and the decoders in Network Interfaces (NIs) are used to correct the errors in any flits of the whole packet. With a very small hardware overhead, the proposed method can guarantee high reachability of packets and greatly decrease End-to-End retransmission. Compared with Hop-to-Hop and End-to-End mechanism, the latency could be highly reduced. Letian Huang, Xinxin Lin, Junshi Wang, Qiang Li 0021 |
ISCAS | 4 |
| 2017 | Non-blocking BIST for continuous reliability monitoring of Networks-on-ChipabstractTo achieve high reliability in on-chip networks, frequent runs of Built-in Self-Test allow the detection of and recovery from faults before they affect packets and the system functionality. However, to test routers, wrappers isolate cores from the network which leads to execution blocking and performance loss. In this paper, we propose a design-for-test reconfigurable router with two alternative bypassing channels. The router architecture allows maintaining the connection between cores and the network during the testing procedure by utilizing the bypassing channels. With the help of an adaptive routing algorithm and a testing strategy, networks can be fully tested at a high testing frequency with <;15% increase of execution time. Junshi Wang, Letian Huang, Masoumeh Ebrahimi, Qiang Li 0021, Guangjun Li, Axel Jantsch |
ISCAS | 4 |
| 2015 | A fast and energy efficient binary-to-pseudo CSD converterabstractThe canonical signed digit (CSD) coding is widely used in digital arithmetic operations due to its property that there is no adjacent nonzero digits in the encoded numbers. However, the benefits of the CSD coding may be faded because of the recursive conversion process from the binary representations. This paper presents a novel pseudo CSD coding method, which takes the merits of CSD, while simplifies the conventional conversion process. The simulation results indicate that the proposed converter can achieve at least 31.8% speed improvement and 42.9% energy reduction for a 16-bit binary operand at 1.2V in a 0.13-μm CMOS technology. It could run even faster than the competitors when the operand length increases. Yajuan He, Ziji Zhang 0001, Bin Ma 0010, Shaowei Zhen, Ping Luo 0005, Qiang Li 0021 |
ISCAS | 7 |
| 2015 | 300mV 50kHz 75.9dB SNDR CT ΔΣ Modulator with Inverter-based Feedforward OTAsabstractA low voltage low power continuous-time (CT) ΔΣ modulator is presented. To obtain a high linearity and power efficient modulator, a novel inverter-based feedforward OTA is proposed. A gain-enhancement technique is utilized to relax the load effect. With a novel complementary differential-difference common mode feedback circuit, the CMRR of OTA achieves 81.8 dB. The circuits were implemented in a 0.13 μm CMOS technology. The post layout simulation demonstrates that the modulator consumes 29.58 μW from a 300 mV supply and achieves a peak SNDR of 75.9 dB, a dynamic range of 79.5 dB, with a 50 kHz bandwidth under 64 oversampling ratio. Lishan Lv, Qiang Li 0021 |
ISCAS | 2 |
| 2014 | A 10-bit 150MS/s SAR ADC with parallel segmented DAC in 65nm CMOSabstractThis paper presents a high speed parallel segmented capacitive DAC that is implemented in a 10-bit 150MSample/s successive approximation register (SAR) ADC. Compared to converters that use the conventional structure, the speed of converting one bit digital code can be 4 times faster while the power remains relatively low. In the switching procedure, a small capacitor array is used to determine the high weight codes. A parallel capacitor array structure is used to reduce the mismatch and kickback noise effect caused by the small capacitor array. A prototype 10-bit 150MS/s SAR ADC with the parallel segmented capacitor array is implemented in 65nm CMOS technology. The ADC achieves an SFDR of 83.77 dB and 9.78-bit ENOB with only 1.476mW power consumption at a 1.2-V supply, resulting in a figure of merit (FOM) of 11.19 fJ/conversion-step. Qiang Li 0021 |
ISCAS | 2 |
| 2013 | A 0.5V rate-resolution scalable SAR ADC with 63.7dB SFDRabstractA 0.5V 6-to-10b rate-resolution scalable SAR ADC with microwatt power consumption is presented. Employing the successive approximation register (SAR) architecture, the proposed ADC exhibits the sampling rates of 125kS/s, 150kS/s and 250kS/s at scalable resolutions of 10b, 8b and 6b, respectively. A low-leakage voltage boosting technique is proposed, which reduces the leakage of MOS switches by 99% as compared to conventional techniques. This has ensured the ADC operating at sampling rates from 175kS/s to 5kS/s with only 0.2b ENOB degradation. Meanwhile, the effect of bridge capacitor on the linearity of merged capacitor switching (MCS) DAC is discussed. Demonstrated in a 0.13μm CMOS process, measured results show the ADC achieves ENOB of 8.51b, 7.42b, and 5.97b at 10b, 8b, and 6b modes, respectively. At 10b 125kS/s, the entire ADC consumes only 3.4μW from a 0.5V supply. Kun Ao, Qiang Li 0021 |
ISCAS | 4 |
| 2013 | Blind-LMS based digital background calibration for a 14-Bit 200-MS/s pipelined ADCabstractA 14-bit and 200-MS/s SHA-less pipelined ADC is implemented by 0.13 μm CMOS process with blind least mean square (BLMS) calibration technique which corrects errors of this pipelined ADC with fast, low gain and inaccurate opamps. Using skip and fill approach, we employ an interpolation filter and a front-end DAC to make the pipelined ADC self-calibratable in the background. Incorporated a 18 stages and 1.5 bit/stage structure, the simulation shows the ADC achieves an SINAD of 86 dB, an SFDR of 107 dB with a 90.55 MHz input signal. Yajuan He, Qiang Li 0021 |
VLSI-SoC | 3 |
| 2011 | Waveform distortion performance evaluation using practical antennas in deterministic multipath impulse radio channelsabstractAs a potential future technology, impulse radio is a promising new reliable data transmission method for time-sensitive applications. However, owing to the complex nature of ultra-wideband (UWB) channels used in practice, there are many design problems that remain to be resolved. This study looks into some UWB channels classified as waveform distortion for which the authors present an implementation-oriented analysis to investigate the performance of impulse radio systems using antennas for a special class of deterministic multipath channels. This study analyses the waveform distortion along signal path to identify the effect of non-idealities in the system and evaluates overall performance of system in terms of bit error ratio observing the waveforms received at the correlator. Through simulation results practical aspects of the impulse-radio transceivers are discussed. From some of the results, the authors provide an antenna selection strategy for some propagation environments. Y. P. Zhang, Qiang Li 0021, Guangjun Li, Habib F. Rashvand |
IET Commun. | 4 |