Jiangwei Zhang

dblp:50/8594 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 WOLF: Weight-Level OutLier and Fault Integration for Reliable LLM Deployment
abstract
The rapid advancement of Transformer-based large language models (LLMs) is presenting significant challenges for their deployment, primarily due to their enormous parameter sizes and intermediate results, which create a bottleneck in memory capacity for effective inference. Compared to traditional DRAM, Non-Volatile Memory (NVM) technologies such as Resistive Random-Access Memory (RRAM) and Phase-Change Memory (PCM) offer higher integration density, making them promising alternatives. However, before NVM can be widely adopted, its reliability issues, particularly manufacturing defects and endurance faults, must be addressed. In response to the limited memory capacity and reliability challenges of deploying LLMs in NVM, we introduce a novel low-overhead weight-level map, namedWolf.Wolfnot only integrates the addresses of faulty weights to support efficient fault tolerance but also includes the addresses of outlier weights in LLMs. This allows for tensor-wise segmented quantization of both outliers and regular weights, enabling lower-bitwidth quantization. TheWolfframework uses a Bloom Filter-based map to efficiently manage outliers and faults. By employing shared hashes for outliers and faults and specific hashes for faults,Wolfsignificantly reduces the area overhead. Building onWolf, we propose a novel fault tolerance method that resolves the observed issue of clustering critical incorrect outliers and fully leverages the inherent resilience of LLMs to improve fault tolerance capabilities. As a result,Wolfachieves segment-wise INT4 quantization with enhanced accuracy. Moreover,Wolfcan adeptly handle Bit Error Rates as high as$1 {\boldsymbol{\times}} 10^{-2}$without compromising accuracy, in stark contrast to the state-of-the-art approach where accuracy declines by more than 20%.
Wanyi Fu, Jiangwei Zhang, Rui Hou 0001, Jian Yang 0011, Yu Wang 0002
IEEE Trans. Computers3
2023 Memory-Efficient and Real-Time SPAD-based dToF Depth Sensor with Spatial and Statistical Correlation
abstract
Single Photon Avalanche Diode (SPAD)-based direct time-of-flight (dToF) depth sensors are widely used in Internet of Things (IoT) devices due to their high accuracy. Existing SPAD-based dToF sensors measure depth by continually accumulating the depth-measured value in a histogram. However, histogram-based methods typically have low convergence speed (~10 frames per second (FPS)) and large memory overhead (MB-level), hindering their use in real-time embedded IoT devices. To overcome these two challenges, we propose SSC, a histogram-free Spatial and Statistical Correlation based depth measurement method. On the one hand, SSC applies the spatial correlation of the adjacent pixels to accelerate the convergence speed. On the other hand, SSC explores the statistical correlation of depth measurements to reduce the memory overhead. In order to implement SSC with small hardware area and low power, we design mert-dToF, a memory-efficient and real-time dToF sensor for efficient execution. mert-dToF abstracts mainly operations in SSC into four basic operators and designs corresponding hardware with a fine-grained pipeline to maximize resource reuse and computational parallelism. Extensive experiments show that compared with state-of-the-art (SOTA) histogram-based dToF sensors, mert-dToF achieves ~8% accuracy improvement and 7.80× speedup (from 6.24 FPS to 48.70 FPS). The memory overhead is reduced by up to 60.91% (from 48 KB to 18.75 KB).
Zhenhua Zhu 0002, Qingpeng Zhu, Jiangwei Zhang, Wenxiu Sun, Guohao Dai 0001, Fei Qiao, Huazhong Yang, Yu Wang 0002
DAC5
2023 Realizing Extreme Endurance Through Fault-aware Wear Leveling and Improved Tolerance
abstract
Phase-change memory (PCM) and resistive memory (RRAM) are promising alternatives to traditional memory technologies. However, both PCM and RRAM suffer from limited write endurance. Wear-leveling (WL) techniques are essential to extend the lifetime of these memories before experiencing endurance faults. Beyond the additional usage afforded by WL, row-sparing and focused error correction can extend the lifetime further after wear faults appear. Unfortunately, the need for extended WL techniques continues to become more pressing as scaling exacerbates process variation. Similarly, scaling causes challenges such as more severe noise and crosstalk to traditional DRAM.In this paper, we propose novel fault-aware WL schemes to allocate write frequencies according to the strength of the rows and handle the imbalance of writes in columns. We use runtime detection schemes to identify weak rows and protect them prior to wear out. In particular, row-level WL, aka RETROFIT, leverages the spare rows provided for redundancy to be used strategically to guard against early cell wear out. RETROFIT is compatible with error correction schemes that guarantee to mitigate hard faults and error-correcting codes (ECC). Rather than discard retired rows, when any spare row completely replaces a retired row, we retarget the retired row to assist with column sparing. It becomes a group of Page Protecting Pointers (PPPs), which utilizes otherwise discarded error correction potential to further enhance the leveling ability of RETROFIT. To relieve column-level imbalance, we apply idle error correction bits before they are used to reduce average bit flips. The evaluation demonstrates that RETROFIT and enhanced RETROFIT with the PPPs improve lifetime by as much as 0.64× and 5.4× in the average case, respectively, over state-of-the-art row-level method while also reducing area overhead. In the worst-case scenario, these improvements further increase to 2.6× and 16.0×. Combined with the proposed column-level WL, enhanced RETROFIT realizes an overall 1.5× memory lifetime improvement over the perfectly uniform wear-leveling with equal storage overhead.
Jiangwei Zhang, Zhenhua Zhu 0002, Donald Kline, Alex K. Jones, Huazhong Yang, Yu Wang 0002
HPCA1
2023 A Three-Step Multi-Resolution Time-to-Digital Converter
abstract
This work proposes a three-step multi-resolution time-to-digital converter (TDC) architecture based on the vernier delay line (VDL). The proposed architecture uses a delay-locked loop (DLL) to control TDC with a smooth coarse-to-fine strategy. In addition, the fine TDC uses a combination of multiple resolutions to reduce the number of delay cells and flip-flops. This architecture helps to reduce the area and power consumption and maintains high resolution. We proposed architecture performs better trade-offs between power consumption, linearity, accuracy, and measurement range. The simulation results show that the 7-bit TDC based on VDL designed in 180 nm CMOS achieves 5 ps of time resolution, 0.76/-0.8 LSB DNL and 1.02/-1.39 LSB INL at 100 MHz clock frequency while consuming 3.1 mW, which corresponds to the figure of merit (FoM) of 0.242 pJ/Conv.
Jiang Yan, Yu Wang 0002, Fei Qiao, Jiangwei Zhang, Qi Wei 0001, Qingpeng Zhu, Wenxiu Sun, Ge Shi 0001
ISCAS6
2023 MMBench: The Match Making Benchmark
abstract
Video gaming has gained huge popularity over the last few decades. As reported, there are about 2.9 billion gamers globally. Among all genres, competitive games are one of the most popular ones.
Yanxing Qi, Jiangwei Zhang, Connie Khor Li Kou, Qiaolin Chen
WSDM3
2021 STMG: Spatial-Temporal Mobility Graph for Location Prediction
Xuan Pan, Xiangrui Cai, Jiangwei Zhang, Yanlong Wen, Ying Zhang 0015, Xiaojie Yuan
DASFAA (1)3
2021 Tuning Memory Fault Tolerance on the Edge
abstract
Error correction and fault tolerance have become pivotal considerations as conventional memories scale and emerging memories come to market. The common thread in these reliability challenges is that deep scaling reveals outliers in the memory system, which are responsible for the vast majority of faults. These cells, which may be attributed to process variation or undetected fabrication defects, tend to be more vulnerable to various forms of crosstalk, read- and write-disturbance, and even radiation-induced faults. By tracking faults in memory cells, identifying the worst offenders, and mitigating their effects accordingly, we can design dramatically improved fault tolerance techniques that are tuned to the fault characteristics of the memory at hand. A critical piece is the development of scalable and fault tolerance registries to track and retain critical information about these faults. The fault registries must be able to function in the faulty memory they protect, operate efficiently at the cell/bit-level, and handle extreme fault rates. Using the knowledge of faulty locations, our fault tolerance techniques applied to conventional main memories like DRAM and endurance-limited memories like flash and phase-change memory improve reliability, endurance, and lifetime by orders of magnitude while maintaining performance and energy efficiency.
Alex K. Jones, Stephen Longofono, Sébastien Ollivier, Donald Kline Jr., Jiangwei Zhang, Rami G. Melhem
ACM Great Lakes Symposium on VLSI5
2020 FLOWER and FaME: A Low Overhead Bit-Level Fault-map and Fault-Tolerance Approach for Deeply Scaled Memories
abstract
To maintain appropriate yields in deeply scaled technologies requires fault-tolerance of increasingly high fault rates. These fault rates far exceed traditional general approaches such as ECC, particularly when faults accrue over time. Effective fault tolerance at such high fault rates requires detailed bit-level knowledge of the location of faulty cells. We provide a solution to this problem in the form of a space efficient, bit-level fault map called FLOWER. FLOWER utilizes Bloom filters to provide detailed fault characterization for a relatively small overhead. We demonstrate how FLOWER can enable improved fault tolerance at high fault rates by enhancing existing fault tolerance proposals and yielding 10–100x improvements. Using in-memory processing, FLOWER can maintain a less than 2% performance overhead at 10E-4 fault rates with less than 2% loss of memory density to report bit-level faults with high accuracy. Using a tuned novel hashing technique called MinCI, FLOWER for memory achieves considerably lower false positives than with disk-level hashing techniques at a fraction of the performance overhead. With a new technique to protect against errors during in-memory operations, PETAL bits, FLOWER can remain resilient against random errors while efficiently targeting predictable errors. Furthermore, we propose a new fault tolerance scheme called FaME, which provides ultra-efficient bit-level sparing by using the FLOWER fault map to identify the location of faults. FLOWER+FaME can achieve 14x longer PCM memory lifetime with half the area overhead versus SECDED ECC.
Donald Kline Jr., Jiangwei Zhang, Rami G. Melhem, Alex K. Jones
HPCA2
2020 LOAD: LSH-Based ℓ 0-Sampling over Stream Data with Near-Duplicates
Dingzhu Lurong, Yanlong Wen, Jiangwei Zhang, Xiaojie Yuan
ECML/PKDD (1)3
2020 PG2S+: Stack Distance Construction Using Popularity, Gap and Machine Learning
abstract
Stack distance characterizes temporal locality of workloads and plays a vital role in cache analysis since the 1970s. However, exact stack distance calculation is too costly, and impractical for online use. Hence, much work was done to optimize the exact computation, or approximate it through sampling or modeling.
Jiangwei Zhang, Y. C. Tay
WWW1
2019 Yielding optimized dependability assurance through bit inversion
Jiangwei Zhang, Donald Kline Jr., Rami G. Melhem, Alex K. Jones
Integr.1
2018 A collaborative framework for tweaking properties in a synthetic dataset
abstract
Researchers and developers use benchmarks to compare their algorithms and products. For database systems, a benchmark must have a dataset D. To be application-specific, this dataset D should be empirical. However, a real D may be too small, or too large, for the benchmarking experiments. Therefore, D must first be scaled to the desired size. Previous related work typically extracts a set of properties Π = { π 1 , . . . , π n } from D, then use Π to generate the synthetic D~. Π may thus ensure D~ is similar to D. This approach of having some monolithic software enforce properties π 1 , . . . , π n becomes increasingly intractable as n increases. Our demonstration will present ASPECT, a framework that takes a different approach. With ASPECT, there is a tool So to first scale the dataset size. The resulting D~ can then be tweaked by tools T 1 , . . . , T n , where T k enforces π k in D~. At the demonstration, a visitor has a choice of (i) D , (ii) size scaler S 0 , (iii) the subset of properties to enforce, and (iv) the order of applying the tools for the chosen properties. The visitor can then see the enforcement error for each π k and the running time for each T k . A video of the demonstration is presented here: http://scaler.d2.comp.nus.edu.sg/
Jiangwei Zhang, Y. C. Tay
Proc. VLDB Endow.1
2018 Data Block Partitioning Methods to Mitigate Stuck-At Faults in Limited Endurance Memories
Jiangwei Zhang, Donald Kline Jr., Rami G. Melhem, Alex K. Jones
IEEE Trans. Very Large Scale Integr. Syst.1
2017 Dynamic partitioning to mitigate stuck-at faults in emerging memories
abstract
Emerging non-volatile memories have many advantages over conventional memory. Unfortunately, many are susceptible to write endurance challenges, resulting in stuck-at faults. Existing mitigation methods statically partition and invert data within a block containing such faults (partition-and-flip) to ensure data is written to match stuck-at cells such that they may remain in service. Unfortunately, these schemes have limited fault tolerance capabilities and require the assumption that their auxiliary bits are fault free. We propose a dynamic partitioning scheme that improves the number of tolerated stuck-at faults and simultaneously protects auxiliary bits. Dynamic partitioning can significantly improve the fault tolerance over existing static partitioning approaches with an equal number of auxiliary bits. Moreover, it can often still improve fault tolerance while reducing the number of auxiliary bits. Compared to flip-N-write and Aegis, a leading mitigation scheme, dynamic partitioning can achieve 7-72% and 5-53 x lower write error rates, respectively, for the same capacity overhead with a stuck-at-fault rate of 10-3.
Jiangwei Zhang, Donald Kline Jr., Rami G. Melhem, Alex K. Jones
ICCAD1
2017 Yoda: Judge Me by My Size, Do You?
abstract
Phase change memory is a promising alternative to conventional memories such as DRAM due to its density and non-volatility. Unfortunately, reliability is still a challenge as limited write endurance, exacerbated by process variation, leads to increasing numbers of stuck-at faults over the memory's lifetime. Error-correcting Pointers (ECP) is a popular proposal to mitigate stuck-at faults by recording the addresses and the values of faulty bits in order to extend the memory lifetime. In this paper, we propose Yoda, a method to extend ECP with one or a small number of additional encoding bits in order to dramatically improve the effectiveness and guaranteed fault correction capability of ECP. Our simulation results demonstrate that Yoda has a 3.0× improvement in fault coverage compared to a fault-aware ECP with a similar overhead, while also providing a 2.5-3.0× improvement over state-of-the-art schemes with comparable complexity.
Jiangwei Zhang, Donald Kline Jr., Rami G. Melhem, Alex K. Jones
ICCD1
2017 Improved BM3D denoising method
abstract
Block matching 3D denoising (BM3D) is an excellent single‐image denoising method. However, it still needs to be improved for solving practical problems. In this study, the authors attempt to improve the method of BM3D. First, one of the problems of BM3D is that some of its references cannot perform self‐adaption when the noise intensity of the images is changed. Therefore, they propose a method using total variation (TV) to calculate the image noise intensity and make the references perform self‐adaption. Second, finding similar blocks in the BM3D method is a time‐consuming procedure. To solve this problem, they analyse the relationship between the numbers of similar blocks and denoising effect, improve the process of searching for similar blocks, and reduce the running time. Third, through the experiment they find that the denoising effect of BM3D method in the domain of complex texture is unsatisfactory. Thus, they proposed a hybrid denoising method for the complex texture area, using the new TV model and BM3D method together to restore the image. Their experimental results show that the improved BM3D method performs better than the original BM3D method.
Yingjiang Li, Jiangwei Zhang, Maoning Wang
IET Image Process.2
2012 Smart Traffic Cloud: An Infrastructure for Traffic Applications
abstract
With rapid development of sensor technologies and wireless network infrastructure, research and development of traffic related applications, such as real time traffic map and on-demand travel route recommendation have attracted much more attentions than ever before. Both archived and real-time data involved in these applications could potentially be very big, depending on the number of deployed sensors. Emerging Cloud infrastructure can elastically handle such big data and conveniently providing nearly unlimited computing and storage resources to hosted applications, to carry out analysis not only for long-term planning and decision making, but also analytics for near real-time decision support. In this paper, we propose Smart Traffic Cloud, a software infrastructure to enable traffic data acquisition, and manage, analyze and present the results in a flexible, scalable and secure manner using a Cloud platform. The proposed infrastructure handles distributed and parallel data management and analysis using ontology database and the popular Map-Reduce framework. We have prototyped the infrastructure in a commercial Cloud platform and we developed a real-time traffic condition map using data collected from commuters' mobile phones.
Jiangwei Zhang, Hock-Beng Lim
ICPADS3