Minge Jing

dblp:402/2732 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0005-0446-5600ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SENTRY-VQA: A Region-Aware Framework for Surveillance Video Quality Assessment
Chenlong He, Leilei Huang, Minge Jing, Yibo Fan
ISCAS6
2026 A Dual-Generalization Low-Light Enhancement Framework for Capsule Endoscopy Image Restoration and Segmentation
abstract
In recent years, deep learning technology has automated the diagnosis of gastrointestinal (GI) tract disease, enabling doctor-machine collaborative diagnosis. However, the images captured by wireless capsule endoscopy (WCE) easily suffer from varying brightness levels of low-light degradation due to the complex structure of GI tract and the limitations of the light source, which impacts both human and machine diagnostic accuracy. Moreover, images may contain varying degrees of structural and semantic details even under a similar brightness level, which still can compromise segmentation accuracy. To address these issues, we propose a dual-generalization framework for low-light WCE images. Our framework includes an Image Guidance and Laplacian Fusion Module (IGLFM), a Brightness Level Generalization Module (BLGM) and a Wavelet Segmentation Generalization Module (WSGM). IGLFM and BLGM can restore low-light images across different brightness levels and WSGM can enhance segmentation accuracy by generalizing the varying degrees of details across images. With BLGM and WSGM, our framework enables two aspects of generalization: generalization to input images with different brightness levels and generalization to images with varying detail levels. Extensive experiments demonstrate that our method achieves significant performance under varying brightness levels and improvements in segmentation accuracy, surpassing the existing state-of-the-art (SOTA) method with gains of 4.70 dB / 0.022 (PSNR/SSIM) on Kvasir-Capsule dataset and 1.61 dB / 0.018 on RLE dataset. WSGM consistently improves segmentation accuracy across six popular networks, achieving up to + 4.7% mIoU and + 5.3% Dice improvements on RLE dataset. Our code will be available at https://github.com/superwsc/Dual-Gen-Frame.
Shuocheng Wang, Ruoxi Zhu, Chengkang Huang, Minge Jing, Yibo Fan
IEEE Trans. Medical Imaging5
2026 EDMOT: A 25-65 ns Latency Event-Driven Multiobject Tracker for Dynamic Vision Sensors
abstract
Dynamic vision sensors (DVSs), renowned for their low latency and sparse event-driven output, have garnered significant attention in machine vision applications, particularly in latency-sensitive applications like object tracking. However, current DVS-based object tracking systems have not fully utilized the advantages of DVS due to their frame-based processing or the high computing intensity of their algorithms. This brief presents EDMOT, a fully event-driven multiobject tracking system for DVS, achieving 25–65 ns latency at 200 MHz. EDMOT introduces three key innovations: 1) a novel event-driven update mechanism that processes only the latest event and the expired oldest event, minimizing computational overhead; 2) a dual-threshold tracking strategy that decouples object formation and motion phases, significantly improving tracking accuracy; and 3) a row–column feature memory with flag registers, enabling object separation within eight clock cycles. The proposed EDMOT is evaluated on public datasets, demonstrating superior tracking accuracy compared to prior methods. Finally, EDMOT was implemented at the HLMC 55 nm, supporting 20–100 Me/s throughput. To the best of our knowledge, this is a multiobject tracking system with minimal delay and the highest event processing throughput.
Feiqiang Li, Yaoyi Chen, Mingyu Wang 0001, Minge Jing, Wenhong Li, Xiaoyang Zeng
IEEE Trans. Very Large Scale Integr. Syst.5
2025 Unicorn: Unified Neural Image Compression with One Number Reconstruction
abstract
Prevalent lossy image compression schemes can be divided into: 1) explicit image compression (EIC), including traditional standards and neural end-to-end algorithms; 2) implicit image compression (IIC) based on implicit neural representations (INR). The former is encountering impasses of leveling off bitrate reduction at a cost of tremendous complexity while the latter suffers from excessive smoothing quality as well as lengthy decoder models. In this paper, we propose an innovative paradigm, which we dub Unicorn (Unified Neural Image Compression with One Nnumber Reconstruction). By conceptualizing the images as index-image pairs and learning the inherent distribution of pairs in a subtle neural network model, Unicorn can reconstruct a visually pleasing image from a randomly generated noise with only one index number. The neural model serves as the unified decoder of images while the noises and indexes corresponds to explicit representations. As a proof of concept, we propose an effective and efficient prototype of Unicorn based on latent diffusion models with tailored model designs. Quantitive and qualitative experimental results demonstrate that our prototype achieves significant bitrates reduction compared with EIC and IIC algorithms. More impressively, benefitting from the unified decoder, our compression ratio escalates as the quantity of images increases. We envision that more advanced model designs will endow Unicorn with greater potential in image compression. The code will be made publicly available upon publication.
Qi Zheng 0004, Haozhi Wang, Zihao Liu 0015, Zhijian Hao, Bu Chen, Min Li 0033, Rui Wan, Peiye Liu, Yanheng Lu, Dimin Niu, Jinjia Zhou, Minge Jing, Yibo Fan
ACM Multimedia13
2025 Defective Pixel Corrector for Line Scan and Area Scan Image Sensors
abstract
The defective pixel corrector is an essential component of the image processor, which detects and corrects defective pixels in the image. Processing discrete defective pixels that exist in the output of area-scan image sensors is the focus of current research. However, the case of defective columns produced by line scan image sensors is a massive challenge for existing algorithms. In this paper, we develop novel algorithms for the detection and correction of columnar defective pixels by modeling the properties of defective columns in the output images of line scan image sensors. In addition, we design the non-extremum verdict and texture adaptive correction for clustered defective pixels in the area scan sensors and apply them to current algorithms to get enhancement, and new algorithms are obtained. Moreover, we propose a generalized hardware architecture for detection and correction. In the experimental stage, the proposed methods are compared with the state-of-the-art methods, both at the algorithmic level and in hardware implementation. The experimental results show that, compared to the current widely used algorithms, our methods for line-scan image sensors improve the detection rate by 28%, the detection precision by 90%, and the image restoration quality by 35% with a 10% increase in hardware consumption. For area-scan image sensors, the proposed non-extremum verdict improves the detection rate by 25%, improves the precision from 0.056 to 0.983, and the texture adaptive correction mitigates the image blurring caused by the correction process, with a 37% improvement in image quality, while incurring only a 1% increase in hardware overhead. In addition, our proposed algorithms achieve comparable detection and correction results with far fewer computational and storage resource requirements than machine learning-based algorithms.
Liyuan Peng, Mingyu Wang 0001, Wenhong Li, Minge Jing, Xiaoyang Zeng
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 Affine Motion Estimation Hardware Implementation With 51.7%/67.5% Internal Bandwidth Reduction for Versatile Video Coding
abstract
Versatile Video Coding (VVC) employs Affine Motion Compensation (AMC) to process scenes with high-order motion. To improve AMC efficiency, the Affine Motion Estimation (AME) process based on the gradient-based iterative algorithm (GIA) and block match algorithm (BMA) is introduced to the VVC Test Model (VTM). However, the AME process is highly complex and difficult for hardware implementation in real-time applications. In this context, this paper proposes a hardware-friendly AME algorithm and implements the corresponding accelerator. Firstly, the weighted least squares regression is used to reduce the iteration of GIA. Then an iteration-free search scheme is proposed to remove the search dependence during the GIA and BMA process. In addition, a motion vector clamping mechanism and four-level memory organization are proposed to solve the problem of reference pixel reading conflict, which reduces 51.7% and 67.5% internal bandwidth of the AME accelerator. Compared with the default AME process of VTM 16.0, experimental results show that the proposed algorithm reduces AME run time by 81.63% while the corresponding Bjontegaard Delta Bit Rate (BDBR) loss is only 0.492%. The proposed AME accelerator can flexibly support AME search tasks in various configurations. Synthesized with the TSMC 28nm process, the proposed architecture has a gate count of 1313K and a power consumption of 156.83 mW. It can achieve$7680\times [email protected]~30fps and the corresponding BDBR loss is 0.492%~1.835%.
Shushi Chen, Leilei Huang, Zhao Zan, Zhijian Hao, Hao Zhang 0126, Xiaoxiang Chen, Minge Jing, Xiaoyang Zeng, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.7
2025 An 8K@120fps Advanced Entropy Coding Hardware Design for AVS3
abstract
The third generation audio video coding standard (AVS3) is the latest video coding standard developed by the China AVS working group. The advanced entropy coding (AEC) tool in AVS3 has critical bin-to-bin data dependencies leading to difficulties in parallelization. The use of a 16384-entries lookup table (LUT) in the AEC context update algorithm poses challenges in balancing area and performance. To address these issues, we propose a high-performance, area-efficient hardware design. Firstly, we introduce a novel multicycle-path parallel architecture and optimize area cost through hardware reuse. Next, we construct a context modeling processing unit to replace the large LUT, significantly reducing area. Finally, we propose a new LUT-free dual-context modeling processing unit, effectively resolving critical paths introduced by parallel context conflicts. As a result, our design processes 2.63 bins per cycle. The synthesis results based on the GlobalFoundries’ 28nm process indicate that its maximum frequency is 990MHz, with an overall throughput of 2604 Mbin/s. Compared to state-of-the-art designs, our design leads in performance by 24.7% while reducing area by 28%.
Wei Li 0257, Leilei Huang, Chenlong He, Minge Jing, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.4
2025 Flips: A Flexible Partitioning Strategy Near Memory Processing Architecture for Recommendation System
abstract
Personalized recommendation systems are massively deployed in production data centers. The memory-intensive embedding layers of recommendation systems are the crucial performance bottleneck, with operations manifesting as sparse memory lookups and simple reduction computations. Recent studies propose near-memory processing (NMP) architectures to speed up embedding operations by utilizing high internal memory bandwidth. However, these solutions typically employ a fixed vector partitioning strategy that fail to adapt to changes in data center deployment scenarios and lack practicality. We propose Flips, aflexiblepartitioningstrategy NMP architecture that accelerates embedding layers. Flips supports more than ten partitioning strategies through hardware-software co-design. Novel hardware architectures and address mapping schemes are designed for the memory-side and host-side. We provide two approaches to determine the optimal partitioning strategy for each embedding table, enabling the architecture to accommodate changes in deployment scenarios. Importantly, Flips is decoupled from the NMP level and can utilize rank-level, bank-group-level and bank-level parallelism. In peer-level NMP evaluations, Flips outperforms state-of-the-art NMP solutions, RecNMP, TRiM, and ReCross by up to 4.0×, 4.1×, and 3.5×, respectively.
Yudi Qiu, Lingfei Lu, Shiyan Yi, Minge Jing, Xiaoyang Zeng, Yibo Fan
IEEE Trans. Parallel Distributed Syst.4