EDBT 2026 Demo / reviewers in the wild / expert
Shengchao Zhou
dblp:194/4326
· DBLP profile ↗
11ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ASSIST-3D: Adapted Scene Synthesis for Class-Agnostic 3D Instance SegmentationabstractClass-agnostic 3D instance segmentation tackles the challenging task of segmenting all object instances, including previously unseen ones, without semantic class reliance. Current methods struggle with generalization due to the scarce annotated 3D scene data or noisy 2D segmentations. While synthetic data generation offers a promising solution, existing 3D scene synthesis methods fail to simultaneously satisfy geometry diversity, context complexity, and layout reasonability, each essential for this task. To address these needs, we propose an Adapted 3D Scene Synthesis pipeline for class-agnostic 3D Instance SegmenTation, termed as ASSIST-3D, to synthesize proper data for model generalization enhancement. Specifically, ASSIST-3D features three key innovations, including 1) Heterogeneous Object Selection from extensive 3D CAD asset collections, incorporating randomness in object sampling to maximize geometric and contextual diversity; 2) Scene Layout Generation through LLM-guided spatial reasoning combined with depth-first search for reasonable object placements; and 3) Realistic Point Cloud Construction via multi-view RGB-D image rendering and fusion from the synthetic scenes, closely mimicking real-world sensor data acquisition. Experiments on ScanNetV2, ScanNet++, and S3DIS benchmarks demonstrate that models trained with ASSIST-3D-generated data significantly outperform existing methods. Further comparisons underscore the superiority of our purpose-built pipeline over existing 3D scene synthesis approaches. Shengchao Zhou, Jiehong Lin, Jiahui Liu 0012, Shizhen Zhao, Chirui Chang, Xiaojuan Qi 0001 |
AAAI | 1 |
| 2026 | BLCIM: An Efficient Radix-16 Booth LUT-Based SRAM-CIM Architecture with Algorithm-Hardware Co-Optimization for NTTabstractLattice-based cryptography relies heavily on the Number Theoretic Transform (NTT), whose performance is dominated by modular multiplication and data movement. This paper proposes BLCIM, the first Radix-16 Booth LUT-based compute-in-memory (CIM) NTT accelerator. We propose an algorithm that precomputes partial modular multiplications with fixed rotation factors and uses input Booth-encoded search results. Then, based on the Booth encoding, it decides whether to perform shifting and inversion to simplify modular multiplication. In addition to the proposed sparsity-aware and stage-skipping schemes, the Radix-16 Booth LUT-based algorithm significantly improves NTT performance at low energy cost. By implementing the algorithm on the SRAM-CIM architecture with a lightweight pipeline design, we achieved nearly 100% utilization for the NTT circuit during acceleration. Simulated in 28 nm CMOS technology, the proposed BLCIM achieves only 4.53% latency and 58.87% energy consumption of the latest work. Qianhua Li, Hongrui Meng, Chunshan Wang, Shengchao Zhou, Teng Zou, Yufeng Xie 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2026 | Horizontal-Parallel ADC-less Sparsity-Clock-Aware RRAM CIM Macro for edge AI devices
Teng Zou, Shengchao Zhou, Hongrui Meng, Yufeng Xie 0001 |
ISCAS | 2 |
| 2026 | A 40-nm Training-Inference STT-MRAM Near-Memory Computing Macro for Memory-Augmented Neural Network AccelerationabstractRecently, memory-augmented neural networks (MANNs) have gained significant attention as a critical solution for few-shot learning (FSL). These networks leverage external memory to store prior knowledge, thereby enhancing classification efficiency. Spin-transfer torque magnetic random access memory (STT-MRAM) is particularly suited for this application due to its compact cell size, excellent data retention, and scalability. In this article, we introduce a STT-MRAM-based near-memory computing (NMC) macro specifically designed for MANNs. Our approach incorporates several key innovations aimed at overcoming challenges in hardware implementation while improving MANN performance as follows: 1) a parallel computing architecture within the NMC to expedite$L1$distance computations; 2) a memory invert coding (MIC) and self-termination write (STW) scheme that reduce write operations and energy consumption, addressing the issues of frequent writes and high write currents during the training phase of MANNs; 3) a dynamic offset-compensation sense amplifier (DOC-SA) and high-throughput switch-capacitor (HTSC) readout scheme to improve read accuracy and throughput, tackling low read margins and limited readout bandwidth; 4) an exploration of MANN architectures validates the reusability of the NMC macro. The optimized matching-networks (MCHnets)-based structure achieves an accuracy exceeding 90% in five-way and eight-way Omniglot classification tasks. Fabricated with a 40-nm CMOS technology, our design achieves classification accuracies of 96.37% for eight-way-five-shot tasks and 93.72% for 16-way-five-shot tasks on the Omniglot dataset utilizing the optimized MCHnet, showcasing an impressive energy efficiency of 6.47 TOPS/W at the basis of 16-bit$L1$distance computing in the classification tasks of MANN. Shengchao Zhou, Hongrui Meng, Yajun Wu, Zizhao Ma, Teng Zou, Tai Min, Shaohao Wang, Yufeng Xie 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | A 40nm STT-MRAM Near-Memory Computing Macro for Memory-Augmented Neural Network AccelerationabstractMemory-augmented neural network (MANN) has gained attention as a pivotal solution for few-shot learning (FSL). Among the candidates for associative memory in MANN accelerators, spin-transfer torque magnetic random-access memory (STT-MRAM) stands out for its compact cell area, long data retention time, and excellent scalability. In this paper, we propose an STT-MRAM near-memory computing (NMC) macro for MANN acceleration. The macro contains following innovations: 1) An array-level parallel computing architecture for L1 distance calculation. 2) A low-area-overhead memory-invert coding technique to reduce write energy consumption. 3) A configurable dynamic offset-compensation sense amplifier (CDOC-SA) to improve classification accuracy. Fabricated in 40nm CMOS process, our macro demonstrates an energy efficiency of 6.47 TOPS/W, achieving the classification accuracy of 98.3% and 93% for 8-way-5-shot tasks and 16-way-5-shot tasks on the Omniglot dataset. Hongrui Meng, Yajun Wu, Shengchao Zhou, Zizhao Ma, Tai Min, Shaohao Wang, Yufeng Xie 0001 |
ISCAS | 3 |
| 2025 | High sensing margin and parallelism 6T-2MTJ SOT-MRAM based TCAM for energy-Efficient Similarity Priority calculation in MANNsabstractWith the development of AI applications, there is a demand for high parallel similarity priority calculation, like in MANNs. TCAM, as a type of memory for high parallel searching work, is suitable to perform the task. However, the CMOS based TCAMs suffer from area overhead, static power consumption and data-loss while power failure. The SOT-MRAM presents a promising alternative for TCAM design due to its nonvolatility, low area overhead and no static power consumption. But, its low on/off ratio results in low sensing margin which limits its processing speed. To address the challenges, we propose a novel TCAM structure with high sensing margin, parallelism and low energy consumption. This work proposes the followings: (1) A novel 6T-2MTJ sot-mram based TCAM structure is proposed. Simulations show it has a 2-3x improvement in sensing margin over other MRAM based ones, 1.37x improvement in searching energy (fJ/bit) over other nonvolatile technologies based ones, and 23% area reduction over other CMOS based ones. (2) A segmented power supply scheme is presented to meet the parallelism need of MANNs which improves the parallelism by 4-8x compared to no segmented one. Shengchao Zhou, Teng Zou, Zeming Wang, Xianwu Hu, Hongrui Meng, Yajun Wu, Chuxin Zhang, Caihua Wan, Yufeng Xie 0001 |
ISCAS | 1 |
| 2023 | Robust Feature Rectification of Pretrained Vision Models for Object RecognitionabstractPretrained vision models for object recognition often suffer a dramatic performance drop with degradations unseen during training. In this work, we propose a RObust FEature Rectification module (ROFER) to improve the performance of pretrained models against degradations. Specifically, ROFER first estimates the type and intensity of the degradation that corrupts the image features. Then, it leverages a Fully Convolutional Network (FCN) to rectify the features from the degradation by pulling them back to clear features. ROFER is a general-purpose module that can address various degradations simultaneously, including blur, noise, and low contrast. Besides, it can be plugged into pretrained models seamlessly to rectify the degraded features without retraining the whole model. Furthermore, ROFER can be easily extended to address composite degradations by adopting a beam search algorithm to find the composition order. Evaluations on CIFAR-10 and Tiny-ImageNet demonstrate that the accuracy of ROFER is 5% higher than that of SOTA methods on different degradations. With respect to composite degradations, ROFER improves the accuracy of a pretrained CNN by 10% and 6% on CIFAR-10 and Tiny-ImageNet respectively. Shengchao Zhou, Gaofeng Meng, Zhaoxiang Zhang 0001, Shiming Xiang |
AAAI | 1 |
| 2023 | UniDistill: A Universal Cross-Modality Knowledge Distillation Framework for 3D Object Detection in Bird's-Eye ViewabstractIn the field of 3D object detection for autonomous driving, the sensor portfolio including multi-modality and single-modality is diverse and complex. Since the multi-modal methods have system complexity while the accuracy of single-modal ones is relatively low, how to make a tradeoff between them is difficult. In this work, we propose a universal cross-modality knowledge distillation framework (UniDistill) to improve the performance of single-modality detectors. Specifically, during training, UniDistill projects the features of both the teacher and the student detector into Bird's-Eye-View (BEV), which is a friendly representation for different modalities. Then, three distillation losses are calculated to sparsely align the foreground features, helping the student learn from the teacher without introducing additional cost during inference. Taking advantage of the similar detection paradigm of different detectors in BEV, UniDistill easily supports LiDAR-to-camera, camera-to-LiDAR, fusion-to-LiDAR and fusion-to-camera distillation paths. Furthermore, the three distillation losses can filter the effect of misaligned background information and balance between objects of different sizes, improving the distillation effectiveness. Extensive experiments on nuScenes demonstrate that UniDistill effectively improves the mAP and NDS of student detectors by 2.0%~3.2%. Shengchao Zhou, Weizhou Liu, Shuchang Zhou 0001 |
CVPR | 1 |
| 2022 | Scheduling a single batch processing machine with non-identical two-dimensional job sizes
Shengchao Zhou, Mingzhou Jin, Huaping Chen 0001 |
Expert Syst. Appl. | 1 |
| 2021 | A Self-Adaptive Differential Evolution Algorithm for Scheduling a Single Batch-Processing Machine With Arbitrary Job Sizes and Release TimesabstractBatch-processing machines (BPMs) can process a number of jobs at a time, which can be found in many industrial systems. This article considers a single BPM scheduling problem with unequal release times and job sizes. The goal is to assign jobs into batches without breaking the machine capacity constraint and then sort the batches to minimize the makespan. A self-adaptive differential evolution algorithm is developed for addressing the problem. In our proposed algorithm, mutation operators are adaptively chosen based on their historical performances. Also, control parameter values are adaptively determined based on their historical performances. Our proposed algorithm is compared to CPLEX, existing metaheuristics for this problem and conventional differential evolution algorithms through comprehensive experiments. The experimental results demonstrate that our proposed self-adaptive algorithm is more effective than other algorithms for this scheduling problem. Shengchao Zhou, Lining Xing 0001, Ni Du, Ling Wang 0001, Qingfu Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2019 | Pseudo transformation mechanism between resource allocation and bin-packing in batching environments
Xinle Liang, Shengchao Zhou, Huaping Chen 0001, Rui Xu 0004 |
Future Gener. Comput. Syst. | 2 |