Chenming Zhang

dblp:120/8467 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 An Efficient Bit-level Sparse MAC-accelerated Architecture with SW/HW Co-design on FPGA
abstract
Exploring bit-level sparsity in the MAC process has been proven to be an important method for improving the efficiency of neural network feedforward processing. The reconfigurable platform offers possibilities for identifying the bitlevel unstructured redundancy during inference with different DNN models. Researchers noticed significant progress in valueaware accelerators on ASICs, yet we are concerned about the few studies on FPGAs. This paper observed the limitations of implementing bit-level sparsity optimizations using FPGA and proposed a software/architecture co-design solution. Specifically, by introducing LUT-friendly encoding with adaptable granularity and hardware structure supporting multiplication time uncertainty, we achieved a better trade-off between potential redundancy and accuracy with compatibility and scalability. Experiments show that under accurate calculation, PEs are up to $2.2 \times$ smaller than bit-parallel ones, and our design boosts performance by $1.04 \times$ to $1.74 \times$ and $1.40 \times$ to $2.79 \times$ over bitparallel and Booth-based designs, respectively.
Chenming Zhang, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou
DAC1
2025 Lightstereo: Channel Boost is All You Need for Efficient 2D Cost Aggregation
abstract
We present LightStereo, a cutting-edge stereomatching network crafted to accelerate the matching process. Departing from conventional methodologies that rely on aggregating computationally intensive 4D costs, LightStereo adopts the 3D cost volume as a lightweight alternative. While similar approaches have been explored previously, our breakthrough lies in enhancing performance through a dedicated focus on the channel dimension of the 3D cost volume, where the distribution of matching costs is encapsulated. Our exhaustive exploration has yielded plenty of strategies to amplify the capacity of the pivotal dimension, ensuring both precision and efficiency. We compare the proposed LightStereo with existing state-of-the-art methods across various benchmarks, which demonstrate its superior performance in speed, accuracy, and resource utilization. LightStereo achieves a competitive EPE metric in the SceneFlow datasets while demanding a minimum of only 22 GFLOPs and 17 ms of runtime, and ranks 1st on KITTI 2015 among real-time models. Our comprehensive analysis reveals the effect of 2 D cost aggregation for stereo matching, paving the way for realworld applications of efficient stereo systems. Code is available at https://github.com/XiandaGuo/OpenStereo.
Xianda Guo, Chenming Zhang, Youmin Zhang 0008, Wenzhao Zheng, Dujun Nie, Matteo Poggi, Long Chen 0005
ICRA2
2025 Adjacent-view Transformers for Supervised Surround-view Depth Estimation
abstract
Depth estimation has been widely studied and serves as the fundamental step of 3D perception for robotics and autonomous driving. Though significant progress has been made in monocular depth estimation in the past decades, these attempts are mainly conducted on the KITTI benchmark with only front-view cameras, which ignores the correlations across surround-view cameras. In this paper, we propose an Adjacent-View Transformer for Supervised Surround-view Depth estimation (AVT-SSDepth), to jointly predict the depth maps across multiple surrounding cameras. Specifically, we employ a global-to-local feature extraction module that combines CNN with transformer layers for enriched representations. Further, the adjacent-view attention mechanism is proposed to enable the intra-view and inter-view feature propagation. The former is achieved by the self-attention module within each view, while the latter is realized by the adjacent attention module, which computes the attention across multi-cameras to exchange the multi-scale representations across surround-view feature maps. In addition, AVT-SSDepth has strong cross-dataset generalization. Extensive experiments show that our method achieves superior performance over existing state-of-the-art methods on both DDAD and nuScenes datasets. Code is available at https://github.com/XiandaGuo/SSDepth.
Xianda Guo, Wenjie Yuan 0004, Chenming Zhang, Qin Zou 0001, Long Chen 0005
IROS5
2025 A PVT-Insensitive 7-Bit Coarse-Fine Ratio-Metric Digital-to-Time Converter for Fractional-N Phase-Locked Loops in 65-nm CMOS
abstract
This paper introduces a coarse-fine ratio-metric digital-to-time converter (DTC) for fractional-N phase-locked loops (PLLs). The proposed DTC exploits the period of a well-defined clock from the PLL as the time-domain reference, thereby achieving PVT-insensitive resolution. The programmable delay is achieved by setting the initial charge on a load capacitor in the pre-discharging phase. It is followed by constant-slope discharging to trigger a comparator and generate the delayed clock edge. The 4-bit MSBs of the DTC is realized by programming the discharging pulse width of the MSB of an IDAC, while the lower 3 bits of the DTC is implemented by controlling the discharging current from the IDAC over a fixed pulse width. Designed in 65-nm CMOS technology, the proposed 7-bit DTC achieves 4.72-ps resolution with a maximum differential nonlinearity (DNL) and integral nonlinearity (INL) of 1.45 ps and 2.26 ps at 600-ps full-scale delay range in simulations. Over PVT variations, the DTC resolution changes by only ±0.85%. The proposed DTC consumes 720uW from 1V supply and occupies 0.0124-mm2silicon area.
Shenjian Zhang, Chenming Zhang, Shengtao Yi, Xuewei Ding, Zhirui Zong
ISCAS2
2025 RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
abstract
Creating recipe images is a key challenge in food computing, with applications in culinary education and multimodal recipe assistants. However, existing datasets lack fine-grained alignment between recipe goals, step-wise instructions, and visual content. We present RecipeGen, the first large-scale, real-world benchmark for recipe-based Text-to-Image (T2I), Image-to-Video (I2V), and Text-to-Video (T2V) generation. RecipeGen contains 26,435 recipes, 196,724 images, and 4,491 videos, covering diverse ingredients, cooking procedures, styles, and dish types. We further propose domain-specific evaluation metrics to assess ingredient fidelity and interaction modeling, benchmark representative T2I, I2V, and T2V models, and provide insights for future recipe generation models. Project page is available at https://wenbin08.github.io/RecipeGen.
Ruoxuan Zhang, Jidong Gao, Bin Wen 0001, Chenming Zhang, Hong-Han Shuai, Wen-Huang Cheng
ACM Multimedia5
2025 SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
abstract
Accurate spatial reasoning in outdoor environments—covering geometry, object pose, and inter-object relationships—is fundamental to downstream tasks such as mapping, motion forecasting, and high-level planning in autonomous driving. We introduce SURDS, a large-scale benchmark designed to systematically evaluate the spatial reasoning capabilities of vision language models (VLMs). Built on the nuScenes dataset, SURDS comprises 41,080 vision–question–answer training instances and 9,250 evaluation samples, spanning six spatial categories: orientation, depth estimation, pixel-level localization, pairwise distance, lateral ordering, and front–behind relations. We benchmark leading general-purpose VLMs, including GPT, Gemini, and Qwen, revealing persistent limitations in fine-grained spatial understanding. To address these deficiencies, we go beyond static evaluation and explore whether alignment techniques can improve spatial reasoning performance. Specifically, we propose a reinforcement learning–based alignment scheme leveraging spatially grounded reward signals—capturing both perception-level accuracy (location) and reasoning consistency (logic). We further incorporate final-answer correctness and output-format rewards to guide fine-grained policy adaptation. Our GRPO-aligned variant achieves overall score of 40.80 in SURDS benchmark. Notably, it outperforms proprietary systems such as GPT-4o (13.30) and Gemini-2.0-flash (35.71). To our best knowledge, this is the first study to demonstrate that reinforcement learning–based alignment can significantly and consistently enhance the spatial reasoning capabilities of VLMs in real-world driving contexts. We release the SURDS benchmark, evaluation toolkit, and GRPO alignment code through: https://github.com/XiandaGuo/Drive-MLLM.
Xianda Guo, Ruijun Zhang, Yiqun Duan, Dujun Nie, Wenke Huang 0003, Chenming Zhang, Shuai Liu 0009, Hao Zhao 0002, Long Chen 0005
NeurIPS7
2025 Feature Disentanglement in GANs for Photorealistic Multi-view Hair Transfer
abstract
Abstract Fast and highly realistic multi‐view hair transfer plays a crucial role in evaluating the effectiveness of virtual hair try‐on systems. However, GAN‐based generation and editing methods face persistent challenges in feature disentanglement. Achieving pixel‐level, attribute‐specific modifications—such as changing hairstyle or hair color without affecting other facial features—remains a long‐standing problem. To address this limitation, we propose a novel multi‐view hair transfer framework that leverages a hair‐only intermediate facial representation and a 3D‐guided masking mechanism. Our approach disentangles tri‐plane facial features into spatial geometric components and global style descriptors, enabling independent and precise control over hairstyle and hair color. By introducing a dedicated intermediate representation focused solely on hair and incorporating a two‐stage feature fusion strategy guided by the generated 3D mask, our framework achieves fine‐grained local editing across multiple viewpoints while preserving facial integrity and improving background consistency. Extensive experiments demonstrate that our method produces visually compelling and natural results in side‐to‐front view hair transfer tasks, offering a robust and flexible solution for high‐fidelity hair reconstruction and manipulation.
Jiayi Xu 0002, Chenming Zhang, Xiaogang Jin 0001, Yaohua Ji
Comput. Graph. Forum3
2025 Learning multi-scale features automatically from food and ingredients
Ruoxuan Zhang, Dantong Ouyang, Ximing Li 0002, Hongtao Bai, Chenming Zhang, Lili He 0002
Multim. Syst.5
2024 GenAD: Generative End-to-End Autonomous Driving
Wenzhao Zheng, Xianda Guo, Chenming Zhang, Long Chen 0005
ECCV (65)4
2024 End-to-End On-Orbit Objects Detection with ConvNets
abstract
As space activities expand, the quantity of space debris also increases, posing significant risks to spacecraft and infrastructure. Space situational awareness (SSA) is essential for avoiding collisions and limiting the generation of extra debris. Accurate and efficient detection of space objects plays a critical role in achieving this goal. Our research focuses on the development of detection algorithms that are both precise and quick, taking into account the real-time and safety of spacecraft operations in orbit. For the first time, we take a fully Convolutional Neural Networks (ConvNets) to run the query-based end-to-end object detection for SSA. We further compare its performance with the newest YOLOv9 algorithm. This is an innovative attempt at SSA. First of all, it does not require predefined a priori anchor boxes or complex post-processing strategies such as Non-Maximum Suppression (NMS), and can directly achieve end-to-end target detection. Secondly, the fully ConvNets are selected as the basic framework, which not only retains the advantages of self-attention mechanism, but also greatly improves the computing efficiency. These methods show outstanding performance on the challenging SPARK data set. The fully ConvNets approach achieves end-to-end detection by utilizing the query attention mechanism, excluding the need for complicated post-processing in traditional object detection methods and with higher efficiency. YOLOv9 involves an enhanced feature pyramid fusion and a more powerful detection head, potentially resulting in higher precision. Following that, we will thoroughly assess the speed, accuracy, and trade-offs of the two algorithms using actual data sets in order to deliver an efficient and dependable solution for detecting targets in aerospace sensing missions.
Bo Ru, Pengrong Hou, Qinghao Chu, Zikang Zeng, Chenming Zhang, Zhelong Wang
SMC6
2024 Personalized hairstyle and hair color editing based on multi-feature fusion
abstract
Abstract In the metaverse era, virtual design of hairstyle becomes very popular for personalized aesthetics. As hair design tasks can be decomposed into hair attribute editing and generation, the development of generative adversarial networks (GANs) has significantly prompted its development. The majority of the existing algorithms focus on transferring the overall hair region from one face to another, which ignore fine control over the color and geometric features. Furthermore, these algorithms may result in unnatural generation results. In this paper, we propose a hair modification framework that learns hairstyle information from a reference face mask and color information from a guidance face image. Firstly, the features of the input face image and reference images are extracted through a group of encoders, and then divided into feature vectors of coarse, medium, and fine levels. Secondly, multi-level feature vectors are fused in the latent space using attention-based modulation modules. Finally, the fused feature vector is passed through a StyleGAN generator to generate face images with specified hairstyle and hair color. Experimental results show that the proposed method can finely simulate the hairstyle transition between long and short hair under the constraint of the reference mask, and can produce realistic fusion effects in the hair-covered regions, such as ears, neck, and forehead. Various hair dyeing effects that adapt to personalized characteristics are demonstrated, as facial features including skin color and hair texture are preserved when transferring the hair color.
Jiayi Xu 0002, Chenming Zhang, Weikang Zhu, Li Li 0014, Xiaoyang Mao
Vis. Comput.2
2020 Analysis of the Inter-Stage Signal Leakage in Wide BW Low OSR and High DR CT MASH ΔΣM
abstract
This paper analyzes the error mechanisms that limit the dynamic range (DR) of wide-bandwidth, low-OSR continuous-time (CT) multi-stage noise-shaping (MASH) ΔΣM and proposes a tool, the Signal Leakage Function (SLF), to optimize the architecture, and hence improving DR. The SLF provides new insights on finding the key parameters which influence the inter-stage signal leakage and thus the inter-stage gain (IG). These insights would lead not only to increasing the overall dynamic range in a very power-efficient way, but also decreasing the performance sensitivity to mismatches and other variations.
Lucien J. Breems, Shagun Bajoria, Muhammed Bolatkale, Chenming Zhang, Georgi I. Radulov
ISCAS5
2018 A 1.9 mW 250 MHz Bandwidth Continuous-Time ΣΔ Modulator for Ultra-Wideband Applications
abstract
This paper proposes an architecture design approach for a wideband continuous-time (CT) ΣΔ modulator with ultra-low oversampling ratio (OSR). The ultra-low OSR is beneficial in terms of power consumption for both the clock distribution network and the subsequent decimation filter. In this work, three signal feedforward paths and an additional feedback path are used to reduce the power consumption. Extensive system-level simulations demonstrate the effectiveness of the proposed solutions. Furthermore, this work verifies the proposed methods by transistor-level design and simulations of a 2 GHz 4th-order CT ΣΔ modulator achieving an SNDR of 46 dB in a signal band of 250 MHz while consuming only 1.91 mW of power in 40 nm CMOS. The proposed solutions enable CT ΣΔ modulators for low power ultra-wideband (UWB) applications.
Marios Neofytou, Meiyi Zhou, Muhammed Bolatkale, Chenming Zhang, Georgi I. Radulov, Peter G. M. Baltus, Lucien J. Breems
ISCAS5
2018 A 2 GHz 0.98 mW 4-bit SAR-Based Quantizer with ELD Compensation in an UWB CT ΣΔ Modulator
abstract
This paper presents a 2 GHz 4-bit asynchronous successive approximation register (SAR) quantizer to enable an ultra-wideband continuous-time (CT) sigma-delta modulator (SDM). Low latency is required for the stability of the SDM. The excess-loop-delay compensation (ELDC) is embedded in the SAR quantizer by adding an extra switched-capacitor DAC segment with two separate reference voltages. To achieve high speed, a gm-boosted StrongARM latch and the monotonic switching scheme are used. This paper presents the transistor-level circuit implementation and the complete verification of the CT SDM. Simulation results show the power consumption of this SAR-based quantizer including ELDC is 0.98 mW, leading to a very competitive Figure-of-Merit of 30.6 fJ/conv.-step.
Meiyi Zhou, Marios Neofytou, Muhammed Bolatkale, Chenming Zhang, Pierluigi Cenci, Georgi I. Radulov, Peter G. M. Baltus, Lucien J. Breems
ISCAS5
2017 Current-mode multi-path excess loop delay compensation for GHz sampling CT ΣΔ ADCs
abstract
This paper proposes both system-level and circuit-level solutions of a current-mode multi-path excess loop delay (ELD) compensation technique for continuous-time (CT) ΣΔ ADCs with multi-bit quantization and several GHz sampling rate. Thanks to the proposed solutions, the amplifier of the loop filter is not in the fast feedback (FB) loop; the delay of the pre-amplifier of the comparator is removed; and the effective regeneration time of the comparator latch is maximized. The proposed novelties enable CT ΣΔ ADCs with wide signal bandwidth and improved power efficiency. Extensive transistor-level simulations demonstrate their effectiveness and robustness. This work validates the proposed methods by transistor level design and simulations of an 8.4 GHz MASH ΣΔ ADC achieving an SNDR of 71 dB in a signal band of 600 MHz. This shows that our proposed solutions enable power-efficient multi-GHz ΣΔ ADC applications.
Chenming Zhang, Lucien J. Breems, Georgi I. Radulov, Muhammed Bolatkale, Hans Hegt, Arthur H. M. van Roermund
ISCAS1
2016 A digital calibration technique for wide-band CT MASH ΣΔ ADCs with relaxed filter requirements
abstract
This paper proposes a simple digital automatic calibration method for wide-band continuous-time (CT) MASH ΣΔ ADCs. The main contribution of this method is the calibration of the errors due to the limited DC gain and 2nd pole of the loop filter integrators. The digital noise cancellation filters are calibrated by successive estimation of the DC gains and the 2nd pole, and updating the coefficients of the FIR filters. Extensive system-level and transistor-level simulations demonstrate the effectiveness and robustness of the proposed method. For an exemplary MASH ΣΔ ADC with BW of 600 MHz and SQNR of 75 dB, it can relax the DC gain requirement by 15 dB, and the 2nd pole requirement by three time, making it an enabling technique for power-efficient GHz-range ΣΔ ADC applications.
Chenming Zhang, Lucien J. Breems, Georgi I. Radulov, Muhammed Bolatkale, Hans Hegt, Arthur H. M. van Roermund
ISCAS1