VLDB 2026 Research / reviewers in the wild / expert
Siming Zheng
dblp:235/1399
· DBLP profile ↗
15ranked-venue papers
3as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Breaking Measurement Barriers: From Compressed Sensing to Deep ReconstructionabstractDeep learning methods have achieved remarkable success in image compressed sensing (CS) task, namely reconstructing a high-fidelity image from its compressed measurement. However, existing methods are deficient in incoherent compressed measurement at sensing phase and implicit measurement representations at reconstruction phase, limiting the overall performance. In this work, we answer two questions: (i) how to improve the measurement incoherence for decreasing the ill-posedness; (ii) how to learn informative representations from measurements. To this end, we propose a novel asymmetric Kronecker CS (AKCS) model and theoretically present its better incoherence than previous Kronecker CS with minimal increase of complexity. Moreover, apart from the explicit measurement representations in gradient descent projection in unfolding networks, we further propose a measurement-aware cross attention (MACA) mechanism to learn implicit measurement representations. We integrate AKCS and MACA into a widely-used unfolding architecture to get a measurement-enhanced unfolding network (MEUNet). Extensive experiments demonstrate that the proposed MEUNet achieves state-of-the-art (SOTA) performance in reconstruction accuracy with high efficiency. Gang Qu 0005, Ping Wang 0029, Siming Zheng, Xin Yuan 0002 |
AAAI | 3 |
| 2026 | Realism Control One-step Diffusion for Real-world Image Super ResolutionabstractPre-trained diffusion models have shown great potential in real-world image super-resolution (Real-ISR) tasks by enabling high-resolution reconstructions. While one-step diffusion (OSD) methods significantly improve efficiency compared to traditional multi-step approaches, they still have limitations in balancing fidelity and realism across diverse scenarios. Since the OSDs for SR are usually trained or distilled by a single timestep, they lack flexible control mechanisms to adaptively prioritize these competing objectives, which are inherently manageable in multi-step methods through adjusting sampling steps. To address this challenge, we propose a Realism Controlled One-step Diffusion (RCOD) framework for Real-ISR. RCOD provides a latent domain grouping strategy that enables explicit control over fidelity-realism trade-offs during the noise prediction phase with minimal training paradigm modifications and original training data. A degradation-aware sampling strategy is also introduced to align distillation regularization with the grouping strategy and enhance the controlling of trade-offs. Moreover, a visual prompt injection module is used to replace conventional text prompts with degradation-aware visual tokens, enhancing both restoration accuracy and semantic consistency. Our method achieves superior fidelity and perceptual quality while maintaining computational efficiency. Extensive experiments demonstrate that RCOD outperforms state-of-the-art OSD methods in both quantitative metrics and visual qualities, with flexible realism control capabilities in the inference stage. Zongliang Wu, Siming Zheng, Peng-Tao Jiang, Xin Yuan 0002 |
AAAI | 2 |
| 2026 | Energy-based haptic rendering for real-time surgical simulation
Mingbo Hu, Wenli Xiu, Siming Zheng, Shuai Li 0001, Aimin Hao |
Comput. Graph. | 5 |
| 2026 | A lightweight and real-time surgical action detection framework using multi-contextual and decoupled representationsabstractAccurate detection of surgical actions in minimally invasive procedures is a critical step toward developing intelligent operative assistance systems. In this work, we propose Surgical You Only Look Once detector (Surg-YOLO), an efficient and high-precision surgical action detection framework built upon the YOLO version 11 (YOLOv11) architecture, specifically optimized for the spatio-temporal complexities of surgical environments. Surg-YOLO integrates three key architectural innovations: the Enhanced Spatial Pyramid Pooling-Fast (ESPPF) module for capturing rich multi-scale spatial features; the Spatio-Temporal Multi-scale Context Aggregation Module (ST-MCAM), which enhances temporal reasoning and contextual awareness across frames; and the Decoupled Dual-Branch Prediction Head (DDPH) for independently refining classification and localization tasks. Extensive experiments on a large-scale surgical action dataset demonstrate that Surg-YOLO significantly outperforms existing baseline models, achieving superior detection accuracy across multiple evaluation thresholds. Qualitative visualizations further validate the model’s ability to localize subtle and concurrent surgical actions with high precision. These results highlight Surg-YOLO’s potential as a reliable solution for real-time surgical action detection. Siming Zheng, A. S. M. Sharifuzzaman Sagar, Yu Chen 0082, Jun Hoong Chan |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Opportunistic ISAC for Rain Microphysical Parameters Estimation Using LOS-MIMO SystemsabstractThe line-of-sight multiple-input multiple-output (LOS-MIMO) has emerged as a potential solution for increasing spectral efficiency in dense urban microwave links, given limited spectrum resources. Phase variation and power attenuation may introduce uncertainty to the channel estimation, affecting channel capacity and the performance of interference cancellation by zero-forcing receivers. This paper investigates the impact of this uncertainty in the presence of rain. Commercial backhaul links (CMLs) in cellular networks are not only essential for data transmission but have also proven useful for rainfall monitoring. They are now emerging as a new opportunistic integrated sensing and communication (OISAC) application for weather sensing. This paper also studies the use of LOS-MIMO backhaul technology for rain rate estimation. The availability of multiple data streams allows the number of rain estimation values to increase linearly with the minimum number of transmit and receive antennas in the MIMO link, Additionally, the potential for LOS-MIMO microwave links to retrieve parameters related to rain drop size distribution based on measurement data is also explored. This new backhaul solution shows great potential to be used for near-ground environmental monitoring and weather prediction studies, particularly with the advent of the big data era. Congzheng Han, Baofeng Ji 0004, Gaoyuan Zhang, Juan Huo, Yongheng Bi, Weidong Nan, Qixing Feng, Guohui Xin, Siming Zheng, Yele Sun |
IEEE Internet Things J. | 9 |
| 2026 | Error Analysis for Deep ReLU Feedforward Density-Ratio Estimation with Bregman DivergenceabstractWe consider the problem of density-ratio estimation using Bregman Divergence with Deep ReLU feedforward neural networks (BDD). We establish non-asymptotic error bounds for BDD density-ratio estimators, which are minimax optimal up to a logarithmic factor when the data distribution has finite support. As an application of our theoretical findings, we propose an estimator for the KL-divergence that is asymptotically normal, leveraging our convergence results for the deep density-ratio estimator and a data-splitting method. We also extend our results to cases with unbounded support and unbounded density ratios. Furthermore, we show that the BDD density-ratio estimator can mitigate the curse of dimensionality when data distributions are supported on an approximately low-dimensional manifold. Our results are applied to investigate the convergence properties of the telescoping density-ratio estimator proposed by Rhodes (2020). We provide sufficient conditions under which it achieves a lower error bound than a single-ratio estimator. Moreover, we conduct simulation studies to validate our main theoretical results and assess the performance of the BDD density-ratio estimator. Siming Zheng, Guohao Shen, Jian Huang 0003 |
J. Mach. Learn. Res. | 1 |
| 2026 | Sparse Transformer for Ultra-Sparse Sampled Video Compressive SensingabstractDigital cameras consume$\sim 0.1$microjoule per pixel to capture and encode video, resulting in a power usage of$\sim 20$W for a 4K sensor operating at 30 fps. Imagining gigapixel cameras operating at 100-1000 fps, the current processing model is unsustainable. To address this, physical layer compressive measurement has been proposed to reduce power consumption per pixel by 10-100×. Video Snapshot Compressive Imaging (SCI) introduces high frequency modulation in the optical sensor layer to increase effective frame rate. A commonly used sampling strategy of video SCI is Random Sampling (RS) where each mask element value is randomly set to be 0 or 1. Similarly, image inpainting (I2P) has demonstrated that images can be recovered from a fraction of the image pixels. Inspired by I2P, we propose Ultra-Sparse Sampling (USS) regime, where at each spatial location, only one sub-frame is set to 1 and all others are set to 0. We then build a Digital Micro-mirror Device (DMD) encoding system to verify the effectiveness of our USS strategy. Ideally, we can decompose the USS measurement into sub-measurements for which we can utilize I2P algorithms to recover high-speed frames. However, due to the mismatch between the DMD and CCD, the USS measurement cannot be perfectly decomposed. To this end, we proposeBSTFormer, a sparse TransFormer that utilizes local Block attention, global Sparse attention, and global Temporal attention to exploit the sparsity of the USS measurement. Extensive results on both simulated and real-world data show that our method significantly outperforms all previous state-of-the-art algorithms. Additionally, an essential advantage of the USS strategy is its higher dynamic range than that of the RS strategy. Finally, from the application perspective, the USS strategy is a good choice to implement a complete video SCI system on chip due to its fixed exposure time. Code is available athttps://github.com/mcao92/BSTFormer. Siming Zheng, Lishun Wang, David J. Brady, Xin Yuan 0002 |
IEEE Trans. Multim. | 2 |
| 2025 | Photography Perspective Composition: Towards Aesthetic Perspective RecommendationabstractTraditional photography composition approaches are dominated by 2D cropping-based methods. However, these methods fall short when scenes contain poorly arranged subjects. Professional photographers often employ perspective adjustment as a form of 3D recomposition, modifying the projected 2D relationships between subjects while maintaining their actual spatial positions to achieve better compositional balance. Inspired by this artistic practice, we propose photography perspective composition (PPC), extending beyond traditional cropping-based methods. However, implementing the PPC faces significant challenges: the scarcity of perspective transformation datasets and undefined assessment criteria for perspective quality. To address these challenges, we present three key contributions: (1) An automated framework for building PPC datasets through expert photographs. (2) A video generation approach that demonstrates the transformation process from less favorable to aesthetically enhanced perspectives. (3) A perspective quality assessment (PQA) model constructed based on human performance. Our approach is concise and requires no additional prompt instructions or camera trajectories, helping and guiding ordinary users to enhance their composition skills. Lujian Yao, Siming Zheng, Xinbin Yuan, Zhuoxuan Cai, Pu Wu, Jinwei Chen 0003, Bo Li 0026, Peng-Tao Jiang |
NeurIPS | 2 |
| 2025 | Maximum local density-driven non-overlapping radial basis function support kernel neural network
Yang Zhao 0014, Siming Zheng, Jihong Pei |
Inf. Sci. | 2 |
| 2024 | The Potential of Precipitation Parameters Retrieval from LoS-MIMO Microwave LinksabstractCommercial microwave links (CMLs) are used to transmit information between base station towers in cellular networks. Opportunistic remote sensing of rainfall, using the signal level measurements from CMLs, has proven to be highly accurate in rain rate estimation. As traditional CML link has single antenna and polarization setup, its capability of retrieving other precipitation parameters is limited. The latest microwave technology line-of-sight multiple-input multiple-output (LoS MIMO) are equipped with multiple transmitters and receivers to increase capacity. In a 2x2 LOS-MIMO system, each of the two antennas employs a different polarization. In this study, we investigate the potential of using LOS-MIMO microwave link to retrieve parameters related to rain drop size distribution based on measurement data. Congzheng Han, Siming Zheng, Juan Huo, Wenying He, Weidong Nan, Yongheng Bi, Shu Duan, Guowei An |
IGARSS | 2 |
| 2024 | A multi-view consistency framework with semi-supervised domain adaptation
Yuting Hong, Li Dong 0006, Xiaojie Qiu, Hui Xiao 0005, Baochen Yao, Siming Zheng, Chengbin Peng 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | Deep Equilibrium Models for Snapshot Compressive ImagingabstractThe ability of snapshot compressive imaging (SCI) systems to efficiently capture high-dimensional (HD) data has led to an inverse problem, which consists of recovering the HD signal from the compressed and noisy measurement. While reconstruction algorithms grow fast to solve it with the recent advances of deep learning, the fundamental issue of accurate and stable recovery remains. To this end, we propose deep equilibrium models (DEQ) for video SCI, fusing data-driven regularization and stable convergence in a theoretically sound manner. Each equilibrium model implicitly learns a nonexpansive operator and analytically computes the fixed point, thus enabling unlimited iterative steps and infinite network depth with only a constant memory requirement in training and testing. Specifically, we demonstrate how DEQ can be applied to two existing models for video SCI reconstruction: recurrent neural networks (RNN) and Plug-and-Play (PnP) algorithms. On a variety of datasets and real data, both quantitative and qualitative evaluations of our results demonstrate the effectiveness and stability of our proposed method. The code and models are available at: https://github.com/IndigoPurple/DEQSCI. Siming Zheng, Xin Yuan 0002 |
AAAI | 2 |
| 2023 | Unfolding Framework with Prior of Convolution-Transformer Mixture and Uncertainty Estimation for Video Snapshot Compressive ImagingabstractWe consider the problem of video snapshot compressive imaging (SCI), where sequential high-speed frames are modulated by different masks and captured by a single measurement. The underlying principle of reconstructing multi-frame images from only one single measurement is to solve an ill-posed problem. By combining optimization algorithms and neural networks, deep unfolding networks (DUNs) score tremendous achievements in solving inverse problems. In this paper, our proposed model is under the DUN framework and we propose a 3D Convolution-Transformer Mixture (CTM) module with a 3D efficient and scalable attention model plugged in, which helps fully learn the correlation between temporal and spatial dimensions by virtue of Transformer. To our best knowledge, this is the first time that Transformer is employed to video SCI reconstruction. Besides, to further investigate the high-frequency information during the reconstruction process which are neglected in previous studies, we introduce variance estimation characterizing the uncertainty on a pixel-by-pixel basis. Extensive experimental results demonstrate that our proposed method achieves state-of-the-art (SOTA) (with a 1.2dB gain in PSNR over previous SOTA algorithm) results. Code can be found on https://github.com/zsm1211/CTM-SCI. Siming Zheng, Xin Yuan 0002 |
ICCV | 1 |
| 2023 | Multiple discriminant preserving support subspace RBFNNs with graph similarity learning
Yang Zhao 0014, Siming Zheng, Jihong Pei |
Inf. Sci. | 2 |
| 2015 | An Exploratory Study on Social Media in China
Siming Zheng |
ENTER | 2 |