EDBT 2026 Demo / reviewers in the wild / expert
Yupeng Shi
dblp:240/7758
· DBLP profile ↗
11ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A model-data fusion approach utilizing residual subset selection and adaptive thresholding for hybrid electric vehicle fault diagnosis
Kaimei Zhang, Shaohua Wang 0004, Dehua Shi, Xingke An, Huanming Huang, Yupeng Shi, Qilin Yao |
Expert Syst. Appl. | 6 |
| 2025 | IDEA-Bench: How Far are Generative Models from Professional Designing?abstractRecent advancements in image generation models enable the creation of high-quality images and targeted modifications based on textual instructions. Some models even support multimodal complex guidance and demonstrate robust task generalization capabilities. However, they still fall short of meeting the nuanced, professional demands of designers. To bridge this gap, we introduce IDEA-Bench, a comprehensive benchmark designed to advance image generation models toward applications with robust task generalization. IDEA-Bench comprises 100 professional image generation tasks and 275 specific cases, categorized into five major types based on the current capabilities of existing models. Furthermore, we provide a representative subset of 18 tasks with enhanced evaluation criteria to facilitate more nuanced and reliable evaluations using Multimodal Large Language Models (MLLMs). By assessing models’ ability to comprehend and execute novel, complex tasks, IDEA-Bench paves the way toward the development of generative models with autonomous and versatile visual generation capabilities. Lianghua Huang, Jingwu Fang, Huanzhang Dou, Wei Wang 0354, Zhi-Fan Wu, Yupeng Shi, Junge Zhang, Xin Zhao 0012, Yu Liu 0063 |
CVPR | 7 |
| 2024 | Graph Matching-Based Spatiotemporal Calibration of Roadside Sensors in Cooperative Vehicle-Infrastructure SystemsabstractSensors, such as cameras, millimeter-wave radar, and LiDAR, are widely deployed in cooperative vehicle-infrastructure systems. The demand for calibration of initial installation, damage replacements, and unstable installation has risen dramatically. Traditional methods require on-site operation and road closure; thus, repeated calibration can severely affect traffic conditions and expose operational personnel to potential safety threats. As more and more autonomous vehicles (AVs) flood the roads, this paper proposes an automatic calibration framework of roadside sensors by leveraging the high-precision positioning and perception data of AVs. First, we design a graph-based target-matching algorithm using an AV’s surrounding traffic perception data to identify the AV of interest from a dataset of multiple target trajectories recorded by roadside sensors. A line search algorithm is then designed to adjust the clock delay between sensors and establish the temporal correspondence, where a Gaussian process is applied to estimate the vehicle state in continuous time. Finally, we develop a least squares optimization model to complete the final calibration with the AV positioning data. The influence of measurement noise and missed detections on the proposed calibration framework are analyzed in simulated scenarios based on a Next Generation SIMulation (NGSIM) dataset, and the practicability is validated based on real-world data collected at Donghai Bridge, Hangzhou Bay Bridge, and DAIR-V2X dataset. It is shown that the proposed target matching algorithm can identify an AV trajectory from roadside sensor data with 20%-90% higher accuracy than baseline models, and the framework can accurately estimate the spatial and temporal parameters even with poor data quality. The mean least squares error of the trajectory alignment reaches centimeter-level accuracy. Delong Ding, Yupeng Shi, Yuxiong Ji, Yuchuan Du |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | A Rapid and Convenient Spatiotemporal Calibration Method of Roadside Sensors Using Floating Connected and Automated Vehicle DataabstractCameras, millimeter-wave radars, and lidars are widely deployed on smart roads to obtain personalized vehicle trajectories for advanced traffic control and risk avoidance. However, these asynchronous roadside sensors need to be spatiotemporally calibrated accurately before they are put into service. Traditional manual manipulation methods are inefficient and will affect traffic operation and safety. A rapid and convenient method has become essential under the trend that large amounts of roadside sensors need to be tested and calibrated frequently. As more and more connected and automated vehicles (CAVs) flood the smart roads, this paper proposes a novel spatiotemporal calibration framework using the positioning and perception data of CAVs. First, a trajectory matching algorithm is designed using motion feature and point feature histogram sequences as the descriptors, which can determine the approximate spatiotemporal correspondence for the CAV from the roadside trajectory dataset. An optimization method is then formulated to tune transformation parameters through the Gaussian Process trajectory representation and Gauss-Newton algorithms, considering the sampling frequency deviation and measurement noise. Based on numerical analysis via the NGSIM and HighD datasets, it is shown that the proposed calibration method can significantly reduce transformation errors and perform robustly in different scenarios. The feasibility and practicability of the calibration method are further validated through real-world experiments at Tongji University and on the Donghai Bridge in Shanghai, China. This study provides an economical and practical way for spatiotemporal calibration of roadside sensors in an era of CAVs. Yupeng Shi, Yuchuan Du, Shengchuan Jiang, Yuxiong Ji, Xiangmo Zhao |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Gesper: A Unified Framework for General Speech RestorationabstractThis paper describes the legends-tencent team’s real-time General Speech Restoration (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This newly proposed system is a two-stage architecture, in which the speech restoration is performed, and then followed by speech enhancement. We propose a complex spectral mapping-based generative adversarial network (CSM-GAN) as the speech restoration module for the first time. For noise suppression and dereverberation, the enhancement module is presented with fullband-wideband parallel processing. On the blind test set of ICASSP 2023 SSI Challenge, the proposed Gesper system, which satisfies the real-time condition, achieves 3.27 P.804 overall mean opinion score (MOS) and 3.35 P.835 overall MOS, ranked 1st in both track 1 and track 2. Jun Chen 0024, Yupeng Shi, Wei Rao 0002, Shulin He, Andong Li, Yannan Wang, Zhiyong Wu 0001, Shidong Shang, Chengshi Zheng |
ICASSP | 2 |
| 2023 | Gesper: A Restoration-Enhancement Framework for General Speech Reconstruction
Yupeng Shi, Jun Chen 0024, Wei Rao 0002, Shulin He, Andong Li, Yannan Wang, Zhiyong Wu 0001 |
INTERSPEECH | 2 |
| 2023 | Multi-mode Neural Speech Coding Based on Deep Generative Networks
Shan Yang 0001, Yupeng Shi, Yuyong Kang, Dan Su 0002, Shidong Shang, Dong Yu 0001 |
INTERSPEECH | 5 |
| 2022 | Retrieval-based Spatially Adaptive Normalization for Semantic Image SynthesisabstractSemantic image synthesis is a challenging task with many practical applications. Albeit remarkable progress has been made in semantic image synthesis with spatiallyadaptive normalization, existing methods usually normalize the feature activations under the coarse-level guidance (e.g., semantic class). However, different parts of a semantic object (e.g., wheel and window of car) are quite different in structures and textures, making blurry synthesis results usually inevitable due to the missing of fine-grained guidance. In this paper, we propose a novel normalization module, termed as REtrieval-based Spatially Adaptive normaLization (RESAIL), for introducing pixel level fine- grained guidance to the normalization architecture. Specifically, we first present a retrieval paradigm by finding a content patch of the same semantic class from training set with the most similar shape to each test semantic mask. Then, the retrieved patches are composited into retrieval-based guidance, which can be used by RESAIL for pixel level fine-grained modulation on feature activations, thereby greatly mitigating blurry synthesis results. Moreover, distorted ground-truth images are also utilized as alternatives of retrieval-based guidance for feature normalization, further benefiting model training and improving visual quality of generated images. Experiments on several challenging datasets show that our RESAIL performs favorably against state-of-the-arts in terms of quantitative metrics, visual quality, and subjective evaluation. The source code is available at https://github.com/Shi-Yupeng/RESAIL-For-SIS. Yupeng Shi, Xiao Liu 0040, Yuxiang Wei 0001, Zhongqin Wu, Wangmeng Zuo |
CVPR | 1 |
| 2022 | Internet Streaming Audio Based Speech Reception Threshold Measurement in Cochlear Implant UsersabstractTraditional face-to-face subjective listening test has become a challenge due to the COVID-19 pandemic. We developed a remote assessment system with Tencent Meeting, a video conferencing application, to address this issue. This paper presents our work on evaluating the reliability of the remote assessment system. Two speech reception threshold (SRT) experiments were conducted to study the effects of noise suppression and maxima selection number on cochlear implant (CI) hearing. Both experiments were conducted locally and remotely, the correlations between the respective results were analyzed. Results showed that remote tests replicated the differences among testing conditions observed in local tests, but the absolute SRT values for individual conditions varied significantly between the two modes. The variations could be attributed to multiple reasons, such as online data transmission issues, audio playback devices, environmental conditions, and the training of participants. In conclusion, the relative variation of SRTs for CIs can be measured reliably, but the absolute SRT values should be carefully compared and explained according to objective and subjective experimental conditions. Yefei Mo, Kang Ouyang, Mingyue Shi, Huali Zhou, Yupeng Shi, Shidong Shang, Nengheng Zheng |
ICASSP | 6 |
| 2021 | A Noise-Robust Signal Processing Strategy for Cochlear Implants Using Neural NetworksabstractSignal processing strategies in most clinical cochlear implants (CIs) extract and transmit speech envelopes to stimulate the auditory neurons. The incomplete representation of the rich fine structures in speech has significantly degraded the CI recipients’ ability in high- level perception, including their speech understanding in noise. This paper presents a noise-robust signal processing strategy to deal with this problem. Neural networks (NN) are built and trained to simulate the advanced combination encoder (ACE, a strategy for CI products of Cochlear Corporation). The NN-based ACE (namely, NNACE) is trained with a sophisticatedly designed loss function to output envelope-like signals that 1) is compatible with ACE-based CI system and can serve as the modulator to generate the electric stimuli, 2) is more noise-robust, and 3) might bear a certain degree of the temporal fine structures of speech. Subjective and objective evaluations with vocoder simulated speech show that NNACE outperforms the other methods and further actual CI experiments are warranted. Nengheng Zheng, Yupeng Shi, Yuyong Kang |
ICASSP | 2 |
| 2021 | Orthogonal Jacobian Regularization for Unsupervised Disentanglement in Image GenerationabstractUnsupervised disentanglement learning is a crucial issue for understanding and exploiting deep generative models. Recently, SeFa tries to find latent disentangled directions by performing SVD on the first projection of a pretrained GAN. However, it is only applied to the first layer and works in a post-processing way. Hessian Penalty minimizes the off-diagonal entries of the output’s Hessian matrix to facilitate disentanglement, and can be applied to multi-layers. However, it constrains each entry of output independently, making it not sufficient in disentangling the latent directions (e.g., shape, size, rotation, etc.) of spatially correlated variations. In this paper, we propose a simple Orthogonal Jacobian Regularization (OroJaR) to encourage deep generative model to learn disentangled representations. It simply encourages the variation of output caused by perturbations on different latent dimensions to be orthogonal, and the Jacobian with respect to the input is calculated to represent this variation. We show that our OroJaR also encourages the output’s Hessian matrix to be diagonal in an indirect manner. In contrast to the Hessian Penalty, our OroJaR constrains the output in a holistic way, making it very effective in disentangling latent dimensions corresponding to spatially correlated variations. Quantitative and qualitative experimental results show that our method is effective in disentangled and controllable image generation, and performs favorably against the state-of-the-art methods. Our code is available at https://github.com/csyxwei/OroJaR. Yuxiang Wei 0001, Yupeng Shi, Xiao Liu 0040, Zhilong Ji, Zhongqin Wu, Wangmeng Zuo |
ICCV | 2 |