EDBT 2026 Demo / reviewers in the wild / expert
Shihao Chen
dblp:61/7845
· DBLP profile ↗
14ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-7646-8003ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 3D-DCASphereNet: 3D dynamic convolutional attention network with spherical representation for high heterogeneity in lung nodule detection
Jingjing Yu 0001, Hanchao Wang, Yiling Wen, Shihao Chen, Yuetong An, Xiaheng Lu, Xuelei He |
Expert Syst. Appl. | 4 |
| 2025 | CSSinger: End-to-End Chunkwise Streaming Singing Voice Synthesis System Based on Conditional Variational AutoencoderabstractSinging Voice Synthesis (SVS) aims to generate singing voices of high fidelity and expressiveness. Conventional SVS systems usually utilize an acoustic model to transform a music score into acoustic features, followed by a vocoder to reconstruct the singing voice. It was recently shown that end-to-end modeling is effective in the fields of SVS and Text to Speech (TTS). In this work, we thus present a fully end-to-end SVS method together with a chunkwise streaming inference to address the latency issue for practical usages. Note that this is the first attempt to fully implement end-to-end streaming audio synthesis using latent representations in VAE. We have made specific improvements to enhance the performance of streaming SVS using latent representations. Experimental results demonstrate that the proposed method achieves synthesized audio with high expressiveness and pitch accuracy in both streaming SVS and TTS tasks. Jianwei Cui 0003, Shihao Chen, Jie Zhang 0042, Li-Rong Dai 0001 |
AAAI | 3 |
| 2025 | Sinba: Singing-To-Accompaniment Generation With Pitch Guidance Via Mamba-Based Language ModelabstractIn this paper, we propose Sinba, a system that can directly generate corresponding background accompaniment music from vocal input, allowing users to create complete songs using only sung vocals. Sinba adopts a decoder-only backbone network architecture. We utilize the Mamba model, which is a linear-time sequence modeling method with selective state spaces and has been proven to achieve more advanced performance than Transformers as a foundation model in long-sequence modeling tasks. However, the Mamba was initially applied to audio tasks by pre-training directly on raw audio waveform samples as the backbone model. In this paper, we convert both the training targets and inputs into discretized tokens for direct training. We also extract pitch information from the vocal input as an additional feature for the model. The proposed model is trained using source-separated data pairs. Subjective and objective experimental results demonstrate that the proposed model can generate high-quality accompaniment that matches the style and rhythm of the vocal input, outperforming the Transformerbased baseline. Synthesized audio samples are available at: https://sounddemos.github.io/sinba. Jianwei Cui 0003, Shihao Chen, Jie Zhang 0042, Chengxing Li, Shan Yang 0001, Li-Rong Dai 0001 |
ASRU | 2 |
| 2025 | A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues PreservationabstractBinaural speech enhancement (BSE) aims to jointly improve the speech quality and intelligibility of noisy signals received by hearing devices and preserve the spatial cues of the target for natural listening. Existing methods often suffer from the compromise between noise reduction (NR) capacity and spatial cues preservation (SCP) accuracy and a high computational demand in complex acoustic scenes. In this work, we present a learning-based lightweight binaural complex convolutional network (LBCCN), which excels in NR by filtering low-frequency bands and keeping the rest. Additionally, our approach explicitly incorporates the estimation of interchannel relative acoustic transfer function to ensure the spatial cues fidelity and speech clarity. Results show that the proposed LBCCN can achieve a comparable NR performance to state-of-the-art methods under fixed-speaker conditions, but with a much lower computational cost and a certain degree of SCP capability. The reproducible code and audio examples are available at https://github.com/jywanng/LBCCN. Shihao Chen |
ICASSP | 3 |
| 2025 | Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMsabstractRecent advancements in multimodal large language models (MLLM) have shown a strong ability in visual perception, reasoning abilities, and vision-language understanding. However, the visual matching ability of MLLMs is rarely studied, despite finding the visual correspondence of objects is essential in computer vision. Our research reveals that the matching capabilities in recent MLLMs still exhibit systematic shortcomings, even with current strong MLLMs models, GPT-4o. In particular, we construct a Multimodal Visual Matching (MMVM) benchmark to fairly benchmark over 30 different MLLMs. The MMVM benchmark is built from 15 open-source datasets and Internet videos with manual annotation. We categorize the data samples of MMVM benchmark into eight aspects based on the required cues and capabilities to more comprehensively evaluate and analyze current MLLMs. In addition, we have designed an automatic annotation pipeline to generate the MMVM SFT dataset, including 220K visual matching data with reasoning annotation. To our knowledge, this is the first visual corresponding dataset and benchmark for the MLLM community. Finally, we present CoLVA, a novel contrastive MLLM with two novel technical designs: fine-grained vision expert with object-level contrastive learning and instruction augmentation strategy. The former learns instance discriminative tokens, while the latter further improves instruction following ability. CoLVA-InternVL2-4B achieves an overall accuracy (OA) of 49.80\% on the MMVM benchmark, surpassing GPT-4o and the best open-source MLLM, Qwen2VL-72B, by 7.15\% and 11.72\% OA, respectively. These results demonstrate the effectiveness of our MMVM SFT dataset and our novel technical designs. Code, benchmark, dataset, and models will be released. Yikang Zhou, Tao Zhang 0042, Shilin Xu 0001, Shihao Chen, Qianyu Zhou 0001, Yunhai Tong, Shunping Ji, Jiangning Zhang, Lu Qi 0001, Xiangtai Li |
ICCV | 4 |
| 2024 | STPose: 6D object pose estimation network based on sparse attention and cross-layer connection
Shihao Chen, Keduo Yan, Yong Li 0028, Dongxu Gao |
BMVC | 1 |
| 2024 | A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech RecognitionabstractAdvanced Audio- Visual Speech Recognition (AVSR) sys-tems have been observed to be sensitive to missing video frames, performing even worse than single-modality mod-els. While applying the common dropout techniques to the video modality enhances robustness to missing frames, it simultaneously results in a performance loss when dealing with complete data input. In this study, we delve into this contrasting phenomenon through the lens of modality bias and uncover that an excessive modality bias towards the audio modality induced by dropout constitutes the fun-damental cause. Next, we present the Modality Bias Hy-pothesis (MBH) to systematically describe the relationship between the modality bias and the robustness against missing modality in multimodal systems. Building on these findings, we propose a novel Multimodal Distribution Approxi-mation with Knowledge Distillation (MDA-KD)framework to reduce over-reliance on the audio modality, maintaining performance and robustness simultaneously. Finally, to address an entirely missing modality, we adopt adapters to dynamically switch decision strategies. The effective-ness of our proposed approach is evaluated through comprehensive experiments on the MISP2021 and MISP2022 datasets. Our code is available at https://github.com/dalision/ModalBiasAV5R. Yusheng Dai, Hang Chen 0001, Jun Du 0002, Ruoyu Wang 0029, Shihao Chen, Chin-Hui Lee 0001 |
CVPR | 5 |
| 2024 | Adversarial Speech for Voice Privacy Protection from Personalized Speech GenerationabstractThe rapid progress in personalized speech generation technology, including personalized text-to-speech (TTS) and voice conversion (VC), poses a challenge in distinguishing between generated and real speech for human listeners, resulting in an urgent demand in protecting speakers' voices from malicious misuse. In this regard, we propose a speaker protection method based on adversarial attacks. The proposed method perturbs speech signals by minimally altering the original speech while rendering downstream speech generation models unable to accurately generate the voice of the target speaker. For validation, we employ the open-source pre-trained YourTTS model for speech generation and protect the target speaker's speech in the white-box scenario. Automatic speaker verification (ASV) evaluations were carried out on the generated speech as the assessment of the voice protection capability. Our experimental results show that we successfully perturbed the speaker encoder of the YourTTS model using the gradient-based I-FGSM adversarial perturbation method. Furthermore, the adversarial perturbation is effective in preventing the YourTTS model from generating the speech of the target speaker. Audio samples can be found in https://voiceprivacy.github.io/Adeversarial-Speech-with-YourTTS. Shihao Chen, Jie Zhang 0042, Kong-Aik Lee, Zhen-Hua Ling, Li-Rong Dai 0001 |
ICASSP | 1 |
| 2024 | A Study of Multichannel Spatiotemporal Features and Knowledge Distillation on Robust Target Speaker ExtractionabstractTarget speaker extraction (TSE) based on direction of arrival (DOA) has a wide range of applications in e.g., remote conferencing, hearing aids, in-car speech interaction. Due to the inherent phase uncertainty, existing TSE methods usually suffer from speaker confusion within specific frequency bands. Imprecise DOA measurements caused by e.g., the calibration of the microphone array and ambient noises, can also deteriorate the TSE performance. In order to improve the robustness of TSE, in this work we propose several new multichannel spatiotemporal features to represent the discriminability of the target speaker. The narrow-band Conformer model is applied in combination with the proposed features to facilitate the extraction of the target speaker. In addition, we consider knowledge distillation for improving the model robustness, particularly in the presence of DOA mis-match. Experimental results on a public dataset verify the efficacy of the proposed method. Yichi Wang 0001, Jie Zhang 0042, Shihao Chen, Weitai Zhang, Zhongyi Ye, Xinyuan Zhou, Li-Rong Dai 0001 |
ICASSP | 3 |
| 2024 | LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
Shihao Chen, Jie Zhang 0042, Rilin Chen, Li-Rong Dai 0001 |
INTERSPEECH | 1 |
| 2024 | Intelligent Energy-Efficient and Fair Resource Scheduling for UAV-Assisted Space-Air-Ground Integrated Networks Under Jamming AttacksabstractThe space-air-ground integrated network (SAGIN) is a crucial technology for sixth-generation (6G) wireless communication networks to achieve seamless coverage and high throughput. In this paper, we propose an unmanned aerial vehicle (UAV)-assisted SAGIN structure, where the UAV is responsible for collecting data from ground users (GUs) and transmitting it to low-earth orbit (LEO) satellites. This paper also formulates a joint energy-efficient and fair resource scheduling optimization problem under jamming attacks and limited energy constraints, where the line-of-sight (LoS) links between the UAV and GUs are susceptible to being jammed. Due to the non-convex problem and dynamic environments, a deep reinforcement learning (DRL)-based twin delayed deep deterministic policy gradient (TD3) is developed to search optimal UAV trajectory to maximize energy efficiency (EE) and fairness against jamming. Simulation results verify that the proposed intelligent resource scheduling algorithm outperforms the baseline algorithms in terms of EE and fairness index in different settings. Shihao Chen, Helin Yang, Liang Xiao 0003, Changyuan Xu, Xianzhong Xie, Zehui Xiong |
VTC Spring | 1 |
| 2020 | NB-IoT Estrus Detection System of Dairy Cows Based on LSTM NetworksabstractTo improve the revenue of dairy farms, cow estrus must be accurately monitored to track mating time. Narrow Band Internet of Things (NB-IoT) is considered as a promising technology to realize cost-effective detection system attributing to its wide coverage and low power consumption. To increase the success rate of real-time detection, machine learning based algorithms have been applied to extract patterns from estrus data. However, due to the lack of multivariate time series data, most previous studies do not consider using the time correlation to guide estrus detection. In this paper, we present a NB-IoT based solution framework where multivariate behavioral time series data collected by neck-mounted sensors and then uploaded to a cloud data center for further analysis through NB-IoT network. Based on the collected data, we propose an estrus prediction algorithm which gives estrus alert by exploiting Long-Short Term Memory (LSTM) and Convolution Neural Network (CNN). Through numerical studies conducted using real data set from the pasture, we show that our proposed solution outperforms exiting detection algorithms in terms of accuracy and efficiency. Shihao Chen, Baoling Liu |
PIMRC | 3 |
| 2018 | Reproducible Interference-Aware Mobile TestingabstractMobile apps are born to work in an environment with ever-changing network connectivity, random hardware interruption, unanticipated task switches, etc. However, such interference cases are often oblivious in traditional mobile testing but happen frequently and sophisticatedly in the field, causing various robustness, responsiveness and consistency problems. In this paper, we propose JazzDroid to introduce interference to mobile testing. JazzDroid adopts a gray-box approach to instrument apps at binary level such that interference logic is inlined with app execution and can be triggered to effectively affect normal execution. Then, JazzDroid repeatedly orchestrates the instrumented app through app developers' existing tests and continuously randomizes interference on the fly to reveal possible faulty executions. Upon discovering problems, JazzDroid generates a test script with the user inputs from developers' tests and the interference injected for developers to reproduce the problems. At a high level, JazzDroid can be seamlessly integrated into app developers' testing procedures, detecting more problems from existing tests. We implement JazzDroid to function on unmodified apps directly from app markets and interface with de facto industrial testing toolchain. JazzDroid improves mobile testing by discovering 6x more problems, including crashes, functional bugs, UI consistency issues and common bug patterns that fail numerous apps. Weilun Xiong, Shihao Chen, Mingyuan Xia 0001, Zhengwei Qi |
ICSME | 2 |
| 2009 | Interactive Inner Structures Visualization in 3D DatasetsabstractThe simultaneous use of images obtained from different sources is common in medical diagnosis. But how to explore the information contents of these data sets is a problem. Real-time methods to directly convey the information contents of large volumetric scalar fields are still a challenge to the computer graphics, especially to large volumetric data sets. So in this paper a GPU based inner structures visualization method is proposed, which on the one hand for the interactive visual exploration of the data, and on the other hand for the generation of high-quality visual representations. The method takes as voxel models and constructs 3D textures, which exploits hardware-assisted texture mapping, re-samples volume data, represented as a stack of 3D texture, onto a sampling surface or so called proxy geometry. For each texel of a slice it perform a fetch to each 3D texture and performs fusion and shading using a fragment shader. The method is very fast and versatile and can provide a good insight into volume dataset. Shihao Chen, Chongyang Hao |
ICIG | 1 |