VLDB 2026 Research / reviewers in the wild / expert
Sitong Cheng
dblp:252/5783
· DBLP profile ↗
6ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0006-6045-4612ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 64% Generative modeling · 36% | |
| Computer networks
1 paper |
Internet of things and sensor networks · 50% Wireless sensing and localization · 50% | |
| Human-computer interaction and pervasive computing
3 papers |
Health and well-being technologies · 84% Wearable and physiological sensing · 16% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
federated learning |
1.4 | 2 | 2024 | ADMarker: A Multi-Modal Federated Learning System for Monitoring Digital Biomarkers of Alzheimer's Disease · MobiCom 2024 Harmony: Heterogeneous Multi-Modal Federated Learning through Disentangled Model Training · MobiSys 2023 |
Machine learning › Efficient and distributed learning › federated learning
multimodal federated learning |
1.4 | 2 | 2024 | ADMarker: A Multi-Modal Federated Learning System for Monitoring Digital Biomarkers of Alzheimer's Disease · MobiCom 2024 Harmony: Heterogeneous Multi-Modal Federated Learning through Disentangled Model Training · MobiSys 2023 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
latent diffusion model |
0.9 | 1 | 2025 | Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation · ICLR 2025 |
Audio and music processing › spatial audio
spatial audio generation |
0.9 | 1 | 2025 | Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation · ICLR 2025 |
Wireless sensing and localization › human activity recognition
activity recognition |
0.9 | 1 | 2025 | AquaScan: A Sonar-based Underwater Sensing System for Human Activity Monitoring · MobiCom 2025 |
Internet of things and sensor networks
underwater sensing |
0.9 | 1 | 2025 | AquaScan: A Sonar-based Underwater Sensing System for Human Activity Monitoring · MobiCom 2025 |
Health and well-being technologies › digital phenotyping
digital biomarkers |
0.8 | 1 | 2024 | ADMarker: A Multi-Modal Federated Learning System for Monitoring Digital Biomarkers of Alzheimer's Disease · MobiCom 2024 |
Machine learning › Efficient and distributed learning › federated learning
heterogeneous federated learning |
0.7 | 1 | 2023 | Harmony: Heterogeneous Multi-Modal Federated Learning through Disentangled Model Training · MobiSys 2023 |
Machine learning › Generative modeling
multimodal generation |
0.3 | 1 | 2025 | Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation · ICLR 2025 |
Health and well-being technologies
health monitoring |
0.2 | 1 | 2023 | Harmony: Heterogeneous Multi-Modal Federated Learning through Disentangled Model Training · MobiSys 2023 |
Methods — techniques the papers use, named apart from their topics
state-transfer-based activity recognition · 1.7spatial-aware encoding · 1.7signal processing · 1.7latent diffusion · 1.7image reconstruction · 1.7azimuth state matrix · 1.7multimodal sensing · 1.5federated learning · 1.5resource allocation · 1.3disentangled model training · 1.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Both Ears Wide Open: Towards Language-Driven Spatial Audio GenerationabstractRecently, diffusion models have achieved great success in mono-channel audio generation.
However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions.
Controlling stereo audio with spatial contexts remains challenging due to high data costs and unstable generative models.
To the best of our knowledge, this work represents the first attempt to address these issues.
We first construct a large-scale, simulation-based, and GPT-assisted dataset, BEWO-1M, with abundant soundscapes and descriptions even including moving and multiple sources.
Beyond text modality, we have also acquired a set of images and rationally paired stereo audios through retrieval to advance multimodal generation.
Existing audio generation models tend to generate rather random and indistinct spatial audio.
To provide accurate guidance for Latent Diffusion Models, we introduce the SpatialSonic model utilizing spatial-aware encoders and azimuth state matrices to reveal reasonable spatial guidance.
By leveraging spatial guidance, our model not only achieves the objective of generating immersive and controllable spatial audio from text but also extends to other modalities as the pioneer attempt.
Finally, under fair settings, we conduct subjective and objective evaluations on simulated and real-world data to compare our approach with prevailing methods.
The results demonstrate the effectiveness of our method, highlighting its capability to generate spatial audio that adheres to physical rules. Peiwen Sun, Sitong Cheng, Xiangtai Li, Zhen Ye 0006, Huadai Liu, Honggang Zhang 0002, Wei Xue 0002, Yike Guo |
ICLR | 2 |
| 2025 | AquaScan: A Sonar-based Underwater Sensing System for Human Activity MonitoringabstractHuman activity monitoring in the water is essential for pool management and drowning prevention. Existing camera-based solutions pose significant concerns about privacy and extra installation costs. Although sonars have been widely used for underwater sensing in open aquatic environments such as oceans and lakes, monitoring human activities with sonars in a pool setup is challenging. In this work, we propose AquaScan, the first scanning sonar-based underwater sensing system for human activity monitoring. To overcome the low frame rate, we propose a novel scanning strategy and apply an image reconstruction method to accelerate the scanning speed without compromising the performance of motion detection. We develop a novel signal processing pipeline based on a physical model to remove noises and localize human subjects. We extract features like motion, time, and spatial information from sonar images and develop a state-transfer-based activity recognition system to recognize five common water activities. We deployed AquaScan on three public swimming pools for a total period of 94 hours. The evaluation results show that AquaScan can successfully recognize the five activities in the water at about 91.5%. Haozheng Hou, Sitong Cheng, Xiaoguang Zhao, Peiheng Wu, Lixing He, Yunqi Guo, Guoliang Xing, Zhenyu Yan 0002 |
MobiCom | 3 |
| 2024 | ADMarker: A Multi-Modal Federated Learning System for Monitoring Digital Biomarkers of Alzheimer's DiseaseabstractAlzheimer's Disease (AD) and related dementia are a growing global health challenge due to the aging population. In this paper, we present ADMarker, the first end-to-end system that integrates multi-modal sensors and new federated learning algorithms for detecting multidimensional AD digital biomarkers in natural living environments. ADMarker features a novel three-stage multi-modal federated learning architecture that can accurately detect digital biomarkers in a privacy-preserving manner. Our approach collectively addresses several major real-world challenges, such as limited data labels, data heterogeneity, and limited computing resources. We built a compact multi-modality hardware system and deployed it in a four-week clinical trial involving 91 elderly participants. The results indicate that ADMarker can accurately detect a comprehensive set of digital biomarkers with up to 93.8% accuracy and identify early AD with an average of 88.9% accuracy. ADMarker offers a new platform that can allow AD clinicians to characterize and track the complex correlation between multidimensional interpretable digital biomarkers, demographic factors of patients, and AD diagnosis in a longitudinal manner. Xiaomin Ouyang, Xian Shuai, Yang Li 0147, Li Pan 0004, Xifan Zhang, Heming Fu, Sitong Cheng, Xinyan Wang 0003, Shihua Cao, Jiang Xin, Hazel Mok, Zhenyu Yan 0002, Doris Sau-Fung Yu, Timothy Kwok, Guoliang Xing |
MobiCom | 7 |
| 2023 | Harmony: Heterogeneous Multi-Modal Federated Learning through Disentangled Model TrainingabstractMulti-modal sensing systems are increasingly prevalent in real-world applications such as health monitoring and autonomous driving. Most multi-modal learning approaches need to access users' raw data, which poses significant concerns to users' privacy. Federated learning (FL) provides a privacy-aware distributed learning framework. However, current FL approaches have not addressed the unique challenges of heterogeneous multi-modal FL systems, such as modality heterogeneity and significantly longer training delay. In this paper, we propose Harmony, a new system for heterogeneous multi-modal federated learning. Harmony disentangles the multi-modal network training in a novel two-stage framework, namely modality-wise federated learning and federated fusion learning. By integrating a novel balance-aware resource allocation mechanism in modality-wise FL and exploiting modality biases in federated fusion learning, Harmony improves the model accuracy under non-i.i.d. data distributions and speeds up system convergence. We implemented Harmony on a real-world multi-modal sensor testbed deployed in the homes of 16 elderly subjects for Alzheimer's Disease monitoring. Our evaluation on the testbed and three large-scale public datasets of different applications show that, Harmony outperforms by up to 46.35% accuracy over state-of-the-art baselines and saves up to 30% training delay. Xiaomin Ouyang, Heming Fu, Sitong Cheng, Li Pan 0004, Neiwen Ling, Guoliang Xing, Jianwei Huang 0001 |
MobiSys | 4 |
| 2020 | CN-Celeb: A Challenging Chinese Speaker Recognition DatasetabstractRecently, researchers set an ambitious goal of conducting speaker recognition in unconstrained conditions where the variations on ambient, channel and emotion could be arbitrary. However, most publicly available datasets are collected under constrained environments, i.e., with little noise and limited channel variation. These datasets tend to deliver over-optimistic performance and do not meet the request of research on speaker recognition in unconstrained conditions.In this paper, we present CN-Celeb, a large-scale speaker recognition dataset collected ‘in the wild’. This dataset contains more than 130,000 utterances from 1,000 Chinese celebrities, and covers 11 different genres in real world. Experiments conducted with two state-of-the-art speaker recognition approaches (i-vector and x-vector) show that the performance on CN-Celeb is far inferior to the one obtained on Vox-Celeb, a widely used speaker recognition dataset. This result demonstrates that in real-life conditions, the performance of existing techniques might be much worse than it was thought. Our database is free for researchers and can be downloaded from http://project.cslt.org. Jiawen Kang 0002, Lantian Li, Kaicheng Li, Sitong Cheng, Pengyuan Zhang, Ziya Zhou, Yunqi Cai, Dong Wang 0013 |
ICASSP | 6 |
| 2020 | ASR-Free Pronunciation AssessmentabstractMost of the pronunciation assessment methods are based on local features derived from automatic speech recognition (ASR), e.g., the Goodness of Pronunciation (GOP) score. In this paper, we investigate an ASR-free scoring approach that is derived from the marginal distribution of raw speech signals. The hypothesis is that even if we have no knowledge of the language (so cannot recognize the phones/words), we can still tell how good a pronunciation is, by comparatively listening to some speech data from the target language. Our analysis shows that this new scoring approach provides an interesting correction for the phone-competition problem of GOP. Experimental results on the ERJ dataset demonstrated that combining the ASR-free score and GOP can achieve better performance than the GOP baseline. Sitong Cheng, Lantian Li, Zhiyuan Tang, Dong Wang 0013, Thomas Fang Zheng |
INTERSPEECH | 1 |