VLDB 2026 Research / reviewers in the wild / expert
Yuchong Gao
dblp:251/8447
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-class segmentation of aortic branches and zones in computed tomography angiography: The AortaSeg24 challenge
Muhammad Imran 0013, Jonathan R. Krebs, Vishal Balaji Sivaraman, Amarjeet Kumar, Walker R. Ueland, Michael J. Fassler, Lisheng Wang, Maximilian Rokuss, Michael Baumgartner 0001, Yannick Kirchhof, Klaus H. Maier-Hein, Fabian Isensee, Shuolin Liu, Bong Thanh Nguyen, Dong-jin Shin, Park Ji-Woo, Matthew Choi, Kwang-Hyun Uhm, Sung-Jea Ko, Chanwoong Lee, Jaehee Chun, Yun Gu, Zhaohong Pan, Xiaokun Liang, Markus Tiefenthaler, Enrique Almar-Munoz, Matthias Schwab, Mikhail Kotyushev, Rostislav Epifanov, Marek Wodzinski, Henning Müller, Abdul Qayyum 0002, Moona Mazher, Steven A. Niederer, Zhiwei Wang 0002, Kaixiang Yang 0004, Jintao Ren, Stine Sofia Korreman, Yuchong Gao, Hongye Zeng, Jinghua Yue, Fugen Zhou, Alexander Cosman, Muxuan Liang, Gilbert R. Upchurch Jr., Yuyin Zhou, Michol A. Cooper, Wei Shao 0008 |
Medical Image Anal. | 50 |
| 2026 | Generative AI for Wireless Communication and Sensing: Toward Unified Foundation Models
Zheng Yang 0002, Guoxuan Chi, Chenshu Wu, Yuchong Gao, Yunhao Liu 0001, Yonina C. Eldar, Jie Xu 0002, Tony Xiao Han |
IEEE Trans. Commun. | 5 |
| 2025 | RF-Prox: Radio-Based Proximity Estimation of Nondirectly Connected DevicesabstractRecent years have witnessed an increasing number of mobile devices, posing a more diversified demand for device localization solutions. Existing methods can locate connected devices but fail to address the spatial proximity between devices lacking direct communication links. This limitation impedes numerous emerging applications, such as implicit control of IoT device and proximity-based autonomous aerial vehicles scheduling. In response to this technical challenge, we introduce RF-Prox, the pioneering system designed for the proximity estimation of nondirectly connected devices. RF-Prox determines the proximity between devices by extracting and analyzing the spatiotemporal correlation between two signals. RF-Prox introduces a multiresolution spatiotemporal encoder (MRSTE) that extracts multiscale features from complex-valued wireless signals, capturing both spatial and dynamic temporal characteristics. Additionally, the proximity metric adaptation network (PMAN) bridges the gap between high-dimensional signal characteristics and physical proximity. To enhance scalability, we leverage a transfer learning framework, significantly reducing the need for extensive data collection and retraining. Extensive experiments demonstrate RF-Prox’s outstanding performance across Wi-Fi and cellular networks, achieving fine-tuned accuracy rates of 98.6% indoors and 91.3% outdoors. Even without fine-tuning, the pretrained model achieves strong zero-shot performance, showcasing its exceptional performance in both proximity estimation accuracy and domain generalizability. Yuchong Gao, Guoxuan Chi, Zheng Yang 0002, Shijie Cheng, Zhiqing Wei |
IEEE Internet Things J. | 1 |
| 2025 | Plugging and Breathing on the Air: A Practical Defense System for Deep Learning-Based Wireless Semantic CommunicationsabstractDeep learning-based semantic communications (DLSC) leverage deep neural networks in transmitters and receivers, pushing the boundaries beyond Shannon limit. However, DLSC is extremely vulnerable to malicious physical-layer adversarial attacks due to the openness of wireless channels. Meanwhile, existing defense approaches still suffer from two challenges for robust DLSC. First, most methods require offline DLSC retraining to defend against various attacks, causing interruptions of online service. Second, they struggle to achieve effective defense in real-world time-varying channels, thus limiting DLSC reliability. We propose PBNet, integrating a pluggable protector and an adaptive protector to respectively address the above two challenges. First, the pluggable protector utilizes a novel denoising module to safeguard the transmitted signals, enabling hot-pluggable deployment without interrupting communication. Second, the adaptive protector leverages a novel alternating adaption strategy to achieve effective defense in time-varying channels, ensuring robust performances under real-world dynamic conditions. Evaluations involving symbols, images, texts, and speeches show the efficacy of our PBNet, which has respectively achieved an impressive 72.22% and 73.71% accuracy improvement in defending against unknown$l_{0}$-norm and$l_{2}$-norm attacks on image-based DLSC. Furthermore, we developed two real-world radio systems of PBNet to perform over-the-air signal generation, integrating hardware and software such as FPGA chips and GNU radio. We also implemented an interactive UI of PBNet based on QT5, aiming to demonstrate the effect of attacks and defense visually. This work achieves robust DLSC performances under various attacks and time-varying channels, taking a significant step towards the practical defense scheme for robust DLSC. Chenyang Qiu 0001, Guoshun Nan, Ruiwen Liang, Wendi Deng, Yuchong Gao, Di Wang 0011, Meng Qu, Zhuoran Duan, Qianlong Sun, Qimei Cui, Xiaodong Xu 0001, Xiaofeng Tao 0001, Tony Q. S. Quek |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | Training-Free Image Style Alignment for Domain Shift on Handheld Ultrasound DevicesabstractHandheld ultrasound devices face usage limitations due to user inexperience and cannot benefit from supervised deep learning without extensive expert annotations. Moreover, the models trained on standard ultrasound device data are constrained by training data distribution and perform poorly when directly applied to handheld device data. In this study, we propose the Training-free Image Style Alignment (TISA) to align the style of handheld device data to those of standard devices. The proposed TISA eliminates the demand for source data, and can transform the image style while preserving spatial context during testing. Furthermore, our TISA avoids continuous updates to the pre-trained model compared to other test-time methods and is suited for clinical applications. We show that TISA performs better and more stably in medical detection and segmentation tasks for handheld device data than other test-time adaptation methods. We further validate TISA as the clinical model for automatic measurements of spinal curvature and carotid intima-media thickness, and the automatic measurements agree well with manual measurements made by human experts. We demonstrate the potential for TISA to facilitate automatic diagnosis on handheld ultrasound devices and expedite their eventual widespread use. Code is available at https://github.com/zenghy96/TISA. Hongye Zeng, Ke Zou, Zhihao Chen 0004, Yuchong Gao, Kang Zhou 0001, Meng Wang 0038, Chang Jiang 0001, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Medical Imaging | 4 |
| 2024 | RoCoSDF: Row-Column Scanned Neural Signed Distance Fields for Freehand 3D Ultrasound Imaging Shape Reconstruction
Yuchong Gao, Shuhang Zhang, Jiangjie Wu, Yuexin Ma |
MICCAI (4) | 2 |
| 2024 | RF-Diffusion: Radio Signal Generation via Time-Frequency DiffusionabstractAlong with AIGC shines in CV and NLP, its potential in the wireless domain has also emerged in recent years. Yet, existing RF-oriented generative solutions are ill-suited for generating high-quality, time-series RF data due to limited representation capabilities. In this work, inspired by the stellar achievements of the diffusion model in CV and NLP, we adapt it to the RF domain and propose RF-Diffusion. To accommodate the unique characteristics of RF signals, we first introduce a novel Time-Frequency Diffusion theory to enhance the original diffusion model, enabling it to tap into the information within the time, frequency, and complex-valued domains of RF signals. On this basis, we propose a Hierarchical Diffusion Transformer to translate the theory into a practical generative DNN through elaborated design spanning network architecture, functional block, and complex-valued operator, making RF-Diffusion a versatile solution to generate diverse, high-quality, and time-series RF data. Performance comparison with three prevalent generative models demonstrates the RF-Diffusion's superior performance in synthesizing Wi-Fi and FMCW signals. We also showcase the versatility of RF-Diffusion in boosting Wi-Fi sensing systems and performing channel estimation in 5G networks. Guoxuan Chi, Zheng Yang 0002, Chenshu Wu, Jingao Xu, Yuchong Gao, Yunhao Liu 0001, Tony Xiao Han |
MobiCom | 5 |
| 2024 | WiViD: Leveraging Wi-Fi and Vision for Depth Estimation via Multimodal DiffusionabstractDepth estimation is crucial for numerous applications, including autonomous driving, robotic navigation and aug-mented reality. Existing solutions based on LiDAR and mm Wave technologies are constrained by high deployment costs, while those utilizing monocular vision suffer from limited accuracy. To address these challenges, this paper proposes WiViD, a diffusion-based depth estimation system that leverages commercial Wi-Fi and vision. Diffusion models, with their ability to iteratively refine predictions, offer significant advantages in producing accurate and detailed estimations. We introduce a Multimodal Conditional Diffusion (MMCD) mechanism and design two encoding modules: the Complex-Valued CSI Encoder (CCE) and the Residual Image Encoder (RIE). These components fully exploit the spatio-temporal information inherent in Wi-Fi CSI and enable the effective fusion of Wi-Fi CSI and RGB image data, which results in high-precision and robust depth estimation. Experimental results in real-world scenarios demonstrate that WiViD out-performs state-of-the-art (SOTA) monocular methods, reducing the Absolute Relative Error (ARE) by 67.2 %, highlighting the advantages of WiViD in terms of accuracy and reliability. Shijie Cheng, Yuchong Gao, Zheng Yang 0002, Guoxuan Chi, Tony Xiao Han |
MSN | 2 |
| 2024 | Anatomical Prior and Inter-Slice Consistency for Semi-Supervised Vertebral Structure Detection in 3D Ultrasound VolumeabstractThree-dimensional (3D) ultrasound imaging technique has been applied for scoliosis assessment, but the current assessment method only uses coronal projection images and cannot illustrate the 3D deformity and vertebra rotation. The vertebra detection is essential to reveal 3D spine information, but the detection task is challenging due to complex data and limited annotations. We propose VertMatch to detect vertebral structures in 3D ultrasound volume containing a detector and classifier. The detector network finds the potential positions of structures on transverse slice globally, and then the local patches are cropped based on detected positions. The classifier is used to distinguish whether the patches contain real vertebral structures and screen the predicted positions from the detector. VertMatch utilizes unlabeled data in a semi-supervised manner, and we develop two novel techniques for semi-supervised learning: 1) anatomical prior is used to acquire high-quality pseudo labels; 2) inter-slice consistency is used to utilize more unlabeled data by inputting multiple adjacent slices. Experimental results demonstrate that VertMatch can detect vertebra accurately in ultrasound volume and outperforms state-of-the-art methods. Moreover, VertMatch is also validated in automatic spinous process angle measurement on forty subjects with scoliosis, and the results illustrate that it can be a promising approach for the 3D assessment of scoliosis. Hongye Zeng, Kang Zhou 0001, Songhan Ge, Yuchong Gao, Jianhao Zhao, Shenghua Gao |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Wi-Prox: Proximity Estimation of Non-Directly Connected Devices via Sim2Real Transfer LearningabstractRecent years have witnessed an increasing number of mobile devices, posing a more diversified demand for device localization solutions. While existing wireless localization solutions can obtain the relative locations of connected devices, they fall short in estimating the spatial relationships between devices that are not directly connected. To address this technical gap, we propose Wi-Prox, the first proximity estimation system for non-directly connected devices. Wi-Prox evaluates the spatial proximity of two devices by analyzing their received wireless signals. It integrates a novel multi-resolution spatial encoder that extracts multi-scale spatial features from complex-valued wireless signals, which are then analyzed and transformed into a domain-adaptive proximity metric. To enhance the general-izability of Wi-Prox, we adopt a simulation-to-reality transfer learning framework. Wi-Prox is pre-trained with a large amount of simulated data and then fine-tuned for real-world deployment, significantly reducing the need for real-world data collection. We implement Wi-Prox and evaluate its performance in both simulated and real environments. Our results indicate that a fine-tuned Wi-Prox achieves an average accuracy of 97.2% in selecting the most proximate device. Even without fine-tuning, a pre-trained Wi-Prox still manages an average accuracy of 93.8%, thereby demonstrating impressive performance in terms of both proximity estimation accuracy and domain generalizability. Yuchong Gao, Guoxuan Chi, Guidong Zhang, Zheng Yang 0002 |
GLOBECOM | 1 |
| 2023 | HACK: Learning a Parametric Head and Neck Model for High-fidelity AnimationabstractSignificant advancements have been made in developing parametric models for digital humans, with various approaches concentrating on parts such as the human body, hand, or face. Nevertheless, connectors such as the neck have been overlooked in these models, with rich anatomical priors often unutilized. In this paper, we introduce HACK (Head-And-neCK), a novel parametric model for constructing the head and cervical region of digital humans. Our model seeks to disentangle the full spectrum of neck and larynx motions, facial expressions, and appearance variations, providing personalized and anatomically consistent controls, particularly for the neck regions. To build our HACK model, we acquire a comprehensive multi-modal dataset of the head and neck under various facial expressions. We employ a 3D ultrasound imaging scheme to extract the inner biomechanical structures, namely the precise 3D rotation information of the seven vertebrae of the cervical spine. We then adopt a multi-view photometric approach to capture the geometry and physically-based textures of diverse subjects, who exhibit a diverse range of static expressions as well as sequential head-and-neck movements. Using the multi-modal dataset, we train the parametric HACK model by separating the 3D head and neck depiction into various shape, pose, expression, and larynx blendshapes from the neutral expression and the rest skeletal pose. We adopt an anatomically-consistent skeletal design for the cervical region, and the expression is linked to facial action units for artist-friendly controls. We also propose to optimize the mapping from the identical shape space to the PCA spaces of personalized blendshapes to augment the pose and expression blendshapes, providing personalized properties within the framework of the generic model. Furthermore, we use larynx blendshapes to accurately control the larynx deformation and force the larynx slicing motions along the vertical direction in the UV-space for precise modeling of the larynx beneath the neck skin. HACK addresses the head and neck as a unified entity, offering more accurate and expressive controls, with a new level of realism, particularly for the neck regions. This approach has significant benefits for numerous applications, including geometric fitting and animation, and enables inter-correlation analysis between head and neck for fine-grained motion synthesis and transfer. Longwen Zhang, Zijun Zhao, Xinzhou Cong, Qixuan Zhang, Shuqi Gu, Yuchong Gao, Wei Yang 0034, Lan Xu 0003, Jingyi Yu 0001 |
ACM Trans. Graph. | 6 |