EDBT 2026 Demo / reviewers in the wild / expert
Boshi Tang
dblp:358/6747
· DBLP profile ↗
11ranked-venue papers
1as first author
11since 2021 · last 2025
0009-0007-6786-4806ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | pFedGPA: Diffusion-based Generative Parameter Aggregation for Personalized Federated LearningabstractFederated Learning (FL) offers a decentralized approach to model training, where data remains local and only model parameters are shared between the clients and the central server. Traditional methods, such as Federated Averaging (FedAvg), linearly aggregate these parameters which are usually trained on heterogeneous data distributions, potentially overlooking the complex, high-dimensional nature of the parameter space. This can result in degraded performance of the aggregated model. While personalized FL approaches can mitigate the heterogeneous data issue to some extent, the limitation of linear aggregation remains unresolved. To alleviate this issue, we investigate the generative approach of diffusion model and propose a novel generative parameter aggregation framework for personalized FL, pFedGPA. In this framework, we deploy a diffusion model on the server to integrate the diverse parameter distributions and propose a parameter inversion method to efficiently generate a set of personalized parameters for each client. This inversion method transforms the uploaded parameters into a latent code, which is then aggregated through denoising sampling to produce the final personalized parameters. By encoding the dependence of a client's model parameters on the specific data distribution using the high-capacity diffusion model, pFedGPA can effectively decouple the complexity of the overall distribution of all clients' model parameters from the complexity of each individual client's parameter distribution. Our experimental results consistently demonstrate the superior performance of the proposed method across multiple datasets, surpassing baseline approaches. Jiahao Lai, Jiaqi Li 0028, Jian Xu 0016, Yanru Wu, Boshi Tang, Wenbo Ding 0001, Yang Li 0104 |
AAAI | 5 |
| 2025 | AdaMesh: Personalized Facial Expressions and Head Poses for Adaptive Speech-Driven 3D Facial AnimationabstractSpeech-driven 3D facial animation aims at generating facial movements that are synchronized with the driving speech, which has been widely explored recently. Existing works mostly neglect the person-specific talking style in generation, including facial expression and head pose styles. Several works intend to capture the personalities by fine-tuning modules. However, limited training data leads to the lack of vividness. In this work, we proposeAdaMesh, a novel adaptive speech-driven facial animation approach, which learns the personalized talking style from a reference video of about 10 seconds and generates vivid facial expressions and head poses. Specifically, we propose mixture-of-low-rank adaptation (MoLoRA) to fine-tune the expression adapter, which efficiently captures the facial expression style. For the personalized pose style, we propose a pose adapter by building a discrete pose prior and retrieving the appropriate style embedding with a semantic-aware pose style matrix without fine-tuning. Extensive experimental results show that our approach outperforms state-of-the-art methods, preserves the talking style in the reference video, and generates vivid facial animation. Liyang Chen, Weihong Bao, Shun Lei, Boshi Tang, Zhiyong Wu 0001, Shiyin Kang, Hao-Zhi Huang 0001, Helen M. Meng |
IEEE Trans. Multim. | 4 |
| 2024 | SimCalib: Graph Neural Network Calibration Based on Similarity between NodesabstractGraph neural networks (GNNs) have exhibited impressive performance in modeling graph data as exemplified in various applications. Recently, the GNN calibration problem has attracted increasing attention, especially in cost-sensitive scenarios. Previous work has gained empirical insights on the issue, and devised effective approaches for it, but theoretical supports still fall short. In this work, we shed light on the relationship between GNN calibration and nodewise similarity via theoretical analysis. A novel calibration framework, named SimCalib, is accordingly proposed to consider similarity between nodes at global and local levels. At the global level, the Mahalanobis distance between the current node and class prototypes is integrated to implicitly consider similarity between the current node and all nodes in the same class. At the local level, the similarity of node representation movement dynamics, quantified by nodewise homophily and relative degree, is considered. Informed about the application of nodewise movement patterns in analyzing nodewise behavior on the over-smoothing problem, we empirically present a possible relationship between over-smoothing and GNN calibration problem. Experimentally, we discover a correlation between nodewise similarity and model calibration improvement, in alignment with our theoretical results. Additionally, we conduct extensive experiments investigating different design factors and demonstrate the effectiveness of our proposed SimCalib framework for GNN calibration by achieving state-of-the-art performance on 14 out of 16 benchmarks. Boshi Tang, Zhiyong Wu 0001, Xixin Wu, Qiaochu Huang, Jun Chen 0024, Shun Lei, Helen M. Meng |
AAAI | 1 |
| 2024 | Explore 3D Dance Generation via Reward Model from Automatically-Ranked DemonstrationsabstractThis paper presents an Exploratory 3D Dance generation framework, E3D2, designed to address the exploration capability deficiency in existing music-conditioned 3D dance generation models. Current models often generate monotonous and simplistic dance sequences that misalign with human preferences because they lack exploration capabilities.The E3D2 framework involves a reward model trained from automatically-ranked dance demonstrations, which then guides the reinforcement learning process. This approach encourages the agent to explore and generate high quality and diverse dance movement sequences. The soundness of the reward model is both theoretically and experimentally validated. Empirical experiments demonstrate the effectiveness of E3D2 on the AIST++ dataset. Zilin Wang 0002, Haolin Zhuang, Yinmin Zhang, Junjie Zhong, Jun Chen 0024, Yu Yang 0016, Boshi Tang, Zhiyong Wu 0001 |
AAAI | 8 |
| 2024 | Unsupervised Human Activity Recognition Via Large Language Models and Iterative EvolutionabstractHuman activity recognition (HAR) is crucial for health monitoring and disease diagnosis in Internet-of-Things environments. However, existing HAR approaches either suffer from poor accuracy or achieve high accuracy at the expense of costly manual annotations. To overcome the challenge above, we propose a novel method named LLMIE-UHAR that that leverages LLMs and Iterative Evolution to realize Unsupervised HAR. Specifically, with our designed prompt engineering mechanism, we employ large language models to fuse both contextual and semantic information, and annotate key samples selected by a clustering algorithm. Moreover, LLMIE-UHAR enhances the recognition accuracy with iterative evolution of clustering algorithm, large language models and the neural network based recognition model. Experiments conducted on the public ARAS datasets show the efficiency of our method, achieving an accuracy of 96.00%. This highlights the practical value of our approach. Jiayuan Gao, Yingwei Zhang 0002, Yiqiang Chen 0001, Tengxiang Zhang, Boshi Tang |
ICASSP | 5 |
| 2024 | Enhancing Expressiveness in Dance Generation Via Integrating Frequency and Music Style InformationabstractDance generation, as a branch of human motion generation, has attracted increasing attention. Recently, a few works attempt to enhance dance expressiveness, which includes genre matching, beat alignment, and dance dynamics, from certain aspects. However, the enhancement is quite limited as they lack comprehensive consideration of the aforementioned three factors. In this paper, we propose ExpressiveBailando, a novel dance generation method designed to generate expressive dances, concurrently taking all three factors into account. Specifically, we mitigate the issue of speed homogenization by incorporating frequency information into VQ-VAE, thus improving dance dynamics. Additionally, we integrate music style information by extracting genre- and beat-related features with a pre-trained music model, hence achieving improvements in the other two factors. Extensive experimental results demonstrate that our proposed method can generate dances with high expressiveness and outperforms existing methods both qualitatively and quantitatively1. Qiaochu Huang, Boshi Tang, Haolin Zhuang, Liyang Chen, Shuochen Gao, Zhiyong Wu 0001, Haozhi Huang 0004, Helen M. Meng |
ICASSP | 3 |
| 2024 | Multi-View Midivae: Fusing Track- and Bar-View Representations for Long Multi-Track Symbolic Music GenerationabstractVariational Autoencoders (VAEs) constitute a crucial component of neural symbolic music generation, among which some works have yielded outstanding results and attracted considerable attention. Nevertheless, previous VAEs still encounter issues with overly long feature sequences and generated results lack contextual coherence, thus the challenge of modeling long multi-track symbolic music still remains unaddressed. To this end, we propose Multi-view MidiVAE, as one of the pioneers in VAE methods that effectively model and generate long multi-track symbolic music. The Multi-view MidiVAE utilizes the two-dimensional (2-D) representation, OctupleMIDI, to capture relationships among notes while reducing the feature sequences length. Moreover, we focus on instrumental characteristics and harmony as well as global and local information about the musical composition by employing a hybrid variational encoding-decoding strategy to integrate both Track- and Bar-view MidiVAE features. Objective and subjective experimental results on the CocoChorales dataset demonstrate that, compared to the baseline, Multi-view MidiVAE exhibits significant improvements in terms of modeling long multi-track symbolic music. Jun Chen 0024, Boshi Tang, Binzhu Sha, Yaolong Ju, Shiyin Kang, Zhiyong Wu 0001, Helen M. Meng |
ICASSP | 3 |
| 2024 | DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D GenerationabstractText-to-image diffusion models pre-trained on billions of image-text pairs have recently enabled 3D content creation by optimizing a randomly initialized differentiable 3D representation with score distillation. However, the optimization process suffers slow convergence and the resultant 3D models often exhibit two limitations: (a) quality concerns such as missing attributes and distorted shape and texture; (b) extremely low diversity comparing to text-guided image synthesis. In this paper, we show that the conflict between the 3D optimization process and uniform timestep sampling in score distillation is the main reason for these limitations. To resolve this conflict, we propose to prioritize timestep sampling with monotonically non-increasing functions, which aligns the 3D optimization process with the sampling process of diffusion model. Extensive experiments show that our simple redesign significantly improves 3D content creation with faster convergence, better quality and diversity. Yukai Shi, Boshi Tang, Xianbiao Qi, Lei Zhang 0001 |
ICLR | 4 |
| 2024 | TOSS: High-quality Text-guided Novel View Synthesis from a Single ImageabstractIn this paper, we present TOSS, which introduces text to the task of novel view synthesis (NVS) from just a single RGB image.
While Zero123 has demonstrated impressive zero-shot open-set NVS capabilities, it treats NVS as a pure image-to-image translation problem. This approach suffers from the challengingly under-constrained nature of single-view NVS: the process lacks means of explicit user control and often result in implausible NVS generations.
To address this limitation, TOSS uses text as high-level semantic information to constrain the NVS solution space.
TOSS fine-tunes text-to-image Stable Diffusion pre-trained on large-scale text-image pairs and introduces modules specifically tailored to image and camera pose conditioning, as well as dedicated training for pose correctness and preservation of fine details.
Comprehensive experiments are conducted with results showing that our proposed TOSS outperforms Zero123 with higher-quality NVS results and faster convergence. We further support these results with comprehensive ablations that underscore the effectiveness and potential of
the introduced semantic guidance and architecture design. Yukai Shi, He Cao, Boshi Tang, Xianbiao Qi, Tianyu Yang 0003, Shilong Liu 0004, Lei Zhang 0001, Harry Shum |
ICLR | 4 |
| 2024 | An End-to-End Approach for Chord-Conditioned Song Generation
Shuochen Gao, Shun Lei, Fan Zhuo, Boshi Tang, Qiaochu Huang, Shiyin Kang |
INTERSPEECH | 6 |
| 2024 | SongCreator: Lyrics-based Universal Song GenerationabstractMusic is an integral part of human culture, embodying human intelligence and creativity, of which songs compose an essential part. While various aspects of song generation have been explored by previous works, such as singing voice, vocal composition and instrumental arrangement, etc., generating songs with both vocals and accompaniment given lyrics remains a significant challenge, hindering the application of music generation models in the real world. In this light, we propose SongCreator, a song-generation system designed to tackle this challenge. The model features two novel designs: a meticulously designed dual-sequence language model (DSLM) to capture the information of vocals and accompaniment for song generation, and a series of attention mask strategies for DSLM, which allows our model to understand, generate and edit songs, making it suitable for various songrelated generation tasks by utilizing specific attention masks. Extensive experiments demonstrate the effectiveness of SongCreator by achieving state-of-the-art or competitive performances on all eight tasks. Notably, it surpasses previous works by a large margin in lyrics-to-song and lyrics-to-vocals. Additionally, it is able to independently control the acoustic conditions of the vocals and accompaniment in the generated song through different audio prompts, exhibiting its potential applicability. Our samples are available at https://thuhcsi.github.io/SongCreator/. Shun Lei, Yixuan Zhou 0002, Boshi Tang, Max W. Y. Lam, Jingcheng Wu, Shiyin Kang, Zhiyong Wu 0001, Helen M. Meng |
NeurIPS | 3 |