EDBT 2026 Demo / reviewers in the wild / expert
Hua Hua
dblp:45/3555
· DBLP profile ↗
15ranked-venue papers
4as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Theory of computation · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiffGraph: Heterogeneous Graph Diffusion ModelabstractRecent advances in Graph Neural Networks (GNNs) have revolutionized graph-structured data modeling, yet traditional GNNs struggle with complex heterogeneous structures prevalent in real-world scenarios. Despite progress in handling heterogeneous interactions, two fundamental challenges persist: noisy data significantly compromising embedding quality and learning performance, and existing methods' inability to capture intricate semantic transitions among heterogeneous relations, which impacts downstream predictions. To address these fundamental issues, we present the Heterogeneous Graph Diffusion Model (DiffGraph), a pioneering framework that introduces an innovative cross-view denoising strategy. This advanced approach transforms auxiliary heterogeneous data into target semantic spaces, enabling precise distillation of task-relevant information. At its core, DiffGraph features a sophisticated latent heterogeneous graph diffusion mechanism, implementing a novel forward and backward diffusion process for superior noise management. This methodology achieves simultaneous heterogeneous graph denoising and cross-type transition, while significantly simplifying graph generation through its latent-space diffusion capabilities. Through rigorous experimental validation on both public and industrial datasets, we demonstrate that DiffGraph consistently surpasses existing methods in link prediction and node classification tasks, establishing new benchmarks for robustness and efficiency in heterogeneous graph processing. The model implementation is publicly available at: https://github.com/HKUDS/DiffGraph. Zongwei Li 0004, Lianghao Xia, Hua Hua, Shuangyang Wang, Chao Huang 0001 |
WSDM | 3 |
| 2025 | Flexpéro: Flexible Expressive Zero-Shot Speech Refinement via In-Context LearningabstractControlling speech expressiveness has emerged as a critical research frontier in speech generation, focusing on synthesizing natural, human-like speech that accurately conveys intended psychological and emotional states. While many large-scale models have demonstrated sufficient zero-shot capability by conditioning the acoustic model on reference speech—which provides cues on speaker identity and style—they often fall short of meeting desired emotional or prosodic targets at fine-grained levels. To address this challenge, we propose a novel speech refinement method based on a zero-shot voice synthesis model that can flexibly and interactively enhance expressiveness on unsatisfactory speech segments. It supports emotion modulation through chunk-wise valence/arousal and flexible keyframe-based prediction of pitch and energy, allowing for the creation of any prosodic patterns. Experimental results show that our method achieves fine-grained control, thus enriching the expressiveness of zero-shot synthetic speech. Hua Hua, Zengqiang Shang, Xuyuan Li, Pengyuan Zhang |
IEEE Signal Process. Lett. | 1 |
| 2024 | A Bayesian Multi-Armed Bandit Algorithm for Bid Shading in Online Display AdvertisingabstractIn real-time bidding systems, ad exchanges and supply-side platforms (SSP) are switching from the second-price auction (SPA) to the first-price auction (FPA), where the advertisers should pay what they bid if they win the auction. To avoid overpaying, advertisers are motivated to conceal their truthful evaluations of impression opportunities through bid shading methods. However, advertisers are consistently facing a trade-off between the probability and cost-saving of winning, due to the information asymmetry, where advertisers lack knowledge about their competitors' bids in the market. To address this challenge, we propose a Bayes ian Multi-Armed Bandit (BayesMAB) algorithm for bid shading when the winning price is unknown to advertisers who lose the impression opportunity. BayesMAB incorporates the mechanism of FPA to infer each price interval's winning rate by progressively updating the market price hidden by SSP. In this way, BayesMAB better approximates the winning rates of price intervals and thus is able to derive the optimal shaded bid that balances the trade-off between the probability and cost-saving of winning the impression opportunity. We conducted large-scale A/B tests on Tencent's online display advertising platform. The cost-per-mile (CPM) and cost-per-action (CPA) decreased by 13.06% and 11.90%, respectively, whereas the return on investment (ROI) increased by 12.31% with only 2.7% sacrifice of the winning rate. We also validated BayesMAB's superior performance in an offline semi-simulated experiment with SPA data sets. BayesMAB has been deployed online and is impacting billions of traffic every day. Codes are available at https://github.com/BayesMAB/BayesMAB. Mengzhuo Guo, Wuqi Zhang, Congde Yuan, Binfeng Jia, Guoqing Song, Hua Hua, Shuangyang Wang, Qingpeng Zhang |
CIKM | 6 |
| 2024 | Expressive paragraph text-to-speech synthesis with multi-step variational autoencoder
Xuyuan Li, Zengqiang Shang, Peiyang Shi, Hua Hua, Ta Li, Pengyuan Zhang |
INTERSPEECH | 4 |
| 2024 | Emilia: An Extensive, Multilingual, and Diverse Speech Dataset For Large-Scale Speech GenerationabstractRecent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and spontaneous speech datasets. In response, we introduce Emilia, the first large-scale, multilingual, and diverse speech generation dataset. Emilia starts with over 101k hours of speech across six languages, covering a wide range of speaking styles to enable more natural and spontaneous speech generation. To facilitate the scale-up of Emilia, we also present Emilia-Pipe, the first open-source preprocessing pipeline designed to efficiently transform raw, in-the-wild speech data into high-quality training data with speech annotations. Experimental results demonstrate the effectiveness of both Emilia and Emilia-Pipe. Demos are available at: https://emilia-dataset.github.io/Emilia-Demo-Page/. Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu 0008, Jiaqi Li 0030, Peiyang Shi, Yuancheng Wang, Kai Chen 0026, Pengyuan Zhang, Zhizheng Wu 0001 |
SLT | 6 |
| 2023 | 3MN: Three Meta Networks for Multi-Scenario and Multi-Task Learning in Online Advertising Recommender SystemsabstractRecommender systems are widely applied on web. For example, online advertising systems rely on recommender systems to accurately estimate the value of display opportunities, which is critical to maximize the profits of advertisers. To reduce computational resource consumption, the core tactic of Multi-Scenario Multi-Task Learning (MSMTL) is to devise a single recommder system that is adapted to all contexts instead of implementing multiple scenario-oriented or task-oriented recommender systems. However, MSMTL is challenging because there are complicated task-task, scenario-scenario, and task-scenario interrelations; the characteristic of different tasks in different scenarios also largely varies; and samples of each context are often unevenly distributed. Previous MSMTL solutions focus on applying scenario knowledge to improve the performance of multi-task learning, while neglecting the complicated interrelations among tasks and scenarios. Moreover, samples derived from different scenarios are transferred into the latent embedding with the same dimension. This static embedding strategy impedes the practicality of model expressiveness, since the scenarios with sufficient samples are underrepresented and those with insufficient samples are over-represented. Hua Hua, Shuangyang Wang, Chongyu Zhong |
CIKM | 2 |
| 2023 | The Impact of Silence on Speech Anti-SpoofingabstractThe current speech anti-spoofing countermeasures (CMs) show excellent performance on specific datasets. However, removing the silence of test speech through Voice Activity Detection (VAD) can severely degrade performance. In this paper, the impact of silence on speech anti-spoofing is analyzed. First, the reasons for the impact are explored, including the proportion of silence duration and the content of silence. The proportion of silence duration in spoof speech generated by text-to-speech (TTS) algorithms is lower than that in bonafide speech. And the content of silence generated by different waveform generators varies compared to bonafide speech. Then the impact of silence on model prediction is explored. Even after retraining, the spoof speech generated by neural network based end-to-end TTS algorithms suffers a significant rise in error rates when the silence is removed. To demonstrate the reasons for the impact of silence on CMs, the attention distribution of a CM is visualized through class activation mapping (CAM). Furthermore, the implementation and analysis of the experiments masking silence or non-silence demonstrates the significance of the proportion of silence duration for detecting TTS and the importance of silence content for detecting voice conversion (VC). Based on the experimental results, improving the robustness of CMs against unknown spoofing attacks by masking silence is also proposed. Finally, the attacks on anti-spoofing CMs through concatenating silence, and the mitigation of VAD and silence attack through low-pass filtering are introduced. Zhuo Li 0020, Jingze Lu, Hua Hua, Pengyuan Zhang |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Towards Explainable Action Recognition by Salient Qualitative Spatial Object Relation ChainsabstractIn order to be trusted by humans, Artificial Intelligence agents should be able to describe rationales behind their decisions. One such application is human action recognition in critical or sensitive scenarios, where trustworthy and explainable action recognizers are expected. For example, reliable pedestrian action recognition is essential for self-driving cars and explanations for real-time decision making are critical for investigations if an accident happens. In this regard, learning-based approaches, despite their popularity and accuracy, are disadvantageous due to their limited interpretability. This paper presents a novel neuro-symbolic approach that recognizes actions from videos with human-understandable explanations. Specifically, we first propose to represent videos symbolically by qualitative spatial relations between objects called qualitative spatial object relation chains. We further develop a neural saliency estimator to capture the correlation between such object relation chains and the occurrence of actions. Given an unseen video, this neural saliency estimator is able to tell which object relation chains are more important for the action recognized. We evaluate our approach on two real-life video datasets, with respect to recognition accuracy and the quality of generated action explanations. Experiments show that our approach achieves superior performance on both aspects to previous symbolic approaches, thus facilitating trustworthy intelligent decision making. Our approach can be used to augment state-of-the-art learning approaches with explainabilities. Hua Hua, Ruiqi Li 0005, Peng Zhang 0021, Jochen Renz, Anthony G. Cohn 0001 |
AAAI | 1 |
| 2022 | Cross-domain Recommendation via Adversarial AdaptationabstractData scarcity, e.g., labeled data being either unavailable or too expensive, is a perpetual challenge of recommendation systems. Cross-domain recommendation leverages the label information in the source domain to facilitate the task in the target domain. However, in many real-world cross-domain recommendation systems, the source domain and the target domain are sampled from different data distributions, which obstructs the cross-domain knowledge transfer. In this paper, we propose to specifically align the data distributions between the source domain and the target domain to alleviate imbalanced sample distribution and thus challenge the data scarcity issue in the target domain. Technically, our proposed approach builds an adversarial adaptation (AA) framework to adversarially train the target model together with a pre-trained source model. A domain discriminator plays the two-player minmax game with the target model and guides the target model to learn domain-invariant features that can be transferred across domains. At the same time, the target model is calibrated to learn domain-specific information of the target domain. With such a formulation, the target model not only learns domain-invariant features for knowledge transfer, but also preserves domain-specific information for target recommendation. We apply the proposed method to address the issues of insufficient data and imbalanced sample distribution in real-world Click-Through Rate (CTR)/Conversion Rate (CVR) predictions on a large-scale dataset. Specifically, we formulate our approach as a plug-and-play module to boost existing recommendation systems. Extensive experiments verify that the proposed method is able to significantly improve the prediction performance on the target domain. For instance, our method can boost PLE with a performance improvement of 13.88% in terms of Area Under Curve (AUC) compared with single-domain PLE. Hongzu Su, Xuejiao Yang, Hua Hua, Shuangyang Wang, Jingjing Li 0001 |
CIKM | 4 |
| 2021 | Unsupervised Novelty Characterization in Physical Environments Using Qualitative Spatial RelationsabstractDetecting, characterizing and adapting to novelty, whether in the form of previously unseen objects or phenomena, or unexpected changes in the behavior of known elements, is essential for Artificial Intelligence agents to operate reliably in unconstrained real-world environments. We propose an automatic, unsupervised approach to novelty characterization for dynamic domains, based on describing the behaviors and interactions of objects in terms of their possible actions. To abstract from the variety of realizations of an action that can occur in physical domains, we model states in terms of qualitative spatial relations (QSRs) between their entities. By first learning a model of actions in the non-novel environment from the state transitions observed as the agent interacts with the world, we can detect novelty by the persistent deviations from this model that it causes, and characterize the novelty by new or modified actions. We also present a new method of learning action models from observation, based on conceptual similarity and hierarchical clustering. Ruiqi Li 0005, Hua Hua, Patrik Haslum, Jochen Renz |
KR | 2 |
| 2019 | Qualitative Place Maps for Landmark-based Localization and Navigation in GPS-denied EnvironmentsabstractGPS-based services (e.g. Google Maps) are very popular in our daily life, while there are still many GPS-denied environments (e.g. indoor and underground scenarios) in which they cannot be used. In these situations, localization and navigation are still important, for example in emergency evacuation or indoor navigation. In this paper we aim to solve the problem of localizing and navigating humans or robots in GPS-denied environments based on landmarks. Our work is inspired by human daily communications about localization and navigation, for example someone who has been to a shopping mall many times can localize and guide another person to get to a certain shop via conversations over the cellphone. Our goal is to build a system with the same capability. We propose a system that relies on qualitative information of places (e.g. the direction relations between landmarks involved in route descriptions), where localization can be achieved in an interactive manner and by analyzing observations provided by users. Our system decides the "best" route from one place to another by three factors: the number of landmarks; the number of ambiguous turns; and the qualitative distance (e.g. near and far). According to the experimental results, the number of requeries in the interactive localization process is acceptable and our route planning algorithm outperforms previous methods in several cases. Hua Hua, Peng Zhang 0021, Jochen Renz |
SIGSPATIAL/GIS | 1 |
| 2018 | Towards Explainable Inference about Object Motion using Qualitative Reasoning
Xiaoyu Ge, Jochen Renz, Hua Hua |
KR | 3 |
| 2018 | Qualitative Representation and Reasoning over Direction Relations across Different Frames of Reference
Hua Hua, Jochen Renz, Xiaoyu Ge |
KR | 1 |
| 2016 | A new framework for remote sensing image super-resolution: Sparse representation-based method by processing dictionaries with multi-type features
Wei Wu 0002, Xiaomin Yang, Kai Liu 0012, Yiguang Liu, Binyu Yan, Hua Hua |
J. Syst. Archit. | 6 |
| 2007 | Research and application of a new predictive control based on state feedback theory in power plant control systemabstractReheat steam temperature of the boiler in power plant is the important parameter of unit security and economical run. This paper proposed a new dynamic matrix control based on state feedback theory, for the large inertia and large delay characteristic of the reheat steam temperature plant. The principle is to compensate the large inertia and delay characteristic of the plant by state feedback theory, then to control the generalized plant by predictive control. In this paper the design of controller is introduced. The simulation results demonstrate that the new control system has good robustness and transient performance. And the practical results show that the satisfactory control results have been achieved. So it is an effective control strategy for large delay industry process. Now it has been successfully applied to reheat steam temperature control system in power plant. Zhigang Hua, Hua Hua, Jianhong Lu |
IEEE Congress on Evolutionary Computation | 2 |