VLDB 2026 Research / reviewers in the wild / expert
Hongxu Zhu
dblp:208/7763
· DBLP profile ↗
19ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0001-6257-7065ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AV-SSAN: Audio-Visual Selective DOA Estimation Through Explicit Multi-Band Semantic-Spatial AlignmentabstractAudio-visual sound source localization (AV-SSL) estimates the position of sound sources by fusing auditory and visual cues. Current AV-SSL methodologies typically require spatially-paired audio-visual data and cannot selectively localize specific target sources. To address these limitations, we introduce Cross-Instance Audio-Visual Localization (CI-AVL), a novel task that localizes target sound sources using visual prompts from different instances of the same semantic class. CI-AVL enables selective localization without spatially paired data. To solve this task, we propose AV-SSAN, a semantic-spatial alignment framework centered on a Multi-Band Semantic-Spatial Alignment Network (MB-SSA Net). MB-SSA Net decomposes the audio spectrogram into multiple frequency bands, aligns each band with semantic visual prompts, and refines spatial cues to estimate the direction-of-arrival (DoA). To facilitate this research, we construct VGGSound-SSL, a large-scale dataset comprising 13,981 spatial audio clips across 296 categories, each paired with visual prompts. AV-SSAN achieves a mean absolute error of 16.59° and an accuracy of 71.29%, significantly outperforming existing AV-SSL methods. Hongxu Zhu, Kainan Chen, Xinyuan Qian 0001 |
AAAI | 2 |
| 2025 | Neural-Symbolic System Control Adjustment Based on Runtime Verification
Hongxu Zhu, Wanwei Liu, Ji Wang 0001 |
ICFEM | 1 |
| 2025 | AUPRC: a metric for evaluating the performance of in-silico perturbation methods in identifying differentially expressed genesabstractIn silico perturbation models, computational methods that can predict cellular responses to perturbations, present an opportunity to reduce the need for costly and time-intensive in vitro experiments. Many recently proposed models predict high-dimensional cellular responses, such as gene or protein expression to perturbations such as gene knockout or drugs. However, evaluating in silico performance has largely relied on metrics such as $R^{2}$, which assess overall prediction accuracy but fail to capture biologically significant outcomes like the identification of differentially expressed (DE) genes. In this study, we present a novel evaluation framework that introduces the AUPRC metric to assess the precision and recall of DE gene predictions. By applying this framework to both single-cell and pseudo-bulked datasets, we systematically benchmark simple and advanced computational models. Our results highlight a significant discrepancy between $R^{2}$ and AUPRC, with models achieving high $R^{2}$ values but struggling to identify DE genes, as reflected in their low AUPRC values. This finding underscores the limitations of traditional evaluation metrics and the importance of biologically relevant assessments. Our framework provides a more comprehensive understanding of model capabilities, advancing the application of computational approaches in cellular perturbation research. Hongxu Zhu, Amir Asiaee, Leila Azinfar, Jun Li 0068, Ehsan Irajizad, Kim-Anh Do, James P. Long |
Briefings Bioinform. | 1 |
| 2024 | An Empirical Study on the Impact of Positional Encoding in Transformer-Based Monaural Speech EnhancementabstractTransformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to enable Transformers to distinguish the order of elements in a sequence. However, it remains unclear how positional encoding exactly impacts speech enhancement based on Transformer architectures. In this paper, we perform a comprehensive empirical study evaluating five positional encoding methods, i.e., Sinusoidal and learned absolute position embedding (APE), T5-RPE, KERPLE, as well as the Transformer without positional encoding (No-Pos), across both causal and noncausal configurations. We conduct extensive speech enhancement experiments, involving spectral mapping and masking methods. Our findings establish that positional encoding is not quite helpful for the models in a causal configuration, which indicates that causal attention may implicitly incorporate position information. In a noncausal configuration, the models significantly benefit from the use of positional encoding. In addition, we find that among the four position embeddings, relative position embeddings outperform APEs. Qiquan Zhang, Meng Ge, Hongxu Zhu, Eliathamby Ambikairajah, Zhaoheng Ni, Haizhou Li 0001 |
ICASSP | 3 |
| 2024 | An Exploration of Length Generalization in Transformer-Based Speech Enhancement
Qiquan Zhang, Hongxu Zhu, Xinyuan Qian 0001, Eliathamby Ambikairajah, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2023 | An Automata-Theoretic Approach to Synthesizing Binarized Neural Networks
Ye Tao 0008, Wanwei Liu, Fu Song, Ji Wang 0001, Hongxu Zhu |
ATVA (1) | 6 |
| 2023 | Ripple Sparse Self-Attention for Monaural Speech EnhancementabstractThe use of Transformer represents a recent success in speech enhancement. However, as its core component, self-attention suffers from quadratic complexity, which is computationally prohibited for long speech recordings. Moreover, it allows each time frame to attend to all time frames, neglecting the strong local correlations of speech signals. This study presents a simple yet effective sparse self-attention for speech enhancement, called ripple attention, which simultaneously performs fine- and coarse-grained modeling for local and global dependencies, respectively. Specifically, we employ local band attention to enable each frame to attend to its closest neighbor frames in a window at fine granularity, while employing dilated attention outside the window to model the global dependencies at a coarse granularity. We evaluate the efficacy of our ripple attention for speech enhancement on two commonly used training objectives. Extensive experimental results consistently confirm the superior performance of the ripple attention design over standard full self-attention, blockwise attention, and dual-path attention (Sep-Former) in terms of speech quality and intelligibility. Qiquan Zhang, Hongxu Zhu, Xinyuan Qian 0001, Zhaoheng Ni, Haizhou Li 0001 |
ICASSP | 2 |
| 2023 | Speech-Oriented Sparse Attention Denoising for Voice User Interface Toward Industry 5.0abstractThe adoption of voice user interface (VUI) will promote network automation with enhanced efficiency with reduced simplicity and operating expense in Industry 5.0. Given the noisy environments, speech denoising is indispensable for the VUI in Internet of Things (IoT) or Industrial IoT (IIoT). Despite Transformer's recent success in speech denoising, the adopted full self-attention suffers from quadratic complexity, which challenges the computational power of the IoT/IIoT components. Considering the strong local correlations of speech signals, a speech-oriented sparse attention denoising scheme is developed to keep the meaningful local and global dependencies while mitigating the redundant attentions, resulting in a significant reduction in computational complexity. With the full self-attention as the baseline, experimental results revealed that the proposed scheme achieves a better denoising performance and yields a lower computational cost, indicating the strong potential for various VUI application scenarios in IoT and IIoT toward Industry 5.0. Hongxu Zhu, Qiquan Zhang, Peng Gao 0005, Xinyuan Qian 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Multistage Deep Transfer Learning for EmIoT-Enabled Human-Computer InteractionabstractEmotional Internet of Things (EmIoT), which provides Internet of Things (IoT) devices cognitive and socialization capabilities, has been regarded as a future direction to improve users’ experiences. With the development of intelligent techniques, the requirement of EmIoT is not only sensing the users’ emotional states but also providing emotional feedbacks. Human–computer interaction has been studied to achieve speech interaction with IoT devices. The recent advances in neural text-to-speech (TTS) have made “human parity” synthesized speech possible for IoT-enabled human–computer interaction. Furthermore, emotion control can be achieved by using the emotional codes in a unified model, referred to as emotional TTS (or ETTS for short). Such ETTS models have achieved promising emotional expressiveness using large-scale emotion-annotated English data set; however, they are not practical in IoT environments with other mainstream languages, especially for Chinese. In fact, the limited available large-scale emotion-annotated data set is challenging the development of Chinese ETTS. To address that we propose a multistage deep transfer learning scheme to design a high-quality Chinese ETTS system under a small-scale training corpus to achieve EmIoT in Mandarin environments. In this scheme, the pretrained knowledge from the former stages corresponding to a large-scale neutral English and a medium-scale emotional English corpora is transferred to a Mandarin ETTS model. Thereby, the trained model can achieve high-quality emotional speech with limited available emotional corpus, which is able to serve various EmIoT-oriented applications. The experiments have been conducted to demonstrate the effectiveness and superiority of the proposed model as compared to other counterparts in terms of naturalness and emotional expressiveness. We refer readers to visit our demo Webpage1enjoy the synthesized speech samples. Rui Liu 0008, Qi Liu 0005, Hongxu Zhu, Hui Cao 0004 |
IEEE Internet Things J. | 3 |
| 2022 | Efficient Load Balancing for Heterogeneous Radio-Replication-Combined LoRaWANabstractLoRa wide area network (LoRaWAN), an emerging IoT protocol, has been popularized in large-scale applications, given its long-range and low-power properties. Hitherto, there is no appropriate traffic model for LoRaWAN to estimate the heterogeneous arriving traffic at the network server cluster (NSC). Inefficient computation power planning or even processing failure might be further caused. Radio replication, commonly existed in the arriving traffic at NSC in LoRaWAN, also causes difficulty estimating the makespan (i.e., mean processing time in NSC). To overcome the abovementioned limitations, a heterogeneous radio-replication-aware traffic aggregation model is proposed to estimate the arriving traffic for LoRaWAN. In addition, a radio-replication-combined supermarket model (RRC-SM), on top of HTAM, is proposed to achieve load balancing among servers in LoRaWAN. Furthermore, a nondominated sorting genetic algorithm based on multiobjective optimization is developed to simultaneously minimize cost and latency on NSC. Experiments reveal that the proposed HTAM and RRC-SM agree well with the simulation outcome. Under the arriving traffic estimated as 6.16 erlangs with four radio replications of each arriving packet on average, the proposed RRC-SM provides more than 50% reduction on the total processing latency and 75% reduction on the number of servers in NSC than other existing models. Yucheng Liu 0001, Kim Fung Tsang, Hongxu Zhu, Hao Ran Chi, Yang Wei 0001, Hao Wang 0055, Chung Kit Wu |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | Extreme RSS Based Indoor Localization for LoRaWAN With Boundary AutocorrelationabstractThe received signal strength (RSS) finger-print-based approaches are widely used for indoor location-based services (LBSs). The emerging long range wide area network (LoRaWAN) is a cost-effective solution for indoor latency-tolerant LBSs attributed to its long-range property. In general, there are serious RSS fluctuations due to fadings along the communication path, thus significantly jeopardizing the localization accuracy. To overcome the challenge, in this article we propose the extreme RSS (ERSS) to stabilize the fingerprint database and formulate boundary autocorrelation to downsize tremendously the searching complexity and thus proliferating localization accuracy. In essence, the RSS fluctuations are modeled as a Bernoulli random process so that the RSS stability can be estimated by a newly defined fluctuation analytic function. To mitigate the impact of the perturbative fluctuation, the ERSS is further defined to cultivate a highly stable and robust fingerprint database which withstands environmental dynamics. In addition, boundary autocorrelation is developed to measure and compare the similarity between the measured RSS values versus the prestored fingerprint database. RSS values with low autocorrelation coefficients are eradicated from the typically lengthy searching. The downsized complexity significantly improves the localization accuracy. Experiments were carried out and the results revealed that the proposed method achieved sub-10-m localization accuracy in indoor environments. Such accuracy is encouraging and superior in contemporary LoRaWAN measurements. Hongxu Zhu, Kim Fung Tsang, Yucheng Liu 0001, Yang Wei 0001, Hao Wang 0055, Chung Kit Wu, Hao Ran Chi |
IEEE Trans. Ind. Informatics | 1 |
| 2018 | Charging Infrastructure Planning for Electric Vehicles in Giant CitiesabstractWith the rapid exhaustion of fossil energy, electric vehicles (EVs) become one of the key candidates for the next generation of transportation. Increasingly perfect technology developed makes EVs grow significantly. Therefore, Charging Stations (CSs), as accessories and necessities of EVs, should form a network with optimal planning. Inappropriate CS network design could cause series of negative effects to the popularization of EVs, the layout of the city traffic network and the financial cost of CS network construction, etc. Besides, the charging infrastructure planning for cities with large population and high EV density becomes even difficult. In this paper, an Effective Planning of CSs Network (CSN) is proposed. Comprehensive environmental elements (e.g. cities' and CSs' information, EV charging status, etc.) are considered in CSN. The CSN deals with complicated city planning in giant cities (i.e. large population, high EV density, etc.). Hong Kong is selected as the case study because it can be regarded as a typical giant city. Results show that the proposed CSN can ensure the EVs can find a CS before it is out of power. Besides, the CSN saves ~18% financial cost for the charging infrastructure planning in giant cities. Hao Ran Chi, Hongxu Zhu, Yucheng Liu 0001, Faan Hei Hung, Kim Fung Tsang, Mo-Yuen Chow, Chengbin Ma |
IECON | 2 |
| 2018 | Packet Loss Analysis for LoRa-Based Heart Monitoring SystemabstractDue to lack of heart monitoring device, heart problems have been causing more than 17 million people death every year. Recently, researches on IoT (Internet of things) system provide feasible solutions to solve the problems. Such as LoRa wireless communication protocol, it can cover more than 1km2area and reliable communication connection. In heart monitoring system, packet transmission frequency should be considered carefully because of the reliability. Therefore, in this paper, the packet loss of heart monitoring system has been analyzed through packet length, transmission frequency and communication distance. Results shows that the packet loss can be less than 1% through specified communication property. The result in the paper can provide reference for reliable heart monitoring system design. Yucheng Liu 0001, Hongxu Zhu, Tsz Tat Yu, Kim Fung Tsang, Chung Kit Wu, Faan Hei Hung |
IECON | 2 |
| 2018 | A Real-Time Drivers' Status Monitoring Scheme with Safety AnalysisabstractSmart transportation and smart healthcare are considered as essential Smart City applications. The emerging light-weight sensors facilitate real-time monitoring drivers' status in various applications especially safety and healthcare. As such, the statistics reveals that >60% of adult drivers felt sleepy while driving, and drunk drivers are found in >40% of traffic accidents. In this paper, an electrocardiogram (ECG) based Drivers' Status Monitoring (ECG-DSM) system is developed to detect drowsy and drunk driving. The proposed ECG-DSM extracted similarities of ECG signals under normal, drowsy and drunk conditions, and the corresponding feature vector was built. The classifier is expected to alert drivers accurately and timely to prevent traffic accidents. Hence, the classifier's trade-off between accuracy and detection time was analysed by adjusting the dimensionality of feature vector. Safety analysis using Monte Carlo simulation was carried out to determine the best classifier under practical working environment. The results demonstrated that the best classifier for ECG-DSM achieves 91 % of average accuracy and 4.2s of detection time, and it can prevent >92 % of vehicle collisions due to drowsy and drunk driving. The proposed work will contribute to road traffic safety and save $50 billion US dollars on the cost of traffic injuries. Wai Hin Wan, Yee Ting Tsang, Hongxu Zhu, Cheon Hoi Koo, Yucheng Liu 0001, Chi Chung Lee 0001 |
IECON | 3 |
| 2018 | Feasiblity Studies on Smart Pole Connectivity Based on LPWA IoT Communication Platform for Industrial ApplicationsabstractIn a metropolitan city, the power consumption of buildings is a great concern for not only economic but also environmental concerns. Researchers is proposing some new low power wide area (LPWA) IoT communication protocols to monitor and control power consumption in buildings. Three common LPWA communication protocols include: NB-IoT, LoRa and Sigfox, which can provide different functions or performance under different conditions. For building management system, these protocols can create platform to monitor and control the “things”, such as energy. Light Poles are used to be built to continuously lighten up the city which are served as passive devices to be controlled by the Highway Department. With the LWPA communication protocols and sensors or actuators, they can provide a low power low cost network nodes to serve local community; and uplink the data to Cloud via mobile communication network or Ethernet, as Smart Poles. In this paper, some fundamental scheme and based on LPWA protocols to design building management system are introduced; and prototype of Smart Poles have been networked and applied in a traditional industrial area. Tsz Tat Yu, Yucheng Liu 0001, Hongxu Zhu, Kim Fung Tsang |
IECON | 3 |
| 2018 | Sleep Apnea Monitoring for Smart HealthcareabstractVocational safety problems have been causing millions of workers dead even with related safety policies launched. Sleep apnea is one of the main causes that renders insufficient sleep and thus becomes a high potential risk of working accidents. To assess sleep apnea, sleep stage monitoring and classification is the main and accurate way. Therefore, in this paper, a new classifier is designed for the sleep stage classification among awake, light sleep and deep sleep. A new kernel is designed for the sleep apnea classification. Results show that the proposed method can achieve an accuracy up to 97% and~18% higher than the previous related works. Such a high accuracy ensures the efficient diagnosis of sleep apnea. Hongxu Zhu, Cheon Hoi Koo, Chung Kit Wu, Wai Hin Wan, Yee Ting Tsang, Kim Fung Tsang |
IECON | 1 |
| 2017 | Review of state-of-the-art wireless technologies and applications in smart citiesabstractThere are increasing preferences to employ wireless communication technologies for high mobility, high scalability and low-cost applications in smart city development. This paper gives a brief synopsis of typical wireless technologies in smart city applications and the comparison analysis between them. The trend for smart city wireless technology is also presented. Examples, for several key applications within smart city development (healthcare, smart grid, localization) are studied and current advanced solutions supporting these applications are summarized with futuristic trends and demands are presented. Hongxu Zhu, Anna S. F. Chang, Roy Kalawsky, Kim Fung Tsang, Gerhard P. Hancke 0002, Lucia Lo Bello, Bingo Wing-Kuen Ling |
IECON | 1 |
| 2017 | A time-synchronized ZigBee building network for smart water managementabstractWater management is an important issue in economics and environment. Recently, amount of water control system has been proposed and developed. For the type of intelligent water control, the related parameters will be the input of the control system. Hence, there is a need of developing a scalable, flexible and reliable sensor network for related parameters monitoring. To install and replace water sensors in building networks, wireless connection will be the first priority. However, improper time synchronization in the network will cause packet loss and long latency which degrades the network performance. In this paper, time-synchronized ZigBee building network (TS-ZBN) is proposed for water management. The node-to-node time synchronization is proposed. The concept is to calculate the clock difference by studying the propagation delay model. The simulation result shows that the mean synchronization error and variance are low. Chung Kit Wu, Hongxu Zhu, Loi Lei Lai, Anna S. F. Chang, Fengjun Li, Kim Fung Tsang, Roy Kalawsky |
INDIN | 2 |
| 2016 | BER performance evaluation of Spatial Modulation via numerical simulationsabstractSpatial Modulation is a recently developed low-complexity MIMO scheme that jointly uses antenna indices and a conventional constellation set to convey information. Different types of developed SM systems have been proposed to mitigate the limitations of basic SM systems. We compared the performance of different types of spatial modulation system in different channel environment and give the key factors that could affect the bit error rate. Hongxu Zhu, Chung Kit Wu, Kim Fung Tsang, Faan Hei Hung |
IECON | 1 |