EDBT 2026 Demo / reviewers in the wild / expert
Bo Hu 0036
dblp:04/2380-36
· DBLP profile ↗
17ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0001-5508-5732ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Graph Mixture of Experts and Memory-augmented Routers for Multivariate Time Series Anomaly DetectionabstractMultivariate time series (MTS) anomaly detection is a critical task that involves identifying abnormal patterns or events in data that consist of multiple interrelated time series. In order to better model the complex interdependence between entities and the various inherent characteristics of each entity, the graph neural network (GNN) based methods are widely adopted by existing methods. In each layer of GNN, node features aggregate information from their neighboring nodes to update their information. In doing so, from shallow layer to deep layer in GNN, original individual node features continue to be weakened and more structural information, i.e., from short-distance neighborhood to long-distance neighborhood, continues to be enhanced. However, research to date has largely ignored the understanding of how hierarchical graph information is represented and their characteristics that can benefit anomaly detection. Existing methods simply leverage the output from the last layer of GNN for anomaly estimation while neglecting the essential information contained in the intermediate GNN layers. To address such limitations, in this paper, we propose a Graph Mixture of Experts (Graph-MoE) network for multivariate time series anomaly detection, which incorporates the mixture of experts (MoE) module to adaptively represent and integrate hierarchical multi-layer graph information into entity representations. It is worth noting that our Graph-MoE can be integrated into any GNN-based MTS anomaly detection method in a plug-and-play manner. In addition, the memory-augmented routers are proposed in this paper to capture the correlation temporal information in terms of the global historical features of MTS to adaptively weigh the obtained entity representations to achieve successful anomaly estimation. Extensive experiments on five challenging datasets prove the superiority of our approach and each proposed module. Weidong Chen 0013, Bo Hu 0036, Zhendong Mao 0001 |
AAAI | 3 |
| 2025 | Same Vaccine, Different Voices: A Cross-Modality Analysis of HPV Vaccine Discourse on Social MediaabstractDespite the proven efficacy of HPV vaccines, uptake remains limited in many regions, including China. This study investigates how health beliefs and emotional responses evolve across text-, audio-, and video-based platforms by analyzing data from three representative platforms in China, including 273,357 posts from Weibo (text-based), 1,228 podcasts from Ximalaya (audio-based), and 1,225 videos from Douyin (video-based) from July 2018 to March 2023. The comparisons are conducted under four dimensions as suggested by the Health Belief Model (HBM), including susceptibility, severity, benefits, and barriers. Our findings reveal distinct modality-specific patterns. For instance, a text-based platform tends to amplify barriers and negativity, an audio-based platform enables balanced and sustained discussions, and a video-based platform highlights personal anecdotes and drives rapid sentiment shifts. By highlighting these modality-specific differences and addressing potential cross-modal incongruities at the content level, we provide actionable insights for public health communicators, policymakers, and platform designers to tailor strategies, foster informed decision-making, and ultimately enhance HPV vaccine uptake in complex social media ecosystems. Mengxiao Zhu 0001, Ruoxiao Su, Bo Hu 0036 |
ICWSM | 6 |
| 2025 | Improving Video Summarization by Exploring the Coherence Between Corresponding CaptionsabstractVideo summarization aims to generate a compact summary of the original video by selecting and combining the most representative parts. Most existing approaches only focus on recognizing key video segments to generate the summary, which lacks holistic considerations. The transitions between selected video segments are usually abrupt and inconsistent, making the summary confusing. Indeed, the coherence of video summaries is crucial to improve the quality and user viewing experience. However, the coherence between video segments is hard to measure and optimize from a pure vision perspective. To this end, we propose a Language-guided Segment Coherence-Aware Network (LS-CAN), which integrates entire coherence considerations into the key segment recognition. The main idea of LS-CAN is to explore the coherence of corresponding text modality to facilitate the entire coherence of the video summary, which leverages the natural property in the language that contextual coherence is easy to measure. In terms of text coherence measures, specifically, we propose the multi-graph correlated neural network module (MGCNN), which constructs a graph for each sentence based on three key components, i.e., subject, attribute, and action words. For each sentence pair, the node features are then discriminatively learned by incorporating neighbors of its own graph and information of its dual graph, reducing the error of synonyms or reference relationships in measuring the correlation between sentences, as well as the error caused by considering each component separately. In doing so, MGCNN utilizes subject agreement, attribute coherence, and action succession to measure text coherence. Besides, with the help of large language models, we augment the original text coherence annotations, improving the ability of MGCNN to judge coherence. Extensive experiments on three challenging datasets demonstrate the superiority of our approach and each proposed module, especially improving the latest records by +3.8%, +14.2% and +12% w.r.t. F1 scores, $\tau $ and $\rho $ metrics on the BLiSS dataset. Cheng Ye 0004, Weidong Chen 0013, Bo Hu 0036, Lei Zhang 0119, Yongdong Zhang 0001, Zhendong Mao 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Gradual Residuals Alignment: A Dual-Stream Framework for GAN Inversion and Image Attribute EditingabstractGAN-based image attribute editing firstly leverages GAN Inversion to project real images into the latent space of GAN and then manipulates corresponding latent codes. Recent inversion methods mainly utilize additional high-bit features to improve image details preservation, as low-bit codes cannot faithfully reconstruct source images, leading to the loss of details. However, during editing, existing works fail to accurately complement the lost details and suffer from poor editability. The main reason is they inject all the lost details indiscriminately at one time, which inherently induces the position and quantity of details to overfit source images, resulting in inconsistent content and artifacts in edited images. This work argues that details should be gradually injected into both the reconstruction and editing process in a multi-stage coarse-to-fine manner for better detail preservation and high editability. Therefore, a novel dual-stream framework is proposed to accurately complement details at each stage. The Reconstruction Stream is employed to embed coarse-to-fine lost details into residual features and then adaptively add them to the GAN generator. In the Editing Stream, residual features are accurately aligned by our Selective Attention mechanism and then injected into the editing process in a multi-stage manner. Extensive experiments have shown the superiority of our framework in both reconstruction accuracy and editing quality compared with existing methods. Hao Li 0189, Mengqi Huang, Lei Zhang 0119, Bo Hu 0036, Yi Liu 0148, Zhendong Mao 0001 |
AAAI | 4 |
| 2024 | Cascade Semantic Prompt Alignment Network for Image CaptioningabstractImage captioning (IC) takes an image as input and generates open-form descriptions in the domain of natural language. IC requires the detection of objects, modeling of relations between them, an assessment of the semantics of the scene and representing the extracted knowledge in a language space. Previous detector-based models suffer from limited semantic perception capability due to predefined object detection classes and semantic inconsistency between visual region features and numeric labels of the detector. Inspired by the fact that text prompts in pre-trained multi-modal models contain specific linguistic knowledge rather than discrete labels, and excel at an open-form semantic understanding of visual inputs and their representation in the domain of natural language. We aim to distill and leverage the transferable language knowledge from the pre-trained RegionCLIP model to remedy the detector for generating rich image captioning. In this paper, we propose a novel Cascade Semantic Prompt Alignment Network (CSA-Net) to produce an aligned fine-grained regional semantic-visual space where rich and consistent textual semantic details are automatically incorporated to region features. Specifically, we first align the object semantic prompt and region features to produce semantic grounded object features. Then, we employ these object features and relation semantic prompt to predict the relations between objects. Finally, these enhanced object and relation features are fed into the language decoder, generating rich descriptions. Extensive experiments conducted on the MSCOCO dataset show that our method achieves a new state-of-the-art performance with 145.2% (single model) and 147.0% (ensemble of 4 models) CIDEr scores on the ‘Karpathy’ split, 141.6% (c5) and 144.1% (c40) CIDEr scores on the official online test server. Significantly, CSA-Net outperforms in generating captions with higher quality and diversity, achieving a RefCLIP-S score of 83.2. Moreover, we expand the testbeds to other challenging captioning benchmarks, i.e., nocaps datasets, CSA-Net demonstrates superior zero-shot capability. Source codes released at https://github.com/CrossmodalGroup/CSA-Net. Lei Zhang 0119, Kun Zhang 0040, Bo Hu 0036, Hongtao Xie 0001, Zhendong Mao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Enhanced Semantic Similarity Learning Framework for Image-Text MatchingabstractImage-text matching is a fundamental task to bridge vision and language. The critical challenge lies in accurately learning the semantic similarity between these two heterogeneous modalities. For visual and textual features, existing methods typically default to a static dimensional correspondence mechanism, i.e., using a single dimension as the measure-unit to perform one-to-one correspondence, to examine semantic similarity, e.g., the cosine/Euclidean distance or the weighted similarity. In this paper, different from the single-dimensional correspondence with limited semantic expressive capability, we propose a novel enhanced semantic similarity learning (ESL), which generalizes both measure-units and their correspondences into a dynamic learnable framework to examine the multi-dimensional enhanced correspondence between visual and textual features. Specifically, we first devise the intra-modal multi-dimensional aggregators with iterative enhancing mechanism, which dynamically captures new measure-units integrated by hierarchical multi-dimensions, producing diverse semantic combinatorial expressive capabilities to provide richer and discriminative information for similarity examination. Then, we devise the inter-modal enhanced correspondence learning with sparse contribution degrees, which comprehensively and efficiently determines the cross-modal semantic similarity. Extensive experiments verify its superiority in achieving state-of-the-art performance. Codes will be released athttps://github.com/CrossmodalGroup/ESL. Kun Zhang 0040, Bo Hu 0036, Huatian Zhang 0001, Zhe Li 0028, Zhendong Mao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Difference-Aware Iterative Reasoning Network for Key Relation DetectionabstractScene graph serves as a crucial visual representation of an image, with salient objects providing richer semantics for detecting key relations. However, most methods use a one-step reasoning manner for key relation detection, which may not utilize potential clues effectively. Humans usually review and revise to achieve the final answer, and semantics of relations offer further linguistic clues. Therefore, we propose the Difference-aware Iterative Reasoning Network (DIRNet) to predict key relations in a multi-step manner. Our model estimates visual saliency, encodes contexts globally with message passing, and then refines predictions iteratively by considering the difference in predicted relation semantics and contextual information across iterations. Extensive experiments show that our model outperforms state-of-the-art methods in key relation prediction on the VG-KR benchmark, and achieves competitive results in common relation prediction on VG, demonstrating its generalization and superiority. Weidong Chen 0013, Bo Hu 0036, Hongtao Xie 0001, Zhendong Mao 0001 |
ICME | 3 |
| 2023 | Unlocking the Power of Cross-Dimensional Semantic Dependency for Image-Text MatchingabstractImage-text matching, as a fundamental cross-modal task, bridges vision and language. The key challenge lies in accurately learning the semantic similarity of these two heterogeneous modalities. To determine the semantic similarity between visual and textual features, existing paradigm typically first maps them into a d-dimensional shared representation space, then independently aggregates all dimensional correspondences of cross-modal features to reflect it, e.g., the inner product. However, in this paper, we are motivated by an insightful finding that dimensions are not mutually independent, but there are intrinsic dependencies among dimensions to jointly represent latent semantics. Ignoring this intrinsic information probably leads to suboptimal aggregation for semantic similarity, impairing cross-modal matching learning. To solve this issue, we propose a novel cross-dimensional semantic dependency-aware model (called X-Dim), which explicitly and adaptively mines the semantic dependencies between dimensions in the shared space, enabling dimensions with joint dependencies to be enhanced and utilized. X-Dim (1) designs a generalized framework to learn dimensions' semantic dependency degrees, and (2) devises the adaptive sparse probabilistic learning to autonomously make the model capture precise dependencies. Theoretical analysis and extensive experiments demonstrate the superiority of X-Dim over state-of-the-art methods, achieving 5.9%-7.3% rSum improvements on Flickr30K and MS-COCO benchmarks. Kun Zhang 0040, Lei Zhang 0119, Bo Hu 0036, Mengxiao Zhu 0001, Zhendong Mao 0001 |
ACM Multimedia | 3 |
| 2023 | GH-DDM: the generalized hybrid denoising diffusion model for medical image generation
Bo Hu 0036, Zhendong Mao 0001 |
Multim. Syst. | 3 |
| 2023 | Intra-Class Adaptive Augmentation With Neighbor Correction for Deep Metric LearningabstractDeep metric learning aims to learn an embedding space, where semantically similar samples are close together and dissimilar ones are repelled against. To explore more hard and informative training signals for augmentation and generalization, recent methods focus on generating synthetic samples to boost metric learning losses. However, these methods just use the deterministic and class-independent generations (e.g., simple linear interpolation), which only can cover the limited part of distribution spaces around original samples. They have overlooked the wide characteristic changes of different classes and can not model abundant intra-class variations for generations. Therefore, generated samples not only lack rich semantics within the certain class, but also might be noisy signals to disturb training. In this paper, we propose a novel intra-class adaptive augmentation (IAA) framework for deep metric learning. We reasonably estimate intra-class variations for every class and generate adaptive synthetic samples to support hard samples mining and boost metric learning losses. Further, for most datasets that have a few samples within the class, we propose the neighbor correction to revise the inaccurate estimations, according to our correlation discovery where similar classes generally have similar variation distributions. Extensive experiments on five benchmarks show our method significantly improves and outperforms the state-of-the-art methods on retrieval performances by 3%-6%. Zheren Fu, Zhendong Mao 0001, Bo Hu 0036, Anan Liu, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Background Layout Generation and Object Knowledge Transfer for Text-to-Image GenerationabstractText-to-Image generation (T2I) aims to generate realistic and semantically consistent images according to the natural language descriptions. Built upon the recent advances in generative adversarial networks (GANs), existing T2I models have made great process. However, a close inspection of their generated images shows two major limitations: 1) the background (e.g., fence, lake) of the generated image with the complicated, real-world scene tends to be unrealistic; 2) the object (e.g., elephant, zebra) in the generated image often presents highly distorted shape or key parts missing. To address these limitations, we propose a two-stage T2I approach, where the first stage redesigns the text-to-layout process to incorporate the background layout with the existing object layout, the second stage transfers the object knowledge from an existing class-to-image model to the layout-to-image process to improve the object fidelity. Specifically, a transformer-based architecture is introduced as the layout generator to learn the mapping from text to layout of object and background, and a Text-attended Layout-aware feature Normalization (TL-Norm) is proposed to adaptively transfer the object knowledge to the image generation. Benefitting from the background layout and transferred object knowledge, the proposed approach significantly surpasses previous state-of-the-art methods in the image quality metric and achieves superior image-text alignment performance. Zhuowei Chen, Zhendong Mao 0001, Shancheng Fang, Bo Hu 0036 |
ACM Multimedia | 4 |
| 2014 | Incentive analysis for cooperative interactive multiview video streaming
Bo Hu 0036, H. Vicky Zhao, Gene Cheung |
Signal Process. Image Commun. | 1 |
| 2013 | Optimizing peer grouping for live free viewpoint video streamingabstractIn free viewpoint video, a user can pull texture and depth videos captured from two nearby reference viewpoints to synthesize his chosen intermediate virtual view for observation via depth-image-based rendering (DIBR). For users who are observing the same video at the same time but not necessarily from the same virtual viewpoint, they have incentive to pull the same reference views so that the streaming cost can be shared. On the other hand, in general distortion of a synthesized virtual view increases with its distance to the reference views, and so a user also has incentive to select reference views that tightly “sandwich” his chosen virtual view, minimizing distortion. In a previous work, reference view sharing strategies-ones that optimally trade off shared streaming costs with synthesized view distortions-were investigated for the case when users are first divided into groups, and each user group independently pulls two reference views and shares the resulting streaming cost. In this paper, we generalize the previous notion of user group, so that a user can simultaneously belong to two groups, and each group shares the streaming cost of a single view. We also aim to find a Nash Equilibrium (NE) solution of reference view selection, which is stable and from which no one has incentive to unilaterally deviate. Specifically, we first derive a lemma based on known properties of synthesized view distortion functions. We then design a search algorithm to find a NE solution, leveraging on the derived lemma to reduce search complexity. Experimental results show that the stable NE solution increases the overall cost only slightly when compared to the unstable optimal reference selection that gives the lowest overall cost. Further, a larger network will give a lower average cost for each user, and thus, users tend to join large networks for cooperation. Yuan Yuan 0007, Bo Hu 0036, Gene Cheung, H. Vicky Zhao |
ICIP | 2 |
| 2012 | Incentive analysis for cooperative distribution of interactive multiview videoabstractIn interactive multiview video streaming (IMVS), users can periodically select one out of many captured views available for observation as video is played back in time. In single-view video streaming, to reduce server's upload burden, cooperative strategies where peers share received packets of the same video have proven to be effective, and incentive mechanisms are designed to stimulate user cooperation. Exploiting user cooperation in high dimensional IMVS, however, is more challenging. First, small number of peers in a local area are likely watching different views among large number of views available, making it difficult for a peer to find partners of the exact same view to cooperate. Second, even if a peer can identify cooperative partners of the same view, they will soon be watching different views after independent view-switching. In this paper, we study the use of a multiview video frame structure for IMVS that facilitates cooperative view switching, where even if peers are observing different views, they can nonetheless help each other. To stimulate user cooperation, we model peers' interaction as an indirect reciprocity game. Using Markov decision process (MDP) as a formalism, each peer makes distributed decisions to maximize his aggregate utilities within his lifetime. Simulation results show that when the cost to help others is much smaller than the utility gained from others' help, users fully cooperate. As the cost-to-gain ratio increases, users tend to behave differently at different views: given peers can predict their future view navigation paths probabilistically, a peer likely to enter a view-switching path not requiring others' help will have less incentive to cooperate. When the cost-to-gain ratio is very large, no users will cooperate. Bo Hu 0036, Gene Cheung, H. Vicky Zhao |
ICASSP | 1 |
| 2011 | Incentive mechanism in wireless multicastabstractIn wireless multicast systems, cooperative multicast has been shown to be effective in dealing with heterogeneous channel conditions and improving the system performance. However, this mechanism requires users' voluntary contributions, which cannot be guaranteed since users are selfish and care only about their own performance. To stimulate user cooperation, in this work, we model the interaction among users in the wireless multicast system as a multi-buyer multi-seller price-based game, where users pay to receive relay service and get paid if they forward packets to others. It is a Stackelberg game, and backward induction is used to find the perfect Nash Equilibrium. We formulate the buyers' game as an evolutionary game and derive the evolutionarily stable strategy. Our simulation results demonstrate the effectiveness of our proposed incentive mechanism. Bo Hu 0036, H. Vicky Zhao, Hai Jiang 0001 |
ICASSP | 1 |
| 2010 | Joint pollution detection and attacker identification in peer-to-peer live streamingabstractIn the emerging peer-to-peer (P2P) live streaming, users cooperate with each other to support efficient delivery of video over networks. Pollution attack is an effective attack against P2P live streaming, where attackers upload useless data to their peers, which may cause distrust among users. To resist pollution attacks and stimulate user cooperation in P2P live streaming, this paper proposes a joint pollution detection and attacker identification system, where polluted chunks are detected as early as possible and trust management is used to identify polluters. We analyze its performance and propose different schemes to address the tradeoff between pollution resistance and system overhead. Our simulation results show that the proposed system can effectively resist pollution attacks while minimizing the user's computation overhead. Bo Hu 0036, H. Vicky Zhao |
ICASSP | 1 |
| 2009 | Pollution-resistant peer-to-peer live streaming using trust managementabstractIn the emerging peer-to-peer (P2P) live streaming, users cooperate with each other to support efficient delivery of video over networks in live streaming applications. Pollution attack is an effective attack against P2P live streaming, where attackers upload bogus multimedia data to their peers. The polluted data can spread over the entire network, and cause severe quality degradation of the videos. To resist pollution attacks in P2P live streaming, this paper proposes a trust management system that identifies attackers and excludes them from further sharing of multimedia data. We investigate possible attacks against the trust management system and analyze the attack resistance of the proposed system. Our simulation results show that the proposed trust management system can efficiently detect attackers and stimulate user cooperation even under attacks. It helps users receive more clean data and improves the performance of P2P live streaming. Bo Hu 0036, H. Vicky Zhao |
ICIP | 1 |