EDBT 2026 Demo / reviewers in the wild / expert
Hong Liu 0007
dblp:29/5010-7
· DBLP profile ↗
28ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-4524-495XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 11 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language ModelsabstractRecent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately understand timestamps for temporal localization (e.g., Temporal Audio Grounding) and are restricted to short audio perception, leading to constrained capabilities on fine-grained tasks. We identify three key aspects that limit their temporal localization and long audio understanding: (i) timestamp representation, (ii) architecture, and (iii) data. To address this, we introduce TimeAudio, a novel method that empowers LALMs to connect their understanding of audio content with precise temporal perception. Specifically, we incorporate unique temporal markers to improve time-sensitive reasoning and apply an absolute time-aware encoding that explicitly grounds the acoustic features with absolute time information. Moreover, to realize end-to-end long audio understanding, we introduce a segment-level token merging module to substantially reduce audio token redundancy and enhance the efficiency of information extraction. Due to the lack of suitable datasets and evaluation metrics, we consolidate existing audio datasets into a new dataset focused on temporal tasks and establish a series of metrics to evaluate the fine-grained performance. Evaluations show strong performance across a variety of fine-grained tasks, such as dense captioning, temporal grounding, and timeline speech summarization, which demonstrates TimeAudio's robust temporal localization and reasoning capabilities. Hualei Wang, Hong Liu 0007 |
AAAI | 4 |
| 2025 | Diverse Audio Caption Generation with Semantic-aware Diffusion ModelabstractAudio captioning aims to perceive sound events in different ways and describe an audio clip from various perspectives. Most existing audio captioning methods tend to generate captions that are deterministic and simple, lacking diversity and limiting their applicability in real-world scenarios. Recently, diffusion-based methods have achieved significant progress in producing diverse captions, but with a potential compromise in accuracy and fluency due to the abstract nature of language and the variable supervised targets. In this work, we propose a semantic-aware diffusion model that leverages its intrinsic stochastic sampling and global context to generate captions. In order to maintain accuracy, we integrate the global CLAP embeddings into the denoising process of the diffusion model to serve as semantic context. To further enhance diversity, we propose an optimized inference process that incorporates a dynamic denoising strategy during the token-restored stage. Extensive experiments on the AudioCaps and Clotho dataset demonstrate that our model achieves superior results on accuracy and diversity metrics compared to state-of-the-art diverse audio caption methods. Hualei Wang, Hong Liu 0007 |
ICME | 3 |
| 2024 | Semi-Supervised Sound Event Detection with Local and Global Consistency RegularizationabstractLearning meaningful frame-wise features on a partially labeled dataset is crucial to semi-supervised sound event detection. Prior works either maintain consistency on frame-level predictions or seek feature-level similarity among neighboring frames, which cannot exploit the potential of unlabeled data. In this work, we design a Local and Global Consistency (LGC) regularization scheme to enhance the model on both label- and feature-level. The audio CutMix is introduced to change the contextual information of clips. Then, the local consistency is adopted to encourage the model to leverage local features for frame-level predictions, and the global consistency is applied to force features to align with global prototypes through a specially designed contrastive loss. Experiments on the DESED dataset indicate the superiority of LGC, surpassing its respective competitors largely under the same settings. Besides, combining LGC with existing methods can obtain further improvements. The code is available at https://github.com/Ming-er/LGC-SED. Hong Liu 0007, Kazushige Ouchi |
ICASSP | 3 |
| 2024 | Leveraging Language Model Capabilities for Sound Event Detection
Hualei Wang, Jianguo Mao, Zhifang Guo, Jiarui Wan, Hong Liu 0007 |
INTERSPEECH | 5 |
| 2024 | Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-trainingabstractRecent advances have been witnessed in audio-language joint learning, such as CLAP, that shows much success in multi-modal understanding tasks. These models usually aggregate uni-modal local representations, namely frame or word features, into global ones, on which the contrastive loss is employed to reach coarse-grained cross-modal alignment. However, frame-level correspondence with texts may be ignored, making it ill-posed on explainability and fine-grained challenges which may also undermine performances on coarse-grained tasks. In this work, we aim to improve both coarse- and fine-grained audio-language alignment in large-scale contrastive pre-training. To unify the granularity and latent distribution of two modalities, a shared codebook is adopted to represent multi-modal global features with common bases, and each codeword is regularized to encode modality-shared semantics, bridging the gap between frame and word features. Based on it, a locality-aware block is involved to purify local patterns, and a hard-negative guided loss is devised to boost alignment. Experiments on eleven zero-shot coarse- and fine-grained tasks suggest that our model not only surpasses the baseline CLAP significantly but also yields superior or competitive results compared to current SOTA works. Zhifang Guo, Hong Liu 0007 |
ACM Multimedia | 4 |
| 2024 | AIROGS: Artificial Intelligence for Robust Glaucoma Screening ChallengeabstractThe early detection of glaucoma is essential in preventing visual impairment. Artificial intelligence (AI) can be used to analyze color fundus photographs (CFPs) in a cost-effective manner, making glaucoma screening more accessible. While AI models for glaucoma screening from CFPs have shown promising results in laboratory settings, their performance decreases significantly in real-world scenarios due to the presence of out-of-distribution and low-quality images. To address this issue, we propose the Artificial Intelligence for Robust Glaucoma Screening (AIROGS) challenge. This challenge includes a large dataset of around 113,000 images from about 60,000 patients and 500 different screening centers, and encourages the development of algorithms that are robust to ungradable and unexpected input data. We evaluated solutions from 14 teams in this paper and found that the best teams performed similarly to a set of 20 expert ophthalmologists and optometrists. The highest-scoring team achieved an area under the receiver operating characteristic curve of 0.99 (95% CI: 0.98-0.99) for detecting ungradable images on-the-fly. Additionally, many of the algorithms showed robust performance when tested on three other publicly available datasets. These results demonstrate the feasibility of robust AI-enabled glaucoma screening. Coen de Vente, Koen A. Vermeer, Nicolas Jaccard, He Wang 0016, Hongyi Sun, Firas Khader, Daniel Truhn, Temirgali Aimyshev, Yerkebulan Zhanibekuly, Tien-Dung Le, Adrian Galdran, Miguel Ángel González Ballester, Gustavo Carneiro 0001, Devika R. G., Hrishikesh Panikkasseril Sethumadhavan, Densen Puthussery, Hong Liu 0007, Zekang Yang, Satoshi Kondo, Satoshi Kasai, Ashritha Durvasula, Jónathan Heras, Miguel Ángel Zapata, Teresa Araujo, Guilherme Aresta, Hrvoje Bogunovic, Mustafa Arikan, Yeong Chan Lee, Hyun Bin Cho, Yoon Ho Choi, Abdul Qayyum 0002, Muhammad Imran Razzak, Bram van Ginneken, Hans G. Lemij, Clara I. Sánchez |
IEEE Trans. Medical Imaging | 17 |
| 2023 | Inferential Knowledge-Enhanced Integrated Reasoning for Video Question AnsweringabstractRecently, video question answering has attracted growing attention. It involves answering a question based on a fine-grained understanding of video multi-modal information. Most existing methods have successfully explored the deep understanding of visual modality. We argue that a deep understanding of linguistic modality is also essential for answer reasoning, especially for videos that contain character dialogues. To this end, we propose an Inferential Knowledge-Enhanced Integrated Reasoning method. Our method consists of two main components: 1) an Inferential Knowledge Reasoner to generate inferential knowledge for linguistic modality inputs that reveals deeper semantics, including the implicit causes, effects, mental states, etc. 2) an Integrated Reasoning Mechanism to enhance video content understanding and answer reasoning by leveraging the generated inferential knowledge. Experimental results show that our method achieves significant improvement on two mainstream datasets. The ablation study further demonstrates the effectiveness of each component of our approach. Jianguo Mao, Wenbin Jiang 0002, Hong Liu 0007, Yajuan Lyu |
AAAI | 3 |
| 2022 | Hierarchical Representation-based Dynamic Reasoning Network for Biomedical Question AnsweringabstractRecently, Biomedical Question Answering (BQA) has attracted growing attention due to its application value and technical challenges. Most existing works treat it as a semantic matching task that predicts answers by computing confidence among questions, options and evidence sentences, which is insufficient for scenarios that require complex reasoning based on a deep understanding of biomedical evidences. We propose a novel model termed Hierarchical Representation-based Dynamic Reasoning Network (HDRN) to tackle this problem. It first constructs the hierarchical representations for biomedical evidences to learn semantics within and among evidences. It then performs dynamic reasoning based on the hierarchical representations of evidences to solve complex biomedical problems. Against the existing state-of-the-art model, the proposed model significantly improves more than 4.5%, 3% and 1.3% on three mainstream BQA datasets, PubMedQA, MedQA-USMLE and NLPEC. The ablation study demonstrates the superiority of each improvement of our model. The code will be released after the paper is published. Jianguo Mao, Zengfeng Zeng, Weihua Peng, Wenbin Jiang 0002, Hong Liu 0007, Yajuan Lyu |
COLING | 7 |
| 2022 | Explainable Question Answering based on Semantic Graph by Global Differentiable Learning and Dynamic Adaptive ReasoningabstractMulti-hop Question Answering is an agent task for testing the reasoning ability.With the development of pre-trained models, the implicit reasoning ability has been surprisingly improved and can even surpass human performance.However, the nature of the black box hinders the construction of explainable intelligent systems.Several researchers have explored explainable neural-symbolic reasoning methods based on question decomposition techniques.The undifferentiable symbolic operations and the error propagation in the reasoning process lead to poor performance.To alleviate it, we propose a simple yet effective Global Differentiable Learning strategy to explore optimal reasoning paths from the latent probability space so that the model learns to solve intermediate reasoning processes without expert annotations.We further design a Dynamic Adaptive Reasoner to enhance the generalization of unseen questions.Our method achieves 17% improvements in F1-score against BreakRC and shows better interpretability.We take a step forward in building interpretable reasoning methods. Jianguo Mao, Wenbin Jiang 0002, Hong Liu 0007, Yajuan Lyu, Qiaoqiao She |
EMNLP | 4 |
| 2022 | MAL: Multi-modal Attention Learning for Tumor Diagnosis Based on Bipartite Graph and Multiple Branches
Menglei Jiao, Hong Liu 0007, Jianfang Liu, Hanqiang Ouyang, Huishu Yuan, Yueliang Qian |
MICCAI (3) | 2 |
| 2022 | Dynamic Multistep Reasoning based on Video Scene Graph for Video Question AnsweringabstractJianguo Mao, Wenbin Jiang, Xiangdong Wang, Zhifan Feng, Yajuan Lyu, Hong Liu, Yong Zhu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jianguo Mao, Wenbin Jiang 0002, Zhifan Feng, Yajuan Lyu, Hong Liu 0007, Yong Zhu 0004 |
NAACL-HLT | 6 |
| 2021 | Speech Synthesis of Chinese Braille with Limited Training DataabstractThis paper describes to our knowledge the first Chinese Braille speech synthesis system. The system consists of modules of Braille front-end processing, prosody prediction, and speech synthesis. The Braille front-end processing includes conversion from the common Braille to Pinyin, and a high-precision Chinese character prediction model. To achieve high precision prosody prediction under limited corpus conditions, we propose a prosody prediction model based on the RoBERTa pre-trained model, which achieves an accuracy of 94.42%. Finally, a real-time TTS system based on Tacotron2 and LPCNet is proposed. We modify Tacotron2, including introducing a forward attention mechanism and extending the autoregressive correlation step size to obtain more natural speech. Jianguo Mao, Hong Liu 0007, Yueliang Qian |
ICME | 4 |
| 2020 | Multi-Branch Learning for Weakly-Labeled Sound Event DetectionabstractThere are two sub-tasks implied in the weakly-supervised SED: audio tagging and event boundary detection. Current methods which combine multi-task learning with SED requires annotations both for these two sub-tasks. Since there are only annotations for audio tagging available in weakly-supervised SED, we design multiple branches with different learning purposes instead of pursuing multiple tasks. Similar to multiple tasks, multiple different learning purposes can also prevent the common feature which the multiple branches share from overfitting to any one of the learning purposes. We design these multiple different learning purposes based on combinations of different MIL strategies and different pooling methods. Experiments on the DCASE 2018 Task 4 dataset and the URBAN-SED dataset both show that our method achieves competitive performance. Hong Liu 0007, Yueliang Qian |
ICASSP | 4 |
| 2020 | Guided Learning for Weakly-Labeled Semi-Supervised Sound Event DetectionabstractWe propose a simple but efficient method termed Guided Learning for weakly-labeled semi-supervised sound event detection (SED). There are two sub-targets implied in weakly-labeled SED: audio tagging and boundary detection. Instead of designing a single model by considering a trade-off between the two sub-targets, we design a teacher model aiming at audio tagging to guide a student model aiming at boundary detection to learn using the unlabeled data. The guidance is guaranteed by the audio tagging performance gap of the two models. In the meantime, the student model liberated from the trade-off is able to provide more excellent boundary detection results. We propose a principle to design such two models based on the relation between the temporal compression scale and the two sub-targets. We also propose an end-to-end semi-supervised learning process for these two models to enable their abilities to rise alternately. Experiments on the DCASE2018 Task4 dataset show that our approach achieves competitive performance. Hong Liu 0007, Yueliang Qian |
ICASSP | 3 |
| 2020 | Specialized Decision Surface and Disentangled Feature for Weakly-Supervised Polyphonic Sound Event DetectionabstractIn this article, a special decision surface for the weakly-supervised sound event detection (SED) and a disentangled feature (DF) for the multi-label problem in polyphonic SED are proposed. We approach SED as a multiple instance learning (MIL) problem and utilize a neural network framework with a pooling module to solve it. General MIL approaches include two kinds: the instance-level approaches and embedding-level approaches. We present a method of generating instance-level probabilities for the embedding level approaches which tend to perform better than the instance-level approaches in terms of bag-level classification but can not provide instance-level probabilities in current approaches. Moreover, we further propose a specialized decision surface (SDS) for the embedding-level attention pooling. We analyze and explained why an embedding-level attention module with SDS is better than other typical pooling modules from the perspective of the high-level feature space. As for the problem of the unbalanced dataset and the co-occurrence of multiple categories in the polyphonic event detection task, we propose a DF to reduce interference among categories, which optimizes the high-level feature space by disentangling it based on class-wise identifiable information and obtaining multiple different subspaces. Experiments on the dataset of DCASE 2018 Task 4 show that the proposed SDS and DF significantly improve the detection performance of the embedding-level MIL approach with an attention pooling module and outperform the first place system in the challenge by $\mathbf {6.6}$ percentage points. Hong Liu 0007, Yueliang Qian |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | CBConv: Service for Automatic Conversion of Chinese Characters into Braille with High AccuracyabstractThe conversion of Chinese Characters into Braille faces the challenges of Braille word segmentation, pronunciation determination, and tone marking. In this paper, we present the CBConv service for automatic conversion of Chinese characters into Braille with high accuracy. The service supports two modes: real-time conversion of Chinese strings and asynchronous conversion of documents. To address the challenges described above, an end-to-end deep learning framework is proposed, which can perform pronunciation determination, word segmentation and tone marking simultaneously using a deep neural network. Accuracy evaluation and user study are conducted, which demonstrate the usability of the proposed system. Jinghua Zhong, Hong Liu 0007, Yueliang Qian |
ASSETS | 4 |
| 2019 | Effective Optical Braille Recognition Based on Two-Stage Learning for Double-Sided Braille Image
Renqiang Li, Hong Liu 0007, Yueliang Qian |
PRICAI (3) | 2 |
| 2018 | Multiple Visual Fields Cascaded Convolutional Neural Network for Breast Cancer Detection
Haomiao Ni, Hong Liu 0007, Zichao Guo, Taijiao Jiang, Kuansong Wang, Yueliang Qian |
PRICAI (1) | 2 |
| 2018 | RGB-D joint modelling with scene geometric information for indoor semantic segmentation
Hong Liu 0007, Wenshan Wu, Yueliang Qian |
Multim. Tools Appl. | 1 |
| 2017 | Efficient Multi-scale Plane Extraction Based RGBD Video Segmentation
Hong Liu 0007, Yueliang Qian |
MMM (1) | 1 |
| 2017 | Improving speech transcription by exploiting user feedback and word repetition
Hong Liu 0007, Yueliang Qian |
Multim. Tools Appl. | 3 |
| 2016 | Action Recognition Based on Optimal Joint Selection and Discriminative Depth Descriptor
Haomiao Ni, Hong Liu 0007, Yueliang Qian |
ACCV (2) | 2 |
| 2016 | A novel obstacle detection method based on distortion of laser patternabstractMost visual based obstacle detection methods usually detect obstacles on the whole image, which are influenced by complex background and various illuminations. This paper proposes a novel and simple obstacle detection method DLP based on analyzing Distortion of Laser Pattern. To reduce computation, a laser pattern with two cross lines is used to project onto the front ground. Firstly, a segmentation algorithm based on color and edge is proposed to extract laser pixels. Then, skeletons of laser pattern are quickly extracted using table scanning strategy to further reduce noise. Secondly, seed-filling is used to extract the region of skeletons, and then laser lines are detected by Hough and merged by region based rule. Finally, according to the distortion of laser pattern projected on different obstacles, we propose a hierarchical discriminant model, which can classify the front scene as free or having obstacles. And the obstacle can be further classified as raised or sunken obstacle, and ascending or descending stairs. Experimental results show the robustness and effectiveness of our DLP method in different scenes, including indoor and outdoor, day and night with 10f/s. Zichao Guo, Hong Liu 0007, Yueliang Qian |
ICME | 2 |
| 2014 | Segment and Label Indoor Scene Based on RGB-D for the Visually Impaired
Hong Liu 0007, Yueliang Qian |
MMM (1) | 2 |
| 2013 | Related HOG Features for Human Detection Using Cascaded Adaboost and SVM Classifiers
Hong Liu 0007, Tao Xu 0029, Yueliang Qian |
MMM (2) | 1 |
| 2012 | A Fast and Robust Pedestrian Detection Framework Based on Static and Dynamic InformationabstractWith the powerful development of pedestrian detection technique based on sliding-window and machine-learning, detection-based tracking systems have become increasingly popular. Most of these systems rely on existing static pedestrian detectors only despite the obvious potential motion information for people detection. This paper proposes a novel pedestrian detection framework fusing static and dynamic features. Motion cue is firstly used to detect potential pedestrian regions. Secondly, static detector scans potential regions to get candidate pedestrian detections. Final detection results are improved by removing false detections based on their motion distribution. The proposed framework significantly raises detection speed and detection performance. Static detector of pedestrian in this paper is trained by AdaBoost with simplified HOG feature (1HOG). Additionally, we introduce a detection-window-pyramid based scanning strategy for quickly extracting 1HOG features. The experimental results on several public data sets show the effectiveness of the proposed approach. Tao Xu 0029, Hong Liu 0007, Yueliang Qian |
ICME | 2 |
| 2007 | HTRDP evaluations on Chinese information processing and intelligent human-machine interface
Qun Liu 0001, Hong Liu 0007, Le Sun 0001, Sheng Tang, Deyi Xiong, Hongxu Hou, Yuanhua Lv, Shouxun Lin, Yueliang Qian |
Frontiers Comput. Sci. China | 3 |
| 2001 | FaceTracker: A Human Face Tracking and Facial Organ Localizing SystemabstractThis paper presents a tracking system based on the technique of gravity-center template to track faces and localize facial organs such as pupil, eyes, nose and mouth. Hongming Zhang 0011, Wen Gao 0001, Hong Liu 0007, Xilin Chen 0001 |
ICCV | 5 |