Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hongfu Liu 0002

dblp:32/9075-2 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-4261-8154ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Trustworthy machine learning · 30% Transfer learning and domain adaptation · 18% Speech recognition and synthesis · 13%
Computer graphics and multimedia
2 papers
Audio and music processing · 100%

Topics — the 21 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
test-time adaptation
1.322024
Advancing Test-Time Adaptation in Wild Acoustic Test Settings · EMNLP 2024
Extrapolative Continuous-time Bayesian Neural Network for Fast Training-free Test-time Adaptation · NeurIPS 2022
Machine learning › Trustworthy machine learning › calibration
confidence calibration
0.912025
On Calibration of LLM-based Guard Models for Reliable Content Moderation · ICLR 2025
Machine learning › Trustworthy machine learning
content moderation
0.912025
On Calibration of LLM-based Guard Models for Reliable Content Moderation · ICLR 2025
Machine learning › Trustworthy machine learning
guard models
0.912025
On Calibration of LLM-based Guard Models for Reliable Content Moderation · ICLR 2025
Machine learning › Trustworthy machine learning
uncertainty and calibration
0.912025
On Calibration of LLM-based Guard Models for Reliable Content Moderation · ICLR 2025
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.812024
Advancing Test-Time Adaptation in Wild Acoustic Test Settings · EMNLP 2024
Natural language and speech › Information extraction and text analysis › discourse analysis
discourse relation recognition
0.812024
Discursive Socratic Questioning: Evaluating the Faithfulness of Language Models' Understanding of Discourse Relations · ACL (1) 2024
Natural language and speech › Language models and text generation › evaluation of language models
faithfulness evaluation
0.812024
Discursive Socratic Questioning: Evaluating the Faithfulness of Language Models' Understanding of Discourse Relations · ACL (1) 2024
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › robust speech recognition
noise-robust speech recognition
0.812024
Advancing Test-Time Adaptation in Wild Acoustic Test Settings · EMNLP 2024
Natural language and speech › Question answering and dialogue systems
question-answering-based evaluation
0.812024
Discursive Socratic Questioning: Evaluating the Faithfulness of Language Models' Understanding of Discourse Relations · ACL (1) 2024
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
0.612022
Extrapolative Continuous-time Bayesian Neural Network for Fast Training-free Test-time Adaptation · NeurIPS 2022
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.612022
Extrapolative Continuous-time Bayesian Neural Network for Fast Training-free Test-time Adaptation · NeurIPS 2022
Machine learning › Deep learning architectures and training › neural differential equations
neural ordinary differential equations
0.512021
STRODE: Stochastic Boundary Ordinary Differential Equation · ICML 2021
Machine learning › Time series and sequential data
time series modeling
0.512021
STRODE: Stochastic Boundary Ordinary Differential Equation · ICML 2021
Audio and music processing
speech recognition
0.512021
STRODE: Stochastic Boundary Ordinary Differential Equation · ICML 2021
Audio and music processing
music generation
0.412019
Mind Band: A Crossmedia AI Music Composing Platform · ACM Multimedia 2019
Security and privacy of machine learning › adversarial attack
jailbreak attack
0.312025
On Calibration of LLM-based Guard Models for Reliable Content Moderation · ICLR 2025
Natural language and speech › Information extraction and text analysis › discourse analysis › discourse processing
discourse interpretation
0.212024
Discursive Socratic Questioning: Evaluating the Faithfulness of Language Models' Understanding of Discourse Relations · ACL (1) 2024
Machine learning › Transfer learning and domain adaptation › model adaptation
training-free adaptation
0.212022
Extrapolative Continuous-time Bayesian Neural Network for Fast Training-free Test-time Adaptation · NeurIPS 2022
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.112021
STRODE: Stochastic Boundary Ordinary Differential Equation · ICML 2021
Machine learning › Generative modeling
variational autoencoder
0.112019
Mind Band: A Crossmedia AI Music Composing Platform · ACM Multimedia 2019

Methods — techniques the papers use, named apart from their topics

temperature scaling · 1.7post-hoc calibration · 1.7contextual calibration · 1.7in-context learning · 0.8fine-tuning · 0.8counterfactual questioning · 0.8consistency regularization · 0.8confidence-aware weighting · 0.8predictive modeling · 0.6bayesian filtering · 0.6variational inference · 0.5stochastic differential equation · 0.5point process · 0.5variational autoencoder · 0.4generative adversarial network · 0.4emotion analysis · 0.4
YearPublicationVenuePosition
2025 On Calibration of LLM-based Guard Models for Reliable Content Moderation
abstract
Large language models (LLMs) pose significant risks due to the potential for generating harmful content or users attempting to evade guardrails. Existing studies have developed LLM-based guard models designed to moderate the input and output of threat LLMs, ensuring adherence to safety policies by blocking content that violates these protocols upon deployment. However, limited attention has been given to the reliability and calibration of such guard models. In this work, we empirically conduct comprehensive investigations of confidence calibration for 9 existing LLM-based guard models on 12 benchmarks in both user input and model output classification. Our findings reveal that current LLM-based guard models tend to 1) produce overconfident predictions, 2) exhibit significant miscalibration when subjected to jailbreak attacks, and 3) demonstrate limited robustness to the outputs generated by different types of response models. Additionally, we assess the effectiveness of post-hoc calibration methods to mitigate miscalibration. We demonstrate the efficacy of temperature scaling and, for the first time, highlight the benefits of contextual calibration for confidence calibration of guard models, particularly in the absence of validation sets. Our analysis and experiments underscore the limitations of current LLM-based guard models and provide valuable insights for the future development of well-calibrated guard models toward more reliable content moderation. We also advocate for incorporating reliability evaluation of confidence calibration when releasing future LLM-based guard models.
Hongfu Liu 0002, Hengguan Huang, Xiangming Gu, Hao Wang 0014, Ye Wang 0007
ICLR1
2025 Interpretable Novel Target Discovery through Open-Set Domain Adaptation
abstract
Open-set domain adaptation (OSDA) considers a special domain adaptation problem in which the target domain contains novel categories that never appear in the well-labeled source domain. Unfortunately, prior efforts on OSDA simply detect and recognize all novel categories as one “unknown” group without further exploration. The demand for exploring these novel categories prompts us to consider the underlying multi-class structure and semantic description of those unknown categories in more detail. In this article, we propose a novel interpretable framework to accurately identify the seen categories in the target domain and effectively recover the semantic knowledge of the unseen categories with attributes and visual interpretations, which is referred to as Semantic Recovery Open-Set Domain Adaptation (SR-OSDA). Specifically, the proposed framework includes an explicit attribute explainable module and an implicit semantic interpretable module, which provide insight into the process of domain adaptation and the discovery of new categories. Furthermore, structure-preserving partial alignment is developed as a method of recognizing and aligning the visible categories across domains with the aid of domain-invariant feature learning. The visual-structural semantic attributes propagation is designed to provide smooth transitions from seen categories to unseen categories via visual-semantic mapping. Three new cross-domain SR-OSDA benchmarks are constructed in order to evaluate the proposed framework in novel and practical challenges. Experimental results and empirical analysis of our proposed solution to open-set recognition and semantic recovery demonstrate its superiority over other state-of-the-art solutions. Our source code is available at https://github.com/scottjingtt/XSROSDA .
Taotao Jing, Haifeng Xia, Hongfu Liu 0002, Zhengming Ding
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Discursive Socratic Questioning: Evaluating the Faithfulness of Language Models' Understanding of Discourse Relations
abstract
While large language models have significantly enhanced the effectiveness of discourse relation classifications, it remains unclear whether their comprehension is faithful and reliable.We provide DISQ, a new method for evaluating the faithfulness of understanding discourse based on question answering.We first employ incontext learning to annotate the reasoning for discourse comprehension, based on the connections among key events within the discourse.Following this, DISQ interrogates the model with a sequence of questions to assess its grasp of core event relations, its resilience to counterfactual queries, as well as its consistency to its previous responses.We then evaluate language models with different architectural designs using DISQ, finding: (1) DISQ presents a significant challenge for all models, with the top-performing GPT model attaining only 41% of the ideal performance in PDTB; (2) DISQ is robust to domain shifts and paraphrase variations; (3) Open-source models generally lag behind their closed-source GPT counterparts, with notable exceptions being those enhanced with chat and code/math features; (4) Our analysis validates the effectiveness of explicitly signalled discourse connectives, the role of contextual information, and the benefits of using historical QA data. Discourse relation: Contingency.Cause.ResultIs "they keep changing their prices" a reason for "it's very frustrating"?Model's Answer: True Is "they keep changing their prices" contrasted with "it's very frustrating"?Is "it's very frustrating" the result of "they keep changing their prices"?🤖 Targeted Score = 1 Counterfactual Score = 0 Consistency Score = 0Arg2: It's very frustrating.Arg1: When I want to buy, they run from you --they keep changing their prices.
Yisong Miao, Hongfu Liu 0002, Wenqiang Lei, Nancy F. Chen, Min-Yen Kan
ACL (1)2
2024 Advancing Test-Time Adaptation in Wild Acoustic Test Settings
abstract
Acoustic foundation models, fine-tuned for Automatic Speech Recognition (ASR), suffer from performance degradation in wild acoustic test settings when deployed in real-world scenarios.Stabilizing online Test-Time Adaptation (TTA) under these conditions remains an open and unexplored question.Existing wild vision TTA methods often fail to handle speech data effectively due to the unique characteristics of high-entropy speech frames, which are unreliably filtered out even when containing crucial semantic content.Furthermore, unlike static vision data, speech signals follow short-term consistency, requiring specialized adaptation strategies.In this work, we propose a novel wild acoustic TTA method tailored for ASR fine-tuned acoustic foundation models.Our method, Confidence-Enhanced Adaptation, performs frame-level adaptation using a confidence-aware weight scheme to avoid filtering out essential information in high-entropy frames.Additionally, we apply consistency regularization during test-time optimization to leverage the inherent short-term consistency of speech signals.Our experiments on both synthetic and real-world datasets demonstrate that our approach outperforms existing baselines under various wild acoustic test settings, including Gaussian noise, environmental sounds, accent variations, and sung speech 1 .
Hongfu Liu 0002, Hengguan Huang, Ye Wang 0007
EMNLP1
2023 Zero-Shot Automatic Pronunciation Assessment
Hongfu Liu 0002, Mingqian Shi, Ye Wang 0007
INTERSPEECH1
2022 Extrapolative Continuous-time Bayesian Neural Network for Fast Training-free Test-time Adaptation
abstract
Human intelligence has shown remarkably lower latency and higher precision than most AI systems when processing non-stationary streaming data in real-time. Numerous neuroscience studies suggest that such abilities may be driven by internal predictive modeling. In this paper, we explore the possibility of introducing such a mechanism in unsupervised domain adaptation (UDA) for handling non-stationary streaming data for real-time streaming applications. We propose to formulate internal predictive modeling as a continuous-time Bayesian filtering problem within a stochastic dynamical system context. Such a dynamical system describes the dynamics of model parameters of a UDA model evolving with non-stationary streaming data. Building on such a dynamical system, we then develop extrapolative continuous-time Bayesian neural networks (ECBNN), which generalize existing Bayesian neural networks to represent temporal dynamics and allow us to extrapolate the distribution of model parameters before observing the incoming data, therefore effectively reducing the latency. Remarkably, our empirical results show that ECBNN is capable of continuously generating better distributions of model parameters along the time axis given historical data only, thereby achieving (1) training-free test-time adaptation with low latency, (2) gradually improved alignment between the source and target features and (3) gradually improved model performance over time during the real-time testing stage.
Hengguan Huang, Xiangming Gu, Hao Wang 0014, Chang Xiao 0004, Hongfu Liu 0002, Ye Wang 0007
NeurIPS5
2021 STRODE: Stochastic Boundary Ordinary Differential Equation
abstract
Perception of time from sequentially acquired sensory inputs is rooted in everyday behaviors of individual organisms. Yet, most algorithms for time-series modeling fail to learn dynamics of random event timings directly from visual or audio inputs, requiring timing annotations during training that are usually unavailable for real-world applications. For instance, neuroscience perspectives on postdiction imply that there exist variable temporal ranges within which the incoming sensory inputs can affect the earlier perception, but such temporal ranges are mostly unannotated for real applications such as automatic speech recognition (ASR). In this paper, we present a probabilistic ordinary differential equation (ODE), called STochastic boundaRy ODE (STRODE), that learns both the timings and the dynamics of time series data without requiring any timing annotations during training. STRODE allows the usage of differential equations to sample from the posterior point processes, efficiently and analytically. We further provide theoretical guarantees on the learning of STRODE. Our empirical results show that our approach successfully infers event timings of time series data. Our method achieves competitive or superior performances compared to existing state-of-the-art methods for both synthetic and real-world datasets.
Hengguan Huang, Hongfu Liu 0002, Hao Wang 0014, Chang Xiao 0004, Ye Wang 0007
ICML2
2019 Mind Band: A Crossmedia AI Music Composing Platform
abstract
Various media information in life can have an important impact on our understanding of music. In this paper, we present a demo, Mind Band, which is a Cross-Media artificial intelligent composing platform using our life elements such as emoji, image and humming. In practice, we base our system on the valence-arousal model. We use emotion analysis of life elements to map them to music pieces, which are generated by a Variational Autoencoder - Generative Adversarial Networks model. We provide users with immersive experience by uploading emoji/image/humming and retrieving emotionally related music pieces back. With this platform, everyone can be a composer.
Zhaolin Qiu, Yufan Ren, Canchen Li, Hongfu Liu 0002, Songruoyao Wu, Hanjia Zheng, Juntao Ji, Jianjia Yu
ACM Multimedia4