Haotian Huang

dblp:332/1287 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 51% Representation and self-supervised learning · 31% Deep learning architectures and training · 18%
Databases, data mining, and information retrieval
1 paper
Graph data management · 50% Data mining · 50%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › structured data mining › graph mining
community detection
1.012026
Biclique Percolation Communities Computation on Temporal Bipartite Graphs · IEEE Trans. Knowl. Data Eng. 2026
Graph data management › temporal graph
temporal bipartite graph
1.012026
Biclique Percolation Communities Computation on Temporal Bipartite Graphs · IEEE Trans. Knowl. Data Eng. 2026
Natural language and speech › Language models and text generation › compositional generalization
length generalization
0.912025
Long-Short Alignment for Effective Long-Context Modeling in LLMs · ICML 2025
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization
long-context modeling
0.912025
Long-Short Alignment for Effective Long-Context Modeling in LLMs · ICML 2025
Machine learning › Deep learning architectures and training › regularization
training regularization
0.912025
Long-Short Alignment for Effective Long-Context Modeling in LLMs · ICML 2025
Machine learning › Representation and self-supervised learning › pre-training
autoregressive pre-training
0.812024
Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining · ICML 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
generative self-supervised learning
0.812024
Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining · ICML 2024
Natural language and speech › Language models and text generation › large language model training › language model pretraining
masked pre-training
0.812024
Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining · ICML 2024

Methods — techniques the papers use, named apart from their topics

biclique percolation · 1.0regularization · 0.9misalignment metric · 0.9variable-length masked objective · 0.8theoretical comparison · 0.8diversity-enhanced autoregressive objective · 0.8
YearPublicationVenuePosition
2026 SemTaint: A scalable taint analysis approach for JavaWeb frameworks and composite containers
Haotian Huang, Ruibin Yan
Comput. Secur.1
2026 Biclique Percolation Communities Computation on Temporal Bipartite Graphs
Zi Chen 0003, Haotian Huang, Long Yuan 0001, Jianqiu Xu, Bolong Zheng, Xuemin Lin 0001
IEEE Trans. Knowl. Data Eng.2
2025 Long-Short Alignment for Effective Long-Context Modeling in LLMs
abstract
Large language models (LLMs) have exhibited impressive performance and surprising emergent properties. However, their effectiveness remains limited by the fixed context window of the transformer architecture, posing challenges for long-context modeling. Among these challenges, length generalization — the ability to generalize to sequences longer than those seen during training — is a classical and fundamental problem. In this work, we propose a fresh perspective on length generalization, shifting the focus from the conventional emphasis on input features such as positional encodings or data structures to the output distribution of the model. Specifically, through case studies on synthetic tasks, we highlight the critical role of **long-short alignment** — the consistency of output distributions across sequences of varying lengths. Extending this insight to natural language tasks, we propose a metric called Long-Short Misalignment to quantify this phenomenon, uncovering a strong correlation between the metric and length generalization performance. Building on these findings, we develop a regularization term that promotes long-short alignment during training. Extensive experiments validate the effectiveness of our approach, offering new insights for achieving more effective long-context modeling in LLMs. Code is available at https://github.com/PKU-ML/LongShortAlignment.
Tianqi Du, Haotian Huang, Yifei Wang 0001, Yisen Wang 0001
ICML2
2025 A Safety-Critical Dynamic System Framework for High-Precision Learning From Demonstration
abstract
Stable dynamic systems enable robotic systems to plan and execute complex geometric motions in unstructured environments. However, in certain scenarios, such as precision assembly tasks, the constrained operational workspace of robotic arms, along with the presence of obstacles, may lead to unintended collisions. Furthermore, motion precision plays a crucial role in determining the success rate of such tasks. To address these challenges, we propose a novel Safe-Critical Dynamic System (SC-DS) framework. The SC-DS framework consists of a stable dynamic system and an obstacle avoidance controller. The stable dynamic system is formulated in a parametric nonlinear form, which enhances performance in terms of accuracy. Additionally, control barrier functions (CBF), corresponding to complex constraint spaces, are learned from demonstration data. By utilizing these learned CBFs, an obstacle avoidance controller is designed to ensure that the system trajectory remains within the learned safety boundaries. Moreover, the controller adaptively extends the effective range of control inputs, thus mitigating replication errors due to input limitations. Experimental results, both in simulation and with a physical robot, demonstrate that the SC-DS framework effectively reproduces trajectories with both stability and safety, outperforming existing methods in terms of overall task performance.Note to Practitioners—This study is motivated by the need to develop a safer and more precise skill-learning framework for practical applications, such as service robots and assembly robots. We propose the SC-DS framework, which integrates the challenges of uncertain environments into the learning of precise motion skills, ensuring both accuracy and safety in trajectory generation. This framework is particularly suitable for applications requiring strict performance and safety standards. By incorporating safety constraints directly into the learning process, our method provides a robust solution that effectively addresses robot-environment interactions, including obstacles, disturbances, and varying conditions. Our research enhances the reliability of dynamic system-based learning frameworks and offers practical tools for real-world applications, ensuring robots perform tasks efficiently while maintaining safety. Practitioners can leverage this framework to improve safety and maintain high-precision skill learning, ensuring effective handling of real-world challenges.
Haotian Huang, Jiayun Fu, Zhehao Jin, Andong Liu, Wen-An Zhang 0001, Chenguang Yang 0001
IEEE Trans Autom. Sci. Eng.1
2024 Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining
abstract
In recent years, the rise of generative self-supervised learning (SSL) paradigms has exhibited impressive performance across visual, language, and multi-modal domains. While the varied designs of generative SSL objectives lead to distinct properties in downstream tasks, a theoretical understanding of these differences remains largely unexplored. In this paper, we establish the first theoretical comparisons between two leading generative SSL paradigms: autoregressive SSL and masked SSL. Through establishing theoretical frameworks, we elucidate the strengths and limitations of autoregressive and masked SSL within the primary evaluation tasks of classification and content generation. Our findings demonstrate that in classification tasks, the flexibility of targeted tokens in masked SSL fosters more inter-sample connections compared to the fixed position of target tokens in autoregressive SSL, which yields superior clustering performance. In content generation tasks, the misalignment between the flexible lengths of test samples and the fixed length of unmasked texts in masked SSL (vs. flexible lengths of conditional texts in autoregressive SSL) hinders its generation performance. To leverage each other’s strengths and mitigate weaknesses, we propose diversity-enhanced autoregressive and variable-length masked objectives, which substantially improve the classification performance of autoregressive SSL and the generation performance of masked SSL. Code is available at https://github.com/PKU-ML/LookAheadLookAround.
Qi Zhang 0067, Tianqi Du, Haotian Huang, Yifei Wang 0001, Yisen Wang 0001
ICML3
2023 Metacognition-driven user-to-project recommendation for online education services
Zezheng Wu, Xinghe Cheng, Haotian Huang
World Wide Web (WWW)4