VLDB 2026 Research / reviewers in the wild / expert
Huanxi Liu
dblp:94/8066
· DBLP profile ↗
19ranked-venue papers
5as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Melody Structure Transfer Network: Generating Music with Separable Self-AttentionabstractMost existing symbolic music generation methods focus on generating short pieces, typically less than 8 bars and occasionally up to 32 bars. Generating long music sequences requires effective representation of coherent musical structures. Vanilla self-attention face challenges in capturing subtle long-term musical structures. We propose an approach to transfer the structural characteristics of training samples for generating music. We introduce a separable self-attention-based model that facilitates the learning and transfer of structural embeddings. It can generate music sequences of up to 100 bars, producing compositions with interpretable structures that closely resemble the structural and compositional techniques of the training set. Experiments show the model’s ability to generate music with targeted structures while maintaining good diversity. Ning Zhang 0040, Boan Chen, Huanxi Liu, Junchi Yan |
ICASSP | 5 |
| 2025 | LaTeXNet: A Specialized Model for Converting Visual Tables and Equations to LaTeX CodeabstractLaTeX provides precise representation of complex elements (i.e., tables and equations) in scientific documents. However, the automated transcription of visual representations into LaTeX code is challenging and prone to errors. This paper introduces LaTeXNet, a specialized model designed to automate the conversion of visual tables and equations into LaTeX code. First, we develop an automated annotation tool that extracts image-LaTeX pairs for tables, equations, and text paragraphs with inline equations from arXiv platform, creating the MM-LaTeX dataset with over 2.5M pairs. Moreover, we design the LaTeXNet model, trained on MM-LaTeX, which unifies the conversion of Tables, Equations, and TextEqs. Our experimental results indicate that LaTeXNet surpasses both open-source and commercial, closed-source models in Table-to-LaTeX, Equation-to-LaTeX and TextEq-to-LaTeX tasks. Renqiu Xia, Hongbin Zhou, Ziming Feng, Huanxi Liu, Boan Chen, Junchi Yan |
ICASSP | 4 |
| 2025 | ALMP: Automatic Layer-By-Layer Mixed-Precision Quantization for Large Language Models
Huanxi Liu, Bo Ding 0001 |
ICIC (23) | 2 |
| 2025 | Extracting Reasoning Patterns from Knowledge Graph to Enhance LLMs' Reasoning CapabilityabstractLarge language models (LLMs) have demonstrated great potential across diverse fields. However, their reasoning capabilities face challenges, especially when dealing with complex tasks. Existing work on enhancing LLMs' reasoning abilities mostly focuses on fields like mathematics and coding. Due to the specificity of domain data, these methods have poor generalization in specific areas such as medicine and materials. In this work, we propose an approach to enhance LLMs' reasoning capabilities in these fields lacking reasoning datasets. We construct reasoning datasets based on the logical relationships contained in their knowledge graphs. First, we clean the metadata, focus on a specific sub-field, and optimize data quality to construct a specialized knowledge graph. Then, we design seven knowledge graph patterns to extract instances, transform them into natural-language questions, and sample using the Deepseek-R1 model to build a reasoning dataset. We fine-tune the Qwen2.5-32B model with this dataset. Experimental results show that in the Relevance dimension, the score of Qwen2.5-32B improves from 45 to 98.57, with a remarkable increase of about 119.04%. In the Logic dimension, the score rises from 76.87 to 87.2, representing an increase of approximately 13.44%. This validates the effectiveness of our method. In these two aspects, the fine-tuned model reaches or approaches the level of advanced LLMs such as Deepseek-R1 and OpenAI 01, indicating that our approach can effectively enhance LLMs' reasoning ability. Tingkai Yang, Yuanzhao Zhai, Huanxi Liu, Huaimin Wang 0001 |
JCC | 3 |
| 2025 | Empowering Large Language Model Agent through Step-Level Self-Critique and Self-TrainingabstractLarge Language Model (LLM) agents frequently produce sub-optimal actions when tackling complex, multi-step decision-making tasks. Employing self-critique to identify flaws and suggest enhancements is an effective strategy for refining actions. Although trajectory-level critique is commonly employed, it often fails to identify flawed steps accurately. In this paper, we introduce SLSC-MCTS, a method that integrates Monte Carlo Tree Search with Step-Level Self-Critique to enhance LLM agents during both testing and self-training phases. During decision tree expansion with SLSC-MCTS, the LLM agent initially generates an action, receives environmental feedback, and subsequently generates further actions via self-critique and refinement. Through multiple episodes of SLSC-MCTS, LLM agents can effectively utilize step-level critiques while disregarding ineffective ones based on node values, thereby incorporating the critiques more robustly. Additionally, our method further empowers LLM agents in a self-training manner, collecting training data from the constructed decision tree to iteratively fine-tune the LLM agents. The self-training data gathered via SLSC-MCTS is diverse and high-quality, which further enhances the reasoning, critiquing, and refining abilities of LLM agents. Experimental results demonstrate that SLSC-MCTS significantly improves LLM agents during testing, surpassing state-of-the-art baselines and achieving shorter task completion trajectories across information retrieval benchmarks such as WebShop and HotPotQA. After three iterations of self-training, LLM agents established by Llama-3.1-8B-Instruct show substantial improvement, even surpassing human experts in WebShop. Yuanzhao Zhai, Huanxi Liu, Zhuo Zhang 0007, Kele Xu, Cheng Yang 0004, Bo Ding 0001, Huaimin Wang 0001 |
SIGIR | 2 |
| 2024 | Nuclear-Norm Maximization for Low-Rank UpdatesabstractPre-trained large language models exhibit significant potential in speech and language processing. Fine-tuning all parameters becomes impractical when confronted with numerous downstream tasks. To address this challenge, various low-rank adaptation techniques have been introduced for parameter-efficient fine-tuning, which freeze the over-parametrized models and learn incremental parameter updates within smaller subspaces. However, our observation reveals that most directions of the learned subspace play a minor role in the incremental updates. Consequently, fine-tuned models may not achieve optimal performance. To bridge this gap, we introduce NNM-LoRA, which strives to harness more meaningful singular directions. Through Nuclear Norm Maximization (NNM), we can better regulate the allocation of singular values. Accordingly, we propose a parameter-free plug-and-play regularizer for low-rank updates. This innovative approach allows us to utilize as many singular directions of the subspace as possible during the training of low-rank updates. To validate the effectiveness of NNM-LoRA, we conduct extensive experiments involving different pre-trained models on various natural language understanding tasks. Results demonstrate that NNM-LoRA exhibits significant improvements compared to baseline methods. Huanxi Liu, Yuanzhao Zhai, Kele Xu, Yiying Li |
ICASSP | 1 |
| 2024 | FP4-Quantization: Lossless 4bit Quantization for Large Language ModelsabstractLarge language models(LLMs) have demonstrated exceptional performance across a wide range of tasks. However, their extensive computational and storage requirements hinder their widespread deployment. To address this, low-bit quantization has emerged as a highly effective approach to reducing the inference cost of LLMs. Nevertheless, the existing repertoire of 4-bit quantization techniques is plagued by a substantial decline in model precision. In this paper, we introduce a novel 4-bit weight quantization method, FP4-Quantization, which leverages a 4-bit floating-point(FP4) representation that aligns better with the weight distribution characteristics of LLMs. Furthermore, it incorporates a Low-Rank Quantization Error Correction(LREC), involving progressive fine-tuning of low-rank parameters to rectify quantization errors, thereby enabling the achievement of precision-preserving 4-bit weight-only quantization. Our Experimental results on multiple zero-shot tasks demonstrate that FP4-Quantization achieves 4-bit weight quantization with an accuracy degradation of less than 0.5%. Huanxi Liu, Bo Ding 0001 |
JCC | 2 |
| 2023 | An empirical study on the structure evolution of deep learning models: taking SAR image processing a case studyabstractWith the continuous improvement on model performance, deep learning models have been widely deployed and achieved promising outcomes in various fields in recent years. However, due to the escalating volumes of training data and the complexity of application problems, it becomes more and more challenging to design a neural network with better performance by hand. Analysing the evolution of typical neural network structures has important reference significance for designing a network structure. In this paper, we select the open source models in SAR image processing for an empirical analysis on the evolution of neural network structures. We analyse the evolution of 239 open source deep learning models from the aspects of framework, computing unit, model computation amount and the combined use of various computing units. Results reveal that preference and co-occurrence exist in computing units, while the average number of convolution, activation and normalization layer increases significantly over time. Model complexity shows an overall upward trend, and the characteristics of SAR image are more and more taken into consideration during the model structure design. Huanxi Liu |
JCC | 1 |
| 2023 | Modeling Event Propagation via Graph Biased Temporal Point ProcessabstractTemporal point process is widely used for sequential data modeling. In this article, we focus on the problem of modeling sequential event propagation in graph, such as retweeting by social network users and news transmitting between websites. Given a collection of event propagation sequences, the conventional point process model considers only the event history, i.e., embed event history into a vector, not the latent graph structure. We propose a graph biased temporal point process (GBTPP) leveraging the structural information from graph representation learning, where the direct influence between nodes and indirect influence from event history is modeled. Moreover, the learned node embedding vector is also integrated into the embedded event history as side information. Experiments on a synthetic data set and two real-world data sets show the efficacy of our model compared with conventional methods and state-of-the-art ones. Weichang Wu, Huanxi Liu, Hongyuan Zha |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Trajectory Prediction from EGO View: A Coordinate Transform and Tail-Light Event Driven ApproachabstractTrajectory prediction plays an important role in modern au-tonomous driving system. The multi-modal characteristics of trajectory prediction makes it difficult to accurately predict the driving intention and future trajectory of vehicles, espe-cially in the complex conditions, such as lane changing and road intersection. To improve the prediction performance, ad-ditional features are needed. High-definition (HD) map fea-ture and tail light feature are employed in this paper, which are fused with trajectory feature to assist vehicle trajectory pre-diction. A prediction model based on temporal convolution network (TCN) and graph convolution network (GCN) is constructed with corresponding loss functions. Experiments are carried out on our simulation dataset and the results show the effectiveness of the proposed feature fusion method as well as the prediction model. Wenlong Liao, Huanxi Liu, Junchi Yan, Yingxin Lou, Shuqi Mei |
ICME | 4 |
| 2022 | Distilldarts: Network Distillation for Smoothing Gradient Distributions in Differentiable Architecture SearchabstractRecent studies show that differentiable architecture search (DARTS) suffers notable instability and collapse issue: skip-connect may gradually dominate the cell, leading to deteri-orating architectures. We conjecture that the domination of skip-connect is due to its superiority in gradient compen-sate. On this foundation, we propose a novel and stable method, called DistillDARTS, to stabilize DARTS by knowl-edge distillation and self-distillation scheme. Specifically, the distillation is able to serve as a substitute for skip-connect and smooth the back-propagated gradient distributions among layers of DARTS. By compensating gradients in shallow lay-ers, our method can relieve the dependence of gradient on skip-connect and hence mitigates the collapse issue. Exten-sive experiments on a range of benchmarks demonstrate that DistillDARTS can obtain sturdy architectures with few skip-connects without additional manual interventions, thus suc-cessfully improving the robustness of DARTS. Due to the im-proved stability, our proposed approach achieves the accuracy of 97.57% on CIFAR-10 and 75.8% on ImageNet. Wenlong Liao, Zhexi Zhang, Xiaoxing Wang, Huanxi Liu, Zhenyu Ren, Jian Yin 0023, Shuzhi Feng |
ICME | 5 |
| 2015 | A Matrix Decomposition Perspective to Multiple Graph MatchingabstractGraph matching has a wide spectrum of real-world applications and in general is known NP-hard. In many vision tasks, one realistic problem arises for finding the global node mappings across a batch of corrupted weighted graphs. This paper is an attempt to connect graph matching, especially multi-graph matching to the matrix decomposition model and its relevant on-the-shelf convex optimization algorithms. Our method aims to extract the common inliers and their synchronized permutations from disordered weighted graphs in the presence of deformation and outliers. Under the proposed framework, several variants can be derived in the hope of accommodating to other types of noises. Experimental results on both synthetic data and real images empirically show that the proposed paradigm exhibits several interesting behaviors and in many cases performs competitively with the state-of-the-arts. Junchi Yan, Hongteng Xu, Hongyuan Zha, Xiaokang Yang 0001, Huanxi Liu, Stephen M. Chu |
ICCV | 5 |
| 2014 | A Correlative Two-Step Approach to Hallucinating FacesabstractFace hallucination is to synthesize high-resolution face image from the input low-resolution one. Although many two-step learning-based face hallucination approaches have been developed, they suffer from the expensive computational cost due to the separate calculation of the global and local models. To overcome this problem, we propose a correlative two-step learning-based face hallucination approach which bridges the gap between the global model and the local model. In the global phase, we build a global face hallucination framework by combining the steerable pyramid decomposition and the reconstruction. In the residue compensation phase, based on the combination weights and constituent samples obtained in the global phase, a residue face image is synthesized by the neighbor reconstruction algorithm to compensate the hallucinated global face image with subtle facial features. The ultimate hallucinated result is synthesized by adding the residue face image to the global face image. Compared with existing methods, in the global phase, our global face image is more similar to the original high-resolution face image. Furthermore, in the residue compensation phase, we use the combination weights and constituent samples obtained in the global phase to compute the residue face image, by which the computational efficiency can be greatly improved without compromising the quality of facial details. The experimental results and comparisons demonstrate that our approach can not only generate convincible high-resolution face images efficiently, but also has high computational efficiency. Furthermore, our proposed approach can be used to restore the damaged face images in image inpainting. The efficacy of our approach is validated by recovering the damaged face images with visually good results. Huanxi Liu, Tianhong Zhu |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2011 | A Double-Layer Model for Foreground Detection from Video SequenceabstractThis paper proposes a method for background modeling and foreground detection in video. This method divides the background into two layers, the dynamic layer and the static layer. An energy descriptor is proposed to analysis the motion state in dynamic layer while a grid filter is proposed to reduce the negative impact of sudden illumination change such as light switching off. Experiment results compared with four typical algorithms show that this method outperforms others in most challenging videos including sudden illumination change and some complex backgrounds. Huanxi Liu, Junchi Yan, Xiaowei Lv, Xiong Li 0004, Tianhong Zhu, Yuncai Liu |
ICIG | 1 |
| 2011 | Sparse Feature Learning for Visual Tracking by Least Absolute Shrinkage and Selection OperatorabstractIn this paper, a robust visual tracking method is proposed by using Least Absolute Shrinkage and Selection Operator (Lasso) in a particle filter framework. First, to locate the tracking target at a new frame, each target candidate is sparsely represented in the space spanned by sampling particle images and sparsity vector. The lasso can replace the sparse approximation problem by a convex problem. Then, the likelihood is evaluated by the sparsity coefficients which is very different from the current tracking scheme using sparse representation. The combination of sampling images with coefficients bigger than a threshold will be taken as the tracking target without any template. After that, tracking is continued using a Bayesian state inference framework in which a particle filter is used for propagating sample distributions over time. The dynamic target updating scheme keeps track of the most representative particles throughout the tracking procedure. The proposed approach shows excellent performance on several image sequences. Minglei Tong, Junchi Yan, Huanxi Liu, Hong Han 0001 |
ICIG | 3 |
| 2010 | Multimodality gender estimation using Bayesian hierarchical modelabstractWe propose to estimate human gender from corresponding fingerprint and face information with the Bayesian hierarchical model. Different from previous works on fingerprint based gender estimation with specially designed features, our method extends to use general local image features. Furthermore, a novel word representation called latent word is designed to work with the Bayesian hierarchical model. The feature representation is embedded to our multimodality model, within which the information from fingerprint and face is fused at the decision level for gender estimation. Experiments on our internal database show the promising performance. Xiong Li 0004, Xu Zhao 0001, Huanxi Liu, Yun Fu 0001, Yuncai Liu |
ICASSP | 3 |
| 2010 | Learning parts-based representation for face transitionabstractThis paper proposes to learn parts-based face representation from real face samples and then applies it to face transition. It differs from previous works in two aspects. First, we learn flexible face decomposition from real faces unsupervisedly instead of designing face template manually, for which two simple priors are embedded into learning procedure through constrained EM formulation. Second, both face representation and transition are derived from an unified probabilistic framework. Based on the learned face representation, the face distance measurement could be defined, which enables us to synthesize face via specifying distance with respect to reference faces and depict the full transition trace of two or more given faces with distinct age, gender and race. Xiong Li 0004, Huanxi Liu, Yuncai Liu |
ACM Multimedia | 3 |
| 2010 | Visual Saliency Detection via Sparsity PursuitabstractSaliency mechanism has been considered crucial in the human visual system and helpful to object detection and recognition. This paper addresses a novel feature-based model for visual saliency detection. It consists of two steps: first, using the learned overcomplete sparse bases to represent image patches; and then, estimating saliency information via low-rank and sparsity matrix decomposition. We compare our model with the previous methods on natural images. Experimental results on both natural images and psychological patterns show that our model performs competitively for visual saliency detection task, and suggest the potential application of matrix decomposition and convex optimization for image analysis. Junchi Yan, Mengyuan Zhu, Huanxi Liu, Yuncai Liu |
IEEE Signal Process. Lett. | 3 |
| 2009 | Fingerprint Orientation Field Estimation: Model of Primary Ridge for Global Structure and Model of Secondary Ridge for Correction
Huanxi Liu, Xiaowei Lv, Xiong Li 0004, Yuncai Liu |
ACCV (3) | 1 |