Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hengxing Cai

dblp:194/2976 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 22% Trustworthy machine learning · 17% Reinforcement learning · 17%
Computer graphics and multimedia
2 papers
Audio and music processing · 54% Multimedia analysis and retrieval · 36% Image and video processing · 11%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.012026
Interpretable Reward Model via Sparse Autoencoder · AAAI 2026
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.012026
Interpretable Reward Model via Sparse Autoencoder · AAAI 2026
Machine learning › Reinforcement learning › reward learning
reward modeling
1.012026
Interpretable Reward Model via Sparse Autoencoder · AAAI 2026
Machine learning › Trustworthy machine learning › interpretability › explainable reinforcement learning
reward model interpretability
1.012026
Interpretable Reward Model via Sparse Autoencoder · AAAI 2026
Robotics › Robot navigation and mapping › mobile robot navigation › 3d navigation
aerial robot navigation
0.912025
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models · EMNLP 2025
Natural language and speech › Language models and text generation
instruction tuning
0.912025
SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding · ICLR 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.912025
Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented Reasoning · NeurIPS 2025
Natural language and speech › Question answering and dialogue systems
retrieval-augmented reasoning
0.912025
Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented Reasoning · NeurIPS 2025
Robotics › Legged, aerial and field robots › aerial robots
UAV navigation
0.912025
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models · EMNLP 2025
Computer vision › Vision and language
vision-and-language navigation
0.912025
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models · EMNLP 2025
Information retrieval
retrieval models
0.912025
Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented Reasoning · NeurIPS 2025
Audio and music processing
sound event detection
0.712023
IFS-SED: Incremental Few-Shot Sound Event Detection Using Explicit Learning and Calibration · ACM Multimedia 2023
Computer vision › Video understanding and tracking
video classification
0.612022
Multiple Temporal Fusion based Weakly-supervised Pre-training Techniques for Video Categorization · ACM Multimedia 2022
Multimedia analysis and retrieval
image retrieval
0.412020
Multi-Scale Generalized Attention-Based Regional Maximum Activation of Convolutions for Beauty Product Retrieval · ACM Multimedia 2020
Computer vision › Vision and language
vision-language model
0.312025
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models · EMNLP 2025
Machine learning › Efficient and distributed learning
model compression
0.212023
IFS-SED: Incremental Few-Shot Sound Event Detection Using Explicit Learning and Calibration · ACM Multimedia 2023
Image and video processing › feature extraction
multi-scale features
0.112020
Multi-Scale Generalized Attention-Based Regional Maximum Activation of Convolutions for Beauty Product Retrieval · ACM Multimedia 2020

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 2.6group relative policy optimization · 1.7model pruning · 1.3explicit learning and calibration · 1.3sparse autoencoder · 1.0feature attribution · 1.0vision-language model · 0.9synthetic instruction generation · 0.9supervised fine-tuning · 0.9continual pre-training · 0.9self-distillation · 0.7multi-scale attention · 0.4generalized attention mechanism · 0.4
YearPublicationVenuePosition
2026 Interpretable Reward Model via Sparse Autoencoder
abstract
Large language models (LLMs) have been widely deployed across numerous fields. Reinforcement Learning from Human Feedback (RLHF) leverages reward models (RMs) as proxies for human preferences to align LLM behaviors with human values, making the accuracy, reliability, and interpretability of RMs critical for effective alignment. However, traditional RMs lack interpretability, offer limited insight into the reasoning behind reward assignments, and are inflexible toward user preference shifts. While recent multidimensional RMs aim for improved interpretability, they often fail to provide feature-level attribution and require costly annotations. To overcome these limitations, we introduce the Sparse Autoencoder-Enhanced Reward Model (SARM), a novel architecture that integrates a pretrained Sparse Autoencoder (SAE) into a reward model. SARM maps the hidden activations of LLM-based RM into an interpretable, sparse, and monosemantic feature space, from which a scalar head aggregates feature activations to produce transparent and conceptually meaningful reward scores. Empirical evaluations demonstrate that SARM facilitates direct feature-level attribution of reward assignments, allows dynamic adjustment to preference shifts, and achieves superior alignment performance compared to conventional reward models.
Sihang Li 0002, Jiayi Liao, Hengxing Cai, Xiang Wang 0010
AAAI6
2025 FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models
abstract
Hengxing Cai, Jinhan Dong, Jingjun Tan, Jingcheng Deng, Sihang Li, Zhifeng Gao, Haidong Wang, Zicheng Su, Agachai Sumalee, Renxin Zhong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Hengxing Cai, Jinhan Dong, Jingjun Tan, Jingcheng Deng, Sihang Li 0002, Zhifeng Gao, Zicheng Su, Agachai Sumalee, Renxin Zhong
EMNLP1
2025 SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding
abstract
Scientific literature understanding is crucial for extracting targeted information and garnering insights, thereby significantly advancing scientific discovery. Despite the remarkable success of Large Language Models (LLMs), they face challenges in scientific literature understanding, primarily due to (1) a lack of scientific knowledge and (2) unfamiliarity with specialized scientific tasks. To develop an LLM specialized in scientific literature understanding, we propose a hybrid strategy that integrates continual pre-training (CPT) and supervised fine-tuning (SFT), to simultaneously infuse scientific domain knowledge and enhance instruction-following capabilities for domain-specific tasks. In this process, we identify two key challenges: (1) constructing high-quality CPT corpora, and (2) generating diverse SFT instructions. We address these challenges through a meticulous pipeline, including PDF text extraction, parsing content error correction, quality filtering, and synthetic instruction creation. Applying this strategy, we present a suite of LLMs: SciLitLLM, specialized in scientific literature understanding. These models demonstrate promising performance on scientific literature understanding benchmarks. (1) We present an effective framework that integrates CPT and SFT to adapt LLMs to scientific literature understanding, which can also be easily adapted to other domains. (2) We propose an LLM-based synthesis method to generate diverse and high-quality scientific instructions, resulting in a new instruction set -- SciLitIns -- for less-represented scientific domains. (3) SciLitLLM achieves promising performance in scientific literature understanding benchmarks.
Sihang Li 0002, Jin Huang 0001, Jiaxi Zhuang, Yaorui Shi, Xiaochen Cai, Mingjun Xu, Xiang Wang 0010, Linfeng Zhang 0002, Guolin Ke, Hengxing Cai
ICLR10
2025 Towards Unified and Lossless Latent Space for 3D Molecular Latent Diffusion Modeling
abstract
3D molecule generation is crucial for drug discovery and material science, requiring models to process complex multi-modalities, including atom types, chemical bonds, and 3D coordinates. A key challenge is integrating these modalities of different shapes while maintaining SE(3) equivariance for 3D coordinates. To achieve this, existing approaches typically maintain separate latent spaces for invariant and equivariant modalities, reducing efficiency in both training and sampling. In this work, we propose **U**nified Variational **A**uto-**E**ncoder for **3D** Molecular Latent Diffusion Modeling (**UAE-3D**), a multi-modal VAE that compresses 3D molecules into latent sequences from a unified latent space, while maintaining near-zero reconstruction error. This unified latent space eliminates the complexities of handling multi-modality and equivariance when performing latent diffusion modeling. We demonstrate this by employing the Diffusion Transformer--a general-purpose diffusion model without any molecular inductive bias--for latent generation. Extensive experiments on GEOM-Drugs and QM9 datasets demonstrate that our method significantly establishes new benchmarks in both *de novo* and conditional 3D molecule generation, achieving leading efficiency and quality. On GEOM-Drugs, it reduces FCD by 72.6% over the previous best result, while achieving over 70% relative average improvements in geometric fidelity. Our code is released at [https://github.com/lyc0930/UAE-3D/](https://github.com/lyc0930/UAE-3D/).
Yanchen Luo, Zhiyuan Liu 0001, Sihang Li 0002, Hengxing Cai, Kenji Kawaguchi, Tat-Seng Chua, Yang Zhang 0094, Xiang Wang 0010
NeurIPS5
2025 Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented Reasoning
abstract
Large language models have demonstrated impressive reasoning capabilities but are inherently limited by their knowledge reservoir. Retrieval-augmented reasoning mitigates this limitation by allowing LLMs to query external resources, but existing methods often retrieve irrelevant or noisy information, hindering accurate reasoning. In this paper, we propose **AutoRefine**, a reinforcement learning post-training framework that adopts a new "search-and-refine-during-think" paradigm. AutoRefine introduces explicit knowledge refinement steps between successive search calls, enabling the model to iteratively filter, distill, and organize evidence before generating an answer. Furthermore, we incorporate tailored retrieval-specific rewards alongside answer correctness rewards using group relative policy optimization. Experiments on single-hop and multi-hop QA benchmarks demonstrate that AutoRefine significantly outperforms existing approaches, particularly in complex, multi-hop reasoning scenarios. Detailed analysis shows that AutoRefine issues frequent, higher-quality searches and synthesizes evidence effectively.
Yaorui Shi, Sihang Li 0002, Chang Wu 0003, Zhiyuan Liu 0001, Junfeng Fang, Hengxing Cai, An Zhang 0003, Xiang Wang 0010
NeurIPS6
2023 IFS-SED: Incremental Few-Shot Sound Event Detection Using Explicit Learning and Calibration
abstract
Sound event detection (SED) refers to recognizing the sound events in a continuous audio signal, which has drawn increasing interest during recent decades. The applications of SED seem to be evident in many fields, ranging from surveillance to monitoring applications. Despite the sustainable efforts that have been made, most of the previous attempts are performed on the closed-set, as only fixed and known sound event classes can be employed during the training. In this paper, we present our incremental few-shot SED framework under the open-set settings, as a practical machine listening system should be able to address unknown sound events. Specifically, an explicit learning and calibration-based multi-stage learning framework is utilized to address the challenges of catastrophic forgetting, and aim to achieve a better trade-off between stability and plasticity. To compress the model efficiently, the model prune and self-distillation paradigm are combined used for the model compression, thus our system can be deployed for the resource-limited devices. Our framework can also provide an uncertain estimation for the inference. Lastly, an interactive interface is presented to demonstrate the functions of our system.
Ming Feng, Kele Xu, Hengxing Cai
ACM Multimedia3
2022 Multiple Temporal Fusion based Weakly-supervised Pre-training Techniques for Video Categorization
abstract
In this paper, we present our solution of the ACM Multimedia 2022 pre-training for video understanding challenge. First, we pre-train the models on large-scale weakly-supervised video datasets with different temporal resolutions, then fine-tune the model for downstream application. Quantitative comparisons are conducted to evaluate the performance of different networks at multiple temporal resolutions. Moreover, we fusion different pre-trained models through weighted averaging. We achieve an accuracy of 62.39% in the testing set, which ranked as the first place in the video categorization track of this challenge.
Xiaochen Cai, Hengxing Cai, Boqing Zhu, Kele Xu, Wei-Wei Tu
ACM Multimedia2
2020 Multi-Scale Generalized Attention-Based Regional Maximum Activation of Convolutions for Beauty Product Retrieval
abstract
The application of beauty and personal-care product retrieval seems to be evident in our daily life, and it has attracted increasing research interests during the last decade. However, the retrieval task is suffered from different image variations and complicated backgrounds. Recent works have demonstrated that Generalized-attention Regional Maximal Activation of Convolutions (GRMAC) descriptor can provide state-of-the-art performance for the retrieval task. However, GRMAC descriptor is restrained from the essentially limited property of the employed feature from a single layer. Features from a single layer are not robust enough for scale variations, shape deformation, and heavy occlusion. In this paper, we propose a novel descriptors, named Multi-Scale Generalized Attention-Based Regional Maximum Activation of Convolutions (MS-GRMAC). This method introduces multi-scale generalized attention mechanism to reduce the influence of scale variations, thus, can boost the performance of the retrieval task. To empirically investigate the effectiveness of the proposed approach, we conduct extensive experiments on the dataset containing more than half-million personal-care products (Perfect-500K) and obtain satisfactory results without ensemble.
Kele Xu, Yuzhong Liu, Ming Feng, Jianqiao Zhao, Huaimin Wang 0001, Hengxing Cai
ACM Multimedia6