EDBT 2026 Demo / reviewers in the wild / expert
Hengxing Cai
dblp:194/2976
· DBLP profile ↗
8ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 22% Trustworthy machine learning · 17% Reinforcement learning · 17% | |
| Computer graphics and multimedia
2 papers |
Audio and music processing · 54% Multimedia analysis and retrieval · 36% Image and video processing · 11% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 17 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
1.0 | 1 | 2026 | Interpretable Reward Model via Sparse Autoencoder · AAAI 2026 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
1.0 | 1 | 2026 | Interpretable Reward Model via Sparse Autoencoder · AAAI 2026 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
1.0 | 1 | 2026 | Interpretable Reward Model via Sparse Autoencoder · AAAI 2026 |
Machine learning › Trustworthy machine learning › interpretability › explainable reinforcement learning
reward model interpretability |
1.0 | 1 | 2026 | Interpretable Reward Model via Sparse Autoencoder · AAAI 2026 |
Robotics › Robot navigation and mapping › mobile robot navigation › 3d navigation
aerial robot navigation |
0.9 | 1 | 2025 | FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models · EMNLP 2025 |
Natural language and speech › Language models and text generation
instruction tuning |
0.9 | 1 | 2025 | SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding · ICLR 2025 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.9 | 1 | 2025 | Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented Reasoning · NeurIPS 2025 |
Natural language and speech › Question answering and dialogue systems
retrieval-augmented reasoning |
0.9 | 1 | 2025 | Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented Reasoning · NeurIPS 2025 |
Robotics › Legged, aerial and field robots › aerial robots
UAV navigation |
0.9 | 1 | 2025 | FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models · EMNLP 2025 |
Computer vision › Vision and language
vision-and-language navigation |
0.9 | 1 | 2025 | FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models · EMNLP 2025 |
Information retrieval
retrieval models |
0.9 | 1 | 2025 | Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented Reasoning · NeurIPS 2025 |
Audio and music processing
sound event detection |
0.7 | 1 | 2023 | IFS-SED: Incremental Few-Shot Sound Event Detection Using Explicit Learning and Calibration · ACM Multimedia 2023 |
Computer vision › Video understanding and tracking
video classification |
0.6 | 1 | 2022 | Multiple Temporal Fusion based Weakly-supervised Pre-training Techniques for Video Categorization · ACM Multimedia 2022 |
Multimedia analysis and retrieval
image retrieval |
0.4 | 1 | 2020 | Multi-Scale Generalized Attention-Based Regional Maximum Activation of Convolutions for Beauty Product Retrieval · ACM Multimedia 2020 |
Computer vision › Vision and language
vision-language model |
0.3 | 1 | 2025 | FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models · EMNLP 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.2 | 1 | 2023 | IFS-SED: Incremental Few-Shot Sound Event Detection Using Explicit Learning and Calibration · ACM Multimedia 2023 |
Image and video processing › feature extraction
multi-scale features |
0.1 | 1 | 2020 | Multi-Scale Generalized Attention-Based Regional Maximum Activation of Convolutions for Beauty Product Retrieval · ACM Multimedia 2020 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 2.6group relative policy optimization · 1.7model pruning · 1.3explicit learning and calibration · 1.3sparse autoencoder · 1.0feature attribution · 1.0vision-language model · 0.9synthetic instruction generation · 0.9supervised fine-tuning · 0.9continual pre-training · 0.9self-distillation · 0.7multi-scale attention · 0.4generalized attention mechanism · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interpretable Reward Model via Sparse AutoencoderabstractLarge language models (LLMs) have been widely deployed across numerous fields. Reinforcement Learning from Human Feedback (RLHF) leverages reward models (RMs) as proxies for human preferences to align LLM behaviors with human values, making the accuracy, reliability, and interpretability of RMs critical for effective alignment. However, traditional RMs lack interpretability, offer limited insight into the reasoning behind reward assignments, and are inflexible toward user preference shifts. While recent multidimensional RMs aim for improved interpretability, they often fail to provide feature-level attribution and require costly annotations. To overcome these limitations, we introduce the Sparse Autoencoder-Enhanced Reward Model (SARM), a novel architecture that integrates a pretrained Sparse Autoencoder (SAE) into a reward model. SARM maps the hidden activations of LLM-based RM into an interpretable, sparse, and monosemantic feature space, from which a scalar head aggregates feature activations to produce transparent and conceptually meaningful reward scores. Empirical evaluations demonstrate that SARM facilitates direct feature-level attribution of reward assignments, allows dynamic adjustment to preference shifts, and achieves superior alignment performance compared to conventional reward models. Sihang Li 0002, Jiayi Liao, Hengxing Cai, Xiang Wang 0010 |
AAAI | 6 |
| 2025 | FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language ModelsabstractHengxing Cai, Jinhan Dong, Jingjun Tan, Jingcheng Deng, Sihang Li, Zhifeng Gao, Haidong Wang, Zicheng Su, Agachai Sumalee, Renxin Zhong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Hengxing Cai, Jinhan Dong, Jingjun Tan, Jingcheng Deng, Sihang Li 0002, Zhifeng Gao, Zicheng Su, Agachai Sumalee, Renxin Zhong |
EMNLP | 1 |
| 2025 | SciLitLLM: How to Adapt LLMs for Scientific Literature UnderstandingabstractScientific literature understanding is crucial for extracting targeted information and garnering insights, thereby significantly advancing scientific discovery.
Despite the remarkable success of Large Language Models (LLMs), they face challenges in scientific literature understanding, primarily due to (1) a lack of scientific knowledge and (2) unfamiliarity with specialized scientific tasks.
To develop an LLM specialized in scientific literature understanding, we propose a hybrid strategy that integrates continual pre-training (CPT) and supervised fine-tuning (SFT), to simultaneously infuse scientific domain knowledge and enhance instruction-following capabilities for domain-specific tasks.
In this process, we identify two key challenges: (1) constructing high-quality CPT corpora, and (2) generating diverse SFT instructions.
We address these challenges through a meticulous pipeline, including PDF text extraction, parsing content error correction, quality filtering, and synthetic instruction creation.
Applying this strategy, we present a suite of LLMs: SciLitLLM, specialized in scientific literature understanding.
These models demonstrate promising performance on scientific literature understanding benchmarks.
(1) We present an effective framework that integrates CPT and SFT to adapt LLMs to scientific literature understanding, which can also be easily adapted to other domains.
(2) We propose an LLM-based synthesis method to generate diverse and high-quality scientific instructions, resulting in a new instruction set -- SciLitIns -- for less-represented scientific domains.
(3) SciLitLLM achieves promising performance in scientific literature understanding benchmarks. Sihang Li 0002, Jin Huang 0001, Jiaxi Zhuang, Yaorui Shi, Xiaochen Cai, Mingjun Xu, Xiang Wang 0010, Linfeng Zhang 0002, Guolin Ke, Hengxing Cai |
ICLR | 10 |
| 2025 | Towards Unified and Lossless Latent Space for 3D Molecular Latent Diffusion Modelingabstract3D molecule generation is crucial for drug discovery and material science, requiring models to process complex multi-modalities, including atom types, chemical bonds, and 3D coordinates. A key challenge is integrating these modalities of different shapes while maintaining SE(3) equivariance for 3D coordinates. To achieve this, existing approaches typically maintain separate latent spaces for invariant and equivariant modalities, reducing efficiency in both training and sampling.
In this work, we propose **U**nified Variational **A**uto-**E**ncoder for **3D** Molecular Latent Diffusion Modeling (**UAE-3D**), a multi-modal VAE that compresses 3D molecules into latent sequences from a unified latent space, while maintaining near-zero reconstruction error. This unified latent space eliminates the complexities of handling multi-modality and equivariance when performing latent diffusion modeling. We demonstrate this by employing the Diffusion Transformer--a general-purpose diffusion model without any molecular inductive bias--for latent generation. Extensive experiments on GEOM-Drugs and QM9 datasets demonstrate that our method significantly establishes new benchmarks in both *de novo* and conditional 3D molecule generation, achieving leading efficiency and quality. On GEOM-Drugs, it reduces FCD by 72.6% over the previous best result, while achieving over 70% relative average improvements in geometric fidelity. Our code is released at [https://github.com/lyc0930/UAE-3D/](https://github.com/lyc0930/UAE-3D/). Yanchen Luo, Zhiyuan Liu 0001, Sihang Li 0002, Hengxing Cai, Kenji Kawaguchi, Tat-Seng Chua, Yang Zhang 0094, Xiang Wang 0010 |
NeurIPS | 5 |
| 2025 | Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented ReasoningabstractLarge language models have demonstrated impressive reasoning capabilities but are inherently limited by their knowledge reservoir.
Retrieval-augmented reasoning mitigates this limitation by allowing LLMs to query external resources, but existing methods often retrieve irrelevant or noisy information, hindering accurate reasoning.
In this paper, we propose **AutoRefine**, a reinforcement learning post-training framework that adopts a new "search-and-refine-during-think" paradigm.
AutoRefine introduces explicit knowledge refinement steps between successive search calls, enabling the model to iteratively filter, distill, and organize evidence before generating an answer.
Furthermore, we incorporate tailored retrieval-specific rewards alongside answer correctness rewards using group relative policy optimization.
Experiments on single-hop and multi-hop QA benchmarks demonstrate that AutoRefine significantly outperforms existing approaches, particularly in complex, multi-hop reasoning scenarios.
Detailed analysis shows that AutoRefine issues frequent, higher-quality searches and synthesizes evidence effectively. Yaorui Shi, Sihang Li 0002, Chang Wu 0003, Zhiyuan Liu 0001, Junfeng Fang, Hengxing Cai, An Zhang 0003, Xiang Wang 0010 |
NeurIPS | 6 |
| 2023 | IFS-SED: Incremental Few-Shot Sound Event Detection Using Explicit Learning and CalibrationabstractSound event detection (SED) refers to recognizing the sound events in a continuous audio signal, which has drawn increasing interest during recent decades. The applications of SED seem to be evident in many fields, ranging from surveillance to monitoring applications. Despite the sustainable efforts that have been made, most of the previous attempts are performed on the closed-set, as only fixed and known sound event classes can be employed during the training. In this paper, we present our incremental few-shot SED framework under the open-set settings, as a practical machine listening system should be able to address unknown sound events. Specifically, an explicit learning and calibration-based multi-stage learning framework is utilized to address the challenges of catastrophic forgetting, and aim to achieve a better trade-off between stability and plasticity. To compress the model efficiently, the model prune and self-distillation paradigm are combined used for the model compression, thus our system can be deployed for the resource-limited devices. Our framework can also provide an uncertain estimation for the inference. Lastly, an interactive interface is presented to demonstrate the functions of our system. Ming Feng, Kele Xu, Hengxing Cai |
ACM Multimedia | 3 |
| 2022 | Multiple Temporal Fusion based Weakly-supervised Pre-training Techniques for Video CategorizationabstractIn this paper, we present our solution of the ACM Multimedia 2022 pre-training for video understanding challenge. First, we pre-train the models on large-scale weakly-supervised video datasets with different temporal resolutions, then fine-tune the model for downstream application. Quantitative comparisons are conducted to evaluate the performance of different networks at multiple temporal resolutions. Moreover, we fusion different pre-trained models through weighted averaging. We achieve an accuracy of 62.39% in the testing set, which ranked as the first place in the video categorization track of this challenge. Xiaochen Cai, Hengxing Cai, Boqing Zhu, Kele Xu, Wei-Wei Tu |
ACM Multimedia | 2 |
| 2020 | Multi-Scale Generalized Attention-Based Regional Maximum Activation of Convolutions for Beauty Product RetrievalabstractThe application of beauty and personal-care product retrieval seems to be evident in our daily life, and it has attracted increasing research interests during the last decade. However, the retrieval task is suffered from different image variations and complicated backgrounds. Recent works have demonstrated that Generalized-attention Regional Maximal Activation of Convolutions (GRMAC) descriptor can provide state-of-the-art performance for the retrieval task. However, GRMAC descriptor is restrained from the essentially limited property of the employed feature from a single layer. Features from a single layer are not robust enough for scale variations, shape deformation, and heavy occlusion. In this paper, we propose a novel descriptors, named Multi-Scale Generalized Attention-Based Regional Maximum Activation of Convolutions (MS-GRMAC). This method introduces multi-scale generalized attention mechanism to reduce the influence of scale variations, thus, can boost the performance of the retrieval task. To empirically investigate the effectiveness of the proposed approach, we conduct extensive experiments on the dataset containing more than half-million personal-care products (Perfect-500K) and obtain satisfactory results without ensemble. Kele Xu, Yuzhong Liu, Ming Feng, Jianqiao Zhao, Huaimin Wang 0001, Hengxing Cai |
ACM Multimedia | 6 |