EDBT 2026 Demo / reviewers in the wild / expert
Houxing Ren
dblp:292/3943
· DBLP profile ↗
18ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0001-9750-1626ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 8 first-author · 17 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-BenchabstractLarge Language Models (LLMs) show significant potential in AI mathematical tutoring, yet current evaluations often rely on simplistic metrics or narrow pedagogical scenarios, failing to assess comprehensive, multi-turn teaching effectiveness. In this paper, we introduce KMP-Bench, a comprehensive K-8 Mathematical Pedagogical Benchmark designed to assess LLMs from two complementary perspectives. The first module, KMP-Dialogue, evaluates holistic pedagogical capabilities against six core principles (e.g., Challenge, Explanation, Feedback), leveraging a novel multi-turn dialogue dataset constructed by weaving together diverse pedagogical components. The second module, KMP-Skills, provides a granular assessment of foundational tutoring abilities, including multi-turn problem-solving, error detection and correction, and problem generation. Our evaluations on KMP-Bench reveal a key disparity: while leading LLMs excel at tasks with verifiable solutions, they struggle with the nuanced application of pedagogical principles. Additionally, we present KMP-Pile, a large-scale (150K) dialogue dataset. Models fine-tuned on KMP-Pile show substantial improvement on KMP-Bench, underscoring the value of pedagogically-rich training data for developing more effective AI math tutors. Weikang Shi, Houxing Ren, Junting Pan, Aojun Zhou, Ke Wang 0036, Zimu Lu, Yunqiao Yang 0002, Linda Wei, Mingjie Zhan, Hongsheng Li 0001 |
AAAI | 2 |
| 2026 | Mind Reader: Latent User Demand-Guided Content Optimization for Generative Search EngineabstractTong Chen, JiaWei Guo, Yuxi Li, Baiming Chen, Houxing Ren, Zhang Zhiwei, Yunxiang Zhang, Hanyang Xia, Kun Liang, Zhaoran Fan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Baiming Chen, Houxing Ren, Zhang Zhiwei, Hanyang Xia, Zhaoran Fan |
ACL (1) | 5 |
| 2026 | Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format ReasoningabstractHouxing Ren, Mingjie Zhan, Zimu Lu, Ke Wang, Yunqiao Yang, Haotian Hou, Hongsheng Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Houxing Ren, Mingjie Zhan, Zimu Lu, Ke Wang 0036, Yunqiao Yang 0002, Haotian Hou, Hongsheng Li 0001 |
ACL (1) | 1 |
| 2026 | MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical ReasoningabstractWeikang Shi, Aldrich Yu, Rongyao Fang, Houxing Ren, Ke Wang, Aojun Zhou, Changyao Tian, Xinyu Fu, Yuxuan Hu, Zimu Lu, Linjiang Huang, Si Liu, Rui Liu, Hongsheng Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Weikang Shi, Aldrich Yu, Rongyao Fang, Houxing Ren, Ke Wang 0036, Aojun Zhou, Changyao Tian, Xinyu Fu 0004, Zimu Lu, Linjiang Huang, Si Liu 0001, Rui Liu 0019, Hongsheng Li 0001 |
ACL (1) | 4 |
| 2025 | ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code GenerationabstractCode generation plays a crucial role in various tasks, such as code auto-completion and mathematical reasoning.Previous work has proposed numerous methods to enhance code generation performance, including integrating feedback from the compiler.Inspired by this, we present ReflectionCoder, a novel approach that effectively leverages reflection sequences constructed by integrating compiler feedback to improve one-off code generation performance.Furthermore, we propose reflection self-distillation and dynamically masked distillation to effectively utilize these reflection sequences.Extensive experiments on three benchmarks, i.e., HumanEval (+), MBPP (+), and MultiPL-E, demonstrate that models fine-tuned with our method achieve state-of-the-art performance.Beyond the code domain, we believe this approach can benefit other domains that focus on final results and require long reasoning paths. Houxing Ren, Mingjie Zhan, Zhongyuan Wu, Aojun Zhou, Junting Pan, Hongsheng Li 0001 |
ACL (1) | 1 |
| 2025 | Alignment with Fill-In-the-Middle for Enhancing Code GenerationabstractHouxing Ren, Zimu Lu, Weikang Shi, Haotian Hou, Yunqiao Yang, Ke Wang, Aojun Zhou, Junting Pan, Mingjie Zhan, Hongsheng Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Houxing Ren, Zimu Lu, Weikang Shi, Haotian Hou, Yunqiao Yang 0002, Ke Wang 0036, Aojun Zhou, Junting Pan, Mingjie Zhan, Hongsheng Li 0001 |
EMNLP | 1 |
| 2025 | MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical CodeabstractCode has been shown to be effective in enhancing the mathematical reasoning abilities of large language models due to its precision and accuracy. Previous works involving continued mathematical pretraining often include code that utilizes math-related packages, which are primarily designed for fields such as engineering, machine learning, signal processing, or module testing, rather than being directly focused on mathematical reasoning. In this paper, we introduce a novel method for generating mathematical code accompanied with corresponding reasoning steps for continued pretraining. Our approach begins with the construction of a high-quality mathematical continued pretraining dataset by incorporating math-related web data, code using mathematical packages, math textbooks, and synthetic data. Next, we construct reasoning steps by extracting LaTeX expressions, the conditions needed for the expressions, and the results of the expressions from the previously collected dataset. Based on this extracted information, we generate corresponding code to accurately capture the mathematical reasoning process. Appending the generated code to each reasoning step results in data consisting of paired natural language reasoning steps and their corresponding code. Combining this data with the original dataset results in a 19.2B-token high-performing mathematical pretraining corpus, which we name MathCode-Pile. Training several popular base models with this corpus significantly improves their mathematical abilities, leading to the creation of the MathCoder2 family of models. All of our data processing and training code is open-sourced, ensuring full transparency and easy reproducibility of the entire data collection and training pipeline. Zimu Lu, Aojun Zhou, Ke Wang 0036, Houxing Ren, Weikang Shi, Junting Pan, Mingjie Zhan, Hongsheng Li 0001 |
ICLR | 4 |
| 2025 | WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from ScratchabstractLLM‑based agents have demonstrated great potential in generating and managing code within complex codebases. In this paper, we introduce WebGen-Bench, a novel benchmark designed to measure an LLM-based agent's ability to create multi-file website codebases from scratch. It contains diverse instructions for website generation, created through the combined efforts of human annotators and GPT-4o. These instructions span three major categories and thirteen minor categories, encompassing nearly all important types of web applications.To assess the quality of the generated websites, we generate test cases targeting each functionality described in the instructions. These test cases are then manually filtered, refined, and organized to ensure accuracy, resulting in a total of 647 test cases. Each test case specifies an operation to be performed on the website and the expected outcome of the operation.To automate testing and improve reproducibility, we employ a powerful web-navigation agent to execute test cases on the generated websites and determine whether the observed responses align with the expected results.We evaluate three high-performance code-agent frameworks—Bolt.diy, OpenHands, and Aider—using multiple proprietary and open-source LLMs as engines. The best-performing combination, Bolt.diy powered by DeepSeek-R1, achieves only 27.8\% accuracy on the test cases, highlighting the challenging nature of our benchmark.Additionally, we construct WebGen-Instruct, a training set consisting of 6,667 website-generation instructions. Training Qwen2.5-Coder-32B-Instruct on Bolt.diy trajectories generated from a subset of the training set achieves an accuracy of 38.2\%, surpassing the performance of the best proprietary model.We release our data-generation, training, and testing code, along with both the datasets and model weights at https://github.com/mnluzimu/WebGen-Bench. Zimu Lu, Yunqiao Yang 0002, Houxing Ren, Haotian Hou, Han Xiao 0010, Ke Wang 0036, Weikang Shi, Aojun Zhou, Mingjie Zhan, Hongsheng Li 0001 |
NeurIPS | 3 |
| 2024 | MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMsabstractZimu Lu, Aojun Zhou, Houxing Ren, Ke Wang, Weikang Shi, Junting Pan, Mingjie Zhan, Hongsheng Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zimu Lu, Aojun Zhou, Houxing Ren, Ke Wang 0036, Weikang Shi, Junting Pan, Mingjie Zhan, Hongsheng Li 0001 |
ACL (1) | 3 |
| 2024 | Empowering Character-level Text Infilling by Eliminating Sub-TokensabstractIn infilling tasks, sub-tokens, representing instances where a complete token is segmented into two parts, often emerge at the boundaries of prefixes, middles, and suffixes.Traditional methods focused on training models at the token level, leading to sub-optimal performance in character-level infilling tasks during the inference stage.Alternately, some approaches considered character-level infilling, but they relied on predicting sub-tokens in inference, yet this strategy diminished ability in character-level infilling tasks due to the large perplexity of the model on sub-tokens.In this paper, we introduce FIM-SE, which stands for Fill-In-the-Middle with both Starting and Ending character constraints.The proposed method addresses character-level infilling tasks by utilizing a line-level format to avoid predicting any sub-token in inference.In addition, we incorporate two special tokens to signify the rest of the incomplete lines, thereby enhancing generation guidance.Extensive experiments demonstrate that our proposed approach surpasses previous methods, offering a significant advantage.Code is available at https://github.com Houxing Ren, Mingjie Zhan, Zhongyuan Wu, Hongsheng Li 0001 |
ACL (1) | 1 |
| 2024 | MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical ReasoningabstractThe recently released GPT-4 Code Interpreter has demonstrated remarkable proficiency in solving challenging math problems, primarily attributed to its ability to seamlessly reason with natural language, generate code, execute code, and continue reasoning based on the execution output. In this paper, we present a method to fine-tune open-source language models, enabling them to use code for modeling and deriving math equations and, consequently, enhancing their mathematical reasoning abilities. We propose a method of generating novel and high-quality datasets with math problems and their code-based solutions, referred to as MathCodeInstruct. Each solution interleaves $\textit{natural language}$, $\textit{code}$, and $\textit{execution results}$. We also introduce a customized supervised fine-tuning and inference approach. This approach yields the MathCoder models, a family of models capable of generating code-based solutions for solving challenging math problems. Impressively, the MathCoder models achieve state-of-the-art scores among open-source LLMs on the MATH (45.2%) and GSM8K (83.9%) datasets, substantially outperforming other open-source alternatives. Notably, the MathCoder model not only surpasses ChatGPT-3.5 and PaLM-2 on GSM8K and MATH but also outperforms GPT-4 on the competition-level MATH dataset. The proposed dataset and models will be released upon acceptance. Ke Wang 0036, Houxing Ren, Aojun Zhou, Zimu Lu, Sichun Luo, Weikang Shi, Renrui Zhang, Linqi Song, Mingjie Zhan, Hongsheng Li 0001 |
ICLR | 2 |
| 2024 | Measuring Multimodal Mathematical Reasoning with MATH-Vision DatasetabstractRecent advancements in Large Multimodal Models (LMMs) have shown promising results in mathematical reasoning within visual contexts, with models exceeding human-level performance on existing benchmarks such as MathVista. However, we observe significant limitations in the diversity of questions and breadth of subjects covered by these benchmarks. To address this issue, we present the MATH-Vision (MATH-V) dataset, a meticulously curated collection of 3,040 high-quality mathematical problems with visual contexts sourced from real math competitions. Spanning 16 distinct mathematical disciplines and graded across 5 levels of difficulty, our dataset provides a comprehensive and diverse set of challenges for evaluating the mathematical reasoning abilities of LMMs. Through extensive experimentation, we unveil a notable performance gap between current LMMs and human performance on \datasetname, underscoring the imperative for further advancements in LMMs. Moreover, our detailed categorization allows for a thorough error analysis of LMMs, offering valuable insights to guide future research and development. The dataset is released at MathLLMs/MathVision Ke Wang 0036, Junting Pan, Weikang Shi, Zimu Lu, Houxing Ren, Aojun Zhou, Mingjie Zhan, Hongsheng Li 0001 |
NeurIPS | 5 |
| 2023 | Self-supervised Trajectory Representation Learning with Temporal Regularities and Travel SemanticsabstractTrajectory Representation Learning (TRL) is a powerful tool for spatial-temporal data analysis and management. TRL aims to convert complicated raw trajectories into low-dimensional representation vectors, which can be applied to various downstream tasks, such as trajectory classification, clustering, and similarity computation. Existing TRL works usually treat trajectories as ordinary sequence data, while some important spatial-temporal characteristics, such as temporal regularities and travel semantics, are not fully exploited. To fill this gap, we propose a novel Self-supervised trajectory representation learning framework with TemporAl Regularities and Travel semantics, namely START. The proposed method consists of two stages. The first stage is a Trajectory Pattern-Enhanced Graph Attention Network (TPE-GAT), which converts the road network features and travel semantics into representation vectors of road segments. The second stage is a Time-Aware Trajectory Encoder (TAT-Enc), which encodes representation vectors of road segments in the same trajectory as a trajectory representation vector, meanwhile incorporating temporal regularities with the trajectory representation. Moreover, we also design two self-supervised tasks, i.e., span-masked trajectory recovery and trajectory contrastive learning, to introduce spatial-temporal characteristics of trajectories into the training process of our START framework. The effectiveness of the proposed method is verified by extensive experiments on two large-scale real-world datasets for three downstream tasks. The experiments also demonstrate that our method can be transferred across different cities to adapt heterogeneous trajectory datasets. Jiawei Jiang 0003, Dayan Pan, Houxing Ren, Chao Li 0001, Jingyuan Wang 0001 |
ICDE | 3 |
| 2022 | RSD: A Reinforced Siamese Network with Domain Knowledge for Early DiagnosisabstractThe availability of electronic health record data makes it possible to develop automatic disease diagnosis approaches. In this paper, we study the early diagnosis of diseases. As being a difficult task (even for experienced doctors), early diagnosis of diseases poses several challenges that are not well solved by prior studies, including insufficient training data, dynamic and complex signs of complications and trade-off between earliness and accuracy. Houxing Ren, Jingyuan Wang 0001, Wayne Xin Zhao |
CIKM | 1 |
| 2022 | Empowering Dual-Encoder with Query Generator for Cross-Lingual Dense RetrievalabstractIn monolingual dense retrieval, lots of works focus on how to distill knowledge from crossencoder re-ranker to dual-encoder retriever and these methods achieve better performance due to the effectiveness of cross-encoder re-ranker.However, we find that the performance of the cross-encoder re-ranker is heavily influenced by the number of training samples and the quality of negative samples, which is hard to obtain in the cross-lingual setting.In this paper, we propose to use a query generator as the teacher in the cross-lingual setting, which is less dependent on enough training samples and highquality negative samples.In addition to traditional knowledge distillation, we further propose a novel enhancement method, which uses the query generator to help the dual-encoder align queries from different languages, but does not need any additional parallel sentences.The experimental results show that our method outperforms the state-of-the-art methods on two benchmark datasets. Houxing Ren, Linjun Shou, Ning Wu 0013, Ming Gong 0001, Daxin Jiang |
EMNLP | 1 |
| 2022 | Unsupervised Context Aware Sentence Representation Pretraining for Multi-lingual Dense RetrievalabstractRecent research demonstrates the effectiveness of using pretrained language models (PLM) to improve dense retrieval and multilingual dense retrieval. In this work, we present a simple but effective monolingual pretraining task called contrastive context prediction (CCP) to learn sentence representation by modeling sentence level contextual relation. By pushing the embedding of sentences in a local context closer and pushing random negative samples away, different languages could form isomorphic structure, then sentence pairs in two different languages will be automatically aligned. Our experiments show that model collapse and information leakage are very easy to happen during contrastive training of language model, but language-specific memory bank and asymmetric batch normalization operation play an essential role in preventing collapsing and information leakage, respectively. Besides, a post-processing for sentence embedding is also very effective to achieve better retrieval performance. On the multilingual sentence retrieval task Tatoeba, our model achieves new SOTA results among methods without using bilingual data. Our model also shows larger gain on Tatoeba when transferring between non-English pairs. On two multi-lingual query-passage retrieval tasks, XOR Retrieve and Mr.TYDI, our model even achieves two SOTA results in both zero-shot and supervised setting among all pretraining models using bilingual data. Ning Wu 0013, Yaobo Liang, Houxing Ren, Linjun Shou, Nan Duan 0001, Ming Gong 0001, Daxin Jiang |
IJCAI | 3 |
| 2022 | Generative Adversarial Networks Enhanced Pre-training for Insufficient Electronic Health Records ModelingabstractIn recent years, automatic computational systems based on deep learning are widely used in medical fields, such as automatic diagnosing and disease prediction. Most of these systems are designed for data sufficient scenarios. However, due to the disease rarity or privacy, the medical data are always insufficient. When applying these data-hungry deep learning models with insufficient data, it is likely to lead to issues of over-fitting and cause serious performance problems. Many data augmentation methods have been proposed to solve the data insufficiency problem, such as using GAN (Generative Adversarial Networks) to generate training data. However, the augmented data usually contains lots of noise. Directly using them to train sensitive medical models is very difficult to achieve satisfactory results. Houxing Ren, Jingyuan Wang 0001, Wayne Xin Zhao |
KDD | 1 |
| 2021 | RAPT: Pre-training of Time-Aware Transformer for Learning Robust Healthcare RepresentationabstractWith the development of electronic health records (EHRs), prenatal care examination records have become available for developing automatic prediction or diagnosis approaches with machine learning methods. In this paper, we study how to effectively learn representations applied to various downstream tasks for EHR data. Although several methods have been proposed in this direction, they usually adapt classic sequential models to solve one specific diagnosis task or address unique EHR data issues. This makes it difficult to reuse these existing methods for the early diagnosis of pregnancy complications or provide a general solution to address the series of health problems caused by pregnancy complications. In this paper, we propose a novel model RAPT, which stands for RepresentAtion by Pre-training time-aware Transformer. To associate pre-training and EHR data, we design an architecture that is suitable for both modeling EHR data and pre-training, namely time-aware Transformer. To handle various characteristics in EHR data, such as insufficiency, we carefully devise three pre-training tasks to handle data insufficiency, data incompleteness and short sequence problems, namely similarity prediction, masked prediction and reasonability check. In this way, our representations can capture various EHR data characteristics. Extensive experimental results for four downstream tasks have shown the effectiveness of the proposed approach. We also introduce sensitivity analysis to interpret the model and design an interface to show results and interpretation for doctors. Finally, we implement a diagnosis system for pregnancy complications based on our pre-training model. Doctors and pregnant women can benefit from the diagnosis system in early diagnosis of pregnancy complications. Houxing Ren, Jingyuan Wang 0001, Wayne Xin Zhao |
KDD | 1 |