Haotian Hou

dblp:407/7925 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%
Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 50% Software testing · 50%
Artificial intelligence
2 papers
Language models and text generation · 56% Multi-agent systems · 44%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation
code generation with language models
1.722025
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch · NeurIPS 2025
Alignment with Fill-In-the-Middle for Enhancing Code Generation · EMNLP 2025
Visual content generation and editing
image editing
1.012026
RAA: Achieving Interactive Remove/Add Anything via Fully Synthetic Data · AAAI 2026
Visual content generation and editing › image editing › text-guided image editing
instruction-based image editing
1.012026
RAA: Achieving Interactive Remove/Add Anything via Fully Synthetic Data · AAAI 2026
Visual content generation and editing › image editing › object-level image editing
object insertion and removal
1.012026
RAA: Achieving Interactive Remove/Add Anything via Fully Synthetic Data · AAAI 2026
Visual content generation and editing
synthetic data generation
1.012026
RAA: Achieving Interactive Remove/Add Anything via Fully Synthetic Data · AAAI 2026
Software testing
test generation
0.912025
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch · NeurIPS 2025
Software testing
web application testing
0.912025
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

fill-in-the-middle · 1.7alignment · 1.7open-set object detection · 1.0multimodal reasoning · 1.0multimodal large language model · 1.0multi-agent reasoning · 1.0large language model · 1.0diffusion transformer · 1.0web-navigation agent · 0.9large language model agent · 0.9
YearPublicationVenuePosition
2026 RAA: Achieving Interactive Remove/Add Anything via Fully Synthetic Data
abstract
Precise and controllable image editing, especially object removal and insertion, represents one of the most common demands in image manipulation. However, existing methods suffer from severe limitations. Mask-based inpainting often introduces visual artifacts and semantic inconsistencies, while instruction-based approaches lack accurate spatial control and tend to unintentionally modify background regions. To address these issues, we propose two key contributions. First, we develop a fully automated and self-improving pipeline for synthetic data generation. This pipeline utilizes a Large Language Model (LLM) to generate diverse prompts, a Diffusion Transformer (DiT) fine-tuned evolutionarily to synthesize high-quality images, and a Multimodal LLM (MLLM) combined with open-set object detector for automated quality control and annotation. This process produces the Remove/Add Dataset (RAD), consisting of over 514,510 high-quality image pairs, each richly annotated with bounding boxes, segmentation masks, and a variety of editing instructions. Second, based on RAD, we introduce Remove/Add Anything (RAA), a novel editing framework with precise spatial control. Built upon a diffusion-based inpainting model, RAA achieves high editing accuracy by conditioning on both textual instructions and an explicitly defined region of interest (ROI), enabling efficient fine-tuning while maintaining global visual coherence. Extensive experiments demonstrate that RAA significantly outperforms existing open-source methods on both addition and removal tasks, and even slightly surpasses costly proprietary models.
Delong Liu, Haotian Hou, Zhaohui Hou, Shihao Han, Mingjie Zhan, Zhicheng Zhao 0001
AAAI2
2026 Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning
abstract
Houxing Ren, Mingjie Zhan, Zimu Lu, Ke Wang, Yunqiao Yang, Haotian Hou, Hongsheng Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Houxing Ren, Mingjie Zhan, Zimu Lu, Ke Wang 0036, Yunqiao Yang 0002, Haotian Hou, Hongsheng Li 0001
ACL (1)6
2025 Alignment with Fill-In-the-Middle for Enhancing Code Generation
abstract
Houxing Ren, Zimu Lu, Weikang Shi, Haotian Hou, Yunqiao Yang, Ke Wang, Aojun Zhou, Junting Pan, Mingjie Zhan, Hongsheng Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Houxing Ren, Zimu Lu, Weikang Shi, Haotian Hou, Yunqiao Yang 0002, Ke Wang 0036, Aojun Zhou, Junting Pan, Mingjie Zhan, Hongsheng Li 0001
EMNLP4
2025 All-in-One Medical Image Restoration with Latent Diffusion-Enhanced Vector-Quantized Codebook Prior
Zhiwen Yang 0001, Haotian Hou, Hui Zhang 0099, Bingzheng Wei, Yan Xu 0001
MICCAI (16)3
2025 WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
abstract
LLM‑based agents have demonstrated great potential in generating and managing code within complex codebases. In this paper, we introduce WebGen-Bench, a novel benchmark designed to measure an LLM-based agent's ability to create multi-file website codebases from scratch. It contains diverse instructions for website generation, created through the combined efforts of human annotators and GPT-4o. These instructions span three major categories and thirteen minor categories, encompassing nearly all important types of web applications.To assess the quality of the generated websites, we generate test cases targeting each functionality described in the instructions. These test cases are then manually filtered, refined, and organized to ensure accuracy, resulting in a total of 647 test cases. Each test case specifies an operation to be performed on the website and the expected outcome of the operation.To automate testing and improve reproducibility, we employ a powerful web-navigation agent to execute test cases on the generated websites and determine whether the observed responses align with the expected results.We evaluate three high-performance code-agent frameworks—Bolt.diy, OpenHands, and Aider—using multiple proprietary and open-source LLMs as engines. The best-performing combination, Bolt.diy powered by DeepSeek-R1, achieves only 27.8\% accuracy on the test cases, highlighting the challenging nature of our benchmark.Additionally, we construct WebGen-Instruct, a training set consisting of 6,667 website-generation instructions. Training Qwen2.5-Coder-32B-Instruct on Bolt.diy trajectories generated from a subset of the training set achieves an accuracy of 38.2\%, surpassing the performance of the best proprietary model.We release our data-generation, training, and testing code, along with both the datasets and model weights at https://github.com/mnluzimu/WebGen-Bench.
Zimu Lu, Yunqiao Yang 0002, Houxing Ren, Haotian Hou, Han Xiao 0010, Ke Wang 0036, Weikang Shi, Aojun Zhou, Mingjie Zhan, Hongsheng Li 0001
NeurIPS4