VLDB 2026 Research / reviewers in the wild / expert
Haoze Sun
dblp:177/9281
· DBLP profile ↗
14ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis
Wenbo Li 0002, Zhongdao Wang, Haoze Sun, Bangzhen Liu, Haoyu Chen 0003, Aoxue Li, Lei Zhu 0003 |
ICCV | 4 |
| 2025 | Fast Image Super-Resolution via Consistency Rectified Flow
Wenbo Li 0002, Haoze Sun, Zhixin Wang, Long Peng 0003, Xiaowei Hu 0001, Renjing Pei, Pheng-Ann Heng |
ICCV | 3 |
| 2025 | Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction TuningabstractLarge Language Models (LLMs) have exhibited significant potential in performing diverse tasks, including the ability to call functions or use external tools to enhance their performance. While current research on function calling by LLMs primarily focuses on single-turn interactions, this paper addresses the overlooked necessity for LLMs to engage in multi-turn function calling—critical for handling compositional, real-world queries that require planning with functions but not only use functions. To facilitate this, we introduce an approach, BUTTON, which generates synthetic compositional instruction tuning data via bottom-up instruction construction and top-down trajectory generation. In the bottom-up phase, we generate simple atomic tasks based on real-world scenarios and build compositional tasks using heuristic strategies based on atomic tasks. Corresponding function definitions are then synthesized for these compositional tasks. The top-down phase features a multi-agent environment where interactions among simulated humans, assistants, and tools are utilized to gather multi-turn function calling trajectories. This approach ensures task compositionality and allows for effective function and trajectory generation by examining atomic tasks within compositional tasks. We produce a dataset BUTTONInstruct comprising 8k data points and demonstrate its effectiveness through extensive experiments across various LLMs. Mingyang Chen 0002, Haoze Sun, Tianpeng Li, Fan Yang 0132, Hao Liang 0017, Keer Lu, Bin Cui 0001, Wentao Zhang 0001, Zenan Zhou, Weipeng Chen |
ICLR | 2 |
| 2025 | SysBench: Can LLMs Follow System Message?abstractLarge Language Models (LLMs) have become instrumental across various applications, with the customization of these models to specific scenarios becoming increasingly critical. System message, a fundamental component of LLMs, is consist of carefully crafted instructions that guide the behavior of model to meet intended goals. Despite the recognized potential of system messages to optimize AI-driven solutions, there is a notable absence of a comprehensive benchmark for evaluating how well LLMs follow system messages. To fill this gap, we introduce SysBench, a benchmark that systematically analyzes system message following ability in terms of three limitations of existing LLMs: constraint violation, instruction misjudgement and multi-turn instability. Specifically, we manually construct evaluation dataset based on six prevalent types of constraints, including 500 tailor-designed system messages and multi-turn user conversations covering various interaction relationships. Additionally, we develop a comprehensive evaluation protocol to measure model performance. Finally, we conduct extensive evaluation across various existing LLMs, measuring their ability to follow specified constraints given in system messages. The results highlight both the strengths and weaknesses of existing models, offering key insights and directions for future research. Yanzhao Qin, Tao Zhang 0194, Wenjing Luo, Haoze Sun, Yan Zhang 0109, Yujing Qiao, Weipeng Chen, Zenan Zhou, Wentao Zhang 0001, Bin Cui 0001 |
ICLR | 6 |
| 2025 | ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningabstractLarge Language Models (LLMs) have shown remarkable capabilities in reasoning, exemplified by the success of OpenAI-o1 and DeepSeek-R1. However, integrating reasoning with external search processes remains challenging, especially for complex multi-hop questions requiring multiple retrieval steps. We propose ReSearch, a novel framework that trains LLMs to Reason with Search via reinforcement learning without using any supervised data on reasoning steps. Our approach treats search operations as integral components of the reasoning chain, where when and how to perform searches is guided by text-based thinking, and search results subsequently influence further reasoning. We train ReSearch on Qwen2.5-7B(-Instruct) and Qwen2.5-32B(-Instruct) models and conduct extensive experiments. Despite being trained on only one dataset, our models demonstrate strong generalizability across various benchmarks. Analysis reveals that ReSearch naturally elicits advanced reasoning capabilities such as reflection and self-correction during the reinforcement learning process. Mingyang Chen 0002, Linzhuang Sun, Tianpeng Li, Haoze Sun, Chenzheng Zhu, Haofen Wang, Jeff Z. Pan, Wen Zhang 0015, Huajun Chen, Fan Yang 0132, Zenan Zhou, Weipeng Chen |
NeurIPS | 4 |
| 2025 | PocketSR: The Super-Resolution Expert in Your Pocket MobilesabstractReal-world image super-resolution (RealSR) aims to enhance the visual quality of in-the-wild images, such as those captured by mobile phones. While existing methods leveraging large generative models demonstrate impressive results, the high computational cost and latency make them impractical for edge deployment. In this paper, we introduce PocketSR, an ultra-lightweight, single-step model that brings generative modeling capabilities to RealSR while maintaining high fidelity. To achieve this, we design LiteED, a highly efficient alternative to the original computationally intensive VAE in SD, reducing parameters by 97.5\% while preserving high-quality encoding and decoding. Additionally, we propose online annealing pruning for the U-Net, which progressively shifts generative priors from heavy modules to lightweight counterparts, ensuring effective knowledge transfer and further optimizing efficiency. To mitigate the loss of prior knowledge during pruning, we incorporate a multi-layer feature distillation loss. Through an in-depth analysis of each design component, we provide valuable insights for future research. PocketSR, with a model size of 146M parameters, processes 4K images in just 0.8 seconds, achieving a remarkable speedup over previous methods. Notably, it delivers performance on par with state-of-the-art single-step and even multi-step RealSR models, making it a highly practical solution for edge-device applications. Haoze Sun, Linfeng Jiang, Renjing Pei, Zhixin Wang, Haoyu Chen 0003, Fenglong Song, Yujiu Yang 0001, Wenbo Li 0002 |
NeurIPS | 1 |
| 2024 | Low-Res Leads the Way: Improving Generalization for Super-Resolution by Self-Supervised LearningabstractFor image super-resolution (SR), bridging the gap between the performance on synthetic datasets and real-world degradation scenarios remains a challenge. This work introduces a novel “Low-Res Leads the Way” (LWay) training framework, merging Supervised Pre-training with Self-supervised Learning to enhance the adaptability of SR models to real-world images. Our approach utilizes a low-resolution (LR) reconstruction network to extract degradation embeddings from LR images, merging them with super-resolved outputs for LR reconstruction. Leveraging unseen LR images for self-supervised learning guides the model to adapt its modeling space to the target domain, facili-tating fine-tuning of SR models without requiring paired high-resolution (HR) images. The integration of Discrete Wavelet Transform (DWT)further refines the focus on high-frequency details. Extensive evaluations show that our method significantly improves the generalization and de-tail restoration capabilities of SR models on unseen real-world datasets, outperforming existing methods. Our training regime is universally compatible, requiring no network architecture modifications, making it a practical solution for real-world SR applications. Haoyu Chen 0003, Wenbo Li 0002, Jinjin Gu, Haoze Sun, Xueyi Zou, Zhensong Zhang, Youliang Yan, Lei Zhu 0003 |
CVPR | 5 |
| 2024 | CoSeR: Bridging Image and Language for Cognitive Super-ResolutionabstractExisting super-resolution (SR) models primarily focus on restoring local texture details, often neglecting the global semantic information within the scene. This oversight can lead to the omission of crucial semantic details or the intro-duction of inaccurate textures during the recovery process. In our work, we introduce the Cognitive Super-Resolution (CoSeR) framework, empowering SR models with the ca-pacity to comprehend low-resolution images. We achieve this by marrying image appearance and language under-standing to generate a cognitive embedding, which not only activates prior information from large text-to-image diffusion models but also facilitates the generation of high-quality reference images to optimize the SR process. To fur-ther improve image fidelity, we propose a novel condition injection scheme called “Ali-in-Attention ”, consolidating all conditional information into a single module. Conse-quently, our method successfully restores semantically cor-rect and photorealistic details, demonstrating state-of-the-art performance across multiple benchmarks. Project page: https://coser-main.github.io/ Haoze Sun, Wenbo Li 0002, Jianzhuang Liu, Haoyu Chen 0003, Renjing Pei, Xueyi Zou, Youliang Yan, Yujiu Yang 0001 |
CVPR | 1 |
| 2024 | Accelerating Diffusion Models for Inverse Problems through Shortcut Sampling
Gongye Liu, Haoze Sun, Jiayi Li 0002, Yujiu Yang 0001 |
IJCAI | 2 |
| 2019 | Armored Target Detection in Battlefield Environment Based on Top-Down Aggregation Network and Hierarchical Scale OptimizationabstractArmored equipment plays a crucial role in the ground battlefield. The fast and accurate detection of enemy armored targets is significant to take the initiative in the battlefield. Comparing to general object detection and vehicle detection, armored target detection in battlefield environment is more challenging due to the long distance of observation and the complicated environment. In this paper, an accurate and robust automatic detection method is proposed to detect armored targets in battlefield environment. Firstly, inspired by Feature Pyramid Network (FPN), we propose a top-down aggregation (TDA) network which enhances shallow feature maps by aggregating semantic information from deeper layers. Then, using the proposed TDA network in a basic Faster R-CNN framework, we explore the further optimization of the approach for armored target detection: for the Region of Interest (RoI) Proposal Network (RPN), we propose a multi-branch RPNs framework to generate proposals that match the scale of armored targets and the specific receptive field of each aggregated layer and design hierarchical loss for the multi-branch RPNs; for RoI Classifier Network (RCN), we apply RoI pooling on the single finest scale feature map and construct a light and fast detection network. To evaluate our method, comparable experiments with state-of-art detection methods were conducted on a challenging dataset of images with armored targets. The experimental results demonstrate the effectiveness of the proposed method in terms of detection accuracy and recall rate. Haoze Sun, Tianqing Chang, Guozhen Yang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2018 | TextDream: Conditional Text Generation by Searching in the Semantic SpaceabstractConditional text generation is a fundamental task in natural language generation. Traditional conditional generative models build conditional probability distributions over the given labels. However, categorical label information is usually very abstract, e.g., sentiment, and it is difficult to be disentangled from the content. Therefore, instead of generating text by modeling conditional probability distribution, we propose a novel text generation method TextDream through searching in the semantic space. Specifically, in this method, a random text seed is initially given and the new text is generated by local search operation. The generation procedure is guided by a fitness function, typically a classification model. Text with higher fitness will be preserved. This procedure loops until the qualified solution is found. Experimental results show that our method is able to generate more diverse text compared with advanced conditional generative models. Weidi Xu, Haoze Sun, Ying Tan 0002 |
CEC | 2 |
| 2018 | An improved artificial bee colony algorithm based on elite group guidance and combined breadth-depth search strategy
Depeng Kong, Tianqing Chang, Wenjun Dai, Quandong Wang, Haoze Sun |
Inf. Sci. | 5 |
| 2017 | Variational Autoencoder for Semi-Supervised Text ClassificationabstractAlthough semi-supervised variational autoencoder (SemiVAE) works in image classification task, it fails in text classification task if using vanilla LSTM as its decoder. From a perspective of reinforcement learning, it is verified that the decoder's capability to distinguish between different categorical labels is essential. Therefore, Semi-supervised Sequential Variational Autoencoder (SSVAE) is proposed, which increases the capability by feeding label into its decoder RNN at each time-step. Two specific decoder structures are investigated and both of them are verified to be effective. Besides, in order to reduce the computational complexity in training, a novel optimization method is proposed, which estimates the gradient of the unlabeled objective function by sampling, along with two variance reduction techniques. Experimental results on Large Movie Review Dataset (IMDB) and AG's News corpus show that the proposed approach significantly improves the classification accuracy compared with pure-supervised classifiers, and achieves competitive performance against previous advanced methods. State-of-the-art results can be obtained by integrating other pretraining-based methods. Weidi Xu, Haoze Sun |
AAAI | 2 |
| 2016 | Multi-digit image synthesis using recurrent conditional variational autoencoderabstractIn the field of deep neural networks, several generative methods have been proposed to address the challenges from generative and discriminative tasks, e.g., natural language process, image caption and image generation. In this paper, a conditional recurrent variational autoencoder is proposed for multi-digit image synthesis. This model is capable of generating multi-digit images from the given number sequences and retaining the generalisation ability to recover different types of background. Our method is evaluated on SVHN dataset and the experimental results show it succeeds to generate multi-digit images with various styles according to the given sequential inputs. The generated images can also be easily identified by both human beings and convolutional neural networks for digit classification. Haoze Sun, Weidi Xu, Ying Tan 0002 |
IJCNN | 1 |