Ermo Hua

dblp:372/0768 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
12since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism
abstract
Yuhua Jiang, Shuang Cheng, Yihao Liu, Ermo Hua, Che Jiang, Weigao Sun, Yu Cheng, Feifei Gao, Biqing Qi, Bowen Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuhua Jiang, Shuang Cheng, Yihao Liu 0008, Ermo Hua, Che Jiang, Weigao Sun, Yu Cheng 0001, Biqing Qi, Bowen Zhou 0002
ACL (1)4
2026 VC-VTON: Toward Across-View and Multi-Posture-Driven Virtual Try-On via Spatiotemporal-Aware View-Consistency Training
abstract
Virtual try-on (VTON) aims to synthesize specific fashion images dressed in given garments, which possesses great potential in real-world scenarios. Existing methods generally stand on the shoulder of the single-view VTON to train a warping model and then fit the given garments onto the human body under a fixed posture and viewpoint, which often fails to preserve the consistent garment characteristics in across-view and multi-pose guided try-on scenarios due to the lack of both across-view data and effective view consistency training. To alleviate this dilemma, we propose a fresh view consistency-driven VTON task (VC-VTON) and release a multi-view virtual try-on dataset with complete annotation (e.g., viewpoint, text, posture, parsing maps, etc.) to encourage across-view training scenarios. Based on this hard-won dataset, we further propose VC-TwinNet, a Twin-UNet baseline based on spatiotemporal-aware View Consistency training, designed specifically for the challenging task. Specifically, to enable view-aware denoising and sparse-to-continuous view generalization, we introduce RoPE and circle embedding to represent the relative and continuous position relation across viewpoints, serving to distinguish their outfitting appearance and warping states. Afterwards, to implicitly learn the interactions across views under given multiple posture conditions, we further contribute a spatiotemporal-aware view attention module to capture the spatial and temporal details for across-view training. Moreover, we utilize an across-view consistency loss to supervise the model training, to ultimately improve the performance of our VC-VTON. Extensive experiments demonstrate the superiority of our approach and state-of-the-art results on various evaluations without declining single-view performance. And as for practicality and timeliness, our proposed components are essentially plug-and-play and remain effective in the new DiT-centered paradigm.
Zhiyuan Ma 0005, Jiabao Wei, Zhihan Cai, Chundi Yang, Ermo Hua, Shulei Xie, Jianjun Li 0010, Bowen Zhou 0002
IEEE Trans. Image Process.5
2025 Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines
abstract
Retrieval-augmented generation (RAG) has emerged to address the knowledge-intensive visual question answering (VQA) task. Current methods mainly employ separate retrieval and generation modules to acquire external knowledge and generate answers, respectively. We propose ReAuSE, an alternative to the previous RAG model for the knowledge-based VQA task, which seamlessly integrates knowledge retriever into the generative multi-modal large language model, serving as a built-in search engine. Specifically, our model functions both as a generative retriever and an accurate answer generator. It not only helps retrieve documents from the knowledge base by producing identifier for each document, but it also answers visual questions based on the retrieved documents. Furthermore, we also propose a reinforced retrieval calibration module from relevance feedback to improve retrieval performance and align with the preferences for accurate answer generation. Extensive experiments on two representative OKVQA and A-OKVQA datasets demonstrate significant improvements ranging from 2.9% to 9.6% across all evaluation metrics when compared to strong baselines.
Xinwei Long, Zhiyuan Ma 0005, Ermo Hua, Biqing Qi, Bowen Zhou 0002
AAAI3
2025 Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process
abstract
Supervised Fine-Tuning (SFT) and Preference Optimization (PO) are key processes for aligning Language Models (LMs) with human preferences post pre-training.While SFT excels in efficiency and PO in effectiveness, they are often combined sequentially without integrating their optimization objectives.This approach ignores the opportunities to bridge their paradigm gap and take the strengths from both.In this paper, we interpret SFT and PO with two subprocesses -Preference Estimation and Transition Optimization -defined at token level within the Markov Decision Process (MDP).This modeling shows that SFT is only a special case of PO with inferior estimation and optimization.PO estimates the model's preference by its entire generation, while SFT only scores model's subsequent predicted tokens based on prior tokens from ground truth answer.These priors deviates from model's distribution, hindering the preference estimation and transition optimization.Building on this view, we introduce Intuitive Fine-Tuning (IFT) to integrate SFT and PO into a single process.Through a temporal residual connection, IFT brings better estimation and optimization by capturing LMs' intuitive sense of its entire answers.But it solely relies on a single policy and the same volume of non-preference-labeled data as SFT.Our experiments show that IFT performs comparably or even superiorly to SFT and some typical PO methods across several tasks, particularly those requires generation, reasoning, and fact-following abilities.An explainable Frozen Lake game further validates the effectiveness of IFT for getting competitive policy.
Ermo Hua, Biqing Qi, Xingtai Lv, Ning Ding 0002, Bowen Zhou 0002
ACL (1)1
2025 OpenPRM: Building Open-domain Process-based Reward Models with Preference Trees
abstract
Scaling inference-time computation is increasingly seen as the next frontier in scaling laws for large language models. Previous work in mathematics and coding has demonstrated the remarkable potential for inference-time scaling. During such scaling, fine-grained supervision through process-based reward models (PRMs) is essential for enhancement. However, exploration of inference-time scaling and PRMs in open-domain problems remains limited, where lacking exact answers and obtaining process supervision prove challenging. In this paper, we explore the construction of PRMs for open-domain tasks, specifically for instruction-following tasks. Utilizing existing outcome-based reward models (ORMs), we develop sentence-level preference trees based on the prefix similarity of parallel sampled candidates from datasets like UltraFeedback. This setup allows us to derive weak supervision for processes via back-propagation from outcome-level rewards. Subsequently, we integrate ORMs and PRMs under the same pairwise ranking objectives, resulting in our newly developed reward models, named OpenPRM. This approach significantly enhances the scalability of process-level supervision in open domains at minimal cost. We assess the performance of OpenPRM across various reward benchmarks, demonstrating its competitive edge over traditional ORMs in open domains and PRMs in specialized domains. Additionally, we investigate the scalability of inference-time computation for open-domain instructions. Our results highlight the limitations of ORMs’ scalability, while OpenPRM shows superior performance in scaled settings. Despite these advances, achieving automatic fine-grained supervision for open-domain inference-time scaling remains a substantial challenge. We hope these findings will spur further development of process supervision reward models in open-domain scenarios.
Jiayuan Zhang 0001, Haoxin Li, Xuekai Zhu, Ermo Hua, Xingtai Lv, Ning Ding 0002, Biqing Qi, Bowen Zhou 0002
ICLR5
2025 Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization
abstract
Extending the context length of Language Models (LMs) by improving Rotary Position Embedding (RoPE) has become a trend. While prior works mainly address RoPE’s limitations within attention, this paper uncovers the adverse effects on length generalization from nearly all parts of LMs. Using Discrete Signal Processing theory, we show that RoPE enables periodic attention by implicitly achieving Non-Uniform Discrete Fourier Transform. However, this periodicity is undermined by the spectrum damage caused by: 1) linear layers and activation functions outside of attention; 2) insufficiently trained frequency components brought by time-domain truncation. Building on our observations, we propose Fourier Position Embedding (FoPE), which enhances attention’s frequency-domain properties to improve both its periodic extension and length generalization. FoPE constructs Fourier Series and zero-outs the destructive frequency components, increasing model robustness against the spectrum damage. Experiments across various model scales and benchmarks show that, within varying context windows, FoPE maintains a more stable performance compared to other baselines. Several analyses and ablations bring further support to our method and theoretical modeling.
Ermo Hua, Che Jiang, Xingtai Lv, Youbang Sun, Yuchen Fan 0001, Xuekai Zhu, Biqing Qi, Ning Ding 0002, Bowen Zhou 0002
ICML1
2025 How to Synthesize Text Data without Model Collapse?
abstract
Model collapse in synthetic data indicates that iterative training on self-generated data leads to a gradual decline in performance. With the proliferation of AI models, synthetic data will fundamentally reshape the web data ecosystem. Future GPT-$\{n\}$ models will inevitably be trained on a blend of synthetic and human-produced data. In this paper, we focus on two questions: what is the impact of synthetic data on language model training, and how to synthesize data without model collapse? We first pre-train language models across different proportions of synthetic data, revealing a negative correlation between the proportion of synthetic data and model performance. We further conduct statistical analysis on synthetic data to uncover distributional shift phenomenon and over-concentration of n-gram features. Inspired by the above findings, we propose token editing on human-produced data to obtain semi-synthetic data. As a proof of concept, we theoretically demonstrate that token-level editing can prevent model collapse, as the test error is constrained by a finite upper bound. We conduct extensive experiments on pre-training from scratch, continual pre-training, and supervised fine-tuning. The results validate our theoretical proof that token-level editing improves data quality and enhances model performance.
Xuekai Zhu, Daixuan Cheng, Hengli Li, Ermo Hua, Xingtai Lv, Ning Ding 0002, Zhouhan Lin, Zilong Zheng, Bowen Zhou 0002
ICML5
2025 MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
abstract
We introduce MedXpertQA, a highly challenging and comprehensive benchmark to evaluate expert-level medical knowledge and advanced reasoning. MedXpertQA includes 4,460 questions spanning 17 specialties and 11 body systems. It includes two subsets, Text for text evaluation and MM for multimodal evaluation. Notably, MM introduces expert-level exam questions with diverse images and rich clinical information, including patient records and examination results, setting it apart from traditional medical multimodal benchmarks with simple QA pairs generated from image captions. MedXpertQA applies rigorous filtering and augmentation to address the insufficient difficulty of existing benchmarks like MedQA, and incorporates specialty board questions to improve clinical relevance and comprehensiveness. We perform data synthesis to mitigate data leakage risk and conduct multiple rounds of expert reviews to ensure accuracy and reliability. We evaluate 18 leading models on MedXpertQA. Moreover, medicine is deeply connected to real-world decision-making, providing a rich and representative setting for assessing reasoning abilities beyond mathematics and code. To this end, we develop a reasoning-oriented subset to facilitate the assessment of o1-like models.
Yuxin Zuo, Shang Qu, Zhang-Ren Chen, Xuekai Zhu, Ermo Hua, Ning Ding 0002, Bowen Zhou 0002
ICML6
2025 TTRL: Test-Time Reinforcement Learning
abstract
This paper investigates Reinforcement Learning (RL) on data without explicit labels for reasoning tasks in Large Language Models (LLMs). The core challenge of the problem is reward estimation during inference while not having access to ground-truth information. While this setting appears elusive, we find that common practices in Test-Time Scaling (TTS), such as majority voting, yield surprisingly effective rewards suitable for driving RL training. In this work, we introduce Test-Time Reinforcement Learning (TTRL), a novel method for training LLMs using RL on unlabeled data. TTRL enables self-evolution of LLMs by utilizing the priors in the pre-trained models. Our experiments demonstrate that TTRL consistently improves performance across a variety of tasks and models. Notably, TTRL boosts the pass@1 performance of Qwen-2.5-Math-7B by approximately 211% on the AIME 2024 with only unlabeled test data. Furthermore, although TTRL is only supervised by the Maj@N metric, TTRL has demonstrated performance to consistently surpass the upper limit of the initial model, and approach the performance of models trained directly on test data with ground-truth labels. Our experimental findings validate the general effectiveness of TTRL across various tasks and highlight TTRL's potential for broader tasks and domains.
Yuxin Zuo, Shang Qu, Ganqu Cui, Xuekai Zhu, Haozhan Li, Xinwei Long, Ermo Hua, Biqing Qi, Youbang Sun, Zhiyuan Ma 0005, Lifan Yuan, Ning Ding 0002, Bowen Zhou 0002
NeurIPS10
2024 CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction Following
abstract
With the advancement of language models (LMs), their exposure to private data is increasingly inevitable, and their deployment (especially for smaller ones) on personal devices, such as PCs and smartphones, has become a prevailing trend.In contexts laden with user information, enabling models to both safeguard user privacy and execute commands efficiently emerges as an essential research imperative.In this paper, we propose CoGenesis, a collaborative generation framework integrating large (hosted on cloud infrastructure) and small models (deployed on local devices) to address privacy concerns logically.Initially, we design a pipeline to create personalized writing instruction datasets enriched with extensive context details as the testbed of this research issue.Subsequently, we introduce two variants of CoGenesis based on sketch and logits respectively.Our experimental findings, based on our synthesized dataset and two additional open-source datasets, indicate that: 1) Large-scale models perform well when provided with user context but struggle in the absence of such context.2) While specialized smaller models fine-tuned on the synthetic dataset show promise, they still lag behind their larger counterparts.3) Our CoGenesis framework, utilizing mixed-scale models, showcases competitive performance, providing a feasible solution to privacy issues.* Corresponding author 1 This paper defines large LMs (LLMs) as both closed and open-source models, designed for universal application and advanced performance, and intended for cloud deployment.Conversely, small LMs (SLMs) refer to models tailored for specific tasks and deployed on local devices.
Jianyu Wang 0012, Ermo Hua, Biqing Qi, Ning Ding 0002, Bowen Zhou 0002
ACL (1)3
2024 Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
abstract
Improving the effectiveness and efficiency of large language models (LLMs) simultaneously is a critical yet challenging research goal.In this paper, we find that low-rank pre-training, normally considered as efficient methods that will compromise performance, can be scalably effective when reduced parameters are precisely targeted.Specifically, applying the low-dimensional module only to the attention layer -resolves this issue and enhances both effectiveness and efficiency.We refer to this structure as Low-dimensional Projected Attention (LPA) and provide an explanatory analysis.Through extensive experimentation at parameter scales of 130M, 370M, and scaling up to 3B, we have validated the effectiveness and scalability of LPA.Our results show that LPA model can save up to 12.4% in time while achieving an approximate 5% improvement in test perplexity (ppl) and on downstream tasks compared with the vanilla Transformer.
Xingtai Lv, Ning Ding 0002, Ermo Hua, Ganqu Cui, Bowen Zhou 0002
EMNLP4
2024 UltraMedical: Building Specialized Generalists in Biomedicine
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains and are moving towards more specialized areas. Recent advanced proprietary models such as GPT-4 and Gemini have achieved significant advancements in biomedicine, which have also raised privacy and security challenges. The construction of specialized generalists hinges largely on high-quality datasets, enhanced by techniques like supervised fine-tuning and reinforcement learning from human or AI feedback, and direct preference optimization. However, these leading technologies (e.g., preference learning) are still significantly limited in the open source community due to the scarcity of specialized data. In this paper, we present the UltraMedical collections, which consist of high-quality manual and synthetic datasets in the biomedicine domain, featuring preference annotations across multiple advanced LLMs. By utilizing these datasets, we fine-tune a suite of specialized medical models based on Llama-3 series, demonstrating breathtaking capabilities across various medical benchmarks. Moreover, we develop powerful reward models skilled in biomedical and general reward benchmark, enhancing further online preference learning within the biomedical LLM community.
Sihang Zeng, Ermo Hua, Ning Ding 0002, Zhang-Ren Chen, Zhiyuan Ma 0005, Haoxin Li, Ganqu Cui, Biqing Qi, Xuekai Zhu, Xingtai Lv, Jinfang Hu, Zhiyuan Liu 0001, Bowen Zhou 0002
NeurIPS3