VLDB 2026 Research / reviewers in the wild / expert
Haonan Lu
dblp:129/0998
· DBLP profile ↗
34ranked-venue papers
6as first author
28since 2021 · last 2026
0000-0001-6332-2785ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 1 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 18 since 2021Computer networks · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 4 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DRIFT: Difference-Aware Reinforcement Through Iterative Fine-Tuning for Language ModelabstractSelf-play fine-tuning has emerged as a promising approach to improve Large Language Models (LLMs) without additional human annotations. However, existing methods struggle with complex generation tasks requiring long context understanding, where models produce partially correct outputs interleaved with errors. Traditional approaches train on entire sequences uniformly, failing to distinguish between well-predicted and erroneous regions, leading to diluted learning signals and slow convergence. We propose DRIFT (Difference-aware Reinforcement through Iterative Fine-Tuning), a novel self-play framework that selectively trains on prediction differences. DRIFT introduces two key innovations: (1) Difference-Aware Masking (DAM) that identifies and masks common subsequences between model outputs and ground truth, focusing training exclusively on error regions; (2) Occurrence-Aware Loss (OAL) that provides position-invariant vocabulary supervision, complementing the position-sensitive adversarial loss. This dual mechanism enables models to correct both positional and lexical errors effectively. Theoretically, we prove that DRIFT converges when masked distributions align. Empirically, we evaluate DRIFT on diverse summarization benchmarks using Qwen2.5-3B and LLaMA-3.1-8B models. Results show that DRIFT significantly outperforms both supervised fine-tuning (SFT) and self-play fine-tuning (SPIN), achieving up to 16\% improvement on SAMSum dialogue summarization tasks while maintaining general capabilities. Notably, DRIFT breaks the performance ceiling of continued SFT and demonstrates superior efficiency compared to holistic self-play methods, validating that targeted optimization on prediction differences is crucial for structured text generation tasks. Wenjie Liao, Haonan Lu |
AAAI | 3 |
| 2026 | X2Edit: Revisiting Arbitrary-Instruction Image Editing Through Self-Constructed Data and Task-Aware Representation LearningabstractExisting open-source datasets for arbitrary-instruction image editing remain suboptimal, while a plug-and-play editing module compatible with community-prevalent generative models is notably absent. In this paper, we first introduce the X2Edit Dataset, a comprehensive dataset covering 14 diverse editing tasks, including subject-driven generation. We utilize the industry-leading unified image generation models and expert models to construct the data. Meanwhile, we design reasonable editing instructions with the VLM and implement various scoring mechanisms to filter the data. As a result, we construct 3.7 million high-quality data with balanced categories. Second, to better integrate seamlessly with community image generation models, we design task-aware MoE-LoRA training based on FLUX.1, with only 8% of the parameters of the full model. To further improve the final performance, we utilize the internal representations of the diffusion model and define positive/negative samples based on image editing types to introduce contrastive learning. Extensive experiments demonstrate that the model's editing performance is competitive among many excellent models. Additionally, the constructed dataset exhibits substantial advantages over existing open-source datasets. Jian Ma 0010, Xujie Zhu, Qirong Peng, Chen Chen 0015, Haonan Lu |
AAAI | 7 |
| 2026 | Aligning Cross-View Visual Geometries in LVLMs Through Human-Like Reasoning LearningabstractSpatial understanding is a critical capability for LVLMs (Large Vision-Language Models) to advance embodied AI applications. Existing works primarily focus on enhancing spatial understanding within a single frame, i.e., injecting 3D spatial concepts into LVLMs under single coordinate system. However, such improvements struggle in real-world tasks that require consistent cross-view spatial reasoning. In this paper, we propose CVVG-Reasoner(Cross-View Visual Geometries) that lifts single-frame spatial comprehension to unified cross-view spatial understanding by mimicking human-like cross-view reasoning mechanisms. First, we introduce MV3DSR(Multi-View 3D Spatial Reasoning), a scalable pipeline for cross-view spatial reasoning data generation, and construct MV3DSR-Dataset, a large-scale dataset with diverse 3D cross-view reasoning tasks. Based on MV3DSR, we propose MV3DSR-Bench, a comprehensive benchmark for evaluating cross-view spatial reasoning capabilities. Second, we design a three-stage training strategy: the first two stages progressively equip the model with (1) fundamental spatial knowledge and (2) human-like cross-view reasoning patterns, while the final stage employs reinforcement learning to further boost its performance. Extensive experiments demonstrate that our CVVG-Reasoner significantly outperforms existing 3D LLMs(Large Language Models) and advanced LVLMs in cross-view tasks while maintaining robust performance on out-of-domain data. Ablations further reveal that injecting human-like reasoning patterns yields 44% performance gain, validating the effectiveness of our design. Yuming Qiao, Dan Meng 0001, Juntuo Wang, Ru Zhen, Haonan Lu |
AAAI | 10 |
| 2026 | OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence RewardabstractVideo captioning aims to generate comprehensive and coherent descriptions of the video content, contributing to the advancement of both video understanding and generation. However, existing methods often suffer from motion-detail imbalance, as models tend to overemphasize one aspect while neglecting the other. This imbalance results in incomplete captions, which in turn leads to a lack of consistency in video understanding and generation. To address this issue, we propose solutions from two aspects: 1) Data aspect: We constructed the Harmonizing Motion-Detail 270K (HMD-270K) dataset through a two-stage pipeline: Motion-Detail Fusion (MDF) and Fine-Grained Examination (FGE). 2) Optimization aspect: We introduce the Caption Set Equivalence Reward (CSER) based on Group Relative Policy Optimization (GRPO). CSER enhances completeness and accuracy in capturing both motion and details through unit-to-set matching and bidirectional validation. Based on the HMD-270K supervised fine-tuning and GRPO post-training with CSER, we developed OwlCap, a powerful video captioning Multi-modal Large Language Model (MLLM) with motion-detail balance. Experimental results demonstrate that OwlCap achieves significant improvements compared to baseline models on two benchmarks: the detail-focused VDC (+4.2 Acc) and the motion-focused DREAM-1K (+4.6 F1). Chunlin Zhong, Qiuxia Hou, Zhangjun Zhou, Shuang Hao 0015, Haonan Lu, He Tang 0002, Xiang Bai |
AAAI | 6 |
| 2026 | Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile AutomationabstractMobile agents show immense potential, yet current state-of-the-art (SoTA) agents exhibit inadequate success rates on real-world, long-horizon, cross-application tasks. We attribute this bottleneck to the agents' excessive reliance on static, internal knowledge within MLLMs, which leads to two critical failure points: 1) strategic hallucinations in high-level planning and 2) operational errors during low-level execution on user interfaces (UI). The core insight of this paper is that high-level planning and low-level UI operations require fundamentally distinct types of knowledge. Planning demands high-level, strategy-oriented experiences, whereas operations necessitate low-level, precise instructions closely tied to specific app UIs. Motivated by these insights, we propose Mobile-Agent-RAG, a novel hierarchical multi-agent framework that innovatively integrates dual-level retrieval augmentation. At the planning stage, we introduce Manager-RAG to reduce strategic hallucinations by retrieving human-validated comprehensive task plans that provide high-level guidance. At the execution stage, we develop Operator-RAG to improve execution accuracy by retrieving the most precise low-level guidance for accurate atomic actions, aligned with the current app and subtask. To accurately deliver these knowledge types, we construct two specialized retrieval-oriented knowledge bases. Furthermore, we introduce Mobile-Eval-RAG, a challenging benchmark for evaluating such agents on realistic multi-app, long-horizon tasks. Extensive experiments demonstrate that Mobile-Agent-RAG significantly outperforms SoTA baselines, improving task completion rate by 11.0% and step efficiency by 10.2%, establishing a robust paradigm for context-aware, reliable multi-agent mobile automation. Jichang Li, Haonan Lu, Guanbin Li |
AAAI | 4 |
| 2026 | CMPF: Harmonizing Cross-Model Prior Fusion for Open-Vocabulary Segmentation
Sicheng Zhao, Xi Chen 0110, Hongxun Yao, Haosen Yang 0003, Yanhao Zhang 0001, Sheng Jin 0002, Xiatian Zhu, Haonan Lu, Kui Jiang, Guiguang Ding |
Int. J. Comput. Vis. | 8 |
| 2026 | LLMI3D: MLLM-Based 3D Perception From a Single 2D ImageabstractRecent advancements in autonomous driving, augmented reality, robotics, and embodied intelligence have necessitated 3D perception algorithms. However, current 3D perception methods, especially specialized small models, exhibit poor generalization in open scenarios. On the other hand, multimodal large language models (MLLMs) excel in general capacity but underperform in 3D tasks, due to weak 3D local spatial object perception, poor text-based geometric numerical output, and inability to handle camera focal variations. To address these challenges, we develop LLMI3D, and propose the following solutions: Spatial-Enhanced Local Feature Mining for better 3D spatial feature extraction, 3D Query Token-Derived Info Decoding for precise geometric regression, and Geometry Projection-Based 3D Reasoning for handling camera focal length variations. We are the first to adapt an MLLM for image-based 3D perception. Additionally, we have constructed the IG3D dataset, which provides fine-grained descriptions and question-answer annotations. Extensive experiments demonstrate that our LLMI3D achieves state-of-the-art performance, outperforming other methods by a large margin. We will publicly release our code, models, and dataset. Fan Yang 0083, Sicheng Zhao, Yanhao Zhang 0001, Hui Chen 0013, Haonan Lu, Jungong Han, Guiguang Ding |
IEEE Trans. Multim. | 5 |
| 2025 | SCott: Accelerating Diffusion Models with Stochastic Consistency DistillationabstractThe iterative sampling procedure employed by diffusion models (DMs) often leads to significant latency. To address this, we propose Stochastic Consistency Distillation (SCott) to enable accelerated text-to-image generation, where high-quality generations can be achieved with just 2-4 sampling steps or even1 step, and further improvements can be obtained by additional cost, e.g., 4 steps. In contrast to vanilla consistency distillation (CD) which distills the ordinary differential equation solvers-based sampling process of a pre-trained teacher model into a student, SCott explores the possibility and validates the efficacy of integrating stochastic differential equation (SDE) solvers into CD to fully unleash the potential of the teacher. SCott is augmented with elaborate strategies to control the noise strength and sampling process of the SDE solver. An adversarial loss is further incorporated to strengthen the sample quality with rare sampling steps. Empirically, on the MSCOCO-2017 5K dataset with a Stable Diffusion-V1.5 teacher, SCott achieves an FID of 21.9, surpassing that of the 1-step InstaFlow (23.4) and the 4-step UFOGen (22.1). Moreover, SCott can yield more diverse samples than other consistency models for high-resolution image generation, with up to 16% improvement in a qualified metric. Hongjian Liu, Qingsong Xie, Tianxiang Ye, Zhijie Deng, Chen Chen 0015, Shixiang Tang, Xueyang Fu, Haonan Lu, Zhengjun Zha |
AAAI | 8 |
| 2025 | GlyphDraw2: Automatic Generation of Complex Glyph Posters with Diffusion Models and Large Language ModelsabstractPosters serve an essential function in marketing and advertising by improving visual communication and brand visibility, thus significantly contributing to industrial design. With the latest developments in controllable T2I diffusion models, research interest has surged in text rendering within synthesized images. Although text rendering accuracy has seen advancements, automatic poster generation remains a relatively untapped area. This paper presents an automatic poster generation framework featuring text rendering capabilities through the use of LLMs. Our framework employs a triple-cross attention mechanism based on alignment learning to achieve precise text placement within detailed contextual backgrounds. Moreover, it supports adjustable fonts, varying image resolutions, and poster rendering with textual prompts in both English and Chinese. Additionally, we present a comprehensive bilingual image-text dataset, GlyphDraw-3M, comprising 3 million image-text pairs, each with OCR annotations and resolutions exceeding 1024. Our method utilizes the SDXL architecture, and extensive experiments confirm its ability to generate posters with intricate and context-rich backgrounds. Jian Ma 0010, Yonglin Deng, Chen Chen 0015, Nanyang Du, Haonan Lu |
AAAI | 5 |
| 2025 | MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and QuantizationabstractMasked Image Modeling (MIM) with Vector Quantization (VQ) has achieved great success in both self-supervised pre-training and image generation. However, most existing methods struggle to address the trade-off in shared latent space for generation quality vs. representation learning and efficiency. To push the limits of this paradigm, we propose MergeVQ, which incorporates token merging techniques into VQ-based generative models to bridge the gap between image generation and visual representation learning in a unified architecture. During pre-training, MergeVQ decouples top-k semantics from latent space with the token merge module after self-attention blocks in the encoder for subsequent Look-up Free Quantization (LFQ) and global alignment and recovers their fine-grained details through cross-attention in the decoder for reconstruction. As for second-stage generation, we introduce MergeAR, which performs KV Cache compression for efficient raster-order prediction. Extensive experiments on ImageNet verify that MergeVQ as an AR generative model achieves competitive performance in both visual representation learning and image generation tasks while maintaining favorable token efficiency and inference speed. Code and model will be available at https://apexgen-x.github.io/MergeVQ. Siyuan Li 0002, Luyuan Zhang, Zedong Wang, Juanxi Tian, Cheng Tan 0012, Zicheng Liu 0006, Chang Yu 0001, Qingsong Xie, Haonan Lu, Haoqian Wang, Zhen Lei 0001 |
CVPR | 9 |
| 2025 | HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility EvaluatorabstractAIGC images are prevalent across various fields, yet they frequently suffer from quality issues like artifacts and unnatural textures. Specialized models aim to predict defect region heatmaps but face two primary challenges: (1) lack of explainability, failing to provide reasons and analyses for subtle defects, and (2) inability to leverage common sense and logical reasoning, leading to poor generalization. Multimodal large language models (MLLMs) promise better comprehension and reasoning but face their own challenges: (1) difficulty in fine-grained defect localization due to the limitations in capturing tiny details, and (2) constraints in providing pixel-wise outputs necessary for precise heatmap generation. To address these challenges, we propose HEIE: a novel MLLM-Based Hierarchical Explainable Image Implausibility Evaluator. We introduce the CoT-Driven Explainable Trinity Evaluator, which integrates heatmaps, scores, and explanation outputs, using CoT to decompose complex tasks into subtasks of increasing difficulty and enhance interpretability. Our Adaptive Hierarchical Implausibility Mapper synergizes low-level image features with high-level mapper tokens from LLMs, enabling precise local-to-global hierarchical heatmap predictions through an uncertainty-based adaptive token approach. Moreover, we propose a new dataset: Expl-AIGI-Eval, designed to facilitate interpretable implausibility evaluation of AIGC images. Our method demonstrates state-of-the-art performance through extensive experiments. Our project is at https://yfthu.github.io/HEIE/. Fan Yang 0083, Ru Zhen, Yanhao Zhang 0001, Haoxiang Chen 0007, Haonan Lu, Sicheng Zhao, Guiguang Ding |
CVPR | 6 |
| 2025 | Advancing Text-to-3D Generation with Linearized Lookahead Variational Score DistillationabstractText-to-3D generation based on score distillation of pre-trained 2D diffusion models has gained increasing interest, with variational score distillation (VSD) as a remarkable example. VSD proves that vanilla score distillation can be improved by introducing an extra score-based model, which characterizes the distribution of images rendered from 3D models, to correct the distillation gradient. Despite the theoretical foundations, VSD, in practice, is likely to suffer from slow and sometimes ill-posed convergence. In this paper, we perform an in-depth investigation of the interplay between the introduced score model and the 3D model, and find that there exists a mismatching problem between LoRA and 3D distributions in practical implementation. We can simply adjust their optimization order to improve the generation quality. By doing so, the score model looks ahead to the current 3D state and hence yields more reasonable corrections. Nevertheless, naive lookahead VSD may suffer from unstable training in practice due to the potential over-fitting. To address this, we propose to use a linearized variant of the model for score distillation, giving rise to the Linearized Lookahead Variational Score Distillation ($L^2$-VSD). $L^2$-VSD can be realized efficiently with forward-mode autodiff functionalities of existing deep learning libraries. Extensive experiments validate the efficacy of $L^2$-VSD, revealing its clear superiority over prior score distillation-based methods. We also show that our method can be seamlessly incorporated into any other VSD-based text-to-3D framework. Bingde Liu, Qingsong Xie, Haonan Lu, Zhijie Deng |
ICCV | 4 |
| 2025 | X2i: Seamless Integration of Multimodal Understanding Into Diffusion Transformer Via Attention Distillation
Jian Ma 0010, Qirong Peng, Chen Chen 0015, Haonan Lu |
ICCV | 5 |
| 2025 | Free-Moref: Instantly Multiplexing Context Perception Capabilities of Video-Mllms Within Single InferenceabstractVideo Multimodal Large Language Models~(Video-MLLM) have achieved remarkable advancements in video understanding tasks. However, constrained by the context length limitation in the underlying LLMs, existing Video-MLLMs typically exhibit suboptimal performance on long video scenarios. To understand extended input frames, common solutions span token compression and streaming inference techniques, which sacrifice feature granularity or inference efficiency. Differently, to efficiently achieve comprehensive understanding of longer frame inputs, we draw ideas from MoE and propose a training-free approach \textbf{Free-MoRef}, which instantly multiplexes the context perception capabilities of Video-MLLMs within one inference pass. Specifically, Free-MoRef reconstructs the vision tokens into several short sequences as multi-references. Subsequently, we introduce MoRef-attention, which gathers clues from the multi-reference chunks in parallel to summarize unified query activations. After the shadow layers in LLMs, a reference fusion step is derived to compose a final mixed reasoning sequence with key tokens from parallel chunks, which compensates the cross-reference vision interactions that are neglected in MoRef-attention. By splitting and fusing the long vision token sequences, Free-MoRef achieves improved performance under much lower computing costs in reasoning multiplexed context length, demonstrating strong efficiency and effectiveness. Experiments on VideoMME, MLVU, LongVideoBench show that Free-MoRef achieves full perception of 2$\times$ to 8$\times$ longer input frames without compression on a single A100 GPU while keeping instant responses, thereby bringing significant performance gains, even surpassing dedicatedly trained long-video-MLLMs. Codes are available at https://github.com/wkfdb/Free-MoRef Quanlong Zheng, Junlin Xie, Jinguo Luo, Haonan Lu, Liang Lin 0004, Guanbin Li |
ICCV | 6 |
| 2025 | MsRAG: Knowledge Augumented Image Captioning with Object-level Multi-source RAGabstractLanguage-Visual Large Models (LVLMs) have made significant strides in enhancing visual understanding capabilities. However, these models often struggle with knowledge-based visual tasks due to constrains in their pre-training data scope and timeliness. Existing Retrieval-Augmented Generation (RAG) methods can effectively solve the problem but primarily rely on user queries, limiting their applicability in scenarios without explicit language input. To overcome these challenges, we introduce MsRAG, a knowledge-augmented captioning framework designed to effectively retrieve and utilize external real-world knowledge, particularly in the absence of user queries, and perform dense captioning for subjects. MsRAG comprises three key components: (1) Parallel Visual Search Module. It retrieves fine-grained object-level knowledge using both online visual search engines and offline domain-knowledge databases, enhancing the robustness and richness of retrieved information. (2) Prompt Templates Pool. The prompt pool dynamically assigns appropriate prompts based on retrieved information, optimizing LVLMs' ability to leverage relevant data under complex RAG conditions. (3) Visual-RAG Alignment Module, which employs a novel visual prompting method to bridge the modality gap between textual RAG content and corresponding visual objects, enabling precise alignment of visual elements with their text-format RAG content. To validate the effectiveness of MsRAG, we conducted a series of qualitative and quantitative experiments. The evaluation results demonstrate the superiority of MsRAG over other methods. Yuming Qiao, Yuechen Wang, Dan Meng 0001, Haonan Lu |
IJCAI | 4 |
| 2025 | InstructHOI: Context-Aware Instruction for Multi-Modal Reasoning in Human-Object Interaction DetectionabstractRecently, Large Foundation Models (LFMs), e.g., CLIP and GPT, have significantly advanced the Human-Object Interaction (HOI) detection, due to their superior generalization and transferability. Prior HOI detectors typically employ single- or multi-modal prompts to generate discriminative representations for HOIs from pretrained LFMs. However, such prompt-based approaches focus on transferring HOI-specific knowledge, but unexplore the potential reasoning capabilities of LFMs, which can provide informative context for ambiguous and open-world interaction recognition. In this paper, we propose InstructHOI, a novel method that leverages context-aware instructions to guide multi-modal reasoning for HOI detection. Specifically, to bridge knowledge gap and enhance reasoning abilities, we first perform HOI-domain fine-tuning on a pretrained multi-modal LFM, using a generated dataset with 140K interaction-reasoning image-text pairs. Then, we develop a Context-aware Instruction Generator (CIG) to guide interaction reasoning. Unlike traditional language-only instructions, CIG first mines visual interactive context at the human-object level, which is then fused with linguistic instructions, forming multi-modal reasoning guidance. Furthermore, an Interest Token Selector (ITS) is adopted to adaptively filter image tokens based on context-aware instructions, thereby aligning reasoning process with interaction regions. Extensive experiments on two public benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones, under both supervised and zero-shot settings. Jinguo Luo, Weihong Ren, Quanlong Zheng, Zhenlong Yuan, Zhiyong Wang 0009, Haonan Lu, Honghai Liu 0001 |
NeurIPS | 7 |
| 2024 | Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion ModelsabstractRecent text-to-image (T2I) diffusion models show outstanding performance in generating high-quality images conditioned on textual prompts. However, they fail to semantically align the generated images with the prompts due to their limited compositional capabilities, leading to attribute leakage, entity leakage, and missing entities. In this paper, we propose a novel attention mask control strategy based on predicted object boxes to address these issues. In particular, we first train a BoxNet to predict a box for each entity that possesses the attribute specified in the prompt. Then, depending on the predicted boxes, a unique mask control is applied to the cross- and self-attention maps. Our approach produces a more semantically accurate synthesis by constraining the attention regions of each token in the prompt to the image. In addition, the proposed method is straightforward and effective and can be readily integrated into existing cross-attention-based T2I generators. We compare our approach to competing methods and demonstrate that it can faithfully convey the semantics of the original text to the generated content and achieve high availability as a ready-to-use plugin. Please refer to https://github.com/OPPO-Mente-Lab/attention-mask-control. Zekang Chen, Chen Chen 0015, Jian Ma 0010, Haonan Lu, Xiaodong Lin 0004 |
AAAI | 5 |
| 2024 | Probing Language Models for Pre-training Data DetectionabstractLarge Language Models (LLMs) have shown their impressive capabilities, while also raising concerns about the data contamination problems due to privacy issues and leakage of benchmark datasets in the pre-training phase.Therefore, it is vital to detect the contamination by checking whether an LLM has been pre-trained on the target texts.Recent studies focus on the generated texts and compute perplexities, which are superficial features and not reliable.In this study, we propose to utilize the probing technique for pre-training data detection by examining the model's internal activations.Our method is simple yet effective and leads to more trustworthy pre-training data detection.Additionally, we propose ArxivMIA, a new challenging benchmark comprising arxiv abstracts from Computer Science and Mathematics categories.Our experiments demonstrate that our method outperforms all the baselines, and achieves state-of-the-art performance on both WikiMIA and ArxivMIA, with additional experiments confirming its efficacy 1 . Tong Zhu 0002, Chuanyuan Tan, Haonan Lu, Wenliang Chen |
ACL (1) | 5 |
| 2024 | PEA-Diffusion: Parameter-Efficient Adapter with Knowledge Distillation in Non-english Text-to-Image Generation
Jian Ma 0010, Chen Chen 0015, Qingsong Xie, Haonan Lu |
ECCV (68) | 4 |
| 2024 | InsCL: A Data-efficient Continual Learning Paradigm for Fine-tuning Large Language Models with InstructionsabstractYifan Wang, Yafei Liu, Chufan Shi, Haoling Li, Chen Chen, Haonan Lu, Yujiu Yang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chufan Shi, Haoling Li, Chen Chen 0015, Haonan Lu, Yujiu Yang 0001 |
NAACL-HLT | 6 |
| 2024 | Accelerating Skewed Workloads With Performance Multipliers in the TurboDB Distributed Database
Jennifer Lam, Jeffrey Helt, Wyatt Lloyd, Haonan Lu |
NSDI | 4 |
| 2024 | Dream360: Diverse and Immersive Outdoor Virtual Scene Creation via Transformer-Based 360° Image Outpaintingabstract360° images, with a field-of-view (FoV) of $180^{\circ}\times 360^{\circ}$, provide immersive and realistic environments for emerging virtual reality (VR) applications, such as virtual tourism, where users desire to create diverse panoramic scenes from a narrow FoV photo they take from a viewpoint via portable devices. It thus brings us to a technical challenge: 'How to allow the users to freely create diverse and immersive virtual scenes from a narrow FoV image with a specified viewport?' To this end, we propose a transformer-based 360° image outpainting framework called Dream360, which can generate diverse, high-fidelity, and high-resolution panoramas from user-selected viewports, considering the spherical properties of 360° images. Compared with existing methods, e.g., [3], which primarily focus on inputs with rectangular masks and central locations while overlooking the spherical property of 360° images, our Dream360 offers higher outpainting flexibility and fidelity based on the spherical representation. Dream360 comprises two key learning stages: (I) codebook-based panorama outpainting via Spherical-VQGAN (S-VQGAN), and (II) frequency-aware refinement with a novel frequency-aware consistency loss. Specifically, S-VQGAN learns a sphere-specific codebook from spherical harmonic (SH) values, providing a better representation of spherical data distribution for scene modeling. The frequency-aware refinement matches the resolution and further improves the semantic consistency and visual fidelity of the generated results. Our Dream360 achieves significantly lower Frechet Inception Distance (FID) scores and better visual fidelity than existing methods. We also conducted a user study involving 15 participants to interactively evaluate the quality of the generated results in VR, demonstrating the flexibility and superiority of our Dream360 framework. Hao Ai, Zidong Cao, Haonan Lu, Chen Chen 0015, Jian Ma 0010, Peng Yuan Zhou, Tae-Kyun Kim 0001, Pan Hui 0001, Lin Wang 0025 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | Interactive Interior Design Recommendation via Coarse-to-fine Multimodal Reinforcement LearningabstractPersonalized interior decoration design often incurs high labor costs. Recent efforts in developing intelligent interior design systems have focused on generating textual requirement-based decoration designs while neglecting the problem of how to mine homeowner's hidden preferences and choose the proper initial design. To fill this gap, we propose an Interactive Interior Design Recommendation System (IIDRS) based on reinforcement learning (RL). IIDRS aims to find an ideal plan by interacting with the user, who provides feedback on the gap between the recommended plan and their ideal one. To improve decision-making efficiency and effectiveness in large decoration spaces, we propose a Decoration Recommendation Coarse-to-Fine Policy Network (DecorRCFN). Additionally, to enhance generalization in online scenarios, we propose an object-aware feedback generation method that augments model training with diversified and dynamic textual feedback. Extensive experiments on a real-world dataset demonstrate our method outperforms traditional methods by a large margin in terms of recommendation accuracy. Further user studies demonstrate that our method reaches higher real-world user satisfaction than baseline methods. He Zhang 0030, Ying Sun 0006, Weiyu Guo, Haonan Lu, Xiaodong Lin 0004, Hui Xiong 0001 |
ACM Multimedia | 5 |
| 2023 | NCC: Natural Concurrency Control for Strictly Serializable Datastores by Avoiding the Timestamp-Inversion Pitfall
Haonan Lu, Shuai Mu 0001, Siddhartha Sen 0001, Wyatt Lloyd |
OSDI | 1 |
| 2022 | GammaE: Gamma Embeddings for Logical Queries on Knowledge GraphsabstractEmbedding knowledge graphs (KGs) for multihop logical reasoning is a challenging problem due to massive and complicated structures in many KGs.Recently, many promising works projected entities and queries into a geometric space to efficiently find answers.However, it remains challenging to model the negation and union operator.The negation operator has no strict boundaries, which generates overlapped embeddings and leads to obtaining ambiguous answers.An additional limitation is that the union operator is non-closure, which undermines the model to handle a series of union operators.To address these problems, we propose a novel probabilistic embedding model, namely Gamma Embeddings (GammaE), for encoding entities and queries to answer different types of FOL queries on KGs.We utilize the linear property and strong boundary support of the Gamma distribution to capture more features of entities and queries, which dramatically reduces model uncertainty.Furthermore, Gam-maE implements the Gamma mixture method to design the closed union operator.The performance of GammaE is validated on three large logical query datasets.Experimental results show that GammaE significantly outperforms state-of-the-art models on public benchmarks. Peijun Qing, Haonan Lu, Xiaodong Lin 0004 |
EMNLP | 4 |
| 2022 | DensE: An enhanced non-commutative representation for knowledge graph embedding with adaptive semantic hierarchy
Haonan Lu, Hailin Hu 0002, Xiaodong Lin 0004 |
Neurocomputing | 1 |
| 2021 | K2: Reading Quickly from Storage Across Many DatacentersabstractThe infrastructure available to large-scale and medium-scale web services now spans dozens of geographically dispersed datacenters. Deploying across many datacenters has the potential to significantly reduce end-user latency by serving users nearer their location. However, deploying across many datacenters requires the backend storage system be partially replicated. In turn, this can sacrifice the low latency benefits of many datacenters, especially when a storage system provides guarantees on what operations will observe. We present the K2 storage system that provides lower latency for large-scale and medium-scale web services using partial replication of data over many datacenters with strong guarantees: causal consistency, read-only transactions, and write-only transactions. K2 provides the best possible worst-case latency for partial replication, a single round trip to remote datacenters, and often avoids sending any requests to far away datacenters using a novel replication approach, write-only transaction algorithm, and read-only transaction algorithm. Khiem Ngo, Haonan Lu, Wyatt Lloyd |
DSN | 2 |
| 2021 | SNOW Revisited: Understanding When Ideal READ Transactions Are PossibleabstractREAD transactions that read data distributed across servers dominate the workloads of real-world distributed storage systems. The SNOW Theorem [13] stated that ideal READ transactions that have optimal latency and the strongest guarantees-i.e., “SNOW” READ transactions-are impossible in one specific setting that requires three or more clients: at least two readers and one writer. However, it left many open questions. We close all of these open questions with new impossibility results and new algorithms. First, we prove rigorously the result from [13] saying that it is impossible to have a READ transactions system that satisfies SNOW properties with three or more clients. The insight we gained from this proof led to teasing out the implicit assumptions that are required to state the results and also, resolving the open question regarding the possibility of SNOW with two clients. We show that it is possible to design an algorithm, where SNOW is possible in a multi-writer, single-reader (MWSR) setting when a client can send messages to other clients; on the other hand, we prove it is impossible to implement SNOW in a multi-writer, single-reader (MWSR) setting-which is more general than the two-client setting-when client-to-client communication is disallowed. We also correct the previous claim in [13] that incorrectly identified one existing system, Eiger [12], as supporting the strongest guarantees (SW) and whose read-only transactions had bounded latency. Thus, there were no previous algorithms that provided the strongest guarantees and had bounded latency. Finally, we introduce the first two algorithms to provide the strongest guarantees with bounded latency. Kishori M. Konwar, Wyatt Lloyd, Haonan Lu, Nancy A. Lynch |
IPDPS | 3 |
| 2020 | Performance-Optimal Read-Only Transactions
Haonan Lu, Siddhartha Sen 0001, Wyatt Lloyd |
OSDI | 1 |
| 2018 | Locora: Practical Wi-Fi packet loss diagnosis with only local observationsabstractIn a Wi-Fi network, packet loss may occur due to bad channel, collision, or hidden terminal. The Wi-Fi protocol however cannot identify the causes of the loss, and may lead the nodes to take unnecessary or even incorrect actions. In this paper, we propose Loss Cause Oracle (Locora), which can diagnose the actual cause of the packet loss and make correct responses with only local observations. Locora is a comprehensive set of solutions that can estimate the collision probability, detect various kinds of hidden terminals, and interact with the rate selection algorithm and assist it to converge to the best data rate. Locora advances Wi-Fi packet loss diagnosis because it can handle all three types of packet loss while modifying only the packet sender and does not depend on any assumptions on the user traffic. We implement Locora in assembly language in the firmware of the BCM4306 wireless card on the OpenFWWF platform and our experiments show that Lococa achieves significant gains over existing schemes. Shuaiyuan Zhou, Haonan Lu, Avishek Mukherjee |
WCNC | 2 |
| 2017 | The record route option is an option!abstractThe IPv4 Record Route (RR) Option instructs routers to record their IP addresses in a packet. RR is subject to a nine hop limit and, traditionally, inconsistent support from routers. Recent changes in interdomain connectivity---the so-called "flattening Internet"---and new best practices for how routers should handle RR packets suggest that now is a good time to reassess the potential of the RR Option. Brian J. Goodchild, Yi-Ching Chiu, Rob Hansen, Haonan Lu, Matt Calder, Matthew J. Luckie, Wyatt Lloyd, David R. Choffnes, Ethan Katz-Bassett |
Internet Measurement Conference | 4 |
| 2016 | The SNOW Theorem and Latency-Optimal Read-Only Transactions
Haonan Lu, Christopher Hodsdon, Khiem Ngo, Shuai Mu 0001, Wyatt Lloyd |
OSDI | 1 |
| 2015 | Existential consistency: measuring and understanding consistency at FacebookabstractReplicated storage for large Web services faces a trade-off between stronger forms of consistency and higher performance properties. Stronger consistency prevents anomalies, i.e., unexpected behavior visible to users, and reduces programming complexity. There is much recent work on improving the performance properties of systems with stronger consistency, yet the flip-side of this trade-off remains elusively hard to quantify. To the best of our knowledge, no prior work does so for a large, production Web service. Haonan Lu, Kaushik Veeraraghavan, Philippe Ajoux, Jim Hunt, Yee Jiun Song, Wendy Tobagus, Wyatt Lloyd |
SOSP | 1 |
| 2012 | Retransmission rate selection for block-based partial packet recoveryabstractIn a Wi-Fi network, partial packets often exist which are packets with some errors. Recent works show that much better efficiency can be achieved by repairing such packets instead of retransmitting such packets. In this paper, we propose a rate selection scheme for the repair packets, refereed to as ReSel. ReSel is developed under the Maranello partial packet recovery framework and is motivated by two observations. First, the repair packet is sent shortly after the original packet and will likely experience similar channel conditions as the original packet. Therefore, given the recent failure on the original packet, a lower data rate should be used for the repair packet. Second, a Maranello receiver will send a NACK packet to the sender, which can be used to carry important information for the selection of the data rate. We build a rate selection table for ReSel with experimental data. We implement ReSel in the firmware and our experiments show that ReSel significantly improves the link throughput and reduces packet jitter. We also find that the repair packet is transmitted only once in the majority of the cases and hence ReSel is capable of selecting the appropriate data rate. Haonan Lu, Shuaiyuan Zhou |
GLOBECOM | 1 |