Guoyuan An

dblp:299/8567 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0008-6233-757XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PFAvatar: Pose-Fusion 3D Personalized Avatar Reconstruction from Real-World Outfit-of-the-Day Photos
abstract
We propose PFAvatar (Pose-Fusion Avatar), a new method that reconstructs high-quality 3D avatars from Outfit of the Day (OOTD) photos, which exhibit diverse poses, occlusions, and complex backgrounds. Our method consists of two stages: (1) fine-tuning a pose-aware diffusion model from few-shot OOTD examples and (2) distilling a 3D avatar represented by a neural radiance field (NeRF). In the first stage, unlike previous methods that segment images into assets (e.g. garments, accessories) for 3D assembly, which is prone to inconsistency, we avoid decomposition and directly model the full-body appearance. By integrating a pre-trained ControlNet for pose estimation and a novel Condition Prior Preservation Loss (CPPL), our method enables end-to-end learning of fine details while mitigating language drift in few-shot training. Our method completes personalization in just 5 minutes, achieving a 48x speed-up compared to previous approaches. In the second stage, we introduce a NeRF-based avatar representation optimized by canonical SMPL-X space sampling and Multi-Resolution 3D-SDS. Compared to mesh-based representations that suffer from resolution-dependent discretization and erroneous occluded geometry, our continuous radiance field can preserve high-frequency textures (e.g., hair) and handle occlusions correctly through transmittance. Experiments demonstrate that PFAvatar outperforms state-of-the-art methods in terms of reconstruction fidelity, detail preservation, and robustness to occlusions/truncations, advancing practical 3D avatar generation from real-world OOTD albums. In addition, the reconstructed 3D avatars support downstream applications such as virtual try-on, animation, and human video reenactment, further demonstrating the versatility and practical value of our approach.
Dianbing Xi, Guoyuan An, Jingsen Zhu, Ruiyuan Zhang, Jiayuan Lu, Yuchi Huo, Rui Wang 0004
AAAI2
2026 Full Retraining, Incremental Fine-tuning, and Hybrid Serving: Model Updating and Serving for Industrial Generative Recommender Systems
abstract
Generative recommendation casts recommendation as conditional sequence generation over text- or token-based representations and has shown strong promise in industrial systems. However, keeping such models up to date in dynamic environments is difficult: full retraining on sliding windows is expensive and slow, while incremental fine-tuning on recent data may introduce distributional bias and catastrophic forgetting.
Qijiong Liu, Zhongzhou Liu, Guoyuan An, Wei Guo 0006, Yong Liu 0020, Xiao-Ming Wu 0003
SIGIR4
2026 Accelerating Generative Recommendation via Simple Categorical User Sequence Compression
abstract
Although generative recommenders demonstrate improved performance with longer sequences, their real-time deployment is hindered by substantial computational costs. To address this challenge, we propose a simple yet effective method for compressing long-term user histories by leveraging inherent item categorical features, thereby preserving user interests while enhancing efficiency. Experiments on two large-scale datasets demonstrate that, compared to the influential HSTU model, our approach achieves up to a 6× reduction in computational cost and up to 39% higher accuracy at comparable cost (i.e., similar sequence length). The source code will be available at https://github.com/Genemmender/CAUSE.
Qijiong Liu, Zhongzhou Liu, Yuankai Luo, Guoyuan An, Nuo Chen 0004, Wei Guo 0006, Yong Liu 0020, Xiao-Ming Wu 0003
WSDM6
2025 Towards Robustness of Person Search Against Corruptions
Woojung Son, Yoonki Cho, Guoyuan An, Chanmi Lee, Sung-Eui Yoon
ICCV3
2025 AniTex: Light-Geometry Consistent PBR Material Generation for Animatable Objects
abstract
High-quality Physically-Based Rendering (PBR) materials are crucial for visual realism in 3D asset creation, yet existing methods primarily target static objects, leading to challenges in maintaining multi-frame consistency for animatable entities. To tackle this issue, we introduce AniTex, the first generative pipeline that utilizes diffusion models to synthesize high-quality PBR materials for animatable objects based on text prompts. The pipeline consists of three key stages: First, sequences of RGB images are generated using a video diffusion model conditioned on depth, normals, irradiance, and motion vectors to ensure temporal coherence and geometric alignment across multiple frames and viewpoints. Second, these RGB image sequences are decomposed into per-view, per-frame PBR material maps (albedo, roughness, metallic) by a specialized Intrinsic Diffusion Model (IDM), which is conditioned on the RGB images along with consistent geometry and lighting cues to disentangle material from illumination. Finally, these per-view, per-frame PBR maps are hierarchically blended. This process first ensures temporal coherence within each view’s frame sequence, then amalgamates these into globally consistent PBR materials for the animatable object, maintaining overall temporal coherence and visual consistency throughout its animation. Extensive experiments show that AniTex produces more realistic PBR materials for both static and animated objects, outperforming baseline methods in visual appeal.
Jieting Xu, Guoyuan An, Rengan Xie, Dianbing Xi, Wenjun Song, Rui Wang 0004, Yuchi Huo
SIGGRAPH Asia3
2025 Ultra-High Resolution Facial Texture Reconstruction from a Single Image
abstract
Advances in mobile cameras have made it easier to capture ultra-high resolution (UHR) portraits. However, existing face reconstruction methods lack specific adaptations for UHR input (e.g., 4096 × 4096), leading to under-use of high-frequency details that are crucial for achieving photorealistic rendering. Our method supports 4096 × 4096 UHR input and utilizes a divide-and-conquer approach for end-to-end 4K albedo, micronormal, and specular texture reconstruction at the original resolution. We employ a two-stage strategy to capture both global distributions and local high-frequency details, effectively mitigating mosaic and seam artifacts common in patch-based prediction. Additionally, we innovatively apply hash encoding to facial U-V coordinates to boost the model’s ability to learn regional high-frequency feature distributions. Our method can be easily incorporated in state-of-the-art facial geometry reconstruction pipelines, significantly improving the texture reconstruction quality, facilitating artistic creation workflows.
Hongxiang Huang, Guoyuan An, Jingzhen Lan, Qi Wang 0111, Rui Wang 0004, Yuchi Huo
Comput. Vis. Media2
2025 OpenSlot: Mixed Open-Set Recognition With Object-Centric Learning
abstract
Existing open-set recognition (OSR) studies typically assume that each image contains only one class label, with the unknown test set (negative) having a disjoint label space from the known test set (positive), a scenario referred to as full-label shift. This paper introduces the mixed OSR problem, where test images contain multiple class semantics, with both known and unknown classes co-occurring in the negatives, leading to a more complex super-label shift that better reflects real-world scenarios. To tackle this challenge, we propose the OpenSlot framework, based on object-centric learning, which uses slot features to represent diverse class semantics and generate class predictions. The proposed anti-noise slot (ANS) technique helps mitigate the impact of noise (invalid or background) slots during classification training, addressing the semantic misalignment between class predictions and ground truth. We evaluate OpenSlot on both mixed and conventional OSR benchmarks. Without elaborate designs, our method not only excels existing approaches in detecting super-label shifts across OSR tasks, but also achieves state-of-the-art performance on conventional benchmarks. Meanwhile, OpenSlot can localize class objects without using bounding boxes during training, demonstrating competitive performance in open-set object detection and potential for generalization.
Xu Yin, Guoyuan An, Yuchi Huo, Sung-Eui Yoon
IEEE Trans. Multim.3
2024 A novel semi-supervised model for pre-impact fall detection with limited fall data
Xiaoqun Yu, Jiansong Wan, Guoyuan An, Xu Yin, Shuping Xiong
Eng. Appl. Artif. Intell.3
2023 Towards Content-based Pixel Retrieval in Revisited Oxford and Paris
abstract
This paper introduces the first two landmark pixel retrieval benchmarks. Pixel retrieval is segmented instance retrieval. Like semantic segmentation extends classification to the pixel level, pixel retrieval is an extension of image retrieval and offers information about which pixels are related to the query object. In addition to retrieving images for the given query, it helps users quickly identify the query object in true positive images and exclude false positive images by denoting the correlated pixels. Our user study results show pixel-level annotation can significantly improve the user experience. Compared with semantic and instance segmentation, pixel retrieval requires a fine-grained recognition capability for variable-granularity targets. To this end, we propose pixel retrieval benchmarks named PROxford and PRParis, which are based on the widely used image retrieval datasets, ROxford and RParis. Three professional annotators label 5,942 images with two rounds of double-checking and refinement. Furthermore, we conduct extensive experiments and analysis on the SOTA methods in image search, image matching, detection, segmentation, and dense matching using our pixel retrieval benchmarks. Results show that the pixel retrieval task is challenging to these approaches and distinctive from existing problems, suggesting that further research can advance the content-based pixel-retrieval and thus user search experience. The datasets can be downloaded from this link.
Guoyuan An, Woo Jae Kim, Saelyne Yang, Yuchi Huo, Sung-Eui Yoon
ICCV1
2023 Topological RANSAC for instance verification and retrieval without fine-tuning
abstract
This paper presents an innovative approach to enhancing explainable image retrieval, particularly in situations where a fine-tuning set is unavailable. The widely-used SPatial verification (SP) method, despite its efficacy, relies on a spatial model and the hypothesis-testing strategy for instance recognition, leading to inherent limitations, including the assumption of planar structures and neglect of topological relations among features. To address these shortcomings, we introduce a pioneering technique that replaces the spatial model with a topological one within the RANSAC process. We propose bio-inspired saccade and fovea functions to verify the topological consistency among features, effectively circumventing the issues associated with SP's spatial model. Our experimental results demonstrate that our method significantly outperforms SP, achieving state-of-the-art performance in non-fine-tuning retrieval. Furthermore, our approach can enhance performance when used in conjunction with fine-tuned features. Importantly, our method retains high explainability and is lightweight, offering a practical and adaptable solution for a variety of real-world applications.
Guoyuan An, Juhyeong Seon, Inkyu An, Yuchi Huo, Sung-Eui Yoon
NeurIPS1
2021 Hypergraph Propagation and Community Selection for Objects Retrieval
abstract
Spatial verification is a crucial technique for particular object retrieval. It utilizes spatial information for the accurate detection of true positive images. However, existing query expansion and diffusion methods cannot efficiently propagate the spatial information in an ordinary graph with scalar edge weights, resulting in low recall or precision. To tackle these problems, we propose a novel hypergraph-based framework that efficiently propagates spatial information in query time and retrieves an object in the database accurately. Additionally, we propose using the image graph's structure information through community selection technique, to measure the accuracy of the initial search result and to provide correct starting points for hypergraph propagation without heavy spatial verification computations. Experiment results on ROxford and RParis show that our method significantly outperforms the existing query expansion and diffusion methods.
Guoyuan An, Yuchi Huo, Sung-Eui Yoon
NeurIPS1
2021 GraphShop: Graph-based Approach for Shop-type Recommendation
abstract
It is essential to predict the popularity of a particular shop type when investors decide which type of shops to open at a given location.Existing shop-type recommender systems have approached this problem by building a regiontype matrix and analyzing the relationship between different regions and shop types.However, these methods make recommendations for each region, thus having difficulty analyzing a specific shop, especially near the two regions' borders.To tackle this challenge, we propose a novel Graph Neural Network (GNN) model, called GraphShop, to represent shops as nodes in a graph and analyze each shop without assigning it to a region.As it is difficult to find the influential neighbors, we propose two aggregation methods, Distance-Module and TypeModule, in GraphShop.DistanceModule aggregates unordered nearby shops in every zone and filters them from the remote zones.TypeModule reorders the nearby shops based on their types and considers the interaction of different types.Furthermore, to address the lack of open shop-type recommendation datasets, we build a qualitative and large-scale dataset collected from a review website and location-based services.It contains most, if not all, shops in a region and is large and diverse, by containing 53,182 shops with 122 types.Our dataset is available at https://github.com/BoSamothrace/GraphShop.Through the experimental results, we demonstrate that our method outperforms the existing state-of-the-art methods for shoptype recommendation by a factor of up to 37 %.
Guoyuan An, Sung-Eui Yoon, Jae Yoon Kim, Myoung Ho Kim
SDM1