VLDB 2026 Research / reviewers in the wild / expert
Jiaxian Guo
dblp:206/6264
· DBLP profile ↗
19ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | One-shot Portrait Stylization via Geometric AlignmentabstractPortrait stylization casts vivid artistic style drawn from style examples to portrait photos. Although recently extensively studied with machine learning algorithms, existing methods still face challenges in stylizing portraits from a single style reference, severely limiting their potential for real-world applications. In this paper, we propose a portrait stylization method that learns style reference from a single artistic portrait image. Unlike previous StyleGAN based methods that heavily rely on the quality of GAN inversion or diffusion based methods that introduce computational expensive operations and fall short of precise control, our method achieves high-quality stylization with small computation and parameter budget. Specifically, we employ geometric alignment to build spatial correlation between content images and style reference. A geometry LoRA and a style LoRA are then jointly optimized based on a pre-trained diffusion backbone respectively, with orthogonal adaptation used to disentangle the geometry and style information. During inference, the style LoRA is integrated into the diffusion backbone and ControlNet is further combined to facilitate better spatial and identity control. We illustrate abundant stylized portraits with multiple styles. Qualitative comparison, quantitative validation and user study prove that our method outperforms existing methods, and ablation study demonstrates the effectiveness of each components. Zilin Guo, Zhuoru Li, Yusuke Iwasawa, Yutaka Matsuo, Jiaxian Guo |
WACV | 8 |
| 2025 | Image Referenced Sketch Colorization Based on Animation Creation WorkflowabstractSketch colorization plays an important role in animation and digital illustration production tasks. However, existing methods still meet problems in that text-guided methods fail to provide accurate color and style reference, hint-guided methods still involve manual operation, and image-referenced methods are prone to cause artifacts. To address these limitations, we propose a diffusion-based framework inspired by real-world animation production work-flows. Our approach leverages the sketch as the spatial guidance and an RGB image as the color reference, and separately extracts foreground and background from the reference image with spatial masks. Particularly, we introduce a split cross-attention mechanism with LoRA (Low-Rank Adaptation) modules. They are trained separately with foreground and background regions to control the corresponding embeddings for keys and values in cross-attention. This design allows the diffusion model to integrate information from foreground and background independently, preventing interference and eliminating the spatial artifacts. During inference, we design switchable inference modes for diverse use scenarios by changing modules activated in the framework. Extensive qualitative and quantitative experiments, along with user studies, demonstrate our advantages over existing methods in generating high-qualigy artifact-free results with geometric mismatched references. Ablation studies further confirm the effectiveness of each component. Codes are available at https://github.com/tellurion-kanata/colorizeDiffusion. Dingkun Yan, Zhuoru Li, Suguru Saito, Yusuke Iwasawa, Yutaka Matsuo, Jiaxian Guo |
CVPR | 7 |
| 2025 | Spark Transformer: Reactivating Sparsity in Transformer FFN and AttentionabstractThe discovery of the *lazy neuron phenomenon* (Li et al., 2022), where fewer than 10% of the feedforward networks (FFN) parameters in trained Transformers are activated per token, has spurred significant interests in *activation sparsity* for enhancing large model efficiency. While notable progress has been made in translating such sparsity to wall-time benefits across CPUs, GPUs, and TPUs, modern Transformers have moved away from the ReLU activation function crucial to this phenomenon. Existing efforts on re-introducing activation sparsity, e.g., by reverting to ReLU or applying top-k masking, often degrade model quality, increase parameter count, or complicate training. Sparse attention, the application of sparse activation to the attention mechanism, often face similar challenges.
This paper introduces the Spark Transformer, a novel architecture that achieves high activation sparsity in both FFN and the attention mechanism while maintaining model quality, parameter count, and standard training procedures. Our method realizes sparsity via top-$k$ masking for explicit control over sparsity level. Crucially, we introduce *statistical top-k*, a hardware-accelerator-friendly, linear-time approximate algorithm that avoids costly sorting and mitigates significant training slowdown from standard top-k operators. Furthermore, Spark Transformer reallocates existing FFN parameters and attention key embeddings to form a low-cost predictor for identifying activated entries. This design not only mitigates quality loss from enforced sparsity, but also enhances wall-time benefit. Pretrained with the Gemma-2 recipe, Spark Transformer demonstrates competitive performance on standard benchmarks while exhibiting significant sparsity: only 8\% of FFN neurons are activated, and each token attends to a maximum of 256 tokens. This translates to a 2.5x reduction in FLOPs, leading to decoding wall-time speedups of up to 1.79x on CPU and 1.40xon GPU. Chong You, Zhipeng Jia, Lin Chen 0003, Srinadh Bhojanapalli, Jiaxian Guo, Utku Evci, Jan Wassenberg, Praneeth Netrapalli, Jeremiah Willcock, Suvinay Subramanian, Felix Chern, Alek Andreev, Shreya Pathak, Felix X. Yu, Prateek Jain 0002, David E. Culler, Henry M. Levy, Sanjiv Kumar |
NeurIPS | 6 |
| 2025 | Real-time data-efficient portrait stylization via geometric alignment
Zhuoru Li, Xuanyu Yin, Yusuke Iwasawa, Yutaka Matsuo, Jiaxian Guo |
Neural Networks | 7 |
| 2024 | Paste and Harmonize via Denoising: Subject-Driven Image Editing with Frozen Pre-Trained Diffusion ModelabstractText-to-Image generative models have shown a remarkable ability to produce high-quality images. However, existing methods still face difficulties in exemplar-guided image editing without destroying the given objects’ identity in the exemplar image. To address this problem, we propose a new framework called Paste and Harmonize via Denoising, which leverages pre-trained diffusion models to facilitate the text-driven transfer of objects from an exemplar image to the edited image while preserving their appearance and characteristics. The framework consists of two main steps: paste and harmonize via denoising. In the paste step, an off-the-shelf text-driven model is utilized to localize the objects in the exemplar image. The editing task is naturally transformed into an image harmonization task by pasting the object patches into the edited image. In the harmonize via denoising step, we introduce an image harmonization module based on pre-trained diffusion models to blend the inserted object with the target image, producing a coherent and realistic image without compromising synthesis quality and preserving the text-driven style transfer editing ability. In the experiments, the qualitative comparisons with baselines demonstrate that our method achieves impressive performance in exemplar-based image editing on both training and in-the-wild images with high fidelity. More qualitative and quantitative results can be found at our website. Jiaxian Guo, Paul Yoo, Yutaka Matsuo, Yusuke Iwasawa |
ICASSP | 2 |
| 2024 | GenDOM: Generalizable One-shot Deformable Object Manipulation with Parameter-Aware PolicyabstractDue to the inherent uncertainty in their deformability during motion, previous methods in deformable object manipulation, such as rope and cloth, often required hundreds of real-world demonstrations to train a manipulation policy for each object, which hinders their applications in our ever-changing world. To address this issue, we introduce GenDOM, a framework that allows the manipulation policy to handle different deformable objects with only a single real-world demonstration. To achieve this, we augment the policy by conditioning it on deformable object parameters and training it with a diverse range of simulated deformable objects so that the policy can adjust actions based on different object parameters. At the time of inference, given a new object, GenDOM can estimate the deformable object parameters with only a single real-world demonstration by minimizing the disparity between the grid density of point clouds of real-world demonstrations and simulations in a differentiable physics simulator. Empirical validations on both simulated and real-world object manipulation setups clearly show that our method can manipulate different objects with a single demonstration and significantly outperforms the baseline in both environments (a 62% improvement for in-domain ropes and a 15% improvement for out-of-distribution ropes in simulation, as well as a 26% improvement for ropes and a 50% improvement for cloths in the real world), demonstrating the effectiveness of our approach in one-shot deformable object manipulation. https://sites.google.com/view/gendom/home. So Kuroki, Jiaxian Guo, Tatsuya Matsushima, Takuya Okubo, Masato Kobayashi 0001, Yuya Ikeda, Ryosuke Takanami, Paul Yoo, Yutaka Matsuo, Yusuke Iwasawa |
ICRA | 2 |
| 2024 | In-N-Out: Lifting 2D Diffusion Prior for 3D Object Removal via Tuning-Free Latents AlignmentabstractNeural representations for 3D scenes have made substantial advancements recently, yet object removal remains a challenging yet practical issue, due to the absence of multi-view supervision over occluded areas. Diffusion Models (DMs), trained on extensive 2D images, show diverse and high-fidelity generative capabilities in the 2D domain. However, due to not being specifically trained on 3D data, their application to multi-view data often exacerbates inconsistency, hence impacting the overall quality of the 3D output. To address these issues, we introduce "In-N-Out", a novel approach that begins by inpainting a prior, i.e., the occluded area from a single view using DMs, followed by outstretching it to create multi-view inpaintings via latents alignments. Our analysis identifies that the variability in DMs' outputs mainly arises from initially sampled latents and intermediate latents predicted in the denoising process. We explicitly align of initial latents using a Neural Radiance Field (NeRF) to establish a consistent foundational structure in the inpainted area, complemented by an implicit alignment of intermediate latents through cross-view attention during the denoising phases, enhancing appearance consistency across views. To further enhance rendering results, we apply a patch-based hybrid loss to optimize NeRF. We demonstrate that our techniques effectively mitigate the challenges posed by inconsistencies in DMs and substantially improve the fidelity and coherence of inpainted 3D representations. Dongting Hu, Huan Fu, Jiaxian Guo, Liuhua Peng, Tingjin Chu, Feng Liu 0003, Tongliang Liu, Mingming Gong |
NeurIPS | 3 |
| 2023 | From Images to Textual Prompts: Zero-shot Visual Question Answering with Frozen Large Language ModelsabstractLarge language models (LLMs) have demonstrated excellent zero-shot generalization to new language tasks. However, effective utilization of LLMs for zero-shot visual question-answering (VQA) remains challenging, primarily due to the modality disconnect and task disconnect between the LLM and VQA tasks. End-to-end training on multimodal data may bridge the disconnects, but is inflexible and computationally expensive. To address this issue, we propose Img2LLM, a plug-and-play module that provides LLM prompts to enable LLMs to perform zeroshot VQA tasks without end-to-end training. We develop LLM-agnostic models describe image content as exemplar question-answer pairs, which prove to be effective LLM prompts. Img2LLM offers the following benefits: 1) It achieves comparable or better performance than methods relying on end-to-end training. For example, we outperform Flamingo [3] by 5.6% on VQAv2. On the challenging A-OKVQA dataset, our method outperforms few-shot methods by as much as 20%. 2) It flexibly interfaces with a wide range of LLMs to perform VQA. 3) It eliminates the need to specialize LLMs using end-to-end finetuning and serve highly specialized LLMs to end users, thereby reducing cost. Code is available via the LAVIS [28] framework at https://github.com/salesforce/LAVIS/tree/main/projects/img2llm-vqa. Jiaxian Guo, Junnan Li 0001, Dongxu Li 0003, Anthony Meng Huat Tiong, Boyang Li 0001, Dacheng Tao, Steven C. H. Hoi |
CVPR | 1 |
| 2023 | An Efficient End-to-End Training Approach for Zero-Shot Human-AI CoordinationabstractThe goal of zero-shot human-AI coordination is to develop an agent that can collaborate with humans without relying on human data. Prevailing two-stage population-based methods require a diverse population of mutually distinct policies to simulate diverse human behaviors. The necessity of such populations severely limits their computational efficiency. To address this issue, we propose E3T, an **E**fficient **E**nd-to-**E**nd **T**raining approach for zero-shot human-AI coordination. E3T employs a mixture of ego policy and random policy to construct the partner policy, making it both coordination-skilled and diverse. In this way, the ego agent is end-to-end trained with this mixture policy without the need of a pre-trained population, thus significantly improving the training efficiency. In addition, a partner modeling module is proposed to predict the partner's action from historical information. With the predicted partner's action, the ego policy is able to adapt its policy and take actions accordingly when collaborating with humans of different behavior patterns. Empirical results on the Overcooked environment show that our method significantly improves the training efficiency while preserving comparable or superior performance than the population-based baselines. Demo videos are available at https://sites.google.com/view/e3t-overcooked. Jiaxian Guo, Xingzhou Lou, Jun Wang 0012, Haifeng Zhang 0002, Yali Du 0001 |
NeurIPS | 2 |
| 2023 | DreamSparse: Escaping from Plato's Cave with 2D Diffusion Model Given Sparse ViewsabstractSynthesizing novel view images from a few views is a challenging but practical problem. Existing methods often struggle with producing high-quality results or necessitate per-object optimization in such few-view settings due to the insufficient information provided. In this work, we explore leveraging the strong 2D priors in pre-trained diffusion models for synthesizing novel view images. 2D diffusion models, nevertheless, lack 3D awareness, leading to distorted image synthesis and compromising the identity. To address these problems, we propose $\textit{DreamSparse}$, a framework that enables the frozen pre-trained diffusion model to generate geometry and identity-consistent novel view images. Specifically, DreamSparse incorporates a geometry module designed to capture features about spatial information from sparse views as a 3D prior. Subsequently, a spatial guidance model is introduced to convert rendered feature maps as spatial information for the generative process. This information is then used to guide the pre-trained diffusion model to
encourage the synthesis of geometrically consistent images without further tuning. Leveraging the strong image priors in the pre-trained diffusion models, DreamSparse is capable of synthesizing high-quality novel views for both object and object-centric scene-level images and generalising to open-set images.
Experimental results demonstrate that our framework can effectively synthesize novel view images from sparse views and outperforms baselines in both trained and open-set category images. More results can be found on our project page: https://sites.google.com/view/dreamsparse-webpage. Paul Yoo, Jiaxian Guo, Yutaka Matsuo, Shixiang Gu |
NeurIPS | 2 |
| 2023 | Disentangled Attribute Features Vision Transformer for Pedestrian Attribute Recognition
Caihua Liu, Jiaxian Guo, Sichu Chen, Xia Feng |
PRCV (6) | 2 |
| 2023 | Prescribed Safety Performance Imitation Learning From a Single Expert DatasetabstractExisting safe imitation learning (safe IL) methods mainly focus on learning safe policies that are similar to expert ones, but may fail in applications requiring different safety constraints. In this paper, we propose the Lagrangian Generative Adversarial Imitation Learning (LGAIL) algorithm, which can adaptively learn safe policies from a single expert dataset under diverse prescribed safety constraints. To achieve this, we augment GAIL with safety constraints and then relax it as an unconstrained optimization problem by utilizing a Lagrange multiplier. The Lagrange multiplier enables explicit consideration of the safety and is dynamically adjusted to balance the imitation and safety performance during training. Then, we apply a two-stage optimization framework to solve LGAIL: (1) a discriminator is optimized to measure the similarity between the agent-generated data and the expert ones; (2) forward reinforcement learning is employed to improve the similarity while considering safety concerns enabled by a Lagrange multiplier. Furthermore, theoretical analyses on the convergence and safety of LGAIL demonstrate its capability of adaptively learning a safe policy given prescribed safety constraints. At last, extensive experiments in OpenAI Safety Gym conclude the effectiveness of our approach. Zhihao Cheng, Li Shen 0008, Miaoxi Zhu, Jiaxian Guo, Liu Liu 0014, Bo Du 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Alleviating Semantics Distortion in Unsupervised Low-Level Image-to-Image Translation via Structure Consistency ConstraintabstractUnsupervised image-to-image (I21) translation aims to learn a domain mapping function that can preserve the semantics of the input images without paired data. However, because the underlying semantics distributions in the source and target domains are often mismatched, current distribution matching-based methods may distort the semantics when matching distributions, resulting in the inconsistency between the input and translated images, which is known as the semantics distortion problem. In this paper, we focus on the low-level I21 translation, where the structure of images is highly related to their semantics. To alleviate semantic distortions in such translation tasks without paired supervision, we propose a novel I21 translation constraint, called Structure Consistency Constraint (SCC), to promote the consistency of image structures by reducing the randomness of color transformation in the translation process. To facilitate estimation and maximization of SCC, we propose an approximate representation of mutual information called relative Squared-loss Mutual Information (rSMI) that enjoys efficient analytic solutions. Our SCC can be easily incorporated into most existing translation models. Quantitative and qualitative comparisons on a range of low-level I21 translation tasks show that translation models with SCC outperform the original models by a significant margin with little additional computational and memory costs. Jiaxian Guo, Huan Fu, Mingming Gong, Kun Zhang 0001, Dacheng Tao |
CVPR | 1 |
| 2022 | Online Continual Learning with Contrastive Vision Transformer
Zhen Wang 0030, Liu Liu 0014, Yajing Kong, Jiaxian Guo, Dacheng Tao |
ECCV (20) | 4 |
| 2022 | A Relational Intervention Approach for Unsupervised Dynamics Generalization in Model-Based Reinforcement Learning
Jiaxian Guo, Mingming Gong, Dacheng Tao |
ICLR | 1 |
| 2020 | LTF: A Label Transformation Framework for Correcting Label ShiftabstractDistribution shift is a major obstacle to the deployment of current deep learning models on real-world problems. Let $Y$ be the class label and $X$ the features. We focus on one type of distribution shift, \emph{ label shift}, where the label marginal distribution $P_Y$ changes but the conditional distribution $P_{X|Y}$ does not. Most existing methods estimate the density ratio between the source- and target-domain label distributions by density matching. However, these methods are either computationally infeasible for large-scale data or restricted to shift correction for discrete labels. In this paper, we propose an end-to-end Label Transformation Framework (LTF) for correcting label shift, which implicitly models the shift of $P_Y$ and the conditional distribution $P_{X|Y}$ using neural networks. Thanks to the flexibility of deep networks, our framework can handle continuous, discrete, and even multi-dimensional labels in a unified way and is scalable to large data. Moreover, for high dimensional $X$, such as images, we find that the redundant information in $X$ severely degrades the estimation accuracy. To remedy this issue, we propose to match the distribution implied by our generative model and the target-domain distribution in a low-dimensional feature space that discards information irrelevant to $Y$. Both theoretical and empirical studies demonstrate the superiority of our method over previous approaches. Jiaxian Guo, Mingming Gong, Tongliang Liu, Kun Zhang 0001, Dacheng Tao |
ICML | 1 |
| 2019 | DNQ: Dynamic Network QuantizationabstractIn this paper, we propose a Dynamic Network Quantization (DNQ) framework. Unlike most existing quantization methods that use a universal quantization bit-width for the whole network, we utilize policy gradient [1] to train an agent to learn the bit-width of each layer by the bit-width controller. Yuhui Xu 0002, Shuai Zhang 0009, Yingyong Qi, Jiaxian Guo, Weiyao Lin, Hongkai Xiong |
DCC | 4 |
| 2018 | Long Text Generation via Adversarial Training with Leaked InformationabstractAutomatically generating coherent and semantically meaningful text has many applications in machine translation, dialogue systems, image captioning, etc. Recently, by combining with policy gradient, Generative Adversarial Nets(GAN) that use a discriminative model to guide the training of the generative model as a reinforcement learning policy has shown promising results in text generation. However, the scalar guiding signal is only available after the entire text has been generated and lacks intermediate information about text structure during the generative process. As such, it limits its success when the length of the generated text samples is long (more than 20 words). In this paper, we propose a new framework, called LeakGAN, to address the problem for long text generation. We allow the discriminative net to leak its own high-level extracted features to the generative net to further help the guidance. The generator incorporates such informative signals into all generation steps through an additional MANAGER module, which takes the extracted features of current generated words and outputs a latent vector to guide the WORKER module for next-word generation.Our extensive experiments on synthetic data and various real-world tasks with Turing test demonstrate that LeakGAN is highly effective in long text generation and also improves the performance in short text generation scenarios. More importantly, without any supervision, LeakGAN would be able to implicitly learn sentence structures only through the interaction between MANAGER and WORKER. Jiaxian Guo, Sidi Lu, Han Cai, Weinan Zhang 0001, Yong Yu 0001, Jun Wang 0012 |
AAAI | 1 |
| 2018 | Texygen: A Benchmarking Platform for Text Generation ModelsabstractWe introduce Texygen, a benchmarking platform to support research on open-domain text generation models. Texygen has not only implemented a majority of text generation models, but also covered a set of metrics that evaluate the diversity, the quality and the consistency of the generated texts. The Texygen platform could help standardize the research on text generation and improve the reproductivity and reliability of future research work in text generation. Yaoming Zhu, Sidi Lu, Lei Zheng 0004, Jiaxian Guo, Weinan Zhang 0001, Jun Wang 0012, Yong Yu 0001 |
SIGIR | 4 |