VLDB 2026 Research / reviewers in the wild / expert
Xiaoyu Kong
dblp:190/2233
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0004-1386-9394ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention HeadsabstractDiffusion Transformers (DiTs) have exhibited robust capabilities in image generation tasks. However, accurate text-guided image editing for multimodal DiTs (MM-DiTs) still poses a significant challenge. Unlike UNet-based structures that could utilize self/cross-attention maps for semantic editing, MM-DiTs inherently lack support for explicit and consistent incorporated text guidance, resulting in semantic misalignment between the edited results and texts. In this study, we disclose the sensitivity of different attention heads to different image semantics within MM-DiTs and introduce HeadRouter , a training-free image editing framework that edits the source image by adaptively routing the text guidance to different attention heads in MM-DiTs. Furthermore, we propose a dual-token refinement module to refine text/image token representations for precise semantic guidance and accurate region expression. Experiments on multiple benchmarks demonstrate HeadRouter’s performance in terms of editing fidelity and image quality. The code is available at https://github.com/ICTMCG/HeadRouter . Fan Tang, Juan Cao 0001, Xiaoyu Kong, Yuxin Zhang 0006, Jintao Li 0001, Oliver Deussen, Tong-Yee Lee |
ACM Trans. Graph. | 4 |
| 2025 | Think before Recommendation: Autonomous Reasoning-enhanced RecommenderabstractThe core task of recommender systems is to learn user preferences from historical user-item interactions. With the rapid development of large language models (LLMs), recent research has explored leveraging the reasoning capabilities of LLMs to enhance rating prediction tasks. However, existing distillation-based methods suffer from limitations such as the teacher model's insufficient recommendation capability, costly and static supervision, and superficial transfer of reasoning ability. To address these issues, this paper proposes RecZero, a reinforcement learning (RL)-based recommendation paradigm that abandons the traditional multi-model and multi-stage distillation approach. Instead, RecZero trains a single LLM through pure RL to autonomously develop reasoning capabilities for rating prediction. RecZero consists of two key components: (1) "Think-before-Recommendation" prompt construction, which employs a structured reasoning template to guide the model in step-wise analysis of user interests, item features, and user-item compatibility; and (2) rule-based reward modeling, which adopts group relative policy optimization (GRPO) to compute rewards for reasoning trajectories and optimize the LLM. Additionally, the paper explores a hybrid paradigm, RecOne, which combines supervised fine-tuning with RL, initializing the model with cold-start reasoning samples and further optimizing it with RL. Experimental results demonstrate that RecZero and RecOne significantly outperform existing baseline methods on multiple benchmark datasets, validating the superiority of the RL paradigm in achieving autonomous reasoning-enhanced recommender systems. Xiaoyu Kong, Junguang Jiang, Ziru Xu, Zhu Han 0001, Jian Xu 0015, Bo Zheng 0007, Jiancan Wu, Xiang Wang 0010 |
NeurIPS | 1 |
| 2024 | Block Image Compressive Sensing with Local and Global Information InteractionabstractBlock image compressive sensing methods, which divide a single image into small blocks for efficient sampling and reconstruction, have achieved significant success. However, these methods process each block locally and thus disregard the global communication among different blocks in the reconstruction step. Existing methods have attempted to address this issue with local filters or by directly reconstructing the entire image, but they have only achieved insufficient communication among adjacent pixels or bypassed the problem. To directly confront the communication problem among blocks and effectively resolve it, we propose a novel approach called Block Reconstruction with Blocks' Communication Network (BRBCN). BRBCN focuses on both local and global information, while further taking their interactions into account. Specifically, BRBCN comprises dual CNN and Transformer architectures, in which CNN is used to reconstruct each block for powerful local processing and Transformer is used to calculate the global communication among all the blocks. Moreover, we propose a global-to-local module (G2L) and a local-to-global module (L2G) to effectively integrate the representations of CNN and Transformer, with which our BRBCN network realizes the bidirectional interaction between local and global information. Extensive experiments show our BRBCN method outperforms existing state-of-the-art methods by a large margin. The code is available at https://github.com/kongxiuxiu/BRBCN Xiaoyu Kong, Yongyong Chen, Feng Zheng 0001, Zhenyu He 0001 |
AAAI | 1 |
| 2024 | Revealing the Two Sides of Data Augmentation: An Asymmetric Distillation-based Win-Win Solution for Open-Set Recognition
Yunbing Jia, Xiaoyu Kong, Fan Tang, Yixing Gao 0001, Weiming Dong |
IJCAI | 2 |
| 2024 | Customizing Language Models with Instance-wise LoRA for Sequential RecommendationabstractSequential recommendation systems predict the next interaction item based on users' past interactions, aligning recommendations with individual preferences. Leveraging the strengths of Large Language Models (LLMs) in knowledge comprehension and reasoning, recent approaches are eager to apply LLMs to sequential recommendation. A common paradigm is converting user behavior sequences into instruction data, and fine-tuning the LLM with parameter-efficient fine-tuning (PEFT) methods like Low-Rank Adaption (LoRA). However, the uniform application of LoRA across diverse user behaviors is insufficient to capture individual variability, resulting in negative transfer between disparate sequences.
To address these challenges, we propose Instance-wise LoRA (iLoRA). We innovatively treat the sequential recommendation task as a form of multi-task learning, integrating LoRA with the Mixture of Experts (MoE) framework. This approach encourages different experts to capture various aspects of user behavior. Additionally, we introduce a sequence representation guided gate function that generates customized expert participation weights for each user sequence, which allows dynamic parameter adjustment for instance-wise recommendations.
In sequential recommendation, iLoRA achieves an average relative improvement of 11.4\% over basic LoRA in the hit ratio metric, with less than a 1\% relative increase in trainable parameters.
Extensive experiments on three benchmark datasets demonstrate the effectiveness of iLoRA, highlighting its superior performance compared to existing methods in mitigating negative transfer and improving recommendation accuracy.
Our data and code are available at https://github.com/AkaliKong/iLoRA. Xiaoyu Kong, Jiancan Wu, An Zhang 0003, Leheng Sheng, Xiang Wang 0010, Xiangnan He 0001 |
NeurIPS | 1 |
| 2024 | When Channel Correlation Meets Sparse Prior: Keeping Interpretability in Image Compressive SensingabstractImage compressive sensing (CS), recovering an unknown image by resorting to a small number of its measurements, has become an increasingly popular topic in multimedia technology and applications. For a better reconstruction, diverse priors, from the original sparse prior to the new deep prior, have been exploited. Despite the powerful learning capability and satisfactory reconstruction performance, the deep prior is known as a black box and loses clear interpretability. In this article, we first revisit image CS with different priors and observe that the method with hand-crafted sparse prior could still outperform state-of-art methods with deep prior or no prior when under the same settings, while the interpretability is well preserved. Then, towards a better performance of the sparse-prior-based method, we propose a Channel Adaptive Thresholding Network, namely CAT-Net. CAT-Net draws the support from channel correlation calculation to extend the single thresholding in the iterative soft thresholding algorithm (ISTA) into channel-wise thresholding. The channel adaptive thresholding conducts soft thresholding operation in each channel of the image features and can be adjusted adaptively to the inputs, which can reconstruct more precisely than a single static thresholding. The careful CAT operation can preserve patterns both in detail and holistically well. Experimental results demonstrate the proposed method outperforms the state-of-the-art image CS methods with both traditional and deep priors. Xiaoyu Kong, Yongyong Chen, Zhenyu He 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Exploring the Temporal Consistency of Arbitrary Style Transfer: A Channelwise PerspectiveabstractArbitrary image stylization by neural networks has become a popular topic, and video stylization is attracting more attention as an extension of image stylization. However, when image stylization methods are applied to videos, unsatisfactory results that suffer from severe flickering effects appear. In this article, we conducted a detailed and comprehensive analysis of the cause of such flickering effects. Systematic comparisons among typical neural style transfer approaches show that the feature migration modules for state-of-the-art (SOTA) learning systems are ill-conditioned and could lead to a channelwise misalignment between the input content representations and the generated frames. Unlike traditional methods that relieve the misalignment via additional optical flow constraints or regularization modules, we focus on keeping the temporal consistency by aligning each output frame with the input frame. To this end, we propose a simple yet efficient multichannel correlation network (MCCNet), to ensure that output frames are directly aligned with inputs in the hidden feature space while maintaining the desired style patterns. An inner channel similarity loss is adopted to eliminate side effects caused by the absence of nonlinear operations such as softmax for strict alignment. Furthermore, to improve the performance of MCCNet under complex light conditions, we introduce an illumination loss during training. Qualitative and quantitative evaluations demonstrate that MCCNet performs well in arbitrary video and image style transfer tasks. Code is available at https://github.com/kongxiuxiu/MCCNetV2. Xiaoyu Kong, Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Yongyong Chen, Zhenyu He 0001, Changsheng Xu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Semantic-Context Graph Network for Point-Based 3D Object DetectionabstractPoint-based indoor 3D object detection has received increasing attention with the large demand for augmented reality, autonomous driving, and robot technology in the industry. However, the detection precision suffers from inputs with semantic ambiguity, i.e., shape symmetries, occlusion, and texture missing, which would lead that different objects appearing similar from different viewpoints and then confusing the detection model. Typical point-based detectors relieve this problem via learning proposal representations with both geometric and semantic information, while the entangled representation may cause a reduction in both semantic and spatial discrimination. In this paper, we focus on alleviating the confusion from entanglement and then enhancing the proposal representation by considering the proposal’s semantics and the context in one scene. A semantic-context graph network (SCGNet) is proposed, which mainly includes two modules: a category-aware proposal recoding module (CAPR) and a proposal context aggregation module (PCAg). To produce semantically clear features from entanglement representation, the CAPR module learns a high-level semantic embedding for each category to extract discriminative semantic clues. In view of further enhancing the proposal representation and leveraging the semantic clues, the PCAg module builds a graph to mine the most relevant context in the scene. With few bells and whistles, the SCGNet achieves SOTA performance and obtains consistent gains when applying to different backbones (0.9% ~ 2.4% on ScanNet V2 and 1.6% ~ 2.2% on SUN RGB-D for [email protected]). Code is available at https://github.com/dsw-jlu-rgzn/SCGNet. Shuwei Dong, Xiaoyu Kong, Xingjia Pan, Fan Tang, Weiming Dong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Deep and Low-Rank Quaternion Priors for Color Image ProcessingabstractDue to the physical nature of color images, color image processing such as denoising and inpainting has shown extensive and versatile possibilities over grayscale image processing. The monochromatic and the concatenation model have been widely used to process color images by processing each color channel independently or concatenating three color channels as one unified one and then used existing grayscale image processing methods directly without specific operations. These above schemes, however, have some limitations: (1) they would destroy the inherent correlation among three color channels since they cannot represent color images holistically; (2) they usually focus on one specific handcrafted prior such as smoothness, low-rankness, or even deep prior and thus failing to fuse deep and handcrafted priors of color images flexibly. To conquer these limitations, we propose one unified model to integrate deep prior and low-rank quaternion prior (DLRQP) for color image processing under the plug-and-play (PnP) framework. Specifically, the quaternion representation with low-rank constraint is introduced to denote the color image in a holistic way and one advanced denoiser is adopted to explore the deep prior in an iterative process. To tightly approximate the quaternion rank, one nonconvex penalty function is further utilized. We derive an alternate iterative approach to tackle the proposed model. We empirically demonstrate that our model can achieve superior performance over existing methods on both color image denoising and inpainting tasks. Xiaoyu Kong, Qiangqiang Shen, Yongyong Chen, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |