VLDB 2026 Research / reviewers in the wild / expert
Xiaogang Xu 0002
dblp:118/2268-2
· DBLP profile ↗
69ranked-venue papers
13as first author
65since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 13 first-author · 50 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 11 first-author · 47 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust-R1: Degradation-Aware Reasoning for Robust Visual UnderstandingabstractMultimodal Large Language Models struggle to maintain reliable performance under extreme real-world visual degradations, which impede their practical robustness. Existing robust MLLMs predominantly rely on implicit training/adaptation that focuses solely on visual encoder generalization, suffering from limited interpretability and isolated optimization. To overcome these limitations, we propose Robust-R1, a novel framework that explicitly models visual degradations through structured reasoning chains. Our approach integrates: (i) supervised fine-tuning for degradation-aware reasoning foundations, (ii) reward-driven alignment for accurately perceiving degradation parameters, and (iii) dynamic reasoning depth scaling adapted to degradation intensity. To facilitate this approach, we introduce a specialized 11K dataset featuring realistic degradations synthesized across four critical real-world visual processing stages, each annotated with structured chains connecting degradation parameters, perceptual influence, pristine semantic reasoning chain, and conclusion. Comprehensive evaluations demonstrate state-of-theart robustness: Robust-R1 outperforms all general and robust baselines on the real-world degradation benchmark R-Bench, while maintaining superior anti-degradation performance under multi-intensity adversarial degradations on MMMB, MMStar, and RealWorldQA. Jiaqi Tang 0005, Jianmin Chen, Wei Wei 0008, Xiaogang Xu 0002, Runtao Liu, Qipeng Xie, Jiafei Wu, Lei Zhang 0001, Qifeng Chen 0001 |
AAAI | 4 |
| 2026 | Class Incremental Medical Image Segmentation via Prototype-Guided Calibration and Dual-Aligned DistillationabstractClass incremental medical image segmentation (CIMIS) aims to preserve knowledge of previously learned classes while learning new ones without relying on old-class annotations. However, existing methods 1) either adopt one-size-fits-all strategies that treat all spatial regions and feature channels equally, which may hinder the preservation of accurate old knowledge, 2) or focus solely on aligning local prototypes with global ones for old classes while overlooking their local representations in new data, leading to knowledge degradation. To mitigate the above issues, we propose Prototype-Guided Calibration Distillation (PGCD) and Dual-Aligned Prototype Distillation (DAPD) for CIMIS in this paper. Specifically, PGCD exploits prototype-to-feature similarity to calibrate class-specific distillation intensity in different spatial regions, effectively reinforcing reliable old knowledge and suppressing misleading cues from old classes. Complementarily, DAPD aligns the local prototypes of old classes extracted from the current model with both global historical prototypes and local prototypes, further enhancing segmentation performance on old categories. Comprehensive evaluations on two widely used multi-organ segmentation benchmarks demonstrate that our method outperforms current state-of-the-art methods, highlighting its robustness and generalization capabilities. Shengqian Zhu, Chengrong Yu, Guangjun Li, Jiafei Wu, Xiaogang Xu 0002, Zhang Yi 0001, Junjie Hu 0004 |
AAAI | 7 |
| 2026 | ACIArena: Toward Unified Evaluation for Agent Cascading InjectionabstractHengyu An, Minxi Li, Jinghuai Zhang, Naen Xu, Chunyi Zhou, Changjiang Li, Xiaogang Xu, Tianyu Du, Shouling Ji. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hengyu An, Minxi Li, Jinghuai Zhang, Naen Xu, Chunyi Zhou 0001, Changjiang Li, Xiaogang Xu 0002, Tianyu Du, Shouling Ji |
ACL (1) | 7 |
| 2026 | GADT: Enhancing transferable adversarial attacks through gradient-guided adversarial data transformation
Yating Ma, Xiaogang Xu 0002, Liming Fang 0001, Jiafei Wu, Lu Zhou 0002 |
Neurocomputing | 2 |
| 2026 | Restoration-Oriented Video Frame Interpolation With Region-Distinguishable Priors From SAM
Xiaogang Xu 0002, Yingqi Lin, Jiafei Wu, Zhe Liu 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Boosting HDR Image Reconstruction via Semantic Knowledge TransferabstractRecovering High Dynamic Range (HDR) images from multiple Standard Dynamic Range (SDR) images becomes challenging when the SDR images exhibit noticeable degradation and missing content. Leveraging scene-specific semantic priors offers a promising solution for restoring heavily degraded regions. However, these priors are typically extracted from sRGB SDR images, the domain/format gap poses a significant challenge when applying it to HDR imaging. To address this issue, we propose a general framework that transfers semantic knowledge derived from SDR domain via self-distillation to boost existing HDR reconstruction. Specifically, the proposed framework first introduces the Semantic Priors Guided Reconstruction Model (SPGRM), which leverages SDR image semantic knowledge to address ill-posed problems in the initial HDR reconstruction results. Subsequently, we leverage a self-distillation mechanism that constrains the color and content information with semantic knowledge, aligning the external outputs between the baseline and SPGRM. Furthermore, to transfer the semantic knowledge of the internal features, we utilize a Semantic Knowledge Alignment Module (SKAM) to fill the missing semantic contents with the complementary masks. Extensive experiments demonstrate that our framework significantly boosts HDR imaging quality for existing methods without altering the network architecture. Tao Hu 0013, Longyao Wu, Wei Dong 0010, Peng Wu 0015, Jinqiu Sun, Xiaogang Xu 0002, Qingsen Yan, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2026 | IAMAgent: Toward an Interactive and Adaptive Multi-Agent System for Image RestorationabstractExisting image restoration and enhancement (IRE) methods suffer from three fundamental limitations: 1) they present a high technical barrier, requiring expert knowledge and lacking intuitive natural language control; 2) they are inflexible and poorly adaptable, as models are typically designed for single, specific degradations and fail on complex or mixed real-world scenarios; and 3) they lack interactivity and ignore subjectivity, operating as "closed-box" tools that cannot incorporate human feedback or understand nuanced user intentions. To overcome these challenges, we pioneer a novel paradigm: a Multi-Agent System (MAS) for interactive and adaptive image restoration. We design and implement a prototype system, Interactive and Adaptive Multi-Agent System (IAMAgent), which orchestrates a team of specialized agents to collaboratively solve complex IRE tasks. At its core, a Manager Agent, driven by a Large Language Model, interprets user commands, devises strategies, and allocates sub-tasks. It directs a Perception Agent for degradation diagnosis, a suite of specialized Execution Agents that encapsulate various low-level vision models, and a Critique Agent for automated quality assessment. This collaborative framework enables an innovative, language-driven, and human-in-the-loop optimization process. Our work is the first to introduce the MAS paradigm to the IRE domain, transforming it from a collection of static tools into a dynamic, user-centric, and intelligent system. We demonstrate that IAMAgent not only significantly enhances restoration performance and adaptability but also bridges the critical gap between high-level human intention and low-level vision tasks. Yanyan Wei, Yilin Zhang 0012, Jiahuan Ren, Xiaogang Xu 0002, Zenglin Shi, Zhao Zhang 0001, Meng Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Exploiting Regional Information Transformer for Single Image DerainingabstractTransformer-based Single Image Deraining (SID) methods have achieved remarkable success, primarily attributed to their robust capability in capturing long-range interactions. However, we've noticed that current methods handle rain-affected and unaffected regions concurrently, overlooking the disparities between these areas, resulting in confusion between rain streaks and background parts, and inabilities to obtain effective interactions, ultimately resulting in suboptimal deraining outcomes. To address the above issue, we introduce the Region Transformer (Regformer), a novel SID method that underlines the importance of independently processing rain-affected and unaffected regions while considering their combined impact for high-quality image reconstruction. The crux of our method is the innovative Region Transformer Block (RTB), which integrates a Region Masked Attention (RMA) mechanism and a Mixed Gate Forward Block (MGFB). Our RTB is used for attention selection of rain-affected and unaffected regions and local modeling of mixed scales. The RMA generates attention maps tailored to these two regions and their interactions, enabling our model to capture comprehensive features essential for rain removal. To better recover high-frequency textures and capture more local details, we develop the MGFB as a compensation module to complete local mixed scale modeling. Extensive experiments demonstrate that our model reaches state-of-the-art performance, significantly improving the image deraining quality. Our code and trained models are publicly available athttps://github.com/ztMotaLee/Regformer. Baiang Li, Zhao Zhang 0001, Xiaogang Xu 0002, Yanyan Wei, Jicong Fan 0001, Meng Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Particle Rendering: Implicitly Aggregating Incident and Outgoing Light Fields for Novel View SynthesisabstractThis paper presents Particle Rendering (PR), a new implicit rendering approach that extends Neural Radiance Fields (NeRF) by incorporating incident light along with traditional outgoing light modeling. In our framework, a 3D scene consists of a mass of particles, each offering a deeper understanding of light interactions by reflecting and emitting light in all directions. Our methodology involves a three-phase training pipeline: 1) Estimating the outgoing light field through a NeRF model; 2) Distilling the incident light field. A simple metric is introduced to assess the quality of the ray for better supervision; 3) Implicit rendering. We propose an implicit method to aggregate incident and outgoing fields that leverages Multilayer Perceptrons (MLP) to directly infer final pixel values, thus avoiding the limitation of traditional physically-based rendering techniques. The effectiveness of PR is demonstrated through state-of-the-art results in various challenging indoor and outdoor scenes, emphasizing its capability to handle complex lighting and reflective materials. Tao Hu 0011, Zhiwen Yan, Xiaogang Xu 0002, Gim Hee Lee |
3DV | 3 |
| 2025 | DiMSOD: A Diffusion-Based Framework for Multi-Modal Salient Object DetectionabstractMulti-modal salient object detection (SOD) through the integration of additional data such as depth or thermal information has become a significant task in computer vision during recent years. Traditionally, the challenges of identifying salient objects in RGB, RGB-D (Depth), and RGB-T (Thermal) images are tackled separately. However, without intricate cross-modal fusion strategies, such approaches struggle to effectively integrate multi-modal information, often resulting in poorly defined object edges or overconfident inaccurate predictions. Recent studies have shown that designing a unified end-to-end framework to handle all three types of SOD tasks simultaneously is both necessary and difficult. To address this need, we propose a novel approach that treats multi-modal SOD as a conditional mask generation task utilizing diffusion models. We introduce DiMSOD, which enables the concurrent use of local (depth maps, thermal maps) and global controls (original images) within a unified model for progressive denoising and refined prediction. DiMSOD is efficient, only requiring fine-tuning of our newly introduced modules on the existing stable diffusion, which not only reduces the fine-tuning cost, making it more viable for practical use, but also enhances the integration of multi-modal conditional controls. Specifically, we have developed modules including SOD-ControlNet, Feature Adaptive Network (FAN), and Feature Injection Attention Network (FIAN) to enhance the model's performance. Extensive experiments demonstrate that DiMSOD efficiently detects salient objects across RGB, RGB-D, and RGB-T datasets, achieving superior performance compared to previous well-established methods. Shuo Zhang 0013, Wenbing Tang 0001, Terrence Hu, Xiaogang Xu 0002, Jing Liu 0012 |
AAAI | 6 |
| 2025 | DR-Encoder: Encode Low-rank Gradients with Random Prior for Large Language Models Differentially PrivatelyabstractThe emergence of the large language model (LLM) has shown its superiority in a wide range of disciplines, including language understanding and translation, relational logic reasoning, and even partial differential equations solving. The transformer is the pervasive backbone architecture for the foundation model construction. It is vital to research how to adjust the Transformer architecture to achieve an end-to-end privacy guarantee in LLM fine-tuning. This paper investigates three potential information leaks during a federated fine-tuning procedure for LLM (FedLLM). Based on the potential information leakage, we insert two-stage randomness into FedLLM to provide an end-to-end privacy guarantee solution. The first stage is to train a gradient auto-encoder with a Gaussian random prior based on the statistical information of the gradients generated by local clients. The second stage is fine-tuning the overall LLM with a differential privacy guarantee by adopting appropriate Gaussian noises. We show our proposed method's efficiency and accuracy gains with several foundation models and two popular evaluation benchmarks. Furthermore, we present a comprehensive privacy analysis with Gaussian Differential Privacy (GDP) and Renyi Differential Privacy (RDP). Huiwen Wu, Deyi Zhang, Xiaogang Xu 0002, Jiafei Wu, Zhe Liu 0001 |
AAAI | 4 |
| 2025 | CG-FedLLM: How to Compress Gradients in Federated Fine-Tuning for Large Language ModelsabstractThe success of current Large-Language Models (LLMs) hinges on extensive training data that are collected and stored centrally, called Centralized Learning (CL). However, such a collection manner poses a privacy threat, and one potential solution is Federated Learning (FL), which transfers gradients, not raw data, among clients. Unlike traditional networks, FL for LLMs incurs significant communication costs due to their tremendous parameters. In this study, we introduce an innovative approach to compress gradients to improve communication efficiency during LLM FL, formulating the new FL pipeline named CG-FedLLM. This approach integrates an encoder on the client side to acquire the compressed gradient features and a decoder on the server side to reconstruct the gradients. We also develop a novel training strategy that comprises Temporal-ensemble Gradient-Aware Pre-training (TGAP) to identify characteristic gradients of the target model and Federated AutoEncoder-Involved Fine-tuning (FAF) to compress gradients adaptively. Extensive experiments confirm that our approach reduces communication costs and improves performance (e.g., average 3 points increment compared with traditional CL- and FL-based fine-tuning with several foundation models on well-recognized benchmarks, MMLU and C-Eval). This is because our encoder-decoder, trained via TGAP and FAF, can filter gradients while selectively preserving critical features. Furthermore, we present a series of experimental analyses that focus on the communication efficiency, accuracy, and generalization ability within this privacy-centric framework, providing insights into the development of more efficient and private LLMs fine-tuning. Huiwen Wu, Xiaogang Xu 0002, Deyi Zhang, Jiafei Wu, Zhe Liu 0001 |
ECAI | 2 |
| 2025 | Diffusion Noise Feature: Accurate and Fast Generated Image DetectionabstractGenerative models now produce images with such stunning realism that they can easily deceive the human eye. While this progress unlocks vast creative potential, it also presents significant risks, such as the spread of misinformation. Consequently, detecting generated images has become a critical research challenge. However, current detection methods are often plagued by low accuracy and poor generalization. In this paper, to address these limitations and enhance the detection of generated images, we propose a novel representation, DIFFUSION NOISE FEATURE (DNF). Derived from the inverse process of diffusion models, DNF effectively amplifies the subtle, high-frequency artifacts that act as fingerprints of artificial generation. Our key insight is that real and generated images exhibit distinct DNF signatures, providing a robust basis for differentiation. By training a simple classifier such as ResNet-50 on DNF, our approach achieves remarkable accuracy, robustness, and generalization in detecting generated images, including those from unseen generators or with novel content. Extensive experiments across four training datasets and five test sets confirm that DNF establishes a new state-of-the-art in generated image detection. The code is available at https://github.com/YichiCS/Diffusion-Noise-Feature. Yichi Zhang 0015, Xiaogang Xu 0002 |
ECAI | 2 |
| 2025 | Exposure-Limited Image Enhancement with Generative Diffusion PriorabstractMany consumer cameras are equipped with 8-bit image sensors, which often struggle to capture scenes with a High Dynamic Range (HDR). This limitation can result in overexposed or underexposed regions, a loss of fine details due to low bit-depth compression, skewed color distributions, and noticeable noise in dark areas. Traditional Standard Dynamic Range (SDR) image enhancement methods typically focus on color mapping by expanding the color range and adjusting brightness. However, they often fail to restore details in dynamic range extremes, i.e. regions where pixel values approach the minimum or maximum limits. We define “exposure-limited image enhancement” as the process of enhancing images with large missing areas due to exposure issues within the SDR space, which differs from existing “mis-exposed image enhancement” methods primarily aimed at correcting color distributions. To enhance these exposure-limited images and overcome the limitations of current models, we propose a novel two-stage approach. In the first stage, we remap color and brightness to a suitable range while preserving existing details. In the second stage, we use a diffusion prior to generate content in severely overexposed or underexposed regions, which are otherwise lost during capture. Notably, this generative refinement module can also serve as a plug-and-play component alongside existing enhancement methods. Extensive experiments demonstrate that our method significantly improves image quality and detail, outperforming state-of-the-art techniques in dynamic range extremes. The project page is at https://Sagiri0208.github.io. Baiang Li, Sizhuo Ma, Yanhong Zeng, Xiaogang Xu 0002, Youqing Fang, Zhao Zhang 0001, Jian Wang 0100, Kai Chen 0026 |
ICCP | 4 |
| 2025 | Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive Injection
Wei Wei 0008, Jiaqi Tang 0005, Jiangtao Nie, Yanyu Ye, Xiaogang Xu 0002, Ying-Cong Chen, Lei Zhang 0001 |
ICCV | 6 |
| 2025 | DiffDoctor: Diagnosing Image Diffusion Models Before TreatingabstractIn spite of recent progress, image diffusion models still produce artifacts. A common solution is to leverage the feedback provided by quality assessment systems or human annotators to optimize the model, where images are generally rated in their entirety. In this work, we believe problem-solving starts with identification, yielding the request that the model should be aware of not just the presence of defects in an image, but their specific locations. Motivated by this, we propose DiffDoctor, a two-stage pipeline to assist image diffusion models in generating fewer artifacts. Concretely, the first stage targets developing a robust artifact detector, for which we collect a dataset of over 1M flawed synthesized images and set up an efficient human-in-the-loop annotation process, incorporating a carefully designed class-balance strategy. The learned artifact detector is then involved in the second stage to optimize the diffusion model by providing pixel-level feedback. Extensive experiments on text-to-image diffusion models demonstrate the effectiveness of our artifact detector as well as the soundness of our diagnose-then-treat design. Xi Chen 0119, Xiaogang Xu 0002, Sihui Ji, Yu Liu 0063, Yujun Shen, Hengshuang Zhao |
ICCV | 3 |
| 2025 | Learnable Feature Patches and Vectors for Boosting Low-Light Image Enhancement Without External Knowledge
Xiaogang Xu 0002, Jiafei Wu, Qingsen Yan, Jiequan Cui, Richang Hong, Bei Yu 0001 |
ICCV | 1 |
| 2025 | Towards Understanding the Robustness of Diffusion-Based Purification: A Stochastic PerspectiveabstractDiffusion-Based Purification (DBP) has emerged as an effective defense mechanism against adversarial attacks. The success of DBP is often attributed to the forward diffusion process, which reduces the distribution gap between clean and adversarial images by adding Gaussian noise. Although this explanation is theoretically grounded, the precise contribution of this process to robustness remains unclear. In this paper, through a systematic investigation, we propose that the intrinsic stochasticity in the DBP procedure is the primary factor driving robustness. To explore this hypothesis, we introduce a novel Deterministic White-Box (DW-box) evaluation protocol to assess robustness in the absence of stochasticity, and analyze attack trajectories and loss landscapes. Our results suggest that DBP models primarily leverage stochasticity to evade effective attack directions, and that their ability to purify adversarial perturbations can be weak. To further enhance the robustness of DBP models, we propose Adversarial Denoising Diffusion Training (ADDT), which incorporates classifier-guided adversarial perturbations into diffusion training, thereby strengthening the models' ability to purify adversarial perturbations. Additionally, we propose Rank-Based Gaussian Mapping (RBGM) to improve the compatibility of perturbations with diffusion models. Experimental results validate the effectiveness of ADDT. In conclusion, our study suggests that future research on DBP can benefit from the perspective of decoupling stochasticity-based and purification-based robustness. Kezhao Liu, Ziyi Dong, Xiaogang Xu 0002, Pengxu Wei, Liang Lin 0004 |
ICLR | 5 |
| 2025 | LARM: Large Auto-Regressive Model for Long-Horizon Embodied IntelligenceabstractRecent embodied agents are primarily built based on reinforcement learning (RL) or large language models (LLMs). Among them, RL agents are efficient for deployment but only perform very few tasks. By contrast, giant LLM agents (often more than 1000B parameters) present strong generalization while demanding enormous computing resources. In this work, we combine their advantages while avoiding the drawbacks by conducting the proposed referee RL on our developed large auto-regressive model (LARM). Specifically, LARM is built upon a lightweight LLM (fewer than 5B parameters) and directly outputs the next action to execute rather than text. We mathematically reveal that classic RL feedbacks vanish in long-horizon embodied exploration and introduce a giant LLM based referee to handle this reward vanishment during training LARM. In this way, LARM learns to complete diverse open-world tasks without human intervention. Especially, LARM successfully harvests enchanted diamond equipment in Minecraft, which demands significantly longer decision-making chains than the highest achievements of prior best methods. Zhuoling Li, Xiaogang Xu 0002, Zhenhua Xu 0003, Ser-Nam Lim, Hengshuang Zhao |
ICML | 2 |
| 2025 | Low-Light Video Enhancement via Spatial-Temporal Consistent DecompositionabstractLow-Light Video Enhancement (LLVE) seeks to restore dynamic or static scenes plagued by severe invisibility and noise. In this paper, we present an innovative video decomposition strategy that incorporates view-independent and view-dependent components to enhance the performance of LLVE. We leverage dynamic cross-frame correspondences for the view-independent term (which primarily captures intrinsic appearance) and impose a scene-level continuity constraint on the view-dependent term (which mainly describes the shading condition) to achieve consistent and satisfactory decomposition results. To further ensure consistent decomposition, we introduce a dual-structure enhancement network featuring a cross-frame interaction mechanism. By supervising different frames simultaneously, this network encourages them to exhibit matching decomposition features. This mechanism can seamlessly integrate with encoder-decoder single-frame networks, incurring minimal additional parameter costs. Extensive experiments are conducted on widely recognized LLVE benchmarks, covering diverse scenarios. Our framework consistently outperforms existing methods, establishing a new SOTA performance. Xiaogang Xu 0002, Kun Zhou 0001, Tao Hu 0011, Jiafei Wu, Ruixing Wang, Hao Peng 0002, Bei Yu 0001 |
IJCAI | 1 |
| 2025 | CFSynthesis: Controllable and Free-view 3D Human Video SynthesisabstractHuman video synthesis aims to create lifelike characters in various environments. While 2D diffusion-based methods have made significant progress, they struggle to generalize to complex 3D poses and varying scene backgrounds. To address these limitations, we introduce CFSynthesis, a novel framework for generating high-quality human videos with customizable attributes, including identity, motion, and scene configurations. Our method leverages a texture-SMPL-based representation to ensure consistent and stable character appearances across free viewpoints. Additionally, we introduce a novel foreground-background separation strategy that effectively decomposes the scene as foreground and background, enabling seamless integration of user-defined backgrounds. Experimental results on multiple datasets show that CFSynthesis not only achieves state-of-the-art performance in complex human animations but also adapts effectively to 3D motions in free-view and user-specified scenarios. Liyuan Cui, Xiaogang Xu 0002, Wenqi Dong, Zesong Yang, Hujun Bao, Zhaopeng Cui |
ICMR | 2 |
| 2025 | PRIME: Prototype-Driven Class Incremental Learning for Medical Image SegmentationabstractClass incremental medical segmentation (CIMS) aims to sequentially learn new classes while preserving knowledge of previously learned categories in the absence of old-class labels. Current methods suffer from performance degradation under class imbalance and require additional segmentation heads to accommodate new categories. Inspired by recent prototype learning that leverages prototypes to achieve robust recognition of new categories under limited-data regimes, we introduce a Prototype-dRIven class increMEntal (PRIME) method. PRIME replaces the incremental segmentation heads with prototypes to mitigate class imbalance, allowing new class learning with the simple addition of new prototypes. Based on prototype learning, PRIME further involves three tailored techniques. First, prototype structure alignment imposes structural constraints on inter-prototype relations to maintain consistent relative distances in the feature space, improving the model's ability to distinguish distinct classes. Second, pixel-wise contrastive loss term groups embeddings of similar samples while separating those of different classes, enhancing segmentation accuracy across all categories. Finally, the consensus-based prototype update mechanism refines the old prototypes during the learning of new classes, preventing performance degradation on the old classes. Extensive experiments on two public multi-organ segmentation datasets demonstrate that our approach significantly outperforms state-of-the-art methods, validating the effectiveness of the proposed PRIME. Shengqian Zhu, Chengrong Yu, Wenbo Qi, Jiafei Wu, Guangjun Li, Zhang Yi 0001, Xiaogang Xu 0002, Junjie Hu 0004 |
ACM Multimedia | 8 |
| 2025 | MiCo: Multi-image Contrast for Reinforcement Visual ReasoningabstractThis work explores enabling Chain-of-Thought (CoT) reasoning to link visual cues across multiple images. A straightforward solution is to adapt rule-based reinforcement learning for Vision-Language Models (VLMs). However, such methods typically rely on manually curated question-answer pairs, which can be particularly challenging when dealing with fine-grained visual details and complex logic across images. Inspired by self-supervised visual representation learning, we observe that images contain inherent constraints that can serve as supervision. Based on this insight, we construct image triplets comprising two augmented views of the same image and a third, similar but distinct image. During training, the model is prompted to generate a reasoning process to compare these images (i.e., determine same or different). Then we optimize the model with rule-based reinforcement learning. Due to the high visual similarity and the presence of augmentations, the model must attend to subtle visual cues and perform logical reasoning to succeed. Experimental results demonstrate that, although trained solely on visual comparison tasks, the learned reasoning ability generalizes effectively to a wide range of questions. Without relying on any human-annotated question-answer pairs, our method achieves significant improvements on multi-image reasoning benchmarks and shows strong performance on general vision tasks. Xi Chen 0119, Mingkang Zhu, Shaoteng Liu, Xiaoyang Wu 0002, Xiaogang Xu 0002, Yu Liu 0063, Xiang Bai, Hengshuang Zhao |
NeurIPS | 5 |
| 2025 | Wan-Move: Motion-controllable Video Generation via Latent Trajectory GuidanceabstractWe present Wan-Move, a simple and scalable framework that brings motion control to video generative models. Existing motion-controllable methods typically suffer from coarse control granularity and limited scalability, leaving their outputs insufficient for practical use. We narrow this gap by achieving precise and high-quality motion control. Our core idea is to directly make the original condition features motion-aware for guiding video synthesis. To this end, we first represent object motions with dense point trajectories, allowing fine-grained control over the scene. We then project these trajectories into latent space and propagate the first frame's features along each trajectory, producing an aligned spatiotemporal feature map that tells how each scene element should move. This feature map serves as the updated latent condition, which is naturally integrated into the off-the-shelf image-to-video model, e.g., Wan-I2V-14B, as motion guidance without any architecture change. It removes the need for auxiliary motion encoders and makes fine-tuning base models easily scalable. Through scaled training, Wan-Move generates 5-second, 480p videos whose motion controllability rivals Kling 1.5 Pro's commercial Motion Brush, as indicated by user studies. To support comprehensive evaluation, we further design MoveBench, a rigorously curated benchmark featuring diverse content categories and hybrid-verified annotations. It is distinguished by larger data volume, longer video durations, and high-quality motion annotations. Extensive experiments on MoveBench and the public dataset consistently show Wan-Move's superior motion quality. Code, models, and benchmark data are made available. Ruihang Chu, Yefei He, Zhekai Chen, Shiwei Zhang 0001, Xiaogang Xu 0002, Dingdong Wang, Hongwei Yi, Xihui Liu, Hengshuang Zhao, Yu Liu 0063, Yingya Zhang, Yujiu Yang 0001 |
NeurIPS | 5 |
| 2025 | DiffCamera: Arbitrary Refocusing on ImagesabstractThe depth-of-field (DoF) effect, which introduces aesthetically pleasing blur, enhances photographic quality but is fixed and difficult to modify once the image has been created. This becomes problematic when the applied blur is undesirable (e.g., the subject is out of focus). To address this, we propose DiffCamera, a model that enables flexible refocusing of a created image conditioned on an arbitrary new focus point and a blur level. Specifically, we design a diffusion transformer framework for refocusing learning. However, the training requires pairs of data with different focus planes and bokeh levels in the same scene, which are hard to acquire. To overcome this limitation, we develop a simulation-based pipeline to generate large-scale image pairs with varying focus planes and bokeh levels. With the simulated data, we find that training with only a vanilla diffusion objective often leads to incorrect DoF behaviors due to the complexity of the task. This requires a stronger constraint during training. Inspired by the photographic principle that photos of different focus planes can be linearly blended into a multi-focus image, we propose a stacking constraint during training to enforce precise DoF manipulation. This constraint enhances model training by imposing physically grounded refocusing behavior that the focusing results should be faithfully aligned with the scene structure and the camera conditions so that they can be combined into the correct multi-focus image. We also construct a benchmark to evaluate the effectiveness of our refocusing model. Extensive experiments demonstrate that DiffCamera supports stable refocusing across a wide range of scenes, providing unprecedented control over DoF adjustments for photography and generative AI applications. Xi Chen 0119, Xiaogang Xu 0002, Yu Liu 0063, Hengshuang Zhao |
SIGGRAPH Asia | 3 |
| 2025 | Toward Unified 3D Object Detection via Algorithm and Data UnificationabstractRealizing unified 3D object detection, including both indoor and outdoor scenes, holds great importance in applications like robot navigation. However, involving various scenarios of data to train models poses challenges due to their significantly distinct characteristics, e.g., diverse geometry properties and heterogeneous domain distributions. In this work, we propose to address the challenges from two perspectives, the algorithm perspective and data perspective. In terms of the algorithm perspective, we first build a monocular 3D object detector based on the bird's-eye-view (BEV) detection paradigm, where the explicit feature projection is beneficial to addressing the geometry learning ambiguity. In this detector, we split the classical BEV detection architecture into two stages and propose an uneven BEV grid design to handle the convergence instability caused by geometry difference between scenarios. Besides, we develop a sparse BEV feature projection strategy to reduce the computational cost and a unified domain alignment method to handle heterogeneous domains. From the data perspective, we propose to incorporate depth information to improve training robustness. Specifically, we build the first unified multi-modal 3D object detection benchmark MM-Omni3D and extend the aforementioned monocular detector to its multi-modal version, which is the first unified multi-modal 3D object detector. We name the designed monocular and multi-modal detectors as UniMODE and MM-UniMODE, respectively. The experimental results reveal several insightful findings highlighting the benefits of multi-modal data and confirm the effectiveness of all the proposed strategies. Zhuoling Li, Xiaogang Xu 0002, Ser-Nam Lim, Hengshuang Zhao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Geometric-Aware Low-Light Image and Video Enhancement via Depth GuidanceabstractLow-Light Enhancement (LLE) is aimed at improving the quality of photos/videos captured under low-light conditions. It is worth noting that most existing LLE methods do not take advantage of geometric modeling. We believe that incorporating geometric information can enhance LLE performance, as it provides insights into the physical structure of the scene that influences illumination conditions. To address this, we propose a Geometry-Guided Low-Light Enhancement Refine Framework (GG-LLERF) designed to assist low-light enhancement models in learning improved features by integrating geometric priors into the feature representation space. In this paper, we employ depth priors as the geometric representation. Our approach focuses on the integration of depth priors into various LLE frameworks using a unified methodology. This methodology comprises two key novel modules. First, a depth-aware feature extraction module is designed to inject depth priors into the image representation. Then, the Hierarchical Depth-Guided Feature Fusion Module (HDGFFM) is formulated with a cross-domain attention mechanism, which combines depth-aware features with the original image features within LLE models. We conducted extensive experiments on public low-light image and video enhancement benchmarks. The results illustrate that our framework significantly enhances existing LLE methods. The source code and pre-trained models are available at https://github.com/Estheryingqi/GG-LLERF. Yingqi Lin, Xiaogang Xu 0002, Jiafei Wu, Zhe Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Parametric Linear Blend Skinning Model for Multiple-Shape 3D GarmentsabstractWe present a novel data-driven Parametric Linear Blend Skinning (PLBS) model meticulously crafted for generalized 3D garment dressing and animation. Previous data-driven methods are impeded by certain challenges including overreliance on human body modeling and limited adaptability across different garment shapes. Our method resolves these challenges via two goals: 1) Develop a model based on garment modeling rather than human body modeling. 2) Separately construct low-dimensional sub-spaces for modeling in-plane deformation (such as variation in garment shape and size) and out-of-plane deformation (such as deformation due to varied body size and motion). Therefore, we formulate garment deformation as a PLBS model controlled by canonical 3D garment mesh, vertex-based skinning weights and associated local patch transformation. Unlike traditional LBS models specialized for individual objects, PLBS model is capable of uniformly expressing varied garments and bodies, the in-plane deformation is encoded on the canonical 3D garment and the out-of-plane deformation is controlled by the local patch transformation. Besides, we propose novel 3D garment registration and skinning weight decomposition strategies to obtain adequate data to build PLBS model under different garment categories. Furthermore, we employ dynamic fine-tuning to complement high-frequency signals missing from LBS for unseen testing data. Experiments illustrate that our method is capable of modeling dynamics for loose-fitting garments, outperforming previous data-driven modeling methods using different sub-space modeling strategies. We showcase that our method can factorize and be generalized for varied body sizes, garment shapes, garment sizes and human motions under different garment categories. Xipeng Chen, Guangrun Wang, Xiaogang Xu 0002, Philip Torr 0001, Liang Lin 0004 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | S2WAT: Image Style Transfer via Hierarchical Vision Transformer Using Strips Window AttentionabstractTransformer's recent integration into style transfer leverages its proficiency in establishing long-range dependencies, albeit at the expense of attenuated local modeling. This paper introduces Strips Window Attention Transformer (S2WAT), a novel hierarchical vision transformer designed for style transfer. S2WAT employs attention computation in diverse window shapes to capture both short- and long-range dependencies. The merged dependencies utilize the "Attn Merge" strategy, which adaptively determines spatial weights based on their relevance to the target. Extensive experiments on representative datasets show the proposed method's effectiveness compared to state-of-the-art (SOTA) transformer-based and other approaches. The code and pre-trained models are available at https://github.com/AlienZhang1996/S2WAT. Chiyu Zhang 0001, Xiaogang Xu 0002, Zaiyan Dai, Jun Yang 0025 |
AAAI | 2 |
| 2024 | Learning to Remove Wrinkled Transparent Film with Polarized PriorabstractIn this paper, we study a new problem, Film Removal (FR), which attempts to remove the interference of wrinkled transparent films and reconstruct the original information under films for industrial recognition systems. We first physically model the imaging of industrial materials covered by the film. Considering the specular highlight from the film can be effectively recorded by the polarized camera, we build a practical dataset with polarization information containing paired data with and without transparent film. We aim to remove interference from the film (specular highlights and other degradations) with an end-to-end framework. To locate the specular highlight, we use an angle estimation network to optimize the polarization angle with the minimized specular highlight. The image with minimized specular highlight is set as a prior for supporting the reconstruction network. Based on the prior and the polarized images, the reconstruction network can decouple all degradations from the film. Extensive experiments show that our framework achieves SOTA performance in both image reconstruction and industrial downstream tasks. Our code will be released at https://github.com/jqtangust/FilmRemoval. Jiaqi Tang 0005, Ruizheng Wu, Xiaogang Xu 0002, Sixing Hu, Ying-Cong Chen |
CVPR | 3 |
| 2024 | UniMODE: Unified Monocular 3D Object DetectionabstractRealizing unified monocular 3D object detection, including both indoor and outdoor scenes, holds great importance in applications like robot navigation. However, involving various scenarios of data to train models poses challenges due to their significantly different characteristics, e.g., di-verse geometry properties and heterogeneous domain distributions. To address these challenges, we build a detector based on the bird's-eye-view (BEV) detection paradigm, where the explicit feature projection is beneficial to ad-dressing the geometry learning ambiguity when employing multiple scenarios of data to train detectors. Then, we split the classical BEV detection architecture into two stages and propose an uneven BEV grid design to handle the convergence instability caused by the aforementioned challenges. Moreover, we develop a sparse BEV feature projection strategy to reduce computational cost and a unified do-main alignment method to handle heterogeneous domains. Combining these techniques, a unified detector UniMODE is derived, which surpasses the previous state-of-the-art on the challenging Omni3D dataset (a large-scale dataset including both indoor and outdoor scenes) by 4.9% AP3D, revealing the first successful generalization of a BEV detector to unified 3D object detection. Zhuoling Li, Xiaogang Xu 0002, Ser-Nam Lim, Hengshuang Zhao |
CVPR | 2 |
| 2024 | LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score MatchingabstractThe recent advancements in text-to-3D generation mark a significant milestone in generative models, unlocking new possibilities for creating imaginative 3D assets across var-ious real-world scenarios. While recent advancements in text-to-3D generation have shown promise, they often fall short in rendering detailed and high-quality 3D models. This problem is especially prevalent as many methods base themselves on Score Distillation Sampling (SDS). This paper identifies a notable deficiency in SDS, that it brings inconsistent and low-quality updating direction for the 3D model, causing the over-smoothing effect. To address this, we propose a novel approach called Interval Score Matching (ISM). ISM employs deterministic diffusing trajectories and utilizes interval-based score matching to counteract over-smoothing. Furthermore, we incorporate 3D Gaussian Splatting into our text-to-3D generation pipeline. Extensive experiments show that our model largely outperforms the state-of-the-art in quality and training efficiency. Our code is available at: EnVision-Research/LucidDreamer Yixun Liang, Xin Yang 0020, Jiantao Lin, Xiaogang Xu 0002, Ying-Cong Chen |
CVPR | 5 |
| 2024 | Boosting Image Restoration via Priors from Pre-Trained ModelsabstractPre-trained models with large-scale training data, such as CLIP and Stable Diffusion, have demonstrated remarkable performance in various high-level computer vision tasks such as image understanding and generation from language descriptions. Yet, their potential for low-level tasks such as image restoration remains relatively unexplored. In this paper, we explore such models to enhance image restoration. As off-the-shelf features (OSF) from pre-trained models do not directly serve image restoration, we propose to learn an additional lightweight module called Pre-Train-Guided Refinement Module (PTG-RM) to refine restoration results of a target restoration network with OSF. PTG-RM consists of two components, Pre-Train-Guided Spatial-Varying Enhancement (PTG-SVE), and Pre-Train-Guided Channel-Spatial Attention (PTG-CSA). PTG-SVE enables optimal short- and long-range neural operations, while PTG-CSA enhances spatial-channel attention for restoration-related learning. Extensive experiments demonstrate that PTG-RM, with its compact size (<1M parameters), effectively enhances restoration performance of various models across different tasks, including low-light enhancement, deraining, deblurring, and denoising. Xiaogang Xu 0002, Shu Kong, Tao Hu 0011, Zhe Liu 0001, Hujun Bao |
CVPR | 1 |
| 2024 | Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataabstractThis work presents Depth Anything11While the grammatical soundness of this name may be questionable, we treat it as a whole and pay homage to Segment Anything [26]., a highly practical solution for robust monocular depth estimation. Without pursuing novel technical modules, we aim to build a simple yet powerful foundation model dealing with any images under any circumstances. To this end, we scale up the dataset by designing a data engine to collect and automatically annotate large-scale unlabeled data (~62M), which significantly enlarges the data coverage and thus is able to reduce the generalization error. We investigate two simple yet effective strategies that make data scaling-up promising. First, a more challenging optimization target is created by leveraging data augmentation tools. It compels the model to actively seek extra visual knowledge and acquire robust representations. Second, an auxiliary supervision is developed to enforce the model to inherit rich semantic priors from pre-trained encoders. We evaluate its zero-shot capabilities extensively, including six public datasets and randomly captured photos. It demonstrates impressive generalization ability (Figure 1). Further, through fine-tuning it with metric depth information from NYUv2 and KITTI, new SOTAs are set. Our better depth model also results in a better depth-conditioned ControlNet. Our models are released here. Lihe Yang, Bingyi Kang, Xiaogang Xu 0002, Jiashi Feng, Hengshuang Zhao |
CVPR | 4 |
| 2024 | An Incremental Unified Framework for Small Defect Inspection
Jiaqi Tang 0005, Hao Lu 0009, Xiaogang Xu 0002, Ruizheng Wu, Sixing Hu, Tong Zhang 0001, Tsz Wa Cheng, Ming Ge, Ying-Cong Chen, Fugee Tsung |
ECCV (31) | 3 |
| 2024 | Refine, Discriminate and Align: Stealing Encoders via Sample-Wise Prototypes and Multi-relational Extraction
Shuchi Wu, Chuan Ma 0001, Kang Wei 0004, Xiaogang Xu 0002, Ming Ding 0001, Yuwen Qian, Di Xiao 0001, Tao Xiang 0001 |
ECCV (34) | 4 |
| 2024 | Unveiling Advanced Frequency Disentanglement Paradigm for Low-Light Image Enhancement
Kun Zhou 0001, Wenbo Li 0002, Xiaogang Xu 0002, Yuanhao Cai, Zhonghang Liu, Xiaoguang Han 0001, Jiangbo Lu |
ECCV (7) | 4 |
| 2024 | HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty DecodingabstractLarge Vision-Language Models (LVLMs) have shown remarkable performance on many visuallanguage tasks.However, these models still suffer from multimodal hallucination, which means the generation of objects or content that violates the images.Many existing work detects hallucination by directly judging whether an object exists in an image, overlooking the association between the object and semantics.To address this issue, we propose Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding (HELPD).This framework incorporates hallucination feedback at both object and sentence semantic levels.Remarkably, even with a marginal degree of training, this approach can alleviate over 15% of hallucination.Simultaneously, HELPD penalizes the output logits according to the image attention window to avoid being overly affected by generated text.HELPD can be seamlessly integrated with any LVLMs.Our experiments demonstrate that the proposed framework yields favorable results across multiple hallucination benchmarks.It effectively mitigates hallucination for different LVLMs and concurrently improves their text generation quality. Chi Qin, Xiaogang Xu 0002, Piji Li |
EMNLP | 3 |
| 2024 | Generative Active Learning for Long-tailed Instance SegmentationabstractRecently, large-scale language-image generative models have gained widespread attention and many works have utilized generated data from these models to further enhance the performance of perception tasks. However, not all generated data can positively impact downstream models, and these methods do not thoroughly explore how to better select and utilize generated data. On the other hand, there is still a lack of research oriented towards active learning on generated data. In this paper, we explore how to perform active learning specifically for generated data in the long-tailed instance segmentation task. Subsequently, we propose BSGAL, a new algorithm that estimates the contribution of the current batch-generated data based on gradient cache. BSGAL is meticulously designed to cater for unlimited generated data and complex downstream segmentation tasks. BSGAL outperforms the baseline approach and effectually improves the performance of long-tailed segmentation. Muzhi Zhu, Chengxiang Fan, Hao Chen 0041, Yang Liu 0357, Weian Mao, Xiaogang Xu 0002, Chunhua Shen |
ICML | 6 |
| 2024 | HAWK: Learning to Understand Open-World Video AnomaliesabstractVideo Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial semantic understanding of scenes and minimal user interaction. Additionally, the prevalent data scarcity in existing datasets restricts their applicability in open-world scenarios.
In this paper, we introduce HAWK, a novel framework that leverages interactive large Visual Language Models (VLM) to interpret video anomalies precisely. Recognizing the difference in motion information between abnormal and normal videos, HAWK explicitly integrates motion modality to enhance anomaly identification. To reinforce motion attention, we construct an auxiliary consistency loss within the motion and video space, guiding the video branch to focus on the motion modality. Moreover, to improve the interpretation of motion-to-language, we establish a clear supervisory relationship between motion and its linguistic representation. Furthermore, we have annotated over 8,000 anomaly videos with language descriptions, enabling effective training across diverse open-world scenarios, and also created 8,000 question-answering pairs for users' open-world questions. The final results demonstrate that HAWK achieves SOTA performance, surpassing existing baselines in both video description generation and question-answering. Our codes/dataset/demo will be released at https://github.com/jqtangust/hawk. Jiaqi Tang 0005, Hao Lu 0009, Ruizheng Wu, Xiaogang Xu 0002, Bin Guo 0001, Jiangbo Lu, Qifeng Chen 0001, Ying-Cong Chen |
NeurIPS | 4 |
| 2024 | Depth Anything V2abstractThis work presents Depth Anything V2. Without pursuing fancy techniques, we aim to reveal crucial findings to pave the way towards building a powerful monocular depth estimation model. Notably, compared with V1, this version produces much finer and more robust depth predictions through three key practices: 1) replacing all labeled real images with synthetic images, 2) scaling up the capacity of our teacher model, and 3) teaching student models via the bridge of large-scale pseudo-labeled real images. Compared with the latest models built on Stable Diffusion, our models are significantly more efficient (more than 10x faster) and more accurate. We offer models of different scales (ranging from 25M to 1.3B params) to support extensive scenarios. Benefiting from their strong generalization capability, we fine-tune them with metric depth labels to obtain our metric depth models. In addition to our models, considering the limited diversity and frequent noise in current test sets, we construct a versatile evaluation benchmark with sparse depth annotations to facilitate future research. Models are available at https://github.com/DepthAnything/Depth-Anything-V2. Lihe Yang, Bingyi Kang, Zhen Zhao 0001, Xiaogang Xu 0002, Jiashi Feng, Hengshuang Zhao |
NeurIPS | 5 |
| 2024 | MalGNE: Enhancing the Performance and Efficiency of CFG-Based Malware Detector by Graph Node Embedding in Low Dimension SpaceabstractThe rich semantic information in Control Flow Graphs (CFGs) of executable programs has made Graph Neural Networks (GNNs) a key focus for malware detection. However, existing CFG-based detection techniques face limitations in node feature extraction, such as information loss, neglect of execution sequence information, and redundancy in representation vectors. These limitations compromise the balance between high efficiency and precision when training detectors. Addressing this, we introduce an innovative Malware CFG Node Embedding (MalGNE) method. This approach utilizes a novel instruction encoding rule to address the Out-Of-Vocabulary(OOV) problem, generates high-quality initial vectors. Then, it employs aggregation layer and sequence layer to extract node aggregation feature and execution sequence feature, in conjunction with GNNs to develop a pre-trained node embedding model. The model maps the semantic information of node assembly instruction sequences into a compact, low-dimensional continuous space, ensuring high-quality feature extraction, and enhancing the performance and efficiency of the detector. We trained the MalGNE model using the BIG 2015 dataset and validated MalGNE-enhanced detector on the SOREL-20M and BODMAS datasets. MalGNE-enhanced detector demonstrates outstanding performance and efficiency in low-dimensional spaces, especially when the dimensionality of the node feature vector is reduced to 16. MalGNE-enhanced detector not only maintains a high detection accuracy of 95.49%. sacrificing only about 1.7% of accuracy to save approximately 73% of training time compared to 128 dimensions. Hao Peng 0002, Jieshuai Yang, Dandan Zhao 0003, Xiaogang Xu 0002, Yuwen Pu, Jianmin Han, Xing Yang 0004, Ming Zhong 0009, Shouling Ji |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | Deep Parametric 3D Filters for Joint Video Denoising and Illumination Enhancement in Video Super ResolutionabstractDespite the quality improvement brought by the recent methods, video super-resolution (SR) is still very challenging, especially for videos that are low-light and noisy. The current best solution is to subsequently employ best models of video SR, denoising, and illumination enhancement, but doing so often lowers the image quality, due to the inconsistency between the models. This paper presents a new parametric representation called the Deep Parametric 3D Filters (DP3DF), which incorporates local spatiotemporal information to enable simultaneous denoising, illumination enhancement, and SR efficiently in a single encoder-and-decoder network. Also, a dynamic residual frame is jointly learned with the DP3DF via a shared backbone to further boost the SR quality. We performed extensive experiments, including a large-scale user study, to show our method's effectiveness. Our method consistently surpasses the best state-of-the-art methods on all the challenging real datasets with top PSNR and user ratings, yet having a very fast run time. The code is available at https://github.com/xiaogang00/DP3DF. Xiaogang Xu 0002, Ruixing Wang, Chi-Wing Fu, Jiaya Jia |
AAAI | 1 |
| 2023 | Point2Pix: Photo-Realistic Point Cloud Rendering via Neural Radiance FieldsabstractSynthesizing photo-realistic images from a point cloud is challenging because of the sparsity of point cloud representation. Recent Neural Radiance Fields and extensions are proposed to synthesize realistic images from 2D input. In this paper, we present Point2Pix as a novel point renderer to link the 3D sparse point clouds with 2D dense image pixels. Taking advantage of the point cloud 3D prior and NeRF rendering pipeline, our method can synthesize high-quality images from colored point clouds, generally for novel indoor scenes. To improve the efficiency of ray sampling, we propose point-guided sampling, which focuses on valid samples. Also, we present Point Encoding to build Multiscale Radiance Fields that provide discriminative 3D point features. Finally, we propose Fusion Encoding to efficiently synthesize high-quality images. Extensive experiments on the ScanNet and ArkitScenes datasets demonstrate the effectiveness and generalization. Tao Hu 0011, Xiaogang Xu 0002, Shu Liu 0005, Jiaya Jia |
CVPR | 2 |
| 2023 | TriVol: Point Cloud Rendering via Triple VolumesabstractExisting learning-based methods for point cloud rendering adopt various 3D representations and feature querying mechanisms to alleviate the sparsity problem of point clouds. However, artifacts still appear in rendered images, due to the challenges in extracting continuous and discriminative 3D features from point clouds. In this paper, we present a dense while lightweight 3D representation, named TriVol, that can be combined with NeRF to render photo-realistic images from point clouds. Our TriVol consists of triple slim volumes, each of which is encoded from the point cloud. TriVol has two advantages. First, it fuses respective fields at different scales and thus extracts local and non-local features for discriminative representation. Second, since the volume size is greatly reduced, our 3D decoder can be efficiently inferred, allowing us to increase the resolution of the 3D space to render more point details. Extensive experiments on different benchmarks with varying kinds of scenes/objects demonstrate our framework's effectiveness compared with current approaches. Moreover, our framework has excellent generalization ability to render a category of scenes/objects without fine-tuning. The source code is available at https://github.com/dvlabresearch/TriVol.git. Tao Hu 0011, Xiaogang Xu 0002, Ruihang Chu, Jiaya Jia |
CVPR | 2 |
| 2023 | Low-Light Image Enhancement via Structure Modeling and GuidanceabstractThis paper proposes a new framework for low-light image enhancement by simultaneously conducting the appearance as well as structure modeling. It employs the structural feature to guide the appearance enhancement, leading to sharp and realistic results. The structure modeling in our framework is implemented as the edge detection in low-light images. It is achieved with a modified generative model via designing a structure-aware feature extractor and generator. The detected edge maps can accurately emphasize the essential structural information, and the edge prediction is robust towards the noises in dark areas. Moreover, to improve the appearance modeling, which is implemented with a simple U-Net, a novel structure-guided enhancement module is proposed with structure-guided feature synthesis layers. The appearance modeling, edge detector, and enhancement module can be trained end-to-end. The experiments are conducted on representative datasets (sRGB and RAW domains), showing that our model consistently achieves SOTA performance on all datasets with the same architecture. The code is available at https://github.com/xiaogangOO/SMG-LLIE. Xiaogang Xu 0002, Ruixing Wang, Jiangbo Lu |
CVPR | 1 |
| 2023 | High Dynamic Range Image Reconstruction via Deep Explicit Polynomial Curve EstimationabstractDue to limited camera capacities, digital images usually have a narrower dynamic illumination range than real-world scene radiance. To resolve this problem, High Dynamic Range (HDR) reconstruction is proposed to recover the dynamic range to better represent real-world scenes. However, due to different physical imaging parameters, the tone-mapping functions between images and real radiance are highly diverse, which makes HDR reconstruction extremely challenging. Existing solutions can not explicitly clarify a corresponding relationship between the tone-mapping function and the generated HDR image, but this relationship is vital when guiding the reconstruction of HDR images. To address this problem, we propose a method to explicitly estimate the tone mapping function and its corresponding HDR image in one network. Firstly, based on the characteristics of the tone mapping function, we construct a model by a polynomial to describe the trend of the tone curve. To fit this curve, we use a learnable network to estimate the coefficients of the polynomial. This curve will be automatically adjusted according to the tone space of the Low Dynamic Range (LDR) image, and reconstruct the real HDR image. Besides, since all current datasets do not provide the corresponding relationship between the tone mapping function and the LDR image, we construct a new dataset with both synthetic and real images. Extensive experiments show that our method generalizes well under different tone-mapping functions and achieves SOTA performance. The code/dataset is available at https://github.com/jqtangust/EPCE-HDR.git. Jiaqi Tang 0005, Xiaogang Xu 0002, Sixing Hu, Ying-Cong Chen |
ECAI | 2 |
| 2023 | Lighting up NeRF via Unsupervised Decomposition and EnhancementabstractNeural Radiance Field (NeRF) is a promising approach for synthesizing novel views, given a set of images and the corresponding camera poses of a scene. However, images photographed from a low-light scene can hardly be used to train a NeRF model to produce high-quality results, due to their low pixel intensities, heavy noise, and color distortion. Combining existing low-light image enhancement methods with NeRF methods also does not work well due to the view inconsistency caused by the individual 2D enhancement process. In this paper, we propose a novel approach, called Low-Light NeRF (or LLNeRF), to enhance the scene representation and synthesize normal-light novel views directly from sRGB low-light images in an unsupervised manner. The core of our approach is a decomposition of radiance field learning, which allows us to enhance the illumination, reduce noise and correct the distorted colors jointly with the NeRF optimization process. Our method is able to produce novel view images with proper lighting and vivid colors and details, given a collection of camera-finished low dynamic range (8-bits/channel) images from a low-light scene. Experiments demonstrate that our method outperforms existing low-light enhancement methods and NeRF methods. Xiaogang Xu 0002, Ke Xu 0010, Rynson W. H. Lau |
ICCV | 2 |
| 2023 | Out-of-domain GAN inversion via Invertibility Decomposition for Photo-Realistic Human Face ManipulationabstractThe fidelity of Generative Adversarial Networks (GAN) inversion is impeded by Out-Of-Domain (OOD) areas (e.g., background, accessories) in the image. Detecting the OOD areas beyond the generation ability of the pre-trained model and blending these regions with the input image can enhance fidelity. The "invertibility mask" figures out these OOD areas, and existing methods predict the mask with the reconstruction error. However, the estimated mask is usually inaccurate due to the influence of the reconstruction error in the In-Domain (ID) area. In this paper, we propose a novel framework that enhances the fidelity of human face in-version by designing a new module to decompose the input images to ID and OOD partitions with invertibility masks. Unlike previous works, our invertibility detector is simultaneously learned with a spatial alignment module. We iteratively align the generated features to the input geometry and reduce the reconstruction error in the ID regions. Thus, the OOD areas are more distinguishable and can be precisely predicted. Then, we improve the fidelity of our results by blending the OOD areas from the input image with the ID GAN inversion results. Our method produces photorealistic results for real-world human face image inversion and manipulation. Extensive experiments demonstrate our method’s superiority over existing methods in the quality of GAN inversion and attribute manipulation. Our code is available at: AbnerVictor/OOD-GAN-inversion Xin Yang 0020, Xiaogang Xu 0002, Ying-Cong Chen |
ICCV | 2 |
| 2023 | Universal Adaptive Data AugmentationabstractExisting automatic data augmentation (DA) methods either ignore updating DA's parameters according to the target model's state during training or adopt update strategies that are not effective enough. In this work, we design a novel data augmentation strategy called ``Universal Adaptive Data Augmentation" (UADA). Different from existing methods, UADA would adaptively update DA's parameters according to the target model's gradient information during training: given a pre-defined set of DA operations, we randomly decide types and magnitudes of DA operations for every data batch during training, and adaptively update DA's parameters along the gradient direction of the loss concerning DA's parameters. In this way, UADA can increase the training loss of the target networks, and the target networks would learn features from harder samples to improve the generalization. Moreover, UADA is very general and can be utilized in numerous tasks, e.g., image classification, semantic segmentation and object detection. Extensive experiments with various models are conducted on CIFAR-10, CIFAR-100, ImageNet, tiny-ImageNet, Cityscapes, and VOC07+12 to prove the significant performance improvements brought by UADA. Xiaogang Xu 0002, Hengshuang Zhao |
IJCAI | 1 |
| 2023 | CorresNeRF: Image Correspondence Priors for Neural Radiance FieldsabstractNeural Radiance Fields (NeRFs) have achieved impressive results in novel view synthesis and surface reconstruction tasks. However, their performance suffers under challenging scenarios with sparse input views. We present CorresNeRF, a novel method that leverages image correspondence priors computed by off-the-shelf methods to supervise NeRF training. We design adaptive processes for augmentation and filtering to generate dense and high-quality correspondences. The correspondences are then used to regularize NeRF training via the correspondence pixel reprojection and depth loss terms. We evaluate our methods on novel view synthesis and surface reconstruction tasks with density-based and SDF-based NeRF models on different datasets. Our method outperforms previous methods in both photometric and geometric metrics. We show that this simple yet effective technique of using correspondence priors can be applied as a plug-and-play module across different NeRF variants. The project page is at https://yxlao.github.io/corres-nerf/. Yixing Lao, Xiaogang Xu 0002, Xihui Liu, Hengshuang Zhao |
NeurIPS | 2 |
| 2023 | FreeMask: Synthetic Images with Dense Annotations Make Stronger Segmentation ModelsabstractSemantic segmentation has witnessed tremendous progress due to the proposal of various advanced network architectures. However, they are extremely hungry for delicate annotations to train, and the acquisition is laborious and unaffordable. Therefore, we present FreeMask in this work, which resorts to synthetic images from generative models to ease the burden of both data collection and annotation procedures. Concretely, we first synthesize abundant training images conditioned on the semantic masks provided by realistic datasets. This yields extra well-aligned image-mask training pairs for semantic segmentation models. We surprisingly observe that, solely trained with synthetic images, we already achieve comparable performance with real ones (e.g., 48.3 vs. 48.5 mIoU on ADE20K, and 49.3 vs. 50.5 on COCO-Stuff). Then, we investigate the role of synthetic images by joint training with real images, or pre-training for real images. Meantime, we design a robust filtering principle to suppress incorrectly synthesized regions. In addition, we propose to inequally treat different semantic masks to prioritize those harder ones and sample more corresponding synthetic images for them. As a result, either jointly trained or pre-trained with our filtered and re-sampled synthesized images, segmentation models can be greatly enhanced, e.g., from 48.7 to 52.0 on ADE20K. Lihe Yang, Xiaogang Xu 0002, Bingyi Kang, Yinghuan Shi, Hengshuang Zhao |
NeurIPS | 2 |
| 2023 | Conditional Temporal Variational AutoEncoder for Action Video Prediction
Xiaogang Xu 0002, Yi Wang 0074, Liwei Wang 0009, Bei Yu 0001, Jiaya Jia |
Int. J. Comput. Vis. | 1 |
| 2022 | Hierarchical Image Generation via Transformer-Based Sequential Patch SelectionabstractTo synthesize images with preferred objects and interactions, a controllable way is to generate the image from a scene graph and a large pool of object crops, where the spatial arrangements of the objects in the image are defined by the scene graph while their appearances are determined by the retrieved crops from the pool. In this paper, we propose a novel framework with such a semi-parametric generation strategy. First, to encourage the retrieval of mutually compatible crops, we design a sequential selection strategy where the crop selection for each object is determined by the contents and locations of all object crops that have been chosen previously. Such process is implemented via a transformer trained with contrastive losses. Second, to generate the final image, our hierarchical generation strategy leverages hierarchical gated convolutions which are employed to synthesize areas not covered by any image crops, and a patch guided spatially adaptive normalization module which is proposed to guarantee the final generated images complying with the crop appearance and the scene graph. Evaluated on the challenging Visual Genome and COCO-Stuff dataset, our experimental results demonstrate the superiority of our proposed method over existing state-of-the-art methods. Xiaogang Xu 0002 |
AAAI | 1 |
| 2022 | SNR-Aware Low-light Image EnhancementabstractThis paper presents a new solution for low-light image enhancement by collectively exploiting Signal-to-Noise-Ratio-aware transformers and convolutional models to dynamically enhance pixels with spatial-varying operations. They are long-range operations for image regions of extremely low Signal-to-Noise-Ratio (SNR) and short-range operations for other regions. We propose to take an SNR prior to guide the feature fusion and formulate the SNR-aware transformer with a new self-attention model to avoid tokens from noisy image regions of very low SNR. Extensive experiments show that our framework consistently achieves better performance than SOTA approaches on seven representative benchmarks with the same structure. Also, we conducted a large-scale user study with 100 participants to verify the superior perceptual quality of our results. The code is available at https://github.com/dvlab-research/SNR-Aware-Low-Light-Enhance. Xiaogang Xu 0002, Ruixing Wang, Chi-Wing Fu, Jiaya Jia |
CVPR | 1 |
| 2022 | DecoupleNet: Decoupled Network for Domain Adaptive Semantic Segmentation
Zhuotao Tian, Xiaogang Xu 0002, Ying-Cong Chen, Shu Liu 0005, Hengshuang Zhao, Liwei Wang 0009, Jiaya Jia |
ECCV (33) | 3 |
| 2022 | MTFormer: Multi-task Learning via Transformer and Cross-Task Reasoning
Xiaogang Xu 0002, Hengshuang Zhao, Vibhav Vineet, Ser-Nam Lim, Antonio Torralba 0001 |
ECCV (27) | 1 |
| 2022 | Text-Guided Human Image Manipulation via Image-Text Shared SpaceabstractText is a new way to guide human image manipulation. Albeit natural and flexible, text usually suffers from inaccuracy in spatial description, ambiguity in the description of appearance, and incompleteness. We in this paper address these issues. To overcome inaccuracy, we use structured information (e.g., poses) to help identify correct location to manipulate, by disentangling the control of appearance and spatial structure. Moreover, we learn the image-text shared space with derived disentanglement to improve accuracy and quality of manipulation, by separating relevant and irrelevant editing directions for the textual instructions in this space. Our model generates a series of manipulation results by moving source images in this space with different degrees of editing strength. Thus, to reduce the ambiguity in text, our model generates sequential output for manual selection. In addition, we propose an efficient pseudo-label loss to enhance editing performance when the text is incomplete. We evaluate our method on various datasets and show its precision and interactiveness to manipulate human images. Xiaogang Xu 0002, Ying-Cong Chen, Xin Tao 0001, Jiaya Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Adversarial CAPTCHAsabstractFollowing the principle of to set one's own spear against one's own shield, we study how to design adversarial completely automated public turing test to tell computers and humans apart (CAPTCHA) in this article. We first identify the similarity and difference between adversarial CAPTCHA generation and existing hot adversarial example (image) generation research. Then, we propose a framework for text-based and image-based adversarial CAPTCHA generation on top of state-of-the-art adversarial image generation techniques. Finally, we design and implement an adversarial CAPTCHA generation and evaluation system, called aCAPTCHA, which integrates 12 image preprocessing techniques, nine CAPTCHA attacks, four baseline adversarial CAPTCHA generation methods, and eight new adversarial CAPTCHA generation methods. To examine the performance of aCAPTCHA, extensive security and usability evaluations are conducted. The results demonstrate that the generated adversarial CAPTCHAs can significantly improve the security of normal CAPTCHAs while maintaining similar usability. To facilitate the CAPTCHA security research, we also open source the aCAPTCHA system, including the source code, trained models, datasets, and the usability evaluation interfaces. Chenghui Shi, Xiaogang Xu 0002, Shouling Ji, Kai Bu, Jianhai Chen, Raheem A. Beyah, Ting Wang 0006 |
IEEE Trans. Cybern. | 2 |
| 2021 | Self-Supervised 3D Mesh Reconstruction From Single ImagesabstractRecent single-view 3D reconstruction methods reconstruct object’s shape and texture from a single image with only 2D image-level annotation. However, without explicit 3D attribute-level supervision, it is still difficult to achieve satisfying reconstruction accuracy. In this paper, we propose a Self-supervised Mesh Reconstruction (SMR) approach to enhance 3D mesh attribute learning process. Our approach is motivated by observations that (1) 3D attributes from interpolation and prediction should be consistent, and (2) feature representation of landmarks from all images should be consistent. By only requiring silhouette mask annotation, our SMR can be trained in an end-to- end manner and generalizes to reconstruct natural objects of birds, cows, motorbikes, etc. Experiments demonstrate that our approach improves both 2D supervised and unsupervised 3D mesh reconstruction on multiple datasets. We also show that our model can be adapted to other image synthesis tasks, e.g., novel view generation, shape transfer, and texture transfer, with promising results. Our code is publicly available at https://github.com/Jia-Research-Lab. Tao Hu 0011, Liwei Wang 0009, Xiaogang Xu 0002, Shu Liu 0005, Jiaya Jia |
CVPR | 3 |
| 2021 | Seeing Dynamic Scene in the Dark: A High-Quality Video Dataset with Mechatronic AlignmentabstractLow-light video enhancement is an important task. Previous work is mostly trained on paired static images or videos. We compile a new dataset formed by our new strategy that contains high-quality spatially-aligned video pairs from dynamic scenes in low- and normal-light conditions. We built it using a mechatronic system to precisely control the dynamics during the video capture process, and further align the video pairs, both spatially and temporally, by identifying the system’s uniform motion stage. Besides the dataset, we propose an end-to-end framework, in which we design a self-supervised strategy to reduce noise, while enhancing the illumination based on the Retinex theory. Extensive experiments based on various metrics and large-scale user study demonstrate the value of our dataset and effectiveness of our method. The dataset and code are available at https://github.com/dvlab-research/SDSD. Ruixing Wang, Xiaogang Xu 0002, Chi-Wing Fu, Jiangbo Lu, Bei Yu 0001, Jiaya Jia |
ICCV | 2 |
| 2021 | Dynamic Divide-and-Conquer Adversarial Training for Robust Semantic SegmentationabstractAdversarial training is promising for improving robustness of deep neural networks towards adversarial perturbations, especially on the classification task. The effect of this type of training on semantic segmentation, contrarily, just commences. We make the initial attempt to explore the defense strategy on semantic segmentation by formulating a general adversarial training procedure that can per-form decently on both adversarial and clean samples. We propose a dynamic divide-and-conquer adversarial training (DDC-AT) strategy to enhance the defense effect, by set-ting additional branches in the target model during training, and dealing with pixels with diverse properties to-wards adversarial perturbation. Our dynamical division mechanism divides pixels into multiple branches automatically. Note all these additional branches can be abandoned during inference and thus leave no extra parameter and computation cost. Extensive experiments with various segmentation models are conducted on PASCAL VOC 2012 and Cityscapes datasets, in which DDC-AT yields satisfying performance under both white- and black-box at-tack. The code is available at https://github.com/dvlab-research/Robust-Semantic-Segmentation. Xiaogang Xu 0002, Hengshuang Zhao, Jiaya Jia |
ICCV | 1 |
| 2021 | Reference-Based Video Colorization With Multi-Scale Semantic Fusion And Temporal AugmentationabstractThe reference-based video colorization method hallucinates a plausible color version for a gray-scale video by referring distributions of possible colors from an input color frame, which has semantic correspondences with the gray-scale frames. The plausibility of colors and the temporal consistency are two significant challenges in this task. In this paper, we propose a novel Generative Adversarial Network (GAN) with a siamese training framework to tackle these challenges. Specifically, the siamese training framework allows us to implement temporal feature augmentation, enhancing temporal consistency. Further, to improve the plausibility of colorization results, we propose a multi-scale fusion module that correlates features of reference frames to source frames accurately. Experiments on various datasets demonstrate that our proposed method performs favorably against the state-of the-art approaches. Xiaoyan Zhang 0002, Xiaogang Xu 0002 |
ICIP | 3 |
| 2021 | Semantic-Aware Video Style Transfer Based on Temporal Consistent Sparse Patch ConstraintabstractThis paper proposes a practical style transfer method to synthesize a temporally smooth video whose style information is semantically consistent with the reference video. Due to the lack of paired videos for training, we extend the structure of CycleGAN with sparse patch and temporal constraints, including a new semantic patch loss and a novel temporal loss. Our approach’s key insights are: (1) the semantically paired sparse patches chosen from synthesized videos and reference frames would promote the semantic meaning of style transfer, the preservation of video content, and the smoothness of results by minimizing the discrepancies between these paired patches. (2) the forward and backward temporal consistency among neighbouring frames can reduce the discontinuity in the synthesized video. Extensive quantitative and qualitative experiments on various metrics demonstrate the superiority of our method over state-of-the-art strategies. Xiaoyan Zhang 0002, Xiaogang Xu 0002 |
ICME | 3 |
| 2021 | Ranking Users in Social Networks with Motif-Based PageRankabstractPageRank has been widely used to measure the authority or the influence of a user in social networks. However, conventional PageRank only makes use of edge-based relations, which represent first-order relations between two connected nodes. It ignores higher-order relations that may exist between nodes. In this article, we propose a novel framework, motif-based PageRank (MPR), to incorporate higher-order relations into the conventional PageRank computation. Motifs are subgraphs consisting of a small number of nodes. We use motifs to capture higher-order relations between nodes in a network and introduce two methods, one linear and one non-linear, to combine first-order and higher-order relations in PageRank computation. We conduct extensive experiments on three real-world networks, namely, DBLP, Epinions, and Ciao. We study different types of motifs, including 3-node simple and anchor motifs, 4-node and 5-node motifs. Besides using single motif, we also run MPR with ensemble of multiple motifs. We also design a learning task to evaluate the abilities of authority prediction with motif-based features. All experimental results demonstrate that MPR can significantly improve the performance of user ranking in social networks compared to the baseline methods. Huan Zhao 0002, Xiaogang Xu 0002, Yangqiu Song, Dik Lun Lee |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Domain Adaptive Image-to-Image TranslationabstractUnpaired image-to-image translation (I2I) has achieved great success in various applications. However, its generalization capacity is still an open question. In this paper, we show that existing I2I models do not generalize well for samples outside the training domain. The cause is twofold. First, an I2I model may not work well when testing samples are beyond its valid input domain. Second, results could be unreliable if the expected output is far from what the model is trained. To deal with these issues, we propose the Domain Adaptive Image-To-Image translation (DAI2I) framework that adapts an I2I model for out-of-domain samples. Our framework introduces two sub-modules -- one maps testing samples to the valid input domain of the I2I model, and the other transforms the output of I2I model to expected results. Extensive experiments manifest that our framework improves the capacity of existing I2I models, allowing them to handle samples that are distinctively different from their primary targets. Ying-Cong Chen, Xiaogang Xu 0002, Jiaya Jia |
CVPR | 2 |
| 2019 | Homomorphic Latent Space Interpolation for Unpaired Image-To-Image TranslationabstractGenerative adversarial networks have achieved great success in unpaired image-to-image translation. Cycle consistency allows modeling the relationship between two distinct domains without paired data. In this paper, we propose an alternative framework, as an extension of latent space interpolation, to consider the intermediate region between two domains during translation. It is based on the fact that in a flat and smooth latent space, there exist many paths that connect two sample points. Properly selecting paths makes it possible to change only certain image attributes, which is useful for generating intermediate images between the two domains. We also show that this framework can be applied to multi-domain and multi-modal translation. Extensive experiments manifest its generality and applicability to various tasks. Ying-Cong Chen, Xiaogang Xu 0002, Zhuotao Tian, Jiaya Jia |
CVPR | 2 |
| 2019 | View Independent Generative Adversarial Network for Novel View SynthesisabstractSynthesizing novel views from a 2D image requires to infer 3D structure and project it back to 2D from a new viewpoint. In this paper, we propose an encoder-decoder based generative adversarial network VI-GAN to tackle this problem. Our method is to let the network, after seeing many images of objects belonging to the same category in different views, obtain essential knowledge of intrinsic properties of the objects. To this end, an encoder is designed to extract view-independent feature that characterizes intrinsic properties of the input image, which includes 3D structure, color, texture etc. We also make the decoder hallucinate the image of a novel view based on the extracted feature and an arbitrary user-specific camera pose. Extensive experiments demonstrate that our model can synthesize high-quality images in different views with continuous camera poses, and is general for various applications. Xiaogang Xu 0002, Ying-Cong Chen, Jiaya Jia |
ICCV | 1 |
| 2018 | Ranking Users in Social Networks With Higher-Order StructuresabstractPageRank has been widely used to measure the authority or the influence of a user in social networks. However, conventional PageRank only makes use of edge-based relations, ignoring higher-order structures captured by motifs, subgraphs consisting of a small number of nodes in complex networks. In this paper, we propose a novel framework, motif-based PageRank (MPR), to incorporate higher-order structures into conventional PageRank computation. We conduct extensive experiments in three real-world networks, i.e., DBLP, Epinions, and Ciao, to show that MPR can significantly improve the effectiveness of PageRank for ranking users in social networks. In addition to numerical results, we also provide detailed analysis for MPR to show how and why incorporating higher-order information works better than PageRank in ranking users in social networks. Huan Zhao 0002, Xiaogang Xu 0002, Yangqiu Song, Dik Lun Lee |
AAAI | 2 |