Xin Jin 0014

dblp:68/3340-14 · DBLP profile ↗
← Back
94ranked-venue papers
23as first author
80since 2021 · last 2026
0000-0002-1820-8358ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 67 · 21 first-author · 54 since 2021Artificial intelligence and machine learning · 60 · 13 first-author · 51 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Temporal Inconsistency Guidance for Super-resolution Video Quality Assessment
abstract
As super-resolution (SR) techniques introduce unique distortions that fundamentally differ from those caused by traditional degradation processes (e.g., compression), there is an increasing demand for specialized video quality assessment (VQA) methods tailored to SR-generated content. One critical factor affecting perceived quality is temporal inconsistency, which refers to irregularities between consecutive frames. However, existing VQA approaches rarely quantify this phenomenon or explicitly investigate its relationship with human perception. Moreover, SR videos exhibit amplified inconsistency levels as a result of enhancement processes. In this paper, we propose Temporal Inconsistency Guidance for Super-resolution Video Quality Assessment (TIG-SVQA) that underscores the critical role of temporal inconsistency in guiding the quality assessment of SR videos. We first design a perception-oriented approach to quantify frame-wise temporal inconsistency. Based on this, we introduce the Inconsistency Highlighted Spatial Module, which localizes inconsistent regions at both coarse and fine scales. Inspired by the human visual system, we further develop an Inconsistency Guided Temporal Module that performs progressive temporal feature aggregation: (1) a consistency-aware fusion stage in which a visual memory capacity block adaptively determines the information load of each temporal segment based on inconsistency levels, and (2) an informative filtering stage for emphasizing quality-related features. Extensive experiments on both single-frame and multi-frame SR video scenarios demonstrate that our method significantly outperforms state-of-the-art VQA approaches.
Xiaoyuan Yang 0003, Weide Liu, Xin Jin 0014, Xu Jia 0012, Yukun Lai, Paul L. Rosin, Hantao Liu, Wei Zhou 0021
AAAI4
2026 Efficient Token Compression for the Understanding and Generation Unified MLLMs
Junyan Lin, Jinming Liu 0001, Shengyang Zhao, Xin Jin 0014
ISCAS4
2026 Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
abstract
Classical visual coding and Multimodal Large Language Model (MLLM) token technology share the core objective - maximizing information fidelity while minimizing computational cost. Therefore, this paper reexamines MLLM token technology, including tokenization, token compression, and token reasoning, through the established principles of long-developed visual coding area. From this perspective, we (1) establish a unified formulation bridging token technology and visual coding, enabling a systematic, module-by-module comparative analysis; (2) synthesize bidirectional insights, exploring how visual coding principles can enhance MLLM token techniques' efficiency and robustness, and conversely, how token technology paradigms can inform the design of next-generation semantic visual codecs; (3) prospect for promising future research directions and critical unsolved challenges. In summary, this study presents the first comprehensive and structured technology comparison of MLLM token and visual coding, paving the way for more efficient multimodal models and more powerful visual codecs simultaneously.
Jinming Liu 0001, Junyan Lin, Yuntao Wei, Kele Shao, Keda Tao, Jianguo Huang, Zhibo Chen 0001, Huan Wang 0014, Xin Jin 0014
ISCAS10
2026 Hierarchical Context Alignment With Disentangled Geometric and Temporal Modeling for Semantic Occupancy Prediction
abstract
Camera-based 3D Semantic Occupancy Prediction (SOP) is crucial for understanding complex 3D scenes from limited 2D image observations. Existing SOP methods typically aggregate contextual features to assist the occupancy representation learning, alleviating issues like occlusion or ambiguity. However, these solutions often face misalignment issues wherein the corresponding features at the same position across different frames may have different semantic meanings during the aggregation process, which leads to unreliable contextual fusion results and an unstable representation learning process. To address this problem, we introduce a new Hierarchical context alignment paradigm for a more accurate SOP (Hi-SOP). Hi-SOP first disentangles the geometric and temporal context for separate alignment, which two branches are then composed to enhance the reliability of SOP. This parsing of the visual input into a local-global alignment hierarchy includes: (I) disentangled geometric and temporal separate alignment, within each leverages depth confidence and camera pose as prior for relevant feature matching respectively; (II) global alignment and composition of the transformed geometric and temporal volumes based on semantics consistency. Our method outperforms SOTAs for semantic scene completion on the SemanticKITTI & NuScenes-Occupancy datasets and LiDAR semantic segmentation on the NuScenes dataset.
Bohan Li 0015, Xin Jin 0014, Jiajun Deng, Yasheng Sun, Wenjun Zeng 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 A Survey of Behavior Foundation Model: Next-Generation Whole-Body Control System of Humanoid Robots
abstract
Humanoid robots are drawing significant attention as versatile platforms for complex motor control, human-robot interaction, and general-purpose physical intelligence. However, achieving efficient whole-body control (WBC) in humanoids remains a fundamental challenge due to sophisticated dynamics, underactuation, and diverse task requirements. While learning-based controllers have shown promise for complex tasks, their reliance on labor-intensive and costly retraining for new scenarios limits real-world applicability. To address these limitations, behavior(al) foundation models (BFMs) have emerged as a new paradigm that leverages large-scale pre-training to learn reusable primitive skills and broad behavioral priors, enabling zero-shot or rapid adaptation to a wide range of downstream tasks. In this paper, we present a comprehensive overview of BFMs for humanoid WBC, tracing their development across diverse pre-training pipelines. Furthermore, we discuss real-world applications, current limitations, urgent challenges, and future opportunities, positioning BFMs as a key approach toward scalable and general-purpose humanoid intelligence. Finally, we provide a curated and regularly updated collection of BFM papers and projects to facilitate further research, which is available at https://github.com/yuanmingqi/awesome-bfm-papers.
Mingqi Yuan, Tao Yu 0012, Wenqi Ge, Xiuyong Yao, Huijiang Wang, Jiayu Chen 0006, Bo Li 0037, Wei Zhang 0262, Wenjun Zeng 0001, Hua Chen 0007, Xin Jin 0014
IEEE Trans. Pattern Anal. Mach. Intell.12
2025 RLLTE: Long-Term Evolution Project of Reinforcement Learning
abstract
We present RLLTE: a long-term evolution, extremely modular, and open-source framework for reinforcement learning (RL) research and application. Beyond delivering top-notch algorithm implementations, RLLTE also serves as a toolkit for developing algorithms. More specifically, RLLTE decouples the RL algorithms completely from the exploitation-exploration perspective, providing a large number of components to accelerate algorithm development and evolution. In particular, RLLTE is the first RL framework to build a comprehensive ecosystem, which includes model training, evaluation, deployment, benchmark hub, and large language model (LLM)-empowered copilot. RLLTE is expected to set standards for RL engineering practice and be highly stimulative for industry and academia. Our documentation, examples, and source code are available at https://github.com/RLE-Foundation/rllte.
Mingqi Yuan, Zequn Zhang, Shihao Luo, Bo Li 0037, Xin Jin 0014, Wenjun Zeng 0001
AAAI6
2025 UniMamba: Unified Spatial-Channel Representation Learning with Group-Efficient Mamba for LiDAR-based 3D Object Detection
abstract
Recent advances in LiDAR 3D detection have demonstrated the effectiveness of Transformer-based frameworks in capturing the global dependencies from point cloud spaces, which serialize the 3D voxels into the flattened 1D sequence for iterative self-attention. However, the spatial structure of 3D voxels will be inevitably destroyed during the serialization process. Besides, due to the considerable number of 3D voxels and quadratic complexity of Transformers, multiple sequences are grouped before feeding to Transformers, leading to a limited receptive field. Inspired by the impressive performance of State Space Models (SSM), in this paper, we propose a novel Unified Mamba (UniMamba), which seamlessly integrates the merits of 3D convolution and SSM in a concise multi-head manner, aiming to perform "local and global" spatial context aggregation efficiently and simultaneously. Specifically, a Uni-Mamba block is designed which mainly consists of spatial locality modeling, complementary Z-order serialization and local-global sequential aggregator. The spatial locality modeling module integrates 3D submanifold convolution to capture the dynamic spatial position embedding before serialization. Then the efficient Z-order curve is adopted for serialization both horizontally and vertically. Furthermore, the local-global sequential aggregator adopts the channel grouping strategy to efficiently encode both "local and global" spatial inter-dependencies using multi-head SSM. Additionally, an encoder-decoder architecture with stacked UniMamba blocks is formed to facilitate multi-scale spatial learning hierarchically. Extensive experiments are conducted on three popular datasets: nuScenes, Waymo and Argoverse 2. Particularly, our UniMamba achieves 70.2 mAP on the nuScenes dataset.
Xin Jin 0014, Haisheng Su, Wei Wu 0021, Fei Hui, Junchi Yan
CVPR1
2025 UniScene: Unified Occupancy-centric Driving Scene Generation
abstract
Generating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles to model the direct layout-to-data distribution. In this paper, we introduce UniScene, the first unified framework for generating three key data forms — semantic occupancy, video, and LiDAR — in driving scenes. UniScene employs a progressive generation process that decomposes the complex task of scene generation into two hierarchical steps: (a) first generating semantic occupancy from a customized scene layout as a meta scene representation rich in both semantic and geometric information, and then (b) conditioned on occupancy, generating video and LiDAR data, respectively, with two novel transfer strategies of Gaussian-based Joint Rendering and Prior-guided Sparse Modeling. This occupancy-centric approach reduces the generation burden, especially for intricate scenes, while providing detailed intermediate representations for the subsequent generation stages. Extensive experiments demonstrate that UniScene outperforms previous SOTAs in the occupancy, video, and LiDAR generation, which also indeed benefits downstream driving tasks. The Project is available at https://arlo0o.github.io/uniscene/.
Bohan Li 0015, Jiazhe Guo, Hongsi Liu, Yingshuang Zou, Yikang Ding, Xiwu Chen, Hu Zhu, Feiyang Tan, Tiancai Wang, Shuchang Zhou 0001, Li Zhang 0040, Xiaojuan Qi 0001, Hao Zhao 0002, Mu Yang, Wenjun Zeng 0001, Xin Jin 0014
CVPR17
2025 Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning
abstract
End-to-end autonomous driving unifies tasks in a differentiable framework, enabling planning-oriented optimization and attracting growing attention. Existing methods aggregate historical information either through dense historical bird's-eye-view (BEV) features or by querying a sparse memory bank, following paradigms inherited from detection. We argue that these paradigms either omit historical information in motion planning or fail to align with its multi-step nature, which requires predicting or planning multiple future time steps. In line with the philosophy of "future is a continuation of past", we propose BridgeAD, which reformulates motion and planning queries as multistep queries to differentiate the queries for each future time step. This design enables the effective use of historical prediction and planning by applying them to the appropriate parts of the end-to-end system based on the time steps, which improves both perception and motion planning. Specifically, historical queries for the current frame are combined with perception, while queries for future frames are integrated with motion planning. In this way, we bridge the gap between past and future by aggregating historical insights at every time step, enhancing the overall coherence and accuracy of the end-to-end autonomous driving pipeline. Extensive experiments on the nuScenes dataset in both open-loop and closed-loop settings demonstrate that BridgeAD achieves state-of-the-art performance.
Bozhou Zhang, Nan Song, Xin Jin 0014, Li Zhang 0040
CVPR3
2025 Multimodal Language Models See Better When They Look Shallower
abstract
Multimodal large language models (MLLMs) typically extract visual features from the final layers of a pretrained Vision Transformer (ViT).This widespread deep-layer bias, however, is largely driven by empirical convention rather than principled analysis.While prior studies suggest that different ViT layers capture different types of information-shallower layers focusing on fine visual details and deeper layers aligning more closely with textual semantics, the impact of this variation on MLLM performance remains underexplored.We present the first comprehensive study of visual layer selection for MLLMs, analyzing representation similarity across ViT layers to establish shallow, middle, and deep layer groupings.Through extensive evaluation of MLLMs (1.4B-7B parameters) across 10 benchmarks encompassing 60+ tasks, we find that while deep layers excel in semantic-rich tasks like OCR, shallow and middle layers significantly outperform them on fine-grained visual tasks including counting, positioning, and object localization.Building on these insights, we propose a lightweight feature fusion method that strategically incorporates shallower layers, achieving consistent improvements over both single-layer and specialized fusion baselines.Our work offers the first principled study of visual layer selection in MLLMs, showing that MLLMs can often see better when they look shallower.
Junyan Lin, Xinghao Chen 0009, Jianfeng Dong, Xin Jin 0014, Hui Su, Jinlan Fu, Xiaoyu Shen 0001
EMNLP6
2025 GeoFormer: Geometry Point Encoder for 3D Object Detection with Graph-Based Transformer
Xin Jin 0014, Haisheng Su, Wei Wu 0021, Fei Hui, Junchi Yan
ICCV1
2025 Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning
abstract
Training visual reinforcement learning (RL) in practical scenarios presents a significant challenge, $\textit{i.e.,}$ RL agents suffer from low sample efficiency in environments with variations. While various approaches have attempted to alleviate this issue by disentangled representation learning, these methods usually start learning from scratch without prior knowledge of the world. This paper, in contrast, tries to learn and understand underlying semantic variations from distracting videos via offline-to-online latent distillation and flexible disentanglement constraints. To enable effective cross-domain semantic knowledge transfer, we introduce an interpretable model-based RL framework, dubbed Disentangled World Models (DisWM). Specifically, we pretrain the action-free video prediction model offline with disentanglement regularization to extract semantic knowledge from distracting videos. The disentanglement capability of the pretrained model is then transferred to the world model through latent distillation. For finetuning in the online environment, we exploit the knowledge from the pretrained model and introduce a disentanglement constraint to the world model. During the adaptation phase, the incorporation of actions and rewards from online environment interactions enriches the diversity of the data, which in turn strengthens the disentangled representation learning. Experimental results validate the superiority of our approach on various benchmarks.
Qi Wang 0080, Baao Xie, Xin Jin 0014, Yunbo Wang, Liaomo Zheng, Xiaokang Yang 0001, Wenjun Zeng 0001
ICCV4
2025 Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
Liang Xu 0012, Chengqun Yang, Zili Lin, Fei Xu 0008, Congsheng Xu, Yiyi Zhang 0002, Jie Qin 0004, Xingdong Sheng, Yunhui Liu 0006, Xin Jin 0014, Yichao Yan, Wenjun Zeng 0001, Xiaokang Yang 0001
ICCV11
2025 ULTHO: Ultra-Lightweight Yet Efficient Hyperparameter Optimization in Deep Reinforcement Learning
abstract
Hyperparameter optimization (HPO) is a billion-dollar problem in machine learning, which significantly impacts the training efficiency and model performance. However, achieving efficient and robust HPO in deep reinforcement learning (RL) is consistently challenging due to its high non-stationarity and computational cost. To tackle this problem, existing approaches attempt to adapt common HPO techniques (e.g., population-based training or Bayesian optimization) to the RL scenario. However, they remain sample-inefficient and computationally expensive, which cannot facilitate a wide range of applications. In this paper, we propose ULTHO, an ultra-lightweight yet powerful framework for fast HPO in deep RL within single runs. Specifically, we formulate the HPO process as a multi-armed bandit with clustered arms (MABC) and link it directly to long-term return optimization. ULTHO also provides a quantified and statistical perspective to filter the HPs efficiently. We test ULTHO on benchmarks including ALE, Procgen, MiniGrid, and PyBullet. Extensive experiments demonstrate that the ULTHO can achieve superior performance with a simple architecture, contributing to the development of advanced and automated RL systems.
Mingqi Yuan, Bo Li 0037, Xin Jin 0014, Wenjun Zeng 0001
ICCV3
2025 Hybrid-Grained Feature Aggregation with Coarse-to-Fine Language Guidance for Self-Supervised Monocular Depth Estimation
abstract
Current self-supervised monocular depth estimation (MDE) approaches encounter performance limitations due to insufficient semantic-spatial knowledge extraction. To address this challenge, we propose Hybrid-depth, a novel framework that systematically integrates foundation models (e.g., CLIP and DINO) to extract visual priors and acquire sufficient contextual information for MDE. Our approach introduces a coarse-to-fine progressive learning framework: 1) Firstly, we aggregate multi-grained features from CLIP (global semantics) and DINO (local spatial details) under contrastive language guidance. A proxy task comparing close-distant image patches is designed to enforce depth-aware feature alignment using text prompts; 2) Next, building on the coarse features, we integrate camera pose information and pixel-wise language alignment to refine depth predictions. This module seamlessly integrates with existing self-supervised MDE pipelines (e.g., Monodepth2, ManyDepth) as a plug-and-play depth encoder, enhancing continuous depth estimation. By aggregating CLIP's semantic context and DINO's spatial details through language guidance, our method effectively addresses feature granularity mismatches. Extensive experiments on the KITTI benchmark demonstrate that our method significantly outperforms SOTA methods across all metrics, which also indeed benefits downstream tasks like BEV perception. Code is available at https://github.com/Zhangwenyao1/Hybrid-depth.
Hongsi Liu, Bohan Li 0015, Jiawei He 0002, Zekun Qi, Yunnan Wang, Shengyang Zhao, Xinqiang Yu, Wenjun Zeng 0001, Xin Jin 0014
ICCV10
2025 Open-World Reinforcement Learning over Long Short-Term Imagination
abstract
Training visual reinforcement learning agents in a high-dimensional open world presents significant challenges. While various model-based methods have improved sample efficiency by learning interactive world models, these agents tend to be “short-sighted”, as they are typically trained on short snippets of imagined experiences. We argue that the primary challenge in open-world decision-making is improving the exploration efficiency across a vast state space, especially for tasks that demand consideration of long-horizon payoffs. In this paper, we present LS-Imagine, which extends the imagination horizon within a limited number of state transition steps, enabling the agent to explore behaviors that potentially lead to promising long-term feedback. The foundation of our approach is to build a $\textit{long short-term world model}$. To achieve this, we simulate goal-conditioned jumpy state transitions and compute corresponding affordance maps by zooming in on specific areas within single images. This facilitates the integration of direct long-term values into behavior learning. Our method demonstrates significant improvements over state-of-the-art techniques in MineDojo.
Jiajian Li, Qi Wang 0080, Yunbo Wang, Xin Jin 0014, Yang Li 0041, Wenjun Zeng 0001, Xiaokang Yang 0001
ICLR4
2025 Representation Disentanglement for Semantic Coding
abstract
The learned image compression methods have achieved advances in both human perception and machine vision. However, previous methods focus on transmitting visual symbols losslessly instead of precisely conveying semantic meaning, resulting in bandwidth waste, especially for AI applications. In this work, we study "Semantic Coding" and propose a novel compression method based on representation disentanglement, which understands images at the attribute level and separates semantic factors into different parts, achieving a semantically structured bitstream for transmission. Specifically, we first leverage a conditional generative diffusion procedure for a disentangled representation learning, which learns meaningful semantic attribute factors in the latent space of the image assisted by the extra language inductive bias. Furthermore, we employ a learned codec to compress the inner disentangled representations as a bitstream, where each part represents a specific semantic and can be used for purposely decoding. Experiments show that our semantic coding method could reconstruct high-quality images and enable encryption by shifting the inner semantics.
Jinming Liu 0001, Junhao Geng, Lexiang Lv, Wenjun Zeng 0001, Xin Jin 0014
ICME5
2025 Electron Density-enhanced Molecular Geometry Learning
abstract
Electron density (ED), which describes the probability distribution of electrons in space, is crucial for accurately understanding the energy and force distribution in molecular force fields (MFF). Existing machine learning force fields (MLFF) focus on mining appropriate physical quantities from the atom-level conformation to enhance the molecular geometry representation while ignoring the unique information from microscopic electrons. In this work, we propose an efficient Electronic Density representation framework to enhance molecular Geometric learning (called EDG), which leverages images rendered from ED to boost molecular geometric representations in MLFF. Specifically, we construct a novel image-based ED representation, which consists of 2 million 6-view images with RGB-D channels, and design an ED representation learning model, called ImageED, to learn ED-related knowledge from these images. We further propose an efficient ED-aware teacher and introduce a cross-modal distillation strategy to transfer knowledge from the image-based teacher to the geometry-based students. Extensive experiments on QM9 and rMD17 demonstrate that EDG can be directly integrated into existing geometry-based models and significantly improves the capabilities of these models (e.g., SchNet, EGNN, SphereNet, ViSNet) for geometry representation learning in MLFF with a maximum average performance increase of 33.7%. Code and appendix are available at https://github.com/HongxinXiang/EDG
Hongxin Xiang, Jun Xia 0001, Xin Jin 0014, Wenjie Du 0003, Xiangxiang Zeng
IJCAI3
2025 Multi-Attribute Continual Learning for Blind Image Quality Assessment
abstract
Blind image quality assessment (BIQA) has evolved into a critical task in visual computing, requiring effective evaluation across multiple quality attributes such as brightness, sharpness, contrast, and colorfulness. Traditional BIQA methods based on single-task learning often suffer from catastrophic forgetting and struggle to generalize across diverse IQ attributes. To address these challenges, we propose a novel Multi-attribute Continual Learning framework, MaC-BIQA, which integrates Gated Attention Mechanism and Knowledge Graph Embedding (KGE) with the Learning without Forgetting (LwF) approach. Specifically, the Gated Attention Mechanism dynamically adjusts attention distribution by focusing on task-specific key regions, while the integration of Knowledge Graph Embedding (KGE) complements LwF by improving the understanding of inter-task relationships, ensuring that relevant information from previous tasks is preserved more effectively, even as the model learns new tasks. Together, they enhance the model’s adaptability and robustness in handling complex multi-attribute tasks, effectively mitigating catastrophic forgetting and significantly improving overall performance in multi-task environments. Extensive experiments on the KonIQ-10K and SPAQ datasets show that our method significantly reduces forgetting rates and improves robustness and generalization across multiple quality attributes. This study presents a key technical contribution by addressing the limitations of catastrophic forgetting and offering a scalable, adaptive solution for real-world multi-attribute BIQA.
Yunhao Luo 0005, Jinming Liu 0001, Wei Zhou 0021, Xin Jin 0014
ISCAS4
2025 Knowledge Regularized Negative Feature Tuning of Vision-Language Models for Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection is crucial for building reliable machine learning models. Although negative prompt tuning has enhanced the OOD detection capabilities of vision-language models, these tuned models often suffer from reduced generalization performance on unseen classes and styles. To address this challenge, we propose a novel method called Knowledge Regularized Negative Feature Tuning (KR-NFT), which integrates an innovative adaptation architecture termed Negative Feature Tuning (NFT) and a corresponding knowledge-regularization (KR) optimization strategy. Specifically, NFT applies distribution-aware transformations to pre-trained text features, effectively separating positive and negative features into distinct spaces. This separation maximizes the distinction between in-distribution (ID) and OOD images. Additionally, we introduce image-conditional learnable factors through a lightweight meta-network, enabling dynamic adaptation to individual images and mitigating sensitivity to class and style shifts. Compared to traditional negative prompt tuning, NFT demonstrates superior efficiency and scalability. To optimize this adaptation architecture, the KR optimization strategy is designed to enhance the discrimination between ID and OOD sets while mitigating pre-trained knowledge forgetting. This enhances OOD detection performance on trained ID classes while simultaneously improving OOD detection on unseen ID datasets. Notably, when trained with few-shot samples from ImageNet dataset, KR-NFT not only improves ID classification accuracy and OOD detection but also significantly reduces the FPR95 by 5.44% under an unexplored generalization setting with unseen ID categories. Codes can be found at https://github.com/ZhuWenjie98/KRNFT.
Wenjie Zhu 0003, Yabin Zhang 0001, Xin Jin 0014, Wenjun Zeng 0001, Lei Zhang 0006
ACM Multimedia3
2025 Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative Prior
abstract
Image compression methods are usually optimized isolatedly for human perception or machine analysis tasks. We reveal fundamental commonalities between these objectives: preserving accurate semantic information is paramount, as it directly dictates the integrity of critical information for intelligent tasks and aids human understanding. Concurrently, enhanced perceptual quality not only improves visual appeal but also, by ensuring realistic image distributions, benefits semantic feature extraction for machine tasks. Based on this insight, we propose Diff-ICMH, a generative image compression framework aiming for harmonizing machine and human vision in image compression. It ensures perceptual realism by leveraging generative priors and simultaneously guarantees semantic fidelity through the incorporation of Semantic Consistency loss (SC loss) during training. Additionally, we introduce the Tag Guidance Module (TGM) that leverages highly semantic image-level tags to stimulate the pre-trained diffusion model's generative capabilities, requiring minimal additional bit rates. Consequently, Diff-ICMH supports multiple intelligent tasks through a single codec and bitstream without any task-specific adaptation, while preserving high-quality visual experience for human perception. Extensive experimental results demonstrate Diff-ICMH's superiority and generalizability across diverse tasks, while maintaining visual appeal for human perception.
Ruoyu Feng 0001, Yunpeng Qi, Jinming Liu 0001, Xin Li 0082, Xin Jin 0014, Zhibo Chen 0001
NeurIPS6
2025 SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
abstract
While spatial reasoning has made progress in object localization relationships, it often overlooks object orientation—a key factor in 6-DoF fine-grained manipulation. Traditional pose representations rely on pre-defined frames or templates, limiting generalization and semantic grounding. In this paper, we introduce the concept of semantic orientation, which defines object orientations using natural language in a reference-frame-free manner (e.g., the ''plug-in'' direction of a USB or the ''handle'' direction of a cup). To support this, we construct OrienText300K, a large-scale dataset of 3D objects annotated with semantic orientations, and develop PointSO, a general model for zero-shot semantic orientation prediction. By integrating semantic orientation into VLM agents, our SoFar framework enables 6-DoF spatial reasoning and generates robotic actions. Extensive experiments demonstrated the effectiveness and generalization of our SoFar, e.g., zero-shot 48.7\% successful rate on Open6DOR and zero-shot 74.9\% successful rate on SIMPLER-Env.
Zekun Qi, Yufei Ding 0002, Runpei Dong, Xinqiang Yu, Baoyu Li, Xialin He, Guofan Fan, Jiazhao Zhang, Jiawei He 0002, Jiayuan Gu, Xin Jin 0014, Kaisheng Ma, Zhizheng Zhang 0011, He Wang 0010, Li Yi 0001
NeurIPS14
2025 EDBench: Large-Scale Electron Density Data for Molecular Modeling
abstract
Existing molecular machine learning force fields (MLFFs) generally focus on the learning of atoms, molecules, and simple quantum chemical properties (such as energy and force), but ignore the importance of electron density (ED) $\rho(r)$ in accurately understanding molecular force fields (MFFs). ED describes the probability of finding electrons at specific locations around atoms or molecules, which uniquely determines all ground state properties (such as energy, molecular structure, etc.) of interactive multi-particle systems according to the Hohenberg-Kohn theorem. However, the calculation of ED relies on the time-consuming first-principles density functional theory (DFT), which leads to the lack of large-scale ED data and limits its application in MLFFs. In this paper, we introduce EDBench, a large-scale, high-quality dataset of ED designed to advance learning-based research at the electronic scale. Built upon the PCQM4Mv2, EDBench provides accurate ED data, covering 3.3 million molecules. To comprehensively evaluate the ability of models to understand and utilize electronic information, we design a suite of ED-centric benchmark tasks spanning prediction, retrieval, and generation. Our evaluation of several state-of-the-art methods demonstrates that learning from EDBench is not only feasible but also achieves high accuracy. Moreover, we show that learning-based methods can efficiently calculate ED with comparable precision while significantly reducing the computational cost relative to traditional DFT calculations. All data and benchmarks from EDBench will be freely available, laying a robust foundation for ED-driven drug discovery and materials science.
Hongxin Xiang, Mingquan Liu, Zhixiang Cheng, Wenjie Du 0003, Jun Xia 0001, Xin Jin 0014, Xiangxiang Zeng
NeurIPS9
2025 DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
abstract
Recent advances in vision-language-action (VLA) models have shown promise in integrating image generation with action prediction to improve generalization and reasoning in robot manipulation. However, existing methods are limited to challenging image-based forecasting, which suffers from redundant information and lacks comprehensive and critical world knowledge, including dynamic, spatial and semantic information. To address these limitations, we propose DreamVLA, a novel VLA framework that integrates comprehensive world knowledge forecasting to enable inverse dynamics modeling, thereby establishing a perception-prediction-action loop for manipulation tasks. Specifically, DreamVLA introduces a dynamic-region-guided world knowledge prediction, integrated with the spatial and semantic cues, which provide compact yet comprehensive representations for action planning. This design aligns with how humans interact with the world by first forming abstract multimodal reasoning chains before acting. To mitigate interference among the dynamic, spatial and semantic information during training, we adopt a block-wise structured attention mechanism that masks their mutual attention, preventing information leakage and keeping each representation clean and disentangled. Moreover, to model the conditional distribution over future actions, we employ a diffusion-based transformer that disentangles action representations from shared latent features. Extensive experiments on both real-world and simulation environments demonstrate that DreamVLA achieves 76.7 success rate on real robot tasks and 4.44 average length on the CALVIN ABC-D benchmarks.
Hongsi Liu, Zekun Qi, Yunnan Wang, Xinqiang Yu, Jiazhao Zhang, Runpei Dong, Jiawei He 0002, He Wang 0010, Zhizheng Zhang 0011, Li Yi 0001, Wenjun Zeng 0001, Xin Jin 0014
NeurIPS13
2025 A human-machine-environment interaction system based on VLA model and brain-computer interface
abstract
The vision language action (VLA) model has shown extraordinary ability in the motion planning of the robots, which has great value in human-machine-environment interaction (HMEI), especially for rehabilitation engineering and robotic arm teleoperation fields. This paper proposes an HMEI system that integrates the brain-computer interface (BCI) with VLA model, enabling robots to interact with surrounding environment according to human intentions. Firstly, neural decoding technique is employed to interpret neural signals into intentions via a time and frequency fusion model. Then, the user’s intentions are converted into a textual instruction, which subsequently guides the VLA model in generating corresponding motion trajectories for robotic arm teleoperation. Experimental results demonstrate smooth and compliant control of the robotic arm, which can be a viable solution for HMEI applications.
Shaoyang Hua, Qile He, Xin Jin 0014
SMC5
2025 Cross-Modal World Models for Offline Visual Reinforcement Learning
abstract
Offline reinforcement learning (RL) with visual pixels encounters two primary challenges: overfitting in representation learning induced by limited data, and value overestimation of out-of-distribution states. Recent work has adopted accessible simulators to mitigate these issues, but rendering images introduces additional computational costs and training challenges. In this paper, we study the problem of anti-exploration with accessible state-based simulators for effective learning of offline visual control. To address these challenges, we introduce a model-based RL framework, dubbed Cross-Modal World Models (X-MWM). Concretely, we build two independent agents: a source model trained on low-dimensional states and a target model that learns from high-dimensional images. Initially, in light of the reward function discrepancy between the domains, we pretrain the source agent with latent disagreement-based intrinsic rewards. Subsequently, to prevent overfitting in offline representation learning, cross-modal latent alignment is employed to close the distance of latent state distributions. In this way, during the target agent training phase, the value of the source critic serves as an anti-exploration constraint to adjust the learning target of the offline RL agent, which encourages more conservative behavior, effectively alleviating value overestimation induced by out-of-distribution states. Experimental results of various robotic manipulation tasks on MetaWorld validate the superiority of our approach.
Qi Wang 0080, Xin Jin 0014, Baao Xie, Xiaokang Yang 0001, Wenjun Zeng 0001
SMC2
2025 Standard Codec is Enough: A Training-Free 4D Gaussian Compression with Dynamic UV Mapping
abstract
4D Gaussian Splatting (4DGS) has demonstrated advances in the dynamic scene representation. However, the time-varying attributes across frames introduce considerable storage and transmission costs, making 4DGS challenging to widely deploy. Existing compression methods struggle to obtain inter-frame residuals due to the unstructured nature of Gaussian representations, making explicit motion estimation and residual modeling inherently challenging. To address these, we propose a Training-Free 4D Gaussian Compression framework, TF4DGC, which transforms 4D Gaussian into a well-structured 2D representation, easy to estimate motion for coding, via a UV mapping. Specifically, we project 3D Gaussians onto a canonical sphere to obtain temporally consistent UV coordinates, and organize per-frame Gaussian attributes into multi-channel video sequences. This design enables the direct use of standard video codecs (e.g., AVC, HEVC) for compression, which is compatible with widespread hardware decoder support on laptops and mobile devices. Experimental results show that our method efficiently compresses both reconstructed and generated Gaussian scenarios, highlighting its general applicability. Our method offers a scalable and practical solution for 4DGS compression and facilitates real-time deployment in bandwidth constrained environments.
Jinming Liu 0001, Shengyang Zhao, Qiang Hu 0003, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014
VCIP7
2025 Quadtree Partitioning-based Visual Token Pruning for MLLMs Considering Information Density
abstract
Multimodal Large Language Models (MLLMs) excel at comprehensive understanding by integrating visual and textual information. However, their inference speed is often bottlenecked by redundant visual token inputs. Existing methods tend to alleviate this issue with a heuristic pruning strategy based on token importance, tailored to certain commonly adopted vision encoders like CLIP. In this paper, we propose a novel training-free token pruning method based on a well-designed metric of information density, where we decide which tokens are retained according to their entropy, following the classic information theory. Based on that, we further propose a quadtree partitioning strategy, in which we retain these tokens with higher entropy so as to preserve the visual spatial structure while allocating more tokens to more informative regions. Experiments on LLaVA-v1.5-7B and 13B across six benchmarks show our method achieves state-of-the-art performance—retaining over 90% of full-token accuracy even at a 6.25% token budget—while cutting TFLOPs by up to 20% compared to FastV and by 81% compared to the original LLaVA-v1.5.
Yuntao Wei, Jinming Liu 0001, Shengyang Zhao, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014
VCIP6
2025 PrismGS: Physically-Grounded Anti-Aliasing for High-Fidelity Large-Scale 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) has recently enabled real-time photorealistic rendering in compact scenes, but scaling to large urban environments introduces severe aliasing artifacts and optimization instability, especially under high-resolution (e.g., 4K) rendering. These artifacts, manifesting as flickering textures and jagged edges, arise from the mismatch between Gaussian primitives and the multi-scale nature of urban geometry. While existing "divide-and-conquer" pipelines address scalability, they fail to resolve this fidelity gap. In this paper, we propose PrismGS, a physically-grounded regularization framework that improves the intrinsic rendering behavior of 3D Gaussians. PrismGS integrates two synergistic regularizers. The first is pyramidal multi-scale supervision, which enforces consistency by supervising the rendering against a pre-filtered image pyramid. This compels the model to learn an inherently anti-aliased representation that remains coherent across different viewing scales, directly mitigating flickering textures. This is complemented by an explicit size regularization that imposes a physically-grounded lower bound on the dimensions of the 3D Gaussians. This prevents the formation of degenerate, view-dependent primitives, leading to more stable and plausible geometric surfaces and reducing jagged edges. Our method is plug-and-play and compatible with existing pipelines. Extensive experiments on MatrixCity, Mill-19, and UrbanScene3D demonstrate that PrismGS achieves state-of-the-art performance, yielding significant PSNR gains around 1.5 dB against CityGaussian, while maintaining its superior quality and robustness under demanding 4K rendering.
Houqiang Zhong, Zhenglong Wu, Sihua Fu, Zihan Zheng, Xin Jin 0014, Xiaoyun Zhang 0001, Li Song 0001, Qiang Hu 0003
VCIP5
2025 Diffusion Models for Image Restoration and Enhancement: A Comprehensive Survey
Xin Li 0082, Yulin Ren, Xin Jin 0014, Cuiling Lan, Xingrui Wang, Wenjun Zeng 0001, Xinchao Wang, Zhibo Chen 0001
Int. J. Comput. Vis.3
2025 Canvas: Compositional Generation for Art Painting With Seamless Subject-Driven Infusion
Yunnan Wang, Lexiang Lv, Zequn Zhang, Xiaoyu Shen 0001, Xin Jin 0014, Wenjun Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.7
2025 MSSF: A 4D Radar and Camera Fusion Framework With Multi-Stage Sampling for 3D Object Detection in Autonomous Driving
abstract
As one of the automotive sensors that have emerged in recent years, 4D millimeter-wave radar has a higher resolution than conventional 3D radar and provides precise elevation measurements. But its point clouds are still sparse and noisy, making it challenging to meet the requirements of autonomous driving. Camera, as another commonly used sensor, can capture rich semantic information. As a result, the fusion of 4D radar and camera can provide an affordable and robust perception solution for autonomous driving systems. However, previous radar-camera fusion methods have not yet been thoroughly investigated, resulting in a large performance gap compared to LiDAR-based methods. Specifically, they ignore the feature-blurring problem and do not deeply interact with image semantic information. To this end, we present a simple but effective multi-stage sampling fusion (MSSF) network based on 4D radar and camera. On the one hand, we design a fusion block that can deeply interact point cloud features with image features, and can be applied to commonly used single-modal backbones in a plug-and-play manner. The fusion block encompasses two types, namely, simple feature fusion (SFF) and multi-scale deformable feature fusion (MSDFF). The SFF is easy to implement, while the MSDFF has stronger fusion abilities. On the other hand, we propose a semantic-guided head to perform foreground-background segmentation on voxels with voxel feature re-weighting, further alleviating the problem of feature blurring. Extensive experiments on the View-of-Delft (VoD) and TJ4DRadset datasets demonstrate the effectiveness of our MSSF. Notably, compared to state-of-the-art methods, MSSF achieves a 7.0% and 4.0% improvement in 3D mean average precision on the VoD and TJ4DRadSet datasets, respectively. It even surpasses classical LiDAR-based methods on the VoD dataset.
Hongsi Liu, Jun Liu 0004, Guangfeng Jiang, Xin Jin 0014
IEEE Trans. Intell. Transp. Syst.4
2025 Exploring Contrastive Pre-Training for Domain Connections in Medical Image Segmentation
abstract
Unsupervised domain adaptation (UDA) in medical image segmentation aims to improve the generalization of deep models by alleviating domain gaps caused by inconsistency across equipment, imaging protocols, and patient conditions. However, existing UDA works remain insufficiently explored and present great limitations: 1) Exhibit cumbersome designs that prioritize aligning statistical metrics and distributions, which limits the model's flexibility and generalization while also overlooking the potential knowledge embedded in unlabeled data; 2) More applicable in a certain domain, lack the generalization capability to handle diverse shifts encountered in clinical scenarios. To overcome these limitations, we introduce MedCon, a unified framework that leverages general unsupervised contrastive pre-training to establish domain connections, effectively handling diverse domain shifts without tailored adjustments. Specifically, it initially explores a general contrastive pre-training to establish domain connections by leveraging the rich prior knowledge from unlabeled images. Thereafter, the pre-trained backbone is fine-tuned using source-based images to ultimately identify per-pixel semantic categories. To capture both intra- and inter-domain connections of anatomical structures, we construct positive-negative pairs from a hybrid aspect of both local and global scales. In this regard, a shared-weight encoder-decoder is employed to generate pixel-level representations, which are then mapped into hyper-spherical space using a non-learnable projection head to facilitate positive pair matching. Comprehensive experiments on diverse medical image datasets confirm that MedCon outperforms previous methods by effectively managing a wide range of domain shifts and showcasing superior generalization capabilities.
Zequn Zhang, Yunnan Wang, Baao Xie, Yuhang Li 0005, Zhen Chen 0013, Xin Jin 0014, Wenjun Zeng 0001
IEEE Trans. Medical Imaging8
2025 Unleash the Power of Vision-Language Models by Visual Attention Prompt and Multimodal Interaction
abstract
Pre-trained vision-language models (VLMs), equipped with parameter-efficient tuning (PET) methods like prompting, have shown impressive knowledge transferability on new downstream tasks, but they are still prone to be limited by catastrophic forgetting and overfitting dilemma due to large gaps among tasks. Furthermore, the underlying physical mechanisms of prompt-based tuning methods (especially for visual prompting) remain largely unexplored. It is unclear why these methods work solely based on learnable parameters as prompts for adaptation. To address the above challenges, we present a new prompt-based framework for vision-language models, termed Uni-prompt. Our framework transfers VLMs to downstream tasks by designing visual prompts from an attention perspective that reduces the transfer/solution space, which enables the vision model to focus on task-relevant regions of the input image while also learning task-specific knowledge. Additionally, Uni-prompt further aligns visual-text prompts learning through a pretext task with masked representation modeling interactions, which implicitly learns a global cross-modal matching between visual and language concepts for consistency. We conduct extensive experiments on the few-shot classification task and achieve significant improvement using our Uni-prompt method while requiring minimal extra parameters cost.
Letian Wu, Zequn Zhang, Tao Yu 0012, Chao Ma 0004, Xin Jin 0014, Xiaokang Yang 0001, Wenjun Zeng 0001
IEEE Trans. Multim.6
2024 SwiftPillars: High-Efficiency Pillar Encoder for Lidar-Based 3D Detection
abstract
Lidar-based 3D Detection is one of the significant components of Autonomous Driving. However, current methods over-focus on improving the performance of 3D Lidar perception, which causes the architecture of networks becoming complicated and hard to deploy. Thus, the methods are difficult to apply in Autonomous Driving for real-time processing. In this paper, we propose a high-efficiency network, SwiftPillars, which includes Swift Pillar Encoder (SPE) and Multi-scale Aggregation Decoder (MAD). The SPE is constructed by a concise Dual-attention Module with lightweight operators. The Dual-attention Module utilizes feature pooling, matrix multiplication, etc. to speed up point-wise and channel-wise attention extraction and fusion. The MAD interconnects multiple scale features extracted by SPE with minimal computational cost to leverage performance. In our experiments, our proposal accomplishes 61.3% NDS and 53.2% mAP in nuScenes dataset. In addition, we evaluate inference time on several platforms (P4, T4, A2, MLU370, RTX3080), where SwiftPillars achieves up to 13.3ms (75FPS) on NVIDIA Tesla T4. Compared with PointPillars, SwiftPillars is on average 26.58% faster in inference speed with equivalent GPUs and a higher mAP of approximately 3.2% in the nuScenes dataset.
Xin Jin 0014, Ruining Yang, Fei Hui, Wei Wu 0021
AAAI1
2024 One at a Time: Progressive Multi-Step Volumetric Probability Learning for Reliable 3D Scene Perception
abstract
Numerous studies have investigated the pivotal role of reliable 3D volume representation in scene perception tasks, such as multi-view stereo (MVS) and semantic scene completion (SSC). They typically construct 3D probability volumes directly with geometric correspondence, attempting to fully address the scene perception tasks in a single forward pass. However, such a single-step solution makes it hard to learn accurate and convincing volumetric probability, especially in challenging regions like unexpected occlusions and complicated light reflections. Therefore, this paper proposes to decompose the complicated 3D volume representation learning into a sequence of generative steps to facilitate fine and reliable scene perception. Considering the recent advances achieved by strong generative diffusion models, we introduce a multi-step learning framework, dubbed as VPD, dedicated to progressively refining the Volumetric Probability in a Diffusion process. Specifically, we first build a coarse probability volume from input images with the off-the-shelf scene perception baselines, which is then conditioned as the basic geometry prior before being fed into a 3D diffusion UNet, to progressively achieve accurate probability distribution modeling. To handle the corner cases in challenging areas, a Confidence-Aware Contextual Collaboration (CACC) module is developed to correct the uncertain regions for reliable volumetric learning based on multi-scale contextual contents. Moreover, an Online Filtering (OF) strategy is designed to maintain representation consistency for stable diffusion sampling. Extensive experiments are conducted on scene perception tasks including multi-view stereo (MVS) and semantic scene completion (SSC), to validate the efficacy of our method in learning reliable volumetric representations. Notably, for the SSC task, our work stands out as the first to surpass LiDAR-based methods on the SemanticKITTI dataset.
Bohan Li 0015, Yasheng Sun, Jingxin Dong 0002, Jinming Liu 0001, Xin Jin 0014, Wenjun Zeng 0001
AAAI6
2024 Multi-Prompts Learning with Cross-Modal Alignment for Attribute-Based Person Re-identification
abstract
The fine-grained attribute descriptions can significantly supplement the valuable semantic information for person image, which is vital to the success of person re-identification (ReID) task. However, current ReID algorithms typically failed to effectively leverage the rich contextual information available, primarily due to their reliance on simplistic and coarse utilization of image attributes. Recent advances in artificial intelligence generated content have made it possible to automatically generate plentiful fine-grained attribute descriptions and make full use of them. Thereby, this paper explores the potential of using the generated multiple person attributes as prompts in ReID tasks with off-the-shelf (large) models for more accurate retrieval results. To this end, we present a new framework called Multi-Prompts ReID (MP-ReID), based on prompt learning and language models, to fully dip fine attributes to assist ReID task. Specifically, MP-ReID first learns to hallucinate diverse, informative, and promptable sentences for describing the query images. This procedure includes (i) explicit prompts of which attributes a person has and furthermore (ii) implicit learnable prompts for adjusting/conditioning the criteria used towards this person identity matching. Explicit prompts are obtained by ensembling generation models, such as ChatGPT and VQA models. Moreover, an alignment module is designed to fuse multi-prompts (i.e., explicit and implicit ones) progressively and mitigate the cross-modal gap. Extensive experiments on the existing attribute-involved ReID datasets, namely, Market1501 and DukeMTMC-reID, demonstrate the effectiveness and rationality of the proposed MP-ReID solution.
Yajing Zhai, Yawen Zeng, Zheng Qin 0001, Xin Jin 0014, Da Cao
AAAI5
2024 Consistency Prior Matters: Biomedical-Prompting Dual Augmentation for Domain Adaptive Medical Image Segmentation
abstract
Existing domain adaptive medical image segmentation works typically rely on style transfer techniques to mitigate the unexpected domain gap, which inevitably suffers from synthesized artifacts or unreasonable stylization. In this paper, we propose to inject biomedical-related prior knowledge (i.e., intensity and anatomical consistency) as regularization in a prompting manner, bridging the domain gap across modalities. Technically, we develop an efficient scheme called Biomedical-Prompting Dual Augmentation (BPDA) to learn domain-invariant representations by enforcing consistent model predictions across different augmented views. BPDA augments unpaired source and target images from intensity and anatomical aspects in a dual manner, while prompting the framework to fully understand the anatomical structure-invariant features. In this way, our method captures discriminative inherent representations on cross-modality scenarios. Furthermore, we also introduce a Cross-Domain Prototype Denoising (CDPD) in BPDA to refine pseudo-labeling results with the class centroids for a reliable augmentation. Extensive experiments on the cross-modality abdominal and cardiac segmentation benchmarks demonstrate the superiority of our method over state-of-the-art alternatives.
Yunnan Wang, Zequn Zhang, Xin Jin 0014, Wenjun Zeng 0001
BIBM3
2024 Inter-X: Towards Versatile Human-Human Interaction Analysis
abstract
The analysis of the ubiquitous human-human interactions is pivotal for understanding humans as social beings. Existing human-human interaction datasets typically suffer from inaccurate body motions, lack of hand gestures and fine- grained textual descriptions. To better perceive and generate human-human interactions, we propose Inter-X, a currently largest human-human interaction dataset with accurate body movements and diverse interaction patterns, together with detailed hand gestures. The dataset includes
Liang Xu 0012, Xintao Lv, Yichao Yan, Xin Jin 0014, Shuwen Wu, Congsheng Xu, Yizhou Zhou, Fengyun Rao, Xingdong Sheng, Yunhui Liu 0006, Wenjun Zeng 0001, Xiaokang Yang 0001
CVPR4
2024 ReGenNet: Towards Human Action-Reaction Synthesis
abstract
Humans constantly interact with their surrounding environments. Current human-centric generative models mainly focus on synthesizing humans plausibly interacting with static scenes and objects, while the dynamic human action-reaction synthesis for ubiquitous causal human-human interactions is less explored. Human-human interactions can be regarded as asymmetric with actors and reactors in atomic interaction periods. In this paper, we compre-hensively analyze the asymmetric, dynamic, synchronous, and detailed nature of human-human interactions and propose the first multi-setting human action-reaction synthe-sis benchmark to generate human reactions conditioned on given human actions. To begin with, we propose to an-notate the actor-reactor order of the interaction sequences for the NTU120, InterHuman, and Chi3D datasets. Based on them, a diffusion-based generative model with a Trans-former decoder architecture called ReGenNet together with an explicit distance-based interaction loss is proposed to predict human reactions in an online manner, where the future states of actors are unavailable to reactors. Quantitative and qualitative results show that our method can gener-ate instant and plausible human reactions compared to the baselines, and can generalize to unseen actor motions and viewpoint changes.
Liang Xu 0012, Yizhou Zhou, Yichao Yan, Xin Jin 0014, Wenhan Zhu, Fengyun Rao, Xiaokang Yang 0001, Wenjun Zeng 0001
CVPR4
2024 Closed-Loop Unsupervised Representation Disentanglement with β-VAE Distillation and Diffusion Probabilistic Feedback
Xin Jin 0014, Bohan Li 0015, Baao Xie, Jinming Liu 0001, Tao Yang 0032, Wenjun Zeng 0001
ECCV (45)1
2024 Hierarchical Temporal Context Learning for Camera-Based Semantic Scene Completion
Bohan Li 0015, Jiajun Deng, Zhujin Liang, Dalong Du, Xin Jin 0014, Wenjun Zeng 0001
ECCV (4)6
2024 Rate-Distortion-Cognition Controllable Versatile Neural Image Compression
Jinming Liu 0001, Ruoyu Feng 0001, Yunpeng Qi, Qiuyu Chen, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014
ECCV (56)7
2024 HIMO: A New Benchmark for Full-Body Human Interacting with Multiple Objects
Xintao Lv, Liang Xu 0012, Yichao Yan, Xin Jin 0014, Congsheng Xu, Shuwen Wu, Lincheng Li, Mengxiao Bi, Wenjun Zeng 0001, Xiaokang Yang 0001
ECCV (4)4
2024 DreamLIP: Language-Image Pre-training with Long Captions
Kecheng Zheng, Wei Wu 0021, Shuailei Ma, Xin Jin 0014, Yujun Shen
ECCV (18)6
2024 Rethinking Domain Adaptation and Generalization in the ERA Of Clip
abstract
In recent studies on domain adaptation, significant emphasis has been placed on the advancement of learning shared knowledge from a source domain to a target domain. Recently, the large vision-language pre-trained model (i.e., CLIP) has shown strong ability on zero-shot recognition, and parameter efficient tuning can further improve its performance on specific tasks. This work demonstrates that a simple domain prior boosts CLIP’s zero-shot recognition in a specific domain. Besides, CLIP’s adaptation relies less on source domain data due to its diverse pre-training dataset. Furthermore, we create a benchmark for zero-shot adaptation and pseudo-labeling based self-training with CLIP. Last but not least, we propose to improve the task generalization ability of CLIP from multiple unlabeled domains, which is a more practical and unique scenario. We believe our findings motivate a rethinking of domain adaptation benchmarks and the associated role of related algorithms in the era of CLIP.
Ruoyu Feng 0001, Tao Yu 0012, Xin Jin 0014, Xiaoyuan Yu, Zhibo Chen 0001
ICIP3
2024 An Explainable Spectral Analysis For Light Field Image Quality Assessment
abstract
The light field (LF) technology is a promising way to achieve the immersive multimedia. The LF image quality assessment is vital for user experience. Compared with traditional twodimensional images, the LF images suffer from not only spatial distortion but also angular distortion, which destroys the light field angular consistency. Many algorithms are proposed to handle the objective light field image quality assessment. However, there is no mathematical description for the LF angular consistency and no discussion of spectral analysis in the angular domain. In this paper we provide an explainable spectral analysis for the LF image quality assessment. First, we formulate the LF as a 4 D signal. It is well-known that the energy of the light field Epipolar Plane Image (EPI) in the frequency domain should concentrate in a double-wedge area, but for the first time we introduce it into the LFI quality assessment. We test and discuss how the different distortions affect the energy distribution. The spectral features for blind LF image quality assessment are further proposed to complement the spatial features extracted by off-the-shelf algorithm. Experimentally we demonstrate that the features based on spectral energy distribution are more sensitive to angular distortion. Consequently, it can be a good guidance for a better LF quality assessment algorithm design.
Shengyang Zhao, Xin Jin 0014
ICIP2
2024 Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion
Bohan Li 0015, Yasheng Sun, Zhujin Liang, Dalong Du, Zhuanghui Zhang, Yunnan Wang, Xin Jin 0014, Wenjun Zeng 0001
IJCAI8
2024 Scene Graph Disentanglement and Composition for Generalizable Complex Image Generation
abstract
There has been exciting progress in generating images from natural language or layout conditions. However, these methods struggle to faithfully reproduce complex scenes due to the insufficient modeling of multiple objects and their relationships. To address this issue, we leverage the scene graph, a powerful structured representation, for complex image generation. Different from the previous works that directly use scene graphs for generation, we employ the generative capabilities of variational autoencoders and diffusion models in a generalizable manner, compositing diverse disentangled visual clues from scene graphs. Specifically, we first propose a Semantics-Layout Variational AutoEncoder (SL-VAE) to jointly derive (layouts, semantics) from the input scene graph, which allows a more diverse and reasonable generation in a one-to-many mapping. We then develop a Compositional Masked Attention (CMA) integrated with a diffusion model, incorporating (layouts, semantics) with fine-grained attributes as generation guidance. To further achieve graph manipulation while keeping the visual content consistent, we introduce a Multi-Layered Sampler (MLS) for an "isolated" image editing effect. Extensive experiments demonstrate that our method outperforms recent competitors based on text, layout, or scene graph, in terms of generation rationality and controllability.
Yunnan Wang, Zequn Zhang, Baao Xie, Xihui Liu, Wenjun Zeng 0001, Xin Jin 0014
NeurIPS8
2024 Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning
abstract
Training offline RL models using visual inputs poses two significant challenges, *i.e.*, the overfitting problem in representation learning and the overestimation bias for expected future rewards. Recent work has attempted to alleviate the overestimation bias by encouraging conservative behaviors. This paper, in contrast, tries to build more flexible constraints for value estimation without impeding the exploration of potential advantages. The key idea is to leverage off-the-shelf RL simulators, which can be easily interacted with in an online manner, as the “*test bed*” for offline policies. To enable effective online-to-offline knowledge transfer, we introduce CoWorld, a model-based RL approach that mitigates cross-domain discrepancies in state and reward spaces. Experimental results demonstrate the effectiveness of CoWorld, outperforming existing RL approaches by large margins.
Qi Wang 0080, Yunbo Wang, Xin Jin 0014, Wenjun Zeng 0001, Xiaokang Yang 0001
NeurIPS4
2024 Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language Models
abstract
Disentangled representation learning (DRL) aims to identify and decompose underlying factors behind observations, thus facilitating data perception and generation. However, current DRL approaches often rely on the unrealistic assumption that semantic factors are statistically independent. In reality, these factors may exhibit correlations, which off-the-shelf solutions have yet to properly address. To tackle this challenge, we introduce a bidirectional weighted graph-based framework, to learn factorized attributes and their interrelations within complex data. Specifically, we propose a $\beta$-VAE based module to extract factors as the initial nodes of the graph, and leverage the multimodal large language model (MLLM) to discover and rank latent correlations, thereby updating the weighted edges. By integrating these complementary modules, our model successfully achieves fine-grained, practical and unsupervised disentanglement. Experiments demonstrate our method's superior performance in disentanglement and reconstruction. Furthermore, the model inherits enhanced interpretability and generalizability from MLLMs.
Baao Xie, Qiuyu Chen, Yunnan Wang, Zequn Zhang, Xin Jin 0014, Wenjun Zeng 0001
NeurIPS5
2024 Tell Codec What Worth Compressing: Semantically Disentangled Image Coding for Machine with LMMs
abstract
We present a new image compression paradigm to achieve "intelligently coding for machine" by cleverly leveraging the common sense of Large Multimodal Models (LMMs). We are motivated by the evidence that large language/multimodal models are powerful general-purpose semantics predictors for understanding the real world. Different from traditional image compression typically optimized for human eyes, the image coding for machines (ICM) framework we focus on requires the compressed bitstream to more comply with different downstream intelligent analysis tasks. To this end, we employ LMM to${\text{tell codec what to compress}}$: 1) first utilize the powerful semantic understanding capability of LMMs w.r.t object grounding, identification, and importance ranking via prompts, to disentangle image content before compression, 2) and then based on these semantic priors we accordingly encode and transmit objects of the image in order with a structured bitstream. In this way, diverse vision benchmarks including image classification, object detection, instance segmentation, etc., can be well supported with such a semantically structured bitstream. We dub our method "SDComp" for "Semantically Disentangled Compression", and compare it with state-of-the-art codecs on a wide variety of different vision tasks. SDComp codec leads to more flexible reconstruction results, promised decoded visual quality, and a more generic/satisfactory intelligent task-supporting ability.
Jinming Liu 0001, Yuntao Wei, Junyan Lin, Shengyang Zhao, Heming Sun, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014
VCIP8
2024 Structure-preserving feature alignment for old photo colorization
Yingxue Pang, Xin Jin 0014, Jun Fu 0007, Zhibo Chen 0001
Pattern Recognit.2
2024 Domain Prompt Tuning via Meta Relabeling for Unsupervised Adversarial Adaptation
abstract
Unsupervised adversarial domain adaptation (ADA) aims to learn domain-invariant features by confusing a domain discriminator. As training goes on, the feature distributions of source and target samples are increasingly aligned/indistinguishable. The discrimination capability of the domain discriminator w.r.t. those aligned samples deteriorates due to the domain label of each sample is still fixed all through the learning process, which thus cannot effectively further drive the feature learning. A recently proposed method named Re-enforceable Adversarial Domain Adaptation (RADA) [1] tend to re-energize the domain discriminator during the training by using dynamic domain labels. Specifically, RADA sets up a heuristic criterion and uses it to relabel the well aligned target domain samples as source domain samples on the fly. In our study, we identify a critical problem of RADA: it is a kind of heuristic domain data re-partition solution without explicitly serving the adaptation task itself, suggesting that the criteria of RADA on which sample should be relabeled is hard to decide. To address the problem, we revisit domain relabeling process from a perspective of prompt tuning, and introduce a meta-optimized learnable prompts into RADA to replace some hand-craft designs in dynamic relabeling process, which scheme is named as RADA-prompt. Particularly, we employ a module of meta-prompter, which learns to adaptively relabel the samples based on the objective of serving UDA task. To train the meta-prompter, we leverage a domain alignment measurement and a classification measurement as the meta optimization objective. Extensive experiments on multiple unsupervised domain adaptation benchmarks demonstrate the effectiveness and superiority of RADA-prompt, this scheme also achieves state-of-the-art performance.
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001
IEEE Trans. Multim.1
2023 Task Residual for Tuning Vision-Language Models
abstract
Large-scale vision-language models (VLMs) pre-trained on billion-level data have learned general visual representations and broad visual concepts. In principle, the welllearned knowledge structure of the VLMs should be inherited appropriately when being transferred to downstream tasks with limited data. However, most existing efficient transfer learning (ETL) approaches for VLMs either damage or are excessively biased towards the prior knowledge, e.g., prompt tuning (PT) discards the pre-trained text-based classifier and builds a new one while adapter-style tuning (AT) fully relies on the pre-trained features. To address this, we propose a new efficient tuning approach for VLMs named Task Residual Tuning (TaskRes), which performs directly on the text-based classifier and explicitly decouples the prior knowledge of the pre-trained models and new knowledge regarding a target task. Specifically, TaskRes keeps the original classifier weights from the VLMs frozen and obtains a new classifier for the target task by tuning a set of prior-independent parameters as a residual to the original one, which enables reliable prior knowledge preservation and flexible task-specific knowledge exploration. The proposed TaskRes is simple yet effective, which significantly outperforms previous ETL methods (e.g., PT and AT) on 11 benchmark datasets while requiring minimal effort for the implementation. Our code is available at https://github.com/geekyutao/TaskRes.
Tao Yu 0012, Zhihe Lu, Xin Jin 0014, Zhibo Chen 0001, Xinchao Wang
CVPR3
2023 Learning Distortion Invariant Representation for Image Restoration from a Causality Perspective
abstract
In recent years, we have witnessed the great advancement of Deep neural networks (DNNs) in image restoration. However, a critical limitation is that they cannot generalize well to real-world degradations with different degrees or types. In this paper, we are the first to propose a novel training strategy for image restoration from the causality perspective, to improve the generalization ability of DNNs for unknown degradations. Our method, termed Distortion Invariant representation Learning (DIL), treats each distortion type and degree as one specific confounder, and learns the distortion-invariant representation by eliminating the harmful confounding effect of each degradation. We derive our DIL with the back-door criterion in causality by modeling the interventions of different distortions from the optimization perspective. Particularly, we introduce counterfactual distortion augmentation to simulate the virtual distortion types and degrees as the confounders. Then, we instantiate the intervention of each distortion with a virtual model updating based on corresponding distorted images, and eliminate them from the meta-learning perspective. Extensive experiments demonstrate the generalization capability of our DIL on unseen distortion types and degrees. Our code will be available at https://github.com/lixinustc/Causal-IR-DIL.
Xin Li 0082, Bingchen Li 0001, Xin Jin 0014, Cuiling Lan, Zhibo Chen 0001
CVPR3
2023 Semantically Structured Image Compression via Irregular Group-Based Decoupling
abstract
Image compression techniques typically focus on compressing rectangular images for human consumption, however, resulting in transmitting redundant content for downstream applications. To overcome this limitation, some previous works propose to semantically structure the bitstream, which can meet specific application requirements by selective transmission and reconstruction. Nevertheless, they divide the input image into multiple rectangular regions according to semantics and ignore avoiding information interaction among them, causing waste of bitrate and distorted reconstruction of region boundaries. In this paper, we propose to decouple an image into multiple groups with irregular shapes based on a customized group mask and compress them independently. Our group mask describes the image at a finer granularity, enabling significant bitrate saving by reducing the transmission of redundant content. Moreover, to ensure the fidelity of selective reconstruction, this paper proposes the concept of group-independent transform that maintain the independence among distinct groups. And we instantiate it by the proposed Group-Independent Swin-Block (GI Swin-Block). Experimental results demonstrate that our framework structures the bitstream with negligible cost, and exhibits superior performance on both visual quality and intelligent task supporting.
Ruoyu Feng 0001, Xin Jin 0014, Runsen Feng, Zhibo Chen 0001
ICCV3
2023 NaviNeRF: NeRF-based 3D Representation Disentanglement by Latent Semantic Navigation
abstract
3D representation disentanglement aims to identify, decompose, and manipulate the underlying explanatory factors of 3D data, which helps AI fundamentally understand our 3D world. This task is currently under-explored and poses great challenges: (i) the 3D representations are complex and in general contains much more information than 2D image; (ii) many 3D representations are not well suited for gradient-based optimization, let alone disentanglement. To address these challenges, we use NeRF as a differentiable 3D representation, and introduce a self-supervised Navigation to identify interpretable semantic directions in the latent space. To our best knowledge, this novel method, dubbed NaviNeRF, is the first work to achieve fine-grained 3D disentanglement without any priors or supervisions. Specifically, NaviNeRF is built upon the generative NeRF pipeline, and equipped with an Outer Navigation Branch and an Inner Refinement Branch. They are complementary —— the outer navigation is to identify global-view semantic directions, and the inner refinement dedicates to fine-grained attributes. A synergistic loss is further devised to coordinate two branches. Extensive experiments demonstrate that NaviNeRF has a superior fine-grained 3D disentanglement ability than the previous 3D-aware models. Its performance is also comparable to editing-oriented models relying on semantic or geometry priors.*
Baao Xie, Bohan Li 0015, Zequn Zhang, Junting Dong, Xin Jin 0014, Jing-Yu Yang 0002, Wenjun Zeng 0001
ICCV5
2023 ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion Generation
abstract
We present a GAN-based Transformer for general action-conditioned 3D human motion generation, including not only single-person actions but also multi-person interactive actions. Our approach consists of a powerful Action-conditioned motion TransFormer (ActFormer) under a GAN training scheme, equipped with a Gaussian Process latent prior. Such a design combines the strong spatio-temporal representation capacity of Transformer, superiority in generative modeling of GAN, and inherent temporal correlations from the latent prior. Furthermore, ActFormer can be naturally extended to multi-person motions by alternately modeling temporal correlations and human interactions with Transformer encoders. To further facilitate research on multi-person motion generation, we introduce a new synthetic dataset of complex multi-person combat behaviors. Extensive experiments on NTU-13, NTU RGB+D 120, BABEL and the proposed combat dataset show that our method can adapt to various human motion representations and achieve superior performance over the state-of-the-art methods on both single-person and multi-person motion generation tasks, demonstrating a promising step towards a general human motion generator. The project website can be found at https://liangxuy.github.io/actformer/.
Liang Xu 0012, Jing Su 0005, Zhicheng Fang, Chenjing Ding, Weihao Gan, Yichao Yan, Xin Jin 0014, Xiaokang Yang 0001, Wenjun Zeng 0001, Wei Wu 0021
ICCV9
2023 Automatic Intrinsic Reward Shaping for Exploration in Deep Reinforcement Learning
abstract
We present AIRS: **A**utomatic **I**ntrinsic **R**eward **S**haping that intelligently and adaptively provides high-quality intrinsic rewards to enhance exploration in reinforcement learning (RL). More specifically, AIRS selects shaping function from a predefined set based on the estimated task return in real-time, providing reliable exploration incentives and alleviating the biased objective problem. Moreover, we develop an intrinsic reward toolkit to provide efficient and reliable implementations of diverse intrinsic reward approaches. We test AIRS on various tasks of MiniGrid, Procgen, and DeepMind Control Suite. Extensive simulation demonstrates that AIRS can outperform the benchmarking schemes and achieve superior performance with simple architecture.
Mingqi Yuan, Bo Li 0037, Xin Jin 0014, Wenjun Zeng 0001
ICML3
2023 Composable Image Coding for Machine via Task-oriented Internal Adaptor and External Prior
abstract
Traditional image coding standards are typically optimized with a focus on human perception, which conflicts with the fact that most of the images are now analyzed by machines. To enable a variety of downstream intelligent tasks, contemporary approaches either utilize traditional codecs for image compression which are then used for task analysis, or develop a unified feature compression paradigm with deep learning techniques. However, they might suffer from accumulative errors and poor compatibility/generalization due to the conflict between standardized codecs and diverse machine tasks. We argue that a favorable image coding for machine (ICM) framework should have highly efficient adaptation capability, and take the ultimate task goals into account. Oriented at this, we propose a composable ICM solution dubbed Com-ICM, which develops plug-and-play lightweight internal adaptors injected into the codec architecture for efficient task transfer, and leverages off-the-shelf (large) models to provide external prior information for further task-oriented semantics learning. The internal adaptors (from the architectural aspect) and external priors (from the precondition aspect) complement each other, resulting in a mutually beneficial effect. We evaluate Com-ICM on diverse vision benchmarks, including image classification, object detection, and semantic segmentation, demonstrating its effectiveness and superiority. We are also actively submitting Com-ICM as a technical proposal to the international organization for standardization.
Jinming Liu 0001, Xin Jin 0014, Ruoyu Feng 0001, Zhibo Chen 0001, Wenjun Zeng 0001
VCIP2
2023 Semantical video coding: Instill static-dynamic clues into structured bitstream for AI tasks
Xin Jin 0014, Ruoyu Feng 0001, Simeng Sun, Runsen Feng, Tianyu He, Zhibo Chen 0001
J. Vis. Commun. Image Represent.1
2023 Extracting 3-D Structural Lines of Building From ALS Point Clouds Using Graph Neural Network Embedded With Corner Information
abstract
The representation quantifies the geometric shape and topology of a building is a necessary procedure for many urban planning applications. A sharp line framework is a high-level structural cue providing a compact building representation. However, accurate and efficient structural line extraction remains a challenging task given the variety and complexity of buildings. This study proposes a general 3-D structural line extraction method from point clouds. The building points are extracted and further divided into various single-building units. In the proposed 3-D structural line extraction method, individual building point cloud is the input. First, the corners are detected by an associative learning module. Next, the curve connection is implemented by a link prediction block based on the graph neural network (GNN) embedded with corner information. After that, the obtained curves are subsequently converted into a topological graph. Finally, the corner points are optimized to achieve precise fitting of the structural lines. The experiments and comparisons on two airborne laser scanning (ALS) point cloud datasets demonstrate the effectiveness of the proposed method and the ability to retrieve ideal structural line results for building point clouds. Furthermore, without reprocessing, the proposed method yielded better results for various dataset types (outdoor building, indoor scene, and furniture point clouds) than the prevalent published methods (i.e., EC-Net, PIE-Net, and PC2WF), verifying its strength and efficacy. To further verify the accuracy of the obtained structural lines, we also introduce a line-based model reconstruction method that employ these lines for building reconstruction.
Tengping Jiang, Zequn Zhang, Yongchao Yang, Xin Jin 0014, Wenjun Zeng 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 RailSeg: Learning Local-Global Feature Aggregation With Contextual Information for Railway Point Cloud Semantic Segmentation
abstract
Incomplete or outdated inventories of railway infrastructures may disrupt the railway sector’s administration and maintenance of transportation infrastructure, thus posing potential threats to the safety of traffic networks. Previous studies have adopted point clouds to accelerate inventory and inspection automation procedures. However, owing to the complexity of the railway scenes, previous studies reveal an imbalance between semantic richness, segmentation accuracy, and processing efficiency. This study aims to advance our understanding by providing a deep-learning framework for railway point cloud semantic segmentation. The proposed framework, named RailSeg, encompasses point cloud downsampling, integrated local-global feature extraction, spatial context aggregation, and semantic regularization. The proposed method, validated using point clouds collected in suburban and rural scenes, generates a point-level railway furniture inventory of 11 categories and achieves competitive performance in overall accuracy and mean intersection over union. In addition, RailSeg achieves better results than the baseline for additional types of point clouds (i.e., plateau railway mobile laser scanning (MLS) point clouds, street MLS point clouds, and urban-scale photogrammetric point clouds), demonstrating the superior generalization capabilities of RailSeg. This study may contribute to the development of 3D semantic segmentation, digital railway, and intelligent transportation.
Tengping Jiang, Bisheng Yang, Qinyu Zhang 0008, Xin Jin 0014, Wenjun Zeng 0001
IEEE Trans. Geosci. Remote. Sens.9
2023 Learning Cross-Scale Weighted Prediction for Efficient Neural Video Compression
abstract
Neural video codecs have demonstrated great potential in video transmission and storage applications. Existing neural hybrid video coding approaches rely on optical flow or Gaussian-scale flow for prediction, which cannot support fine-grained adaptation to diverse motion content. Towards more content-adaptive prediction, we propose a novel cross-scale prediction module that achieves more effective motion compensation. Specifically, on the one hand, we produce a reference feature pyramid as prediction sources and then transmit cross-scale flows that leverage the feature scale to control the precision of prediction. On the other hand, for the first time, a weighted prediction mechanism is introduced even if only a single reference frame is available, which can help synthesize a fine prediction result by transmitting cross-scale weight maps. In addition to the cross-scale prediction module, we further propose a multi-stage quantization strategy, which improves the rate-distortion performance with no extra computational penalty during inference. We show the encouraging performance of our efficient neural video codec (ENVC) on several benchmark datasets. In particular, the proposed ENVC can compete with the latest coding standard H.266/VVC in terms of sRGB PSNR on UVG dataset for the low-latency mode. We also analyze in detail the effectiveness of the cross-scale prediction module in handling various video content, and provide a comprehensive ablation study to analyze those important components. Test code is available at https://github.com/USTC-IMCL/ENVC.
Zongyu Guo, Runsen Feng, Zhizheng Zhang 0004, Xin Jin 0014, Zhibo Chen 0001
IEEE Trans. Image Process.4
2022 Reusing the Task-specific Classifier as a Discriminator: Discriminator-free Adversarial Domain Adaptation
abstract
Adversarial learning has achieved remarkable performances for unsupervised domain adaptation (UDA). Existing adversarial UDA methods typically adopt an additional discriminator to play the min-max game with a feature extractor. However, most of these methods failed to effectively leverage the predicted discriminative information, and thus cause mode collapse for generator. In this work, we address this problem from a different perspective and design a simple yet effective adversarial paradigm in the form of a discriminator-free adversarial learning network (DALN), wherein the category classifier is reused as a discriminator, which achieves explicit domain alignment and category distinguishment through a unified objective, enabling the DALN to leverage the predicted discriminative information for sufficient feature alignment. Basically, we introduce a Nuclear-norm Wasserstein discrepancy (NWD) that has definite guidance meaning for performing discrimination. Such NWD can be coupled with the classifier to serve as a discriminator satisfying the K-Lipschitz constraint without the requirements of additional weight clipping or gradient penalty strategy. Without bells and whistles, DALN compares favorably against the existing state-of-the-art (SOTA) methods on a variety of public datasets. Moreover, as a plug-and-play technique, NWD can be directly used as a generic regularizer to benefit existing UDA algorithms. Code is available at https://github.com/xiaoachen98/DALN.
Lin Chen 0026, Huaian Chen, Zhixiang Wei, Xin Jin 0014, Xiao Tan 0004, Yi Jin 0002, Enhong Chen
CVPR4
2022 Cloth-Changing Person Re-identification from A Single Image with Gait Prediction and Regularization
abstract
Cloth-Changing person re-identification (CC-ReID) aims at matching the same person across different locations over a long-duration, e.g., over days, and therefore inevitably has cases of changing clothing. In this paper, we focus on handling well the CC-ReID problem under a more challenging setting, i.e., just from a single image, which enables an efficient and latency-free person identity matching for surveillance. Specifically, we introduce Gait recognition as an auxiliary task to drive the Image ReID model to learn cloth-agnostic representations by leveraging personal unique and cloth-independent gait information, we name this framework as GI-ReID. GI-ReID adopts a two-stream architecture that consists of an image ReID-Stream and an auxiliary gait recognition stream (Gait-Stream). The Gait-Stream, that is discarded in the inference for high efficiency, acts as a regulator to encourage the ReID-Stream to capture cloth-invariant biometric motion features during the training. To get temporal continuous motion cues from a single image, we design a Gait Sequence Prediction (GSP) module for Gait-Stream to enrich gait information. Finally, a semantics consistency constraint over two streams is enforced for effective knowledge regularization. Extensive experiments on multiple image-based Cloth-Changing ReID benchmarks, e.g., LTCC, PRCC, Real28, and VC-Clothes, demonstrate that GI-ReID performs favorably against the state-of-the-art methods.
Xin Jin 0014, Tianyu He, Kecheng Zheng, Zhiheng Yin, Xu Shen 0001, Zhen Huang 0007, Ruoyu Feng 0001, Jianqiang Huang 0001, Zhibo Chen 0001, Xian-Sheng Hua 0001
CVPR1
2022 Unleashing Potential of Unsupervised Pre-Training with Intra-Identity Regularization for Person Re-Identification
abstract
Existing person reidentification (ReID) methods typically load the pretrained ImageNet weights for initialization directly. However, as a fine-grained classification task, ReID is more challenging and there exists a large domain gap between ImageNet classification. Inspired by the great success of self-supervised representation learning with contrastive objectives, in this paper, we design an Unsupervised Pretraining framework for reidentification (UP-ReID) based on the contrastive learning (CL) pipeline. During the pre-training, we attempt to address two critical issues for learning fine-grained ReID features: (1) the augmentations in the CL pipeline usually distort the discriminative clues in person images, and (2) the fine-grained local features of person images are not fully-explored. Therefore, we introduce an intra-identity (12-) regularization in the UP-ReID, which is instantiated as two constraints coming from the global image and local patch aspects, respectively. A global consistency constraint is enforced between augmented and original person images to increase robustness to augmentation, while an intrinsic contrastive constraint among local patches of each image is employed to fully explore the local discriminative clues. Extensive experiments on multiple popular reid datasets, PersonX, Market1501, CUHK03, and MSMT17, demonstrate that our UP-ReID pretrained model can significantly benefit the downstream ReID fine-tuning and achieve state-of-the-art performance.
Zizheng Yang, Xin Jin 0014, Kecheng Zheng, Feng Zhao 0004
CVPR2
2022 Image Coding for Machines with Omnipotent Feature Learning
Ruoyu Feng 0001, Xin Jin 0014, Zongyu Guo, Runsen Feng, Tianyu He, Zhizheng Zhang 0004, Simeng Sun, Zhibo Chen 0001
ECCV (37)2
2022 Learning with Recoverable Forgetting
Jingwen Ye, Yifang Fu, Jie Song 0011, Xingyi Yang, Songhua Liu, Xin Jin 0014, Mingli Song, Xinchao Wang
ECCV (11)6
2022 Meta Clustering Learning for Large-scale Unsupervised Person Re-identification
abstract
Unsupervised Person Re-identification (U-ReID) with pseudo labeling recently reaches a competitive performance compared to fully-supervised ReID methods based on modern clustering algorithms. However, such clustering-based scheme becomes computationally prohibitive for large-scale datasets, making it infeasible to be applied in real-world application. How to efficiently leverage endless unlabeled data with limited computing resources for better U-ReID is under-explored. In this paper, we make the first attempt to the large-scale U-ReID and propose a "small data for big task" paradigm dubbed Meta Clustering Learning (MCL). MCL only pseudo-labels a subset of the entire unlabeled data via clustering to save computing for the first-phase training. After that, the learned cluster centroids, termed as meta-prototypes in our MCL, are regarded as a proxy annotator to softly annotate the rest unlabeled data for further polishing the model. To alleviate the potential noisy labeling issue in the polishment phase, we enforce two well-designed loss constraints to promise intra-identity consistency and inter-identity strong correlation. For multiple widely-used U-ReID benchmarks, our method significantly saves computational cost while achieving a comparable or even better performance compared to prior works.
Xin Jin 0014, Tianyu He, Xu Shen 0001, Tongliang Liu, Xinchao Wang, Jianqiang Huang 0001, Zhibo Chen 0001, Xian-Sheng Hua 0001
ACM Multimedia1
2022 Deliberated Domain Bridging for Domain Adaptive Semantic Segmentation
abstract
In unsupervised domain adaptation (UDA), directly adapting from the source to the target domain usually suffers significant discrepancies and leads to insufficient alignment. Thus, many UDA works attempt to vanish the domain gap gradually and softly via various intermediate spaces, dubbed domain bridging (DB). However, for dense prediction tasks such as domain adaptive semantic segmentation (DASS), existing solutions have mostly relied on rough style transfer and how to elegantly bridge domains is still under-explored. In this work, we resort to data mixing to establish a deliberated domain bridging (DDB) for DASS, through which the joint distributions of source and target domains are aligned and interacted with each in the intermediate space. At the heart of DDB lies a dual-path domain bridging step for generating two intermediate domains using the coarse-wise and the fine-wise data mixing techniques, alongside a cross-path knowledge distillation step for taking two complementary models trained on generated intermediate samples as ‘teachers’ to develop a superior ‘student’ in a multi-teacher distillation manner. These two optimization steps work in an alternating way and reinforce each other to give rise to DDB with strong adaptation power. Extensive experiments on adaptive segmentation tasks with different settings demonstrate that our DDB significantly outperforms state-of-the-art methods.
Lin Chen 0026, Zhixiang Wei, Xin Jin 0014, Huaian Chen, Miao Zheng, Kai Chen 0026, Yi Jin 0002
NeurIPS3
2022 Learned Block-Based Hybrid Image Compression
abstract
Recent works on learned image compression perform encoding and decoding processes in a full-resolution manner, resulting in two problems when deployed for practical applications. First, parallel acceleration of the autoregressive entropy model cannot be achieved due to serial decoding. Second, full-resolution inference often causes the out-of-memory (OOM) problem with limited GPU resources, especially for high-resolution images. Block partition is a good choice to handle the above issues, but it brings about new challenges in reducing the redundancy between blocks and eliminating block effects. To tackle the above challenges, this paper provides a learned block-based hybrid image compression (LBHIC) framework. Specifically, we introduce explicit intra prediction into a learned image compression framework to utilize the relation among adjacent blocks. Superior to context modeling by linear weighting of neighbor pixels in traditional codecs, we propose a contextual prediction module (CPM) to better capture long-range correlations by utilizing the strip pooling to extract the most relevant information in neighboring latent space, thus achieving effective information prediction. Moreover, to alleviate blocking artifacts, we further propose a boundary-aware postprocessing module (BPM) with the edge importance taken into account. Extensive experiments demonstrate that the proposed LBHIC codec outperforms the VVC, with a bit-rate conservation of 4.1%, and reduces the decoding time by approximately 86.7% compared with that of state-of-the-art learned image compression methods.
Yaojun Wu 0001, Xin Li 0082, Zhizheng Zhang 0004, Xin Jin 0014, Zhibo Chen 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Dual Prior Learning for Blind and Blended Image Restoration
abstract
Unsupervised single image restoration approach, Deep Image Prior (DIP), aims to restore images by learning enough raw image statistic priors from the corrupted observation. However, it is not uncommon that an image is contaminated by the multiple unknown distortions. Thus it is hard to disentangle the clean and the hybrid distortion signals by solely relying on image prior learning to restore the images. To overcome this problem, we propose the Dual Prior Learning (DPL) method by taking both image and distortion priors into account. DPL goes beyond DIP by considering an additional step to explicitly learn the blended distortion prior. Furthermore, to coordinate the learning of two priors and avoid them learning the same knowledge, we exploit unpaired training data to enforce a weakly supervision in an adversarial manner to encourage disentangling two priors. Extensive experiments show the effectiveness and appealing performance of the proposed DPL on restoring images with challenging unknown blended distortions.
Xin Jin 0014, Li Zhang 0040, Chaowei Shan, Xin Li 0082, Zhibo Chen 0001
IEEE Trans. Image Process.1
2022 Style Normalization and Restitution for Domain Generalization and Adaptation
abstract
For many computer vision applications, the learned models usually have high performance on the training datasets but suffer from significant performance degradation when deployed in new environments, where there are usually style differences between the training images and the testing images. For high-level vision tasks, an effective domain generalizable model is expected to be able to learn feature representations that are both generalizable and discriminative. In this paper, we design a novel Style Normalization and Restitution module (SNR) to simultaneously ensure high generalization and discrimination capability of the networks. In SNR, particularly, we filter out the style variations (e.g., illumination, color contrast) by performing Instance Normalization (IN) to obtain style normalized features, where the discrepancy among different samples/domains is reduced. However, such a process is task-ignorant and inevitably removes some task-relevant discriminative information, which may hurt the performance. To remedy this, we propose to distill task-relevant discriminative features from the residual (i.e., the difference between the original feature and the style normalized feature) and add them back to the network to ensure high discrimination. Moreover, for better disentanglement, we enforce a dual restitution loss constraint to encourage the better separation of task-relevant and task-irrelevant features. We validate the effectiveness of our SNR on different vision tasks, including classification, semantic segmentation, and object detection. Experiments demonstrate that our SNR is capable of improving the performance of networks for domain generalization (DG) and unsupervised domain adaptation (UDA).
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001
IEEE Trans. Multim.1
2021 Learning Omni-Frequency Region-adaptive Representations for Real Image Super-Resolution
abstract
Traditional single image super-resolution (SISR) methods that focus on solving single and uniform degradation (i.e., bicubic down-sampling), typically suffer from poor performance when applied into real-world low-resolution (LR) images due to the complicated realistic degradations. The key to solving this more challenging real image super-resolution (RealSR) problem lies in learning feature representations that are both informative and content-aware. In this paper, we propose a Omni-frequency Region-adaptive Network (OR-Net) to address both challenges, here we call features of all low, middle and high frequencies omni-frequency features. Specifically, we start from the frequency perspective and design a Frequency Decomposition (FD) module to separate different frequency components to comprehensively compensate the information lost for real LR image. Then, considering the different regions of real LR image have different frequency information lost, we further design a Region-adaptive Frequency Aggregation (RFA) module by leveraging dynamic convolution and spatial attention to adaptively restore frequency components for different regions. The extensive experiments endorse the high-efficient, effective, and scenario-agnostic nature of our OR-Net for RealSR.
Xin Li 0082, Xin Jin 0014, Tao Yu 0012, Simeng Sun, Yingxue Pang, Zhizheng Zhang 0004, Zhibo Chen 0001
AAAI2
2021 Re-energizing Domain Discriminator with Sample Relabeling for Adversarial Domain Adaptation
abstract
Many unsupervised domain adaptation (UDA) methods exploit domain adversarial training to align the features to reduce domain gap, where a feature extractor is trained to fool a domain discriminator in order to have aligned feature distributions. The discrimination capability of the domain classifier w.r.t. the increasingly aligned feature distributions deteriorates as training goes on, thus cannot effectively further drive the training of feature extractor. In this work, we propose an efficient optimization strategy named Re-enforceable Adversarial Domain Adaptation (RADA) which aims to re-energize the domain discriminator during the training by using dynamic domain labels. Particularly, we relabel the well aligned target domain samples as source domain samples on the fly. Such relabeling makes the less separable distributions more separable, and thus leads to a more powerful domain classifier w.r.t. the new data distributions, which in turn further drives feature alignment. Extensive experiments on multiple UDA benchmarks demonstrate the effectiveness and superiority of our RADA.
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001
ICCV1
2021 Dense Interaction Learning for Video-based Person Re-identification
abstract
Video-based person re-identification (re-ID) aims at matching the same person across video clips. Efficiently exploiting multi-scale fine-grained features while building the structural interaction among them is pivotal for its success. In this paper, we propose a hybrid framework, Dense Interaction Learning (DenseIL), that takes the principal advantages of both CNN-based and Attention-based architectures to tackle video-based person re-ID difficulties. DenseIL contains a CNN encoder and a Dense Interaction (DI) decoder. The CNN encoder is responsible for efficiently extracting discriminative spatial features while the DI decoder is designed to densely model spatial-temporal inherent interaction across frames. Different from previous works, we additionally let the DI decoder densely attends to intermediate fine-grained CNN features and that naturally yields multi-grained spatial-temporal representation for each video clip. Moreover, we introduce Spatio-TEmporal Positional Embedding (STEP-Emb) into the DI decoder to investigate the positional relation among the spatial-temporal inputs. Our experiments consistently and significantly outperform all the state-of-the-art methods on multiple standard video-based person re-ID datasets.
Tianyu He, Xin Jin 0014, Xu Shen 0001, Jianqiang Huang 0001, Zhibo Chen 0001, Xian-Sheng Hua 0001
ICCV2
2021 CASINet: Content-Adaptive Scale Interaction Networks for scene parsing
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhizheng Zhang 0004, Zhibo Chen 0001
Neurocomputing1
2021 Multi-task Learning-based All-in-one Collaboration Framework for Degraded Image Super-resolution
abstract
In this article, we address the degraded image super-resolution problem in a multi-task learning (MTL) manner. To better share representations between multiple tasks, we propose an all-in-one collaboration framework (ACF) with a learnable “junction” unit to handle two major problems that exist in MTL—“How to share” and “How much to share.” Specifically, ACF consists of a sharing phase and a reconstruction phase. Considering the intrinsic characteristic of multiple image degradations, we propose to first deal with the compression artifact, motion blur, and spatial structure information of the input image in parallel under a three-branch architecture in the sharing phase. Subsequently, in the reconstruction phase, we up-sample the previous features for high-resolution image reconstruction with a channel-wise and spatial attention mechanism. To coordinate two phases, we introduce a learnable “junction” unit with a dual-voting mechanism to selectively filter or preserve shared feature representations that come from sharing phase, learning an optimal combination for the following reconstruction phase. Finally, a curriculum learning-based training scheme is further proposed to improve the convergence of the whole framework. Extensive experimental results on synthetic and real-world low-resolution images show that the proposed all-in-one collaboration framework not only produces favorable high-resolution results while removing serious degradation, but also has high computational efficiency, outperforming state-of-the-art methods. We also have applied ACF to some image-quality sensitive practical task, such as pose estimation, to improve estimation accuracy of low-resolution images.
Xin Jin 0014, Kazuyuki Tasaka, Zhibo Chen 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2020 Uncertainty-Aware Multi-Shot Knowledge Distillation for Image-Based Object Re-Identification
abstract
Object re-identification (re-id) aims to identify a specific object across times or camera views, with the person re-id and vehicle re-id as the most widely studied applications. Re-id is challenging because of the variations in viewpoints, (human) poses, and occlusions. Multi-shots of the same object can cover diverse viewpoints/poses and thus provide more comprehensive information. In this paper, we propose exploiting the multi-shots of the same identity to guide the feature learning of each individual image. Specifically, we design an Uncertainty-aware Multi-shot Teacher-Student (UMTS) Network. It consists of a teacher network (T-net) that learns the comprehensive features from multiple images of the same object, and a student network (S-net) that takes a single image as input. In particular, we take into account the data dependent heteroscedastic uncertainty for effectively transferring the knowledge from the T-net to S-net. To the best of our knowledge, we are the first to make use of multi-shots of an object in a teacher-student learning manner for effectively boosting the single image based re-id. We validate the effectiveness of our approach on the popular vehicle re-id and person re-id datasets. In inference, the S-net alone significantly outperforms the baselines and achieves the state-of-the-art performance.
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001
AAAI1
2020 Semantics-Aligned Representation Learning for Person Re-Identification
abstract
Person re-identification (reID) aims to match person images to retrieve the ones with the same identity. This is a challenging task, as the images to be matched are generally semantically misaligned due to the diversity of human poses and capture viewpoints, incompleteness of the visible bodies (due to occlusion), etc. In this paper, we propose a framework that drives the reID network to learn semantics-aligned feature representation through delicate supervision designs. Specifically, we build a Semantics Aligning Network (SAN) which consists of a base network as encoder (SA-Enc) for re-ID, and a decoder (SA-Dec) for reconstructing/regressing the densely semantics aligned full texture image. We jointly train the SAN under the supervisions of person re-identification and aligned texture generation. Moreover, at the decoder, besides the reconstruction loss, we add Triplet ReID constraints over the feature maps as the perceptual losses. The decoder is discarded in the inference and thus our scheme is computationally efficient. Ablation studies demonstrate the effectiveness of our design. We achieve the state-of-the-art performances on the benchmark datasets CUHK03, Market1501, MSMT17, and the partial person reID dataset Partial REID.
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Guoqiang Wei, Zhibo Chen 0001
AAAI1
2020 Region Normalization for Image Inpainting
abstract
Feature Normalization (FN) is an important technique to help neural network training, which typically normalizes features across spatial dimensions. Most previous image inpainting methods apply FN in their networks without considering the impact of the corrupted regions of the input image on normalization, e.g. mean and variance shifts. In this work, we show that the mean and variance shifts caused by full-spatial FN limit the image inpainting network training and we propose a spatial region-wise normalization named Region Normalization (RN) to overcome the limitation. RN divides spatial pixels into different regions according to the input mask, and computes the mean and variance in each region for normalization. We develop two kinds of RN for our image inpainting network: (1) Basic RN (RN-B), which normalizes pixels from the corrupted and uncorrupted regions separately based on the original inpainting mask to solve the mean and variance shift problem; (2) Learnable RN (RN-L), which automatically detects potentially corrupted and uncorrupted regions for separate normalization, and performs global affine transformation to enhance their fusion. We apply RN-B in the early layers and RN-L in the latter layers of the network respectively. Experiments show that our method outperforms current state-of-the-art methods quantitatively and qualitatively. We further generalize RN to other inpainting networks and achieve consistent performance improvements.
Tao Yu 0012, Zongyu Guo, Xin Jin 0014, Shilin Wu, Zhibo Chen 0001, Weiping Li 0003, Zhizheng Zhang 0004, Sen Liu 0001
AAAI3
2020 Style Normalization and Restitution for Generalizable Person Re-Identification
abstract
Existing fully-supervised person re-identification (ReID) methods usually suffer from poor generalization capability caused by domain gaps. The key to solving this problem lies in filtering out identity-irrelevant interference and learning domain-invariant person representations. In this paper, we aim to design a generalizable person ReID framework which trains a model on source domains yet is able to generalize/perform well on target domains. To achieve this goal, we propose a simple yet effective Style Normalization and Restitution (SNR) module. Specifically, we filter out style variations (e.g., illumination, color contrast) by Instance Normalization (IN). However, such a process inevitably removes discriminative information. We propose to distill identity-relevant feature from the removed information and restitute it to the network to ensure high discrimination. For better disentanglement, we enforce a dual causal loss constraint in SNR to encourage the separation of identity-relevant features and identity-irrelevant features. Extensive experiments demonstrate the strong generalization capability of our framework. Our models empowered by the SNR modules significantly outperform the state-of-the-art domain generalization approaches on multiple widely-used person ReID benchmarks, and also show superiority on unsupervised domain adaptation.
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001, Li Zhang 0040
CVPR1
2020 Relation-Aware Global Attention for Person Re-Identification
abstract
For person re-identification (re-id), attention mechanisms have become attractive as they aim at strengthening discriminative features and suppressing irrelevant ones, which matches well the key of re-id, i.e., discriminative feature learning. Previous approaches typically learn attention using local convolutions, ignoring the mining of knowledge from global structure patterns. Intuitively, the affinities among spatial positions/nodes in the feature map provide clustering-like information and are helpful for inferring semantics and thus attention, especially for person images where the feasible human poses are constrained. In this work, we propose an effective Relation-Aware Global Attention (RGA) module which captures the global structural information for better attention learning. Specifically, for each feature position, in order to compactly grasp the structural information of global scope and local appearance information, we propose to stack the relations, i.e., its pairwise correlations/affinities with all the feature positions (e.g., in raster scan order), and the feature itself together to learn the attention with a shallow convolutional model. Extensive ablation studies demonstrate that our RGA can significantly enhance the feature representation power and help achieve the state-of-the-art performance on several popular benchmarks. The source code is available at https://github.com/microsoft/Relation-Aware-Global-Attention-Networks.
Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Xin Jin 0014, Zhibo Chen 0001
CVPR4
2020 Global Distance-Distributions Separation for Unsupervised Person Re-identification
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001
ECCV (7)1
2020 Learning Disentangled Feature Representation for Hybrid-Distorted Image Restoration
Xin Li 0082, Xin Jin 0014, Sen Liu 0001, Yaojun Wu 0001, Tao Yu 0012, Wei Zhou 0021, Zhibo Chen 0001
ECCV (29)2
2020 AI-GAN: Asynchronous interactive generative adversarial network for single image rain removal
Xin Jin 0014, Zhibo Chen 0001, Weiping Li 0003
Pattern Recognit.1
2020 Learning for Video Compression
abstract
One key challenge to learning-based video compression is that motion predictive coding, a very effective tool for video compression, can hardly be trained into a neural network. In this paper, we propose the concept of Pixel-MotionCNN (PMCNN) which includes motion extension and hybrid prediction networks. PMCNN can model spatiotemporal coherence to effectively perform predictive coding inside the learning network. On the basis of PMCNN, we further explore a learning-based framework for video compression with additional components of iterative analysis/synthesis and binarization. The experimental results demonstrate the effectiveness of the proposed scheme. Although entropy coding and complex configurations are not employed in this paper, we still demonstrate superior performance compared with MPEG-2 and achieve comparable results with H.264 codec. The proposed learning-based scheme provides a possible new direction to further improve compression efficiency and functionalities of future video coding.
Zhibo Chen 0001, Tianyu He, Xin Jin 0014, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2019 Unsupervised Single Image Deraining with Self-Supervised Constraints
abstract
Most existing single image deraining methods require learning supervised models from a large set of paired synthetic training data, which limits their generality and practicality in real-world multimedia applications. Besides, due to lack of labeled-supervised constraints, directly applying existing unsupervised frameworks to the image deraining task will suffer from low-quality recovery. Therefore, we propose an Unsupervised Deraining Generative Adversarial Network (UD-GAN) to tackle above problems by introducing self-supervised constraints from the intrinsic statistics of unpaired rainy and clean images. Specifically, we design two collaboratively optimized modules, namely Rain Guidance Module (RGM) and Background Guidance Module (BGM), to take full advantage of rainy image characteristics. UD-GAN outperforms state-of-the-art methods on various benchmarking datasets in both quantitative and qualitative comparisons.
Xin Jin 0014, Zhibo Chen 0001, Wei Zhou 0021
ICIP1
2018 A Decomposed Dual-Cross Generative Adversarial Network for Image Rain Removal
Xin Jin 0014, Zhibo Chen 0001, Jiale Chen 0001, Wei Zhou 0021, Chaowei Shan
BMVC1
2018 Augmented Coarse-to-Fine Video Frame Synthesis with Semantic Loss
Xin Jin 0014, Zhibo Chen 0001, Sen Liu 0001, Wei Zhou 0021
PRCV (1)1
2018 Multiscale Progressive Image Compression Network Guided by Learnable Just Noticeable Distortion
abstract
One key challenge to the learning-based image compression is that adaptive bit allocation is crucial for compression effectiveness but can hardly be trained into a neural network. Hereby, in this work, We presents an end-to-end trainable image compression framework, named Multi-scale Progressive Network (MPN) to achieve spatially variant bit allocation and rate control through the guidance of a novel learnable just noticeable distortion (JND) map. Specifically, MPN's encoder archives multi-scale feature representation through a three-branched structure. Each branch employs an independent feature extraction strategy for the specific receptive field and merge progressively under the guidance of corresponding learnable JND maps that generated by our proposed Bit-Allocation sub-Network (BAN), which make MPN focus on the areas where attract the human visual system (HVS) and preserve more texture of the image during the compression procedure. Finally, a hybrid objective function is introduced to further make MPN more efficient and mimic the discriminative characteristics of the human visual system (HVS). Experiments show that MPN significantly outperforms traditional JPEG, JPEG 2000 and few state-of-art learning-based methods by multi-scale structural similarity (MS-SSIM) index, and has the ability to produce the much better visual result with rich textures, sharp edges, and fewer artifacts.
Xin Jin 0014, Runchun Ye, Zhibo Chen 0001
VCIP1
2005 H.264-compatible spatially scalable video coding with in-band prediction
abstract
In this paper, a H.264 compatible spatially scalable video coding method with in-band prediction is proposed which taking advantages from both the high coding efficiency of H.264 coding scheme and the attractive performance of in-band overcomplete discrete wavelet transform (ODWT) in wavelet-domain motion estimation and motion compensation. Four MV prediction modes are proposed for INTER prediction of high frequency subbands. The intra prediction modes of H.264 are also simplified for each high band according to the directional features inherited inside. Finally, a H.264 compatible scheme based on one of the MV prediction modes is presented to provide better tradeoff among standard compatibility, low complexity and high performance.
Xin Jin 0014, Xiaoyan Sun 0001, Feng Wu 0001, Guangxi Zhu, Shipeng Li 0001
ICIP (1)1