Jian Jia

dblp:89/2939 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 9 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 B3CT: Three-Branch Learning with Unlabeled Target Signals for Domain-Robust Semantic Segmentation
Xin Zhao 0012, Jian Jia, Junyan Wang 0001, Lijun Cao
Int. J. Comput. Vis.3
2026 ASR-Enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval
abstract
E-commerce is increasinglymultimedia-enriched, with products exhibited in a broad-domain manner as images, short videos, or live stream promotions. A unified and vectorized cross-domain production representation is essential. Due to large intra-product variance and high inter-product similarity in the broad-domain scenario, a visual-only representation is inadequate. While Automatic Speech Recognition (ASR) text derived from the short or live-stream videos is readily accessible, how to de-noise the excessively noisy text for multimodal representation learning is mostly untouched. We proposeASR-enhancedMultimodalProduct Representation Learning (AMPere). In order to extract product-specific information from the raw ASR text,AMPereuses an easy-to-implement LLM-based ASR text summarizer. The LLM-summarized text, together with visual data, is then fed into a multi-branch network to generate compact multimodal embeddings. Extensive experiments on a large-scale tri-domain dataset verify the effectiveness ofAMPerein obtaining a unified multimodal product representation that clearly improves cross-domain product retrieval.
Ruixiang Zhao, Jian Jia, Yan Li 0043, Xuehan Bai, Quan Chen 0006, Han Li 0005, Peng Jiang 0002, Xirong Li 0001
IEEE Trans. Multim.2
2025 LEARN: Knowledge Adaptation from Large Language Model to Recommendation for Practical Industrial Application
abstract
Contemporary recommendation systems predominantly rely on ID embedding to capture latent associations among users and items. However, this approach overlooks the wealth of semantic information embedded within textual descriptions of items, leading to suboptimal performance and poor generalizations. Leveraging the capability of large language models to comprehend and reason about textual content presents a promising avenue for advancing recommendation systems. To achieve this, we propose an Llm-driven knowlEdge Adaptive RecommeNdation (LEARN) framework that synergizes open-world knowledge with collaborative knowledge. We address computational complexity concerns by utilizing pretrained LLMs as item encoders and freezing LLM parameters to avoid catastrophic forgetting and preserve open-world knowledge. To bridge the gap between the open-world and collaborative domains, we design a twin-tower structure supervised by the recommendation task and tailored for practical industrial application. Through experiments on the real large-scale industrial dataset and online A/B tests, we demonstrate the efficacy of our approach in industry application. We also achieve state-of-the-art performance on six Amazon Review datasets to verify the superiority of our method.
Jian Jia, Yan Li 0043, Honggang Chen, Xuehan Bai, Zhaocheng Liu, Jian Liang 0001, Quan Chen 0006, Han Li 0005, Peng Jiang 0002, Kun Gai
AAAI1
2025 SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
Zhentao Tan, Ben Xue, Jian Jia, Wencai Ye, Shaoyun Shi, Mingjie Sun, Wenjin Wu, Quan Chen 0006, Peng Jiang 0002
ICCV3
2025 MOTION: Multi-object Video Editing with Training-Free Attention Guidance
Qitong Yan, Jian Jia, Bo Wang 0071, Quan Chen 0006, Peng Jiang 0002, Minfeng Zhu 0001, Linchao Zhu, Wei Chen 0001
ICIC (18)2
2025 Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
abstract
We introduce Orthus, a unified multimodal model that excels in generating interleaved images and text from mixed-modality inputs by simultaneously handling discrete text tokens and continuous image features under the AR modeling principle. The continuous treatment of visual signals minimizes the information loss while the fully AR formulation renders the characterization of the correlation between modalities straightforward. Orthus leverages these advantages through its modality-specific heads—one regular language modeling (LM) head predicts discrete text tokens and one diffusion head generates continuous image features. We devise an efficient strategy for building Orthus—by substituting the Vector Quantization (VQ) operation in the existing unified AR model with a soft alternative, introducing a diffusion head, and tuning the added modules to reconstruct images, we can create an Orthus-base model effortlessly (e.g., within 72 A100 GPU hours). Orthus-base can further embrace post-training to craft lengthy interleaved image-text, reflecting the potential for handling intricate real-world tasks. For visual understanding and generation, Orthus achieves a GenEval score of 0.58 and an MME-P score of 1265.8 using 7B parameters, outperforming competing baselines including Show-o and Chameleon.
Siqi Kou, Jiachun Jin, Jian Jia, Quan Chen 0006, Peng Jiang 0002, Zhijie Deng
ICML6
2024 Spatiotemporal Fine-grained Video Description for Short Videos
abstract
In the mobile internet era, short videos are inundating people's lives. However, research on visual language models specifically designed for short videos has not yet received sufficient attention. Short videos are not just videos of limited duration. The prominent visual details and high information density of short videos differentiate them to long videos. In this paper, we propose the SpatioTemporal Fine-grained Description (STFVD) emphasizing on the uniqueness of short videos, which entails capturing the intricate details of the main subject and fine-grained movements. To this end, we create a comprehensive Short Video Advertisements Description (SVAD) dataset, comprising 34,930 clips from 5,046 videos. The dataset covers a range of topics, including 191 sub-industries, 649 popular products, and 470 trending games. Various efforts have been made in the data annotation process to ensure the inclusion of fine-grained spatiotemporal information, resulting in 34,930 high-quality annotations. Compared to existing datasets, samples in SVAD exhibit a superior text information density, suggesting that SVAD is more appropriate for the analysis of short videos. Based on the SVAD dataset, we develop a visual language model (SVAD-VLM) to generate spatiotemporal fine-grained description for short videos. We use a prompt-guided keyword generation task to efficiently learn key visual information. Moreover, we also utilize dual visual alignment to exploit the advantage of mixed-datasets training. Experiments on SVAD dataset demonstrate the challenge of STFVD and the competitive performance of proposed method compared to previous ones.
Te Yang, Jian Jia, Bo Wang 0071, Yanhua Cheng, Yan Li 0043, Dongze Hao, Xipeng Cao, Quan Chen 0006, Han Li 0005, Peng Jiang 0002, Xiangyu Zhu 0001, Zhen Lei 0001
ACM Multimedia2
2023 Beyond Appearance: A Semantic Controllable Self-Supervised Learning Framework for Human-Centric Visual Tasks
abstract
Human-centric visual tasks have attracted increasing research attention due to their widespread applications. In this paper, we aim to learn a general human representation from massive unlabeled human images which can benefit downstream human-centric tasks to the maximum extent. We call this method SOLIDER, a Semantic cOntrollable seLf-supervIseD lEaRning framework. Unlike the existing self-supervised learning methods, prior knowledge from human images is utilized in SOLIDER to build pseudo semantic labels and import more semantic information into the learned representation. Meanwhile, we note that different downstream tasks always require different ratios of semantic information and appearance information. For example, human parsing requires more semantic information, while person re-identification needs more appearance information for identification purpose. So a single learned representation cannot fit for all requirements. To solve this problem, SOLIDER introduces a conditional network with a semantic controller. After the model is trained, users can send values to the controller to produce representations with different ratios of semantic information, which can fit different needs of downstream tasks. Finally, SOLIDER is verified on six downstream human-centric visual tasks. It outperforms state of the arts and builds new baselines for these tasks. The code is released in https://github.com/tinyvision/SOLIDER.
Jian Jia, Hao Luo 0004, Fan Wang 0019, Rong Jin 0001, Xiuyu Sun
CVPR3
2023 An Improved Global-Local Fusion Network for Depression Detection Telemedicine Framework
abstract
In recent years, remote depression detection is becoming a new favorite of the Internet of Things because of its promise. However, deployable and sensitive information security has not been properly addressed. Thus, a novel depression detection framework is proposed to successfully solve the above problems. This article concentrates solely on the development and execution of our proposed core algorithm, LKCT. Further details concerning the system’s construction will be disclosed separately as future work. The contribution of this article can be summarized as follows: 1) proposing a depression detection framework that is easy to promote and can effectively solve the security of sensitive information; 2) proposing a novel two-stream deep network LKCT based on the global–local concept; 3) designing a novel network called purified vision transformer (PIT) to enhance the model’s ability to capture local details in images; and 4) introducing decoupled knowledge distillation to distill LKCT into the lightweight network MobileNet for deployability. To evaluate the true performance of our proposed method, we conducted sufficient experiments on the Chinese Academy of Sciences Institute of Automation (CASIA), facial expression recognition challenge (FERC), AVEC2014, and private depression data sets. LKCT achieved an accuracy of 93.4% on CASIA, demonstrating its strong ability to distinguish facial features. The accuracy on FERC reached 99%, proving its ability to accurately distinguish simple emotions. On AVEC2014, the model achieved state-of-the-art results with an RMSE of 7.41 and MAE of 5.48, validating its ability to identify depression. The distilled model obtained an average accuracy of 83% on the private data set, indicating good generalization ability.
Jian Zhao 0002, Jian Jia, Xianjia Meng
IEEE Internet Things J.4
2022 QueryProp: Object Query Propagation for High-Performance Video Object Detection
abstract
Video object detection has been an important yet challenging topic in computer vision. Traditional methods mainly focus on designing the image-level or box-level feature propagation strategies to exploit temporal information. This paper argues that with a more effective and efficient feature propagation framework, video object detectors can gain improvement in terms of both accuracy and speed. For this purpose, this paper studies object-level feature propagation, and proposes an object query propagation (QueryProp) framework for high-performance video object detection. The proposed QueryProp contains two propagation strategies: 1) query propagation is performed from sparse key frames to dense non-key frames to reduce the redundant computation on non-key frames; 2) query propagation is performed from previous key frames to the current key frame to improve feature representation by temporal context modeling. To further facilitate query propagation, an adaptive propagation gate is designed to achieve flexible key frame selection. We conduct extensive experiments on the ImageNet VID dataset. QueryProp achieves comparable accuracy with state-of-the-art methods and strikes a decent accuracy/speed trade-off.
Naiyu Gao, Jian Jia, Xin Zhao 0012, Kaiqi Huang
AAAI3
2022 Learning Disentangled Attribute Representations for Robust Pedestrian Attribute Recognition
abstract
Although various methods have been proposed for pedestrian attribute recognition, most studies follow the same feature learning mechanism, \ie, learning a shared pedestrian image feature to classify multiple attributes. However, this mechanism leads to low-confidence predictions and non-robustness of the model in the inference stage. In this paper, we investigate why this is the case. We mathematically discover that the central cause is that the optimal shared feature cannot maintain high similarities with multiple classifiers simultaneously in the context of minimizing classification loss. In addition, this feature learning mechanism ignores the spatial and semantic distinctions between different attributes. To address these limitations, we propose a novel disentangled attribute feature learning (DAFL) framework to learn a disentangled feature for each attribute, which exploits the semantic and spatial characteristics of attributes. The framework mainly consists of learnable semantic queries, a cascaded semantic-spatial cross-attention (SSCA) module, and a group attention merging (GAM) module. Specifically, based on learnable semantic queries, the cascaded SSCA module iteratively enhances the spatial localization of attribute-related regions and aggregates region features into multiple disentangled attribute features, used for classification and updating learnable semantic queries. The GAM module splits attributes into groups based on spatial distribution and utilizes reliable group attention to supervise query attention maps. Experiments on PETA, RAPv1, PA100k, and RAPv2 show that the proposed method performs favorably against state-of-the-art methods.
Jian Jia, Naiyu Gao, Xiaotang Chen, Kaiqi Huang
AAAI1
2022 PanopticDepth: A Unified Framework for Depth-aware Panoptic Segmentation
abstract
This paper presents a unified framework for depth-aware panoptic segmentation (DPS), which aims to reconstruct 3D scene with instance-level semantics from one single image. Prior works address this problem by simply adding a dense depth regression head to panoptic segmentation (PS) networks, resulting in two independent task branches. This neglects the mutually-beneficial relations between these two tasks, thus failing to exploit handy instance-level semantic cues to boost depth accuracy while also producing sub-optimal depth maps. To overcome these limitations, we propose a unified framework for the DPS task by applying a dynamic convolution technique to both the PS and depth prediction tasks. Specifically, instead of predicting depth for all pixels at a time, we generate instance-specific kernels to predict depth and segmentation masks for each instance. Moreover, leveraging the instance-wise depth estimation scheme, we add additional instance-level depth cues to assist with supervising the depth learning via a new depth loss. Extensive experiments on Cityscapes-DPS and SemKITTI-DPS show the effectiveness and promise of our method. We hope our unified solution to DPS can lead a new paradigm in this area. Code is available at https://github.com/NaiyuGao/PanopticDepth.
Naiyu Gao, Jian Jia, Yanhu Shan, Xin Zhao 0012, Kaiqi Huang
CVPR3
2022 InsPro: Propagating Instance Query and Proposal for Online Video Instance Segmentation
abstract
Video instance segmentation (VIS) aims at segmenting and tracking objects in videos. Prior methods typically generate frame-level or clip-level object instances first and then associate them by either additional tracking heads or complex instance matching algorithms. This explicit instance association approach increases system complexity and fails to fully exploit temporal cues in videos. In this paper, we design a simple, fast and yet effective query-based framework for online VIS. Relying on an instance query and proposal propagation mechanism with several specially developed components, this framework can perform accurate instance association implicitly. Specifically, we generate frame-level object instances based on a set of instance query-proposal pairs propagated from previous frames. This instance query-proposal pair is learned to bind with one specific object across frames through conscientiously developed strategies. When using such a pair to predict an object instance on the current frame, not only the generated instance is automatically associated with its precursors on previous frames, but the model gets a good prior for predicting the same object. In this way, we naturally achieve implicit instance association in parallel with segmentation and elegantly take advantage of temporal clues in videos. To show the effectiveness of our method InsPro, we evaluate it on two popular VIS benchmarks, i.e., YouTube-VIS 2019 and YouTube-VIS 2021. Without bells-and-whistles, our InsPro with ResNet-50 backbone achieves 43.2 AP and 37.6 AP on these two benchmarks respectively, outperforming all other online VIS methods.
Naiyu Gao, Jian Jia, Yanhu Shan, Xin Zhao 0012, Kaiqi Huang
NeurIPS4
2021 Spatial and Semantic Consistency Regularizations for Pedestrian Attribute Recognition
abstract
While recent studies on pedestrian attribute recognition have shown remarkable progress in leveraging complicated networks and attention mechanisms, most of them neglect the inter-image relations and an important prior: spatial consistency and semantic consistency of attributes under surveillance scenarios. The spatial locations of the same attribute should be consistent between different pedestrian images, e.g., the "hat" attribute and the "boots" attribute are always located at the top and bottom of the picture respectively. In addition, the inherent semantic feature of the "hat" attribute should be consistent, whether it is a baseball cap, beret, or helmet. To fully exploit inter-image relations and aggregate human prior in the model learning process, we construct a Spatial and Semantic Consistency (SSC) framework that consists of two complementary regularizations to achieve spatial and semantic consistency for each attribute. Specifically, we first propose a spatial consistency regularization to focus on reliable and stable attribute-related regions. Based on the precise attribute locations, we further propose a semantic consistency regularization to extract intrinsic and discriminative semantic features. We conduct extensive experiments on popular benchmarks including PA100K, RAP, and PETA. Results show that the proposed method performs favorably against state- of-the-art methods without increasing parameters.
Jian Jia, Xiaotang Chen, Kaiqi Huang
ICCV1
2014 Multi-focus image fusion based on non-negative matrix factorization and difference images
Jian Jia, Zhihua Zhao 0001
Signal Process.3