Jiajun Ding

dblp:178/9437 · DBLP profile ↗
← Back
29ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0002-7497-7485ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SRSplat: Feed-Forward Super-Resolution Gaussian Splatting from Sparse Multi-View Images
abstract
Feed-forward 3D reconstruction from sparse, low-resolution (LR) images is a crucial capability for real-world applications, such as autonomous driving and embodied AI. However, existing methods often fail to recover fine texture details. This limitation stems from the inherent lack of high-frequency information in LR inputs. To address this, we propose SRSplat, a feed-forward framework that reconstructs high-resolution 3D scenes from only a few LR views. Our main insight is to compensate for the deficiency of texture information by jointly leveraging external high-quality reference images and internal texture cues. We first construct a scene-specific reference gallery, generated for each scene using Multimodal Large Language Models (MLLMs) and diffusion models. To integrate this external information, we introduce the Reference-Guided Feature Enhancement (RGFE) module, which aligns and fuses features from the LR input images and their reference twin image. Subsequently, we train a decoder to predict the Gaussian primitives using the multi-view fused feature obtained from RGFE. To further refine predicted Gaussian primitives, we introduce Texture-Aware Density Control (TADC), which adaptively adjusts Gaussian density based on the internal texture richness of the LR inputs. Extensive experiments demonstrate that our SRSplat outperforms existing methods on various datasets, including RealEstate10K, ACID, and DTU, and exhibits strong cross-dataset and cross-resolution generalization capabilities.
Changyue Shi, Chuxiao Yang, Jiajun Ding, Zhou Yu 0001, Min Tan 0005
AAAI5
2026 Sparse4DGS: 4D Gaussian Splatting for Sparse-Frame Dynamic Scene Reconstruction
abstract
Dynamic Gaussian Splatting approaches have achieved remarkable performance for 4D scene reconstruction. However, these approaches rely on dense-frame video sequences for photorealistic reconstruction. In real-world scenarios, due to equipment constraints, sometimes only sparse frames are accessible. In this paper, we propose Sparse4DGS, the first method for sparse-frame dynamic scene reconstruction. We observe that dynamic reconstruction methods fail in both canonical and deformed spaces under sparse-frame settings, especially in areas with high texture richness. Sparse4DGS tackles this challenge by focusing on texture-rich areas. For the deformation network, we propose Texture-Aware Deformation Regularization, which introduces a texture-based depth alignment loss to regulate Gaussian deformation. For the canonical Gaussian field, we introduce Texture-Aware Canonical Optimization, which incorporates texture-based noise into the gradient descent process of canonical Gaussians. Extensive experiments show that when taking sparse frames as inputs, our method outperforms existing dynamic or few-shot techniques on NeRF-Synthetic, HyperNeRF, NeRF-DS, and our iPhone-4D datasets.
Changyue Shi, Chuxiao Yang, Wenwen Pan 0003, Jiajun Ding, Zhou Yu 0001, Jun Yu 0002
AAAI7
2026 KF-GS: Kalman filter-guided Gaussian splatting for real-time high-quality dynamic scene reconstruction
Qingyuan Tang, Yufei Yin, Yanming Zhu 0001, Zhou Yu 0001, Zhenzhong Kuang, Jiajun Ding, Jifa He
J. Vis. Commun. Image Represent.6
2026 GC-GS: Gradient control Gaussian splatting with various image degradation
Qida Cao, Jiajun Ding, Qingyuan Tang, Tianning Zhao, Xiaoling Gu, Jianping Fan 0001, Zhou Yu 0001
Pattern Recognit.2
2026 Fuzzy Language Gaussian Splatting
abstract
Recent advancements in open-vocabulary 3D querying have achieved remarkable progress. However, existing approaches such as LERF and LangSplat still rely heavily on specific query text for accurate 3D target identification. They fail to accurately comprehend fuzzy query text (which describes a target's properties, functions, or traits rather than naming it directly), causing localization mistakes and poor 3D segmentation results. In this work, we introduce a novel task—open-vocabulary 3D fuzzy query—which aims to locate and generate precise 3D target masks based on fuzzy query text, a capability that is crucial for building a 3D intelligent system. To address this challenge, we proposeFuzzy Language Gaussian Splatting (FL-GS), a framework consisting of three key stages: We first leverage a multimodal large language model (MLLM) to identify and localize potential targets in representative views based on fuzzy query text, thereby generating initial masks using the segment anything model (SAM). Subsequently, we compute the pairwise similarity between all target masks and construct an undirected, unweighted graph, from which the maximum clique is identified, thus obtaining the most probable correct masks, referred to as refined masks. Finally, we propose single-dimensional mask encoding to efficiently achieve precise 3D target masks through supervision with refined masks. Furthermore, we manually annotated two key datasets and established a benchmark for this new task. Experimental results clearly demonstrate that FL-GS outperforms existing methods in open-vocabulary 3D fuzzy querying.
Jiajun Ding, Yaowei Liu, Hongxi Zhu, Min Tan 0005, Zhou Yu 0001
IEEE Trans. Fuzzy Syst.1
2025 Gaussian Splatting-based Scene Reconstruction with Adaptive Sampling and Region-based Rendering
abstract
Recently, 3D Gaussian Splatting (3DGS) based methods, such as Scaffold-GS, have exhibited state-of-the-art performance for high-fidelity scene reconstruction. However, existing methods may suffer from degraded details because they usually pay the same attention to the whole image set and their contents. In fact, some parts of the image data may contain more information than others. Thus, it is reasonable to treat them differently. Motivated by this, in this paper, we propose a new 3DGS-based method to adaptively choose the data that are more valuable for training. Our method consists of two key parts: Adaptive Sampling (AS) and Region-based Rendering (RR). AS focuses on collecting difficult data for 3DGS training. RR focuses on enabling training with partial image data. Our method can produce more fine-grained details for scene reconstruction and has good scalability. On multiple public datasets, we verify the effectiveness of the proposed method by conducting comparative and ablation experiments. The quantitative and qualitative experimental results show that our method has achieved state-of-the-art results compared with the baselines.
Takahiko Furuya, Zhenzhong Kuang, Xiaoling Gu, Jiajun Ding
IJCNN5
2025 DiSCo: Disentangled Attribute Manipulation Retrieval via Semantic Reconstruction and Consistency Regularization
abstract
The rapid evolution of the online fashion industry has intensified the demand for interactive fashion retrieval systems capable of precise and flexible searches based on user-specified attribute modifications. However, prevailing fashion retrieval methods often overlook the distinctive distributional properties of fashion images and struggle to preserve semantic consistency during attribute manipulation. To address these limitations, we propose DiSCo, a novel disentangled attribute manipulation retrieval framework via semantic reconstruction and consistency regularization. Our approach comprises three key components: (1) An attribute-aware manipulation network that constructs target fashion embeddings through cross-modal attribute modification deltas, leveraging dedicated fashion attribute encoders; (2) A cross-modal semantic reconstruction network that synthesizes target images directly from modified attribute descriptions, supervised by adversarial and attribute classification losses to ensure interpretable edits; (3) An adaptive fusion mechanism that dynamically integrates attribute-modified embeddings with reconstructed image features. Extensive evaluations on two benchmark datasets (DeepFashion and Shopping100K) demonstrate that DiSCo achieves superior retrieval accuracy over state-of-the-arts while maintaining high-fidelity editing. Quantitative and qualitative analyses further confirm that DiSCo generates more realistic fashion representations, underscoring its effectiveness in attribute-aware retrieval tasks.
Min Tan 0005, Guanhao Liu, Huijing Zhan, Yuyu Yin, Zhou Yu 0001, Jiajun Ding, Yinfu Feng
ACM Multimedia6
2025 FutureGS: Structured Gaussian Fields for Future-Aware Dynamic Scene Modeling
Mingyang Ding, Tingting Han 0003, Jiajun Ding, Min Tan 0005, Zhenzhong Kuang
ACM Multimedia6
2025 ViewCloud: A lightweight multi-view point cloud representation for efficient 3D recognition and cross-domain retrieval
Zhihe Wu, Yaomin Wang, Zhenzhong Kuang, Jiajun Ding, Min Tan 0005, Xuefei Yin, Yanming Zhu 0001
Comput. Aided Des.4
2025 MMGS: Multi-Model Synergistic Gaussian Splatting for Sparse View Synthesis
Changyue Shi, Chuxiao Yang, Jiajun Ding, Min Tan 0005
Image Vis. Comput.5
2025 SR4D: Dynamic scene super resolution from monocular videos
Chuxiao Yang, Changyue Shi, Suguo Zhu, Jiajun Ding, Yeru Wang, Min Tan 0005
Knowl. Based Syst.5
2025 DE-NAF: decoupled neural attenuation fields for sparse-view CBCT reconstruction
Tianning Zhao, Guoping Ding, Zhenyang Liu, Hangping Wei, Min Tan 0005, Jiajun Ding
Pattern Anal. Appl.7
2025 Spatio-Temporal and Retrieval-Augmented Modeling for Chest X-Ray Report Generation
abstract
Chest X-ray report generation has attracted increasing research attention. However, most existing methods neglect the temporal information and typically generate reports conditioned on a fixed number of images. In this paper, we propose STREAM: Spatio-Temporal and REtrieval-Augmented Modelling for automatic chest X-ray report generation. It mimics clinical diagnosis by integrating current and historical studies to interpret the present condition (temporal), with each study containing images from multi-views (spatial). Concretely, our STREAM is built upon an encoder-decoder architecture, utilizing a large language model (LLM) as the decoder. Overall, spatio-temporal visual dynamics are packed as visual prompts and regional semantic entities are retrieved as textual prompts. First, a token packer is proposed to capture condensed spatio-temporal visual dynamics, enabling the flexible fusion of images from current and historical studies. Second, to augment the generation with existing knowledge and regional details, a progressive semantic retriever is proposed to retrieve semantic entities from a preconstructed knowledge bank as heuristic text prompts. The knowledge bank is constructed to encapsulate anatomical chest X-ray knowledge into structured entities, each linked to a specific chest region. Extensive experiments on public datasets have shown the state-of-the-art performance of our method. Related codes and the knowledge bank are available at https://github.com/yangyan22/STREAM.
Xiaoxing You, Ke Zhang 0029, Zhenqi Fu, Xianyun Wang, Jiajun Ding, Jiamei Sun, Zhou Yu 0001, Qingming Huang, Weidong Han 0001, Jun Yu 0002
IEEE Trans. Medical Imaging6
2025 Imp: Highly Capable Large Multimodal Models for Mobile Devices
abstract
By harnessing the capabilities of large language models (LLMs), recent large multimodal models (LMMs) have shown remarkable versatility in open-world multimodal understanding. Nevertheless, they are usually parameter-heavy and computation-intensive, thus hindering their applicability in resource-constrained scenarios. To this end, several lightweight LMMs have been proposed successively to maximize the capabilities under constrained scale (e.g., 3B). Despite the encouraging results achieved by these methods, most of them only focus on one or two aspects of the design space, and the key design choices that influence model capability have not yet been thoroughly investigated. In this paper, we conduct a systematic study for lightweight LMMs from the aspects of model architecture, training strategy, and training data. Based on our findings, we obtain Imp—a family of highly capable LMMs at the 2B$\sim$4B scales. Notably, our Imp-3B model steadily outperforms all the existing lightweight LMMs of similar size, and even surpasses the state-of-the-art LMMs at the 13B scale. With low-bit quantization and resolution reduction techniques, our Imp model can be deployed on a Qualcomm Snapdragon 8Gen3 mobile chip with a high inference speed of about 13 tokens/s.
Zhenwei Shao, Zhou Yu 0001, Jun Yu 0002, Xuecheng Ouyang, Lihao Zheng 0001, Zhenbiao Gai, Zhenzhong Kuang, Jiajun Ding
IEEE Trans. Multim.9
2025 VC-GS: view-consistent deblurring Gaussian splatting via alternating branch optimization
Qida Cao, Jiajun Ding, Zhenyang Liu, Zhenzhong Kuang, Yijie Shao, Yilan Shen
Vis. Comput.2
2025 Collaborative neural radiance fields for novel view synthesis
Junqing Yuan, Mengting Fan, Zhenyang Liu, Tongxuan Han, Zhenzhong Kuang, Chihao Pan, Jiajun Ding
Vis. Comput.7
2024 Multi-Domain Deep Learning from a Multi-View Perspective for Cross-Border E-commerce Search
abstract
Building click-through rate (CTR) and conversion rate (CVR) prediction models for cross-border e-commerce search requires modeling the correlations among multi-domains. Existing multi-domain methods would suffer severely from poor scalability and low efficiency when number of domains increases. To this end, we propose a Domain-Aware Multi-view mOdel (DAMO), which is domain-number-invariant, to effectively leverage cross-domain relations from a multi-view perspective. Specifically, instead of working in the original feature space defined by different domains, DAMO maps everything to a new low-rank multi-view space. To achieve this, DAMO firstly extracts multi-domain features in an explicit feature-interactive manner. These features are parsed to a multi-view extractor to obtain view-invariant and view-specific features. Then a multi-view predictor inputs these two sets of features and outputs view-based predictions. To enforce view-awareness in the predictor, we further propose a lightweight view-attention estimator to dynamically learn the optimal view-specific weights w.r.t. a view-guided loss. Extensive experiments on public and industrial datasets show that compared with state-of-the-art models, our DAMO achieves better performance with lower storage and computational costs. In addition, deploying DAMO to a large-scale cross-border e-commence platform leads to 1.21%, 1.76%, and 1.66% improvements over the existing CGC-based model in the online AB-testing experiment in terms of CTR, CVR, and Gross Merchandises Value, respectively.
Yinfu Feng, Yunan Ye, Min Tan 0005, Rong Xiao 0005, Haihong Tang, Jiajun Ding, Jun Yu 0002
AAAI8
2024 ZS-SRT: An efficient zero-shot super-resolution training method for Neural Radiance Fields
Yongbo He, Chengkai Wang, Zhenzhong Kuang, Jiajun Ding, Fei-wei Qin, Jun Yu 0002, Jianping Fan 0001
Neurocomputing6
2024 Confidence correction for trained graph convolutional networks
abstract
Adopting Graph Convolutional Networks (GCNs) for transductive node classification is a hot research direction in artificial intelligence . Vanilla GCNs are primarily under-confident and struggle to clarify the final classification results explicitly due to the lack of supervision. Existing works mainly alleviated this issue by improving annotation deficiency and introducing addition regularization terms. However, these methods need to re-train the model from the beginning, which is computationally expensive for large dataset and model. To deal with this problem, a novel confidence correction mechanism (CCM) for trained GCNs is proposed in this work. Such mechanism aims at calibrating the confidence output of each node in the inference stage by jointly inferring the feature and predicted pseudo label. Specifically, in the inference stage, it uses the predicted pseudo label to select target-related features over all network to obtain a more confident and better result. Such selectivity is formulated as an optimization problem to maximize the category score of each node. In addition, the greedy optimization strategy is utilized to solve this problem and we have mathematically proven that the proposed mechanism can reach the local optimum by mathematical induction . Note that such mechanism is flexible and can be introduced to most GCN-based model. Extensive experimental results on benchmark datasets show that the proposed method can promote the confidence of the final target category and improve the performance of GCNs in the inference stage.
Junqing Yuan, Huanlei Guo, Chenyi Zhou, Jiajun Ding, Zhenzhong Kuang, Zhou Yu 0001
Pattern Recognit.4
2023 Follow-me: Deceiving Trackers with Fabricated Paths
abstract
Convolutional Neural Networks (CNNs) are vulnerable to adversarial attacks in which visually imperceptible perturbations can deceive CNN-based models. While current research on adversarial attacks in single object tracking exists, it overlooks a critical aspect of manipulating predicted trajectories to follow user-defined paths regardless of the actual location of the targeted object. To address this, we propose the very first white-box attack algorithm that is capable of deceiving victim trackers by compelling them to generate trajectories that adhere to predetermined counterfeit paths. Specifically, we focus on Siamese-based trackers as our victim models. Given an arbitrary counterfeit path, we first decompose it into discrete target locations in each frame, with the assumption of constant velocity. These locations are converted to heatmap anchors, which represent the offset of their location from the target object's location in the previous frame. Later on, we design a novel loss function to minimize the gap between above-mentioned anchors and our predicted ones. Finally, the gradients computed by such loss are used to update the original video, resulting in our adversarial video. To validate our ideas, we design three sets of counterfeit paths as well as novel evaluation metrics to measure the path-following properties. Experiments with two victim models on three publicly available datasets, OTB100, VOT2018, and VOT2016, demonstrate that our algorithm not only outperforms SOTA methods significantly under conventional evaluation metrics, e.g. 90% and 68.4% precision and successful rate drop on OTB100, but also follows the counterfeit paths well, which is beyond any existing attack methods. The source code is available at https://github.com/loushengtao/Follow-me.
Shengtao Lou, Buyu Liu, Jun Bao, Jiajun Ding, Jun Yu 0002
ACM Multimedia4
2023 EGRA-NeRF: Edge-Guided Ray Allocation for Neural Radiance Fields
Zhenbiao Gai, Zhenyang Liu, Min Tan 0005, Jiajun Ding, Jun Yu 0002, Mingzhao Tong, Junqing Yuan
Image Vis. Comput.4
2023 An efficient multi-path structure with staged connection and multi-scale mechanism for text-to-image synthesis
Jiajun Ding, Beili Liu, Jun Yu 0002, Huanlei Guo, Kenong Shen
Multim. Syst.1
2023 Import vertical characteristic of rain streak for single image deraining
Zhexin Zhang, Jiajun Ding, Jun Yu 0002, Yiming Yuan, Jianping Fan 0001
Multim. Syst.2
2021 Contrastive learning of graph encoder for accelerating pedestrian trajectory prediction training
abstract
Abstract In the area of pedestrian trajectory prediction, the hybrid structures of temporal feature extractor or spatial feature extractor have paved the way for the precise prediction model, and they are in larger and larger scale. Learning of specific feature encoding model not only influenced by the structure of the network, but also by the learning manners such as supervised learning and unsupervised learning. Previous works concentrated on more comprehensive encoders and more delicate designs of feature extractors. However, the mutual influence factors from the neighbour pedestrians associate with the distance to the centre pedestrian seldomly noticed. Most of the existed feature extractors in prediction models trained in the way of supervised learning other than unsupervised manners caused the problem that the extracted features are always handcrafted without the natural distinction of obscure situations. The graph contrastive accelerating encoder is proposed, which accelerates the pedestrian trajectory prediction training process of the state of the art method of spatio‐temporal graph transformer networks. Employing the unsupervised contrastive learning process and the graph of neighbours representing distance affection of nearest and farthest pedestrian to the centre pedestrian, the graph contrastive accelerating encoder significantly shrinked the training time. Holding the final performance on to state of the art level, the proposed method let the lowest pedestrian trajectory prediction error show up in the obviously earlier training steps.
Zonggui Yao, Jun Yu 0002, Jiajun Ding
IET Image Process.3
2021 Distributed feedback network for single-image deraining
Jiajun Ding, Huanlei Guo, Jun Yu 0002, Xiongxiong He, Bo Jiang 0016
Inf. Sci.1
2018 On Collaborative Compressive Sensing Systems: The Framework, Design, and Algorithm
abstract
Based on the maximum likelihood estimation principle, we derive a collaborative estimation framework that fuses several different estimators and yields a better estimate. Applying it to compressive sensing (CS), we propose a collaborative CS (CCS) scheme consisting of a bank of $K$ CS systems that share the same sensing matrix but have different sparsifying dictionaries. This CCS system is expected to yield better performance than each individual CS system, while requiring the same time as that needed for each individual CS system when a parallel computing strategy is used. We then provide an approach to designing optimal CCS systems by utilizing a measure that involves both the sensing matrix and dictionaries and hence allows us to simultaneously optimize the sensing matrix and all the $K$ dictionaries. An alternating minimization-based algorithm is derived for solving the corresponding optimal design problem. With a rigorous convergence analysis, we show that the proposed algorithm is convergent. Experiments are carried out to confirm the theoretical results and show that the proposed CCS system yields significant improvements over the existing CS systems in terms of the signal recovery accuracy.
Zhihui Zhu, Gang Li 0010, Jiajun Ding, Qiuwei Li, Xiongxiong He
SIAM J. Imaging Sci.3
2018 A novel multi-dictionary framework with global sensing matrix design for compressed sensing
Jiajun Ding, Donghai Bao, Qingpei Wang, Xiongxiong He, Huang Bai, Sheng Li 0005
Signal Process.1
2018 Automatic clustering based on density peak detection using generalized extreme value distribution
Jiajun Ding, Xiongxiong He, Junqing Yuan, Bo Jiang 0016
Soft Comput.1
2016 An improved adaboost face detection algorithm based on the different sample weights
abstract
An improved face detection method is proposed on the basis of traditional adaboost algorithm. The training samples are not distinguished in the traditional face detection based on adaboost algorithm, which results in ignoring face samples in the process of training and the face feature information can't be fully shown. In addition, because face samples and non-face samples are treated equally, all samples must be calculated and the time of training classifier is extended. In order to improve the bad results, this paper proposes an improved strategy for implementation of algorithm. Face samples and non-face samples are set different initial weights when training classifier, so they attract different attention. And face and non-face samples are handled separately in order to reduce the complexity of the time. Compared with traditional methods, the improved method spends less time on training classifier.
Xingqiang Zhang, Jiajun Ding
CSCWD2