Shilong Jin

dblp:248/2203 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Debiasing Diffusion Priors via 3D Attention for Consistent Gaussian Splatting
abstract
Versatile 3D tasks (e.g., generation or editing) distilling Text-to-Image (T2I) diffusion models have attracted significant research interest for not relying on extensive 3D training data. However, T2I models exhibit limitations resulting from prior view bias, which produces conflicting appearances between different views of an object. This bias causes subject-words to preferentially activate prior view features during cross-attention (CA) computation, regardless of the target view condition. To overcome this limitation, we conduct a comprehensive mathematical analysis to reveal the root cause of the prior view bias in T2I models. Moreover, we find different UNet-Layers show different effects of prior view in CA. Therefore, we propose a novel framework, TD-Attn, which addresses multi-view inconsistency via two key components: (1) the 3D-Aware Attention Guidance Module 3D-AAG constructs a view-consistent 3D attention Gaussian for subject-words to enforce spatial consistency across attention-focused regions, thereby compensating for the limited spatial information in 2D individual view CA maps; (2) the Hierarchical Attention Modulation Module (HAM) utilizes a semantic guidance tree to direct the Semantic Response Profiler (SRP) in localizing and modulating CA layers that are highly responsive to view conditions, where the enhanced CA maps further support the construction of more consistent 3D attention Gaussians. Notably, HAM facilitates semantic-specific interventions, enabling controllable and precise 3D editing. Extensive experiments firmly establish that TD-Attn has the potential to serve as a transformative, universal plugin, significantly enhancing multi-view consistency across a wide range of 3D tasks.
Shilong Jin, Haoran Duan 0001, Litao Hua, Yuan Zhou 0023
AAAI1
2026 ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding
abstract
Keyframe selection has become essential for video understanding with vision-language models (VLMs) due to limited input tokens and the temporal sparsity of relevant information across video frames. Video understanding often relies on effective keyframes that are not only informative but also causally decisive. To this end, we propose Reinforced Causal Search with Information Bottleneck (ReaSon), a framework that formulates keyframe selection as an optimization problem with the help of a novel Causal Information Bottleneck (CIB), which explicitly defines keyframes as those satisfying both predictive sufficiency and causal necessity. Specifically, ReaSon employs a learnable policy network to select keyframes from a visually relevant pool of candidate frames to capture predictive sufficiency, and then assesses causal necessity via counterfactual interventions. Finally, a composite reward aligned with the CIB principle is designed to guide the selection policy through reinforcement learning. Extensive experiments on NExT-QA, EgoSchema, and Video-MME demonstrate that ReaSon consistently outperforms existing state-of-the-art methods under limited-frame settings, validating its effectiveness and generalization ability.
Yuan Zhou 0023, Litao Hua, Shilong Jin, Haoran Duan 0001
AAAI3
2026 ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation
abstract
Recent advances in zero-shot text-to-3D generation have revolutionised 3D content creation by enabling direct synthesis from textual descriptions. While state-of-the-art methods leverage 3D Gaussian Splatting with score distillation to enhance multi-view rendering through pre-trained text-to-image (T2I) models, they suffer from inherent prior view biases in T2I Models. These biases lead to inconsistent 3D generation, particularly manifesting as the multi-face Janus problem, where objects exhibit conflicting features across views. To address this fundamental challenge, we propose ConsDreamer, a novel method that mitigates view bias by refining both the conditional and unconditional terms in the score distillation process: (1) a View Disentanglement Module (VDM) that eliminates viewpoint biases in conditional prompts by decoupling irrelevant view components and injecting precise view control; and (2) a similarity-based partial order loss that enforces geometric consistency in the unconditional term by aligning cosine similarities with azimuth relationships. Extensive experiments demonstrate that ConsDreamer can be seamlessly integrated into various 3D representations and score distillation paradigms, effectively mitigating the multi-face Janus problem.
Yuan Zhou 0023, Shilong Jin, Litao Hua, Wanjun Lv, Haoran Duan 0001, Jungong Han
IEEE Trans. Image Process.2
2025 Multiagent Confrontation Method Based on Three-Party Dynamic Multistrategy Evolutionary Game
abstract
Unmanned agents represent a significant advancement in unmanned control and constitute an important element in the future agent warfare. Their autonomous decision-making capabilities are integral to accomplishing tasks independently. To address challenges inherent in multiparty game scenarios that traditional method struggle with and enhance the applicability and accuracy of game decision-making, this article proposes a novel multiagent confrontation method for unmanned vessels, tailored to a three-party dynamic multistrategy evolutionary game in incomplete information scenarios. The approach introduces a new incentive mechanism designed to enhance both individual and collective profits of agents. Using evolutionary game theory, a three-party model is developed, incorporating interactions among player, enemy, and neutral agents. The model tracks the evolution of strategies to identify stable equilibria across various perceptual conditions. Simulations validate the effectiveness of the proposed method in selecting optimal strategies for unmanned vessels in complex battlefield scenarios, demonstrating its potential for improving autonomous decision-making in multiparty confrontations.
Shilong Jin, Yingjie Wang 0002, Peiyong Duan, Haijing Zhang, Gang Li 0005, Zhipeng Cai 0001
IEEE Trans. Comput. Soc. Syst.1
2024 A Self-calibration Kalman Filter Algorithm for Dual-axis RINS Based on the Transverse Ellipsoidal Earth Model
abstract
The Kalman filter method plays a crucial role in enhancing the navigation accuracy of the dual-axis rotational inertial navigation system (RINS) through periodic estimation and compensation of device errors. Due to the particularity of polar geography, the traditional RINS mechanism in the local-level geographic frame loses efficacy in the polar region. This paper proposes a self-calibration Kalman filter algorithm based on the transverse ellipsoidal earth model to solve the self-calibration problem of RINS in polar region. This method firstly transforms the state of the local-level geographic frame to the transverse frame, and then constructs the prediction model and observation model of the Kalman filter based on the carrier state and error parameters in the transverse frame. In the self-calibration stage, a suitable rotation strategy is employed to stimulate the errors of RINS, and the proposed algorithm is utilized to estimate and compensate for the resulting errors. In addition, the traditional spherical earth model is improved to ellipsoidal earth model in this algorithm to avoid additional errors in the polar region. Monte Carlo experiments are carried out with simulation data at high latitudes, and then experiments at middle latitudes are carried out with ring laser gyro-based dual-axis RINS. The results demonstrate that the proposed method enables precise calibration of all error parameters, which aligns consistently with results obtained within the local-level geographic frame.
Pengcheng Mu, Shilong Jin, Zhikun Liao, Zhonghong Liang, Yuanhan Wang, Lin Wang 0103
FUSION2
2024 An Incentive Algorithm for Cross-region Task Allocation based on Worker Coalition Under Mobile Crowdsourcing
abstract
Mobile crowdsourcing is rapidly growing with Artificial Intelligent of Things. At the same time, the type and complexity of the tasks requested by the requester change and diversify. Therefore, how to design allocation algorithms for the situation of increasing task complexity is particularly critical. In this paper, to cope with this problem, the idea of worker coalition collaboration and reputation evaluation mechanisms are introduced into it. A two-stage allocation based on same-region and cross-region is performed in the divided regional grid. In the first stage, multi-worker and multi-task allocation is realized by combining the reverse auction theory based on the workers’ historical reputation value, which motivates the workers with high reputation value to choose their tasks and contribute data with high sensed quality. The second stage utilizes genetic algorithms to select a coalition of workers for cross-region sensed execution for tasks that do not meet sensed quality requirements. This process will provide additional payoff incentives to compensate for travel costs within the worker coalition, increasing the number of tasks completed and maximizing social welfare. Finally, multiple comparison experiments on the real dataset Yelp are conducted for validation.
Kaige Jiang, Yang Gao 0028, Peng Wang 0190, Zhaolong Gao, Xiangrong Tong, Yingjie Wang 0002, Zhipeng Cai 0001, Yingxin Li, Shilong Jin
ICWS9
2024 Bidirectional Choice for Many-to-many Online Task Assignment in Mobile Crowdsourcing
abstract
The evolution of 5G and 6G technologies has boosted mobile network speed, reduced delays, and widened coverage, empowering Mobile Crowd Sensing (MCS) to overcome surface and terrain obstacles. However, this advancement brings new hurdles for online task assignment. While most MCS methods suit surface applications, they struggle with allocation in complex environments. This paper focuses on MCS in varied settings like surface, air, and high altitude. Currently, planning-based task assignment works better for one-to-one or one-to-many scenarios, with limited options for many-to-many situations. Improving platform utility, attracting top-quality crowd workers, and enhancing task completion efficiency are vital. To tackle these challenges, the paper introduces a spatial division algorithm using 3D Voronoi diagrams for complex environments. This algorithm utilizes task coordinates to delineate assignment spaces. Additionally, it introduces a two-stage many-to-many online task assignment algorithm (MOTA) that forecasts crowd workers’ arrival probabilities and combines auction-based incentives with differential evolution algorithms. MOTA ensures efficient matching of workers and tasks within spatio-temporal constraints, balancing both parties’ interests. Finally, comparative experiments on real datasets assess the proposed MOTA algorithm’s usability and effectiveness based on overall gain, running time, task count, and assignment rate.
Yingjie Wang 0002, Yang Gao 0028, Chunxiao Mu, Zhipeng Cai 0001, Yingxin Li, Shilong Jin
ICWS8