Jinsong Lan

dblp:146/8009 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0000-6890-4960ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 REVISION:Reflective Intent Mining and Online Reasoning Auxiliary for E-Commerce Visual Search System Optimization
Qiuyu Zhao, Zenghui Sun, Jinsong Lan, Xiaoyong Zhu, Bo Zheng 0007
ICDE4
2026 Disentangling Representations from Search Behaviors for Recommendation via Counterfactual Learning
abstract
For recommender systems in internet platforms, search activities provide additional insights into user interest through query-click interactions with items, and are thus widely used for enhancing personalized recommendation. However, these interacted items have not only transferable features that match users’ interests and are beneficial to the recommendation domain, but also have features related to users’ unique intents in the search domain. Such a domain gap of item features is neglected by most current search-enhanced recommendation methods. They directly incorporate these search behaviors into recommendation, and thus introduce partial negative transfer. Tackling this problem is challenging due to the lack of explicit supervision signals to disentangle features matching search-specific intent or general interest. To address this, we propose ClardRec, a c ounterfactual l e a rning-driven r epresentation d isentanglement framework for search-enhanced recommendation, based on the common belief that a user would click an item under a query not solely because of the item-query match but also due to the item’s query-independent general features (e.g., color or style) that interest the user. These general features exclude the reflection of search-specific intents contained in queries, ensuring a pure match to users’ underlying interests to complement recommendation. We perform the disentanglement based on a counterfactual thinking idea, how would user preferences and query match change for items if we removed their query-related features in search. Specifically, we leverage search queries to construct counterfactual signals to disentangle item representations, isolating only query-independent general features. These representations subsequently enable feature augmentation and data augmentation for the recommendation scenario. Comprehensive experiments on real datasets demonstrate that ClardRec is effective in both collaborative filtering and sequential recommendation scenarios. The source code is available at https://github.com/JJCui96/ClardRec .
Jiajun Cui, Xu Chen 0026, Shuai Xiao 0002, Chen Ju, Jinsong Lan, Jianyong Wang 0001, Wei Zhang 0056
ACM Trans. Inf. Syst.5
2025 Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal Model
abstract
Large Multimodal Models (LMMs) have significantly progressed by extending large language models. Building on this progress, the latest developments in LMMs demonstrate the ability to generate dense pixel-wise segmentation by integrating segmentation models. Despite the innovations, existing works’ textual responses and segmentation masks remain at the instance level, showing limited ability to perform fine-grained understanding and segmentation even provided with detailed textual cues. To overcome this limitation, we introduce a Multi-Granularity Large Multimodal Model (MGLMM), which is capable of seamlessly adjusting the granularity of Segmentation and Captioning (SegCap) following user instructions, from panoptic SegCap to fine-grained SegCap. We name such a new task Multi-Granularity Segmentation and Captioning (MGSC). Observing the lack of a benchmark for model training and evaluation over the MGSC task, we establish a benchmark with aligned masks and captions in multi-granularity using our customized automated annotation pipeline. This benchmark comprises 10K images and more than 30K image-question pairs. We will release our dataset along with the implementation of our automated dataset annotation pipeline for further research. Besides, we propose a novel unified SegCap data format to unify heterogeneous segmentation datasets; it effectively facilitates learning to associate object concepts with visual features during multi-task training. Extensive experiments demonstrate that our MGLMM excels at tackling more than eight downstream tasks and achieves state-of-the-art performance in MGSC, GCG, image captioning, referring segmentation, multiple/empty segmentation, and reasoning segmentation. The great properties and versatility of MGLMM underscore its potential impact on advancing multimodal research.
Li Zhou 0017, Zenghui Sun, Zikun Zhou, Jinsong Lan
AAAI5
2025 Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training
abstract
In rapidly evolving field of vision-language models (VLMs), contrastive language-image pre-training (CLIP) has made significant strides, becoming foundation for various downstream tasks. However, relying on one-to-one (image, text) contrastive paradigm to learn alignment from large-scale messy web data, CLIP faces a serious myopic dilemma, resulting in biases towards monotonous short texts and shallow visual expressivity. To overcome these issues, this paper advances CLIP into one novel holistic paradigm, by updating both diverse data and alignment optimization. To obtain colorful data with low cost, we use image-to-text captioning to generate multi-texts for each image, from multiple perspectives, granularities, and hierarchies. Two gadgets are proposed to encourage textual diversity. To match such (image, multi-texts) pairs, we modify the CLIP image encoder into multi-branch, and propose multi-to-multi contrastive optimization for image-text part-to-part matching. As a result, diverse visual embeddings are learned for each image, bringing good interpretability and generalization. Extensive experiments and ablations across over ten benchmarks indicate that our holistic CLIP significantly outperforms existing myopic CLIP, including image-text retrieval, open-vocabulary classification, and dense visual tasks. Project page is available to further promote the prosperity of VLMs: https://voide1220.github.io/Holism/.
Haicheng Wang, Chen Ju, Weixiong Lin, Shuai Xiao 0002, Mingshuai Yao, Jinsong Lan, Ying Chen 0011, Qingwen Liu 0002
CVPR9
2025 Inter: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling
abstract
Hallucinations in large vision-language models (LVLMs) pose significant challenges for real-world applications, as LVLMs may generate responses that appear plausible yet remain inconsistent with the associated visual content. This issue rarely occurs in human cognition. We argue that this discrepancy arises from humans' ability to effectively leverage multimodal interaction information in data samples. Specifically, humans typically first gather multimodal information, analyze the interactions across modalities for understanding, and then express their understanding through language. Motivated by this observation, we conduct extensive experiments on popular LVLMs and obtained insights that surprisingly reveal human-like, though less pronounced, cognitive behavior of LVLMs on multimodal samples. Building on these findings, we further propose \textbf{INTER}: \textbf{Inter}action Guidance Sampling, a novel training-free algorithm that mitigate hallucinations without requiring additional data. Specifically, INTER explicitly guides LVLMs to effectively reapply their understanding of multimodal interaction information when generating responses, thereby reducing potential hallucinations. On six benchmarks including VQA and image captioning tasks, INTER achieves an average improvement of up to 3.4\% on five LVLMs compared to the state-of-the-art decoding strategy. The code will be released when the paper is accepted.
Zenghui Sun, Lihua Jing, Jinsong Lan, Xiaoyong Zhu, Bo Zheng 0007
ICCV8
2024 Wear-Any-Way: Manipulable Virtual Try-on via Sparse Correspondence Alignment
Zhonghua Zhai, Chen Ju, Xuewen Hong, Jinsong Lan
ECCV (14)6
2024 Turbo: Informativity-Driven Acceleration Plug-In for Vision-Language Large Models
Chen Ju, Haicheng Wang, Haozhe Cheng, Xu Chen 0026, Zhonghua Zhai, Jinsong Lan, Shuai Xiao 0002, Bo Zheng 0007
ECCV (46)7
2024 Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in Videos
abstract
Video try-on is challenging and has not been well tackled in previous works. The main obstacle lies in preserving the clothing details and modeling the coherent motions simultaneously. Faced with those difficulties, we address video try-on by proposing a diffusion-based framework named ''Tunnel Try-on.'' The core idea is excavating a ''focus tunnel'' in the input video that gives close-up shots around the clothing regions. We zoom in on the region in the tunnel to better preserve the fine details of the clothing. To generate coherent motions, we leverage the Kalman filter to smooth the tunnel and inject its position embedding into attention layers to improve the continuity of the generated videos. In addition, we develop an environment encoder to extract the context information outside the tunnels. Equipped with these techniques, Tunnel Try-on keeps fine clothing details and synthesizes stable and smooth videos. Demonstrating significant advancements, Tunnel Try-on could be regarded as the first attempt toward the commercial-level application of virtual try-on in videos. The project page is https://mengtingchen.github.io/tunnel-try-on-page/.
Zhengze Xu, Linyu Xing, Zhonghua Zhai, Nong Sang, Jinsong Lan, Shuai Xiao 0002, Changxin Gao
ACM Multimedia7
2014 A New Framework for Traffic Anomaly Detection
abstract
Trajectory data is becoming more and more popular nowadays and extensive studies have been conducted on trajectory data. One important research direction about trajectory data is the anomaly detection which is to find all anomalies based on trajectory patterns in a road network. In this paper, we introduce a road segment-based anomaly detection problem, which is to detect the abnormal road segments each of which has its “real” traffic deviating from its “expected” traffic and to infer the major causes of anomalies on the road network. First, a deviation-based method is proposed to quantify the anomaly of reach road segment. Second, based on the observation that one anomaly from a road segment can trigger other anomalies from the road segments nearby, a diffusion-based method based on a heat diffusion model is proposed to infer the major causes of anomalies on the whole road network. To validate our methods, we conduct intensive experiments on a large real-world GPS dataset of about 23,000 taxis in Shenzhen, China to demonstrate the performance of our algorithms.
Jinsong Lan, Cheng Long 0001, Raymond Chi-Wing Wong, Youyang Chen, Yanjie Fu, Danhuai Guo, Yong Ge 0001, Yuanchun Zhou
SDM1
2013 FEDCVS: A fair and efficient scheduling scheme for dynamic cooperative video streaming on smartphones
abstract
As video applications are increasingly popular over smartphones, many cooperative video streaming mechanisms have been proposed. These mechanisms use cellular link as well device-to-device links simultaneously to provide higher quality video streaming to mobile users. However current works solely focus on throughput enhancement in static scenarios. Consequently these mechanisms result in unfairness since smartphones with higher download rate expend more cellular traffic and monetary costs. Additionally, previous works assume a static scenario that all smartpone users start to watch the same video at the same time. Obviously, the static scenario is unrealistic in actual mobile environments. Based on these insights, in this paper, we focus on a more practical dynamic cooperation scenario and propose a scheduling scheme to achieve efficient cooperative video streaming and guarantee fluent user experience. More importantly, the proposed scheduling scheme achieves a significant improvement in fairness among cooperators. Through extensive simulations across a wide range of scenarios, we show that the proposed scheme significantly outperforms other works by 52%, 24% and 27% respectively in terms of fairness, without sacrificing efficiency.
Anfu Zhou, Min Liu 0001, Jinsong Lan, Zhongcheng Li
GLOBECOM4