Jianting Tang

dblp:400/6950 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Artificial intelligence
2 papers
Vision and language · 61% Knowledge representation and reasoning · 30% Representation and self-supervised learning · 9%
Network and information security
1 paper
Security and privacy of machine learning · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 9 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › retrieval models › neural retrieval
dense retrieval
1.012026
Large Reasoning Embedding Models: Towards Next-Generation Dense Retrieval Paradigm · WWW 2026
Information retrieval › retrieval models › neural retrieval
embedding-based retrieval
1.012026
Large Reasoning Embedding Models: Towards Next-Generation Dense Retrieval Paradigm · WWW 2026
Information retrieval
retrieval models
1.012026
Large Reasoning Embedding Models: Towards Next-Generation Dense Retrieval Paradigm · WWW 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models · ICCV 2025
Computer vision › Vision and language › cross-modal alignment
visual alignment
0.912025
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models · ICCV 2025
Bioinformatics and computational biology › molecular informatics
molecular representation learning
0.912025
CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs · ACM Multimedia 2025
Security and privacy of machine learning
adversarial attack
0.912025
Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial Images · ICLR 2025
Security and privacy of machine learning
model intellectual property protection
0.912025
Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial Images · ICLR 2025
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning
0.312025
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models · ICCV 2025

Methods — techniques the papers use, named apart from their topics

cross-view prefix tuning · 1.7SMILES guided resampling · 1.7large reasoning models · 1.0parameter learning · 0.9logit distribution matching · 0.9fine-tuning · 0.9direct visual supervision · 0.9auto-regressive supervision · 0.9adversarial attack · 0.9
YearPublicationVenuePosition
2026 Large Reasoning Embedding Models: Towards Next-Generation Dense Retrieval Paradigm
Jianting Tang, Dongshuai Li, Tao Wen 0018, Fuyu Lv, Dan Ou, Linli Xu 0002
WWW1
2026 Multimodal Deep Learning for Online Review Helpfulness Prediction
abstract
ABSTRACT Review helpfulness prediction (RHP) is critical for alleviating information overload and supporting reliable decision making on e‐commerce platforms, yet it is challenged by unstructured and multimodal review content as well as distorted helpfulness signals. Prior studies have examined textual, structural, and reviewer‐related determinants and built models based on manually engineered features or deep learning, but most approaches remain text‐centric, make limited use of review images and structured metadata, and offer little interpretability regarding how different information sources contribute to predictions. To address these issues, we propose Attentive Gated Multimodal Fusion (AGMF), a deep learning framework that jointly models review text, review images, and structured metadata. AGMF employs a pre‐trained language model for textual representations, a convolutional neural network for visual features, and metadata features capturing reviewer characteristics and behavioural signals, and fuses them through cross‐modal attention and a gated multimodal unit that adaptively weights each modality at the instance level, providing modality‐level interpretability. Experiments on a large‐scale dataset of Chinese e‐commerce reviews show that AGMF consistently outperforms traditional machine‐learning methods, strong single‐modality deep learning baselines, and competitive multimodal fusion models in terms of accuracy, F1‐score, and AUC, and ablation studies confirm the effectiveness of each modality and the proposed fusion mechanism. Overall, this study contributes an interpretable multimodal architecture for RHP that effectively integrates text, images, and metadata and offers practical guidance for designing intelligent review filtering systems on online platforms.
Jianting Tang, Yongxiang Sheng
Expert Syst. J. Knowl. Eng.3
2025 BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
abstract
Mainstream Multimodal Large Language Models (MLLMs) achieve visual understanding by using a vision projector to bridge well-pretrained vision encoders and large language models (LLMs). The inherent gap between visual and textual modalities makes the embeddings from the vision projector critical for visual comprehension. However, current alignment approaches treat visual embeddings as contextual cues and merely apply auto-regressive supervision to textual outputs, neglecting the necessity of introducing equivalent direct visual supervision, which hinders the potential finer alignment of visual embeddings. In this paper, based on our analysis of the refinement process of visual embeddings in the LLM's shallow layers, we propose BASIC, a method that utilizes refined visual embeddings within the LLM as supervision to directly guide the projector in generating initial visual embeddings. Specifically, the guidance is conducted from two perspectives: (i) optimizing embedding directions by reducing angles between initial and supervisory embeddings in semantic space; (ii) improving semantic matching by minimizing disparities between the logit distributions of both visual embeddings. Without additional supervisory models or artificial annotations, BASIC significantly improves the performance of MLLMs across a wide range of benchmarks, demonstrating the effectiveness of our introduced direct visual supervision.
Jianting Tang, Yubo Wang 0010, Haoyu Cao 0004, Linli Xu 0002
ICCV1
2025 Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial Images
abstract
Large vision-language models (LVLMs) have demonstrated remarkable image understanding and dialogue capabilities, allowing them to handle a variety of visual question answering tasks. However, their widespread availability raises concerns about unauthorized usage and copyright infringement, where users or individuals can develop their own LVLMs by fine-tuning published models. In this paper, we propose a novel method called Parameter Learning Attack (PLA) for tracking the copyright of LVLMs without modifying the original model. Specifically, we construct adversarial images through targeted attacks against the original model, enabling it to generate specific outputs. To ensure these attacks remain effective on potential fine-tuned models to trigger copyright tracking, we allow the original model to learn the trigger images by updating parameters in the opposite direction during the adversarial attack process. Notably, the proposed method can be applied after the release of the original model, thus not affecting the model’s performance and behavior. To simulate real-world applications, we fine-tune the original model using various strategies across diverse datasets, creating a range of models for copyright verification. Extensive experiments demonstrate that our method can more effectively identify the original copyright of fine-tuned models compared to baseline methods. Therefore, this work provides a powerful tool for tracking copyrights and detecting unlicensed usage of LVLMs.
Yubo Wang 0010, Jianting Tang, Chaohu Liu, Linli Xu 0002
ICLR2
2025 CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs
abstract
Recent advances in molecular science have been propelled significantly by large language models (LLMs). However, their effectiveness is limited when relying solely on molecular sequences, which fail to capture the complex structures of molecules. Beyond sequence representation, molecules exhibit two complementary structural views: the first focuses on the topological relationships between atoms, as exemplified by the graph view; and the second emphasizes the spatial configuration of molecules, as represented by the image view. The two types of views provide unique insights into molecular structures. To leverage these views collaboratively, we propose the CROss-view Prefixes (CROP) to enhance LLMs' molecular understanding through efficient multi-view integration. CROP possesses two advantages: (i) efficiency: by jointly resampling multiple structural views into fixed-length prefixes, it avoids excessive consumption of the LLM's limited context length and allows easy expansion to more views; (ii) effectiveness: by utilizing the LLM's self-encoded molecular sequences to guide the resampling process, it boosts the quality of the generated prefixes. Specifically, our framework features a carefully designed SMILES Guided Resampler for view resampling, and a Structural Embedding Gate for converting the resulting embeddings into LLM's prefixes. Extensive experiments demonstrate the superiority of CROP in tasks including molecule captioning, IUPAC name prediction and molecule property prediction.
Jianting Tang, Yubo Wang 0010, Haoyu Cao 0001, Linli Xu 0002
ACM Multimedia1