VLDB 2026 Research / reviewers in the wild / expert
Jianting Tang
dblp:400/6950
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Artificial intelligence
2 papers |
Vision and language · 61% Knowledge representation and reasoning · 30% Representation and self-supervised learning · 9% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 9 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › retrieval models › neural retrieval
dense retrieval |
1.0 | 1 | 2026 | Large Reasoning Embedding Models: Towards Next-Generation Dense Retrieval Paradigm · WWW 2026 |
Information retrieval › retrieval models › neural retrieval
embedding-based retrieval |
1.0 | 1 | 2026 | Large Reasoning Embedding Models: Towards Next-Generation Dense Retrieval Paradigm · WWW 2026 |
Information retrieval
retrieval models |
1.0 | 1 | 2026 | Large Reasoning Embedding Models: Towards Next-Generation Dense Retrieval Paradigm · WWW 2026 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models · ICCV 2025 |
Computer vision › Vision and language › cross-modal alignment
visual alignment |
0.9 | 1 | 2025 | BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models · ICCV 2025 |
Bioinformatics and computational biology › molecular informatics
molecular representation learning |
0.9 | 1 | 2025 | CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs · ACM Multimedia 2025 |
Security and privacy of machine learning
adversarial attack |
0.9 | 1 | 2025 | Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial Images · ICLR 2025 |
Security and privacy of machine learning
model intellectual property protection |
0.9 | 1 | 2025 | Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial Images · ICLR 2025 |
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning |
0.3 | 1 | 2025 | BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models · ICCV 2025 |
Methods — techniques the papers use, named apart from their topics
cross-view prefix tuning · 1.7SMILES guided resampling · 1.7large reasoning models · 1.0parameter learning · 0.9logit distribution matching · 0.9fine-tuning · 0.9direct visual supervision · 0.9auto-regressive supervision · 0.9adversarial attack · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Large Reasoning Embedding Models: Towards Next-Generation Dense Retrieval Paradigm
Jianting Tang, Dongshuai Li, Tao Wen 0018, Fuyu Lv, Dan Ou, Linli Xu 0002 |
WWW | 1 |
| 2026 | Multimodal Deep Learning for Online Review Helpfulness PredictionabstractABSTRACT Review helpfulness prediction (RHP) is critical for alleviating information overload and supporting reliable decision making on e‐commerce platforms, yet it is challenged by unstructured and multimodal review content as well as distorted helpfulness signals. Prior studies have examined textual, structural, and reviewer‐related determinants and built models based on manually engineered features or deep learning, but most approaches remain text‐centric, make limited use of review images and structured metadata, and offer little interpretability regarding how different information sources contribute to predictions. To address these issues, we propose Attentive Gated Multimodal Fusion (AGMF), a deep learning framework that jointly models review text, review images, and structured metadata. AGMF employs a pre‐trained language model for textual representations, a convolutional neural network for visual features, and metadata features capturing reviewer characteristics and behavioural signals, and fuses them through cross‐modal attention and a gated multimodal unit that adaptively weights each modality at the instance level, providing modality‐level interpretability. Experiments on a large‐scale dataset of Chinese e‐commerce reviews show that AGMF consistently outperforms traditional machine‐learning methods, strong single‐modality deep learning baselines, and competitive multimodal fusion models in terms of accuracy, F1‐score, and AUC, and ablation studies confirm the effectiveness of each modality and the proposed fusion mechanism. Overall, this study contributes an interpretable multimodal architecture for RHP that effectively integrates text, images, and metadata and offers practical guidance for designing intelligent review filtering systems on online platforms. Jianting Tang, Yongxiang Sheng |
Expert Syst. J. Knowl. Eng. | 3 |
| 2025 | BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language ModelsabstractMainstream Multimodal Large Language Models (MLLMs) achieve visual understanding by using a vision projector to bridge well-pretrained vision encoders and large language models (LLMs). The inherent gap between visual and textual modalities makes the embeddings from the vision projector critical for visual comprehension. However, current alignment approaches treat visual embeddings as contextual cues and merely apply auto-regressive supervision to textual outputs, neglecting the necessity of introducing equivalent direct visual supervision, which hinders the potential finer alignment of visual embeddings. In this paper, based on our analysis of the refinement process of visual embeddings in the LLM's shallow layers, we propose BASIC, a method that utilizes refined visual embeddings within the LLM as supervision to directly guide the projector in generating initial visual embeddings. Specifically, the guidance is conducted from two perspectives: (i) optimizing embedding directions by reducing angles between initial and supervisory embeddings in semantic space; (ii) improving semantic matching by minimizing disparities between the logit distributions of both visual embeddings. Without additional supervisory models or artificial annotations, BASIC significantly improves the performance of MLLMs across a wide range of benchmarks, demonstrating the effectiveness of our introduced direct visual supervision. Jianting Tang, Yubo Wang 0010, Haoyu Cao 0004, Linli Xu 0002 |
ICCV | 1 |
| 2025 | Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial ImagesabstractLarge vision-language models (LVLMs) have demonstrated remarkable image understanding and dialogue capabilities, allowing them to handle a variety of visual question answering tasks. However, their widespread availability raises concerns about unauthorized usage and copyright infringement, where users or individuals can develop their own LVLMs by fine-tuning published models. In this paper, we propose a novel method called Parameter Learning Attack (PLA) for tracking the copyright of LVLMs without modifying the original model. Specifically, we construct adversarial images through targeted attacks against the original model, enabling it to generate specific outputs. To ensure these attacks remain effective on potential fine-tuned models to trigger copyright tracking, we allow the original model to learn the trigger images by updating parameters in the opposite direction during the adversarial attack process. Notably, the proposed method can be applied after the release of the original model, thus not affecting the model’s performance and behavior. To simulate real-world applications, we fine-tune the original model using various strategies across diverse datasets, creating a range of models for copyright verification. Extensive experiments demonstrate that our method can more effectively identify the original copyright of fine-tuned models compared to baseline methods. Therefore, this work provides a powerful tool for tracking copyrights and detecting unlicensed usage of LVLMs. Yubo Wang 0010, Jianting Tang, Chaohu Liu, Linli Xu 0002 |
ICLR | 2 |
| 2025 | CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMsabstractRecent advances in molecular science have been propelled significantly by large language models (LLMs). However, their effectiveness is limited when relying solely on molecular sequences, which fail to capture the complex structures of molecules. Beyond sequence representation, molecules exhibit two complementary structural views: the first focuses on the topological relationships between atoms, as exemplified by the graph view; and the second emphasizes the spatial configuration of molecules, as represented by the image view. The two types of views provide unique insights into molecular structures. To leverage these views collaboratively, we propose the CROss-view Prefixes (CROP) to enhance LLMs' molecular understanding through efficient multi-view integration. CROP possesses two advantages: (i) efficiency: by jointly resampling multiple structural views into fixed-length prefixes, it avoids excessive consumption of the LLM's limited context length and allows easy expansion to more views; (ii) effectiveness: by utilizing the LLM's self-encoded molecular sequences to guide the resampling process, it boosts the quality of the generated prefixes. Specifically, our framework features a carefully designed SMILES Guided Resampler for view resampling, and a Structural Embedding Gate for converting the resulting embeddings into LLM's prefixes. Extensive experiments demonstrate the superiority of CROP in tasks including molecule captioning, IUPAC name prediction and molecule property prediction. Jianting Tang, Yubo Wang 0010, Haoyu Cao 0001, Linli Xu 0002 |
ACM Multimedia | 1 |