Shengjie Ma

dblp:18/8541 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MVP-Net: Multi-oriented Vessel Detail Preservation Network
Shengjie Ma, Yang Liu 0262
ICIC (15)2
2025 Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented Generation
abstract
Retrieval-augmented generation (RAG) has improved large language models (LLMs) by using knowledge retrieval to overcome knowledge deficiencies. However, current RAG methods often fall short of ensuring the depth and completeness of retrieved information, which is necessary for complex reasoning tasks. In this work, we introduce Think-on-Graph 2.0 (ToG-2), a hybrid RAG framework that iteratively retrieves information from both unstructured and structured knowledge sources in a tight-coupling manner. Specifically, ToG-2 leverages knowledge graphs (KGs) to link documents via entities, facilitating deep and knowledge-guided context retrieval. Simultaneously, it utilizes documents as entity contexts to achieve precise and efficient graph retrieval. ToG-2 alternates between graph retrieval and context retrieval to search for in-depth clues relevant to the question, enabling LLMs to generate answers. We conduct a series of well-designed experiments to highlight the following advantages of ToG-2: 1) ToG-2 tightly couples the processes of context retrieval and graph retrieval, deepening context retrieval via the KG while enabling reliable graph retrieval based on contexts; 2) it achieves deep and faithful reasoning in LLMs through an iterative knowledge retrieval process of collaboration between contexts and the KG; and 3) ToG-2 is training-free and plug-and-play compatible with various LLMs. Extensive experiments demonstrate that ToG-2 achieves overall state-of-the-art (SOTA) performance on 6 out of 7 knowledge-intensive datasets with GPT-3.5, and can elevate the performance of smaller models (e.g., LLAMA-2-13B) to the level of GPT-3.5’s direct reasoning. The source code is available on https://anonymous.4open.science/r/ToG2.
Shengjie Ma, Chengjin Xu, Xuhui Jiang, Muzhi Li 0001, Huaren Qu, Cehao Yang, Jiaxin Mao, Jian Guo 0016
ICLR1
2025 Relightable Gaussian blendshapes for head avatar animation
Shengjie Ma, Youyi Zheng, Yanlin Weng, Kun Zhou 0001
Comput. Graph.1
2024 Mixed Pixel Spatial Unmixing with Hyperspectral LiDAR Echo Waveform Analysis
abstract
Hyperspectral LiDAR (HSL) has demonstrated significant promise in achieving super range resolution, yet the effects of measurement variables like multi-target spectral relationships and signal-to-noise ratio (SNR) on its limits are unclear. This research establishes a mathematical model for HSL’s multi-layer target detection, validated by a 94% match with measurements. A novel method for spatial unmixing is introduced, informed by prior-knowledge acquisition and waveform decomposition, demonstrating that with SNR over 10 dB, HSL can resolve targets 10 cm apart with distinct spectral signatures, using a 4 ns pulse width. These findings advance the understanding of HSL’s capabilities for detailed environmental awareness and object discrimination.
Yuhao Xia, Shilong Xu, Shengjie Ma, Wenxin Tian, Yihua Hu 0001
IGARSS3
2023 Meta-Path Based Social Relation Reasoning in a Deep and Robust Way
Xuhui Jiang, Yinghan Shen, Yuanzhuo Wang, Huawei Shen, Chengjin Xu, Shengjie Ma
DASFAA (3)6
2023 Session Search with Pre-trained Graph Classification Model
abstract
Session search is a widely adopted technique in search engines that seeks to leverage the complete interaction history of a search session to better understand the information needs of users and provide more relevant ranking results. The vast majority of existing methods model a search session as a sequence of queries and previously clicked documents. However, if we simply represent a search session as a sequence we will lose the topological information in the original search session. It is non-trivial to model the intra-session interactions and complicated structural patterns among the previously issued queries, clicked documents, as well as the terms or entities that appeared in them. To solve this problem, in this paper, we propose a novel Session Search with Graph Classification Model (SSGC), which regards session search as a graph classification task on a heterogeneous graph that represents the search history in each session. To improve the performance of the graph classification, we design a specific pre-training strategy for our proposed GNN-based classification model. Extensive experiments on two public session search datasets demonstrate the effectiveness of our model in the session search task.
Shengjie Ma, Chong Chen 0001, Jiaxin Mao, Qi Tian 0001, Xuhui Jiang
SIGIR1
2022 A multiresolution network architecture for deferred neural lighting
abstract
Abstract We present a novel multiresolution network architecture for deferred neural lighting. The key idea is to explicitly separate the processing of appearance at different spatial resolutions, leading to considerably improved high‐frequency details as well as temporal stability in animation sequences with varying view conditions. Moreover, our network is only half the size of the original one, and requires less training data to converge to satisfactory results. The network is tested over five captured datasets from deferred neural lighting and may be extended to other neural appearance techniques, such as NeRF or neural textures.
Shengjie Ma, Hongzhi Wu, Zhong Ren 0004, Kun Zhou 0001
Comput. Animat. Virtual Worlds1
2022 NEAWalk: Inferring missing social interactions via topological-temporal embeddings of social groups
Yinghan Shen, Xuhui Jiang, Zijian Li 0014, Yuanzhuo Wang, Xiaolong Jin 0001, Shengjie Ma, Xueqi Cheng 0001
Knowl. Inf. Syst.6
2021 Neural compositing for real-time augmented reality rendering in low-frequency lighting environments
Shengjie Ma, Qiming Hou, Zhong Ren 0001, Kun Zhou 0001
Sci. China Inf. Sci.1
2019 Knowing User Better: Jointly Predicting Click-Through and Playtime for Micro-Video
abstract
Most micro-video recommender systems use the click-through to measure user satisfaction. However, the amount of time that users spend on a video, the playtime, measures user engagement on video contents and should be used as a complement to the click based signals. In this paper, we propose a coarse-to-fine multi-task jointly optimizing model to predict click-through and playtime. Following the click-through prediction, the playtime is first discretized into several intervals and classified into a specific one with a proposed ordered-balanced cross entropy loss. Then, to further improve upon coarse estimates, we learn a subtle offset with a regressor and produce a fine-grained playtime estimation. To make mutual promotion between click-through and playtime predictors, we optimize them jointly in a multi-task manner. Experimental results show that we achieve state-of-the-art performance on recommendation task and demonstrate effectiveness on playtime prediction at the same time.
Shengjie Ma, Zhengjun Zha, Feng Wu 0001
ICME1
2018 Content-Based Video Relevance Prediction with Second-Order Relevance and Attention Modeling
abstract
This paper describes our proposed method for the Content-Based Video Relevance Prediction (CBVRP) challenge. Our method is based on deep learning, i.e. we train a deep network to predict the relevance between two video sequences from their features. We explore the usage of second-order relevance, both in preparing training data, and in extending the deep network. Second-order relevance refers to e.g. the relevance between x and z if x is relevant to y and y is relevant to z. In our proposed method, we use second-order relevance to increase positive samples and decrease negative samples, when preparing training data. We further extend the deep network with an attention module, where the attention mechanism is designed for second-order relevant video sequences. We verify the effectiveness of our method on the validation set of the CBVRP challenge.
Xusong Chen, Rui Zhao 0001, Shengjie Ma, Dong Liu 0002, Zhengjun Zha
ACM Multimedia3