VLDB 2026 Research / reviewers in the wild / expert
Jing Sha
dblp:96/5272
· DBLP profile ↗
19ranked-venue papers
3as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 12 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Computer networks · 2Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Mathematical Reasoning Through Autonomously Learning KnowledgeabstractEnabling machines to solve mathematical problems is a vital endeavor in developing intelligence that emulates human-like thinking and reasoning. However, most existing approaches focus on reconstructing human comprehension of problems, which are still far from enough since they neglect the fundamental human ability to learn knowledge from experiences. In this article, we focus on empowering models with the cognitive capacity to autonomously learn knowledge from mathematical problem-solving. We first propose a Cognitive Solver (CogSolver) that contains an intelligent BRAIN-ARM framework as the cognitive structure and operates the knowledge learning process in Store-Apply-Update steps inspired by two cognitive science theories. The BRAIN system stores three basic types of mathematical knowledge, and the ARM system applies them organically in answer reasoning process. After solving problems, the BRAIN updates its stored knowledge based on the ARM's feedback, with knowledge filters to eliminate redundancies and foster a more rational knowledge base. Our CogSolver carries out the above three steps iteratively, emulating a more human-like behavior. Furthermore, in order to overcome knowledge forgetting during the learning process, we extend CogSolver to CogSolver+ by incorporating an essential knowledge Recall mechanism, which is inspired by another prominent cognitive theory. We first discuss and fuse three crucial factors in simulating human memory replay. Then, we propose a influenced-based method with a theoretical guarantee of efficiency to consolidate the updated knowledge. Experiments on three math word problem benchmarks demonstrate the improvements of our CogSolver and CogSolver+ in answer reasoning and clearly illustrate how they acquire knowledge, leading to superior interpretability. Jiayu Liu 0001, Zhenya Huang, Enhong Chen, Qi Liu 0003, Hongke Zhao, Xin Lin 0005, Jing Sha, Shijin Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Enhancing Chain-of-Thought Reasoning via Neuron Activation Differential AnalysisabstractDespite the impressive chain-of-thought (CoT) reasoning ability of large language models (LLMs), its underlying mechanisms remains unclear.In this paper, we explore the inner workings of LLM's CoT ability via the lens of neurons in the feed-forward layers.We propose an efficient method to identify reasoningcritical neurons by analyzing their activation patterns under reasoning chains of varying quality.Based on it, we devise a rather simple intervention method that directly stimulates these reasoning-critical neurons, to guide the generation of high-quality reasoning chains.Extended experiments validate the effectiveness of our method and demonstrate the critical role these identified neurons play in CoT reasoning. Yiru Tang, Kun Zhou 0002, Yingqian Min, Wayne Xin Zhao, Jing Sha, Zhichao Sheng, Shijin Wang 0001 |
EMNLP | 5 |
| 2025 | CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive PerspectiveabstractAlthough large language models (LLMs) show promise in solving complex mathematical tasks, existing evaluation paradigms rely solely on a coarse measure of overall answer accuracy, which are insufficient for assessing their authentic capabilities. In this paper, we propose CogMath, which comprehensively assesses LLMs’ mathematical abilities through the lens of human cognition. Specifically, inspired by psychological theories, CogMath formalizes human reasoning process into 3 stages: problem comprehension, problem solving, and solution summarization. Within these stages, we investigate perspectives such as numerical calculation, knowledge, and counterfactuals, and design a total of 9 fine-grained evaluation dimensions. In each dimension, we develop an “Inquiry-Judge-Reference” multi-agent system to generate inquiries that assess LLMs’ mastery from this dimension. An LLM is considered to truly master a problem only when excelling in all inquiries from the 9 dimensions. By applying CogMath on three benchmarks, we reveal that the mathematical capabilities of 7 mainstream LLMs are overestimated by 30%-40%. Moreover, we locate their strengths and weaknesses across specific stages/dimensions, offering in-depth insights to further enhance their reasoning abilities. Jiayu Liu 0001, Zhenya Huang, Jing Sha, Qi Liu 0003, Shijin Wang 0001, Enhong Chen |
ICML | 6 |
| 2025 | DPFA-UNet: Dual-Path Fusion Attention for Accurate Brain Tumor SegmentationabstractABSTRACT Gliomas are the most common primary brain tumors within the central nervous system, typically observed through magnetic resonance imaging (MRI). Precise segmentation of brain tumor in MRI is highly significant for both clinical diagnosis and treatment. However, due to complexity of tumor structures, existing deep‐learning‐based methods for brain tumor segmentation still face challenges in accurately delineating tumor core (TC) and enhancing tumor (ET) regions, which are primary targets for actual treatment. To address this problem, this work proposes dual‐path fusion attention‐based UNet (DPFA‐UNet) that leverages a dual‐path attention block (DPA) and a concurrent attention fusion block (CAF) within a U‐shaped architecture. Specifically, DPA enhances adaptability to lesions of varying sizes by using multi‐scale branches that capture fine details and global features. CAF fuses high‐ and low‐level semantic features using a parallel attention mechanism, effectively focusing on the focal regions. It also incorporates a mask generated by deep supervision mechanism to further guide feature fusion. Additionally, to reduce demand for hardware resources, we incorporate depthwise separable convolution into the model. Experiments are conducted on public BraTS 2021 and BraTS 2019 datasets. The results verify that DPFA‐UNet outperforms existing brain tumor segmentation methods, particularly in segmenting TC and ET regions. Jing Sha |
IET Image Process. | 1 |
| 2024 | Learning to Solve Geometry Problems via Simulating Human Dual-Reasoning Process
Tong Xiao 0004, Jiayu Liu 0001, Zhenya Huang, Jing Sha, Shijin Wang 0001, Enhong Chen |
IJCAI | 5 |
| 2024 | Item-Difficulty-Aware Learning Path Recommendation: From a Real Walking PerspectiveabstractLearning path recommendation aims to provide learners with a reasonable order of items to achieve their learning goals. Intuitively, the learning process on the learning path can be metaphorically likened to walking. Despite extensive efforts in this area, most previous methods mainly focus on the relationship among items but overlook the difficulty of items, which may raise two issues from a real walking perspective: (1) The path may be rough: When learners tread the path without considering item difficulty, it's akin to walking a dark, uneven road, making learning harder and dampening interest. (2) The path may be inefficient: Allowing learners only a few attempts on very challenging items before switching, or persisting with a difficult item despite numerous attempts without mastery, can result in inefficiencies in the learning journey. To conquer the above limitations, we propose a novel method named Difficulty-constrained Learning Path Recommendation (DLPR), which is aware of item difficulty. Specifically, we first explicitly categorize items into learning items and practice items, then construct a hierarchical graph to model and leverage item difficulty adequately. Then we design a Difficulty-driven Hierarchical Reinforcement Learning (DHRL) framework to facilitate learning paths with efficiency and smoothness. Finally, extensive experiments on three different simulators demonstrate our framework achieves state-of-the-art performance. Haotian Zhang 0007, Shuanghong Shen, Bihan Xu, Zhenya Huang, Jing Sha, Shijin Wang 0001 |
KDD | 6 |
| 2024 | SocraticLM: Exploring Socratic Personalized Teaching with Large Language ModelsabstractLarge language models (LLMs) are considered a crucial technology for advancing intelligent education since they exhibit the potential for an in-depth understanding of teaching scenarios and providing students with personalized guidance. Nonetheless, current LLM-based application in personalized teaching predominantly follows a "Question-Answering" paradigm, where students are passively provided with answers and explanations. In this paper, we propose SocraticLM, which achieves a Socratic "Thought-Provoking" teaching paradigm that fulfills the role of a real classroom teacher in actively engaging students in the thought process required for genuine problem-solving mastery. To build SocraticLM, we first propose a novel "Dean-Teacher-Student" multi-agent pipeline to construct a new dataset, SocraTeach, which contains $35$K meticulously crafted Socratic-style multi-round (equivalent to $208$K single-round) teaching dialogues grounded in fundamental mathematical problems. Our dataset simulates authentic teaching scenarios, interacting with six representative types of simulated students with different cognitive states, and strengthening four crucial teaching abilities. SocraticLM is then fine-tuned on SocraTeach with three strategies balancing its teaching and reasoning abilities. Moreover, we contribute a comprehensive evaluation system encompassing five pedagogical dimensions for assessing the teaching quality of LLMs. Extensive experiments verify that SocraticLM achieves significant improvements in the teaching performance, outperforming GPT4 by more than 12\%. Our dataset and code is available at https://github.com/Ljyustc/SocraticLM. Jiayu Liu 0001, Zhenya Huang, Tong Xiao 0004, Jing Sha, Qi Liu 0003, Shijin Wang 0001, Enhong Chen |
NeurIPS | 4 |
| 2024 | JiuZhang3.0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis ModelsabstractMathematical reasoning is an important capability of large language models~(LLMs) for real-world applications.
To enhance this capability, existing work either collects large-scale math-related texts for pre-training, or relies on stronger LLMs (\eg GPT-4) to synthesize massive math problems. Both types of work generally lead to large costs in training or synthesis.
To reduce the cost, based on open-source available texts, we propose an efficient way that trains a small LLM for math problem synthesis, to efficiently generate sufficient high-quality pre-training data.
To achieve it, we create a dataset using GPT-4 to distill its data synthesis capability into the small LLM.
Concretely, we craft a set of prompts based on human education stages to guide GPT-4, to synthesize problems covering diverse math knowledge and difficulty levels.
Besides, we adopt the gradient-based influence estimation method to select the most valuable math-related texts.
The both are fed into GPT-4 for creating the knowledge distillation dataset to train the small LLM.
We leverage it to synthesize 6 million math problems for pre-training our JiuZhang3.0 model. The whole process only needs to invoke GPT-4 API 9.3k times and use 4.6B data for training.
Experimental results have shown that JiuZhang3.0 achieves state-of-the-art performance on several mathematical reasoning datasets, under both natural language reasoning and tool manipulation settings.
Our code and data will be publicly released in \url{https://github.com/RUCAIBox/JiuZhang3.0}. Kun Zhou 0002, Beichen Zhang 0003, Zhipeng Chen 0001, Wayne Xin Zhao, Jing Sha, Zhichao Sheng, Shijin Wang 0001, Ji-Rong Wen |
NeurIPS | 6 |
| 2024 | Graph-based Student Knowledge Profile for Online Intelligent EducationabstractStudent knowledge profile is the basis for adaptive learning applications in online learning resulting from modeling the student mastery of knowledge concepts. In recent years, typical works based on knowledge tracing (KT) expect to profile students and have achieved significant success for the next performance prediction. However, in practical online learning scenarios, current methods tend to suffer from the following challenges: 1) Prediction inconsistency: The accuracy of the next performance prediction is inconsistent with the accuracy of student knowledge profile prediction, which is the more required result. 2) Cold start of knowledge: In online learning scenarios, it is often necessary to profile some knowledge concepts without learning records in advance. In this paper, we propose a novel Graph-based Student Knowledge Profile Model (GSKPM), along with a new end-to-end training objective, to tackle these challenges. We first define a new training objective to ensure the model is capable of inferring consistent student knowledge profiles. Then in this model, a two-stage hyper-aggregation process is employed to make full use of the topological relations between knowledge concepts and knowledge domains to provide information during profiling, especially for cold start knowledge concepts. Finally, through extensive experiments on real-world datasets, we will show that GSKPM achieves better prediction performances on student knowledge profiles and well deals with the cold start problem. Haotian Zhang 0007, Zhenya Huang, Qi Liu 0003, Jing Sha, Enhong Chen, Shijin Wang 0001 |
SDM | 6 |
| 2024 | Bit-mask Robust Contrastive Knowledge Distillation for Unsupervised Semantic HashingabstractUnsupervised semantic hashing has emerged as an indispensable technique for fast image search, which aims to convert images into binary hash codes without relying on labels. Recent advancements in the field demonstrate that employing large-scale backbones (e.g., ViT) in unsupervised semantic hashing models can yield substantial improvements. However, the inference delay has become increasingly difficult to overlook. Knowledge distillation provides a means for practical model compression to alleviate this delay. Nevertheless, the prevailing knowledge distillation approaches are not explicitly designed for semantic hashing. They ignore the unique search paradigm of semantic hashing, the inherent necessities of the distillation process, and the property of hash codes. In this paper, we propose an innovative Bit-mask Robust Contrastive knowledge Distillation (BRCD) method, specifically devised for the distillation of semantic hashing models. To ensure the effectiveness of two kinds of search paradigms in the context of semantic hashing, BRCD first aligns the semantic spaces between the teacher and student models through a contrastive knowledge distillation objective. Additionally, to eliminate noisy augmentations and ensure robust optimization, a cluster-based method within the knowledge distillation process is introduced. Furthermore, through a bit-level analysis, we uncover the presence of redundancy bits resulting from the bit independence property. To mitigate these effects, we introduce a bit mask mechanism in our knowledge distillation objective. Finally, extensive experiments not only showcase the noteworthy performance of our BRCD method in comparison to other knowledge distillation methods but also substantiate the generality of our methods across diverse semantic hashing models and backbones. The code for BRCD is available at https://github.com/hly1998/BRCD. Liyang He, Zhenya Huang, Jiayu Liu 0001, Enhong Chen, Fei Wang 0063, Jing Sha, Shijin Wang 0001 |
WWW | 6 |
| 2023 | Exploiting Non-Interactive Exercises in Cognitive DiagnosisabstractCognitive Diagnosis aims to quantify the proficiency level of students on specific knowledge concepts. Existing studies merely leverage observed historical students-exercise interaction logs to access proficiency levels. Despite effectiveness, observed interactions usually exhibit a power-law distribution, where the long tail consisting of students with few records lacks supervision signals. This phenomenon leads to inferior diagnosis among few records students. In this paper, we propose the Exercise-aware Informative Response Sampling (EIRS) framework to address the long-tail problem. EIRS is a general framework that explores the partial order between observed and unobserved responses as auxiliary ranking-based training signals to supplement cognitive diagnosis. Considering the abundance and complexity of unobserved responses, we first design an Exercise-aware Candidates Selection module, which helps our framework produce reliable potential responses for effective supplementary training. Then, we develop an Expected Ability Change-weighted Informative Sampling strategy to adaptively sample informative potential responses that contribute greatly to model training. Experiments on real-world datasets demonstrate the supremacy of our framework in long-tailed data. Fangzhou Yao, Qi Liu 0003, Min Hou 0004, Shiwei Tong, Zhenya Huang, Enhong Chen, Jing Sha, Shijin Wang 0001 |
IJCAI | 7 |
| 2023 | JiuZhang 2.0: A Unified Chinese Pre-trained Language Model for Multi-task Mathematical Problem SolvingabstractAlthough pre-trained language models~(PLMs) have recently advanced the research progress in mathematical reasoning, they are not specially designed as a capable multi-task solver, suffering from high cost for multi-task deployment (e.g. a model copy for a task) and inferior performance on complex mathematical problems in practical applications. To address these issues, we propose JiuZhang 2.0, a unified Chinese PLM specially for multi-task mathematical problem solving. Our idea is to maintain a moderate-sized model and employ the cross-task knowledge sharing to improve the model capacity in a multi-task setting. Specially, we construct a Mixture-of-Experts (MoE) architecture for modeling mathematical text, to capture the common mathematical knowledge across tasks. For optimizing the MoE architecture, we design multi-task continual pre-training and multi-task fine-tuning strategies for multi-task adaptation. These training strategies can effectively decompose the knowledge from the task data and establish the cross-task sharing via expert networks. To further improve the general capacity of solving different complex tasks, we leverage large language models (LLMs) as complementary models to iteratively refine the generated solution by our PLM, via in-context learning. Extensive experiments have demonstrated the effectiveness of our model. Wayne Xin Zhao, Kun Zhou 0002, Beichen Zhang 0003, Zheng Gong 0001, Zhipeng Chen 0001, Yuanhang Zhou, Ji-Rong Wen, Jing Sha, Shijin Wang 0001, Cong Liu 0006 |
KDD | 8 |
| 2023 | Evaluating and Improving Tool-Augmented Computation-Intensive Math ReasoningabstractChain-of-thought prompting (CoT) and tool augmentation have been validated in recent work as effective practices for improving large language models (LLMs) to perform step-by-step reasoning on complex math-related tasks.However, most existing math reasoning datasets may not be able to fully evaluate and analyze the ability of LLMs in manipulating tools and performing reasoning, as they often only require very few invocations of tools or miss annotations for evaluating intermediate reasoning steps, thus supporting only outcome evaluation.To address the issue, we construct CARP, a new Chinese dataset consisting of 4,886 computation-intensive algebra problems with formulated annotations on intermediate steps, facilitating the evaluation of the intermediate reasoning process.In CARP, we test four LLMs with CoT prompting, and find that they are all prone to make mistakes at the early steps of the solution, leading to incorrect answers.Based on this finding, we propose a new approach that can facilitate the deliberation on reasoning steps with tool interfaces, namely DELI.In DELI, we first initialize a step-by-step solution based on retrieved exemplars, then iterate two deliberation procedures that check and refine the intermediate steps of the generated solution, from both tool manipulation and natural language reasoning perspectives, until solutions converge or the maximum iteration is achieved.Experimental results on CARP and six other datasets show that the proposed DELI mostly outperforms competitive baselines, and can further boost the performance of existing CoT methods.Our data and code are available at https://github.com/RUCAIBox/CARP. Beichen Zhang 0003, Kun Zhou 0002, Xilin Wei, Wayne Xin Zhao, Jing Sha, Shijin Wang 0001, Ji-Rong Wen |
NeurIPS | 5 |
| 2023 | A light-weight and accurate pig detection method based on complex scenes
Jing Sha, Gong-Li Zeng |
Multim. Tools Appl. | 1 |
| 2022 | Continual Pre-training of Language Models for Math Problem Understanding with Syntax-Aware Memory NetworkabstractIn this paper, we study how to continually pretrain language models for improving the understanding of math problems.Specifically, we focus on solving a fundamental challenge in modeling math problems, i.e., how to fuse the semantics of textual description and formulas, which are highly different in essence.To address this issue, we propose a new approach called COMUS to continually pre-train language models for math problem understanding with syntax-aware memory network.In this approach, we first construct the math syntax graph to model the structural semantic information, by combining the parsing trees of the text and formulas, and then design the syntax-aware memory networks to deeply fuse the features from the graph and text.With the help of syntax relations, we can model the interaction between the token from the text and its semantic-related nodes within the formulas, which is helpful to capture fine-grained semantic correlations between texts and formulas.Besides, we devise three continual pre-training tasks to further align and fuse the representations of the text and math syntax graph.Experimental results on four tasks in the math domain demonstrate the effectiveness of our approach.Our code and data are publicly available at the link: https: //github.com/RUCAIBox/COMUS. Zheng Gong 0001, Kun Zhou 0002, Wayne Xin Zhao, Jing Sha, Shijin Wang 0001, Ji-Rong Wen |
ACL (1) | 4 |
| 2022 | JiuZhang: A Chinese Pre-trained Language Model for Mathematical Problem UnderstandingabstractThis paper aims to advance the mathematical intelligence of machines by presenting the first Chinese mathematical pre-trained language model (PLM) for effectively understanding and representing mathematical problems. Unlike other standard NLP tasks, mathematical texts are difficult to understand, since they involve mathematical terminology, symbols and formulas in the problem statement. Typically, it requires complex mathematical logic and background knowledge for solving mathematical problems. Wayne Xin Zhao, Kun Zhou 0002, Zheng Gong 0001, Beichen Zhang 0003, Yuanhang Zhou, Jing Sha, Zhigang Chen 0003, Shijin Wang 0001, Cong Liu 0006, Ji-Rong Wen |
KDD | 6 |
| 2020 | A method for virtual machine migration in cloud computing using a collective behavior-based metaheuristics algorithmabstractSummary Due to the growth of applications and the integration of new customers into the world of computing systems, computing needs to be changed and to become more powerful and flexible than before. Meanwhile, cloud computing is presented as a model beyond a system that is currently capable of answering most request needs. Flexible infrastructure for cloud computing and virtualization technology provide new features to support business activities. Clouds are a very important topic that used secure management tools for storage, security and securing data centers in a flexible manner. One of the important matters in cloud technologies is virtual machine migration (VMM). There are different ways to implement the VMM, but because of the limitation of resources' energy, energy management is also very important and challenging. Due to the NP‐hard nature of this problem, this paper presents an energy‐aware VMM engine for cloud computing using the discrete bacterial foraging algorithm as a new collective behavior‐based metaheuristics algorithm. The CloudSim simulator is employed to investigate the efficiency of this method. The obtained results have shown that the proposed method improves the energy consumption and the migration count. Jing Sha, Abdol Ghaffar Ebadi, Dinesh Mavaluru, Mohammed Alshehri, Osama Alfarraj, Lila Rajabion |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | ASSCA: API sequence and statistics features combined architecture for malware detection
Fangshuo Jiang, Shengwei Yi, Jing Sha, Pietro Liò |
Comput. Networks | 5 |
| 2018 | Terminal Sensitive Data Protection by Adjusting Access Time Bidirectionally and AutomaticallyabstractAlong with the rapid development of the Internet, people more incline to access to the network for life needs through intelligent terminal. Once the devices are beyond control, the privacy information are leaked. Hence it's important to protect the data obtained by client. From above view, a terminal sensitive data protection method by adjusting access time automatically is proposed. This method achieves fine grit visit control based on sliding time window, applies dynamic authorization rules based on temporary parameters, employs temporary parameters to identify users and control their visits, and uses the data desensitization to protect sensitive privacy. It can ensure the security of obtained data by client, and also meets the availability of the data. The proposed method employing access control, time authorization and desensitization was compared with those existing data protection methods, as well as its validity was verified. Shengfei Zhang, Pietro Liò, Jing Sha |
ICCCN | 5 |