EDBT 2026 Demo / reviewers in the wild / expert
Jianhao Shen
dblp:217/2324
· DBLP profile ↗
15ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4abstractChengwu Liu, Yichun Yin, Ye Yuan, Jiaxuan Xie, Botao Li, Siqi Li, Jianhao Shen, Yan Xu, Lifeng Shang, Ming Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chengwu Liu 0001, Yichun Yin, Ye Yuan 0016, Jiaxuan Xie, Botao Li, Jianhao Shen, Lifeng Shang, Ming Zhang 0004 |
ACL (1) | 7 |
| 2025 | Cluster-guided Contrastive Class-imbalanced Graph ClassificationabstractThis paper studies the problem of class-imbalanced graph classification, which aims at effectively classifying the graph categories in scenarios with imbalanced class distributions. While graph neural networks (GNNs) have achieved remarkable success, their modeling ability on imbalanced graph-structured data remains suboptimal, which typically leads to predictions biased towards the majority classes. On the other hand, existing class-imbalanced learning methods in vision may overlook the rich graph semantic substructures of the majority classes and excessively emphasize learning from the minority classes. To address these challenges, we propose a simple yet powerful approach called C3GNN that integrates the idea of clustering into contrastive learning to enhance class-imbalanced graph classification. Technically, C3GNN clusters graphs from each majority class into multiple subclasses, with sizes comparable to the minority class, mitigating class imbalance. It also employs the Mixup technique to generate synthetic samples, enriching the semantic diversity of each subclass. Furthermore, supervised contrastive learning is used to hierarchically learn effective graph representations, enabling the model to thoroughly explore semantic substructures in majority classes while avoiding excessive focus on minority classes. Extensive experiments on real-world graph benchmark datasets verify the superior performance of our proposed method against competitive baselines. Wei Ju 0001, Zhengyang Mao, Siyu Yi, Yifang Qin, Yiyang Gu, Zhiping Xiao 0001, Jianhao Shen, Ziyue Qiao, Ming Zhang 0004 |
AAAI | 7 |
| 2025 | AGSPN: Efficient attention-gated spatial propagation network for depth completion
Huajie Wen, Jianhao Shen, Qiaohui Feng |
Expert Syst. Appl. | 2 |
| 2025 | Single image dehazing via fine-scale perception network
Huajie Wen, Jianhao Shen, Yuanzhi Deng, Liting Chen, Qiaohui Feng |
Neurocomputing | 2 |
| 2024 | Measuring Vision-Language STEM Skills of Neural ModelsabstractWe introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding of multimodal vision-language information of STEM. Our dataset features one of the largest and most comprehensive datasets for the challenge. It includes $448$ skills and $1,073,146$ questions spanning all STEM subjects. Compared to existing datasets that often focus on examining expert-level ability, our dataset includes fundamental skills and questions designed based on the K-12 curriculum. We also add state-of-the-art foundation models such as CLIP and GPT-3.5-Turbo to our benchmark. Results show that the recent model advances only help master a very limited number of lower grade-level skills ($2.5$% in the third grade) in our dataset. In fact, these models are still well below (averaging $54.7$%) the performance of elementary students, not to mention near expert-level performance. To understand and increase the performance on our dataset, we teach the models on a training split of our dataset.
Even though we observe improved performance, the model performance remains relatively low compared to average elementary students. To solve STEM problems, we will need novel algorithmic innovations from the community. Jianhao Shen, Ye Yuan 0016, Srbuhi Mirzoyan, Ming Zhang 0004, Chenguang Wang 0001 |
ICLR | 1 |
| 2024 | Dense frustum-aware fusion for 3D object detection in perception systems
Yuanzhi Deng, Jianhao Shen, Huajie Wen |
Expert Syst. Appl. | 2 |
| 2024 | A Comprehensive Survey on Deep Graph Representation Learning
Wei Ju 0001, Zheng Fang 0007, Yiyang Gu, Zequn Liu, Qingqing Long, Ziyue Qiao, Yifang Qin, Jianhao Shen, Zhiping Xiao 0001, Jingyang Yuan, Yusheng Zhao, Yifan Wang 0014, Xiao Luo 0001, Ming Zhang 0004 |
Neural Networks | 8 |
| 2023 | DT-Solver: Automated Theorem Proving with Dynamic-Tree Sampling Guided by Proof-level Value FunctionabstractHaiming Wang, Ye Yuan, Zhengying Liu, Jianhao Shen, Yichun Yin, Jing Xiong, Enze Xie, Han Shi, Yujun Li, Lin Li, Jian Yin, Zhenguo Li, Xiaodan Liang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhengying Liu, Jianhao Shen, Yichun Yin, Enze Xie, Jian Yin 0001, Zhenguo Li, Xiaodan Liang |
ACL (1) | 4 |
| 2023 | ComSearch: Equation Searching with Combinatorial Strategy for Solving Math Word Problems with Weak SupervisionabstractPrevious studies have introduced a weaklysupervised paradigm for solving math word problems requiring only the answer value annotation.While these methods search for correct value equation candidates as pseudo labels, they search among a narrow sub-space of the enormous equation space.To address this problem, we propose a novel search algorithm with combinatorial strategy ComSearch, which can compress the search space by excluding mathematically equivalent equations.The compression allows the searching algorithm to enumerate all possible equations and obtain high-quality data.We investigate the noise in the pseudo labels that hold wrong mathematical logic, which we refer to as the false-matching problem, and propose a ranking model to denoise the pseudo labels.Our approach holds a flexible framework to utilize two existing supervised math word problem solvers to train pseudo labels, and both achieve state-of-the-art performance in the weak supervision task. 1 Qianying Liu, Wenyu Guan, Jianhao Shen, Fei Cheng 0002, Sadao Kurohashi |
EACL | 3 |
| 2023 | TRIGO: Benchmarking Formal Mathematical Proof Reduction for Generative Language ModelsabstractJing Xiong, Jianhao Shen, Ye Yuan, Haiming Wang, Yichun Yin, Zhengying Liu, Lin Li, Zhijiang Guo, Qingxing Cao, Yinya Huang, Chuanyang Zheng, Xiaodan Liang, Ming Zhang, Qun Liu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jianhao Shen, Ye Yuan 0016, Yichun Yin, Zhengying Liu, Zhijiang Guo, Qingxing Cao, Yinya Huang, Chuanyang Zheng, Xiaodan Liang, Ming Zhang 0004, Qun Liu 0001 |
EMNLP | 2 |
| 2022 | Joint Language Semantic and Structure Embedding for Knowledge Graph CompletionabstractThe task of completing knowledge triplets has broad downstream applications. Both structural and semantic information plays an important role in knowledge graph completion. Unlike previous approaches that rely on either the structures or semantics of the knowledge graphs, we propose to jointly embed the semantics in the natural language description of the knowledge triplets with their structure information. Our method embeds knowledge graphs for the completion task via fine-tuning pre-trained language models with respect to a probabilistic structured loss, where the forward pass of the language models captures semantics and the loss reconstructs structures. Our extensive experiments on a variety of knowledge graph benchmarks have demonstrated the state-of-the-art performance of our method. We also show that our method can significantly improve the performance in a low-resource regime, thanks to the better use of semantics. The code and datasets are available at https://github.com/pkusjh/LASS. Jianhao Shen, Chenguang Wang 0001, Linyuan Gong, Dawn Song |
COLING | 1 |
| 2022 | HE-SNE: Heterogeneous Event Sequence-based Streaming Network Embedding for Dynamic BehaviorsabstractLarge amounts of user behavior data provide opportunities for user behavior modeling and have great potential in many downstream applications such as advertising and anomaly detection. Compared with traditional methods, embedding-based methods are used more often recently because of their efficiency and scalability. These methods build a “behavior-entity” bipartite graph and learn static embeddings for nodes in the graph. However, behavior patterns in the real world could not be static because entity properties such as user interests usually evolve along with time. In this paper, we formulate user behaviors as a temporal event sequence and propose a stream network embedding approach to capture the evolving nature of user behaviors. Representation of each event is built and used to update the embeddings of nodes. Two contextual behavior modeling tasks are studied for dynamic user behaviors, and experimental results with real-world data demonstrate the effectiveness of our proposed approach over several competitive baselines. Yifan Wang 0014, Jianhao Shen, Yiping Song, Sheng Wang 0012, Ming Zhang 0004 |
IJCNN | 2 |
| 2022 | KGNN: Harnessing Kernel-based Networks for Semi-supervised Graph ClassificationabstractThis paper studies semi-supervised graph classification, which is an important problem with various applications in social network analysis and bioinformatics. This problem is typically solved by using graph neural networks (GNNs), which yet rely on a large number of labeled graphs for training and are unable to leverage unlabeled graphs. We address the limitations by proposing the Kernel-based Graph Neural Network (KGNN). A KGNN consists of a GNN-based network as well as a kernel-based network parameterized by a memory network. The GNN-based network performs classification through learning graph representations to implicitly capture the similarity between query graphs and labeled graphs, while the kernel-based network uses graph kernels to explicitly compare each query graph with all the labeled graphs stored in a memory for prediction. The two networks are motivated from complementary perspectives, and thus combing them allows KGNN to use labeled graphs more effectively. We jointly train the two networks by maximizing their agreement on unlabeled graphs via posterior regularization, so that the unlabeled graphs serve as a bridge to let both networks mutually enhance each other. Experiments on a range of well-known benchmark datasets demonstrate that KGNN achieves impressive performance over competitive baselines. Wei Ju 0001, Meng Qu, Weiping Song, Jianhao Shen, Ming Zhang 0004 |
WSDM | 5 |
| 2020 | Multi-task Learning via Adaptation to Similar Tasks for Mortality Prediction of Diverse Rare Diseases
Luchen Liu, Zequn Liu, Haoxian Wu, Zichang Wang, Jianhao Shen, Yiping Song, Ming Zhang 0004 |
AMIA | 5 |
| 2018 | Learning the Joint Representation of Heterogeneous Temporal Events for Clinical Endpoint PredictionabstractThe availability of a large amount of electronic health records (EHR) provides huge opportunities to improve health care service by mining these data. One important application is clinical endpoint prediction, which aims to predict whether a disease, a symptom or an abnormal lab test will happen in the future according to patients' history records. This paper develops deep learning techniques for clinical endpoint prediction, which are effective in many practical applications. However, the problem is very challenging since patients' history records contain multiple heterogeneous temporal events such as lab tests, diagnosis, and drug administrations. The visiting patterns of different types of events vary significantly, and there exist complex nonlinear relationships between different events. In this paper, we propose a novel model for learning the joint representation of heterogeneous temporal events. The model adds a new gate to control the visiting rates of different events which effectively models the irregular patterns of different events and their nonlinear correlations. Experiment results with real-world clinical data on the tasks of predicting death and abnormal lab tests prove the effectiveness of our proposed approach over competitive baselines. Luchen Liu, Jianhao Shen, Ming Zhang 0004, Zichang Wang, Jian Tang 0005 |
AAAI | 2 |