Lei Zhang 0199

dblp:97/8704-199 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0002-5808-5313ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledge
abstract
Chaoyue He, Xin Zhou, Yi Wu, Xinjia Yu, Yan Zhang, Lei Zhang, Di Wang, Shengfei Lyu, Hong Xu, Wang Xiaoqiao, Wei Liu, Chunyan Miao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Chaoyue He, Xin Zhou 0008, Xinjia Yu, Lei Zhang 0199, Di Wang 0004, Shengfei Lyu, Hong Xu 0004, Xiaoqiao Wang, Chunyan Miao
EMNLP6
2025 MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
abstract
Environmental, Social, and Governance (ESG) reports are essential for assessing sustainability, regulatory compliance, and financial transparency. However, these documents are typically long, multimodal, and structurally complex, combining dense text, tables, figures, and layout-sensitive semantics. Existing AI systems often struggle to perform reliable document-level reasoning in such settings, and no dedicated benchmark currently exists in ESG domain. To fill the gap, we introduce MMESGBench, a first-of-its-kind benchmark dataset targeted to evaluate multimodal understanding and reasoning across multi-source ESG documents. This dataset is constructed via a human-AI collaborative, multi-stage pipeline. First, a multimodal LLM generates candidate question-answer (QA) pairs by jointly interpreting textual, tabular, and visual information from layout-aware document pages. Second, an LLM verifies the semantic accuracy, completeness, and reasoning complexity of each QA pair. This automated process is followed by an expert-in-the-loop validation, where domain specialists validate and calibrate QA pairs to ensure quality, relevance, and diversity. MMESGBench comprises 933 validated QA pairs derived from 45 ESG documents, spanning across seven distinct document types and three major ESG source categories. Questions are categorized as single-page, cross-page, or unanswerable, with each accompanied by fine-grained multimodal evidence. Initial experiments validate that multimodal and retrieval-augmented models substantially outperform text-only baselines. MMESGBench is publicly available as an open-source dataset at https://github.com/Zhanglei1103/MMESGBench.
Lei Zhang 0199, Xin Zhou 0008, Chaoyue He, Di Wang 0004, Hong Xu 0004, Chunyan Miao
ACM Multimedia1
2024 Estimating package arrival time via heterogeneous hypergraph neural network
Lei Zhang 0199, Yong Liu 0020, Xin Zhou 0008, Li-Zhen Cui 0001, Chunyan Miao
Expert Syst. Appl.1
2024 Package Arrival Time Prediction via Knowledge Distillation Graph Neural Network
abstract
Accurately estimating packages’ arrival time in e-commerce can enhance users’ shopping experience and improve the placement rate of products. This problem is often formalized as an Origin-Destination (OD)-based ETA (i.e., estimated time of arrival) prediction task, where the delivery time is estimated mainly based on sender and receiver addresses and other context information. One inherent challenge of the OD-based ETA problem is that the delivery time highly depends on the actual delivery trajectory which is unknown at the time of prediction. In this article, we tackle this challenge by effectively exploiting historical delivery trajectories. We propose a novel Knowledge Distillation Graph neural network-based package ETA prediction (KDG-ETA) model, which uses knowledge distillation in the training phase to distill the knowledge of historical trajectories into OD pair embeddings. In KDG-ETA, a multi-level trajectory graph representation model is proposed to fully exploit trajectory information at the node-level, edge-level, and path-level. Then, the OD representations embedded with trajectory knowledge are combined with context embeddings from feature extraction module for delivery time prediction using an adaptive attention module. KDG-ETA consistently outperforms existing state-of-the-art OD-based ETA prediction methods on three real-world Alibaba datasets, reducing the Mean Absolute Error (MAE) by 3.0%–39.1% as demonstrated in our extensive empirical evaluation.
Lei Zhang 0199, Yong Liu 0020, Zhiqi Shen 0001, Li-Zhen Cui 0001
ACM Trans. Knowl. Discov. Data1
2023 MMTN: Multi-Modal Memory Transformer Network for Image-Report Consistent Medical Report Generation
abstract
Automatic medical report generation is an essential task in applying artificial intelligence to the medical domain, which can lighten the workloads of doctors and promote clinical automation. The state-of-the-art approaches employ Transformer-based encoder-decoder architectures to generate reports for medical images. However, they do not fully explore the relationships between multi-modal medical data, and generate inaccurate and inconsistent reports. To address these issues, this paper proposes a Multi-modal Memory Transformer Network (MMTN) to cope with multi-modal medical data for generating image-report consistent medical reports. On the one hand, MMTN reduces the occurrence of image-report inconsistencies by designing a unique encoder to associate and memorize the relationship between medical images and medical terminologies. On the other hand, MMTN utilizes the cross-modal complementarity of the medical vision and language for the word prediction, which further enhances the accuracy of generating medical reports. Extensive experiments on three real datasets show that MMTN achieves significant effectiveness over state-of-the-art approaches on both automatic metrics and human evaluation.
Li-Zhen Cui 0001, Lei Zhang 0199, Fuqiang Yu, Zhen Li 0049
AAAI3
2023 CMT: Cross-modal Memory Transformer for Medical Image Report Generation
Li-Zhen Cui 0001, Lei Zhang 0199, Fuqiang Yu, Zhen Li 0049, Chunyan Miao
DASFAA (3)3
2023 Dual Graph Multitask Framework for Imbalanced Delivery Time Estimation
Lei Zhang 0199, Xin Zhou 0008, Li-Zhen Cui 0001, Zhiqi Shen 0001
DASFAA (4)1
2023 Delivery Time Prediction Using Large-Scale Graph Structure Learning Based on Quantile Regression
abstract
Predicting Estimated Time of Arrival (ETA) for packages is a critical problem in e-commerce. The prediction is often made based on spatial (sending and receiving addresses), temporal (payment time), and context (merchants) attributes. Existing methods usually formalize this task as an Origin-Destination (OD) ETA prediction problem and exploit the attribute relations with graph learning. However, most existing methods make use of fixed and manually defined graph structures, which are often not optimal for downstream ETA task and hence lead to unsatisfactory prediction results. In addition, current ETA models tend to focus on prediction accuracy without considering fulfillment rate. This may lead to a low fulfillment rate in practice, i.e., actual delivery time is much longer than estimations provided by models, which consequently exacerbates the frustrating experiences for users. To address these issues, we propose a novel Graph Structure Learning-based Quantile Regression (GSL-QR) model for e-commerce ETA prediction in this paper. Specifically, we utilize graph structure learning to dynamically update the spatial and temporal relation graphs of orders and learn optimal graph structures and graph embeddings guided by downstream ETA prediction task. To guarantee both prediction accuracy and order fulfillment rate, we design a multi-objective quantile regression in GSL-QR that can find the Pareto solution of the problem. In order to extend GSL to large-scale real-world graphs, we devise a Fast Sampling-based Graph Structure Learning (FS-GSL) method, which can significantly reduce the computational complexity of graph structure learning. Finally, we conduct comprehensive experiments on three industrial datasets collected from Alibaba e-commerce platform. The results demonstrate that the proposed model can significantly outperform baselines on both ETA prediction accuracy and order fulfillment rate.
Lei Zhang 0199, Xin Zhou 0008, Yong Liu 0020, Li-Zhen Cui 0001, Zhiqi Shen 0001
ICDE1
2022 KdINet: Knowledge-driven Interpretable Network for Medical Imaging Diagnosis
abstract
Automatic diagnosis for medical images is a significant research problem of computer-aided diagnosis to reduce the workload of doctors in recent years. However, existing deep learning approaches for diagnoses are usually black-box models with implicit decision-making processes that make them inexplainable. To alleviate the issue, in this paper, we propose a Knowledge-driven Interpretable Network (KdINet) for interpretable disease classification of medical images. KdINet first exploits the pretrained CNN module and the hierarchical representation module to learn two different disease feature representations (i.e., visual disease features and hierarchical disease features). Subsequently, the joint training of two disease features by KdINet’s disease classifier learns a hierarchical classification criterion to infer the diagnosis and generate corresponding interpretable justifications (i.e., ancestral disease paths of the diagnosis). Extensive experiments on three datasets demonstrate that our KdINet achieves significantly higher effectiveness than state-of-the-art approaches on disease classification metrics.
Li-Zhen Cui 0001, Lei Zhang 0199, Fuqiang Yu, Zhen Li 0049, Chunyan Miao
BIBM3
2022 KdTNet: Medical Image Report Generation via Knowledge-Driven Transformer
Li-Zhen Cui 0001, Fuqiang Yu, Lei Zhang 0199, Zhen Li 0049, Ning Liu 0014
DASFAA (3)4
2019 Standard Condition Number of Hessian Matrix for Neural Networks
abstract
Neural networks are becoming more and more important for intelligent communications and their theoretical research has become a top priority. Loss surfaces are crucial to understand and improve performance in neural networks. In this paper, the Hessian matrix of second order optimization method is analyzed through the analytical framework of random matrix theory (RMT) in order to understand the geometry of loss surfaces. The limited spectrum distribution, extreme eigenvalue distribution, and standard condition number (SCN) of Hessian matrix are analyzed to understand their asymptotic characteristics. Moreover, the relationships among the extreme eigenvalue distribution, SCN, and the convergence of loss surfaces are investigated. The above analyses give insight into utilizing RMT to analyze the neural network theory.
Lei Zhang 0199, Wensheng Zhang 0004, Yunzeng Li, Jian Sun 0013, Cheng-Xiang Wang 0001
ICC1