Mingzheng Li

dblp:210/3321 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 HybridSparse: An End-to-End Hybrid Framework for Efficient Large-Scale Retrieval
abstract
Large-scale retrieval systems must operate under strict latency constraints while maintaining high recall. Sparse retrieval offers efficiency and interpretability, whereas dense retrieval provides stronger semantic matching. Although hybrid approaches combine both signals, their interaction is often limited, especially under intersection-based retrieval. We introduce HybridSparse, an end-to-end hybrid retrieval framework that strengthens sparse--dense interaction across modeling, training, and serving. It adopts a unified encoder with a shared backbone and jointly optimizes lexical and semantic representations through co-training. To further improve alignment, we incorporate hybrid score regularization and consistency distillation, enabling more stable and effective hybrid scoring. Experiments on public benchmarks demonstrate consistent improvements over strong sparse, dense, and hybrid baselines. In large-scale production deployment for Bing advertisement retrieval, HybridSparse delivers a +1.30% RPM gain, highlighting its practical impact.
Haotong Bao, Jianjin Zhang, Weihao Han, Qi Chen 0009, Dongzhe Jiang, Zhengxin Zeng, Mingzheng Li, Hao Sun 0015, Feng Sun 0008, Qi Zhang 0066
SIGIR8
2026 CausalSymptom: Learning Causal Disentangled Representation for Depression Severity Estimation on Transcribed Clinical Interviews
Mingzheng Li, Xiao Sun 0003, Xinke Wang, Feng-Qi Cui, Xun Yang 0001
IEEE Trans. Affect. Comput.1
2026 Modeling Long-Term Emotional Support Through Causal World Modeling With Imitation Learning
abstract
Emotional support conversation systems have emerged as a promising complement to traditional mental health consultations, offering context-aware dialogue to support seekers’ emotional well-being. Despite their potential, two fundamental challenges remain unresolved: 1) modeling long-term emotional trajectory beyond short-term relief; and 2) adapting support strategies to context in a psychologically coherent manner. To address these challenges, we propose CAIWO, a novel framework that integrates world modeling and causality-enhanced imitation learning to systematically support seekers, alleviate psychological stress, and restore emotional balance. Specifically, CAIWO comprises two core components. The first is an emotional world model, which captures long-term emotional trajectories from historical interactions to inform anticipatory guidance. The second is a causality-enhanced imitation learning module, which infers latent causal dependencies to facilitate coherent strategy transitions and mitigate compounding errors typical of conventional imitation learning. By incorporating the final latent variables into the response decoder, CAIWO dynamically adjusts the strategies and generates emotionally resonant responses. Extensive experiments on the ESConv benchmark demonstrate that CAIWO outperforms state-of-the-art baselines by 8.6%, significantly improving the generation of responses that align with seekers’ emotional development and psychological needs.
Mingzheng Li, Fei Wang 0073, Kun Li 0008, Yanyan Wei, Yiqi Nie, Yanbin Hao, Xun Yang 0001, Meng Wang 0001
IEEE Trans. Comput. Soc. Syst.1
2025 When Graph Meets Multimodal: Benchmarking and Meditating on Multimodal Attributed Graph Learning
abstract
Multimodal Attributed Graphs (MAGs) are ubiquitous in real-world applications, encompassing extensive knowledge through multimodal attributes attached to nodes (e.g., texts and images) and topological structure representing node interactions. Despite its potential to advance diverse research fields like social networks and e-commerce, MAG representation learning (MAGRL) remains underexplored due to the lack of standardized datasets and evaluation frameworks. In this paper, we first propose MAGB, a comprehensive MAG benchmark dataset, featuring curated graphs from various domains with both textual and visual attributes. Based on the MAGB dataset, we further systematically evaluate two mainstream MAGRL paradigms: GNN-as-Predictor, which integrates multimodal attributes via Graph Neural Networks (GNNs), and VLM-as-Predictor, which harnesses Vision Language Models (VLMs) for zero-shot reasoning. Extensive experiments on MAGB reveal the following critical insights: (i) Modality significances fluctuate drastically with specific domain characteristics. (ii) Multimodal embeddings can elevate the performance ceiling of GNNs. However, intrinsic biases among modalities may impede effective training, particularly in low-data scenarios. (iii) VLMs are highly effective at generating multimodal embeddings that alleviate the imbalance between textual and visual attributes. These discoveries, which illuminate the synergy between multimodal attributes and graph topologies, contribute to reliable benchmarks, paving the way for future research.
Hao Yan 0004, Chaozhuo Li, Jun Yin 0005, Weihao Han, Mingzheng Li, Zhengxin Zeng, Hao Sun 0015, Senzhang Wang
KDD (2)6
2025 Unleash LLMs Potential for Sequential Recommendation by Coordinating Dual Dynamic Index Mechanism
abstract
Owing to the unprecedented capability in semantic understanding and logical reasoning, large language models (LLMs) have shown fantastic potential in developing next-generation sequential recommender systems (RSs). However, existing LLM-based sequential RSs mostly separate index generation from sequential recommendation, leading to insufficient integration between semantic information and collaborative information. On the other hand, the neglect of user-related information hinders LLM-based sequential RSs from exploiting high-order user-item interaction patterns. In this paper, we propose the End-to-End Dual Dynamic (ED2) recommender, the first LLM-based sequential RS which adopts dual dynamic index mechanism, targeting resolving the above limitations simultaneously. The dual dynamic index mechanism can not only assembly index generation and sequential recommendation into a unified LLM-backbone pipeline, but also make it practical for LLM-based sequential recommender to take advantage of user-related information. Specifically, to facilitate the LLM comprehension ability to dual dynamic index, we propose a multigrained token regulator which constructs alignment supervision based on LLMs semantic knowledge across multiple representation granularities. Moreover, the associated user collection data and a series of novel instruction tuning tasks are specially customized to capture the high-order user-item interaction patterns. Extensive experiments on three public datasets demonstrate the superiority of ED2, achieving an average improvement of 19.62% in Hit-Rate and 21.11% in NDCG.
Jun Yin 0005, Zhengxin Zeng, Mingzheng Li, Hao Yan 0004, Chaozhuo Li, Weihao Han, Jianjin Zhang, Ruochen Liu 0001, Hao Sun 0015, Feng Sun 0008, Qi Zhang 0066, Shirui Pan, Senzhang Wang
WWW3
2025 Facial Depression Estimation via Multi-Cue Contrastive Learning
abstract
Vision-based depression estimation is an emerging yet impactful task, whose challenge lies in predicting the severity of depression from facial videos lasting at least several minutes. Existing methods primarily focus on fusing frame-level features to create comprehensive representations. However, they often overlook two crucial aspects: 1) inter- and intra-cue correlations, and 2) variations among samples. Hence, simply characterizing sample embeddings while ignoring to mine the relation among multiple cues leads to limitations. To address this problem, we propose a novel Multi-Cue Contrastive Learning (MCCL) framework to mine the relation among multiple cues for discriminative representation. Specifically, we first introduce a novel cross-characteristic attentive interaction module to model the relationship among multiple cues from four facial features (e.g., 3D landmarks, head poses, gazes, FAUs). Then, we propose a temporal segment attentive interaction module to capture the temporal relationships within each facial feature over time intervals. Moreover, we integrate contrastive learning to leverage the variations among samples by regarding the embeddings of inter-cue and intra-cue as positive pairs while considering embeddings from other samples as negative. In this way, the proposed MCCL framework leverages the relationships among the facial features and the variations among samples to enhance the process of multi-cue mining, thereby achieving more accurate facial depression estimation. Extensive experiments on public datasets, DAIC-WOZ, CMDC, and E-DAIC, demonstrate that our model not only outperforms the advanced depression methods but that the discriminative representations of facial behaviors provide potential insights about depression. Our code is available at:https://github.com/xkwangcn/MCCL.git
Xinke Wang, Xiao Sun 0003, Mingzheng Li, Bin Hu 0001, Dan Guo 0001, Meng Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 DTBIA: An Immersive Visual Analytics System for Brain-Inspired Research
abstract
The Digital Twin Brain (DTB) is an advanced artificial intelligence framework that integrates spiking neurons to simulate complex cognitive functions and collaborative behaviors. For domain experts, visualizing the DTB's simulation outcomes is essential to understanding complex cognitive activities. However, this task poses significant challenges due to DTB data's inherent characteristics, including its high-dimensionality, temporal dynamics, and spatial complexity. To address these challenges, we developed DTBIA, an Immersive Visual Analytics System for Brain-Inspired Research. In collaboration with domain experts, we identified key requirements for effectively visualizing spatiotemporal and topological patterns at multiple levels of detail. DTBIA incorporates a hierarchical workflow - ranging from brain regions to voxels and slice sections - along with immersive navigation and a 3D edge bundling algorithm to enhance clarity and provide deeper insights into both functional (BOLD) and structural (DTI) brain data. The utility and effectiveness of DTBIA are validated through two case studies involving with brain research experts. The results underscore the system's role in enhancing the comprehension of complex neural behaviors and interactions.
Jun-Hsiang Yao, Mingzheng Li, Yuxiao Li 0002, Jielin Feng, Jun Han 0010, Qibao Zheng, Jianfeng Feng, Siming Chen 0001
IEEE Trans. Vis. Comput. Graph.2
2024 Detecting Depression With Heterogeneous Graph Neural Network in Clinical Interview Transcript
abstract
Depression has an intense impact on individuals, yet many cases go undiagnosed. Thus, it is imperative to design an effective model for the automated diagnosis of depression. However, existing methods do not adequately capture contextual information in a clinical interview. Inspired by the depression diagnosis process, we propose a new perspective on detecting depression as a dialog information extraction task. Specifically, this article constructs a heterogeneous graph that models the participant’s depression state and uses the graph attention network to aggregate the pieces of depressive clues. In addition, we use the focal loss as a loss function for dealing with class imbalance by reshaping the standard cross-entropy loss. Experimental results demonstrate that our proposed model depression state extraction with heterogeneous graph attention neural network (DSE-HGAT) surpasses the baseline models on the Distress Analysis Interview Corpus-Wizard of Oz (DAIC-WOZ) dataset. Meanwhile, agreement analysis between our proposed model and the gold standard shows that it is moderate ($k$= 0.528,$p 0.05$). Overall, our model is very effective in identifying depression in the clinical interview transcript, which has the potential to assist doctors with medical conditions.
Mingzheng Li, Xiao Sun 0003, Meng Wang 0001
IEEE Trans. Comput. Soc. Syst.1
2023 An Adaptive Graph Pre-training Framework for Localized Collaborative Filtering
abstract
Graph neural networks (GNNs) have been widely applied in the recommendation tasks and have achieved very appealing performance. However, most GNN-based recommendation methods suffer from the problem of data sparsity in practice. Meanwhile, pre-training techniques have achieved great success in mitigating data sparsity in various domains such as natural language processing (NLP) and computer vision (CV) . Thus, graph pre-training has the great potential to alleviate data sparsity in GNN-based recommendations. However, pre-training GNNs for recommendations faces unique challenges. For example, user-item interaction graphs in different recommendation tasks have distinct sets of users and items, and they often present different properties. Therefore, the successful mechanisms commonly used in NLP and CV to transfer knowledge from pre-training tasks to downstream tasks such as sharing learned embeddings or feature extractors are not directly applicable to existing GNN-based recommendations models. To tackle these challenges, we delicately design an adaptive graph pre-training framework for localized collaborative filtering (ADAPT) . It does not require transferring user/item embeddings, and is able to capture both the common knowledge across different graphs and the uniqueness for each graph simultaneously. Extensive experimental results have demonstrated the effectiveness and superiority of ADAPT.
Yiqi Wang 0001, Chaozhuo Li, Zheng Liu 0011, Mingzheng Li, Jiliang Tang, Xing Xie 0001, Lei Chen 0002, Philip S. Yu
ACM Trans. Inf. Syst.4
2022 Localized Graph Collaborative Filtering
abstract
User-item interactions in recommendations can be naturally denoted as a user-item bipartite graph. Given the success of graph neural networks (GNNs) in graph representation learning, GNN-based Collaborative Filtering (CF) methods have been proposed to advance recommender systems. These methods often make recommendations based on the learned user and item embeddings. However, we found that they do not perform well with sparse user-item graphs which are quite common in real-world recommendations. Therefore, in this work, we introduce a novel perspective to build GNN-based CF methods for recommendations which leads to the proposed framework Localized Graph Collaborative Filtering (LGCF). One key advantage of LGCF is that it does not need to learn embeddings for each user and item, which is challenging in sparse scenarios. Alternatively, LGCF aims at encoding useful CF information into a localized graph and making recommendations based on such graph. Extensive experiments on various datasets validate the effectiveness of LGCF, especially in sparse scenarios. Furthermore, empirical results demonstrate that LGCF provides complementary information to the embedding-based CF model which can be utilized to boost recommendation performance.
Yiqi Wang 0001, Chaozhuo Li, Mingzheng Li, Wei Jin 0009, Hao Sun 0015, Xing Xie 0001, Jiliang Tang
SDM3
2021 Sentiment analysis of Chinese stock reviews based on BERT model
Mingzheng Li
Appl. Intell.1
2019 SPARC: Towards a Scalable Distributed Control Plane Architecture for Protocol-Oblivious SDN Networks
abstract
High-level programming abstraction and large-scale deployment have become two important trends of software-defined networking (SDN) in the past decade. Using high-level program to manage the large-scale network faces more serious control plane extensibility problem due to its complex intermediate representation and protocol-independent feature. To address this problem, we propose SPARC, a programmable and scalable controller architecture, which employs a hybrid hierarchical structure to maximize flexibility with regard to control plane distribution. Our architecture also allows for pushing down control decision making closer to the data plane and localize network event processing to lower the latency of control plane operations while exploiting SDN's global visibility to build optimal policy decisions. Furthermore, we investigate the feasibility of SPARC by exemplifying the case of delivering ICN mobility services and then conduct evaluations to demonstrate the efficacy of our design.
Mingzheng Li, Xiaodong Wang 0013, Haojie Tong, Ye Tian 0004
ICCCN1
2018 Design and Implementation of a Novel SDN-Based Architecture for Wi-Fi Networks
Mingzheng Li, Lei Mei, Ye Tian 0004
PDCAT2
2018 PNPL: Simplifying programming for protocol-oblivious SDN networks
Xiaodong Wang 0013, Ye Tian 0004, Mingzheng Li, Lei Mei, Xinming Zhang 0001
Comput. Networks4
2017 Quality-Aware Movie Recommendation System on Big Data
abstract
The movie recommendation is one of the most active application domain for recommendation systems (RS). However, with the rapid growth in the number of films, users have vastly different needs for the quality of the movie. In addition, facing big data, the traditional stand-alone RS is incapable to meet the need of an accurate and prompt recommendation. Aiming at solving these challenges, in this paper, we first parallelize the collaborative filtering to improve the computational efficiency, then we propose a quality-aware big data based movie recommendation system.
Mingzheng Li, Wangsong Wang, Pengcheng Xuan, Kun Geng
BDCAT2