Mingjun Zhao

dblp:207/0270 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity Mixture-of-Experts
abstract
Zhenyu Liu, Yunxin li, Xuanyu Zhang, Qixun Teng, Shenyuan Jiang, Xinyu Chen, Haoyuan Shi, Haolan Chen, Fanbo Meng, Mingjun Zhao, Yu Xu, Yancheng He, Baotian Hu, Haizhou Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yunxin Li, Xuanyu Zhang 0006, Qixun Teng, Shenyuan Jiang, Xinyu Chen 0003, Haolan Chen, Mingjun Zhao, Yancheng He, Baotian Hu, Haizhou Li 0001, Min Zhang 0005
ACL (1)10
2026 Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs
abstract
Zhenyu Liu, Xuanyu Zhang, Yunxin li, Qixun Teng, Shenyuan Jiang, Haolan Chen, Mingjun Zhao, Fanbo Meng, Yu Xu, Yancheng He, Baotian Hu, Haizhou Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xuanyu Zhang 0006, Yunxin Li, Qixun Teng, Shenyuan Jiang, Haolan Chen, Mingjun Zhao, Yancheng He, Baotian Hu, Haizhou Li 0001, Min Zhang 0005
ACL (1)7
2026 Early Warning of Dangerous Electricity Consumption Behavior Based on Unsupervised Clustering and Multiscale Ensemble Learning
abstract
The dangerous electricity consumption behaviors of enterprises may lead to energy-related safety incidents during the production process, like overloading. To address these risks, the government has increasingly relied on Internet of Things (IoT)-enabled monitoring systems to analyze electricity consumption curves in real-time. The essence of government monitoring electricity consumption behaviors is to conduct a comparative analysis of electricity consumption curves in combination with historical data. However, there are two difficulties in the working process: 1. Clustering based on time-series data is the core of electricity curve analysis, but existing methods are only suitable for clustering discrete data points. 2. It is prone to misjudge behaviors not existing in historical data, lacking of common feature capture for behaviors. Therefore, this paper proposes an IoT-integrated early warning method operated within a multi-layered IoT architecture (perception → network → edge → cloud). First of all, designing a time-series feature capture enhancement module based on Self-Organizing Mapping (SOM) and Hierarchical Clustering algorithm (HCA). The SOM is utilized to capture the features of the power consumption curves, and the capture effect is enhanced through reintegration by HCA. Secondly, constructing a multi-scale feature extraction ensemble fusion framework. Exploring the common features of electricity consumption behavior avoids the situation of misjudgment during the work process. Finally, through the practical application in the regional power grid of a certain area in China, the proposed method achieves an early warning accuracy of 97.99%, which is more than 20% higher than that of the current methods.
Tianhao Ma, Zhifang Yang, Mingjun Zhao, Mao Fan
IEEE Internet Things J.4
2025 Time-Resolved Laser Speckle Contrast Imaging (TR-LSCI) of Cerebral Blood Flow
abstract
To address many of the deficiencies in optical neuroimaging technologies, such as poor tempo-spatial resolution, low penetration depth, contact-based measurement, and time-consuming image reconstruction, a novel, noncontact, portable, time-resolved laser speckle contrast imaging (TR-LSCI) technique has been developed for continuous, fast, and high-resolution 2D mapping of cerebral blood flow (CBF) at different depths of the head. TR-LSCI illuminates the head with picosecond-pulsed, coherent, widefield near-infrared light and synchronizes a fast, high-resolution, gated single-photon avalanche diode camera to selectively collect diffuse photons with longer pathlengths through the head, thus improving the accuracy of CBF measurement in the deep brain. The reconstruction of a CBF map was dramatically expedited by incorporating convolution functions with parallel computations. The performance of TR-LSCI was evaluated using head-simulating phantoms with known properties and in-vivo rodents with varied hemodynamic challenges to the brain. TR-LSCI enabled mapping CBF variations at different depths with a sampling rate of up to 1 Hz and spatial resolutions ranging from tens/hundreds of micrometers on rodent head surfaces to 1-2 millimeters in deep brains. With additional improvements and validation in larger populations against established methods, we anticipate offering a noncontact, fast, high-resolution, portable, and affordable brain imager for fundamental neuroscience research in animals and for translational studies in humans.
Faraneh Fathi, Siavash Mazdeyasna, Dara Singh, Chong Huang 0003, Mehrana Mohtasebi, Xuhui Liu, Samaneh Rabienia Haratbar, Mingjun Zhao, Arin C. Ulku, Paul Mos, Claudio Bruschini, Edoardo Charbon, Guoqiang Yu
IEEE Trans. Medical Imaging8
2024 Multi-task online allocation based on path planning strategy in Spatial crowdsourcing environment
abstract
This article takes the spatiotemporal crowdsourcing task allocation in a multi worker and multi task environment as the background, and designs a candidate task set algorithm based on region partitioning model for scenarios where workers can accept multiple tasks at the same time. Firstly, the tasks are stored in different regions, named and indexed using the GEOhash method. Secondly, based on the network maximum cost flow model for task allocation, four strategies are proposed to maximize the utility of online multi task allocation. The results show that the CFTA algorithm out-performs the greedy algorithm and random threshold method in terms of total utility, task allocation success rate, and average worker income. In terms of runtime, the CFTA algorithm can effectively solve the multi worker and multi task allocation problem without sacrificing a small amount of time cost.
Yiduo Cheng, Dunhui Yu, Mingjun Zhao
CSCWD4
2024 DimReg: Embedding Dimension Search via Regularization for Recommender Systems
abstract
Modern recommender systems aim to identify items that are most pertinent to a particular user and are particularly useful when an overwhelming number of items are present. Feature embedding is essential to deep recommender systems, which constructs memory-efficient and semantically meaningful representations by mapping high-dimensional sparse feature vectors into low-dimensional dense vectors. Most existing systems assign a unified dimension to all feature fields, regardless of the diverse importance of different features, which usually results in sub-optimal performance and high memory usage. In this paper, we propose a low-cost embedding dimension search approach named DimReg for recommender systems, by assessing information overlapping between the dimensions within each feature field and pruning unimportant and redundant dimensions progressively during model training via a two-level polarization regularizer, while introducing minimum overhead. Moreover, our method does not require retraining after embedding dimension search, which significantly reduces the computational cost and is more friendly to deployment in real-world recommender systems. Extensive experiments conducted on multiple CTR (Click Through Rate) prediction tasks demonstrate that our method can efficiently reduce the model parameters up to 98.6%, and achieve strong recommendation performance outperforming existing automated embedding dimension search methods.
Mingjun Zhao, Liyao Jiang, Yakun Yu, Xinmin Wang, Zheng Wei 0004, Di Niu 0002
SDM1
2023 Search-Map-Search: A Frame Selection Paradigm for Action Recognition
abstract
Despite the success of deep learning in video understanding tasks, processing every frame in a video is computationally expensive and often unnecessary in real-time applications. Frame selection aims to extract the most informative and representative frames to help a model better understand video content. Existing frame selection methods either individually sample frames based on per-frame importance prediction, without considering interaction among frames, or adopt reinforcement learning agents to find representative frames in succession, which are costly to train and may lead to potential stability issues. To overcome the limitations of existing methods, we propose a Search-Map-Search learning paradigm which combines the advantages of heuristic search and supervised learning to select the best combination of frames from a video as one entity. By combining search with learning, the proposed method can better capture frame interactions while incurring a low inference overhead. Specifically, we first propose a hierarchical search method conducted on each training video to search for the optimal combination of frames with the lowest error on the downstream task. A feature mapping function is then learned to map the frames of a video to the representation of its target optimal frame combination. During inference, another search is performed on an unseen video to select a combination of frames whose feature representation is close to the projected feature representation. Extensive experiments based on several action recognition benchmarks demonstrate that our frame selection method effectively improves performance of action recognition models, and significantly outperforms a number of competitive baselines.
Mingjun Zhao, Yakun Yu, Xiaoli Wang 0004, Di Niu 0002
CVPR1
2023 Online Volume Optimization for Notifications via Long Short-Term Value Modeling
Mingjun Zhao, Weiyu Tou, Haolan Chen, Di Niu 0002, Cunxiang Yin, Yancheng He
PAKDD (3)2
2023 BDA: Bandit-based Transferable AutoAugment
abstract
AutoAugment is an automatic method to design data augmentation policies for deep learning, and has achieved significant improvements on computer vision tasks. However, since early AutoAugment approaches cost thousands of GPU hours, there is a recent demand to investigate low-cost search methods that can still find effective augmentation policies. In this paper, we propose a multi-armed bandit algorithm, named Bandit Data Augment (BDA), to efficiently search for optimal and transferable data augmentation policies. We leverage Successive Halving to make the bandit model progressively focus on more promising augmentation operations during the search, leading to sparse selection of operations and more generalizable augmentation policies. We also propose a computationally efficient rewarding scheme to reduce the evaluation cost of augmentation policies. Extensive experiments demonstrate that BDA can achieve comparable or better performance than prior Auto Augment methods on a wide range of models on CIFAR-10/100 and ImageNet benchmarks. Besides, BDA is 555 times and 536 times faster than AutoAugment on CIFAR-10 and ImageNet, respectively. In addition, BDA is 16 times faster than Fast Auto Augment on ImageNet. More importantly, BDA can discover policies that are transferable across datasets and models, and achieve similar performance to policies found directly on the target dataset.
Mingjun Zhao, Songling Yuan, Xiaoli Wang 0004, Di Niu 0002
SDM2
2023 CEIL: A General Classification-Enhanced Iterative Learning Framework for Text Clustering
abstract
Text clustering, as one of the most fundamental challenges in unsupervised learning, aims at grouping semantically similar text segments without relying on human annotations. With the rapid development of deep learning, deep clustering has achieved significant advantages over traditional clustering methods. Despite the effectiveness, most existing deep text clustering methods rely heavily on representations pre-trained in general domains, which may not be the most suitable solution for clustering in specific target domains. To address this issue, we propose CEIL, a novel Classification-Enhanced Iterative Learning framework for short text clustering, which aims at generally promoting the clustering performance by introducing a classification objective to iteratively improve feature representations. In each iteration, we first adopt a language model to retrieve the initial text representations, from which the clustering results are collected using our proposed Category Disentangled Contrastive Clustering (CDCC) algorithm. After strict data filtering and aggregation processes, samples with clean category labels are retrieved, which serve as supervision information to update the language model with the classification objective via a prompt learning approach. Finally, the updated language model with improved representation ability is used to enhance clustering in the next iteration. Extensive experiments demonstrate that the CEIL framework significantly improves the clustering performance over iterations, and is generally effective on various clustering algorithms. Moreover, by incorporating CEIL on CDCC, we achieve the state-of-the-art clustering performance on a wide range of short text clustering benchmarks outperforming other strong baseline methods.
Mingjun Zhao, Mengzhen Wang, Yinglong Ma 0001, Di Niu 0002, Haijiang Wu
WWW1
2022 LA3: Efficient Label-Aware AutoAugment
Mingjun Zhao, Xiaoli Wang 0004, Di Niu 0002
ECCV (21)1
2022 RecGURU: Adversarial Learning of Generalized User Representations for Cross-Domain Recommendation
abstract
Cross-domain recommendation can help alleviate the data sparsity issue in traditional sequential recommender systems. In this paper, we propose the RecGURU algorithm framework to generate a Generalized User Representation (GUR) incorporating user information across domains in sequential recommendation, even when there is minimum or no common users in the two domains. We propose a self-attentive autoencoder to derive latent user representations, and a domain discriminator, which aims to predict the origin domain of a generated latent representation. We propose a novel adversarial learning method to train the two modules to unify user embeddings generated from different domains into a single global GUR for each user. The learned GUR captures the overall preferences and characteristics of a user and thus can be used to augment the behavior data and improve recommendations in any single domain in which the user is involved. Extensive experiments have been conducted on two public cross-domain recommendation datasets as well as a large dataset collected from real-world applications. The results demonstrate that RecGURU boosts performance and outperforms various state-of-the-art sequential recommendation and cross-domain recommendation methods. The collected data will be released to facilitate future research.
Mingjun Zhao, Huanming Zhang, Chenyun Yu, Lei Cheng 0005, Guoqiang Shu, Beibei Kong, Di Niu 0002
WSDM2
2021 Verdi: Quality Estimation and Error Detection for Bilingual Corpora
Mingjun Zhao, Haijiang Wu, Di Niu 0002, Xiaoli Wang 0004
WWW1
2021 QBSUM: A large-scale query-based document summarization dataset from real-world applications
Mingjun Zhao, Shengli Yan, Bang Liu 0003, Xinwang Zhong, Qian Hao, Haolan Chen, Di Niu 0002, Bowei Long, Weidong Guo
Comput. Speech Lang.1
2020 Reinforced Curriculum Learning on Pre-Trained Neural Machine Translation Models
abstract
The competitive performance of neural machine translation (NMT) critically relies on large amounts of training data. However, acquiring high-quality translation pairs requires expert knowledge and is costly. Therefore, how to best utilize a given dataset of samples with diverse quality and characteristics becomes an important yet understudied question in NMT. Curriculum learning methods have been introduced to NMT to optimize a model's performance by prescribing the data input order, based on heuristics such as the assessment of noise and difficulty levels. However, existing methods require training from scratch, while in practice most NMT models are pre-trained on big data already. Moreover, as heuristics, they do not generalize well. In this paper, we aim to learn a curriculum for improving a pre-trained NMT model by re-selecting influential data samples from the original training set and formulate this task as a reinforcement learning problem. Specifically, we propose a data selection framework based on Deterministic Actor-Critic, in which a critic network predicts the expected change of model performance due to a certain sample, while an actor network learns to select the best sample out of a random batch of samples presented to it. Experiments on several translation datasets show that our method can further improve the performance of NMT when original batch training reaches its ceiling, without using additional new training data, and significantly outperforms several strong baseline methods.
Mingjun Zhao, Haijiang Wu, Di Niu 0002, Xiaoli Wang 0004
AAAI1
2019 Learning to Generate Questions by LearningWhat not to Generate
abstract
Automatic question generation is an important technique that can improve the training of question answering, help chatbots to start or continue a conversation with humans, and provide assessment materials for educational purposes. Existing neural question generation models are not sufficient mainly due to their inability to properly model the process of how each word in the question is selected, i.e., whether repeating the given passage or being generated from a vocabulary. In this paper, we propose our Clue Guided Copy Network for Question Generation (CGC-QG), which is a sequence-to-sequence generative model with copying mechanism, yet employing a variety of novel components and techniques to boost the performance of question generation. In CGC-QG, we design a multi-task labeling strategy to identify whether a question word should be copied from the input passage or be generated instead, guiding the model to learn the accurate boundaries between copying and generation. Furthermore, our input passage encoder takes as input, among a diverse range of other features, the prediction made by a clue word predictor, which helps identify whether each word in the input passage is a potential clue to be copied into the target question. The clue word predictor is designed based on a novel application of Graph Convolutional Networks onto a syntactic dependency tree representation of each passage, thus being able to predict clue words only based on their context in the passage and their relative positions to the answer in the tree. We jointly train the clue prediction as well as question generation with multi-task learning and a number of practical strategies to reduce the complexity. Extensive evaluations show that our model significantly improves the performance of question generation and out-performs all previous state-of-the-art neural question generation models by a substantial margin.
Bang Liu 0003, Mingjun Zhao, Di Niu 0002, Kunfeng Lai, Yancheng He, Haojie Wei
WWW2
2017 Noncontact 3-D Speckle Contrast Diffuse Correlation Tomography of Tissue Blood Flow Distribution
abstract
Recent advancements in near-infrared diffuse correlation techniques and instrumentation have opened the path for versatile deep tissue microvasculature blood flow imaging systems. Despite this progress there remains a need for a completely noncontact, noninvasive device with high translatability from small/testing (animal) to large/target (human) subjects with trivial application on both. Accordingly, we discuss our newly developed setup which meets this demand, termed noncontact speckle contrast diffuse correlation tomography (nc_scDCT). The nc_scDCT provides fast, continuous, portable, noninvasive, and inexpensive acquisition of 3-D tomographic deep (up to 10 mm) tissue blood flow distributions with straightforward design and customization. The features presented include a finite-element-method implementation for incorporating complex tissue boundaries, fully noncontact hardware for avoiding tissue compression and interactions, rapid data collection with a diffuse speckle contrast method, reflectance-based design promoting experimental translation, extensibility to related techniques, and robust adjustable source and detector patterns and density for high resolution measurement with flexible regions of interest enabling unique application-specific setups. Validation is shown in the detection and characterization of both high and low contrasts in flow relative to the background using tissue phantoms with a pump-connected tube (high) and phantom spheres (low). Furthermore, in vivo validation of extracting spatiotemporal 3-D blood flow distributions and hyperemic response during forearm cuff occlusion is demonstrated. Finally, the success of instrument feasibility in clinical use is examined through the intraoperative imaging of mastectomy skin flap.
Chong Huang 0003, Daniel Irwin, Mingjun Zhao, Nneamaka Agochukwu, Lesley Wong, Guoqiang Yu
IEEE Trans. Medical Imaging3