Ming Pang

dblp:59/5011 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 ADORE: Autonomous Domain-Oriented Relevance Engine for E-commerce
abstract
Relevance modeling in e-commerce search remains challenged by semantic gaps in term-matching methods (e.g., BM25) and neural models' reliance on the scarcity of domain-specific hard samples. We propose ADORE, a self-sustaining framework that synergizes three innovations: (1) A Rule-aware Relevance Discrimination module, where a Chain-of-Thought LLM generates intent-aligned training data, refined via Kahneman-Tversky Optimization (KTO) to align with user behavior; (2) An Error-type-aware Data Synthesis module that auto-generates adversarial examples to harden robustness; and (3) A Key-attribute-enhanced Knowledge Distillation module that injects domain-specific attribute hierarchies into a deployable student model. ADORE automates annotation, adversarial generation, and distillation, overcoming data scarcity while enhancing reasoning. Large-scale experiments and online A/B testing verify the effectiveness of ADORE. The framework establishes a new paradigm for resource-efficient, cognitively aligned relevance modeling in industrial applications.
Donghao Xie, Ming Pang, Chunyuan Yuan, Changping Peng, Zhangang Lin
SIGIR3
2025 Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising
abstract
The retrieval system is a crucial module in e-commerce search advertising that matches user queries with ads. The diverse expressions of users often produce massive tail queries that cannot match merchant bidwords, leading to poor retrieval efficiency. Existing methods, such as query log mining and vector matching, fail to optimize relevance, authenticity, and ad revenue of the rewrite.
Zhenhui Liu, Chunyuan Yuan, Ming Pang, Li Yuan 0007, Changping Peng, Zhangang Lin, Jingping Shao
SIGIR3
2025 Adaptive Sub-Nanometer Control of a Piezoelectric Positioning Platform
abstract
This paper reports an asymmetric Bouc-Wen (ABW) hysteresis model and a hybrid control algorithm based on multi-modal Bayesian gradient optimization (MBGO) for trajectory tracking in the micro-positioning phase of a piezoelectric positioning platform. First, a system-level dynamic model capable of expressing hysteresis nonlinearity is established based on the asymmetric Bouc-Wen model. Second, an MBGO parameter identification algorithm based on Particle Swarm Optimization (PSO) is proposed to improve the characterization capability of the hysteresis model. Subsequently, a feedforward adaptive fuzzy PID (FF-AFPID) composite controller is designed by compensating the hysteresis nonlinearity through the ABW inverse model while dynamically adjusting PID parameters with adaptive fuzzy rules. Through triangular and sinusoidal trajectory tracking experiments, the root mean square errors were reduced to 0.112 nm and 0.103 nm by the FF-AFPID, with an improvement of 74.944%, 57.088% (triangular) and 77.511%, 57.083% (sinusoidal) over the FF-PID and FF-FPID algorithms, respectively. The results demonstrate that the trajectory tracking of performance the positioning platform in micro-positioning phase was significantly enhanced by FF-AFPID, with the maximum error being suppressed to sub-nanometer levels.
Siyuan Meng, Jiankang Jiang, Qian Ju, Dongmei Wu, Wei Dong 0004, Ming Pang, Changhai Ru
IEEE Trans Autom. Sci. Eng.7
2024 Towards Better Seach Query Classification with Distribution-Diverse Multi-Expert Knowledge Distillation in JD Ads Search
abstract
In the dynamic landscape of online advertising, decoding user intent remains a pivotal challenge, particularly in the context of query classification. Swift classification models, exemplified by FastText, cater to the demand for real-time responses but encounter limitations in handling intricate queries. Conversely, accuracy-centric models like BERT introduce challenges associated with increased latency. This paper undertakes a nuanced exploration, navigating the delicate balance between efficiency and accuracy. It unveils FastText's latent potential as an 'online dictionary' for historical queries while harnessing the semantic robustness of BERT for novel and complex scenarios. The proposed Distribution-Diverse Multi-Expert (DDME) framework employs multiple teacher models trained from diverse data distributions. Through meticulous data categorization and enrichment, it elevates the classification performance across the query spectrum. Empirical results within the JD ads search system validate the superiority of our proposed approaches.
Kun-Peng Ning, Ming Pang, Xiwei Zhao, Changping Peng, Zhangang Lin, Jinghe Hu, Jingping Shao, Li Yuan 0007
CIKM2
2024 TentISSA-BPNN: a novel evaluation model for cloud service providers for petroleum enterprises
Ke Hou, Jianping Sun, Mingcheng Guo, Ming Pang
J. Supercomput.4
2024 Data Provenance via Differential Auditing
abstract
With the rising awareness of data assets, data governance, which is to understand where data comes from, how it is collected, and how it is used, has been assuming evergrowing importance. One critical component of data governance gaining increasing attention is auditing machine learning models to determine if specific data has been used for training. Existing auditing techniques, like shadow auditing methods, have shown feasibility under specific conditions such as having access to label information and knowledge of training protocols. However, these conditions are often not met in most real-world applications. In this paper, we introduce a practical framework for auditing data provenance based on a differential mechanism, i.e., after carefully designed transformation, perturbed input data from the target model's training set would result in much more drastic changes in the output than those from the model's non-training set. Our framework is data-dependent and does not require distinguishing training data from non-training data or training additional shadow models with labeled output data. Furthermore, our framework extends beyond point-based data auditing to group-based data auditing, aligning with the needs of realworld applications. Our theoretical analysis of the differential mechanism and the experimental results on real-world data sets verify the proposal's effectiveness. The codes have been uploaded in an anonymous link.
Xin Mu, Ming Pang, Feida Zhu 0001
IEEE Trans. Knowl. Data Eng.2
2022 Protein structure prediction based on particle swarm optimization and tabu search strategy
abstract
BACKGROUND: The stability of protein sequence structure plays an important role in the prevention and treatment of diseases. RESULTS: In this paper, particle swarm optimization and tabu search are combined to propose a new method for protein structure prediction. The experimental results show that: for four groups of artificial protein sequences with different lengths, this method obtains the lowest potential energy value and stable structure prediction results, and the effect is obviously better than the other two comparison methods. Taking the first group of protein sequences as an example, our method improves the prediction of minimum potential energy by 127% and 7% respectively. CONCLUSIONS: Therefore, the method proposed in this paper is more suitable for the prediction of protein structural stability.
Shuchun Yu, Li Xianxiang, Tian Xue, Ming Pang
BMC Bioinform.4
2022 Improving Deep Forest by Screening
abstract
Most studies about deep learning are based on neural network models, where many layers of parameterized nonlinear differentiable modules are trained by backpropagation. Recently, it has been shown that deep learning can also be realized by non-differentiable modules without backpropagation training called deep forest. We identify that deep forest has high time costs and memory requirements—this has inhibited its use on large-scale datasets. In this paper, we propose a simple and effective approach with three main strategies for efficient learning of deep forest. First, it substantially reduces the number of instances that needs to be processed through redirecting instances having high predictive confidence straight to the final level for prediction, by-passing all the intermediate levels. Second, many non-informative features are screened out, and only the informative ones are used for learning at each level. Third, an unsupervised feature transformation procedure is proposed to replace the supervised multi-grained scanning procedure. Our theoretical analysis supports the proposed approach in varying the model complexity from low to high as the number of levels increases in deep forest. Experiments show that our approach achieves highly competitive predictive performance with reduced time cost and memory requirement by one to two orders of magnitude.
Ming Pang, Kai Ming Ting, Peng Zhao 0006, Zhi-Hua Zhou
IEEE Trans. Knowl. Data Eng.1
2021 Reconstruction-based Anomaly Detection with Completely Random Forest
Yi-Xuan Xu, Ming Pang, Ji Feng, Kai Ming Ting, Yuan Jiang 0001, Zhi-Hua Zhou
SDM2
2020 Monitoring of human body running training with wireless sensor based wearable devices
Fuyu Guan, Ming Pang, Shuangling Li
Comput. Commun.3
2018 Improving Deep Forest by Confidence Screening
abstract
Most studies about deep learning are based on neural network models, where many layers of parameterized nonlinear differentiable modules are trained by backpropagation. Recently, it has been shown that deep learning can also be realized by non-differentiable modules without backpropagation training called deep forest. The developed representation learning process is based on a cascade of cascades of decision tree forests, where the high memory requirement and the high time cost inhibit the training of large models. In this paper, we propose a simple yet effective approach to improve the efficiency of deep forest. The key idea is to pass the instances with high confidence directly to the final stage rather than passing through all the levels. We also provide a theoretical analysis suggesting a means to vary the model complexity from low to high as the level increases in the cascade, which further reduces the memory requirement and time cost. Our experiments show that the proposed approach achieves highly competitive predictive performance with significantly reduced time cost and memory requirement by up to one order of magnitude.
Ming Pang, Kai Ming Ting, Peng Zhao 0006, Zhi-Hua Zhou
ICDM1
2018 Unorganized Malicious Attacks Detection
abstract
Recommender systems have attracted much attention during the past decade. Many attack detection algorithms have been developed for better recommendations, mostly focusing on shilling attacks, where an attack organizer produces a large number of user profiles by the same strategy to promote or demote an item. This work considers another different attack style: unorganized malicious attacks, where attackers individually utilize a small number of user profiles to attack different items without organizer. This attack style occurs in many real applications, yet relevant study remains open. We formulate the unorganized malicious attacks detection as a matrix completion problem, and propose the Unorganized Malicious Attacks detection (UMA) algorithm, based on the alternating splitting augmented Lagrangian method. We verify, both theoretically and empirically, the effectiveness of the proposed approach.
Ming Pang, Wei Gao 0008, Zhi-Hua Zhou
NeurIPS1
2006 A Streaming Implementation of Transform and Quantization in H.264
Chunyuan Zhang, Li Li 0005, Ming Pang
HPCC4