Mingyuan Tao

dblp:289/5997 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 11 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 APPNet: Automatic Feature Partitioning-Based Parameter Personalized Network for Conversion Prediction in E-commerce
abstract
Traditional conversion rate prediction models suffer from suboptimal performance due to sharing the same network parameters for all instances, failing to capture heterogeneous underlying distributions across instances. Recent parameter personalized network based models address this by grouping instances and adjust parameters for each group. However, existing parameter personalization methods face challenges: (1) taking prior information features as grouping condition for model parameter personalization leads to suboptimal performance due to human's limited understanding of data distribution, or (2) using all the features for both parameter generation module and deep neural network (DNN) of conversion prediction tasks causes gradient conflicts during backpropagation. A better approach is to automatically select features as grouping condition based on data distribution through iterative learning. Therefore, we propose Automatic Feature Partitioning-Based Parameter Personalized Network (APPNet), which consists of two components: Automatic Feature Partitioning (AFP) and Parameter Personalized Network (PPNet). The AFP module automatically partitions all the features into two parts: one part for DNN of conversion prediction tasks, and the other part for PPNet module to generate weights to adjust DNN parameters of conversion prediction tasks. Specifically, we implemented two versions of AFP: feature-wise AFP and bit-wise AFP. The feature-wise AFP partitions features at the feature field granularity, while the bit-wise AFP partitions each bit of the feature embeddings. The PPNet module adjusts model parameters of conversion prediction task for each group of instances by applying element-wise multiplication to the DNN parameters of conversion tasks. Extensive offline experiments demonstrate APPNet outperforms previous parameter personalized models. Furthermore, online A/B testing in production system achieved a 1.09% improvement on conversion rate, validating its practical effectiveness.
Mingyuan Tao, Maofei Que, Pan Li 0008, Zhuoran Zhuang
WSDM2
2025 NAM: A Normalization Attention Model for Personalized Product Search In Fliggy
abstract
Personalized product search provides significant benefits to e-commerce platforms by extracting more accurate user preferences from historical behaviors. Previous studies largely focused on the user factors when personalizing the search query, while ignoring the item perspective, which leads to the following two challenges that we summarize in this paper: First, previous approaches relying only on co-occurrence frequency tend to overestimate the conversion rates for popular items and underestimate those for long-tail items, resulting in inaccurate item similarities; Second, user purchasing propensity is highly heterogeneous according to the popularity of the target item: it is less correlated with the user's historical behavior for a popular item and more correlated for a long-tail item. To address these challenges, in this paper we propose NAM, a Normalization Attention Model, which optimizes ''when to personalize'' by utilizing Inverse Item Frequency (IIF) and employing a gating mechanism, as well as optimizes ''how to personalize'' by normalizing the attention mechanism from a global perspective. Through comprehensive experiments, we demonstrate that our proposed NAM model significantly outperforms state-of-the-art baseline models. Furthermore, we conducted an online A/B test at Fliggy, and obtained a significant improvement of 0.8% over the latest production system in conversion rate.
Mingyuan Tao, Maofei Que, Pan Li 0008, Dong Li 0037, Shenghua Ni, Zhuoran Zhuang
SIGIR2
2024 Changenet: Multi-Temporal Asymmetric Change Detection Dataset
abstract
Change Detection (CD) has been attracting extensive interests with the availability of bi-temporal datasets. However, due to the huge cost of multi-temporal images acquisition and labeling, existing change detection datasets are small in quantity, short in temporal, and low in practicability. Therefore, a large-scale practical-oriented dataset covering wide temporal phases is urgently needed to facilitate the community. To this end, the ChangeNet dataset is presented especially for multi-temporal change detection, along with the new task of "Asymmetric Change Detection". Specifically, ChangeNet consists of 31,000 multi-temporal images pairs, a wide range of complex scenes from 100 cities, and 6 pixel-level annotated categories, which is far superior to all the existing change detection datasets including LEVIR-CD, WHU Building CD, etc.. In addition, ChangeNet contains amounts of real-world perspective distortions in different temporal phases on the same areas, which is able to promote the practical application of change detection algorithms. The ChangeNet dataset is suitable for both binary change detection (BCD) and semantic change detection (SCD) tasks. Accordingly, we benchmark the ChangeNet dataset on six BCD methods and two SCD methods, and extensive experiments demonstrate its challenges and great significance. The dataset is available at https://github.com/jankyee/ChangeNet.
Deyi Ji, Mingyuan Tao, Hongtao Lu 0001, Feng Zhao 0004
ICASSP3
2024 INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection
abstract
Knowledge hallucination have raised widespread concerns for the security and reliability of deployed LLMs. Previous efforts in detecting hallucinations have been employed at logit-level uncertainty estimation or language-level self-consistency evaluation, where the semantic information is inevitably lost during the token-decoding procedure. Thus, we propose to explore the dense semantic information retained within LLMs' \textbf{IN}ternal \textbf{S}tates for halluc\textbf{I}nation \textbf{DE}tection (\textbf{INSIDE}). In particular, a simple yet effective \textbf{EigenScore} metric is proposed to better evaluate responses' self-consistency, which exploits the eigenvalues of responses' covariance matrix to measure the semantic consistency/diversity in the dense embedding space. Furthermore, from the perspective of self-consistent hallucination detection, a test time feature clipping approach is explored to truncate extreme activations in the internal states, which reduces overconfident generations and potentially benefits the detection of overconfident hallucinations. Extensive experiments and ablation studies are performed on several popular LLMs and question-answering (QA) benchmarks, showing the effectiveness of our proposal.
Chao Chen 0026, Kai Liu 0023, Ze Chen 0001, Mingyuan Tao, Zhihang Fu, Jieping Ye
ICLR6
2024 Water Salinity Sensing with UAV-Mounted IR-UWB Radar
abstract
The quality of surface water is closely related to human’s production and livelihood. Water salinity is one of the key indicators of water quality assessment. Recently, there has been an increased salinization problem of surface water in many regions of the world, making it necessary to timely monitor the salinity of surface water. Water salinity sensing could be challenging when it comes to surface water with complicated basin and tributaries, where existing methods fail to satisfy both efficiency and accuracy requirements. To address this problem, we propose a novel water salinity sensing system USalt, which leverages the high mobility of UAV and the contactless sensing ability of IR-UWB radar, and realizes fast and accurate water salinity sensing for surface water. Specifically, we design novel methods to eliminate the contamination in raw received radar signals and extract salinity-related features from radar signals. Furthermore, we adopt a neural network model ssNet to precisely estimate water salinity using the extracted features. To efficiently adapt ssNet to different environments, we customize meta learning and design a meta-learning framework mssNet. Extensive real-world experiments carried out by our UAV-based system illustrate that USalt can accurately sense the salinity of water with an MAE of 0.39 g/100 mL.
Guiyun Fan, Haiming Jin, Wentian Hao, Mingyuan Tao
ACM Trans. Sens. Networks6
2023 Ultra-High Resolution Segmentation with Ultra-Rich Context: A Novel Benchmark
abstract
With the increasing interest and rapid development of methods for Ultra-High Resolution (UHR) segmentation, a large-scale benchmark covering a wide range of scenes with full fine-grained dense annotations is urgently needed to facilitate the field. To this end, the URUR dataset is introduced, in the meaning of Ultra-High Resolution dataset with Ultra-Rich Context. As the name suggests, URUR contains amounts of images with high enough resolution (3,008 images of size 5,120 × 5,120), a wide range of complex scenes (from 63 cities), rich-enough context (1 million instances with 8 categories) and fine-grained annotations (about 80 billion manually annotated pixels), which is far superior to all the existing UHR datasets including DeepGlobe, Inria Aerial, UDD, etc.. Moreover, we also propose WSDNet, a more efficient and effective framework for UHR segmentation especially with ultra-rich context. Specifically, multi-level Discrete Wavelet Transform (DWT) is naturally integrated to release computation burden while preserve more spatial details, along with a Wavelet Smooth Loss (WSL) to reconstruct original structured context and texture with a smooth constrain. Experiments on several UHR datasets demonstrate its state-of-the-art performance. The dataset is available at https://github.com/jankyee/URUR.
Deyi Ji, Feng Zhao 0004, Hongtao Lu 0001, Mingyuan Tao, Jieping Ye
CVPR4
2023 Optimal Parameter and Neuron Pruning for Out-of-Distribution Detection
abstract
For a machine learning model deployed in real world scenarios, the ability of detecting out-of-distribution (OOD) samples is indispensable and challenging. Most existing OOD detection methods focused on exploring advanced training skills or training-free tricks to prevent the model from yielding overconfident confidence score for unknown samples. The training-based methods require expensive training cost and rely on OOD samples which are not always available, while most training-free methods can not efficiently utilize the prior information from the training data. In this work, we propose an \textbf{O}ptimal \textbf{P}arameter and \textbf{N}euron \textbf{P}runing (\textbf{OPNP}) approach, which aims to identify and remove those parameters and neurons that lead to over-fitting. The main method is divided into two steps. In the first step, we evaluate the sensitivity of the model parameters and neurons by averaging gradients over all training samples. In the second step, the parameters and neurons with exceptionally large or close to zero sensitivities are removed for prediction. Our proposal is training-free, compatible with other post-hoc methods, and exploring the information from all training data. Extensive experiments are performed on multiple OOD detection tasks and model architectures, showing that our proposed OPNP consistently outperforms the existing methods by a large margin.
Chao Chen 0026, Zhihang Fu, Kai Liu 0023, Ze Chen 0001, Mingyuan Tao, Jieping Ye
NeurIPS5
2023 Category-Extensible Out-of-Distribution Detection via Hierarchical Context Descriptions
abstract
The key to OOD detection has two aspects: generalized feature representation and precise category description. Recently, vision-language models such as CLIP provide significant advances in both two issues, but constructing precise category descriptions is still in its infancy due to the absence of unseen categories. This work introduces two hierarchical contexts, namely perceptual context and spurious context, to carefully describe the precise category boundary through automatic prompt tuning. Specifically, perceptual contexts perceive the inter-category difference (e.g., cats vs apples) for current classification tasks, while spurious contexts further identify spurious (similar but exactly not) OOD samples for every single category (e.g., cats vs panthers, apples vs peaches). The two contexts hierarchically construct the precise description for a certain category, which is, first roughly classifying a sample to the predicted category and then delicately identifying whether it is truly an ID sample or actually OOD. Moreover, the precise descriptions for those categories within the vision-language framework present a novel application: CATegory-EXtensible OOD detection (CATEX). One can efficiently extend the set of recognizable categories by simply merging the hierarchical contexts learned under different sub-task settings. And extensive experiments are conducted to demonstrate CATEX’s effectiveness, robustness, and category-extensibility. For instance, CATEX consistently surpasses the rivals by a large margin with several protocols on the challenging ImageNet-1K dataset. In addition, we offer new insights on how to efficiently scale up the prompt engineering in vision-language models to recognize thousands of object categories, as well as how to incorporate large language models (like GPT-3) to boost zero-shot applications.
Kai Liu 0023, Zhihang Fu, Chao Chen 0026, Sheng Jin 0002, Ze Chen 0001, Mingyuan Tao, Rongxin Jiang 0001, Jieping Ye
NeurIPS6
2022 Structural and Statistical Texture Knowledge Distillation for Semantic Segmentation
abstract
Existing knowledge distillation works for semantic seg-mentation mainly focus on transfering high-level contextual knowledge from teacher to student. However, low-level texture knowledge is also of vital importance for characterizing the local structural pattern and global statistical prop-erty, such as boundary, smoothness, regularity and color contrast, which may not be well addressed by high-level deep features. In this paper, we are intended to take full advantage of both structural and statistical texture knowledge and propose a novel Structural and Statistical Texture Knowledge Distillation (SSTKD) framework for Semantic Segmentation. Specifically, for structural texture knowledge, we introduce a Contourlet Decomposition Module (CDM) that decomposes low-level features with iterative laplacian pyramid and directional filter bank to mine the structural texture knowledge. For statistical knowledge, we propose a Denoised Texture Intensity Equalization Module (DTIEM) to adaptively extract and enhance statistical texture knowledge through heuristics iterative quantization and denoised operation. Finally, each knowledge learning is supervised by an individual loss function, forcing the student network to mimic the teacher better from a broader perspective. Experiments show that the proposed method achieves state-of-the-art performance on Cityscapes, Pascal VOC 2012 and ADE20K datasets.
Deyi Ji, Mingyuan Tao, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Hongtao Lu 0001
CVPR3
2022 Invariant Feature Learning for Generalized Long-Tailed Classification
Kaihua Tang, Mingyuan Tao, Jiaxin Qi, Zhenguang Liu, Hanwang Zhang
ECCV (24)2
2022 G2NET: A General Geography-Aware Representation Network for Hotel Search Ranking
abstract
Hotel search ranking is the core function of Online Travel Platforms (OTPs), while geography information of location entities involved in it plays a critically important role in guaranteeing its ranking quality. The closest line of works to the hotel search ranking problem is thus the next POI (or location) recommendation problem, which has extensive works but fails to cope with two new challenges, i.e., consideration of two more location entities and effective utilization of geographical information, in a hotel search ranking scenario. To this end, we propose a General Geography-aware representation NETwork (G2NET for short) to better represent geography information of location entities so as to optimize the hotel search ranking. In G2NET, to address the first challenge, we first propose the concept of Geography Interaction Schema (GIS) which is a meta template for representing the arbitrary number of location entity types and their interactions. Then, a novel geography interaction encoder is devised providing general representation ability for an instance of GIS, followed by an attentive operation that aggregates representations of instances corresponding to all historically interacted hotels of a user in a weighted manner. The second challenge is handled by the combined application of three proposed geography embedding modules in G2NET, each of which focuses on computing embeddings of location entities based on a certain aspect of geographical information of location entities. Moreover, a self-attention layer is deployed in G2NET, to capture correlations among historically interacted hotels of a user which provides non-trivial functionality of understanding the user's behaviors. Both offline and online experiments show that G2NET outperforms the state-of-the-art methods. G2NET has now been successfully deployed to provide the high-quality hotel search ranking service at Fliggy, one of the most popular OTPs in China, serving tens of millions of users.
Jia Xu 0005, Zulong Chen, Mingyuan Tao, Liangyue Li
KDD4
2022 Dynamic supervisor for cross-dataset object detection
Ze Chen 0001, Zhihang Fu, Jianqiang Huang 0001, Mingyuan Tao, Rongxin Jiang 0001, Xiang Tian 0002, Yaowu Chen, Xian-Sheng Hua 0001
Neurocomputing4
2021 Deep Inclusion Relation-aware Network for User Response Prediction at Fliggy
abstract
User response prediction plays a crucial role in many applications (e.g. search ranking and personalized recommendation) at online travel platforms. Although existing methods have made a great success by focusing on feature interaction or user behaviors, they cannot synthetically exploit item inclusion relations describing relationships of an item including or being included by another one, which are important components among travel items. To this end, in this paper, we propose a novel Deep Inclusion Relation-aware Network (DIRN) for user response prediction by synthetically exploiting inclusion relations among travel items. Specifically, on the item graph constructed with inclusion relations, we first leverage a node embedding approach to learn the item graph-based embedding. Then, we design Representation-based Interest Layer and Relation Path Interest Layer to extract user latent interest with user behaviors in two ways. Representation-based Interest Layer models the item-to-item similarity based on item representations containing the graph-based embedding with an attention mechanism and obtains user temporal interest by summing up representations of interacted items with similarities. Relation Path Interest Layer measures item-to-item realistic associations to extract user interest with inclusion relation paths. Offline experiments on a real-world data from Fliggy clearly validate the effectiveness of DIRN. Furthermore, DIRN has been successfully deployed online in search ranking at Fliggy and achieves significant improvement.
Zai Huang, Mingyuan Tao, Bufeng Zhang
KDD2
2021 Deep User Match Network for Click-Through Rate Prediction
abstract
Click-through rate (CTR) prediction is a crucial task in many applications (e.g. recommender systems). Recently deep learning based models have been proposed and successfully applied for CTR prediction by focusing on feature interaction or user interest based on the item-to-item relevance between user behaviors and candidate item. However, these existing models neglect the user-to-user relevance between the target user and those who like the candidate item, which can reflect the preference of target user. To this end, in this paper, we propose a novel Deep User Match Network (DUMN) which measures the user-to-user relevance for CTR prediction. Specifically, in DUMN, we design a User Representation Layer to learn a unified user representation which contains user latent interest based on user behaviors. Then, User Match Layer is designed to measure the user-to-user relevance by matching the target user and those who have interacted with candidate item and modeling their similarities in user representation space. Extensive experimental results on three public real-world datasets validate the effectiveness of DUMN compared with state-of-the-art methods.
Zai Huang, Mingyuan Tao, Bufeng Zhang
SIGIR2
2021 Spatial likelihood voting with self-knowledge distillation for weakly supervised object detection
Ze Chen 0001, Zhihang Fu, Jianqiang Huang 0001, Mingyuan Tao, Rongxin Jiang 0001, Xiang Tian 0002, Yaowu Chen, Xian-Sheng Hua 0001
Image Vis. Comput.4