EDBT 2026 Demo / reviewers in the wild / expert
Qi Yu 0001
dblp:58/6957-1
· DBLP profile ↗
16ranked-venue papers in the field
5as first author
5since 2021 · last 2024
0000-0002-0426-5407ORCID · conflict
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 9 (1 first)Database Systems & Data Management · 3 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 2 (1 first)Information Retrieval & Web Search · 1 (1 first)Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Reinforced Compressive Neural Architecture Search for Versatile Adversarial RobustnessabstractPrior research on neural architecture search (NAS) for adversarial robustness has revealed that a lightweight and adversarially robust sub-network could exist in a non-robust large teacher network. Such a sub-network is generally discovered based on heuristic rules to perform neural architecture search. However, heuristic rules are inadequate to handle diverse adversarial attacks and different "teacher" network capacity. To address this key challenge, we propose Reinforced Compressive Neural Architecture Search (RC-NAS), aiming to achieve Versatile Adversarial Robustness. Specifically, we define novel task settings that compose datasets, adversarial attacks, and teacher network configuration. Given diverse tasks, we develop an innovative dual-level training paradigm that consists of a meta-training and a fine-tuning phase to effectively expose the RL agent to diverse attack scenarios (in meta-training), and make it adapt quickly to locate an optimal sub-network (in fine-tuning) for previously unseen scenarios. Experiments show that our framework could achieve adaptive compression towards different initial teacher networks, datasets, and adversarial attacks, resulting in more lightweight and adversarially robust architectures. We also provide a theoretical analysis to explain why the reinforcement learning (RL)-guided adversarial architectural search helps adversarial robustness over standard adversarial training methods. Dingrong Wang, Hitesh Sapkota, Zhiqiang Tao, Qi Yu 0001 |
KDD | 4 |
| 2022 | Balancing Bias and Variance for Active Weakly Supervised LearningabstractAs a widely used weakly supervised learning scheme, modern multiple instance learning (MIL) models achieve competitive performance at the bag level. However, instance-level prediction, which is essential for many important applications, remains largely unsatisfactory. We propose to conduct novel active deep multiple instance learning that samples a small subset of informative instances for annotation, aiming to significantly boost the instance-level prediction. A variance regularized loss function is designed to properly balance the bias and variance of instance-level predictions, aiming to effectively accommodate the highly imbalanced instance distribution in MIL and other fundamental challenges. Instead of directly minimizing the variance regularized loss that is non-convex, we optimize a distributionally robust bag level likelihood as its convex surrogate. The robust bag likelihood provides a good approximation of the variance based MIL loss with a strong theoretical guarantee. It also automatically balances bias and variance, making it effective to identify the potentially positive instances to support active sampling. The robust bag likelihood can be naturally integrated with a deep architecture to support deep model training using mini-batches of positive-negative bag pairs. Finally, a novel P-F sampling function is developed that combines a probability vector and predicted instance scores, obtained by optimizing the robust bag likelihood. By leveraging the key MIL assumption, the sampling function can explore the most challenging bags and effectively detect their positive instances for annotation, which significantly improves the instance-level prediction. Experiments conducted over multiple real-world datasets clearly demonstrate the state-of-the-art instance-level prediction achieved by the proposed model. Hitesh Sapkota, Qi Yu 0001 |
KDD | 2 |
| 2021 | Uncertainty-Aware Multiple Instance Learning from Large-Scale Long Time Series DataabstractWe propose a novel framework to classify large-scale time series data with long duration. Long time series classification (L-TSC) is a challenging problem because the data often contains a large amount of irrelevant information to the classification target. The irrelevant period degrades the classification performance while the relevance is unknown to the system. This paper proposes an uncertainty-aware multiple instance learning (MIL) framework to identify the most relevant period automatically. The predictive uncertainty enables designing an attention mechanism that forces the MIL model to learn from the possibly discriminant period. Moreover, the predicted uncertainty yields a principled estimator to identify whether a prediction is trustworthy or not. We further incorporate another modality to accommodate unreliable predictions by training a separate model based on its availability and conduct uncertainty aware fusion to produce the final prediction. Systematic evaluation is conducted on the Automatic Identification System (AIS) data, which is collected to identify and track real-world vessels. Empirical results demonstrate that the proposed method can effectively detect the types of vessels based on the trajectory and the uncertainty-aware fusion with other available data modality (Synthetic-Aperture Radar or SAR imagery is used in our experiments) can further improve the detection accuracy. Yuansheng Zhu, Weishi Shi, Deep Shankar Pandey, Xiaofan Que, Daniel E. Krutz, Qi Yu 0001 |
IEEE BigData | 7 |
| 2021 | MetaEDL: Meta Evidential Learning For Uncertainty-Aware Cold-Start RecommendationsabstractRecommender systems have been widely used to predict users’ interests and filter information from a large number of candidate items. However, accurately capturing the interests of users having limited interactions with a system remains a long-lasting challenge. Furthermore, existing recommender systems primarily focus on predicting user preferences without quantifying the prediction uncertainty. Uncertainty can help to quantify the model confidence when making a recommendation where low model confidence could serve as a more accurate indicator of a user’s cold-start level than simply using the number of interactions. We present a novel recommendation model that seamlessly integrates a meta-learning module with an evidential learning approach. The former module generalizes meta knowledge to tackle cold-start recommendations by exploiting fast adaptation. The latter quantifies both aleatoric and epistemic uncertainty without performing expensive posterior inference. Evidential learning achieves this by placing evidential priors and treating the output of the meta-learning module as evidence-based pseudo counts and learns a function to directly predict the evidence of a target interaction. Experiments on four benchmark datasets justify that our proposed model captures the uncertainty of users and demonstrates its superior performance over the state-of-the-art recommendation models. Krishna Prasad Neupane, Ervine Zheng, Qi Yu 0001 |
ICDM | 3 |
| 2021 | Deep Reinforced Attention Regression for Partial Sketch Based Image RetrievalabstractFine-Grained Sketch-Based Image Retrieval (FG-SBIR) aims at finding a specific image from a large gallery given a query sketch. Despite the widespread applicability of FG-SBIR in many critical domains (e.g., crime activity tracking), existing approaches still suffer from a low accuracy while being sensitive to external noises such as unnecessary strokes in the sketch. The retrieval performance will further deteriorate under a more practical on-the-fly setting, where only a partially complete sketch with only a few (noisy) strokes are available to retrieve corresponding images. We propose a novel framework that leverages a uniquely designed deep reinforcement learning model that performs a dual-level exploration to deal with partial sketch training and attention region selection. By enforcing the model’s attention on the important regions of the original sketches, it remains robust to unnecessary stroke noises and improve the retrieval accuracy by a large margin. To sufficiently explore partial sketches and locate the important regions to attend, the model performs bootstrapped policy gradient for global exploration while adjusting a standard deviation term that governs a locator network for local exploration. The training process is guided by a hybrid loss that integrates a reinforcement loss and a supervised loss. A dynamic ranking reward is developed to fit the on-the-fly image retrieval process using partial sketches. The extensive experimentation performed on three public datasets shows that our proposed approach achieves the state-of-the-art performance on partial sketch based image retrieval. Dingrong Wang, Hitesh Sapkota, Xumin Liu, Qi Yu 0001 |
ICDM | 4 |
| 2020 | Integrating reinforcement learning and skyline computing for adaptive service composition
Xingguo Hu, Qi Yu 0001, Mingzhu Gu, Tianjing Hong |
Inf. Sci. | 3 |
| 2018 | An Efficient Many-Class Active Learning Framework for Knowledge-Rich DomainsabstractThe high cost for labeling data instances is a key bottleneck for training effective supervised learning models. This is especially the case in domains such as medicine and bioinformatics, where expert knowledge is required for understanding and extracting the underlying semantics of data. Active learning provides a means to reduce human labeling efforts by identifying the most informative data instances. In this paper, we propose a cost-effective active learning framework to further lessen human efforts, especially in knowledge-rich domains where a large number of classes may be subject to scrutiny during decision making. In particular, this framework employs a novel many-class sampling model, MC-S, for data sample selection. MC-S is further augmented with convex hull-based sampling to achieve faster convergence of active learning. Evaluation studies conducted over multiple real-world datasets with many classes demonstrate that the proposed framework significantly reduces the overall labeling efforts through fast convergence and early stop of active learning. Weishi Shi, Qi Yu 0001 |
ICDM | 2 |
| 2018 | Log sequence clustering for workflow mining in multi-workflow systems
Xumin Liu, Moayad Alshangiti, Chen Ding 0004, Qi Yu 0001 |
Data Knowl. Eng. | 4 |
| 2016 | An Expert-in-the-loop Paradigm for Learning Medical Image Grouping
Qi Yu 0001, Rui Li 0002, Cecilia O. Alm, Cara Calvelli, Anne R. Haake |
PAKDD (1) | 2 |
| 2016 | Incorporating Heterogeneous Information for Mashup Discovery with Consistent Regularization
Yao Wan 0001, Liang Chen 0001, Qi Yu 0001, Tingting Liang, Jian Wu 0001 |
PAKDD (1) | 3 |
| 2015 | CloudRec: a framework for personalized service Recommendation in the Cloud
Qi Yu 0001 |
Knowl. Inf. Syst. | 1 |
| 2013 | Maximizing influence of viral marketing via evolutionary user selectionabstractViral marketing, which uses the "word of mouth" marketing technique over virtual networks, relies on the selection of a small subset of most influential users in the network for efficient marketing. Nonetheless, most existing viral marketing techniques ignore the dynamic nature of the virtual network. In this paper, we develop a novel framework that exploits the temporal dynamics of the network to select an optimal subset of users that maximize the marketing influence over the network. Sanket Anil Naik, Qi Yu 0001 |
ASONAM | 2 |
| 2013 | Efficient Large-Scale Service Clustering via Sparse Functional Representation and Accelerated OptimizationabstractClustering techniques offer a systematic approach to organize the diverse and fast increasing Web services by assigning relevant services into homogeneous service communities. However, the ever increasing number of Web services poses key challenges for building large-scale service communities. In this paper, we tackle the scalability issue in service clustering, aiming to accurately and efficiently discover service communities over very large-scale services. A key observation is that service descriptions are usually represented by long but very sparse term vectors as each service is only described by a limited number of terms. This inspires us to seek a new service representation that is economical to store, efficient to process, and intuitive to interpret. This new representation enables service clustering to scale to massive number of services. More specifically, a set of anchor services are identified that allows each service to represent as a linear combination of a small number of anchor services. In this way, the large number of services are encoded with a much more compact anchor service space. Despite service clustering can be performed much more efficiently in the compact anchor service space, discovery of anchor services from large-scale service descriptions may incur high computational cost. We develop principled optimization strategies for efficient anchor service discovery. Extensive experiments are conducted on real-world service data to assess both the effectiveness and efficiency of the proposed approach. Results on a dataset with over 3,700 Web services clearly demonstrate the good scalability of sparse functional representation and the efficiency of the optimization algorithms for anchor service discovery. Qi Yu 0001 |
Int. J. Cooperative Inf. Syst. | 1 |
| 2013 | Efficient Service Skyline Computation for Composite Service SelectionabstractService composition is emerging as an effective vehicle for integrating existing web services to create value-added and personalized composite services. As web services with similar functionality are expected to be provided by competing providers, a key challenge is to find the “best” web services to participate in the composition. When multiple quality aspects (e.g., response time, fee, etc.) are considered, a weighting mechanism is usually adopted by most existing approaches, which requires users to specify their preferences as numeric values. We propose to exploit the dominance relationship among service providers to find a set of “best” possible composite services, referred to as a composite service skyline. We develop efficient algorithms that allow us to find the composite service skyline from a significantly reduced searching space instead of considering all possible service compositions. We propose a novel bottom-up computation framework that enables the skyline algorithm to scale well with the number of services in a composition. We conduct a comprehensive analytical and experimental study to evaluate the effectiveness, efficiency, and scalability of the composite skyline computation approaches. Qi Yu 0001, Athman Bouguettaya |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2008 | Framework for Web service query algebra and optimizationabstractWe present a query algebra that supports optimized access of Web services through service-oriented queries. The service query algebra is defined based on a formal service model that provides a high-level abstraction of Web services across an application domain. The algebra defines a set of algebraic operators. Algebraic service queries can be formulated using these operators. This allows users to query their desired services based on both functionality and quality. We provide the implementation of each algebraic operator. This enables the generation of Service Execution Plans (SEPs) that can be used by users to directly access services. We present an optimization algorithm by extending the Dynamic Programming (DP) approach to efficiently select the SEPs with the best user-desired quality. The experimental study validates the proposed algorithm by demonstrating significant performance improvement compared with the traditional DP approach. Qi Yu 0001, Athman Bouguettaya |
ACM Trans. Web | 1 |
| 2008 | Deploying and managing Web services: issues, solutions, and directions
Qi Yu 0001, Xumin Liu, Athman Bouguettaya, Brahim Medjahed |
VLDB J. | 1 |