EDBT 2026 Demo / reviewers in the wild / expert
Qing Xie 0002
dblp:98/2931-2
· DBLP profile ↗
41ranked-venue papers in the field
5as first author
24since 2021 · last 2026
0000-0003-4530-588XORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 17 (3 first)Database Systems & Data Management · 10 (2 first)Data Mining & Knowledge Discovery · 6Knowledge Engineering, Semantic Web & Information Systems · 4Other / Interdisciplinary · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cluster-Guided Disentangled Representation for Cold-Start Cross-Domain Recommendation
Huping Yu, Yuhan Wang 0004, Qing Xie 0002, Mengzi Tang, Lin Li 0001, Yongjian Liu |
DASFAA (1) | 3 |
| 2026 | A Reproducibility Study of Bundle Editing and Bundle RecommendationabstractBundle recommender system is divided into two main stages: bundle editing and bundle recommendation. While substantial research progress has been made in each stage, in practical application scenarios, bundle compositions and the final recommended bundles mutually influence each other: the continuously editing bundle compositions affect the recommendation results, while user feedback on recommended bundles in turn guides the refinement of bundle compositions. This paper presents the first comprehensive reproducibility study of the complete bundle recommendation pipeline. We implement eight bundle-level editing methods, nine item-level editing methods, and seven state-of-the-art bundle recommendation models, and evaluate their performance across six real-world datasets. Our empirical analysis reveals several key findings. First, bundle-level editing faces the challenge of generating high-quality bundles. Second, in the item-level editing, the replacement operation emerges as a universal bottleneck across all methods. Third, in the recommendation stage, recommendation models exhibit varying performance across different interaction density scenarios (e.g., cold-start). Finally, bundle recommendation suffers degraded performance when integrating item-level editing and bundle recommendation within a unified pipeline. Overall, there is the systemic limitation of bundle recommendation: prior work has focused on optimizing individual stages independently, disregarding the interdependencies throughout the entire recommendation system. These findings highlight the urgent need to develop end-to-end solutions that can holistically address the bundle editing and recommendation workflow. Our repository is now available for public access via https://github.com/anyr123/Bundle_Edit_Rec_SIGIR26. Yiran An, Lin Li 0001, Ming Li 0072, Wenxin Ye, Qing Xie 0002, Jimmy Huang 0001 |
SIGIR | 5 |
| 2026 | A Sketch+Text Composed Image Retrieval Dataset for ThangkaabstractComposed Image Retrieval (CIR) enables image retrieval by combining multiple query modalities, but existing benchmarks predominantly focus on general-domain imagery and rely on reference images with short textual modifications. As a result, they provide limited support for retrieval scenarios that require fine-grained semantic reasoning, structured visual understanding, and domain-specific knowledge. In this work, we introduce CIRThan, a sketch+text composed image retrieval dataset for Thangka imagery, a culturally grounded and knowledge-specific visual domain characterized by complex structures, dense symbolic elements, and domain-dependent semantic conventions. CIRThan contains 2,287 high-quality Thangka images, each paired with a human-drawn sketch and hierarchical textual descriptions at three semantic levels, enabling composed queries that jointly express structural intent and multi-level semantic specification. We provide standardized data splits, comprehensive dataset analysis, and benchmark evaluations of representative supervised and zero-shot CIR methods. Experimental results reveal that existing CIR approaches, largely developed for general-domain imagery, struggle to effectively align sketch-based abstractions and hierarchical textual semantics with fine-grained Thangka images, particularly without in-domain supervision. We believe CIRThan offers a valuable benchmark for advancing sketch+text CIR, hierarchical semantic modeling, and multimodal retrieval in cultural heritage and other knowledge-specific visual domains. The dataset is publicly available at https://github.com/jinyuxu-whut/CIRThan. Jinyu Xu 0001, Jiangling Zhang, Qing Xie 0002, Daomin Ji, Zhifeng Bao, Jiachen Li 0002, Yanchun Ma, Yongjian Liu |
SIGIR | 4 |
| 2026 | SDR-CIR: Semantic Debias Retrieval Framework for Training-Free Zero-Shot Composed Image RetrievalabstractComposed Image Retrieval (CIR) aims to retrieve a target image from a query composed of a reference image and modification text. Recent training-free zero-shot methods often employ Multimodal Large Language Models (MLLMs) with Chain-of-Thought (CoT) to compose a target image description for retrieval. However, due to the fuzzy matching nature of ZS-CIR, the generated description is prone to semantic bias relative to the target image. We propose SDR-CIR, a training-free Semantic Debias Ranking method based on CoT reasoning. First, Selective CoT guides the MLLM to extract visual content relevant to the modification text during image understanding, thereby reducing visual noise at the source. We then introduce a Semantic Debias Ranking with two steps, Anchor and Debias, to mitigate semantic bias. In the Anchor step, we fuse reference image features with target description features to reinforce useful semantics and supplement omitted cues. In the Debias step, we explicitly model the visual semantic contribution of the reference image to the description and incorporate it into the similarity score as a penalty term. By supplementing omitted cues while suppressing redundancy, SDR-CIR mitigates semantic bias and improves retrieval performance. Experiments on three standard CIR benchmarks show that SDR-CIR achieves state-of-the-art results among one-stage methods while maintaining high efficiency. The code is publicly available at https://github.com/suny105/SDR-CIR. Jinyu Xu 0001, Qing Xie 0002, Jiachen Li 0002, Yanchun Ma, Yongjian Liu |
WWW | 3 |
| 2026 | Learning resource recommendation models based on learning behaviors and hierarchical structure graph
Lihua Bai, Qing Xie 0002, Mengzi Tang |
J. Intell. Inf. Syst. | 4 |
| 2025 | TOVect: Topology-Optimized Vectorization for Intangible Cultural Heritage Thangka Element Line ArtabstractThangka art, part of the UNESCO Intangible Cultural Heritage of Humanity, is visually characterized by complex junctions and intricate corners, demand high-fidelity vectorization to preserve its structural integrity and smooth curvilinear aesthetics. Conventional line art vectorization algorithms applied to Thangka element line art face challenges: (1) hard to fit complex junctions that leads to spurious spikes and discontinuous strokes; and (2) unnatural distortions in long curves due to insufficient smoothness constraints. To address these challenges, we propose a skeleton-guided vectorization framework to optimize the topology of vectorized Thangka element line art, and a multilayer perceptual loss as a smoothness regulation to improve curve continuity. Experimental results on manually annotated Thangka element line art dataset demonstrate that our method surpasses state-of-the-art approaches in preserving topological integrity and achieving visual smoothness, offering a robust foundation for digitizing cultural heritage artworks with complex topologies and similar aesthetic requirements. Anshu Hu, Yifei Sun 0018, Jiachen Li 0002, Yanchun Ma, Qing Xie 0002, Yongjian Liu |
MMAsia | 5 |
| 2025 | Robust Dual Embedding Contrastive Learning for Text-to-Image Person Re-identification with Noisy CorrespondenceabstractText-to-Image person re-identification (TIReID) aims to retrieve pedestrian images from a gallery based on textual descriptions, thus bridging vision and language modalities for practical retrieval scenarios. Despite recent advances leveraging various cross-modal alignment strategies, existing methods typically assume all image-text pairs in training datasets are correctly matched, overlooking the pervasive Noisy Correspondence (NC) problem—erroneous image-text associations that degrade model robustness. Prior approaches either lack noise identification mechanisms or rely on direct filtering of detected noisy samples, which only partially mitigates the adverse effects of noise and cannot fully prevent overfitting to incorrect correspondences during training. Addressing this challenge, we propose Robust Dual Embedding Contrastive Learning (RDECL), which consists of two main components: 1) A Dual-View Cumulative Trust Division (DCTD) progressively constructs a high-confidence clean sample repository via adaptive sample selection, ensuring reliable image-text correspondence learning under uncertain noise detection.2) A Robust Generalized Contrastive Loss (RGCL) further enhances robustness by leveraging all negative samples and maximizing the loss distribution discrepancy between clean and noisy samples, thereby suppressing overfitting to noisy labels. We conduct extensive experiments on three public benchmark datasets, namely CUHK-PEDES, ICFG-PEDES, and RSTPReID, to evaluate the performance and robustness of our RDECL. Jingjie Zhang, Lingli Tang, Jiachen Li 0002, Jinyu Xu 0001, Yanchun Ma, Qing Xie 0002 |
MMAsia | 6 |
| 2025 | Enhancing Transferability and Consistency in Cross-Domain Recommendations via Supervised DisentanglementabstractCross-domain recommendation (CDR) aims to alleviate the data sparsity by transferring knowledge across domains.Disentangled representation learning provides an effective solution to model complex user preferences by separating intra-domain features (domainshared and domain-specific features), thereby enhancing robustness and interpretability.However, disentanglement-based CDR methods employing generative modeling or GNNs with contrastive objectives face two key challenges: (i) pre-separation strategies decouple features before extracting collaborative signals, disrupting intra-domain interactions and introducing noise; (ii) unsupervised disentanglement objectives lack explicit task-specific guidance, resulting in limited consistency and suboptimal alignment.To address these challenges, we propose DGCDR, a GNN-enhanced encoder-decoder framework.To handle challenge (i), DGCDR first applies GNN to extract high-order collaborative signals, providing enriched representations as a robust foundation for disentanglement.The encoder then dynamically disentangles features into domain-shared and -specific spaces, preserving collaborative information during the separation process.To handle challenge (ii), the Yuhan Wang 0004, Qing Xie 0002, Zhifeng Bao, Mengzi Tang, Lin Li 0001, Yongjian Liu |
RecSys | 2 |
| 2025 | Bibliometric analysis and review of AI-based video generation: research dynamics and application trends (2020-2025)abstractAI-based video generation is rapidly advancing field with significant research focus and broad applications across industries like entertainment, education, and healthcare. To gain a structured understanding of the field, this paper systematically reviews the research dynamics and application trends in video generation from 2020 to April 2025. Utilizing a combined approach of content analysis and bibliometric analysis on 422 research publications selected from the Web of Science Core Collection database, we investigate key developments. The content analysis examines state-of-the-art algorithmic innovations, generation strategies, and prominent application domains. Bibliometric analysis, employing tools like CiteSpace and VOSviewer, maps publication trends, identifies influential authors, institutions, and sources, visualizes collaboration networks, and analyzes the evolution of research hotspots through keywords. Our findings the rise of specific architectures, pinpoint high-activity application areas, reveal evolving thematic clusters, and also remarks persistent challenges, including technical hurdles and critical ethical issues such as deepfake detection, algorithmic transparency, and data privacy. By synthesizing these analyses, this study offers a structured overview of the recent landscape, and highlights areas that may warrant further exploration in AI video generation. Anshu Hu, Qing Xie 0002, Ruoyu Wan, Yuhan Liu 0001 |
Discov. Comput. | 3 |
| 2025 | Erratum: A Dual Perspective Framework of Knowledge-correlation for Cross-domain RecommendationabstractThis is an erratum for the article "A Dual Perspective Framework of Knowledge-correlation for Cross-domain Recommendation" published in ACM Trans. Knowl. Discov. Data 18(6): 152:1-152:28 (2024). Yuhan Wang 0004, Qing Xie 0002, Mengzi Tang, Lin Li 0001, Jingling Yuan, Yongjian Liu |
ACM Trans. Knowl. Discov. Data | 2 |
| 2024 | CrimeAlarm: Towards Intensive Intent Dynamics in Fine-Grained Crime Prediction
Kaixi Hu, Lin Li 0001, Qing Xie 0002, Xiaohui Tao 0001, Guandong Xu |
DASFAA (7) | 3 |
| 2024 | Dlpp-Net: Degradation Location Prior Prediction Network for Image Restoration
Yongjian Liu, Shunwei Zhang, Jinyu Xu 0001, Jiachen Li 0002, Yanchun Ma, Qing Xie 0002 |
MMAsia | 6 |
| 2024 | Amazon-KG: A Knowledge Graph Enhanced Cross-Domain Recommendation DatasetabstractCross-domain recommendation (CDR) aims to utilize the information from relevant domains to guide the recommendation task in the target domain, and shows great potential in alleviating the data sparsity and cold-start problems of recommender systems. Most existing methods utilize the interaction information (e.g., ratings and clicks) or consider auxiliary information (e.g., tags and comments) to analyze the users' cross-domain preferences, but such kinds of information ignore the intrinsic semantic relationship of different domains. In order to effectively explore the inter-domain correlations, encyclopedic knowledge graphs (KG) involving different domains are highly desired in cross-domain recommendation tasks because they contain general information covering various domains with structured data format. However, there are few datasets containing KG information for CDR tasks, so in order to enrich the available data resource, we build a KG-enhanced cross-domain recommendation dataset, named Amazon-KG, based on the widely used Amazon dataset for CDR and the well-known KG DBpedia. In this work, we analyze the potential of KG applying in cross-domain recommendations, and describe the construction process of our dataset in detail. Finally, we perform quantitative statistical analysis on the dataset. We believe that datasets like Amazon-KG contribute to the development of knowledge-aware cross-domain recommender systems. Our dataset has been released at https://github.com/WangYuhan-0520/Amazon-KG-v2.0-dataset. Yuhan Wang 0004, Qing Xie 0002, Mengzi Tang, Lin Li 0001, Jingling Yuan, Yongjian Liu |
SIGIR | 2 |
| 2024 | Boosting Healthiness Exposure in Category-Constrained Meal Recommendation Using Nutritional StandardsabstractFood computing, a newly emerging topic, is closely linked to human life through computational methodologies. Meal recommendation, a food-related study about human health, aims to provide users a meal with courses constrained from specific categories (e.g., appetizers, main dishes) that can be enjoyed as a service. Historical interaction data, important user information, is often used by existing models to learn user preferences. However, if a user’s preferences favor less healthy meals, the model will follow that preference and make similar recommendations, potentially negatively impacting the user’s long-term health. This emphasizes the necessity for health-oriented and responsible meal recommendation systems. In this article, we propose a healthiness-aware and category-wise meal recommendation model called CateRec, which boosts healthiness exposure by using nutritional standards as knowledge to guide the model training. Two fundamental questions are raised and answered: (1) How can the healthiness of meals be evaluated? Two well-known nutritional standards from the World Health Organization and the United Kingdom Food Standards Agency are used to calculate the healthiness score of the meal. (2) How can the model training be guided in a health-oriented manner? We construct category-wise personalization partial rankings and category-wise healthiness partial rankings, and theoretically analyze that they meet the necessary properties and assumptions required to be trained by the maximum posterior estimator under Bayesian probability. The data analysis confirms the existence of user preferences leaning towards less healthy meals in two public datasets. A comprehensive experiment demonstrates that our CateRec effectively boosts healthiness exposure in terms of mean healthiness score and ranking exposure while being comparable to the state-of-the-art model in terms of recommendation accuracy. Ming Li 0072, Lin Li 0001, Xiaohui Tao 0001, Zhongwei Xie, Qing Xie 0002, Jingling Yuan |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2024 | A Dual Perspective Framework of Knowledge-correlation for Cross-domain RecommendationabstractRecommender System provides users with online services in a personalized way. The performance of traditional recommender systems may deteriorate because of problems such as cold-start and data sparsity. Cross-domain Recommendation System utilizes the richer information from auxiliary domains to guide the task in the target domain. However, direct knowledge transfer may lead to a negative impact due to data heterogeneity and feature mismatch between domains. In this article, we innovatively explore the cross-domain correlation from the perspectives of content semanticity and structural connectivity to fully exploit the information of Knowledge Graph. First, we adopt domain adaptation that automatically extracts transferable features to capture cross-domain semantic relations. Second, we devise a knowledge-aware graph neural network to explicitly model the high-order connectivity across domains. Third, we develop feature fusion strategies to combine the advantages of semantic and structural information. By simulating the cold-start scenario on two real-world datasets, the experimental results show that our proposed method has superior performance in accuracy and diversity compared with the SOTA methods. It demonstrates that our method can accurately predict users’ expressed preferences while exploring their potential diverse interests. Yuhan Wang 0004, Qing Xie 0002, Mengzi Tang, Lin Li 0001, Jingling Yuan, Yongjian Liu |
ACM Trans. Knowl. Discov. Data | 2 |
| 2024 | Decoupled Progressive Distillation for Sequential Prediction with Interaction DynamicsabstractSequential prediction has great value for resource allocation due to its capability in analyzing intents for next prediction. A fundamental challenge arises from real-world interaction dynamics where similar sequences involving multiple intents may exhibit different next items. More importantly, the character of volume candidate items in sequential prediction may amplify such dynamics, making deep networks hard to capture comprehensive intents. This article presents a sequential prediction framework with Decoupled Progressive Distillation (DePoD), drawing on the progressive nature of human cognition. We redefine target and non-target item distillation according to their different effects in the decoupled formulation. This can be achieved through two aspects: (1) Regarding how to learn, our target item distillation with progressive difficulty increases the contribution of low-confidence samples in the later training phase while keeping high-confidence samples in the earlier phase. And, the non-target item distillation starts from a small subset of non-target items from which size increases according to the item frequency. (2) Regarding whom to learn from, a difference evaluator is utilized to progressively select an expert that provides informative knowledge among items from the cohort of peers. Extensive experiments on four public datasets show DePoD outperforms state-of-the-art methods in terms of accuracy-based metrics. Kaixi Hu, Lin Li 0001, Qing Xie 0002, Jianquan Liu, Xiaohui Tao 0001, Guandong Xu |
ACM Trans. Inf. Syst. | 3 |
| 2023 | L2QA: Long Legal Article Question Answering with Cascaded Key Segment Learning
Shugui Xie, Lin Li 0001, Jingling Yuan, Qing Xie 0002, Xiaohui Tao 0001 |
DASFAA (3) | 4 |
| 2023 | A Multi-scale and Dense Object Detector for Tibetan Thangka ImagesabstractThangka cultural elements detection aims to locate and identify instances in Thangka. However, as a unique form of pictorial art, Thangka exhibits distinct spatial structures that deviate significantly from general images in scale and density. Therefore, it is challenging for most state-of-the-art detectors designed for natural scenes to handle Thangka cultural elements detection effectively. To overcome this issue, we propose a multi-scale and dense object detector referred as MDDet. It embeds a multi-scale receptive field fusion module (MRF) that enlarges the receptive field while capturing the spatial and channel relationships at different scales, which significantly enriches the multi-scale features extracted from the backbone. In addition, we introduce a threshold-slicing aided hyper inference (T-SAHI) scheme, which adaptively slices images in dense scenarios to aid with dense object detection in the test time. We thoroughly evaluate our method, and MDDet outperforms the prior art by a clear margin on the Thangka dataset, achieving an absolute improvement of 1.9% in average precision (AP). For the challenging medium and small objects in Thangka, MDDet obtains wide margins of 12% and 3.7% in accuracy improvement, respectively. It also shows strong generalization ability when evaluated on general scenarios, e.g., Pascal VOC 2007 and MS COCO, validating the role of MDDet in object detection. Gaohuan Dong, Qing Xie 0002, Jiachen Li 0002, Yanchun Ma, Yuhan Liu 0001, Yongjian Liu |
MMAsia | 2 |
| 2022 | A Two-Tower Spatial-Temporal Graph Neural Network for Traffic Speed Prediction
Yansong Shen, Lin Li 0001, Qing Xie 0002, Xin Li 0064, Guandong Xu |
PAKDD (1) | 3 |
| 2021 | What is Next when Sequential Prediction Meets Implicitly Hard Interaction?abstractHard interaction learning between source sequences and their next targets is challenging, which exists in a myriad of sequential prediction tasks. During the training process, most existing methods focus on explicitly hard interactions caused by wrong responses. However, a model might conduct correct responses by capturing a subset of learnable patterns, which results in implicitly hard interactions with some unlearned patterns. As such, its generalization performance is weakened. The problem gets more serious in sequential prediction due to the interference of substantial similar candidate targets. Kaixi Hu, Lin Li 0001, Qing Xie 0002, Jianquan Liu, Xiaohui Tao 0001 |
CIKM | 3 |
| 2021 | Multi-subspace Implicit Alignment for Cross-modal Retrieval on Cooking Recipes and Food ImagesabstractCross-modal retrieval technology can help people quickly achieve mutual information between cooking recipes and food images. Both the embeddings of the image and the recipe consist of multiple representation subspaces. We argue that multiple aspects in the recipe are related to multiple regions in the food image. It is challenging to improve the cross-modal retrieval quality by making full use of the implicit connection between multiple subspaces of recipes and images. In this paper, we propose a multi-subspace implicit alignment cross-modal retrieval framework of recipes and images. Our framework learns multi-subspace information about cooking recipes and food images with multi-head attention networks; the implicit alignment at the subspace level promotes narrowing the semantic gap between recipe embeddings and food image embeddings; triple loss and adversarial loss are combined to help our framework for cross-modal learning. The experimental results show that our framework significantly outperforms to state-of-the-art methods in terms of MedR and [email protected] on Recipe 1M. Lin Li 0001, Ming Li 0072, Zichen Zan, Qing Xie 0002, Jianquan Liu |
CIKM | 4 |
| 2021 | An Empirical Study on Effect of Semantic Measures in Cross-Domain Recommender System in User Cold-Start Scenario
Yuhan Wang 0004, Qing Xie 0002, Lin Li 0001, Yongjian Liu |
KSEM | 2 |
| 2021 | Visible-infrared Person Re-identification with Human Body Parts AssistanceabstractPerson re-identification (re-id) has received ever-increasing research focus, because of its important role in video surveillance applications. This paper addresses the re-id problem between visible images of color cameras and infrared images of infrared cameras, which is significant in case that the appearance information is insufficient in poor illumination conditions. In this field, there are two key challenges, i.e., the difficulty to locate the discriminative information to re-identify the same person between visible and infrared images, and the difficulty to learn a robust metric for such large-scale cross-modality retrieval. In this paper, we propose a novel human body parts assistance network (BANet) to tackle the two challenges above. BANet mainly focuses on extracting discriminative information and learning robust features by leveraging the human body part cues. Extensive experiments demonstrate that the proposed approach outperforms the baseline and the state-of-the-art methods. Huangpeng Dai, Qing Xie 0002, Jiachen Li 0002, Yanchun Ma, Lin Li 0001, Yongjian Liu |
ICMR | 2 |
| 2021 | C2-Guard: A Cross-Correlation Gaining Framework for Urban Air Quality Prediction
Yu Chu, Lin Li 0001, Qing Xie 0002, Guandong Xu |
PAKDD (1) | 3 |
| 2020 | MOOCRec: An Attention Meta-path Based Model for Top-K Recommendation in MOOC
Deming Sheng, Jingling Yuan, Qing Xie 0002, Pei Luo |
KSEM (1) | 3 |
| 2018 | A Hybrid Model Reuse Training Approach for Multilingual OCR
Zhongwei Xie, Lin Li 0001, Xian Zhong, Luo Zhong, Qing Xie 0002, Jianwen Xiang |
WISE (1) | 5 |
| 2017 | Co-training an Improved Recurrent Neural Network with Probability Statistic Models for Named Entity Recognition
Yueqing Sun, Lin Li 0001, Zhongwei Xie, Qing Xie 0002, Xin Li 0064, Guandong Xu |
DASFAA (2) | 4 |
| 2016 | CB-CAS: A CAS-Based Cross-Browser SSO System
Yongjian Liu, Qing Xie 0002 |
APWeb (2) | 3 |
| 2016 | TagTour: A Personalized Tourist Resource Recommendation System
Tian Han 0002, Yongjian Liu, Qing Xie 0002 |
APWeb (2) | 3 |
| 2016 | Personalized Resource Recommendation Based on Regular Tag and User Operation
Yongjian Liu, Qing Xie 0002 |
APWeb (2) | 3 |
| 2016 | CrowdAidRepair: A Crowd-Aided Interactive Data Repairing Method
Zhixu Li, Binbin Gu, Qing Xie 0002, Jia Zhu 0003, Xiangliang Zhang 0001, Guoliang Li 0001 |
DASFAA (1) | 4 |
| 2016 | Exploiting link structure for web page genre identification
Jia Zhu 0003, Qing Xie 0002, Shoou-I Yu, Wai-Hung Collin Wong |
Data Min. Knowl. Discov. | 2 |
| 2016 | Optimizing Cost of Continuous Overlapping Queries over Data Streams by Filter AdaptionabstractThe problem we aim to address is the optimization of cost management for executing multiple continuous queries on data streams, where each query is defined by several filters, each of which monitors certain status of the data stream. Specially, the filter can be shared by different queries and expensive to evaluate. The conventional objective for such a problem is to minimize the overall execution cost to solve all queries, by planning the order of filter evaluation in shared strategy. However, in the streaming scenario, the characteristics of data items may change in process, which can bring some uncertainty to the outcome of individual filter evaluation, and affect the plan of query execution as well as the overall execution cost. In our work, considering the influence of the uncertain variation of data characteristics, we propose a framework to deal with the dynamic adjustment of filter ordering for query execution on data stream, and focus on the issues of cost management. By incrementally monitoring and analyzing the results of filter evaluation, our proposed approach can be effectively adaptive to the varied stream behavior and adjust the optimal ordering of filter evaluation, so as to optimize the execution cost. In order to achieve satisfactory performance and efficiency, we also discuss the trade-off between the adaptivity of our framework and the overhead incurred by filter adaption. The experimental results on synthetic and two real data sets (traffic and multimedia) show that our framework can effectively reduce and balance the overall query execution cost and keep high adaptivity in streaming scenario. Qing Xie 0002, Xiangliang Zhang 0001, Zhixu Li, Xiaofang Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | Addressing Instance Ambiguity in Web HarvestingabstractWeb Harvesting enables the enrichment of incomplete data sets by retrieving required information from the Web. However, the ambiguity of instances may greatly decrease the quality of the harvested data, given that any instance in the local data set may become ambiguous when attempting to identify it on the Web. Although plenty of disambiguation methods have been proposed to deal with the ambiguity problems in various settings, none of them are able to handle the instance ambiguity problem in Web Harvesting. In this paper, we propose to do instance disambiguation in Web Harvesting with a novel disambiguation method inspired by the idea of collaborative identity recognition. In particular, we expect to find some common properties in forms of latent shared attribute values among instances in the list, such that these shared attribute values can differentiate instances within the list against those ambiguous ones on the Web. Our extensive experimental evaluation illustrates the utility of collaborative disambiguation for a popular Web Harvesting application, and shows that it substantially improves the accuracy of the harvested data. Zhixu Li, Xiangliang Zhang 0001, Hai Huang 0003, Qing Xie 0002, Jia Zhu 0003, Xiaofang Zhou 0001 |
WebDB | 4 |
| 2015 | An improved early detection method of type-2 diabetes mellitus using multiple classifier system
Jia Zhu 0003, Qing Xie 0002, Kai Zheng 0001 |
Inf. Sci. | 2 |
| 2014 | Cost Reduction for Web-Based Data Imputation
Zhixu Li, Shuo Shang, Qing Xie 0002, Xiangliang Zhang 0001 |
DASFAA (2) | 3 |
| 2014 | Maximum error-bounded Piecewise Linear Representation for online stream approximation
Qing Xie 0002, Chaoyi Pang, Xiaofang Zhou 0001, Xiangliang Zhang 0001 |
VLDB J. | 1 |
| 2013 | Local correlation detection with linearity enhancement in streaming dataabstractThis paper addresses the challenges in detecting the potential correlation between numerical data streams, which facilitates the research of data stream mining and pattern discovery. We focus on local correlation with delay, which may occur in burst at different time in different streams, and last for a limited period. The uncertainty on the correlation occurrence and the time delay make it difficult to monitor the correlation online. Furthermore, the conventional correlation measure lacks the ability of reflecting visual linearity, which is more desirable in reality. This paper proposes effective methods to continuously detect the correlation between data streams. Our approach is based on the Discrete Fourier Transform to make rapid cross-correlation calculation with time delay allowed. In addition, we introduce a shape-based similarity measure into the framework, which refines the results by representative trend patterns to enhance the significance of linearity. The similarity of proposed linear representations can quickly estimate the correlation, and the window sliding strategy in segment level improves the efficiency for online detection. The empirical study demonstrates the accuracy of our detection approach, as well as more than $30\%$ improvement of efficiency. Qing Xie 0002, Shuo Shang, Bo Yuan 0003, Chaoyi Pang, Xiangliang Zhang 0001 |
CIKM | 1 |
| 2012 | Efficient buffer management for piecewise linear representation of multiple data streamsabstractPiecewise Linear Representation (PLR) has been a widely used method for approximating data streams in the form of compact line segments. The buffer-based approach to PLR enables a semi-global approximation which relies on the aggregated processing of batches of streamed data so that to adjust and improve the approximation results. However, one challenge towards applying the buffer-based approach is allocating the necessary memory resources for stream buffering. This challenge is further complicated in a multi-stream environment where multiple data streams are competing for the available memory resources, especially in resource-constrained systems such as sensors and mobile devices. Qing Xie 0002, Jia Zhu 0003, Mohamed A. Sharaf, Xiaofang Zhou 0001, Chaoyi Pang |
CIKM | 1 |
| 2012 | A Hybrid Time-Series Link Prediction Framework for Large Social Network
Jia Zhu 0003, Qing Xie 0002, Eun Jung Chin |
DEXA (2) | 2 |
| 2010 | Efficient and Continuous Near-duplicate Video DetectionabstractOnline video steam data is surging to an unprecedented level. Massive video publishing and sharing impose heavy demands on continuous video near-duplicate detection for many novel video applications. This paper presents an accurate and accelerated system for video near-duplicate detection over continuous video streams. We propose to transform a high-dimensional video stream into a one-dimensional Video Trend Stream (VTS) to monitor the continuous luminance changes of consecutive frames, based on which video similarity is derived. In order to do fast comparison and effective early pruning, a compact auxiliary signature named CutSig is proposed to approximate the video structure. CutSig explores cut distribution feature of the video structure and contributes to filter candidates quickly. To scan along a video stream in a rapid way, shot cuts with local maximum AI (average information) in a query video are used as reference cuts, and a skipping approach based on reference cut alignment is embedded for efficient acceleration. Extensive experimental results on detecting diverse near-duplicates in real video streams show the effectiveness and efficiency of our method. Qing Xie 0002, Zi Huang, Heng Tao Shen, Xiaofang Zhou 0001, Chaoyi Pang |
APWeb | 1 |