EDBT 2026 Demo / reviewers in the wild / expert
Tengjiao Wang 0003
dblp:39/1084-3
· DBLP profile ↗
92ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0002-7395-5567ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 63 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 33 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorSystems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Soft Contrastive Learning for Spatio-Temporal Forecasting
Hanzhi Deng, Heyuan Wang 0001, Tengjiao Wang 0003, Kam-Fai Wong |
DASFAA (4) | 4 |
| 2025 | Clear Up Confusion: Iterative Differential Generation for Fine-grained Intent Detection with Contrastive FeedbackabstractFine-grained intent detection involves identifying a large number of classes with subtle variations. Recently, generating pseudo samples via large language models has attracted increasing attention to alleviate the data scarcity caused by emerging new intents. However, these methods generate samples for each class independently and neglect the relationships between classes, leading to ambiguity in pseudo samples, particularly for fine-grained labels. And, they typically rely on one-time generation and overlook feedback from pseudo samples. In this paper, we propose an iterative differential generation framework with contrastive feedback to generate high-quality pseudo samples and accurately capture the crucial nuances in target class distribution. Specifically, we propose differential guidelines that include potential ambiguous labels to reduce confusion for similar labels. Then we conduct rubric-driven refinement, ensuring the validity and diversity of pseudo samples. Finally, despite one generation, we propose to iteratively generate new samples with contrastive feedback to achieve accurate identification and distillation of target knowledge. Extensive experiments in zero/few-shot and full-shot settings on three datasets verify the effectiveness of our method. Feng Zhang 0027, Wei Chen 0056, Tengjiao Wang 0003, Jiahui Yao, Jiabin Zheng |
COLING | 5 |
| 2025 | PR-KGC: Text-enhanced Knowledge Graph Completion with Pair-wise Re-rankingabstractRecent advancements in Knowledge Graph Completion (KGC) often adopt a two-stage pipeline that combines triple-based retrieval with text-based re-ranking. However, point-wise re-rankers, which score candidates individually, often fail to capture subtle distinctions between similar candidates due to the lack of direct comparisons. While list-wise re-rankers address this by evaluating all candidates simultaneously, generating permutations of all candidates is computationally challenging for pre-trained language models (PLMs) and can result in issues such as omissions, rejections, and especially inconsistencies in the output. To address these challenges, this paper introduces a Pairwise Re-ranking method for Knowledge Graph Completion (PRKGC), which mitigates the complexities of point-wise calibrated scoring and list-wise permutation outputs. It reduces the burden on PLMs by requiring them to perform nuanced comparisons between pairs of candidates, rather than all candidates at once. During inference, our approach processes all possible permutations of the top k candidate pairs, ensuring a thorough evaluation and consistency in ranking. Extensive experiments on link prediction tasks demonstrate that the proposed strategy effectively elevates much smaller PLMs (∼100M parameters) to achieve state-of-the-art performance, outperforming models based on 3x RoBERTa-Large and 70x LLaMA2-7B. Additionally, case studies reveal that these improvements stem from the model’s enhanced ability to discern subtle differences between similar candidates. Under nearly identical performance in Hits@3, it outperforms the most competitive baselines by approximately 1.6-3.8% in Hits@1. Feng Zhang 0027, Wei Chen 0056, Tengjiao Wang 0003, Jiabin Zheng, Jiahui Yao |
ICASSP | 4 |
| 2025 | Less is Enough: Relation Graph Guided Few-shot Learning for Multi-label Aspect Category DetectionabstractFew-shot Multi-label Aspect Category Detection (FMACD) is an essential task, which aims to identify multiple aspect categories in a given sentence with limited data. Recently, the prototypical network as a mainline has been used for the task due to its powerful capacity. However, existing methods mostly rely on intra-cluster samples to generate prototypes, and they struggle to extract robust prototype features in very few data cases (e.g., 1-shot). Therefore, these methods may fail to estimate label-query relevance during multi-label prediction. To solve the above issues, we propose a novel relation graph guided learning method for FMACD by considering all intra- and inter-cluster samples. Specifically, the proposed method explicitly models a relation graph to generate more robust prototypes by exploring sample relations among intra- and inter-cluster. Then, a multi-label inference strategy is proposed to enhance label-query relevance for multi-label prediction. Besides, graph contrastive learning enhances intra-cluster commonality and inter-cluster uniqueness to improve performance. Experiments show that the proposed method achieves significant performance, esp., it obtains an average of 1.55% AUC and 5.01% Macro-F1 improvement in 1-shot scenarios. Shiman Zhao, Wei Chen 0056, Tengjiao Wang 0003, Jiahui Yao, Jiabin Zheng |
ICASSP | 3 |
| 2025 | Instance Relation Learning Network with Label Knowledge Propagation for Few-shot Multi-label Intent DetectionabstractFew-shot Multi-label Intent Detection (MID) is crucial for dialogue systems, aiming to detect multiple intents of utterances in low-resource dialogue domains. Previous studies focus on a two-stage pipeline. They first learn representations of utterances with multiple labels and then use a threshold-based strategy to identify multi-label results. However, these methods rely on representation classification and ignore instance relations, leading to error propagation. To solve the above issues, we propose a multi-label joint learning method for few-shot MID in an end-to-end manner, which constructs an instance relation learning network with label knowledge propagation to eliminate error propagation. Concretely, we learn the interaction relations between instances with class information to propagate label knowledge between a few labeled (support set) and unlabeled (query set) instances. With label knowledge propagation, the relation strength between instances directly indicates whether two utterances belong to the same intent for multi-label prediction. Besides, a dual relation-enhanced loss is developed to optimize support- and query-level relation strength to improve performance. Experiments show that we outperform strong baselines by an average of 9.54% AUC and 11.19% Macro-F1 in 1-shot scenarios. Shiman Zhao, Shangyuan Li, Wei Chen 0056, Tengjiao Wang 0003, Jiahui Yao, Jiabin Zheng, Kam-Fai Wong |
IJCAI | 4 |
| 2025 | Enhancer: A Distribution-Aware Framework with Temporal-Relational Meta-Learning for Stock PredictionabstractAccurate stock prediction is critical for portfolio management, where learning to adapt to market changes is the key to sustainable profitability. Financial markets, as complex interactive systems, exhibit evolution in both temporal dynamics and relational structures. While current temporal-relational models have achieved remarkable success in stock prediction, they face fundamental challenges in learning and adapting to market changes, particularly the systematic shifts in temporal and relational distributions that challenge the i.i.d. assumption underlying model training. In this study, we pioneer the research of temporal-relational distribution shifts in stock prediction and introduce Enhancer, a model-agnostic framework that can be applied to any downstream predictor. Enhancer adopts a meta-learning architecture featuring both a Temporal Meta-Learner (TML) and a Relational Meta-Learner (RML). Specifically, we introduce Reactive Point Processes Attention (RPPsAtt) within TML to overcome the limitations of missing fine-grained temporal point information, a common issue with prior methods that rely on distribution inference for mitigating temporal distribution shift. To enhance relational generalization, we introduce the Approximation-Intervention (Ant) mechanism within RML, marking the first method to mitigate relational distribution shift for quantitative investment. We conduct experiments on four long-term stock datasets, categorizing them into two tasks: stock trend prediction and stock investment recommendation. Our experimental results show that Enhancer achieves an average increase of 29.3% in profit ratio and 18.54% in the Sharpe ratio compared to the baselines across two tasks. Weijun Chen 0002, Shun Li 0001, Heyuan Wang 0001, Tengjiao Wang 0003 |
KDD (2) | 4 |
| 2025 | Basis is also explanation: Interpretable Legal Judgment Reasoning prompted by multi-source knowledge
Shangyuan Li, Shiman Zhao, Zhuoran Zhang 0003, Zihao Fang, Wei Chen 0056, Tengjiao Wang 0003 |
Inf. Process. Manag. | 6 |
| 2025 | ReranKGC: A cooperative retrieve-and-rerank framework for multi-modal knowledge graph completion
Wei Chen 0056, Feng Zhang 0027, Tengjiao Wang 0003, Jiahui Yao, Jiabin Zheng, Kam-Fai Wong |
Neural Networks | 6 |
| 2024 | Meta-Prompt Tuning Vision-Language Model for Multi-Label Few-Shot Image RecognitionabstractMulti-label few-shot image recognition aims to identify multiple unseen objects using only a handful of examples. Recent methods typically tune pre-trained vision-language models with shared or class-specific prompts. However, they still have drawbacks. Tuning a shared prompt is insufficient for all samples especially when the tasks are complex and tuning specific prompts for each class is inevitable to lose generalization ability, thus failing to capture diverse visual knowledge. To address these issues, we propose to meta-tune a generalized prompt pool, enabling each prompt to act as an expert for multi-label few-shot image recognition. Specifically, we first construct a diverse prompt pool to handle complex samples and tasks effectively. Then, the meta-tuning strategy is designed to learn meta-knowledge and transfer it from source tasks to target tasks, enhancing the generalization of prompts. Extensive experimental results on two widely used multi-label image recognition datasets demonstrate the effectiveness of our method. Feng Zhang 0027, Wei Chen 0056, Tengjiao Wang 0003, Jiabin Zheng |
CIKM | 4 |
| 2024 | Adaptive Image-Enhanced Knowledge Graph CompletionabstractImage-enhanced knowledge graph completion aims to predict potential linkings in KGs by utilizing auxiliary visual information from entity’s images. Ideally, the image selection process should be dynamic, considering the context of inferred relations. However, current methods often assess image importance in a static way, based on how well they represent each entity. To solve above issues, we propose an Adaptive Image-enhanced Knowledge Graph Completion model (AI-KGC). The core idea behind this model is to switch image selection strategies adaptively according to the entity’s role. Specially, for a query entity, we employ a relation-guided dynamic image selection mechanism to select the image that is most appropriate for the given relation context. The chosen image may offer insightful clues about the ground-truth entity. For a candidate entity, we use a name-guided static image selection mechanism to capture its most representative visual features. By exploiting a similarity-based scoring function, the ranking of entities exhibiting similar visual features therefore will be enhanced. This adaptive selection strategy can extract valuable insights from images that previous methods may have overlooked. Extensive experiments conducted on link prediction tasks show that AI-KGC outperforms the state-of-the-art KGC methods remarkably. Wei Chen 0056, Tengjiao Wang 0003, Jiabin Zheng |
ICASSP | 3 |
| 2024 | Automatic De-Biased Temporal-Relational Modeling for Stock Investment Recommendation
Weijun Chen 0002, Shun Li 0001, Xipu Yu, Heyuan Wang 0001, Wei Chen 0021, Tengjiao Wang 0003 |
IJCAI | 6 |
| 2024 | Metric-Free Learning Network with Dual Relations Propagation for Few-Shot Aspect Category Sentiment AnalysisabstractAbstract Few-shot Aspect Category Sentiment Analysis (ACSA) is a crucial task for aspect-based sentiment analysis, which aims to detect sentiment polarity for a given aspect category in a sentence with limited data. However, few-shot learning methods focus on distance metrics between the query and support sets to classify queries, heavily relying on aspect distributions in the embedding space. Thus, they suffer from overlapping distributions of aspect embeddings caused by irrelevant sentiment noise among sentences with multiple sentiment aspects, leading to misclassifications. To solve the above issues, we propose a metric-free method for few-shot ACSA, which models the associated relations among the aspects of support and query sentences by Dual Relations Propagation (DRP), addressing the passive effect of overlapping distributions. Specifically, DRP uses the dual relations (similarity and diversity) among the aspects of support and query sentences to explore intra-cluster commonality and inter-cluster uniqueness for alleviating sentiment noise and enhancing aspect features. Additionally, the dual relations are transformed from support-query to class-query to promote query inference by learning class knowledge. Experiments show that we achieve convincing performance on few-shot ACSA, especially an average improvement of 2.93% accuracy and 2.10% F1 score in the 3-way 1-shot setting. Shiman Zhao, Wei Chen 0056, Tengjiao Wang 0003, Jiahui Yao, Jiabin Zheng |
Trans. Assoc. Comput. Linguistics | 4 |
| 2024 | Agree to Disagree: Personalized Temporal Embedding and Routing for Stock ForecastabstractStock forecast is a crucial yet challenging task in modern quantitative trading. Given theoretical and investment merits, recently a variety of deep learning methods have been proposed for automatically simulating stock movements from historical time series. However, these methods typically follow the i.i.d. assumption that actually contradicts the complex trading environment. In reality, individual stocks often exhibit diverse volatility patterns, while macro market scenarios may also change over time, jointly resulting in distribution shifts and weak generalization. To combat these bottlenecks, in this paper we propose a new learning architecture calledPersonalized Temporal Embedding and Routing(PTER) to improve stock forecast by forming a relaxed weight-sharing paradigm. The key of PTER is introducing hypernetworks to guide tailoring target network parameters, such that stock time series are embedded adapting to multi-object multi-scenario data disparities. Specifically, in the encoding stage, PTER first captures hyper-knowledge characterizing the similarity and peculiarity of different stocks and market scenarios. The knowledge space is then projected onto the temporal parameter space, enabling the customization of protruded features from chaotic observation signals. In the inference stage, each sample is dispatched to orthogonal predictor heads to dynamically output expected returns based on market conditions. Through experiments on benchmark datasets spanning over five years on four of the world's largest exchange markets, we show that PTER improves the cumulative and risk-adjusted revenue performance by a significant margin. Heyuan Wang 0001, Tengjiao Wang 0003, Shun Li 0001, Weijun Chen 0002, Wei Chen 0056 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Dual Class Knowledge Propagation Network for Multi-label Few-shot Intent DetectionabstractMulti-label intent detection aims to assign multiple labels to utterances and attracts increasing attention as a practical task in task-oriented dialogue systems.As dialogue domains change rapidly and new intents emerge fast, the lack of annotated data motivates multi-label few-shot intent detection.However, previous studies are confused by the identical representation of the utterance with multiple labels and overlook the intrinsic intra-class and inter-class interactions.To address these two limitations, we propose a novel dual class knowledge propagation network in this paper.In order to learn well-separated representations for utterances with multiple intents, we first introduce a labelsemantic augmentation module incorporating class name information.For better consideration of the inherent intra-class and inter-class relations, an instance-level and a class-level graph neural network are constructed, which not only propagate label information but also propagate feature structure.And we use a simple yet effective method to predict the intent count of each utterance.Extensive experimental results on two multi-label intent datasets have demonstrated that our proposed method outperforms strong baselines by a large margin. Feng Zhang 0027, Wei Chen 0056, Tengjiao Wang 0003 |
ACL (1) | 4 |
| 2023 | Learning Few-shot Sample-set Operations for Noisy Multi-label Aspect Category DetectionabstractMulti-label Aspect Category Detection (MACD) is essential for aspect-based sentiment analysis, which aims to identify multiple aspect categories in a given sentence. Few-shot MACD is critical due to the scarcity of labeled data. However, MACD is a high-noise task, and existing methods fail to address it with only two or three training samples per class, which limits the application in practice. To solve above issues, we propose a group of Few-shot Sample-set Operations (FSO) to solve noisy MACD in fewer sample scenarios by identifying the semantic contents of samples. Learning interactions among intersection, subtraction, and union networks, the FSO imitates arithmetic operations on samples to distinguish relevant and irrelevant aspect contents. Eliminating the negative effect caused by noises, the FSO extracts discriminative prototypes and customizes a dedicated query vector for each class. Besides, we design a multi-label architecture, which integrates with score-wise loss and multi-label loss to optimize the FSO for multi-label prediction, avoiding complex threshold training or selection. Experiments show that our method achieves considerable performance. Significantly, it improves by 11.01% at most and an average of 8.59% Macro-F in fewer sample scenarios. Shiman Zhao, Tengjiao Wang 0003 |
IJCAI | 3 |
| 2023 | HATR-I: Hierarchical Adaptive Temporal Relational Interaction for Stock Trend PredictionabstractStock trend prediction is a hot issue in theFintechfield. Effective stock profiling is challenging due to highly non-stationary dynamics and complex interplays. Existing methods usually regard each stock independently or detect simplistic homogeneous structures. Practically, stock correlation originates from diverse aspects and underlying relationship signals are implicit in comprehensive graphs. Besides, RNNs are extensively used to simulate stock volatility while inadequate in capturing fine-granular patterns across local time snippets. To this end, in this paper we propose HATR-I, a Hierarchical Adaptive Temporal-Relational Interaction model to characterize and predict stock evolutions. Specifically, we grasp short- and long-term transition regularities of stock dynamics based on cascaded dilated convolutions and gating paths. By formulating different views of domain adjacency graphs into a unified multiplex network with edge attributes, we inject node- and semantic-level dual attention to refine the propagation of inter-stock collaborative information. Particularly, the stock pair matching is proceeding along each time-stage rather than until final compressed representations, meanwhile significant feature points and scales are identified considering the effect of time attenuation. Finally, we deduce latent shared clusters as global regularization to optimize the stock representations. Experiments on three real-world stock market datasets demonstrate the effectiveness of our proposed model. Heyuan Wang 0001, Tengjiao Wang 0003, Shun Li 0001, Shijie Guan |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Heterogeneous Interactive Snapshot Network for Review-Enhanced Stock Profiling and RecommendationabstractStock recommendation plays a critical role in modern quantitative trading. The large volumes of social media information such as investment reviews that delegate emotion-driven factors, together with price technical indicators formulate a “snapshot” of the evolving stock market profile. However, previous studies usually model the temporal trajectories of price and media modalities separately while losing their interrelated influences. Moreover, they mainly extract review semantics via sequential or attentive models, whereas the rich text associated knowledge is largely neglected. In this paper, we propose a novel heterogeneous interactive snapshot network for stock profiling and recommendation. We model investment reviews in each snapshot as a heterogeneous document graph, and develop a flexible hierarchical attentive propagation framework to capture fine-grained proximity features. Further, to learn stock embedding for ranking, we introduce a novel twins-GRU method, which tightly couples the media and price parallel sequences in a cross-interactive fashion to catch dynamic dependencies between successive snapshots. Our approach excels state-of-the-arts over 7.6% in terms of cumulative and risk-adjusted returns in trading simulations on both English and Chinese benchmarks. Heyuan Wang 0001, Tengjiao Wang 0003, Shun Li 0001, Shijie Guan, Wei Chen 0021 |
IJCAI | 2 |
| 2022 | Adaptive Long-Short Pattern Transformer for Stock Investment SelectionabstractStock investment selection is a hard issue in the Fintech field due to non-stationary dynamics and complex market interdependencies. Existing studies are mostly based on RNNs, which struggle to capture interactive information among fine granular volatility patterns. Besides, they either treat stocks as isolated, or presuppose a fixed graph structure heavily relying on prior domain knowledge. In this paper, we propose a novel Adaptive Long-Short Pattern Transformer (ALSP-TF) for stock ranking in terms of expected returns. Specifically, we overcome the limitations of canonical self-attention including context and position agnostic, with two additional capacities: (i) fine-grained pattern distiller to contextualize queries and keys based on localized feature scales, and (ii) time-adaptive modulator to let the dependency modeling among pattern pairs sensitive to different time intervals. Attention heads in stacked layers gradually harvest short- and long-term transition traits, spontaneously boosting the diversity of representations. Moreover, we devise a graph self-supervised regularization, which helps automatically assimilate the collective synergy of stocks and improve the generalization ability of overall model. Experiments on three exchange market datasets show ALSP-TF’s superiority over state-of-the-art stock forecast methods. Heyuan Wang 0001, Tengjiao Wang 0003, Shun Li 0001, Shijie Guan, Wei Chen 0021 |
IJCAI | 2 |
| 2022 | Stable Contrastive Learning for Self-Supervised Sentence Embeddings With Pseudo-Siamese Mutual LearningabstractLearning semantic sentence embeddings is beneficial to a variety of natural language processing tasks. Recently, methods using the contrastive learning framework to fine-tune pre-trained language models have been proposed and have achieved significant performance on sentence embeddings. However, sentence embeddings are easy to “overfit” to the contrastive learning goal. With the training of contrastive learning, the gap between contrastive learning and test tasks leads to unstable even declining performance on test tasks. For this reason, existing methods rely on the labeled development set to frequently evaluate the performance on test tasks and get the best checkpoints. In such a way, models are limited when the labeled data is unavailable or extremely scarce. To address this problem, we proposePseudo-Siamese networkMutualLearning (PSML) for self-supervised sentence embeddings to reduce the gap between contrastive learning and test tasks. Consisting of the main encoder and the auxiliary encoder, PSML utilizes mutual learning as the basic framework. Between the two encoders, two mutual learning losses are constructed to share learning signals. The proposed model framework and losses of PSML help the model be optimized more stably and generalize better to test tasks, such as semantic textual similarity. Extensive experiments on seven public semantic textual similarity datasets show that PSML performs better than previous unsupervised contrastive methods for sentence embeddings. Besides, PSML also gives a stable performance curve on test tasks with training and is able to get the comparative performance without frequent evaluation on the labeled development set. Qiyu Wu 0001, Wei Chen 0056, Tengjiao Wang 0003 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Hierarchical Adaptive Temporal-Relational Modeling for Stock Trend PredictionabstractStock trend prediction is a challenging task due to the non-stationary dynamics and complex market dependencies. Existing methods usually regard each stock as isolated for prediction, or simply detect their correlations based on a fixed predefined graph structure. Genuinely, stock associations stem from diverse aspects, the underlying relation signals should be implicit in comprehensive graphs. On the other hand, the RNN network is mainly used to model stock historical data, while is hard to capture fine-granular volatility patterns implied in different time spans. In this paper, we propose a novel Hierarchical Adaptive Temporal-Relational Network (HATR) to characterize and predict stock evolutions. By stacking dilated causal convolutions and gating paths, short- and long-term transition features are gradually grasped from multi-scale local compositions of stock trading sequences. Particularly, a dual attention mechanism with Hawkes process and target-specific query is proposed to detect significant temporal points and scales conditioned on individual stock traits. Furthermore, we develop a multi-graph interaction module which consolidates prior domain knowledge and data-driven adaptive learning to capture interdependencies among stocks. All components are integrated seamlessly in a unified end-to-end framework. Experiments on three real-world stock market datasets validate the effectiveness of our model. Heyuan Wang 0001, Shun Li 0001, Tengjiao Wang 0003 |
IJCAI | 3 |
| 2020 | Incorporating Expert-Based Investment Opinion Signals in Stock Prediction: A Deep Learning FrameworkabstractInvestment messages published on social media platforms are highly valuable for stock prediction. Most previous work regards overall message sentiments as forecast indicators and relies on shallow features (bag-of-words, noun phrases, etc.) to determine the investment opinion signals. These methods neither capture the time-sensitive and target-aware characteristics of stock investment reviews, nor consider the impact of investor's reliability. In this study, we provide an in-depth analysis of public stock reviews and their application in stock movement prediction. Specifically, we propose a novel framework which includes the following three key components: time-sensitive and target-aware investment stance detection, expert-based dynamic stance aggregation, and stock movement prediction. We first introduce our stance detection model named MFN, which learns the representation of each review by integrating multi-view textual features and extended knowledge in financial domain to distill bullish/bearish investment opinions. Then we show how to identify the validity of each review, and enhance stock movement prediction by incorporating expert-based aggregated opinion signals. Experiments on real datasets show our framework can effectively improve the performance of both investment opinion mining and individual stock forecasting. Heyuan Wang 0001, Tengjiao Wang 0003 |
AAAI | 2 |
| 2020 | Stance Detection with Stance-Wise Convolution Network
Dechuan Yang, Qiyu Wu 0001, Wei Chen 0021, Tengjiao Wang 0003, Yingbao Cui |
NLPCC (1) | 4 |
| 2019 | Type Sequence Preserving Heterogeneous Information Network EmbeddingabstractLacking in sequence preserving mechanism, existing heterogeneous information network (HIN) embedding discards the essential type sequence information during embedding. We propose a Type Sequence Preserving HIN Embedding model (SeqHINE) which expands the HIN embedding to sequence level. SeqHINE incorporates the type sequence information via type-aware GRU and preserves representative sequence information by decay function. Abundant experiments show that SeqHINE can outperform state-of-the-art even with 50% less labeled data. Tengjiao Wang 0003, Wei Chen 0021 |
AAAI | 2 |
| 2019 | An Environment-Aware Market Strategy for Data Allocation and Dynamic Migration in Cloud DatabaseabstractCurrently, a cloud database is employed to serve on-line query-intensive applications. It inevitably happens that some cloud data nodes storing hot records are facing high frequent query requests while others are rarely visited or even idle. Therefore, how data are dynamically allocated and migrated at runtime has significant impact on query load distribution and system performance. Existing system adopt centralized approaches, and they face two main challenges: (1) Query load on individual node cannot be always balancing even if the data are fairly distributed; (2) For each node, the dynamic changes of configuration resources cannot be captured during the runtime. To this end, this paper presents an environment-aware market strategy based system, named e-MARS, for reasonable data migration to achieve query load balance in cloud database. In e-MARS, cloud database is modeled as a cloudDB market, while data nodes are regarded as intelligent traders and the query load as commodity. Each trader is aware of its local environmental re-sources, such as computing capacity, disk volume, based on which the trader itself decides how to trade the query load and migrates the corresponding data. In this way the cloudDB market will achieve equilibrium. Experiments are conducted on the real communication data, and e-MARS significantly enhances the efficiency. Compared with HBase Balancer, more than 65% improvement is achieved in terms of query response time. Tengjiao Wang 0003, Binyang Li, Wei Chen 0021, Jinzhong Niu, Kam-Fai Wong |
ICDE | 1 |
| 2018 | ROSIE: Runtime Optimization of SPARQL Queries over RDF Using Incremental Evaluation
Lei Gai, Tengjiao Wang 0003 |
KSEM (2) | 3 |
| 2018 | The UIR Uncertainty Corpus for Chinese: Annotating Chinese Microblog Corpus for Uncertainty Identification from Social Media
Binyang Li, Ruifeng Xu 0001, Tengjiao Wang 0003, Kam-Fai Wong |
LREC | 7 |
| 2017 | Category-Level Transfer Learning from Knowledge Base to Microblog Stream for Accurate Event Detection
Weijing Huang, Tengjiao Wang 0003, Wei Chen 0021 |
DASFAA (1) | 2 |
| 2017 | Efficient Topic Modeling on Phrases via SparsityabstractTopic modeling on phrases is important in understanding documents by providing interpretable topics. But existing methods are not as efficient as the topic modeling methods on words, which may limit their potential application.Towards providing a more efficient method, we propose a novel topic model SparseTP, which (1) models the words and phrases by linking them in Markov Random Field when necessary; (2) provides a well-formed lower bound of the model for Gibbs sampling; (3) utilizes the sparse distribution of words and phrases on topics to speed up the inference. The experiments demonstrate that it can achieve the high efficiency without sacrificing the effectiveness. Weijing Huang, Wei Chen 0021, Tengjiao Wang 0003, Shibo Tao |
ICTAI | 3 |
| 2017 | Opinion-aware Knowledge Graph for Political Ideology DetectionabstractIdentifying individual's political ideology from their speeches and written texts is important for analyzing political opinions and user behavior on social media. Traditional opinion mining methods rely on bag-of-words representations to classify texts into different ideology categories. Such methods are too coarse for understanding political ideologies. The key to identify different ideologies is to recognize different opinions expressed toward a specific topic. To model this insight, we classify ideologies based on the distribution of opinions expressed towards real-world entities or topics. Specifically, we propose a novel approach to political ideology detection that makes predictions based on an opinion-aware knowledge graph. We show how to construct such graph by integrating the opinions and targeted entities extracted from text into an existing structured knowledge base, and show how to perform ideology inference by information propagation on the graph. Experimental results demonstrate that our method achieves high accuracy in detecting ideologies compared to baselines including LR, SVM and RNN. Wei Chen 0021, Tengjiao Wang 0003, Bishan Yang |
IJCAI | 3 |
| 2016 | An Efficient Online Event Detection Method for Microblogs via User Modeling
Weijing Huang, Wei Chen 0021, Lamei Zhang, Tengjiao Wang 0003 |
APWeb (1) | 4 |
| 2016 | Online Learning for Accurate Real-Time Map Matching
Biwei Liang, Tengjiao Wang 0003, Shun Li 0001, Wei Chen 0021, Hongyan Li 0002, Kai Lei |
PAKDD (2) | 2 |
| 2016 | Dboost: A Fast Algorithm for DBSCAN-based Clustering on High Dimensional Data
Xiaorong Wang, Wei Chen 0021, Tengjiao Wang 0003, Kai Lei |
PAKDD (2) | 5 |
| 2016 | Valuable Group Trajectory Pattern Mining Directed by Adaptable Value Measuring Model
Tengjiao Wang 0003, Shun Li 0001, Wei Chen 0021 |
WAIM (2) | 2 |
| 2015 | ASEM: Mining Aspects and Sentiment of Events from MicroblogabstractMicroblogs contain the most up-to-date and abundant opinion information on current events. Aspect-based opinion mining is a good way to get a comprehensive summarization of events. The most popular aspect based opinion mining models are used in the field of product and service. However, existing models are not suitable for event mining. In this paper we propose a novel probabilistic generative model (ASEM) to simultaneously discover aspects and the specified opinions. ASEM incorporate a sequence labeling model(CRF) into a generative topic model. Additionally, we adopt a set of features for separating aspects and sentiments. Moreover, we novelly present a continuously learning model. It can utilize the knowledge of one event to learn another, and get a better performance. We use five real world events to do experiment. The experimental results show that ASEM extracts aspects and sentiments well, and ASEM outperforms other state-of-art models and the intuitive two-step method. Ruhui Wang, Weijing Huang, Wei Chen 0021, Tengjiao Wang 0003, Kai Lei |
CIKM | 4 |
| 2015 | Overlapping Community Detection in Directed Heterogeneous Social Network
Changhe Qiu, Wei Chen 0021, Tengjiao Wang 0003, Kai Lei |
WAIM | 3 |
| 2015 | An Adaptive Skew Handling Join Algorithm for Large-scale Data Analysis
Tengjiao Wang 0003, Shun Li 0001, Hongyan Li 0002, Kai Lei |
WAIM | 2 |
| 2015 | SparkRDF: In-Memory Distributed RDF Management Framework for Large-Scale Social Data
Wei Chen 0021, Lei Gai, Tengjiao Wang 0003 |
WAIM | 4 |
| 2014 | Towards Efficient Path Query on Social Network with Hybrid RDF Management
Lei Gai, Wei Chen 0021, Changhe Qiu, Tengjiao Wang 0003 |
APWeb | 5 |
| 2014 | An Adaptive Skew Insensitive Join Algorithm for Large Scale Data Analytics
Wenjing Liao, Tengjiao Wang 0003, Hongyan Li 0002, Dongqing Yang, Kai Lei |
APWeb | 2 |
| 2014 | Topic-Based Sentiment Analysis Incorporating User Interactions
Wei Chen 0021, Gaoyan Ou, Tengjiao Wang 0003, Dongqing Yang, Kai Lei |
APWeb | 4 |
| 2014 | CLUSM: An Unsupervised Model for Microblog Sentiment Analysis Incorporating Link Information
Gaoyan Ou, Wei Chen 0021, Binyang Li, Tengjiao Wang 0003, Dongqing Yang, Kam-Fai Wong |
DASFAA (1) | 4 |
| 2014 | Exploiting Community Emotion for Microblog Event DetectionabstractMicroblog has become a major platform for information about real-world events.Automatically discovering realworld events from microblog has attracted the attention of many researchers.However, most of existing work ignore the importance of emotion information for event detection.We argue that people's emotional reactions immediately reflect the occurring of real-world events and should be important for event detection.In this study, we focus on the problem of communityrelated event detection by community emotions.To address the problem, we propose a novel framework which include the following three key components: microblog emotion classification, community emotion aggregation and community emotion burst detection.We evaluate our approach on real microblog data sets.Experimental results demonstrate the effectiveness of the proposed framework. Gaoyan Ou, Wei Chen 0021, Tengjiao Wang 0003, Zhongyu Wei, Binyang Li, Dongqing Yang, Kam-Fai Wong |
EMNLP | 3 |
| 2014 | BF-Matrix: A Secondary Index for the Cloud Storage
Hongyan Li 0002, Yue Wang 0014, Tengjiao Wang 0003, Dongqing Yang |
WAIM | 4 |
| 2014 | Sarcasm Detection in Social Media Based on Imbalanced Classification
Wei Chen 0021, Gaoyan Ou, Tengjiao Wang 0003, Dongqing Yang, Kai Lei |
WAIM | 4 |
| 2014 | Finding Vacant Taxis Using Large Scale GPS Traces
Hongyan Li 0002, Shenda Hong, Yiyong Lin 0003, Nana Fan, Gaoyan Ou, Tengjiao Wang 0003, Lilue Fan |
WAIM | 7 |
| 2014 | Shortest Path Computing in Relational DBMSsabstractThis paper takes the shortest path discovery to study efficient relational approaches to graph search queries. We first abstract three enhanced relational operators, based on which we introduce an FEM framework to bridge the gap between relational operations and graph operations. We show new features introduced by recent SQL standards, such as window function and merge statement, can improve the performance of the FEM framework. Second, we propose an edge weight aware graph partitioning schema and design a bi-directional restrictive BFS (breadth-first-search)over partitioned tables, which improves the scalability and performance without extra indexing overheads. The final extensive experimental results illustrate our relational approach with optimization strategies can achieve high scalability and performance. Jun Gao 0003, Jiashuai Zhou, Jeffrey Xu Yu, Tengjiao Wang 0003 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2013 | Logistic Regression Bias Correction for Large Scale Data with Rare Events
Hongyan Li 0002, Hanchen Su, Gaoyan Ou, Tengjiao Wang 0003 |
ADMA (2) | 5 |
| 2013 | Massively parallel learning of Bayesian networks with MapReduce for factor relationship analysisabstractBayesian Network (BN) is one of the most popular models in data mining technologies. Most of the algorithms of BN structure learning are developed for the centralized datasets, where all the data are gathered into a single computer node. They are often too costly or impractical for learning BN structures from large scale data. Through a simple interface with two functions, map and reduce, MapReduce facilitates parallel implementation of many real-world tasks such as data processing for search engines and machine learning. In this paper, we present a parallel algorithm for BN structure leaning from large-scale dateset by using a MapReduce cluster. We discuss the benefits of using MapReduce for BN structure learning, and demonstrate the performance of this approach by applying it to a real world financial factor relationships learning task from the domain of financial analysis. Wei Chen 0021, Tengjiao Wang 0003, Dongqing Yang, Kai Lei, Yueqin Liu |
IJCNN | 2 |
| 2013 | Aspect-Specific Polarity-Aware Summarization of Online Reviews
Gaoyan Ou, Wei Chen 0021, Tengjiao Wang 0003, Dongqing Yang, Kai Lei, Yueqin Liu |
WAIM | 4 |
| 2013 | Influence maximization with limit cost in social network
Yue Wang 0014, Weijing Huang, Lang Zong, Tengjiao Wang 0003, Dongqing Yang |
Sci. China Inf. Sci. | 4 |
| 2013 | Outsourcing shortest distance computing with privacy protection
Jun Gao 0003, Jeffrey Xu Yu, Ruoming Jin, Jiashuai Zhou, Tengjiao Wang 0003, Dongqing Yang |
VLDB J. | 5 |
| 2012 | MBA: A market-based approach to data allocation and dynamic migration for cloud database
Tengjiao Wang 0003, Ziyu Lin, Bishan Yang, Jun Gao 0003, Allen Huang, Dongqing Yang, Shiwei Tang, Jinzhong Niu |
Sci. China Inf. Sci. | 1 |
| 2012 | Holistic Top-k Simple Shortest Path Join in GraphsabstractMotivated by the needs such as group relationship analysis, this paper introduces a new operation on graphs, named top-k path join, which discovers the top-k simple shortest paths between two given node sets. Rather than discovering the top-k simple paths between each node pair, this paper proposes a holistic join method which answers the top-k path join by finding constrained top-k simple shortest paths between two nodes, and then devises an efficient method to handle the latter problem. Specifically, we transform the graph by encoding the precomputed shortest paths to the target node, and use the transformed graph in the candidate path searching. We show that the candidate path searching on the transformed graph not only has the same result as that on the original graph but also can be terminated much earlier with the aid of precomputed results. We also discuss two other optimization strategies, including considering the join constraint in the candidate path generation as early as possible, and pruning search space in each candidate path generation with an adaptively determined threshold. The final extensive experimental results also show that our method offers a significant performance improvement over existing ones. Jun Gao 0003, Jeffrey Xu Yu, Huida Qiu, Tengjiao Wang 0003, Dongqing Yang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2011 | Efficient Subject-Oriented Evaluating and Mining Methods for Data with Schema Uncertainty
Yue Wang 0014, Changjie Tang, Tengjiao Wang 0003, Dongqing Yang |
ADMA (1) | 3 |
| 2011 | Preface to the 2nd International Workshop on Unstructured Data Management (USDM 2011)
Tengjiao Wang 0003 |
APWeb | 1 |
| 2011 | Neighborhood-privacy protected shortest distance computing in cloudabstractWith the advent of cloud computing, it becomes desirable to utilize cloud computing to efficiently process complex operations on large graphs without compromising their sensitive information. This paper studies shortest distance computing in the cloud, which aims at the following goals: i) preventing outsourced graphs from neighborhood attack, ii) preserving shortest distances in outsourced graphs, iii) minimizing overhead on the client side. The basic idea of this paper is to transform an original graph G into a link graph Gl kept locally and a set of outsourced graphs Go. Each outsourced graph should meet the requirement of a new security model called 1-neighborhood-d-radius. In addition, the shortest distance query can be answered using Gl and Go. Our objective is to minimize the space cost on the client side when both security and utility requirements are satisfied. We devise a greedy method to produce Gl and Go, which can exactly answer the shortest distance queries. We also develop an efficient transformation method to support approximate shortest distance answering under a given additive error bound. The final experimental results illustrate the effectiveness and efficiency of our method. Jun Gao 0003, Jeffrey Xu Yu, Ruoming Jin, Jiashuai Zhou, Tengjiao Wang 0003, Dongqing Yang |
SIGMOD Conference | 5 |
| 2011 | Informed Prediction with Incremental Core-Based Friend Cycle Discovering
Yue Wang 0014, Weijing Huang, Wei Chen 0021, Tengjiao Wang 0003, Dongqing Yang |
WAIM | 4 |
| 2011 | Relational Approach for Shortest Path Discovery over Large GraphsabstractWith the rapid growth of large graphs, we cannot assume that graphs can still be fully loaded into memory, thus the disk-based graph operation is inevitable. In this paper, we take the shortest path discovery as an example to investigate the technique issues when leveraging existing infrastructure of relational database (RDB) in the graph data management. Based on the observation that a variety of graph search queries can be implemented by iterative operations including selecting frontier nodes from visited nodes, making expansion from the selected frontier nodes, and merging the expanded nodes into the visited ones, we introduce a relational FEM framework with three corresponding operators to implement graph search tasks in the RDB context. We show new features such as window function and merge statement introduced by recent SQL standards can not only simplify the expression but also improve the performance of the FEM framework. In addition, we propose two optimization strategies specific to shortest path discovery inside the FEM framework. First, we take a bi-directional set Dijkstra's algorithm in the path finding. The bi-directional strategy can reduce the search space, and set Dijkstra's algorithm finds the shortest path in a set-at-a-time fashion. Second, we introduce an index named SegTable to preserve the local shortest segments, and exploit SegTable to further improve the performance. The final extensive experimental results illustrate our relational approach with the optimization strategies achieves high scalability and performance. Jun Gao 0003, Ruoming Jin, Jiashuai Zhou, Jeffrey Xu Yu, Tengjiao Wang 0003 |
Proc. VLDB Endow. | 6 |
| 2010 | A General Multi-relational Classification Approach Using Feature Generation and Selection
Miao Zou, Tengjiao Wang 0003, Hongyan Li 0002, Dongqing Yang |
ADMA (2) | 2 |
| 2010 | A Heuristic Method for Unstructured Pattern Management over Data StreamsabstractPattern management is an important task in data stream mining and has attracted increasing attention recently. Variations of data stream patterns typically imply some fundamental changes of underlying objects and possess significant domain meanings. Many database applications require investigating the history information to get the knowledge about the evolving process of data streams. However, in most circumstances, the data stream patterns are unstructured: limited memory space cannot record all the patterns discovered online, no training sets or predefined models are available, and large numbers of noises bring another non-trivial challenge. This paper presents our research effort in online pattern management over such streams. A novel algorithm is proposed to detect stream changes, organize meaningful patterns and distinguish useful variations from noises. It extracts new trends from unstructured data heuristically, and involves a special parameter to identify whether the current event should be treated as significant. Several experiments are performed and the results prove this new method feasible and efficient. Gaoshan Miao, Hongyan Li 0002, Tengjiao Wang 0003 |
APWeb | 3 |
| 2010 | Fast top-k simple shortest paths discovery in graphsabstractWith the wide applications of large scale graph data such as social networks, the problem of finding the top-k shortest paths attracts increasing attention. This paper focuses on the discovery of the top-k simple shortest paths (paths without loops). The well known algorithm for this problem is due to Yen, and the provided worstcase bound O(kn(m + nlogn)), which comes from O(n) times single-source shortest path discovery for each of k shortest paths, remains unbeaten for 30 years, where n is the number of nodes and m is the number of edges. In this paper, we observe that there are shared sub-paths among O(kn) single-source shortest paths. The basic idea behind our method is to pre-compute the shortest paths to the target node, and utilize them to reduce the discovery cost at running time. Specifically, we transform the original graph by encoding the pre-computed paths, and prove that the shortest path discovered over the transformed graph is equivalent to that in the original graph. Most importantly, the path discovery over the transformed graph can be terminated much earlier than before. In addition, two optimization strategies are presented. One is to reduce the total iteration times for shortest path discovery, and the other is to prune the search space in each iteration with an adaptively-determined threshold. Although the worst-case complexity cannot be lowered, our method is proven to be much more efficient in a general case. The final extensive experimental results (on both real and synthetic graphs) also show that our method offers a significant performance improvement over the existing ones. Jun Gao 0003, Huida Qiu, Tengjiao Wang 0003, Dongqing Yang |
CIKM | 4 |
| 2010 | Multiple Sensitive Association Protection in the Outsourced Database
Jun Gao 0003, Tengjiao Wang 0003, Dongqing Yang |
DASFAA (2) | 3 |
| 2010 | Efficient evaluation of query rewriting plan over materialized XML view
Jun Gao 0003, Jiaheng Lu, Tengjiao Wang 0003, Dongqing Yang |
J. Syst. Softw. | 3 |
| 2009 | Effective multi-label active learning for text classificationabstractLabeling text data is quite time-consuming but essential for automatic text classification. Especially, manually creating multiple labels for each document may become impractical when a very large amount of data is needed for training multi-label text classifiers. To minimize the human-labeling efforts, we propose a novel multi-label active learning approach which can reduce the required labeled data without sacrificing the classification accuracy. Traditional active learning algorithms can only handle single-label problems, that is, each data is restricted to have one label. Our approach takes into account the multi-label information, and select the unlabeled data which can lead to the largest reduction of the expected model loss. Specifically, the model loss is approximated by the size of version space, and the reduction rate of the size of version space is optimized with Support Vector Machines (SVM). An effective label prediction method is designed to predict possible labels for each unlabeled data point, and the expected loss for multi-label data is approximated by summing up losses on all labels according to the most confident result of label prediction. Experiments on several real-world data sets (all are publicly available) demonstrate that our approach can obtain promising classification result with much fewer labeled data than state-of-the-art methods. Bishan Yang, Jian-Tao Sun, Tengjiao Wang 0003, Zheng Chen 0001 |
KDD | 3 |
| 2009 | MobileMiner: a real world case study of data mining in mobile communicationabstractMobile communication data analysis has been often used as a background application to motivate many data mining problems. However, very few data mining researchers have a chance to see a working data mining system on real mobile communication data. In this demo, we showcase our new system MobileMiner on a real mobile communication data set, which presents a case study of business solutions using state-of-the-art data mining techniques. MobileMiner adaptively profiles users' behavior from their calling and moving record streams. Customer segmentation and social community analysis can be conducted based on user profiles. We show how data mining techniques can help in mobile communication data analysis. Moreover, we also show some interesting observations which still cannot be mined by the current techniques, and thus may motivate new research and development. Tengjiao Wang 0003, Bishan Yang, Jun Gao 0003, Dongqing Yang, Shiwei Tang, Kedong Liu, Jian Pei 0001 |
SIGMOD Conference | 1 |
| 2009 | Efficient algorithms for incremental maintenance of closed sequential patterns in large databases
Lei Chang, Tengjiao Wang 0003, Dongqing Yang, Hua Luan, Shiwei Tang |
Data Knowl. Eng. | 2 |
| 2008 | Effective Data Distribution and Reallocation Strategies for Fast Query Response in Distributed Query-Intensive Data Environments
Tengjiao Wang 0003, Bishan Yang, Jun Gao 0003, Dongqing Yang |
APWeb | 1 |
| 2008 | Discovery of Frequent Query Patterns in XML Pattern Graph with DTD Cardinality ConstraintsabstractCommon query patterns of multiple XML queries can be stored and shared to accelerate the query execution efficiently. Such common patterns typically arise in many applications. In this paper we present a new technique for efficiently mining frequent XML query patterns in the XML pattern graph with DTD cardinality constraints. First we propose a new method of finding all connected sub-graphs that appear frequently in a large XML query pattern graph. Then we ldquopushrdquo the DTD cardinality constraints deep into the mining process to prune the search space and still ensure the completeness of the answers. At last, we propose an algorithm FESG for effectively mining frequent XML query patterns in the XML pattern graph. We validate the effectiveness and efficiency of the new technique in two ways. The experimental results generated from the real data reveal that the algorithm works well in practice. Tengjiao Wang 0003 |
CISIS | 2 |
| 2008 | SeqStream: Mining Closed Sequential Patterns over Stream Sliding WindowsabstractPrevious studies have shown mining closed patterns provides more benefits than mining the complete set of frequent patterns, since closed pattern mining leads to more compact results and more efficient algorithms. It is quite useful in a data stream environment where memory and computation power are major concerns. This paper studies the problem of mining closed sequential patterns over data stream sliding windows. A synopsis structure IST (Inverse Closed Sequence Tree) is designed to keep inverse closed sequential patterns in current window. An efficient algorithm SeqStream is developed to mine closed sequential patterns in stream windows incrementally, and various novel strategies are adopted in SeqStream to prune search space aggressively. Extensive experiments on both real and synthetic data sets show that SeqStream outperforms PrefixSpan, CloSpan and BIDE by a factor of about one to two orders of magnitude. Lei Chang, Tengjiao Wang 0003, Dongqing Yang, Hua Luan |
ICDM | 2 |
| 2008 | Road Network Based Adaptive Query Evaluation in VANETabstractIn the Vehicle Ad-hoc NETwork (VANET), moving vehicles organize into a mobile wireless Ad-hoc network to share online traffic information. Each vehicle can issue a declarative query for aggregating the traffic information from others in order to facilitate the navigation and avoid traffic jam. Existing query methods suffer from high latency, incomplete results, and large messages due to the movement of the vehicles in VANET. In this paper, we propose an adaptive query evaluation method based on the road network. In order to overcome the problems incurred by the movement, a relative static query evaluation plan is constructed based on the road network, and each vehicle can participate in the query evaluation plan autonomously. We also introduce control messages to notify the changed location of the query originator to other vehicles involved in the evaluation plan. In addition, we propose an one time message transferring based results collecting method to reduce the message cost. The optimization over the multiple queries is also discussed to reduce the messages further. We evaluate the performance of our method by extensive simulations. Experimental results show that our method can provide complete results within a short response time and small traffic overhead. Jun Gao 0003, Jinsong Han, Dongqing Yang, Tengjiao Wang 0003 |
MDM | 4 |
| 2008 | BOAI: Fast Alternating Decision Tree Induction Based on Bottom-Up Evaluation
Bishan Yang, Tengjiao Wang 0003, Dongqing Yang, Lei Chang |
PAKDD | 2 |
| 2008 | A new similarity computing method based on concept similarity in Chinese text processing
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Jun Gao 0003 |
Sci. China Ser. F Inf. Sci. | 4 |
| 2008 | XFlat: Query-friendly encrypted XML view publishing
Jun Gao 0003, Tengjiao Wang 0003, Dongqing Yang |
Inf. Sci. | 2 |
| 2007 | MQTree Based Query Rewriting over Multiple XML Views
Jun Gao 0003, Tengjiao Wang 0003, Dongqing Yang |
DEXA | 2 |
| 2007 | A New Text Clustering Method Using Hidden Markov Model
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Aiqiang Gao |
NLDB | 4 |
| 2006 | Mining Compressed Sequential Patterns
Lei Chang, Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003 |
ADMA | 4 |
| 2006 | XFlat: Query Friendly Encrypted XML View Publishing
Jun Gao 0003, Tengjiao Wang 0003, Dongqing Yang |
APWeb | 2 |
| 2006 | CCWrapper: Adaptive Predefined Schema Guided Web Extraction
Jun Gao 0003, Dongqing Yang, Tengjiao Wang 0003 |
WAIM | 3 |
| 2006 | KCAM: Concentrating on Structural Similarity for XML Fragments
Lingbo Kong, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Jun Gao 0003 |
WAIM | 4 |
| 2006 | Cardinality Computing: A New Step Towards Fully Representing Multi-sets by Bloom Filters
Jiakui Zhao, Dongqing Yang, Lijun Chen 0002, Jun Gao 0003, Tengjiao Wang 0003 |
WISE | 5 |
| 2005 | Validating key constraints over XML document using XPath and structure checking
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Jun Gao 0003 |
Future Gener. Comput. Syst. | 4 |
| 2004 | WIEAS: Helping to Discover Web Information Sources and Extract Data from Them
Liyu Li, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Zhi-Hong Deng 0001, Zhihua Su |
APWeb | 4 |
| 2004 | EGA: An Algorithm for Automatic Semi-structured Web Documents Extraction
Liyu Li, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Zhihua Su |
DASFAA | 4 |
| 2004 | Incremental Maintenance of Discovered Mobile User Maximal Moving Sequential Patterns
Shuai Ma 0001, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Chanjun Yang |
DASFAA | 4 |
| 2004 | Discovering and Generating Materialized XML Views in Data Integration Systems
Tengjiao Wang 0003, Dongqing Yang, Shiwei Tang |
IDEAS | 1 |
| 2004 | Combining Clustering with Moving Sequential Pattern Mining: A Novel and Efficient Technique
Shuai Ma 0001, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Jinqiang Han |
PAKDD | 4 |
| 2004 | QReduction: Synopsizing XPath Query Set Efficiently under Resource Constraint
Jun Gao 0003, Xiuli Ma, Dongqing Yang, Tengjiao Wang 0003, Shiwei Tang |
WAIM | 4 |
| 2004 | Extracting Key Value and Checking Structural Constraints for Validating XML Key Constraints
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Jun Gao 0003 |
WAIM | 4 |
| 2004 | Discovery of Frequent XML Query Patterns with DTD Cardinality Constraints
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Jun Gao 0003 |
WAIM | 4 |
| 2003 | A New Fast Clustering Algorithm Based on Reference and Density
Shuai Ma 0001, Tengjiao Wang 0003, Shiwei Tang, Dongqing Yang, Jun Gao 0003 |
WAIM | 2 |
| 2002 | COMMIX: towards effective web information extraction, integration and query answeringabstractAs WWW becomes more and more popular and powerful, how to search information on the web in database way becomes an important research topic. COMMIX, which is developed in the DB group in Peking University (China), is a system towards building very large database using data from the Web for information extraction, integration and query answering. COMMIX has some innovative features, such as ontology-based wrapper generation, XML-based information integration, view-based query answering, and QBE-style XML query interface. Tengjiao Wang 0003, Shiwei Tang, Dongqing Yang, Jun Gao 0003, Yuqing Wu, Jian Pei 0001 |
SIGMOD Conference | 1 |
| 2001 | Extracting Local Schema from Semistructured Data Based on Graph-Oriented Semantic Model
Tengjiao Wang 0003, Shiwei Tang, Dongqing Yang |
J. Comput. Sci. Technol. | 1 |