VLDB 2026 Research / reviewers in the wild / expert
Tianlei Hu
dblp:02/3803
· DBLP profile ↗
40ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0003-0744-6454ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 since 2021Artificial intelligence and machine learning · 14 · 5 since 2021Databases, data management, data science and information retrieval · 13 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ACT: A Unified Framework for Rigging and Animating Characters with Arbitrary TopologiesabstractRecent advances in generative models have democratized the creation of high-quality static 3D assets, yet animating these meshes remains a labor-intensive bottleneck. Traditional pipelines fracture this process into sequential stages—rigging, skinning, and motion synthesis—ignoring the inherent coupling between morphological structure and motor function. To bridge this gap, we introduce ACT, a unified generative framework that reformulates rigging and animation not as independent tasks, but as complementary views of a single hyper-kinematic process. Our key insight is to model the joint distribution of skeletal topology and temporal motion within a shared latent space. ACT utilizes a Vision Language Model (VLM) to extract semantic topological priors from arbitrary meshes, which then condition a Diffusion Transformer (DiT) backbone. By treating static rest poses and dynamic trajectories as a unified sequence, our model employs a task-aware masking strategy to flexibly perform zero-shot rigging, text-guided motion generation, and motion completion within a single end-to-end architecture. Furthermore, a geometry-guided decoder ensures that surface deformations are tightly coupled with the generated kinematics. Extensive experiments demonstrate that ACT generalizes robustly to diverse, non-humanoid characters without retraining. By replacing brittle cascaded pipelines with a holistic prior, our method enables novel applications such as semantic-driven topology editing and generative in-betweening, offering a versatile and efficient solution for automating 3D character animation. Pengyu Long, Weirui Wang, Qingcheng Zhao, Qixuan Zhang, Jiaqing Zhou, Tianlei Hu, Wei Yang 0034, Lan Xu 0003, Jingyi Yu 0001 |
ACM Trans. Graph. | 8 |
| 2025 | Multi-Instance Multi-Label Classification from Crowdsourced LabelsabstractMulti-instance multi-label classification (MIML) is a fundamental task in machine learning, where each data sample comprises a bag containing several instances and multiple binary labels. Despite its wide applications, the data collection process involves matching multiple instances and labels, typically resulting in high annotation costs. In this paper, we study a novel yet practical crowdsourced multi-instance multi-label classification (CMIML) setup, where labels are collected from multiple crowd sources. To address this problem, we first propose a novel data generation process for CMIML, i.e., cross-label transition, where cross-label annotation error is more likely to appear rather than previous single-label transition assumption, due to the inherent similarity of localized instances from different classes. Then, we formally define the cross-label transition by cross-label transition matrices which are dependent across classes. Subsequently, we establish the first unbiased risk estimator for CMIML and further improve it through aggregation techniques, along with a rigorous generalization error bound. We also provide a practical implementation of cross-label transition matrix estimation. Comprehensive experiments on six benchmark datasets under various scenarios demonstrate that our algorithm outperforms the baselines by a large margin, validating its effectiveness in handling the CMIML problem. Ziquan Wang, Mingxuan Xia, Jiaqing Zhou, Gengyu Lyu, Tianlei Hu, Haobo Wang 0001 |
AAAI | 6 |
| 2025 | Ensembling Prompting Strategies for Zero-Shot Hierarchical Text Classification with Large Language ModelsabstractHierarchical text classification aims to classify documents into multiple labels within a hierarchical taxonomy, making it an essential yet challenging task in natural language processing.Recently, using Large Language Models (LLM) to tackle hierarchical text classification in a zero-shot manner has attracted increasing attention due to their cost-efficiency and flexibility.Given the challenges of understanding the hierarchy, various HTC prompting strategies have been explored to elicit the best performance from LLMs.However, our empirical study reveals that LLMs are highly sensitive to these prompting strategies-(i) within a task, different strategies yield substantially different results, and (ii) across various tasks, the relative effectiveness of a given strategy varies significantly.To address this, we propose a novel ensemble method, HiEPS, which integrates the results of diverse prompting strategies to promote LLMs' reliability.We also introduce a path-valid voting mechanism for ensembling, which selects a valid result with the highest path frequency score.Extensive experiments on three benchmark datasets show that HiEPS boosts the performance of single prompting strategies and achieves SOTA results.The source code is available at https: //github.com/MingxuanXia/HiEPS. Mingxuan Xia, Zhijie Jiang, Haobo Wang 0001, Junbo Zhao 0002, Tianlei Hu, Gang Chen 0001 |
EMNLP | 5 |
| 2025 | A Timestep-Adaptive Frequency-Enhancement Framework for Diffusion-based Image Super-ResolutionabstractImage super-resolution (ISR) is a classic and challenging problem in computer vision because of complex and unknown degradation patterns in the data collection process. Leveraging powerful generative priors, diffusion-based methods have recently established new state-of-the-art ISR performance, but their characteristics in the frequency domain are still underexplored. In this paper, we innovatively investigate their frequency-domain behaviors from a sampling timestep perspective. Experimentally, we find that current diffusion-based ISR algorithms exhibit insufficiency in different frequency components in distinct groups of timesteps during the sampling. To address this, we first propose a Timestep Division Controller that is able to adaptively divide the timesteps into groups based on the performance gradient across different components. Next, we design two dedicated modules --- the Amplitude and Phase Enhancement Module (APEM) and the High- and Low-Frequency Enhancement Module (HLEM), to regulate the information flow of distinct frequency-domain features. By adaptively enhancing specific frequency components at different stages of the sampling process, the two modules effectively compensate for the insufficient frequency-domain perception of diffusion-based ISR models. Extensive experiments on three benchmark datasets verify the superior ISR performance of our method, e.g., achieving an average 5.40% improvement on CLIP-IQA compared to the best diffusion-based ISR baseline. Hanbin Zhao, Jiaqing Zhou, Guozhi Xu, Tianlei Hu, Gang Chen 0001, Haobo Wang 0001 |
IJCAI | 5 |
| 2025 | GMCoT: a graph-augmented multimodal chain-of-thought reasoning framework for multi-label zero-shot learningabstractIn recent years, multi-label zero-shot learning (ML-ZSL) has garnered increasing attention because of its wide range of potential applications, such as image annotation, text classification, and bioinformatics. The central challenge in ML-ZSL lies in predicting multiple labels for unseen classes without requiring any labeled training data, which contrasts with conventional supervised learning paradigms. However, existing methods face several significant challenges. These include the substantial semantic gap between different modalities, which impedes effective knowledge transfer, and the intricate and typically complex relationships among multiple labels, making it difficult to model them in a meaningful and accurate manner. To overcome these challenges, we propose a graph-augmented multimodal chain-of-thought (GMCoT) reasoning approach. The proposed method combines the strengths of multimodal large language models with graph-based structures, significantly enhancing the reasoning process involved in multi-label prediction. First, a novel multimodal chain-of-thought reasoning framework is presented which imitates human-like step-by-step reasoning to produce multi-label predictions. Second, a technique is presented for integrating label graphs into the reasoning process. This technique enables the capture of complex semantic relationships among labels, thereby improving the accuracy and consistency of multi-label generation. Comprehensive experiments on benchmark datasets demonstrate that the proposed GMCoT approach outperforms state-of-the-art methods in ML-ZSL. Haobo Wang 0001, Ke Chen 0005, Tianlei Hu, Gang Chen 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2024 | A Separation and Alignment Framework for Black-Box Domain AdaptationabstractBlack-box domain adaptation (BDA) targets to learn a classifier on an unsupervised target domain while assuming only access to black-box predictors trained from unseen source data. Although a few BDA approaches have demonstrated promise by manipulating the transferred labels, they largely overlook the rich underlying structure in the target domain. To address this problem, we introduce a novel separation and alignment framework for BDA. Firstly, we locate those well-adapted samples via loss ranking and a flexible confidence-thresholding procedure. Then, we introduce a novel graph contrastive learning objective that aligns under-adapted samples to their local neighbors and well-adapted samples. Lastly, the adaptation is finally achieved by a nearest-centroid-augmented objective that exploits the clustering effect in the feature space. Extensive experiments demonstrate that our proposed method outperforms best baselines on benchmark datasets, e.g. improving the averaged per-class accuracy by 4.1% on the VisDA dataset. The source code is available at: https://github.com/MingxuanXia/SEAL. Mingxuan Xia, Junbo Zhao 0002, Gengyu Lyu, Zenan Huang, Tianlei Hu, Gang Chen 0001, Haobo Wang 0001 |
AAAI | 5 |
| 2024 | On the Value of Head Labels in Multi-Label Text ClassificationabstractA formidable challenge in the multi-label text classification (MLTC) context is that the labels often exhibit a long-tailed distribution, which typically prevents deep MLTC models from obtaining satisfactory performance. To alleviate this problem, most existing solutions attempt to improve tail performance by means of sampling or introducing extra knowledge. Data-rich labels, though more trustworthy, have not received the attention they deserve. In this work, we propose a multiple-stage training framework to exploit both model- and feature-level knowledge from the head labels, to improve both the representation and generalization ability of MLTC models. Moreover, we theoretically prove the superiority of our framework design over other alternatives. Comprehensive experiments on widely used MLTC datasets clearly demonstrate that the proposed framework achieves highly superior results to state-of-the-art methods, highlighting the value of head labels in MLTC. Haobo Wang 0001, Cheng Peng 0011, Hede Dong, Lei Feng 0006, Weiwei Liu 0003, Tianlei Hu, Ke Chen 0005, Gang Chen 0001 |
ACM Trans. Knowl. Discov. Data | 6 |
| 2023 | Deep Partial Multi-Label Learning with Graph DisambiguationabstractIn partial multi-label learning (PML), each data example is equipped with a candidate label set, which consists of multiple ground-truth labels and other false-positive labels. Recently, graph-based methods, which demonstrate a good ability to estimate accurate confidence scores from candidate labels, have been prevalent to deal with PML problems. However, we observe that existing graph-based PML methods typically adopt linear multi-label classifiers and thus fail to achieve superior performance. In this work, we attempt to remove several obstacles for extending them to deep models and propose a novel deep Partial multi-Label model with grAph-disambIguatioN (PLAIN). Specifically, we introduce the instance-level and label-level similarities to recover label confidences as well as exploit label dependencies. At each training epoch, labels are propagated on the instance and label graphs to produce relatively accurate pseudo-labels; then, we train the deep model to fit the numerical labels. Moreover, we provide a careful analysis of the risk functions to guarantee the robustness of the proposed model. Extensive experiments on various synthetic datasets and three real-world PML datasets demonstrate that PLAIN achieves significantly superior results to state-of-the-art methods. Haobo Wang 0001, Shisong Yang, Gengyu Lyu, Weiwei Liu 0003, Tianlei Hu, Ke Chen 0005, Songhe Feng, Gang Chen 0001 |
IJCAI | 5 |
| 2023 | Uncertainty-aware complementary label queries for active learningabstractIn this paper, we tackle the problem of ALCL (Liu et al., 2023). The objective of ALCL is to directly reduce the cost of annotation actions in AL, while providing a feasible approach for obtaining complementary labels. To solve ALCL, we design a sampling strategy USD, which uses the uncertainty in deep learning to guide the queries of active learning in this novel setup. Moreover, we upgrade the WEBB method to suit this sampling strategy. Comprehensive experimental results validate the performance of our proposed approaches. In the future, we plan to investigate the applicability of our approaches to large-scale datasets and account for noise in the feedback of annotators. Ke Chen 0005, Tianlei Hu, Yunqing Mao |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2023 | Decision Boundary-Aware Data Augmentation for Adversarial TrainingabstractAdversarial training (AT) is a typical method to learn adversarially robust deep neural networks via training on the adversarial variants generated by their natural examples. However, as training progresses, the training data becomes less attackable, which may undermine the enhancement of model robustness. A straightforward remedy is to incorporate more training data, but it may incur an unaffordable cost. To mitigate this issue, in this paper, we propose a deCisiOn bounDary-aware data Augmentation framework (CODA): in each epoch, the CODA directly employs the meta information of the previous epoch to guide the augmentation process and generate more data that are close to the decision boundary, i.e., attackable data. Compared with the vanilla mixup, our proposed CODA can provide a higher ratio of attackable data, which is beneficial to enhance model robustness; it meanwhile mitigates the model's linear behavior between classes, where the linear behavior is favorable to the standard training for generalization but not to the adversarial training for robustness. As a result, our proposed CODA encourages the model to predict invariantly in the cluster of each class. Experiments demonstrate that our proposed CODA can indeed enhance adversarial robustness across various adversarial training methods and multiple datasets. Chen Chen 0043, Jingfeng Zhang, Xilie Xu, Lingjuan Lyu, Chaochao Chen 0001, Tianlei Hu, Gang Chen 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2023 | Multi-Source Multi-Label Learning for User Profiling in Online GamesabstractIn online games, user profiling plays a vital role in a variety of personalized services. Current solutions typically treat different dimensions or labels (e.g., willing to pay or not, high, medium, or low appetite for some gameplays) of the full user profiles as independent multi-class/binary classification tasks. However, such one-by-one profiling strategy clearly overlooks the implicitly correlations among profiling tasks, which results in a degraded performance. To cope with this issue, we make the first attempt to formalize this problem as a multi-label learning task. Accordingly, we develop a unified Multi-Source Multi-Label learning framework~(MSML) that well utilizes semantically rich features and labels for boosted user profiling in online games. Specifically, we first introduce a multi-source user representation network that exploits multi-source data in online games to obtain informative user representations. Subsequently, to handle multiple labels, we propose a novel embedding-based multi-label network that consists of two variational autoencoders with disentangled latent spaces. Note that our framework can guarantee the consistency of the training and testing phases by a novel dual-tower design to overcome the limitation of existing approaches that use one coupled decoder for both features and labels. Extensive experiments on six public multi-label datasets and one real-world online game dataset from Justice demonstrate that the proposed framework outperforms the state-of-the-art baseline methods. Moreover, our proposed framework has been successfully deployed in several online games, yielding a significant boost in multi-label user profiling. Haobo Wang 0001, Runze Wu 0001, Manhu Qu, Tianlei Hu, Gang Chen 0001, Jianrong Tao, Changjie Fan |
IEEE Trans. Multim. | 6 |
| 2020 | Incorporating Label Embedding and Feature Augmentation for Multi-Dimensional ClassificationabstractFeature augmentation, which manipulates the feature space by integrating the label information, is one of the most popular strategies for solving Multi-Dimensional Classification (MDC) problems. However, the vanilla feature augmentation approaches fail to consider the intra-class exclusiveness, and may achieve degenerated performance. To fill this gap, a novel neural network based model is proposed which seamlessly integrates the Label Embedding and Feature Augmentation (LEFA) techniques to learn label correlations. Specifically, based on attentional factorization machine, a cross correlation aware network is introduced to learn a low-dimensional label representation that simultaneously depicts the inter-class correlations and the intra-class exclusiveness. Then the learned latent label vector can be used to augment the original feature space. Extensive experiments on seven real-world datasets demonstrate the superiority of LEFA over state-of-the-art MDC approaches. Haobo Wang 0001, Chen Chen 0043, Weiwei Liu 0003, Ke Chen 0005, Tianlei Hu, Gang Chen 0001 |
AAAI | 5 |
| 2020 | Learning From Multi-Dimensional Partial LabelsabstractMulti-dimensional classification has attracted huge attention from the community. Though most studies consider fully annotated data, in real practice obtaining fully labeled data in MDC tasks is usually intractable. In this paper, we propose a novel learning paradigm: MultiDimensional Partial Label Learning (MDPL) where the ground-truth labels of each instance are concealed in multiple candidate label sets. We first introduce the partial hamming loss for MDPL that incurs a large loss if the predicted labels are not in candidate label sets, and provide an empirical risk minimization (ERM) framework. Theoretically, we rigorously prove the conditions for ERM learnability of MDPL in both independent and dependent cases. Furthermore, we present two MDPL algorithms under our proposed ERM framework. Comprehensive experiments on both synthetic and real-world datasets validate the effectiveness of our proposals. Haobo Wang 0001, Weiwei Liu 0003, Yang Zhao 0022, Tianlei Hu, Ke Chen 0005, Gang Chen 0001 |
IJCAI | 4 |
| 2020 | Collaboration Based Multi-Label Propagation for Fraud DetectionabstractDetecting fraud users, who fraudulently promote certain target items, is a challenging issue faced by e-commerce platforms. Generally, many fraud users have different spam behaviors simultaneously, e.g. spam transactions, clicks, reviews and so on. Existing solutions have two main limitations: 1) the correlations among multiple spam behaviors are neglected; 2) large-scale computations are intractable when dealing with an enormous user set. To remedy these problems, this work proposes a collaboration based multi-label propagation (CMLP) algorithm. We first introduce a general-purpose version that involves collaboration technique to exploit label correlations. Specifically, it breaks the final prediction into two parts: 1) its own prediction part; 2) the prediction of others, i.e. collaborative part. Then, to accelerate it on large-scale e-commerce data, we propose a heterogeneous graph based variant that detects communities on the user-item graph directly. Both theoretical analysis and empirical results clearly validate the effectiveness and scalability of our proposals. Haobo Wang 0001, Zhao Li 0007, Pengrui Hui, Weiwei Liu 0003, Tianlei Hu, Gang Chen 0001 |
IJCAI | 6 |
| 2020 | Off-Policy Recommendation System Without Exploration
Chengwei Wang, Chen Chen 0043, Tianlei Hu, Gang Chen 0001 |
PAKDD (1) | 4 |
| 2020 | Online Partial Label Learning
Haobo Wang 0001, Yuzhou Qiang, Chen Chen 0043, Weiwei Liu 0003, Tianlei Hu, Zhao Li 0007, Gang Chen 0001 |
ECML/PKDD (2) | 5 |
| 2020 | HAM: a deep collaborative ranking method incorporating textual informationabstractThe recommendation task with a textual corpus aims to model customer preferences from both user feedback and item textual descriptions. It is highly desirable to explore a very deep neural network to capture the complicated nonlinear preferences. However, training a deeper recommender is not as effortless as simply adding layers. A deeper recommender suffers from the gradient vanishing/exploding issue and cannot be easily trained by gradient-based methods. Moreover, textual descriptions probably contain noisy word sequences. Directly extracting feature vectors from them can harm the recommender’s performance. To overcome these difficulties, we propose a new recommendation method named the HighwAy recoMmender (HAM). HAM explores a highway mechanism to make gradient-based training methods stable. A multi-head attention mechanism is devised to automatically denoise textual information. Moreover, a block coordinate descent method is devised to train a deep neural recommender. Empirical studies show that the proposed method outperforms state-of-the-art methods significantly in terms of accuracy. Chengwei Wang, Chen Chen 0043, Tianlei Hu, Gang Chen 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2019 | Two-Stage Label Embedding via Neural Factorization Machine for Multi-Label ClassificationabstractLabel embedding has been widely used as a method to exploit label dependency with dimension reduction in multilabel classification tasks. However, existing embedding methods intend to extract label correlations directly, and thus they might be easily trapped by complex label hierarchies. To tackle this issue, we propose a novel Two-Stage Label Embedding (TSLE) paradigm that involves Neural Factorization Machine (NFM) to jointly project features and labels into a latent space. In encoding phase, we introduce a Twin Encoding Network (TEN) that digs out pairwise feature and label interactions in the first stage and then efficiently learn higherorder correlations with deep neural networks (DNNs) in the second stage. After the codewords are obtained, a set of hidden layers is applied to recover the output labels in decoding phase. Moreover, we develop a novel learning model by leveraging a max margin encoding loss and a label-correlation aware decoding loss, and we adopt the mini-batch Adam to optimize our learning model. Lastly, we also provide a kernel insight to better understand our proposed TSLE. Extensive experiments on various real-world datasets demonstrate that our proposed model significantly outperforms other state-ofthe-art approaches. Chen Chen 0043, Haobo Wang 0001, Weiwei Liu 0003, Xingyuan Zhao, Tianlei Hu, Gang Chen 0001 |
AAAI | 5 |
| 2019 | CAMO: A Collaborative Ranking Method for Content Based RecommendationabstractIn real-world recommendation tasks, feedback data are usually sparse. Therefore, a recommender’s performance is often determined by how much information that it can extract from textual contents. However, current methods do not make full use of the semantic information. They encode the textual contents either by “bag-of-words” technique or Recurrent Neural Network (RNN). The former neglects the order of words while the latter ignores the fact that textual contents can contain multiple topics. Besides, there exists a dilemma in designing a recommender. On the one hand, we shall use a sophisticated model to exploit every drop of information in item contents; on the other hand, we shall adopt a simple model to prevent itself from over-fitting when facing the sparse feedbacks. To fill the gaps, we propose a recommender named CAMO 1. CAMO employs a multi-layer content encoder for simultaneously capturing the semantic information of multitopic and word order. Moreover, CAMO makes use of adversarial training to prevent the complex encoder from overfitting. Extensive empirical studies show that CAMO outperforms state-of-the-art methods in predicting users’ preferences. Chengwei Wang, Chen Chen 0043, Tianlei Hu, Gang Chen 0001 |
AAAI | 4 |
| 2019 | Discriminative and Correlative Partial Multi-Label LearningabstractIn partial label learning (PML), each instance is associated with a candidate label set that contains multiple relevant labels and other false positive labels. The most challenging issue for the PML is that the training procedure is prone to be affected by the labeling noise. We observe that state-of-the-art PML methods are either powerless to disambiguate the correct labels from the candidate labels or incapable of extracting the label correlations sufficiently. To fill this gap, a two-stage DiscRiminative and correlAtive partial Multi-label leArning (DRAMA) algorithm is presented in this work. In the first stage, a confidence value is learned for each label by utilizing the feature manifold, which indicates how likely a label is correct. In the second stage, a gradient boosting model is induced to fit the label confidences. Specifically, to explore the label correlations, we augment the feature space by the previously elicited labels on each boosting round. Extensive experiments on various real-world datasets clearly validate the superiority of our proposed method. Haobo Wang 0001, Weiwei Liu 0003, Yang Zhao 0022, Chen Zhang 0020, Tianlei Hu, Gang Chen 0001 |
IJCAI | 5 |
| 2018 | Runtime Shader Simplification via Instant Search in Reduced Optimization SpaceabstractAbstract Traditional automatic shader simplification simplifies shaders in an offline process, which is typically carried out in a context‐oblivious manner or with the use of some example contexts, e.g., certain hardware platforms, scenes, and uniform parameters, etc. As a result, these pre‐simplified shaders may fail at adapting to runtime changes of the rendering context that were not considered in the simplification process. In this paper, we propose a new automatic shader simplification technique, which explores two key aspects of a runtime simplification framework: the optimization space and the instant search for optimal simplified shaders with runtime context. The proposed technique still requires a preprocess stage to process the original shader. However, instead of directly computing optimal simplified shaders, the proposed preprocess generates a reduced shader optimization space. In particular, two heuristic estimates of the quality and performance of simplified shaders are presented to group similar variants into representative ones, which serve as basic graph nodes of the simplification dependency graph (SDG), a new representation of the optimization space. At the runtime simplification stage, a parallel discrete optimization algorithm is employed to instantly search in the SDG for optimal simplified shaders. New data‐driven cost models are proposed to predict the runtime quality and performance of simplified shaders on the basis of data collected during runtime. Results show that the selected simplifications of complex shaders achieve 1.6 to 2.5 times speedup and still retain high rendering quality. Yazhen Yuan, Rui Wang 0004, Tianlei Hu, Hujun Bao |
Comput. Graph. Forum | 3 |
| 2018 | Where Do Local Experts Go? Evaluating User Geo-Topical Similarity for Top-N Place Recommendation
Tianlei Hu, Gang Chen 0001 |
J. Comput. Sci. Technol. | 2 |
| 2016 | Real-time rendering on a power budgetabstractWith recent advances on mobile computing, power consumption has become a significant limiting constraint for many graphics applications. As a result, rendering on a power budget arises as an emerging demand. In this paper, we present a real-time, power-optimal rendering framework to address this problem, by finding the optimal rendering settings that minimize power consumption while maximizing visual quality. We first introduce a novel power-error, multi-objective cost space, and formally formulate power saving as an optimization problem. Then, we develop a two-step algorithm to efficiently explore the vast power-error space and leverage optimal Pareto frontiers at runtime. Finally, we show that our rendering framework can be generalized across different platforms, desktop PC or mobile device, by demonstrating its performance on our own OpenGL rendering framework, as well as the commercially available Unreal Engine. Rui Wang 0004, Julio Marco, Tianlei Hu, Diego Gutierrez, Hujun Bao |
ACM Trans. Graph. | 4 |
| 2016 | Adaptive matrix column sampling and completion for rendering participating mediaabstractSeveral scalable many-light rendering methods have been proposed recently for the efficient computation of global illumination. However, gathering contributions of virtual lights in participating media remains an inefficient and time-consuming task. In this paper, we present a novel sparse sampling and reconstruction method to accelerate the gathering step of the many-light rendering for participating media. Our technique explores the observation that the scattered lightings are usually locally coherent and of low rank even in heterogeneous media. In particular, we first introduce a matrix formation with light segments as columns and eye ray segments as rows, and formulate the gathering step into a matrix sampling and reconstruction problem. We then propose an adaptive matrix column sampling and completion algorithm to efficiently reconstruct the matrix by only sampling a small number of elements. Experimental results show that our approach greatly improves the performance, and obtains up to one order of magnitude speedup compared with other state-of-the-art methods of many-light rendering for participating media. Yuchi Huo, Rui Wang 0004, Tianlei Hu, Wei Hua 0002, Hujun Bao |
ACM Trans. Graph. | 3 |
| 2014 | BestPeer++: A Peer-to-Peer BasedLarge-Scale Data Processing PlatformabstractThe corporate network is often used for sharing information among the participating companies and facilitating collaboration in a certain industry sector where companies share a common interest. It can effectively help the companies to reduce their operational costs and increase the revenues. However, the inter-company data sharing and processing poses unique challenges to such a data management system including scalability, performance, throughput, and security. In this paper, we present BestPeer++, a system which delivers elastic data sharing services for corporate network applications in the cloud based on BestPeer - a peer-to-peer (P2P) based data management platform. By integrating cloud computing, database, and P2P technologies into one system, BestPeer++ provides an economical, flexible and scalable platform for corporate network applications and delivers data sharing services to participants based on the widely accepted pay-as-you-go business model. We evaluate BestPeer++ on Amazon EC2 Cloud platform. The benchmarking results show that BestPeer++ outperforms HadoopDB, a recently proposed large-scale data processing system, in performance when both systems are employed to handle typical corporate network workloads. The benchmarking results also demonstrate that BestPeer++ achieves near linear scalability for throughput with respect to the number of peer nodes. Gang Chen 0001, Tianlei Hu, Dawei Jiang, Peng Lu 0013, Kian-Lee Tan, Hoang Tam Vo, Sai Wu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2013 | Competence-based song recommendationabstractSinging is a popular social activity and a good way of expressing one's feelings. One important reason for unsuccessful singing performance is because the singer fails to choose a suitable song. In this paper, we propose a novel singing competence-based song recommendation framework. It is distinguished from most existing music recommendation systems which rely on the computation of listeners' interests or similarity. We model a singer's vocal competence as singer profile, which takes voice pitch, intensity, and quality into consideration. Then we propose techniques to acquire singer profiles. We also present a song profile model which is used to construct a human annotated song database. Finally, we propose a learning-to-rank scheme for recommending songs by singer profile. The experimental study on real singers demonstrates the effectiveness of our approach and its advantages over two baseline methods. To the best of our knowledge, our work is the first to study competence-based song recommendation. Lidan Shou, Kuang Mao, Xinyuan Luo, Ke Chen 0005, Gang Chen 0001, Tianlei Hu |
SIGIR | 6 |
| 2012 | BestPeer++: A Peer-to-Peer Based Large-Scale Data Processing PlatformabstractThe corporate network is often used for sharing information among the participating companies and facilitating collaboration in a certain industry sector where companies share a common interest. It can effectively help the companies to reduce their operational costs and increase the revenues. However, the inter-company data sharing and processing poses unique challenges to such a data management system including scalability, performance, throughput, and security. In this paper, we present Best Peer++, a system which delivers elastic data sharing services for corporate network applications in the cloud based on Best Peer -- a peer-to-peer (P2P) based data management platform. By integrating cloud computing, database, and P2P technologies into one system, Best Peer++ provides an economical, flexible and scalable platform for corporate network applications and delivers data sharing services to participants based on the widely accepted pay-as-you-go business model. We evaluate Best Peer++ on Amazon EC2 Cloud platform. The benchmarking results show that Best Peer++ outperforms Hadoop DB, a recently proposed large-scale data processing system, in performance when both systems are employed to handle typical corporate network workloads. The benchmarking results also demonstrate that Best Peer++ achieves near linear scalability for throughput with respect to the number of peer nodes. Gang Chen 0001, Tianlei Hu, Dawei Jiang, Peng Lu 0013, Kian-Lee Tan, Hoang Tam Vo, Sai Wu |
ICDE | 2 |
| 2010 | (k, P)-anonymity: towards pattern-preserving anonymity of time-series dataabstractThe challenges with privacy protection of time series are mainly due to the complex nature of the data and the queries performed on them. We study the anonymization of time series while trying to support complex queries, such as range and pattern similarity queries, on the published data. The conventional k-anonymity cannot effectively address this problem as it may suffer severe pattern loss. We propose a novel anonymization model called (k,P)-anonymity for pattern-rich time series. This model publishes both the attribute values and the patterns of time series in separate data forms. We demonstrate that our model can prevent linkage attacks on the published data while effectively support a wide variety of queries on the anonymized data. We also design an efficient algorithm for enforcing (k,P)-anonymity on time series data. Xuan Shang, Ke Chen 0005, Lidan Shou, Gang Chen 0001, Tianlei Hu |
CIKM | 5 |
| 2010 | Update Migration: An Efficient B+ Tree for Flash Storage
Lidan Shou, Gang Chen 0001, Tianlei Hu |
DASFAA (2) | 5 |
| 2010 | Towards Efficient Concurrent Scans on Flash Disks
Lidan Shou, Gang Chen 0001, Tianlei Hu, Ke Chen 0005 |
DEXA (1) | 5 |
| 2008 | Query Triggered Crawling Strategy: Build a Time Sensitive Vertical Search EngineabstractIn today's information society, it is important to retrieve fresh information. Many of the vertical search results are valid to users in only a short period of time. But due to resource constraints, it is not possible to keep the entire local storage synchronized with the Web. We implemented a TSVS (time sensitive vertical search engine) prototype named Velocisaurus focused on time-critical airfare discount information search to investigate the time critical requirements of vertical search and proposed a QTC (query triggered crawling) strategy to coordinate the crawling systems by real-time user queries and solve this problem. Experiment shows that QTC driven crawlers significantly improves the freshness of the search results and utilizes the resources more efficiently compared to regular search engine crawlers.. Lidan Shou, Tianlei Hu, Gang Chen 0001 |
CW | 3 |
| 2008 | Pivotbrowser: a tag-space image searching prototypeabstractWe propose a novel iterative searching and refining prototype for tagged images. This prototype, named PivotBrowser, captures semantically similar tag sets in a structure called pivot. By constructing a pivot for a textual query, PivotBrowser first selects candidate images possibly relevant to the query. The tags contained in these candidate images are then selected in terms of their tag relevances to the pivot. The shortlisted tags are clustered and one of the tag clusters is used to select the results from the candidate images. Ranking of the images in each partition is based on their relevance to the tag cluster. With the guidance of the tag clusters presented, a user is able to perform searching and iterative query refinement. Lidan Shou, Gang Chen 0001, Xiaolong Zhang 0008, Tianlei Hu, Jinxiang Dong |
WWW | 5 |
| 2008 | Modeling Image Data for Effective Indexing and Retrieval in Large General Image DatabasesabstractIn this paper, we propose an image semantic model based on the knowledge and criteria in the field of linguistics and taxonomy. Our work bridges the "semantic gap" by seamlessly exploiting the synergy of both visual feature processing and semantic relevance computation in a new way, and provides improved query efficiency and effectiveness for large general image databases. Our main contributions are as follows: we design novel data structures, namely, a lexical hierarchy, an image-semantic hierarchy, and a number of atomic semantic domains, to capture the semantics and the features of the database, and to provide the indexing scheme. We present a novel image query algorithm based on the proposed structures. In addition, we propose a novel term expansion mechanism to improve the lexical processing. Our extensive experiments indicate that our proposed techniques are effective in achieving high runtime performance with improved retrieval accuracy. The experiments also show that the proposed method has good scalability. Lidan Shou, Gang Chen 0001, Tianlei Hu, Jinxiang Dong |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2007 | A Caching System for XML Queries Using Frequent Query PatternsabstractIn this paper, we present an optimization framework for querying XML data by caching frequent query patterns. In this framework, frequent query patterns are mined online from user queries, and these query results are cached for future use. The mining process of frequent query patterns is launched automatically when user queries meet predefined requirements. To deal with queries that are similar to but not exactly same as the cached ones, a novel technique named query rewriting is adopted. This technique is able to handle four kinds of similar queries namely exact matching, exact containment, semantic matching and semantic containment. A cache replacement scheme that utilizes both the query pattern support and query pattern accessing time is employed to perform a fine-grained region purging. Experiments are carried out on the XMARK datasets. The results show that the proposed methods are both effective and efficient in improving the performance of XML queries. Yijun Bei, Gang Chen 0001, Tianlei Hu, Jinxiang Dong |
CSCWD | 3 |
| 2006 | Resilient Rights Protection for Product DataabstractWith the development of advance manufacture applications, enterprise must interact with their suppliers and customers to share more product data information. Consequently, copyright and integrity protection over product data is becoming an urgent requirement. Product data information consists of two parts: product structure and part content, besides protecting the three-dimensional (3D) model of part by watermarking, we also need to protect product structure information, which describes the relations among product parts and can be represented as a tree. In this paper, based on the 3D polygonal meshes watermark technology, the structure/semi-structure data watermark technologies and the idea of Fractal, we propose a novel watermark scheme for tree-structure product data with the value lying both in product structure and in part node content, which gives a more comprehensive and resilient right protection for both part node content and product structure. Moreover it is particularly effective when resisting invertibility attacks Ke Chen 0005, Gang Chen 0001, Tianlei Hu, Jinxiang Dong |
CSCWD | 3 |
| 2006 | Collaborative Agents Supported Automatic Physical Database Design Based on Description Logics ReasoningabstractPhysical database design is a key phase in devising and deploying a database system to improve overall system performance, and automatic physical database design becomes an important value-added feature in a database system to decrease the TCO (total cost of ownership). A GAPDD (general automatic physical database design) framework, with two phases, the workload mining phase and the design advising phase, is provided to support automatic physical database design for relational database systems. Three collaborative agents, the workload mining agent (WMA), the feature generation agent (FGA), and the optimal search agent (OSA) work together in GAPDD. The feature generation procedure in FGA is further studied, and a reasoner based on description logics (DL) is offered to reason useful features. An index advisor for OSCAR relational engine is prototyped, and experiment results show its feasibility and efficiency Tianlei Hu, Gang Chen 0001, Yin-Jie Hong, Xiaolong Zhang 0008, Jinxiang Dong |
CSCWD | 1 |
| 2005 | A logical replication-based intrusion detection approach for CSCW systemabstractWith the rapid development of CSCW techniques, CSCW system security has been of much attention. Its major challenge comes from malicious attacks, which cannot be handled by traditional security mechanisms, such as authorization, access control, encryption, etc. Although there existed some intrusion detection systems, current researches on intrusion detection are still insufficient in accuracy and efficiency for two reasons: firstly, most of them implement malicious transaction detection only by matching samples with predefined patterns; secondly, most current researches mainly resort to semantic analyzing or even only manual interventions to judge the malicious attacks. Inspired by the replication technique, we propose and implement a novel logical replication-based intrusion detection solution for CSCW system in this paper. It provides an additional layer of defense against application attacks, especially program attacks. The approach entitles the system to detect malicious intrusions efficiently at low false positive rate. Ke Chen 0005, Gang Chen 0001, Tianlei Hu, Jinxiang Dong |
CSCWD (2) | 3 |
| 2005 | An adaptive load balancing framework for parallel database systems based on collaborative agentsabstractLoad balancing is a key technique in parallel computer supported collaborative work (CSCW) systems, parallel database system and P2P system for instance, to boost performance and improve scalability. In order to reduce total cost of ownership (TCO), adaptive/self-tuning administration techniques are gradually and extensively expected in the cyberspace. In parallel database systems, adaptive load balancing techniques are proposed to face the change in data storage patterns and access patterns in a dynamic real environment. The techniques utilized in both shared-nothing and shared-disk parallel database systems are discussed, and a general flexible framework based on collaborative agents is studied to support these techniques in both architectures. The framework supports two kinds of load balancing - one is passively executing query statements balancedly, and the other one is proactively adjusting data placement and task execution scheme, by means of data and task migration, whenever load unbalance is detected. Three categories of agents, scheduling agents, monitoring agents and task agents, are identified in the framework. The collaboration protocols and scheduling algorithms to support adaptive load balancing are described. The framework also applies to other parallel systems such as P2P systems and shared file processing systems due to their underlying commonness. Tianlei Hu, Gang Chen 0001, Ke Chen 0005, Jinxiang Dong |
CSCWD (1) | 1 |
| 2005 | Watermarking Abstract Tree-Structured Data
Gang Chen 0001, Ke Chen 0005, Tianlei Hu, Jinxiang Dong |
WAIM | 3 |
| 2005 | GARWM: Towards a Generalized and Adaptive Watermark Scheme for Relational Data
Tianlei Hu, Gang Chen 0001, Ke Chen 0005, Jinxiang Dong |
WAIM | 1 |