VLDB 2026 Research / reviewers in the wild / expert
Tingting Wu 0007
dblp:20/815-7
· DBLP profile ↗
12ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0003-3437-8899ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Chain-of-Detection enables robust and efficient jailbreak defense
Tingting Wu 0007, Hao Zhang 0016 |
Neural Networks | 1 |
| 2026 | Model-Heterogeneous Federated Learning With Bidirectional Knowledge Distillation
Hao Zhang 0016, Yaolin Zhu, Tingting Wu 0007, Siyao Cheng, Jie Liu 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | CC-FedAvg: Computationally Customized Federated AveragingabstractFederated learning (FL) is an emerging paradigm to train model with distributed data from numerous Internet of Things (IoT) devices. It inherently assumes a uniform capacity among participants. However, due to different conditions such as differing energy budgets or executing parallel unrelated tasks, participants have diverse computational resources in practice. Participants with insufficient computation budgets must plan for the use of restricted computational resources appropriately; otherwise, they would be unable to complete the entire training procedure, resulting in model performance decline. To address this issue, we propose a strategy for estimating local models without computationally intensive iterations. Based on it, we propose computationally customized federated averaging (CC-FedAvg), which allows participants to determine whether to perform traditional local training or model estimation in each round based on their current computational budgets. Both theoretical analysis and exhaustive experiments indicate that CC-FedAvg has the same convergence rate and comparable performance as FedAvg without resource constraints. Furthermore, CC-FedAvg can be viewed as a computation-efficient version of FedAvg that retains model performance while considerably lowering computation overhead. Hao Zhang 0016, Tingting Wu 0007, Siyao Cheng, Jie Liu 0001 |
IEEE Internet Things J. | 2 |
| 2024 | Reasoning subevent relation over heterogeneous event graph
Tingting Wu 0007, Bing Qin 0001, Ting Liu 0001 |
Knowl. Inf. Syst. | 1 |
| 2024 | Uncertainty-guided label correction with wavelet-transformed discriminative representation enhancement
Tingting Wu 0007, Hao Zhang 0016, Minji Tang, Bing Qin 0001, Ting Liu 0001 |
Neural Networks | 1 |
| 2024 | DiscrimLoss: A Universal Loss for Hard Samples and Incorrect Samples DiscriminationabstractGiven data with label noise (i.e., incorrect data), deep neural networks would gradually memorize the label noise and impair model performance. To relieve this issue, curriculum learning is proposed to improve model performance and generalization by ordering training samples in a meaningful (e.g., easy to hard) sequence. Previous work takes incorrect samples as generic hard ones without discriminating between hard samples (i.e., hard samples in correct data) and incorrect samples. Indeed, a model should learn from hard samples to promote generalization rather than overfit to incorrect ones. In this article, we address this problem by appending a novel loss functionDiscrimLoss, on top of the existing task loss. Its main effect is to automatically and stably estimate the importance of easy samples and difficult samples (including hard and incorrect samples) at the early stages of training to improve the model performance. Then, during the following stages, DiscrimLoss is dedicated to discriminating between hard and incorrect samples to improve the model generalization. Such a training strategy can be formulated dynamically in a self-supervised manner, effectively mimicking the main principle of curriculum learning. Experiments on image classification, image regression, text sequence regression, and event relation reasoning demonstrate the versatility and effectiveness of our method, particularly in the presence of diversified noise levels. Tingting Wu 0007, Hao Zhang 0016, Jinglong Gao, Minji Tang, Bing Qin 0001, Ting Liu 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Dynamic layer-wise sparsification for distributed deep learning
Hao Zhang 0016, Tingting Wu 0007, Zhifeng Ma, Jie Liu 0001 |
Future Gener. Comput. Syst. | 2 |
| 2023 | Data-Augmentation-Based Federated LearningabstractWith the rapid growth of the number of devices generating and collecting data, dispersion becomes an important feature of data in Internet of Things. Federated learning (FL) provides a feasible way to mine information in such distributed data. It involves training machine learning models over multiple distributed participants without raw data transmission. However, due to the data heterogeneity among participants, the performance of the FL model degrades dramatically. Currently, improved methods mainly reduce data heterogeneity from the perspective of modifying the process of model training, which usually have problems, such as high-resource consumption or the need for auxiliary data. In this article, we enhance FL model from another perspective, focusing on data rather than model training. We reduce data heterogeneity by enhancing the trained local data to improve FL performance. Specifically, we propose an FL method based on data augmentation (abbreviated as FedM-UNE), implementing the classic data augmentation method MixUp in federated scenarios without transferring raw data. Furthermore, in order to adapt this method to regression tasks, we first modify MixUp by bilateral neighborhood expansion (MixUp-BNE), and then propose a federated data augmentation method named FedM-BNE based on it. Compared with the conventional FL method, both FedM-UNE and FedM-BNE increase negligible overhead. To demonstrate the effectiveness, we conduct exhaustive experiments on six data sets employing a variety of loss functions. The results indicate that FedM-UNE and FedM-BNE consistently improve the performance of the FL model. Moreover, our methods are compatible with existing FL enhancements, which yield further improvements in performance. Hao Zhang 0016, Qingying Hou, Tingting Wu 0007, Siyao Cheng, Jie Liu 0001 |
IEEE Internet Things J. | 3 |
| 2023 | FedCos: A Scene-Adaptive Enhancement for Federated LearningabstractFederated learning (FL) training global machine learning models over distributed edge devices has attracted sustained attentions. However, the heterogeneity of client data severely degrades the performance of FL compared with that in centralized training. On the one hand, it slows down or even stalls global updates, leading to inefficient communication. On the other hand, it enlarges the distances between local models, resulting in an aggregated global model with poor performance. Fortunately, these shortcomings can be mitigated by reducing the angle between the directions in which a local model move. Based on this observation, we propose FedCos, which reduces the directional inconsistency of local models by introducing a cosine-similarity penalty. It promotes local model iterations toward an auxiliary global direction. Moreover, our approach is auto-adapted to various non-identically and independently distributed (IID) settings without an elaborate selection of hyperparameters. Experimental results on both vision and language tasks with a variety of models (including CNN, ResNet, LSTM, etc.) show that FedCos outperforms the well-known baselines and can enhance them under a variety of FL scenes, including varying degrees of data heterogeneity, different number of participants, and cross-silo and cross-device settings. Besides, FedCos improves the communication efficiency by 2–5 times. With the help of FedCos, multiple FL methods require significantly fewer communication rounds than before to obtain a comparable model. Hao Zhang 0016, Tingting Wu 0007, Siyao Cheng, Jie Liu 0001 |
IEEE Internet Things J. | 2 |
| 2022 | STGN: an Implicit Regularization Method for Learning with Noisy Labels in Natural Language ProcessingabstractNoisy labels are ubiquitous in natural language processing (NLP) tasks.Existing work, namely learning with noisy labels in NLP, is often limited to dedicated tasks or specific training procedures, making it hard to be widely used.To address this issue, SGD noise has been explored to provide a more general way to alleviate the effect of noisy labels by involving benign noise in the process of stochastic gradient descent.However, previous studies exert identical perturbation for all samples, which may cause overfitting on incorrect ones or optimizing correct ones inadequately.To facilitate this, we propose a novel stochastic tailor-made gradient noise (STGN), mitigating the effect of inherent label noise by introducing tailor-made benign noise for each sample.Specifically, we investigate multiple principles to precisely and stably discriminate correct samples from incorrect ones and thus apply different intensities of perturbation to them.A detailed theoretical analysis shows that STGN has good properties, beneficial for model generalization.Experiments on three different NLP tasks demonstrate the effectiveness and versatility of STGN.Also, STGN can boost existing robust training methods. 1 Tingting Wu 0007, Minji Tang, Hao Zhang 0016, Bing Qin 0001, Ting Liu 0001 |
EMNLP | 1 |
| 2022 | Aperiodic Local SGD: Beyond Local SGDabstractVariations of stochastic gradient decedent (SGD) methods are at the core of training deep neural network models. However, in distributed deep learning, where multiple computing devices and data segments are employed in the training process, the performance of SGD can be significantly limited by the overhead of gradient communication. Local SGD methods are designed to overcome this bottleneck by averaging individual gradients trained over parallel workers after multiple local iterations. Currently, both for theoretical analyses and for practical applications, most studies employ periodic synchronization scheme by default, while few of them focus on the aperiodic schemes to obtain better performance models with limited computation and communication overhead. In this paper, we investigate local SGD with an arbitrary synchronization scheme to answer two questions: (1) Is the periodic synchronization scheme best? (2) If not, what is the optimal one? First, for any synchronization scheme, we derive the performance boundary with fixed overhead, and formulate the performance optimization under given computation and communication constraints. Then we find a succinct property of the optimal scheme that the local iteration number decreases as training continues, which indicates the periodic one is suboptimal. Furthermore, with some reasonable approximations, we obtain an explicit form of the optimal scheme and propose Aperiodic Local SGD (ALSGD) as an improved substitute for local SGD without any overhead increment. Our experiments also confirm that with the same computation and communication overhead, ALSGD outperforms local SGD in performance, especially for heterogeneous data. Hao Zhang 0016, Tingting Wu 0007, Siyao Cheng, Jie Liu 0001 |
ICPP | 2 |
| 2014 | Quadtree-based optimal path routing with the smallest routing table sizeabstractRouting schemes play an important role in the network. In terms of the information utilized by routing, the existing schemes can be classified into two main groups: topology-based routing and geographical routing. The former can always guarantee the optimal path, but its routing tables often contain massive entries which seriously impact the algorithm's efficiency. The latter can not guarantee the optimal path, but its routing table size is fairly small. Based on the characteristics of above routing mechanisms, we present a novel geographical routing mechanism which can guarantee the optimal path with the minimum overhead. By utilizing the geographical location information and the quadtree data structure, the routing table size can be reduced to its information-theoretic lower bound. Our theoretical analysis suggests that the performance of the routing table size of our proposed scheme is better than the best IP-based routing table compression result in the literature. Tingting Wu 0007, Chi Zhang 0001, Nenghai Yu, Miao Pan |
GLOBECOM | 1 |