EDBT 2026 Demo / reviewers in the wild / expert
Tong Lin 0002
dblp:74/5719-2
· DBLP profile ↗
32ranked-venue papers
8as first author
17since 2021 · last 2025
0000-0002-0000-834XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 6 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Large Model EnhancementabstractVision-Language Models (VLMs) bring powerful understanding and reasoning capabilities to multimodal tasks. Meanwhile, the great need for capable aritificial intelligence on mobile devices also arises, such as the AI assistant software. Some efforts try to migrate VLMs to edge devices to expand their application scope. Simplifying the model structure is a common method, but as the model shrinks, the trade-off between performance and size becomes more and more difficult. Knowledge distillation (KD) can help models improve comprehensive capabilities without increasing size or data volume. However, most of the existing large model distillation techniques only consider applications on single-modal LLMs, or only use teachers to create new data environments for students. None of these methods takes into account the distillation of the most important cross-modal alignment knowledge in VLMs. We propose a method called Align-KD to guide the student model to learn the cross-modal matching that occurs at the shallow layer. The teacher also helps student learn the projection of vision token into text embedding space based on the focus of text. Under the guidance of Align-KD, the 1.7B MobileVLM V2 model can learn rich knowledge from the 7B teacher model with light design of training loss, and achieve an average score improvement of 2.0 across 6 benchmarks under two training subsets respectively. Qianhan Feng, Tong Lin 0002, Xinghao Chen 0001 |
CVPR | 3 |
| 2025 | A new adaptive gradient method with gradient decomposition
Zhou Shao, Tong Lin 0002 |
Mach. Learn. | 3 |
| 2025 | Full-Stage Pseudo Label Quality Enhancement for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised Temporal Action Localization (WSTAL) aims to localize actions in untrimmed videos using only video-level supervision. Latest method introduce a pseudo label learning framework to bridge the gap between classification-based training and inference targets at localization. Typically, this framework employs a classification-based teacher model to generate pseudo labels, which are then used to train a regression-based student model for precise boundary prediction. However, the quality of these pseudo labels—critical to the student model’s performance—has not been systematically investigated, leading to suboptimal localization accuracy. In this paper, we propose a set of simple yet efficient mechanisms for pseudo label quality enhancement to build our FuSTAL framework. Unlike previous one or two stages methods, FuSTAL decomposes the learning process into three stages and enhances pseudo label quality at each one: cross-video contrastive learning for more informative initiative pseudo labels at theGeneration-Stage, prior-based filtering to remove the false positive proposals at theSelection-Stageand EMA-based distillation for smoother pseudo labels at theTraining-Stage. These designs supplement each other, and enhance action proposals’ quality with respect to the accuracy, true positive rate and smoothness. With the help of these comprehensive designs at all three stages, FuSTAL achieves an average mAP of 50.8% on the benchmark data THUMOS’14, outperforming the previous best method by 1.2%. Qianhan Feng, Tong Lin 0002, Xinghao Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | BaCon: Boosting Imbalanced Semi-supervised Learning via Balanced Feature-Level Contrastive LearningabstractSemi-supervised Learning (SSL) reduces the need for extensive annotations in deep learning, but the more realistic challenge of imbalanced data distribution in SSL remains largely unexplored. In Class Imbalanced Semi-supervised Learning (CISSL), the bias introduced by unreliable pseudo-labels can be exacerbated by imbalanced data distributions. Most existing methods address this issue at instance-level through reweighting or resampling, but the performance is heavily limited by their reliance on biased backbone representation. Some other methods do perform feature-level adjustments like feature blending but might introduce unfavorable noise. In this paper, we discuss the bonus of a more balanced feature distribution for the CISSL problem, and further propose a Balanced Feature-Level Contrastive Learning method (BaCon). Our method directly regularizes the distribution of instances' representations in a well-designed contrastive manner. Specifically, class-wise feature centers are computed as the positive anchors, while negative anchors are selected by a straightforward yet effective mechanism. A distribution-related temperature adjustment is leveraged to control the class-wise contrastive degrees dynamically. Our method demonstrates its effectiveness through comprehensive experiments on the CIFAR10-LT, CIFAR100-LT, STL10-LT, and SVHN-LT datasets across various settings. For example, BaCon surpasses instance-level method FixMatch-based ABC on CIFAR10-LT with a 1.21% accuracy improvement, and outperforms state-of-the-art feature-level method CoSSL on CIFAR100-LT with a 0.63% accuracy improvement. When encountering more extreme imbalance degree, BaCon also shows better robustness than other methods. Qianhan Feng, Lujing Xie, Shijie Fang, Tong Lin 0002 |
AAAI | 4 |
| 2024 | VCC-INFUSE: Towards Accurate and Efficient Selection of Unlabeled Examples in Semi-supervised Learning
Shijie Fang, Qianhan Feng, Tong Lin 0002 |
IJCAI | 3 |
| 2023 | Gradient Descent Optimizes Normalization-Free ResNetsabstractRecent empirical studies observe that even without normalization, a deep residual network can be trained reliably. We call such a structure as normalization-free Residual Networks (N-F ResNets), which add a learnable parameter$\alpha$to control the scale of the residual block instead of normalization. However, the theoretical understanding on N-F ResNets is still limited despite their empirical success. In this paper, we provide the first theoretical understanding of N-F ResNets from two perspectives. Firstly, we prove that the gradient descent (GD) algorithm can find the global minimum of the training loss at a linear rate for over-parameterized N-F ResNets. Secondly, we prove that N-F ResNets can avoid the gradient exploding or vanishing problem, by initializing the key parameter$\alpha$to be a small constant. Notably, we demonstrate that the gradients of N-F ResNets are more stable than those of ResNets with Kaiming initialization. Moreover, empirical experiments on benchmark datasets verify our theoretical results. Zongpeng Zhang, Zenan Ling, Tong Lin 0002, Zhouchen Lin |
IJCNN | 3 |
| 2023 | Hessian regularization of deep neural networks: A novel approach based on stochastic estimators of Hessian trace
Yucong Liu, Shixing Yu, Tong Lin 0002 |
Neurocomputing | 3 |
| 2022 | Convolutional Transformer Networks for Epileptic Seizure DetectionabstractEpilepsy is a chronic neurological disease that affects many people in the world. Automatic epileptic seizure detection based on electroencephalogram (EEG) signals is of great significance and has been widely studied. The current deep learning epilepsy detection algorithms are often designed to be relatively simple and seldom consider the characteristics of EEG signals. In this paper, we propose a promising epilepsy detection model based on convolutional transformer networks. We demonstrate that integrating convolution and transformer modules can achieve higher detection performance. Our convolutional transformer model is composed of two branches: one extracts time-domain features from multiple inputs of channel-exchanged EEG signals, and the other handle frequency-domain representations. Experiments on two EEG datasets show that our model offers state-of-the-art performance. Particularly on the CHB-MIT dataset, our model achieves 96.02% in average sensitivity and 97.94% in average specificity, outperforming other existing methods with clear margins. Nan Ke, Tong Lin 0002, Zhouchen Lin, Xiao-Hua Zhou, Taoyun Ji |
CIKM | 2 |
| 2022 | Efficient Meta-Learning for Continual Learning with Taylor Expansion ApproximationabstractContinual learning aims to alleviate catastrophic forgetting when handling consecutive tasks under non-stationary distributions. Gradient-based meta-learning algorithms have shown the capability to implicitly solve the transfer-interference trade-off problem between different examples. However, they still suffer from the catastrophic forgetting problem in the setting of continual learning, since the past data of previous tasks are no longer available. In this work, we propose a novel efficient meta-learning algorithm for solving the online continual learning problem, where the regularization terms and learning rates are adapted to the Taylor approximation of the parameter's importance to mitigate forgetting. The proposed method expresses the gradient of the meta-loss in closed-form and thus avoid computing second-order derivative which is computationally inhibitable. We also use Proximal Gradient Descent to further improve computational efficiency and accuracy. Experiments on diverse benchmarks show that our method achieves better or on-par performance and much higher efficiency compared to the state-of-the-art approaches. Xiaohan Zou, Tong Lin 0002 |
IJCNN | 2 |
| 2021 | Principal Gradient Direction and Confidence Reservoir Sampling for Continual Learning
Tong Lin 0002 |
ICANN (2) | 2 |
| 2021 | Adversarial Variational Knowledge Distillation
Tong Lin 0002 |
ICANN (3) | 2 |
| 2021 | Channel Capacity of Neural Networks
Gen Ye, Tong Lin 0002 |
ICANN (4) | 2 |
| 2021 | An Efficient non-Backpropagation Method for Training Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) have recently attracted significant research interest and have been regarded as the next generation of artificial neural networks due to their suitability for energy-efficient event-driven neuromorphic computing. However, the existing SNNs error backpropagation (BP) method may have severe difficulties in non-differentiable spiking generation functions and vanishing or exploding gradients. In this paper, we introduce an efficient method for training SNNs without backpropagation. The information bottleneck (IB) principle is leveraged to learn synaptic weights and neuron thresholds of an SNN. The membrane potential state for information representation is learned in real-time for higher time and space efficiency compared with the conventional BP method. Experimental results show that the proposed biologically plausible method achieves comparable accuracy and considerable steps/memory reduction in training SNN on MNIST/FashionMNIST datasets. Shiqi Guo, Tong Lin 0002 |
ICTAI | 2 |
| 2021 | Channel-Weighted Squeeze-and-Excitation Networks For Epileptic Seizure DetectionabstractEpilepsy is a chronic neurological disorder that affects many people in the world. Automatic epileptic detection based on multi-channel electroencephalogram (EEG) signals is of great significance and has been widely studied. Recent deep learning models fail to consider the weights of different EEG channels since a few channels can play more important roles than other ones. In this paper, we propose an end-to-end epilepsy detection model, CW-SRNet, to solve this problem. We design a novel channel-weighted block (CW-Block) to capture the different importance of EEG channels automatically and dynamically. We combine our novel CW-Block with the squeeze-and-excitation residual network to improve epilepsy detection performance. Experiments on two public EEG datasets show that our model achieves state-of-the-art performance. Particularly on the CHB-MIT dataset, our model achieves an average sensitivity of 96.84% and an average specificity of 99.68%, outperforming other methods with clear margins. Nan Ke, Tong Lin 0002, Zhouchen Lin |
ICTAI | 2 |
| 2021 | A Supervisory Mask Attentional Network for Person Re-Identification in Uniform Dress ScenesabstractPerson re-identification (Re-ID) aims at retrieving a person's identification across multiple non-overlapping cameras. While recent Re-ID methods have achieved significant success on a number of benchmark datasets, most of them are still insufficient in the scenes with unified or highly similar dressing like construction sites, schools and factories. To address this problem, we propose a supervisory mask attention network (SMA-Net). Our approach combines two key components: (1) ROI mask mapping (RMM) is a supervisory branch to provide ROI mappings that divide a person region into several parts; (2) Partial mask attention (PMA) integrates channel and space attention mechanisms that focus on local features and different accessories in each ROI. Therefore, the network can pay more attention to the local features and different accessories herein. Compared with the state of the art methods, SMA-Net demonstrates excellent performance on our dataset of construction scenes, with improvement of 6.65% in mAP and 3.8% in Rank-1 accuracy. Ling Bai, Tong Lin 0002 |
ICTAI | 5 |
| 2021 | Stochastic Euler Heavy Ball MethodabstractStochastic Heavy Ball method (SHB) has been widely used in various machine learning and deep learning tasks due to its superior generalization performance. However, a large number of effort need to be spent in tuning the learning rates of SHB, which is costly and inefficient in practical applications. Towards this end, this paper proposes the Stochastic Euler Heavy Ball method (SEHB), which simultaneously achieves good generalization like SHB and obtains rapid convergence. Our method adopts new adaptive learning rates which is different from classical adaptive methods like Adam. Convergence analysis is discussed in both convex and non-convex situations. Furthermore, we conduct numerical experiments and deep learning experiments to test the performance of SEHB. Empirical results demonstrate that our method shows better generalization performance than classical stochastic optimization methods such as SHB and Adam. Zhou Shao, Tong Lin 0002 |
ICTAI | 3 |
| 2021 | Intra-Model Collaborative Learning of Neural NetworksabstractRecently, collaborative learning proposed by Song and Chai has achieved remarkable improvements in image classification tasks by simultaneously training multiple classifier heads. However, huge memory footprints required by such multihead structures may hinder the training of large-capacity baseline models. The natural question is how to achieve collaborative learning within a single network without duplicating any modules. In this paper, we propose four ways of collaborative learning among different parts of a single network with negligible engineering efforts. To improve the robustness of the network, we leverage the consistency of the output layer and intermediate layers for training under the collaborative learning framework. Besides, the similarity of intermediate representation and convolution kernel is also introduced to reduce the reduce redundant in a neural network. Compared to the method of Song and Chai, our framework further considers the collaboration inside a single model and takes smaller overhead. Extensive experiments on Cifar-10, Cifar-100, ImageNet32 and STL-10 corroborate the effectiveness of these four ways separately while combining them leads to further improvements. In particular, test errors on the STL-10 dataset are decreased by 9.28% and 5.45% for ResNet-18 and VGG-16 respectively. Moreover, our method is proven to be robust to label noise with experiments on Cifar-10 dataset. For example, our method has 3.53% higher performance under 50% noise ratio setting. Shijie Fang, Tong Lin 0002 |
IJCNN | 2 |
| 2020 | A Dense-Gated U-Net for Brain Lesion SegmentationabstractBrain lesion segmentation plays a crucial role in diagnosis and monitoring of disease progression. DenseNets have been widely used for medical image segmentation, but much redundancy arises in dense-connected feature maps and the training process becomes harder. In this paper, we address the brain lesion segmentation task by proposing a Dense-Gated U-Net (DGNet), which is a hybrid of Dense-gated blocks and U-Net. The main contribution lies in the dense-gated blocks that explicitly model dependencies among concatenated layers and alleviate redundancy. Based on dense-gated blocks, DGNet can achieve weighted concatenation and suppress useless features. Extensive experiments on MICCAI BraTS 2018 challenge and our collected intracranial hemorrhage dataset demonstrate that our approach outperforms a powerful backbone model and other state-of-the-art methods. Zhongyi Ji, Tong Lin 0002, Wenmin Wang 0001 |
VCIP | 3 |
| 2019 | A preliminary geometric structure simplification for Principal Component Analysis
Huamao Gu, Tong Lin 0002, Xun Wang 0007 |
Neurocomputing | 2 |
| 2017 | Factorization for projective and metric reconstruction via truncated nuclear normabstractStructure from motion (SfM) is a crucial and widely studied problem in computer vision. Recently, the factorization framework for SfM was formulated as a low rank approximation problem: the rank of rescaled measurement matrix is always smaller than four. Since the rank function is non-convex, a common practice is to replace with its convex surrogate, i.e., the nuclear norm. However, nuclear norm sometimes gets unsatisfactory results. In this paper, we apply the recently proposed truncated nuclear norm to handle the factorization framework in a non-convex way, which heavily penalizes the singular values beyond the desired rank. We further introduce weighted ℓ1-norm to handle missing data and outliers uniformly. Based on truncated nuclear norm, we propose two factorization models for projective reconstruction and metric reconstruction, respectively. We also proposed an extremely efficient algorithm to tackle one of the optimization sub-problems. Extensive experiments on synthetic and real datasets verify the effectiveness of our method for projective and metric reconstructions. Our method achieves higher accuracy in 3D reconstruction and is more robust to missing data and outliers. Zhouchen Lin, Tong Lin 0002, Hongbin Zha |
IJCNN | 4 |
| 2015 | Supervised learning via Euler's Elastica models
Tong Lin 0002, Hanlin Xue, Hongbin Zha |
J. Mach. Learn. Res. | 1 |
| 2012 | Total Variation and Euler's Elastica for Supervised Learning
Tong Lin 0002, Hanlin Xue, Hongbin Zha |
ICML | 1 |
| 2012 | Incoherent dictionary learning for sparse representation
Tong Lin 0002, Hongbin Zha |
ICPR | 1 |
| 2010 | CDP Mixture Models for Data ClusteringabstractIn Dirichlet process (DP) mixture models, the number of components is implicitly determined by the sampling parameters of Dirichlet process. However, this kind of models usually produces lots of small mixture components when modeling real-world data, especially high-dimensional data. In this paper, we propose a new class of Dirichlet process mixture models with some constrained principles, named constrained Dirichlet process (CDP) mixture models. Based on general DP mixture models, we add a resampling step to obtain latent parameters. In this way, CDP mixture models can suppress noise and generate the compact patterns of the data. Experimental results on data clustering show the remarkable performance of the CDP mixture models. Yangfeng Ji, Tong Lin 0002, Hongbin Zha |
ICPR | 2 |
| 2009 | Refined Exponential Filter with Applications to Image Restoration and Interpolation
Yanlin Geng, Tong Lin 0002, Zhouchen Lin, Pengwei Hao |
ACCV (3) | 2 |
| 2009 | Mahalanobis Distance Based Non-negative Sparse Representation for Face RecognitionabstractSparse representation for machine learning has been exploited in past years. Several sparse representation based classification algorithms have been developed for some applications, for example, face recognition. In this paper, we propose an improved sparse representation based classification algorithm. Firstly, for a discriminative representation, a non-negative constraint of sparse coefficient is added to sparse representation problem. Secondly, Mahalanobis distance is employed instead of Euclidean distance to measure the similarity between original data and reconstructed data. The proposed classification algorithm for face recognition has been evaluated under varying illumination and pose using standard face databases. The experimental results demonstrate that the performance of our algorithm is better than that of the up-to-date face recognition algorithm based on sparse representation. Yangfeng Ji, Tong Lin 0002, Hongbin Zha |
ICMLA | 2 |
| 2008 | Tumor Targeting for Lung Cancer Radiotherapy Using Machine Learning TechniquesabstractAccurate lung tumor targeting in real time plays a fundamental role in image-guide radiotherapy of lung cancers. Precise tumor targeting is required for both respiratory gating and tracking. Gating is considered as the current state of the art for precise lung cancer radiotherapy, which irradiates the tumor when it moves into a predefined gating window. Tracking seems to be a next-generation technique, and it operates in a more aggressive fashion by following the tumor position with radiation beam in real time. Existing methods for gating and tracking often rely on observed motion patterns of external surrogates or implanted fiducial markers. However, external surrogates suffer from certain degrees of inaccuracy, and implanted fiducial markers are in limited uses due to the risk of pneumothorax. Therefore, direct tumor targeting techniques without implanting fiducial markers are desired. Previous studies in fluoroscopic markerless targeting are mainly based on template matching methods, which may fail when tumor boundary is unclear in fluoroscopic images. In this paper, we propose a novel framework of markerless gating and tracking based on machine learning algorithms. Specifically, gating is treated as a two-class classification problem, which is solved by principal component analysis (PCA) and artificial neural network (ANN). Further, we formulate the tracking problem as a regression task, which employs the correlation between the tumor position and nearby surrogate anatomic features in the image. Four regression methods were tested in this study: 1-degree and 2-degree linear regression, artificial neural network (ANN), and support vector machine (SVM). Finally, we demonstrate the superb performance of the proposed markerless gating and tracking algorithms on 10 fluoroscopic image sequences of 9 patients. For gating, the target coverage (the precision) ranges from 90% to 99%, with mean of 96.5%. For tracking, the mean localization error is about 2.1 pixels and the maximum error at 95% confidence level is about 4.6 pixels (pixel size is about 0.5 mm). Tong Lin 0002, Laura I. Cervino, Nuno Vasconcelos, Steve B. Jiang |
ICMLA | 1 |
| 2008 | Towards On-line Treatment Verification Using cine EPID for Hypofractionated Lung RadiotherapyabstractWe propose a novel approach for on-line treatment verification using cine EPID (electronic portal imaging device) images for hypofractionated lung radiotherapy based on a machine learning algorithm. Hypofractionated lung radiotherapy has high precision requirement, and it is essential to effectively monitor the target making sure the tumor is within beam aperture. We model the treatment verification problem as a two-class classification problem and apply artificial neural network (ANN) to classify the cine EPID images acquired during the treatment into corresponding classes-tumor inside or outside of the beam aperture. Training samples of ANN are generated using digitally reconstructed radiograph (DRR) with artificially added shifts in tumor location-to simulate cine EPID images with different tumor locations. Principal component analysis (PCA) is used to reduce the dimensionality of the training samples and cine EPID images acquired during the treatment. The proposed treatment verification algorithm has been tested on six hypofrationated lung patients in a retrospective fashion. On average, our proposed algorithm achieved 94.66% classification accuracy, 94.50% recall rate, and 99.79% precision rate. Tong Lin 0002, Steve B. Jiang |
ICMLA | 2 |
| 2008 | Riemannian Manifold LearningabstractRecently, manifold learning has been widely exploited in pattern recognition, data analysis, and machine learning. This paper presents a novel framework, called Riemannian manifold learning (RML), based on the assumption that the input high-dimensional data lie on an intrinsically low-dimensional Riemannian manifold. The main idea is to formulate the dimensionality reduction problem as a classical problem in Riemannian geometry, i.e., how to construct coordinate charts for a given Riemannian manifold? We implement the Riemannian normal coordinate chart, which has been the most widely used in Riemannian geometry, for a set of unorganized data points. First, two input parameters (the neighborhood size k and the intrinsic dimension d) are estimated based on an efficient simplicial reconstruction of the underlying manifold. Then, the normal coordinates are computed to map the input high-dimensional data into a low-dimensional space. Experiments on synthetic data as well as real world images demonstrate that our algorithm can learn intrinsic geometric structures of the data, preserve radial geodesic distances, and yield regular embeddings. Tong Lin 0002, Hongbin Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | Integrating color and spatial features for content-based video retrievalabstractWe present a novel scheme for content-based video retrieval by exploring the spatio-temporal information. A shot with significant content changes can be segmented into several subshots that are of coherent content, and a shot similarity measure for video retrieval can be computed from the similarity between corresponding subshots. To characterize the temporal content variations in one shot, we developed two descriptors: dominant color histograms (DCH) and spatial structure histograms (SSH). By fusing temporal information into color content, DCHs for a "group of frames" (GoF) are trying to capture the dominant colors with long durations, which would be the colors of the focused objects or background. SSH is a set of features extracted from color-blob maps to describe spatial information for one individual frame. Experimental results on real-world sports videos prove that our proposed approach achieves the best performance on the average recall (AR) and average normalized modified retrieval rank (ANMRR) for video shot retrievals. Tong Lin 0002, Chong-Wah Ngo, HongJiang Zhang, Qing-Yun Shi |
ICIP (3) | 1 |
| 2001 | Video Scene Extraction by Force CompetitionabstractIn this paper, we present a novel scheme for automatic video scene extraction. A pseudo-object-based shot representation containing more semantics is proposed to measure shot similarity and an force competition approach is proposed to group shots into scene based on content coherences between shots. Two content descriptors, color objects: Dominant Color Histograms (DCH) and Spatial Structure Histograms (SSH), are introduced. To represent temporal content variations, a shot can be segmented into several subshots that are of coherent content, and shot similarity measure is formulated as subshot similarity measure. With this shot representation, scene structure can be extracted by analyzing the splitting and merging force competitions at each shot boundary. Experiment on MPEG-7 test videos achieves promising results by the proposed algorithm. 1. Tong Lin 0002, HongJiang Zhang, Qing-Yun Shi |
ICME | 1 |
| 2000 | Automatic Video Scene Extraction by Shot GroupingabstractFor more efficient organizing, browsing, and retrieving digital video content, it is important to extract video structure information at both scene and shot levels. The paper presents an effective approach to video scene segmentation based on a pseudo-object-based shot correlation analysis. A measure of the semantic correlation of consecutive shots based on dominant color grouping and tracking is proposed. A shot grouping method called expanding window is designed to cluster correlated consecutive shots into one scene. Evaluations based on real-world sports video programs validate the efficiency and effectiveness of our shot correlation measure and scene structure construction. Tong Lin 0002, HongJiang Zhang |
ICPR | 1 |