Tao You

dblp:83/7371 · DBLP profile ↗
← Back
30ranked-venue papers
6as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DSAD: A dual-stream aligned detection framework for LLM-generated exam answers
Yongjun Li 0006, Tao You, Linxin Liu
Expert Syst. Appl.3
2026 Knowledge annotation of multi-modal educational resources adaptable to new category discovery
Zeyuan Qu, Tao You
Inf. Sci.3
2025 CCEGAN: Enhancing GAN clustering through contrastive clustering ensemble
Jing Liu 0077, Tao You, Zhong-Yuan Zhang
Inf. Sci.4
2025 The significance of Kappa and F-score in clustering ensemble: a comprehensive analysis
Tao You, Zhong-Yuan Zhang
Knowl. Inf. Syst.4
2025 Dual-Centralized Q-Network-Based Reinforcement Learning for Cooperative Path Planning of Multiple UAVs
Jinchao Chen, Chongde Ren, Yujiao Hu, Ying Zhang 0060, Yantao Lu, Qing Li 0022, Tao You, Joel J. P. C. Rodrigues
IEEE Trans. Intell. Transp. Syst.7
2025 Vision-Based Geometric Model for Accurate and Fast Lane Recognition in Complex Conditions
abstract
Lane recognition is an important component of autonomous driving system and advanced driving assistance system (ADAS) for intelligent vehicles. In complex driving conditions, accurate and fast lane recognition is a challenging issue. In this paper, a vision-based geometric model (VBGM) is proposed for accurate and fast lane recognition in complex conditions. The framework of the VBGM includes an image preprocessing stage and a lane recognition stage. In the image preprocessing stage, the region of interest (ROI) is extracted from the original image, and the original image is transformed into an undistorted greyscale image. In the lane recognition stage, the lane contour is first extracted using the Roberts operator. Then, to accurately and quickly recognize the lane marking, a lane recognition coordinate system (LRCS) and a rotational LRCS (R-LRCS) are constructed. The distracting contours in abnormal regions are padded based on the LRCS using a contextual frames correlation (CFC) strategy, and the midpoints of the lane contour are identified based on the R-LRCS. Finally, an adaptive-order polynomial fitting model is built to fit the lane marking according to the midpoints in the LRCS. To evaluate the effectiveness of the proposed method, two state-of-the-art methods are selected for comparison. The comparative results indicate that the proposed method possesses a higher recognition rate and speed for lane recognition in complex conditions.
Ying Zhang 0060, Shuaishuai Ge, Tingyi Zhao, Jinchao Chen, Tao You, Yantao Lu, Chenglie Du
IEEE Trans. Intell. Transp. Syst.6
2024 DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
abstract
Audio-visual saliency prediction can draw support from diverse modality complements, but further performance enhancement is still challenged by customized architectures as well as task-specific loss functions. In recent studies, denoising diffusion models have shown more promising in unifying task frameworks owing to their inherent ability of generalization. Following this motivation, a novel Diffusion architecture for generalized audio-visual Saliency prediction (DiffSal) is proposed in this work, which formulates the prediction problem as a conditional generative task of the saliency map by utilizing input audio and video as the conditions. Based on the spatiotemporal audio-visual features, an extra network Saliency-UNet is designed to perform multimodal attention modulation for progressive refinement of the ground-truth saliency map from the noisy map. Extensive experiments demonstrate that the proposed DiffSal can achieve excellent performance across six challenging audio-visual benchmarks, with an average relative improvement of 6.3% over the previous state-of-the-art results by six metrics. The project url is htt ps: //junwenxiong. github.io/DiffSal.
Junwen Xiong, Peng Zhang 0005, Tao You, Chuanyue Li, Wei Huang 0013, Yufei Zha
CVPR3
2024 LLM-Assisted Multi-Teacher Continual Learning for Visual Question Answering in Robotic Surgery
abstract
Visual question answering (VQA) can be fundamentally crucial for promoting robotic-assisted surgical education. In practice, the needs of trainees are constantly evolving, such as learning more surgical types and adapting to new surgical instruments/techniques. Therefore, continually updating the VQA system by a sequential data stream from multiple resources is demanded in robotic surgery to address new tasks. In surgical scenarios, the privacy issue of patient data often restricts the availability of old data when updating the model, necessitating an exemplar-free continual learning (CL) setup. However, prior studies overlooked two vital problems of the surgical domain: i) large domain shifts from diverse surgical operations collected from multiple departments or clinical centers, and ii) severe data imbalance arising from the uneven presence of surgical instruments or activities during surgical procedures. This paper proposes to address these two problems with a multimodal large language model (LLM) and an adaptive weight assignment methodology. We first develop a new multi-teacher CL framework that leverages a multimodal LLM as the additional teacher. The strong generalization ability of the LLM can bridge the knowledge gap when domain shifts and data imbalances occur. We then put forth a novel data processing method that transforms complex LLM embeddings into logits compatible with our CL framework. We also design an adaptive weight assignment approach that balances the generalization ability of the LLM and the domain expertise of the old CL model. Finally, we construct a new dataset for surgical VQA tasks. Extensive experimental results demonstrate the superiority of our method to other advanced CL models.
Kexin Chen 0003, Yuyang Du 0001, Tao You, Mobarakol Islam, Yueming Jin, Guangyong Chen, Pheng-Ann Heng
ICRA3
2024 An Iterative Framework for Document-Level Event Argument Extraction Assisted by Long Short-Term Memory
Tao You, Zhihao Fan, Cunxiang Yin, Yancheng He, Jinhua Fu, Zhongyu Wei
NLPCC (2)1
2024 Accurate prediction of antibody function and structure using bio-inspired antibody language model
abstract
In recent decades, antibodies have emerged as indispensable therapeutics for combating diseases, particularly viral infections. However, their development has been hindered by limited structural information and labor-intensive engineering processes. Fortunately, significant advancements in deep learning methods have facilitated the precise prediction of protein structure and function by leveraging co-evolution information from homologous proteins. Despite these advances, predicting the conformation of antibodies remains challenging due to their unique evolution and the high flexibility of their antigen-binding regions. Here, to address this challenge, we present the Bio-inspired Antibody Language Model (BALM). This model is trained on a vast dataset comprising 336 million 40% nonredundant unlabeled antibody sequences, capturing both unique and conserved properties specific to antibodies. Notably, BALM showcases exceptional performance across four antigen-binding prediction tasks. Moreover, we introduce BALMFold, an end-to-end method derived from BALM, capable of swiftly predicting full atomic antibody structures from individual sequences. Remarkably, BALMFold outperforms those well-established methods like AlphaFold2, IgFold, ESMFold and OmegaFold in the antibody benchmark, demonstrating significant potential to advance innovative engineering and streamline therapeutic antibody development by reducing the need for unnecessary trials. The BALMFold structure prediction server is freely available at https://beamlab-sh.com/models/BALMFold.
Hongtai Jing, Zhengtao Gao, Zhangzhi Peng, Shwai He, Tao You, Shuang Ye, Wei Lin 0003
Briefings Bioinform.7
2024 BioDynGrap: Biomedical event prediction via interpretable learning framework for heterogeneous dynamic graphs
Qing Li 0022, Tao You, Jinchao Chen, Ying Zhang 0060, Chenglie Du
Expert Syst. Appl.2
2024 How does Layer Normalization improve Batch Normalization in self-supervised sound source localization?
Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
Neurocomputing5
2024 TransLSTD: Augmenting hierarchical disease risk prediction model with time and context awareness via disease clustering
Tao You, Qiaodong Dang, Qing Li 0022, Guanzhong Wu
Inf. Syst.1
2024 Anomaly detection with dual-channel heterogeneous graph based on hypersphere learning
Qing Li 0022, Guanzhong Wu, Hang Ni, Tao You
Inf. Sci.4
2024 Closed-loop unified knowledge distillation for dense object detection
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
Pattern Recognit.5
2024 Dual ODE: Spatial-Spectral Neural Ordinary Differential Equations for Hyperspectral Image Super-Resolution
abstract
Significant advancements have been made in hyperspectral image (HSI) super-resolution with the development of deep-learning techniques. However, the current application of deep neural network architectures to HSI super-resolution heavily relies on empirical design strategies, which can potentially impede the improvement of image reconstruction performance and introduce distortions in the results. To address this, we propose an innovative HSI super-resolution network called dual ordinary differential equations (Dual ODEs). Drawing inspiration from ordinary differential equations (ODEs), our approach offers reliable guidelines for the design of HSI super-resolution networks. The Dual ODE model leverages a spatial ODE block to extract spatial information and a spectral ODE block to capture internal spectral features. This is accomplished by redefining the conventional residual module using the multiple ODE functions method. To evaluate the performance of our model, we conducted extensive experiments on four benchmark HSI datasets. The results conclusively demonstrate the superiority of our Dual ODE approach over state-of-the-art models. Moreover, our approach incorporates a small number of parameters while maintaining an interpretable model design, thereby reducing model complexity.
Xiao Zhang 0058, Chongxing Song, Tao You, Qicheng Bai, Wei Wei 0008, Lei Zhang 0054
IEEE Trans. Geosci. Remote. Sens.3
2024 LI-EMRSQL: Linking Information Enhanced Text2SQL Parsing on Complex Electronic Medical Records
abstract
Converting natural language text into executable SQL queries significantly impacts the healthcare domain, specifically when applied to electronic medical records. Given that electronic medical records store extensive patient information in a relational multitable database, developing a Text-to-SQL parser would enable the correlation of intricate medical terminology through semantic parsing. A major challenge is designing a versatile Text2SQL parser applicable to new databases. A critical step towards this goal involves schema linking - accurately identifying references to previously unseen columns or tables during SQL creation. In response to these key challenges, we propose a novel framework—Linking Information Enhanced Text2SQL Parsing on Complex Electronic Medical Records (LI-EMRSQL). This model leverages the Poincaré distance metric detection procedure, utilizing induced relations to enhance the performance of pre-existing graph-based parsers and improve schema linkage. To enhance the generalizability of LI-EMRSQL, the detection process is completely unsupervised and does not necessitate additional parameters. On two conventional Text2SQL datasets and two EMRs Text2SQL datasets, the system delivers SOTA performance. Furthermore, notable enhancements in the model's comprehension and alignment of schemas are observed.
Qing Li 0022, Tao You, Jinchao Chen, Ying Zhang 0060, Chenglie Du
IEEE Trans. Reliab.2
2023 Integrated Local and Global Information for Health Risk Prediction Model
abstract
Electronic health record (EHR) data has been widely used in health risk prediction models, and it has an important preventive and intervention role in healthcare. Existing approaches typically regard EHR data in a monolayer observational model, and they assume that visits are monotonically decreasing in importance over time. However, in healthcare practice, clinical experts usually focus on diseases and visits that are closely related to the target disease. In addition, the duration of different categories of diseases has a fixed model, as chronic diseases are usually consistently diagnosed during patient visits. To make full use of this disease category information, a hierarchical self-attentive model is proposed that can model patient representations at both the local and global levels. Specifically, a disease duration matrix with multiple times is constructed for disease clustering. We combine the category information to compute dependencies between diseases and disease embeddings. We further explore the pattern of patient health development from a spatio-temporal perspective. Visit embeddings are updated by learning the effects between different visits via a self-attentive mechanism. In addition, the time interval, a special kind of medical event, is introduced to enhance visit sequence temporal modeling. Extensive experiments on two real-world datasets demonstrate the sota performance of the model. At the same time, we demonstrate the plausibility and interpretability of the model through case studies.
Tao You, Qiaodong Dang
BIBM1
2023 Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source Localization
abstract
Self-supervised sound source localization is usually challenged by the modality inconsistency. In recent studies, contrastive learning based strategies have shown promising to establish such a consistent correspondence between audio and sound sources in visual scenarios. Unfortunately, the insufficient attention to the heterogeneity influence in the different modality features still limits this scheme to be further improved, which also becomes the motivation of our work. In this study, an Induction Network is proposed to bridge the modality gap more effectively. By decoupling the gradients of visual and audio modalities, the discriminative visual representations of sound sources can be learned with the designed Induction Vector in a bootstrap manner, which also enables the audio modality to be aligned with the visual modality consistently. In addition to a visual weighted contrastive loss, an adaptive threshold selection strategy is introduced to enhance the robustness of the Induction Network. Substantial experiments conducted on SoundNet-Flickr and VGG-Sound Source datasets have demonstrated a superior performance compared to other state-of-the-art works in different challenging scenarios. The code is available at https://github.com/Tahy1/AVIN.
Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
ACM Multimedia5
2023 Work-in-Progress: Time-Aware Formation Control of Connected and Automated Vehicle Platoon Based on Weighted Graph Theory
abstract
The regulation time is an important index for formation switching control of connected and automated vehicle (CA V) platoon. This paper proposes a time-aware formation control (T AFC) strategy to improve the formation switching performance of CA V platoon. To construct an effective information sharing mechanism among the vehicles in the platoon, a unidirectional weighted graph is designed to construct the relation of the CA V platoon and calculate the impact factor between two different vehicles. Based on the unidirectional weighted graph, the time-aware requirement is converted to the regulation order problem, and the regulation order which corresponding to the minimum time is designed. According to the T AFC, the qualitative regulation strategy of the CA V platoon and the quantitative tune-up strategy of the vehicles are determined. In order to analyze the performance of the TAFC strategy, two state-of-art methods are selected as the benchmarked methods. The validation results demonstrate the proposed method possesses better performance for formation switching control compared with the benchmarked methods.
Ying Zhang 0060, Tingyi Zhao, Tao You, Yantao Lu, Jinchao Chen
RTSS4
2023 Scheduling energy consumption-constrained workflows in heterogeneous multi-processor embedded systems
Jinchao Chen, Pengcheng Han, Ying Zhang 0060, Tao You, Pengyi Zheng
J. Syst. Archit.4
2023 Object detection based on cortex hierarchical activation in border sensitive mechanism and classification-GIou joint representation
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
Pattern Recognit.5
2022 BioELM: Integrating Biomedical Knowledge into Language Model with Entity-Linking
abstract
Pretrained language models have achieved widespread success on various natural language processing tasks. In the biomedical domain, one line of research is to utilize a large amount of in-domain corpus for pre-training.While these models achieved remarkable improvement on in-domain tasks, they do not take into account the positive role of large-scale in-domain knowledge bases. Integrating biomedical knowledge in the knowledge base like the Unified Medical Language System(UMLS) into these models can further benefit in-domain downstream tasks, such as biomedical named entities and relation extraction. To this end, we proposed BioELM, a pre-trained language model based on entity linking that explicitly leverages knowledge from the UMLS knowledge base. We utilize a two-layer entity-linking structure to integrate entity representations. To optimize the pre-training process, we optimized the masked language modeling and added two training objectives as named entity recognition and entity linking. We validate the performance of our BioELM on named entity recognition and relation extraction tasks on the BLURB benchmark. The experimental results demonstrate that the pre-training tasks and entity-linking strategy on BioELM can improve the performance on both biomedical named entity recognition and relation extraction tasks.
Guanzhong Wu, Tao You
BIBM3
2022 Accelerated Frequent Closed Sequential Pattern Mining for uncertain data
Tao You, Ying Zhang 0060, Jinchao Chen
Expert Syst. Appl.1
2022 BioKnowPrompt: Incorporating imprecise knowledge into prompt-tuning verbalizer with biomedical text for relation extraction
Qing Li 0022, Tao You, Yantao Lu
Inf. Sci.3
2022 Energy-Saving Optimization and Control of Autonomous Electric Vehicles With Considering Multiconstraints
abstract
The energy utilization efficiency of autonomous electric vehicles is seriously affected by the longitudinal motion control performance. However, the longitudinal motion control is constrained by the driving scene. This article proposes an energy-saving optimization and control (ESOC) method to improve the energy utilization efficiency of autonomous electric vehicles. In ESOC, the constraints from the driving scene are thoroughly considered, and the autonomous driving scene constraints are mapped to the vehicle dynamics and control domain. On this basis, the efficiency self-searching method and the multiconstraint energy-saving control strategy are designed. The main ideology of the proposed ESOC is that the energy utilization efficiency of an autonomous electric vehicle can be improved by optimizing and controlling the operation point distribution of the powertrain efficiency. The experimental results demonstrate that the operation point distribution of the autonomous electric vehicle's powertrain efficiency can be well optimized by the proposed ESOC, and the energy consumption results indicate that the proposed ESOC outperforms the state-of-the-art methods.
Ying Zhang 0060, Zhaoyang Ai, Jinchao Chen, Tao You, Chenglie Du
IEEE Trans. Cybern.4
2022 An Adaptive Clustering-Based Algorithm for Automatic Path Planning of Heterogeneous UAVs
abstract
Due to the high maneuverability and strong adaptability, autonomous unmanned aerial vehicles (UAVs) are of high interest to many civilian and military organizations around the world. Automatic path planning which autonomously finds a good enough path that covers the whole area of interest, is an essential aspect of UAV autonomy. In this study, we focus on the automatic path planning of heterogeneous UAVs with different flight and scan capabilities, and try to present an efficient algorithm to produce appropriate paths for UAVs. First, models of heterogeneous UAVs are built, and the automatic path planning is abstracted as a multi-constraint optimization problem and solved by a linear programming formulation. Then, inspired by the density-based clustering analysis and symbiotic interaction behaviours of organisms, an adaptive clustering-based algorithm with a symbiotic organisms search-based optimization strategy is proposed to efficiently settle the path planning problem and generate feasible paths for heterogeneous UAVs with a view to minimizing the time consumption of the search tasks. Experiments on randomly generated regions are conducted to evaluate the performance of the proposed approach in terms of task completion time, execution time and deviation ratio.
Jinchao Chen, Ying Zhang 0060, Lianwei Wu, Tao You, Xin Ning 0001
IEEE Trans. Intell. Transp. Syst.4
2021 Multiple object tracking based on multi-task learning with strip attention
abstract
Abstract Multiple object tracking (MOT) framework based on bifurcate strategy was usually challenged by data association of different model path, which work for object localisation and appearance embedding independently. By incorporating the re‐identification (re‐ID) as appearance embedding model, more recent studies on task combination of a single network have made a great progress in tracking performance. Unfortunately, the contributive improvement from re‐ID model is hard to balance the accuracy and efficiency for the whole framework. For more effective enhancement of the overall tracking performance, a real‐time detection needs to be taken into consideration with other auxiliary means for MOT modelling. Therefore, in this study, a one‐shot multiple object tracking is proposed based on multi‐task learning to obtain satisfactory performance in both speed and robustness. With updated re‐training strategy for the backbone model of detection, a D2LA network is proposed to achieve more characteristic fine‐grained feature extraction in branching task of pedestrian recognition. Additionally, a strip attention module is also introduced to further strengthen the feature discriminative capability of the tracking framework in occlusion. Experiments on the 2DMOT15, MOT16, MOT17, and MOT20 benchmark data sets have shown a superior performance in comparison to other state‐of‐the‐art tracking approaches.
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
IET Image Process.5
2014 Time scale analysis of receptor enzyme activity: Irreversible inhibition sometimes exhibits incubation-time independence
abstract
At early drug discovery, purified protein-based assays are often used to characterise compound potency. As far as dose response is concerned, it is often thought that a time-independent inhibitor is reversible and a time-dependent inhibitor is irreversible. Using a simple kinetics model, we investigate the legitimacy of this. Our model-based analytical analysis and numerical studies reveal that dose response of an irreversible inhibitor may appear time-independent under certain parametric conditions. Hence, time-independence cannot be used as evidence for inhibitor reversibility. Furthermore, we also analysed how the synthesis and degradation of a target receptor affect drug inhibition in an in vitro cell-based assay setting. Indeed, these processes may also influence dose response of an irreversible inhibitor in such a way that it appears time-independent under certain conditions. Hence, time-independent dose response in a cell assay also needs careful considerations. It is necessary to formulate a suitable model for analysis of protein-based assay and in vitro cell assay data to ensure a consistent understanding.
Tao You, Hong Yue
BIBM1
2009 Research on Real-Time Software Sensors Based on Aspect Reconstruction and Reentrant
Tao You, Chenglie Du
ICIC (1)1