Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Daniel Tang

dblp:42/7442 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 46% Software maintenance and evolution · 27% Program analysis · 27%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program analysis
code representation learning
0.912025
Hedgecode: A Multi-Task Hedging Contrastive Learning Framework for Code Search · ICSE 2025
Software maintenance and evolution
code search
0.912025
Hedgecode: A Multi-Task Hedging Contrastive Learning Framework for Code Search · ICSE 2025
Program synthesis and code generation
code generation with language models
0.812024
DataRecipe - How to Cook the Data for CodeLLM? · ASE 2024
Program synthesis and code generation › code generation with language models
training data quality
0.812024
DataRecipe - How to Cook the Data for CodeLLM? · ASE 2024
Data integration and cleaning
data quality
0.212024
DataRecipe - How to Cook the Data for CodeLLM? · ASE 2024

Methods — techniques the papers use, named apart from their topics

semantic correlation analysis · 1.5metadata analysis · 1.5multi-task learning · 0.9contrastive learning · 0.9
YearPublicationVenuePosition
2025 Hedgecode: A Multi-Task Hedging Contrastive Learning Framework for Code Search
abstract
Code search is a vital activity in software engineering, focused on identifying and retrieving the correct code snippets based on a query provided in natural language. Approaches based on deep learning techniques have been increasingly adopted for this task, enhancing the initial representations of both code and its natural language descriptions. Despite this progress, there remains an unexplored gap in ensuring consistency between the representation spaces of code and its descriptions. Furthermore, existing methods have not fully leveraged the potential relevance between code snippets and their descriptions, presenting a challenge in discerning fine-grained semantic distinctions among similar code snippets. To address these challenges, we introduce a multi-task hedging contrastive Learning framework for Code Search, referred to as HedgeCode. HedgeCode is structured around two primary training phases. The first phase, known as the representation alignment stage, proposes a hedging contrastive learning approach. This method aims to detect subtle differences between code and natural language text, thereby aligning their representation spaces by identifying relevance. The subsequent phase involves multi-task joint learning, wherein the previously trained model serves as the encoder. This stage optimizes the model through a combination of supervised and self-supervised contrastive learning tasks. Our framework's effectiveness is demonstrated through its performance on the CodeSearchNet benchmark, showcasing HedgeCode's ability to address the mentioned limitations in code search tasks.
Gong Chen 0007, Xiaoyuan Xie, Daniel Tang, Qi Xin 0001
ICSE3
2024 Anti-hate Speech Framework: Leveraging Hedging Hyperbolic Learning
Hongyi Zhao, Daniel Tang, Fanliang Bu
ICANN (7)4
2024 DataRecipe - How to Cook the Data for CodeLLM?
abstract
Despite the proliferation of language models, a lack of transparency persists regarding the training datasets used. Security concerns are often cited, but identifying high-quality training data is crucial for optimal model performance. Yet, while significant efforts have been made to improve model performance, dataset quality remains an under-explored area. Our study addresses this gap by comprehensively investigating data-quality properties and processing strategies used to train code generation models. We focus on identifying dataset features that impact model performance and leverage these insights to optimize datasets and enhance model efficacy. Our approach involves a multifaceted analysis encompassing metadata, statistics, data quality issues, semantic correlations between intent and code, and design choices. By manipulating these features, we explore their influence on model performance. Our findings reveal that dataset design choices significantly impact the performance of code generation models. Additionally, semantic correlations between intent and code can also affect performance, although to varying degrees.
Kisub Kim, Jounghoon Kim, Byeongjo Park, Dongsun Kim 0001, Chun Yong Chong, Tiezhu Sun, Daniel Tang, Jacques Klein, Tegawendé F. Bissyandé
ASE8
2024 VisLingInstruct: Elevating Zero-Shot Learning in Multi-Modal Language Models with Autonomous Instruction Optimization
abstract
Dongsheng Zhu, Daniel Tang, Weidong Han, Jinghui Lu, Yukun Zhao, Guoliang Xing, Junfeng Wang, Dawei Yin. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Dongsheng Zhu, Daniel Tang, Weidong Han 0002, Jinghui Lu, Yukun Zhao, Guoliang Xing, Junfeng Wang 0009, Dawei Yin 0001
NAACL-HLT2
2024 A large-scale microblog dataset and stock movement prediction based on Supervised Contrastive Learning model
Daniel Tang
Neurocomputing2
2022 HieNet: Bidirectional Hierarchy Framework for Automated ICD Coding
Daniel Tang, Luchen Zhang, Ding Han
DASFAA (2)2
2022 Emotion Detection in Unfix-Length-Context Conversation
Daniel Tang
ICONIP (2)2
2021 A Large-Scale Hierarchical Structure Knowledge Enhanced Pre-training Framework for Automatic ICD Coding
Daniel Tang, Luchen Zhang
ICONIP (6)2
2021 AIR: Personalized Product Recommender System for Nike's Digital Transformation
abstract
Share on AIR: Personalized Product Recommender System for Nike's Digital Transformation Authors: Steven Essinger Nike, Inc., USA Nike, Inc., USAView Profile , Dave Huber Nike, Inc., USA Nike, Inc., USAView Profile , Daniel Tang Nike, Inc., USA Nike, Inc., USAView Profile Authors Info & Claims RecSys '21: Fifteenth ACM Conference on Recommender SystemsSeptember 2021 Pages 530–532https://doi.org/10.1145/3460231.3474621Online:13 September 2021Publication History 0citation626DownloadsMetricsTotal Citations0Total Downloads626Last 12 Months626Last 6 weeks43 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Steven Essinger, Dave Huber, Daniel Tang
RecSys3