EDBT 2026 Demo / reviewers in the wild / expert
Daniel Tang
dblp:42/7442
· DBLP profile ↗
9ranked-venue papers
0as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Program synthesis and code generation · 46% Software maintenance and evolution · 27% Program analysis · 27% | |
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program analysis
code representation learning |
0.9 | 1 | 2025 | Hedgecode: A Multi-Task Hedging Contrastive Learning Framework for Code Search · ICSE 2025 |
Software maintenance and evolution
code search |
0.9 | 1 | 2025 | Hedgecode: A Multi-Task Hedging Contrastive Learning Framework for Code Search · ICSE 2025 |
Program synthesis and code generation
code generation with language models |
0.8 | 1 | 2024 | DataRecipe - How to Cook the Data for CodeLLM? · ASE 2024 |
Program synthesis and code generation › code generation with language models
training data quality |
0.8 | 1 | 2024 | DataRecipe - How to Cook the Data for CodeLLM? · ASE 2024 |
Data integration and cleaning
data quality |
0.2 | 1 | 2024 | DataRecipe - How to Cook the Data for CodeLLM? · ASE 2024 |
Methods — techniques the papers use, named apart from their topics
semantic correlation analysis · 1.5metadata analysis · 1.5multi-task learning · 0.9contrastive learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hedgecode: A Multi-Task Hedging Contrastive Learning Framework for Code SearchabstractCode search is a vital activity in software engineering, focused on identifying and retrieving the correct code snippets based on a query provided in natural language. Approaches based on deep learning techniques have been increasingly adopted for this task, enhancing the initial representations of both code and its natural language descriptions. Despite this progress, there remains an unexplored gap in ensuring consistency between the representation spaces of code and its descriptions. Furthermore, existing methods have not fully leveraged the potential relevance between code snippets and their descriptions, presenting a challenge in discerning fine-grained semantic distinctions among similar code snippets. To address these challenges, we introduce a multi-task hedging contrastive Learning framework for Code Search, referred to as HedgeCode. HedgeCode is structured around two primary training phases. The first phase, known as the representation alignment stage, proposes a hedging contrastive learning approach. This method aims to detect subtle differences between code and natural language text, thereby aligning their representation spaces by identifying relevance. The subsequent phase involves multi-task joint learning, wherein the previously trained model serves as the encoder. This stage optimizes the model through a combination of supervised and self-supervised contrastive learning tasks. Our framework's effectiveness is demonstrated through its performance on the CodeSearchNet benchmark, showcasing HedgeCode's ability to address the mentioned limitations in code search tasks. Gong Chen 0007, Xiaoyuan Xie, Daniel Tang, Qi Xin 0001 |
ICSE | 3 |
| 2024 | Anti-hate Speech Framework: Leveraging Hedging Hyperbolic Learning
Hongyi Zhao, Daniel Tang, Fanliang Bu |
ICANN (7) | 4 |
| 2024 | DataRecipe - How to Cook the Data for CodeLLM?abstractDespite the proliferation of language models, a lack of transparency persists regarding the training datasets used. Security concerns are often cited, but identifying high-quality training data is crucial for optimal model performance. Yet, while significant efforts have been made to improve model performance, dataset quality remains an under-explored area. Our study addresses this gap by comprehensively investigating data-quality properties and processing strategies used to train code generation models. We focus on identifying dataset features that impact model performance and leverage these insights to optimize datasets and enhance model efficacy. Our approach involves a multifaceted analysis encompassing metadata, statistics, data quality issues, semantic correlations between intent and code, and design choices. By manipulating these features, we explore their influence on model performance. Our findings reveal that dataset design choices significantly impact the performance of code generation models. Additionally, semantic correlations between intent and code can also affect performance, although to varying degrees. Kisub Kim, Jounghoon Kim, Byeongjo Park, Dongsun Kim 0001, Chun Yong Chong, Tiezhu Sun, Daniel Tang, Jacques Klein, Tegawendé F. Bissyandé |
ASE | 8 |
| 2024 | VisLingInstruct: Elevating Zero-Shot Learning in Multi-Modal Language Models with Autonomous Instruction OptimizationabstractDongsheng Zhu, Daniel Tang, Weidong Han, Jinghui Lu, Yukun Zhao, Guoliang Xing, Junfeng Wang, Dawei Yin. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Dongsheng Zhu, Daniel Tang, Weidong Han 0002, Jinghui Lu, Yukun Zhao, Guoliang Xing, Junfeng Wang 0009, Dawei Yin 0001 |
NAACL-HLT | 2 |
| 2024 | A large-scale microblog dataset and stock movement prediction based on Supervised Contrastive Learning model
Daniel Tang |
Neurocomputing | 2 |
| 2022 | HieNet: Bidirectional Hierarchy Framework for Automated ICD Coding
Daniel Tang, Luchen Zhang, Ding Han |
DASFAA (2) | 2 |
| 2022 | Emotion Detection in Unfix-Length-Context Conversation
Daniel Tang |
ICONIP (2) | 2 |
| 2021 | A Large-Scale Hierarchical Structure Knowledge Enhanced Pre-training Framework for Automatic ICD Coding
Daniel Tang, Luchen Zhang |
ICONIP (6) | 2 |
| 2021 | AIR: Personalized Product Recommender System for Nike's Digital TransformationabstractShare on AIR: Personalized Product Recommender System for Nike's Digital Transformation Authors: Steven Essinger Nike, Inc., USA Nike, Inc., USAView Profile , Dave Huber Nike, Inc., USA Nike, Inc., USAView Profile , Daniel Tang Nike, Inc., USA Nike, Inc., USAView Profile Authors Info & Claims RecSys '21: Fifteenth ACM Conference on Recommender SystemsSeptember 2021 Pages 530–532https://doi.org/10.1145/3460231.3474621Online:13 September 2021Publication History 0citation626DownloadsMetricsTotal Citations0Total Downloads626Last 12 Months626Last 6 weeks43 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Steven Essinger, Dave Huber, Daniel Tang |
RecSys | 3 |