EDBT 2026 Demo / reviewers in the wild / expert
Jinan Sun
dblp:16/10588
· DBLP profile ↗
24ranked-venue papers
2as first author
19since 2021 · last 2025
0000-0001-8138-9262ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | R3-Bench: Reproducible Real-world Reverse Engineering Dataset for Symbol RecoveryabstractSymbol recovery in reverse engineering is crucial for restoring variable and data structure information in compiled binaries. While learning-based methods have shown promise in recovering both semantic information (names and types) and syntactic information (shapes), they require comprehensive datasets where expressions in binary code are precisely aligned with their source code equivalents. Current techniques for generating such alignments struggle with complex data access patterns, resulting in incomplete training data and consequently hampering model performance and recovery accuracy. We present AST-Align, a novel technique unifying alignment of variables and struct access expressions across multiple architectures (x86 and ARM) and languages (C/C++/Rust). AST-Align significantly improves the number of generated ground truths, capturing four times more struct fields than previous methods. Using this algorithm, we develop R3-Bench, a metadata-rich, extensible dataset with explicit project inclusion criteria and reproducible processing pipeline, comprising over 10 million functions across multiple architectures. Our evaluation establishes baseline performance by testing various approaches from n-gram models to Large Language Models. The results show that while general LLMs initially perform poorly, their effectiveness dramatically improves with proper demonstration. R3-Bench provides a robust foundation for assessing model capabilities and serves as a valuable reference for future symbol recovery research. Muzhi Yu, Zhengran Zeng, Wei Ye 0004, Jinan Sun, Xiaolong Bai, Shikun Zhang |
ASE | 4 |
| 2025 | Learning Resistant Binary Descriptors Against Noise for Efficient Image RetrievalabstractHashing aims to learn a binary-output function that maps an image to a binary vector, which has received increasing attention with its potential in large-scale visual similarity search. Recently, supervised hashing methods have shown remarkable performance, but they assume that all examples are properly labeled. While in reality, it is unsurprising that we may encounter a range of label noise, which may significantly degrade retrieval performance. In response, we propose a noise-resistant Hashing Contrastive learning with hybrid selection (STAR). Specifically, STAR first develops noise-resistant hashing contrastive learning to preserve the similarity structure against label noise. In addition, we propose a hybrid sample selection strategy from the view of both Hamming distance and output uncertainty, which identifies reliable clean examples. Finally, to get rid of potential memorizing of noisy data, we incorporate both clean samples and noisy samples into selective centroid learning, which minimizes distances between clean samples and their centroids while pushing noisy samples away from negative centroids. Extensive experiments validate the efficacy of STAR. Qingqing Long, Haixin Wang 0003, Jinan Sun, Yijia Xiao, Yusheng Zhao, Xiao Luo 0001 |
SIGIR | 3 |
| 2024 | LION: Implicit Vision Prompt TuningabstractDespite recent promising performances across a range of vision tasks, vision Transformers still have an issue of high computational costs. Recently, vision prompt learning has provided an economical solution to this problem without fine-tuning the whole large-scale model. However, the efficiency and effectiveness of existing models are still far from satisfactory due to the parameter cost of extensive prompt blocks and tricky prompt framework designs. In this paper, we propose a light-weight prompt framework named impLicit vIsion prOmpt tuNing (LION), which is motivated by deep implicit models with stable low memory costs for various complex tasks. In particular, we merely insect two equilibrium implicit layers in two ends of the pre-trained backbone with parameters frozen. Moreover, according to the lottery hypothesis, we further prune the parameters to relieve the computation burden in implicit layers. Various experiments have validated that our LION obtains promising performances on a wide range of datasets. Most importantly, LION reduces up to 11.5 % of training parameter numbers while obtaining higher performance than the state-of-the-art VPT, especially under challenging scenes. Furthermore, we find that our proposed LION has an excellent generalization performance, making it an easy way to boost transfer learning in the future. Haixin Wang 0003, Jianlong Chang, Yihang Zhai, Xiao Luo 0001, Jinan Sun, Zhouchen Lin, Qi Tian 0001 |
AAAI | 5 |
| 2024 | ROSE: Relational and Prototypical Structure Learning for Universal Domain Adaptive HashingabstractAs an important problem in searching system development, domain adaptive retrieval seeks to train a retrieval model with both labeled source samples and unlabeled target samples. Although several domain adaptive hashing algorithms have been proposed to handle the problem with high efficiency, they often presume that source and target domains share all classes. However, prior knowledge about the label space on the target domain is hard to obtain in reality. To tackle this, we study a novel and challenging problem of universal domain adaptive retrieval, which evidently increases the difficulty of effective domain alignment. In this paper, we propose a hashing method namedRelational and prOtotypicalStructure lEarning (ROSE) to solve the problem. In particular, to overcome domain shift, we construct a relational structure depicting cross-domain similar pairs based on ranking statistics, then learn from the structure by maximizing the similarity between similar pairs compared with challenging negatives. Moreover, target private samples are detected using the min-max criterion, which helps to construct hashing prototypes in the Hamming space. On this basis, we combine prototypical structure learning with online clustering in the Hamming space, which improves target semantic learning under label deficiency. Extensive experiments on several benchmarks demonstrate that our proposed ROSE significantly outperforms a wide range of state-of-the-art methods. Our source code is available athttps://github.com/WillDreamer/Rose.git. Xinlong Yang, Haixin Wang 0003, Jinan Sun, Yijia Xiao, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | DIOR: Learning to Hash With Label Noise Via Dual Partition and Contrastive LearningabstractDue to the excellent computing efficiency, learning to hash has acquired broad popularity for Big Data retrieval. Although supervised hashing methods have achieved promising performance recently, they presume that all training samples are appropriately annotated. Unfortunately, label noise is ubiquitous owing to erroneous annotations in real-world applications, which could seriously deteriorate the retrieval performance due to imprecise supervised guidance and severe memorization of noisy data. Here we propose a comprehensive method DIOR to handle the difficulties of learning to hash with label noise. DIOR performs partitions from two complementary levels, namely sample level and parameter level. On the one hand, DIOR divides the dataset into a labeled set with clean samples and an unlabeled set with noisy samples using an ensemble of perturbed views. Then we train the network in a contrastive semi-supervised manner by reconstructing label embeddings for both reliable supervision of clean data and sufficient exploration of noisy data. On the other hand, inspired by recent pruning techniques, DIOR divides the parameters in the hashing network into crucial parameters and non-crucial parameters, and then optimizes them separately to reduce the overfitting of noisy data. Extensive experiments on four popular benchmark datasets demonstrate the effectiveness of DIOR. Haixin Wang 0003, Huiyu Jiang, Jinan Sun, Shikun Zhang, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Look Into Gradients: Learning Compact Hash Codes for Out-of-Distribution RetrievalabstractHashing aims to compress raw data into compact binary descriptors, which has drawn increasing interest for efficient large-scale image retrieval. Current deep hashing often employs evaluation protocols where usually query data and training data are from similar distributions. However, more realistic evaluations should take into account a broad spectrum of distribution shifts with varying degrees. Therefore, we study the problem of out-of-distribution generalization in image retrieval, which seeks to learn a retrieval model from a source domain and generalize to unseen target domains. However, this problem is challenging owing to data scarcity in target domains and the potential overfitting of domain-specific patterns. Here, we propose a novel hashing model namedLooking-into-gradients (LOG) for image retrieval under out-of-distribution shifts, which comprehensively explores gradients for both data generation and model optimization. Specifically, to overcome data deficiency in target domains, we formalize the worst-case problem to generate challenging virtue samples via adversarial gradient ascend. Besides, to further enhance model generalization capability, we not only identify non-crucial parameters with minor gradients and values and shrink them to zero, but also modify the inconsistent gradients across domains to prevent learning domain-specific patterns. Extensive experiments on various datasets demonstrate that LOG outperforms state-of-the-art methods by up to 8.54%. Haixin Wang 0003, Xinlong Yang, Jinan Sun, Shikun Zhang, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Leveraging Imitation Learning on Pose Regulation Problem of a Robotic FishabstractIn this article, the pose regulation control problem of a robotic fish is investigated by formulating it as a Markov decision process (MDP). Such a typical task that requires the robot to arrive at the desired position with the desired orientation remains a challenge, since two objectives (position and orientation) may be conflicted during optimization. To handle the challenge, we adopt the sparse reward scheme, i.e., the robot will be rewarded if and only if it completes the pose regulation task. Although deep reinforcement learning (DRL) can achieve such an MDP with sparse rewards, the absence of immediate reward hinders the robot from efficient learning. To this end, we propose a novel imitation learning (IL) method that learns DRL-based policies from demonstrations with inverse reward shaping to overcome the challenge raised by extremely sparse rewards. Moreover, we design a demonstrator to generate various trajectory demonstrations based on one simple example from a nonexpert helper, which greatly reduces the time consumption of collecting robot samples. The simulation results evaluate the effectiveness of our proposed demonstrator and the state-of-the-art (SOTA) performance of our proposed IL method. Furthermore, we deploy the trained IL policy on a physical robotic fish to perform pose regulation in a swimming tank without/with external disturbances. The experimental results verify the effectiveness and robustness of our proposed methods in real world. Therefore, we believe this article is a step forward in the field of biomimetic underwater robot learning. Lu Yue, Chen Wang 0005, Jinan Sun, Shikun Zhang, Airong Wei, Guangming Xie |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Prototypical Mixing and Retrieval-based Refinement for Label Noise-resistant Image RetrievalabstractLabel noise is pervasive in real-world applications, which influences the optimization of neural network models. This paper investigates a realistic but understudied problem of image retrieval under label noise, which could lead to severe overfitting or memorization of noisy samples during optimization. Moreover, identifying noisy samples correctly is still a challenging problem for retrieval models. In this paper, we propose a novel approach called Prototypical Mixing and Retrieval-based Refinement (TITAN) for label noise-resistant image retrieval, which corrects label noise and mitigates the effects of the memorization simultaneously. Specifically, we first characterize numerous prototypes with Gaussian distributions in the hidden space, which would direct the Mixing procedure in providing synthesized samples. These samples are fed into a similarity learning framework with varying emphasis based on the prototypical structure to learn semantics with reduced overfitting. In addition, we retrieve comparable samples for each prototype from simple to complex, which refine noisy samples in an accurate and class-balanced manner. Comprehensive experiments on five benchmark datasets demonstrate the superiority of our proposed TITAN compared with various competing baselines. Xinlong Yang, Haixin Wang 0003, Jinan Sun, Shikun Zhang, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
ICCV | 3 |
| 2023 | Parameter-efficient Tuning of Large-scale Multimodal Foundation ModelabstractDriven by the progress of large-scale pre-training, parameter-efficient transfer learning has gained immense popularity across different subfields of Artificial Intelligence. The core is to adapt the model to downstream tasks with only a small set of parameters. Recently, researchers have leveraged such proven techniques in multimodal tasks and achieve promising results. However, two critical issues remain unresolved: how to further reduce the complexity with lightweight design and how to boost alignment between modalities under extremely low parameters. In this paper, we propose A gracefUl pRompt framewOrk for cRoss-modal trAnsfer (AURORA) to overcome these challenges. Considering the redundancy in existing architectures, we first utilize the mode approximation to generate 0.1M trainable parameters to implement the multimodal parameter-efficient tuning, which explores the low intrinsic dimension with only 0.04% parameters of the pre-trained model. Then, for better modality alignment, we propose the Informative Context Enhancement and Gated Query Transformation module under extremely few parameters scenes. A thorough evaluation on six cross-modal benchmarks shows that it not only outperforms the state-of-the-art but even outperforms the full fine-tuning approach. Our code is available at: https://github.com/WillDreamer/Aurora. Haixin Wang 0003, Xinlong Yang, Jianlong Chang, Dian Jin 0004, Jinan Sun, Shikun Zhang, Xiao Luo 0001, Qi Tian 0001 |
NeurIPS | 5 |
| 2023 | IDEA: An Invariant Perspective for Efficient Domain Adaptive Image RetrievalabstractIn this paper, we investigate the problem of unsupervised domain adaptive hashing, which leverage knowledge from a label-rich source domain to expedite learning to hash on a label-scarce target domain. Although numerous existing approaches attempt to incorporate transfer learning techniques into deep hashing frameworks, they often neglect the essential invariance for adequate alignment between these two domains. Worse yet, these methods fail to distinguish between causal and non-causal effects embedded in images, rendering cross-domain retrieval ineffective. To address these challenges, we propose an Invariance-acquired Domain AdaptivE HAshing (IDEA) model. Our IDEA first decomposes each image into a causal feature representing label information, and a non-causal feature indicating domain information. Subsequently, we generate discriminative hash codes using causal features with consistency learning on both source and target domains. More importantly, we employ a generative model for synthetic samples to simulate the intervention of various non-causal effects, ultimately minimizing their impact on hash codes for domain invariance. Comprehensive experiments conducted on benchmark datasets validate the superior performance of our IDEA compared to a variety of competitive baselines. Haixin Wang 0003, Hao Wu 0094, Jinan Sun, Shikun Zhang, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
NeurIPS | 3 |
| 2023 | DANCE: Learning A Domain Adaptive Framework for Deep HashingabstractThis paper studies unsupervised domain adaptive hashing, which aims to transfer a hashing model from a label-rich source domain to a label-scarce target domain. Current state-of-the-art approaches generally resolve the problem by integrating pseudo-labeling and domain adaptation techniques into deep hashing paradigms. Nevertheless, they usually suffer from serious class imbalance in pseudo-labels and suboptimal domain alignment caused by the neglection of the intrinsic structures of two domains. To address this issue, we propose a novel method named unbiaseD duAl hashiNg Contrastive lEarning (DANCE) for domain adaptive image retrieval. The core of our DANCE is to perform contrastive learning on hash codes from both instance level and prototype level. To begin, DANCE utilizes label information to guide instance-level hashing contrastive learning in the source domain. To generate unbiased and reliable pseudo-labels for semantic learning in the target domain, we uniformly select samples around each label embedding in the Hamming space. A momentum-update scheme is also utilized to smooth the optimization process. Additionally, we measure the semantic prototype representations in both source and target domains and incorporate them into a domain-aware prototype-level contrastive learning paradigm, which enhances domain alignment in the Hamming space while maximizing the model capacity. Experimental results on a number of well-known domain adaptive retrieval benchmarks validate the effectiveness of our proposed DANCE compared to a variety of competing baselines in different settings. Haixin Wang 0003, Jinan Sun, Shikun Zhang, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
WWW | 2 |
| 2023 | Toward Effective Domain Adaptive RetrievalabstractThis paper studies the problem of unsupervised domain adaptive hashing, which is less-explored but emerging for efficient image retrieval, particularly for cross-domain retrieval. This problem is typically tackled by learning hashing networks with pseudo-labeling and domain alignment techniques. Nevertheless, these approaches usually suffer from overconfident and biased pseudo-labels and inefficient domain alignment without sufficiently exploring semantics, thus failing to achieve satisfactory retrieval performance. To tackle this issue, we present PEACE, a principled framework which holistically explores semantic information in both source and target data and extensively incorporates it for effective domain alignment. For comprehensive semantic learning, PEACE leverages label embeddings to guide the optimization of hash codes for source data. More importantly, to mitigate the effects of noisy pseudo-labels, we propose a novel method to holistically measure the uncertainty of pseudo-labels for unlabeled target data and progressively minimize them through alternative optimization under the guidance of the domain discrepancy. Additionally, PEACE effectively removes domain discrepancy in the Hamming space from two views. In particular, it not only introduces composite adversarial learning to implicitly explore semantic information embedded in hash codes, but also aligns cluster semantic centroids across domains to explicitly exploit label information. Experimental results on several popular domain adaptive retrieval benchmarks demonstrate the superiority of our proposed PEACE compared with various state-of-the-art methods on both single-domain and cross-domain retrieval tasks. Our source codes are available at https://github.com/WillDreamer/PEACE. Haixin Wang 0003, Jinan Sun, Xiao Luo 0001, Shikun Zhang, Chong Chen 0002, Xian-Sheng Hua 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Frequency-Aware Contrastive Learning for Neural Machine TranslationabstractLow-frequency word prediction remains a challenge in modern neural machine translation (NMT) systems. Recent adaptive training methods promote the output of infrequent words by emphasizing their weights in the overall training objectives. Despite the improved recall of low-frequency words, their prediction precision is unexpectedly hindered by the adaptive objectives. Inspired by the observation that low-frequency words form a more compact embedding space, we tackle this challenge from a representation learning perspective. Specifically, we propose a frequency-aware token-level contrastive learning method, in which the hidden state of each decoding step is pushed away from the counterparts of other target words, in a soft contrastive way based on the corresponding word frequencies. We conduct experiments on widely used NIST Chinese-English and WMT14 English-German translation tasks. Empirical results show that our proposed methods can not only significantly improve the translation quality but also enhance lexical diversity and optimize word representation space. Further investigation reveals that, comparing with related adaptive training strategies, the superiority of our method on low-frequency word prediction lies in the robustness of token-level recall across different frequencies without sacrificing precision. Tong Zhang 0001, Wei Ye 0004, Baosong Yang, Long Zhang 0012, Xingzhang Ren, Dayiheng Liu, Jinan Sun, Shikun Zhang, Haibo Zhang 0013 |
AAAI | 7 |
| 2022 | HEART: Towards Effective Hash Codes under Label NoiseabstractHashing, which encodes raw data into compact binary codes, has grown in popularity for large-scale image retrieval due to its storage and computation efficiency. Although deep supervised hashing has lately shown promising performance, they mostly assume that the semantic labels of training data are ideally noise-free, which is often unrealistic in real-world applications. In this paper, considering the practical application, we focus on the problem of learning to hash with label noise and propose a novel method called HEART to address the problem. HEART is a holistic framework which explores latent semantic distributions to select both clean samples and pairs of high confidence for mitigating the impacts of label noise. From a statistical perspective, our HEART characterizes each image by its multiple augmented views that can be considered as examples from its latent distribution and then calculates semantic distances between images using energy distances between their latent distributions. With semantic distances, we can select confident similar pairs to guide hashing contrastive learning for high-quality hash codes. Moreover, to prevent the memorization of noisy examples, we propose a novel strategy to identify clean samples which have small variations of losses on the latent distributions and train the network on clean samples using a pointwise loss. Experimental results on several popular benchmark datasets demonstrate the effectiveness of our HEART compared with a wide range of baselines. Jinan Sun, Haixin Wang 0003, Xiao Luo 0001, Shikun Zhang, Chong Chen 0002, Xian-Sheng Hua 0001 |
ACM Multimedia | 1 |
| 2022 | TencentCLS: The Cloud Log Service with High Query PerformancesabstractWith the trend of cloud computing, the cloud log service is becoming increasingly important, as it plays a critical role in tasks such as root cause analysis, service monitoring and security audition. To meet these needs, we provide Tencent Cloud Log Service (TencentCLS), a one-stop solution for log collection, storage, analysis and dumping. It currently hosts more than a million tenants, of which the largest ones can generate up to PB-level logs per day. The most important challenge that TencentCLS faces is to support both low-latency and resource-efficient queries on such large quantities of log data. To address that challenge, we propose a novel search engine based upon Lucene. The system features a novel procedure for querying logs within a time range, an indexing technique for the time field, as well as optimized query algorithms dedicated to multiple critical and common query types. As a result, the search engine at TencentCLS gains significant performance improvements against Lucene. It achieves 20x performance increase with standard queries, and 10x performance increase with histogram queries in massive log query scenarios. In addition, TencentCLS also supports storing and querying with microsecond-level time precision, as well as the microsecond-level time order preservation capability. Muzhi Yu, Zhaoxiang Lin, Jinan Sun, Runyun Zhou, Guoqiang Jiang, Shikun Zhang |
Proc. VLDB Endow. | 3 |
| 2022 | From Simulation to Reality: A Learning Framework for Fish-Like Robots to Perform Control TasksabstractThe fish-like robot is one of the typical underwater robots, which has the advantage of high maneuverability with low noise due to its bioinspired structure and biomimetic locomotion. However, it is challenging to efficiently design motion controllers for such robots to achieve satisfactory performance on specific control tasks in the real underwater environment, since the complex fluid-structure interaction exists during their swimming and exact dynamic models are absent. In this article, we propose a learning framework, incorporating a simulation system and a training methodology, to autonomously and fast train in simulation to create control policies that are capable of directly applying to a type of physical fish-like robots to perform motion control tasks. First, we construct a simulation system combining a data-driven environment and a computational fluid dynamics (CFD)-based environment, thus well balancing the simulation accuracy and the calculation speed. Second, we design a training methodology to train deep reinforcement learning (DRL)-based policies for the robot in our constructed simulation system to perform a specific control task. Then, we use two typical motion control tasks to verify our proposed framework. One is the path-following control task, which is a one-objective problem with dense rewards, while the other is the pose control task which is a two-objective problem with sparse rewards. For each task, the DRL-based control policy trained by our learning framework is directly deployed on the physical fish-like robot to perform the task in the real world. Experimental results show that the policies trained in simulation still work well in the real world, and perform even better in terms of control accuracy and stability compared with the traditional control methods, thus demonstrating the effectiveness of our learning framework. Runyu Tian, Hongqi Yang, Chen Wang 0005, Jinan Sun, Shikun Zhang, Guangming Xie |
IEEE Trans. Robotics | 5 |
| 2021 | Point, Disambiguate and Copy: Incorporating Bilingual Dictionaries for Neural Machine TranslationabstractTong Zhang, Long Zhang, Wei Ye, Bo Li, Jinan Sun, Xiaoyu Zhu, Wen Zhao, Shikun Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Tong Zhang 0001, Long Zhang 0012, Wei Ye 0004, Bo Li 0099, Jinan Sun, Shikun Zhang |
ACL/IJCNLP (1) | 5 |
| 2021 | TransVae: A Novel Variational Sequence-to-Sequence Framework for Semi-supervised Learning and Diversity ImprovementabstractText generation tasks require that the generated text have certain diversity while ensuring the relevance. Traditional Seq2Seq models usually use cross entropy as the objective function. It demands the results keep strictly consistent with the ground truth texts, which easily leads to the lack of variability in generated texts. In this paper, we propose a novel framework, TransVAE, which applies Variational Auto-Encoder (VAE) to improve the Seq2Seq architecture. We design the Translator module to transform the latent variable spaces of origin input to target output, thus enhancing the diversity of generated texts and supporting semi-supervised learning. Moreover, we add attention and copy mechanisms to the TransVAE model to balance the relevance and diversity. Abundant experiments are carried out on three different string transduction tasks: dialogue generation, machine translation, and text summarization. The experiment results verify the effectiveness of our method. Tianxiang Hu, Xingzhang Ren, Jinan Sun, Kai Liu 0028 |
IJCNN | 4 |
| 2021 | Exploiting Method Names to Improve Code Summarization: A Deliberation Multi-Task Learning ApproachabstractCode summaries are brief natural language descriptions of source code pieces. The main purpose of code summarization is to assist developers in understanding code and to reduce documentation workload. In this paper, we design a novel multi-task learning (MTL) approach for code summarization through mining the relationship between method code summaries and method names. More specifically, since a method's name can be considered as a shorter version of its code summary, we first introduce the tasks of generation and informativeness prediction of method names as two auxiliary training objectives for code summarization. A novel two-pass deliberation mechanism is then incorporated into our MTL architecture to generate more consistent intermediate states fed into a summary decoder, especially when informative method names do not exist. To evaluate our deliberation MTL approach, we carried out a large-scale experiment on two existing datasets for Java and Python. The experiment results show that our technique can be easily applied to many state-of-the-art neural models for code summarization and improve their performance. Meanwhile, our approach shows significant superiority when generating summaries for methods with non-informative names. Rui Xie 0003, Wei Ye 0004, Jinan Sun, Shikun Zhang |
ICPC | 3 |
| 2020 | Deep Dynamic Boosted ForestabstractRandom forest is widely exploited as an ensemble learning method. In many practical applications, however, there is still a significant challenge to learn from imbalanced data. To alleviate this limitation, we propose a deep dynamic boosted forest (DDBF), a novel ensemble algorithm that incorporates the notion of hard example mining into random forest. Specifically, we propose to measure the quality of each leaf node of every decision tree in the random forest to determine hard examples. By iteratively training and then removing easy examples from training data, we evolve the random forest to focus on hard examples dynamically so as to balance the proportion of samples and learn decision boundaries better. Data can be cascaded through these random forests learned in each iteration in sequence to generate more accurate predictions. Our DDBF outperforms random forest on 5 UCI datasets, MNIST and SATIMAGE, and achieved state-of-the-art results compared to other deep models. Moreover, we show that DDBF is also a new way of sampling and can be very useful and efficient when learning from imbalanced data. Haixin Wang 0003, Xingzhang Ren, Jinan Sun, Wei Ye 0004, Muzhi Yu, Shikun Zhang |
ACML | 3 |
| 2020 | Stacking Networks Dynamically for Image Restoration Based on the Plug-and-Play Framework
Haixin Wang 0003, Muzhi Yu, Jinan Sun, Wei Ye 0004, Chen Wang 0005, Shikun Zhang |
ECCV (13) | 4 |
| 2018 | Runtime Resource Management for Microservices-Based Applications: A Congestion Game Approach (Short Paper)
Ruici Luo, Wei Ye 0004, Jinan Sun, Xueyang Liu, Shikun Zhang |
CollaborateCom | 3 |
| 2018 | Refining Traceability Links Between Vulnerability and Software Component in a Vulnerability Knowledge Graph
Dongdong Du, Xingzhang Ren, Jien Chen, Wei Ye 0004, Jinan Sun, Xiangyu Xi, Shikun Zhang |
ICWE | 6 |
| 2011 | A Novel Method for Formally Detecting RFID Event Using Petri Nets
Jinan Sun, Yu Huang 0004, Shikun Zhang, Chong-Yi Yuan |
SEKE | 1 |