Yangyu Huang

dblp:52/9052 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1Computer networks · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Demystifying Data Organization for Enhanced LLM Training
abstract
Yalun Dai, Yangyu Huang, Tongshen Yang, Yonghan Wang, Xin Zhang, Wenshan Wu, Qihao Zhao, Hao Li, Yuanyuan Gao, Kim-Hui Yap, Scarlett Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yalun Dai, Yangyu Huang, Tongshen Yang, Yonghan Wang, Wenshan Wu, Qihao Zhao, Hao Li 0069, Kim-Hui Yap, Scarlett Li
ACL (1)2
2026 Family Dynamics with Smart Voice Assistants and Implications for Child-Centered AI Design CSCW015
abstract
AI-powered technologies are becoming increasingly integral to children’s digital experiences through devices like interactive toys, home automation systems, and apps, offering rich, personalized, and dynamic interactions. Despite their growing prevalence, how these AI-powered platforms can be designed to address the unique needs of children remains largely underexplored. Leveraging family interactions with Smart Voice Assistants (VAs) as a case study, we aim to explore how to approach child-centered AI (CCAI) design from a family perspective in this work. Specifically, we interviewed 20 parents and observed children’s VA interactions in eight households in a non-Western context. Using the theoretical lenses of agency and family functioning, we provide empirical insights into family dynamics when interacting with VAs in a less studied cultural setting, such as variations in family interaction types around VAs, the autonomy exercised by different parties, and the family functional roles VAs played. Based on these findings, we argue that CCAI design should be understood as balancing children’s agency, the roles and goals of other involved actors, and the contexts in which AI is used, and that it should focus on creating AI technologies that support positive outcomes for children in ethical ways while thoughtfully considering other stakeholders and their varying purposes for engaging with AI. In doing so, we offer a reconceptualization of CCAI and point to design directions for AI technologies that more meaningfully center child users in family contexts.
Yangyu Huang, Kaiyue Jia, Zhibin Zhou 0002, Junnan Yu
Proc. ACM Hum. Comput. Interact.1
2025 FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
abstract
Wei Li, Xin Zhang, Zhongxin Guo, Shaoguang Mao, Wen Luo, Guangyue Peng, Yangyu Huang, Houfeng Wang, Scarlett Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Wei Li 0232, Xin Zhang 0099, Zhongxin Guo, Shaoguang Mao, Wen Luo 0001, Guangyue Peng, Yangyu Huang, Houfeng Wang, Scarlett Li
ACL (1)7
2025 MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
abstract
Multiple-choice question (MCQ) datasets like Massive Multitask Language Understanding (MMLU) are widely used to evaluate the commonsense, understanding, and problem-solving abilities of large language models (LLMs). However, the open-source nature of these benchmarks and the broad sources of training data for LLMs have inevitably led to benchmark contamination, resulting in unreliable evaluation. To alleviate this issue, we propose the contamination-free MCQ benchmark called MMLU-CF, which reassesses LLMs’ understanding of world knowledge by averting both unintentional and malicious data contamination. To mitigate unintentional data contamination, we source questions from a broader domain of over 200 billion webpages and apply three specifically designed decontamination rules. To prevent malicious data contamination, we divide the benchmark into validation and test sets with similar difficulty and subject distributions. The test set remains closed-source to ensure reliable results, while the validation set is publicly available to promote transparency and facilitate independent evaluation. The performance gap between these two sets of LLMs will indicate the contamination degree on the validation set in the future. We evaluated over 40 mainstream LLMs on the MMLU-CF. Compared to the original MMLU, not only LLMs’ performances significantly dropped but also the performance rankings of them changed considerably. This indicates the effectiveness of our approach in establishing a contamination-free and fairer evaluation standard.
Qihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui 0001, Qinzheng Sun, Shaoguang Mao, Qiufeng Yin, Scarlett Li, Furu Wei
ACL (1)2
2025 PEACE: Empowering Geologic Map Holistic Understanding with MLLMs
abstract
Geologic map, as a fundamental diagram in geology science, provides critical insights into the structure and composition of Earth’s subsurface and surface. These maps are indispensable in various fields, including disaster assessment, resource exploration, and civil engineering. Despite their significance, current Multimodal Large Language Models (MLLMs) often fall short in geologic map understanding. This gap is primarily due to the challenging nature of cartographic generalization, which involves handling high-resolution map, managing multiple associated components, and requiring domain-specific knowledge. To quantify this gap, we construct GeoMap-Bench, the first-ever benchmark for evaluating MLLMs in geologic map understanding, which assesses the full-scale abilities in extracting, referring, grounding, reasoning, and analyzing. To bridge this gap, we introduce GeoMap-Agent, the inaugural agent designed for geologic map understanding, which features three modules: Hierarchical Information Extraction (HIE), Domain Knowledge Injection (DKI), and Prompt-enhanced Question Answering (PEQA). Inspired by the interdisciplinary collaboration among human scientists, an AI expert group acts as consultants, utilizing a diverse tool pool to comprehensively analyze questions. Through comprehensive experiments, GeoMap-Agent achieves an overall score of 0.811 on GeoMap-Bench, significantly outperforming 0.369 of GPT-4o. Our work, emPowering gEologic mAp holistiC undErstanding (PEACE) with MLLMs, paves the way for advanced AI applications in geology, enhancing the efficiency and accuracy of geological investigations. The code and data are available at https://github.com/microsoft/PEACE.
Yangyu Huang, Qihao Zhao, Zhipeng Gui, Tengchao Lv, Lei Cui 0001, Scarlett Li, Furu Wei
CVPR1
2025 Teaching Your Models to Understand Code via Focal Preference Alignment
abstract
Jie Wu, Haoling Li, Xin Zhang, Xiao Liu, Yangyu Huang, Jianwen Luo, Yizhen Zhang, Zuchao Li, Ruihang Chu, Yujiu Yang, Scarlett Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jie Wu 0001, Haoling Li, Xin Zhang 0099, Xiao Liu 0029, Yangyu Huang, Zuchao Li, Ruihang Chu, Yujiu Yang 0001, Scarlett Li
EMNLP5
2025 EpiCoder: Encompassing Diversity and Complexity in Code Generation
abstract
Existing methods for code generation use code snippets as seed data, restricting the complexity and diversity of the synthesized data. In this paper, we introduce a novel feature tree-based synthesis framework, which revolves around hierarchical code features derived from high-level abstractions of code. The feature tree is constructed from raw data and refined iteratively to increase the quantity and diversity of the extracted features, which captures and recognizes more complex patterns and relationships within the code. By adjusting the depth and breadth of the sampled subtrees, our framework provides precise control over the complexity of the generated code, enabling functionalities that range from function-level operations to multi-file scenarios. We fine-tuned widely-used base models to obtain EpiCoder series, achieving state-of-the-art performance on multiple benchmarks at both the function and file levels. In particular, empirical evidence indicates that our approach shows significant potential in the synthesizing of repository-level code data. Our code and data are publicly available.
Yaoxiang Wang, Haoling Li, Xin Zhang 0099, Jie Wu 0001, Xiao Liu 0029, Wenxiang Hu, Zhongxin Guo, Yangyu Huang, Yujiu Yang 0001, Jinsong Su, Qi Chen 0009, Scarlett Li
ICML8
2024 WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning
abstract
Zhaojian Yu, Xin Zhang, Ning Shang, Yangyu Huang, Can Xu, Yishujie Zhao, Wenxiang Hu, Qiufeng Yin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhaojian Yu, Xin Zhang 0099, Yangyu Huang, Yishujie Zhao, Wenxiang Hu, Qiufeng Yin
ACL (1)4
2023 FreeEnricher: Enriching Face Landmarks without Additional Cost
abstract
Recent years have witnessed significant growth of face alignment. Though dense facial landmark is highly demanded in various scenarios, e.g., cosmetic medicine and facial beautification, most works only consider sparse face alignment. To address this problem, we present a framework that can enrich landmark density by existing sparse landmark datasets, e.g., 300W with 68 points and WFLW with 98 points. Firstly, we observe that the local patches along each semantic contour are highly similar in appearance. Then, we propose a weakly-supervised idea of learning the refinement ability on original sparse landmarks and adapting this ability to enriched dense landmarks. Meanwhile, several operators are devised and organized together to implement the idea. Finally, the trained model is applied as a plug-and-play module to the existing face alignment networks. To evaluate our method, we manually label the dense landmarks on 300W testset. Our method yields state-of-the-art accuracy not only in newly-constructed dense 300W testset but also in the original sparse 300W and WFLW testsets without additional cost.
Yangyu Huang, Jongyoo Kim, Hao Yang 0036, Jiaolong Yang, Dong Chen 0003
AAAI1
2023 MixPro: Data Augmentation with MaskMix and Progressive Attention Labeling for Vision Transformer
Qihao Zhao, Yangyu Huang, Wei Hu 0004, Fan Zhang 0007, Jun Liu 0036
ICLR2
2022 General Facial Representation Learning in a Visual-Linguistic Manner
abstract
How to learn a universal facial representation that boosts all face analysis tasks? This paper takes one step toward this goal. In this paper, we study the transfer performance of pre-trained models on face analysis tasks and introduce a framework, called FaRL, for general facial representation learning. On one hand, the framework involves a contrastive loss to learn high-level semantic meaning from image-text pairs. On the other hand, we propose exploring low-level information simultaneously to further enhance the face representation by adding a masked image modeling. We perform pre-training on LAION-FACE, a dataset containing a large amount of face image-text pairs, and evaluate the representation capability on multiple downstream tasks. We show that FaRL achieves better transfer performance compared with previous pre-trained models. We also verify its superiority in the low-data regime. More importantly, our model surpasses the state-of-the-art methods on face analysis tasks including face parsing and face alignment.
Yinglin Zheng, Hao Yang 0036, Ting Zhang 0002, Jianmin Bao, Dongdong Chen 0001, Yangyu Huang, Lu Yuan 0001, Dong Chen 0003, Ming Zeng 0008, Fang Wen 0001
CVPR6
2021 ADNet: Leveraging Error-Bias Towards Normal Direction in Face Alignment
abstract
The recent progress of CNN has dramatically improved face alignment performance. However, few works have paid attention to the error-bias with respect to error distribution of facial landmarks. In this paper, we investigate the error-bias issue in face alignment, where the distributions of landmark errors tend to spread along the tangent line to landmark curves. This error-bias is not trivial since it is closely connected to the ambiguous landmark labeling task. Inspired by this observation, we seek a way to leverage the error-bias property for better convergence of CNN model. To this end, we propose anisotropic direction loss (ADL) and anisotropic attention module (AAM) for coordinate and heatmap regression, respectively. ADL imposes strong binding force in normal direction for each landmark point on facial boundaries. On the other hand, AAM is an attention module which can get anisotropic attention mask focusing on the region of point and its local edge connected by adjacent points, it has a stronger response in tangent than in normal, which means relaxed constraints in the tangent. These two methods work in a complementary manner to learn both facial structures and texture details. Finally, we integrate them into an optimized end-to-end training pipeline named ADNet. Our ADNet achieves state-of-the-art results on 300W, WFLW and COFW datasets, which demonstrates the effectiveness and robustness.
Yangyu Huang, Hao Yang 0036, Jongyoo Kim, Fangyun Wei
ICCV1
2021 SeqFace: Learning discriminative features by using face sequences
abstract
Abstract Deep convolutional neural networks (CNNs) have greatly improved the Face Recognition (FR) performance in recent years. Almost all CNNs in FR are trained on the carefully labeled datasets containing plenty of identities. However, such high‐quality datasets are very expensive to collect, which restricts many researchers to achieve state‐of‐the‐art performance. In this paper, a framework, called SeqFace, for learning discriminative face features is proposed. Besides a traditional identity training dataset, the designed SeqFace can train CNNs by using an additional dataset which includes a large number of face sequences collected from videos. Moreover, the label smoothing regularization (LSR) and a new proposed discriminative sequence agent (DSA) loss are employed to enhance the discrimination power of deep face features via making full use of the sequence data. Only with a single ResNet model, the method achieves very competitive performance on several face recognition benchmarks, including LFW, YTF, CFP, AgeDB, and MegaFace. The code and model are publicly available at the website https://github.com/huangyangyu/SeqFace .
Wei Hu 0004, Yangyu Huang, Fan Zhang 0007, Ruirui Li 0001, Heng-Chao Li 0001
IET Image Process.2
2021 P-DIFF+: Improving learning classifier with noisy labels by Noisy Negative Learning loss
Qihao Zhao, Wei Hu 0004, Yangyu Huang, Fan Zhang 0007
Neural Networks3
2020 P-DIFF: Learning Classifier with Noisy Labels based on Probability Difference Distributions
abstract
Learning deep neural network (DNN) classifier with noisy labels is a challenging task because the DNN can easily overfit on these noisy labels due to its high capability. In this paper, we present a very simple but effective training paradigm called P-DIFF, which can train DNN classifiers but obviously alleviate the adverse impact of noisy labels. Our proposed probability difference distribution implicitly reflects the probability of a training sample to be clean, then this probability is employed to re-weight the corresponding sample during the training process. P-DIFF can also achieve good performance even without prior-knowledge on the noise rate of training samples. Experiments on benchmark datasets also demonstrate that P-DIFF is superior to the state-of-the-art sample selection methods.
Wei Hu 0004, Qihao Zhao, Yangyu Huang, Fan Zhang 0007
ICPR3
2020 A flexible n/2 adversary node resistant and halting recoverable blockchain sharding protocol
abstract
Summary Blockchain sharding is a promising approach to solving the dilemma between decentralization and high performance (transaction throughput) for blockchain. The main challenge of blockchain sharding systems is how to reach a decision on a statement among a subgroup (shard) of people while ensuring the whole population recognizes this statement. Namely, the challenge is to prevent an adversary who does not have the majority of nodes globally but have the majority of nodes inside a shard. Most blockchain sharding approaches can only reach a correct consensus inside a shard with at most n/3 evil nodes in a n node system. There is a blockchain sharding approach which can prevent an incorrect decision to be reached when the adversary does not have n/2 nodes globally. However, the system can be stopped from reaching consensus (become deadlocked) if the adversary controls a smaller number of nodes. In this article, we present an improved Blockchain sharding approach that can withstand n/2 adversarial nodes and recover from deadlocks. The recovery is made by dynamically adjusting the number of shards and the shard size. A performance analysis suggests our approach has a high performance (transaction throughput) while requiring little bandwidth for synchronization.
Yibin Xu, Yangyu Huang, Jianhua Shao 0001, George Theodorakopoulos 0001
Concurr. Comput. Pract. Exp.2
2019 Noise-Tolerant Paradigm for Training Face Recognition CNNs
abstract
Benefit from large-scale training datasets, deep Convolutional Neural Networks(CNNs) have achieved impressive results in face recognition(FR). However, tremendous scale of datasets inevitably lead to noisy data, which obviously reduce the performance of the trained CNN models. Kicking out wrong labels from large-scale FR datasets is still very expensive, although some cleaning approaches are proposed. According to the analysis of the whole process of training CNN models supervised by angular margin based loss(AM-Loss) functions, we find that the distribution of training samples implicitly reflects their probability of being clean. Thus, we propose a novel training paradigm that employs the idea of weighting samples based on the above probability. Without any prior knowledge of noise, we can train high performance CNN models with largescale FR datasets. Experiments demonstrate the effectiveness of our training paradigm. The codes are available at https://github.com/huangyangyu/NoiseFace.
Wei Hu 0004, Yangyu Huang, Fan Zhang 0007, Ruirui Li 0001
CVPR2
2019 Contract-connection: An efficient communication protocol for Distributed Ledger Technology
abstract
Distributed Ledger Technology (DLT) is promising to become the foundation of many decentralised systems. However, the unbalanced and unregulated network layout contributes to the inefficiency of DLT especially in Internet of Things (IoT) environments, where nodes connect to only a limited number of peers. The data communication speed globally is unbalanced and does not live up to the constraints of efficient real-time distributed systems. In this paper, we introduce a new communication protocol, which enables nodes to calculate the tradeoff between connecting/disconnecting a peer in a completely decentralised manner. The network layout globally is continuously re-balancing and optimising along with nodes adjusting their peers. This communication protocol weakened the inequality of the communication network. The experiment suggests this communication protocol is stable and efficient.
Yibin Xu, Yangyu Huang
IPCCC2
2019 MWPoW: Multiple Winners Proof of Work Protocol, a Decentralisation Strengthened Fast-Confirm Blockchain Protocol
abstract
Blockchain mining should not be a game among power oligarchs. In this paper, we present the Multiple Winners Proof of Work Protocol (MWPoW), a mining-pool-like decentralised blockchain consensus protocol. MWPoW enables disadvantaged nodes which post only a small amount of calculation resource in the mining game to create blocks together and compete with power oligarchs without centralised representatives. A precise Support Rate of blocks can be determined through the mining process; the mechanism of the mainchain determination is therefore changed and has become faster and more straightforward. A method that periodically adjusts the block size and the block interval is introduced into MWPoW, which increases the system flexibility in the changes of network conditions and data flow. Experiments suggest, without lifting calculation and bandwidth requirements, MWPoW is more attractive to disadvantaged nodes due to its mostly increased reward expectation for disadvantaged nodes. The transaction pending time is shortened chiefly, and either the block interval or the block size can be adapted amid the changes of overall network conditions.
Yibin Xu, Yangyu Huang
Secur. Commun. Networks2
2018 3dRPC: a web server for 3D RNA-protein structure prediction
abstract
RNA-protein interactions occur in many biological processes. To understand the mechanism of these interactions one needs to know three-dimensional (3D) structures of RNA-protein complexes. 3dRPC is an algorithm for prediction of 3D RNA-protein complex structures and consists of a docking algorithm RPDOCK and a scoring function 3dRPC-Score. RPDOCK is used to sample possible complex conformations of an RNA and a protein by calculating the geometric and electrostatic complementarities and stacking interactions at the RNA-protein interface according to the features of atom packing of the interface. 3dRPC-Score is a knowledge-based potential that uses the conformations of nucleotide-amino-acid pairs as statistical variables and that is used to choose the near-native complex-conformations obtained from the docking method above. Recently, we built a web server for 3dRPC. The users can easily use 3dRPC without installing it locally. RNA and protein structures in PDB (Protein Data Bank) format are the only needed input files. It can also incorporate the information of interface residues or residue-pairs obtained from experiments or theoretical predictions to improve the prediction. Availability and implementation: The address of 3dRPC web server is http://biophy.hust.edu.cn/3dRPC. Contact: [email protected].
Yangyu Huang
Bioinform.1
2014 Ray tracing via GPU rasterization
Wei Hu 0004, Yangyu Huang, Fan Zhang 0007, Guodong Yuan, Wei Li 0032
Vis. Comput.2
2011 ASPDock: protein-protein docking algorithm using atomic solvation parameters model
abstract
BACKGROUND: Atomic Solvation Parameters (ASP) model has been proven to be a very successful method of calculating the binding free energy of protein complexes. This suggests that incorporating it into docking algorithms should improve the accuracy of prediction. In this paper we propose an FFT-based algorithm to calculate ASP scores of protein complexes and develop an ASP-based protein-protein docking method (ASPDock). RESULTS: The ASPDock is first tested on the 21 complexes whose binding free energies have been determined experimentally. The results show that the calculated ASP scores have stronger correlation (r ≈ 0.69) with the binding free energies than the pure shape complementarity scores (r ≈ 0.48). The ASPDock is further tested on a large dataset, the benchmark 3.0, which contain 124 complexes and also shows better performance than pure shape complementarity method in docking prediction. Comparisons with other state-of-the-art docking algorithms showed that ASP score indeed gives higher success rate than the pure shape complementarity score of FTDock but lower success rate than Zdock3.0. We also developed a softly restricting method to add the information of predicted binding sites into our docking algorithm. The ASP-based docking method performed well in CAPRI rounds 18 and 19. CONCLUSIONS: ASP may be more accurate and physical than the pure shape complementarity in describing the feature of protein docking.
Dachuan Guo, Yangyu Huang, Shiyong Liu
BMC Bioinform.3