Haitian Sun

dblp:185/6000 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 7 first-author · 10 since 2021
YearPublicationVenuePosition
2026 YOLO-pineapple: enhanced pineapple detection in UAV images using an optimized YOLOv8 model
abstract
Accurate pre-harvest yield estimation is fundamental to pineapple production, as it facilitates optimized harvest scheduling, informs market-responsive pricing strategies, and supports data-driven decision-making in smart agriculture. The rapid advancement of object detection algorithms, coupled with the deployment of miniaturized cameras on Unmanned Aerial Vehicle (UAV) platforms, has made high-throughput pineapple counting for precise yield estimation increasingly feasible. Nonetheless, accurate detection of pineapples in UAV imagery remains challenging due to factors such as the small size of individual fruits, significant scale variations, and complex background textures, all of which impede precise localization and identification. To address these challenges, the present study introduces a novel object detection framework, termed YOLO-Pineapple, designed for accurate pineapple detection in high-resolution color images captured by UAVs. YOLO-Pineapple enhances the baseline YOLOv8 model through several key innovations: (1) the Dynamic Interactive Task Alignment Head (DITAH) is proposed to resolve feature inconsistency and task misalignment between localization and classification branches by integrating interactive and independent features, thereby improving detection accuracy; (2) the incorporation of the Grouped Multi-Scale Convolution (GMSC) module reduces redundant feature computations while capturing richer multi-scale features, enhancing performance in cluttered and heavily occluded field environments; (3) the introduction of the Spatial and Channel Synergistic Attention (SCSA) mechanism facilitates enhanced semantic feature interaction via the combined application of multi-semantic spatial attention and channel self-attention; and (4) the integration of a sample-adaptive weighting mechanism derived from Focaler-IoU with the angle-aware distance component of SIoU culminates in a novel loss function, designated Focaler_SIoU, which achieves more precise bounding box regression, particularly for small objects. Experimental evaluations demonstrate the effectiveness of the proposed YOLO-Pineapple model, which attains a mean average precision (mAP) of 94.4%, a Recall rate of 88.9%, and a precision of 94.6%. The optimized YOLO-Pineapple algorithm constitutes a significant advancement in overcoming the challenges associated with pineapple detection from UAV imagery, while exhibiting promising potential for yield estimation and pineapple field management.
Zhong Xue, Yehong Liu, Yuyin Chen, Mengyao Dong, Xiaying Hao, Weihua Shen, Haitian Sun
Expert Syst. Appl.8
2026 Multi-source heterogeneous domain adaptation with dual-adversarial feature alignment
Yun Zhang 0009, Haitian Sun
Expert Syst. Appl.3
2024 SEMQA: Semi-Extractive Multi-Source Question Answering
abstract
Tal Schuster, Adam Lelkes, Haitian Sun, Jai Gupta, Jonathan Berant, William Cohen, Donald Metzler. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Tal Schuster, Ádám Dániel Lelkes, Haitian Sun, Jai Gupta 0001, Jonathan Berant, William W. Cohen, Donald Metzler
NAACL-HLT3
2023 Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
abstract
Pre-trained vision and language models (Chen et al., 2023b,a;Dai et al., 2023; Li et al., 2023b) have demonstrated state-of-the-art capabilities over existing tasks involving images and texts, including visual question answering.However, it remains unclear whether these models possess the capability to answer questions that are not only querying visual content but knowledge-intensive and informationseeking.In this study, we introduce INFOS-EEK 1 , a visual question answering dataset tailored for information-seeking questions that cannot be answered with only common sense knowledge.Using INFOSEEK, we analyze various pre-trained visual question answering models and gain insights into their characteristics.Our findings reveal that state-of-the-art pre-trained multi-modal models (e.g., PaLI-X, BLIP2, etc.) face challenges in answering visual information-seeking questions, but finetuning on the INFOSEEK dataset elicits models to use fine-grained knowledge that was learned during their pre-training.Furthermore, we show that accurate visual entity recognition can be used to improve performance on INFOSEEK by retrieving relevant documents, showing a significant space for improvement.* Work done when interned at Google 1 Our dataset is available at https:// open-vision-language.github.io/infoseek/.Dataset OK-VQA ViQuAE INFOSEEK PaLM (Q-only) 23.8 31.5 5.6 Current SotA 66.1 22.1 18.2 Require Knowledge † 29.2% 95.2% 95.6% † :% of questions that require knowledge to answer.PaLM (Q-only): a question-only baseline using PaLM.
Yang Chen 0065, Hexiang Hu, Yi Luan, Haitian Sun, Soravit Changpinyo, Alan Ritter, Ming-Wei Chang
EMNLP4
2023 Scenario-based Question Answering with Interacting Contextual Properties
Haitian Sun, William W. Cohen, Ruslan Salakhutdinov
ICLR1
2022 ConditionalQA: A Complex Reading Comprehension Dataset with Conditional Answers
abstract
We describe a Question Answering (QA) dataset that contains complex questions with conditional answers, i.e. the answers are only applicable when certain conditions apply.Answering the questions requires compositional logical reasoning across complex context.We call this dataset ConditionalQA.In addition to conditional answers, the dataset also features:(1) long context documents with information that is related in logically complex ways; (2) multi-hop questions that require compositional logical reasoning; (3) a combination of extractive questions, yes/no questions, questions with multiple answers, and not-answerable questions; (4) questions asked without knowing the answers.We show that ConditionalQA is challenging for many of the existing QA models, especially in selecting answer conditions.We believe that this dataset will motivate further research in understanding complex documents to answer hard questions. 1
Haitian Sun, William W. Cohen, Ruslan Salakhutdinov
ACL (1)1
2021 LEGO: Latent Execution-Guided Reasoning for Multi-Hop Question Answering on Knowledge Graphs
abstract
Answering complex natural language questions on knowledge graphs (KGQA) is a challenging task. It requires reasoning with the input natural language questions as well as a massive, incomplete heterogeneous KG. Prior methods obtain an abstract structured query graph/tree from the input question and traverse the KG for answers following the query tree. However, they inherently cannot deal with missing links in the KG. Here we present LEGO, a Latent Execution-Guided reasOning framework to handle this challenge in KGQA. LEGO works in an iterative way, which alternates between (1) a Query Synthesizer, which synthesizes a reasoning action and grows the query tree step-by-step, and (2) a Latent Space Executor that executes the reasoning action in the latent embedding space to combat against the missing information in KG. To learn the synthesizer without step-wise supervision, we design a generic latent execution guided bottom-up search procedure to find good execution traces efficiently in the vast query space. Experimental results on several KGQA benchmarks demonstrate the effectiveness of our framework compared with previous state of the art.
Hongyu Ren, Hanjun Dai, Bo Dai 0001, Michihiro Yasunaga, Haitian Sun, Dale Schuurmans, Jure Leskovec, Denny Zhou
ICML6
2021 Reasoning Over Virtual Knowledge Bases With Open Predicate Relations
abstract
We present the Open Predicate Query Language (OPQL); a method for constructing a virtual KB (VKB) trained entirely from text. Large Knowledge Bases (KBs) are indispensable for a wide-range of industry applications such as question answering and recommendation. Typically, KBs encode world knowledge in a structured, readily accessible form derived from laborious human annotation efforts. Unfortunately, while they are extremely high precision, KBs are inevitably highly incomplete and automated methods for enriching them are far too inaccurate. Instead, OPQL constructs a VKB by encoding and indexing a set of relation mentions in a way that naturally enables reasoning and can be trained without any structured supervision. We demonstrate that OPQL outperforms prior VKB methods on two different KB reasoning tasks and, additionally, can be used as an external memory integrated into a language model (OPQL-LM) leading to improvements on two open-domain question answering tasks.
Haitian Sun, Patrick Verga, Bhuwan Dhingra, Ruslan Salakhutdinov, William W. Cohen
ICML1
2021 Differentiable Open-Ended Commonsense Reasoning
abstract
Bill Yuchen Lin, Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Xiang Ren, William Cohen. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Bill Y. Lin, Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Xiang Ren 0001, William W. Cohen
NAACL-HLT2
2021 Adaptable and Interpretable Neural MemoryOver Symbolic Knowledge
abstract
Pat Verga, Haitian Sun, Livio Baldini Soares, William Cohen. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Patrick Verga, Haitian Sun, Livio B. Soares, William W. Cohen
NAACL-HLT2
2020 Scalable Neural Methods for Reasoning With a Symbolic Knowledge Base
William W. Cohen, Haitian Sun, R. Alex Hofer, Matthew Siegler
ICLR2
2020 Faithful Embeddings for Knowledge Base Queries
abstract
The deductive closure of an ideal knowledge base (KB) contains exactly the logical queries that the KB can answer. However, in practice KBs are both incomplete and over-specified, failing to answer some queries that have real-world answers. \emph{Query embedding} (QE) techniques have been recently proposed where KB entities and KB queries are represented jointly in an embedding space, supporting relaxation and generalization in KB inference. However, experiments in this paper show that QE systems may disagree with deductive reasoning on answers that do not require generalization or relaxation. We address this problem with a novel QE method that is more faithful to deductive reasoning, and show that this leads to better performance on complex queries to incomplete KBs. Finally we show that inserting this new QE module into a neural question-answering system leads to substantial improvements over the state-of-the-art.
Haitian Sun, Andrew O. Arnold, Tania Bedrax-Weiss, Fernando Pereira 0003, William W. Cohen
NeurIPS1
2019 PullNet: Open Domain Question Answering with Iterative Retrieval on Knowledge Bases and Text
abstract
Haitian Sun, Tania Bedrax-Weiss, William Cohen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Haitian Sun, Tania Bedrax-Weiss, William W. Cohen
EMNLP/IJCNLP (1)1
2018 Open Domain Question Answering Using Early Fusion of Knowledge Bases and Text
abstract
Open Domain Question Answering (QA) is evolving from complex pipelined systems to end-to-end deep neural networks.Specialized neural models have been developed for extracting answers from either text alone or Knowledge Bases (KBs) alone.In this paper we look at a more practical setting, namely QA over the combination of a KB and entitylinked text, which is appropriate when an incomplete KB is available with a large text corpus.Building on recent advances in graph representation learning we propose a novel model, GRAFT-Net, for extracting answers from a question-specific subgraph containing text and KB entities and relations.We construct a suite of benchmark tasks for this problem, varying the difficulty of questions, the amount of training data, and KB completeness.We show that GRAFT-Net is competitive with the state-of-the-art when tested using either KBs or text alone, and vastly outperforms existing methods in the combined setting.
Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Kathryn Mazaitis, Ruslan Salakhutdinov, William W. Cohen
EMNLP1
2018 Semi-Supervised Learning with Declaratively Specified Entropy Constraints
abstract
We propose a technique for declaratively specifying strategies for semi-supervised learning (SSL). SSL methods based on different assumptions perform differently on different tasks, which leads to difficulties applying them in practice. In this paper, we propose to use entropy to unify many types of constraints. Our method can be used to easily specify ensembles of semi-supervised learners, as well as agreement constraints and entropic regularization constraints between these learners, and can be used to model both well-known heuristics such as co-training, and novel domain-specific heuristics. Besides, our model is flexible as to the underlying learning mechanism. Compared to prior frameworks for specifying SSL techniques, our technique achieves consistent improvements on a suite of well-studied SSL benchmarks, and obtains a new state-of-the-art result on a difficult relation extraction task.
Haitian Sun, William W. Cohen, Lidong Bing
NeurIPS1