Huan Sun 0001

dblp:33/2952-1 · DBLP profile ↗
← Back
88ranked-venue papers
7as first author
41since 2021 · last 2026
0000-0001-6436-4813ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 60 · 3 first-author · 36 since 2021Databases, data management, data science and information retrieval · 28 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Computer networks · 2Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2026 The Treacherous Envoy Problem: Trust, Collusion, and Accountability in Multi-Agent Workflows [Blue Sky Paper]
abstract
LLM-based agents increasingly execute workflows on behalf of principals with conflicting interests: booking hotels, approving procurement, coordinating payments. We call an agent treacherous when its realized effects violate the delegating principal's intent while the evidence it releases is locally consistent with every check the principal's disclosure policy allows. We formulate the Treacherous Envoy Problem (TEP): given a natural-language delegation, publicly observable tool-side effects, and released artifacts, can a principal decide whether the workflow is unsafe, and in particular whether an envoy has acted treacherously, in time to abort before irreversible harm where possible and otherwise with evidence sufficient for adjudication and accountability, when other agents may collude and no universally trusted mediator exists? We argue TEP is structurally hard. Detection requires three coupled capabilities (expressive natural-language negotiation, verifiable conformance of effects to intent, and bounded disclosure of private context), and these capabilities resist independent resolution: any mechanism that strengthens one tightens the constraints on the others. We give formal evidence for two of the coupling edges, identify five tensions where the trilemma bites in practice, five workflow exploitation patterns ranging from treachery proper to harmful-but-compliant value degradation, and research directions toward the detection and deterrence infrastructure that agent workflows currently lack.
Zhiqiang Lin 0001, Huan Sun 0001
SACMAT3
2025 AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection
abstract
Weidi Luo, Shenghong Dai, Xiaogeng Liu, Suman Banerjee, Huan Sun, Muhao Chen, Chaowei Xiao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Weidi Luo, Shenghong Dai, Xiaogeng Liu, Suman Banerjee 0001, Huan Sun 0001, Muhao Chen 0001, Chaowei Xiao
ACL (1)5
2025 MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
abstract
Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, Graham Neubig. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang 0019, Kai Zhang 0033, Shengbang Tong, Yuxuan Sun 0002, Botao Yu, Ge Zhang 0009, Huan Sun 0001, Yu Su 0001, Wenhu Chen, Graham Neubig
ACL (1)10
2025 AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists
abstract
Yifei Li, Hanane Nour Moussa, Ziru Chen, Shijie Chen, Botao Yu, Mingyi Xue, Benjamin Burns, Tzu-Yao Chiu, Vishal Dey, Zitong Lu, Chen Wei, Qianheng Zhang, Tianyu Zhang, Song Gao, Xuhui Huang, Xia Ning, Nesreen K. Ahmed, Ali Payani, Huan Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yifei Li 0005, Hanane Nour Moussa, Ziru Chen, Botao Yu, Mingyi Xue 0001, Benjamin Burns, Tzu-Yao Chiu, Vishal Dey, Zitong Lu, Qianheng Zhang, Song Gao 0001, Xuhui Huang, Xia Ning, Nesreen K. Ahmed, Ali Payani, Huan Sun 0001
EMNLP19
2025 ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
abstract
The advancements of language language models (LLMs) have piqued growing interest in developing LLM-based language agents to automate scientific discovery end-to-end, which has sparked both excitement and skepticism about the true capabilities of such agents. In this work, we argue that for an agent to fully automate scientific discovery, it must be able to complete all essential tasks in the workflow. Thus, we call for rigorous assessment of agents on individual tasks in a scientific workflow before making bold claims on end-to-end automation. To this end, we present ScienceAgentBench, a new benchmark for evaluating language agents for data-driven scientific discovery. To ensure the scientific authenticity and real-world relevance of our benchmark, we extract 102 tasks from 44 peer-reviewed publications in four disciplines and engage nine subject matter experts to validate them. We unify the target output for every task to a self-contained Python program file and employ an array of evaluation metrics to examine the generated programs, execution results, and costs. Each task goes through multiple rounds of manual validation by annotators and subject matter experts to ensure its annotation quality and scientific plausibility. We also propose two effective strategies to mitigate data contamination concerns. Using our benchmark, we evaluate five open-weight and proprietary LLMs, each with three frameworks: direct prompting, OpenHands, and self-debug. Given three attempts for each task, the best-performing agent can only solve 32.4% of the tasks independently and 34.3% with expert-provided knowledge. These results underscore the limited capacities of current language agents in generating code for data-driven discovery, let alone end-to-end automation for scientific research.
Ziru Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li 0005, Zeyi Liao, Zitong Lu, Vishal Dey, Mingyi Xue 0001, Frazier N. Baker, Benjamin Burns, Daniel Adu-Ampratwum, Xuhui Huang, Xia Ning, Song Gao 0001, Yu Su 0001, Huan Sun 0001
ICLR20
2025 Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
abstract
Multimodal large language models (MLLMs) are transforming the capabilities of graphical user interface (GUI) agents, facilitating their transition from controlled simulations to complex, real-world applications across various platforms. However, the effectiveness of these agents hinges on the robustness of their grounding capability. Current GUI agents predominantly utilize text-based representations such as HTML or accessibility trees, which, despite their utility, often introduce noise, incompleteness, and increased computational overhead. In this paper, we advocate a human-like embodiment for GUI agents that perceive the environment entirely visually and directly perform pixel-level operations on the GUI. The key is visual grounding models that can accurately map diverse referring expressions of GUI elements to their coordinates on the GUI across different platforms. We show that a simple recipe, which includes web-based synthetic data and slight adaptation of the LLaVA architecture, is surprisingly effective for training such visual grounding models. We collect the largest dataset for GUI visual grounding so far, containing 10M GUI elements and their referring expressions over 1.3M screenshots, and use it to train UGround, a strong universal visual grounding model for GUI agents. Empirical results on six benchmarks spanning three categories (grounding, offline agent, and online agent) show that 1) UGround substantially outperforms existing visual grounding models for GUI agents, by up to 20\% absolute, and 2) agents with UGround outperform state-of-the-art agents, despite the fact that existing agents use additional text-based input while ours only uses visual perception. These results provide strong support for the feasibility and promises of GUI agents that navigate the digital world as humans do.
Boyu Gou, Boyuan Zheng 0001, Yanan Xie, Cheng Chang 0001, Yiheng Shu, Huan Sun 0001, Yu Su 0001
ICLR7
2025 Eia: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
abstract
Recently, generalist web agents have demonstrated remarkable potential in autonomously completing a wide range of tasks on real websites, significantly boosting human productivity. However, web tasks, such as booking flights, usually involve users' personally identifiable information (PII), which may be exposed to potential privacy risks if web agents accidentally interact with compromised websites—a scenario that remains largely unexplored in the literature. In this work, we narrow this gap by conducting the first study on the privacy risks of generalist web agents in adversarial environments. First, we present a realistic threat model for attacks on the website, where we consider two adversarial targets: stealing users' specific PII or the entire user request. Then, we propose a novel attack method, termed Environmental Injection Attack (EIA). EIA injects malicious content designed to adapt well to environments where the agents operate and our work instantiates EIA specifically for privacy scenarios in web environments. We collect 177 action steps that involve diverse PII categories on realistic websites from the Mind2Web dataset, and conduct experiments using one of the most capable generalist web agent frameworks to date. The results demonstrate that EIA achieves up to 70\% attack success rate (ASR) in stealing users' specific PII and 16\% ASR in stealing a full user request at an action step. Additionally, by evaluating the detectability and testing defensive system prompts, we indicate that EIA is challenging to detect and mitigate. Notably, attacks that are not well adapted for a webpage can be detected through careful human inspection, leading to our discussion about the trade-off between security and autonomy. However, extra attackers' efforts can make EIA seamlessly adapted, rendering such human supervision ineffective. Thus, we further discuss the implications on defenses at the pre- and post-deployment stages of the websites without relying on human supervision and call for more advanced defense strategies.
Zeyi Liao, Lingbo Mo, Chejian Xu, Mintong Kang, Jiawei Zhang 0013, Chaowei Xiao, Bo Li 0026, Huan Sun 0001
ICLR9
2025 AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
abstract
Jailbreak attacks serve as essential red-teaming tools, proactively assessing whether LLMs can behave responsibly and safely in adversarial environments. Despite diverse strategies (e.g., cipher, low-resource language, persuasions, and so on) that have been proposed and shown success, these strategies are still manually designed, limiting their scope and effectiveness as a red-teaming tool. In this paper, we propose AutoDAN-Turbo, a black-box jailbreak method that can automatically discover as many jailbreak strategies as possible from scratch, without any human intervention or predefined scopes (e.g., specified candidate strategies), and use them for red-teaming. As a result, AutoDAN-Turbo can significantly outperform baseline methods, achieving a 74.3% higher average attack success rate on public benchmarks. Notably, AutoDAN-Turbo achieves an 88.5 attack success rate on GPT-4-1106-turbo. In addition, AutoDAN-Turbo is a unified framework that can incorporate existing human-designed jailbreak strategies in a plug-and-play manner. By integrating human-designed strategies, AutoDAN-Turbo can even achieve a higher attack success rate of 93.4 on GPT-4-1106-turbo.
Xiaogeng Liu, G. Edward Suh, Yevgeniy Vorobeychik, Z. Morley Mao, Somesh Jha, Patrick McDaniel, Huan Sun 0001, Bo Li 0026, Chaowei Xiao
ICLR8
2025 AdvAgent: Controllable Blackbox Red-teaming on Web Agents
abstract
Foundation model-based agents are increasingly used to automate complex tasks, enhancing efficiency and productivity. However, their access to sensitive resources and autonomous decision-making also introduce significant security risks, where successful attacks could lead to severe consequences. To systematically uncover these vulnerabilities, we propose AdvAgent, a black-box red-teaming framework for attacking web agents. Unlike existing approaches, AdvAgent employs a reinforcement learning-based pipeline to train an adversarial prompter model that optimizes adversarial prompts using feedback from the black-box agent. With careful attack design, these prompts effectively exploit agent weaknesses while maintaining stealthiness and controllability. Extensive evaluations demonstrate that AdvAgent achieves high success rates against state-of-the-art GPT-4-based web agents across diverse web tasks. Furthermore, we find that existing prompt-based defenses provide only limited protection, leaving agents vulnerable to our framework. These findings highlight critical vulnerabilities in current web agents and emphasize the urgent need for stronger defense mechanisms. We release code at https://ai-secure.github.io/AdvAgent/.
Chejian Xu, Mintong Kang, Jiawei Zhang 0013, Zeyi Liao, Lingbo Mo, Huan Sun 0001, Bo Li 0026
ICML7
2025 GroundCocoa: A Benchmark for Evaluating Compositional & Conditional Reasoning in Language Models
abstract
Harsh Kohli, Sachin Kumar, Huan Sun. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Harsh Kohli, Sachin Kumar 0009, Huan Sun 0001
NAACL (Long Papers)3
2025 Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
abstract
Agentic search such as Deep Research systems-where agents autonomously browse the web, synthesize information, and return comprehensive citation-backed answers-represents a major shift in how users interact with web-scale information. While promising greater efficiency and cognitive offloading, the growing complexity and open-endedness of agentic search have outpaced existing evaluation benchmarks and methodologies, which largely assume short search horizons and static answers. In this paper, we introduce Mind2Web 2, a benchmark of 130 realistic, high-quality, and long-horizon tasks that require real-time web browsing and extensive information synthesis, constructed with over 1000 hours of human labor. To address the challenge of evaluating time-varying and complex answers, we propose a novel Agent-as-a-Judge framework. Our method constructs task-specific judge agents based on a tree-structured rubric design to automatically assess both answer correctness and source attribution. We conduct a comprehensive evaluation of ten frontier agentic search systems and human performance, along with a detailed error analysis to draw insights for future development. The best-performing system, OpenAI Deep Research, can already achieve 50-70% of human performance while spending half the time, highlighting its great potential. Altogether, Mind2Web 2 provides a rigorous foundation for developing and benchmarking the next generation of agentic search systems.
Boyu Gou, Zanming Huang, Yuting Ning, Yu Gu 0016, Michael Lin, Weijian Qi, Andrei Kopanev, Botao Yu, Bernal Jimenez Gutierrez, Yiheng Shu, Chan Hee Song, Jiaman Wu, Hanane Nour Moussa, Tianshu Zhang 0001, Yifei Li 0005, Tianci Xue, Zeyi Liao, Kai Zhang 0033, Boyuan Zheng 0001, Zhaowei Cai, Viktor Rozgic, Morteza Ziyadi, Huan Sun 0001, Yu Su 0001
NeurIPS25
2025 EVOSCHEMA: TOWARDS TEXT-TO-SQL ROBUSTNESS AGAINST SCHEMA EVOLUTION
abstract
Neural text-to-SQL models, which translate natural language questions (NLQs) into SQL queries given a database schema, have achieved remarkable performance. However, database schemas frequently evolve to meet new requirements. Such schema evolution often leads to performance degradation for models trained on static schemas. Existing work either mainly focuses on simply paraphrasing some syntactic or semantic mappings among NLQ, DB and SQL, or lacks a comprehensive and controllable way to investigate the model robustness issue under the schema evolution, which is insufficient when facing the increasingly complex and rich database schema changes in reality, especially in the LLM era. To address the challenges posed by schema evolution, we present EvoSchema, a comprehensive benchmark designed to assess and enhance the robustness of text-to-SQL systems under real-world schema changes. EvoSchema introduces a novel schema evolution taxonomy, encompassing ten perturbation types across column-level and table-level modifications, systematically simulating the dynamic nature of database schemas. Through EvoSchema, we conduct an in-depth evaluation spanning different open-source and closed-source LLMs, revealing that table-level perturbations have a significantly greater impact on model performance compared to column-level changes. Furthermore, EvoSchema inspires the development of more resilient text-to-SQL systems, in terms of both model training and database design. The models trained on EvoSchema's diverse schema designs can force the model to distinguish the schema difference for the same questions to avoid learning spurious patterns, which demonstrate remarkable robustness compared to those trained on unperturbed data on average. This benchmark offers valuable insights into model behavior and a path forward for designing systems capable of thriving in dynamic, real-world environments.
Tianshu Zhang 0001, Kun Qian 0002, Siddhartha Sahai, Shaddy Garg, Huan Sun 0001, Yunyao Li 0001
Proc. VLDB Endow.6
2024 When is Tree Search Useful for LLM Planning? It Depends on the Discriminator
abstract
In this paper, we examine how large language models (LLMs) solve multi-step problems under a language agent framework with three components: a generator, a discriminator, and a planning method.We investigate the practical utility of two advanced planning methods, iterative correction and tree search.We present a comprehensive analysis of how discrimination accuracy affects the overall performance of agents when using these two methods or a simpler method, re-ranking.Experiments on two tasks, text-to-SQL parsing and mathematical reasoning, show that: (1) advanced planning methods demand discriminators with at least 90% accuracy to achieve significant improvements over re-ranking; (2) current LLMs' discrimination abilities have not met the needs of advanced planning methods to achieve such improvements; (3) with LLM-based discriminators, advanced planning methods may not adequately balance accuracy and efficiency.For example, compared to the other two methods, tree search is at least 10-20 times slower but leads to negligible performance gains, which hinders its real-world applications.1
Ziru Chen, Michael White 0001, Raymond J. Mooney, Ali Payani, Yu Su 0001, Huan Sun 0001
ACL (1)6
2024 MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
abstract
We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected multimodal questions from college exams, quizzes, and text-books, covering six core disciplines: Art & Design, Busi-ness, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering. These questions span 30 subjects and 183 subfields, comprising 30 highly het-erogeneous image types, such as charts, diagrams, maps, tables, music sheets, and chemical structures. Unlike existing benchmarks, MMMU focuses on advanced perception and reasoning with domain-specific knowledge, challenging models to perform tasks akin to those faced by experts. The evaluation of 28 open-source LMMs as well as the propri-etary GPT-4V(ision) and Gemini highlights the substantial challenges posed by MMMU. Even the advanced GPT-4V and Gemini Ultra only achieve accuracies of 56% and 59% respectively, indicating significant room for improvement. We believe MMMU will stimulate the community to build next-generation multimodal foundation models towards expert artificial general intelligence.
Xiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 0033, Ruoqi Liu, Ge Zhang 0009, Samuel Stevens 0001, Dongfu Jiang, Weiming Ren, Yuxuan Sun 0002, Cong Wei 0001, Botao Yu, Ruibin Yuan, Renliang Sun, Boyuan Zheng 0001, Zhenzhu Yang, Wenhao Huang 0001, Huan Sun 0001, Yu Su 0001, Wenhu Chen
CVPR20
2024 AgentBench: Evaluating LLMs as Agents
abstract
The potential of Large Language Model (LLM) as agents has been widely acknowledged recently. Thus, there is an urgent need to quantitatively evaluate LLMs as agents on challenging tasks in interactive environments. We present AgentBench, a multi-dimensional benchmark that consists of 8 distinct environments to assess LLM-as-Agent's reasoning and decision-making abilities. Our extensive test over 29 API-based and open-sourced (OSS) LLMs shows that, while top commercial LLMs present a strong ability of acting as agents in complex environments, there is a significant disparity in performance between them and many OSS competitors that are no larger than 70B. We identify the typical reasons of failures in environments and LLMs, showing that poor long-term reasoning, decision-making, and instruction following abilities are the main obstacles for developing usable LLM agents. Improving instruction following and training on high quality multi-round alignment data could improve agent performance. And different from existing assumptions, training on code present ambivalent impacts on different agent tasks. Datasets, environments, and an integrated evaluation package for AgentBench are released at https://github.com/THUDM/AgentBench.
Xiao Liu 0036, Hao Yu 0030, Hanchen Zhang, Yifan Xu 0014, Xuanyu Lei, Hanyu Lai, Yu Gu 0016, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng 0001, Aohan Zeng, Zhengxiao Du, Sheng Shen 0001, Tianjun Zhang, Yu Su 0001, Huan Sun 0001, Minlie Huang, Yuxiao Dong, Jie Tang 0001
ICLR19
2024 MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning
abstract
We introduce MAmmoTH, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving. The MAmmoTH models are trained on MathInstruct, our meticulously curated instruction tuning dataset. MathInstruct is compiled from 13 math datasets with intermediate rationales, six of which have rationales newly curated by us. It presents a unique hybrid of chain-of-thought (CoT) and program-of-thought (PoT) rationales, and also ensures extensive coverage of diverse fields in math. The hybrid of CoT and PoT not only unleashes the potential of tool use but also allows different thought processes for different math problems. As a result, the MAmmoTH series substantially outperform existing open-source models on nine mathematical reasoning datasets across all scales with an average accuracy gain between 16% and 32%. Remarkably, our MAmmoTH-7B model reaches 33% on MATH (a competition-level dataset), which exceeds the best open-source 7B model (WizardMath) by 23%, and the MAmmoTH-34B model achieves 44% accuracy on MATH, even surpassing GPT-4’s CoT result. Our work underscores the importance of diverse problem coverage and the use of hybrid rationales in developing superior math generalist models.
Xiang Yue, Xingwei Qu, Ge Zhang 0009, Wenhao Huang 0001, Huan Sun 0001, Yu Su 0001, Wenhu Chen
ICLR6
2024 eCeLLM: Generalizing Large Language Models for E-commerce from Large-scale, High-quality Instruction Data
abstract
With tremendous efforts on developing effective e-commerce models, conventional e-commerce models show limited success in generalist e-commerce modeling, and suffer from unsatisfactory performance on new users and new products – a typical out-of-domain generalization challenge. Meanwhile, large language models (LLMs) demonstrate outstanding performance in generalist modeling and out-of-domain generalizability in many fields. Toward fully unleashing their power for e-commerce, in this paper, we construct ECInstruct, the first open-sourced, large-scale, and high-quality benchmark instruction dataset for e-commerce. Leveraging ECInstruct, we develop eCeLLM, a series of e-commerce LLMs, by instruction-tuning general-purpose LLMs. Our comprehensive experiments and evaluation demonstrate that eCeLLM models substantially outperform baseline models, including the most advanced GPT-4, and the state-of-the-art task-specific models in in-domain evaluation. Moreover, eCeLLM exhibits excellent generalizability to out-of-domain settings, including unseen products and unseen instructions, highlighting its superiority as a generalist e-commerce model. Both the ECInstruct dataset and the eCeLLM models show great potential in empowering versatile and effective LLMs for e-commerce. ECInstruct and eCeLLM models are publicly accessible through this link.
Bo Peng 0009, Xinyi Ling, Ziru Chen, Huan Sun 0001, Xia Ning
ICML4
2024 GPT-4V(ision) is a Generalist Web Agent, if Grounded
abstract
The recent development on large multimodal models (LMMs), especially GPT-4V(ision) and Gemini, has been quickly expanding the capability boundaries of multimodal models beyond traditional tasks like image captioning and visual question answering. In this work, we explore the potential of LMMs like GPT-4V as a generalist web agent that can follow natural language instructions to complete tasks on any given website. We propose SEEACT, a generalist web agent that harnesses the power of LMMs for integrated visual understanding and acting on the web. We evaluate on the recent MIND2WEB benchmark. In addition to standard offline evaluation on cached websites, we enable a new online evaluation setting by developing a tool that allows running web agents on live websites. We show that GPT-4V presents a great potential for web agents—it can successfully complete 51.1% of the tasks on live websites if we manually ground its textual plans into actions on the websites. This substantially outperforms text-only LLMs like GPT-4 or smaller models (FLAN-T5 and BLIP-2) specifically fine-tuned for web agents. However, grounding still remains a major challenge. Existing LMM grounding strategies like set-of-mark prompting turns out to be not effective for web agents, and the best grounding strategy we develop in this paper leverages both the HTML structure and visuals. Yet, there is still a substantial gap with oracle grounding, leaving ample room for further improvement. All code, data, and evaluation tools are available at https://github.com/OSU-NLP-Group/SeeAct.
Boyuan Zheng 0001, Boyu Gou, Jihyung Kil, Huan Sun 0001, Yu Su 0001
ICML4
2024 How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
abstract
Lingbo Mo, Boshi Wang, Muhao Chen, Huan Sun. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Lingbo Mo, Boshi Wang, Muhao Chen 0001, Huan Sun 0001
NAACL-HLT4
2024 TableLlama: Towards Open Large Generalist Models for Tables
abstract
Tianshu Zhang, Xiang Yue, Yifei Li, Huan Sun. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Tianshu Zhang 0001, Xiang Yue, Yifei Li 0005, Huan Sun 0001
NAACL-HLT4
2024 Grokking of Implicit Reasoning in Transformers: A Mechanistic Journey to the Edge of Generalization
abstract
We study whether transformers can learn to *implicitly* reason over parametric knowledge, a skill that even the most capable language models struggle with. Focusing on two representative reasoning types, composition and comparison, we consistently find that transformers *can* learn implicit reasoning, but only through *grokking*, i.e., extended training far beyond overfitting. The levels of generalization also vary across reasoning types: when faced with out-of-distribution examples, transformers fail to systematically generalize for composition but succeed for comparison. We delve into the model's internals throughout training, conducting analytical experiments that reveal: 1) the mechanism behind grokking, such as the formation of the generalizing circuit and its relation to the relative efficiency of generalizing and memorizing circuits, and 2) the connection between systematicity and the configuration of the generalizing circuit. Our findings guide data and training setup to better induce implicit reasoning and suggest potential improvements to the transformer architecture, such as encouraging cross-layer knowledge sharing. Furthermore, we demonstrate that for a challenging reasoning task with a large search space, GPT-4-Turbo and Gemini-1.5-Pro based on non-parametric memory fail badly regardless of prompting styles or retrieval augmentation, while a fully grokked transformer can achieve near-perfect accuracy, showcasing the power of parametric memory for complex reasoning.
Boshi Wang, Xiang Yue, Yu Su 0001, Huan Sun 0001
NeurIPS4
2023 Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters
abstract
Boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer, Huan Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Boshi Wang, Sewon Min, Xiang Deng 0001, You Wu 0001, Luke Zettlemoyer, Huan Sun 0001
ACL (1)7
2023 Synthetic Text Generation with Differential Privacy: A Simple and Practical Recipe
abstract
Xiang Yue, Huseyin Inan, Xuechen Li, Girish Kumar, Julia McAnallen, Hoda Shajari, Huan Sun, David Levitan, Robert Sim. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Xiang Yue, Huseyin A. Inan, Julia McAnallen, Hoda Shajari, Huan Sun 0001, David Levitan, Robert Sim
ACL (1)7
2023 Federated Learning for Semantic Parsing: Task Formulation, Evaluation Setup, New Algorithms
abstract
This paper studies a new task of federated learning (FL) for semantic parsing, where multiple clients collaboratively train one global model without sharing their semantic parsing data.By leveraging data from multiple clients, the FL paradigm can be especially beneficial for clients that have little training data to develop a data-hungry neural semantic parser on their own.We propose an evaluation setup to study this task, where we re-purpose widely-used single-domain text-to-SQL datasets as clients to form a realistic heterogeneous FL setting and collaboratively train a global model.As standard FL algorithms suffer from the high client heterogeneity in our realistic setup, we further propose a novel LOss Reduction Adjusted Reweighting (Lorar) mechanism to mitigate the performance degradation, which adjusts each client's contribution to the global model update based on its training loss reduction during each round.Our intuition is that the larger the loss reduction, the further away the current global model is from the client's local optimum, and the larger weight the client should get.By applying Lorar to three widely adopted FL algorithms (FedAvg, FedOPT and FedProx), we observe that their performance can be improved substantially on average (4%-20% absolute gain under MacroAvg) and that clients with smaller datasets enjoy larger performance gains.In addition, the global model converges faster for almost all the clients. 1
Tianshu Zhang 0001, Changchang Liu, Wei-Han Lee, Yu Su 0001, Huan Sun 0001
ACL (1)5
2023 DP-Forward: Fine-tuning and Inference on Language Models with Differential Privacy in Forward Pass
abstract
Differentially private stochastic gradient descent (DP-SGD) adds noise to gradients in back-propagation, safeguarding training data from privacy leakage, particularly membership inference. It fails to cover (inference-time) threats like embedding inversion and sensitive attribute inference. It is also costly in storage and computation when used to fine-tune large pre-trained language models (LMs).
Minxin Du, Xiang Yue, Sherman S. M. Chow, Tianhao Wang 0001, Huan Sun 0001
CCS6
2023 Exploring Chain of Thought Style Prompting for Text-to-SQL
abstract
In-context learning with large language models (LLMs) has recently caught increasing attention due to its superior few-shot performance on various tasks.However, its performance on text-to-SQL parsing still has much room for improvement.In this paper, we hypothesize that a crucial aspect of LLMs to improve for text-to-SQL parsing is their multi-step reasoning ability.Thus, we systematically study how to enhance LLMs' reasoning ability through chain of thought (CoT) style prompting, including the original chain-of-thought prompting (Wei et al., 2022b) and least-to-most prompting (Zhou et al., 2023).Our experiments demonstrate that iterative prompting as in Zhou et al. (2023) may be unnecessary for text-to-SQL parsing, and using detailed reasoning steps tends to have more error propagation issues.Based on these findings, we propose a new CoT-style prompting method for text-to-SQL parsing.It brings 5.2 and 6.5 point absolute gains on the Spider development set and the Spider Realistic set, respectively, compared to the standard prompting method without reasoning steps; 2.4 and 1.5 point absolute gains, compared to the least-to-most prompting method 1 .
Chang-Yu Tai, Ziru Chen, Tianshu Zhang 0001, Xiang Deng 0001, Huan Sun 0001
EMNLP5
2023 Multitask Prompt Tuning Enables Parameter-Efficient Transfer Learning
Zhen Wang 0041, Rameswar Panda, Leonid Karlinsky, Rogério Feris, Huan Sun 0001
ICLR5
2023 Mind2Web: Towards a Generalist Agent for the Web
abstract
We introduce Mind2Web, the first dataset for developing and evaluating generalist agents for the web that can follow language instructions to complete complex tasks on any website. Existing datasets for web agents either use simulated websites or only cover a limited set of websites and tasks, thus not suitable for generalist web agents. With over 2,000 open-ended tasks collected from 137 websites spanning 31 domains and crowdsourced action sequences for the tasks, Mind2Web provides three necessary ingredients for building generalist web agents: 1) diverse domains, websites, and tasks, 2) use of real-world websites instead of simulated and simplified ones, and 3) a broad spectrum of user interaction patterns. Based on Mind2Web, we conduct an initial exploration of using large language models (LLMs) for building generalist web agents. While the raw HTML of real-world websites are often too large to be fed to LLMs, we show that first filtering it with a small LM significantly improves the effectiveness and efficiency of LLMs. Our solution demonstrates a decent level of performance, even on websites or entire domains the model has never seen before, but there is still a substantial room to improve towards truly generalizable agents. We open-source our dataset, model implementation, and trained models (https://osu-nlp-group.github.io/Mind2Web) to facilitate further research on building a generalist agent for the web.
Xiang Deng 0001, Yu Gu 0016, Boyuan Zheng 0001, Samual Stevens, Boshi Wang, Huan Sun 0001, Yu Su 0001
NeurIPS7
2023 MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
abstract
Text-guided image editing is widely needed in daily life, ranging from personal use to professional applications such as Photoshop.However, existing methods are either zero-shot or trained on an automatically synthesized dataset, which contains a high volume of noise.Thus, they still require lots of manual tuning to produce desirable outcomes in practice.To address this issue, we introduce MagicBrush, the first large-scale, manually annotated dataset for instruction-guided real image editing that covers diverse scenarios: single-turn, multi-turn, mask-provided, and mask-free editing.MagicBrush comprises over 10K manually annotated triplets (source image, instruction, target image), which supports trainining large-scale text-guided image editing models.We fine-tune InstructPix2Pix on MagicBrush and show that the new model can produce much better images according to human evaluation.We further conduct extensive experiments to evaluate current image editing baselines from multiple dimensions including quantitative, qualitative, and human evaluations.The results reveal the challenging nature of our dataset and the gap between current baselines and real-world editing needs.
Kai Zhang 0033, Lingbo Mo, Wenhu Chen, Huan Sun 0001, Yu Su 0001
NeurIPS4
2023 Roll Up Your Sleeves: Working with a Collaborative and Engaging Task-Oriented Dialogue System
abstract
Lingbo Mo, Shijie Chen, Ziru Chen, Xiang Deng, Ashley Lewis, Sunit Singh, Samuel Stevens, Chang-You Tai, Zhen Wang, Xiang Yue, Tianshu Zhang, Yu Su, Huan Sun. Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2023.
Lingbo Mo, Ziru Chen, Xiang Deng 0001, Ashley Lewis, Sunit Singh, Samuel Stevens 0001, Chang-You Tai, Zhen Wang 0041, Xiang Yue, Tianshu Zhang 0001, Yu Su 0001, Huan Sun 0001
SIGDIAL13
2023 Sanitizing Sentence Embeddings (and Labels) for Local Differential Privacy
abstract
Differentially private (DP) learning, notably DP stochastic gradient descent (DP-SGD), has limited applicability in fine-tuning gigantic pre-trained language models (LMs) for natural language processing tasks. The culprit is the perturbation of gradients (as gigantic as entire models), leading to significant efficiency and accuracy drops.
Minxin Du, Xiang Yue, Sherman S. M. Chow, Huan Sun 0001
WWW4
2022 Synthetic Question Value Estimation for Domain Adaptation of Question Answering
abstract
Synthesizing QA pairs with a question generator (QG) on the target domain has become a popular approach for domain adaptation of question answering (QA) models.Since synthetic questions are often noisy in practice, existing work adapts scores from a pretrained QA (or QG) model as criteria to select highquality questions.However, these scores do not directly serve the ultimate goal of improving QA performance on the target domain.In this paper, we introduce a novel idea of training a question value estimator (QVE) that directly estimates the usefulness of synthetic questions for improving the target-domain QA performance.By conducting comprehensive experiments, we show that the synthetic questions selected by QVE can help achieve better target-domain QA performance, in comparison with existing techniques.We additionally show that by using such questions and only around 15% of the human annotations on the target domain, we can achieve comparable performance to the fully-supervised baselines.1
Xiang Yue, Ziyu Yao 0002, Huan Sun 0001
ACL (1)3
2022 Iteratively Prompt Pre-trained Language Models for Chain of Thought
abstract
While Pre-trained Language Models (PLMs) internalize a great amount of world knowledge, they have been shown incapable of recalling these knowledge to solve tasks requiring complex & multi-step reasoning.Similar to how humans develop a "chain of thought" for these tasks, how can we equip PLMs with such abilities?In this work, we explore an iterative prompting framework, a new prompting paradigm which progressively elicits relevant knowledge from PLMs for multi-step inference.We identify key limitations of existing prompting methods, namely they are either restricted to queries with a single identifiable relation/predicate, or being agnostic to input contexts, which makes it difficult to capture variabilities across different inference steps.We propose an iterative context-aware prompter, which addresses these limitations by learning to dynamically synthesize prompts conditioned on the current step's contexts.Experiments on three datasets involving multi-step reasoning show the effectiveness of the iterative scheme and the context-aware prompter design. 1
Boshi Wang, Xiang Deng 0001, Huan Sun 0001
EMNLP3
2021 CliniQG4QA: Generating Diverse Questions for Domain Adaptation of Clinical Question Answering
abstract
Clinical question answering (QA) aims to automatically answer questions from medical professionals based on clinical texts. Studies show that neural QA models trained on one corpus may not generalize well to new clinical texts from a different institute or a different patient group, where largescale QA pairs are not readily available for model retraining. To address this challenge, we propose a simple yet effective framework, CliniQG4QA, which leverages question generation (QG) to synthesize QA pairs on new clinical contexts and boosts QA models without requiring manual annotations. In order to generate diverse types of questions that are essential for training QA models, we further introduce a seq2seq-based question phrase prediction (QPP) module that can be used together with most existing QG models to diversify the generation. Our comprehensive experiment results show that the QA corpus generated by our framework can improve QA models on the new contexts (up to 8% absolute gain in terms of Exact Match), and that the QPP module plays a crucial role in achieving the gain.11Our dataset and code are available at: https://github.com/sunlabosu/CliniQG4QA/.
Xiang Yue, Xinliang Frederick Zhang, Ziyu Yao 0002, Simon M. Lin, Huan Sun 0001
BIBM5
2021 ReasonBERT: Pre-trained to Reason with Distant Supervision
abstract
We present ReasonBERT, a pre-training method that augments language models with the ability to reason over long-range relations and multiple, possibly hybrid, contexts.Unlike existing pre-training methods that only harvest learning signals from local contexts of naturally occurring texts, we propose a generalized notion of distant supervision to automatically connect multiple pieces of text and tables to create pre-training examples that require long-range reasoning.Different types of reasoning are simulated, including intersecting multiple pieces of evidence, bridging from one piece of evidence to another, and detecting unanswerable cases.We conduct a comprehensive evaluation on a variety of extractive question answering datasets ranging from single-hop to multi-hop and from text-only to table-only to hybrid that require various reasoning capabilities and show that ReasonBERT achieves remarkable improvement over an array of strong baselines.Fewshot experiments further demonstrate that our pre-training method substantially improves sample efficiency. 1
Xiang Deng 0001, Yu Su 0001, Alyssa Lees, You Wu 0001, Cong Yu 0001, Huan Sun 0001
EMNLP (1)6
2021 COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval
abstract
We present a large, challenging dataset, COUGH, for COVID-19 FAQ retrieval.Similar to a standard FAQ dataset, COUGH consists of three parts: FAQ Bank, Query Bank and Relevance Set.The FAQ Bank contains ∼16K FAQ items scraped from 55 credible websites (e.g., CDC and WHO).For evaluation, we introduce Query Bank and Relevance Set, where the former contains 1,236 human-paraphrased queries while the latter contains ∼32 humanannotated FAQ items for each query.We analyze COUGH by testing different FAQ retrieval models built on top of BM25 and BERT, among which the best model achieves 48.8 under P@5, indicating a great challenge presented by COUGH and encouraging future research for further improvement.Our COUGH dataset is available at https://github. com/sunlab-osu/covid-faq. *Work was done when the first two authors were at OSU. 1 q and a are question and answer fields in an FAQ item.Question1: Should children wear masks?Answer1: In general, children 2 years and older should wear a mask...Appropriate and consistent use of masks...FAQ Bank Question2: Coping with Self-Quarantine Answer2: Remind yourself that difficult emotions are normal during self-quarantine... Query1: Is it possible for human beings to get sick with COVID-19 transmitted to them from animals?Query2: Is it possible to get infected by COVID 19 if I touch food surface packaging?Query Bank Question3: COVID-19是如何在⼈与⼈之间传播的? (How does COVID-19 spread between people?) Answer3: .
Xinliang Frederick Zhang, Heming Sun, Xiang Yue, Simon M. Lin, Huan Sun 0001
EMNLP (1)5
2021 Learning Structural Edits via Incremental Tree Transformations
Ziyu Yao 0002, Frank F. Xu, Huan Sun 0001, Graham Neubig
ICLR4
2021 From Tables to Knowledge: Recent Advances in Table Understanding
abstract
A wealth of human knowledge is expressed in structured tables, across web pages, scientific articles, spreadsheets, and databases. This wealth of knowledge is mirrored by diversity in the vast number of layout structures, content types, formats, and surface forms used to express tables. Recent advances in representation learning and knowledge representation have made progress in exploiting structural regularities in tabular data to unlock this knowledge. In this tutorial, we provide a survey of these advances for a host of table understanding tasks, including table segmentation, semantic typing of cells, transforming tables to knowledge graphs, entity linking, and table retrieval tasks for question answering.
Jay Pujara, Pedro A. Szekely, Huan Sun 0001, Muhao Chen 0001
KDD3
2021 TopNet: Learning from Neural Topic Model to Generate Long Stories
abstract
Long story generation (LSG) is one of the coveted goals in natural language processing. Different from most text generation tasks, LSG requires to output a long story of rich content based on a much shorter text input, and often suffers from information sparsity. In this paper, we propose TopNet to alleviate this problem, by leveraging the recent advances in neural topic modeling to obtain high-quality skeleton words to complement the short input. In particular, instead of directly generating a story, we first learn to map the short text input to a low-dimensional topic distribution (which is pre-assigned by a topic model). Based on this latent topic distribution, we can use the reconstruction decoder of the topic model to sample a sequence of inter-related words as a skeleton for the story. Experiments on two benchmark datasets show that our proposed framework is highly effective in skeleton word selection and significantly outperforms the state-of-the-art models in both automatic evaluation and human evaluation.
Yazheng Yang, Boyuan Pan, Deng Cai 0001, Huan Sun 0001
KDD4
2021 Structure-Grounded Pretraining for Text-to-SQL
abstract
Xiang Deng, Ahmed Hassan Awadallah, Christopher Meek, Oleksandr Polozov, Huan Sun, Matthew Richardson. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Xiang Deng 0001, Ahmed Awadallah 0001, Christopher Meek, Oleksandr Polozov, Huan Sun 0001, Matthew Richardson
NAACL-HLT5
2021 Modeling Context Pair Interaction for Pairwise Tasks on Graphs
abstract
Predicting pairwise relationships between nodes in graphs is a fundamental task in data mining with many real-world applications, such as link prediction on social networks, relation prediction on knowledge graphs, etc. A dominating methodology is to first use advanced graph representation methods to learn generic node representations and then build a pairwise prediction classifier with the target nodes' vectors concatenated as input. However, such methods suffer from low interpretability, as it is difficult to explain why certain relationships are predicted only based on their prediction scores. In this paper, we propose to model the pairwise interactions between neighboring nodes (i.e., contexts) of target pairs. The new formulation enables us to build more appropriate representations for node pairs and gain better model interpretability (by highlighting meaningful interactions). To this end, we introduce a unified framework with two general perspectives, node-centric and pair-centric, about how to model context pair interactions. We also propose a novel pair-centric context interaction model and a new pre-trained embedding, which represents the pair semantics and shows many attractive properties. We test our models on two common pairwise prediction tasks: link prediction task and relation prediction task, and compare them with graph feature-based, embedding-based, and Graph Neural Network (GNN)-based baselines. Our experimental results show the superior performance of the pre-trained pair embeddings and that the pair-centric interaction model outperforms all baselines by a large margin.
Zhen Wang 0041, Bo Zong, Huan Sun 0001
WSDM3
2020 Question-Driven Purchasing Propensity Analysis for Recommendation
abstract
Merchants of e-commerce Websites expect recommender systems to entice more consumption which is highly correlated with the customers' purchasing propensity. However, most existing recommender systems focus on customers' general preference rather than purchasing propensity often governed by instant demands which we deem to be well conveyed by the questions asked by customers. A typical recommendation scenario is: Bob wants to buy a cell phone which can play the game PUBG. He is interested in HUAWEI P20 and asks “can PUBG run smoothly on this phone?” under it. Then our system will be triggered to recommend the most eligible cell phones to him. Intuitively, diverse user questions could probably be addressed in reviews written by other users who have similar concerns. To address this recommendation problem, we propose a novel Question-Driven Attentive Neural Network (QDANN) to assess the instant demands of questioners and the eligibility of products based on user generated reviews, and do recommendation accordingly. Without supervision, QDANN can well exploit reviews to achieve this goal. The attention mechanisms can be used to provide explanations for recommendations. We evaluate QDANN in three domains of Taobao. The results show the efficacy of our method and its superiority over baseline methods.
Long Chen 0007, Ziyu Guan, Qibin Xu, Huan Sun 0001, Guangyue Lu, Deng Cai 0001
AAAI5
2020 Rationalizing Medical Relation Prediction from Corpus-level Statistics
abstract
Nowadays, the interpretability of machine learning models is becoming increasingly important, especially in the medical domain.Aiming to shed some light on how to rationalize medical relation prediction, we present a new interpretable framework inspired by existing theories on how human memory works, e.g., theories of recall and recognition.Given the corpus-level statistics, i.e., a global cooccurrence graph of a clinical text corpus, to predict the relations between two entities, we first recall rich contexts associated with the target entities, and then recognize relational interactions between these contexts to form model rationales, which will contribute to the final prediction.We conduct experiments on a real-world public clinical dataset and show that our framework can not only achieve competitive predictive performance against a comprehensive list of neural baseline models, but also present rationales to justify its prediction.We further collaborate with medical experts deeply to verify the usefulness of our model rationales for clinical decision making 1 .
Zhen Wang 0041, Jennifer Lee, Simon M. Lin, Huan Sun 0001
ACL4
2020 Clinical Reading Comprehension: A Thorough Analysis of the emrQA Dataset
abstract
Machine reading comprehension has made great progress in recent years owing to largescale annotated datasets.In the clinical domain, however, creating such datasets is quite difficult due to the domain expertise required for annotation.Recently, Pampari et al. (2018) tackled this issue by using expert-annotated question templates and existing i2b2 annotations to create emrQA, the first large-scale dataset for question answering (QA) based on clinical notes.In this paper, we provide an indepth analysis of this dataset and the clinical reading comprehension (CliniRC) task.From our qualitative analysis, we find that (i) emrQA answers are often incomplete, and (ii) emrQA questions are often answerable without using domain knowledge.From our quantitative experiments, surprising results include that (iii) using a small sampled subset (5%-20%), we can obtain roughly equal performance compared to the model trained on the entire dataset, (iv) this performance is close to human expert's performance, and (v) BERT models do not beat the best performing base model.Following our analysis of the emrQA, we further explore two desired aspects of CliniRC systems: the ability to utilize clinical domain knowledge and to generalize to unseen questions and contexts.We argue that both should be considered when creating future datasets.1
Xiang Yue, Bernal Jimenez Gutierrez, Huan Sun 0001
ACL3
2020 Clinical Phrase Mining with Language Models
abstract
A vast amount of vital clinical data is available within unstructured texts such as discharge summaries and procedure notes in Electronic Medical Records (EMRs). Automatically transforming such unstructured data into structured units is crucial for effective data analysis in the field of clinical informatics. Recognizing phrases that reveal important medical information in a concise and thorough manner is a fundamental step in this process. Existing systems that are built for opendomain texts are designed to detect mostly non-medical phrases, while tools designed specifically for extracting concepts from clinical texts are not scalable to large corpora and often leave out essential context surrounding those detected clinical concepts. We address these issues by proposing a framework, CliniPhrase, which adapts domain-specific deep neural network based language models (such as ClinicalBERT) to effectively and efficiently extract high-quality phrases from clinical documents with a limited amount of training data. Experimental results on the MIMIC-III dataset show that our method can outperform the current state-of-the-art techniques by up to 18% in terms of F1measure while being very efficient (up to 48 times faster).
Kaushik Mani, Xiang Yue, Bernal Jimenez Gutierrez, Yungui Huang, Simon M. Lin, Huan Sun 0001
BIBM6
2020 Learning a Cost-Effective Annotation Policy for Question Answering
abstract
State-of-the-art question answering (QA) relies upon large amounts of training data for which labeling is time consuming and thus expensive.For this reason, customizing QA systems is challenging.As a remedy, we propose a novel framework for annotating QA datasets that entails learning a cost-effective annotation policy and a semi-supervised annotation scheme.The latter reduces the human effort: it leverages the underlying QA system to suggest potential candidate annotations.Human annotators then simply provide binary feedback on these candidates.Our system is designed such that past annotations continuously improve the future performance and thus overall annotation cost.To the best of our knowledge, this is the first paper to address the problem of annotating questions with minimal annotation cost.We compare our framework against traditional manual annotations in an extensive set of experiments.We find that our approach can reduce up to 21.1% of the annotation cost.
Bernhard Kratzwald, Stefan Feuerriegel, Huan Sun 0001
EMNLP (1)3
2020 An Imitation Game for Learning Semantic Parsers from User Interaction
abstract
Despite the widely successful applications, building a semantic parser is still a tedious process in practice with challenges from costly data annotation and privacy risks.We suggest an alternative, human-in-the-loop methodology for learning semantic parsers directly from users.A semantic parser should be introspective of its uncertainties and prompt for user demonstrations when uncertain.In doing so it also gets to imitate the user behavior and continue improving itself autonomously with the hope that eventually it may become as good as the user in interpreting their questions.To combat the sparsity of demonstrations, we propose a novel annotation-efficient imitation learning algorithm, which iteratively collects new datasets by mixing demonstrated states and confident predictions and retrains the semantic parser in a Dataset Aggregation fashion (Ross et al., 2011).We provide a theoretical analysis of its cost bound and also empirically demonstrate its promising performance on the text-to-SQL problem. 1
Ziyu Yao 0002, Yiqi Tang, Scott Yih, Huan Sun 0001, Yu Su 0001
EMNLP (1)4
2020 EndCold: An End-to-End Framework for Cold Question Routing in Community Question Answering Services
abstract
Routing newly posted questions (a.k.a cold questions) to potential answerers with suitable expertise in Community Question Answering sites (CQAs) is an important and challenging task. The existing methods either focus only on embedding the graph structural information and are less effective for newly posted questions, or adopt manually engineered feature vectors that are not as representative as the graph embedding methods. Therefore, we propose to address the challenge of leveraging heterogeneous graph and textual information for cold question routing by designing an end-to-end framework that jointly learns CQA node embeddings and finds best answerers for cold questions. We conducted extensive experiments to confirm the usefulness of incorporating the textual information from question tags and demonstrate that an end-2-end framework can achieve promising performances on routing newly posted questions asked by both existing users and newly registered users.
Jiankai Sun, Jie Zhao 0013, Huan Sun 0001, Srinivasan Parthasarathy 0001
IJCAI3
2020 Graph embedding on biomedical networks: methods, applications and evaluations
abstract
MOTIVATION: Graph embedding learning that aims to automatically learn low-dimensional node representations, has drawn increasing attention in recent years. To date, most recent graph embedding methods are evaluated on social and information networks and are not comprehensively studied on biomedical networks under systematic experiments and analyses. On the other hand, for a variety of biomedical network analysis tasks, traditional techniques such as matrix factorization (which can be seen as a type of graph embedding methods) have shown promising results, and hence there is a need to systematically evaluate the more recent graph embedding methods (e.g. random walk-based and neural network-based) in terms of their usability and potential to further the state-of-the-art. RESULTS: We select 11 representative graph embedding methods and conduct a systematic comparison on 3 important biomedical link prediction tasks: drug-disease association (DDA) prediction, drug-drug interaction (DDI) prediction, protein-protein interaction (PPI) prediction; and 2 node classification tasks: medical term semantic type classification, protein function prediction. Our experimental results demonstrate that the recent graph embedding methods achieve promising results and deserve more attention in the future biomedical graph analysis. Compared with three state-of-the-art methods for DDAs, DDIs and protein function predictions, the recent graph embedding methods achieve competitive performance without using any biological features and the learned embeddings can be treated as complementary representations for the biological features. By summarizing the experimental results, we provide general guidelines for properly selecting graph embedding methods and setting their hyper-parameters for different biomedical tasks. AVAILABILITY AND IMPLEMENTATION: As part of our contributions in the paper, we develop an easy-to-use Python package with detailed instructions, BioNEV, available at: https://github.com/xiangyue9607/BioNEV, including all source code and datasets, to facilitate studying various graph embedding methods on biomedical tasks. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiang Yue, Zhen Wang 0041, Jingong Huang, Srinivasan Parthasarathy 0001, Soheil Moosavinasab, Yungui Huang, Simon M. Lin, Wen Zhang 0008, Ping Zhang 0016, Huan Sun 0001
Bioinform.10
2020 TURL: Table Understanding through Representation Learning
abstract
Relational tables on the Web store a vast amount of knowledge. Owing to the wealth of such tables, there has been tremendous progress on a variety of tasks in the area of table understanding. However, existing work generally relies on heavily-engineered task-specific features and model architectures. In this paper, we present TURL, a novel framework that introduces the pre-training/fine-tuning paradigm to relational Web tables. During pre-training, our framework learns deep contextualized representations on relational tables in an unsupervised manner. Its universal model design with pre-trained representations can be applied to a wide range of tasks with minimal task-specific fine-tuning. Specifically, we propose a structure-aware Transformer encoder to model the row-column structure of relational tables, and present a new Masked Entity Recovery (MER) objective for pre-training to capture the semantics and knowledge in large-scale unlabeled data. We systematically evaluate TURL with a benchmark consisting of 6 different tasks for table understanding (e.g., relation extraction, cell filling). We show that TURL generalizes well to all tasks and substantially outperforms existing methods in almost all instances.
Xiang Deng 0001, Huan Sun 0001, Alyssa Lees, You Wu 0001, Cong Yu 0001
Proc. VLDB Endow.2
2020 Discerning Influence Patterns with Beta-Poisson Factorization in Microblogging Environments
abstract
Social influence analysis in microblogging services has attracted much attention in recent years. However, most previous studies were focused on measuring users' (topical) influence. Little effort has been made to discern and quantify how a user is influenced. Specifically, the fact that user i retweets a tweet from author j could be either because i is influenced by j (i.e., j is a topical authority), or simply because he is “influenced” by the content (interested in the content). To mine such influence patterns, we propose a novel Bayesian factorization model, dubbed Influence Beta-Poisson Factorization (IBPF). IBPF jointly factorizes the retweet data and tweet content to quantify latent topical factors of user preference, author influence and content influence. It generates every retweet record according to the sum of two causing terms: one representing author influence, and the other one derived from content influence. To control the impact of the two terms, for each user IBPF generates a probability for each latent topic by Beta distribution, indicating how strongly the user cares about the topical authority of the author. We develop an efficient variational inference algorithm for IBPF. We demonstrate the efficacy of IBPF on two public microblogging datasets.
Wei Zhao 0019, Ziyu Guan, Yuhui Huang, Ting-ting Xi, Huan Sun 0001, Zhiheng Wang 0001, Xiaofei He 0001
IEEE Trans. Knowl. Data Eng.5
2019 Answer Identification from Product Reviews for User Questions by Multi-Task Attentive Networks
abstract
Online Shopping has become a part of our daily routine, but it still cannot offer intuitive experience as store shopping. Nowadays, most e-commerce Websites offer a Question Answering (QA) system that allows users to consult other users who have purchased the product. However, users still need to wait patiently for others’ replies. In this paper, we investigate how to provide a quick response to the asker by plausible answer identification from product reviews. By analyzing the similarity and discrepancy between explicit answers and reviews that can be answers, a novel multi-task deep learning method with carefully designed attention mechanisms is developed. The method can well exploit large amounts of user generated QA data and a few manually labeled review data to address the problem. Experiments on data collected from Amazon demonstrate its effectiveness and superiority over competitive baselines.
Long Chen 0007, Ziyu Guan, Wei Zhao 0019, Wanqing Zhao, Zhou Zhao 0001, Huan Sun 0001
AAAI7
2019 Interactive Semantic Parsing for If-Then Recipes via Hierarchical Reinforcement Learning
abstract
Given a text description, most existing semantic parsers synthesize a program in one shot. However, it is quite challenging to produce a correct program solely based on the description, which in reality is often ambiguous or incomplete. In this paper, we investigate interactive semantic parsing, where the agent can ask the user clarification questions to resolve ambiguities via a multi-turn dialogue, on an important type of programs called “If-Then recipes.” We develop a hierarchical reinforcement learning (HRL) based agent that significantly improves the parsing performance with minimal questions to the user. Results under both simulation and human evaluation show that our agent substantially outperforms non-interactive semantic parsers and rule-based agents.1
Ziyu Yao 0002, Xiujun Li, Jianfeng Gao 0001, Brian M. Sadler, Huan Sun 0001
AAAI5
2019 Reinforced Dynamic Reasoning for Conversational Question Generation
abstract
This paper investigates a new task named Conversational Question Generation (CQG) which is to generate a question based on a passage and a conversation history (i.e., previous turns of question-answer pairs).CQG is a crucial task for developing intelligent agents that can drive question-answering style conversations or test user understanding of a given passage.Towards that end, we propose a new approach named Reinforced Dynamic Reasoning (ReDR) network, which is based on the general encoder-decoder framework but incorporates a reasoning procedure in a dynamic manner to better understand what has been asked and what to ask next about the passage.To encourage producing meaningful questions, we leverage a popular question answering (QA) model to provide feedback and fine-tune the question generator using a reinforcement learning mechanism.Empirical results on the recently released CoQA dataset demonstrate the effectiveness of our method in comparison with various baselines and model variants.Moreover, to show the applicability of our method, we also apply it to create multiturn question-answering conversations for passages in SQuAD. * Work done while visiting the Ohio State University.Shelly is in second grade.She is a new student at her school.Shelly's family has lived in many different places.Shelly was born in Florida.Her family moved to Tennessee when she was two years old.When she was four years old, they moved to Texas.They moved from there to Arizona, where they now live.Q1: What grade is Shelly in ?A1: second R1: Shelly is in second grade.Q2: Was she a
Boyuan Pan, Hao Li 0009, Ziyu Yao 0002, Deng Cai 0001, Huan Sun 0001
ACL (1)5
2019 Dynamic Bayesian Metric Learning for Personalized Product Search
abstract
In this paper, we study the problem of personalized product search under streaming scenarios. We address the problem by proposing a Dynamic Bayesian Metric Learning model, abbreviated as DBML, which can collaboratively track the evolutions of latent semantic representations of different categories of entities (i.e., users, products and words) over time in a joint metric space. In particular, unlike previous work using inner-product metric to model the affinities between entities, our DBML is a novel probabilistic metric learning approach that is able to avoid the contradicts, keep the triangle inequality in the latent space, and correctly utilize implicit feedbacks. For inferring dynamic embeddings of the entities, we propose a scalable online inference algorithm, which can jointly learn the latent representations of entities and smooth their changes across time, based on amortized inference. The inferred dynamic semantic representations of entities collaboratively inferred in a unified form by our DBML can benefit not only for improving personalized product search, but also for capturing the affinities between users, products and words. Experimental results on large datasets over a number of applications demonstrate that our DBML outperforms the state-of-the-art algorithms, and can effectively capture the evolutions of semantic representations of different categories of entities over time.
Teng Xiao, Zaiqiao Meng, Huan Sun 0001, Shangsong Liang
CIKM4
2019 Leveraging 2-hop Distant Supervision from Table Entity Pairs for Relation Extraction
abstract
Xiang Deng, Huan Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xiang Deng 0001, Huan Sun 0001
EMNLP/IJCNLP (1)2
2019 Model-based Interactive Semantic Parsing: A Unified Framework and A Text-to-SQL Case Study
abstract
Ziyu Yao, Yu Su, Huan Sun, Wen-tau Yih. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ziyu Yao 0002, Yu Su 0001, Huan Sun 0001, Scott Yih
EMNLP/IJCNLP (1)3
2019 SurfCon: Synonym Discovery on Privacy-Aware Clinical Data
abstract
Unstructured clinical texts contain rich health-related information. To better utilize the knowledge buried in clinical texts, discovering synonyms for a medical query term has become an important task. Recent automatic synonym discovery methods leveraging raw text information have been developed. However, to preserve patient privacy and security, it is usually quite difficult to get access to large-scale raw clinical texts. In this paper, we study a new setting named synonym discovery on privacy-aware clinical data (i.e., medical terms extracted from the clinical texts and their aggregated co-occurrence counts, without raw clinical texts). To solve the problem, we propose a new framework SurfCon that leverages two important types of information in the privacy-aware clinical data, i.e., the surface form information, and the global context information for synonym discovery. In particular, the surface form module enables us to detect synonyms that look similar while the global context module plays a complementary role to discover synonyms that are semantically similar but in different surface forms, and both allow us to deal with the OOV query issue (i.e., when the query is not found in the given data). We conduct extensive experiments and case studies on publicly available privacy-aware clinical data, and show that SurfCon can outperform strong baseline methods by large margins under various settings.
Zhen Wang 0041, Xiang Yue, Soheil Moosavinasab, Yungui Huang, Simon M. Lin, Huan Sun 0001
KDD6
2019 Riker: Mining Rich Keyword Representations for Interpretable Product Question Answering
abstract
This work studies product question answering (PQA) which aims to answer product-related questions based on customer reviews. Most recent PQA approaches adopt end2end semantic matching methodologies, which map questions and answers to a latent vector space to measure their relevance. Such methods often achieve superior performance but it tends to be difficult to interpret why. On the other hand, simple keyword-based search methods exhibit natural interpretability through matched keywords, but often suffer from the lexical gap problem. In this work, we develop a new PQA framework (named Riker) that enjoys the benefits of both interpretability and effectiveness. Riker mines rich keyword representations of a question with two major components, internal word re-weighting and external word association, which predict the importance of each question word and associate the question with outside relevant keywords respectively, and can be jointly trained under weak supervision with large-scale QA pairs. The keyword representations from Riker can be directly used as input to a keyword-based search module, enabling the whole process to be effective while preserving good interpretability. We conduct extensive experiments using Amazon QA and review datasets from 5 different departments, and our results show that Riker substantially outperforms previous state-of-the-art methods in both synthetic settings and real user evaluations. In addition, we compare keyword representations from Riker and those from attention mechanisms popularly used for deep neural networks through case studies, showing that the former are more effective and interpretable.
Jie Zhao 0013, Ziyu Guan, Huan Sun 0001
KDD3
2019 CoaCor: Code Annotation for Code Retrieval with Reinforcement Learning
abstract
To accelerate software development, much research has been performed to help people understand and reuse the huge amount of available code resources. Two important tasks have been widely studied: code retrieval, which aims to retrieve code snippets relevant to a given natural language query from a code base, and code annotation, where the goal is to annotate a code snippet with a natural language description. Despite their advancement in recent years, the two tasks are mostly explored separately. In this work, we investigate a novel perspective of Code annotation for Code retrieval (hence called “CoaCor”), where a code annotation model is trained to generate a natural language annotation that can represent the semantic meaning of a given code snippet and can be leveraged by a code retrieval model to better distinguish relevant code snippets from others. To this end, we propose an effective framework based on reinforcement learning, which explicitly encourages the code annotation model to generate annotations that can be used for the retrieval task. Through extensive experiments, we show that code annotations generated by our framework are much more detailed and more useful for code retrieval, and they can further improve the performance of existing code retrieval models significantly.1
Ziyu Yao 0002, Jayavardhan Reddy Peddamail, Huan Sun 0001
WWW3
2018 Global Relation Embedding for Relation Extraction
abstract
Yu Su, Honglei Liu, Semih Yavuz, Izzeddin Gür, Huan Sun, Xifeng Yan. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Yu Su 0001, Honglei Liu 0001, Semih Yavuz, Izzeddin Gur, Huan Sun 0001, Xifeng Yan
NAACL-HLT5
2018 StaQC: A Systematically Mined Question-Code Dataset from Stack Overflow
abstract
Stack Overflow (SO) has been a great source of natural language questions and their code solutions (i.e., question-code pairs), which are critical for many tasks including code retrieval and annotation. In most existing research, question-code pairs were collected heuristically and tend to have low quality. In this paper, we investigate a new problem of systematically mining question-code pairs from Stack Overflow (in contrast to heuristically collecting them). It is formulated as predicting whether or not a code snippet is a standalone solution to a question. We propose a novel Bi-View Hierarchical Neural Network which can capture both the programming content and the textual context of a code snippet (i.e., two views) to make a prediction. On two manually annotated datasets in Python and SQL domain, our framework substantially outperforms heuristic methods with at least 15% higher F1 and accuracy. Furthermore, we present StaQC (Stack Overflow Question-Code pairs), the largest dataset to date of ~148K Python and ~120K SQL question-code pairs, automatically mined from SO using our framework. Under various case studies, we demonstrate that StaQC can greatly help develop data-hungry models for associating natural language with programming language
Ziyu Yao 0002, Daniel S. Weld, Wei-Peng Chen, Huan Sun 0001
WWW4
2017 An End-to-End Deep Framework for Answer Triggering with a Novel Group-Level Objective
abstract
Given a question and a set of answer candidates, answer triggering determines whether the candidate set contains any correct answers.If yes, it then outputs a correct one.In contrast to existing pipeline methods which first consider individual candidate answers separately and then make a prediction based on a threshold, we propose an end-to-end deep neural network framework, which is trained by a novel group-level objective function that directly optimizes the answer triggering performance.Our objective function penalizes three potential types of error and allows training the framework in an end-to-end manner.Experimental results on the WIKIQA benchmark show that our framework outperforms the state of the arts by a 6.6% absolute gain under F 1 measure 1 .
Jie Zhao 0013, Yu Su 0001, Ziyu Guan, Huan Sun 0001
EMNLP4
2017 Multi-Panel Based Hybrid Beamforming for Multi-User Massive MIMO
abstract
Targeting massive Multiple Input Multiple Output (MIMO) systems, we consider an efficient multi- panel array architecture and propose hybrid beamforming techniques for millimeter wave frequencies. Taking advantages of the multi-panel array, hybrid two-stage processing including coarse and fine beamforming can be carried out in a distributed fashion. Panel-coordinated analog beamforming as well as decoupling only requires channels measured at each panel itself. Non- linear precoding is then applied on a per-panel basis. It is shown that the proposed schemes relax the antenna dimensionality constraint and exhibit a good performance even in an overloaded case, which are practically feasible in both the single-/multi-cell scenarios.
Nuan Song, Pingping Wen, Huan Sun 0001, Tao Yang 0012
GLOBECOM3
2017 Reliable Medical Diagnosis from Crowdsourcing: Discover Trustworthy Answers from Non-Experts
abstract
Nowadays, increasingly more people are receiving medical diagnoses from healthcare-related question answering platforms as people can get diagnoses quickly and conveniently. However, such diagnoses from non-expert crowdsourcing users are noisy or even wrong due to the lack of medical domain knowledge, which can cause serious consequences. To unleash the power of crowdsourcing on healthcare question answering, it is important to identify trustworthy answers and filter out noisy ones from user-generated data. Truth discovery methods estimate user reliability degrees and infer trustworthy information simultaneously, and thus these methods can be adopted to discover trustworthy diagnoses from crowdsourced answers. However, existing truth discovery methods do not take into account the rich semantic meanings of the answers. In the light of this challenge, we propose a method to automatically capture the semantic meanings of answers, where answers are represented as real-valued vectors in the semantic space. To learn such vector representations from noisy user-generated data, we tightly combine the truth discovery and vector learning processes. In this way, the learned vector representations enable truth discovery method to model the semantic relations among answers, and the information trustworthiness inferred by truth discovery can help the procedure of vector representation learning. To demonstrate the effectiveness of the proposed method, we collect a large-scale real-world dataset that involves 219,527 medical diagnosis questions and 23,657 non-expert users. Experimental results show that the proposed method improves the accuracy of identified trustworthy answers due to the successful consideration of answers' semantic meanings. Further, we demonstrate the fast convergence and good scalability of the proposed method, which makes it practical for real-world applications.
Yaliang Li, Nan Du 0001, Chaochun Liu, Yusheng Xie, Wei Fan 0001, Qi Li 0012, Jing Gao 0004, Huan Sun 0001
WSDM8
2017 Overlapped Subarray Based Hybrid Beamforming for Millimeter Wave Multiuser Massive MIMO
abstract
For massive multiple input multiple output systems at millimeter wave (mmWave) bands, we consider an efficient hybrid array architecture, namely, overlapped subarray (OSA), and develop a Unified Low Rank Sparse (ULoRaS) recovery algorithm for hybrid beamforming in downlink multiuser scenarios. The ULoRaS scheme takes advantage of the transmit-receive coordinated beamforming procedure to achieve large array gains. It has no dimensionality constraint and can be applied to the generalized OSA architecture, including both the fully connected and the widely discussed non-OSA cases. It is shown that the proposed ULoRaS algorithm for the novel OSA design is a good compromise of the performance and the required hardware complexity.
Nuan Song, Tao Yang 0012, Huan Sun 0001
IEEE Signal Process. Lett.3
2016 On Generating Characteristic-rich Question Sets for QA Evaluation
abstract
We present a semi-automated framework for constructing factoid question answering (QA) datasets, where an array of question characteristics are formalized, including structure complexity, function, commonness, answer cardinality, and paraphrasing.Instead of collecting questions and manually characterizing them, we employ a reverse procedure, first generating a kind of graph-structured logical forms from a knowledge base, and then converting them into questions.Our work is the first to generate questions with explicitly specified characteristics for QA evaluation.We construct a new QA dataset with over 5,000 logical form-question pairs, associated with answers from the knowledge base, and show that datasets constructed in this way enable finegrained analyses of QA systems.The dataset can be found in https://github.com/ysu1989/GraphQuestions.
Yu Su 0001, Huan Sun 0001, Brian M. Sadler, Mudhakar Srivatsa, Izzeddin Gur, Zenghui Yan, Xifeng Yan
EMNLP2
2016 Coordinated Hybrid Beamforming for Millimeter Wave Multi-User Massive MIMO Systems
abstract
Massive Millimeter Wave (mmWave) Multiple Input Multiple Output (MIMO) systems utilize hybrid beamforming techniques to alleviate the implementation complexity of combining a large number of antennas. To solve the optimization problem in the hybrid design supporting multiple users and multiple data streams per user, we first propose a coordinated Radio Frequency (RF) beamforming technique based on the Generalized Low Rank Approximation of Matrices (GLRAM) approach and then develop an efficient Modified GLRAM (MGLRAM) algorithm. The proposed scheme only requires the information of the composite channel, instead of the complete physical channel matrix which is assumed to be known in the existing literature. It takes advantage of the coordination between the base station and users to achieve a maximal array gain and has no dimensionality constraint. The multiplexing gain is then exploited by applying the Block Diagonalization (BD) technique. It is shown that our proposed scheme is a practical and competing solution, which can be easily applied to both the Time Division Duplex (TDD) and Frequency Division Duplex (FDD) systems.
Nuan Song, Huan Sun 0001, Tao Yang 0012
GLOBECOM2
2016 Augmented LSTM Framework to Construct Medical Self-Diagnosis Android
abstract
Given a health-related question (such as "I have a bad stomach ache. What should I do?"), a medical self-diagnosis Android inquires further information from the user, diagnoses the disease, and ultimately recommend best solutions. One practical challenge to build such an Android is to ask correct questions and obtain most relevant information, in order to correctly pinpoint the most likely causes of health conditions. In this paper, we tackle this challenge, named "relevant symptom question generation": Given a limited set of patient described symptoms in the initial question (e.g., "stomach ache"), what are the most critical symptoms to further ask the patient, in order to correctly diagnose their potential problems? We propose an augmented long short-term memory (LSTM) framework, where the network architecture can naturally incorporate the inputs from embedding vectors of patient described symptoms and an initial disease hypothesis given by a predictive model. Then the proposed framework generates the most important symptom questions. The generation process essentially models the conditional probability to observe a new and undisclosed symptom, given a set of symptoms from a patient as well as an initial disease hypothesis. Experimental results show that the proposed model obtains improvements over alternative methods by over 30% (both precision and mean ordinal distance).
Chaochun Liu, Huan Sun 0001, Nan Du 0001, Shulong Tan, Hongliang Fei, Wei Fan 0001, Tao Yang 0012, Yaliang Li
ICDM2
2016 Distributed Representations of Expertise
abstract
Collaborative networks are common in real life, where domain experts work together to solve tasks issued by customers. How to model the proficiency of experts is critical for us to understand and optimize collaborative networks. Traditional expertise models, such as topic model based methods, cannot capture two aspects of human expertise simultaneously: Specialization (what area an expert is good at?) and Proficiency Level (to what degree?). In this paper, we propose new models to overcome this problem. We embed all historical task data in a lower dimension space and learn vector representations of expertise based on both solved and unsolved tasks. Specifically, in our first model, we assume that each expert will only handle tasks whose difficulty level just matches his/her proficiency level, while experts in the second model accept tasks whose levels are equal to or lower than his/her proficiency level. Experiments on real world datasets show that both models outperform topic model based approaches and standard classifiers such as logistic regression and support vector machine in terms of prediction accuracy. The learnt vector representations can be used to compare expertise in a large organization and optimize expert allocation.
Fangqiu Han, Shulong Tan, Huan Sun 0001, Mudhakar Srivatsa, Deng Cai 0001, Xifeng Yan
SDM3
2016 Design of a Wideband and Dual-Polarized CPW-Fed Monopole Antenna for Future 5G Communications
abstract
As fast development in communication field, sub-6 GHz spectrum band becomes ultra crowd. Even though Carrier Aggregation (CA) and higher order modulation used, available spectrum is still the bottleneck of communication speed and network latency. Nowadays, more and more companies and institutes are paying more attention to over 6GHz communication system research. Recently released news indicate frequency points are mainly on 12GHz, 15GHz, 28GHz, 38GHz even 70GHz, and accompanied working bandwidth is about 400MHz to 1GHz. As working band goes higher, antenna conceptions will also be updated. A CPW-Fed monopole antenna working in KU-band with more than 20% relative bandwidth is hereby introduced, which might be a good candidate for future 5G cellular networks and also satellite based DBS reception system. Proposed antenna has small profile lower than λg/4 with dual-polarization, which is thus suitable for base station and mobile device conceptions.
Huan Sun 0001, Tao Yang 0012, Yann Mahe, Tchanguiz Razban
VTC Fall2
2016 Entity Disambiguation with Linkless Knowledge Bases
abstract
Named Entity Disambiguation is the task of disambiguating named entity mentions in natural language text and link them to their corresponding entries in a reference knowledge base (e.g. Wikipedia). Such disambiguation can help add semantics to plain text and distinguish homonymous entities. Previous research has tackled this problem by making use of two types of context-aware features derived from the reference knowledge base, namely, the context similarity and the semantic relatedness. Both features heavily rely on the cross-document hyperlinks within the knowledge base: the semantic relatedness feature is directly measured via those hyperlinks, while the context similarity feature implicitly makes use of those hyperlinks to expand entity candidates' descriptions and then compares them against the query context. Unfortunately, cross-document hyperlinks are rarely available in many closed domain knowledge bases and it is very expensive to manually add such links. Therefore few algorithms can work well on linkless knowledge bases. In this work, we propose the challenging Named Entity Disambiguation with Linkless Knowledge Bases (LNED) problem and tackle it by leveraging the useful disambiguation evidences scattered across the reference knowledge base. We propose a generative model to automatically mine such evidences out of noisy information. The mined evidences can mimic the role of the missing links and help boost the LNED performance. Experimental results show that our proposed method substantially improves the disambiguation accuracy over the baseline approaches.
Yang Li 0150, Shulong Tan, Huan Sun 0001, Jiawei Han 0001, Dan Roth 0001, Xifeng Yan
WWW3
2016 Table Cell Search for Question Answering
abstract
Tables are pervasive on the Web. Informative web tables range across a large variety of topics, which can naturally serve as a significant resource to satisfy user information needs. Driven by such observations, in this paper, we investigate an important yet largely under-addressed problem: Given millions of tables, how to precisely retrieve table cells to answer a user question. This work proposes a novel table cell search framework to attack this problem. We first formulate the concept of a relational chain which connects two cells in a table and represents the semantic relation between them. With the help of search engine snippets, our framework generates a set of relational chains pointing to potentially correct answer cells. We further employ deep neural networks to conduct more fine-grained inference on which relational chains best match the input question and finally extract the corresponding answer cells. Based on millions of tables crawled from the Web, we evaluate our framework in the open-domain question answering (QA) setting, using both the well-known WebQuestions dataset and user queries mined from Bing search engine logs. On WebQuestions, our framework is comparable to state-of-the-art QA systems based on knowledge bases (KBs), while on Bing queries, it outperforms other systems with a 56.7% relative gain. Moreover, when combined with results from our framework, KB-based QA performance can obtain a relative improvement of 28.1% to 66.7%, demonstrating that web tables supply rich knowledge that might not exist or is difficult to be identified in existing KBs.
Huan Sun 0001, Hao Ma 0001, Xiaodong He 0001, Scott Yih, Yu Su 0001, Xifeng Yan
WWW1
2015 Exploiting Relevance Feedback in Knowledge Graph Search
abstract
The big data era is witnessing a prevalent shift of data from homogeneous to heterogeneous, from isolated to linked. Exemplar outcomes of this shift are a wide range of graph data such as information, social, and knowledge graphs. The unique characteristics of graph data are challenging traditional search techniques like SQL and keyword search. Graph query is emerging as a promising complementary search form. In this paper, we study how to improve graph query by relevance feedback. Specifically, we focus on knowledge graph query, and formulate the graph relevance feedback (GRF) problem. We propose a general GRF framework that is able to (1) tune the original ranking function based on user feedback and (2) further enrich the query itself by mining new features from user feedback. As a consequence, a query-specific ranking function is generated, which is better aligned with the user search intent. Given a newly learned ranking function based on user feedback, we further investigate whether we shall re-rank the existing answers, or choose to search from scratch. We propose a strategy to train a binary classifier to predict which action will be more beneficial for a given query. The GRF framework is applied to searching DBpedia with graph queries derived from YAGO and Wikipedia. Experiment results show that GRF can improve the mean average precision by 80% to 100%.
Yu Su 0001, Shengqi Yang, Huan Sun 0001, Mudhakar Srivatsa, Sue Kase, Michelle Vanni, Xifeng Yan
KDD3
2015 A low complexity scheme for realistic multi-cell downlink coherent joint transmission
abstract
As one of the key LTE-A technologies, coordinated multi-point transmission (CoMP) has been extensively investigated. Among various CoMP schemes, coherent joint transmission scheme (CJT) shows the largest CoMP gains in both cell average performance and cell edge performance. However, for pursuit of the upper bound of CoMP gains, centralized scheduling was adopted in those proposed CJT schemes. With CoMP cluster size increasing, those proposed CoMP schemes face multiple severe challenges, such as the unacceptable computational complexity and more strict latency requirements. In this paper, a novel CJT scheme is proposed, which is based on distributed scheduling and joint transmit precoder design. The complexity of the new CJT scheme is theoretically analyzed and compared with that of centralized scheduling CJT schemes. Moreover, comprehensive performance evaluations of the new CJT scheme are performed by system level simulations in various transmission scenarios. Taking the performance of non-CoMP scheme as baseline, the new CJT scheme can achieve over 35% cell average gain and over 32% cell edge gain, although there is around 10% cell edge CoMP gain loss than that of centralized scheduling CJT scheme.
Huan Sun 0001, Tao Yang 0012
PIMRC1
2015 Performance Evaluation of Distributed Scheduling for Downlink Coherent Joint Transmission
abstract
Coordinated multi-point transmission (CoMP) has been regarded as one of the key LTE-A technologies. Various CoMP schemes has been extensively investigated,and coherent joint transmission scheme (CJT) has been demonstrated to extract the largest CoMP gains in both cell average performance and cell edge performance. In CJT scheme, multiple coordinated cells constitute a supper cell adopting centralized multi-user scheduling and resource allocation. However, with CoMP cluster size increasing, traditional CJT scheme faces multiple severe challenges, such as the unacceptable computational complexity and more strict latency requirements. In this paper, Distributed scheduling scheme is proposed for downlink CJT to alleviate the computational complexity and other system strict requirements. The complexity of the proposed scheme is analyzed. Comprehensive performance evaluations of the proposed scheme are performed by system level simulations in various transmission scenarios. Taking the performance of single cell as the baseline, the proposed scheme can achieve over 35% cell average gain and over 32% cell edge gain with low complexity.
Huan Sun 0001, Tao Yang 0012
VTC Fall1
2015 Open Domain Question Answering via Semantic Enrichment
abstract
Most recent question answering (QA) systems query large-scale knowledge bases (KBs) to answer a question, after parsing and transforming natural language questions to KBs-executable forms (e.g., logical forms). As a well-known fact, KBs are far from complete, so that information required to answer questions may not always exist in KBs. In this paper, we develop a new QA system that mines answers directly from the Web, and meanwhile employs KBs as a significant auxiliary to further boost the QA performance. Specifically, to the best of our knowledge, we make the first attempt to link answer candidates to entities in Freebase, during answer candidate generation. Several remarkable advantages follow: (1) Redundancy among answer candidates is automatically reduced. (2) The types of an answer candidate can be effortlessly determined by those of its corresponding entity in Freebase. (3) Capitalizing on the rich information about entities in Freebase, we can develop semantic features for each answer candidate after linking them to Freebase. Particularly, we construct answer-type related features with two novel probabilistic models, which directly evaluate the appropriateness of an answer candidate's types under a given question. Overall, such semantic features turn out to play significant roles in determining the true answers from the large answer candidate pool. The experimental results show that across two testing datasets, our QA system achieves an 18%~54% improvement under F_1 metric, compared with various existing QA systems.
Huan Sun 0001, Hao Ma 0001, Scott Yih, Chen-Tse Tsai, Jingjing Liu 0001, Ming-Wei Chang
WWW1
2015 Fine-Grained Knowledge Sharing in Collaborative Environments
abstract
In collaborative environments, members may try to acquire similar information on the web in order to gain knowledge in one domain. For example, in a company several departments may successively need to buy business intelligence software and employees from these departments may have studied online about different business intelligence tools and their features independently. It will be productive to get them connected and share learned knowledge. We investigate fine-grained knowledge sharing in collaborative environments. We propose to analyze members' web surfing data to summarize the fine-grained knowledge acquired by them. A two-step framework is proposed for mining fine-grained knowledge: (1) web surfing data is clustered into tasks by a nonparametric generative model; (2) a novel discriminative infinite Hidden Markov Model is developed to mine fine-grained aspects in each task. Finally, the classic expert search method is applied to the mined results to find proper members for knowledge sharing. Experiments on web surfing data collected from our lab at UCSB and IBM show that the fine-grained aspect mining framework works as expected and outperforms baselines. When it is integrated with expert search, the search accuracy improves significantly, in comparison with applying the classic expert search method directly on web surfing data.
Ziyu Guan, Shengqi Yang, Huan Sun 0001, Mudhakar Srivatsa, Xifeng Yan
IEEE Trans. Knowl. Data Eng.3
2014 Analyzing expert behaviors in collaborative networks
abstract
Collaborative networks are composed of experts who cooperate with each other to complete specific tasks, such as resolving problems reported by customers. A task is posted and subsequently routed in the network from an expert to another until being resolved. When an expert cannot solve a task, his routing decision (i.e., where to transfer a task) is critical since it can significantly affect the completion time of a task. In this work, we attempt to deduce the cognitive process of task routing, and model the decision making of experts as a generative process where a routing decision is made based on mixed routing patterns.
Huan Sun 0001, Mudhakar Srivatsa, Shulong Tan, Yang Li 0150, Lance M. Kaplan, Shu Tao, Xifeng Yan
KDD1
2014 Network mining and analysis for social applications
abstract
The recent blossom of social network and communication services in both public and corporate settings have generated a staggering amount of network data of all kinds. Unlike the bio-networks and the chemical compound graph data often used in traditional network mining and analysis, the new network data grown out of the social applications are characterized by their rich attributes, high heterogeneity, enormous sizes and complex patterns of various semantic meanings, all of which have posed significant research challenges to the graph/network mining community. In this tutorial, we aim to examine some recent advances in network mining and analysis for social applications, covering a diverse collection of methodologies and applications from the perspectives of event, relationship, collaboration, and network pattern. We would present the problem settings, the challenges, the recent research advances and some future directions for each perspective. Topics include but are not limited to correlation mining, iceberg finding, anomaly detection, relationship discovery, information flow, task routing, and pattern mining.
Feida Zhu 0001, Huan Sun 0001, Xifeng Yan
KDD2
2014 A Probabilistic Approach to Uncovering Attributed Graph Anomalies
abstract
Uncovering subgraphs with an abnormal distribution of attributes reveals much insight into network behaviors. For example in social or communication networks, diseases or intrusions usually do not propagate uniformly, which makes it critical to find anomalous regions with high concentrations of a specific disease or intrusion. In this paper, we introduce a probabilistic model to identify anomalous subgraphs containing a significantly different percentage of a certain vertex attribute, such as a specific disease or an intrusion, compared to the rest of the graph. Our framework, gAnomaly, models generative processes of vertex attributes and divides the graph into regions that are governed by background and anomaly processes. Two types of regularizers are employed to smoothen the regions and to facilitate vertex assignment. We utilize deterministic annealing EM to learn the model parameters, which is less initialization-dependent and better at avoiding local optima. In order to find fine-grained anomalies, an iterative procedure is further proposed. Experiments show gAnomaly outperforms a state-of-the-art algorithm at uncovering anomalous subgraphs in attributed graphs.
Huan Sun 0001, Kyle C. Chipman, Jemin George, Xifeng Yan
SDM2
2014 SLQ: a user-friendly graph querying system
abstract
Querying complex graph databases such as knowledge graphs is a challenging task for non-professional users. In this demo, we present SLQ, a user-friendly graph querying system enabling schemales and structures graph querying, where a user need not describe queries precisely as required by most databases. SLQ system combines searching and ranking: it leverages a set of transformation functions, including abbreviation, ontology, synonym, etc., that map keywords and linkages from a query to their matches in a data graph, based on an automatically learned ranking model. To help users better understand search results at different levels of granularity, it supports effective result summarization with "drill-down" and "roll-up" operations. Better still, the architecture of SLQ is elastic for new transformation functions, query logs and user feedback, to iteratively refine the ranking model. SLQ significantly improves the usability of graph querying. This demonstration highlights (1) SLQ can automatically learn an effective ranking model, without assuming manually labeled training examples, (2) it can efficiently return top ranked matches over noisy, large data graphs, (3) it can summarize the query matches to help users easily access, explore and understand query results, and (4) its GUI can interact with users to help them construct queries, explore data graphs and inspect matches in a user-friendly manner.
Shengqi Yang, Yanan Xie, Yinghui Wu 0001, Huan Sun 0001, Jian Wu 0001, Xifeng Yan
SIGMOD Conference5
2014 Schemaless and Structureless Graph Querying
abstract
Querying complex graph databases such as knowledge graphs is a challenging task for non-professional users. Due to their complex schemas and variational information descriptions, it becomes very hard for users to formulate a query that can be properly processed by the existing systems. We argue that for a user-friendly graph query engine, it must support various kinds of transformations such as synonym, abbreviation, and ontology. Furthermore, the derived query results must be ranked in a principled manner. In this paper, we introduce a novel framework enabling schemaless and structureless graph querying (SLQ), where a user need not describe queries precisely as required by most databases. The query engine is built on a set of transformation functions that automatically map keywords and linkages from a query to their matches in a graph. It automatically learns an effective ranking model, without assuming manually labeled training examples, and can efficiently return top ranked matches using graph sketch and belief propagation. The architecture of SLQ is elastic for "plug-in" new transformation functions and query logs. Our experimental results show that this new graph querying paradigm is promising: It identifies high-quality matches for both keyword and graph queries over real-life knowledge graphs, and outperforms existing methods significantly in terms of effectiveness and efficiency.
Shengqi Yang, Yinghui Wu 0001, Huan Sun 0001, Xifeng Yan
Proc. VLDB Endow.3
2014 Interpreting the Public Sentiment Variations on Twitter
abstract
Millions of users share their opinions on Twitter, making it a valuable platform for tracking and analyzing public sentiment. Such tracking and analysis can provide critical information for decision making in various domains. Therefore it has attracted attention in both academia and industry. Previous research mainly focused on modeling and tracking public sentiment. In this work, we move one step further to interpret sentiment variations. We observed that emerging topics (named foreground topics) within the sentiment variation periods are highly related to the genuine reasons behind the variations. Based on this observation, we propose a Latent Dirichlet Allocation (LDA) based model, Foreground and Background LDA (FB-LDA), to distill foreground topics and filter out longstanding background topics. These foreground topics can give potential interpretations of the sentiment variations. To further enhance the readability of the mined reasons, we select the most representative tweets for foreground topics and develop another generative model called Reason Candidate and Background LDA (RCB-LDA) to rank them with respect to their “popularity” within the variation period. Experimental results show that our methods can effectively find foreground topics and rank reason candidates. The proposed models can also be applied to other tasks such as finding topic differences between two sets of documents.
Shulong Tan, Yang Li 0150, Huan Sun 0001, Ziyu Guan, Xifeng Yan, Jiajun Bu, Chun Chen 0001, Xiaofei He 0001
IEEE Trans. Knowl. Data Eng.3
2013 Noise-Resistant Bicluster Recognition
abstract
Biclustering is crucial in finding co-expressed genes and their associated conditions in gene expression data. While various biclustering algorithms (e.g., combinatorial, probabilistic modelling, and matrix factorization) have been proposed and constantly improved in the past decade, data noise and bicluster overlaps make biclustering a still challenging task. It becomes difficult to further improve biclustering performance, without resorting to a new approach. Inspired by the recent progress in unsupervised feature learning using deep neural networks, in this work, we propose a novel model for biclustering, named Auto Decoder (AD), by relating biclusters to features and leveraging a neural network that is able to automatically learn features from the input data. To suppress severe noise present in gene expression data, we introduce a non-uniform signal recovery mechanism: Instead of reconstructing the whole input data to capture the bicluster patterns, AD weighs the zero and non-zero parts of the input data differently and is more flexible in dealing with different types of noise. AD is also properly regularized to deal with bicluster overlaps. To the best of our knowledge, this is the first biclustering algorithm that leverages neural network techniques to recover overlapped biclusters hidden in noisy gene expression data. We compared our approach with four state-of-the-art biclustering algorithms on both synthetic and real datasets. On three out of the four real datasets, AD significantly outperforms the other approaches. On controlled synthetic datasets, AD performs the best when noise level is beyond 15%.
Huan Sun 0001, Gengxin Miao, Xifeng Yan
ICDM1
2013 Synthetic review spamming and defense
abstract
Online reviews have been popularly adopted in many applications. Since they can either promote or harm the reputation of a product or a service, buying and selling fake reviews becomes a profitable business and a big threat. In this paper, we introduce a very simple, but powerful review spamming technique that could fail the existing feature-based detection algorithms easily. It uses one truthful review as a template, and replaces its sentences with those from other reviews in a repository. Fake reviews generated by this mechanism are extremely hard to detect: Both the state-of-the-art computational approaches and human readers acquire an error rate of 35%-48%, just slightly better than a random guess. While it is challenging to detect such fake reviews, we have made solid progress in suppressing them. A novel defense method that leverages the difference of semantic flows between synthetic and truthful reviews is developed, which is able to reduce the detection error rate to approximately 22%, a significant improvement over the performance of existing approaches. Nevertheless, it is still a challenging research task to further decrease the error rate.
Huan Sun 0001, Alex Morales, Xifeng Yan
KDD1
2012 Enhanced Multiuser Eigenmode Transmission for Joint Frequency-Spatial Resource Allocation in OFDM-MIMO Downlink Systems
abstract
Orthogonal frequency-division multiplexing (OFDM) and multiple-input-multiple-output (MIMO) are two key technologies adopted in future wireless communication systems, such as LTE/LTE-A. There are many papers addressed the resource allocation problems for OFDM and MIMO systems independently. However few discussions targets on the joint frequency-spatial resource allocation problem, particularly on eigenmode scheduling level. In this paper, we try to propose novel resource allocation schemes on eigenmode selection level, with the joint consideration of precoding problem in spatial domain, eigenmode scheduling and power allocation cross spatial and frequency domains.
Fanglei Sun, Huan Sun 0001, Mingli You, Tao Yang 0012
VTC Spring2
2011 A priority-aware hybrid multi-hop energy saving strategy for inter-eNB scenario 2
abstract
In this paper, we investigate the problem of energy saving (ES) for wireless access networks, especially for the inter-eNB scenario 2 which is a typical non-overlapped scenario specified by 3rd generation partnership project (3GPP). Due to the uneven distribution of the traffic load both over time and over cells, energy efficiency can be improved by switching off some eNBs or reducing their transmission power according to the actual traffic demand. Considering this characteristic of wireless networks, we propose a flexible and efficient solution based on the load information of each cell and the idea of multi-hop load release. The proposed solution is flexible to support different heterogeneous network architecture and different optimization cases with less signaling overhead. Moreover, the gradual multi-hop load release mechanism can efficiently release the load to the whole network and then maximize the total energy efficiency. Simulation results demonstrate that the ES gain of the proposed solution can reach about 25% with appropriate values of different thresholds for event triggering.
Chongxian Zhong, Tao Yang 0012, Huan Sun 0001
PIMRC3