Heng-Tze Cheng

dblp:30/8739 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
6since 2021 · last 2024
0009-0007-3845-8796ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorComputer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 57% Reinforcement learning · 16% Efficient and distributed learning · 12%
Databases, data mining, and information retrieval
7 papers
Information retrieval · 57% Recommender systems · 24% Machine learning and data management · 13%
Human-computer interaction and pervasive computing
3 papers
Ubiquitous computing and smart environments · 83% Wearable and physiological sensing · 17%

Topics — the 30 heaviest of 36, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
prompting
1.522024
SELF-DISCOVER: Large Language Models Self-Compose Reasoning Structures · NeurIPS 2024
Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models · ICLR 2024
Information retrieval
personalized search
1.022022
Multi-Resolution Attention for Personalized Item Search · WSDM 2022
End-to-End Deep Attentive Personalized Item Retrieval for Online Content-sharing Platforms · WWW 2020
Natural language and speech › Language models and text generation
alignment
0.812024
Controlled Decoding from Language Models · ICML 2024
Natural language and speech › Language models and text generation › decoding
controlled decoding
0.812024
Controlled Decoding from Language Models · ICML 2024
Machine learning › Reinforcement learning › policy optimization
KL-regularized RL
0.812024
Controlled Decoding from Language Models · ICML 2024
Natural language and speech › Language models and text generation
large language model reasoning
0.812024
SELF-DISCOVER: Large Language Models Self-Compose Reasoning Structures · NeurIPS 2024
Machine learning › Learning paradigms
multi-task learning
0.612022
HyperPrompt: Prompt-based Task-Conditioning of Transformers · ICML 2022
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.612022
HyperPrompt: Prompt-based Task-Conditioning of Transformers · ICML 2022
Natural language and speech › Language models and text generation
prompt tuning
0.612022
HyperPrompt: Prompt-based Task-Conditioning of Transformers · ICML 2022
Recommender systems › neural recommendation
attention-based recommendation
0.612022
Multi-Resolution Attention for Personalized Item Search · WSDM 2022
Information retrieval › retrieval models › neural retrieval
neural ranking model
0.612022
Multi-Resolution Attention for Personalized Item Search · WSDM 2022
Information retrieval › e-commerce search
personalized product search
0.612022
Multi-Resolution Attention for Personalized Item Search · WSDM 2022
Information retrieval
ranking
0.612022
Multi-Resolution Attention for Personalized Item Search · WSDM 2022
Ubiquitous computing and smart environments › context recognition
activity recognition
0.532014
Nonparametric discovery of human routines from sensor data · PerCom 2014
NuActiv: recognizing unseen new activities using semantic attribute-based learning · MobiSys 2013
Towards zero-shot learning for human activity recognition using semantic attribute sequence model · UbiComp 2013
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
error correction
0.512021
Mondegreen: A Post-Processing Solution to Speech Recognition Error Correction for Voice Search Queries · KDD 2021
Information retrieval
query processing
0.512021
Mondegreen: A Post-Processing Solution to Speech Recognition Error Correction for Voice Search Queries · KDD 2021
Information retrieval › e-commerce search
item retrieval
0.412020
End-to-End Deep Attentive Personalized Item Retrieval for Online Content-sharing Platforms · WWW 2020
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
0.412019
SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets · IJCAI 2019
Machine learning › Reinforcement learning
value-based reinforcement learning
0.412019
SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets · IJCAI 2019
Recommender systems
reinforcement-learning-based recommendation
0.412019
SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets · IJCAI 2019
Recommender systems › interactive recommendation
slate recommendation
0.412019
SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets · IJCAI 2019
Machine learning › Efficient and distributed learning
distributed training
0.312017
TensorFlow Estimators: Managing Simplicity vs. Flexibility in High-Level Machine Learning Frameworks · KDD 2017
Machine learning › Efficient and distributed learning
production machine learning
0.312017
TensorFlow Estimators: Managing Simplicity vs. Flexibility in High-Level Machine Learning Frameworks · KDD 2017
Machine learning and data management
machine learning lifecycle management
0.312017
TFX: A TensorFlow-Based Production-Scale Machine Learning Platform · KDD 2017
Machine learning and data management › machine learning systems
machine learning platform
0.312017
TFX: A TensorFlow-Based Production-Scale Machine Learning Platform · KDD 2017
Natural language and speech › Language models and text generation › large language model inference
inference-time decoding
0.212024
Controlled Decoding from Language Models · ICML 2024
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering
0.212024
Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models · ICLR 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
reasoning structure
0.212024
SELF-DISCOVER: Large Language Models Self-Compose Reasoning Structures · NeurIPS 2024
Data mining › text mining › topic modeling
nonparametric topic model
0.212014
Nonparametric discovery of human routines from sensor data · PerCom 2014
Data mining › text mining
topic modeling
0.212014
Nonparametric discovery of human routines from sensor data · PerCom 2014

Methods — techniques the papers use, named apart from their topics

value function · 0.8self-discovery · 0.8self-consistency · 0.8reinforcement learning · 0.8prefix scorer · 0.8large language model prompting · 0.8in-context learning · 0.8chain-of-thought · 0.8soft-thresholding · 0.6prompt tuning · 0.6multi-head attention · 0.6hypernetwork · 0.6text-space correction · 0.5extreme multi-class softmax · 0.4embedding · 0.4attention mechanism · 0.4temporal difference learning · 0.4q-learning · 0.4
YearPublicationVenuePosition
2024 Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models
abstract
We present STEP-BACK PROMPTING, a simple prompting technique that enables LLMs to do abstractions to derive high-level concepts and first principles from instances containing specific details. Using the concepts and principles to guide reasoning, LLMs significantly improve their abilities in following a correct reasoning path towards the solution. We conduct experiments of STEP-BACK PROMPTING with PaLM-2L, GPT-4 and Llama2-70B models, and observe substantial performance gains on various challenging reasoning-intensive tasks including STEM, Knowledge QA, and Multi-Hop Reasoning. For instance, STEP-BACK PROMPTING improves PaLM-2L performance on MMLU (Physics and Chemistry) by 7% and 11% respectively, TimeQA by 27%, and MuSiQue by 7%.
Huaixiu Steven Zheng, Swaroop Mishra, Heng-Tze Cheng, Ed H. Chi, Quoc V. Le, Denny Zhou
ICLR4
2024 Controlled Decoding from Language Models
abstract
KL-regularized reinforcement learning (RL) is a popular alignment framework to control the language model responses towards high reward outcomes. We pose a tokenwise RL objective and propose a modular solver for it, called *controlled decoding (CD)*. CD exerts control through a separate *prefix scorer* module, which is trained to learn a value function for the reward. The prefix scorer is used at inference time to control the generation from a frozen base model, provably sampling from a solution to the RL objective. We empirically demonstrate that CD is effective as a control mechanism on popular benchmarks. We also show that prefix scorers for multiple rewards may be combined at inference time, effectively solving a multi-objective RL problem with no additional training. We show that the benefits of applying CD transfer to an unseen base model with no further tuning as well. Finally, we show that CD can be applied in a blockwise decoding fashion at inference-time, essentially bridging the gap between the popular best-of-$K$ strategy and tokenwise control through reinforcement learning. This makes CD a promising approach for alignment of language models.
Sidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li, Yanping Huang, Heng-Tze Cheng, Trevor Strohman, Jilin Chen, Alex Beutel, Ahmad Beirami
ICML8
2024 SELF-DISCOVER: Large Language Models Self-Compose Reasoning Structures
abstract
We introduce SELF-DISCOVER, a general framework for LLMs to self-discover the task-intrinsic reasoning structures to tackle complex reasoning problems that are challenging for typical prompting methods. Core to the framework is a self-discovery process where LLMs select multiple atomic reasoning modules such as critical thinking and step-by-step thinking, and compose them into an explicit reasoning structure for LLMs to follow during decoding. SELF-DISCOVER substantially improves GPT-4 and PaLM 2’s performance on challenging reasoning benchmarks such as BigBench-Hard, grounded agent reasoning, and MATH, by as much as 32% compared to Chain of Thought (CoT). Furthermore, SELF-DISCOVER outperforms inference-intensive methods such as CoT-Self-Consistency by more than 20%, while requiring 10-40x fewer inference compute. Finally, we show that the self-discovered reasoning structures are universally applicable across model families: from PaLM 2-L to GPT-4, and from GPT-4 to Llama2, and share commonalities with human reasoning patterns.
Jay Pujara, Xiang Ren 0001, Heng-Tze Cheng, Quoc V. Le, Ed H. Chi, Denny Zhou, Swaroop Mishra, Huaixiu Steven Zheng
NeurIPS5
2022 HyperPrompt: Prompt-based Task-Conditioning of Transformers
abstract
Prompt-Tuning is a new paradigm for finetuning pre-trained language models in a parameter efficient way. Here, we explore the use of HyperNetworks to generate hyper-prompts: we propose HyperPrompt, a novel architecture for prompt-based task-conditioning of self-attention in Transformers. The hyper-prompts are end-to-end learnable via generation by a HyperNetwork. HyperPrompt allows the network to learn task-specific feature maps where the hyper-prompts serve as task global memories for the queries to attend to, at the same time enabling flexible information sharing among tasks. We show that HyperPrompt is competitive against strong multi-task learning baselines with as few as 0.14% of additional task-conditioning parameters, achieving great parameter and computational efficiency. Through extensive empirical experiments, we demonstrate that HyperPrompt can achieve superior performances over strong T5 multi-task learning baselines and parameter-efficient adapter variants including Prompt-Tuning and HyperFormer++ on Natural Language Understanding benchmarks of GLUE and SuperGLUE across many model sizes.
Huaixiu Steven Zheng, Yi Tay, Jai Gupta 0001, Vamsi Aribandi, Zhe Zhao 0001, YaGuang Li, Donald Metzler, Heng-Tze Cheng, Ed H. Chi
ICML11
2022 Multi-Resolution Attention for Personalized Item Search
abstract
Personalized item search has become an essential tool for online platforms---where users interact with a large corpus of items (e.g., click, purchase, like) via a search query---to provide their users with a more satisfactory search experience. The record (or history) of users' past interactions serves as a valuable asset to achieve personalization. While user history data can span over a long period of time, only a part of the history is relevant to a user's current search intent. Moreover, since historical interactions take place at aperiodic points in time, modeling their relevance to the current search query entangles complex temporal dependencies. We propose multi-resolution attention to address these challenges for personalized item search. Our approach captures higher-order temporal relations between user queries and their history across several temporal subspaces (i.e., resolutions), each corresponding to distinct temporal ranges with adaptive time boundaries that are also learned directly from data. We achieve this by coupling the conventional multi-head attention module with a differentiable soft-thresholding mechanism, which essentially operates as a masking function in the temporal domain. Comparisons with strong baselines on an open-source benchmark dataset confirm the efficacy of our approach.
Furkan Kocayusufoglu, Anima Singh, George Roumpos, Heng-Tze Cheng, Sagar Jain, Ed H. Chi, Ambuj K. Singh
WSDM5
2021 Mondegreen: A Post-Processing Solution to Speech Recognition Error Correction for Voice Search Queries
abstract
As more and more online search queries come from voice, automatic speech recognition becomes a key component to deliver relevant search results. Errors introduced by automatic speech recognition (ASR) lead to irrelevant search results returned to the user, thus causing user dissatisfaction. In this paper, we introduce an approach, "Mondegreen", to correct voice queries in text space without depending on audio signals, which may not always be available due to system constraints or privacy or bandwidth (for example, some ASR systems run on-device) considerations. We focus on voice queries transcribed via several proprietary commercial ASR systems. These queries come from users making internet, or online service search queries. We first present an analysis showing how different the language distribution coming from user voice queries is from that in traditional text corpora used to train off-the-shelf ASR systems. We then demonstrate that Mondegreen can achieve significant improvements in increased user interaction by correcting user voice queries in one of the largest search systems in Google. Finally, we see Mondegreen as complementing existing highly-optimized production ASR systems, which may not be frequently retrained and thus lag behind due to vocabulary drifts.
Sukhdeep S. Sodhi, Ellie Ka In Chio, Ambarish Jash, Santiago Ontañón, Ajit Apte, Ayooluwakunmi Jeje, Dima Kuzmin, Harry Fung, Heng-Tze Cheng, Jon Effrat, Tarush Bali, Nitin Jindal, Sarvjeet Singh, Senqiang Zhou, Tameen Khan, Amol Wankhede, Moustafa Farid Alzantot, Allen Wu, Tushar Chandra
KDD10
2020 Zero-Shot Heterogeneous Transfer Learning from Recommender Systems to Cold-Start Search Retrieval
abstract
Many recent advances in neural information retrieval models, which predict top-K items given a query, learn directly from a large training set of (query, item) pairs. However, they are often insufficient when there are many previously unseen (query, item) combinations, often referred to as the cold start problem. Furthermore, the search system can be biased towards items that are frequently shown to a query previously, also known as the 'rich get richer' (a.k.a. feedback loop) problem. In light of these problems, we observed that most online content platforms have both a search and a recommender system that, while having heterogeneous input spaces, can be connected through their common output item space and a shared semantic representation. In this paper, we propose a new Zero-Shot Heterogeneous Transfer Learning framework that transfers learned knowledge from the recommender system component to improve the search component of a content platform. First, it learns representations of items and their natural-language features by predicting (item, item) correlation graphs derived from the recommender system as an auxiliary task. Then, the learned representations are transferred to solve the target search retrieval task, performing query-to-item prediction without having seen any (query, item) pairs in training. We conduct online and offline experiments on one of the world's largest search and recommender systems from Google, and present the results and lessons learned. We demonstrate that the proposed approach can achieve high performance on offline search retrieval tasks, and more importantly, achieved significant improvements on relevance and user interactions over the highly-optimized production system in online experiments.
Ellie Ka In Chio, Heng-Tze Cheng, Steffen Rendle, Dima Kuzmin, Ritesh Agarwal, Li Zhang 0001, John R. Anderson, Sarvjeet Singh, Tushar Chandra, Ed H. Chi, Alex Soares, Nitin Jindal
CIKM3
2020 End-to-End Deep Attentive Personalized Item Retrieval for Online Content-sharing Platforms
abstract
Modern online content-sharing platforms host billions of items like music, videos, and products uploaded by various providers for users to discover items of their interests. To satisfy the information needs, the task of effective item retrieval (or item search ranking) given user search queries has become one of the most fundamental problems to online content-sharing platforms. Moreover, the same query can represent different search intents for different users, so personalization is also essential for providing more satisfactory search results. Different from other similar research tasks, such as ad-hoc retrieval and product retrieval with copious words and reviews, items in content-sharing platforms usually lack sufficient descriptive information and related meta-data as features. In this paper, we propose the end-to-end deep attentive model (EDAM) to deal with personalized item retrieval for online content-sharing platforms using only discrete personal item history and queries. Each discrete item in the personal item history of a user and its content provider are first mapped to embedding vectors as continuous representations. A query-aware attention mechanism is then applied to identify the relevant contexts in the user history and construct the overall personal representation for a given query. Finally, an extreme multi-class softmax classifier aggregates the representations of both query and personal item history to provide personalized search results. We conduct extensive experiments on a large-scale real-world dataset with hundreds of million users from a large video media platform at Google. The experimental results demonstrate that our proposed approach significantly outperforms several competitive baseline methods. It is also worth mentioning that this work utilizes a massive dataset from a real-world commercial content-sharing platform for personalized item retrieval to provide more insightful analysis from the industrial aspects.
Jyun-Yu Jiang, George Roumpos, Heng-Tze Cheng, Xinyang Yi, Ed H. Chi, Harish Ganapathy, Nitin Jindal, Wei Wang 0010
WWW4
2019 SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets
abstract
Reinforcement learning methods for recommender systems optimize recommendations for long-term user engagement. However, since users are often presented with slates of multiple items---which may have interacting effects on user choice---methods are required to deal with the combinatorics of the RL action space. We develop SlateQ, a decomposition of value-based temporal-difference and Q-learning that renders RL tractable with slates. Under mild assumptions on user choice behavior, we show that the long-term value (LTV) of a slate can be decomposed into a tractable function of its component item-wise LTVs. We demonstrate our methods in simulation, and validate the scalability and effectiveness of decomposed TD-learning on YouTube.
Eugene Ie, Vihan Jain, Sanmit Narvekar, Ritesh Agarwal, Rui Wu 0020, Heng-Tze Cheng, Tushar Chandra, Craig Boutilier
IJCAI7
2017 TFX: A TensorFlow-Based Production-Scale Machine Learning Platform
abstract
Creating and maintaining a platform for reliably producing and deploying machine learning models requires careful orchestration of many components---a learner for generating models based on training data, modules for analyzing and validating both data as well as models, and finally infrastructure for serving models in production. This becomes particularly challenging when data changes over time and fresh models need to be produced continuously. Unfortunately, such orchestration is often done ad hoc using glue code and custom scripts developed by individual teams for specific use cases, leading to duplicated effort and fragile systems with high technical debt.
Denis Baylor, Eric Breck, Heng-Tze Cheng, Noah Fiedel, Chuan Yu Foo, Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc 0001, Chiu Yuen Koo, Lukasz Lew, Clemens Mewald, Akshay Naresh Modi, Neoklis Polyzotis, Sukriti Ramesh, Sudip Roy 0002, Steven Euijong Whang, Martin Wicke, Jarek Wilkiewicz, Martin Zinkevich
KDD3
2017 TensorFlow Estimators: Managing Simplicity vs. Flexibility in High-Level Machine Learning Frameworks
abstract
We present a framework for specifying, training, evaluating, and deploying machine learning models. Our focus is on simplifying cutting edge machine learning for practitioners in order to bring such technologies into production. Recognizing the fast evolution of the field of deep learning, we make no attempt to capture the design space of all possible model architectures in a domain-specific language (DSL) or similar configuration language. We allow users to write code to define their models, but provide abstractions that guide developers to write models in ways conducive to productionization. We also provide a unifying Estimator interface, making it possible to write downstream infrastructure (e.g. distributed training, hyperparameter tuning) independent of the model implementation.
Heng-Tze Cheng, Zakaria Haque, Lichan Hong, Mustafa Ispir, Clemens Mewald, Illia Polosukhin, George Roumpos, D. Sculley, Jamie Smith, David Soergel, Yuan Tang 0001, Philipp Tucker, Martin Wicke, Cassandra Xia, Jianwei Xie
KDD1
2014 Nonparametric discovery of human routines from sensor data
abstract
People engage in routine behaviors. Automatic routine discovery goes beyond low-level activity recognition such as sitting or standing and analyzes human behaviors at a higher level (e.g., commuting to work). With recent developments in ubiquitous sensor technologies, it becomes easier to acquire a massive amount of sensor data. One main line of research is to mine human routines from sensor data using parametric topic models such as latent Dirichlet allocation. The main shortcoming of parametric models is that it assumes a fixed, pre-specified parameter regardless of the data. Choosing an appropriate parameter usually requires an inefficient trial-and-error model selection process. Furthermore, it is even more difficult to find optimal parameter values in advance for personalized applications. In this paper, we present a novel nonparametric framework for human routine discovery that can infer high-level routines without knowing the number of latent topics beforehand. Our approach is evaluated on public datasets in two routine domains: a 34-daily-activity dataset and a transportation mode dataset. Experimental results show that our nonparametric framework can automatically learn the appropriate model parameters from sensor data without any form of model selection procedure and can outperform traditional parametric approaches for human routine discovery tasks.
Feng-Tso Sun, Heng-Tze Cheng, Cynthia Kuo, Martin L. Griss
PerCom3
2013 Towards zero-shot learning for human activity recognition using semantic attribute sequence model
abstract
Understanding human activities is important for user-centric and context-aware applications. Previous studies showed promising results using various machine learning algorithms. However, most existing methods can only recognize the activities that were previously seen in the training data. In this paper, we present a new zero-shot learning framework for human activity recognition that can recognize an unseen new activity even when there are no training samples of that activity in the dataset. We propose a semantic attribute sequence model that takes into account both the hierarchical and sequential nature of activity data. Evaluation on datasets in two activity domains show that the proposed zero-shot learning approach achieves 70-75% precision and recall recognizing unseen new activities, and outperforms supervised learning with limited labeled data for the new classes.
Heng-Tze Cheng, Martin L. Griss, Di You
UbiComp1
2013 NuActiv: recognizing unseen new activities using semantic attribute-based learning
abstract
We study the problem of how to recognize a new human activity when we have never seen any training example of that activity before. Recognizing human activities is an essential element for user-centric and context-aware applications. Previous studies showed promising results using various machine learning algorithms. However, most existing methods can only recognize the activities that were previously seen in the training data. A previously unseen activity class cannot be recognized if there were no training samples in the dataset. Even if all of the activities can be enumerated in advance, labeled samples are often time consuming and expensive to get, as they require huge effort from human annotators or experts. In this paper, we present NuActiv, an activity recognition system that can recognize a human activity even when there are no training data for that activity class. Firstly, we designed a new representation of activities using semantic attributes, where each attribute is a human readable term that describes a basic element or an inherent characteristic of an activity. Secondly, based on this representation, a two-layer zero-shot learning algorithm is developed for activity recognition. Finally, to reinforce recognition accuracy using minimal user feedback, we developed an active learning algorithm for activity recognition. Our approach is evaluated on two datasets, including a 10-exercise-activity dataset we collected, and a public dataset of 34 daily life activities. Experimental results show that using semantic attribute-based learning, NuActiv can generalize knowledge to recognize unseen new activities. Our approach achieved up to 79% accuracy in unseen activity recognition.
Heng-Tze Cheng, Feng-Tso Sun, Martin L. Griss, Di You
MobiSys1