VLDB 2026 Research / reviewers in the wild / expert
Heng-Tze Cheng
dblp:30/8739
· DBLP profile ↗
14ranked-venue papers
3as first author
6since 2021 · last 2024
0009-0007-3845-8796ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorComputer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 57% Reinforcement learning · 16% Efficient and distributed learning · 12% | |
| Databases, data mining, and information retrieval
7 papers |
Information retrieval · 57% Recommender systems · 24% Machine learning and data management · 13% | |
| Human-computer interaction and pervasive computing
3 papers |
Ubiquitous computing and smart environments · 83% Wearable and physiological sensing · 17% |
Topics — the 30 heaviest of 36, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
prompting |
1.5 | 2 | 2024 | SELF-DISCOVER: Large Language Models Self-Compose Reasoning Structures · NeurIPS 2024 Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models · ICLR 2024 |
Information retrieval
personalized search |
1.0 | 2 | 2022 | Multi-Resolution Attention for Personalized Item Search · WSDM 2022 End-to-End Deep Attentive Personalized Item Retrieval for Online Content-sharing Platforms · WWW 2020 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | Controlled Decoding from Language Models · ICML 2024 |
Natural language and speech › Language models and text generation › decoding
controlled decoding |
0.8 | 1 | 2024 | Controlled Decoding from Language Models · ICML 2024 |
Machine learning › Reinforcement learning › policy optimization
KL-regularized RL |
0.8 | 1 | 2024 | Controlled Decoding from Language Models · ICML 2024 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.8 | 1 | 2024 | SELF-DISCOVER: Large Language Models Self-Compose Reasoning Structures · NeurIPS 2024 |
Machine learning › Learning paradigms
multi-task learning |
0.6 | 1 | 2022 | HyperPrompt: Prompt-based Task-Conditioning of Transformers · ICML 2022 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.6 | 1 | 2022 | HyperPrompt: Prompt-based Task-Conditioning of Transformers · ICML 2022 |
Natural language and speech › Language models and text generation
prompt tuning |
0.6 | 1 | 2022 | HyperPrompt: Prompt-based Task-Conditioning of Transformers · ICML 2022 |
Recommender systems › neural recommendation
attention-based recommendation |
0.6 | 1 | 2022 | Multi-Resolution Attention for Personalized Item Search · WSDM 2022 |
Information retrieval › retrieval models › neural retrieval
neural ranking model |
0.6 | 1 | 2022 | Multi-Resolution Attention for Personalized Item Search · WSDM 2022 |
Information retrieval › e-commerce search
personalized product search |
0.6 | 1 | 2022 | Multi-Resolution Attention for Personalized Item Search · WSDM 2022 |
Information retrieval
ranking |
0.6 | 1 | 2022 | Multi-Resolution Attention for Personalized Item Search · WSDM 2022 |
Ubiquitous computing and smart environments › context recognition
activity recognition |
0.5 | 3 | 2014 | Nonparametric discovery of human routines from sensor data · PerCom 2014 NuActiv: recognizing unseen new activities using semantic attribute-based learning · MobiSys 2013 Towards zero-shot learning for human activity recognition using semantic attribute sequence model · UbiComp 2013 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
error correction |
0.5 | 1 | 2021 | Mondegreen: A Post-Processing Solution to Speech Recognition Error Correction for Voice Search Queries · KDD 2021 |
Information retrieval
query processing |
0.5 | 1 | 2021 | Mondegreen: A Post-Processing Solution to Speech Recognition Error Correction for Voice Search Queries · KDD 2021 |
Information retrieval › e-commerce search
item retrieval |
0.4 | 1 | 2020 | End-to-End Deep Attentive Personalized Item Retrieval for Online Content-sharing Platforms · WWW 2020 |
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning |
0.4 | 1 | 2019 | SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets · IJCAI 2019 |
Machine learning › Reinforcement learning
value-based reinforcement learning |
0.4 | 1 | 2019 | SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets · IJCAI 2019 |
Recommender systems
reinforcement-learning-based recommendation |
0.4 | 1 | 2019 | SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets · IJCAI 2019 |
Recommender systems › interactive recommendation
slate recommendation |
0.4 | 1 | 2019 | SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets · IJCAI 2019 |
Machine learning › Efficient and distributed learning
distributed training |
0.3 | 1 | 2017 | TensorFlow Estimators: Managing Simplicity vs. Flexibility in High-Level Machine Learning Frameworks · KDD 2017 |
Machine learning › Efficient and distributed learning
production machine learning |
0.3 | 1 | 2017 | TensorFlow Estimators: Managing Simplicity vs. Flexibility in High-Level Machine Learning Frameworks · KDD 2017 |
Machine learning and data management
machine learning lifecycle management |
0.3 | 1 | 2017 | TFX: A TensorFlow-Based Production-Scale Machine Learning Platform · KDD 2017 |
Machine learning and data management › machine learning systems
machine learning platform |
0.3 | 1 | 2017 | TFX: A TensorFlow-Based Production-Scale Machine Learning Platform · KDD 2017 |
Natural language and speech › Language models and text generation › large language model inference
inference-time decoding |
0.2 | 1 | 2024 | Controlled Decoding from Language Models · ICML 2024 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering |
0.2 | 1 | 2024 | Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models · ICLR 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
reasoning structure |
0.2 | 1 | 2024 | SELF-DISCOVER: Large Language Models Self-Compose Reasoning Structures · NeurIPS 2024 |
Data mining › text mining › topic modeling
nonparametric topic model |
0.2 | 1 | 2014 | Nonparametric discovery of human routines from sensor data · PerCom 2014 |
Data mining › text mining
topic modeling |
0.2 | 1 | 2014 | Nonparametric discovery of human routines from sensor data · PerCom 2014 |
Methods — techniques the papers use, named apart from their topics
value function · 0.8self-discovery · 0.8self-consistency · 0.8reinforcement learning · 0.8prefix scorer · 0.8large language model prompting · 0.8in-context learning · 0.8chain-of-thought · 0.8soft-thresholding · 0.6prompt tuning · 0.6multi-head attention · 0.6hypernetwork · 0.6text-space correction · 0.5extreme multi-class softmax · 0.4embedding · 0.4attention mechanism · 0.4temporal difference learning · 0.4q-learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Take a Step Back: Evoking Reasoning via Abstraction in Large Language ModelsabstractWe present STEP-BACK PROMPTING, a simple prompting technique that enables LLMs to do abstractions to derive high-level concepts and first principles from instances containing specific details. Using the concepts and principles to guide reasoning, LLMs significantly improve their abilities in following a correct reasoning path towards the solution. We conduct experiments of STEP-BACK PROMPTING with PaLM-2L, GPT-4 and Llama2-70B models, and observe substantial performance gains on various challenging reasoning-intensive tasks including STEM, Knowledge QA, and Multi-Hop Reasoning. For instance, STEP-BACK PROMPTING improves PaLM-2L performance on MMLU (Physics and Chemistry) by 7% and 11% respectively, TimeQA by 27%, and MuSiQue by 7%. Huaixiu Steven Zheng, Swaroop Mishra, Heng-Tze Cheng, Ed H. Chi, Quoc V. Le, Denny Zhou |
ICLR | 4 |
| 2024 | Controlled Decoding from Language ModelsabstractKL-regularized reinforcement learning (RL) is a popular alignment framework to control the language model responses towards high reward outcomes. We pose a tokenwise RL objective and propose a modular solver for it, called *controlled decoding (CD)*. CD exerts control through a separate *prefix scorer* module, which is trained to learn a value function for the reward. The prefix scorer is used at inference time to control the generation from a frozen base model, provably sampling from a solution to the RL objective. We empirically demonstrate that CD is effective as a control mechanism on popular benchmarks. We also show that prefix scorers for multiple rewards may be combined at inference time, effectively solving a multi-objective RL problem with no additional training. We show that the benefits of applying CD transfer to an unseen base model with no further tuning as well. Finally, we show that CD can be applied in a blockwise decoding fashion at inference-time, essentially bridging the gap between the popular best-of-$K$ strategy and tokenwise control through reinforcement learning. This makes CD a promising approach for alignment of language models. Sidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li, Yanping Huang, Heng-Tze Cheng, Trevor Strohman, Jilin Chen, Alex Beutel, Ahmad Beirami |
ICML | 8 |
| 2024 | SELF-DISCOVER: Large Language Models Self-Compose Reasoning StructuresabstractWe introduce SELF-DISCOVER, a general framework for LLMs to self-discover the task-intrinsic reasoning structures to tackle complex reasoning problems that are challenging for typical prompting methods. Core to the framework is a self-discovery process where LLMs select multiple atomic reasoning modules such as critical thinking and step-by-step thinking, and compose them into an explicit reasoning structure for LLMs to follow during decoding. SELF-DISCOVER substantially improves GPT-4 and PaLM 2’s performance on challenging reasoning benchmarks such as BigBench-Hard, grounded agent reasoning, and MATH, by as much as 32% compared to Chain of Thought (CoT). Furthermore, SELF-DISCOVER outperforms inference-intensive methods such as CoT-Self-Consistency by more than 20%, while requiring 10-40x fewer inference compute. Finally, we show that the self-discovered reasoning structures are universally applicable across model families: from PaLM 2-L to GPT-4, and from GPT-4 to Llama2, and share commonalities with human reasoning patterns. Jay Pujara, Xiang Ren 0001, Heng-Tze Cheng, Quoc V. Le, Ed H. Chi, Denny Zhou, Swaroop Mishra, Huaixiu Steven Zheng |
NeurIPS | 5 |
| 2022 | HyperPrompt: Prompt-based Task-Conditioning of TransformersabstractPrompt-Tuning is a new paradigm for finetuning pre-trained language models in a parameter efficient way. Here, we explore the use of HyperNetworks to generate hyper-prompts: we propose HyperPrompt, a novel architecture for prompt-based task-conditioning of self-attention in Transformers. The hyper-prompts are end-to-end learnable via generation by a HyperNetwork. HyperPrompt allows the network to learn task-specific feature maps where the hyper-prompts serve as task global memories for the queries to attend to, at the same time enabling flexible information sharing among tasks. We show that HyperPrompt is competitive against strong multi-task learning baselines with as few as 0.14% of additional task-conditioning parameters, achieving great parameter and computational efficiency. Through extensive empirical experiments, we demonstrate that HyperPrompt can achieve superior performances over strong T5 multi-task learning baselines and parameter-efficient adapter variants including Prompt-Tuning and HyperFormer++ on Natural Language Understanding benchmarks of GLUE and SuperGLUE across many model sizes. Huaixiu Steven Zheng, Yi Tay, Jai Gupta 0001, Vamsi Aribandi, Zhe Zhao 0001, YaGuang Li, Donald Metzler, Heng-Tze Cheng, Ed H. Chi |
ICML | 11 |
| 2022 | Multi-Resolution Attention for Personalized Item SearchabstractPersonalized item search has become an essential tool for online platforms---where users interact with a large corpus of items (e.g., click, purchase, like) via a search query---to provide their users with a more satisfactory search experience. The record (or history) of users' past interactions serves as a valuable asset to achieve personalization. While user history data can span over a long period of time, only a part of the history is relevant to a user's current search intent. Moreover, since historical interactions take place at aperiodic points in time, modeling their relevance to the current search query entangles complex temporal dependencies. We propose multi-resolution attention to address these challenges for personalized item search. Our approach captures higher-order temporal relations between user queries and their history across several temporal subspaces (i.e., resolutions), each corresponding to distinct temporal ranges with adaptive time boundaries that are also learned directly from data. We achieve this by coupling the conventional multi-head attention module with a differentiable soft-thresholding mechanism, which essentially operates as a masking function in the temporal domain. Comparisons with strong baselines on an open-source benchmark dataset confirm the efficacy of our approach. Furkan Kocayusufoglu, Anima Singh, George Roumpos, Heng-Tze Cheng, Sagar Jain, Ed H. Chi, Ambuj K. Singh |
WSDM | 5 |
| 2021 | Mondegreen: A Post-Processing Solution to Speech Recognition Error Correction for Voice Search QueriesabstractAs more and more online search queries come from voice, automatic speech recognition becomes a key component to deliver relevant search results. Errors introduced by automatic speech recognition (ASR) lead to irrelevant search results returned to the user, thus causing user dissatisfaction. In this paper, we introduce an approach, "Mondegreen", to correct voice queries in text space without depending on audio signals, which may not always be available due to system constraints or privacy or bandwidth (for example, some ASR systems run on-device) considerations. We focus on voice queries transcribed via several proprietary commercial ASR systems. These queries come from users making internet, or online service search queries. We first present an analysis showing how different the language distribution coming from user voice queries is from that in traditional text corpora used to train off-the-shelf ASR systems. We then demonstrate that Mondegreen can achieve significant improvements in increased user interaction by correcting user voice queries in one of the largest search systems in Google. Finally, we see Mondegreen as complementing existing highly-optimized production ASR systems, which may not be frequently retrained and thus lag behind due to vocabulary drifts. Sukhdeep S. Sodhi, Ellie Ka In Chio, Ambarish Jash, Santiago Ontañón, Ajit Apte, Ayooluwakunmi Jeje, Dima Kuzmin, Harry Fung, Heng-Tze Cheng, Jon Effrat, Tarush Bali, Nitin Jindal, Sarvjeet Singh, Senqiang Zhou, Tameen Khan, Amol Wankhede, Moustafa Farid Alzantot, Allen Wu, Tushar Chandra |
KDD | 10 |
| 2020 | Zero-Shot Heterogeneous Transfer Learning from Recommender Systems to Cold-Start Search RetrievalabstractMany recent advances in neural information retrieval models, which predict top-K items given a query, learn directly from a large training set of (query, item) pairs. However, they are often insufficient when there are many previously unseen (query, item) combinations, often referred to as the cold start problem. Furthermore, the search system can be biased towards items that are frequently shown to a query previously, also known as the 'rich get richer' (a.k.a. feedback loop) problem. In light of these problems, we observed that most online content platforms have both a search and a recommender system that, while having heterogeneous input spaces, can be connected through their common output item space and a shared semantic representation. In this paper, we propose a new Zero-Shot Heterogeneous Transfer Learning framework that transfers learned knowledge from the recommender system component to improve the search component of a content platform. First, it learns representations of items and their natural-language features by predicting (item, item) correlation graphs derived from the recommender system as an auxiliary task. Then, the learned representations are transferred to solve the target search retrieval task, performing query-to-item prediction without having seen any (query, item) pairs in training. We conduct online and offline experiments on one of the world's largest search and recommender systems from Google, and present the results and lessons learned. We demonstrate that the proposed approach can achieve high performance on offline search retrieval tasks, and more importantly, achieved significant improvements on relevance and user interactions over the highly-optimized production system in online experiments. Ellie Ka In Chio, Heng-Tze Cheng, Steffen Rendle, Dima Kuzmin, Ritesh Agarwal, Li Zhang 0001, John R. Anderson, Sarvjeet Singh, Tushar Chandra, Ed H. Chi, Alex Soares, Nitin Jindal |
CIKM | 3 |
| 2020 | End-to-End Deep Attentive Personalized Item Retrieval for Online Content-sharing PlatformsabstractModern online content-sharing platforms host billions of items like music, videos, and products uploaded by various providers for users to discover items of their interests. To satisfy the information needs, the task of effective item retrieval (or item search ranking) given user search queries has become one of the most fundamental problems to online content-sharing platforms. Moreover, the same query can represent different search intents for different users, so personalization is also essential for providing more satisfactory search results. Different from other similar research tasks, such as ad-hoc retrieval and product retrieval with copious words and reviews, items in content-sharing platforms usually lack sufficient descriptive information and related meta-data as features. In this paper, we propose the end-to-end deep attentive model (EDAM) to deal with personalized item retrieval for online content-sharing platforms using only discrete personal item history and queries. Each discrete item in the personal item history of a user and its content provider are first mapped to embedding vectors as continuous representations. A query-aware attention mechanism is then applied to identify the relevant contexts in the user history and construct the overall personal representation for a given query. Finally, an extreme multi-class softmax classifier aggregates the representations of both query and personal item history to provide personalized search results. We conduct extensive experiments on a large-scale real-world dataset with hundreds of million users from a large video media platform at Google. The experimental results demonstrate that our proposed approach significantly outperforms several competitive baseline methods. It is also worth mentioning that this work utilizes a massive dataset from a real-world commercial content-sharing platform for personalized item retrieval to provide more insightful analysis from the industrial aspects. Jyun-Yu Jiang, George Roumpos, Heng-Tze Cheng, Xinyang Yi, Ed H. Chi, Harish Ganapathy, Nitin Jindal, Wei Wang 0010 |
WWW | 4 |
| 2019 | SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation SetsabstractReinforcement learning methods for recommender systems optimize recommendations for long-term user engagement. However, since users are often presented with slates of multiple items---which may have interacting effects on user choice---methods are required to deal with the combinatorics of the RL action space. We develop SlateQ, a decomposition of value-based temporal-difference and Q-learning that renders RL tractable with slates. Under mild assumptions on user choice behavior, we show that the long-term value (LTV) of a slate can be decomposed into a tractable function of its component item-wise LTVs. We demonstrate our methods in simulation, and validate the scalability and effectiveness of decomposed TD-learning on YouTube. Eugene Ie, Vihan Jain, Sanmit Narvekar, Ritesh Agarwal, Rui Wu 0020, Heng-Tze Cheng, Tushar Chandra, Craig Boutilier |
IJCAI | 7 |
| 2017 | TFX: A TensorFlow-Based Production-Scale Machine Learning PlatformabstractCreating and maintaining a platform for reliably producing and deploying machine learning models requires careful orchestration of many components---a learner for generating models based on training data, modules for analyzing and validating both data as well as models, and finally infrastructure for serving models in production. This becomes particularly challenging when data changes over time and fresh models need to be produced continuously. Unfortunately, such orchestration is often done ad hoc using glue code and custom scripts developed by individual teams for specific use cases, leading to duplicated effort and fragile systems with high technical debt. Denis Baylor, Eric Breck, Heng-Tze Cheng, Noah Fiedel, Chuan Yu Foo, Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc 0001, Chiu Yuen Koo, Lukasz Lew, Clemens Mewald, Akshay Naresh Modi, Neoklis Polyzotis, Sukriti Ramesh, Sudip Roy 0002, Steven Euijong Whang, Martin Wicke, Jarek Wilkiewicz, Martin Zinkevich |
KDD | 3 |
| 2017 | TensorFlow Estimators: Managing Simplicity vs. Flexibility in High-Level Machine Learning FrameworksabstractWe present a framework for specifying, training, evaluating, and deploying machine learning models. Our focus is on simplifying cutting edge machine learning for practitioners in order to bring such technologies into production. Recognizing the fast evolution of the field of deep learning, we make no attempt to capture the design space of all possible model architectures in a domain-specific language (DSL) or similar configuration language. We allow users to write code to define their models, but provide abstractions that guide developers to write models in ways conducive to productionization. We also provide a unifying Estimator interface, making it possible to write downstream infrastructure (e.g. distributed training, hyperparameter tuning) independent of the model implementation. Heng-Tze Cheng, Zakaria Haque, Lichan Hong, Mustafa Ispir, Clemens Mewald, Illia Polosukhin, George Roumpos, D. Sculley, Jamie Smith, David Soergel, Yuan Tang 0001, Philipp Tucker, Martin Wicke, Cassandra Xia, Jianwei Xie |
KDD | 1 |
| 2014 | Nonparametric discovery of human routines from sensor dataabstractPeople engage in routine behaviors. Automatic routine discovery goes beyond low-level activity recognition such as sitting or standing and analyzes human behaviors at a higher level (e.g., commuting to work). With recent developments in ubiquitous sensor technologies, it becomes easier to acquire a massive amount of sensor data. One main line of research is to mine human routines from sensor data using parametric topic models such as latent Dirichlet allocation. The main shortcoming of parametric models is that it assumes a fixed, pre-specified parameter regardless of the data. Choosing an appropriate parameter usually requires an inefficient trial-and-error model selection process. Furthermore, it is even more difficult to find optimal parameter values in advance for personalized applications. In this paper, we present a novel nonparametric framework for human routine discovery that can infer high-level routines without knowing the number of latent topics beforehand. Our approach is evaluated on public datasets in two routine domains: a 34-daily-activity dataset and a transportation mode dataset. Experimental results show that our nonparametric framework can automatically learn the appropriate model parameters from sensor data without any form of model selection procedure and can outperform traditional parametric approaches for human routine discovery tasks. Feng-Tso Sun, Heng-Tze Cheng, Cynthia Kuo, Martin L. Griss |
PerCom | 3 |
| 2013 | Towards zero-shot learning for human activity recognition using semantic attribute sequence modelabstractUnderstanding human activities is important for user-centric and context-aware applications. Previous studies showed promising results using various machine learning algorithms. However, most existing methods can only recognize the activities that were previously seen in the training data. In this paper, we present a new zero-shot learning framework for human activity recognition that can recognize an unseen new activity even when there are no training samples of that activity in the dataset. We propose a semantic attribute sequence model that takes into account both the hierarchical and sequential nature of activity data. Evaluation on datasets in two activity domains show that the proposed zero-shot learning approach achieves 70-75% precision and recall recognizing unseen new activities, and outperforms supervised learning with limited labeled data for the new classes. Heng-Tze Cheng, Martin L. Griss, Di You |
UbiComp | 1 |
| 2013 | NuActiv: recognizing unseen new activities using semantic attribute-based learningabstractWe study the problem of how to recognize a new human activity when we have never seen any training example of that activity before. Recognizing human activities is an essential element for user-centric and context-aware applications. Previous studies showed promising results using various machine learning algorithms. However, most existing methods can only recognize the activities that were previously seen in the training data. A previously unseen activity class cannot be recognized if there were no training samples in the dataset. Even if all of the activities can be enumerated in advance, labeled samples are often time consuming and expensive to get, as they require huge effort from human annotators or experts. In this paper, we present NuActiv, an activity recognition system that can recognize a human activity even when there are no training data for that activity class. Firstly, we designed a new representation of activities using semantic attributes, where each attribute is a human readable term that describes a basic element or an inherent characteristic of an activity. Secondly, based on this representation, a two-layer zero-shot learning algorithm is developed for activity recognition. Finally, to reinforce recognition accuracy using minimal user feedback, we developed an active learning algorithm for activity recognition. Our approach is evaluated on two datasets, including a 10-exercise-activity dataset we collected, and a public dataset of 34 daily life activities. Experimental results show that using semantic attribute-based learning, NuActiv can generalize knowledge to recognize unseen new activities. Our approach achieved up to 79% accuracy in unseen activity recognition. Heng-Tze Cheng, Feng-Tso Sun, Martin L. Griss, Di You |
MobiSys | 1 |