Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yoshi Suhara

dblp:322/0272 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 44% Deep learning architectures and training · 31% Language models and text generation · 22%
Databases, data mining, and information retrieval
2 papers
Data integration and cleaning · 54% Information retrieval · 46%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
attention mechanism
0.912025
Hymba: A Hybrid-head Architecture for Small Language Models · ICLR 2025
Machine learning › Deep learning architectures and training › attention mechanism
hybrid attention
0.912025
Hymba: A Hybrid-head Architecture for Small Language Models · ICLR 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression
pruning
0.912025
Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression › lightweight neural network
small language models
0.912025
Hymba: A Hybrid-head Architecture for Small Language Models · ICLR 2025
Data integration and cleaning
entity matching
0.712023
Effective entity matching with transformers · VLDB J. 2023
Data integration and cleaning › entity matching
transformer-based entity matching
0.712023
Effective entity matching with transformers · VLDB J. 2023
Natural language and speech › Language models and text generation
text summarization
0.612022
Summarizing Community-based Question-Answer Pairs · EMNLP 2022
Information retrieval › text summarization
opinion summarization
0.612022
Beyond Opinion Mining: Summarizing Opinions of Customer Reviews · SIGIR 2022
Information retrieval
text summarization
0.612022
Beyond Opinion Mining: Summarizing Opinions of Customer Reviews · SIGIR 2022
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.312025
Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning · NeurIPS 2025
Machine learning › Deep learning architectures and training
state space model
0.312025
Hymba: A Hybrid-head Architecture for Small Language Models · ICLR 2025
Natural language and speech › Question answering and dialogue systems
community question answering
0.212022
Summarizing Community-based Question-Answer Pairs · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

state space model pruning · 0.9meta tokens · 0.9knowledge distillation · 0.9group-aware pruning · 0.9cross-layer key-value sharing · 0.9transformer · 0.7variational inference · 0.6few-shot learning · 0.6extractive summarization · 0.6controllable text generation · 0.6auto-encoding · 0.6abstractive summarization · 0.6
YearPublicationVenuePosition
2025 Hymba: A Hybrid-head Architecture for Small Language Models
abstract
We propose Hymba, a family of small language models featuring a hybrid-head parallel architecture that integrates attention mechanisms and state space models (SSMs) within the same layer, offering parallel and complementary processing of the same inputs. In this hybrid-head module, attention heads provide high-resolution recall, while SSM heads facilitate efficient context summarization. Additionally, we introduce learnable meta tokens, which are prepended to prompts to store critical meta information, guiding subsequent tokens and alleviating the “forced-to-attend” burden associated with attention mechanisms. Thanks to the global context summarized by SSMs, the attention heads in our model can be further optimized through cross-layer key-value (KV) sharing and a mix of global and local attention, resulting in a compact cache size without compromising accuracy. Notably, Hymba achieves state-of-the-art performance among small LMs: Our Hymba-1.5B-Base model surpasses all sub-2B public models and even outperforms Llama-3.2-3B, achieving 1.32\% higher average accuracy, an 11.67$\times$ reduction in cache size, and 3.49$\times$ higher throughput.
Xin Dong 0009, Yonggan Fu, Shizhe Diao, Wonmin Byeon, Zijia Chen, Ameya Mahabaleshwarkar, Shih-Yang Liu, Matthijs Van Keirsbilck, Min-Hung Chen, Yoshi Suhara, Yingyan (Celine) Lin, Jan Kautz, Pavlo Molchanov 0001
ICLR10
2025 When2Call: When (not) to Call Tools
abstract
Hayley Ross, Ameya Sunil Mahabaleshwarkar, Yoshi Suhara. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Hayley Ross, Ameya Mahabaleshwarkar, Yoshi Suhara
NAACL (Long Papers)3
2025 Nemotron-CLIMB: Clustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
abstract
Pre-training datasets are typically collected from web content and lack inherent domain divisions. For instance, widely used datasets like Common Crawl do not include explicit domain labels, while manually curating labeled datasets such as The Pile is labor-intensive. Consequently, identifying an optimal pre-training data mixture remains a challenging problem, despite its significant benefits for pre-training performance. To address these challenges, we propose CLustering-based Iterative Data Mixture Bootstrapping (Nemotron-CLIMB), an automated framework that discovers, evaluates, and refines data mixtures in a pre-training setting. Specifically, Nemotron-CLIMB embeds and clusters large-scale datasets in a semantic space and then iteratively searches for optimal mixtures using a smaller proxy model and a predictor. This strategy enables effective domain adaptation without relying solely on curated data. When continuously trained on 400B tokens with this mixture, our 1B model exceeds the state-of-the-art Llama-3.2-1B by 2.0%. Moreover, we observe that optimizing for a specific domain (e.g., Social Sciences) yields a 5% improvement over random sampling. Finally, we introduce Nemotron-ClimbLab, a filtered 1.2-trillion-token corpus with 20 clusters as a research playground, and Nemotron-ClimbMix, a compact yet powerful 400-billion-token dataset designed for efficient pre-training that delivers superior performance under an equal token budget. We analyze the final data mixture, elucidating the characteristics of an optimal data mixture.
Shizhe Diao, Yonggan Fu, Xin Dong 0009, Dan Su 0003, Markus Kliegl, Zijia Chen, Peter Belcak, Yoshi Suhara, Hongxu Yin, Mostofa Patwary, Yingyan (Celine) Lin, Jan Kautz, Pavlo Molchanov 0001
NeurIPS9
2025 Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
abstract
Hybrid language models that combine Attention and State Space Models (SSMs) have been shown to achieve state-of-the-art accuracy and runtime performance. Recent work has also demonstrated that applying pruning and distillation to Attention-only models yields smaller, more accurate models at a fraction of the training cost. In this work, we explore the effectiveness of compressing Hybrid architectures. To this end, we introduce a novel group-aware pruning method for Mamba layers that preserves the structural integrity of SSM blocks and their sequence modeling capabilities. We combine this method with FFN, embedding dimension, and layer pruning, along with knowledge distillation-based retraining to obtain a unified compression recipe for hybrid models. Using this recipe, we compress the Nemotron-H 8B Hybrid model down to 4B parameters with up to $40\times$ fewer training tokens compared to similarly-sized models. The resulting model surpasses the accuracy of similarly-sized models while achieving $\sim2\times$ faster inference throughput, significantly advancing the Pareto frontier.
Ali Taghibakhshi, Sharath Turuvekere Sreenivas, Saurav Muralidharan, Marcin Chochowski, Yashaswi Karnati, Raviraj Joshi, Ameya Mahabaleshwarkar, Zijia Chen, Yoshi Suhara, Oluwatobi Olabiyi, Daniel Korzekwa, Mostofa Patwary, Mohammad Shoeybi, Jan Kautz, Bryan Catanzaro, Ashwath Aithal, Nima Tajbakhsh, Pavlo Molchanov 0001
NeurIPS9
2024 Noisy Pairing and Partial Supervision for Stylized Opinion Summarization
abstract
Opinion summarization research has primarily focused on generating summaries reflecting important opinions from customer reviews without paying much attention to the writing style.In this paper, we propose the stylized opinion summarization task, which aims to generate a summary of customer reviews in the desired (e.g., professional) writing style.To tackle the difficulty in collecting customer and professional review pairs, we develop a non-parallel training framework, Noisy Pairing and Partial Supervision (Napa ), which trains a stylized opinion summarization system from non-parallel customer and professional review sets.We create a benchmark PRO-SUM by collecting customer and professional reviews from Yelp and Michelin.Experimental results on PROSUM and FewSum demonstrate that our non-parallel training framework consistently improves both automatic and human evaluations, successfully building a stylized opinion summarization model that can generate professionally-written summaries from customer reviews. 1
Hayate Iso, Xiaolan Wang 0001, Yoshi Suhara
INLG3
2023 Effective entity matching with transformers
Yuliang Li 0001, Yoshi Suhara, AnHai Doan, Wang Chiew Tan
VLDB J.3
2022 Summarizing Community-based Question-Answer Pairs
abstract
Community-based Question Answering (CQA), which allows users to acquire their desired information, has increasingly become an essential component of online services in various domains such as E-commerce, travel, and dining.However, an overwhelming number of CQA pairs makes it difficult for users without particular intent to find useful information spread over CQA pairs.To help users quickly digest the key information, we propose the novel CQA summarization task that aims to create a concise summary from CQA pairs.To this end, we first design a multi-stage data annotation process and create a benchmark dataset, CO-QASUM, based on the Amazon QA corpus.We then compare a collection of extractive and abstractive summarization methods and establish a strong baseline approach DedupLED for the CQA summarization task.Our experiment further confirms two key challenges, sentencetype transfer and deduplication removal, towards the CQA summarization task.Our data and code are publicly available.1 * Work done while at Megagon Labs. 1 https://github.com/megagonlabs/ qa-summarization Q: Is this actually a rigid board or more of a floppy mat?A: It is rigid.themain board is rigid,the two sides are semi.Q: Is this actually a rigid board or more of a floppy mat?A: The main area is very sturdy.Then there are two work area pads that are more flexible so when moving those I keep two hands on them.Q: how wide is each Side piece?"A: 16 inches wide (there are two).Q: will this mat hold 1000 piece puzzle?A: most certainly will.… … (omitted 27 QAs) Summary: This puzzle board comes with a rigid main board.You can arrange pieces in the middle and on two side pieces, and then pick up those side pieces to place them atop the middle area before folding the wings in.The dimension of the puzzle space is 32"x21.75".The closed unit is almost the same size as the puzzle workspace (32"x21.75").There are two 16" wide side inserts.The mat holds most 1000 pieces puzzles.It is too big to use on you lap and definitely needs a table.(a).QAs for a puzzle board product (Input) (b).Summary of QAs (Output) Q: what is the storage size when case is fully closed for storage?A: Closed size is 32.25 x 22.75".Q: What size is the closed unit?A: Closed is almost the same size as the puzzle workspace.32.25 x 22.
Ting-Yao Hsu, Yoshi Suhara, Xiaolan Wang 0001
EMNLP2
2022 Beyond Opinion Mining: Summarizing Opinions of Customer Reviews
abstract
Customer reviews are vital for making purchasing decisions in the Information Age. Such reviews can be automatically summarized to provide the user with an overview of opinions. In this tutorial, we present various aspects of opinion summarization that are useful for researchers and practitioners. First, we will introduce the task and major challenges. Then, we will present existing opinion summarization solutions, both pre-neural and neural. We will discuss how summarizers can be trained in the unsupervised, few-shot, and supervised regimes. Each regime has roots in different machine learning methods, such as auto-encoding, controllable text generation, and variational inference. Finally, we will discuss resources and evaluation methods and conclude with the future directions. This three-hour tutorial will provide a comprehensive overview over major advances in opinion summarization. The listeners will be well-equipped with the knowledge that is both useful for research and practical applications.
Reinald Kim Amplayo, Arthur Brazinskas, Yoshi Suhara, Xiaolan Wang 0001, Bing Liu 0001
SIGIR3