Jiantao Huang

dblp:90/5407 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
3since 2021 · last 2024
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Knowledge representation and reasoning · 35% Representation and self-supervised learning · 22% Language models and text generation · 17%
Databases, data mining, and information retrieval
2 papers
Data mining · 60% Web and social media mining · 40%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
commonsense evaluation
0.812024
CORECODE: A Common Sense Annotated Dialogue Dataset with Benchmark Tasks for Chinese Large Language Models · AAAI 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.812024
CORECODE: A Common Sense Annotated Dialogue Dataset with Benchmark Tasks for Chinese Large Language Models · AAAI 2024
Natural language and speech › Question answering and dialogue systems
dialogue understanding
0.812024
CORECODE: A Common Sense Annotated Dialogue Dataset with Benchmark Tasks for Chinese Large Language Models · AAAI 2024
Natural language and speech › Language models and text generation
large language model evaluation
0.812024
CORECODE: A Common Sense Annotated Dialogue Dataset with Benchmark Tasks for Chinese Large Language Models · AAAI 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
Twin Contrastive Learning for Online Clustering · Int. J. Comput. Vis. 2022
Data mining
clustering
0.612022
Twin Contrastive Learning for Online Clustering · Int. J. Comput. Vis. 2022
Data mining › clustering
deep clustering
0.612022
Twin Contrastive Learning for Online Clustering · Int. J. Comput. Vis. 2022
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › semantic embedding
topic embedding
0.412019
A Novel Generative Topic Embedding Model by Introducing Network Communities · WWW 2019
Natural language and speech › Information extraction and text analysis
topic model
0.412019
A Novel Generative Topic Embedding Model by Introducing Network Communities · WWW 2019
Web and social media mining › community analysis
community structure
0.412019
A Novel Generative Topic Embedding Model by Introducing Network Communities · WWW 2019
Web and social media mining › information networks
document networks
0.412019
A Novel Generative Topic Embedding Model by Introducing Network Communities · WWW 2019

Methods — techniques the papers use, named apart from their topics

twin contrastive learning · 1.1zero-shot evaluation · 0.8variational inference · 0.8probabilistic generative model · 0.8crowdsourcing annotation · 0.8
YearPublicationVenuePosition
2024 CORECODE: A Common Sense Annotated Dialogue Dataset with Benchmark Tasks for Chinese Large Language Models
abstract
As an indispensable ingredient of intelligence, commonsense reasoning is crucial for large language models (LLMs) in real-world scenarios. In this paper, we propose CORECODE, a dataset that contains abundant commonsense knowledge manually annotated on dyadic dialogues, to evaluate the commonsense reasoning and commonsense conflict detection capabilities of Chinese LLMs. We categorize commonsense knowledge in everyday conversations into three dimensions: entity, event, and social interaction. For easy and consistent annotation, we standardize the form of commonsense knowledge annotation in open-domain dialogues as "domain: slot = value". A total of 9 domains and 37 slots are defined to capture diverse commonsense knowledge. With these pre-defined domains and slots, we collect 76,787 commonsense knowledge annotations from 19,700 dialogues through crowdsourcing. To evaluate and enhance the commonsense reasoning capability for LLMs on the curated dataset, we establish a series of dialogue-level reasoning and detection tasks, including commonsense knowledge filling, commonsense knowledge generation, commonsense conflict phrase detection, domain identification, slot identification, and event causal inference. A wide variety of existing open-source Chinese LLMs are evaluated with these tasks on our dataset. Experimental results demonstrate that these models are not competent to predict CORECODE's plentiful reasoning content, and even ChatGPT could only achieve 0.275 and 0.084 accuracy on the domain identification and slot identification tasks under the zero-shot setting. We release the data and codes of CORECODE at https://github.com/danshi777/CORECODE to promote commonsense reasoning evaluation and study of LLMs in the context of daily conversations.
Dan Shi 0001, Chaobin You, Jiantao Huang, Taihao Li, Deyi Xiong
AAAI3
2024 NumHG: A Dataset for Number-Focused Headline Generation
abstract
Headline generation, a key task in abstractive summarization, strives to condense a full-length article into a succinct, single line of text. Notably, while contemporary encoder-decoder models excel based on the ROUGE metric, they often falter when it comes to the precise generation of numerals in headlines. We identify the lack of datasets providing fine-grained annotations for accurate numeral generation as a major roadblock. To address this, we introduce a new dataset, the NumHG, and provide over 27,000 annotated numeral-rich news articles for detailed investigation. Further, we evaluate five well-performing models from previous headline-generation tasks using human evaluation in terms of numerical accuracy, reasonableness, and readability. Our study reveals a need for improvement in numerical accuracy, demonstrating the potential of the NumHG dataset to drive progress in number-focused headline generation and stimulate further discussions in numeral-focused text generation.
Jiantao Huang, Chung-Chi Chen 0001, Hen-Hsen Huang, Hsin-Hsi Chen
LREC/COLING1
2022 Twin Contrastive Learning for Online Clustering
Yunfan Li 0003, Mouxing Yang, Dezhong Peng, Taihao Li, Jiantao Huang, Xi Peng 0001
Int. J. Comput. Vis.5
2019 A Novel Generative Topic Embedding Model by Introducing Network Communities
abstract
Topic models have many important applications in fields such as Natural Language Processing. Topic embedding modelling aims at introducing word and topic embeddings into topic models to describe correlations between topics. Existing topic embedding methods use documents alone, which suffer from the topical fuzziness problem brought by the introduction of embeddings of semantic fuzzy words, e.g. polysemous words or some misleading academic terms. Links often exist between documents which form document networks. The use of links may alleviate this semantic fuzziness, but they are sparse and noisy which may meanwhile mislead topics. In this paper, we utilize community structure to solve these problems. It can not only alleviate the topical fuzziness of topic embeddings since communities are often believed to be topic related, but also can overcome the drawbacks brought by the sparsity and noise of networks (because community is a high-order network information). We give a new generative topic embedding model which incorporates documents (with topics) and network (with communities) together, and uses probability transition to describe the relationship between topics and communities to make it robust when topics and communities do not match. An efficient variational inference algorithm is then proposed to learn the model. We validate the superiority of our new approach on two tasks, document classifications and visualization of topic embeddings, respectively.
Di Jin 0001, Jiantao Huang, Pengfei Jiao, Liang Yang 0002, Dongxiao He, Françoise Fogelman-Soulié
WWW2
1999 Spatio-Temporal Tracking of Myocardial Deformation with a 4D B-Spline Model from Tagged MRI
abstract
Accurate delineation of the volumetric motion of the left ventricle (LV) of the heart from tagged magnetic resonance imaging (MRI) is an important area of research. We have built a system that takes extracted tag line features from short axis (SA) and long axis (LA) image sequences as input and fits a four-dimensional (4-D) time-varying B-spline model to the data by simultaneously fitting the model knot solids to MRI frames via matching three sequences of solid knot planes to the LV tag planes for 4-D tracking. Important advantages of the model are that reconstruction of tag surfaces, three-dimensional (3-D) material point localization, as well as displacement reconstruction are all achieved in a single step. The generated 3-D displacement fields are validated with a cardiac motion simulator, and 3-D motion fields capturing in vivo deformations in a porcine model with posterolateral myocardial infarction are illustrated.
Jiantao Huang, Dana Abendschein, Victor G. Dávila-Román, Amir A. Amini
IEEE Trans. Medical Imaging1
1998 Flexible Shapes for Segmentation and Tracking of Cardiovascular Data
abstract
In this invited paper, an overview of techniques developed at the Cardiovascular Image Analysis Laboratory at Washington University is discussed. At the core of the authors' methodologies lie flexible shape models, which are employed in automated as well as semi-automated analysis of cardiac MRI and X-ray angiography images. The mathematical bases used for the flexible templates are of the B-spline variety, providing compact representation and interactive capabilities for manipulation of curves, surfaces, and volumes.
Amir A. Amini, Jiantao Huang, Andreas K. Klein, Petia Radeva, Mohamed Elayyadi
ICIP (2)2
1998 Anatomical Object Volumes from Deformable B-spline Surface Models
abstract
Accurate delineation of anatomical objects in 3D volumetric data is a significant problem in medical imaging. The authors have built a system that takes as input a spatial stack of 2D image slices being studied. The output of the system is a smooth 3D surface delineating the medical structure of interest, and the volume enclosed by the 3D surface. This paper presents a new method to represent a 3D anatomical object by a deformable B-spline tube, that facilitates the calculation of the volume enclosed based on a closed-form volume calculation formula. In order to validate the deformable B-spline surface fitting and associated volume calculation, 19 simulated 3D images were considered and theoretical volumes were compared with volumes produced by the authors' system. The errors were less than 4.9%.
Jiantao Huang, Amir A. Amini
ICIP (1)1
1997 Deformable B-Solids and Implicit Snakes for 3D Localization and Tracking of SPAMM MRI Data
Petia Radeva, Amir A. Amini, Jiantao Huang
Comput. Vis. Image Underst.3