Jingyue Huang

dblp:203/9630 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 43% Question answering and dialogue systems · 28% Information extraction and text analysis · 14%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › retrieval-augmented generation
multimodal retrieval-augmented generation
0.912025
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing · NeurIPS 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.912025
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing · NeurIPS 2025
Multimedia analysis and retrieval
video summarization
0.912025
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing · NeurIPS 2025
Natural language and speech › Information extraction and text analysis
entity typing
0.612022
Generative Entity Typing with Curriculum Learning · EMNLP 2022
Natural language and speech › Question answering and dialogue systems
multimodal question answering
0.612022
Weakly Supervised Learning for Textbook Question Answering · IEEE Trans. Image Process. 2022
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
textbook question answering
0.612022
Weakly Supervised Learning for Textbook Question Answering · IEEE Trans. Image Process. 2022
Machine learning › Transfer learning and domain adaptation
zero-shot learning
0.612022
Generative Entity Typing with Curriculum Learning · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

retrieval model · 1.7large language model · 1.7weakly supervised learning · 0.6text matching · 0.6self-paced learning · 0.6relation detection · 0.6pre-trained language model · 0.6object detection · 0.6multi-task learning · 0.6curriculum learning · 0.6
YearPublicationVenuePosition
2025 REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
abstract
Short videos are an effective tool for promoting contents and improving knowledge accessibility. While existing extractive video summarization methods struggle to produce a coherent narrative, existing abstractive methods cannot `quote' from the input videos, i.e., inserting short video clips in their outputs. In this work, we explore novel video editing models for generating shorts that feature a coherent narrative with embedded video insertions extracted from a long input video. We propose a novel retrieval-embedded generation framework that allows a large language model to quote multimodal resources while maintaining a coherent narrative. Our proposed REGen system first generates the output story script with quote placeholders using a finetuned large language model, and then uses a novel retrieval model to replace the quote placeholders by selecting a video clip that best supports the narrative from a pool of candidate quotable video clips. We examine the proposed method on the task of documentary teaser generation, where short interview insertions are commonly used to support the narrative of a documentary. Our objective evaluations show that the proposed method can effectively insert short video clips while maintaining a coherent narrative. In a subjective survey, we show that our proposed method outperforms existing abstractive and extractive approaches in terms of coherence, alignment, and realism in teaser generation.
Weihan Xu, Yimeng Ma, Jingyue Huang, Wenye Ma, Taylor Berg-Kirkpatrick, Julian J. McAuley, Paul Pu Liang, Hao-Wen Dong
NeurIPS3
2025 Two-Hop Partial Task Offloading and Resource Allocation in Air-Ground Integrated Mobile Edge Computing Network: A DRL-Based Method
abstract
The integration of mobile edge computing (MEC) and air-ground integrated network is viewed as a crucial technology for Internet of Remote Things (IoRT) devices. It provides widespread service coverage and allows the tasks of IoRT devices to be executed by the uncrewed aerial vehicles (UAVs) and the high altitude platforms (HAPs). In this article, we investigate a joint partial task offloading, resource allocation, and UAV trajectory design problem to minimize the total task offloading delay of all IoRT devices in the air-ground integrated MEC network. Given that the problem is nonconvex and hard to solve by the traditional methods, we convert it into a Markov decision process (MDP) and leverage the deep reinforcement learning method to address it. Considering the complexity of the MDP grows with the number of the IoRT devices and the UAVs increasing, the primal problem is decomposed into two subproblems: 1) the UAV trajectory design and IoRT device power control subproblem, and 2) the partial task offloading and resource allocation subproblem. To address these two subproblems, we apply the basic concepts of the multiagent deep deterministic policy gradient (MADDPG) and the independent proximal policy optimization (IPPO) methods, respectively. Additionally, we introduce the enhanced prioritized experience replay and noise value to improve both the convergence performance and rate. This leads to the development of the MADDPG-improved prioritized experience replay (MADDPG-IPER) algorithm and noise value-IPPO (NV-IPPO) algorithm. Based on the solution of these two subproblems, a joint partial task offloading, resource allocation, and UAV trajectory design (JPTORAUTD) algorithm is proposed. Simulation results present that the proposed JPTORAUTD algorithm outperforms other benchmark algorithms in terms of reducing the total task offloading delay.
Shichao Li 0001, Bingji Lu, Laha Ale, Hongbin Chen 0001, Fangqing Tan, Jingyue Huang
IEEE Internet Things J.6
2024 Text Keyword Extraction Based on GPT
abstract
This study investigates text keyword and phrase extraction methods based on the GPT-3.5 model,and validates their effectiveness through comparative analysis. Initially, researchers employ the GPT-3.5 model for extracting keywords and phrases from text to uncover crucial information within the text. Subsequently, the extracted data from GPT-3.5 is compared with the key text from the original dataset to assess extraction performance and consistency. Lastly, extracted keywords are utilized for sentiment analysis, conducting comparative experiments with the BERT-TextCNN model, and validation across diverse datasets. The research findings demonstrate the GPT-3.5 model's capability to efficiently extract crucial textual information and significantly enhance sentiment classification precision. This enhancement contributes to improved performance and interpretability in text analysis, thereby providing substantial support for the field of natural language processing.
Pinyao He, Jingyue Huang
CSCWD2
2022 Generative Entity Typing with Curriculum Learning
abstract
Entity typing aims to assign types to the entity mentions in given texts.The traditional classification-based entity typing paradigm has two unignorable drawbacks: 1) it fails to assign an entity to the types beyond the predefined type set, and 2) it can hardly handle few-shot and zero-shot situations where many long-tail types only have few or even no training instances.To overcome these drawbacks, we propose a novel generative entity typing (GET) paradigm: given a text with an entity mention, the multiple types for the role that the entity plays in the text are generated with a pre-trained language model (PLM).However, PLMs tend to generate coarse-grained types after finetuning upon the entity typing dataset.In addition, only the heterogeneous training data consisting of a small portion of human-annotated data and a large portion of auto-generated but low-quality data are provided for model training.To tackle these problems, we employ curriculum learning (CL) to train our GET model on heterogeneous data, where the curriculum could be self-adjusted with the self-paced learning according to its comprehension of the type granularity and data heterogeneity.Our extensive experiments upon the datasets of different languages and downstream tasks justify the superiority of our GET model over the state-ofthe-art entity typing models.The code has been released on https://github.com/siyuyuan/GET.
Deqing Yang, Jiaqing Liang, Zhixu Li, Jinxi Liu, Jingyue Huang, Yanghua Xiao
EMNLP6
2022 A simplified iteratively regularized projection method for nonlinear ill-posed problems
Jingyue Huang, Xingjun Luo
J. Complex.1
2022 Weakly Supervised Learning for Textbook Question Answering
abstract
Textbook Question Answering (TQA) is the task of answering diagram and non-diagram questions given large multi-modal contexts consisting of abundant text and diagrams. Deep text understandings and effective learning of diagram semantics are important for this task due to its specificity. In this paper, we propose a Weakly Supervised learning method for TQA (WSTQ), which regards the incompletely accurate results of essential intermediate procedures for this task as supervision to develop Text Matching (TM) and Relation Detection (RD) tasks and then employs the tasks to motivate itself to learn strong text comprehension and excellent diagram semantics respectively. Specifically, we apply the result of text retrieval to build positive as well as negative text pairs. In order to learn deep text understandings, we first pre-train the text understanding module of WSTQ on TM and then fine-tune it on TQA. We build positive as well as negative relation pairs by checking whether there is any overlap between the items/regions detected from diagrams using object detection. The RD task forces our method to learn the relationships between regions, which are crucial to express the diagram semantics. We train WSTQ on RD and TQA simultaneously, i.e., multitask learning, to obtain effective diagram semantics and then improve the TQA performance. Extensive experiments are carried out on CK12-QA and AI2D to verify the effectiveness of WSTQ. Experimental results show that our method achieves significant accuracy improvements of 5.02% and 4.12% on test splits of the above datasets respectively than the current state-of-the-art baseline. We have released our code on https://github.com/dr-majie/WSTQ.
Jie Ma 0001, Qi Chai, Jingyue Huang, Jun Liu 0002, Yang You 0001
IEEE Trans. Image Process.3
2021 Large-Scale Multi-granular Concept Extraction Based on Machine Reading Comprehension
Deqing Yang, Jiaqing Liang, Jilun Sun, Jingyue Huang, Kaiyan Cao, Yanghua Xiao, Rui Xie 0005
ISWC5
2021 Communication-avoiding kernel ridge regression on parallel and distributed systems
Yang You 0001, Jingyue Huang, Cho-Jui Hsieh, Richard W. Vuduc, James Demmel
CCF Trans. High Perform. Comput.2
2018 SWISS: Spectrum weighted identification of signal sources for mmWave systems
abstract
This paper considers the channel estimation problem in millimeter-wave (mmWave) systems where a single-antenna user communicates with a massive multiple-input multiple-output (MIMO) base station (BS) in the uplink. Unlike many existing works which estimate the channel gain under the assumption that the number of channel paths is given a priori, we address first the problem of path-number identification. By taking the weighted discrete Fourier transform (WDFT) of the received noisy signal, we formulate an optimization problem to determine the optimum combination of DFT components in this weighted spectrum that leads to a time-domain reconstructed signal (the channel vector) that is at the minimum Euclidean distance from the received signal. Our algorithm, called SWISS (Spectrum Weighted Identification of Signal Sources), is an accurate and computationally efficient means for identifying the paths in the channel vector, providing the information needed for BS beamforming. Once the paths are identified, their individual directions-of-arrival (DoAs) and complex fading gains can be obtained easily. Simulation results for the case of no power leakage in the DFT are presented to demonstrate the effectiveness of SWISS.
Ziming Cheng, Jingyue Huang, Meixia Tao, Pooi Yuen Kam
WCNC2
2017 Low-complexity hybrid analog/digital beamforming for multicast transmission in mmwave systems
abstract
This paper studies multi-group multicast beamforming with a hybrid large-scale antenna array in millimeter wave (mmWave) communication systems. A low-complexity hybrid structure is adopted, where each RF chain is only connected to part of the antenna elements. We formulate a hybrid analog and digital beamforming design problem for multi-group multicast transmission with the objective of minimizing the total transmit power at the base station, subject to an individual signal-to-interference-plus-noise ratio constraint for each multicast group. The problem is very challenging and its global optimal solution is difficult to obtain. We first adopt alternating minimization method to design the analog and digital beamformer alternatively. Then we solve each of the analog and digital subproblems through solving a sequence of convex problems via concave-convex procedure (CCP). Each convex CCP subproblem is reformulated as a novel alternating direction method of multipliers (ADMM) form. Our ADMM reformulation enables that each updating step can be decomposed into multiple subproblems with much smaller size, which can be solved optimally in parallel with closed-form expressions. Simulation results show that our algorithm can achieve favorable performance with very low complexity compared with the state-of-art methods.
Jingyue Huang, Ziming Cheng, Erkai Chen, Meixia Tao
ICC1