Jianzhong Ju

dblp:289/4197 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-0581-7066ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Vision and language · 29% Language models and text generation · 25% Multi-agent systems · 14%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
1.222026
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution · AAAI 2026
Computer vision › 3D vision › 3d scene understanding
3d visual grounding
1.012026
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution · AAAI 2026
Knowledge, reasoning and agents › Multi-agent systems › autonomous agents
embodied agent
1.012026
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution · AAAI 2026
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › scheduling
task scheduling
1.012026
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution · AAAI 2026
Natural language and speech › Language models and text generation › large language model training
post-training
0.912025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model training › post-training
reinforcement learning post-training
0.912025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025
Computer vision › Vision and language
temporal grounding
0.912025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards
0.312025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

scheduling token mechanism · 1.0verifiable reward · 0.9supervised fine-tuning · 0.9reinforcement learning · 0.9
YearPublicationVenuePosition
2026 Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution
abstract
Task scheduling has become increasingly critical for embodied AI, where agents need to follow natural language instructions and execute actions efficiently in 3D physical worlds. Existing datasets for task planning in 3D environments often simplify the problem, lacking operations research knowledge for task scheduling and 3D grounding for real-world applications. In this work, we propose Operations Research Knowledge-based 3D Grounded Task Scheduling (OKS3D), a new task that requires synerization of language understanding, 3D grounding, and efficiency optimization for embodied agents. OKS3D reflects real-world demands by requiring agents to generate efficient, step-by-step schedules that are grounded in 3D space. To facilitate research on OKS3D, we construct a large-scale dataset called OKS3D-60K, comprising 60K tasks across 4K real-world scenes. Furthermore, we propose GRANT, an embodied multi-modal large language model equipped with a simple yet effective scheduling token mechanism to generate efficient task schedules and grounded actions. Extensive experiments on the OKS3D-60K dataset validate the effectiveness of GRANT across language understanding, 3D grounding, and scheduling efficiency.
Dingkang Liang, Cheng Zhang 0020, Xiaopeng Xu, Jianzhong Ju, Zhenbo Luo, Xiang Bai
AAAI4
2025 LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
abstract
Recent advances in large vision-language models (LVLMs) typically employ vision encoders based on the Vision Transformer (ViT) architecture. The division of the images into patches by ViT results in a fragmented perception, thereby hindering the visual understanding capabilities of LVLMs. In this paper, we propose an innovative enhancement to address this limitation by introducing a Scene Graph Expression (SGE) module in LVLMs. This module extracts and structurally expresses the complex semantic information within images, thereby improving the foundational perception and understanding abilities of LVLMs. Extensive experiments demonstrate that integrating our SGE module significantly enhances the LVLM’s performance in vision-language tasks, indicating its effectiveness in preserving intricate semantic details and facilitating better visual understanding.
Jianzhong Ju, Jian Luan 0001, Zhidong Deng
ICASSP2
2025 Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains
abstract
Large Language Models (LLMs) achieve superior performance through Chain-of-Thought (CoT) reasoning, but these token-level reasoning chains are computationally expensive and inefficient. In this paper, we introduce Compressed Latent Reasoning (CoLaR), a novel framework that dynamically compresses reasoning processes in latent space through a two-stage training approach. First, during supervised fine-tuning, CoLaR extends beyond next-token prediction by incorporating an auxiliary next compressed embedding prediction objective. This process merges embeddings of consecutive tokens using a compression factor $c$ randomly sampled from a predefined range, and trains a specialized latent head to predict distributions of subsequent compressed embeddings. Second, we enhance CoLaR through reinforcement learning (RL) that leverages the latent head's non-deterministic nature to explore diverse reasoning paths and exploit more compact ones. This approach enables CoLaR to: i) **perform reasoning at a dense latent level** (i.e., silently), substantially reducing reasoning chain length, and ii) **dynamically adjust reasoning speed** at inference time by simply prompting the desired compression factor. Extensive experiments across four mathematical reasoning datasets demonstrate that CoLaR achieves 14.1% higher accuracy than latent-based baseline methods at comparable compression ratios, and reduces reasoning chain length by 53.3% with only 4.8% performance degradation compared to explicit CoT method. Moreover, when applied to more challenging mathematical reasoning tasks, our RL-enhanced CoLaR demonstrates performance gains of up to 5.4% while dramatically reducing latent reasoning chain length by 82.8%. The code and models will be released upon acceptance.
Wenhui Tan, Jianzhong Ju, Zhenbo Luo, Ruihua Song, Jian Luan 0001
NeurIPS3
2025 Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
abstract
Temporal Video Grounding (TVG), the task of locating specific video segments based on language queries, is a core challenge in long-form video understanding. While recent Large Vision-Language Models (LVLMs) have shown early promise in tackling TVG through supervised fine-tuning (SFT), their ability to generalize remains limited. To address this, we propose a novel post-training framework that enhances the generalization capabilities of LVLMs via reinforcement learning (RL). Specifically, our contributions span three key directions: (1) Time-R1: we introduce a reasoning-guided post-training framework via RL with verifiable reward to enhance capabilities of LVLMs on the TVG task. (2) TimeRFT: we explore post-training strategies on our curated RL-friendly dataset, which trains the model to progressively comprehend more difficult samples, leading to better generalization. (3) TVGBench: we carefully construct a small but comprehensive and balanced benchmark suitable for LVLM evaluation, which is sourced from available public benchmarks. Extensive experiments demonstrate that Time-R1 achieves state-of-the-art performance across multiple downstream datasets using significantly less training data than prior LVLM approaches, while improving its general video understanding capabilities. Project Page: https://xuboshen.github.io/Time-R1/.
Boshen Xu, Yang Du 0011, Kejun Lin, Zihan Xiao 0001, Zihao Yue, Jianzhong Ju, Dingyi Yang, Xiangnan Fang, Zewen He, Zhenbo Luo, Wenxuan Wang 0001, Junqi Lin, Jian Luan 0001, Qin Jin
NeurIPS8
2023 Toward Ultrasonic Wire Bonding for High Power Device: A Vector Based Resonant Frequency Tracking and Constant Amplitude Control
abstract
The electrical packaging for high power devices such as IGBT or MOSFET is achieved by heavy aluminum ultrasonic wire bonding. In this process, the amplitude stability of ultrasonic transducer determines the bonding quality, which demands more accuracy in resonant frequency tracking (RFT) and constant amplitude control (CAC). In this investigation, we proposed a vector based resonant frequency (VRFT) tracking and vector based constant amplitude control (VCAC) to control the ultrasonic transducer. The current of dynamic branch was conducted by calculating the voltage and current vector of ultrasonic transducer. Small signal model of ultrasonic transducer was established and applied into RFT and CAC. A control strategy of fuzzy self-tuning PID closed loop was utilized to RFT and CAC. The PID parameters was calculated by the linear control theories and modified by the fuzzy controller which improved the vibration stability of the transducer. A prototype of 100Watt/60kHz ultrasonic driver was self-developed to verify the VRFT and VCAC. It demonstrated that the locked frequency of VRFT can be controlled within less than 0.04 millisecond when staring from the swept frequency, and the vibration response is fast with the varying bonding pressure. The amplitude of fuzzy PID VCAC is more stable and robust compared to the classical PID control, which improved the bonded joint quality in heavy aluminum wire packaging. Note to Practitioners—This article was motivated by the reliability and efficiency of IGBT/ MOSFET device packaging through ultrasonic aluminum wire bonding. This bonding process requires more amplitude stability of the ultrasonic transducer to achieve acceptable bonded quality. Currently, the resonant frequency tracking (RFT) by classical PID control is in slow speed, and the amplitude of the transducer cannot be controlled constantly due to the inherent static capacitor of the transducer. These issues for the ultrasonic control are originated from the engineer’s experience or simple heuristic rule instead of advanced approaches. To improve the vibration performance, we proposed a vector based model consisting of the relationship the dynamic branch current and the total current of the transducer, to control the RFT and constant amplitude control (CAC). Moreover, the vector based is utilized into the fuzzy PID control, where the calculation is reduced by area coefficient defuzzy. The verification experiments prove that the response time of the proposed vector based RFT and CAC, and the bonded joint quality is optimized obviously. It is noted that in our control strategy, the RMS value, the phase of voltage and current of the transducer are acquired by DSP processor, and the calculation of fuzzy PID controller is not complicated, which is feasible to implement in application. This significantly benefits the demand of accurate control to the ultrasonic power in electronic packaging.
Shuyuan Ye, Zhili Long, Jianzhong Ju, Tianyu Peng, Heng Zhao 0006
IEEE Trans Autom. Sci. Eng.3
2023 Investigation to Dual-Frequency Direct Digital Synthesis and Resonance Frequency Tracking in Power Ultrasonic Generator
abstract
We propose a novel dual-frequency direct digital synthesis (DDS) module and dual-frequency resonant frequency tracking (RFT) algorithm which can be applied to the power ultrasonic such as micro hole drilling applications. To attain the M-type vibration trajectory of the drilling tool, the dual frequency, which the first and the third order frequency of the transducer, is excited at the same time, and the vibration amplitude of the transducer is controlled in independent or coupling mode. Due to the resonant frequency of the transducer is shifted during the drilling process, the resonant frequency tracking (RFT) of the first-order frequency is developed. In order to accurately achieve the phase between the voltage and current, we propose a fourth order Butter-worth low pass filter to the zero-crossing circuits to realize RFT of the first-order frequency of the transducer. Moreover, a dual-frequency direct digital synthesis (DDS) which can generate a mixed signal output with independent frequency and vibration amplitude control is proposed. Finally, a 100 Watt prototype of dual-frequency ultrasonic generator is established to verify the proposed dual-frequency DDS and RFT.
Shuyuan Ye, Zhili Long, Heng Zhao 0006, Jianzhong Ju, Mian Yao, Xiangqing Li, Xuezhi Zhang
IEEE Trans. Circuits Syst. I Regul. Pap.4