Jo-Ku Cheng

dblp:382/4643 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0004-4165-1384ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Geometric modeling and processing · 100%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%

Topics — the 1 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computing education
problem generation
0.312025
GeoUni: A Unified Model for Generating Geometry Diagrams, Problems and Problem Solutions · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

unified model · 2.6
YearPublicationVenuePosition
2025 Enhancing Large Language Models on Domain-specific Tasks: A Novel Training Strategy via Domain Adaptation and Preference Alignment
abstract
In handling complex, domain-specific tasks, particularly in the context of state-owned assets and enterprises (SOAEs), general LLMs suffer from the knowledge gap due to insufficient exposure to domain-specific corpora, and the value disagreement, as they are aligned with universal values rather than domain-specific ones. To tackle these challenges, we propose a novel training strategy tailored for the SOAEs domain. This strategy includes a improved domain-adaptive pretraining (DAP) phase with a replay mechanism to mitigate catastrophic forgetting. Following DAP, we utilize a selective portion of domain-specific data for supervised fine-tuning (SFT), and innovatively integrate low-quality data with the remaining SFT data to curate tailored preference datasets, leveraging the Kahneman-Tversky Optimization technique to align our LLMs. Our proposed approach effectively utilizes the data that is often discarded in conventional training procedures, highlighting the substantial improvements in model performance and the importance of training methodologies for domain-specific tasks.
Jingyang Deng, Zeren Zhang, Jo-Ku Cheng, Jinwen Ma
ICASSP3
2025 Diagram Formalization Enhanced Multi-Modal Geometry Problem Solver
abstract
Mathematical reasoning remains an ongoing challenge for AI models, especially for geometry problems, which require both linguistic and visual signals. As the vision encoders of most MLLMs are trained on natural scenes, they often struggle to understand geometric diagrams, performing no better in geometry problem-solving than LLMs that only process text. This limitation is further amplified by the lack of effective methods for representing geometric relationships. To address these issues, we introduce the Diagram Formalization Enhanced Geometry Problem Solver (DFE-GPS), a new framework that integrates visual features, geometric formal language, and natural language representations. Specifically, we propose a novel synthetic data approach and construct a large-scale geometric dataset, SynthGeo228K, annotated with formal and natural language captions, designed to enhance the vision encoder to understand geometric structures better. Our framework improves MLLMs’ ability to process geometric diagrams and extends their application to open-ended tasks on the formalgeo7k dataset.
Zeren Zhang, Jo-Ku Cheng, Jingyang Deng, Jinwen Ma, Ziran Qin, Tuo Leng
ICASSP2
2025 SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space
abstract
Combining face-swapping with lip synchronization offers a cost-effective solution for generating customized talking faces. However, directly cascading existing models can introduce significant interference and reduce video clarity due to limited interaction space in the low-level RGB domain. To solve this, we propose SwapTalk, a unified framework that performs face-swapping and lip synchronization within the same latent VQ-embedding space, known for its editability and fidelity. We enhance generalization to unseen identities with identity loss in the face-swapping module and improve synchronization quality with expert discriminator supervision. To better approximate real-world applications, we expand the evaluation scope to asynchronous audio-video scenarios. Furthermore, we introduce a novel identity consistency metric to more comprehensively assess the identity consistency over time series in generated facial videos. Experiments on HDTF show that SwapTalk outperforms existing methods in video quality, lip synchronization accuracy, face-swapping fidelity, and identity consistency.
Zeren Zhang, Haibo Qin, Jo-Ku Cheng, Yitao Duan, Jinwen Ma
ICASSP4
2025 GeoUni: A Unified Model for Generating Geometry Diagrams, Problems and Problem Solutions
Jo-Ku Cheng, Zeren Zhang, Ran Chen 0002, Jingyang Deng, Ziran Qin, Jinwen Ma
ACM Multimedia1
2024 A Fusion Framework of Whitespace Smear Cutting and Swin Transformer for Document Layout Analysis
Ran Chen 0002, Jo-Ku Cheng, Jinwen Ma
ICIC (6)2