EDBT 2026 Demo / reviewers in the wild / expert
Yujun Cai
dblp:227/4399
· DBLP profile ↗
49ranked-venue papers
7as first author
41since 2021 · last 2026
0000-0002-0993-4024ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 5 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAGabstractScaling multimodal large language models (MLLMs) to long videos is constrained by limited context windows.While retrievalaugmented generation (RAG) is a promising remedy by organizing query-relevant visual evidence into a compact context, most existing methods (i) flatten videos into independent segments, breaking their inherent spatio-temporal structure, and (ii) depend on explicit semantic matching, which can miss cues that are implicitly relevant to the query's intent.To overcome these limitations, we propose VideoStir, a structured and intent-aware long-video RAG framework.It firstly structures a video as a spatiotemporal graph at clip level, and then performs multi-hop retrieval to aggregate evidence across distant yet contextually related events.Furthermore, it introduces an MLLM-backed intentrelevance scorer that retrieves frames based on their alignment with the query's reasoning intent.To support this capability, we curate IR-600K, a large-scale dataset tailored for learning frame-query intent alignment.Experiments show that VideoStir is competitive with stateof-the-art baselines without relying on auxiliary information, highlighting the promise of shifting long-video RAG from flattened semantic matching to structured, intent-aware reasoning.Codes and checkpoints are available at https: //github.com/RomGai/VideoStir. Honghao Fu, Yiwei Wang 0001, Dailing Zhang, Jun Liu 0036, Yujun Cai |
ACL (1) | 6 |
| 2026 | Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language ModelsabstractDiffusion large language models (dLLMs) offer bidirectional attention and parallel generation, enabling them to exploit global context and naturally support format-constrained tasks like parseable JSON or reasoning templates.While straightforward fixed anchors can enforce such constraints, they often impose rigid spans, leading to truncated reasoning or redundant content.To overcome this, we propose Dynamic Infilling Anchors (DIA), a training-free method that dynamically estimates end-anchor positions to adjust generation length before iterative infilling.This flexible mechanism ensures structural correctness and semantic coherence, avoiding the inefficiencies of fixed-span methods.Experiments on reasoning benchmarks demonstrate that DIA substantially improves format compliance and answer accuracy, achieving significant zero-shot gains on GSM8K and MATH.These results establish DIA as a robust pathway toward reliable, structure-aware generation. Boyan Han, Yiwei Wang 0001, Yujun Cai, Chi Zhang 0007 |
ACL (1) | 4 |
| 2026 | Federated-Learning-Assisted RIS Active and Passive Beamforming With ADMM for IoT DevicesabstractFederated learning (FL) and reconfigurable intelligent surfaces (RIS) are pivotal technologies for future Internet of Things (IoT) networks, enhancing user privacy and system efficiency. However, realizing their full potential necessitates a cohesive and synergistic integration, challenging the traditional view of them as disparate components. This paper tackles the complex problem of maximizing energy efficiency (EE)—a critical yet under-explored metric insuch tightly coupled FL-RIS systems. We address this gap by formulating ajoint optimization problem that intrinsically links the FL process with physical layer resource allocation. Our framework maximizes the system’s global EE by concurrently designing the base station’s active beamforming and the RIS’s passive phase shifts,with an FL aggregation mechanism that is explicitly channel-aware and adaptive to the RIS-optimized wireless environment. This co-design ensures RIS actively facilitates FL by establishing robust communication, while FL intelligently leverages these improved channels for efficient and accelerated learning, all under practical FL performance constraints. Simulation results demonstrate that our proposed framework significantly enhances system energy efficiency compared to several benchmark schemes and exhibits robust convergence properties. Yujun Cai, Shufeng Li, Qianyun Zhang 0001, Zhijin Qin, Xinruo Zhang |
IEEE Internet Things J. | 1 |
| 2025 | Vulnerability of LLMs to Vertically Aligned Text ManipulationsabstractZhecheng Li, Yiwei Wang, Bryan Hooi, Yujun Cai, Zhen Xiong, Nanyun Peng, Kai-Wei Chang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhecheng Li, Yiwei Wang 0001, Bryan Hooi, Yujun Cai, Zhen Xiong, Nanyun Peng 0001, Kai-Wei Chang 0001 |
ACL (1) | 4 |
| 2025 | Con-ReCall: Detecting Pre-training Data in LLMs via Contrastive DecodingabstractThe training data in large language models is key to their success, but it also presents privacy and security risks, as it may contain sensitive information. Detecting pre-training data is crucial for mitigating these concerns. Existing methods typically analyze target text in isolation or solely with non-member contexts, overlooking potential insights from simultaneously considering both member and non-member contexts. While previous work suggested that member contexts provide little information due to the minor distributional shift they induce, our analysis reveals that these subtle shifts can be effectively leveraged when contrasted with non-member contexts. In this paper, we propose Con-ReCall, a novel approach that leverages the asymmetric distributional shifts induced by member and non-member contexts through contrastive decoding, amplifying subtle differences to enhance membership inference. Extensive empirical evaluations demonstrate that Con-ReCall achieves state-of-the-art performance on the WikiMIA benchmark and is robust against various text manipulation techniques. Yiwei Wang 0001, Bryan Hooi, Yujun Cai, Nanyun Peng 0001, Kai-Wei Chang 0001 |
COLING | 4 |
| 2025 | Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language ModelsabstractDespite inheriting security measures from underlying language models, Vision-Language Models (VLMs) may still be vulnerable to safety alignment issues. Through empirical analysis, we uncover two critical findings: scenario- matched images can significantly amplify harmful outputs, and contrary to common assumptions in gradient-based attacks, minimal loss values do not guarantee optimal attack effectiveness. Building on these insights, we introduce MLAI (Multi-Loss Adversarial Images), a novel jailbreak framework that leverages scenario-aware image generation for semantic alignment, exploits flat minima theory for robust adversarial image selection, and employs multi- image collaborative attacks for enhanced effectiveness. Extensive experiments demonstrate MLAI’s significant impact, achieving attack success rates of 77.75% on MiniGPT-4 and 82.80% on LLaVA-2, substantially outperforming existing methods by margins of 34.37% and 12.77% respectively. Furthermore, MLAI shows considerable transferability to commercial black-box VLMs, achieving up to 60.11% success rate. Our work reveals fundamental visual vulnerabilities in current VLMs safety mechanisms and underscores the need for stronger defenses. Warning: This paper contains potentially harmful example text. Shuyang Hao, Bryan Hooi, Jun Liu 0036, Kai-Wei Chang 0001, Zi Huang, Yujun Cai |
CVPR | 6 |
| 2025 | LatentHOI: On the Generalizable Hand Object Motion Generation with Latent Hand DiffusionabstractCurrent research on generating 3D hand-object interaction motion primarily focuses on in-domain objects. Generalization to unseen objects is essential for practical applications, yet it remains both challenging and largely unexplored. In this paper, we propose LatentHOI, a novel approach designed to tackle the challenges of generalizing hand-object interaction synthesis to unseen objects. Our main insight lies in decoupling high-level temporal motion from fine-grained spatial hand-object interactions via a latent diffusion model coupled with a Grasping Variational Autoencoder (Grasp-VAE). This configuration introduces regularization by enforcing a conditional dependency between spatial grasping and temporal motion, as well as through the regularized latent space for better generalization ability. We conducted extensive experiments in an unseen-object setting on both single-hand grasping and bi-manual motion datasets, including GRAB, DexYCB†, and OakInk. Quantitative and qualitative evaluations demonstrate that our method significantly enhances the realism and physical plausibility of generated motions for unseen objects, both in single and bimanual manipulations, compared to the state-of-the-art. Muchen Li, Sammy Joe Christen, Chengde Wan, Yujun Cai, Renjie Liao 0001, Leonid Sigal, Shugao Ma |
CVPR | 4 |
| 2025 | VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for MinecraftabstractLarge language models (LLMs) have shown significant promise in embodied decisionmaking tasks within virtual open-world environments.Nonetheless, their performance is hindered by the absence of domain-specific knowledge.Methods that finetune on largescale domain-specific data entail prohibitive development costs.This paper introduces Vista-Wise, a cost-effective agent framework that integrates cross-modal domain knowledge and finetunes a dedicated object detection model for visual analysis.It reduces the requirement for domain-specific training data from millions of samples to a few hundred.VistaWise integrates visual information and textual dependencies into a cross-modal knowledge graph (KG), enabling a comprehensive and accurate understanding of multimodal environments.We also equip the agent with a retrieval-based pooling strategy to extract task-related information from the KG, and a desktop-level skill library to support direct operation of the Minecraft desktop client via mouse and keyboard inputs.Experimental results demonstrate that VistaWise achieves state-of-the-art performance across various open-world tasks, highlighting its effectiveness in reducing development costs while enhancing agent performance. Honghao Fu, Junlong Ren, Qi Chai, Deheng Ye, Yujun Cai |
EMNLP | 5 |
| 2025 | SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global ThinkingabstractVision-language models (VLMs) excel in semantic tasks but falter at a core human capability: detecting hidden content in optical illusions or AI-generated images through perceptual adjustments like zooming.We introduce HC-Bench, a benchmark of 112 images with hidden texts, objects, and illusions, revealing that leading VLMs achieve near-zero accuracy (0-5.36%) even with explicit prompting.Humans resolve such ambiguities instinctively, yet VLMs fail due to an overreliance on high-level semantics.Strikingly, we propose SemVink (Semantic Visual Thinking) by simply scaling images to low resolutions, which unlocks over 99% accuracy by eliminating redundant visual noise.This exposes a critical architectural flaw: VLMs prioritize abstract reasoning over lowlevel visual operations crucial for real-world robustness.Our work urges a shift toward hybrid models integrating multi-scale processing, bridging the gap between computational vision and human cognition for applications in medical imaging, security, and beyond. Yujun Cai, Yiwei Wang 0001 |
EMNLP | 2 |
| 2025 | DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual ReasoningabstractGrounding natural language queries in graphical user interfaces (GUIs) poses unique challenges due to the diversity of visual elements, spatial clutter, and the ambiguity of language. In this paper, we introduce DiMo-GUI, a training-free framework for GUI grounding that leverages two core strategies: dynamic visual grounding and modality-aware optimization. Instead of treating the GUI as a monolithic image, our method splits the input into textual elements and iconic elements, allowing the model to reason over each modality independently using general-purpose vision-language models. When predictions are ambiguous or incorrect, DiMo-GUI dynamically focuses attention by generating candidate focal regions centered on the model’s initial predictions and incrementally zooms into subregions to refine the grounding result. This hierarchical refinement process helps disambiguate visually crowded layouts without the need for additional training or annotations. We evaluate our approach on standard GUI grounding benchmarks and demonstrate consistent improvements over baseline inference pipelines, highlighting the effectiveness of combining modality separation with region-focused reasoning. Yujun Cai, Chang Liu 0072, Qingwen Ye, Ming-Hsuan Yang 0001, Yiwei Wang 0001 |
EMNLP | 3 |
| 2025 | Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLMsabstractRecent advances in test-time scaling have enabled Large Language Models (LLMs) to display sophisticated reasoning abilities via extended Chain-of-Thought (CoT) generation.Despite their impressive reasoning abilities, Large Reasoning Models (LRMs) frequently display unstable behaviors, e.g., hallucinating unsupported premises, overthinking simple tasks, and displaying higher sensitivity to prompt variations.This raises a deeper research question: How can we represent the reasoning process of LRMs to map their minds?To address this, we propose a unified graph-based analytical framework for fine-grained modeling and quantitative analysis of LRM reasoning dynamics.Our method first clusters long, verbose CoT outputs into semantically coherent reasoning steps, then constructs directed reasoning graphs to capture contextual and logical dependencies among these steps.Through a comprehensive analysis of derived reasoning graphs, we also reveal that key structural properties, such as exploration density, branching, and convergence ratios, strongly correlate with models' performance.The proposed framework enables quantitative evaluation of internal reasoning structure and quality beyond conventional metrics and also provides practical insights for prompt engineering and cognitive analysis of LLMs.Code and resources will be released to facilitate future research in this direction. Zhen Xiong, Yujun Cai, Zhecheng Li, Yiwei Wang 0001 |
EMNLP | 2 |
| 2025 | Learning Few-Step Diffusion Models by Trajectory Distribution MatchingabstractAccelerating diffusion model sampling is crucial for efficient AIGC deployment. While diffusion distillation methods -- based on distribution matching and trajectory matching -- reduce sampling to as few as one step, they fall short on complex tasks like text-to-image generation. Few-step generation offers a better balance between speed and quality, but existing approaches face a persistent trade-off: distribution matching lacks flexibility for multi-step sampling, while trajectory matching often yields suboptimal image quality. To bridge this gap, we propose learning few-step diffusion models by Trajectory Distribution Matching (TDM), a unified distillation paradigm that combines the strengths of distribution and trajectory matching. Our method introduces a data-free score distillation objective, aligning the student's trajectory with the teacher's at the distribution level. Further, we develop a sampling-steps-aware objective that decouples learning targets across different steps, enabling more adjustable sampling. This approach supports both deterministic sampling for superior image quality and flexible multi-step adaptation, achieving state-of-the-art performance with remarkable efficiency. Our model, TDM, outperforms existing methods on various backbones, such as SDXL and PixArt-$α$, delivering superior quality and significantly reduced training costs. In particular, our method distills PixArt-$α$ into a 4-step generator that outperforms its teacher on real user preference at 1024 resolution. This is accomplished with 500 iterations and 2 A800 hours -- a mere 0.01% of the teacher's training cost. In addition, our proposed TDM can be extended to accelerate text-to-video diffusion. Notably, TDM can outperform its teacher model (CogVideoX-2B) by using only 4 NFE on VBench, improving the total score from 80.91 to 81.65. Project page: https://tdm-t2x.github.io/ Yihong Luo, Tianyang Hu 0001, Yujun Cai, Jing Tang 0004 |
ICCV | 4 |
| 2025 | Tricking Retrievers with Influential Tokens: An Efficient Black-Box Corpus Poisoning AttackabstractCheng Wang, Yiwei Wang, Yujun Cai, Bryan Hooi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yiwei Wang 0001, Yujun Cai, Bryan Hooi |
NAACL (Long Papers) | 3 |
| 2025 | HAIF-GS: Hierarchical and Induced Flow-Guided Gaussian Splatting for Dynamic SceneabstractReconstructing dynamic 3D scenes from monocular videos remains a fundamental challenge in 3D vision. While 3D Gaussian Splatting (3DGS) achieves real-time rendering in static settings, extending it to dynamic scenes is challenging due to the difficulty of learning structured and temporally consistent motion representations. This challenge often manifests as three limitations in existing methods: redundant Gaussian updates, insufficient motion supervision, and weak modeling of complex non-rigid deformations. These issues collectively hinder coherent and efficient dynamic reconstruction. To address these limitations, we propose HAIF-GS, a unified framework that enables structured and consistent dynamic modeling through sparse anchor-driven deformation. It first identifies motion-relevant regions via an Anchor Filter to suppress redundant updates in static areas. A self-supervised Induced Flow-Guided Deformation module induces anchor motion using multi-frame feature aggregation, eliminating the need for explicit flow labels. To further handle fine-grained deformations, a Hierarchical Anchor Propagation mechanism increases anchor resolution based on motion complexity and propagates multi-level transformations. Extensive experiments on synthetic and real-world benchmarks validate that HAIF-GS significantly outperforms prior dynamic 3DGS methods in rendering quality, temporal coherence, and reconstruction efficiency. Jianing Chen 0007, Yujun Cai, Hao Jiang 0013, Chengxuan Qian, Juyuan Kang, Shuqin Gao, Honglong Zhao, Tianlu Mao |
NeurIPS | 3 |
| 2025 | Federated Learning for Semantic Communication Based on CNNs and TransformerabstractThis study focuses on the latest research advancements in the field of semantic communication. Traditional communication systems prioritize the transmission of raw data, whilst semantic communication emphasizes conveying the meaning represented by the data. However, the extracted semantic information is often ambiguous and subject to subjective evaluation. To address this problem, this study proposes a model that combines a convolutional neural network (CNN) with a Transformer, called DeepSC‐CT. The model utilizes a CNN to extract semantic information from the data, followed by a Transformer model to capture spatial relationships and contextual information within the semantic content. We utilize federated learning to train the model and propose an adaptive aggregation algorithm to accelerate the convergence process. Moreover, we expand the single‐modality semantic communication model to encompass multiple modalities, such as texts, audio, and images. Furthermore, this study introduces a learnable position‐encoding method for the Transformer. The experimental results and visual effects of audio and image restoration demonstrate that the proposed method exhibits impressive performance and that the proposed model shows robust data restoration capabilities under various signal‐to‐noise ratio conditions. Shufeng Li, Yujun Cai, Zhaokai Deng, Xinran Ba, Qinghe Zheng, Xinruo Zhang, Baoxin Su |
Int. J. Intell. Syst. | 2 |
| 2025 | SED-MVS: Segmentation-Driven and Edge-Aligned Deformation Multi-View Stereo With Depth Restoration and Occlusion ConstraintabstractRecently, patch-deformation methods have exhibited significant effectiveness in multi-view stereo owing to the deformable and expandable patches in reconstructing textureless areas. However, existing approaches neglect to address the problem of deformation instability caused by easily overlooked edge-skipping, potentially leading to matching distortions, thus leaving room for further improvement. To fill this gap, we propose SED-MVS, which adopts panoptic segmentation and multi-trajectory diffusion strategy for segmentation-driven and edge-aligned patch deformation. Specifically, to prevent unanticipated edge-skipping, we first employ SAM2 for panoptic segmentation as depth-edge guidance to guide patch deformation, followed by multi-trajectory diffusion strategy to ensure patches are comprehensively aligned with depth edges. Moreover, to avoid potential inaccuracy of random initialization, we combine both sparse points from LoFTR and monocular depth map from DepthAnything V2 to restore reliable and realistic depth map for initialization and supervised guidance. Finally, we integrate the segmentation image with the monocular depth map to exploit inter-instance occlusion relationship, then further regard them as occlusion map to implement two distinct edge constraint, thereby facilitating occlusion-aware patch deformation. Extensive results on ETH3D, Tanks & Temples, BlendedMVS, Strecha and DL3DV-10K datasets validate the state-of-the-art performance and robust generalization capability of our proposed method. Zhenlong Yuan, Zhidong Yang, Yujun Cai, Kuangxin Wu, Mufan Liu, Hao Jiang 0013, Zhaoxin Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | STMG: A Machine Learning Microgesture Recognition System for Supporting Thumb-Based VR/AR InputabstractAR/VR devices have started to adopt hand tracking, in lieu of controllers, to support user interaction. However, today’s hand input rely primarily on one gesture: pinch. Moreover, current mappings of hand motion to use cases like VR locomotion and content scrolling involve more complex and larger arm motions than joystick or trackpad usage. STMG increases the gesture space by recognizing additional small thumb-based microgestures from skeletal tracking running on a headset. We take a machine learning approach and achieve a 95.1% recognition accuracy across seven thumb gestures performed on the index finger surface: four directional thumb swipes (left, right, forward, backward), thumb tap, and fingertip pinch start and pinch end. We detail the components to our machine learning pipeline and highlight our design decisions and lessons learned in producing a well generalized model. We then demonstrate how these microgestures simplify and reduce arm motions for hand-based locomotion and scrolling interactions. Kenrick Kin, Chengde Wan, Ken Koh, Andrei Marin, Necati Cihan Camgöz, Yujun Cai, Fedor Kovalev, Moshe Ben-Zacharia, Shannon Hoople, Marcos Nunes-Ueno, Mariel Sanchez-Rodriguez, Ayush Bhargava, Robert Wang 0002, Eric Sauser, Shugao Ma |
CHI | 7 |
| 2024 | LLMs are Good Action RecognizersabstractSkeleton-based action recognition has attracted lots of research attention. Recently, to build an accurate skeleton-based action recognizer, a variety of works have been pro-posed. Among them, some works use large model architectures as backbones of their recognizers to boost the skeleton data representation capability, while some other works pre-train their recognizers on external data to enrich the knowl-edge. In this work, we observe that large language models which have been extensively used in various natural language processing tasks generally hold both large model ar-chitectures and rich implicit knowledge. Motivated by this, we propose a novel LLM-AR framework, in which we in-vestigate treating the Large Language Model as an Action Recognizer. In our framework, we propose a linguistic pro-jection process to project each input action signal (i.e., each skeleton sequence) into its “sentence format” (i.e., an “action sentence”). Moreover, we also incorporate our frame-work with several designs to further facilitate this linguistic projection process. Extensive experiments demonstrate the efficacy of our proposed framework. Haoxuan Qu, Yujun Cai, Jun Liu 0036 |
CVPR | 2 |
| 2024 | 6D-Diff: A Keypoint Diffusion Framework for 6D Object Pose EstimationabstractEstimating the 6D object pose from a single RGB image often involves noise and indeterminacy due to challenges such as occlusions and cluttered backgrounds. Mean-while, diffusion models have shown appealing performance in generating high-quality images from random noise with high indeterminacy through step-by-step denoising. Inspired by their denoising capability, we propose a novel diffusion-based framework (6D-Diff) to handle the noise and indeterminacy in object pose estimation for better performance. In our framework, to establish accurate 2D-3D correspondence, we formulate 2D keypoints detection as a reverse diffusion (denoising) process. To facilitate such a denoising process, we design a Mixture-of-Cauchy-based forward diffusion process and condition the reverse process on the object appearance features. Extensive experiments on the LM-O and YCB-V datasets demonstrate the effectiveness of our framework. Haoxuan Qu, Yujun Cai, Jun Liu 0036 |
CVPR | 3 |
| 2024 | Energy-Calibrated VAE with Test Time Free Lunch
Yihong Luo, Siya Qiu, Xingjian Tao, Yujun Cai, Jing Tang 0004 |
ECCV (85) | 4 |
| 2024 | DisC-GS: Discontinuity-aware Gaussian SplattingabstractRecently, Gaussian Splatting, a method that represents a 3D scene as a collection of Gaussian distributions, has gained significant attention in addressing the task of novel view synthesis. In this paper, we highlight a fundamental limitation of Gaussian Splatting: its inability to accurately render discontinuities and boundaries in images due to the continuous nature of Gaussian distributions. To address this issue, we propose a novel framework enabling Gaussian Splatting to perform discontinuity-aware image rendering. Additionally, we introduce a B\'ezier-boundary gradient approximation strategy within our framework to keep the ``differentiability'' of the proposed discontinuity-aware rendering process. Extensive experiments demonstrate the efficacy of our framework. Haoxuan Qu, Zhuoling Li, Hossein Rahmani 0001, Yujun Cai, Jun Liu 0036 |
NeurIPS | 4 |
| 2024 | emg2pose: A Large and Diverse Benchmark for Surface Electromyographic Hand Pose EstimationabstractHands are the primary means through which humans interact with the world. Reliable and always-available hand pose inference could yield new and intuitive control schemes for human-computer interactions, particularly in virtual and augmented reality. Computer vision is effective but requires one or multiple cameras and can struggle with occlusions, limited field of view, and poor lighting. Wearable wrist-based surface electromyography (sEMG) presents a promising alternative as an always-available modality sensing muscle activities that drive hand motion. However, sEMG signals are strongly dependent on user anatomy and sensor placement; existing sEMG models have thus required hundreds of users and device placements to effectively generalize for tasks other than pose inference. To facilitate progress on sEMG pose inference, we introduce the emg2pose benchmark, which is to our knowledge the first publicly available dataset of high-quality hand pose labels and wrist sEMG recordings. emg2pose contains 2kHz, 16 channel sEMG and pose labels from a 26-camera motion capture rig for 193 users, 370 hours, and 29 stages with diverse gestures - a scale comparable to vision-based hand pose datasets. We provide competitive baselines and challenging tasks evaluating real-world generalization scenarios: held-out users, sensor placements, and stages. This benchmark provides the machine learning community a platform for exploring complex generalization problems, holding potential to significantly enhance the development of sEMG-based human-computer interactions. Sasha Salter, Richard Warren, Collin Schlager, Adrian Spurr, Shangchen Han, Rohin Bhasin, Yujun Cai, Peter Walkington, Anuoluwapo Bolarinwa, Robert Wang 0002, Nathan Danielson, Josh Merel, Eftychios A. Pnevmatikakis, Jesse Marshall |
NeurIPS | 7 |
| 2024 | RIS-Assisted Federated Learning Algorithm Based on Device Selection and Weighted AveragingabstractTo protect user privacy and improve the transmitting environment of wireless communication, federated learning (FL) and reconfigurable intelligent surface (RIS) are proposed as promising technologies for future communication. Meanwhile, studies have proved that the combination of FL and RIS guarantees better performance for system models. However, the combined model still has problems such as high communication overhead and slow convergence speed. Therefore, in this paper, we proposed a channel quality based device selection and weighted averaging algorithm in a RIS-assisted federated learning model. Simulation results proved that the proposed algorithm outperforms the classic federated averaging (FedAvg) algorithm in convergence speed, test accuracy, and training loss. Yujun Cai, Shufeng Li, Deyou Zhang |
VTC Spring | 1 |
| 2023 | How Fragile is Relation Extraction under Entity Replacements?abstractYiwei Wang, Bryan Hooi, Fei Wang, Yujun Cai, Yuxuan Liang, Wenxuan Zhou, Jing Tang, Manjuan Duan, Muhao Chen. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL). 2023. Yiwei Wang 0001, Bryan Hooi, Fei Wang 0060, Yujun Cai, Yuxuan Liang 0002, Wenxuan Zhou 0002, Jing Tang 0004, Manjuan Duan, Muhao Chen 0001 |
CoNLL | 4 |
| 2023 | A Characteristic Function-Based Method for Bottom-Up Human Pose EstimationabstractMost recent methods formulate the task of human pose estimation as a heatmap estimation problem, and use the overall L2 loss computed from the entire heatmap to optimize the heatmap prediction. In this paper, we show that in bottom-up human pose estimation where each heatmap often contains multiple body joints, using the overall L2 loss to optimize the heatmap prediction may not be the optimal choice. This is because, minimizing the overall L2 loss cannot always lead the model to locate all the body joints across different sub-regions of the heatmap more accurately. To cope with this problem, from a novel perspective, we propose a new bottom-up human pose estimation method that optimizes the heatmap prediction via minimizing the distance between two characteristic functions respectively constructed from the predicted heatmap and the groundtruth heatmap. Our analysis presented in this paper indicates that the distance between these two characteristic functions is essentially the upper bound of the L2 losses w.r.t. sub-regions of the predicted heatmap. Therefore, via minimizing the distance between the two characteristic functions, we can optimize the model to provide a more accurate localization result for the body joints in different sub-regions of the predicted heatmap. We show the effectiveness of our proposed method through extensive experiments on the COCO dataset and the CrowdPose dataset. Haoxuan Qu, Yujun Cai, Lin Geng Foo, Ajay Kumar 0001, Jun Liu 0036 |
CVPR | 2 |
| 2023 | Primacy Effect of ChatGPTabstractInstruction-tuned large language models (LLMs), such as ChatGPT, have led to promising zero-shot performance in discriminative natural language understanding (NLU) tasks.This involves querying the LLM using a prompt containing the question, and the candidate labels to choose from.The question-answering capabilities of ChatGPT arise from its pre-training on large amounts of human-written text, as well as its subsequent fine-tuning on human preferences, which motivates us to ask: Does ChatGPT also inherit humans' cognitive biases?In this paper, we study the primacy effect of ChatGPT: the tendency of selecting the labels at earlier positions as the answer.We have two main findings: i) ChatGPT's decision is sensitive to the order of labels in the prompt; ii) ChatGPT has a clearly higher chance to select the labels at earlier positions as the answer.We hope that our experiments and analyses provide additional insights into building more reliable ChatGPT-based solutions.We release the source code at https: //github.com/wangywUST/PrimacyEffectGPT. Yiwei Wang 0001, Yujun Cai, Muhao Chen 0001, Yuxuan Liang 0002, Bryan Hooi |
EMNLP | 2 |
| 2023 | Social Diffusion: Long-term Multiple Human Motion AnticipationabstractWe propose Social Diffusion, a novel method for short-term and long-term forecasting of the motion of multiple persons as well as their social interactions. Jointly forecasting motions for multiple persons involved in social activities is inherently a challenging problem due to the interdependencies between individuals. In this work, we leverage a diffusion model conditioned on motion histories and causal temporal convolutional networks to forecast individually and contextually plausible motions for all participants. The contextual plausibility is achieved via an order-invariant aggregation function. As a second contribution, we design a new evaluation protocol that measures the plausibility of social interactions which we evaluate on the Haggling dataset, which features a challenging social activity where people are actively taking turns to talk and switching their attention. We evaluate our approach on four datasets for multi-person forecasting where our approach outperforms the state-of-the-art in terms of motion realism and contextual plausibility. Julian Tanke, Linguang Zhang, Amy Zhao, Chengcheng Tang, Yujun Cai, Lezi Wang, Po-Chen Wu, Juergen Gall, Cem Keskin |
ICCV | 5 |
| 2023 | LMC: Large Model Collaboration with Cross-assessment for Training-Free Open-Set Object RecognitionabstractOpen-set object recognition aims to identify if an object is from a class that has been encountered during training or not. To perform open-set object recognition accurately, a key challenge is how to reduce the reliance on spurious-discriminative features. In this paper, motivated by that different large models pre-trained through different paradigms can possess very rich while distinct implicit knowledge, we propose a novel framework named Large Model Collaboration (LMC) to tackle the above challenge via collaborating different off-the-shelf large models in a training-free manner. Moreover, we also incorporate the proposed framework with several novel designs to effectively extract implicit knowledge from large models. Extensive experiments demonstrate the efficacy of our proposed framework. Code is available \href{https://github.com/Harryqu123/LMC}{here}. Haoxuan Qu, Xiaofei Hui, Yujun Cai, Jun Liu 0036 |
NeurIPS | 3 |
| 2023 | DeepEMD: Differentiable Earth Mover's Distance for Few-Shot LearningabstractIn this work, we develop methods for few-shot image classification from a new perspective of optimal matching between image regions. We employ the Earth Mover's Distance (EMD) as a metric to compute a structural distance between dense image representations to determine image relevance. The EMD generates the optimal matching flows between structural elements that have the minimum matching cost, which is used to calculate the image distance for classification. To generate the important weights of elements in the EMD formulation, we design a cross-reference mechanism, which can effectively alleviate the adverse impact caused by the cluttered background and large intra-class appearance variations. To implement k-shot classification, we propose to learn a structured fully connected layer that can directly classify dense image representations with the EMD. Based on the implicit function theorem, the EMD can be inserted as a layer into the network for end-to-end training. Our extensive experiments validate the effectiveness of our algorithm which outperforms state-of-the-art methods by a significant margin on five widely used few-shot classification benchmarks, namely, miniImageNet, tieredImageNet, Fewshot-CIFAR100 (FC100), Caltech-UCSD Birds-200-2011 (CUB), and CIFAR-FewShot (CIFAR-FS). We also demonstrate the effectiveness of our method on the image retrieval task in our experiments. Chi Zhang 0007, Yujun Cai, Guosheng Lin, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Geometry-Guided Progressive NeRF for Generalizable and Efficient Neural Human Rendering
Mingfei Chen, Xiangyu Xu 0002, Yujun Cai, Jiashi Feng, Shuicheng Yan |
ECCV (23) | 5 |
| 2022 | Time-Aware Neighbor Sampling on Temporal GraphsabstractWe present a new neighbor sampling method on temporal graphs. In a temporal graph, predicting different nodes' time-varying properties can require the receptive neigh-borhood of various temporal scales. In this work, we propose the TNS (Time-aware Neighbor Sampling) method: TNS learns from temporal information to provide an adaptive receptive neighborhood for every node at any time. Learning how to sample neighbors is non-trivial, since the neighbor indices in time order are discrete and not differentiable. To address this challenge, we transform neighbor indices from discrete values to continuous ones by interpolating the neighbors' messages. TNS can be flexibly incorporated into popular temporal graph networks to improve their effectiveness without increasing their time complexity. TNS can be trained in an end-to-end manner. It requires no extra supervision and is automatically and implicitly guided to sample the neighbors that are most beneficial for prediction. Empirical results on multiple standard datasets show that TNS yields significant gains on edge prediction and node classification. Yiwei Wang 0001, Yujun Cai, Yuxuan Liang 0002, Henghui Ding, Changhu Wang, Bryan Hooi |
IJCNN | 2 |
| 2022 | Should We Rely on Entity Mentions for Relation Extraction? Debiasing Relation Extraction with Counterfactual AnalysisabstractYiwei Wang, Muhao Chen, Wenxuan Zhou, Yujun Cai, Yuxuan Liang, Dayiheng Liu, Baosong Yang, Juncheng Liu, Bryan Hooi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yiwei Wang 0001, Muhao Chen 0001, Wenxuan Zhou 0002, Yujun Cai, Yuxuan Liang 0002, Dayiheng Liu, Baosong Yang, Bryan Hooi |
NAACL-HLT | 4 |
| 2022 | Heatmap Distribution Matching for Human Pose EstimationabstractFor tackling the task of 2D human pose estimation, the great majority of the recent methods regard this task as a heatmap estimation problem, and optimize the heatmap prediction using the Gaussian-smoothed heatmap as the optimization objective and using the pixel-wise loss (e.g. MSE) as the loss function. In this paper, we show that optimizing the heatmap prediction in such a way, the model performance of body joint localization, which is the intrinsic objective of this task, may not be consistently improved during the optimization process of the heatmap prediction. To address this problem, from a novel perspective, we propose to formulate the optimization of the heatmap prediction as a distribution matching problem between the predicted heatmap and the dot annotation of the body joint directly. By doing so, our proposed method does not need to construct the Gaussian-smoothed heatmap and can achieve a more consistent model performance improvement during the optimization of the heatmap prediction. We show the effectiveness of our proposed method through extensive experiments on the COCO dataset and the MPII dataset. Haoxuan Qu, Yujun Cai, Lin Geng Foo, Jun Liu 0036 |
NeurIPS | 3 |
| 2022 | UmeTrack: Unified multi-view end-to-end hand tracking for VRabstractReal-time tracking of 3D hand pose in world space is a challenging problem and plays an important role in VR interaction. Existing work in this space are limited to either producing root-relative (versus world space) 3D pose or rely on multiple stages such as generating heatmaps and kinematic optimization to obtain 3D pose. Moreover, the typical VR scenario, which involves multi-view tracking from wide field of view (FOV) cameras is seldom addressed by these methods. In this paper, we present a unified end-to-end differentiable framework for multi-view, multi-frame hand tracking that directly predicts 3D hand pose in world space. We demonstrate the benefits of end-to-end differentiabilty by extending our framework with downstream tasks such as jitter reduction and pinch prediction. To demonstrate the efficacy of our model, we further present a new large-scale egocentric hand pose dataset that consists of both real and synthetic data. Experiments show that our system trained on this dataset handles various challenging interactive motions, and has been successfully applied to real-time VR applications. Shangchen Han, Po-Chen Wu, Linguang Zhang, Weiguang Si, Peizhao Zhang, Yujun Cai, Tomas Hodan, Randi Cabezas, Luan Tran, Muzaffer Akbay, Tsz-Ho Yu, Cem Keskin, Robert Wang 0002 |
SIGGRAPH Asia | 9 |
| 2022 | MonaGO: a novel gene ontology enrichment analysis visualisation systemabstractBACKGROUND: Gene ontology (GO) enrichment analysis is frequently undertaken during exploration of various -omics data sets. Despite the wide array of tools available to biologists to perform this analysis, meaningful visualisation of the overrepresented GO in a manner which is easy to interpret is still lacking. RESULTS: Monash Gene Ontology (MonaGO) is a novel web-based visualisation system that provides an intuitive, interactive and responsive interface for performing GO enrichment analysis and visualising the results. MonaGO supports gene lists as well as GO terms as inputs. Visualisation results can be exported as high-resolution images or restored in new sessions, allowing reproducibility of the analysis. An extensive comparison between MonaGO and 11 state-of-the-art GO enrichment visualisation tools based on 9 features revealed that MonaGO is a unique platform that simultaneously allows interactive visualisation within one single output page, directly accessible through a web browser with customisable display options. CONCLUSION: MonaGO combines dynamic clustering and interactive visualisation as well as customisation options to assist biologists in obtaining meaningful representation of overrepresented GO terms, producing simplified outputs in an unbiased manner. MonaGO will facilitate the interpretation of GO analysis and will assist the biologists into the representation of the results. Ziyin Xin, Yujun Cai, Louis T. Dang, Hannah M. S. Burke, Jerico Revote, Natalie Charitakis, Denis Bienroth, Hieu T. Nim, Yuan-Fang Li, Mirana Ramialison |
BMC Bioinform. | 2 |
| 2021 | A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto-Encoder∗abstractWe present a unified and flexible framework to address the generalized problem of 3D motion synthesis that covers the tasks of motion prediction, completion, interpolation, and spatial-temporal recovery. Since these tasks have different input constraints and various fidelity and diversity requirements, most existing approaches only cater to a specific task or use different architectures to address various tasks. Here we propose a unified framework based on Conditional Variational Auto-Encoder (CVAE), where we treat any arbitrary input as a masked motion series. Notably, by considering this problem as a conditional generation process, we estimate a parametric distribution of the missing regions based on the input conditions, from which to sample and synthesize the full motion series. To further allow the flexibility of manipulating the motion style of the generated series, we design an Action-Adaptive Modulation (AAM) to propagate the given semantic guidance through the whole sequence. We also introduce a cross-attention mechanism to exploit distant relations among decoder and encoder features for better realism and global consistency. We conducted extensive experiments on Human 3.6M and CMU-Mocap. The results show that our method produces coherent and realistic results for various motion synthesis tasks, with the synthesized motions distinctly adapted by the given action labels. Yujun Cai, Yiwei Wang 0001, Yiheng Zhu 0003, Tat-Jen Cham, Jianfei Cai 0001, Junsong Yuan 0001, Jun Liu 0036, Chuanxia Zheng, Sijie Yan, Henghui Ding, Xiaohui Shen, Ding Liu 0001, Nadia Magnenat-Thalmann |
ICCV | 1 |
| 2021 | Adaptive Data Augmentation on Temporal GraphsabstractTemporal Graph Networks (TGNs) are powerful on modeling temporal graph data based on their increased complexity. Higher complexity carries with it a higher risk of overfitting, which makes TGNs capture random noise instead of essential semantic information. To address this issue, our idea is to transform the temporal graphs using data augmentation (DA) with adaptive magnitudes, so as to effectively augment the input features and preserve the essential semantic information. Based on this idea, we present the MeTA (Memory Tower Augmentation) module: a multi-level module that processes the augmented graphs of different magnitudes on separate levels, and performs message passing across levels to provide adaptively augmented inputs for every prediction. MeTA can be flexibly applied to the training of popular TGNs to improve their effectiveness without increasing their time complexity. To complement MeTA, we propose three DA strategies to realistically model noise by modifying both the temporal and topological features. Empirical results on standard datasets show that MeTA yields significant gains for the popular TGN models on edge prediction and node classification in an efficient manner. Yiwei Wang 0001, Yujun Cai, Yuxuan Liang 0002, Henghui Ding, Changhu Wang, Siddharth Bhatia 0001, Bryan Hooi |
NeurIPS | 2 |
| 2021 | Direct Multi-view Multi-person 3D Pose EstimationabstractWe present Multi-view Pose transformer (MvP) for estimating multi-person 3D poses from multi-view images. Instead of estimating 3D joint locations from costly volumetric representation or reconstructing the per-person 3D pose from multiple detected 2D poses as in previous methods, MvP directly regresses the multi-person 3D poses in a clean and efficient way, without relying on intermediate tasks. Specifically, MvP represents skeleton joints as learnable query embeddings and let them progressively attend to and reason over the multi-view information from the input images to directly regress the actual 3D joint locations. To improve the accuracy of such a simple pipeline, MvP presents a hierarchical scheme to concisely represent query embeddings of multi-person skeleton joints and introduces an input-dependent query adaptation approach. Further, MvP designs a novel geometrically guided attention mechanism, called projective attention, to more precisely fuse the cross-view information for each joint. MvP also introduces a RayConv operation to integrate the view-dependent camera geometry into the feature representations for augmenting the projective attention. We show experimentally that our MvP model outperforms the state-of-the-art methods on several benchmarks while being much more efficient. Notably, it achieves 92.3% AP25 on the challenging Panoptic dataset, improving upon the previous best approach [35] by 9.8%. MvP is general and also extendable to recovering human mesh represented by the SMPL model, thus useful for modeling multi-person body shapes. Code and models are available at https://github.com/sail-sg/mvp. Tao Wang 0053, Yujun Cai, Shuicheng Yan, Jiashi Feng |
NeurIPS | 3 |
| 2021 | CurGraph: Curriculum Learning for Graph ClassificationabstractGraph neural networks (GNNs) have achieved state-of-the-art performance on graph classification tasks. Existing work usually feeds graphs to GNNs in random order for training. However, graphs can vary greatly in their difficulty for classification, and we argue that GNNs can benefit from an easy-to-difficult curriculum, similar to the learning process of humans. Evaluating the difficulty of graphs is challenging due to the high irregularity of graph data. To address this issue, we present the CurGraph (Curriculum Learning for Graph Classification) framework, that analyzes the graph difficulty in the high-level semantic feature space. Specifically, we use the infomax method to obtain graph-level embeddings and a neural density estimator to model the embedding distributions. Then we calculate the difficulty scores of graphs based on the intra-class and inter-class distributions of their embeddings. Given the difficulty scores, CurGraph first exposes a GNN to easy graphs, before gradually moving on to hard ones. To provide a soft transition from easy to hard, we propose a smooth-step method, which utilizes a time-variant smooth function to filter out hard graphs. Thanks to CurGraph, a GNN learns from the graphs at the border of its capability, neither too easy or too hard, to gradually expand its border at each training step. Empirically, CurGraph yields significant gains for popular GNN models on graph classification and enables them to achieve superior performance on miscellaneous graphs. Yiwei Wang 0001, Wei Wang 0059, Yuxuan Liang 0002, Yujun Cai, Bryan Hooi |
WWW | 4 |
| 2021 | Mixup for Node and Graph ClassificationabstractMixup is an advanced data augmentation method for training neural network based image classifiers, which interpolates both features and labels of a pair of images to produce synthetic samples. However, devising the Mixup methods for graph learning is challenging due to the irregularity and connectivity of graph data. In this paper, we propose the Mixup methods for two fundamental tasks in graph learning: node and graph classification. To interpolate the irregular graph topology, we propose the two-branch graph convolution to mix the receptive field subgraphs for the paired nodes. Mixup on different node pairs can interfere with the mixed features for each other due to the connectivity between nodes. To block this interference, we propose the two-stage Mixup framework, which uses each node’s neighbors’ representations before Mixup for graph convolutions. For graph classification, we interpolate complex and diverse graphs in the semantic space. Qualitatively, our Mixup methods enable GNNs to learn more discriminative features and reduce over-fitting. Quantitative results show that our method yields consistent gains in terms of test accuracy and F1-micro scores on standard datasets, for both node and graph classification. Overall, our method effectively regularizes popular graph neural networks for better generalization without increasing their time complexity. Yiwei Wang 0001, Wei Wang 0059, Yuxuan Liang 0002, Yujun Cai, Bryan Hooi |
WWW | 4 |
| 2021 | 3D Hand Pose Estimation Using Synthetic Data and Weakly Labeled RGB ImagesabstractCompared with depth-based 3D hand pose estimation, it is more challenging to infer 3D hand pose from monocular RGB images, due to the substantial depth ambiguity and the difficulty of obtaining fully-annotated training data. Different from the existing learning-based monocular RGB-input approaches that require accurate 3D annotations for training, we propose to leverage the depth images that can be easily obtained from commodity RGB-D cameras during training, while during testing we take only RGB inputs for 3D joint predictions. In this way, we alleviate the burden of the costly 3D annotations in real-world dataset. Particularly, we propose a weakly-supervised method, adaptating from fully-annotated synthetic dataset to weakly-labeled real-world single RGB dataset with the aid of a depth regularizer, which serves as weak supervision for 3D pose prediction. To further exploit the physical structure of 3D hand pose, we present a novel CVAE-based statistical framework to embed the pose-specific subspace from RGB images, which can then be used to infer the 3D hand joint locations. Extensive experiments on benchmark datasets validate that our proposed approach outperforms baselines and state-of-the-art methods, which proves the effectiveness of the proposed depth regularizer and the CVAE-based framework. Yujun Cai, Liuhao Ge, Jianfei Cai 0001, Nadia Magnenat-Thalmann, Junsong Yuan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | DeepEMD: Few-Shot Image Classification With Differentiable Earth Mover's Distance and Structured ClassifiersabstractIn this paper, we address the few-shot classification task from a new perspective of optimal matching between image regions. We adopt the Earth Mover's Distance (EMD) as a metric to compute a structural distance between dense image representations to determine image relevance. The EMD generates the optimal matching flows between structural elements that have the minimum matching cost, which is used to represent the image distance for classification. To generate the important weights of elements in the EMD formulation, we design a cross-reference mechanism, which can effectively minimize the impact caused by the cluttered background and large intra-class appearance variations. To handle k-shot classification, we propose to learn a structured fully connected layer that can directly classify dense image representations with the EMD. Based on the implicit function theorem, the EMD can be inserted as a layer into the network for end-to-end training. We conduct comprehensive experiments to validate our algorithm and we set new state-of-the-art performance on four popular few-shot classification benchmarks, namely miniImageNet, tieredImageNet, Fewshot-CIFAR100 (FC100) and Caltech-UCSD Birds-200-2011 (CUB). Chi Zhang 0007, Yujun Cai, Guosheng Lin, Chunhua Shen |
CVPR | 2 |
| 2020 | Learning Progressive Joint Propagation for Human Motion Prediction
Yujun Cai, Lin Huang 0004, Yiwei Wang 0001, Tat-Jen Cham, Jianfei Cai 0001, Junsong Yuan 0001, Jun Liu 0036, Xu Yang 0021, Yiheng Zhu 0003, Xiaohui Shen, Ding Liu 0001, Jing Liu 0050, Nadia Magnenat-Thalmann |
ECCV (7) | 1 |
| 2020 | Detecting Implementation Bugs in Graph Convolutional Network based Node ClassifiersabstractGraph convolutional networks (GCNs) have achieved state-of-the-art performance on the task of node classification. However, the performance of GCNs is prone to implementation bugs that do not explicitly produce compile-time or run-time errors but degrade their effectiveness heavily. These bugs are hard to detect, since the way in which the node attributes and graph structures contribute to the outputs is complicated, non-transparent, and not traceable by humans. To address this issue, we propose a systematic approach with formal justifications to detect implementation bugs in GCN based node classifiers. Our approach is based on the idea of Metamorphic Testing, which does not check input-output relations for a single input, but for input-output pairs. To speed up our approach, we design a pipeline system, which synchronizes the workload on CPUs and GPUs adaptively and processes them simultaneously. Our empirical study shows that our approach is able to identify over 80% of the synthetic mutants and two real-world bugs in GCN implementations. In addition, our pipeline system can achieve more than 10× speedup over the sequential system that leaves the CPU/GPU idle when using the other. Yiwei Wang 0001, Wei Wang 0059, Yujun Cai, Bryan Hooi, Beng Chin Ooi |
ISSRE | 3 |
| 2020 | NodeAug: Semi-Supervised Node Classification with Data AugmentationabstractBy using Data Augmentation (DA), we present a new method to enhance Graph Convolutional Networks (GCNs), that are the state-of-the-art models for semi-supervised node classification. DA for graph data remains under-explored. Due to the connections built by edges, DA for different nodes influence each other and lead to undesired results, such as uncontrollable DA magnitudes and changes of ground-truth labels. To address this issue, we present the NodeAug (Node-Parallel Augmentation) scheme, that creates a 'parallel universe' for each node to conduct DA, to block the undesired effects from other nodes. NodeAug regularizes the model prediction of every node (including unlabeled) to be invariant with respect to changes induced by Data Augmentation (DA), so as to improve the effectiveness. To augment the input features from different aspects, we propose three DA strategies by modifying both node attributes and the graph structure. In addition, we introduce the subgraph mini-batch training for the efficient implementation of NodeAug. The approach takes the subgraph corresponding to the receptive fields of a batch of nodes as the input per iteration, rather than the whole graph that the prior full-batch training takes. Empirically, NodeAug yields significant gains for strong GCN models on the Cora, Citeseer, Pubmed, and two co-authorship networks, with a more efficient training process thanks to the proposed subgraph mini-batch training approach. Yiwei Wang 0001, Wei Wang 0059, Yuxuan Liang 0002, Yujun Cai, Bryan Hooi |
KDD | 4 |
| 2020 | Progressive Supervision for Node Classification
Yiwei Wang 0001, Wei Wang 0059, Yuxuan Liang 0002, Yujun Cai, Bryan Hooi |
ECML/PKDD (1) | 4 |
| 2019 | Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksabstractDespite great progress in 3D pose estimation from single-view images or videos, it remains a challenging task due to the substantial depth ambiguity and severe self-occlusions. Motivated by the effectiveness of incorporating spatial dependencies and temporal consistencies to alleviate these issues, we propose a novel graph-based method to tackle the problem of 3D human body and 3D hand pose estimation from a short sequence of 2D joint detections. Particularly, domain knowledge about the human hand (body) configurations is explicitly incorporated into the graph convolutional operations to meet the specific demand of the 3D pose estimation. Furthermore, we introduce a local-to-global network architecture, which is capable of learning multi-scale features for the graph-based representations. We evaluate the proposed method on challenging benchmark datasets for both 3D hand pose estimation and 3D body pose estimation. Experimental results show that our method achieves state-of-the-art performance on both tasks. Yujun Cai, Liuhao Ge, Jun Liu 0036, Jianfei Cai 0001, Tat-Jen Cham, Junsong Yuan 0001, Nadia Magnenat-Thalmann |
ICCV | 1 |
| 2018 | Hand PointNet: 3D Hand Pose Estimation Using Point SetsabstractConvolutional Neural Network (CNN) has shown promising results for 3D hand pose estimation in depth images. Different from existing CNN-based hand pose estimation methods that take either 2D images or 3D volumes as the input, our proposed Hand PointNet directly processes the 3D point cloud that models the visible surface of the hand for pose regression. Taking the normalized point cloud as the input, our proposed hand pose regression network is able to capture complex hand structures and accurately regress a low dimensional representation of the 3D hand pose. In order to further improve the accuracy of fingertips, we design a fingertip refinement network that directly takes the neighboring points of the estimated fingertip location as input to refine the fingertip location. Experiments on three challenging hand pose datasets show that our proposed method outperforms state-of-the-art methods. Liuhao Ge, Yujun Cai, Junwu Weng, Junsong Yuan 0001 |
CVPR | 2 |
| 2018 | Weakly-Supervised 3D Hand Pose Estimation from Monocular RGB Images
Yujun Cai, Liuhao Ge, Jianfei Cai 0001, Junsong Yuan 0001 |
ECCV (6) | 1 |