EDBT 2026 Demo / reviewers in the wild / expert
Bang Zhang
dblp:11/4046
· DBLP profile ↗
63ranked-venue papers
11as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 7 first-author · 21 since 2021Artificial intelligence and machine learning · 33 · 4 first-author · 22 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-authorComputer networks · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Policy Extraction-Based Adversarial Attack in Multi-Agent Reinforcement Learning
Bang Zhang, Wenjian Luo, Kesheng Chen, Yujiang Liu, Shuhan Qi, Xuan Wang 0002 |
ICIC (2) | 1 |
| 2026 | MaTe3D: Mask-Guided Text-based 3D-Aware Portrait Editing
Kangneng Zhou, Daiheng Gao, Xuan Wang 0009, Jie Zhang 0090, Peng Zhang 0080, Xusen Sun, Longhao Zhang, Shiqi Yang 0002, Bang Zhang, Liefeng Bo, Yaxing Wang, Ming-Ming Cheng |
Int. J. Comput. Vis. | 9 |
| 2025 | VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid PriorabstractAudio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, facial motion, head pose generation, and video quality. However, no model has yet led or tied on all these metrics due to the one-to-many mapping between audio and motion. In this paper, we propose VividTalk, a two-stage generic framework that supports generating high-visual quality talking head videos with all the above properties. Specifically, in the first stage, we map the audio to mesh by learning two motions, including non-rigid facial motion and rigid head motion. For facial motion, both blendshape and vertex are adopted as the intermediate representation to maximize the representation ability of the model. For head motion, a novel learnable head pose codebook with a two-phase training mechanism is proposed. In the second stage, we proposed a dual branch motion-vae and a generator to transform the meshes into dense motion and synthesize high-quality video frame-by-frame. Extensive experiments show that the proposed VividTalk can generate high-visual quality talking head videos with lip-sync and realistic enhanced by a large margin, and outperforms previous state-of-the-art works in objective and subjective comparisons. The code will be publicly released upon publication. Xusen Sun, Longhao Zhang, Hao Zhu 0004, Peng Zhang 0080, Bang Zhang, Xinya Ji, Kangneng Zhou, Daiheng Gao, Liefeng Bo, Xun Cao |
3DV | 5 |
| 2025 | S²R: Teaching LLMs to Self-verify and Self-correct via Reinforcement LearningabstractRecent studies have demonstrated the effectiveness of LLM test-time scaling. However, existing approaches to incentivize LLMs’ deep thinking abilities generally require large-scale data or significant training efforts. Meanwhile, it remains unclear how to improve the thinking abilities of less powerful base models. In this work, we introduce S^2R, an efficient framework that enhances LLM reasoning by teaching models to self-verify and self-correct during inference. Specifically, we first initialize LLMs with iterative self-verification and self-correction behaviors through supervised fine-tuning on carefully curated data. The self-verification and self-correction skills are then further strengthened by outcome-level and process-level reinforcement learning with minimized resource requirements. Our results demonstrate that, with only 3.1k behavior initialization samples, Qwen2.5-math-7B achieves an accuracy improvement from 51.0% to 81.6%, outperforming models trained on an equivalent amount of long-CoT distilled data. We also discuss the effect of different RL strategies on enhancing LLMs’ deep reasoning. Extensive experiments and analysis based on three base models across both in-domain and out-of-domain benchmarks validate the effectiveness of S^2R. Ruotian Ma, Peisong Wang 0002, Xingyan Liu, Bang Zhang, Jia Li 0009 |
ACL (1) | 6 |
| 2025 | Exploring Timeline Control for Facial Motion GenerationabstractThis paper introduces a new control signal for facial motion generation: timeline control. Compared to audio and text signals, timelines provide more fine-grained control, such as generating specific facial motions with precise timing. Users can specify a multi-track timeline of facial actions arranged in temporal intervals, allowing precise control over the timing of each action. To model the timeline control capability, We first annotate the time intervals of facial actions in natural facial motion sequences at a frame-level granularity. This process is facilitated by Toeplitz Inverse Covariance-based Clustering to minimize human labor. Based on the annotations, we propose a diffusion-based generation model capable of generating facial motions that are natural and accurately aligned with input timelines. Our method supports text-guided motion generation by using ChatGPT to convert text into timelines. Experimental results show that our method can annotate facial action intervals with satisfactory accuracy, and produces natural facial motions accurately aligned with timelines. Yifeng Ma 0001, Jinwei Qi, Chaonan Ji, Peng Zhang 0080, Bang Zhang, Zhidong Deng, Liefeng Bo |
CVPR | 5 |
| 2025 | Animate Anyone 2: High-Fidelity Character Image Animation with Environment AffordanceabstractRecent character image animation methods based on diffusion models, such as Animate Anyone, have made significant progress in generating consistent and generalizable character animations. However, these approaches fail to produce reasonable associations between characters and their environments. To address this limitation, we introduce Animate Anyone 2, aiming to animate characters with environment affordance. Beyond extracting motion signals from source video, we additionally capture environmental representations as conditional inputs. The environment is formulated as the region with the exclusion of characters and our model generates characters to populate these regions while maintaining coherence with the environmental context. We propose a shape-agnostic mask strategy that more effectively characterizes the relationship between character and environment. Furthermore, to enhance the fidelity of object interactions, we leverage an object guider to extract features of interacting objects and employ spatial blending for feature injection. We also introduce a pose modulation strategy that enables the model to handle more diverse motion patterns. Experimental results demonstrate the superior performance of the proposed method. Guangyuan Wang, Dechao Meng, Lian Zhuo, Peng Zhang 0080, Bang Zhang, Liefeng Bo |
ICCV | 8 |
| 2025 | Controllable and Expressive One-Shot Video Head SwappingabstractIn this paper, we propose a novel diffusion-based multi-condition controllable framework for video head swapping, which seamlessly transplant a human head from a static image into a dynamic video, while preserving the original body and background of target video, and further allowing to tweak head expressions and movements during swapping as needed. Existing face-swapping methods mainly focus on localized facial replacement neglecting holistic head morphology, while head-swapping approaches struggling with hairstyle diversity and complex backgrounds, and none of these methods allow users to modify the transplanted head expressions after swapping. To tackle these challenges, our method incorporates several innovative strategies through a unified latent diffusion paradigm. 1) Identity-preserving context fusion: We propose a shape-agnostic mask strategy to explicitly disentangle foreground head identity features from background/body contexts, combining hair enhancement strategy to achieve robust holistic head identity preservation across diverse hair types and complex backgrounds. 2) Expression-aware landmark retargeting and editing: We propose a disentangled 3DMM-driven retargeting module that decouples identity, expression, and head poses, minimizing the impact of original expressions in input images and supporting expression editing. While a scale-aware retargeting strategy is further employed to minimize cross-identity expression distortion for higher transfer precision. Experimental results demonstrate that our method excels in seamless background integration while preserving the identity of the source portrait, as well as showcasing superior expression transfer capabilities applicable to both real and virtual characters. Chaonan Ji, Jinwei Qi, Bang Zhang, Liefeng Bo |
ICCV | 4 |
| 2025 | MIR: Efficient Exploration in Episodic Multi-agent Reinforcement Learning via Mutual Intrinsic Reward
Kesheng Chen, Wenjian Luo, Bang Zhang, Zeping Yin, Zipeng Ye |
ICIC (14) | 3 |
| 2025 | Beyond Sliders: Mastering the Art of Diffusion-based Image ManipulationabstractIn the realm of image generation, the quest for realism and customization has never been more pressing. While existing methods like concept sliders have made strides, they often falter when it comes to non-AIGC images, particularly images captured in real-world settings. To bridge this gap, we introduce Beyond Sliders, an innovative framework that integrates GANs and diffusion models to facilitate sophisticated image manipulation across diverse image categories. Improved upon concept sliders, our method refines the image through fine-grained guidance—both textual and visual—in an adversarial manner, leading to a marked enhancement in image quality and realism. Extensive experimental validation confirms the robustness and versatility of Beyond Sliders across a spectrum of applications. Yufei Tang, Daiheng Gao, Pingyu Wu, Wenbo Zhou 0004, Bang Zhang, Weiming Zhang 0001 |
ICME | 5 |
| 2025 | ExtPose: Robust and Coherent Pose Estimation by Extending ViTsabstractVision Transformers (ViT) are remarkable at 3D pose estimation, yet they still encounter certain challenges. One issue is that the popular ViT architecture for pose estimation is limited to images and lacks temporal information. Another challenge is that the prediction often fails to maintain pixel alignment with the original images. To address these issues, we propose a systematic framework for 3D pose estimation, called ExtPose. ExtPose extends image ViT to the challenging scenario and video setting by taking in additional 2D pose evidence and capturing temporal information in a full attention-based manner. We use 2D human skeleton images to integrate structured 2D pose information. By sharing parameters and attending across modalities and frames, we enhance the consistency between 3D poses and 2D videos without introducing additional parameters. We achieve state-of-the-art (SOTA) performance on multiple human and hand pose estimation benchmarks with substantial improvements to 34.0mm (-23%) on 3DPW and 4.9mm (-18%) on FreiHAND in PA-MPJPE over the other ViT-based methods respectively. Rongyu Chen, Lian Zhuo, Linlin Yang 0001, Qi Wang 0148, Liefeng Bo, Bang Zhang, Angela Yao |
ICML | 6 |
| 2025 | EraseAnything: Enabling Concept Erasure in Rectified Flow TransformersabstractRemoving unwanted concepts from large-scale text-to-image (T2I) diffusion models while maintaining their overall generative quality remains an open challenge. This difficulty is especially pronounced in emerging paradigms, such as Stable Diffusion (SD) v3 and Flux, which incorporate flow matching and transformer-based architectures. These advancements limit the transferability of existing concept-erasure techniques that were originally designed for the previous T2I paradigm (e.g., SD v1.4). In this work, we introduce EraseAnything, the first method specifically developed to address concept erasure within the latest flow-based T2I framework. We formulate concept erasure as a bi-level optimization problem, employing LoRA-based parameter tuning and an attention map regularizer to selectively suppress undesirable activations. Furthermore, we propose a self-contrastive learning strategy to ensure that removing unwanted concepts does not inadvertently harm performance on unrelated ones. Experimental results demonstrate that EraseAnything successfully fills the research gap left by earlier methods in this new T2I paradigm, achieving state-of-the-art performance across a wide range of concept erasure tasks. Daiheng Gao, Shilin Lu, Wenbo Zhou 0004, Jiaming Chu, Jie Zhang 0073, Mengxi Jia, Bang Zhang, Zhaoxin Fan, Weiming Zhang 0001 |
ICML | 7 |
| 2025 | Black-Box Adversarial Robustness Testing with Partial Observation for Multi-Agent Reinforcement LearningabstractMulti-Agent Reinforcement Learning (MARL) has shown great success in many aspects. However, the cooperative policy trained by MARL is vulnerable to adversarial attacks towards agents' observations, which could cause immeasurable damage to the agent team. A few techniques have been developed to test the robustness of MARL, but the feasibility of real implementation is not carefully considered. In this work, we propose a two-step framework to conduct destructive and sparse attacks under realistic conditions. The first step contains two attack scenes: Ally Observation Attack (AOA) and Enemy Observation Attack (EOA), which have limitations on the use of agents' observations. The first step selects victim agent from the team and screens the attack time, while the second step generates adversarial perturbation and adds it to the observation of the chosen victim in a black-box environment. To the best of our knowledge, this is the first work to test the robustness with partial observation and conduct the test in a black-box environment. Experiments on SMAC environments demonstrate that our methods show great performance in reducing the win rate and the team reward of the agent team trained by QMIX algorithm with lower perturbation steps. Bang Zhang, Wenjian Luo, Kesheng Chen, Yujiang Liu, Shuhan Qi, Xuan Wang 0002 |
ICPADS | 1 |
| 2025 | SPC: Evolving Self-Play Critic via Adversarial Games for LLM ReasoningabstractEvaluating the step-by-step reliability of large language model (LLM) reasoning, such as Chain-of-Thought, remains challenging due to the difficulty and cost of obtaining high-quality step-level supervision. In this paper, we introduce Self-Play Critic (SPC), a novel approach where a critic model evolves its ability to assess reasoning steps through adversarial self-play games, eliminating the need for manual step-level annotation. SPC involves fine-tuning two copies of a base model to play two roles, namely a "sneaky generator" that deliberately produces erroneous steps designed to be difficult to detect, and a "critic" that analyzes the correctness of reasoning steps. These two models engage in an adversarial game in which the generator aims to fool the critic, while the critic model seeks to identify the generator's errors. Using reinforcement learning based on the game outcomes, the models iteratively improve; the winner of each confrontation receives a positive reward and the loser receives a negative reward, driving continuous self-evolution. Experiments on three reasoning process benchmarks (ProcessBench, PRM800K, DeltaBench) demonstrate that our SPC progressively enhances its error detection capabilities (e.g., accuracy increases from 70.8% to 77.7% on ProcessBench) and surpasses strong baselines, including distilled R1 model. Furthermore, SPC can guide the test-time search of diverse LLMs and significantly improve their mathematical reasoning performance on MATH500 and AIME2024, surpassing those guided by state-of-the-art process reward models. Bang Zhang, Ruotian Ma, Peisong Wang 0002, Xiaodan Liang, Zhaopeng Tu, Kwan-Yee Kenneth Wong |
NeurIPS | 2 |
| 2025 | OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style MimickingabstractAlthough significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video content from input text while emulating the target identity's speaking and facial movement styles, including speech characteristics, head motion, and facial dynamics. Our framework adopts a dual-branch diffusion transformer (DiT) architecture, with one branch dedicated to audio generation and the other to video synthesis.
At the shallow layers, cross-modal fusion modules are introduced to integrate information between the two modalities. In deeper layers, each modality is processed independently, with the generated audio decoded by a vocoder and the video rendered using a GAN-based high-quality visual renderer. Leveraging DiT’s in-context learning capability through a masked-infilling strategy, our model can simultaneously capture both audio and visual styles without requiring explicit style extraction modules. Thanks to the efficiency of the DiT backbone and the optimized visual renderer, OmniTalker achieves real-time inference at 25 FPS.
To the best of our knowledge, OmniTalker is the first one-shot framework capable of jointly modeling speech and facial styles in real time. Extensive experiments demonstrate its superiority over existing methods in terms of generation quality, particularly in preserving style consistency and ensuring precise audio-video synchronization, all while maintaining efficient inference. Zhongjian Wang, Peng Zhang 0080, Jinwei Qi, Sheng Xu 0007, Bang Zhang |
NeurIPS | 6 |
| 2024 | Cloth2Tex: A Customized Cloth Texture Generation Pipeline for 3D Virtual Try-OnabstractFabricating and designing 3D garments has become extremely demanding with the increasing need for synthesizing realistic dressed persons for a variety of applications, e.g. 3D virtual try-on, digitalization of 2D clothes into 3D apparel, and cloth animation. It thus necessitates a simple and straightforward pipeline to obtain high-quality texture from simple input, such as 2D reference images. Since traditional warping-based texture generation methods require a significant number of control points to be manually selected for each type of garment, which can be a time-consuming and tedious process. We propose a novel method, called Cloth2Tex, which eliminates the human burden in this process. Cloth2Tex is a self-supervised method that generates texture maps with reasonable layout and structural consistency. Another key feature of Cloth2Tex is that it can be used to support high-fidelity texture inpainting. This is done by combining Cloth2Tex with a prevailing latent diffusion model. We evaluate our approach both qualitatively and quantitatively and demonstrate that Cloth2Tex can generate high-quality texture maps and achieve the best visual effects in comparison to other methods. For more details and animated results, please see https://tomguluson92. github.io/projects/cloth2tex/. Daiheng Gao, Xindi Zhang 0003, Qi Wang 0148, Bang Zhang, Liefeng Bo, Qixing Huang |
3DV | 6 |
| 2024 | Cross-Sentence Gloss Consistency for Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) aims to recognize gloss sequences from continuous sign videos. Recent works enhance the gloss representation consistency by mining correlations between visual and contextual modules within individual sentences. However, there still remain much richer correlations among glosses across different sentences. In this paper, we present a simple yet effective Cross-Sentence Gloss Consistency (CSGC), which enforces glosses belonging to a same category to be more consistent in representation than those belonging to different categories, across all training sentences. Specifically, in CSGC, a prototype is maintained for each gloss category and benefits the gloss discrimination in a contrastive way. Thanks to the well-distinguished gloss prototype, an auxiliary similarity classifier is devised to enhance the recognition clues, thus yielding more accurate results. Extensive experiments conducted on three CSLR datasets show that our proposed CSGC significantly boosts the performance of CSLR, surpassing existing state-of-the-art works by large margins (i.e., 1.6% on PHOENIX14, 2.4% on PHOENIX14-T, and 5.7% on CSL-Daily). Qi Rao, Bang Zhang |
AAAI | 5 |
| 2024 | EMO: Emote Portrait Alive Generating Expressive Portrait Videos with Audio2Video Diffusion Model Under Weak Conditions
Linrui Tian, Bang Zhang, Liefeng Bo |
ECCV (83) | 3 |
| 2024 | SKIM: Skeleton-Based Isolated Sign Language Recognition With Part MixingabstractIn this article, we present skeleton-based isolated sign language recognition (IsoSLR) with part mixing - SKIM. An IsoSLR model that solely takes the skeleton representation of the human body as input. Previous skeleton-based works either perform worse when compared to RGB-based counterparts or require fusion with other modalities to obtain competitive results. With SKIM, a single skeleton-based model without complex pre-training can obtain similar or even higher accuracy than current state-of-the-art methods. This margin can be further increased by simple late fusion within the same modality. To achieve this, we first develop a novel data augmentation technique called part mixing. It swaps the corresponding keypoints within one region (e.g. hand) between two randomly selected samples and combines their labels linearly as the new label. As regions like hand and face are key articulators for sign language, direct swapping of such parts creates a believable pseudo sign that promotes the model to recognize the true pairs. Secondly, following current advances in skeleton-based action recognition, we devise a channel-wise graph neural network with multi-scale awareness and per-keypoint temporal re-weighting. With this design, the backbone is capable of leveraging both manual and non-manual features. The combination of hand mixing and the channel-wise multi-scale GCN backbone allows us to achieve state-of-the-art accuracy on both WLASL and NMFs-CSL benchmarks. Kezhou Lin, Linchao Zhu, Bang Zhang, Yi Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Jointly Harnessing Prior Structures and Temporal Consistency for Sign Language Video GenerationabstractSign language provides a way for differently-abled individuals to express their feelings and emotions. However, learning sign language can be challenging and time consuming. An alternative approach is to animate user photos using sign language videos of specific words, which can be achieved using existing image animation methods. However, the finger motions in the generated videos are often not ideal. To address this issue, we propose the Structure-aware Temporal Consistency Network (STCNet), which jointly optimizes the prior structure of humans with temporal consistency to produce sign language videos. We use a fine-grained skeleton detector to acquire knowledge of body structure and introduce both short- and long-term cycle loss to ensure the continuity of the generated video. The two losses and keypoint detector network are optimized in an end-to-end manner. Quantitative and qualitative evaluations on three widely used datasets, namely LSA64, Phoenix-2014T, and WLASL-2000, demonstrate the effectiveness of the proposed method. It is our hope that this work can contribute to future studies on sign language production. Yucheng Suo, Zhedong Zheng, Bang Zhang, Yi Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Gloss-Free End-to-End Sign Language TranslationabstractIn this paper, we tackle the problem of sign language translation (SLT) without gloss annotations.Although intermediate representation like gloss has been proven effective, gloss annotations are hard to acquire, especially in large quantities.This limits the domain coverage of translation datasets, thus handicapping real-world applications.To mitigate this problem, we design the Gloss-Free End-to-end sign language translation framework (GloFE).Our method improves the performance of SLT in the gloss-free setting by exploiting the shared underlying semantics of signs and the corresponding spoken translation.Common concepts are extracted from the text and used as a weak form of intermediate representation.The global embedding of these concepts is used as a query for cross-attention to find the corresponding information within the learned visual features.In a contrastive manner, we encourage the similarity of query results between samples containing such concepts and decrease those that do not.We obtained state-of-the-art results on large-scale datasets, including OpenASL and How2Sign. 1 Kezhou Lin, Linchao Zhu, Bang Zhang, Yi Yang 0001 |
ACL (1) | 5 |
| 2023 | One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance FieldabstractTalking head generation aims to generate faces that maintain the identity information of the source image and imitate the motion of the driving image. Most pioneering methods rely primarily on 2D representations and thus will inevitably suffer from face distortion when large head rotations are encountered. Recent works instead employ explicit 3D structural representations or implicit neural rendering to improve performance under large pose changes. Nevertheless, the fidelity of identity and expression is not so desirable, especially for novel-view synthesis. In this paper, we propose HiDe-NeRF, which achieves high-fidelity and free-view talking-head synthesis. Drawing on the recently proposed Deformable Neural Radiance Fields, HiDe-NeRF represents the 3D dynamic scene into a canonical appearance field and an implicit deformation field, where the former comprises the canonical source face and the latter models the driving pose and expression. In particular, we improve fidelity from two aspects: (i) to enhance identity expressiveness, we design a generalized appearance module that leverages multi-scale volume features to preserve face shape and details; (ii) to improve expression preciseness, we propose a lightweight deformation module that explicitly decouples the pose and expression to enable precise expression modeling. Extensive experiments demonstrate that our proposed approach can generate better results than previous works. Project page: https://www.waytron.net/hidenerf/ Weichuang Li, Longhao Zhang, Dong Wang 0028, Bin Zhao 0001, Zhigang Wang 0002, Mulin Chen, Bang Zhang, Zhongjian Wang, Liefeng Bo, Xuelong Li 0001 |
CVPR | 7 |
| 2023 | Towards Stable Human Pose Estimation via Cross-View Fusion and Foot StabilizationabstractTowards stable human pose estimation from monocular images, there remain two main dilemmas. On the one hand, the different perspectives, i.e., front view, side view, and top view, appear the inconsistent performances due to the depth ambiguity. On the other hand, foot posture plays a significant role in complicated human pose estimation, i.e., dance and sports, and foot-ground interaction, but unfortunately, it is omitted in most general approaches and datasets. In this paper, we first propose the Cross-View Fusion (CVF) module to catch up with better 3D intermediate representation and alleviate the view inconsistency based on the vision transformer encoder. Then the optimization-based method is introduced to reconstruct the foot pose and foot-ground contact for the general multi-view datasets including AIST++ and Human3.6M. Besides, the reversible kinematic topology strategy is innovated to utilize the contact information into the full-body with foot pose regressor. Extensive experiments on the popular benchmarks demonstrate that our method outperforms the state-of-the-art approaches by achieving 40.1mm PA-MPJPE on the 3DPW test set and 43.8mm on the AIST++ test set. Lian Zhuo, Qi Wang 0148, Bang Zhang, Liefeng Bo |
CVPR | 4 |
| 2023 | RenderIH: A Large-scale Synthetic Dataset for 3D Interacting Hand Pose EstimationabstractThe current interacting hand (IH) datasets are relatively simplistic in terms of background and texture, with hand joints being annotated by a machine annotator, which may result in inaccuracies, and the diversity of pose distribution is limited. However, the variability of background, pose distribution, and texture can greatly influence the generalization ability. Therefore, we present a large-scale synthetic dataset –RenderIH– for interacting hands with accurate and diverse pose annotations. The dataset contains 1M photo-realistic images with varied backgrounds, perspectives, and hand textures. To generate natural and diverse interacting poses, we propose a new pose optimization algorithm. Additionally, for better pose estimation accuracy, we introduce a transformer-based pose estimation network, TransHand, to leverage the correlation between interacting hands and verify the effectiveness of RenderIH in improving results. Our dataset is model-agnostic and can improve more accuracy of any hand pose estimation method in comparison to other real or synthetic datasets. Experiments have shown that pretraining on our synthetic data can significantly decrease the error from 6.76mm to 5.79mm, and our Transhand surpasses contemporary methods. Our dataset and code are available at https://github.com/adwardlee/RenderIH. Linrui Tian, Xindi Zhang 0003, Qi Wang 0148, Bang Zhang, Liefeng Bo, Chen Chen 0001 |
ICCV | 5 |
| 2023 | Multi-view Consistent Generative Adversarial Networks for Compositional 3D-Aware Image SynthesisabstractAbstract This paper studies compositional 3D-aware image synthesis for both single-object and multi-object scenes. We observe that two challenges remain in this field: existing approaches (1) lack geometry constraints and thus compromise the multi-view consistency of the single object, and (2) can not scale to multi-object scenes with complex backgrounds. To address these challenges coherently, we propose multi-view consistent generative adversarial networks (MVCGAN) for compositional 3D-aware image synthesis. First, we build the geometry constraints on the single object by leveraging the underlying 3D information. Specifically, we enforce the photometric consistency between pairs of views, encouraging the model to learn the inherent 3D shape. Second, we adapt MVCGAN to multi-object scenarios. In particular, we formulate the multi-object scene generation as a “decompose and compose” process. During training, we adopt the top-down strategy to decompose training images into objects and backgrounds. When rendering, we deploy a reverse bottom-up manner by composing the generated objects and background into the holistic scene. Extensive experiments on both single-object and multi-object datasets show that the proposed method achieves competitive performance for 3D-aware image synthesis. Xuanmeng Zhang, Zhedong Zheng, Daiheng Gao, Bang Zhang, Yi Yang 0001, Tat-Seng Chua |
Int. J. Comput. Vis. | 4 |
| 2022 | Recurrent Dynamic Embedding for Video Object SegmentationabstractSpace-time memory (STM) based video object segmentation (VOS) networks usually keep increasing memory bank every several frames, which shows excellent performance. However, 1) the hardware cannot withstand the ever-increasing memory requirements as the video length increases. 2) Storing lots of information inevitably introduces lots of noise, which is not conducive to reading the most important information from the memory bank. In this paper, we propose a Recurrent Dynamic Embedding (RDE) to build a memory bank of constant size. Specifically, we explicitly generate and update RDE by the proposed Spatio-temporal Aggregation Module (SAM), which exploits the cue of historical information. To avoid error accumulation owing to the recurrent usage of SAM, we propose an unbiased guidance loss during the training stage, which makes SAM more robust in long videos. Moreover, the predicted masks in the memory bank are inaccurate due to the inaccurate network inference, which affects the seg-mentation of the query frame. To address this problem, we design a novel self-correction strategy so that the network can repair the embeddings of masks with different qualities in the memory bank. Extensive experiments show our method achieves the best tradeoff between performance and speed. Code is available at https://github.com/Limingxing00/RDE-VOS-CVPR2022. Mingxing Li 0003, Zhiwei Xiong, Bang Zhang, Dong Liu 0002 |
CVPR | 4 |
| 2022 | Multi-View Consistent Generative Adversarial Networks for 3D-aware Image Synthesisabstract3D-aware image synthesis aims to generate images of objects from multiple views by learning a 3D representation. However, one key challenge remains: existing approaches lack geometry constraints, hence usually fail to generate multi-view consistent images. To address this challenge, we propose Multi-View Consistent Generative Adversarial Networks (MVCGAN) for high-quality 3D-aware image synthesis with geometry constraints. By leveraging the underlying 3D geometry information of generated images, i.e., depth and camera transformation matrix, we explicitly establish stereo correspondence between views to perform multi-view joint optimization. In particular, we enforce the photometric consistency between pairs of views and integrate a stereo mixup mechanism into the training process, encouraging the model to reason about the correct 3D shape. Besides, we design a two-stage training strategy with feature-level multi-view joint optimization to improve the image quality. Extensive experiments on three datasets demonstrate that MVCGAN achieves the state-of-the-art performance for 3D-aware image synthesis. Xuanmeng Zhang, Zhedong Zheng, Daiheng Gao, Bang Zhang, Yi Yang 0001 |
CVPR | 4 |
| 2022 | A Speech-driven Sign Language Avatar Animation System for Hearing Impaired ApplicationsabstractSign language is the communication language used in hearing impaired community. Recently, the research of sign language production has made great progress but still need to cope with some critical challenges. In this paper, we propose a system-level scheme and push forward the implementation of sign language production for practical usage. We build a system capable of translating speech into sign language avatar. Different from previous approach only focusing on single technology, we systematically combine algorithms of language translation, body gesture animation and facial avatar generation. We also develop two applications: Sign Language Interpretation APP and Virtual Sign Language Anchor, to facilitate easy and clear communication for hearing impaired people. Jiahui Li 0009, Bang Zhang |
IJCAI | 5 |
| 2022 | Text/Speech-Driven Full-Body AnimationabstractDue to the increasing demand in films and games, synthesizing 3D avatar animation has attracted much attention recently. In this work, we present a production-ready text/speech-driven full-body animation synthesis system. Given the text and corresponding speech, our system synthesizes face and body animations simultaneously, which are then skinned and rendered to obtain a video stream output. We adopt a learning-based approach for synthesizing facial animation and a graph-based approach to animate the body, which generates high-quality avatar animation efficiently and robustly. Our results demonstrate the generated avatar animations are realistic, diverse and highly text/speech-correlated. Wenlin Zhuang, Jinwei Qi, Peng Zhang 0080, Bang Zhang |
IJCAI | 4 |
| 2022 | CycleHand: Increasing 3D Pose Estimation Ability on In-the-wild Monocular Image through Cyclic FlowabstractCurrent methods for 3D hand pose estimation fail to generalize well to in-the-wild new scenarios due to varying camera viewpoints, self-occlusions, and complex environments. To address this problem, we propose CycleHand to improve the generalization ability of the model in a self-supervised manner. Our motivation is based on an observation: if one globally rotates the whole hand and reversely rotates it back, the estimated 3D poses of fingers should keep consistent before and after the rotation because the wrist-relative hand poses stay unchanged during global 3D rotation. Hence, we propose arbitrary-rotation self-supervised consistency learning to improve the model's robustness for varying viewpoints. Another innovation of CycleHand is that we propose a high-fidelity texture map to render the photorealistic rotated hand with different lighting conditions, backgrounds, and skin tones to further enhance the effectiveness of our self-supervised task. To reduce the potential negative effects brought by the domain shift of synthetic images, we use the idea of contrastive learning to learn a synthetic-real consistent feature extractor in extracting domain-irrelevant hand representations. Experiments show that CycleHand can largely improve the hand pose estimation performance in both canonical datasets and real-world applications. Daiheng Gao, Xindi Zhang 0003, Xingyu Chen 0002, Andong Tan, Bang Zhang, Ping Tan 0002 |
ACM Multimedia | 5 |
| 2022 | DART: Articulated Hand Model with Diverse Accessories and Rich TexturesabstractHand, the bearer of human productivity and intelligence, is receiving much attention due to the recent fever of digital twins. Among different hand morphable models, MANO has been widely used in vision and graphics community. However, MANO disregards textures and accessories, which largely limits its power to synthesize photorealistic hand data. In this paper, we extend MANO with Diverse Accessories and Rich Textures, namely DART. DART is composed of 50 daily 3D accessories which varies in appearance and shape, and 325 hand-crafted 2D texture maps covers different kinds of blemishes or make-ups. Unity GUI is also provided to generate synthetic hand data with user-defined settings, e.g., pose, camera, background, lighting, textures, and accessories. Finally, we release DARTset, which contains large-scale (800K), high-fidelity synthetic hand images, paired with perfect-aligned 3D labels. Experiments demonstrate its superiority in diversity. As a complement to existing hand datasets, DARTset boosts the generalization in both hand pose estimation and mesh recovery tasks. Raw ingredients (textures, accessories), Unity GUI, source code and DARTset are publicly available at dart2022.github.io. Daiheng Gao, Yuliang Xiu, Kailin Li 0001, Lixin Yang 0001, Feng Wang 0072, Peng Zhang 0080, Bang Zhang, Cewu Lu |
NeurIPS | 7 |
| 2021 | Learning Position and Target Consistency for Memory-Based Video Object SegmentationabstractThis paper studies the problem of semi-supervised video object segmentation(VOS). Multiple works have shown that memory-based approaches can be effective for video object segmentation. They are mostly based on pixel-level matching, both spatially and temporally. The main shortcoming of memory-based approaches is that they do not take into account the sequential order among frames and do not exploit object-level knowledge from the target. To address this limitation, we propose to Learn position and target Consistency framework for Memory-based video object segmentation, termed as LCM. It applies the memory mechanism to retrieve pixels globally, and meanwhile learns position consistency for more reliable segmentation. The learned location response promotes a better discrimination between target and distractors. Besides, LCM introduces an object-level relationship from the target to maintain target consistency, making LCM more robust to error drifting. Experiments show that our LCM achieves state-of-the-art performance on both DAVIS and Youtube-VOS benchmark. And we rank the 1st in the DAVIS 2020 challenge semi-supervised VOS task. Peng Zhang 0080, Bang Zhang, Rong Jin 0001 |
CVPR | 3 |
| 2021 | Text-driven 3D Avatar Animation with Emotional and Expressive BehaviorsabstractText-driven 3D avatar animation has been an essential part of virtual human techniques, which has a wide range of applications in movie, digital games and video streaming. In this work, we introduce a practical system which drives both facial and body movements of 3D avatar by text input. Our proposed system first converts text input to speech signal and conducts text analysis to extract semantic tags simultaneously. Then we generate the lip movements from the synthetic speech, and meanwhile facial expression and body movement are generated by the joint modeling of speech and textual information, which can drive our virtual 3D avatar talking and acting like a real human. Jinwei Qi, Bang Zhang |
ACM Multimedia | 3 |
| 2021 | A Virtual Character Generation and Animation System for E-Commerce Live StreamingabstractVirtual character has been widely adopted in many areas, such as virtual assistant, virtual customer service, robotics and etc. In this paper, we focus on its application in e-commerce live streaming. Particularly, we propose a virtual character generation and animation system that supports e-commerce live streaming with virtual characters as anchors. The system offers a virtual character face generation tool based on a weakly supervised 3D face reconstruction method. The method takes a single photo as input and generates a 3D face model with both similarity and aesthetics considered. It does not require 3D face annotation data due to the assist of differentiable neural rendering technique which seamlessly integrates rendering into a deep learning based 3D face reconstruction framework. Moreover, the system provides two animation approaches which support two different ways of live stream respectively. The first approach is based on real-time motion capture. An actor's performance is captured in real-time via a monocular camera, and then utilized for animating a virtual anchor. The second approach is text driven animation, in which the human-like animation is automatically generated based on a text script. The relationship between text script and animation is learned based on the training data which can be accumulated via the motion capture based animation. To our best knowledge, the presented work is the first sophisticated virtual character generation and animation system that is designed for e-commerce live streaming and actually deployed on an online shopping platform with millions of daily audiences. Bang Zhang, Peng Zhang 0080, Jinwei Qi, Daiheng Gao, Haiming Zhao, Xiaoduan Feng, Qi Wang 0148, Lian Zhuo |
ACM Multimedia | 2 |
| 2019 | Finding Temporal Influential Users Over Evolving Social NetworksabstractInfluence maximization (IM) continues to be a key research problem in social networks. The goal is to find a small seed set of target users that have the greatest influence in the network under various stochastic diffusion models. While significant progress has been made on the IM problem in recent years, several interesting challenges remain. For example, social networks in reality are constantly evolving, and "important" users with the most influence also change over time. As a result, several recent studies have proposed approaches to update the seed set as the social networks evolve. However, this seed set is not guaranteed to be the best seed set over a period of time. In this paper we study the problem of Distinct Influence Maximization (DIM) where the goal is to identify a seed set of influencers who maximize the number of distinct users influenced over a predefined window of time. Our new approach allows social network providers to make fewer incremental changes to targeted advertising while still maximizing the coverage of the advertisements. It also provides finer grained control over service level agreements where a certain number of impressions for an advertisement must be displayed in a specific time period. We propose two different strategies HCS and VCS with novel graph compression techniques to solve this problem. Additionally, VCS can also be applied directly to the traditional IM problem. Extensive experiments on real-world datasets verify the efficiency, accuracy and scalability of our solutions on both the DIM and IM problems. Shixun Huang, Zhifeng Bao, J. Shane Culpepper, Bang Zhang |
ICDE | 4 |
| 2019 | Recovering DTW Distance Between Noise Superposed NHPP
Yongzhe Chang, Zhidong Li, Bang Zhang, Ling Luo 0002, Arcot Sowmya, Yang Wang 0002, Fang Chen 0001 |
PAKDD (2) | 3 |
| 2018 | MaxBRkNN Queries for Streaming Geo-Data
Hui Luo 0001, Farhana Murtaza Choudhury, Zhifeng Bao, J. Shane Culpepper, Bang Zhang |
DASFAA (1) | 5 |
| 2018 | Simultaneous Urban Region Function Discovery and Popularity Estimation via an Infinite Urbanization Process ModelabstractUrbanization is a global trend that we have all witnessed in the past decades. It brings us both opportunities and challenges. On the one hand, urban system is one of the most sophisticated social-economic systems that is responsible for efficiently providing supplies meeting the demand of residents in various of domains, e.g., dwelling, education, entertainment, healthcare, etc. On the other hand, significant diversity and inequality exist in the development patterns of urban systems, which makes urban data analysis difficult. Different urban regions often exhibit diverse urbanization patterns and provide distinct urban functions, e.g., commercial and residential areas offer significantly different urban functions. It is desired to develop the data analytic capabilities for discovering the underlying cross-domain urbanization patterns, clustering urban regions based on their function similarity and predicting region popularity in specified domains. Previous studies in the urban data analysis area often just focus on individual domains and rarely consider cross-domain urban development patterns hidden in different urban regions. In this paper, we propose the infinite urbanization process (IUP) model for simultaneous urban region function discovery and region popularity prediction. The IUP model is a generative Bayesian nonparametric process that is capable of describing a potentially infinite number of urbanization patterns. It is developed within the supervised topic modelling framework and is supported by a novel hierarchical spatial distance dependent Bayesian nonparametric prior over the spatial region partition space. The empirical study conducted on the real-world datasets shows promising outcome compared with the state-of-the-art techniques. Bang Zhang, Lelin Zhang, Ting Guo 0005, Yang Wang 0002, Fang Chen 0001 |
KDD | 1 |
| 2018 | A Linear-Time Algorithm for Finding Induced Planar SubgraphsabstractIn this paper we study the problem of efficiently and effectively extracting induced planar subgraphs. Edwards and Farr proposed an algorithm with O(mn) time complexity to find an induced planar subgraph of at least 3n/(d+1) vertices in a graph of maximum degree d. They also proposed an alternative algorithm with O(mn) time complexity to find an induced planar subgraph graph of at least 3n/(bar{d}+1) vertices, where bar{d} is the average degree of the graph. These two methods appear to be best known when d and bar{d} are small. Unfortunately, they sacrifice accuracy for lower time complexity by using indirect indicators of planarity. A limitation of those approaches is that the algorithms do not implicitly test for planarity, and the additional costs of this test can be significant in large graphs. In contrast, we propose a linear-time algorithm that finds an induced planar subgraph of n-nu vertices in a graph of n vertices, where nu denotes the total number of vertices shared by the detected Kuratowski subdivisions. An added benefit of our approach is that we are able to detect when a graph is planar, and terminate the reduction. The resulting planar subgraphs also do not have any rigid constraints on the maximum degree of the induced subgraph. The experiment results show that our method achieves better performance than current methods on graphs with small skewness. Shixun Huang, Zhifeng Bao, J. Shane Culpepper, Bang Zhang |
SEA | 5 |
| 2018 | Multi-objective optimization controller placement problem in internet-oriented software defined network
Bang Zhang, Xingwei Wang 0001, Min Huang 0001 |
Comput. Commun. | 1 |
| 2016 | Interaction Point Processes via Infinite Branching ModelabstractMany natural and social phenomena can be modeled by interaction point processes (IPPs) (Diggle et al. 1994), stochastic point processes considering the interaction between points. In this paper, we propose the infinite branching model (IBM), a Bayesian statistical model that can generalize and extend some popular IPPs, e.g., Hawkes process (Hawkes 1971; Hawkes and Oakes 1974). It treats IPP as a mixture of basis point processes with the aid of a distance dependent prior over branching structure that describes the relationship between points. The IBM can estimate point event intensity, interaction mechanism and branching structure simultaneously. A generic Metropolis-within-Gibbs sampling method is also developed for model parameter inference. The experiments on synthetic and real-world data demonstrate the superiority of the IBM. Bang Zhang, Ting Guo 0005, Yang Wang 0002, Fang Chen 0001 |
AAAI | 2 |
| 2016 | Infinite Hidden Semi-Markov Modulated Interaction Point ProcessabstractThe correlation between events is ubiquitous and important for temporal events modelling. In many cases, the correlation exists between not only events' emitted observations, but also their arrival times. State space models (e.g., hidden Markov model) and stochastic interaction point process models (e.g., Hawkes process) have been studied extensively yet separately for the two types of correlations in the past. In this paper, we propose a Bayesian nonparametric approach that considers both types of correlations via unifying and generalizing hidden semi-Markov model and interaction point process model. The proposed approach can simultaneously model both the observations and arrival times of temporal events, and determine the number of latent states from data. A Metropolis-within-particle-Gibbs sampler with ancestor resampling is developed for efficient posterior inference. The approach is tested on both synthetic and real-world data with promising outcomes. Bang Zhang, Ting Guo 0005, Yang Wang 0002, Fang Chen 0001 |
NIPS | 2 |
| 2016 | Robust Bayesian non-parametric dictionary learning with heterogeneous Gaussian noise
Yi Wang 0041, Bin Li 0015, Yang Wang 0002, Fang Chen 0001, Bang Zhang, Zhidong Li |
Comput. Vis. Image Underst. | 5 |
| 2015 | Data Driven Water Pipe Failure Prediction: A Bayesian Nonparametric ApproachabstractWater pipe failures can cause significant economic and social costs, hence have become the primary challenge to water utilities. In this paper, we propose a Bayesian nonparametric approach, namely the Dirichlet process mixture of hierarchical beta process model, for water pipe failure prediction. It can select high-risk pipes for physical condition assessment, thereby preventing disastrous failures proactively. Bang Zhang, Yi Wang 0041, Zhidong Li, Bin Li 0015, Yang Wang 0002, Fang Chen 0001 |
CIKM | 2 |
| 2015 | On Damage Identification in Civil Structures Using Tensor Analysis
Khoa L. D. Nguyen, Bang Zhang, Yang Wang 0002, Wei Liu 0007, Fang Chen 0001, Samir Mustapha, Peter Runcie |
PAKDD (1) | 2 |
| 2014 | Stable Learning in Coding Space for Multi-class Decoding and Its Extension for Multi-class Hypothesis Transfer LearningabstractMany prevalent multi-class classification approaches can be unified and generalized by the output coding framework which usually consists of three phases: (1) coding, (2) learning binary classifiers, and (3) decoding. Most of these approaches focus on the first two phases and predefined distance function is used for decoding. In this paper, however, we propose to perform learning in coding space for more adaptive decoding, thereby improving overall performance. Ramp loss is exploited for measuring multi-class decoding error. The proposed algorithm has uniform stability. It is insensitive to data noises and scalable with large scale datasets. Generalization error bound and numerical results are given with promising outcomes. Bang Zhang, Yi Wang 0041, Yang Wang 0002, Fang Chen 0001 |
CVPR | 1 |
| 2014 | Water pipe condition assessment: a hierarchical beta process approach for sparse incident data
Zhidong Li, Bang Zhang, Yang Wang 0002, Fang Chen 0001, Ronnie Taib, Vicky Whiffin, Yi Wang 0041 |
Mach. Learn. | 2 |
| 2014 | Multilabel Image Classification Via High-Order Label Correlation Driven Active LearningabstractSupervised machine learning techniques have been applied to multilabel image classification problems with tremendous success. Despite disparate learning mechanisms, their performances heavily rely on the quality of training images. However, the acquisition of training images requires significant efforts from human annotators. This hinders the applications of supervised learning techniques to large scale problems. In this paper, we propose a high-order label correlation driven active learning (HoAL) approach that allows the iterative learning algorithm itself to select the informative example-label pairs from which it learns so as to learn an accurate classifier with less annotation efforts. Four crucial issues are considered by the proposed HoAL: 1) unlike binary cases, the selection granularity for multilabel active learning need to be fined from example to example-label pair; 2) different labels are seldom independent, and label correlations provide critical information for efficient learning; 3) in addition to pair-wise label correlations, high-order label correlations are also informative for multilabel active learning; and 4) since the number of label combinations increases exponentially with respect to the number of labels, an efficient mining method is required to discover informative label correlations. The proposed approach is tested on public data sets, and the empirical results demonstrate its effectiveness. Bang Zhang, Yang Wang 0002, Fang Chen 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | Mutual information-based method for selecting informative feature sets
Gunawan Herman, Bang Zhang, Yang Wang 0002, Getian Ye, Fang Chen 0001 |
Pattern Recognit. | 2 |
| 2012 | Batch mode active learning for multi-label image classification with informative label correlation miningabstractThe performances of supervised learning techniques on image classification problems heavily rely on the quality of their training images. But the acquisition of high quality training images requires significant efforts from human annotators. In this paper, we propose a novel multi-label batch model active learning (MLBAL) approach that allows the learning algorithm to actively select a batch of informative example-label pairs from which it learns at each learning iteration, so as to learn accurate classifiers with less annotation efforts. Unlike existing methods, the proposed approach fines the active selection granularity from example to example-label pair, and takes into account the informative label correlations for active learning. And the empirical studies demonstrate its effectiveness. Bang Zhang, Yang Wang 0002, Wei Wang 0011 |
WACV | 1 |
| 2012 | Multiple-Instance learning from multiple perspectives: Combining models for Multiple-Instance learningabstractMultiple-Instance learning (MIL), which relaxes training annotation granularity from instance level to instance collection (bag) level by applying bag concept, obtains increasing attentions from computer vision community. Due to its flexible annotation mechanism, MIL has been naturally utilized on a variety of computer vision problems. And numerous models have been proposed, each of which is ingeniously designed to catch certain characteristics of MIL. However different models only perform well on certain tasks, and further improvement can hardly be achieved. In this paper, we propose a framework that combines multiple complementary models for solving MIL. Multiple-kernel learning as well as boosting based ensemble learning are utilized to achieve optimal combination. Moreover, the framework is extended to integrate active learning, so as to further reduce the annotation costs on acquiring an accurate image classifier. Experimental studies demonstrate the effectiveness of the proposed methods. Bang Zhang, Yang Wang 0002, Wei Wang 0011 |
WACV | 1 |
| 2011 | Feature fusion for vehicle detection and tracking with low-angle camerasabstractIn this paper, we address the problem of vehicle detection and tracking with low-angle cameras by combining windshield detection and feature points clustering, effectively fusing several primitive image features such as color, edge and interest point. By exploring various heterogenous features and multiple vehicle models, we achieve at least two improvements over the existing methods: higher detection accuracy and the ability to distinguish different vehicle types. Our experiments on real-world traffic video sequences demonstrate the benefits of feature fusion and the improved performance. Jun Yang 0033, Yang Wang 0002, Arcot Sowmya, Zhidong Li, Bang Zhang, Jie Xu 0008 |
WACV | 5 |
| 2010 | Spatial-Temporal Affinity Propagation for Feature Clustering with Application to Traffic Video Analysis
Jun Yang 0033, Yang Wang 0002, Arcot Sowmya, Jie Xu 0008, Zhidong Li, Bang Zhang |
ACCV (2) | 6 |
| 2010 | Affinity Propagation Feature Clustering with Application to Vehicle Detection and Tracking in Road Traffic SurveillanceabstractIn this paper, we investigate the applicability of the newly proposed data clustering method, affinity propagation, in feature points clustering and the task of vehicle detection and tracking in road traffic surveillance. We propose a model-based temporal association scheme and novel preprocessing and postprocessing operations which together with affinity propagation make a quite successful method for the given task. Our experiments demonstrate the effectiveness and efficiency of our method and its superiority over the state-of-the-art algorithm. Jun Yang 0033, Yang Wang 0002, Arcot Sowmya, Bang Zhang, Jie Xu 0008, Zhidong Li |
AVSS | 4 |
| 2010 | Multi-class Graph Boosting with Subgraph Sharing for Object RecognitionabstractIn this paper, we propose a novel multi-class graph boosting algorithm to recognize different visual objects. The proposed method treats subgraph as feature to construct base classifier, and utilizes popular error correcting output code scheme to solve multi-class problem. Both factors, base classifier and error-correcting coding matrix are considered simultaneously. And subgragphs, which are shareable by different classes, are wisely used to improve the classification performance. The experimental results on multi-class object recognition show the effectiveness of the proposed algorithm. Bang Zhang, Getian Ye, Yang Wang 0002, Wei Wang 0011, Jie Xu 0008, Gunawan Herman |
ICPR | 1 |
| 2009 | Incremental EM for Probabilistic Latent Semantic Analysis on Human Action RecognitionabstractHuman action recognition is a significant task in automatic understanding systems for video surveillance. Probabilistic Latent Semantic Analysis (PLSA) model has been used to learn and recognize human actions in videos. Specifically, PLSA employs the expectation maximization (EM) algorithm for parameter estimation during the training. The EM algorithm is an iterative estimation scheme that is guaranteed to find a local maximum of the likelihood function. However its convergence usually takes a large number of iterations. For action recognition with large amount of training data, this would result in long training time. This paper presents an incremental version of EM to speed up the training of PLSA without sacrificing performance accuracy. The proposed algorithm is tested on two challenging human action datasets. Experimental results demonstrate that the proposed algorithm converges with fewer number of full passes compared with the batch EM algorithm. And the trained PLSA models achieve comparable or better recognition accuracies than those using batch EM training. Jie Xu 0008, Getian Ye, Yang Wang 0002, Gunawan Herman, Bang Zhang, Jun Yang 0033 |
AVSS | 5 |
| 2009 | Finding shareable informative patterns and optimal coding matrix for multiclass boostingabstractA multiclass classification problem can be reduced to a collection of binary problems using an error-correcting coding matrix that specifies the binary partitions of the classes. The final classifier is an ensemble of base classifiers learned on binary problems and its performance is affected by two major factors: the qualities of the base classifiers and the coding matrix. Previous studies either focus on one of these factors or consider two factors separately. In this paper, we propose a new multiclass boosting algorithm called AdaBoost.SIP that considers both two factors simultaneously. In this algorithm, informative patterns, which are shareable by different classes rather than only discriminative on specific single class, are generated at first. Then the binary partition preferred by each pattern is found by performing stage-wise functional gradient descent on a margin-based cost function. Finally, base classifiers and coding matrix are optimized simultaneously by maximizing the negative gradient of such cost function. The proposed algorithm is applied to scene and event recognition and experimental results show its effectiveness in multiclass classification. Bang Zhang, Getian Ye, Yang Wang 0002, Jie Xu 0008, Gunawan Herman |
ICCV | 1 |
| 2009 | Feature clustering for vehicle detection and tracking in road traffic surveillanceabstractIn this paper, we formulate the feature clustering problem for vehicle detection and tracking as a general MAP problem and solve it using MCMC. The proposed approach exhibits two advantages over existing methods: general Bayesian model can handle arbitrary objective functions and MCMC guarantees global optimal solution. Our algorithm is validated on real-world traffic video sequences, and is shown to outperform the state-of-the-art approach. Jun Yang 0033, Yang Wang 0002, Getian Ye, Arcot Sowmya, Bang Zhang, Jie Xu 0008 |
ICIP | 5 |
| 2009 | Informative frequent assembled feature for face detectionabstractIn this paper, we propose a novel approach to automatically generating, instead of manually designing, discriminative visual features for face detection. The features are composed by multiple local features (e.g., Haar features), and such features can capture not only the local texture information but also their spatial configurations. Therefore, the proposed feature contains rich semantic information so that the classifier built on a set of such features can achieve high accuracy and high efficiency. Experimental results show that the proposed approach outperforms the techniques based on local features and the state-of-the-art discriminative features for face detection. Bang Zhang, Getian Ye, Yang Wang 0002, Wei Wang 0011, Jie Xu 0008, Gunawan Herman, Jun Yang 0033 |
ICIP | 1 |
| 2009 | Multi-instance learning with relational information of instancesabstractMulti-instance learning (MIL) has many applications, including image and text categorization. One of the most effective approaches to MIL is by using support vector machines with multi-instance kernels. In this paper we propose a multi-instance kernel, called MIR-kernel, that takes into account the relational information of instances when computing similarities between bags. The relational information of instances are derived from the statistics of the distances between instances in feature space. The aim of MIR-kernel is to efficiently capture the context in which instances occur within bags, so that it is able to better compute the similarities between bags. Experimental results on image and text categorization demonstrate the effectiveness of the proposed method compared to other methods. Gunawan Herman, Getian Ye, Yang Wang 0002, Jie Xu 0008, Bang Zhang |
WACV | 5 |
| 2008 | An efficient approach to detecting pedestrians in videoabstractIn this paper, we propose an efficient approach to moving pedestrian detection in video. This approach incorporates both motion and shape information and learns a codebook of shape context descriptors from a very small number of training samples. During the testing process, moving edgelets are firstly identified between adjacent frames using a local search method. Shape context descriptors for numerous sample points on identified edgelets are then produced and are matched against the instances of the learned codebook to generate initial hypotheses. The final hypotheses for pedestrians are obtained by pruning initial hypotheses. The proposed approach has the following advantages by comparison with the existing techniques: (1) lower computational cost, (2) lower false positive rate, and (3) fewer training samples. Experiments with a publicly available dataset confirm the performance of the proposed approach. Jie Xu 0008, Getian Ye, Gunawan Herman, Bang Zhang |
ACM Multimedia | 4 |
| 2008 | Region-based image categorization with reduced feature setabstractIn this paper we propose a new algorithm for region-based image categorization that is formulated as a multiple instance learning (MIL) problem. The proposed algorithm transforms the MIL problem into a traditional supervised learning problem, and solves it using a standard supervised learning method. The features used in the proposed algorithm are the hyperclique patterns which are ldquocondensedrdquo into a small set of discriminative features. Each hyperclique pattern consists of multiple strongly-correlated instances (i.e., features). As a result, hyperclique patterns are able to capture the information that are not shared by individual features. The advantages of the proposed algorithm over existing algorithms are threefold: (i) unlike some existing algorithms which use learning methods that are specifically designed for MIL or for certain datasets, the proposed algorithm uses a general-purpose standard supervised learning method, (ii) it uses a significantly small set of features which are empirically more discriminative than the PCA features (i.e. principal components), and (iii) it is simple and efficient and achieves a comparable performance to most state-of-the-art algorithms. The efficiency and good performance of the proposed algorithm make it a practical solution to general MIL problems. In this paper, we apply the proposed algorithm to both drug activity prediction and image categorization, and promising results are obtained. Gunawan Herman, Getian Ye, Jie Xu 0008, Bang Zhang |
MMSP | 4 |
| 2008 | Detecting and recognizing moving pedestrians in videoabstractDetecting and recognizing pedestrians in video footages are two essential and significant tasks in many automatic video understanding systems. In this paper, we propose an efficient approach to moving pedestrian detection and recognition in video. The testing process of this approach involves two main steps: moving edge detection and hypotheses generation. Moving edges are firstly extracted by comparing the edges identified in adjacent frames. Shape context descriptors are then produced for the edge points sampled from the moving edges and matched against the instances of a codebook that is learned from a set of training samples to generate initial hypotheses. Final hypotheses are formed by pruning initial hypotheses with large overlaps. Experiments with a publicly available dataset show that the proposed approach can reliably detect and recognize moving pedestrians in real scenes that contain either different viewing angles or different degrees of occlusions. Jie Xu 0008, Getian Ye, Gunawan Herman, Bang Zhang |
MMSP | 4 |
| 2008 | A practical approach to multiple super-resolution sprite generationabstractThe MPEG-4 video coding standard introduces a novel concept of sprite or mosaic that is a large image composed of pixels belonging to a video object visible throughout a video segment. The sprite captures spatio-temporal information in a very compact way and makes it possible for efficient object-based video compression. In this paper, we propose a practical approach to generating multiple super-resolution sprites for sprite coding. In order to construct super-resolution sprites and reduce coding cost, we firstly partition a video sequence into multiple independent sprites and group the images covering a similar scene into the same sprite. We then propose efficient and practical algorithms for cumulative global motion estimation and super-resolution sprite construction. Experiments with real video sequences show that the proposed approach outperforms the previous single sprite and multiple sprite techniques. Getian Ye, Yang Wang 0002, Jie Xu 0008, Gunawan Herman, Bang Zhang |
MMSP | 5 |