EDBT 2026 Demo / reviewers in the wild / expert
Shuang Liang 0001
dblp:20/1080-1
· DBLP profile ↗
43ranked-venue papers
15as first author
25since 2021 · last 2026
0000-0003-0457-6093ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 23 · 9 first-author · 10 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bi-Level Keypoint Relation Helps Versatile and Occluded Human Pose EstimationabstractRecently, there has been significant progress in 2D pose estimation. However, accurately localizing limb keypoints and occluded keypoints is still challenging. To tackle these difficulties, prior in human body structure has been leveraged in previous studies. One approach involves localizing a challenging keypoint by utilizing its neighbor keypoint. A previous study successfully employed neighbor-joint spatial relation (SR), which transfers features from a neighbor keypoint to the target keypoint being predicted. Building upon this idea, our work extends the keypoint relation-based method by incorporating another level of keypoint relation, namely channel-wise feature relation. This additional feature relation (FR) module assists in selecting more suitable neighbor keypoint feature channels and enhances the effectiveness of SR. By combining FR and SR, we develop a simple and intuitive bi-level keypoint relation module that can be trained end-to-end with existing methods. Through comprehensive experimental results and ablation studies, we demonstrate the effectiveness of our approach. Shuang Liang 0001, Chi Xie 0001, Jiewen Wang, Gang Chu, Shuwei Yan |
FG | 1 |
| 2026 | Targeted Mining of Time-Interval Related PatternsabstractCompared with frequent pattern mining, sequential pattern mining emphasizes the temporal aspect and finds broad applications across various fields. However, numerous studies treat temporal events as single time points, neglecting their durations. Time-interval-related pattern (TIRP) mining is introduced to address this issue and has been applied to healthcare analytics, stock prediction, etc. Typically, mining all patterns is not only computationally challenging for accurate forecasting but also resource-intensive in terms of time and memory. Targeting the extraction of TIRPs based on specific criteria can improve data analysis efficiency and better align with customer pReferences. Therefore, this article proposes a novel algorithm called TaTIRP to discover targeted TIRP. In addition, we develop multiple pruning strategies to eliminate redundant extension operations, thereby enhancing performance on large-scale datasets. Finally, we conduct experiments on various real-world and synthetic datasets to validate the accuracy and efficiency of the proposed algorithm. Shuang Liang 0001, Wensheng Gan, Philip S. Yu, Shengjie Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Classifier Recalibration for Human-Object Interaction Detection
Shuwei Yan, Shuang Liang 0001, Kenan Ye, Baihua Liu, Chi Xie 0001, Shengjie Zhao 0001 |
ICIC (6) | 2 |
| 2025 | Rethinking sketch-based 3D shape retrieval: A simple baseline and benchmark reconstruction
Shuang Liang 0001, Weidong Dai, Changmao Cheng, Yiyang Cai |
Neurocomputing | 1 |
| 2025 | Front-door causal attention for unbiased panoptic scene graph generation
Shuang Liang 0001, Baihua Liu, Shuwei Yan, Hongming Zhu |
Neurocomputing | 1 |
| 2025 | RelationLMM: Large Multimodal Model as Open and Versatile Visual Relationship GeneralistabstractVisual relationships are crucial for visual perception and reasoning, and cover tasks like Scene Graph Generation, Human-Object Interaction, and object affordance. Despite significant efforts, this field still suffers from the following limitations: specialists for a specific task without considering similar ones, strict and complex task formulations with limited flexibility, and underexploited reasoning with language and knowledge. To solve these limitations, we seek to build a new framework, one model for all tasks, over Large Multimodal Models (LMMs). LMMs offer the potential of unifying tasks, flexible forms, and reasoning with language. However, they fail to handle visual relationship tasks well. We find the obstacles include the conflicts between different tasks and insufficient instance-level information. We solve these problems by reforming the data for LMMs, rather than architectures, considering their strong language-in language-out capability. We propose to disassemble tasks into simple and common sub-tasks, verbally estimate instance confidence, and augment instance diversity, all without additional modules. These strategies help us build a visual relationship generalist, RelationLMM, with a simple architecture. Exhaustive experiments demonstrate RelationLMM is strong, generalizable and flexible to different tasks, with one model and one suite of weight. Chi Xie 0001, Shuang Liang 0001, Zhao Zhang 0018, Feng Zhu 0006, Rui Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Self-Supervised 3-D Action Recognition by Contrasting Context-Enhanced Action Embeddingsabstract3-D action recognition become a fast-pacing field in recent years. However, traditional approaches have limitations. They either focus on modeling overly detailed yet redundant information by reconstructing the coordinates of each body joint. Alternatively, they treat actions as a whole, overlooking the spatial and temporal variations in the semantic locality of actions. To address these limitations, we propose representing long-term actions as contexts of short-term actions organized by locality-aware graphs. In our framework, we take the inspiration that the continuity of motion and pose variations generate higher correlations. These correlations occur among spatio-temporally adjacent joints. Built upon this, we craft short-term actions as embeddings using spatio-temporal graph convolutions. This graph-based encoding not only captures richer high-level semantics but also maintains an awareness of the topology. To capture long-term action dynamics effectively, we integrate a graph convolutional gated recurrent unit (GraphGRU) for the fusion of action embeddings. Additionally, we introduce the context-aware topological attention (CTA) mechanism. Positioned between embedding encoding and context aggregation phases, CTA amplifies the features of context-relevant nodes. Lastly, we create self-supervision by contrasting predicted embeddings with actual encoded embeddings. This approach explicitly learns changes in dynamics to obtain distinct embeddings. Empirical evaluations demonstrate that our approach outperforms mainstream unsupervised 3-D action recognition methods. Kenan Ye, Brian Nlong Zhao, Shuang Liang 0001, Wenzhen Jia |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Label Correction For Sketch-Based 3d Shape RetrievalabstractSketch-based 3D shape retrieval has become an intense topic in the multimedia society. Recent progress in this field can be mainly attributed to well-designed loss functions. However, two essential aspects remain to be considered by previous works: Firstly, the inherent abstract, sparse, and diverse nature of the sketch leads to inevitable label noise, which significantly impairs model efficacy. Secondly, a deficient focus on hard sample mining can lead to model performance degradation. To address these issues, we propose a label self-correction method. Starting from the margin-based loss function, we relate the ground truth class center of the sample to its nearest negative class center. By leveraging the decision boundary, this module bolsters the learning capacity within similar classes, ensuring the separability of the feature space. Extensive experiments on two benchmarks shows that our method surpasses previous state-of-the-art methods. Shuang Liang 0001, Jiaming Lu, Yiyang Cai |
ICASSP | 1 |
| 2024 | Causal Intervention for Panoptic Scene Graph GenerationabstractPanoptic Scene Graph Generation (PSG) is a computer vision task that involves recognizing objects and their relationships within a given image. However, due to the long-tail distribution of the dataset, current PSG methods face bias issues. Previous debiasing methods usually rely on resampling and reweighting, which are unable to address the issue fundamentally, and may lead to underfitting for the head categories to some extent. Causal intervention, as a method to identify confounders from a causal inference perspective, can serve as a suitable means of debiasing. In this work, we propose a novel debiasing method based on causal intervention, which treats the long-tail distribution prior of the dataset as a confounder and eliminates it using the backdoor criterion. An uncertainty estimation module is further employed to assist in determining hard samples. Experiments demonstrate a significant improvement in the competitiveness of our approach compared to the baseline. Moreover, our method exhibits notable enhancements in accuracy and generalization on tail categories. Shuang Liang 0001, Chi Xie 0001 |
ICME | 1 |
| 2024 | Sketch-based 3D shape retrieval via teacher-student learning
Shuang Liang 0001, Weidong Dai, Yiyang Cai, Chi Xie 0001 |
Comput. Vis. Image Underst. | 1 |
| 2024 | Scribble-based complementary graph reasoning network for weakly supervised salient object detection
Shuang Liang 0001, Zhiqi Yan, Chi Xie 0001, Hongming Zhu, Jiewen Wang |
Comput. Vis. Image Underst. | 1 |
| 2024 | Relation with Free Objects for Action RecognitionabstractRelevant objects are widely used for aiding human action recognition in still images. Such objects are founded by a dedicated and pre-trained object detector in all previous methods. Such methods have two drawbacks. First, training an object detector requires intensive data annotation. This is costly and sometimes unaffordable in practice. Second, the relation between objects and humans are not fully taken into account in training. This work proposes a systematic approach to address the two problems. We propose two novel network modules. The first is an object extraction module that automatically finds relevant objects for action recognition, without requiring annotations. Thus, it is free . The second is a human-object relation module that models the pairwise relation between humans and objects, and enhances their features. Both modules are trained in the action recognition network, end-to-end. Comprehensive experiments and ablation studies on three datasets for action recognition in still images demonstrate the effectiveness of the proposed approach. Our method yields state-of-the-art results. Specifically, on the HICO dataset, it achieves 44.9% mAP, which is 12% relative improvement over the previous best result. In addition, this work makes an observational contribution that it is no longer necessary to rely on a pre-trained object detector for this task. Relevant objects can be found via end-to-end learning with only action labels. This is encouraging for action recognition in the wild. Models and code will be released. Shuang Liang 0001, Wentao Ma 0004, Chi Xie 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Patch excitation network for boxless action recognition in still images
Shuang Liang 0001, Jiewen Wang, Zikun Zhuang |
Vis. Comput. | 1 |
| 2023 | Category Query Learning for Human-Object Interaction ClassificationabstractUnlike most previous HOI methods that focus on learning better human-object features, we propose a novel and complementary approach called category query learning. Such queries are explicitly associated to interaction categories, converted to image specific category representation via a transformer decoder, and learnt via an auxiliary image-level classification task. This idea is motivated by an earlier multi-label image classification method, but is for the first time applied for the challenging human-object interaction classification task. Our method is simple, general and effective. It is validated on three representative HOI baselines and achieves new state-of-the-art results on two benchmarks. Code will be available at https://github.com/charles-xie/CQL. Chi Xie 0001, Fangao Zeng, Yue Hu 0011, Shuang Liang 0001 |
CVPR | 4 |
| 2023 | Uncertainty-Aware Cross-Modal Transfer Network for Sketch-Based 3D Shape RetrievalabstractIn recent years, sketch-based 3D shape retrieval has attracted growing attention. While many previous studies have focused on cross-modal matching between hand-drawn sketches and 3D shapes, the critical issue of how to handle low-quality and noisy samples in sketch data has been largely neglected. This paper presents an uncertainty-aware cross-modal transfer network (UACTN) that addresses this issue. UACTN decouples the representation learning of sketches and 3D shapes into two separate tasks: classification-based sketch uncertainty learning and 3D shape feature transfer. We first introduce an end-to-end classification-based approach that simultaneously learns sketch features and uncertainty, allowing uncertainty to prevent overfitting noisy sketches by assigning different levels of importance to clean and noisy sketches. Then, 3D shape features are mapped into the pre-learned sketch embedding space for feature alignment. Extensive experiments and ablation studies on two benchmarks demonstrate the superiority of our proposed method compared to state-of-the-art methods. Yiyang Cai, Jiaming Lu, Jiewen Wang, Shuang Liang 0001 |
ICME | 4 |
| 2023 | Compositional Learning in Transformer-Based Human-Object Interaction DetectionabstractHuman-object interaction (HOI) detection is an important part of understanding human activities and visual scenes. The long-tailed distribution of labeled instances is a primary challenge in HOI detection, promoting research in few-shot and zero-shot learning. Inspired by the combinatorial nature of HOI triplets, some existing approaches adopt the idea of compositional learning, in which object and action features are learned individually and re-composed as new training samples. However, these methods follow the CNN-based two-stage paradigm with limited feature extraction ability, and often rely on auxiliary information for better performance. Without introducing any additional information, we creatively propose a transformer-based framework for compositional HOI learning. Human-object pair representations and interaction representations are re-composed across different HOI instances, which involves richer contextual information and promotes the generalization of knowledge. Experiments show our simple but effective method achieves state-of-the-art performance, especially on rare HOI classes. Zikun Zhuang, Ruihao Qian, Chi Xie 0001, Shuang Liang 0001 |
ICME | 4 |
| 2023 | Described Object Detection: Liberating Object Detection with Flexible ExpressionsabstractDetecting objects based on language information is a popular task that includes Open-Vocabulary object Detection (OVD) and Referring Expression Comprehension (REC). In this paper, we advance them to a more practical setting called *Described Object Detection* (DOD) by expanding category names to flexible language expressions for OVD and overcoming the limitation of REC only grounding the pre-existing object. We establish the research foundation for DOD by constructing a *Description Detection Dataset* ($D^3$). This dataset features flexible language expressions, whether short category names or long descriptions, and annotating all described objects on all images without omission. By evaluating previous SOTA methods on $D^3$, we find some troublemakers that fail current REC, OVD, and bi-functional methods. REC methods struggle with confidence scores, rejecting negative instances, and multi-target scenarios, while OVD methods face constraints with long and complex descriptions. Recent bi-functional methods also do not work well on DOD due to their separated training procedures and inference strategies for REC and OVD tasks. Building upon the aforementioned findings, we propose a baseline that largely improves REC methods by reconstructing the training data and introducing a binary classification sub-task, outperforming existing methods. Data and code are available at https://github.com/shikras/d-cube and related works are tracked in https://github.com/Charles-Xie/awesome-described-object-detection. Chi Xie 0001, Zhao Zhang 0018, Feng Zhu 0006, Rui Zhao 0001, Shuang Liang 0001 |
NeurIPS | 6 |
| 2023 | SVFNeXt: Sparse Voxel Fusion for LiDAR-Based 3D Object Detection
Deze Zhao, Shengjie Zhao 0001, Shuang Liang 0001 |
PRICAI (3) | 3 |
| 2023 | Temporal Dropout for Weakly Supervised Action LocalizationabstractWeakly supervised action localization is a challenging problem in video understanding and action recognition. Existing models usually formulate the training process as direct classification using video-level supervision. They tend to only locate the most discriminative parts of action instances and produce temporally incomplete detection results. A natural solution for this problem, the adversarial erasing strategy, is to remove such parts from training so that models can attend to complementary parts. Previous works do it in an offline and heuristic way. They adopt a multi-stage pipeline, where discriminative regions are determined and erased under the guidance of detection results from last stage. Such a pipeline can be both ineffective and inefficient, possibly hindering the overall performance. On the contrary, we combine adversarial erasing with dropout mechanism and propose a Temporal Dropout Module that learns where to remove in a data-driven and online manner. This plug-and-play module is trained without iterative stages, which not only simplifies the pipeline but also makes the regularization during training easier and more adaptive. Experiments show that the proposed method outperforms previous erasing-based methods by a large margin. More importantly, it achieves universal improvement when plugged into various direct classification methods and obtains state-of-the-art performance. Chi Xie 0001, Zikun Zhuang, Shengjie Zhao 0001, Shuang Liang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Structural Attention for Channel-Wise Adaptive Graph Convolution in Skeleton-Based Action RecognitionabstractIn skeleton-based action recognition, graph convolutions to model human action dynamics have been widely implemented and achieved remarkable results. Among these convolutions, channel-wise adaptive graph convolution shows outstanding performance. However, this method focuses too much on capturing correlation between joints within each channel and lacks the capability of learning structural features, which are generally hidden in geometric property of the skeleton on spatial domain. Our proposed method (SA-GCN) introduces symmetry trajectory attention module to measure the relation between left and right part of body and part relation attention module for exploration of the attention on general relation of each part. Both modules are intended to make full use of structural features in skeleton, further strengthening advantages of graph convolution. Experiments on three datasets (NW-UCLA, NTU-RGB+D and NTU-RGB+D 120) demonstrate state-of-the-art performance of our model, especially on joint modality. Ruihao Qian, Jiewen Wang, Jianxiu Wang, Shuang Liang 0001 |
ICME | 4 |
| 2022 | Pose-Enhanced Relation Feature for Action Recognition in Still Images
Jiewen Wang, Shuang Liang 0001 |
MMM (1) | 2 |
| 2022 | Joint relation based human pose estimation
Shuang Liang 0001, Gang Chu, Chi Xie 0001, Jiewen Wang |
Vis. Comput. | 1 |
| 2021 | Recurrent Graph Convolutional Autoencoder for Unsupervised Skeleton-Based Action RecognitionabstractSkeleton-based action recognition is a significant task in computer vision due to its robustness and wide application. Most unsupervised methods do not employ topological information of skeleton graphs, which actually ignore the spatial dependencies of action sequences. In this paper, we introduce a Recurrent Graph Convolutional Autoencoder (RGCA) for unsupervised action recognition from skeleton data. Our method explicitly exploits the spatial relationships among every frame’s joints while preserving the long-term temporal dynamics in whole sequences. Moreover, a Spatial Joints Attention Module is employed to measure the importance of joints in the input sequence automatically. We conduct experiments on three datasets (NTU RGB+D 60, NW-UCLA, and UWA3D) and exceed the state-of-the-art performance. Sev-Jue Zhao, Chi Xie 0001, Kenan Ye, Shuang Liang 0001 |
ICME | 5 |
| 2021 | Automatic Pose Quality Assessment for Adaptive Human Pose Refinement
Gang Chu, Chi Xie 0001, Shuang Liang 0001 |
MMM (1) | 3 |
| 2021 | Uncertainty Learning for Noise Resistant Sketch-Based 3D Shape RetrievalabstractRecently, sketch-based 3D shape retrieval has received growing attention in the community of computer graphics and computer vision. Most previous works focus on the problem of how to reduce the large cross-modality difference between 2D sketch and 3D shape data and make significant progress. Nevertheless, little attention has been paid to another important problem of how to deal with noise in the sketch data. For the first time, this work investigates the problem of noisy sketch data. It firstly provides qualitative and insightful analysis on the impact of noise, revealing that the noisy data are a key factor for unsatisfactory retrieval performance, as they cause severe over fitting and impair feature learning. Thus, the issue is worthy of serious treatment. Then, we propose to estimate sketch noise as data uncertainty, motivated by existing ideas that model data uncertainty with a distributional representation. We present methods with simple network structure and loss functions. They achieve strong results and establish new state-of-the-art on two benchmarks. Comprehensive experiment results, ablation studies, and insightful analysis validate the effectiveness of our methods, revealing that sketch feature learning with uncertainty is crucial for noise resistant sketch based 3D shape retrieval. Shuang Liang 0001, Weidong Dai |
IEEE Trans. Image Process. | 1 |
| 2020 | Cross-Modal Guidance Network For Sketch-Based 3d Shape RetrievalabstractThe main challenge of sketch-based 3D shape retrieval is the large cross-modal differences between 2D sketches and 3D shapes. Most recent works employed two heterogeneous networks and a shared loss to directly map the features from different modalities to a common feature space, which failed to reduce the cross-modal differences effectively. In this paper, we propose a novel method that adopts a teacher-student strategy to learn an aligned cross-modal feature space indirectly. Specifically, our method first employs a classification network to learn the discriminative features of 3D shapes. Then, the pre-learned features are considered as a teacher to guide the feature learning of 2D sketches. In order to align the cross-modal features, 2D sketch features are transferred to the pre-learned 3D feature space. Our experiments on two benchmark datasets demonstrate that our method obtains superior retrieval performance than the state-of-the-art approaches. Weidong Dai, Shuang Liang 0001 |
ICME | 2 |
| 2020 | Human-Object Relation Network For Action Recognition In Still ImagesabstractSurrounding object information has been widely used for action recognition. However, the relation between human and object, as an important cue, is usually ignored in the still image action recognition field. In this paper, we propose a novel approach for action recognition. The key to ours is a human-object relation module. By using the appearance as well as the spatial location of human and object, the module can compute (b) the pair-wise relation information between human and object to enhance features for action classification and can be trained jointly with our action recognition network. Experimental results on two popular datasets demonstrate the effectiveness of the proposed approach. Moreover, our method yields the new state-of-the-art results of 92.8% and 94.6% mAP on the PASCAL VOC 2012 Action and Stanford 40 Actions datasets respectively. Ablation study and visualization confirm the proposed method can model and utilize the human-object relation for action recognition. Wentao Ma 0004, Shuang Liang 0001 |
ICME | 2 |
| 2018 | Pseudo Mask Augmented Object DetectionabstractIn this work, we present a novel and effective framework to facilitate object detection with the instance-level segmentation information that is only supervised by bounding box annotation. Starting from the joint object detection and instance segmentation network, we propose to recursively estimate the pseudo ground-truth object masks from the instance-level object segmentation network training, and then enhance the detection network with top-down segmentation feedbacks. The pseudo ground truth mask and network parameters are optimized alternatively to mutually benefit each other. To obtain the promising pseudo masks in each iteration, we embed a graphical inference that incorporates the low-level image appearance consistency and the bounding box annotations to refine the segmentation masks predicted by the segmentation network. Our approach progressively improves the object detection performance by incorporating the detailed pixel-wise information learned from the weakly-supervised segmentation network. Extensive evaluation on the detection task in PASCAL VOC 2007 and 2012 [12] verifies that the proposed approach is effective. Xiangyun Zhao, Shuang Liang 0001 |
CVPR | 2 |
| 2018 | Integral Human Pose Regression
Xiao Sun 0001, Fangyin Wei, Shuang Liang 0001 |
ECCV (6) | 4 |
| 2018 | Immersing Web3D Furniture into Real Interior ImagesabstractPlatforms for interior DIY(Do It Yourself) design should hold sufficient realistic sense and light manual operation in interior modeling, besides more flexibility and adaptation of online editing are also essential. But current pure Web3D or pure image based platforms are hardly meet those goals. Therefore this paper presents a lightweight and immersive solutions by editing virtual 3D furniture into captured 2D interior pictures interactively, placing them at the optimal location automatically and rendering them in real time to consistently harmonize them with real interior pictures in terms of geometric layout and lighting visual effects. Our contributions consist in: (1)cuboid modeling from camera captured interior pictures interactively without loss of realistic sense and with light manual operation; (2)lightweight automatic furniture arrangement method with enough flexibility and adaptation for online editing; (3)lightweight IBL(Image Based Lighting) and PBR(Physic Based Rendering) to make virtual furniture immerse into real interior image more visually authentic. Compared with those existing online systems, this solution can provide low-cost, convenient, pervasive online services for interior DIY design over mobile Internet. Shuang Liang 0001, Jinyuan Jia 0002 |
VR | 2 |
| 2018 | Compositional Human Pose Regression
Shuang Liang 0001, Xiao Sun 0001 |
Comput. Vis. Image Underst. | 1 |
| 2017 | Compositional Human Pose RegressionabstractRegression based methods are not performing as well as detection based methods for human pose estimation. A central problem is that the structural information in the pose is not well exploited in the previous regression methods. In this work, we propose a structure-aware regression approach. It adopts a reparameterized pose representation using bones instead of joints. It exploits the joint connection structure to define a compositional loss function that encodes the long range interactions in the pose. It is simple, effective, and general for both 2D and 3D pose estimation in a unified setting. Comprehensive evaluation validates the effectiveness of our approach. It significantly advances the state-of-the-art on Human3.6M [20] and is competitive with state-of-the-art results on MPII [3]. Xiao Sun 0001, Jiaxiang Shang, Shuang Liang 0001 |
ICCV | 3 |
| 2016 | Attribute Recognition from Adaptive Parts
Luwei Yang, Ligen Zhu, Shuang Liang 0001 |
BMVC | 4 |
| 2016 | 3D tree skeletonization from multiple images based on PyrLK optical flow
Dejia Zhang, Ning Xie 0003, Shuang Liang 0001, Jinyuan Jia 0002 |
Pattern Recognit. Lett. | 3 |
| 2015 | Cascaded hand pose regressionabstractWe extends the previous 2D cascaded object pose regression work [9] in two aspects so that it works better for 3D articulated objects. Our first contribution is 3D pose-indexed features that generalize the previous 2D parameterized features and achieve better invariance to 3D transformations. Our second contribution is a principled hierarchical regression that is adapted to the articulated object structure. It is therefore more accurate and faster. Comprehensive experiments verify the state-of-the-art accuracy and efficiency of the proposed approach on the challenging 3D hand pose estimation problem, on a public dataset and our new dataset. Xiao Sun 0001, Shuang Liang 0001, Xiaoou Tang, Jian Sun 0001 |
CVPR | 3 |
| 2015 | Object proposal by multi-branch hierarchical segmentationabstractHierarchical segmentation based object proposal methods have become an important step in modern object detection paradigm. However, standard single-way hierarchical methods are fundamentally flawed in that the errors in early steps cannot be corrected and accumulate. In this work, we propose a novel multi-branch hierarchical segmentation approach that alleviates such problems by learning multiple merging strategies in each step in a complementary manner, such that errors in one merging strategy could be corrected by the others. Our approach achieves the state-of-the-art performance for both object proposal and object detection tasks, comparing to previous object proposal methods. Chaoyang Wang 0001, Long Zhao 0003, Shuang Liang 0001, Liqing Zhang 0001, Jinyuan Jia 0002 |
CVPR | 3 |
| 2015 | Sketch Matching on Topology Product GraphabstractSketch matching is the fundamental problem in sketch based interfaces. After years of study, it remains challenging when there exists large irregularity and variations in the hand drawn sketch shapes. While most existing works exploit topology relations and graph representations for this problem, they are usually limited by the coarse topology exploration and heuristic (thus suboptimal) similarity metrics between graphs. We present a new sketch matching method with two novel contributions. We introduce a comprehensive definition of topology relations, which results in a rich and informative graph representation of sketches. For graph matching, we propose topology product graph that retains the full correspondence for matching two graphs. Based on it, we derive an intuitive sketch similarity metric whose exact solution is easy to compute. In addition, the graph representation and new metric naturally support partial matching, an important practical problem that received less attention in the literature. Extensive experimental results on a real challenging dataset and the superior performance of our method show that it outperforms the state-of-the-art. Shuang Liang 0001, Wenyin Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Learning best views of 3D shapes from sketch contour
Long Zhao 0003, Shuang Liang 0001, Jinyuan Jia 0002 |
Vis. Comput. | 2 |
| 2014 | Size and Location Matter: A New Baseline for Salient Object Detection
Long Zhao 0003, Shuang Liang 0001, Jinyuan Jia 0002 |
ACCV (3) | 2 |
| 2014 | Saliency Optimization from Robust Background DetectionabstractRecent progresses in salient object detection have exploited the boundary prior, or background information, to assist other saliency cues such as contrast, achieving state-of-the-art results. However, their usage of boundary prior is very simple, fragile, and the integration with other cues is mostly heuristic. In this work, we present new methods to address these issues. First, we propose a robust background measure, called boundary connectivity. It characterizes the spatial layout of image regions with respect to image boundaries and is much more robust. It has an intuitive geometrical interpretation and presents unique benefits that are absent in previous saliency measures. Second, we propose a principled optimization framework to integrate multiple low level cues, including our background measure, to obtain clean and uniform saliency maps. Our formulation is intuitive, efficient and achieves state-of-the-art results on several benchmark datasets. Wangjiang Zhu, Shuang Liang 0001, Jian Sun 0001 |
CVPR | 2 |
| 2013 | Online stroke segmentation by quick penalty-based dynamic programmingabstractA stroke segmentation method named quick penalty‐based dynamic programming is proposed for splitting a sketchy stroke into several regular primitive shapes, such as line segments and elliptical arcs. The authors extend the dynamic programming framework with a customisable penalty function, which measures the correctness of splitting a stroke at a particular point. With the help of the penalty function, the proposed dynamic programming framework can finish the stroke segmentation process without any prior knowledge of the number and/or the type of segments contained in the sketchy stroke. Its response time is sufficiently short for online applications, even for long strokes. Experiments show that the proposed method is robust for strokes with arbitrary shape and size. Wenyin Liu, Tong Lu 0002, Yajie Yu, Shuang Liang 0001, Rui Zhang 0031 |
IET Comput. Vis. | 4 |
| 2010 | A creative try: composing weaving patterns by playing on a multi-input deviceabstractWoven fabrics are widely used in clothing because of their parallel and interlaced properties, which are formed by weaving. Creating a weaving pattern, especially hand weaving for interlacing yarns is a cumbersome task in the textile industry. In this paper, we propose two kinds of playing for creating weaving patterns on multi-input devices: the tie-up plan and the lift plan. Discrete notes on the treble staff are translated into signatures of treadling sequences and discrete notes in the bass staff are translated into signatures of theadling sequences. Artists can use their right hand to compose a treadling sequence for weft yarns and their left hand to play a threading sequence for warp yarns. The treadling and threading sequences become the notes on the full gamut of shafts and treadles. Our result shows that we are able to compose a family of weaving patterns in a similar way to playing the piano in a short time. George Baciu, Shuang Liang 0001 |
VRST | 3 |
| 2008 | Sketch retrieval and relevance feedback with biased SVM classification
Shuang Liang 0001, Zhengxing Sun |
Pattern Recognit. Lett. | 1 |