VLDB 2026 Research / reviewers in the wild / expert
Dongming Yang
dblp:225/6778
· DBLP profile ↗
19ranked-venue papers
9as first author
10since 2021 · last 2025
0000-0002-1478-9642ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rethinking erasing strategy on weakly supervised object localization
Yuming Fan, Shikui Wei, Chuangchuang Tan, Dongming Yang, Yao Zhao 0001 |
Signal Process. Image Commun. | 5 |
| 2024 | Fake-GPT: Detecting Fake Image via Large Language Model
Yuming Fan, Dongming Yang, Jiguang Zhang, Bang Yang, Yuexian Zou |
PRCV (8) | 2 |
| 2024 | Dynamic recommender system for chronic disease-focused online health community
Junruo Gao, Yuan Zhao 0014, Dongming Yang |
Expert Syst. Appl. | 3 |
| 2022 | All You Need Is a Second Look: Towards Arbitrary-Shaped Text DetectionabstractArbitrary-shaped text detection is a challenging task since curved texts in the wild are of the complex geometric layouts. Existing mainstream methods follow the instance segmentation pipeline to obtain the text regions. However, arbitrary-shaped texts are difficult to be depicted through one single segmentation network because of the varying scales. In this paper, we propose a two-stage segmentation-based detector, termed as NASK (Need A Second looK), for arbitrary-shaped text detection. Compared to the traditional single-stage segmentation network, our NASK conducts the detection in a coarse-to-fine manner with the first stage segmentation spotting the rectangle text proposals and the second one retrieving compact representations. Specifically, NASK is composed of a Text Instance Segmentation (TIS) network ($1^{st}$stage), a Geometry-aware Text RoI Alignment (GeoAlign) module, and a Fiducial pOint eXpression (FOX) module ($2^{nd}$stage). Firstly, TIS extracts the augmented features with a novel Group Spatial and Channel Attention (GSCA) module and conducts instance segmentation to obtain rectangle proposals. Then, GeoAlign converts these rectangles into the fixed size and encodes RoI-wise feature representations. Finally, FOX disintegrates the text instance into serval pivotal geometrical attributes to refine the detection results. Extensive experimental results on four public benchmarks including Total-Text, SCUT-CTW1500, ICDAR 2015 and ICDAR 2017 MLT verify that our NASK outperforms recent state-of-the-art methods. Meng Cao 0002, Can Zhang 0001, Dongming Yang, Yuexian Zou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | RR-Net: Relation Reasoning for End-to-End Human-Object Interaction DetectionabstractThe task of Human-Object Interaction (HOI) detection devotes to learn how humans interact with surrounding objects via inferring fine-grained triplets of$\langle $human, verb, object$\rangle $. Most HOI feature learning techniques are dependent on pre-detected instance regions or human body-part regions, which are computationally expensive and hardly applicable to end-to-end detectors in real applications. In this paper, based on an end-to-end HOI detector, we make a first try to explore region-independent relation reasoning for HOI detection. We first present a Relation-aware Frame, which brings a progressive structure for interaction inference. Upon the Relation-aware Frame, an Interaction Intensifier Module and a Correlation Parsing Module are carefully designed, where: a) interactive semantics from humans can be exploited and passed to objects to intensify interactions, b) interactive correlations among humans, objects and interactions are integrated to promote predictions. Based on modules above, we construct a fully differentiable and end-to-end trainable network named Relation Reasoning Network (abbr. RR-Net). Extensive experiments show that our proposed RR-Net leads to competitive results compared with the state-of-the-art methods on both V-COCO and HICO-DET benchmarks and improves the baseline about 7.6% and 11.1% relatively, validating that this first effort in exploring region-independent relation reasoning has brought obvious improvement for end-to-end HOI detection. Dongming Yang, Yuexian Zou, Can Zhang 0001, Meng Cao 0002, Jie Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | CoLA: Weakly-Supervised Temporal Action Localization With Snippet Contrastive LearningabstractWeakly-supervised temporal action localization (WS-TAL) aims to localize actions in untrimmed videos with only video-level labels. Most existing models follow the "localization by classification" procedure: locate temporal regions contributing most to the video-level classification. Generally, they process each snippet (or frame) individually and thus overlook the fruitful temporal context relation. Here arises the single snippet cheating issue: "hard" snippets are too vague to be classified. In this paper, we argue that learning by comparing helps identify these hard snip-pets and we propose to utilize snippet Contrastive learning to Localize Actions, CoLA for short. Specifically, we propose a Snippet Contrast (SniCo) Loss to refine the hard snippet representation in feature space, which guides the network to perceive precise temporal boundaries and avoid the temporal interval interruption. Besides, since it is in-feasible to access frame-level annotations, we introduce a Hard Snippet Mining algorithm to locate the potential hard snippets. Substantial analyses verify that this mining strategy efficaciously captures the hard snippets and SniCo Loss leads to more informative feature representation. Extensive experiments show that CoLA achieves state-of-the-art results on THUMOS’14 and ActivityNet v1.2 datasets. Can Zhang 0001, Meng Cao 0002, Dongming Yang, Jie Chen 0001, Yuexian Zou |
CVPR | 3 |
| 2021 | RR-Net: Injecting Interactive Semantics in Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection devotes to learn how humans interact with surrounding objects. Latest end-to-end HOI detectors are short of relation reasoning, which leads to inability to learn HOI-specific interactive semantics for predictions. In this paper, we therefore propose novel relation reasoning for HOI detection. We first present a progressive Relation-aware Frame, which brings a new structure and parameter sharing pattern for interaction inference. Upon the frame, an Interaction Intensifier Module and a Correlation Parsing Module are carefully designed, where: a) interactive semantics from humans can be exploited and passed to objects to intensify interactions, b) interactive correlations among humans, objects and interactions are integrated to promote predictions. Based on modules above, we construct an end-to-end trainable framework named Relation Reasoning Network (abbr. RR-Net). Extensive experiments show that our proposed RR-Net sets a new state-of-the-art on both V-COCO and HICO-DET benchmarks and improves the baseline about 5.5% and 9.8% relatively, validating that this first effort in exploring relation reasoning and integrating interactive semantics has brought obvious improvement for end-to-end HOI detection. Dongming Yang, Yuexian Zou, Can Zhang 0001, Meng Cao 0002, Jie Chen 0001 |
IJCAI | 1 |
| 2021 | GID-Net: Detecting human-object interaction with global and instance dependency
Dongming Yang, Yuexian Zou, Jian Zhang 0002, Ge Li 0002 |
Neurocomputing | 1 |
| 2021 | Synergic learning for noise-insensitive webly-supervised temporal action localization
Can Zhang 0001, Meng Cao 0002, Dongming Yang, Ji Jiang, Yuexian Zou |
Image Vis. Comput. | 3 |
| 2021 | Learning Human-Object Interaction via Interactive Semantic ReasoningabstractHuman-Object Interaction (HOI) detection devotes to learn how humans interact with surrounding objects via inferring triplets of 〈 human, verb, object 〉 . Recent HOI detection methods infer HOIs by directly extracting appearance features and spatial configuration from related visual targets of human and object, but neglect powerful interactive semantic reasoning between these targets. Meanwhile, existing spatial encodings of visual targets have been simply concatenated to appearance features, which is unable to dynamically promote the visual feature learning. To solve these problems, we first present a novel semantic-based Interactive Reasoning Block, in which interactive semantics implied among visual targets are efficiently exploited. Beyond inferring HOIs using discrete instance features, we then design a HOI Inferring Structure to parse pairwise interactive semantics among visual targets in scene-wide level and instance-wide level. Furthermore, we propose a Spatial Guidance Model based on the location of human body-parts and object, which serves as a geometric guidance to dynamically enhance the visual feature learning. Based on the above modules, we construct a framework named Interactive-Net for HOI detection, which is fully differentiable and end-to-end trainable. Extensive experiments show that our proposed framework outperforms existing HOI detection methods on both V-COCO and HICO-DET benchmarks and improves the baseline about 5.9% and 17.7% relatively, validating its efficacy in detecting HOIs. Dongming Yang, Yuexian Zou, Zhu Li 0001, Ge Li 0002 |
IEEE Trans. Image Process. | 1 |
| 2020 | Semanticgan: Generative Adversarial Networks For Semantic Image To Photo-Realistic Image TranslationabstractGenerative Adversarial Networks (GANs) have shown remarkable success in Semantic label map to Photo-realistic image Translation (S2PT) task. However, the results of the state-of-the-art approaches are often limited to blurriness and artifacts, and still far from realistic, since these methods lack effective semantic constrains to preserve the semantic information and ignore the structural correlations between the textures. To address those problems, we propose a SemanticGAN to synthesize high resolution image with fine details and realistic textures from the semantic label map. Specifically, we propose a Semantic Information Preserved Loss (SIPL) to maintain semantic information in the process of the generation via a segmentation model. Furthermore, we develop a novel generator to obtain the correlations between the image textures using newly-designed Correlated Residual Block (CRB). Experiments evaluated on Cityscapes dataset show that SemanticGAN outperforms many recent state-of-the-art methods in terms of qualitative and quantitative performance. Junling Liu, Yuexian Zou, Dongming Yang |
ICASSP | 3 |
| 2020 | A Graph-based Interactive Reasoning for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection devotes to learn how humans interact with surrounding objects via inferring triplets of < human, verb, object >. However, recent HOI detection methods mostly rely on additional annotations (e.g., human pose) and neglect powerful interactive reasoning beyond convolutions. In this paper, we present a novel graph-based interactive reasoning model called Interactive Graph (abbr. in-Graph) to infer HOIs, in which interactive semantics implied among visual targets are efficiently exploited. The proposed model consists of a project function that maps related targets from convolution space to a graph-based semantic space, a message passing process propagating semantics among all nodes and an update function transforming the reasoned nodes back to convolution space. Furthermore, we construct a new framework to assemble in-Graph models for detecting HOIs, namely in-GraphNet. Beyond inferring HOIs using instance features respectively, the framework dynamically parses pairwise interactive semantics among visual targets by integrating two-level in-Graphs, i.e., scene-wide and instance-wide in-Graphs. Our framework is end-to-end trainable and free from costly annotations like human pose. Extensive experiments show that our proposed framework outperforms existing HOI detection methods on both V-COCO and HICO-DET benchmarks and improves the baseline about 9.4% and 15% relatively, validating its efficacy in detecting HOIs. Dongming Yang, Yuexian Zou |
IJCAI | 1 |
| 2020 | A Novel Application of Educational Management Information System based on Micro FrontendsabstractWith the launch of the Education Informatization 2.0 action plan by the Ministry of Education, a large number of college information systems have been born in China. Most of these systems are single page web applications (SPA) based on traditional MVC structures. Due to the complex logic and high coupling between educational businesses, developers need to write a lot of code. The education information system has many businesses and high coupling between businesses that the system often face problems such as bloated frontend businesses, iterative system updates, and difficult incremental function developments. Combined with the idea of service-oriented architecture, this paper proposes a micro frontends solution and applies it to the new generation of graduate information platform of East China Normal University, which has better agile development capabilities. From the aspects of service separation, efficient development, and incremental upgrade, this paper verifies that the architecture can well adapt to the needs of future educational management information system. The design of the micro frontends provides a new idea for the development of a new generation of education information system. Daojiang Wang, Dongming Yang, Daocheng Hong, Qiwen Dong, Shubing Song |
KES | 2 |
| 2020 | DevOps in Practice for Education Management Information System at ECNUabstractWith the rapid development of the Internet, the education information systems have become more prevalent aligning with better management to produce better education. However, the limitations of prior education systems development are gradually exposed, which ignore the changing requirements, the high concurrency bottlenecks and lean development of education information systems. Therefore, we develop and build a novel education information system at ECNU based on DevOps and related techniques. This paper reveals the practice of DevOps for new education information system from four aspects: CI (Continuous Integration), CD (Continuous Deployment), log management, and code quality. Meanwhile, brief technical explanations include Git, Jenkins, Kubernetes, ELK, SonarQube, etc. Through our continuous engineering practice, the new education information system has been developed and implemented at ECNU. The DevOps practice for information system establishes that it is so convenient for developing, testing and release of education information systems, and it also improves reliability, availability and scalability of information platform especially considering the guarantee of efficiency. Daojiang Wang, Dongming Yang, Qiwen Dong, Daocheng Hong |
KES | 3 |
| 2020 | A novel application integration architecture for the education industryabstractSince the Ministry of Education launched Education Informatization 2.0, the digitalization of colleges and universities has entered a stage of rapid growth. However, after more than 20 years of construction, problems such as system barriers and information islands have emerged in the digital construction of university systems. In order to solve such problems between the university systems, this paper proposes an easily expandable and configurable open information integration architecture by considering traditional information integration methods and combining with Web service technology. The architecture handles user service invocation information through a service layer, and manages the registration and invocation of services through a service module. The permission module manages user permissions to prevent information leakage and security issues. The data module abstracts data-related services to provide a basis for the deep use of data. And other optional development services are designed to satisfy special requirements for different platforms. The architecture proposed in this paper can integrate different heterogeneous subsystems in colleges and universities, eliminating the problem of system barriers and information islands, and providing specifications for the construction of new applications. Dongming Yang, Daojiang Wang, Shubing Song, Qiwen Dong |
KES | 1 |
| 2019 | Cascade Region Proposal Networks for Object Detection in the WildabstractAlthough significant progresses have been made in object detection on common benchmarks (i.e., Pascal VOC), object detection in the wild is still challenging due to the serious data inadequacy and imbalance. To address this challenge, we construct a cascade framework which consists of multiple region proposal networks, referred to as C-RPNs. The essence of C-RPNs is adopting multiple stages to mine hard samples and learn better classifiers. Meanwhile, a feature chain and a score chain are proposed to help learning more discriminative representations for proposals. Moreover, a loss function of cascade stages is designed to train cascade classifiers through backpropagation. Our newly proposed object detection method is evaluated on Pascal VOC and a challenging dataset of littoral birds named BSBDV 2017. Our method outperforms baseline by an obvious margin, validating its efficacy for detection in the wild. Dongming Yang, Yuexian Zou |
ICME | 1 |
| 2019 | Enhancing Scene Text Detection via Fused Semantic Segmentation Network with Attention
Yuexian Zou, Dongming Yang |
MMM (1) | 3 |
| 2019 | C-RPNs: Promoting object detection in real world via a cascade structure of Region Proposal Networks
Dongming Yang, Yuexian Zou, Jian Zhang 0002, Ge Li 0002 |
Neurocomputing | 1 |
| 2018 | Real-time pedestrian detection via hierarchical convolutional feature
Dongming Yang, Jiguang Zhang, Shibiao Xu, Shuiying Ge, G. Hemanth Kumar, Xiaopeng Zhang 0001 |
Multim. Tools Appl. | 1 |