Abstract
Remote sensing vision tasks require extensive labeled data across multiple, interconnected domains. However, current generative data augmentation frameworks are task-isolated, i.e., each vision task requires training an independent generative model, and ignore the modeling of geographical information and spatial constraints. To address these issues, we propose TerraGen, a unified layout-to-image generation framework that enables flexible, spatially controllable synthesis of remote sensing imagery for various high-level vision tasks, e.g., detection, segmentation, and extraction. Specifically, Terra-Gen introduces a geographic-spatial layout encoder that unifies bounding box and segmentation mask inputs, combined with a multi-scale injection scheme and mask-weighted loss to explicitly encode spatial constraints, from global structures to fine details. Moreover, we construct the first large-scale multi-task remote sensing layout generation dataset and establish a standardized evaluation protocol for this task. Experimental results show that TerraGen achieves the best image generation quality across diverse tasks. Additionally, TerraGen can be used as a universal data-augmentation generator, enhancing downstream task performance significantly and demonstrating robust cross-task generalization in both full-data and few-shot scenarios.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Geoscience and Remote Sensing |
| DOIs | |
| State | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- data augmentation
- diffusion models
- layout-to-image generation
- multi-task learning
- Remote sensing
Fingerprint
Dive into the research topics of 'TerraGen: A Unified Multi-Task Layout Generation Framework for Remote Sensing Data Augmentation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver