What is semantic segmentation?
Semantic segmentation labels every image pixel. Pixels of the same class share the same label, so touching buildings may form one building region. Choose Instance Segmentation when separate object identities matter.
Inputs and outputs
| Input | Output |
|---|---|
| Image tiles and aligned pixel-level class masks, with consistent dimensions and no-data handling. | A raster class map, optionally with class scores. Retain georeferencing when assembling tiles. |
Where it helps in GIS
Map land cover, road surfaces, water or building regions. Boundary quality depends on resolution and reference-mask quality.
Model families and examples
- U-Net: an encoder-decoder with skip connections, originally developed for biomedical segmentation.
- DeepLabV3: a semantic segmentation architecture available in Torchvision.
- Geospatial foundation encoder plus decoder: an option when the input bands and pretraining fit the task.
These examples explain the task. The sidebar shows only models currently published on GISSchools.
How to start training
Document the class index and ignore value. Split geographically before creating tiles. Train with suitable augmentation, monitor per-class overlap, and inspect full-scene mosaics for seams and boundary errors.
How to judge the result
Use intersection-over-union or Dice by class. Examine thin roads and small patches separately. Exclude no-data pixels from loss and evaluation; a smooth map can still contain systematic mistakes.
Learn more
The category image is an illustration, not an evaluated model prediction.
