
Exact checkpoint
google/vit-base-patch16-224 at revision 3f49326eb077187dfe1c2a2bb15fbd74e6ab91e3. This checkpoint was pretrained on ImageNet-21k and fine-tuned on ImageNet-1k. The provider card credits conversion through timm from JAX into PyTorch; it is not a claim that the original authors directly released this safetensors serialization.
How it works
The image becomes a sequence of square patch embeddings. Positional information retains their arrangement, while attention allows distant patches to exchange information. A classification token summarizes the scene for the prediction head.
Typical input and output
This selected variant uses RGB 224 by 224 input and 16 by 16 patches. Use its image processor. Output is an ImageNet class-score vector for the whole image, not a spatial label grid.
Strengths
Attention provides a different way to combine context across a scene. Keeping a Transformer and a CNN baseline helps separate architecture effects from dataset or split effects.
GIS and remote-sensing use
Fine-tune for scene classification of spatially defined chips. Decide whether the desired label represents dominant cover, land use or presence of a feature before annotating. A scene classifier cannot replace a segmentation model when individual pixels need labels.
Limitations
The provided classes describe ImageNet images rather than geospatial land-use categories. Small objects can be diluted when a large image is resized. Attention visualizations are not automatically reliable explanations or GIS boundaries.
Training and fine-tuning
Replace the classifier with the target label set. Use consistent chip extent and processor settings, then compare frozen-backbone and fine-tuned runs. Keep neighboring or overlapping chips in the same split. If experimenting with larger image sizes, follow the implementation guidance for positional embeddings instead of silently changing resolution.
Framework and hardware
PyTorch and Hugging Face Transformers; original JAX implementation is also linked. Load the exact revision and its processor. A GPU is useful for fine-tuning; practical batch size depends on resolution and device memory. No local inference speed was measured.
Files, license and download
Official external checkpoint: model.safetensors is 346,293,852 bytes, above the existing 100 MiB resource ceiling. The provider declares Apache-2.0; size is the reason this checkpoint is not mirrored here.
Before using a result
- Keep imagery rights, acquisition dates and preprocessing with each experiment.
- Separate locations before tiling to avoid train/test leakage.
- Inspect failures and uncertain outputs alongside the original imagery.
- Preserve coordinate transforms and verify units before exporting a GIS layer.
Validation status: Source provenance and website delivery are checked. GISSchools has not run model inference, training or an accuracy benchmark for this record. Feature images and gallery diagrams are explanatory illustrations, not predictions.
Official references
Examples and images
Downloads and resources
Official external checkpoint: model.safetensors is 346,293,852 bytes, above the existing 100 MiB resource ceiling. The provider declares Apache-2.0; size is the reason this checkpoint is not mirrored here.



Comments
No comments yet. Start the discussion.
Leave a reply
Your email address will not be published. Required fields are marked.