Image Classification
Transformers
Safetensors
English
vit
vision-transformer
stanford-40-actions
computer-vision
Eval Results (legacy)
Instructions to use lucasddmc/vit-b16-stanford40-actions with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lucasddmc/vit-b16-stanford40-actions with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-classification", model="lucasddmc/vit-b16-stanford40-actions") pipe("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/hub/parrots.png")# pip install -U transformers accelerate # Load model directly from transformers import AutoImageProcessor, AutoModelForImageClassification processor = AutoImageProcessor.from_pretrained("lucasddmc/vit-b16-stanford40-actions") model = AutoModelForImageClassification.from_pretrained("lucasddmc/vit-b16-stanford40-actions", device_map="auto") - Notebooks
- Google Colab
- Kaggle
ViT-B/16 fine-tuned on Stanford 40 Actions
Vision Transformer ViT-B/16 fine-tuned on the Stanford 40 Actions dataset for action recognition in still images.
This model reaches 91.5% accuracy on the held-out test split.
Model description
- Architecture: ViT-Base, patch size 16 (224ร224 resolution)
- Base model:
google/vit-base-patch16-224-in21k - Task: Single-label image classification (40 human actions)
- Input: RGB image with a person performing an action
- Output: One of 40 action labels (e.g.
applauding,rowing a boat,texting message)
How to use
from PIL import Image
import requests
from transformers import AutoImageProcessor, AutoModelForImageClassification
model_id = "lucasddmc/vit-b16-stanford40-actions"
processor = AutoImageProcessor.from_pretrained(model_id)
model = AutoModelForImageClassification.from_pretrained(model_id)
url = "https://example.com/some_action_image.jpg"
image = Image.open(requests.get(url, stream=True).raw).convert("RGB")
inputs = processor(images=image, return_tensors="pt")
outputs = model(**inputs)
logits = outputs.logits
predicted_class_idx = logits.argmax(-1).item()
print(model.config.id2label[predicted_class_idx])
Training details (summary)
- Dataset: Stanford 40 Actions
- Framework: Trained with PyTorch + timm; exported to Hugging Face Transformers format (config.json/model.safetensors) for inference
- Preprocessing / Augmentation (typical for ViT):
- Resize to 256ร256
- RandomResizedCrop to 224ร224
- RandomHorizontalFlip
- Normalization with ImageNet mean/std
- Loss: CrossEntropyLoss (with/without label smoothing, depending on your setup)
- Optimizer: AdamW
- Body LR:
1e-5โ3e-5 - Head LR:
1e-4โ3e-4 - Weight decay:
0.01โ0.05
- Body LR:
- Scheduler: CosineAnnealingLR
- Epochs: ~20โ50 (adjust based on your actual training)
- Hardware: 1x GPU (specify if desired: V100, A100, etc.)
Update these values with the exact numbers from your experiment.
Evaluation
- Metric: Top-1 accuracy
- Split: Stanford 40 Actions test
- Result: 91.5%
| Split | Metric | Value |
|---|---|---|
| Test | Accuracy | 0.915 |
Intended uses & limitations
Intended uses
- Research in action recognition on still images
- Baseline / feature extractor for computer vision experiments
- Educational use (example of ViT fine-tuning on a human actions dataset)
Limitations
- Trained only on Stanford 40 Actions (limited context/cultural diversity).
- Assumes a single primary action per image.
- Not recommended for use in sensitive scenarios (surveillance, critical decision-making, etc.) without careful evaluation.
Citation
If you use this model, please cite the dataset:
@INPROCEEDINGS{6126386,
author={Yao, Bangpeng and Jiang, Xiaoye and Khosla, Aditya and Lin, Andy Lai and Guibas, Leonidas and Fei-Fei, Li},
booktitle={2011 International Conference on Computer Vision},
title={Human action recognition by learning bases of action attributes and parts},
year={2011},
volume={},
number={},
pages={1331-1338},
keywords={Humans;Detectors;Image reconstruction;Vectors;Feature extraction;Noise;Image recognition},
doi={10.1109/ICCV.2011.6126386}
}
And this model's page on Hugging Face:
@misc{vit_b16_stanford40_actions,
title={ViT-B/16 fine-tuned on Stanford 40 Actions},
author={Lucas Dantas de Moura Carvalho},
year={2025},
howpublished={\url{https://huggingface.co/lucasddmc/vit-b16-stanford40-actions}}
}
- Downloads last month
- 8
Model tree for lucasddmc/vit-b16-stanford40-actions
Base model
google/vit-base-patch16-224-in21kEvaluation results
- Accuracy on Stanford 40 Actionstest set self-reported0.915