egrace479 commited on
Commit
74129d1
·
verified ·
1 Parent(s): ce901ab

initial model card update

Browse files

set for 50 epoch change, eval needs to be updated

Files changed (1) hide show
  1. README.md +26 -16
README.md CHANGED
@@ -5,7 +5,7 @@ language:
5
  - en
6
  library_name: open_clip
7
  model_name: BioCLIP
8
- model_description: "Foundation model for the tree of life, built using CLIP architecture as a vision model for general organismal biology. It is trained on TreeOfLife-10M, our specially-created dataset covering over 450K taxa--the most biologically diverse ML-ready dataset available to date."
9
  tags:
10
  - zero-shot-image-classification
11
  - clip
@@ -13,6 +13,8 @@ tags:
13
  - CV
14
  - images
15
  - animals
 
 
16
  - species
17
  - taxonomy
18
  - rare species
@@ -30,11 +32,14 @@ datasets:
30
 
31
  # Model Card for BioCLIP
32
 
 
 
 
33
  <!--
34
  This modelcard has been generated using [this raw template](https://github.com/huggingface/huggingface_hub/blob/main/src/huggingface_hub/templates/modelcard_template.md?plain=1). And further altered to suit Imageomics Institute needs -->
35
 
36
  BioCLIP is a foundation model for the tree of life, built using CLIP architecture as a vision model for general organismal biology.
37
- It is trained on [TreeOfLife-10M](https://huggingface.co/datasets/imageomics/TreeOfLife-10M), our specially-created dataset covering over 450K taxa--the most biologically diverse ML-ready dataset available to date.
38
  Through rigorous benchmarking on a diverse set of fine-grained biological classification tasks, BioCLIP consistently outperformed existing baselines by 16% to 17% absolute.
39
  Through intrinsic evaluation, we found that BioCLIP learned a hierarchical representation aligned to the tree of life, which demonstrates its potential for robust generalizability.
40
 
@@ -47,7 +52,7 @@ Through intrinsic evaluation, we found that BioCLIP learned a hierarchical repre
47
  BioCLIP is based on OpenAI's [CLIP](https://openai.com/research/clip).
48
  We trained the model on [TreeOfLife-10M](https://huggingface.co/datasets/imageomics/TreeOfLife-10M) from OpenAI's ViT-B/16 checkpoint, using [OpenCLIP's](https://github.com/mlfoundations/open_clip) code.
49
  BioCLIP is trained with the standard CLIP objective to imbue the model with an understanding, not just of different species, but of the hierarchical structure that relates species across the tree of life.
50
- In this way, BioCLIP offers potential to aid biologists in discovery of new and related creatures, since it does not see the 454K different taxa as distinct classes, but as part of an interconnected hierarchy.
51
 
52
 
53
  - **Developed by:** Samuel Stevens, Jiaman Wu, Matthew J. Thompson, Elizabeth G. Campolongo, Chan Hee Song, David Edward Carlyn, Li Dong, Wasila M. Dahdul, Charles Stewart, Tanya Berger-Wolf, Wei-Lun Chao, and Yu Su
@@ -60,8 +65,9 @@ This model was developed for the benefit of the community as an open-source prod
60
  ### Model Sources
61
 
62
  - **Repository:** [BioCLIP](https://github.com/Imageomics/BioCLIP)
63
- - **Paper:** BioCLIP: A Vision Foundation Model for the Tree of Life ([arXiv](https://doi.org/10.48550/arXiv.2311.18803))
64
  - **Demo:** [BioCLIP Demo](https://huggingface.co/spaces/imageomics/bioclip-demo)
 
65
 
66
  ## Uses
67
 
@@ -72,7 +78,7 @@ The ViT-B/16 vision encoder is recommended as a base model for any computer visi
72
  ### Direct Use
73
 
74
  See the demo [here](https://huggingface.co/spaces/imageomics/bioclip-demo) for examples of zero-shot classification.
75
- It can also be used in a few-shot setting with a KNN; please see [our paper](https://doi.org/10.48550/arXiv.2311.18803) for details for both few-shot and zero-shot settings without fine-tuning.
76
 
77
 
78
  ## Bias, Risks, and Limitations
@@ -102,9 +108,9 @@ tokenizer = open_clip.get_tokenizer('hf-hub:imageomics/bioclip')
102
 
103
  ### Compute Infrastructure
104
 
105
- Training was performed on 8 NVIDIA A100-80GB GPUs distributed over 2 nodes on [OSC's](https://www.osc.edu/) Ascend HPC Cluster with global batch size 32,768 for 4 days.
106
 
107
- Based on [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://doi.org/10.48550/arXiv.1910.09700), that's 132.71 kg of CO<sub>2</sub> eq., or 536km driven by an average ICE car.
108
 
109
  ### Training Data
110
 
@@ -115,7 +121,7 @@ This model was trained on [TreeOfLife-10M](https://huggingface.co/datasets/image
115
  - **Training regime:** fp16 mixed precision.
116
 
117
  We resize images to 224 x 224 pixels.
118
- We use a maximum learning rate of 1e4 with 1000 linear warm-up steps, then use cosine decay to 0 over 100 epochs.
119
  We also use a weight decay of 0.2 and a batch size of 32K.
120
 
121
  ## Evaluation
@@ -127,7 +133,7 @@ We tested BioCLIP on the following collection of 10 biologically-relevant tasks.
127
  - [Birds 525](https://www.kaggle.com/datasets/gpiosenka/100-bird-species): We evaluated on the 2,625 test images provided with the dataset.
128
  - [Rare Species](https://huggingface.co/datasets/imageomics/rare-species): A new dataset we curated for the purpose of testing this model and to contribute to the ML for Conservation community. It consists of 400 species labeled Near Threatened through Extinct in the Wild by the [IUCN Red List](https://www.iucnredlist.org/), with 30 images per species. For more information, see our dataset, [Rare Species](https://huggingface.co/datasets/imageomics/rare-species).
129
 
130
- For more information about the contents of these datasets, see Table 2 and associated sections of [our paper](https://doi.org/10.48550/arXiv.2311.18803).
131
 
132
  ### Metrics
133
 
@@ -137,7 +143,7 @@ We use top-1 and top-5 accuracy to evaluate models, and validation loss to choos
137
 
138
  We compare BioCLIP to OpenAI's CLIP and OpenCLIP's LAION-2B checkpoint.
139
  Here are the zero-shot classification results on our benchmark tasks.
140
- Please see [our paper](https://doi.org/10.48550/arXiv.2311.18803) for few-shot results.
141
 
142
  <table cellpadding="0" cellspacing="0">
143
  <thead>
@@ -227,7 +233,7 @@ BioCLIP outperforms general-domain baselines by 17% on average for zero-shot.
227
 
228
  ### Model Examination
229
 
230
- We encourage readers to see Section 4.6 of [our paper](https://doi.org/10.48550/arXiv.2311.18803).
231
  In short, BioCLIP forms representations that more closely align to the taxonomic hierarchy compared to general-domain baselines like CLIP or OpenCLIP.
232
 
233
 
@@ -238,14 +244,16 @@ In short, BioCLIP forms representations that more closely align to the taxonomic
238
  ```
239
  @software{bioclip2023,
240
  author = {Samuel Stevens and Jiaman Wu and Matthew J. Thompson and Elizabeth G. Campolongo and Chan Hee Song and David Edward Carlyn and Li Dong and Wasila M. Dahdul and Charles Stewart and Tanya Berger-Wolf and Wei-Lun Chao and Yu Su},
241
- doi = {10.57967/hf/1511},
242
- month = nov,
243
  title = {BioCLIP},
244
- version = {v0.1},
245
- year = {2023}
246
  }
247
  ```
248
 
 
 
 
249
  Please also cite our paper:
250
 
251
  ```
@@ -288,7 +296,9 @@ Please also consider citing OpenCLIP, iNat21 and BIOSCAN-1M:
288
 
289
  ## Acknowledgements
290
 
291
- The authors would like to thank Josef Uyeda, Jim Balhoff, Dan Rubenstein, Hank Bart, Hilmar Lapp, Sara Beery, and colleagues from the Imageomics Institute and the OSU NLP group for their valuable feedback. We also thank the BIOSCAN-1M team and the iNaturalist team for making their data available and easy to use, and Jennifer Hammack at EOL for her invaluable help in accessing EOL’s images.
 
 
292
 
293
  The [Imageomics Institute](https://imageomics.org) is funded by the US National Science Foundation's Harnessing the Data Revolution (HDR) program under [Award #2118240](https://www.nsf.gov/awardsearch/showAward?AWD_ID=2118240) (Imageomics: A New Frontier of Biological Information Powered by Knowledge-Guided Machine Learning). Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation.
294
 
 
5
  - en
6
  library_name: open_clip
7
  model_name: BioCLIP
8
+ model_description: "Foundation model for the tree of life, built using CLIP architecture (ViT-B/16) as a vision model for general organismal biology. It is trained on TreeOfLife-10M, our specially-created dataset covering over 390K taxa--the most biologically diverse ML-ready dataset available at release."
9
  tags:
10
  - zero-shot-image-classification
11
  - clip
 
13
  - CV
14
  - images
15
  - animals
16
+ - plants
17
+ - fungi
18
  - species
19
  - taxonomy
20
  - rare species
 
32
 
33
  # Model Card for BioCLIP
34
 
35
+ If you are looking for the original BioCLIP model presented in the paper, please see [Revision 7b4abf1](https://huggingface.co/imageomics/bioclip/tree/7b4abf1f6ee747c15de00c7d28a5e62990b5dabc) (the most accurate documentation for the original version at BioCLIP [Revision ce901ab](https://huggingface.co/imageomics/bioclip/tree/ce901ab3c6a913f9e9ef94ce6d27761069f4f01c)), trained on TreeOfLife-10M [Revision ffa2a31](https://huggingface.co/datasets/imageomics/TreeOfLife-10M/tree/ffa2a318a1396f2f9e456ba171d3b5b5d8b4f051).
36
+
37
+
38
  <!--
39
  This modelcard has been generated using [this raw template](https://github.com/huggingface/huggingface_hub/blob/main/src/huggingface_hub/templates/modelcard_template.md?plain=1). And further altered to suit Imageomics Institute needs -->
40
 
41
  BioCLIP is a foundation model for the tree of life, built using CLIP architecture as a vision model for general organismal biology.
42
+ It is trained on [TreeOfLife-10M](https://huggingface.co/datasets/imageomics/TreeOfLife-10M), our specially-created dataset covering over 390K taxa--the most biologically diverse ML-ready dataset available at its release.
43
  Through rigorous benchmarking on a diverse set of fine-grained biological classification tasks, BioCLIP consistently outperformed existing baselines by 16% to 17% absolute.
44
  Through intrinsic evaluation, we found that BioCLIP learned a hierarchical representation aligned to the tree of life, which demonstrates its potential for robust generalizability.
45
 
 
52
  BioCLIP is based on OpenAI's [CLIP](https://openai.com/research/clip).
53
  We trained the model on [TreeOfLife-10M](https://huggingface.co/datasets/imageomics/TreeOfLife-10M) from OpenAI's ViT-B/16 checkpoint, using [OpenCLIP's](https://github.com/mlfoundations/open_clip) code.
54
  BioCLIP is trained with the standard CLIP objective to imbue the model with an understanding, not just of different species, but of the hierarchical structure that relates species across the tree of life.
55
+ In this way, BioCLIP offers potential to aid biologists in discovery of new and related creatures, since it does not see the 394K different taxa as distinct classes, but as part of an interconnected hierarchy.
56
 
57
 
58
  - **Developed by:** Samuel Stevens, Jiaman Wu, Matthew J. Thompson, Elizabeth G. Campolongo, Chan Hee Song, David Edward Carlyn, Li Dong, Wasila M. Dahdul, Charles Stewart, Tanya Berger-Wolf, Wei-Lun Chao, and Yu Su
 
65
  ### Model Sources
66
 
67
  - **Repository:** [BioCLIP](https://github.com/Imageomics/BioCLIP)
68
+ - **Paper:** [BioCLIP: A Vision Foundation Model for the Tree of Life](https://openaccess.thecvf.com/content/CVPR2024/papers/Stevens_BioCLIP_A_Vision_Foundation_Model_for_the_Tree_of_Life_CVPR_2024_paper.pdf)
69
  - **Demo:** [BioCLIP Demo](https://huggingface.co/spaces/imageomics/bioclip-demo)
70
+ <!-- [arXiv](https://doi.org/10.48550/arXiv.2311.18803) -->
71
 
72
  ## Uses
73
 
 
78
  ### Direct Use
79
 
80
  See the demo [here](https://huggingface.co/spaces/imageomics/bioclip-demo) for examples of zero-shot classification.
81
+ It can also be used in a few-shot setting with a KNN; please see [our paper](https://openaccess.thecvf.com/content/CVPR2024/papers/Stevens_BioCLIP_A_Vision_Foundation_Model_for_the_Tree_of_Life_CVPR_2024_paper.pdf) for details for both few-shot and zero-shot settings without fine-tuning.
82
 
83
 
84
  ## Bias, Risks, and Limitations
 
108
 
109
  ### Compute Infrastructure
110
 
111
+ Training was performed on 8 NVIDIA H100-100GB GPUs distributed over 2 nodes on [OSC's](https://www.osc.edu/) Cardinal HPC Cluster with global batch size 32,768 for 23 hours.
112
 
113
+ Based on [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://doi.org/10.48550/arXiv.1910.09700), that's 3.97 kg eq. CO<sub>2</sub> eq., or 16km driven by an average ICE car.
114
 
115
  ### Training Data
116
 
 
121
  - **Training regime:** fp16 mixed precision.
122
 
123
  We resize images to 224 x 224 pixels.
124
+ We use a maximum learning rate of 1e4 with 500 linear warm-up steps, then use cosine decay to 0 over 50 epochs.
125
  We also use a weight decay of 0.2 and a batch size of 32K.
126
 
127
  ## Evaluation
 
133
  - [Birds 525](https://www.kaggle.com/datasets/gpiosenka/100-bird-species): We evaluated on the 2,625 test images provided with the dataset.
134
  - [Rare Species](https://huggingface.co/datasets/imageomics/rare-species): A new dataset we curated for the purpose of testing this model and to contribute to the ML for Conservation community. It consists of 400 species labeled Near Threatened through Extinct in the Wild by the [IUCN Red List](https://www.iucnredlist.org/), with 30 images per species. For more information, see our dataset, [Rare Species](https://huggingface.co/datasets/imageomics/rare-species).
135
 
136
+ For more information about the contents of these datasets, see Table 2 and associated sections of [our paper](https://openaccess.thecvf.com/content/CVPR2024/papers/Stevens_BioCLIP_A_Vision_Foundation_Model_for_the_Tree_of_Life_CVPR_2024_paper.pdf).
137
 
138
  ### Metrics
139
 
 
143
 
144
  We compare BioCLIP to OpenAI's CLIP and OpenCLIP's LAION-2B checkpoint.
145
  Here are the zero-shot classification results on our benchmark tasks.
146
+ Please see [our paper](https://openaccess.thecvf.com/content/CVPR2024/papers/Stevens_BioCLIP_A_Vision_Foundation_Model_for_the_Tree_of_Life_CVPR_2024_paper.pdf) for few-shot results.
147
 
148
  <table cellpadding="0" cellspacing="0">
149
  <thead>
 
233
 
234
  ### Model Examination
235
 
236
+ We encourage readers to see Section 4.6 of [our paper](https://openaccess.thecvf.com/content/CVPR2024/papers/Stevens_BioCLIP_A_Vision_Foundation_Model_for_the_Tree_of_Life_CVPR_2024_paper.pdf).
237
  In short, BioCLIP forms representations that more closely align to the taxonomic hierarchy compared to general-domain baselines like CLIP or OpenCLIP.
238
 
239
 
 
244
  ```
245
  @software{bioclip2023,
246
  author = {Samuel Stevens and Jiaman Wu and Matthew J. Thompson and Elizabeth G. Campolongo and Chan Hee Song and David Edward Carlyn and Li Dong and Wasila M. Dahdul and Charles Stewart and Tanya Berger-Wolf and Wei-Lun Chao and Yu Su},
247
+ doi = {<update-on-generation>},
248
+ month = feb,
249
  title = {BioCLIP},
250
+ year = {2026}
 
251
  }
252
  ```
253
 
254
+ Note that this version is updated from the original BioCLIP model presented in the paper, please see [Revision 7b4abf1](https://huggingface.co/imageomics/bioclip/tree/7b4abf1f6ee747c15de00c7d28a5e62990b5dabc) (the most accurate documentation for the original version at BioCLIP [Revision ce901ab](https://huggingface.co/imageomics/bioclip/tree/ce901ab3c6a913f9e9ef94ce6d27761069f4f01c)), trained on TreeOfLife-10M [Revision ffa2a31](https://huggingface.co/datasets/imageomics/TreeOfLife-10M/tree/ffa2a318a1396f2f9e456ba171d3b5b5d8b4f051). This updated version of the model uses an updated TreeOfLife-10M which resolves taxonomic alignment issues discovered in the first version. The taxonomic resolution was completed using [TaxonoPy](https://github.com/Imageomics/TaxonoPy), which was developed for [TreeOfLife-200M](https://huggingface.co/datasets/imageomics/TreeOfLife-200M).
255
+
256
+
257
  Please also cite our paper:
258
 
259
  ```
 
296
 
297
  ## Acknowledgements
298
 
299
+ The authors would like to thank Josef Uyeda, Jim Balhoff, Dan Rubenstein, Hank Bart, Hilmar Lapp, Sara Beery, and colleagues from the Imageomics Institute and the OSU NLP group for their valuable feedback. We also thank the BIOSCAN-1M team and the iNaturalist team for making their data available and easy to use, and Jennifer Hammock at EOL for her invaluable help in accessing EOL’s images.
300
+
301
+ Additionally, we thank Ziheng Zhang and Jianyang Gu for running the training and evaluation of the revised model (with the taxonomic fixes) while they were training [BioCAP](https://huggingface.co/imageomics/biocap).
302
 
303
  The [Imageomics Institute](https://imageomics.org) is funded by the US National Science Foundation's Harnessing the Data Revolution (HDR) program under [Award #2118240](https://www.nsf.gov/awardsearch/showAward?AWD_ID=2118240) (Imageomics: A New Frontier of Biological Information Powered by Knowledge-Guided Machine Learning). Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation.
304