Data
As of 7 August 2026, we have processed 2465 original data sets containing a total of 3,220,958 records. The map below shows all locations for which we have at least one observation, one thematic group at a time.
For ease of organization we divide the data into thematic groups. These are not mutually exclusive. For example, the first place to look for crop response to fertilizer data would be in the “agronomy” group. However, the “survey”, and “varieties” groups may also contain fertilizer application data. Likewise, the “varieties” group has data for comparing crop varieties, but variety names are also reported in the “agronomy” group. This means that you may want to consider using data from multiple groups.
The table below shows the current groups and the number of original datasets and records in each group. We also show these numbers for the datasets that have a Creative Commons (CC) license.
| Group | Datasets | Records | CC-Datasets | CC-Records |
|---|---|---|---|---|
| agronomy | 310 | 474 107 | 237 | 308 355 |
| pest_disease | 8 | 3 225 | 6 | 2 593 |
| soil_samples | 38 | 239 484 | 14 | 228 462 |
| survey | 94 | 545 153 | 64 | 446 197 |
| varieties | 111 | 246 839 | 107 | 245 863 |
| varieties_cassava | 1467 | 229 179 | 1467 | 229 173 |
| varieties_cowpea | 76 | 23 193 | 76 | 23 193 |
| varieties_maize | 79 | 79 527 | 62 | 69 703 |
| varieties_potato | 57 | 30 830 | 57 | 30 830 |
| varieties_wheat | 225 | 1 349 421 | 4 | 19 234 |
Below, you can download the compiled standardized data that come with a Creative Commons license. You can create the full datasets yourself by following these instructions.
You can download data by group, or, if you want all available data, select “everything”. If you want data for a single data set, you can find these here. You can use R package caramba to integrate the data download into an R workflow.
Please note that we have currently only partially processed much of the survey data, and the original data sources may contains many more variables. The data available here are our first attempt to standardize widely variable data with lots of data quality issues. The data still contain errors from the original data that remain, and likely also errors that we have introduced.