Contributing Datasets
Contributing Guidelines
Thank you for choosing to contribute to AgML!
Contributing Data
If you've found (or already have) a new dataset and you want to contribute the dataset to AgML, then the instructions below will help you format the data to the AgML standard and publish it through the Hugging Face Hub, which is how all new datasets are distributed.
Dataset Formats
Currently, we have image classification, object detection, and semantic segmentation datasets available
in AgML. Every dataset is published as a Hugging Face datasets repository
under the Project-AgML organization, backed by Parquet, with a dataset_info config describing its columns and
splits. The column layout depends on the task:
- Image Classification: an
imagecolumn plus alabelcolumn of typeClassLabel. - Object Detection: an
imagecolumn plus anobjectscolumn holding COCO-style bounding boxes and category IDs. - Semantic Segmentation: an
imagecolumn plus amaskcolumn (single-channelL-mode image).
HuggingFaceDataLoader (in agml/data/hf_loader.py) relies on these names to verify that columns were cast correctly correctly upon loading, so new datasets must follow this schema
for the loader to work.
Image Classification
Build a datasets.Dataset (or DatasetDict, if you have predefined splits) with two columns:
image: the image, as a Hugging FaceImagefeature.label: aClassLabelfeature naming each class (e.g."Healthy Leaf","Bacterial Blight").
from datasets import Dataset, Features, Image, ClassLabel
features = Features({
"image": Image(),
"label": ClassLabel(names=["label_1", "label_2"]),
})
ds = Dataset.from_dict({"image": image_paths, "label": labels}, features=features)
HuggingFaceDataLoader reads the ClassLabel names as the mapping between each label and its numerical value —
you don't need to build that mapping yourself.
Object Detection
Store bounding boxes as an objects column: a dictionary (per row) with parallel arrays for the boxes and their
category IDs, alongside an image column. For example:
{
"image": <PIL.Image>,
"objects": {
"bbox": [[x_center, y_center, width, height], ...],
"category": [0, 0, 1, ...],
},
}
The categories should map back to human-readable names — either via a ClassLabel feature on objects.category,
or documented in the dataset card's README.md. If you're starting from a COCO-formatted annotations.json, use
datasets' built-in COCO conversion utilities (or imagefolder with a metadata.jsonl) to build the objects
column rather than uploading the raw COCO file directly.
Semantic Segmentation
Use an image column alongside a mask column, where the mask is a single-channel (mode: L) image of the same
size as its corresponding image:
from datasets import Dataset, Features, Image
features = Features({
"image": Image(),
"mask": Image(), # mode "L": each pixel is a numerical class label, or 0 for background
})
ds = Dataset.from_dict({"image": image_paths, "mask": mask_paths}, features=features)
Contributing a Dataset
If you've found a new dataset that isn't already being used in AgML and you want to add it, there's a few things you
need to do. New datasets are hosted as their own repositories on the Hugging Face Hub
under the Project-AgML organization, and loaded through HuggingFaceDataLoader rather than being bundled into AgML's
own zip storage.
Some Things to Check
- Make sure each image is in the range of 0-255 in integers as opposed to 0-1 as floats. This will prevent any loss of data that could adversely affect training.
- For a semantic segmentation dataset, ensure the masks are in a
pngformat as opposed tojpgor other.
Structuring the Hugging Face Repository
Once the dataset is built into a datasets.Dataset/DatasetDict with the columns described above, push it to a new
dataset repository on the Hub (for example,
Project-AgML/apple_flower_segmentation) with ds.push_to_hub("<org-name>/<dataset name>"). This uploads the data
as Parquet shards and writes the dataset_info config (features, splits, sizes) into the repository's README.md
front matter automatically — you don't need to hand-write that config. Add to the generated README.md:
- A description of the dataset, its original source, and its license. Hugging Face renders this on the dataset's page.
- A
LICENSEfile, if the source dataset ships with one.
If the dataset has an augmented counterpart (e.g. extra synthetic samples), upload it as a separate config within the
same repository rather than a separate repository, since HuggingFaceDataLoader selects configs by name:
ds.push_to_hub("Project-AgML/my_new_dataset", config_name="augmented")
# ... which is later loaded with:
loader = HuggingFaceDataLoader("Project-AgML/my_new_dataset", "augmented")
Registering the Dataset
Once the dataset is live on the Hub, it needs to be registered in this website's catalog so it shows up in the
dataset search and leaderboard pages. This repository (project-agml.github.io) stores that catalog as a JSON array
at static/data/hf_datasets.json.
Add a new object to that array for your dataset:
{
"name": "my_new_dataset",
"source": "huggingface",
"hf_link": "https://huggingface.co/datasets/Project-AgML/my_new_dataset",
"machine_learning_task": "object_detection",
"agricultural_task": "fruit_detection",
"location": ["United States"],
"environment": "field",
"crop_types": ["apple"],
"sensor_modality": "rgb",
"real_or_synthetic": "real",
"platform": "ground",
"input_data_format": "image_folder",
"annotation_format": "boundingBox",
"num_images": 1200,
"documentation": "my_new_dataset",
"examples_image_url": "/img/agml/sample_images/my_new_dataset_sample.webp",
"license": "cc-by-4.0",
"citation": null,
"parent_dataset": null,
"zip_size_bytes": 214000000,
"stats_mean": null,
"stats_std": null
}
The table below describes every field the catalog reads (see the Dataset type in
src/lib/datasets.ts).
name is the only field that is strictly required for an entry to load at all, but the search, filter, and detail
views depend on the rest — leave a field null rather than omitting it if the value is genuinely unknown.
| Field | Required | Description |
|---|---|---|
name | Yes | The dataset's identifier. Must match the Hugging Face repo name after Project-AgML/. |
source | Yes | Set to "huggingface" for datasets hosted on the Hub. |
hf_link | Yes | Full URL to the dataset's page on the Hugging Face Hub. |
machine_learning_task | Yes | One of image_classification, object_detection, semantic_segmentation (see table below). |
agricultural_task | Recommended | The agricultural task the dataset supports (e.g. disease_classification, fruit_detection). Keep it broad but agriculture-specific — see the guidance below. |
location | Recommended | Array of countries the data was collected in, e.g. ["Italy"]. Use ["worldwide"] if collected across many countries or unknown. |
environment | Recommended | field, greenhouse, lab, or similar. |
crop_types | Recommended | Array of crop names covered by the dataset, lowercase (e.g. ["grapes"]). |
sensor_modality | Recommended | Usually rgb, but can include others such as lidar, multispectral. |
real_or_synthetic | Recommended | real or synthetic. |
platform | Recommended | How the images were captured, e.g. handheld, ground, drone. |
input_data_format | Yes | See the table below. |
annotation_format | Yes | See the table below. |
num_images | Yes | Total number of images in the dataset. Drives the "Labeled images" stat on the homepage. |
documentation | Recommended | A short doc slug/description, or a URL if the clearest documentation lives elsewhere. |
examples_image_url | Recommended | Path to a sample-image thumbnail under static/img/agml/sample_images/, used in dataset cards and the homepage marquee. |
license | Yes | The dataset's license identifier (e.g. cc-by-4.0). Leave null only if truly unlicensed. |
citation | Recommended | The paper or library to cite. null if none exists. |
parent_dataset | Only for variants | Set to the parent's name if this entry is an augmented/derived variant of another catalog entry. |
zip_size_bytes | Optional | Approximate size of the dataset, used for download-size display. |
stats_mean / stats_std | Optional | Per-channel normalization statistics, if precomputed. |
augmented_num_images / augmented_zip_size_bytes | Only for variants | Image count / size of the augmented counterpart, if applicable. |
machine_learning_task, input_data_format, and annotation_format
These three fields are coupled and should be taken from the same row:
| Dataset Format | machine_learning_task | input_data_format | annotation_format | HF column |
|---|---|---|---|---|
| Image Classification | image_classification | image_folder | classLabel | label |
| Object Detection | object_detection | image_folder | boundingBox | objects |
| Semantic Segmentation | semantic_segmentation | image_folder | segmentationMask | mask |
input_data_format is image_folder for all Hugging Face-hosted datasets — it refers to how the loader accesses the
images (via the Hub, regardless of the underlying Parquet storage), not a literal local folder.
agricultural_task
This field is more broadly defined than machine_learning_task — it should describe the main agricultural task the
dataset supports. For instance, any of the *_leaf_disease_classification datasets are associated with the
disease_classification task, while a fruit-segmentation dataset would use fruit_segmentation. Browse
static/data/hf_datasets.json for existing examples of valid agricultural_task values before inventing a new one.
Keep it broad enough to describe the dataset in general (e.g. fruit rather than a specific cultivar), but not as
broad as machine_learning_task — it should always carry an agricultural component.
Note: If there is no explicit documentation for the dataset, then reach out to the AgML team regarding what you should put. It is important that we have references to as many datasets as possible, to allow users to acquire raw data as they desire.
Opening the Pull Request
Once the dataset is uploaded to the Hub and its entry is added to static/data/hf_datasets.json, open a pull request
against the project-agml.github.io repository with the JSON change (and the sample-image asset, if you added one).
We'll review the entry and confirm the Hugging Face repository is publicly accessible before merging.
Quality Checks
When contributing a dataset, you should abide by the following guidelines to ensure compatibility with AgML and ensure that there are no problems for users who are working with the datasets:
- Check that the dataset can be properly downloaded and loaded. It is best to instantiate a
HuggingFaceDataLoaderfor the dataset and inspect a few samples to validate that the images and annotations are in the right format. - Confirm the entry you added to
static/data/hf_datasets.jsonis valid JSON and thatnum_images,machine_learning_task,input_data_format, andannotation_formatall match the uploaded data. - Confirm the dataset appears correctly in the dataset search page (
/datasets) after building the site locally, including its sample image and metadata modal.