Skip to main content

Contributing Datasets

Contributing Guidelines

Thank you for choosing to contribute to AgML!

Contributing Data

If you've found (or already have) a new dataset and you want to contribute the dataset to AgML, then the instructions below will help you format the data to the AgML standard and publish it through the Hugging Face Hub, which is how all new datasets are distributed.

Dataset Formats

Currently, we have image classification, object detection, and semantic segmentation datasets available in AgML. Every dataset is published as a Hugging Face datasets repository under the Project-AgML organization, backed by Parquet, with a dataset_info config describing its columns and splits. The column layout depends on the task:

  • Image Classification: an image column plus a label column of type ClassLabel.
  • Object Detection: an image column plus an objects column holding COCO-style bounding boxes and category IDs.
  • Semantic Segmentation: an image column plus a mask column (single-channel L-mode image).

HuggingFaceDataLoader (in agml/data/hf_loader.py) relies on these names to verify that columns were cast correctly correctly upon loading, so new datasets must follow this schema for the loader to work.

Image Classification

Build a datasets.Dataset (or DatasetDict, if you have predefined splits) with two columns:

  • image: the image, as a Hugging Face Image feature.
  • label: a ClassLabel feature naming each class (e.g. "Healthy Leaf", "Bacterial Blight").
from datasets import Dataset, Features, Image, ClassLabel

features = Features({
"image": Image(),
"label": ClassLabel(names=["label_1", "label_2"]),
})
ds = Dataset.from_dict({"image": image_paths, "label": labels}, features=features)

HuggingFaceDataLoader reads the ClassLabel names as the mapping between each label and its numerical value — you don't need to build that mapping yourself.

Object Detection

Store bounding boxes as an objects column: a dictionary (per row) with parallel arrays for the boxes and their category IDs, alongside an image column. For example:

{
"image": <PIL.Image>,
"objects": {
"bbox": [[x_center, y_center, width, height], ...],
"category": [0, 0, 1, ...],
},
}

The categories should map back to human-readable names — either via a ClassLabel feature on objects.category, or documented in the dataset card's README.md. If you're starting from a COCO-formatted annotations.json, use datasets' built-in COCO conversion utilities (or imagefolder with a metadata.jsonl) to build the objects column rather than uploading the raw COCO file directly.

Semantic Segmentation

Use an image column alongside a mask column, where the mask is a single-channel (mode: L) image of the same size as its corresponding image:

from datasets import Dataset, Features, Image

features = Features({
"image": Image(),
"mask": Image(), # mode "L": each pixel is a numerical class label, or 0 for background
})
ds = Dataset.from_dict({"image": image_paths, "mask": mask_paths}, features=features)

Contributing a Dataset

If you've found a new dataset that isn't already being used in AgML and you want to add it, there's a few things you need to do. New datasets are hosted as their own repositories on the Hugging Face Hub under the Project-AgML organization, and loaded through HuggingFaceDataLoader rather than being bundled into AgML's own zip storage.

Some Things to Check

  • Make sure each image is in the range of 0-255 in integers as opposed to 0-1 as floats. This will prevent any loss of data that could adversely affect training.
  • For a semantic segmentation dataset, ensure the masks are in a png format as opposed to jpg or other.

Structuring the Hugging Face Repository

Once the dataset is built into a datasets.Dataset/DatasetDict with the columns described above, push it to a new dataset repository on the Hub (for example, Project-AgML/apple_flower_segmentation) with ds.push_to_hub("<org-name>/<dataset name>"). This uploads the data as Parquet shards and writes the dataset_info config (features, splits, sizes) into the repository's README.md front matter automatically — you don't need to hand-write that config. Add to the generated README.md:

  • A description of the dataset, its original source, and its license. Hugging Face renders this on the dataset's page.
  • A LICENSE file, if the source dataset ships with one.

If the dataset has an augmented counterpart (e.g. extra synthetic samples), upload it as a separate config within the same repository rather than a separate repository, since HuggingFaceDataLoader selects configs by name:

ds.push_to_hub("Project-AgML/my_new_dataset", config_name="augmented")

# ... which is later loaded with:
loader = HuggingFaceDataLoader("Project-AgML/my_new_dataset", "augmented")

Registering the Dataset

Once the dataset is live on the Hub, it needs to be registered in this website's catalog so it shows up in the dataset search and leaderboard pages. This repository (project-agml.github.io) stores that catalog as a JSON array at static/data/hf_datasets.json. Add a new object to that array for your dataset:

{
"name": "my_new_dataset",
"source": "huggingface",
"hf_link": "https://huggingface.co/datasets/Project-AgML/my_new_dataset",
"machine_learning_task": "object_detection",
"agricultural_task": "fruit_detection",
"location": ["United States"],
"environment": "field",
"crop_types": ["apple"],
"sensor_modality": "rgb",
"real_or_synthetic": "real",
"platform": "ground",
"input_data_format": "image_folder",
"annotation_format": "boundingBox",
"num_images": 1200,
"documentation": "my_new_dataset",
"examples_image_url": "/img/agml/sample_images/my_new_dataset_sample.webp",
"license": "cc-by-4.0",
"citation": null,
"parent_dataset": null,
"zip_size_bytes": 214000000,
"stats_mean": null,
"stats_std": null
}

The table below describes every field the catalog reads (see the Dataset type in src/lib/datasets.ts). name is the only field that is strictly required for an entry to load at all, but the search, filter, and detail views depend on the rest — leave a field null rather than omitting it if the value is genuinely unknown.

FieldRequiredDescription
nameYesThe dataset's identifier. Must match the Hugging Face repo name after Project-AgML/.
sourceYesSet to "huggingface" for datasets hosted on the Hub.
hf_linkYesFull URL to the dataset's page on the Hugging Face Hub.
machine_learning_taskYesOne of image_classification, object_detection, semantic_segmentation (see table below).
agricultural_taskRecommendedThe agricultural task the dataset supports (e.g. disease_classification, fruit_detection). Keep it broad but agriculture-specific — see the guidance below.
locationRecommendedArray of countries the data was collected in, e.g. ["Italy"]. Use ["worldwide"] if collected across many countries or unknown.
environmentRecommendedfield, greenhouse, lab, or similar.
crop_typesRecommendedArray of crop names covered by the dataset, lowercase (e.g. ["grapes"]).
sensor_modalityRecommendedUsually rgb, but can include others such as lidar, multispectral.
real_or_syntheticRecommendedreal or synthetic.
platformRecommendedHow the images were captured, e.g. handheld, ground, drone.
input_data_formatYesSee the table below.
annotation_formatYesSee the table below.
num_imagesYesTotal number of images in the dataset. Drives the "Labeled images" stat on the homepage.
documentationRecommendedA short doc slug/description, or a URL if the clearest documentation lives elsewhere.
examples_image_urlRecommendedPath to a sample-image thumbnail under static/img/agml/sample_images/, used in dataset cards and the homepage marquee.
licenseYesThe dataset's license identifier (e.g. cc-by-4.0). Leave null only if truly unlicensed.
citationRecommendedThe paper or library to cite. null if none exists.
parent_datasetOnly for variantsSet to the parent's name if this entry is an augmented/derived variant of another catalog entry.
zip_size_bytesOptionalApproximate size of the dataset, used for download-size display.
stats_mean / stats_stdOptionalPer-channel normalization statistics, if precomputed.
augmented_num_images / augmented_zip_size_bytesOnly for variantsImage count / size of the augmented counterpart, if applicable.

machine_learning_task, input_data_format, and annotation_format

These three fields are coupled and should be taken from the same row:

Dataset Formatmachine_learning_taskinput_data_formatannotation_formatHF column
Image Classificationimage_classificationimage_folderclassLabellabel
Object Detectionobject_detectionimage_folderboundingBoxobjects
Semantic Segmentationsemantic_segmentationimage_foldersegmentationMaskmask

input_data_format is image_folder for all Hugging Face-hosted datasets — it refers to how the loader accesses the images (via the Hub, regardless of the underlying Parquet storage), not a literal local folder.

agricultural_task

This field is more broadly defined than machine_learning_task — it should describe the main agricultural task the dataset supports. For instance, any of the *_leaf_disease_classification datasets are associated with the disease_classification task, while a fruit-segmentation dataset would use fruit_segmentation. Browse static/data/hf_datasets.json for existing examples of valid agricultural_task values before inventing a new one. Keep it broad enough to describe the dataset in general (e.g. fruit rather than a specific cultivar), but not as broad as machine_learning_task — it should always carry an agricultural component.

Note: If there is no explicit documentation for the dataset, then reach out to the AgML team regarding what you should put. It is important that we have references to as many datasets as possible, to allow users to acquire raw data as they desire.

Opening the Pull Request

Once the dataset is uploaded to the Hub and its entry is added to static/data/hf_datasets.json, open a pull request against the project-agml.github.io repository with the JSON change (and the sample-image asset, if you added one). We'll review the entry and confirm the Hugging Face repository is publicly accessible before merging.

Quality Checks

When contributing a dataset, you should abide by the following guidelines to ensure compatibility with AgML and ensure that there are no problems for users who are working with the datasets:

  • Check that the dataset can be properly downloaded and loaded. It is best to instantiate a HuggingFaceDataLoader for the dataset and inspect a few samples to validate that the images and annotations are in the right format.
  • Confirm the entry you added to static/data/hf_datasets.json is valid JSON and that num_images, machine_learning_task, input_data_format, and annotation_format all match the uploaded data.
  • Confirm the dataset appears correctly in the dataset search page (/datasets) after building the site locally, including its sample image and metadata modal.