Data Packaging Examples and Best Practices


These examples are from HEAL-funded studies that have submitted data to a HEAL-compliant repository. The datasets are publicly accessible, and the Principal Investigators have given the HEAL Stewards permission to link to their data packages. Some data types below do not have data package examples available yet. Examples will be added as they become available. In the meantime, general data sharing guidance materials are provided to help investigators prepare their data packages. While reviewing the examples below, look for the symbols that identify which core (✅) and additional (✔️) components each data package includes.

 

Omics Data


Data Package example:

Why this is a good example:

This SPARC dataset provides an example of how long-read sequencing data can be packaged for transparency and reuse. The data include both raw and processed files, detailed documentation of experimental methods and sequencing workflows, and clear metadata describing instruments, file types, and analysis tools.

Components of the Data Package:

✅ Data file(s)
✅ README or Summary file
✅ Variable-level Metadata documentation
✅ Repository-specific documentation
✔️ Study Protocol
✔️ Context or explanatory documents

Data Package example:

Why this is a good example:

These connected data packages demonstrate short-read sequencing data sharing through linked GEO and SRA records. The GEO submission includes both raw and processed files with clear metadata describing experimental design and analysis methods, while the SRA record provides access to the underlying sequencing reads in an open, standardized format. Together, they illustrate how coordinated repository submissions can support transparency, reproducibility, and long-term reuse.

Components of the Data Package:

✅ Data file(s)
✅ Summary or README file
✅ Variable-level Metadata documentation
✅ Repository-specific documentation
✔️ Publication Citation(s)

Genomics/Sequencing-Related Resources:

  • NIH Genomic Data Submission and Release Expectations outlines NIH expectations for timely submission, data access, and genomic dataset releases to ensure compliance with the Genomic Data Sharing Policy.
  • HEAL Stewards Guidance provides information on Genomic Data as a Sensitive Data Type.
  • GEO Submission Guidance provides step-by-step instructions for submitting functional genomics data to the Gene Expression Omnibus (GEO) repository, including file preparation, metadata, and repository-specific requirements.
  • GEO Templates offers downloadable spreadsheet templates to help researchers organize and format GEO submissions in a consistent, machine-readable structure.
  • SRA Submission Guidance explains how to prepare, validate, and submit sequencing data to the Sequence Read Archive (SRA), covering accepted file types, metadata, and submission tools.

Best Practices:

Share raw and processed data files, organized in a consistent folder structure, using open formats such as mzML or mzIdentML and accompanied by complete metadata, describing instruments, software, and analytical methods. Include version information, a descriptive README, and identifiers that link related files.

Proteomic Resources:

  • The Human Proteome Organization Proteomics Standards Initiative (HUPO-PSI) develops and maintains community-driven standards, file formats, and controlled vocabularies to support interoperable, reusable proteomics and mass-spectrometry datasets.
  • MassIVE, a HEAL-compliant repository, offers a dedicated platform to archive, browse, and re-analyze mass-spectrometry proteomics data, supporting community reuse and transparency through structured submission workflows.