Data Sharing Throughout the Research Lifecycle


Decisions and activities at each stage of the study’s lifecycle impact data sharing. Select a tab to learn more about key topics in a lifecycle stage, why they matter, and what actions you can take.

Share


Storing data in a trusted research repository supports reproducibility, access, and reuse. HEAL studies are expected to share data through one or more of the 29 HEAL-compliant data repositories, evaluated for HEAL-funded data. Note: the HEAL Data Platform is a catalog that links to HEAL data in compliant repositories; it is not a data repository.

Lessons learned: Repositories provide a range of services to researchers during submission; review protocols and available curation support before submitting data. Use different repositories for different types of study data (e.g. one repository for sequence or imaging data, and another for code/scripts) if needed. Include VLMD with your data. Your organization may have membership benefits with some data repositories.

What to do:

Additional resources:

According to the NNLM, “Data curation is composed of research data management and digital preservation and involves processes such as adding metadata to make data more findable and understandable, ingesting data into a [data] repository, … validating file checksums and file fixity checks, and other tasks for organizing, cleaning, describing, enhancing, storing, and preserving data.” Data curation transforms data used by the study team into forms appropriate for external sharing and long-term preservation.

Lessons learned: Under the NIH DMS Policy and HEAL Public Access and Data Sharing policy, HEAL studies must appropriately share scientific data underlying research findings, but not all raw data needs to be shared. Curation often involves HEAL studies transforming data (e.g. de-identification) and generating metadata. Starting early reduces the workload later.

What to do:

  • Determine what data to share, including data supporting publications, high-value data, or data required by the NIH HEAL Initiative.
  • Prepare data, determining if any data should be held back for privacy/security reasons, de-identifying/anonymizing data, generating or enhancing SLMD and VLMD, and reformatting data into open formats (e.g., CSV, PDF, PNG).
  • Choose an appropriate license for your shared data. Creative Commons licenses commonly define terms of use for published datasets.
  • Consult a data curator at your organization or data repository. Some will curate or de-identify data at cost, while others offer guidance.

Additional resources:

Open access allows anyone to freely access and reuse shared data (aka public access). Controlled and restricted access limit findability, accessibility, and reusability. Available access controls vary across data repositories. Data licenses or contracts can define allowable uses of shared data.

Lessons learned: HEAL studies may use open, controlled, or restricted access approaches. For sensitive data, consider repositories with secure platform protections and controlled access options.

What to do:

  • Decide on the best access option for your shared data (open, controlled, or restricted).
  • Choose a data repository with adequate access controls, such as access request mechanisms, enclaves (virtual/physical), temporary access/”visiting”, security controls, or others.
  • Define restrictions and establish licenses (Creative Commons for open access; Terms of service, Terms of use, DUAs or DSAs for controlled access), and define requirements around IP, citations, and allowable uses.

Additional resources:

Publishing research data with journal articles supports replication and re-use and is often required by the publisher. It can also boost citation counts and enhance research impact. Persistent identifiers (PIDs) allow research artifacts to be linked and referenced across different locations, promoting findability.

Lessons learned: HEAL policy expects that “Underlying Primary Data for the Publications will be made broadly available through an appropriate data repository.” “Available upon request” statements are not HEAL compliant.

What to do:

  • Assign persistent identifiers, such as DOIs (Digital Object Identifiers), to datasets, publications, code, and other research outputs to ensure long-term findability and reference.
  • Include a publication data availability statement that points to the repository data deposit. For example: "The data that support the findings of this study are available from [repository name] with the following DOI: [DOI].”
  • Include associated publication DOIs with repository metadata.
  • Use data embargos if needed, before the publication is released. An embargo delays data visibility to others. Be sure to lift the embargo at the end of the performance period for HEAL compliance.

Sharing research data fosters collaboration and knowledge-building, cross-disciplinary discoveries, and public health progress. It enhances researcher visibility, supports career advancement, and increases publication/citation opportunities.

Lessons learned: Repositories often track dataset reuse metrics (views, downloads, citations), which may support tenure and other promotion considerations. Studies show that publications linked to repository-hosted data are up to 25% more likely to be cited.

What to do:

  • Link published data to your ORCID or researcher profile and include your ORCID in the repository deposit metadata.
  • Clearly specify copyright and/or citation requirements to ensure correct attribution upon re-use.
  • Measure your research impact, leveraging persistent identifiers and dataset reuse metrics to track data access and use.

Additional resources: