Metadata Resources


The resources below provide an introduction to metadata and different metadata annotation levels, and include explanations of why researchers should annotate their studies at each level.

Other Data Standards


Overview

In addition to SLMD, VLMD, and CDEs, other data standards may enhance data interoperability. There are many types of data standards, including data models, ontologies, and controlled vocabularies, which provide a common framework for encoding information. These standards are usually maintained by formal communities of practice. They enhance interoperability and harmonization between variables and datasets from one or more different studies, thereby facilitating more robust data integration and secondary analysis. Below are some examples of data standards widely used in medical research.

Variable Standards Finder: A New Tool to Support Your Study Planning

Use the HEAL Variable Standards Finder to identify variable standards appropriate for your study and data type. Based on your answers to eight questions, the Finder highlights data required standards that may apply to your data and recommended standards to increase data interoperability.

Data Standards Example

  • PhenX Toolkit: While not a formal standard, the PhenX Toolkit is a freely available, online resource that provides standardized measurement protocols for biomedical research. By offering over 800+ rigorously vetted protocols, the PhenX Toolkit ensures that researchers collect data in a consistent and harmonized manner across studies. These protocols have been selected by domain experts to represent best practices and facilitate comparability of data across different research projects. By adopting PhenX protocols, researchers can reduce variability in data collection methods, enhance the interoperability of datasets, and enable integration of findings across studies, which is critical for advancing scientific discovery and maximizing the utility of shared data. The Toolkit is particularly useful for investigators working outside their primary expertise, providing a foundation for designing studies with high-quality, standardized data collection.
  • LOINC (Logical Observation Identifiers Names and Codes): LOINC is a controlled vocabulary and document ontology used for health measurements, observations, and documents.
  • SNOMED CT (Systematized Nomenclature of Medicine Clinical Terms): SNOMED CT is a controlled vocabulary and ontology containing comprehensive, multilingual healthcare terminology that provides a standardized way to represent clinical content, assisting to harmonize health data across various healthcare settings.
  • CDISC (Clinical Data Interchange Standards Consortium): CDISC standards, such as SDTM (Study Data Tabulation Model) and ADaM (Analysis Data Model), provide a standardized approach for representing clinical and non-clinical research data, enabling consistent data collection, sharing, and analysis. CDISC maintains its own controlled vocabulary and allows use of other vocabularies, such as SNOMED CT and LOINC.
  • MeSH (Medical Subject Headings): MeSH is a comprehensive controlled vocabulary used to index journal articles and books in the life sciences, facilitating a common language for research topics. The National Library of Medicine (NLM) maintains MeSH.

Repository Requirements

When preparing to deposit data into a repository, be aware that certain repositories may have specific metadata or data standard requirements. These may include adherence to particular metadata standards, controlled vocabularies, or data models to ensure the data's reusability and discoverability. For example, repositories like the National Institute of Mental Health’s (NIMH) Data Archive (NDA) and dbGaP require data to conform to specific formats for acceptance.

Repositories may require SLMD and/or VLMD submission as part of data deposition. They may require metadata in a format different from the format required by the HEAL Data Ecosystem. In most cases, HEAL SLMD and VLMD can simply be reformatted for repository submission.

To ensure compliance, always review the repository’s guidelines before depositing your data. This proactive step can save time and effort, ensuring your data is ready for submission and meets the standards for reuse by the broader research community.