This resource offers guidance for preparing variable-level metadata (often in the form of a data dictionary) to support clarity, consistency, machine processing, reuse, and alignment with the variable-level metadata schema used by the HEAL Data Ecosystem. Variable-level metadata (VLMD) is a core component of a complete HEAL data package and, along with key supporting documentation, helps others understand how your research defines, measures, and encodes variables for reuse and analysis. Following these practices can also help ensure your file is ready for use with the Platform’s VLMD tool, enabling extraction of HEAL-compliant VLMD and validation against the HEAL VLMD schema.
Each variable should be clearly defined, self-contained, and understandable without prior study knowledge. Well-documented variables improve interpretability, support reuse, and reduce ambiguity for both humans and machines.
Variables that represent similar concepts should follow consistent naming, structure, and encoding patterns. Consistency improves interpretability, supports cross-variable comparisons, and enables more efficient data harmonization across studies and instruments.
Data dictionaries should be structured in a simple, consistent, and machine-readable format to support automated processing, validation, and reuse. Clean structure reduces parsing errors, enables tools (like the HEAL VLMD tool) to interpret data reliably, and ensures compatibility across systems and workflows.
Data dictionaries should accurately reflect the original source data and preserve each variable’s full meaning. Maintaining source system fidelity reduces errors introduced through manual handling, ensures completeness, and supports reliable interpretation and reuse.