Research data require context: Only through good organization and documentation do they become understandable, traceable, and reusable in the long term. Documentation covers both the data themselves and their creation, processing, and structure.

On this page you will learn about:
- Why document data?
- What belongs in data documentation?
- How can research data be organized effectively?
- Metadata and standards
- Tools for organization and documentation
- Documentation for publication
- Further resources and support
Why document data?
Good documentation:
- enables the long-term reuse of research data,
- creates transparency,
- facilitates collaboration,
- preserves knowledge about methods and workflows.
Without documentation, important information about variables, methods, or processing steps can be lost. As a result, research data can quickly lose their value.

The Research Data Scary Tale humorously illustrates why careful documentation of research data is essential. Although documenting data requires time, it is far more time-consuming to prepare poorly documented data that are several years old.
Go to the full Research Data Scary Tale
Use of the illustration with kind permission from Sandruschka.
What belongs in data documentation?
Data documentation describes the context, structure, and processing of research data. Among other things, it answers the following questions:
- Data creation
- Who created the data?
- When and where were the data collected or generated?
- Under what conditions were they created?
- Methods and processing
- Which methods, measuring instruments, software, or models were used?
- Which processing steps or analyses were carried out?
- Structure and description
- Which file formats, variables, and units are used?
- Which file formats, variables, and units are used?
- How are files and folders organized?
- Access and use
- Which access rights apply?
- Are there restrictions or data protection requirements?
- Under which license can the data be used?
Forms of documentation
Documentation can take various forms:
- ReadMe files
Describe a dataset or project and provide information about content, structure, file formats, variables, usage, and special features. Templates and recommendations help to record the most important information in a structured way (e.g., the ReadMe Template for Data and ReadMe Template for Software provided by Cornell University). - Laboratory notebooks and protocols
Document the progress of experiments, work steps, and observations. At TUHH, a shared electronic laboratory notebook (ELN) based on openBIS is currently being established. - Codebooks and data dictionaries
Describe the structure of datasets, including variables, value ranges, units, and coding schemes. - Machine-readable metadata
Describe research data in standardized formats so that they can be processed by repositories and information systems.
How can research data be organized effectively?
A structured organization makes daily work easier and supports later reuse. The following practices have proven effective:
- consistent file names and folder structures,
- separate storage of raw, processed, and result data,
- traceable versioning,
- joint management of data and scripts,
- continuous maintenance of project documentation.
Practical tips:
- The 5S Data Method provides guidance for the structured organization of research data and helps to systematically sort, structure, and prepare data for reuse.
- ReadMe files placed at the top level of a project folder provide initial guidance and make it easier to understand the data structure of a project.
Metadata and standards
Metadata describe research data in a structured and machine-readable way. They improve the discoverability, comprehensibility, and reusability of research data. Which information should be documented using standardized metadata depends on the discipline, the type of data, and the intended publication venue.
This educational video briefly explains what metadata are and where they occur in the research data lifecycle (German, 06:58 min).
When publishing research data, repositories often specify certain metadata requirements and provide standardized input forms. Therefore, it is recommended to check the requirements of the intended repository at an early stage.
Research data management tools may also include structured metadata templates. For example, electronic laboratory notebooks (ELNs) support consistent documentation by providing templates that help capture relevant information during data collection and documentation.
Whenever possible, established discipline-specific metadata standards and controlled vocabularies should be used. Suitable standards can be found through the following resources:
- NFDI4Ing Terminology Service – Provides access to engineering ontologies and supports consistent descriptions of research data.
- FAIRsharing – Database for standards, databases, and policies.
- TIB Terminology Service – Provides access to specialized terminologies and ontologies from various scientific fields.
Tools for organization and documentation
Various open-source tools are available for organizing, documenting, and preparing research data. Some are provided centrally at TUHH, while others are used within individual institutes or by the University Library.
Electronic laboratory notebook (ELN) for documenting experiments, samples, and research processes. At TUHH, openBIS is currently being introduced as a shared ELN.
Questions regarding the current status of the openBIS implementation can be addressed to the Research Data Team at research-data@tuhh.de.
Further links:

Supports version control and collaborative work on code, scripts, and text-based documentation.
Further links:

Enables collaborative storage, sharing, and joint editing of files within project teams.
Further links:

Used for reproducible documentation of data processing, analyses, and visualizations. At TUHH, Jupyter Notebooks are already used in individual institutes and by the University Library.
If you need support using Jupyter Notebooks, please contact research-data@tuhh.de.
Further links:
- Project Jupyter
- JupyterLab Documentation
- Jupyter4NFDI – Jupyter-based services and offerings for the National Research Data Infrastructure

Open-source tool for cleaning, standardizing, and preparing research data. It supports structured preparation of data before analysis or publication.
If you need support using OpenRefine, please contact research-data@tuhh.de.
Further links:

Documentation for publication
When publishing research data, a ReadMe file often complements internal documentation. It provides an overview of the dataset and supports its understanding and reuse. Many repositories expect or recommend such a description.
Further Resources
Introduction to Data Documentation
- forschungsdaten.info: Datendokumentation – Warum, was, wie
- ELIXIR RDMkit: Documentation and Metadata
- Schmitz et al. (2019): Forschungsdaten und ihre Metadaten [Video]
- Gerlach et al. (2021): Coffee Lecture – Datendokumentation
Organizing Research Data
- Lang, K. et al. (2021). The 5S Methodology in Research Data Management
- ELIXIR RDMkit: Data Organization
- Data Carpentry: File Organization for Reproducible Research
- Demerdash et al. (2025): Data Organisation Made Easy
ReadMe Files
Contact and Support

Research Data Management Specialist
(Central RDM contact)
Email: forschungsdaten@tuhh.de
Phone: +49 40 30601-3311

Data Steward, Cluster of Excellence BlueMat
(RDM support for BlueMat projects)
Email: bluemat-rdm@tuhh.de
Phone: +49 40 30601-2647

Data Steward, CRC 1615
(RDM support for projects within CRC 1615)
Email: smart-rdm@tuhh.de
Phone: +49 40 30601-2539
Detailed information on research data can be found on the information platform forschungsdaten.info.

