Home » Publishing » Research Data » Organizing and Documenting Research Data

Organizing and Documenting Research Data

Research data require context: Only through good organization and documentation do they become understandable, traceable, and reusable in the long term. Documentation covers both the data themselves and their creation, processing, and structure.

Source: pixabay.com. Free for use under the Pixabay Content License

On this page you will learn about:

Why document data?

Good documentation:

  • enables the long-term reuse of research data,
  • creates transparency,
  • facilitates collaboration,
  • preserves knowledge about methods and workflows.

Without documentation, important information about variables, methods, or processing steps can be lost. As a result, research data can quickly lose their value.

The Research Data Scary Tale humorously illustrates why careful documentation of research data is essential. Although documenting data requires time, it is far more time-consuming to prepare poorly documented data that are several years old.

Go to the full Research Data Scary Tale

Use of the illustration with kind permission from Sandruschka.

What belongs in data documentation?

Data documentation describes the context, structure, and processing of research data. Among other things, it answers the following questions:

  • Data creation
    • Who created the data?
    • When and where were the data collected or generated?
    • Under what conditions were they created?
  • Methods and processing
    • Which methods, measuring instruments, software, or models were used?
    • Which processing steps or analyses were carried out?
  • Structure and description
    • Which file formats, variables, and units are used?
    • Which file formats, variables, and units are used?
    • How are files and folders organized?
  • Access and use
    • Which access rights apply?
    • Are there restrictions or data protection requirements?
    • Under which license can the data be used?

Forms of documentation

Documentation can take various forms:

  • ReadMe files
    Describe a dataset or project and provide information about content, structure, file formats, variables, usage, and special features. Templates and recommendations help to record the most important information in a structured way (e.g., the ReadMe Template for Data and ReadMe Template for Software provided by Cornell University).
  • Laboratory notebooks and protocols
    Document the progress of experiments, work steps, and observations. At TUHH, a shared electronic laboratory notebook (ELN) based on openBIS is currently being established.
  • Codebooks and data dictionaries
    Describe the structure of datasets, including variables, value ranges, units, and coding schemes.
  • Machine-readable metadata
    Describe research data in standardized formats so that they can be processed by repositories and information systems.

How can research data be organized effectively?

A structured organization makes daily work easier and supports later reuse. The following practices have proven effective:

  • consistent file names and folder structures,
  • separate storage of raw, processed, and result data,
  • traceable versioning,
  • joint management of data and scripts,
  • continuous maintenance of project documentation.

Practical tips:

  • The 5S Data Method provides guidance for the structured organization of research data and helps to systematically sort, structure, and prepare data for reuse.
  • ReadMe files placed at the top level of a project folder provide initial guidance and make it easier to understand the data structure of a project.

Metadata and standards

Metadata describe research data in a structured and machine-readable way. They improve the discoverability, comprehensibility, and reusability of research data. Which information should be documented using standardized metadata depends on the discipline, the type of data, and the intended publication venue.

This educational video briefly explains what metadata are and where they occur in the research data lifecycle (German, 06:58 min).

When publishing research data, repositories often specify certain metadata requirements and provide standardized input forms. Therefore, it is recommended to check the requirements of the intended repository at an early stage.

Research data management tools may also include structured metadata templates. For example, electronic laboratory notebooks (ELNs) support consistent documentation by providing templates that help capture relevant information during data collection and documentation.

Whenever possible, established discipline-specific metadata standards and controlled vocabularies should be used. Suitable standards can be found through the following resources:

  • NFDI4Ing Terminology Service – Provides access to engineering ontologies and supports consistent descriptions of research data.
  • FAIRsharing – Database for standards, databases, and policies.
  • TIB Terminology Service – Provides access to specialized terminologies and ontologies from various scientific fields.

Tools for organization and documentation

Various open-source tools are available for organizing, documenting, and preparing research data. Some are provided centrally at TUHH, while others are used within individual institutes or by the University Library.

Electronic laboratory notebook (ELN) for documenting experiments, samples, and research processes. At TUHH, openBIS is currently being introduced as a shared ELN.

Questions regarding the current status of the openBIS implementation can be addressed to the Research Data Team at research-data@tuhh.de.

Further links:

Supports version control and collaborative work on code, scripts, and text-based documentation.

Further links:

Enables collaborative storage, sharing, and joint editing of files within project teams.

Further links:

Used for reproducible documentation of data processing, analyses, and visualizations. At TUHH, Jupyter Notebooks are already used in individual institutes and by the University Library.

If you need support using Jupyter Notebooks, please contact research-data@tuhh.de.

Further links:

Open-source tool for cleaning, standardizing, and preparing research data. It supports structured preparation of data before analysis or publication.

If you need support using OpenRefine, please contact research-data@tuhh.de.

Further links:

Documentation for publication

When publishing research data, a ReadMe file often complements internal documentation. It provides an overview of the dataset and supports its understanding and reuse. Many repositories expect or recommend such a description.

Contact and Support

Francesca Schulze

Francesca Schulze

Research Data Management Specialist
(Central RDM contact)

Email: forschungsdaten@tuhh.de
Phone: +49 40 30601-3311

Dr.-Ing. Seoyun Sohn, Data Steward, Cluster of Excellence BlueMat

Dr.-Ing. Seoyun Sohn

Data Steward, Cluster of Excellence BlueMat
(RDM support for BlueMat projects)

Email:  bluemat-rdm@tuhh.de
Phone: +49 40 30601-2647

Nimet Eylem Tas

Data Steward, CRC 1615
(RDM support for projects within CRC 1615)

Email: smart-rdm@tuhh.de
Phone: +49 40 30601-2539


Detailed information on research data can be found on the information platform forschungsdaten.info.

Scroll to Top