Home » Publishing » Research Data » Publish Research Data

Publish Research Data

Research data should – where appropriate and feasible – be published alongside the associated research findings, so that the research remains traceable and reusable. Suitable research data repositories should be used for long-term preservation and archiving. On this page, you can find out what publication options are available and how to select a suitable repository.

Verknüpfte Strukturen im Gehirn mit Knotenpunkten (Künstliche Intelligenz)
Source: pixabay.com. Free for use under the Pixabay Content License

On this page you will find out:

Why publish research data?

Publishing research data offers many advantages::

  • Transparency and traceability: Availabe data facilitate the verification and better contextualisation of scientific findints.
  • Reuse and new research: Existing data can be used for new research questions, supplementary analyses and alternative interpretations.
  • Visibility and collaboration: Published and citable data increase the visibility of research and facilitate scientific exchange.
  • Good scientific practice: The publication and long-term preservation of research data are in line with the principles of good scientific practice.

Public access to research results

The publication of research data promotes open and transparent science. Guideline 13 of the DFG Code: ‘Providing public access to research results’ stipulates that research results and the underlying research data should be made publicly available, taking into account legal, ethical and disciplinary considerations.

Archiving and long-term preservation

Regardless of whether research data can be published, data relevant to the research must, as a general rule, be preserved for at least ten years in accordance with Guideline 17: ‘Archiving’ of the DFG Code. This applies to the research data on which the published results are based, as well as key materials and, where applicable, the research software used. The data should be stored using a suitable infrastructure.

What does FAIR stand for?

Research data should be published and documented in such a way that it is findable, accessible, interoperable and reusable. The FAIR Data Principles provide a framework for this.

Illustration der FAIR Data Principles

In practice, FAIR means, amongst other things:

  • unambiguous and permanent identification of the data, e.g. via a DOI (Digital Object Identifier),
  • the data being discoverable, e.g. via a research data repository,
  • meaningful and standardised metadata,
  • suitable and, where possible, open data formats,
  • clear information on access and use.

FAIR does not mean that all research data must be publicly accessible. Legal, ethical or other protection considerations may necessitate restricted access. The key is to make data as accessible and reusable as possible within the relevant framework.

Before publication, research data should be appropriately prepared and documented. Further information can be found under ‘Documenting research data’.

There are also corresponding principles for research software: the FAIR Principles for Research Software (FAIR4S).

Which research data should be published and archived?

Not all research data needs to be published. The key factor is which data is relevant to the traceability of the research or its subsequent reuse. Which data should be published and preserved in the long term may vary depending on the research project and the disciplinary context.

  • Scientific or societal value
  • Contribution to the traceability and reproducibility of research results
  • Potential for re-use
  • Long-term significance or difficulty in reproducing the results
  • Legal and ethical frameworks
  • Costs and benefits of processing and archiving

Suitable examples include final or processed datasets, raw data that is difficult to reproduce, derived data, anonymised datasets and code with associated data.

The Digital Curation Centre’s checklist ‘Five steps to decide what data to keep’ provides guidance on selecting which research data to preserve.

What publication channels are available?

Research data can be published in various ways. It can be published as a standalone research output in a research data repository, linked to a publication as ‘Supplementary Data’, or described in a data paper. Research software and code can also be published as standalone outputs and described in a scientific context.

Research data can be published and archived long-term in a suitable research data repository. It can be made available independently of a publication or linked to an associated article or data paper. The data is made available as a standalone research output and can be referenced and cited independently via a persistent identifier (PID).

You can find out more about selecting a suitable research data repository in the section ‘Where can research data be published and made available in the long term?’.

Supplementary Data refers to research data published to complement a scientific article. Many journals allow for this form of data publication. The data can be made available directly alongside the article or deposited in a research data repository and linked to the article.

If the data are made available directly alongside the article, they are often not discoverable or citable independently of the article. It is therefore advisable to also publish the underlying research data in a suitable research data repository and make them available in the long term. Where possible, open, established data formats and standards should be used.

Data Availability Statement

When publishing a research article, many journals require a Data Availability Statement. It specifies where the research data underlying the article are available and under what conditions they can be accessed.

Please follow the guidelines of the respective journal or publisher. If the research data have been published in a repository, the statement should include the dataset’s DOI or other persistent identifier. If access to the data is restricted, the applicable access conditions should be specified.

Guidance and examples:

Linking a publication to research data

When research data is published in a repository, the DOI of the dataset should be provided in the relevant link field of the journal. Conversely, the dataset should also be linked to the associated article. In this way, the article and the research data are unambiguously linked to one another and can each be found in the context of the other resource.

In a data paper, the dataset itself is the focus of the publication. In particular, the data paper describes the data collection, methodology, processing, quality and potential re-use of the data.

Data papers are published in data journals and are usually subject to peer review. The research data itself is made available in a research data repository and linked to the data paper.

Selection of data journals

Various directories and overviews can help you find suitable data journals:

Before submitting, check the reliability and quality of any journal you are unfamiliar with. The Directory of Open Access Journals (DOAJ), for example, provides guidance.

Where can research data be published and made available in the long term?

Research data repositories are the key venues for the long-term preservation and archiving of research data. They ensure that research data can be found, accessed and cited in the long term.

Different research data repositories are available depending on the discipline, research context or data type. If a suitable discipline-specific repository is available, this should be used in preference to others. If no suitable option is available, TUHH Open Research (TORE) is available at the TUHH.

TUHH Open Research (TORE)

The TUHH Open Research (TORE) research data collection is the TUHH’s institutional research data repository. It supports researchers in publishing their research data in a FAIR-compliant, long-term and citable manner.

Research datasets published in TORE are described using standardised metadata and are assigned a DOI to ensure they can be cited permanently. Long-term provision and archiving are carried out using the infrastructure of the University of Hamburg’s Regional Computing Centre.

Through the TORE Dashboard, research data publications and their links to scholarly publications are presented in a clear and accessible way. This makes it easy to identify, for example, the research data associated with a particular article at a glance. An example is the TORE Dashboard for CRC 1615

TORE is particularly suitable when:

  • there is no suitable subject-specific repository available,
  • research data spans several disciplines and cannot be assigned to a single subject-specific repository,
  • research data is published, for example, in connection with a doctoral thesis.

  • Advice and support from the TUHH University Library on the preparation and publication of your research data,
  • Persistent identifiers (PIDs) for unambiguous identification: DOIs (Digital Object Identifiers) for research data, as well as ORCID iDs for individuals and ROR IDs for organisations,
  • Support for the DataCite metadata standard and established licensing models, e.g. Creative Commons,
  • long-term provision and archiving of data,
  • Support for the provision of data with open or restricted access,
  • Automatic inclusion of your research data publications in the TUHH Research Report,
  • Display of your research data citations in the publication lists on the department webpages.

Research data should be permanently and reliably citable when published as a standalone research output. Through DataCite, datasets are assigned a unique DOI, which enables them to be found and cited permanently.

Example:

Aberle, Christoph (2019). Mobility as a Service: ein Angebot auch für Einkommensarme? (Geo-Datensatz). TUHH Universitätsbibliothek. https://doi.org/10.15480/336.2396.2

Datasets from TUHH Open Research can also be imported into an ORCID profile and thus uniquely linked to the individual.

In TORE, you can also cite research data that has already been published in another repository, such as Zenodo or the TIB AV Portal. There is no need to re-upload the data to TORE; it is simply referenced within TORE. This means that research data published in another suitable repository can also be made visible as TUHH research output and cited in the TUHH Research Report.


Subject-specific repositories

Where a suitable subject-specific or data-type-specific repository is available, it should be used in preference to other options. It supports specific data and metadata standards and makes it easier to locate and re-use research data within the relevant research community. Many of these repositories also offer subject-specific advice and support in preparing data for publication.

  • NOMAD: Repository for data from materials science and related fields, with an integrated electronic laboratory notebook.
  • Chemotion: Research data repository with an integrated electronic laboratory notebook for chemistry.
  • PANGAEA: Repository for georeferenced data from the geosciences and earth sciences.
  • EMPIAR: Repository for electron microscopy data.
  • Repo4Cat: Repository of the National Research Data Infrastructure for catalysis data.
  • TIB AV Portal: Repository for scientific audiovisual media.
  • GESIS Data Services: Repository for research data from the social and economic sciences.

Selecting a suitable repository

If you are looking for a suitable subject-specific repository, you can use re3data.org. The directory provides an overview of research data repositories from various subject areas. Using search filters and the DFG subject classification, you can find suitable repositories and compare them based on various characteristics.

When selecting a repository, you should pay particular attention to ensuring that it:

  • enables the assignment of a persistent identifier (PID), such as a DOI,
  • supports appropriate metadata standards,
  • offers suitable access and licensing models,
  • provides a transparent policy on the operation,
  • access and long-term preservation of the data,
  • guarantees the long-term availability and archiving of the data.

Science Europe’s Practical Guide to the International Alignment of Research Data Management provides detailed guidance on selecting trustworthy repositories, particularly the chapter entitled ‘Criteria for the Selection of Trustworthy Repositories’.

Publishing research software and code

Auch Forschungssoftware und Code können als Forschungsoutput eigenständig veröffentlicht und wissenschaftlich beschrieben werden.

In a repository: Zenodo

Software is frequently developed at Hamburg University of Technology (TUHH) using GitLab or GitHub. To ensure long-term availability and citability, it can, for example, be archived in the interdisciplinary repository Zenodo. When releases are published, each version is assigned its own DOI. The software is also promoted as a research output of the TUHH via the TUHH Community.

When publishing research software, a suitable licence should be chosen. Choose a Licence helps you select an appropriate open-source licence.

In a software journal

Furthermore, they can be described and published as standalone academic articles in a software journal. A selection of peer-reviewed open-access journals for research software:

FAIR Principles

  • FAIR Principles – GO FAIR
  • Wilkinson, M. D. et al. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3, 160018. DOI
  • Chue Hong, N. P. et al. (2022). FAIR Principles for Research Software (FAIR4RS Principles). DOI
  • Jones, S. & Grootveld, M. (2017). How FAIR are your data? – Checkliste zur Selbsteinschätzung der FAIRness von Forschungsdaten. DOI

Research Data

Research Software

  • Tas, E., Ruprecht, D. & Knopp, T. (2026). Information Sheet: Publishing Research Software as a FAIR Research Output (Version V01). Zenodo. DOI
  • Choose a Licence – Selecting an open-source licence

Consultation and Contact

Francesca Schulze

Francesca Schulze

RDM Officer
(Central RDM contact)

Email: research-data@tuhh.de
Phone: +49 40 30601-3311

Dr.-Ing. Seoyun Sohn, Data Steward, Cluster of Excellence BlueMat

Dr.-Ing. Seoyun Sohn

Data Steward, Cluster of Excellence BlueMat
(RDM support for BlueMat projects)

Email:  bluemat-rdm@tuhh.de
Phone: +49 40 30601-2647

Nimet Eylem Tas

Data Steward, CRC 1615
(RDM support for projects within CRC 1615)

Email: smart-rdm@tuhh.de
Phone: +49 40 30601-2539


Detailed information on research data can be found on the information platform forschungsdaten.info.

Scroll to Top