Illustration of connected search, communication and digital information

What exactly is Wikidata and why should we know about it?

A text about a database that quietly shapes parts of the information landscape around us. And why it might deserve a place in courses on academic research and information literacy.

How this text came about

This text is the result of an exchange during an Erasmus visit: my colleague Patricia Caló from Bari (Italy) was visiting us at TUHH to learn more about library work in Germany.

As we talked about „openness“ at TUHH, topics like Open Access and Open Education, we quickly got onto Patricia’s thesis on Wikidata and the parallels to our everyday work at TUHH: from classic advice at the service point to metadata work in backend systems to Open Access. We realised: Wikidata is the perfect connecting piece between all these worlds. We brought these threads together and thought about how Wikidata could be practically integrated into existing teaching formats. In the spirit of Open Education, we have designed this text itself as an Open Educational Resource (OER).

You already used Wikidata today. You just don’t know it.

Think of the information boxes that appear alongside many Google searches: they are built from structured knowledge about people, places, and concepts. Wikidata organises knowledge in a similarly structured, machine-readable way. Its data can also be used by voice assistants and other digital services to answer factual questions. And the structured information shared across Wikipedia’s more than 300 language versions? Wikidata plays a key role in connecting it.

Wikidata feels like it is almost everywhere. And yet almost nobody knows its name.

What actually is Wikidata?

Wikidata is a free, structured knowledge base hosted by the Wikimedia Foundation, which also hosts Wikipedia. While Wikipedia consists of articles written in natural language, Wikidata stores facts in machine-readable form: as so-called statements made up of an item, a property, and a value.

Put simply, the item is the thing being described, the property specifies what kind of information is being expressed, and the value provides that information.

A simple example:

The item „Mona Lisa“ has the property „creator“ with the value „Leonardo da Vinci“.

  • Simplified Wikidata statement Mona Lisa
  • Wikidata item for Mona Lisa showing the statement creator: Leonardo da Vinci.

A classic Wikidata example:

Marie Curie is often used in Wikidata documentation to illustrate how structured statements work.

The item „Marie Curie“ has the property „place of birth“ with the value „Warsaw“. And „Warsaw“ is itself an item that can be connected to further information, such as coordinates. In this way, individual statements become part of a larger network of structured knowledge.

  • Simplified Wikidata statement Marie Curie
  • Wikidata item for Marie Curie showing selected statements, including place of birth: Warsaw

This kind of linking is at the heart of Linked Open Data: instead of existing in isolation, data is connected to other data in a way that machines can read, query, and process. Wikidata is therefore not an encyclopaedia. It is more like the structured backbone behind many of the knowledge products we use every day.

And all of it is open: the data is published under the CC0 licence, meaning it is entirely in the public domain and can be used, adapted, and embedded by anyone. More details on CC0 and other Creative Commons licences are available on the TUHH University Library’s Creative Commons information page.

Why is this more than a technical detail?

At first glance, Wikidata might seem like a topic mainly for computer scientists, data stewards, or other specialists who work heavily with data. But that reading falls short.

Wikidata has become a very large structured knowledge base. It is used by research organisations, museums, libraries, and cultural institutions around the world. Artificial intelligence, particularly large language models and search systems, can also draw on Wikidata as a source of structured knowledge.

Where Wikidata is used, its contents can therefore influence what information digital services surface and how entities and facts are represented. This is more than background knowledge. It is information infrastructure.

And as with any infrastructure: those who do not know it cannot critically evaluate it – or actively shape it.

What does this have to do with academic work?

In our seminar „Academic Research“ at TUHH, we deal with questions like: Where do I find literature? How do I evaluate sources? How do I cite correctly? How do I keep my knowledge organised and traceable?

Wikidata touches many of these questions in ways that are not immediately obvious, but on closer inspection, surprisingly concrete.

Source evaluation and information literacy

A key topic in the seminar is critical engagement with sources. Wikipedia is often considered non-citable in academic contexts. Depending on the context, that is quite fair when it is used instead of the underlying scholarly or primary sources. But the underlying question is more interesting than the prohibition:

  • Where does what is written there actually come from?
  • How is it updated?
  • Who can change it?

Wikidata makes exactly this process visible and traceable. Every statement in Wikidata can be backed by sources. And every change is documented. That is a good starting point for a conversation about how knowledge is created and verified in digital spaces.

Identifiers and authority data

Anyone working with academic databases will sooner or later encounter concepts like DOI, ORCID, or GND. Wikidata links all of these identifiers together. An entry for a researcher can contain not just their name, but also their ORCID, their GND number, their affiliations, their publications. This makes Wikidata an unexpected bridge between the open web and the structured world of academic databases and helps illustrate why identifiers and structured metadata are so important for reference management tools like Zotero or Citavi.

The „Where does background knowledge come from?“ conversation

In practice, students often start their research on Wikipedia or use AI tools to get an initial overview of a new topic. For getting started, that is not a problem at all, as long as they understand what these sources can and cannot do. Wikidata offers a tangible entry point here: you can trace in real time how structured knowledge is built, which sources back it up, and where the gaps are.

Why should librarians and educators know about Wikidata?

You do not need to be a Wikidata expert to work meaningfully with it. But a basic understanding of what Wikidata is and how it fits into the information landscape is valuable for anyone who supports students with literature research and information literacy.

Ryan Hughes, Digital Initiatives Librarian, puts it well:

For library professionals Wikidata is an extension of the work that we do in our institutions.“

For people working in libraries, there is a direct connection: concepts like authority data, subject indexing, and authority control, which have long been central to library catalogues, have an open, collaborative counterpart in Wikidata. Knowing Wikidata does not mean learning something entirely new. It means recognising familiar concepts in a new context.

A look ahead: Wikidata in the seminar?

Could we meaningfully integrate Wikidata into a seminar session on academic research? We think so. Not as a detour into the world of data, but as a concrete hook for topics already part of the course.

A possible 90-minute session might look like this:

  • Introduction (10 min): A short demonstration of what happens when you search for a person or concept on Google. What kind of structured knowledge appears alongside the search results? What lies behind such information?
  • Brief overview (15 min): What is Wikidata? How is it structured? What are items, properties, statements?
  • Hands-on (30 min): Students look up a term or person from their own field in Wikidata. What information is there? What is missing? What sources are cited? What is well-evidenced and what is not?
  • Discussion (25 min): How does this knowledge change the way we look at information we use every day? What does it mean for source research?
  • Wrap-up (10 min): A brief transfer: where else do we encounter Wikidata? And what do we take away for our own academic work?

This would not be compulsory content for everyone. But it would be a session that shows: information literacy does not end with telling a good source from a bad one. It begins with understanding the infrastructure on which our information landscape is built. And that is exactly what Wikidata can make tangible.

Want to try it yourself?

Anyone who wants to explore Wikidata directly, without working through lengthy documentation, should check out the Wikidata Tours. Reading about Wikidata is one thing, but seeing how it works in practice can make its underlying principles much easier to understand. The interactive guided tours offer a quick and accessible way to get started and run directly in a browser. No registration or account needed.

Conclusion

Wikidata is not a curiosity for data enthusiasts. It is one of the central knowledge infrastructures of our time that operates almost entirely out of sight. That is precisely why it is worth making it visible.

For librarians supporting students with information research, Wikidata is a useful point of reference: it connects familiar concepts from the library world with the realities of the open web. And for seminars on academic research, it could be a tangible way to make abstract questions about knowledge, sources, and information infrastructure concrete and discussable.

Perhaps information literacy starts exactly there: not with a ban on citing Wikipedia, but with an understanding of why Wikipedia works the way it does and what lies behind it.

Do you have experience with Wikidata, in teaching, at school, in your work with information, or in library practice? Or questions this post has raised? Share your thoughts in the comments, on Mastodon or via mail!

Further reading


CC BY 4.0
Reuse as OER explicitly permitted: This work and its contents are — unless otherwise stated — licensed under CC BY 4.0. Attribution according to the TULLU rule please as follows: What exactly is Wikidata – and why should we know about it? by Patrizia Caló and Florian Hagen, Licence: CC BY 4.0. Der Beitrag und dazugehörige Materialien stehen auch in nachnutzbaren Formaten sowie als PDF zum Download zur Verfügung.

Schreibe einen Kommentar

Deine E-Mail-Adresse wird nicht veröffentlicht. Erforderliche Felder sind mit * markiert