Let’s talk FAIR: The role of the SSH FAIR Vocabulary Registry in supporting (re)use of vocabularies in social sciences and humanities research

17 February 2026

Written by Angelica Maineri in collaboration with Liliana Melgar-Estrada, Kerim Meijer, Menzo Windhouwer, and Tadas Bernotas.

Why vocabularies matter for FAIR research in the Social Sciences and Humanities

In the Social Sciences and Humanities (SSH) field, we spend a lot of time debating concepts: What exactly do we mean by “household”? How do we define educational attainment? Which terms are used to describe occupations, languages, or countries? 

These choices are often captured in controlled vocabularies: curated lists of terms, codes, or concepts that help researchers describe their data in a consistent way. Vocabularies can be created for substantive terms (i.e., for concepts such as “timmerman”)  as well as for describing a schema or data model (e.g., for a column in your spreadsheet called “occupation”). When used well, vocabularies make data easier to find, interpret, and reuse—not only by humans, but also by machines. We previously explored this in more detail in the ODISSEI FAIR blog post Controlled vocabularies for the social sciences.

Human (and machine-readable) vocabularies play a crucial role in making research data FAIR by enabling Interoperability. Reusing existing vocabularies, fully or partially, instead of inventing new ones from scratch, helps data connect to the wider research ecosystem, reduces ambiguity, and saves time and resources. But a practical question remains: where do SSH researchers actually find these vocabularies?

The SSH FAIR Vocabulary Registry: a one-stop place for SSH vocabularies

To address this challenge, the SSH FAIR Vocabulary Registry was developed within the CLARIAH infrastructure project and is now maintained under SSHOC-NL (the Social Sciences and Humanities Open Cloud for the Netherlands), as part of the task Build SSH FAIR & Linked Open Data Model (learn more in this short video).

The registry provides a single access point where SSH researchers, data managers, and developers can discover vocabularies that are relevant to their domain. Rather than hosting vocabularies centrally, the registry acts as a node in a distributed network of semantic resources, pointing users to vocabularies wherever they are maintained.

The registry addresses several challenges in the field :

  • SSH vocabularies already exist, but are often hard to find.
  • Researchers regularly reinvent vocabularies because they are unaware of existing ones.
  • Data interoperability suffers when similar concepts are described in incompatible ways.

By improving the visibility, transparency, and reuse of vocabularies, the SSH FAIR Vocabulary Registry supports open science practices and helps researchers make more informed modelling choices. Next to the vocabulary search, the registry will also include Tutorials and Guidelines on creating, reusing, adapting, and publishing vocabularies.

How does the registry work?

The registry can be used by different people for various purposes. For instance, researchers may use it to find a suitable vocabulary for describing their data, and data stewards can use it to advise researchers on which vocabularies and standards to use.  The SSH FAIR Vocabulary Registry is designed to be useful without requiring technical expertise.

What the registry deliberately does not do is impose one “correct” vocabulary. Instead, it supports informed choice and transparency—key ingredients for meaningful interoperability.

Figure 1: Main page of the SSH FAIR Vocabulary Registry (version 1.1, 20251119)

Figure 2: Example of a vocabulary page, with vocabulary metadata (version 1.1, 20251119)

In terms of main functionalities,  the registry:

  • Allows users to search and browse vocabularies relevant to the SSH domain, using either text search or filter facets;
  • Provides clear contextual information: what a vocabulary is about, where it can be found, who maintains it, how many versions there are, and how it can be used;
  • Points to where the vocabularies are hosted, while respecting community ownership. It also keeps a stored copy of each vocabulary to ensure continued access if the original source becomes unavailable.
  • Allows users to query the vocabularies using SPARQL;
  • Encourages the reuse of vocabularies that follow the FAIR principles.

The SSH FAIR Vocabulary Registry will continue to evolve during the remaining period of the SSHOC-NL project, guided by user feedback. A major upcoming feature is a recommender tool that helps researchers identify relevant vocabularies for their data: by uploading a spreadsheet, the tool analyses the column headers and suggests appropriate vocabularies, schemas, or properties for modelling. Experimental versions of this tool are already in use.

Help grow the registry

The SSH FAIR Vocabulary Registry is meant to be a community resource, and its value grows with use and participation. For instance, it already includes resources listed in the Awesome Ontologies for Digital Humanities and its Social Science equivalent

If you know of a vocabulary relevant to the Social Sciences or Humanities that should be included, the team would love to hear from you. You can contribute by sending a message to fairsupport@odissei-data.nl with a link and a brief description of the vocabulary. You can also share questions or suggestions for the Guidelines (a draft is available here: https://registry.vocabs.clariah.nl/guidelines/) or get involved in helping us develop them.

By sharing knowledge about existing vocabularies, you help others avoid duplication and strengthen the FAIRness of SSH research data—one concept at a time.

Relevant links

DH Benelux presentation: https://zenodo.org/records/15696647

Photo by Joshua Hoehne on Unsplash