YourMetadata@DataCite

Much of the work being done on metadata completeness is focused on measuring the completeness of a repository or a specific subset of a repository related to a specific institution. Results are conceived in terms of repository behaviors. Recent presentations of the National Institutes of Health S-index Challenge included the idea that the methods of determining the S-index could be applied at many levels, in particular, individuals, labs, and institutions can have S-index scores. We can explore what this looks like in DataCite using the selection query capability in the Metadata Game Changers Repository Tools.

Figure 1 shows the selection box for the metadata completeness tool. It takes advantage of the amazing DataCite query capability to retrieve all metadata records from DataCite that include my ORCID as a creator identifier with two inputs:

  1. the special repository ID “datacite.all” tells the tool to search all of dataCite and

  2. the query box includes the DataCite query for my ORCID as the Creator Identifier: “creators.nameIdentifiers.nameIdentifier:*0000-0003-3585-6733”.

Paste your ORCID into this query to see your metadata. Don’t forget the ‘*’ that covers the cases where ‘https://orcid.org/” is included as part of the ORCID.

The API URL used for the retrieval (https://api.datacite.org/dois?page[size]=250&affiliation=true&publisher=true&query=creators.nameIdentifiers.nameIdentifier%3A*0000-0003-3585-6733&page[number]=1) is shown below the selection along with the number of records matching the query (181) and the number analyzed (all of them in this case).

Figure 1. The selection box for Repository Tools showing the inputs for an ORCID search across all of DataCite.

How Complete Are My Metadata?

The primary focus of this tool set is the completeness and connectivity of these metadata, so this is the first question we can answer. Figure 2 shows completeness of my DataCite metadata for four FAIR use cases. The radar plots shown here are a overview strip activated with the Plots button and designed for sharing, i.e. pasting into emails, reports, or or blogs. The tool shows the metadata elements around the edge of each plot and how they are mapped to DataCite metadata elements here. The overall completeness of 34% puts my metadata in the top 25% of all of DataCite!

Figure 2. Metadata completeness scores and radar plots for four FAIR use cases for these metadata

These FAIR use cases are interesting because they have been used for many years to characterize completeness of many DataCite Repositories and have a well-established empirical corpus. Of course, there are other community recommendations that individuals might be interested in exploring. The Repository Tools support 1) a DataCite approximation of the SHARE Score proposed by the ConductScience Foundation, one of the three winners of the NIH S-index challenge, 2) a project metadata recommendation I proposed earlier this year, and a couple of extras that include all DataCite relationTypes and contributorTypes if you are interested in connections supported by DataCite metadata.  

Where Are My Metadata?

A second obvious question is “where are my DataCite metadata?”. This question can be answered using the values feature of the metadata completeness tool. The metadata elements in each use case are listed next to the radar plots below the result summary and plot strip described above. Clicking on those concepts shows the distribution of values for that concept in the selected metadata. Figure 3 shows the values of Resource Publisher in my metadata. Two of the Generalist Repository Ecosystem Initiative repositories are at the top (Zenodo and Figshare) with many other organizations I have worked with also showing up in the long-tail of the distribution. Note that the Zenodo count is larger than the others partially because it includes multiple version DOIs for each resource object.

Figure 3. Values from my DataCite metadata for the concept Resource Publisher that maps to the publisher.name DataCite element. These values can be copied into the clipboard or saved to a CSV file for comparisons or analysis.

One of the primary benefits of using ORCIDs to identify people is disambiguating multiple people with the same name and also multiple spellings of my name. We can get a picture of how this is working by using the same interface to search for query=creator.name:Ted+Habermann. This query returns more results, 250 vs 181, as expected, and a much wider selection of publishers, 23 vs. 14. Comparison of these two groups provides an opportunity for some metadata triage - diagnosing potential errors in various metadata records.

Limitations

There are several limitations of this approach to finding your metadata that need to be kept in mind:

1.    Despite increasing adoption of ORCID across the global research infrastructure, there is still a long way to go. Comparing the results of the ORCID search and the name search described above may shed light on metadata with your name and not your ORCID, and the data shown here might be a small nudge towards increased ORCID adoption, a community-wide challenge.

2.    This tool is limited to DataCite metadata. The DataCite Commons and OpenAlex provide a broader view (274 and 289 works) of my metadata that includes Crossref and some other sources, along with some facets. They are both great discovery tools that have large, growing corpora. Like any discovery mechanism, they both have idiosyncrasies and data journeys that need to be considered when you use them.

Metadata Improvement

The focus described here is metadata understanding, measurement and improvement rather than discovery so these tools facilitate a deeper look into your metadata than the discovery tools. For example, the values display described above goes beyond the facet displays that are generally just the top ten values. Sometimes, the devil is in the long tail!

The connectivity tool, linked to the discovery tool in the top bar as “Share with Metadata Connectivity”, focuses on identifiers for creators, contributors, publishers, and funders and integrates ORCID and ROR searches that can help you and your repositories increase identifier coverage. As such, they are another tool in your belt for helping you find the identifiers that improve results and accuracy of the discovery tools and ensure that these connectors are in your metadata. Working with DataCite metadata has the benefit that DataCite metadata are sometimes more accessible to researchers than other metadata with more intermediary organizations. Finding a viable path for re-curating new content into the metadata in response to problems identified with these tools may be easier in the DataCite case.

These tools are freely available and can be applied to any DataCite Repository by entering the name or the id into the explore box shown in Figure 1. Hopefully you find them interesting and useful for you and your repository. They might even help you improve your S-index!

Please let me know if you have questions or suggestions for improvements.