Synced automatically from infinity-project/rulebook (main) at build time – edit the source repository, not this file.

Software and data management

← Back to index

Metadata file for publish & find software

The software should be findable through a structured metadata file and by following best practices for repository discovery and SEO. The recommended metadata file is codemeta.json, a standard JSON metadata schema for software projects, supported across repositories, Zenodo, and software registries. It:

  • Provides machine-readable metadata for your software
  • Makes it easier for search engines, package managers, and registries to index your software
  • Can be used to automatically populate Zenodo or other software catalogs

Minimal example:

{
  "@context": "DOI of the current release at Zenodo",
  "@type": "Component type",
  "name": "Software title",
  "description": "a short abstract (goal, motivation, problem, future and current evaluation methods) max 200 words",
  "version": "Release versioning number",
  "license": "License type",
  "author": [
    {"%CHAMPION%": ""}
  ],
  "repository": "URL of the repository as part of the INFINITY ecosystem",
  "datePublished": "Date of latest release"
}

It belongs in the root of your repository. Optionally, also include a CITATION.cff, also used by Zenodo, providing a standardised research citation for your repository. Update the citation at each release (which should have a new Zenodo DOI) so it’s possible to track which version was used in which work.

Metadata for publish & find data (FAIR principles)

All datasets published in the INFINITY ecosystem must align with FAIR data principles: findable, accessible, interoperable and reusable (cf. INFINITY Data Management Plan, Chapter 4).

We also conform to the CARE principles — Collective Benefit, Authority to Control, Responsibility, and Ethics — for datasets whose production and reuse engage communities with an interest in how their cultural heritage is described and made available.

Findable

All data components registered in the INFINITY Research Ecosystem are assigned persistent identifiers (PIDs) at the point of first public release. The PID scheme is differentiated by component type. Datasets, annotated corpora, and AI models are deposited in Zenodo, where a DOI is minted automatically. They should use Schema.org annotations, enabling indexing by general-purpose search engines.

MUST use JSON-LD embedded in HTML with the Schema.org Dataset schema:

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Dataset",
  "name": "INFINITY Climate Dataset",
  "description": "Dataset of European climate measurements.",
  "url": "https://your-pages-url",
  "keywords": ["climate", "temperature", "Europe"],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "creator": {
    "@type": "Organization",
    "name": "INFINITY Labs"
  },
  "datePublished": "2026-03-24"
}
</script>

SHOULD additionally use a metadata schema such as DCAT, which provides richer catalog metadata. Established mappings exist between Schema.org and DCAT.

Accessible

Metadata for all INFINITY components remains publicly accessible regardless of whether the underlying data are open, restricted, or under embargo.

Interoperable

INFINITY reuses established controlled vocabularies and reference ontologies to maximise cross-collection alignment. All components are described using Dublin Core Terms for bibliographic metadata and DCAT-AP 2.1 for dataset catalogue records. Named entity annotations are aligned to Wikidata, GeoNames, and the Getty Vocabularies (AAT, TGN, ULAN) using owl:sameAs and skos:exactMatch assertions where equivalence can be established. Rights statements are expressed using the RightsStatements.org vocabulary and ODRL 2.2 policies, enabling machine-actionable licence reasoning.

Reusable

Open licenses are the default (cf. Table 4.1 of the Data Management Plan; see Licensing below). Licence assertions are recorded in machine-readable form in each component’s metadata record using the dct:license property with a URI identifying the licence, and as SPDX identifiers in software repository headers. For AI models, the model card schema includes a dedicated licence field.

FAIR compliance is assessed and documented at each DMP revision using the F-UJI automated FAIR data assessment tool, evaluating compliance against the RDA FAIR Data Maturity Model indicators. More detail in the INFINITY Data Management Plan.

Licensing

The project will adopt an open-source licensing strategy aligned with European Commission guidance on interoperability and software reuse, selecting appropriate licenses according to the intended exploitation model, ecosystem integration needs, and sustainability objectives. All contributions must use an approved* open-source license such as:

  • MIT
  • Apache 2.0
  • GNU GPL v3
  • BSD 3-Clause
  • EUPL-1.2

A description of the differences between usual open source licenses can be found here.

Can derivatives become closed source?

License Closed-source derivatives?
MIT ✔ Yes
Apache 2.0 ✔ Yes
EUPL ⚠ Usually no
GPL ❌ No

It is noted that EUPL is the license supported by the EC for software developed in EU-funded projects. See the compatibility matrix to help choose between license options.

* Approved is anything approved by the Open Source Initiative (OSI), such as the four license types above. Generally, code can be re-used and further modified; new distributions of the code (same or modified) may have to preserve the open source character; the code creators do not provide any warranty and have no liability for it.

License information is shared in the repository through the LICENSE file.

Dataset licensing will be selected to balance openness, reuse, and responsible exploitation, enabling broad access and interoperability while ensuring appropriate protections for sensitive data, provenance requirements, and downstream commercial or research use where applicable. Project datasets must be released under open data licences such as Creative Commons Attribution 4.0 (CC BY 4.0), enabling broad reuse, interoperability, and downstream innovation while ensuring appropriate attribution and alignment with FAIR and Open Science principles.


← Previous: Code documentation · Back to index · Next: Semantic layer development and documentation →