From DDI-C XML to Graph Databases: Modelling Research Metadata as Connected Knowledge

Research metadata is rarely flat. A single study can be connected to different authors, organizations, funders, topics, keywords and many other fields. While these connections can be hidden within hierarchical XML structures, graph databases make relationships explicit, offering a promising way to explore and reuse DDI-C metadata.

While in a traditional XML document these numerous connections are hidden inside of a hierarchical structure, in a relational database the same connections may instead be split across several tables and join tables. A graph database offers a different option: it treats relationships as first-class data, just like entities themselves. This distinction makes it a promising model, especially when the goal is to explore how datasets, people, organizations and other metadata elements relate to one another.

This article describes the modelling decisions of a practical, experimental demo in which FSD's DDI-C metadata was converted into a graph and imported into a Neo4j graph database using Python and the Neo4j Python driver. The demo was motivated by a broader question: could graph databases offer added value for managing and reusing research metadata? The demo showed that DDI-C metadata can be transformed into a graph structure quite naturally while preserving much of the original XML's structure. However, the process also revealed that the XML to graph process is not purely a one-to-one, mechanical conversion. It requires decisions about graph modelling, missing fields, duplicated entities and validation.

Why convert DDI-C metadata into a graph database?

The DDI-C format is a structured standard to describe research metadata. DDI-C utilizes XML which is well suited for representing such data due to its hierarchical structure. However, research metadata is not purely hierarchical. It is also relational, or in other words, interconnected. A single study may have several authors, an author may be linked to several other studies, a keyword may belong to a controlled vocabulary, and a study may be connected to other related studies. These are not simple hierarchical parent-child relationships. They instead form a connected, complex network of relationships.

Large graph visualisation showing a network of pink study nodes connected to orange keyword nodes, illustrating shared keywords and connections between studies.
A network of pink nodes representing studies connected to orange nodes representing keywords.

This is where graph databases excel as they are designed exactly for this kind of networked information. In a graph database, entities are represented as nodes and connections between them as relationships. Both can be further enriched with metadata using properties, which are key-value pairs similar to relational databases. Unlike relational databases, where relationships are often reconstructed through joins at query time, graph databases store these relationships explicitly. This makes them especially useful when the relationships between entities are as important, or more, than the entities themselves.

This difference matters for DDI-C metadata, because many useful questions tend to be relationship-oriented. For example, consider the following: which studies share the same keywords? Which authors does a study have? Which studies are related to a specific topic, country or time period? By using a graph model, answering these questions is possible by simply traversing relationships between nodes, rather than repeatedly joining relational tables or parsing XML structures.

Another motivation for using graph databases is future reuse. Graph databases can support knowledge graph use cases, visualizing metadata, exploratory searches through relationships and potentially graphRAG (graph retrieval augmented generation) for LLM (large language model) applications. Knowledge graphs and LLMs also complement each other: a graph can provide structured and validated contextual knowledge for the LLM, while language models can help users query or interpret that knowledge. For research data archives such as FSD, this opens up interesting possibilities for richer discovery interfaces or metadata-aware AI applications. But these possibilities depend on the underlying graph model: the graph needs to be designed to support the required use cases. This raises practical modelling questions of which parts of a DDI-C XML file should become nodes, relationships or properties in the graph.

Metadata entities in a DDI-C XML file

The demo's graph model was built around the central concept of a study. In a DDI-C file, a singular study description contains many different metadata elements, but not all of them are equally important for a graph model. The most important entities are those that either describe the study directly or connect it to other reusable concepts.

The central node in the demo model was the Study node. It acts as the root of the graph representation for one DDI-C file and contains properties such as a title, URI, and a unique FSD id. From the Study node, relationships lead to most of the other metadata entities. This study-centric structure was chosen because it gives each dataset a natural anchor point. Most queries about research data begin with a study or aim to find studies that match given criteria.

Graph visualisation with a central Study node connected to numerous DDI-C metadata keywords and other entities.
A study node surrounded by various nodes and relationships representing DDI-C elements.

Beyond the central Study node, the graph contains a wide range of metadata entities parsed from a DDI-C file including descriptive elements (e.g. countries, collection methods, time periods), contributors (authors, funders, distributors), controlled vocabulary concepts (keywords, topics, vocabularies), and lower-level metadata such as variables and files. These entities are important because they contain meaningful metadata that describes studies from different perspectives while also serving as reusable connection points across the graph database. Rather than being tied to a single study, the same organization, country, keyword or vocabulary can be linked to many different studies, creating a rich network of relationships to be queried, while avoiding unnecessary duplication.

Converting XML elements into graph nodes

The basic modelling rule to convert from XML to graph was relatively straightforward: meaningful XML elements were converted into graph nodes, relationships between XML elements became graph relationships, and XML attributes and text fields were stored as node or relationship properties. This largely preserved the structure of the original DDI-C structure while taking advantage of a graph database's features.

However, not every XML element should be converted into a graph. Some DDI-C elements function mainly as structural containers. For example, elements like docDscr, stdyDscr, or titlStmt. These organize the XML document, but do not contain meaningful information by themselves and as such, are irrelevant for the graph. If every such container was imported as a node, the graph would contain several empty nodes with no added benefit to querying it. As such, it is important to focus on importing elements with actual metadata content or elements that serve as meaningful connections and avoid importing empty elements.

Excerpt of a DDI-C XML document showing the structural elements containing level and URI attributes.
An example of structural XML elements. Notice how docDscr and stdyDscr do not have their own attributes or text content. In contrast, otherMat contains attributes and therefore is included in the graph.

This is one of the key differences between preserving XML and designing a useful graph. XML is a tree, and a tree is technically a kind of graph, but a direct copy of the XML tree wouldn't be a particularly useful graph model. The graph model should support the queries users want to utilize. That means some of the XML hierarchy can be flattened, some elements can become properties, and some values can become shared nodes.

An example of a central design decision was whether a metadata value should become a node or a property. Elements with meaningful content became nodes when they represented entities that might be queried independently. For example, a keyword is more useful as a node than as a plain text field property if users want to find all studies linked to a particular keyword. Similarly, organizations, authors, vocabularies and many other fields likely to be reused across many studies are strong candidates for nodes. Meanwhile, values that simply describe a singular entity, such as a title string, URI, or an identifier, are better stored as properties for that entity. This keeps the graph more compact and avoids creating a new node for every literal value. Regardless, the resulting graph model was not entirely unproblematic.

Conversion challenges

The first major challenge was deciding how closely the graph should follow the original XML. Keeping the graph close to DDI-C helps traceability and allows users to see how graph elements directly relate to XML elements. But following XML too literally could produce a graph that is overly hierarchical and inefficient to query for most of the common use cases. This required careful consideration for the compromises between XML structural fidelity and graph usability.

Another major challenge was missing and inconsistent metadata. This does not refer to a fault in the metadata itself, but an intended consequence of DDI-C structure. Some fields are optional, and even mandatory fields may vary in practice depending on the type of study (qualitative vs quantitative). Due to this the import process had to check for field existence and avoid creating empty nodes, which is an important requirement for any potential production pipelines.

Finally, validating the graph remains an open challenge. A production-ready import process would need systematic validation to compare the graph against the original XML files. This could include checking whether the number of imported studies, variables, keywords etc. matches the source files, whether required values were imported, and whether relationships were created in the right places. This kind of validation system would likely require separate development.

Conclusion

The demo showed that DDI-C metadata can be successfully imported into a graph database and modelled in a way that makes relationships more visible and queryable. However, the process also demonstrated that converting DDI-C XML into a graph is not just a format conversion. It also required careful decisions about modelling during the conversion process. These decisions on which elements are converted into nodes, relationships and properties directly affect how useful the resulting graph will be. For the experimental demo a graph model resembling the DDI-C's structure worked well, while production use could benefit from a more bespoke approach.

The broader implication is that graph databases are most valuable when metadata is used not only as storage, but as connected knowledge. For research data archives, graph databases could support better data discovery, richer metadata analysis and potentially AI-assisted services grounded in graph information. Rather than replacing XML, graph databases can serve as a complementary technology alongside it. XML remains highly suitable for long-term preservation, structured metadata exchange and standards-based archival storage, while graph databases excel at exploring relationships and supporting complex discovery use cases. Used together, XML can function as the definitive preservation format and graph databases as a layer for analysis, exploration and knowledge discovery.

Text and images: Väinö Mäkelä