Druckansicht | © 2002 – 2026 tcworld GmbH | Seite drucken

Think outside the folder

Technical documentation doesn’t have to live in rigid folders. Metadata can create a flexible information space that makes content easy to find, reuse, and deliver.

Text by Fritz Lülf

Inhaltsübersicht

Image: Manojkumar Madhusoodananpillai/istockphoto.com

Metadata for clearly defined data objects

The findability of any technical information begins with a clear understanding of what is to be found. Data objects define units which can be a topic, an instruction, a figure, a term, a table, a module, or a complete document. Such data objects have a clear beginning and end. This allows them to be reused, moved, combined, and enriched with metadata. In this sense, metadata is not simply “data about data”. It is additional information that describes a clearly defined data object and helps locate it within an information space. Metadata is what makes the data object retrievable.

This is relevant to both content creation and content delivery. Authors of data objects need to make content findable for colleagues as well as for their own future work. Users, in turn, need a transparent way to reach precisely the information that applies to their situation.

Machine processing adds a further requirement: AI systems also need explicit structures if they are to classify content not merely by linguistic similarity but in accordance with the right product, target audience, and usage context.
  

From folder structures to information spaces

Traditional folder structures provide access to data objects through a fixed hierarchy. Every filing decision imposes an order: Content might first be organized by life-cycle phase and then by target audience – or the other way around. As more characteristics are added, parallel paths emerge at lower levels. Objects must be stored more than once, linked, or made accessible through additional conventions. The more levels a hierarchy contains, the more duplicates, special cases, and ultimately effort are needed.

Thinking about metadata in terms of linear algebra takes a different approach. Instead of defining a fixed path through a folder tree, it creates an information space. The individual metadata dimensions form the directions into which the information space expands. The associated values act as coordinates. The location of a data object is not defined by a folder tree but by the combination of its coordinates.

For example, the dimensions [target audience] and [life-cycle phase] might contain the values {maintenance crew} and {end user}, and {commissioning} and {decommissioning}, respectively. Content about commissioning for end users can then be located at the point defined by these two coordinates: [end user; commissioning] (Figure 1).

The key advantage is that access no longer depends on a fixed sequence. A product manager might start with a product and then select a market or region. A legal department might start with the target market and select the product afterwards. Both routes lead to the same data object.

Metadata as an information space therefore enables different professions to approach the same content repository from different perspectives without having to structure the content separately for each one.
  

Distinguish dimensions from values

A robust metadata concept relies on a clear separation between dimensions and values. A dimension describes a type of metadata, e.g., [language], [target audience], [product], [life-cycle phase], or [region]. A value is a specific manifestation within that dimension, e.g., {German}, {end user}, {Ultra}, {commissioning}, or {North America}. If a single value is modeled as a dimension in its own right – such as [pink] with the values {yes} and {no} alongside an existing dimension [color] – the model quickly becomes inconsistent and harder to maintain.

This distinction is also essential for reliable filters and queries. Systems can retrieve and combine data reliably only when data objects are described according to the same principle. A metadata concept relies on a limited number of clearly named characteristics with controlled and comprehensible values, rather than an unstructured collection of tags.

These values do not have to be organized as a flat list. Groups can be formed within one dimension. For example, <British, American, Australian, South African> English can be part of the broader category {English} in the dimension [Language]. This preserves the semantic relationship within one dimension while allowing users to select either the entire language group or a specific variant.
  

Only independent dimensions add informational value

Not every additional dimension makes the information space more meaningful. Consider, for example, the dimensions [target audience] and [responsible role]. If the person responsible for a task is also its target audience, adding a separate dimension for responsible role provides little additional information while increasing the maintenance effort. In terms of linear algebra, such dimensions are not orthogonal.

By contrast, the dimensions [target audience] and [life-cycle phase] can vary independently. Different audiences can need information during different phases of a product's life cycle: Logistians are needed during commissioning and decommissioning. The two dimensions therefore provide distinct information and can be considered orthogonal.

As a technical communicator, consider orthogonality as a conceptual test rather than a mathematical measurement: Can a dimension be combined with the values of the other dimensions? Does every dimension provide a new distinction? The more certainly this question can be answered with “yes”, the more useful the information space becomes.

Orthogonality is indispensable for adding and removing dimensions and values without having to reorganize the entire information space. A new value – for example a new language – or a new dimension – such as release state – can be added without forcing the rest of the model to be reorganized. This creates a flexibility that rigid folder hierarchies cannot provide.
  

Align the axes with actual information needs

A metadata model should not simply reproduce existing organizational structures or filing systems. What matters is how content is actually distinguished, searched, and delivered in everyday work.

Imagine that the product models {Ultra}, {Mega} and {Giga} are permanently linked to the sizes {2000}, {3000}, and {4000}. The dimensions [model] and [size] then do not describe completely independent characteristics. A simpler dimension such as [type], with values {A, B, C}, might express the relevant relationship more effectively.

Linear algebra describes such realignment as a change of basis. The same data object can be addressed in different coordinate systems. Suitable principal axes shorten search paths, reduce redundant characteristics, and directly reveal the decisive distinctions. Frequent search issues and manual workarounds indicate that the dimensions may be unsuitable. The points where users fail to find content reveal which metadata is missing or which dimensions are poorly designed.
  

Scalability through many dimensions with few values

The number of possible combinations grows with the number of values and dimensions. With w values in each dimension and d dimensions, the model can in principle describe wᵈ positions. Three dimensions with three values each yield 27 possible combinations. Six dimensions with three values each already yield 729 combinations. Adding independent dimensions therefore expands the information space far more than merely extending individual value lists.

This leads to an important design principle: Many clearly separated dimensions with a small number of values each are preferable to a few overloaded dimensions containing long and heterogeneous value lists. The resulting combinations are not an end in themselves. Their value lies in allowing content to be described in a differentiated way and then narrowed down efficiently 
through a few meaningful decisions (Figure 2).

Filtering reverses combinatorial growth

Searching within an information space reverses the process of combinatorial growth. In a space with three dimensions and three values per dimension, 27 combinations are possible. Once a value is fixed for one dimension, nine combinations remain. Fixing a second value reduces the set to three, and the third decision reduces it to exactly one combination. Each filter fixes one coordinate and reduces the remaining information space.

This is how a content delivery portal works in practice. A service technician may select a machine type, a size, and, where relevant, a customer. A few intuitive choices reduce a repository of several thousand data objects to a small set of relevant information. The quality of the result therefore depends not only on its search technology but largely on metadata quality. Only clearly discriminating dimensions and consistently assigned values allow reliable filtering.
  

Subspaces enable flexible outputs

Selecting values creates smaller subspaces within the overall information space. Consider a three-dimensional space consisting of [product], [size], and [language]. Selecting a product leaves all sizes and languages for that product. Selecting a particular size leaves the languages for that specific product and size only. This makes it possible to generate both individual data objects and useful subsets: all content for one product, all language versions of a particular variant, or all sizes available in a particular market (Figure 3).

This principle remains valid independent of the number of dimensions. An information space may contain 80 dimensions which cannot be visualized. If values are fixed for 68 of them, a twelve-dimensional subspace remains. This is what makes the approach suitable for complex documentation landscapes: The entire information space does not need to be visualized as long as the dimensions are clearly defined and orthogonal.
  

Model the information space before choosing the tool

The model of the information space is not tied to a particular software. The relevant dimensions and values are independent of any tool. The properties of the data objects should therefore be defined first in the information space. Only then should the concept be implemented in a CCMS, CDP, or other software.

The same principle applies across different roles in a company – from engineering and R&D to legal and marketing. Moving from hierarchical thinking to a multidimensional perspective requires effort, because folder trees are familiar and appear to provide a clear order. The benefits of an information space become apparent when different roles need to access the same repository from different perspectives. Everyone in the company can then define their own route to the information without interfering with anyone else's.
  

Metadata for AI, standards, and requirements

For AI applications, the separation of content and context is particularly important. A statement such as "The operating pressure is 2 bar" provides domain information. It does not indicate that the statement applies to a particular product, variant, or operating situation. AI cannot assign this content to a product (unless all the engineering knowledge has been provided as training data). The relationship therefore needs to be represented explicitly in the content repository or supplied through metadata. Metadata can support AI training, retrieval, and classification processes by defining the information space in which a statement is valid and findable. It does not, however, replace professional knowledge of the underlying relationships.

Standards and regulatory requirements can also be linked to data objects through metadata, at least to a limited extent. A data object may be assigned to several clauses of one or more standards. Such mappings can become difficult to maintain, however, because new editions of standards may move sections or change requirements. For extensive and highly regulated mappings, coordination with professional requirements management is advisable.
  

Develop, test, and improve

A metadata concept does not have to be designed in full in a single step. A structure based on orthogonal dimensions can be developed iteratively. Start with a small number of central dimensions and test them in real search and retrieval situations before adjusting and expanding them. Add new dimensions when they provide independent informational value. Unsuitable dimensions can be removed without rebuilding the entire structure.

Practical implementation relies on four guidelines:

  • Data objects must be clearly delimited.
  • Dimensions and values must be consistently distinguished.
  • Each dimension should be as independent as possible from all others.
  • Many lean dimensions are preferable to a few overloaded categories.

Search difficulties should not be seen merely as user errors. They can also indicate missing metadata or poorly chosen dimensions. The metadata model becomes a learning system that evolves with an organization's requirements and experience.
  

Conclusion

Viewing metadata as an information space connects an abstract mathematical model with practical tasks in technical communication. Dimensions define the information space, and values locate the data objects. Orthogonal dimensions allow for easy access along all routes to a shared content repository and an iterative definition and refinement of the entire information space. Filters create manageable subspaces. This approach overcomes the limitations of rigid folder hierarchies.

A good metadata model does not represent every conceivable property, but those characteristics that are genuinely relevant to authoring, reuse, search, delivery, and AI processing. When metadata is organized around these principal characteristics and developed incrementally, it can provide a scalable foundation for CCMS, CDP, and other data-driven applications.