Data dictionary =============== The data dictionary is the resolved form of a project: every data object with its owner, its users, its shape, its limits and its scaling already worked out. It is the contract between the checking front end of DDD and its output backends, and it is the only thing they share. Everything before it - loading the description files, resolving the declarations, running the :doc:`consistency checks ` - produces a dictionary; everything after it - the c backend, the a2l backend, whatever is added next - consumes one and nothing else. A backend therefore never reaches into the loader or into the analysis, and if it needs to know something, that something is a field of the dictionary rather than a second calculation performed in a template. That arrangement is worth having for two reasons. The first is that two backends cannot disagree about what a project contains: the shape a c array is declared with and the ``MATRIX_DIM`` written into the a2l are read from one field, so they cannot drift apart - each backend only has to order the indices the way its own format wants them, c with the last index running fastest and ASAP2 with the first. The second is that the work of resolving a project is done once, in a place that reports findings, instead of once per output format in a place that can only crash. DDD publishes the dictionary rather than keeping it to itself. ``ddd dump`` writes it out as json, and ``ddd schema dictionary`` prints its json schema, so a generator DDD does not ship - a report, a database importer, a header for a language DDD knows nothing about - can consume a checked project without importing python and without depending on any of the implementation: .. code-block:: bash ddd dump examples/demo/demo.ddd.json > dictionary.json ddd schema dictionary > dictionary.schema.json .. code-block:: text $ ddd dump examples/demo/demo.ddd.json { "format": 9, "name": "DemoDevice", "description": "Demonstration project showing every kind of data object", "source": "demo.ddd.json", "components": [ { "name": "Controller", "description": "Consumes the raw values and produces the derived ones", "source": "controller.ddd.json", "declarations": [ { "name": "ValueA", "scope": "input", "condition": null }, ... ``ddd dump`` is the one command whose standard output *is* the payload, so its findings go to standard error and the redirection above works whether or not the project has any. A script writes the file with ``-o dictionary.json`` instead: the same text, as the same bytes on every platform, and a file whose content would not change is left untouched, so that what reads it does not run again for nothing - a redirection leaves the bytes to the shell, and empties the file before the tool has even started. A run that generates anyway writes it beside the artefacts with ``ddd generate --dictionary``, in the same write as them, which is what the :doc:`cmake integration ` does. The same file is what :doc:`comparing two deliveries ` needs: archive the dictionary of a delivery next to its binary, and a later ``ddd compare`` can answer whether the next delivery may replace it, long after the sources of the first one have moved on. What is already resolved ------------------------ The value of the dictionary is what a backend no longer has to work out. By the time a backend sees it: The **limits are filled in.** ``limits`` is never absent. A declaration that states its physical limits keeps them; one that does not gets the full range its datatype and conversion imply, computed once. A backend writing ``LOWER_LIMIT`` and ``UPPER_LIMIT`` into an a2l therefore never has to decide what a missing limit means, and cannot decide it differently from anybody else. The **shape is a tuple of dimensions**, empty for a scalar, and it is complete even when the description file never stated it. A measurement or a value block declares its ``dimensions`` and an axis its ``size``, but a curve and a map deliberately do not repeat the size of their tables: they name the axes they are interpolated over, and the shape follows from those. The demo project declares ``CurveA`` like this - no limits, no dimensions: .. code-block:: json { "scope": "local", "definition": { "kind": "curve", "name": "CurveA", "id": "6ywadt4ekhj3", "description": "Calibratable curve over AxisA", "datatype": "uint16", "unit": "ms", "conversion": { "factor": 0.01 }, "axis": "AxisA", "init": [1200, 900, 800, 750, 700, 650], "volatile": false } } and the dictionary hands the backends this: .. code-block:: json { "name": "CurveA", "id": "6ywadt4ekhj3", "extensions": {}, "kind": "curve", "datatype": "uint16", "description": "Calibratable curve over AxisA", "unit": "ms", "conversion": { "kind": "linear", "factor": 0.01, "offset": 0.0 }, "limits": { "min": 0.0, "max": 655.35 }, "shape": [6], "dimensions": [6], "init": [1200, 900, 800, 750, 700, 650], "section": null, "raster": null, "point_counts": "none", "volatile": false, "condition": null, "references": { "axis": "AxisA" }, "owner": "Controller", "consumers": [], "local": true, "a2l": { "export": true, "format": null, "display_identifier": null } } The six points come from ``AxisA``, and the limits from the full ``uint16`` range through the linear conversion: 65535 raw counts of 0.01 ms are 655.35 ms. ``dimensions`` is the same shape again, spelled the way the project spells it: here the number, and for an array dimensioned by a :doc:`declared constant ` the constant's name, so a generator can declare the array by the name while sizing it by the number. The ``kind`` of the conversion, left out of the description because a block carrying a ``factor`` can only be linear, is spelled out. Both blocks above are folded onto fewer lines than the files themselves use; the values are exactly the ones in ``examples/demo`` and in the dump of it. The **owner and the consumers are worked out.** ``owner`` names the component whose declaration was taken as the authoritative one, ``consumers`` lists the components that declared the object as an input, and ``local`` says whether the owner keeps it to itself. This is what lets the c backend group the definitions by owning component and emit a header per component that contains that component's interface and nothing else, and what lets the a2l backend build one ``GROUP`` per component with something to export - without either of them knowing anything about scopes, ownership rules or how a disagreement between two components is settled. Where components disagreed, the **producing component's declaration is the one that survives**: the analysis reports the disagreement against the deviating consumer and puts the producer's definition into the dictionary, so a backend never sees two versions of one object. The **condition is the producing declaration's.** A variable that only exists when a preprocessor symbol is defined carries that expression here, which is what the c backend wraps in ``#if`` and what the a2l backend notes in a comment, a2l having no notion of conditional compilation. The **objects are sorted by name** and the enumerations are collected, de-duplicated and sorted by name as well, so that a generated file depends on the content of a project and not on the order in which its files happened to be read. Together with the include patterns being expanded in sorted order and the generated files carrying no time stamp, that is what makes a regeneration without an input change produce a byte identical result - and therefore what lets a build system skip the recompilation. The components, by contrast, keep the order in which the project included them, and the declarations of a component keep the order the author wrote them in, because that order is information: it is how the interface of a component reads in its own file, and it is how it reads in its generated header. The **plugin blocks are resolved too.** Every object, every structured instance and the dictionary itself carry ``extensions``, validated and dumped back with each plugin's own defaults filled in, and the dictionary names the plugins in play under ``plugins``. That is what keeps a plugin's questions answerable from the dictionary alone, long after the project that named the plugin has moved on: the archived dump is what ``ddd compare --plugin`` reads back a plugin's comparison rules against, and what ``ddd generate `` hands to a plugin's own backend. See :doc:`plugins` for what a plugin does with them. .. note:: ``owner`` is ``null`` for an input that no component produces, which reaches the dictionary whenever ``missing-producer`` is relaxed - a component checked or dumped on its own, or an inconsistent project generated with ``ddd generate --force``; the c backend files such an object under ````. Every other field is always present. Point counts ahead of a table ----------------------------- Some firmware stores the number of axis points *inside* every curve, map and axis, ahead of its data, in the object's own type - a 16 by 16 ``uint32`` map begins with its two counts:: 8093841c 10000000 10000000 00000000 ... a 16 x 16 map of uint32 nx = 16 ny = 16 then the 256 values ``point_counts`` says whether an object is stored that way. It takes ``"none"``, the ordinary layout, or ``"leading"``. It may be stated once, on the project, as the default for every curve, map and axis; a component may state it again to override that default for the curves, maps and axes *it defines* - the producing declaration, exactly as a component's :doc:`raster ` does. A reader of the object never influences it. A component read on its own - a standalone ``ddd dump``, ``list`` or ``generate`` of a component file - has no project to take a default from, so its tables resolve to ``"none"`` unless the component states the key itself. A table resolved to ``"leading"`` is declared flat, counts first, in its own datatype: .. code-block:: c const uint32_t M[2 + (16) * (16)] = { 16U, 16U, ... }; and the a2l describes it with a record layout of its own, the counts ahead of the values it governs: .. code-block:: text /begin RECORD_LAYOUT RL_MAP_COUNTED_ULONG NO_AXIS_PTS_X 1 ULONG NO_AXIS_PTS_Y 2 ULONG FNC_VALUES 3 ULONG ROW_DIR DIRECT /end RECORD_LAYOUT Two :doc:`consistency checks ` come with the setting: ``point-counts-unrepresentable`` refuses a table whose datatype cannot hold one of its counts, and ``point-counts-mismatch`` warns when a curve or a map resolves to one convention and one of its axes to the other. From format 9 on, every resolved object carries the ``point_counts`` it was given, ``"none"`` for every kind but a curve, a map or an axis. The format field ---------------- A dumped dictionary is meant to be archived next to a delivery and read back by a later version of DDD, possibly years later. The ``format`` field stamps the shape of the document - currently ``9`` - and changes only when that shape changes, not with every release of the tool. It exists so that a later reader can say *this file is newer than I understand* rather than misread it. DDD accepts a dictionary whose format is the one it knows or older, and refuses one that is newer: .. code-block:: text $ ddd compare baseline.json demo.ddd.json baseline.json#format: error[schema]: in the baseline: this dictionary is in format 10, and this DDD understands up to 9; use a newer DDD to read it 1 error Refusing is the only safe answer: reading the file anyway would compare a delivery against fields this version does not know about, and quietly report every one of them as unchanged - which is precisely the verdict that would let a broken delivery out of the door. The rule is one-directional on purpose, so that a new DDD keeps reading the dictionaries archived by older ones. That is also why a field of the dictionary may keep a default the description files no longer allow. ``volatile`` has to be stated by every definition an author writes, but the dictionary still defaults it to ``false``, so a dictionary dumped by an older DDD still reads back and can still be compared against, instead of a required field turning every archived document into a file this version refuses. .. warning:: A dictionary is a snapshot of a project, not a description of it. ``ddd generate`` and ``ddd check`` read description files; the dictionary is what ``ddd compare`` reads back and what a foreign generator consumes. Treat it as an artefact of a build - archive it, do not edit it, and do not maintain a project in it. Consuming it from another tool ------------------------------ The document is plain json, described by a published schema, and it round-trips: a dictionary written by ``ddd dump`` and read back is the same dictionary, and a backend fed the reloaded document produces byte identical output. Both properties are asserted by the test suite (``tests/test_backends.py``), because they are the whole point of publishing the contract - a third party generating from a dumped dictionary has to get what DDD would have got. One field asks something of its reader: ``init`` is carried as the declaration wrote it, so a list is nested one level per dimension and a scalar written for an array is still a scalar - ``"init": 0`` beside ``"shape": [4]`` means four zeros. It is not expanded here because expanding it would write one literal per element for a value the description states once, and an array of ten thousand elements would carry ten thousand of them. A generator broadcasts the scalar over ``shape``, which is what the c backend does before rendering the initialiser. ``ddd schema dictionary`` prints the whole thing, definitions included; its top level, which is where a consumer starts, is this (the per field documentation the schema also carries is elided here for space): .. code-block:: json { "additionalProperties": false, "description": "The resolved data of one project.", "properties": { "format": { "default": 9, "minimum": 1, "title": "Format", "type": "integer" }, "name": { "maxLength": 128, "minLength": 1, "pattern": "^[A-Za-z_][A-Za-z0-9_]*$", "title": "Name", "type": "string" }, "description": { "default": "", "title": "Description", "type": "string" }, "source": { "default": "", "title": "Source", "type": "string" }, "components": { "default": [], "items": { "$ref": "#/$defs/ResolvedComponent" }, "title": "Components", "type": "array" }, "objects": { "default": [], "items": { "$ref": "#/$defs/ResolvedObject" }, "title": "Objects", "type": "array" }, "enums": { "default": [], "items": { "$ref": "#/$defs/EnumConversion" }, "title": "Enums", "type": "array" }, "constants": { "default": [], "items": { "$ref": "#/$defs/ConstantDeclaration" }, "title": "Constants", "type": "array" }, "rasters": { "default": [], "items": { "$ref": "#/$defs/ResolvedRaster" }, "title": "Rasters", "type": "array" }, "types": { "default": [], "items": { "$ref": "#/$defs/ResolvedStruct" }, "title": "Types", "type": "array" }, "instances": { "default": [], "items": { "$ref": "#/$defs/ResolvedInstance" }, "title": "Instances", "type": "array" }, "leaves": { "default": [], "items": { "$ref": "#/$defs/ResolvedLeaf" }, "title": "Leaves", "type": "array" }, "plugins": { "default": [], "items": { "type": "string" }, "title": "Plugins", "type": "array" }, "extensions": { "additionalProperties": { "additionalProperties": true, "type": "object" }, "title": "Extensions", "type": "object" } }, "required": [ "name" ], "title": "DDD data dictionary", "type": "object" } ``additionalProperties`` is ``false`` here, as it is on every object DDD describes - the one open door is ``extensions``, whose blocks belong to the plugins that own them - so a consumer validating against the schema finds a key it was not expecting instead of skipping it. ``enums`` carries the distinct enumerations the objects use, so a consumer that wants to emit a type per enumeration - which is what the c backend offers its templates as ``model.enums`` - does not have to walk every object and de-duplicate them itself. ``constants`` records the :doc:`declared constants ` whole - name, value and description - so a dimension an object spells by name stays resolvable from the document alone. ``rasters`` records the :doc:`declared measurement rasters ` the same way - name, event channel, period and description - so that a generator can write the XCP event list itself instead of reconstructing it from the events the objects happen to name. The conversions themselves are the same models the description files use, and they are documented with the other :doc:`data contracts `. ``types``, ``instances`` and ``leaves`` describe the structured variables: the declared structures, the variables instantiating one, and the member objects each instance flattens into. Reference --------- The dictionary is a pydantic model like every other contract, which means it is validated when the analysis hands it over: a bug in the front end surfaces at that boundary rather than half way through a jinja template. .. autopydantic_model:: ddd.ir.DataDictionary :field-show-constraints: False .. autopydantic_model:: ddd.ir.ResolvedObject :field-show-constraints: False .. autopydantic_model:: ddd.ir.ResolvedComponent :field-show-constraints: False .. autopydantic_model:: ddd.ir.ComponentDeclaration :field-show-constraints: False .. autopydantic_model:: ddd.ir.ResolvedRaster :field-show-constraints: False .. autopydantic_model:: ddd.ir.ResolvedStruct :field-show-constraints: False .. autopydantic_model:: ddd.ir.ResolvedMember :field-show-constraints: False .. autopydantic_model:: ddd.ir.ResolvedInstance :field-show-constraints: False .. autopydantic_model:: ddd.ir.ResolvedLeaf :field-show-constraints: False