Zum Inhalt

Element <generate>

Zweck: Generiert, transformiert, iteriert und exportiert optional einen Produktstrom.

Warum: Verwende dieses Element für einen begrenzten Produktstrom, der Datensätze generiert, transformiert oder exportiert.

Beispiel

1
<generate name="customers" count="1"/>

Entscheidungshilfe

Fachlicher Nutzen: Überführt synthetische oder quellabgeleitete Datensätze in einen begrenzten, optional exportierten Produktstrom.

  • Verwenden, wenn

    • Beim Erzeugen eines festen synthetischen Datensatzes.
    • Beim Transformieren einer begrenzten Quelle in ein oder mehrere Ziele.
    • Beim Erzeugen eines deterministischen Zeitfensters.
  • Anderen Ansatz wählen, wenn

    • Beim Durchlaufen einer Quelle ohne Ausgabeziel.
    • Bei der Mutation vorhandener Daten über einen Operation-Control-Workflow.
  • Voraussetzungen

    • Gib einen Produktnamen und genau eine eindeutige Ausführungsbasis an: count, source oder Zeitfenster.
  • Alternativen

    • Verwende iterate für reines Quelldurchlaufen ohne Ausgabeziel. (Siehe: <iterate>)
    • Verwende operate für geprüfte Mutationen vorhandener Daten. (Siehe: <operate>)

Vollständige Beispiele

Ein synthetisches CSV-Artefakt erzeugen

Verwende dieses minimale vollständige Modell, wenn keine aus Produktionsdaten abgeleiteten Quelldaten benötigt werden.

synthetic-csv/datamimic.xml
1
2
3
4
5
6
<setup>
    <generate name="customers" count="10" target="LogExporter">
        <id name="id" generator="IncrementGenerator"/>
        <key name="email" generator="EmailAddressGenerator"/>
    </generate>
</setup>
Eine geordnete Entity-CSV-Quelle transformieren

Verwende diese Form, wenn die Quellreihenfolge beim Erzeugen eines Platform-Artefakts erhalten bleiben muss.

ordered-csv-source/data/customers.ent.csv
1
2
3
4
id|name
1|Ada
2|Grace
3|Linus
ordered-csv-source/datamimic.xml
1
2
3
4
5
6
7
<setup defaultSeparator="|">
    <generate name="ordered_customers"
              source="data/customers.ent.csv"
              count="3"
              distribution="ordered"
              target="LogExporter"/>
</setup>
Eine begrenzte Memstore-Quelle zyklisch wiederverwenden

Verwende cyclic nur, wenn bewusste Wiederholung dem Stoppen nach Erschöpfung des Pools vorzuziehen ist.

cyclic-memstore/datamimic.xml
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
<setup>
    <generate name="seed_customers" count="3" target="mem">
        <id name="id" generator="IncrementGenerator"/>
    </generate>
    <generate name="cycled_customers"
              type="seed_customers"
              source="mem"
              count="8"
              cyclic="true"
              distribution="ordered"
              target="LogExporter"/>
</setup>
Eine begrenzte Zeitreihe erzeugen

Verwende start/end/interval, wenn Datensatzpositionen deterministische Zeitpunkte darstellen.

time-series/datamimic.xml
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
<setup>
    <generate name="events"
              start="2026-01-01T00:00:00+00:00"
              end="2026-01-02T00:00:00+00:00"
              interval="PT1H"
              target="LogExporter">
        <key name="timestamp" script="ts.now.isoformat()"/>
        <key name="step" script="ts.step"/>
    </generate>
</setup>

Regeln und ungültige Kombinationen

Generierung erfordert eine Anzahlstrategie, source, script, ein Zeitreihenintervall oder ein Operationsziel.

Attribute: count, minCount, maxCount, source, script, interval, target

Warum: Ohne Eingabe oder begrenzte Generierungsstrategie kann die Runtime weder Arbeit noch Kardinalität bestimmen.

Gültige Kombination
1
<generate name="customers" count="10"/>
Ungültige Kombination
1
<generate name="customers"/>
Optionen zur Quellenauswahl erfordern source.

Attribute: source, selector, separator, sourceScripted, cyclic, weightColumn, stratifyBy

Warum: selector, separator, sourceScripted, cyclic, weightColumn und stratifyBy verändern ausschließlich einen Quellenzugriff.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" count="10"/>
Ungültige Kombination
1
<generate name="customers" count="10" selector="select * from customers"/>
Eine source darf type oder selector angeben, aber nicht beides.

Attribute: source, type, selector

Warum: selector bestimmt die Abfrageform, während type eine nicht selektierte Quell-Entity bezeichnet.

Gültige Kombination
1
<generate name="customers" source="customerDb" type="Customer" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="customerDb" type="Customer" selector="select * from customers" count="10"/>
Verwende entweder count oder die zufällige minCount/maxCount-Strategie, niemals beides.

Attribute: count, minCount, maxCount

Warum: Beide Strategien bestimmen die Datensatzanzahl; ihre Kombination würde die Ausführung mehrdeutig machen.

Gültige Kombination
1
<generate name="customers" count="10"/>
Ungültige Kombination
1
<generate name="customers" count="10" minCount="2"/>
minCount muss kleiner oder gleich maxCount sein.

Attribute: minCount, maxCount

Warum: Die zufällige Anzahl wird aus dem inklusiven Intervall zwischen diesen Grenzen gezogen.

Gültige Kombination
1
<generate name="customers" minCount="2" maxCount="10"/>
Ungültige Kombination
1
<generate name="customers" minCount="10" maxCount="2"/>
minCount/maxCount dürfen nicht mit start/end/interval kombiniert werden.

Attribute: minCount, maxCount, start, end, interval

Warum: Die Länge einer Zeitreihe ergibt sich aus Zeitfenster und Intervall, nicht aus einer zufälligen Datensatzanzahl.

Gültige Kombination
1
<generate name="events" start="2026-01-01T00:00:00Z" end="2026-01-02T00:00:00Z" interval="PT1H"/>
Ungültige Kombination
1
<generate name="events" minCount="2" start="2026-01-01T00:00:00Z" end="2026-01-02T00:00:00Z" interval="PT1H"/>
offset erfordert source.

Attribute: offset, source

Warum: Ein Offset überspringt Quelldatensätze und hat deshalb für rein synthetische Generierung keine Bedeutung.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" offset="2" count="10"/>
Ungültige Kombination
1
<generate name="customers" count="10" offset="2"/>
resume='group' erfordert resumeGroup.

Attribute: resume, resumeGroup

Warum: Die Runtime benötigt den gemeinsamen Gruppennamen, um den Fortsetzungszustand aufzulösen.

Gültige Kombination
1
<generate name="customers" count="10" resume="group" resumeGroup="customer-import"/>
Ungültige Kombination
1
<generate name="customers" count="10" resume="group"/>
resumeGroup erfordert resume='group'.

Attribute: resume, resumeGroup

Warum: Der Gruppenname hat für statement- oder tabellenbezogene Fortsetzung keine Bedeutung.

Gültige Kombination
1
<generate name="customers" count="10" resume="group" resumeGroup="customer-import"/>
Ungültige Kombination
1
<generate name="customers" count="10" resume="stmt" resumeGroup="customer-import"/>
unique erfordert eine endliche source.

Attribute: unique, source

Warum: Eine eindeutige Stichprobe ohne Zurücklegen benötigt einen Quellpool.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" unique="true" count="10"/>
Ungültige Kombination
1
<generate name="customers" count="10" unique="true"/>
unique darf nicht mit cyclic kombiniert werden.

Attribute: unique, cyclic

Warum: Unique-Auswahl verbraucht Datensätze ohne Zurücklegen, während cyclic erschöpfte Datensätze wiederholt.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" unique="true" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" count="10" unique="true" cyclic="true"/>
unique lässt sich explizit nur mit distribution='random' kombinieren.

Attribute: unique, distribution

Warum: Geordnete oder gewichtete Auswahl widerspricht einer eindeutigen Zufallsstichprobe ohne Zurücklegen.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" unique="true" distribution="random" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" count="10" unique="true" distribution="ordered"/>
cyclic erfordert eine Distribution, die Wiederholung unterstützt.

Attribute: cyclic, distribution

Warum: Vollständige gewichtete und stratifizierte Ziehungen besitzen keine stabile Erschöpfungsgrenze für einen Neustart.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" cyclic="true" distribution="ordered" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" count="10" cyclic="true" distribution="weighted" weightColumn="weight"/>
distribution='weighted' erfordert weightColumn.

Attribute: distribution, weightColumn

Warum: Die Auswahl benötigt eine nicht negative numerische Spalte zur Berechnung der Stichprobenwahrscheinlichkeiten.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" distribution="weighted" weightColumn="weight" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" count="10" distribution="weighted"/>
distribution='stratified' erfordert stratifyBy.

Attribute: distribution, stratifyBy

Warum: Die Auswahl benötigt eine Quellspalte, die das jeweilige Stratum identifiziert.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" distribution="stratified" stratifyBy="segment" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" count="10" distribution="stratified"/>
Eine semantische Projektdatei-Quelle erfordert ihr katalogisiertes Metadatenattribut.

Attribute: source, weightColumn

Warum: Gewichtete Entity-Zeilen benötigen weightColumn, damit die Auswahl die Metadatenspalte verwenden und anschließend entfernen kann.

Gültige Kombination
1
<generate name="customers" source="data/customers.wgt.ent.csv" distribution="weighted" weightColumn="weight" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.wgt.ent.csv" count="10"/>
Eine semantische Projektdatei-Quelle erfordert ihre katalogisierte Distribution.

Attribute: source, distribution

Warum: Der Dateisuffix deklariert Auswahlsemantik, die eine widersprechende explizite Distribution nicht überschreiben darf.

Gültige Kombination
1
<generate name="customers" source="data/customers.wgt.ent.csv" distribution="weighted" weightColumn="weight" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.wgt.ent.csv" count="10" weightColumn="weight" distribution="ordered"/>
Eine gewichtete Entity-Projektquelle erfordert explizites count.

Attribute: source, count

Warum: Gewichtete Auswahl leitet keine begrenzte Ausgabegröße aus der Quellenlänge ab.

Gültige Kombination
1
<generate name="customers" source="data/customers.wgt.ent.csv" distribution="weighted" weightColumn="weight" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.wgt.ent.csv" weightColumn="weight"/>
W004 — Deprecated XML Attribute Ignored

<{element}> attribute '{attribute}' in descriptor '{descriptor}' is deprecated and ignored; execution continues. {migration}

Warum: The descriptor uses a known legacy attribute that no longer controls execution.

Lösung: Follow the migration hint when updating the model. Removing the attribute is not required to run it.

Vollständige Regel

I883 — Time-Series Bad Start

Invalid time-series configuration: {detail}

Warum: The 'start' attribute is not a valid ISO 8601 datetime.

Lösung: Set start to an ISO 8601 datetime, e.g. '2026-01-01T00:00:00Z'.

Vollständige Regel

I884 — Time-Series Bad End

Invalid time-series configuration: {detail}

Warum: The 'end' attribute is not a valid ISO 8601 datetime.

Lösung: Set end to an ISO 8601 datetime, e.g. '2026-01-02T00:00:00Z'.

Vollständige Regel

I885 — Time-Series Bad Interval

Invalid time-series configuration: {detail}

Warum: The 'interval' attribute is not a valid ISO 8601 duration.

Lösung: Set interval to an ISO 8601 duration, e.g. 'PT1H'.

Vollständige Regel

I886 — Time-Series End Not After Start

Invalid time-series configuration: {detail}

Warum: The 'end' attribute is not strictly after the 'start' attribute.

Lösung: Set end to a datetime strictly after start.

Vollständige Regel

I887 — Time-Series Interval Too Fine

Invalid time-series configuration: {detail}

Warum: The 'interval' duration is non-positive, sub-microsecond, or a months/years Duration.

Lösung: Set interval to a positive constant-length duration of at least 1us, e.g. 'PT1S'.

Vollständige Regel

I888 — Time-Series Namespace Collision

is not allowed inside a time-series : 'ts' is reserved for the time-iterator namespace (ts.now/ts.step/ts.series). Rename the variable, e.g. 'ts_meta'.

Warum: A would shadow the reserved time-iterator namespace.

Lösung: Rename the variable to something other than 'ts', e.g. 'ts_meta'.

Vollständige Regel

I889 — Time-Series Incomplete Config

Time-series attributes start/end/interval must be set together; missing: {missing}

Warum: Only some of the time-series attributes start/end/interval were provided.

Lösung: Provide all three attributes (start, end, interval) or none of them.

Vollständige Regel

I949 — Source ML Model Option Unsupported

ML model source '{source}' does not support {option}={value}. Supported behavior: {supported}.

Warum: The requested source-selection option cannot be preserved by generated ml:// model samples.

Lösung: Remove the unsupported option and use bounded ordered ML model generation.

Vollständige Regel

I192 — Source Reference Identifier Empty

Explicit source URI for family '{family}' requires a non-empty identifier

Warum: A recognized source-family URI was declared without the identifier needed to resolve its source.

Lösung: Add the source identifier after the URI scheme and retry.

Vollständige Regel

I194 — Source Reference Client Family Mismatch

Explicit source family '{family}' does not match configured client '{client_id}' of type '{actual_client_type}'

Warum: The referenced client exists but does not implement the source family declared by the URI.

Lösung: Use the URI scheme matching the configured client or reference a client of the declared family.

Vollständige Regel

I195 — Source Reference Identifier Conflict

Explicit source identifier '{identifier}' conflicts with {attribute}='{configured_identifier}'

Warum: Two source attributes select different identities for the same explicit source reference.

Lösung: Remove the legacy override or make it equal to the identifier in source.

Vollständige Regel

Erlaubte Elternelemente / Erlaubte Kindelemente

Erlaubte Elternelemente: else, else-if, generate, if, iterate, setup, while

Erlaubte Kindelemente:

array, assert, condition, echo, generate, id, include, iterate, key, list, mapping, nestedKey, reference, rule, sourceConstraints, targetConstraints, variable, while

Attribute

Alle 50 Attribute anzeigen

bucket

Bucket name for the external source or target.

optional; string; Standardwert: null.

container

Container name for the external source or target.

optional; string; Standardwert: null.

converter

Converter for element data transformation.

optional; string; Standardwert: null.

count

Number of records to generate.

optional; string; Standardwert: null.

cyclic

Enable or disable cyclic generation.

optional; boolean; Standardwert: null.

device

Computation device to use.

optional; string; Standardwert: null.

distribution

Distribution type for data generation.

optional; string; Standardwert: null; Werte: ordered, random, weighted, stratified, round_robin, reservoir, cumulated.

encoding

Override encoding for this generate task.

optional; string; Standardwert: null.

end

Time-series window end (ISO 8601 datetime).

optional; string; Standardwert: null.

exportUri

Explicit URI for exporters that support file paths.

optional; string; Standardwert: null.

fairness

Fairness configuration (JSON).

optional; string; Standardwert: null.

generationBatchSize

Batch size during generation.

optional; integer; Standardwert: null.

imputation

Imputation configuration (JSON).

optional; string; Standardwert: null.

interval

Time-series tick interval (ISO 8601 duration).

optional; string; Standardwert: null.

iterationSelector

Selector evaluated per iteration.

optional; string; Standardwert: null.

maxCount

Maximum count for randomized generate/iterate length (mutually exclusive with 'count').

optional; integer; Standardwert: null.

minCount

Minimum count for randomized generate/iterate length (mutually exclusive with 'count').

optional; integer; Standardwert: null.

mpPlatform

Multiprocessing platform override.

optional; string; Standardwert: null; Werte: multiprocessing, fork, spawn, forkserver.

multiprocessing

Beliebiger alter Parallelitäts-String; wird mit Warnung akzeptiert und ignoriert. Verwende numProcess und mpPlatform.

optional; string; Standardwert: null.

name

Name of the generation task.

erforderlich; string.

numProcess

Specify the number of processes to use.

optional; integer; Standardwert: null.

offset

Skip the first N rows of the source.

optional; integer; Standardwert: null.

pageSize

Page size for processing data.

optional; integer; Standardwert: null.

page_bytes_cap

Hard cap on page size in bytes.

optional; integer; Standardwert: null.

page_memory_cap_mb

Cap in megabytes for page memory usage.

optional; integer; Standardwert: null.

rareCategoryReplacementMethod

Method for handling rare categories.

optional; string; Standardwert: null; Werte: constant, sample.

rebalancing

Class/feature rebalancing configuration (JSON).

optional; string; Standardwert: null.

resume

Optional per-statement resume scope override.

optional; string; Standardwert: null; Werte: stmt, table, group.

resumeGroup

Logical resume group key used when resume='group'.

optional; string; Standardwert: null.

samplingTemperature

Sampling temperature used for generation (0-2).

optional; number; Standardwert: null.

samplingTopP

Top-p (nucleus) sampling threshold (0-1).

optional; number; Standardwert: null.

script

Script driving generation logic.

optional; string; Standardwert: null.

selector

Selector for data generation.

optional; string; Standardwert: null.

separator

Separator for generated data.

optional; string; Standardwert: null.

source

Canonical source URI (for example file://data/orders.csv, database://sourceDb, or ml://customer_model); legacy raw source values remain accepted.

optional; string; Standardwert: null.

sourceClient

Override the client used to read when source is ambiguous.

optional; string; Standardwert: null.

sourceEntity

Explicit physical entity to read (table/collection/product). Overrides selector or inferred name.

optional; string; Standardwert: null.

sourceScripted

Enable or disable scripted sources.

optional; boolean; Standardwert: null.

sourceUri

URI backing the source (file/object storage).

optional; string; Standardwert: null.

start

Time-series window start (ISO 8601 datetime).

optional; string; Standardwert: null.

storageId

Object storage client id.

optional; string; Standardwert: null.

stratifyBy

Stratum column for distribution='stratified' source generation.

optional; string; Standardwert: null.

target

Target output or explicit client operation for generated data.

optional; string; Standardwert: null.

targetClient

Override the client used to write/operate when target is ambiguous.

optional; string; Standardwert: null.

targetEntity

Explicit physical entity to write or operate on (table, collection, or product).

optional; string; Standardwert: null.

type

Type of data generation.

optional; string; Standardwert: null.

unique

Emit each source row at most once (distinct selection without replacement).

optional; boolean; Standardwert: null.

variablePrefix

Prefix before field's name for query select data in selector element

optional; string; Standardwert: null.

variableSuffix

Suffix after field's name for query select data in selector element

optional; string; Standardwert: null.

weightColumn

Weight column for distribution='weighted' source generation; required for a .wgt.ent.csv source.

optional; string; Standardwert: null.