Zum Inhalt

Element <iterate>

Zweck: Durchläuft Datensätze aus einer erforderlichen Quelle und stellt den aktuellen Datensatz ohne eigenes Ausgabeziel in Kindkontexten bereit.

Warum: Verwende iterate, wenn untergeordnete Elemente jeden aktuellen Quelldatensatz über ihren Elternkontext verarbeiten.

Beispiel

1
<generate name="customers" source="data/customers.ent.csv"/>

Entscheidungshilfe

Fachlicher Nutzen: Macht die Quelldurchquerung explizit, während Kindkontexte den aktuellen Datensatz verwenden.

  • Verwenden, wenn

    • Wenn Kindelemente Felder jedes aktuellen Quelldatensatzes über ihren Elternkontext benötigen.
  • Anderen Ansatz wählen, wenn

    • Wenn die Operation direkt ein Produkt erzeugt oder exportiert.
  • Voraussetzungen

    • Gib eine Quelle und, falls vom Quellvertrag verlangt, eine explizite Begrenzung an.
  • Alternativen

    • Verwende generate, wenn die Operation ein Ausgabeziel besitzt oder ein Produkt erzeugt. (Siehe: <generate>)

Vollständige Beispiele

Ein In-Memory-Produkt durchlaufen und ein angereichertes Kindprodukt exportieren

Verwende iterate, wenn Kindkontexte jeden aktuellen Quelldatensatz verarbeiten; das verschachtelte generate besitzt das angereicherte Ausgabeprodukt und dessen Ziel.

iterate-and-enrich/datamimic.xml
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
<setup>
    <generate name="source_rows" count="4" target="mem">
        <key name="n" generator="IncrementGenerator"/>
    </generate>
    <iterate name="source_row"
             source="mem"
             type="source_rows"
             count="4"
             distribution="ordered">
        <generate name="enriched_rows" count="1" target="LogExporter">
            <key name="n" script="parent.n"/>
            <key name="doubled" script="n * 2"/>
            <nestedKey name="metrics" type="dict">
                <variable name="n" script="parent.n"/>
                <key name="squared" script="n * n"/>
            </nestedKey>
        </generate>
    </iterate>
</setup>

Regeln und ungültige Kombinationen

Optionen zur Quellenauswahl erfordern source.

Attribute: source, selector, separator, sourceScripted, cyclic, weightColumn, stratifyBy

Warum: selector, separator, sourceScripted, cyclic, weightColumn und stratifyBy verändern ausschließlich einen Quellenzugriff.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" count="10"/>
Ungültige Kombination
1
<generate name="customers" count="10" selector="select * from customers"/>
Eine source darf type oder selector angeben, aber nicht beides.

Attribute: source, type, selector

Warum: selector bestimmt die Abfrageform, während type eine nicht selektierte Quell-Entity bezeichnet.

Gültige Kombination
1
<generate name="customers" source="customerDb" type="Customer" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="customerDb" type="Customer" selector="select * from customers" count="10"/>
Verwende entweder count oder die zufällige minCount/maxCount-Strategie, niemals beides.

Attribute: count, minCount, maxCount

Warum: Beide Strategien bestimmen die Datensatzanzahl; ihre Kombination würde die Ausführung mehrdeutig machen.

Gültige Kombination
1
<generate name="customers" count="10"/>
Ungültige Kombination
1
<generate name="customers" count="10" minCount="2"/>
minCount muss kleiner oder gleich maxCount sein.

Attribute: minCount, maxCount

Warum: Die zufällige Anzahl wird aus dem inklusiven Intervall zwischen diesen Grenzen gezogen.

Gültige Kombination
1
<generate name="customers" minCount="2" maxCount="10"/>
Ungültige Kombination
1
<generate name="customers" minCount="10" maxCount="2"/>
minCount/maxCount dürfen nicht mit start/end/interval kombiniert werden.

Attribute: minCount, maxCount, start, end, interval

Warum: Die Länge einer Zeitreihe ergibt sich aus Zeitfenster und Intervall, nicht aus einer zufälligen Datensatzanzahl.

Gültige Kombination
1
<generate name="events" start="2026-01-01T00:00:00Z" end="2026-01-02T00:00:00Z" interval="PT1H"/>
Ungültige Kombination
1
<generate name="events" minCount="2" start="2026-01-01T00:00:00Z" end="2026-01-02T00:00:00Z" interval="PT1H"/>
offset erfordert source.

Attribute: offset, source

Warum: Ein Offset überspringt Quelldatensätze und hat deshalb für rein synthetische Generierung keine Bedeutung.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" offset="2" count="10"/>
Ungültige Kombination
1
<generate name="customers" count="10" offset="2"/>
resume='group' erfordert resumeGroup.

Attribute: resume, resumeGroup

Warum: Die Runtime benötigt den gemeinsamen Gruppennamen, um den Fortsetzungszustand aufzulösen.

Gültige Kombination
1
<generate name="customers" count="10" resume="group" resumeGroup="customer-import"/>
Ungültige Kombination
1
<generate name="customers" count="10" resume="group"/>
resumeGroup erfordert resume='group'.

Attribute: resume, resumeGroup

Warum: Der Gruppenname hat für statement- oder tabellenbezogene Fortsetzung keine Bedeutung.

Gültige Kombination
1
<generate name="customers" count="10" resume="group" resumeGroup="customer-import"/>
Ungültige Kombination
1
<generate name="customers" count="10" resume="stmt" resumeGroup="customer-import"/>
unique erfordert eine endliche source.

Attribute: unique, source

Warum: Eine eindeutige Stichprobe ohne Zurücklegen benötigt einen Quellpool.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" unique="true" count="10"/>
Ungültige Kombination
1
<generate name="customers" count="10" unique="true"/>
unique darf nicht mit cyclic kombiniert werden.

Attribute: unique, cyclic

Warum: Unique-Auswahl verbraucht Datensätze ohne Zurücklegen, während cyclic erschöpfte Datensätze wiederholt.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" unique="true" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" count="10" unique="true" cyclic="true"/>
unique lässt sich explizit nur mit distribution='random' kombinieren.

Attribute: unique, distribution

Warum: Geordnete oder gewichtete Auswahl widerspricht einer eindeutigen Zufallsstichprobe ohne Zurücklegen.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" unique="true" distribution="random" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" count="10" unique="true" distribution="ordered"/>
cyclic erfordert eine Distribution, die Wiederholung unterstützt.

Attribute: cyclic, distribution

Warum: Vollständige gewichtete und stratifizierte Ziehungen besitzen keine stabile Erschöpfungsgrenze für einen Neustart.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" cyclic="true" distribution="ordered" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" count="10" cyclic="true" distribution="weighted" weightColumn="weight"/>
distribution='weighted' erfordert weightColumn.

Attribute: distribution, weightColumn

Warum: Die Auswahl benötigt eine nicht negative numerische Spalte zur Berechnung der Stichprobenwahrscheinlichkeiten.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" distribution="weighted" weightColumn="weight" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" count="10" distribution="weighted"/>
distribution='stratified' erfordert stratifyBy.

Attribute: distribution, stratifyBy

Warum: Die Auswahl benötigt eine Quellspalte, die das jeweilige Stratum identifiziert.

Gültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" distribution="stratified" stratifyBy="segment" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.ent.csv" count="10" distribution="stratified"/>
Eine semantische Projektdatei-Quelle erfordert ihr katalogisiertes Metadatenattribut.

Attribute: source, weightColumn

Warum: Gewichtete Entity-Zeilen benötigen weightColumn, damit die Auswahl die Metadatenspalte verwenden und anschließend entfernen kann.

Gültige Kombination
1
<generate name="customers" source="data/customers.wgt.ent.csv" distribution="weighted" weightColumn="weight" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.wgt.ent.csv" count="10"/>
Eine semantische Projektdatei-Quelle erfordert ihre katalogisierte Distribution.

Attribute: source, distribution

Warum: Der Dateisuffix deklariert Auswahlsemantik, die eine widersprechende explizite Distribution nicht überschreiben darf.

Gültige Kombination
1
<generate name="customers" source="data/customers.wgt.ent.csv" distribution="weighted" weightColumn="weight" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.wgt.ent.csv" count="10" weightColumn="weight" distribution="ordered"/>
Eine gewichtete Entity-Projektquelle erfordert explizites count.

Attribute: source, count

Warum: Gewichtete Auswahl leitet keine begrenzte Ausgabegröße aus der Quellenlänge ab.

Gültige Kombination
1
<generate name="customers" source="data/customers.wgt.ent.csv" distribution="weighted" weightColumn="weight" count="10"/>
Ungültige Kombination
1
<generate name="customers" source="data/customers.wgt.ent.csv" weightColumn="weight"/>
W004 — Deprecated XML Attribute Ignored

<{element}> attribute '{attribute}' in descriptor '{descriptor}' is deprecated and ignored; execution continues. {migration}

Warum: The descriptor uses a known legacy attribute that no longer controls execution.

Lösung: Follow the migration hint when updating the model. Removing the attribute is not required to run it.

Vollständige Regel

I883 — Time-Series Bad Start

Invalid time-series configuration: {detail}

Warum: The 'start' attribute is not a valid ISO 8601 datetime.

Lösung: Set start to an ISO 8601 datetime, e.g. '2026-01-01T00:00:00Z'.

Vollständige Regel

I884 — Time-Series Bad End

Invalid time-series configuration: {detail}

Warum: The 'end' attribute is not a valid ISO 8601 datetime.

Lösung: Set end to an ISO 8601 datetime, e.g. '2026-01-02T00:00:00Z'.

Vollständige Regel

I885 — Time-Series Bad Interval

Invalid time-series configuration: {detail}

Warum: The 'interval' attribute is not a valid ISO 8601 duration.

Lösung: Set interval to an ISO 8601 duration, e.g. 'PT1H'.

Vollständige Regel

I886 — Time-Series End Not After Start

Invalid time-series configuration: {detail}

Warum: The 'end' attribute is not strictly after the 'start' attribute.

Lösung: Set end to a datetime strictly after start.

Vollständige Regel

I887 — Time-Series Interval Too Fine

Invalid time-series configuration: {detail}

Warum: The 'interval' duration is non-positive, sub-microsecond, or a months/years Duration.

Lösung: Set interval to a positive constant-length duration of at least 1us, e.g. 'PT1S'.

Vollständige Regel

I888 — Time-Series Namespace Collision

is not allowed inside a time-series : 'ts' is reserved for the time-iterator namespace (ts.now/ts.step/ts.series). Rename the variable, e.g. 'ts_meta'.

Warum: A would shadow the reserved time-iterator namespace.

Lösung: Rename the variable to something other than 'ts', e.g. 'ts_meta'.

Vollständige Regel

I889 — Time-Series Incomplete Config

Time-series attributes start/end/interval must be set together; missing: {missing}

Warum: Only some of the time-series attributes start/end/interval were provided.

Lösung: Provide all three attributes (start, end, interval) or none of them.

Vollständige Regel

I949 — Source ML Model Option Unsupported

ML model source '{source}' does not support {option}={value}. Supported behavior: {supported}.

Warum: The requested source-selection option cannot be preserved by generated ml:// model samples.

Lösung: Remove the unsupported option and use bounded ordered ML model generation.

Vollständige Regel

I192 — Source Reference Identifier Empty

Explicit source URI for family '{family}' requires a non-empty identifier

Warum: A recognized source-family URI was declared without the identifier needed to resolve its source.

Lösung: Add the source identifier after the URI scheme and retry.

Vollständige Regel

I194 — Source Reference Client Family Mismatch

Explicit source family '{family}' does not match configured client '{client_id}' of type '{actual_client_type}'

Warum: The referenced client exists but does not implement the source family declared by the URI.

Lösung: Use the URI scheme matching the configured client or reference a client of the declared family.

Vollständige Regel

I195 — Source Reference Identifier Conflict

Explicit source identifier '{identifier}' conflicts with {attribute}='{configured_identifier}'

Warum: Two source attributes select different identities for the same explicit source reference.

Lösung: Remove the legacy override or make it equal to the identifier in source.

Vollständige Regel

Erlaubte Elternelemente / Erlaubte Kindelemente

Erlaubte Elternelemente: else, else-if, generate, if, iterate, setup, while

Erlaubte Kindelemente:

array, assert, condition, echo, generate, id, include, iterate, key, list, mapping, nestedKey, reference, rule, sourceConstraints, targetConstraints, variable, while

Attribute

Alle 46 Attribute anzeigen

bucket

Bucket name for the external source or target.

optional; string; Standardwert: null.

container

Container name for the external source or target.

optional; string; Standardwert: null.

converter

Converter for element data transformation.

optional; string; Standardwert: null.

count

Number of records to generate.

optional; string; Standardwert: null.

cyclic

Enable or disable cyclic generation.

optional; boolean; Standardwert: null.

device

Computation device to use.

optional; string; Standardwert: null.

distribution

Distribution type for data generation.

optional; string; Standardwert: null; Werte: ordered, random, weighted, stratified, round_robin, reservoir, cumulated.

encoding

Override encoding for this generate task.

optional; string; Standardwert: null.

end

Time-series window end (ISO 8601 datetime).

optional; string; Standardwert: null.

fairness

Fairness configuration (JSON).

optional; string; Standardwert: null.

generationBatchSize

Batch size during generation.

optional; integer; Standardwert: null.

imputation

Imputation configuration (JSON).

optional; string; Standardwert: null.

interval

Time-series tick interval (ISO 8601 duration).

optional; string; Standardwert: null.

iterationSelector

Selector evaluated per iteration.

optional; string; Standardwert: null.

maxCount

Maximum count for randomized generate/iterate length (mutually exclusive with 'count').

optional; integer; Standardwert: null.

minCount

Minimum count for randomized generate/iterate length (mutually exclusive with 'count').

optional; integer; Standardwert: null.

mpPlatform

Multiprocessing platform override.

optional; string; Standardwert: null; Werte: multiprocessing, fork, spawn, forkserver.

multiprocessing

Deprecated: accepted with a warning and ignored. Use numProcess and mpPlatform instead.

optional; string; Standardwert: null.

name

Name of the generation task.

erforderlich; string.

numProcess

Specify the number of processes to use.

optional; integer; Standardwert: null.

offset

Skip the first N rows of the source.

optional; integer; Standardwert: null.

pageSize

Page size for processing data.

optional; integer; Standardwert: null.

page_bytes_cap

Hard cap on page size in bytes.

optional; integer; Standardwert: null.

page_memory_cap_mb

Cap in megabytes for page memory usage.

optional; integer; Standardwert: null.

rareCategoryReplacementMethod

Method for handling rare categories.

optional; string; Standardwert: null; Werte: constant, sample.

rebalancing

Class/feature rebalancing configuration (JSON).

optional; string; Standardwert: null.

resume

Optional per-statement resume scope override.

optional; string; Standardwert: null; Werte: stmt, table, group.

resumeGroup

Logical resume group key used when resume='group'.

optional; string; Standardwert: null.

samplingTemperature

Sampling temperature used for generation (0-2).

optional; number; Standardwert: null.

samplingTopP

Top-p (nucleus) sampling threshold (0-1).

optional; number; Standardwert: null.

script

Script driving generation logic.

optional; string; Standardwert: null.

selector

Selector for data generation.

optional; string; Standardwert: null.

separator

Separator for generated data.

optional; string; Standardwert: null.

source

Canonical source URI (for example file://data/orders.csv, database://sourceDb, or ml://customer_model); legacy raw source values remain accepted.

erforderlich; string; Standardwert: null.

sourceClient

Override the client used to read when source is ambiguous.

optional; string; Standardwert: null.

sourceEntity

Explicit physical entity to read (table/collection/product). Overrides selector or inferred name.

optional; string; Standardwert: null.

sourceScripted

Enable or disable scripted sources.

optional; boolean; Standardwert: null.

sourceUri

URI backing the source (file/object storage).

optional; string; Standardwert: null.

start

Time-series window start (ISO 8601 datetime).

optional; string; Standardwert: null.

storageId

Object storage client id.

optional; string; Standardwert: null.

stratifyBy

Stratum column for distribution='stratified' source generation.

optional; string; Standardwert: null.

type

Type of data generation.

optional; string; Standardwert: null.

unique

Emit each source row at most once (distinct selection without replacement).

optional; boolean; Standardwert: null.

variablePrefix

Prefix before field's name for query select data in selector element

optional; string; Standardwert: null.

variableSuffix

Suffix after field's name for query select data in selector element

optional; string; Standardwert: null.

weightColumn

Weight column for distribution='weighted' source generation; required for a .wgt.ent.csv source.

optional; string; Standardwert: null.