Zum Inhalt

Element <ml-train>

Zweck: Trainiert und veröffentlicht ein konfiguriertes Machine-Learning-Artefakt aus Projektdaten.

Warum: Verwende dieses Element, um ein konfiguriertes Modellartefakt aus projekteigenen Daten zu trainieren und zu veröffentlichen.

Beispiel

1
<ml-train name="customer_churn" source="analyticsDb"/>

Entscheidungshilfe

Fachlicher Nutzen: Erzeugt ein versioniertes tabulares Modellartefakt für Use Cases mit gelernten Verteilungen.

  • Verwenden, wenn

    • Wenn ein gesteuertes Projekt ausdrücklich Training aus einem eigenen Datensatz verlangt.
  • Anderen Ansatz wählen, wenn

    • Wenn deterministische Regeln oder registrierte Generatoren den Datenvertrag ausdrücken können.
  • Voraussetzungen

    • Liefere autorisierte Quelle, Privacy-Ausschlüsse, Ressourcengrenzen und das ML-Testprofil.
  • Alternativen

    • Verwende deterministische generate-Erzeugung, wenn fachliche Regeln die Werte definieren. (Siehe: <generate>)

Vollständige Beispiele

Ein begrenztes tabellarisches Modell aus einem eigenen relationalen Dataset trainieren

Verwende ml-train nur, wenn benannte Quelldaten, auszuschließende sensible Felder und eine begrenzte Trainingspolicy explizite Modellanforderungen sind.

ml-training/data/customers.ent.csv
1
2
3
4
5
6
customer_id|segment|age
C-001|retail|31
C-002|business|48
C-003|retail|26
C-004|business|55
C-005|retail|39
ml-training/datamimic.xml
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
<setup defaultSeparator="|" numProcess="1">
    <ml-train name="customer_model"
              source="data/customers.ent.csv"
              maxTrainingTime="1"/>
    <generate name="synthetic_customers"
              count="10"
              source="ml://customer_model"
              distribution="ordered"
              target="LogExporter">
        <key name="model_source" constant="customer_model"/>
    </generate>
</setup>

Regeln und ungültige Kombinationen

W004 — Deprecated XML Attribute Ignored

<{element}> attribute '{attribute}' in descriptor '{descriptor}' is deprecated and ignored; execution continues. {migration}

Warum: The descriptor uses a known legacy attribute that no longer controls execution.

Lösung: Follow the migration hint when updating the model. Removing the attribute is not required to run it.

Vollständige Regel

Erlaubte Elternelemente / Erlaubte Kindelemente

Erlaubte Elternelemente: else, else-if, if, setup, while

Erlaubte Kindelemente:

Keine

Attribute

Alle 35 Attribute anzeigen

analyzeDropColumns

Columns to drop during analyze stage.

optional; string; Standardwert: null.

analyzeValueProtection

Enable analysis for value protection.

optional; boolean; Standardwert: null.

batchSize

Batch size for training.

optional; integer; Standardwert: null.

contextSource

Optional relational context source for two-table training.

optional; string; Standardwert: null.

contextType

Type of the relational context.

optional; string; Standardwert: null.

ctxPrimaryKey

Primary key column(s) in context; supports composite.

optional; string; Standardwert: null.

device

Computation device selector passed to the ML engine.

optional; string; Standardwert: null.

differentialPrivacy

Differential privacy configuration (JSON).

optional; string; Standardwert: null.

encodingTypes

Per-column encoding overrides.

optional; string; Standardwert: null; Werte: AUTO, TABULAR_CATEGORICAL, TABULAR_NUMERIC_AUTO, TABULAR_NUMERIC_DISCRETE, TABULAR_NUMERIC_BINNED, TABULAR_NUMERIC_DIGIT, TABULAR_CHARACTER, TABULAR_DATETIME, TABULAR_DATETIME_RELATIVE, TABULAR_LAT_LONG, LANGUAGE_TEXT, LANGUAGE_CATEGORICAL, LANGUAGE_NUMERIC, LANGUAGE_DATETIME.

fairness

Fairness configuration (JSON).

optional; string; Standardwert: null.

flexibleGeneration

Enable flexible generation mode.

optional; boolean; Standardwert: null.

generationBatchSize

Batch size during generation.

optional; integer; Standardwert: null.

gradientAccumulationSteps

Steps for gradient accumulation.

optional; integer; Standardwert: null.

imputation

Imputation configuration (JSON).

optional; string; Standardwert: null.

maxEpochs

Maximum number of training epochs.

optional; number; Standardwert: null.

maxSequenceWindow

Maximum sequence window size.

optional; integer; Standardwert: null.

maxTrainingTime

Maximum training time in minutes.

optional; string; Standardwert: null.

mode

Veralteter Trainingsmodus; beliebige alte mode-Werte werden mit Warnung akzeptiert und ignoriert und steuern keine Persistenz.

optional; string; Standardwert: null.

modelStateStrategy

Strategy for managing model state.

optional; string; Standardwert: null; Werte: reset, resume, reuse.

modelType

Model type/preset used by the engine.

optional; string; Standardwert: null; Werte: TABULAR, LANGUAGE.

name

Name of the ML model.

erforderlich; string.

rareCategoryReplacementMethod

Method for handling rare categories.

optional; string; Standardwert: null; Werte: constant, sample.

rebalancing

Class/feature rebalancing configuration (JSON).

optional; string; Standardwert: null.

samplingTemperature

Sampling temperature used for generation (0-2).

optional; number; Standardwert: null.

samplingTopP

Top-p (nucleus) sampling threshold (0-1).

optional; number; Standardwert: null.

sensitiveColumns

Columns to drop for privacy (comma-separated).

optional; string; Standardwert: null.

separator

Separator for generated data.

optional; string; Standardwert: null.

source

Training data source: a project file or configured source identifier.

erforderlich; string.

splitPartitions

Number of partitions for dataset split.

optional; integer; Standardwert: null.

targetColumn

Target column for supervised training.

optional; string; Standardwert: null.

tgtContextKey

Join key(s) linking target to context; supports composite.

optional; string; Standardwert: null.

tgtPrimaryKey

Primary key column(s) in target; supports composite.

optional; string; Standardwert: null.

trainModel

Training model identifier or preset.

optional; string; Standardwert: null.

trainValSplit

Train/validation split ratio (0-1).

optional; number; Standardwert: null.

type

Type of data to train on.

optional; string; Standardwert: null.