Element <ml-train>¶
Zweck: Trainiert und veröffentlicht ein konfiguriertes Machine-Learning-Artefakt aus Projektdaten.
Warum: Verwende dieses Element, um ein konfiguriertes Modellartefakt aus projekteigenen Daten zu trainieren und zu veröffentlichen.
Beispiel¶
1 | |
Entscheidungshilfe¶
Fachlicher Nutzen: Erzeugt ein versioniertes tabulares Modellartefakt für Use Cases mit gelernten Verteilungen.
-
Verwenden, wenn
- Wenn ein gesteuertes Projekt ausdrücklich Training aus einem eigenen Datensatz verlangt.
-
Anderen Ansatz wählen, wenn
- Wenn deterministische Regeln oder registrierte Generatoren den Datenvertrag ausdrücken können.
-
Voraussetzungen
- Liefere autorisierte Quelle, Privacy-Ausschlüsse, Ressourcengrenzen und das ML-Testprofil.
-
Alternativen
- Verwende deterministische generate-Erzeugung, wenn fachliche Regeln die Werte definieren. (Siehe:
<generate>)
- Verwende deterministische generate-Erzeugung, wenn fachliche Regeln die Werte definieren. (Siehe:
Vollständige Beispiele¶
Ein begrenztes tabellarisches Modell aus einem eigenen relationalen Dataset trainieren
Verwende ml-train nur, wenn benannte Quelldaten, auszuschließende sensible Felder und eine begrenzte Trainingspolicy explizite Modellanforderungen sind.
| ml-training/data/customers.ent.csv | |
|---|---|
1 2 3 4 5 6 | |
| ml-training/datamimic.xml | |
|---|---|
1 2 3 4 5 6 7 8 9 10 11 12 | |
Regeln und ungültige Kombinationen¶
W004 — Deprecated XML Attribute Ignored
<{element}> attribute '{attribute}' in descriptor '{descriptor}' is deprecated and ignored; execution continues. {migration}
Warum: The descriptor uses a known legacy attribute that no longer controls execution.
Lösung: Follow the migration hint when updating the model. Removing the attribute is not required to run it.
Erlaubte Elternelemente / Erlaubte Kindelemente¶
Erlaubte Elternelemente: else, else-if, if, setup, while
Erlaubte Kindelemente:
Keine
Attribute¶
Alle 35 Attribute anzeigen
analyzeDropColumns
Columns to drop during analyze stage.
optional; string; Standardwert: null.
analyzeValueProtection
Enable analysis for value protection.
optional; boolean; Standardwert: null.
batchSize
Batch size for training.
optional; integer; Standardwert: null.
contextSource
Optional relational context source for two-table training.
optional; string; Standardwert: null.
contextType
Type of the relational context.
optional; string; Standardwert: null.
ctxPrimaryKey
Primary key column(s) in context; supports composite.
optional; string; Standardwert: null.
device
Computation device selector passed to the ML engine.
optional; string; Standardwert: null.
differentialPrivacy
Differential privacy configuration (JSON).
optional; string; Standardwert: null.
encodingTypes
Per-column encoding overrides.
optional; string; Standardwert: null; Werte: AUTO, TABULAR_CATEGORICAL, TABULAR_NUMERIC_AUTO, TABULAR_NUMERIC_DISCRETE, TABULAR_NUMERIC_BINNED, TABULAR_NUMERIC_DIGIT, TABULAR_CHARACTER, TABULAR_DATETIME, TABULAR_DATETIME_RELATIVE, TABULAR_LAT_LONG, LANGUAGE_TEXT, LANGUAGE_CATEGORICAL, LANGUAGE_NUMERIC, LANGUAGE_DATETIME.
fairness
Fairness configuration (JSON).
optional; string; Standardwert: null.
flexibleGeneration
Enable flexible generation mode.
optional; boolean; Standardwert: null.
generationBatchSize
Batch size during generation.
optional; integer; Standardwert: null.
gradientAccumulationSteps
Steps for gradient accumulation.
optional; integer; Standardwert: null.
imputation
Imputation configuration (JSON).
optional; string; Standardwert: null.
maxEpochs
Maximum number of training epochs.
optional; number; Standardwert: null.
maxSequenceWindow
Maximum sequence window size.
optional; integer; Standardwert: null.
maxTrainingTime
Maximum training time in minutes.
optional; string; Standardwert: null.
mode
Veralteter Trainingsmodus; beliebige alte mode-Werte werden mit Warnung akzeptiert und ignoriert und steuern keine Persistenz.
optional; string; Standardwert: null.
modelStateStrategy
Strategy for managing model state.
optional; string; Standardwert: null; Werte: reset, resume, reuse.
modelType
Model type/preset used by the engine.
optional; string; Standardwert: null; Werte: TABULAR, LANGUAGE.
name
Name of the ML model.
erforderlich; string.
rareCategoryReplacementMethod
Method for handling rare categories.
optional; string; Standardwert: null; Werte: constant, sample.
rebalancing
Class/feature rebalancing configuration (JSON).
optional; string; Standardwert: null.
samplingTemperature
Sampling temperature used for generation (0-2).
optional; number; Standardwert: null.
samplingTopP
Top-p (nucleus) sampling threshold (0-1).
optional; number; Standardwert: null.
sensitiveColumns
Columns to drop for privacy (comma-separated).
optional; string; Standardwert: null.
separator
Separator for generated data.
optional; string; Standardwert: null.
source
Training data source: a project file or configured source identifier.
erforderlich; string.
splitPartitions
Number of partitions for dataset split.
optional; integer; Standardwert: null.
targetColumn
Target column for supervised training.
optional; string; Standardwert: null.
tgtContextKey
Join key(s) linking target to context; supports composite.
optional; string; Standardwert: null.
tgtPrimaryKey
Primary key column(s) in target; supports composite.
optional; string; Standardwert: null.
trainModel
Training model identifier or preset.
optional; string; Standardwert: null.
trainValSplit
Train/validation split ratio (0-1).
optional; number; Standardwert: null.
type
Type of data to train on.
optional; string; Standardwert: null.