Bỏ qua

Custom Worker Configuration

Warning

The current page still doesn't have a translation for this language.

But you can help translating it: Contributing.

DATAMIMIC runs all background work as Celery tasks. Every task type has its own queue, and each worker process consumes a configurable set of queues. This page explains the task/queue model, lists every supported task, and shows how to assign queues to workers so that no task type is left unconsumed.

How task routing works

  • Each task type maps to exactly one queue. Use the short task-type name in WORKER_QUEUES, for example data_source_scan; the worker launcher adds the datamimic_ broker-queue prefix automatically.
  • When WORKER_QUEUES is an explicit comma-separated list, a worker only processes those queues. Use short names with no spaces. Setting WORKER_QUEUES=all or leaving it unset subscribes the worker to all configured task queues.
  • A task whose queue is not consumed by any running worker stays enqueued and never runs. Splitting work across workers is therefore a deliberate routing decision — every queue you intend to use must appear in at least one running worker's WORKER_QUEUES.

Custom topologies must cover every queue you use

The default deployment ships values that already cover all queues, so it needs no action. If you define your own worker layout, make sure every task type your users can trigger is covered — otherwise those tasks never run. Task types added in a release (for example the Workbench scan family data_source_browse, data_source_scan, and scan_result_generate_model in 3.5.0) must be added to an existing worker's WORKER_QUEUES.

Upgrading to 4.0

Relevant when: your Helm values, container environment or .env pins workers to explicit queue lists. An older list does not automatically subscribe to newly introduced queues. A running catch-all worker (WORKER_QUEUES=all or unset) already covers them.

Why this matters: separate queues let you assign discovery and model-authoring work independently from generation jobs. Defining a queue does not create a consumer: without a subscribed worker, requests remain queued and the caller may time out.

The task catalog adds exactly two queues between 3.5.2 and 4.0. No existing task-type queue is removed in that comparison.

Introduced in Queue to cover User-visible operation
4.0 authoring_manifest_fetch Refresh Authoring discovery metadata from EE.
4.0 authoring Compile/decompile models and perform bounded verification for agent authoring.
3.5, also needed when skipping that release data_source_browse, data_source_scan, scan_result_generate_model Browse sources, scan them, and generate a model from a saved scan.

How to migrate:

  1. Keep your existing queue assignments. Append authoring_manifest_fetch,authoring to an appropriate worker's list, or add a dedicated Authoring worker as shown in the 4.0 release notes. Do not replace a full worker list with a migration fragment.
  2. If your configuration predates 3.5, also assign the three Workbench queues above. Queues such as runtime_capabilities_fetch, healthcheck, Git, backup and migration queues already existed in 3.5.2; retain them rather than treating them as new 4.0 tasks.
  3. Use short names, not datamimic_authoring or the UI labels. WORKER_QUEUES keeps its existing name; do not rename it to DM_WORKER_QUEUES.
  4. Roll out the changed workers. Confirm in their startup logs that they subscribe to the intended queues, then test Authoring discovery and a small model build. Test Workbench browse/scan/model generation too if you changed those assignments.

For a source-managed deployment, the Platform repository includes a read-only validator. Run it from the repository root with your complete values file, not just the additional-worker fragment:

1
uv run --no-sync python script/tools/helm_worker_queues.py validate --values /path/to/values.yaml

The validator reports missing and unknown queues across the configured workers. It checks configuration coverage, not whether the deployed workers are running or successfully executing tasks. Use the supported tasks below to review the complete list.

Supported tasks

The Queue name column is exactly the value you put in WORKER_QUEUES. The Label is the short form shown in TaskView. Scope is one of: project (user-triggered work tied to a project), system (global maintenance, no project), or backup_infrastructure (backup/restore, runs outside backup scope).

Synthetic data generation

Queue name Label Scope Description
standard GEN-SDN project Standard generation run (soft/hard timeout ~15 min).
infinite GEN-INF project Long-running generation with no timeout.
timed_5min GEN-5M project Generation capped at 5 minutes.
timed_30min GEN-30M project Generation capped at 30 minutes.
timed_1hour GEN-1H project Generation capped at 1 hour.
timed_4hour GEN-4H project Generation capped at 4 hours.
timed_8hour GEN-8H project Generation capped at 8 hours.
timed_24hour GEN-24H project Generation capped at 24 hours.
cron GEN-SCH project Scheduled (cron) generation run, no timeout.

Model and template generation from a source

Queue name Label Scope Description
database_generate_model DB-GEN-MOD project Generate a DATAMIMIC model from a scanned SQL database.
database_generate_weighting DB-GEN-WGT project Generate distribution weighting for a database model.
json_generate_model JSON-GEN-MOD project Generate a model from a JSON sample.
json_generate_template JSON-GEN-TMP project Generate a template from a JSON sample.
xml_generate_model XML-GEN-MOD project Generate a model from an XML sample.
xml_generate_template XML-GEN-TMP project Generate a template from an XML sample.
csv_generate_model CSV-GEN-MOD project Generate a model from a CSV sample.

Data Workbench scan family

Queue name Label Scope Description
db_metadata_scan DB-META-SCAN project Scan SQL database metadata (Workbench Database mode).
data_source_browse DS-BROWSE project Browse a data-source location (databases, collections, directories, files).
data_source_scan DS-SCAN project Scan a selected table, collection, or object into a snapshot.
scan_result_generate_model SCAN-GEN-MOD project Generate a model from a persisted scan snapshot.

Agent model authoring

Queue name Label Scope Description
authoring_manifest_fetch AUTH-MNFST system Refresh the verified EE Authoring manifest used by discovery.
authoring AUTHORING project Compile or decompile a DM JSON model and run bounded verification. The server commits a Ready result separately.

Connections, migration, and Git

Queue name Label Scope Description
healthcheck CHK-ENV project Test an environment connection (SQL, MongoDB, Kafka, object storage).
gwa_migration GWA-MIG project Migrate a GWA project into DATAMIMIC.
ama_migration AMA-MIG project Migrate an AMA/EDI definition into DATAMIMIC.
project_git_push GIT-PUSH project Push project changes to the configured Git remote.
project_git_update GIT-UPDATE project Update the project from its Git remote.

System, maintenance, and backup

Queue name Label Scope Description
runtime_capabilities_fetch RUNTIME-CAPS system Fetch runtime capability metadata from the engine.
housekeeping HOUSEKEEPING system Periodic internal maintenance.
clean_up_project_storage CLEAN UP PROJECT STORAGE project Remove orphaned project storage.
system_backup SYSTEM_BACKUP backup_infrastructure Back up the DATAMIMIC data instance.
system_restore SYSTEM_RESTORE backup_infrastructure Restore the DATAMIMIC data instance from a backup.

A practical layout separates short interactive generation, long-running generation, and operational/source tasks across three worker groups. This is an example layout; verify the queue list against the chart and release version used by your deployment:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
celeryworker:
  install: true

  # Each worker listens to a specific set of queues via WORKER_QUEUES.
  # The queue names are the TaskType values defined by the platform.
  workers:
    - name: operation-worker
      replicaCount: 1
      extraEnvs:
        WORKER_QUEUES: "healthcheck,housekeeping,clean_up_project_storage,runtime_capabilities_fetch,authoring_manifest_fetch,authoring,gwa_migration,ama_migration,project_git_push,project_git_update,db_metadata_scan,database_generate_model,database_generate_weighting,json_generate_model,json_generate_template,xml_generate_model,xml_generate_template,csv_generate_model,data_source_browse,data_source_scan,scan_result_generate_model"
      resources:
        limits:
          cpu: 300m
          memory: 500Mi
        requests:
          cpu: 100m
          memory: 256Mi
    - name: short-worker
      replicaCount: 1
      extraEnvs:
        WORKER_QUEUES: "standard,timed_5min,timed_30min,timed_1hour,cron"
      resources:
        limits:
          cpu: 1500m
          memory: 1.5Gi
        requests:
          cpu: 500m
          memory: 1Gi
    - name: long-worker
      replicaCount: 1
      extraEnvs:
        WORKER_QUEUES: "infinite,timed_4hour,timed_8hour,timed_24hour,system_backup,system_restore"
      resources:
        limits:
          cpu: 500m
          memory: 1Gi
        requests:
          cpu: 200m
          memory: 512Mi

Include database_generate_weighting

The database_generate_weighting (DB-GEN-WGT) queue must be consumed wherever database model generation is used. If you copied an older reference values.yaml, confirm this queue is present on a worker — otherwise weighting generation tasks are enqueued but never run.

Adding a queue to a worker

  1. Pick the worker group that should run the task (short interactive, long-running, or operational).
  2. Add the queue name to that worker's WORKER_QUEUES — comma-separated, no spaces.
  3. Roll out the worker so the new configuration takes effect.

Queue names always equal the task type value (the Queue name column above). Do not invent queue names; only the values defined by the platform are routed.

For non-Helm deployments (for example Docker Compose), set the same WORKER_QUEUES environment variable on the worker container or in its .env file.

Task Monitor is not a worker

Since 4.0.0 the Task Monitor runs as its own service, not as a Celery worker. It consumes no WORKER_QUEUES and must be deployed as a separate workload in custom topologies. Chart-specific names depend on the chart release; see Helm Deployment.

Mongo Workbench scan timeout

DM_DATA_SOURCE_MONGO_SCAN_TIMEOUT_MS (renamed from DATA_SOURCE_MONGO_SCAN_TIMEOUT_MS in 4.0.0) bounds MongoDB connection, server-selection, and collection-read waits for Workbench scans. The default is 15000 ms; supported values are 1000 through 120000 ms. Mongo schema inference additionally reads at most 10,000 documents and 64 MiB in deterministic _id order, so a collection with millions of documents is never read in full.