Train Models on Ray

Zipline can run your model training as a Ray job, triggered from Hub like any other workflow. You write the trainer as ordinary Python, declare a Model alongside your features, and Hub takes care of computing the training data first, launching the job, and tracking it to completion.

The key property is that training sits inside the dependency graph. When you ask Hub to train, it recomputes the Join and the labels first if they are stale, so the model never trains on data that quietly went out of date.

When to use this

Use Ray training when your trainer is Python, you want it to run on the same cluster as the rest of your jobs, and you want retraining to follow feature changes automatically.

You do not need Ray for the feature computation itself — that stays on Spark. Ray runs the training step only.

How a training run works

  1. You submit a backfill for the model conf.
  2. Hub walks the dependency graph: StagingQuery, Join, GroupBys — whatever produces your training table — and runs anything that is missing or stale.
  3. When the training data is ready, Hub hands the training step to the Ray layer.
  4. Ray fetches your trainer archive from object storage and unpacks it as the job's working directory.
  5. Your trainer runs, reads the training table, and writes its artifacts back to object storage.
  6. Hub polls until the job finishes and records the result.

Prerequisites

  • A training table. Usually a StagingQuery that joins your features to a label column.
  • A blob storage location for artifacts (s3://, gs:// or abfss://), set as the model's model_artifact_base_uri.
  • A trainer package, described below.

Writing the trainer

The trainer is a normal Python package with a setup.py and a module Ray can execute. The minimum shape:

my_model/
  setup.py
  trainer/
    __init__.py
    cli.py
    train.py

setup.py declares the package name and the dependencies Ray installs into the job's environment:

from setuptools import find_packages, setup
 
setup(
    name="listing_ctr_model",
    version="1.0",
    packages=find_packages(),
    install_requires=[
        "xgboost==2.1.3",
        "pandas==2.2.3",
        "pyiceberg==0.12.0",
    ],
)

Your entry point receives its configuration as command-line flags. Three of them are supplied by Zipline on every run and you should not set them yourself:

Flag Supplied by Meaning
--input-table Zipline the training table to read
--start-ds Zipline first partition of the training window
--end-ds Zipline last partition of the training window
--model-dir Zipline where to write artifacts (see Where artifacts land)

Everything else comes from the job_configs map on your model, turned into --key=value pairs. So {"max-depth": "4"} arrives as --max-depth=4.

Packaging and versioning the trainer

Trainer code is shipped as a zip in blob storage, and its version is the hash of the zip's own contents — a content address. Identical code always produces the same version, and different code cannot share one. That means a version can never point at code other than exactly what you built.

Zipline looks for the archive at:

{model_artifact_base_uri}/builds/{model_name}-{version}.zip

Build the archive deterministically (sorted entries, fixed timestamps) so the same source always produces the same hash, take the first twelve characters of its SHA-256, and upload it to that path. Then set that same value as the model's version.

Note: Because the version is a content address, any change inside the archive changes it — including renaming the package in setup.py. That is the mechanism working as intended. Rebuild, upload under the new name, and update version. The old archive can stay where it is; anything pinned to the old version keeps resolving.

Declaring the model

from ai.chronon.types import (
    EventSource, InferenceSpec, Model, ModelBackend, Query,
    ResourceConfig, TimeUnit, TrainingSpec, Window, selects,
)
 
listing_ctr_model = Model(
    version="afce3c0e00af",          # the trainer archive's content address
    model_artifact_base_uri="s3://your-warehouse",
    inference_spec=InferenceSpec(
        model_backend=ModelBackend.RAY,
        model_backend_params={
            "model_name": "listing_ctr_model",   # names the archive and the output path
            "model_type": "custom",
        },
    ),
    training_conf=TrainingSpec(
        training_data_source=EventSource(
            table=ctr_labels.v1.table,           # the exact name its producer emits
            query=Query(selects=selects(...), start_partition="2026-07-01"),
        ),
        training_data_window=Window(length=4, time_unit=TimeUnit.DAYS),
        python_module="trainer.train",
        resource_config=ResourceConfig(min_replica_count=0, max_replica_count=0),
        job_configs={
            "features": "price_log:price_log,inventory_count:inventory_count",
            "label-column": "label",
            "partition-column": "ds",
            "max-depth": "4",
            "eval-days": "1",
        },
    ),
)

Careful: Use the exact table name your producer emits for training_data_source. Hub matches a model to its upstream producer by comparing table names literally. If the names differ — even by a catalog qualifier — Hub cannot link them, assumes the table is external and already up to date, and trains on whatever is already there. The run succeeds and the model silently trains on stale data.

model_name matters more than it looks: it names both the trainer archive and the output directory. Give it a business-friendly name that describes what the model predicts, not the runtime that executes it. The versioned metadata name is what identifies a particular run.

Passing credentials

Never put secrets in job_configs — those values become command-line arguments and are visible in process listings and logs. Pass the names of environment variables instead, and let the platform inject the values into the Ray pods:

job_configs={
    "client-id-env": "DATABRICKS_CLIENT_ID",
    "client-secret-env": "DATABRICKS_CLIENT_SECRET",
}

Your trainer reads os.environ[...] for those names at runtime.

Where artifacts land

Each model version writes to its own directory, derived from the same coordinates as the archive:

{model_artifact_base_uri}/builds/{model_name}-{version}.zip     # code going in
{model_artifact_base_uri}/models/{model_name}/{version}/        # artifacts coming out

Nothing has to be told either path — a caller holding the model definition can compute both. Because the version is part of the output path, a new version cannot overwrite an earlier run's artifacts.

If you set model-dir yourself in job_configs, your value is used instead and the derived one is skipped. Prefer the derived path unless you have a reason not to.

Evaluating the model

Split by date, not at random. A random split lets the model learn from rows adjacent in time to the ones it is scored on, which flatters the result — this is leakage, and it is the most common way an offline metric ends up lying to you. A date split mirrors how the model is actually used: trained on the past, predicting the future.

Set eval-days to hold out the most recent partitions:

job_configs={"eval-days": "1"}

A trainer that scores a holdout should write its metrics alongside the model, so downstream tooling can read them without parsing logs. A useful report contains:

  • metrics for both the training and held-out sets, so you can see the gap
  • the row counts and date ranges of each split
  • feature importance, so you can tell which inputs the model actually used
  • a precision-recall curve, if the problem is classification
  • a status field explaining why a metric is missing when it is — an undefined AUC because the holdout has one class is a different problem from a degenerate training set, and a reader needs to know which

Note: Interpreting the result: an AUC near 0.5 means the model is no better than chance. If your labels are synthetic or randomly generated, that is the correct result and a high score would indicate a leak rather than success. Check what your labels actually are before celebrating or panicking.

Resources

ResourceConfig controls Ray worker counts:

resource_config=ResourceConfig(min_replica_count=0, max_replica_count=0)

Zero workers means the job runs on the Ray head pod alone, which is right for small or single-node training. Raise both for distributed training. An unset maximum defaults to the minimum, so you cannot accidentally ask for more minimum workers than maximum.

Running it

Compile and submit like any other conf:

zipline compile
zipline hub backfill compiled/models/{team}/{your_model} --start-ds 2026-08-24 --end-ds 2026-08-24

Hub returns a workflow id. Training runs as one step inside that workflow, after whatever upstream jobs it needs.

Troubleshooting

The job starts, succeeds, and the model does not reflect a feature you just added. Almost always the table-name mismatch described above. Check that training_data_source.table is exactly what the producing conf emits, then confirm the workflow actually contains steps for the Join and the labels rather than only the training step.

Ray cannot fetch the trainer archive. The version and the uploaded file must agree. Confirm an object exists at {base}/builds/{model_name}-{version}.zip. If you changed anything inside the package — including its name in setup.py — the hash moved and you need to rebuild and re-upload.

The version is rejected. Versions must be content addresses: 12 to 64 lowercase hex characters. Mutable tags like 1.0 or latest are refused, because two different builds sharing one version would also share an output directory.

Metrics are missing from the report. Check the status field. A holdout with only one class cannot produce an AUC; neither can a training set with only one class. These are different problems and the report should say which.

Training data is empty. The window asked for partitions that do not exist. Check training_data_window against the partitions actually present in the training table.

Reference: job config keys

These are the keys used by the reference example. Your trainer defines its own — anything in job_configs becomes a flag.

Key Purpose
features ordered source:alias pairs selecting the feature columns
label-column column holding the target
partition-column the date partition column, usually ds
eval-days partitions to hold out for evaluation
model-dir output location; omit to use the derived path
seed fixed seed for repeatable runs
catalog-type, catalog-uri, warehouse, scope catalog connection for reading the training table
client-id-env, client-secret-env env var names holding credentials
eta, max-depth, num-boost-round example hyperparameters
Edit this page