Challenges

Iterating on a model that’s already in production is a common challenge. The hardest part of this is typically adding, changing, or removing features.

Specifically, teams often encounter the following pain points:

  1. Integrating new features to a large feature set is slow and expensive.
  2. Deploying experiments on data pipelines is not supported by existing infrastructure.
  3. Deprecating old features is difficult due to unclear lineage.

These are explained in more detail below:

Integrating new features

The main challenge in this step is speed, efficiency and correctness of training data generation.

Because it’s common to add just a few features to a training set of many hundreds or thousands of existing features, an optimal data pipeline for generating the new training set should only compute the new features, and re-using existing computation for unchanged features.

However, doing this in a way that guarantees the correctness of data is a complex data engineering problem.

The costs of not solving this challenge include:

  • Inflated compute costs – recomputing everything for each new version
  • More time spent battling data pipelines instead of training and deploying
  • Fewer iterations on critical models due to high cost and time overhead

Deploying experiments

Almost all teams have a way to deploy code changes from a branch for testing, but rarely are able to do so with data pipelines. What makes the problem even more challenging is that the dev pipeline needs isolation from production to ensure that launching the test doesn’t cause unexpected regressions.

Because of the difficulty in orchestrating branched data flows, most teams simply opt to skip this step and merge new data pipelines into production when launching a test version of a new model. Oftentimes this ends up being a copy of the original pipeline (duplicated features) with incremental additions, or changes-in-place to production feature sets if the existing infrastructure allows it.

The costs of not solving this challenge include:

  • Latency degradation in production due to bloating payload size
  • Pipeline SLA issues due to issues in new feature computation
  • Model performance degradation in production if changes aren’t handled carefully
  • Lack of clarity on what features are in use for which models, which can create further observability and governance challenges

Deprecating old features

When every feature makes it to the main repository without clear versioning, it can be hard to know which features are still used and which have been deprecated. This problem can be exacerbated by unclear ownership that’s common in jumbled feature pipelines.

As a result, feature repositories get bloated and compute and storage costs bloat along with them.

This is also related to the data-to-model lineage challenge, which will be covered in another post.

The costs of not solving this challenge include:

  • Bloated compute and storage costs
  • Difficult observability due to unclear lineage
  • Difficult/impossible repository governance

Zipline’s Approach

Zipline natively supports compute versioning and makes experiment tracking easy and safe by associating versions with git branches. This approach guarantees that production never mutates during experiments and surgically recompute features based on their column lineage.

Every feature defined in Zipline has an associated version. When a user wishes to modify a feature set, they can make the change and bump the version in their definition.

The interaction between versions and data computation, orchestration and serving are outlined in more detail below.

Branching

Once a user makes a code change and a corresponding version bump on their branch, they can deploy that dev version directly from their branch with a single command.

For example, let’s say that a v0 of a particular feature set is in production, and the user makes some modifications to it and bumps the version to v1. All they would then have to do is call zipline deploy features_v1 and that would be enough to deploy the v1 pipeline.

This would automatically create the following resources:

  1. Pipelines to regularly compute training data
  2. A serving index to power low-latency serving of features_v1 for inference
  3. Pipelines and streaming jobs to keep the feature values in the serving index up to date

This unblocks them to both train and serve a model from this branch without interfering with production inference or training.

Zipline is able to guarantee that orchestration of development branches does not interfere with production by keeping the compute and serving resources separated. Compute sharing (explained in more detail below) is strictly single-directional, using production data where possible in development, but never vice-versa.

Another benefit of this approach is that users can also go to production without any data downtime. As soon as the PR is merged, Zipline identifies V1 as the current production version, and data flow orchestration continues with the only difference being that it is now considered production.

Merging a branch updates the production pointers for datasets, pipelines, and serving endpoints.
Merging a branch updates the production pointers for datasets, pipelines, and serving endpoints.

At this point, V0 will be considered deprecated, meaning that orchestration will be automatically halted and eventually data assets will be cleaned up by janitor jobs.

See the Lineage section below on how Zipline ensures that there are no further dependencies on V0 at this point.

Surgical Evolution

With the code change made, all a user needs to do is run a zipline backfill command on the new version of their entity.

Under the hood, Zipline knows how to interpret these versions at the feature level by using semantic hashing, meaning that on every version change, the system knows which features were modified, added or deleted as part of that change.

For example, say you had an existing training set with features: feature_1, feature_2, feature_3, and you modified it to include feature_4, feature_5, feature_6 in a new version (v0 -> v1).

Zipline keeps track of the semantic hash (a hash of the entire definition of each feature at the column-level) of each feature, so it would know that features 1, 2, and 3 were unchanged.

Then when computing a backfill, it looks to reuse as much computation as possible. In this case it would pull features 1, 2, and 3 from the V0 and only compute features 4 and 5, illustrated below.

Zipline performs these operations under the hood, removing the data engineering complexity from the user’s workflow.

Furthermore, let’s say that the user realizes that there’s something wrong with one of their new feature definitions, for example f6, and makes a change to it.

Now Zipline will detect that there was a semantic change to this feature, and the next time the user calls zipline backfill, the engine will know to only recompute that feature and reuse the rest from the existing table.

Lineage

When making changes to features in a shared repository, it becomes critical to ensure that all downstream consumers are properly handled.

For example, if a user wishes to deprecate a feature, they need to be confident that they can do so without breaking another model. Ideally, these issues should be caught before any jobs are run.

Zipline handles this with the concept of repo compilation. Every time a user makes a change, the repo needs to compile before they can run their new feature definitions. A change that explicitly breaks a downstream consumer will fail to compile.

A successful compilation:

✅ Compile completed successfully
📋 ======== EVAL SUMMARY FOR gcp.search_demo_labels.v1__7 ======== 📋
 
✅ Query
 
📊 Output Schema:
  user_id_user_event_struct_last128_7d: ArrayType(StructType(StructField(event_type,StringType,true),StructField(listing_id,LongType,true),StructField(timestamp,LongType,true)),true)
  user_id_view_event_sum_30d: LongType
  user_id_is_mobile_sum_30d: LongType
  user_id_click_event_sum_1d: LongType
  user_id_add_to_cart_event_average_1d: DoubleType
  user_id_user_event_struct_last128_30d: ArrayType(StructType(StructField(event_type,StringType,true),StructField(listing_id,LongType,true),StructField(timestamp,LongType,true)),true)
  user_id_add_to_cart_event_average_14d: DoubleType
  checkout_label: IntegerType
  ds: TimestampType
 
📊 ======== LINEAGE ANALYSIS ======== 📊
 
✅ gcp.search_demo_labels.v1__7
├── ✅ gcp.search_demo.v1__5
│   ├── ✅ gcp.exports.user_activities__0
│   ├── ✅ gcp.user_activities.v1__1
│   │   └── ✅ gcp.exports.user_activities__0
│   ├── ✅ gcp.dim_listings.v1__0
│   │   └── ✅ gcp.exports.dim_listings__3
│   └── ✅ gcp.dim_merchants.v1__0
│       └── ✅ gcp.exports.dim_merchants__3
└── ✅ gcp.exports.checkouts__0

A failed compilation due to removing a feature user_id_user_event_struct_last128_7d that is used downstream:

📋 ======== EVAL SUMMARY FOR gcp.search_demo_labels.v1__7 ======== 📋
 
❌ Query: Failed to evaluate staging query: [UNRESOLVED_COLUMN.WITH_SUGGESTION] A column or function parameter with name `features`.`user_id_user_event_struct_last128_7d` cannot be resolved. Did you mean one of the following? [`features`.`user_id_view_event_sum_14d`, `features`.`user_id_view_event_sum_1d`, `features`.`user_id_view_event_sum_7d`, `features`.`user_id_view_event_average_7d`, `features`.`user_id_view_event_sum_30d`].;
 
📊 ======== LINEAGE ANALYSIS ======== 📊
 
❌ gcp.search_demo_labels.v1__7
├── ✅ gcp.search_demo.v1__4
│   ├── ✅ gcp.exports.user_activities__0
│   ├── ✅ gcp.user_activities.v1__1
│   │   └── ✅ gcp.exports.user_activities__0
│   ├── ✅ gcp.dim_listings.v1__0
│   │   └── ✅ gcp.exports.dim_listings__3
│   └── ✅ gcp.dim_merchants.v1__0
│       └── ✅ gcp.exports.dim_merchants__3
└── ✅ gcp.exports.checkouts__0

Conclusion

Zipline's approach to versioning and branching ensures that iterating on AI systems is safe, efficient, and cost-effective. By providing isolated environments for testing, intelligent compute re-use, and clear lineage tracking, Zipline empowers practitioners to rapidly innovate without the common pitfalls of data engineering complexity. If you'd like to discuss how Zipline can streamline your AI development workflow, book a meeting with us using the link on our homepage.

Get in touch
Contact us with any questions.
[f] Talk to founders