Skip to main content
Blog

SDP-META: Pipelines at Scale

Managing pipelines using a metadata configuration driven framework

Databricks has officially released SDP-META v0.1.0. SDP-META is an upgrade and renaming of the former DLT-META project. This blog post discusses the major improvements in this release and why you should consider learning and/or incorporating SDP-META into your pipeline development.

Why the rename?

DLT-Meta was built for “DLT” (Delta Live Tables), which has since been rebranded to Lakeflow Spark Declarative Pipelines. This release reflects the name change and incorporates some significant changes. For existing DLT-Meta pipelines there's a migration process documented in the “DLT-Meta to SDP-META migration guide”.

New Features

  • SDP-META Agent – Agentic support for pipeline development
  • New Databricks Declarative Automation Bundles (DAB) templates
  • Full release of the SDP-META Databricks App
  • Support for YAML configuration files as well as JSON
  • Multi-source AUTO CDC
  • Pipeline Filtering at the silver layer
  • Support for Automatic liquid clustering in pipeline definitions

Introduction

SDP-META is a pipeline framework for defining, deploying, and managing Bronze and Silver pipelines on Databricks Lakeflow Spark Declarative Pipelines (SDP). Its purpose is to simplify developing and maintaining your Bronze and Silver pipelines at scale. Instead of creating many individual jobs and notebooks for each of your pipelines, there's a single notebook or generic pipeline that consolidates and standardizes your process for pipeline development. To create a new pipeline, you define a JSON or YAML configuration file and SDP-META reads that metadata and takes care of building and deploying the full processing graph.

SDP-META is for data engineering teams that wish to standardize a repeatable process for Bronze and Silver pipelines across many datasets. Engineers with a Data Warehousing background can see how this framework simplifies the process of getting data from Raw to Bronze to Silver, where the interesting queries happen. Engineers with a Software Engineering background can see how this framework modularizes pipeline development to simplify testing and apply consistent standardized quality controls to your pipelines.

Whether this process is for a hundred pipelines or thousands, we recommend reviewing this framework and either adopting it directly into your process or using it to find new techniques and procedures that you can incorporate into your existing framework.

Getting Started

Rearc has recently worked with Databricks to introduce SDP-META in a top-tier financial institution for managing and migrating their many dataflow pipelines to Spark Declarative Pipelines. The ability to easily share and review data quality assertions across pipelines, teams, and business units was instrumental in developing and maintaining a consistent data lake.

For those who are interested in a quick start, Rearc developed a public SDP-META Learning Lab. This lab is entirely Databricks Workspace notebook-driven, and is a great resource for getting started with SDP-META and familiarizing yourself with its architecture and DAB-based deployment. It's designed to be a focused and iterative mini-project without throwing you into the deep end of the SDP-META architecture right away. This lab quickly onboards and trains new developers and gets them up-to-speed and familiar with the development and deployment process.

The 30,000-Foot Overview

sdp-meta architecture Image from https://databrickslabs.github.io/sdp-meta/docs/intro

The components of a SDP-META pipeline:

  • SDP-META Wheel file – The wheel file is available on PyPI or can be built and deployed from the SDP-META repository. The wheel file must be available to both the Onboarding job and deployed Pipeline clusters.

  • Onboarding file – A JSON or YAML file that defines the pipeline. The onboarding file describes the data source, how to access source data, where it's located, and which files to load. It also defines target bronze and silver tables and specifies which data quality expectations and silver transform files to use. Data quality expectations are data quality rules that are applied to your tables at the Bronze or Silver level. Rules can be specified to log or quarantine records or fail the pipeline. A Silver Transformations file defines a set of transformations and filters that are applied to records from your bronze table before they are saved to your silver table.

  • SDP-META Onboard job – This job is responsible for reading the pipeline specification files and populating the DataflowSpec table. Run this job whenever pipeline specification files are changed (onboarding file, data quality expectations, or silver transformation files).

  • DataflowSpec tables – This artifact is created and maintained by the onboarding job. When an SDP-META pipeline is run, it reads the DataflowSpec table for its defined data_flow_group, adds or updates pipeline tables if needed, and then runs the pipeline to read from the data source and populate tables. There's a DataflowSpec table for each stage of the pipeline: bronze and silver.

  • Generic Declarative Pipeline – All pipelines use the same generic declarative pipeline notebook. This notebook is responsible for installing the SDP-META wheel file to the cluster and invoking the execution module. Each pipeline is responsible for a single data_flow_group, or collection of data flows, and runs all data feeds defined in the onboarding file that share that data_flow_group key value.

  • Source Data – The source data in its raw format. This can be files in cloud storage, a Databricks Volume, a Delta table, Kafka, or any of the other sources that are supported by Autoloader or Lakeflow Connect.

  • Bronze and Silver – The pipeline creates and populate the target tables. The Bronze layer contains the raw data ingested into Delta format and also includes Quarantine tables for records which fail data quality rules. The Silver layer contains cleaned and enriched data ready for use in analytics or populating downstream Gold tables.

Supported Pipelines

SDP-META supports, out of the box, in basic configurations, the following types of pipelines:

  • Bronze and Silver pipelines
  • Bronze only pipelines
  • Autoloader/Cloudfiles sources
  • Azure Event Hub sources
  • Apache Kafka sources
  • Kafka sinks
  • Lakeflow Connector sources such as SFTP, Postgres, or custom APIs through Community Connectors, and many others
  • Snapshot sources
  • Change Data Capture (CDC) using AUTO CDC
  • Silver fanout (one bronze to many silver)
  • Multi-source fan-in
  • Row Filtering for permission based access to table rows

In addition, The SDP-META pipeline object, DataflowPipeline, includes injection points for custom data transformation functions and more complicated snapshot filter and assembling scenarios.

Pipeline Deployment

SDP-META includes a host of deployment methods suitable for any environment, from quick one-off exploratory deployments, code-free deployments, and deployments in tightly locked down CI/CD production environments.

This release introduces Agentic Deployments with the new MCP Server and improves on the existing deployment options available in the command line interface (CLI), the web-based app, and manual deployments via directly invoking methods of the SDP-META library from a notebook.

Here's a summary of the available deployment options in SDP-META:

Deployment Options:

MethodGit-trackableTouches workspaceInterfaceBest for
DABYesYesCommand LineProduction, teams, CI/CD
Interactive CLINoYesCommand LineOne-off, quick start
Databricks AppNoYesBrowserNon-CLI users, demos
Manual Job SetupVariesYesConsole or NotebookNo CLI access, custom orchestration
MCP (scaffolding)Yes (bundle)NoAgentAI-assisted DAB workflows

Deployment Methods

Define your pipelines, infrastructure, and deployment options using Databricks bundle files, (databricks.yml, variables.yml, onboarding job, and pipelines). You can either create and manage these files directly through an IDE or manage them using command line utilities, such as ‘bundle-init’ or ‘bundle-add-flow’. This option is best suited for teams using CI/CD pipelines, multi-environment deployments, and Git-tracked infrastructure as code.

2. Interactive Command Line

SDP-META provides command line utilities for onboarding and deploying pipelines directly without writing any YAML. This is great for quick one-off deployments and demos and for a quick start using SDP-META in development and demo environments. For production and regulated environments, it's recommended to use the DAB approach and maintain your configuration files with a version control system such as Git.

3. Databricks App

A Flask web app that wraps the full onboarding workflow in a browser GUI. This is a click-based and code-free method that makes it easy to configure and deploy pipelines without using the command line.

4. Manual Job Setup

For teams that need a more custom deployment approach than DAB. SDP-META provides the ability to directly configure the onboarding job and pipeline through the Databricks Workflows UI or a custom notebook. Either deploy a Python Wheel Job using the Databricks UI or invoke the Onboarding commands directly from a custom notebook using OnboardDataflowspec(...).onboard_dataflow_specs().

5. MCP Server – New in this release

An AI agent uses MCP tools to scaffold and validate bundles locally, then hands off to the command line or DAB for live workspace deployment.

Developing With the New SDP-META Agent

SDP-META v0.1.0 introduces first-class agentic support for pipeline development through two complementary components: an Agent Skill and an MCP Server. Together, they allow AI coding agents, such as Claude Code and Cursor, to scaffold, inspect, and validate SDP-META bundles using natural language.

The Agent Skill and MCP Server

  • The Agent Skill (skills/sdp-meta) is a structured Markdown operating manual for AI agents. It gives the agent the correct procedures for using the MCP Server. The skill is used automatically in skill-aware clients when your request mentions SDP-META or metadata-driven pipelines.
  • The MCP Server is what actually performs the work. This a local tool server that exposes five tools that an agent can call to scaffold bundles. This is currently limited to scaffolding and validation, with future releases expected to add tools for onboarding and deployment.

The Five Tools of the MCP server:

ToolWhat the agent can do
sdp_meta_bundle_initScaffold a complete new DAB – job, pipeline, variables, runner notebook, and flow recipes
sdp_meta_bundle_add_flowAppend typed flow entries to a bundle's onboarding file
sdp_meta_bundle_validateRun databricks bundle validate plus SDP-META sanity checks
sdp_meta_list_templatesList every packaged onboarding, DQE, and silver-transformation template
sdp_meta_get_onboarding_templateReturn the raw content of any packaged template by name

Packaged templates are also exposed as MCP resources under the sdp‑meta://templates/ URI prefix, so agents can read them directly as reference material when constructing onboarding files.

Development Flow

The MCP tools are designed to be composed in sequence. A well-configured agent follows this pattern without prompting:

  1. Scaffold – Initialize or update a DAB deployment bundle
  2. Add flows – Creates typed flow definitions for each source
  3. Validate – Validates the bundle and SDP-META configurations and inspects any errors
  4. Reference templates – If needed, the agent can pull in canonical field examples rather than invent new pipeline configurations
  5. Hand off – Once validation passes, the agent hands-off deployment to the developer

For teams managing many onboarding tables, this unlocks a vastly different workflow: describe your sources in a natural language, and let the agents produce the onboarding files and data quality validations. The agent performs the tedious structural work, such as ordering fields, incrementing IDs, and ensuring group-name consistency. This eliminates the kind of repetitive, schema-heavy work that leads to costly time-consuming errors in deployment and validation.

Here are a few example prompts, ranging from simple to more complete, for invoking the SDP-META MCP server tools:

Minimal – Just scaffold a bundle
"Scaffold an SDP-META bundle for my project using quickstart defaults."

Realistic – New pipeline from a description
"I need to onboard three Autoloader tables from /Volumes/main/landing/files/ – orders, customers, and transactions. Each needs bronze and silver layers with basic data quality rules. Scaffold an SDP-META bundle and add flows for all three."

Template-first – When you want to see a reference template first
"Show me an example SDP-META onboarding template for an Event Hubs source, then scaffold a bundle using those settings for my telemetry topic."

Validate an existing bundle
"Validate my SDP-META bundle at ./my_sdp_meta_pipeline against the dev target and fix any issues you find."

End-to-end with deployment
"Set up an SDP-META pipeline for the tables in my UC schema raw.landing. Bronze and silver layers, YAML format, split pipelines. Scaffold the bundle, add all the flows, validate it, then deploy to my workspace using the dev profile."

The key trigger phrases that cause the agent to load the SDP-META skill and reach for the MCP tools are: "sdp-meta", "dlt-meta", "onboard a dataflowspec", "metadata-driven pipeline", "bronze/silver pipeline from config", and "scaffold an sdp-meta bundle".

Support for SDP-META Projects

One final note regarding SDP-META: as stated by Databricks, “all projects released under Databricks Labs are provided for exploration only, and are not formally supported by Databricks with Service Level Agreements (SLAs).” SDP-META is an open source project that can be extended and modified to fit the needs of your deployment environment. Please keep this in mind if your environment requires a formal support agreement. Rearc may be able to assist you with your development, pipeline onboarding, training, or support needs. Please contact us if you are interested in pursuing any of these in your Databricks pipeline environments.

Next steps

Ready to talk about your next project?

1

Tell us more about your custom needs.

2

We’ll get back to you, really fast

We will evaluate your query and respond within 2 business days.

3

Kick-off meeting

We will schedule a quick meeting to further understand your use case and start working toward a solution together!

Let's Talk