dbt Core vs dbt Cloud: Features, Costs and Choosing
Chat2DB Teamdbt — the data build tool — is the transformation layer of the modern analytics stack. It does not move data and it does not compute anything itself. It compiles SQL you write into dependency-ordered CREATE TABLE AS and CREATE VIEW statements and runs them against your warehouse.
The confusing part for anyone evaluating it is that "dbt" refers to two products with the same core: dbt Core, the open source Python package, and dbt Cloud, the hosted platform built around it. This article explains what each gives you and how to choose.
What dbt actually does
A dbt model is a .sql file containing a SELECT:
-- models/marts/fct_orders.sql
{{ config(materialized='incremental', unique_key='order_id') }}
SELECT
o.order_id,
o.customer_id,
o.ordered_at,
o.status,
sum(oi.quantity * oi.unit_price) AS order_total
FROM {{ ref('stg_orders') }} o
JOIN {{ ref('stg_order_items') }} oi ON oi.order_id = o.order_id
{% if is_incremental() %}
WHERE o.ordered_at > (SELECT max(ordered_at) FROM {{ this }})
{% endif %}
GROUP BY 1, 2, 3, 4ref() is the key idea. It resolves to the fully qualified name of another model and declares a dependency, so dbt builds a DAG and runs models in the right order, in parallel where possible. Change one staging model and dbt run --select stg_orders+ rebuilds it and everything downstream.
Around that sit the pieces that make it a framework rather than a SQL runner:
- Materialisations —
view,table,incremental,ephemeral,materialized_view, plus snapshots for slowly changing dimensions. - Tests — declarative assertions in YAML, and singular tests as SQL files.
- Sources — declared raw tables with freshness thresholds.
- Documentation — generated from YAML plus the DAG, served as a static site.
- Macros — Jinja functions for reusable SQL.
- Packages — reusable dbt projects, installed from the hub.
Tests are declared next to the model:
# models/marts/schema.yml
version: 2
models:
- name: fct_orders
description: One row per order, with computed totals.
columns:
- name: order_id
description: Primary key.
tests: [unique, not_null]
- name: customer_id
tests:
- not_null
- relationships:
to: ref('dim_customers')
field: customer_id
- name: order_total
tests:
- dbt_utils.accepted_range:
min_value: 0
inclusive: true
sources:
- name: raw
freshness:
warn_after: {count: 12, period: hour}
error_after: {count: 24, period: hour}
tables:
- name: orders
loaded_at_field: _ingested_atAnd the everyday commands:
dbt deps # install packages
dbt build # run + test, in DAG order
dbt run --select fct_orders+ # a model and everything downstream
dbt test --select tag:critical
dbt source freshness
dbt docs generate && dbt docs serve
dbt run --select state:modified+ --defer --state ./prod-artifacts # CI: build only what changedThat last one — slim CI — is the single biggest cost saver on a large project. It builds only modified models and defers unchanged references to production, so a pull request touching two models does not rebuild four hundred.
dbt Core
dbt Core is Apache 2.0 licensed, installed with pip, and runs anywhere:
pip install dbt-core dbt-postgres # or dbt-snowflake, dbt-bigquery, dbt-databricks, dbt-trino
dbt init my_projectConnection details live in ~/.dbt/profiles.yml:
my_project:
target: dev
outputs:
dev:
type: postgres
host: "{{ env_var('PG_HOST') }}"
user: "{{ env_var('PG_USER') }}"
password: "{{ env_var('PG_PASSWORD') }}"
port: 5432
dbname: analytics
schema: dbt_dev
threads: 8You get the entire transformation framework: models, tests, snapshots, seeds, macros, packages, docs, exposures, contracts and the full CLI. Nothing about the modelling capability is held back.
What you supply yourself:
- A scheduler. Airflow, Dagster, Prefect, GitHub Actions, or cron.
- CI/CD. A pipeline that runs
dbt buildon pull requests, ideally with slim CI and a separate schema per PR. - Secrets management for warehouse credentials.
- Docs hosting, since
dbt docs generateproduces a static site someone has to serve. - Alerting on failures.
- An editor. Most teams use VS Code with the dbt Power User extension.
A minimal GitHub Actions run:
name: dbt
on: [push]
jobs:
build:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: {python-version: '3.12'}
- run: pip install dbt-core dbt-snowflake
- run: dbt deps
- run: dbt build --target ci
env:
SNOWFLAKE_ACCOUNT: ${{ secrets.SNOWFLAKE_ACCOUNT }}
SNOWFLAKE_PASSWORD: ${{ secrets.SNOWFLAKE_PASSWORD }}None of this is hard. It is a few days of setup and some ongoing maintenance.
dbt Cloud
dbt Cloud is the hosted platform. It runs the same dbt Core underneath and adds:
- Managed scheduler with job definitions, dependencies, retries and run history.
- Browser IDE with autocomplete, lineage preview and a compile/preview pane — the main draw for analysts who do not want a local Python environment.
- CI integration that creates a temporary schema per pull request, runs slim CI automatically, and posts results back to GitHub, GitLab or Azure DevOps.
- Hosted documentation and lineage, always current, with SSO in front of it.
- The Semantic Layer, built on MetricFlow: define a metric once and query it consistently from BI tools.
- dbt Explorer for column-level lineage, model performance history and metadata search.
- The discovery / metadata API for programmatic access to run and lineage metadata.
- SSO, RBAC and audit logs on paid tiers.
- dbt Fusion, the Rust-based engine introduced in 2025, which parses and compiles projects dramatically faster than the Python implementation and enables live SQL validation in the IDE.
Pricing has three shapes: a free Developer tier for a single developer with limited runs; a per-seat Team tier; and Enterprise with SSO, RBAC, the Semantic Layer and custom terms. Consumption on higher tiers is metered in successful model builds, so cost scales with project size and run frequency rather than purely with headcount. Check current pricing directly — it has changed more than once.
The comparison that matters
| dbt Core | dbt Cloud | |
|---|---|---|
| Modelling, tests, snapshots, macros | Full | Full (same engine) |
| Cost | Free (you pay for compute and infrastructure) | Free single seat, then per seat + usage |
| Scheduling | Bring your own | Included |
| CI with slim builds | Build it yourself | Built in |
| Development environment | Local IDE + CLI | Browser IDE or CLI |
| Hosted docs and lineage | Self-host | Included, plus column-level lineage |
| Semantic Layer | MetricFlow locally, no served API | Served API for BI tools |
| SSO / RBAC / audit | Not applicable | Enterprise tier |
| Support | Community Slack | Contracted |
| Portability | Total | Project is portable; platform features are not |
The decision usually reduces to one question: do you already have a mature orchestration and CI platform, and engineers who own it?
If yes — you run Airflow or Dagster, you have CI conventions, you have a platform team — dbt Core slots in as one more task type and Cloud's headline features duplicate what you already operate.
If no — you are a small data team, mostly analysts, without a platform engineer — Cloud replaces an infrastructure project you would otherwise have to staff.
Common middle grounds
Few teams are purely one or the other:
- Core plus a managed orchestrator. dbt Core tasks inside Dagster or Airflow, with dbt's artifacts feeding asset-level lineage in the orchestrator UI. Very common at companies with existing data engineering.
- Cloud for scheduling, CLI for development. Engineers work locally with dbt Core and the same repository; Cloud runs production jobs. The browser IDE is optional.
- Core plus open source tooling. Elementary or re_data for observability and alerting, a static docs site behind your own auth, Great Expectations or dbt tests for quality.
Whatever you choose, the transformations end up as tables and views in a warehouse that someone has to query and validate. Being able to run the compiled SQL from target/compiled/ directly against Snowflake, BigQuery, Postgres or Trino while debugging a model is part of the daily loop — Chat2DB (opens in a new tab) connects to those engines and 20+ others, with a browser version at app.chat2db.ai (opens in a new tab).
Practical advice for starting
- Start with Core locally, even if you expect to buy Cloud. A day with the CLI teaches you the model far better than the IDE does, and it makes you a better judge of what Cloud is worth.
- Get the project structure right early:
staging/for one-model-per-source cleanups,intermediate/for reusable joins,marts/for business-facing tables. Restructuring later is tedious. - Test from day one.
uniqueandnot_nullon every primary key,relationshipson every foreign key. These catch the majority of real pipeline bugs. - Set up slim CI before the project grows. Retrofitting
state:modified+onto a 500-model project during a cost review is unpleasant. - Decide on Cloud when a specific gap hurts — usually scheduling reliability, analyst onboarding, or the need for a served semantic layer. Buying it before you feel the gap makes it hard to evaluate.
Summary
dbt Core is the complete transformation framework — models, ref()-driven DAG, materialisations, tests, snapshots, docs, macros — under Apache 2.0, free forever. dbt Cloud runs that same engine and sells the platform around it: scheduling, CI with slim builds and per-PR schemas, a browser IDE, hosted lineage, the Semantic Layer, governance, and the faster Fusion engine. Teams with existing orchestration and platform engineering usually get more from Core plus Airflow or Dagster; small analyst-led teams usually get more from Cloud than the licence costs them. Start with Core, and buy Cloud when you can name the specific thing it fixes.
