For dbt and analytics engineering teams
Your dbt models,
current on every commit.
eddy is a hosted analytics database. It reads your dbt project’s compiled manifest, follows your Postgres through one replication slot, and keeps every model current under inserts, updates and deletes: joins, window functions and recursive CTEs included. Every model is checked row for row against Postgres or DuckDB.
Illustration on a 432-row table with a 3–6 row join fan-out. The ratio on your tables follows their size.
┌─ sources ─────────┐ ┌─ eddy ───────────────────────────────────┐ ┌─ readers ───────────┐ │ Postgres (CDC) │ │ │ │ dashboards │ │ Files (next) │──changes──▶ │ SQL ──▶ plan ──▶ incremental circuit │──views────▶ │ services │ │ Kafka │ │ state on object storage │ (pgwire) │ agents │ └───────────────────┘ └──────────────────────────────────────────┘ └─────────────────────┘
“Incremental” usually means append-only.
Most streaming SQL engines assume rows only ever arrive. Real data gets corrected: a refund posts, a customer moves region, an order is cancelled a week later. Each one has to undo work the view already emitted — and that is exactly where append-only engines stop supporting your SQL, or quietly go wrong.
eddy is built on DBSP, an algebra for incremental computation in which every change carries a weight. An update is a retraction and an insertion; a delete is a row with weight −1. The circuit does only the work proportional to the change, and the result is the same multiset a full rebuild would produce.
Inserts
New rows flow through joins and aggregates and land in the view within one tick.
Updates
A dimension change rewrites history: every fact row that joined to it is retracted and re-emitted.
Deletes
Removed rows come back out of sums, ranks and recursive closures — with nothing left behind.
Watch one change cross the pipeline.
The demo is a glass pipeline: Postgres on one side, your dbt marts on the other, and every row that moves visible on the edge it crossed. Below, a customer places an order. One row is retracted, one is inserted, and every mart in the project is current before a plain view would have finished re-reading the first table.
| w | customer_id | name | orders | last_amount |
|---|---|---|---|---|
| −1 | 1 | Ada Lovelace | 6 | 49.99 |
| +1 | 1 | Ada Lovelace | 7 | 49.99 |
Nothing else in the table was touched. Every operator speaks in these weighted rows, which is why an update is never a rebuild.
category_month_on_hand: 2.2 s as a plain view. pg_trickle on the same report: 13.6 s, a full recompute every time.with lines as (
select order_id, sum(qty * unit_price) as gross
from order_lines group by order_id),
refs as (
select order_id, sum(amount) as refunded
from refunds group by order_id)
select o.id, o.order_date, c.region,
l.gross,
coalesce(r.refunded, 0) as refunded,
l.gross - coalesce(r.refunded, 0) as net
from orders o
join customers c on c.id = o.customer_id
join lines l on l.order_id = o.id
left join refs r on r.order_id = o.id
where o.status <> 'cancelled'
You write the model. eddy does the bookkeeping.
Plain SQL in, the same file dbt already runs. eddy reads the compiled manifest, plans each model at runtime, and feeds it from Postgres’s write-ahead log. No triggers, no extension, nothing installed in your database.
Reads cost nothing.
So concurrency does too.
A plain view re-runs the query for every reader. eddy does the work once, when the data changes, and every reader takes a snapshot of the result. The hundredth dashboard costs what the first did, and none of them contend with the writes, or with each other.
Many tenants, one node
A tenant costs a fixed 3 MiB before it holds data, and idle tenants are nearly free in CPU: five hundred of them shared 3% of one core.
Evict and wake
An idle tenant is evicted to S3, where it costs storage and nothing else. A write wakes it: 290 ms to open its state and apply the change.
Serve from the pipeline
There is no second copy to keep in sync. Reads come from the pipeline’s own state, over the Postgres protocol, at a staleness of at most one tick.
How it works
- 01
Write SQL
The model you would write anyway. No annotations, no watermarks, no DAG configuration. If DataFusion can plan it, eddy renders it operator by operator; what it can’t render yet,
preview_modeltells you with the reason, before anything is built. - 02
eddy plans it, at runtime
The optimized plan is rendered into incremental operators one node at a time. No code generation, no compiler in the loop, no per-model binary.
- 03
Read the view
Views are served straight from the pipeline’s own state over the Postgres protocol. A snapshot costs microseconds at any size, and reads never block updates.
The SQL other engines refuse.
The jobs your models actually do, and what a correction does to each. Fifty-six models, from jaffle-shop marts to TPC-H shapes, are fed changes and diffed against DuckDB after every chunk. These eight are the ones append-only engines turn away.
Running balances
window sum over a ledger
Stock on hand for every product on every day. Correct one receipt and only the days after it change.
agrees with DuckDBSessionization
lag, then a running sum of session starts
Sessions per product from a stream of views. A late or deleted event merges or splits a session, and the counts follow.
agrees with DuckDBRolling windows
RANGE frame over time
Trailing seven-day revenue per shop at each order. Amend one order and only the windows that contain it move.
agrees with DuckDBFunnels and cohorts
full outer joins · min-then-join
Views, orders and units per product, kept even when a product has only one of them. Monthly cohorts by first order.
agrees with DuckDBLatest state per key
row_number = 1 · correlated top-1
Each customer’s most recent order, or the current row of anything. Delete it and the previous one takes its place.
agrees with DuckDBTop-N per group
rank filter over a window
Top three orders per customer, top products per store. An update re-ranks its partition and only its partition.
agrees with DuckDBFacts with dimensions that change
multi-way join · outer join
Orders, customers, lines and an outer join on refunds. A dimension change rewrites history in every fact row it reaches.
agrees with DuckDBHierarchy rollups
recursive CTE
Revenue rolled up a category tree of any depth. Move a product and both ancestors’ totals change.
agrees with DuckDBChecked, not claimed.
Every report eddy keeps is recomputed from scratch by Postgres or DuckDB after every transaction and compared row for row. These are the numbers that survive that.
Measured on the demo dataset and eddy’s own test harness, with Postgres and DuckDB as independent judges. The method: every model is recomputed from scratch by Postgres or DuckDB after every change and compared row for row; 45% of the test changes are updates and deletes; a kill mid-stream and a restore must match end to end. We’re glad to walk through it on a call.
eddy is built by Matt Helm at Hollyburn Analytics Inc.
How it stays correct.
No generated code
One byte-ordered row type serves every operator, so a new model is a plan, not a build. Adding a view never means shipping a binary.
An oracle as the judge
Correctness is checked against a different SQL implementation, not against our own planner. It has already caught two bugs a smoke test would have missed.
Exact recovery
State checkpoints to S3. Kill the process mid-stream, restore, replay: the view matches the oracle end to end.
Hosted: nothing to deploy or scale
Compute is proportional to the change, state lives on S3 with kilobytes resident, and reads are snapshots. An idle customer is evicted and costs storage until a write wakes it.
Get early access.
Leave a work email. We reply by hand, and if you have a model your current stack can’t keep incremental we’ll show it staying correct under updates and deletes on your own data shapes.
- A reply from a founder within two business days
- A walkthrough on your SQL if you want one, no slides
- No mailing list