Real-time feature store

A feature store expects a stream. Signals brings one.

Signals brings the collection and the streaming engine with it, and computes the feature where it serves it, so a value is current <1s after the event and readable in 6ms.

  • 6ms p50 read, in-region
  • <1s event to feature current
  • Free tier, no card
Try it on yourself

Below is your own visit to this page, computed by Signals as you read.

Signals, about you

connecting

See live events firing as you browse this page on the left; what Signals computed about you on the right, refreshed every few seconds.

this page · events0 writes
nothing sent yet
write
attribute store · your sessionno read yet
reading_now
pages_last_5_min
seconds_since_last_action
sections_read
seconds_engaged
pricing_views
menus_explored
features_wanted
cta_clicks
arrived_from
waiting for the first computed value…
Your first events are in flight through a real Snowplow pipeline. Attributes appear as the stream computes them, usually within a few seconds.
Where it stalls

The store was never the hard part. The stream feeding it was.

Before

A store waiting on a pipeline

The online store is provisioned and empty. Two quarters later the project is still upstream, building collection, schemas and the streaming jobs that were supposed to fill it.

features_served = 0
With Signals

Collection and computation in one product

An SDK in your application, and one definition that computes inside the serving layer: a streaming engine for now, a batch engine for warehouse history. No job of yours in between.

sessions_since_purchase = 3
Result

Production reads what training saw

The same definition serves the request path and builds the point-in-time correct training set, so the lift you measured offline is the lift you get online.

one definition, both paths
Why it matters

What changes when the events and the serving layer are the same product.

An online store is only as current as the pipeline somebody keeps running into it. Here the compute is the serving layer, so a feature is current <1s after the event that moved it.

Train on history you have not computed yet

Attributes are computed on demand from the events you already collected, so a feature defined this morning can train on your whole history this afternoon rather than waiting for a backfill to accumulate.

time to first feature

No train and serve skew to chase

One definition feeds both paths, which removes the quarter usually spent tracing why the online numbers disagree with the notebook.

offline-to-online lift gap

Key on things that are not users

Baskets, listings, stores, sessions, accounts. The entity does not have to exist in modeled warehouse data before you can aggregate over it.

entities per model

Application engineers read it too

One call returns a whole service inside a render budget, so the storefront and the service desk read the same values the model server does.

consumers per definition
How it works

One definition. Live and against history.

The streaming engine keeps the value current, the batch engine fills it from warehouse history, and both come from the definition you wrote once.

  1. 01Collect. Add a Snowplow SDK to your application and behavioral events start flowing in. No stream to stand up first.
  2. 02Define. Say what you want to know about the customer, in the console or in code, versioned through CI/CD.
  3. 03Serve. The streaming engine keeps the value current within <1s of the event. Your model server and your application read it at 6ms p50.
Signals → your app6ms p50
sessions_last_7d4
categories_viewed_10m3
seconds_since_last_event12
failed_searches_session2
basket_value184.00
Trigger churn_risk_highdelivered to your model server
Questions

Before you sign up

Is this a feature store?

It overlaps on serving and on definition parity. It differs in bringing the collection and the compute with it, rather than assuming a stream and a job you already run, and in being aimed at application engineers as much as at ML teams. We file it as real-time customer context, because what it serves is behavior rather than any feature you care to define.

Can we keep the feature store we already run?

Yes. The common pattern is Signals for the application path and the existing store for model training and serving. Or build the training set with Signals from behavior and keep the store for everything else.

What about point-in-time correctness?

The dataset builder computes each attribute from only the events before the moment you are predicting from, with the same definitions the live surface reads. Nothing that happened afterwards can leak into training.

Do I need an event stream already?

No. Collection is part of Signals, and it is usually the half of a feature store project that takes the longest. Add a Snowplow SDK to your web, mobile or server application and events start flowing in. If you already run Snowplow, Signals reads the pipeline you have.

What does the free tier include?

The full product against your own traffic, up to 5 million events a month, with no card and no sales call. You should see an attribute updating against your own traffic in the first session.

Define one feature and read it from your own traffic today.

Add the SDK, write one definition, then click around your own product and watch the value move, the same way the panel at the top of this page moved for you.

The real engine · no card · no sales call
Or hand it to your coding agent
npx plugins add snowplow/skills
Then: Let's add Snowplow Signals to this app. Understand the app first, then define a behavioral feature for session engagement and read it back before the model scores.
Signals, about you
connecting...