Building Eventra · Part 4 of 103 min read

The rollup job: the quiet heart of the backend, and the part I had to fix the most

  • buildinpublic
  • postgres
  • backend
  • nestjs
raw eventsrollupevery minuteTOTAL USES12,408dashboard numbers

If you open Eventra's dashboard, almost every number you see comes from one background job that runs once a minute. It is called the rollup. It is the least visible part of the backend, and the part I have fixed more times than any other. The backend audit came back to it more than once.

What it does

When your app sends an event, Eventra stores it first. That is fast and safe, but a raw event log is not something a dashboard can read directly: counting millions of rows every time someone opens a page would be slow and expensive.

So a job processes new events in the background and turns them into the numbers the dashboard reads: usage per feature, unique users, first and last use, daily and hourly activity, per-user activity. The rollup is the bridge between "something happened" and "here is your chart".

Two principles kept it dependable from early on. Only one copy of the job may work on a given set of events at a time. And the raw data always remains the source of truth: if the rollup falls behind, nothing is lost, the numbers are only late.

April 4

April 4 has eleven commits with "rollup" in the message, and April 7 has five more. The day before, a commit had already touched the unique-user counter and a race condition in the analytics service.

A job like this is hard to debug because of how it runs: in the background, once a minute, over data that keeps changing. When a number on the dashboard is wrong, the cause can be in what was written, what was read, when it ran, or what ran at the same time. You cannot step through it. You reason about states that exist for milliseconds.

The failure that day: raw events arrived in a shape the job did not expect, the job crashed, and nothing reported it. There was no alert and no error on the dashboard. The events sat in storage and never reached the numbers, so the dashboard stopped moving. A job that fails quietly is much harder to fix than one that fails loudly, and that became the first rule for the rollup.

QUIET FAILUREjob crashesnumbers frozen, nobody knowsLOUD FAILUREjob crashesalertfixed within minutes
A quiet failure leaves the numbers frozen with no signal. A loud one raises an alert.

It kept coming back, in different forms

Every fix to the rollup revealed a different kind of problem. It was never the same bug twice.

  • February: getting the job to run reliably without two copies interfering with each other.
  • April: correctness under load: race conditions and a unique-user count that was not exact.
  • July: a review found a case where some detailed statistics were silently dropped once a batch got large enough.
  • August: the fix for that case only worked for a single server process, so it had to be reworked to hold up with several.
  • September: the CLI's own bookkeeping events were counted as real usage, which made every feature look used at least once. I wrote about that one in a separate dev.to article.

Each of those bugs appeared only because the system had grown into a new situation: more load, more servers, a new kind of event. A job that is correct for one server and a quiet database is not automatically correct for two servers and a busy one.

Two rules

Two rules came out of the rollup.

Make failure loud and progress visible. A crash with no alert is the worst case, so a job like this has to tell you when it stops. Now the job reports how far behind it is, and the ops page in the dashboard shows it. The first time I could look at one number and know whether the rollup was healthy, a whole class of guessing stopped.

OperationsHealth of the background jobs.ROLLUP LAG9 sHealthyINGEST BUFFER0 pendingHealthyLAST ROLLUP42 s agoHealthyEVENTS PER MINUTE
The ops page in the dashboard, where the state of the background jobs is visible at a glance.

Design for more than one of everything from the start. Almost every serious rollup bug came from an assumption of one process, one batch or one kind of event. Those assumptions are invisible when you write them and obvious a month later. On April 4 I kept working until the numbers moved again, which is why that day has so many commits.