I audited my own backend and found a billing bug that charged the wrong customer
- buildinpublic
- backend
- postgres
- quality
At the end of the first week of Eventra I made a commit called "audit". It was the first of several: another in May, a bigger one on July 7, and a follow-up on August 4. The first one was in the first week. The reason is simple: when you build alone, nobody else reviews the code, so the review has to be scheduled. Each round means reading the code again, slowly, as if it belonged to someone else.
The July 7 round is the one worth writing about, because of what it found. Thirteen fixes, and the first seven were the serious ones. None of them were things I would have caught by clicking around the app. All of them were places where the code looked like it worked.
How I ran it: first by hand, registering a project from scratch and checking that the data was accurate, and then with AI help to go through the code systematically.
The rule I held myself to: no finding goes on the list unless I can point at the exact place in the code, and no fix counts until I have run it against a real database and watched the number change. "I read the code and it seems fine" does not count in either direction.
The bug that charged the wrong customer
Eventra processes events from many customers at the same time. In one rare combination of timing, usage could be counted against the wrong customer's account: one customer billed for another's events, and the other for nothing.
It never threw an error. The totals were quietly wrong, and the first place anyone would have noticed is a customer asking why they reached their limit early. The fix was to make sure every unit of usage is attributed to the account it belongs to, and I verified it with two accounts sending at the same moment and checking that each one's total matched exactly what it sent.
The cleanup job that never cleaned anything
Raw events are kept for 12 months, and a nightly job removes older data. That job called a database function that did not exist. It failed every night, the error was caught and swallowed, and nothing in the logs looked alarming. Retention, a promise I make on the privacy page, had never run once.
The same review found that a setup step needed by fresh installs was applied by hand and not as part of the deploy. Now the deploy runs it every time and it is safe to re-run.
When fixing one thing breaks another
On August 4 I moved shared state, such as rate limiting and the billing queue, out of a single process's memory so that more than one API instance could run safely. That was the right call.
The follow-up finding was more interesting than the move itself. The storage service I moved to had been set up earlier as a disposable cache, and its defaults were right for throwaway counters and wrong for a billing queue: under memory pressure it could discard queued data without any error. I changed its settings before deploying. The lesson I took is that a component's defaults carry assumptions about what it is for, and when you give it a new job, those assumptions come with it.
Code that looked like it was doing its job
The rest of the list has a pattern I keep running into: code whose shape is right but which never connects to anything. A cleanup branch that could never run. A value that was computed and then never read. Checks that existed in one place and were missing in the place next to it. Each of these is small, and each looked fine in a code review.
What I took from it
The audit closed the findings above.
What changed is how I treat "it works". If the failure is silent, if the only symptom is a number that is slightly wrong or a job that quietly does nothing, then testing the happy path will never show it. You have to go and ask, for each piece, what would happen if this stopped working, and how I would know.