A Metrics System with S3 Tables
When working at a large tech company, one thing I really enjoyed and found useful was that if you had a question about some aspect of the system (like say the cache ratio), all you had to do to get an answer was emit a metric in the code with a library, deploy and see the live data flow into some internal embeddable version of cloudwatch. The developer experience was particularly good since from my POV all I had to do was use the library in my service and emit a metric and everything else worked out of the box. My sense is that this was an internal implementation of canonical log lines1 since I recollect seeing these metrics in the boxes where the logs were collected as well. Ideally someone should be able to add a single line like CanonicalLog.set(Metric.PAYMENT_SUCCEEDED, 1); in the control flow and everything should just work.
Now, I’ve always wanted a similar metrics system and felt like the shape of problem fit well into “parquet on s3” type bucket. After a couple of hours working with Claude, I was able to build what I think is a simple production ready system that fits my use case.
The full design doc I wrote with Claude is here: Canonical log lines to S3 Tables — design doc. It goes into a lot more details about the system like the Java interfaces, a sample log line, the table schema, the DuckDB queries, a line-by-line monthly cost breakdown, the Iceberg vs. Parquet query math, etc.
A high level description of the system is that for each API request, we log a JSON with details like userid, metric key-value pairs, endpoint, etc if at least one metric is emitted in the codepath. Our Cloudwatch log group has a subscription filter that forwards these to Firehose (some mild massaging through a Lambda function), which buffers for 15 minutes and puts it in a single metrics table in S3 Table (basically managed Apache Iceberg). I (and Claude Code through a skill) query from Duckdb from my laptop and I use Duckdb metabase connector to run queries from metabase. That’s it. I estimate this will cost us about $1 a month at our prod traffic.
I went with S3 Tables rather than raw Parquet because it compacts and sorts the files automatically, which the design doc estimates makes queries ~90× cheaper than reading Firehose’s small 15-minute files directly.
A few advantages off the top of my head for using something like this :-
- Generally having a metric system like this makes asking questions of the system much easier. It’s especially useful when doing rollouts via feature gate to ensure we don’t break anything.
- Claude code loves that the data sits so cleanly in S3 Tables ( which can easily be read by Duckdb client). It’s incredibly powerful to ask Claude to get the answer to a specific question (especially after it instrumented the code).
- It’s simple and dirt cheap (So agents can go crazy running multiple queries against it).
- Duckdb can be embedded in metabase which is the source of our product analytics and acts complementary.