S3 Table Metrics System
When working at a large tech company, one thing I really enjoyed and found useful was that if you had a question about some aspect of the system (like say the cache ratio), all you had to do to get an answer was emit a metric in the code with a library, deploy and see the live data flow into some internal embeddable version of cloudwatch. The developer experience was particularly good since from my POV all I had to do was use the library in my service and emit a metric and everything else worked out of the box. My sense is that this was an internal implementation of canonical log lines1 since I recollect seeing these metrics in the boxes where the logs were collected as well.
Now, I’ve always wanted a similar metrics system and felt like the shape of problem fit well into “parquet on s3” type bucket. After a couple of hours working with Claude, I was able to build what I think is a simple production ready system that fits my use case.
A high level description of the system is that for each API request, we log a JSON with details like userid, metric key-value pairs, endpoint, etc if at least one metric is emitted in the codepath. Our Cloudwatch log group has a log filter that filter forwards these to Firehose (some mild massaging through a Lambda function), which 15 mins buffers and puts it in a single metrics table in S3 Table (basically managed Apache Iceberg). I (and Claude Code through a skill) query from Duckdb from my laptop and I use duckb metabase connector to run queries from metabase. That’s it.
I got Claude Sonnet 5.5 to write a proper design artifact before letting it implement and tested it in staging before pushing to prod - https://claude.ai/artifact/NS7q4DFxWeHzLsXxJRgnmB .
A few advantages off the top of my head for using something like this :-
- Generally having a metric system like this makes asking questions of the system much easier. It’s especially useful when doing rollouts via feature gate to ensure we don’t break anything.
- Claude code loves that the data sits so cleanly in a S3 Tables ( which can easily be read by duckdb client). It’s incredibly powerful to ask Claude to get the answer to a specific question (especially after it instrumented the code).
- It’s simple and dirt cheap (So agents can go crazy running multiple queries against it).
- Duckdb can be embedded in metabase which is the source of our product analytics and acts complementary.