Just to start us off (Disclaimer: I am one of the judges)
I remember moving from a world of on-prem systems, to that of one based in the clouds. Effectively someone else’s computer. We thought we were pretty nifty with lambdas, and serverless helping connecting multiple services together, until something went wrong and we were trying to sift through the logs.
There were a-lot of them, and it was like looking for a needle in a haystack, with welding goggles on. We very quickly learned about and implemented distributed tracing.
That wasn’t the end of our troubles. We had an MI system, that would get populated by way of a csv uploaded to an S3 bucket. This process started via our Postgres database setup with logical replication. Wal2json provided json changeset messages which were sent on a kinesis stream, enriched and transformed, and uploaded to the S3 bucket.
We got a call from management, that the payout system has been relatively quiet for the last few days. It transpired that we had run a large database update, pushing our changeset message over 1mb, which was the limit for kinesis messages.
Our code was unaware of this error condition, so blindly would try and re-attempt. In the end, we had to fix forward by rescuing from the error and then splitting our message into chunks.
The morale of this story, is bugs are everywhere, and you can only fix what you can see, so try to capture as much as you can from your running systems, so that you can lift off those welding goggles, blow away the haystack and be left with a treasure map to the root cause.