Reliability over noise
Monitoring should catch real issues, not generate alerts nobody trusts. I'd rather have fewer, meaningful alerts than a dashboard full of things everyone's learned to ignore.
The principles behind how I approach infrastructure, wherever I'm working.
Monitoring should catch real issues, not generate alerts nobody trusts. I'd rather have fewer, meaningful alerts than a dashboard full of things everyone's learned to ignore.
Not automation for its own sake — the point is removing genuinely repetitive toil, handling edge cases properly, and leaving something the next person can actually maintain.
If I'm the only person who knows how something works, that's a risk, not job security. Good docs mean the team isn't stuck waiting on one person at 3am.
Provisioning access, reviewing what's actually still needed, and closing gaps before they become incidents — not glamorous, but it's most of what actually keeps a business safe.
The same two things, regardless of what I'm actually working on.
Experience running Linux at scale: monitoring (Zabbix + Pingdom), config management (Ansible), and making changes that don't break things at 2am.
Status updates and human explanations, not just logs and graphs. Whoever I'm working with knows what's happening and why.
Same process whether it's a five-minute fix or a multi-week project.
What's actually broken, who it affects, and how urgent it really is — before touching anything.
What I'm going to do, what could go wrong, and how to roll it back if it does.
Make the change, verify it actually worked, watch for knock-on effects.
So the next person — possibly future me — isn't starting from zero.
Always happy to chat infrastructure, automation, or anything on this site.
Get in touch