Staff Infrastructure Engineer · GCP · reliability · system design
Writing
Distributed Consensus Explained: From Paxos Theory to Real-World Systems
Imagine you’re building a banking system where multiple servers need to agree on account balances. Server A thinks Alice has $100, but Server B thinks she has $50 because it missed a recent transaction.
Cloud Cost Optimization: A Senior Engineer’s Guide
Cloud costs are spiraling out of control for most organizations. The market for cost optimization is exploding because companies are burning cash on oversized instances, idle resources, and poor architectural choices.
Building Resilient Cloud Applications: How Reactive Architecture Solves Modern Scaling Challenges
In today’s cloud-native world, traditional request-response architectures are hitting their limits. Applications struggling under load, cascading failures bringing down entire systems, and the complexity of coordinating asynchronous operations across microservices highlight…
Service Mesh: Taming the Complexity of Service-to-Service Communication
As microservices architectures have evolved, service-to-service communication has become increasingly complex. Different teams often implement their own approaches to handling retries, timeouts, and circuit breakers — some using language-specific libraries, others building…
Modular Monoliths: The Architecture That Dares to Stay Together
Modular monoliths offer the modularity and clear boundaries of microservices with the simplicity of monolithic deployment. Companies like Shopify, GitHub, and Basecamp have built massive systems this way, achieving team autonomy and clear boundaries without distributed…
From Monoliths to Microservices and Back: A Decade of Hard-Won Lessons
When Amazon Prime Video announced they’d moved their video monitoring and analytics services back to a monolithic architecture to reduce latency and operational complexity, it sent shockwaves through the tech industry. This was particularly striking coming from Amazon — one…
Observability and SRE: Building Systems You Can Actually Debug
In distributed systems, traditional monitoring often fails when you need it most. A service can appear healthy on infrastructure metrics while users experience degraded performance due to complex interactions between components.
Circuit Breakers: Preventing Cascade Failures in Distributed Systems
The circuit breaker pattern protects distributed systems from the cascade effects of downstream failures. Named after electrical circuit breakers that prevent electrical fires by cutting power during overloads, software circuit breakers monitor interactions with external…
Bulkheads: Isolating Failure Domains
The bulkhead principle isolates different parts of a system to prevent failures in one area from affecting others. Named after the watertight compartments in ships that prevent a single breach from sinking the entire vessel, software bulkheads create isolated failure domains…
Beyond Java Beans: Tactical Domain-Driven Design for Richer Domain Models
How to move from anemic objects to expressive domain models that capture business complexity
Every architectural decision is a trade-off.
Every architectural decision is a trade-off. When you choose one path, you’re implicitly saying no to others.
The Art of Drawing Boundaries: Mastering Decomposition in Software Architecture
Software architecture, at its core, is about defining components and how they relate to each other. But what makes an architect draw a rectangle on their diagram?
Kubernetes Resource Hierarchy Guide
Kubernetes orchestrates containerized applications through a rich ecosystem of interconnected resources. Understanding how these resources relate to each other is crucial for effective cluster management and application deployment.
Kubernetes Internal Architecture: Deep Dive
Kubernetes has revolutionized how we think about deploying and managing containerized applications at scale. But beneath its elegant command-line interface lies a sophisticated distributed system that exemplifies many of the core principles we study in distributed systems design.
Kafka Exactly-Once Semantics: How It Really Works
Exactly-once message processing has long been the “holy grail” of distributed systems. For decades, developers were forced to choose between losing messages (at-most-once) or handling duplicates (at-least-once).
Load Balancers, API Gateways, Proxies & CDNs Explained: The Network Components That Scale Your App…
Modern web applications rely on critical network components that enable scaling, resilience, and global performance. This article covers the essential building blocks every system architect should understand: load balancers, API gateways, proxies, CDNs, and related…
Kubernetes Networking: A Complete Guide from Basics to Advanced
Networking is truly the backbone of Kubernetes. And understanding Kubernetes networking isn’t just about memorizing concepts — it’s about building the foundational knowledge that helps you reason about complex distributed systems, troubleshoot issues quickly, and design…
Stop Getting Lost in AWS Networking: A Developer’s Guide to VPCs, Security Groups, and Route Tables
Networking forms the foundation of cloud infrastructure, with virtually every service depending on it for communication and connectivity. While many engineers assume networking is solely the domain of system administrators or DevOps teams, this perspective can severely limit…
OAuth 2.0, Passkeys, and Zero Trust: The Modern Authentication Stack Explained
Let’s be honest — every time you log into an app, there’s a complex dance happening behind the scenes that most people never think about. But if you’re building systems that handle user data (and let’s face it, what system doesn’t these days?), understanding authentication…
Sharding vs Replication: The Mental Model You’ll Actually Understand
After consuming numerous books, videos, courses and solving real production problems all related to distributed systems, I think I’ve come to a conclusion how to better present the fundamental ideas of distributed computing, so that it clicks in one’s mind and starts to make…
From Message Queues to Global Streams: The Rise of Event-Driven Architectures
The journey from simple message queues to sophisticated streaming platforms represents one of the most significant architectural shifts in distributed systems. This evolution started back in the 1980s and 1990s when the first commercial message queue systems entered the…
Rate Limiting at Scale: Lessons From Production SaaS Systems
How should we design a SaaS application to deal with a rise in incoming traffic? Of course, the answer is to design it for scalability to handle increased load, be it sudden burst of traffic or steady usage growth.
Beyond the Buzzwords: What ACID and BASE Really Tell Us About System Design
ACID and BASE are two acronyms that offer a convenient but somewhat reductionist categorization of systems (e.g., databases) and their guarantees. They reflect broader design philosophies governing how systems handle consistency, availability, and reliability .
Transaction Isolation Levels: Build Understanding from the Ground Up
How many times have you tried to get your head around and memorize transaction isolation levels? It’s a topic that would pop up in almost every single book on DBs or architecture.
Implementing Distributed Locks Correctly
Distributed locks allow multiple processes to coordinate access to shared resources by ensuring mutual exclusion (only one holder at a time). They are essential for consistency and coordination in distributed systems.
Understanding the Domain of DDD: A Strategic and Tactical Design Mental Map
Let’s get our head around the Domain of DDD . I’ve always found it to be a valuable concept that offers a systematic approach to analyzing business problems and modeling solutions in software.
Scalability in Databases. Exploring Different Approaches Across Relational, NoSQL, and OLAP systems
As I’ve navigated through various engineering roles and architectural decisions over the years, I’ve come to realize that understanding database scalability isn’t just about memorizing features or benchmarks — it’s about building a mental map of how different systems think…