Infrastructure · Reliability · Automation

Santosh Bitra

Senior DevOps/SRE Engineer

I design and operate secure, reliable cloud platforms using AWS, Azure, Kubernetes, infrastructure as code, CI/CD, and observability.

Santosh Bitra monogram

Engineering evidence

  • ReversibleKubernetes chaos experiments designed with policy checks and rollback paths
  • Private by defaultEKS worker nodes with explicit network, IAM, and security boundaries
  • Recovery readyEncrypted incremental backups with retention and a documented restore path

Published work

Project case studies

View all projects
Chaos Sensei project mark showing a stable system around a controlled disruption
MVPKubernetes · Policy-driven automation

Chaos Sensei

Teams need realistic failure practice, but ad hoc chaos experiments can be difficult to understand, approve, and reverse safely.

Outcome: Established a provider-oriented design for adding failure mechanisms without coupling them to repository analysis.

Read the case study
Homelab recovery project mark showing encrypted backups moving from a server to cloud storage
PRODUCTIONRestic · Amazon S3

Homelab disaster recovery

A homelab can be rebuilt from code only if its stateful data and recovery procedure survive host or storage failure.

Outcome: Automated encrypted incremental backups to off-host Amazon S3 storage.

Read the case study
Production EKS project mark showing a Kubernetes cluster inside a protected AWS network
PRODUCTIONAmazon Web Services · Amazon EKS

Production-grade EKS infrastructure

Kubernetes teams need repeatable AWS infrastructure that makes network boundaries, identity, and validation explicit instead of relying on one-off cluster setup.

Outcome: Produced a reusable Terraform structure for EKS, networking, IAM, and security controls.

Read the case study

Capabilities

Engineering with evidence

Field notes

All writing

Browse all writing

Medium

We Ditched IRSA for Pod Identity When Setting Up Karpenter — Here’s Exactly How We Did It

A practical account of moving from Cluster Autoscaler to Karpenter on Amazon EKS with Pod Identity, including IAM setup, authentication decisions, and troubleshooting notes.

Read on Medium

Medium

Kafka Architecture — From Basics to Quick References of Kafka (1)

An approachable introduction to Kafka producers, clusters, consumers, data flow, and message-delivery semantics, supported by quick-reference diagrams.

Read on Medium

Medium

Grocery Shopping with HTTP Codes

A light-hearted explanation of common HTTP status codes using a familiar grocery-shopping scenario.

Read on Medium

Medium

Kubernetes — Practical HandNotes for DevOps Engineer (Part-1)

Practical notes on Kubernetes architecture, control-plane and worker-node components, request flow, foundational kubectl commands, and cluster setup resources.

Read on Medium

Medium

Kubernetes — Practical HandNotes for DevOps Engineer (Part-2)

The second set of practical Kubernetes notes, continuing the series with workload and day-to-day cluster concepts for DevOps engineers.

Read on Medium

Professional experience

Operating systems that teams depend on

My work focuses on secure cloud foundations, Kubernetes platforms, delivery automation, observability, and practical reliability engineering.

Explore experience and skills

Contact

Let's build reliable systems.

For roles, technical conversations, or collaboration, reach me directly.