Platform Engineering · Chennai, India
I build the platforms other engineers build on.
Platform engineering lead with 8 years turning unstable, hand-run systems into declarative, self-service platforms. I write here about the real work: GitOps, internal developer platforms, SRE, and keeping the cloud bill honest.
open to talkSelected impact
Numbers from the last four years
Availability
of 100% uptime
sustained on mixed node pools
MTTR
reduction
observability + incident response
Cloud cost
reduction
spot + reserved blending
Config drift
reduction
terraform standardisation
Experience
Eight years, three teams
Senior SDE (Manager), Platform Engineering
SRE → Senior SRE → Technical Lead
Software Engineer (Associate → SE)
8y 7m across 3 teams
Recent blogs
Latest from the log
OTel agent to Datadog: logs without the Datadog agent
Shipping logs to Datadog through an OpenTelemetry Collector: the exporter config, the batching that causes 413s, the hostname bug that restarts your pods, and keeping the bill in your control.
OTel agent to Loki: logs without a log shipper
The Loki exporter is gone from the Collector. What replaced it, how OTLP attributes become Loki labels, and the two limits that will reject your logs in production.
Monitoring setup for PySpark applications
Spark applications are ephemeral, so the usual scrape model does not fit. Wiring the Prometheus servlet, catching the Python memory that JVM metrics never show, and keeping evidence after the driver exits.
Monitoring setup for Airflow
Airflow 3 emits metrics through StatsD or OpenTelemetry, and the choice changes your Prometheus config. The pipeline, the metrics that predict failures, and the alerts worth paging on.
Work
Platforms I own or built
Internal Developer Platform
Self-service infrastructure on Kubernetes-native APIs, with golden paths that let product teams provision without a ticket in sight.
GitOps Auto-PR Agents
Automation that raises pull requests for platform updates across every product repo: consistent propagation, human-in-the-loop merges.
Release Self-Service Skill
A Claude-powered assistant that walks product teams through platform upgrades step by step, cutting onboarding friction and support load.
FinOps Cost Engine
Cost dashboards and spend-leak detection feeding a Spot + On-Demand + Reservation strategy that cut compute spend ~40%.
About
Who you're reading
I'm Praveen, a platform engineer based in Chennai. I like the unglamorous middle of the stack: the paved roads, the reconciliation loops, the guardrails that turn “please raise a ticket” into “just push to main.”
Over eight years I've moved from firefighting SRE to building platforms product teams actually want to use, treating the platform as a product, with developer experience as the metric that matters.
This site is where I think out loud. New posts land most days: short, specific, and drawn from real work rather than trend cycles.