SVS recruiting recruiting
← Back to jobs

Principal Engineer, DevOps

⌂ Plane ▤ 10.0+ years ◉ Hyderabad ◎ devops

Principal DevOps Engineer with 10+ years owning infrastructure for high-availability distributed systems. Lead design and automation of Kubernetes, GitOps, CI/CD, and secure multi-environment deployments across cloud, self-hosted, and air-gapped infrastructure.

०१ At a glance
Location
Hyderabad
Experience
10.0+ years
Published
Sep 17, 2026
०२ Description

About Plane

Plane’s mission is to build the infrastructure the world’s work runs on. Every organization runs on three things: the projects it’s driving, the knowledge it keeps, and the requests it fields. Plane brings all three into one open, adaptable platform: simple enough for any team to adopt, dependable enough for organizations to build on. And we are building it for a future where humans and AI agents do that work together.

Plane began in public on GitHub at the end of 2022. Since then it has grown into a work management platform used by teams around the world: 58,000+ stars, 5,000+ forks, and a contributor community that reads our code and files our issues. Organizations run Plane as a managed Cloud service, on their own infrastructure, or inside fully isolated environments. Building in the open keeps us close to users and raises the standard for everything we ship.

Plane is the #1 work infrastructure in aerospace, defense, financial services, and other regulated industries: organizations whose requirements for control, auditability, and data residency rule most software out. When the strictest buyers pick a system of record, that choice means something. Adoption is growing fastest on Plane Cloud and in sovereign clouds, deployments that keep everything inside a country’s own borders and rules.

Plane is backed by top investors and built across San Francisco, London, and Hyderabad. We work in tightly knit teams, stay close to users, and care about the visible product as much as the unglamorous details that make software dependable. People own problems end to end, and we add process only when it helps the work.

Humans and agents

We believe the next decade of work will be done by humans and AI agents together. Not agents replacing people, and not a chat window bolted onto software built for humans clicking around, but both working in one system of action, where an agent’s work is as visible and as accountable as a person’s.

Most software treats AI as a feature. We treat agents as a kind of bot/worker, and that changes what the system underneath has to be. Agents are only as good as the context they can see and the state they are allowed to change, so shared context, explicit state, durable history, and accountable action are not items on our roadmap. They are the product. Plane is built so that when an agent acts, the humans responsible can see what happened, why, and on whose authority, and the record survives.

This is what the infrastructure is for: making the future where humans and agents work together useful, legible, and fully within the organization’s control. Every role at Plane is some part of building that.

About the role

Plane runs across environments we control and environments we never see: Plane Cloud, sovereign clouds, customer-managed infrastructure, self-hosted clusters, and fully air-gapped networks. As a Principal DevOps Engineer, you will own the infrastructure and operational architecture that makes this possible. You will lead how we build, deploy, secure, observe, and operate Plane across these environments, from Kubernetes, GitOps, and CI/CD to databases, incident response, compliance, and infrastructure cost. As Plane expands its agentic capabilities, you will also help build the infrastructure required to securely run AI workloads and open-weight models across cloud and customer data centers. You will work across engineering teams to make deploying and operating Plane reliable, secure, and increasingly automated at every scale.

What you’ll do

  • Experienced in operating high-availability, fault-tolerant, scalable, distributed software across cloud and self-hosted deployments.
  • Hands-on experience in handling cloud infrastructure on AWS.
  • Design and automate infrastructure, CI/CD, GitOps, secrets management, AMIs, and one-click deployments across AWS, DigitalOcean, Heroku, and other platforms.
  • Set up processes to improve visibility across cloud infrastructure. Maintain incident management, disaster recovery, database migrations, and production reliability.
  • Experienced in handling infrastructure security, vulnerability management, SAST, compliance, data governance, and security standards.
  • Drive infrastructure performance and cost optimization while improving deployment speed, reliability, and developer experience.

What you’ll bring

  • You have 10+ years in DevOps, SRE, platform, or infrastructure engineering, and you have owned the systems you built in production.
  • You have run AWS, Kubernetes, and Docker in production and managed secrets.
  • You have managed infrastructure as code with Terraform and Ansible, and shipped changes through CI/CD and GitOps with Argo CD.
  • You have shipped and supported enterprise software across multiple clouds, self-hosted setups, and air-gapped (fully offline) environments.
  • You have managed production Postgres and Redis and moved them between clusters without losing data.
  • You write Python and Bash well.
  • You understand networking, reverse proxies, monitoring, and infrastructure security well enough to debug them in production.

Nice to haves

  • You have run bare-metal deployments and worked in restricted enterprise environments.
  • You have deployed and operated open-weight LLM models across public cloud and on-premise data centers.
  • You have worked with serverless architectures and sandboxed execution environments.
  • You have built infrastructure for AI agents or agentic development environments.
  • You have operated infrastructure under compliance, data residency, or sovereign-cloud requirements.

Tech

Cloud and infrastructure: AWS, Kubernetes, Rancher, Docker, Terraform, Ansible, Helm, AMIs

Delivery: Argo CD, GitOps, GitHub Actions, CI/CD, DigitalOcean, Heroku

Observability: Grafana, Prometheus, Datadog, Sentry

Networking: Nginx, Caddy, Kong, Traefik, DNS, load balancing

Data and security: Postgres, Redis, secrets management, SAST, vulnerability and container scanning.

Ready when you are. Apply