Senior DevOps / Infrastructure Engineer
The vacancy is well-structured and informative, providing clarity on tasks, compensation, and company background.
Check Match — Just drop your CV
See your fit for Senior DevOps / Infrastructure Engineer in seconds.
Overview
Join Category Labs as a Senior DevOps Engineer to operate and enhance the infrastructure behind Monad, focusing on AI-driven operations. Competitive salary and benefits included. Category Labs (formerly known as Monad Labs) is a team of systems engineers and researchers on a mission to design and build at the frontier of decentralized technology. We strive to deliver significant improvements over existing blockchain solutions. After raising $225M in series A funding, led by Paradigm, we are growing our team. We’re the team behind Monad, a high-performance, EVM-compatible Layer 1 whose public mainnet is now live. We write the core software that runs it: a parallel-execution EVM, a custom state database, and a BFT consensus client, all developed in the open.
What You'll Do
- •Operate the Monad node fleet: health, sync, upgrades, and recovery across validators, full nodes, archive/historical, and indexer nodes on mainnet and testnet, including safe, staged rollouts and incident response.
- •Own our infrastructure-as-code: Ansible for fleet configuration, Terraform + Atlantis for cloud and DNS, and Kubernetes/Flux (GitOps) for platform services.
- •Build and operate observability and alerting (Prometheus, Grafana, Loki); create dashboards and alerts that catch problems before they page while minimizing false positives.
- •Automate the release pipeline: node upgrades, canary rollouts, snapshot/restore, and the guardrails that bound blast radius (e.g., protecting validators from automated changes).
- •Design and build agentic operations: develop AI agents, tooling (e.g., MCP servers), and runbooks-as-code that let agents safely investigate, diagnose, and execute routine operations, with deterministic guardrails and human oversight.
- •Codify operational knowledge into tools and automation that the whole team, and its agents, can reuse.
- •Harden nodes and services, manage secrets, and continuously drive down manual toil.
Why Work with Us
- •Challenging problems: You’ll work on extremely challenging problems with massive impact. See our **Blogs** and **Publications & Talks** for a flavor of the problems we are solving in the real world.
- •Huge opportunity: The Ethereum Virtual Machine (EVM) standard is ubiquitous, but existing EVM-compatible chains are very slow. Monad’s core innovations offer developers the best of both worlds (portability and performance) and are a game-changer for mass user adoption in crypto.
- •The right team: You’ll be part of a small, exceptional team (engineers and researchers make up 90% of the team).
- •Open by default: Our core software is public on GitHub. You’ll build in the open, and your work ships where the whole ecosystem can see it.
- •Culture: We’re a lean team working together to achieve very ambitious goals. We are united in our culture of collaboration, low ego, and high-quality output. As an early member of our team, you’ll help to shape our culture.
- •Compensation: You’ll receive a competitive salary and equity package.
- •Resources and growth: We’re well-capitalized, with backing from leading venture funds like Paradigm, Electric Capital, Greenoaks, Dragonfly, and Coinbase Ventures. We keep a lean team, and this is a rare opportunity to join. You’ll learn a lot and grow as our company scales.
Who You Are
- •You have 5+ years in DevOps, SRE, or Infrastructure Engineering, operating production systems at scale.
- •You have strong Linux, systemd, networking, and shell fundamentals, and you're comfortable debugging live systems over SSH.
- •You have deep, hands-on infrastructure-as-code experience with Ansible and Terraform.
- •You have experience with observability stacks (Prometheus, Grafana, Loki, or equivalents).
- •You have hands-on fluency with AI-assisted engineering: you use coding agents and LLM tooling in your daily workflow and have judgment on where it helps and where it's risky.
- •You have experience designing automation with safe guardrails, and you bring calm, methodical incident response.
- •You have programming and scripting experience (e.g., Python, bash).
- •Experience with Kubernetes and GitOps (Flux or Argo) is a plus.
- •Experience building AI agent tooling, MCP servers, or agent orchestration frameworks is a plus.
- •Experience serving inference, either locally or as a service is a plus.
- •Previous experience with blockchain clients or node operations is a plus.
- •A Bachelor of Science in Computer Science, Engineering, or a related field is a plus.