Cohere

Engineering Manager, GPU Infrastructure

8.0/10
Cohere
Not specified
Office / on-site
mid
about 3 hours ago
AI SummaryVerified by Aipplify AI

The vacancy is well-structured with clear responsibilities and company information, but lacks specific salary details.

AI quality score8.1 / 10

Check Match — Just drop your CV

See your fit for Engineering Manager, GPU Infrastructure in seconds.

Overview

Cohere is seeking an Engineering Manager for GPU Infrastructure to lead a team focused on building and operating superclusters for AI models. This role involves technical strategy, team leadership, and cross-functional collaboration. Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!

Team Leadership & Development

  • Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement
  • Manage performance, career development, and hiring for team members
  • Conduct regular 1:1s and team meetings to ensure alignment and address challenges
  • Provide technical guidance and support to team members on complex infrastructure problems

Technical Strategy & Execution

  • Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling
  • Oversee the implementation of workload scheduling and queuing, hardware fault detection, and performance optimization systems
  • Collaborate with cloud providers and MLEs to adapt our training and inference stack to bleeding-edge GPU architectures
  • Ensure infrastructure reliability, scalability, and security across all GPU environments

Cross-Functional Collaboration

  • Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions
  • Work with the research teams on training software stack adaptation for new GPU architectures
  • Coordinate with Capacity EPM and Finance to manage capacity of a rapidly growing compute footprint
  • Interface with Legal and Security teams on compliance requirements
  • Collaborate with other infrastructure teams on shared goals and dependencies

Operational Excellence

  • Establish observability and monitoring frameworks for GPU utilization, performance, and reliability
  • Drive practices and policies to automate cluster provisioning and management
  • Drive cost optimization initiatives while maintaining performance standards
  • Manage vendor relationships and contract negotiations for hardware and cloud services

Full-Time Employees at Cohere Enjoy These Perks

  • A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.
  • Full health and dental benefits, including a separate budget for mental health.
  • RRSP matching, 401K, Pension Scheme.
  • 100% Parental Leave top-up for up to 6 months, for either parent.
  • Annual enrichment benefits:
  • Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.
  • Education & learning stipend for conferences, courses, and coaching.
  • 6 weeks of paid vacation (30 working days!)
  • Budget for traveling to other offices if you are remote, plus an annual company offsite.

Leadership & Management Skills

  • Experience managing engineering or SRE teams with a focus on technical mentorship and growth
  • Strong communication skills to translate complex technical concepts for diverse audiences
  • Ability to make data-informed decisions under pressure
  • Experience working in remote, distributed teams
  • Commitment to fostering an inclusive and collaborative team culture

Technical Expertise

  • Deep expertise in ML/HPC infrastructure: GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and high-performance computing environments
  • Proven experience with Kubernetes at scale: deployment, management, and troubleshooting cloud-native clusters for AI workloads in multi-cloud environments
  • Knowledge of infrastructure monitoring tools (Prometheus, Grafana)
  • Familiarity with Terraform, ArgoCD, or other IaC tools
  • Experience with cost optimization and capacity planning for GPU infrastructure
  • Track record of collaborating with AI researchers or ML engineers to solve infrastructure challenges

Personal Qualities

  • Strong problem-solving abilities with a data-driven approach
  • Passion for enabling AI research through robust infrastructure
  • Collaborative mindset with a focus on cross-team success
  • Willingness to learn and adapt in a fast-paced, evolving environment
Loading similar jobs...