← Zurück zur Liste
Stelle

Cluster Operations Software Engineer

Softwareentwickler • Remote • Vollzeit • Europäische Union EU/EMEA

Cerebras builds the world's largest AI chip, the Wafer-Scale Engine, and operates high-speed AI training and inference infrastructure. This role is responsible for operating, monitoring, and building internal tooling for large-scale AI compute clusters.

Responsibilities

  • ▹Deploy, configure and debug Docker-based container services
  • ▹Build monitoring platforms, workflow automation and operational dashboards
  • ▹Develop APIs and automation services for operations
  • ▹Manage and operate multiple large-scale AI compute clusters
  • ▹Monitor cluster health and proactively resolve issues
  • ▹Provide 24/7 monitoring and troubleshooting with automated tools

Requirements

  • ▹6-8 years of experience operating complex compute infrastructure
  • ▹Strong proficiency in Python and Go
  • ▹Experience with distributed systems
  • ▹Deep knowledge of Linux-based systems
  • ▹Expertise in Docker and Kubernetes-based orchestration
  • ▹Experience with monitoring and alerting systems

Soft skills

Proactive problem-solving mindsetReliability and customer-success orientationAbility to work autonomously in complex environments

About the company

Cerebras builds the world's largest AI chip and delivers AI training and inference services significantly faster than traditional GPU-based solutions, serving leading model labs and enterprises including OpenAI.

Ähnliche Stellen