Cerebras builds the world's largest AI chip, the Wafer-Scale Engine, and operates high-speed AI training and inference infrastructure. This role is responsible for operating, monitoring, and building internal tooling for large-scale AI compute clusters.
Responsibilities
- ▹Deploy, configure and debug Docker-based container services
- ▹Build monitoring platforms, workflow automation and operational dashboards
- ▹Develop APIs and automation services for operations
- ▹Manage and operate multiple large-scale AI compute clusters
- ▹Monitor cluster health and proactively resolve issues
- ▹Provide 24/7 monitoring and troubleshooting with automated tools
Requirements
- ▹6-8 years of experience operating complex compute infrastructure
- ▹Strong proficiency in Python and Go
- ▹Experience with distributed systems
- ▹Deep knowledge of Linux-based systems
- ▹Expertise in Docker and Kubernetes-based orchestration
- ▹Experience with monitoring and alerting systems
Soft skills
Proactive problem-solving mindsetReliability and customer-success orientationAbility to work autonomously in complex environments
About the company
Cerebras builds the world's largest AI chip and delivers AI training and inference services significantly faster than traditional GPU-based solutions, serving leading model labs and enterprises including OpenAI.
Similar jobs

Job
Software Engineer Cloud EOS Team (Poland/Romania/Hungary/Greece)
Arista Networks
Cpp
+3
💰 Salary: not specified
🌍 Remote
🗣️ EN

Job
Software Engineering Technical Leader (Remote)
Cisco
+4
💰 Salary: not specified
🌍 Remote
49 Locations
🗣️ EN

Job
Software Engineer, Infrastructure - Self Managed Experience (SMX)
Camunda
+8
💰 Salary: not specified
🌍 Remote
🗣️ EN

Job
Software Engineer - Crypto
Alpaca
Cpp
+6
💰 Salary: not specified
🌍 Remote
EU - LATAM - NA
🗣️ EN

Job
Software Engineer
Cloudflare
+6
💰 Salary: not specified
🏢 On-site
🗣️ EN

Job
Distributed Systems Testing Software Engineer, Python / Go
Canonical Ltd.
AI/MLData Science
+10
💰 Salary: not specified
🌍 Remote
🗣️ EN
