Cloudflare is becoming critical infrastructure for banks, governments, media companies, and health systems; the Customer Reliability Engineer applies the SRE discipline outward — from rapid incident response to proactive reliability engineering.
Stack
Responsibilities
- ▹Own the most complex, high-severity customer issues end-to-end, from first signal through confirmed resolution
- ▹Lead deep-dive debugging across the full stack: edge, network, DNS, transport, APIs, application, customer-side configuration
- ▹Reproduce defects, validate fixes with Engineering, and confirm customer-side resolution
- ▹Produce postmortems other engineers rely on
- ▹Hold on-call for high-severity incidents as part of a global rotation that includes weekends
- ▹Analyze support and telemetry signals across the customer base to find systemic risks before they become incidents
- ▹Contribute monitoring, detection, and diagnostic capability to the core product and engineering systems for early Customer Support visibility
- ▹Define customer-facing reliability metrics (error rates, resolution times, repeat-contact rates) and drive measurable improvement
- ▹Write automation that reduces mean-time-to-detect and mean-time-to-resolve
- ▹Manage the technical escalation lifecycle with clear ownership and timely communication
- ▹Partner with Product Engineering to drive fixes, workarounds, and configuration changes
- ▹Represent the customer reliability perspective in engineering syncs, incident reviews, and post-mortem processes
- ▹Raise the technical floor of Customer Support through pair-debugging, structured knowledge transfer, and shared tooling
- ▹Document diagnostic procedures and resolution patterns in runbooks, internal knowledge bases, and AI skills
- ▹Maintain deep, current expertise across Cloudflare's product portfolio: edge networking, DNS, CDN, WAF, DDoS mitigation, Zero Trust, Workers, and the developer platform
- ▹Anticipate customer impact from new releases and architecture changes
- ▹Help build agents and tooling that pre-diagnose incidents
Requirements
- ▹Minimum 5 years of hands-on experience in site reliability engineering, escalation engineering, systems engineering, or a comparable deeply technical support/operations role, with at least 2 years in customer-facing environments
- ▹Strong foundation in networking and security: TCP/IP, OSI model, IPv4/IPv6 addressing, subnetting, routing, switching
- ▹Core protocols: DNS, HTTP/S, TLS/SSL, SMTP, SNMP, NTP
- ▹Routing protocols: BGP, OSPF, including path selection and route propagation
- ▹Firewall concepts: stateful/stateless inspection, rule sets, NAT, ACLs
- ▹VPN and encryption: IPSec, SSL/TLS tunnels, GRE
- ▹Zero Trust architecture, network segmentation, modern security models
- ▹Proficiency with observability and diagnostic tooling: packet capture (Wireshark, tcpdump), log aggregation (Kibana, Elasticsearch), metrics dashboards (Grafana), distributed tracing
- ▹Strong scripting and automation skills (Bash, Python) with a track record of shipping reliability-improving tooling
- ▹Experience with incident management, postmortem culture, and SLO/SLI-based reliability metrics
About the company
Cloudflare runs one of the world's largest networks, protecting and accelerating internet applications for millions of websites without requiring new hardware or code changes. The company's mission is increasingly to be critical infrastructure: banks, governments, media companies, and health systems run on it. Cloudflare also runs mission-driven programs like Project Galileo and the Athenian Project, plus the free 1.1.1.1 public DNS resolver.
Ähnliche Stellen

Stelle
Delivery Engineer - Cloud & Automation
Bluestaq US External
+4
83 428–122 946 €/Jahr
brutto
🏢 Vor Ort
📍 Colorado Springs
🗣️ EN
Stelle
Senior Site Reliability Engineer
alpenlabs
+1
💰 Gehalt: keine Angabe
🏢 Vor Ort
📍 NAMER
🗣️ EN

Stelle
Staff Engineer (Product)
Later
AI/ML
+6
153 683–219 547 €/Jahr
brutto
🌍 Remote
📍 Remote
🗣️ EN

Stelle
Staff Engineer (Platform)
Later
AI/ML
+5
153 683–219 547 €/Jahr
brutto
🌍 Remote
📍 Remote
🗣️ EN

Stelle
Staff Software Engineer
BambooHR
+6
💰 Gehalt: keine Angabe
🏢 Vor Ort
📍 Utah | Hybrid
🗣️ EN

Stelle
Staff TDI Site Reliability Engineer, Okta Federal
Okta
AI/ML
💰 Gehalt: keine Angabe
🏢 Vor Ort
📍 Washington
🗣️ EN
