Cloudflare is becoming critical infrastructure for banks, governments, media companies, and health systems; the Customer Reliability Engineer applies the SRE discipline outward — from rapid incident response to proactive reliability engineering.
Stack
Responsibilities
- ▹Own the most complex, high-severity customer issues end-to-end, from first signal through confirmed resolution
- ▹Lead deep-dive debugging across the full stack: edge, network, DNS, transport, APIs, application, customer-side configuration
- ▹Reproduce defects, validate fixes with Engineering, and confirm customer-side resolution
- ▹Produce postmortems other engineers rely on
- ▹Hold on-call for high-severity incidents as part of a global rotation that includes weekends
- ▹Analyze support and telemetry signals across the customer base to find systemic risks before they become incidents
- ▹Contribute monitoring, detection, and diagnostic capability to the core product and engineering systems for early Customer Support visibility
- ▹Define customer-facing reliability metrics (error rates, resolution times, repeat-contact rates) and drive measurable improvement
- ▹Write automation that reduces mean-time-to-detect and mean-time-to-resolve
- ▹Manage the technical escalation lifecycle with clear ownership and timely communication
- ▹Partner with Product Engineering to drive fixes, workarounds, and configuration changes
- ▹Represent the customer reliability perspective in engineering syncs, incident reviews, and post-mortem processes
- ▹Raise the technical floor of Customer Support through pair-debugging, structured knowledge transfer, and shared tooling
- ▹Document diagnostic procedures and resolution patterns in runbooks, internal knowledge bases, and AI skills
- ▹Maintain deep, current expertise across Cloudflare's product portfolio: edge networking, DNS, CDN, WAF, DDoS mitigation, Zero Trust, Workers, and the developer platform
- ▹Anticipate customer impact from new releases and architecture changes
- ▹Help build agents and tooling that pre-diagnose incidents
Requirements
- ▹Minimum 5 years of hands-on experience in site reliability engineering, escalation engineering, systems engineering, or a comparable deeply technical support/operations role, with at least 2 years in customer-facing environments
- ▹Strong foundation in networking and security: TCP/IP, OSI model, IPv4/IPv6 addressing, subnetting, routing, switching
- ▹Core protocols: DNS, HTTP/S, TLS/SSL, SMTP, SNMP, NTP
- ▹Routing protocols: BGP, OSPF, including path selection and route propagation
- ▹Firewall concepts: stateful/stateless inspection, rule sets, NAT, ACLs
- ▹VPN and encryption: IPSec, SSL/TLS tunnels, GRE
- ▹Zero Trust architecture, network segmentation, modern security models
- ▹Proficiency with observability and diagnostic tooling: packet capture (Wireshark, tcpdump), log aggregation (Kibana, Elasticsearch), metrics dashboards (Grafana), distributed tracing
- ▹Strong scripting and automation skills (Bash, Python) with a track record of shipping reliability-improving tooling
- ▹Experience with incident management, postmortem culture, and SLO/SLI-based reliability metrics
About the company
Cloudflare runs one of the world's largest networks, protecting and accelerating internet applications for millions of websites without requiring new hardware or code changes. The company's mission is increasingly to be critical infrastructure: banks, governments, media companies, and health systems run on it. Cloudflare also runs mission-driven programs like Project Galileo and the Athenian Project, plus the free 1.1.1.1 public DNS resolver.
Similar jobs

Job
Delivery Engineer - Cloud & Automation
Bluestaq US External
+4
$95,000–$140,000/yr
gross
🏢 On-site
📍 Colorado Springs
🗣️ EN
Job
Senior Site Reliability Engineer
alpenlabs
+1
💰 Salary: not specified
🏢 On-site
📍 NAMER
🗣️ EN

Job
Staff Engineer (Product)
Later
AI/ML
+6
$175,000–$250,000/yr
gross
🌍 Remote
📍 Remote
🗣️ EN

Job
Staff Engineer (Platform)
Later
AI/ML
+5
$175,000–$250,000/yr
gross
🌍 Remote
📍 Remote
🗣️ EN

Job
Staff Software Engineer
BambooHR
+6
💰 Salary: not specified
🏢 On-site
📍 Utah | Hybrid
🗣️ EN

Job
Staff TDI Site Reliability Engineer, Okta Federal
Okta
AI/ML
💰 Salary: not specified
🏢 On-site
📍 Washington
🗣️ EN
