← Back to list

Job
· Senior
Senior Site Reliability Engineer
DevOps / SRE
• Senior
• On-site
• Full-time
•
Dublin, Ireland
A senior SRE in Salesforce's Dublin Site Reliability organization, improving cloud service reliability through automation, observability and AI-powered tooling, and acting as a technical leader in incident management.
Stack
Responsibilities
- ▹Lead incident detection, response and resolution, driving root cause analysis, postmortems and proactive measures for high uptime
- ▹Lead post-incident reviews and drive systemic fixes through corrective actions
- ▹Independently design and implement complex automation platforms, self-healing systems and AI-powered operational tooling using durable workflow engines (Temporal, Airflow, Argo Workflows)
- ▹Architect and build production-grade observability solutions (monitoring, logging, alerting, tracing) enabling proactive detection and autonomous remediation
- ▹Design and implement AI/ML-powered operations tools: anomaly detection, predictive analysis pipelines, intelligent runbook automation and prompt-engineered operational agents (MCP-based)
- ▹Optimise system performance, reliability and cost-effectiveness through proactive monitoring and tuning
- ▹Ensure the Site Reliability team's work complies with the company's internal compliance policy and directives
- ▹Create comprehensive technical epics with clear problem statements, project documentation and measurable business outcomes
- ▹Provide technical coaching to junior team members through pair programming, design reviews and code reviews
- ▹Collaborate with engineering and product teams to define and uphold SLAs/SLOs
- ▹Build and ship production-grade software with modern engineering practices, using AI as a core part of the development workflow
- ▹Design and orchestrate systems where AI agents integrate into human workflows
- ▹Critically evaluate human or AI-generated code for correctness, quality, security and performance
- ▹Contribute to building and maintaining the shared system context
Requirements
- ▹5+ years of experience in systems and software engineering for large-scale, internet-facing services
- ▹Hands-on expertise with containerized architectures (Docker, Kubernetes) and orchestration platforms
- ▹Strong knowledge of distributed systems and Linux/Unix internals, with performance tuning and troubleshooting at scale
- ▹Familiarity with large-scale internet service architectures (DNS, HTTP, load balancing, caching, etc.)
- ▹Proven proficiency in Python and Go with strong software engineering practices (testing, code review, CI/CD)
- ▹Production experience building and operating observability platforms (Grafana, Prometheus, ELK, Splunk, Datadog or similar)
- ▹Solid background in incident management, including on-call participation, root cause analysis and postmortem practices
- ▹Strong understanding of SRE principles: SLIs/SLOs, error budgets, toil reduction, blameless culture, capacity planning
- ▹Hands-on experience with workflow/orchestration engines (Temporal, Airflow, Argo Workflows or similar)
- ▹Experience applying AI/ML to operations, including anomaly detection, predictive analysis, LLM-based automation and prompt engineering for operational agents
- ▹Excellent communication skills, with ability to lead during high-pressure incidents, present technical designs to leadership and mentor junior engineers
- ▹Track record of mentoring and technically coaching other engineers
- ▹Ability to work in a 24/7 global operations model, managing multiple priorities under time pressure
- ▹Growth mindset and curiosity to explore new technologies
- ▹A genuine AI-first approach to engineering, using AI to move faster and build fluency beyond your core specialty
- ▹Experience using AI tools (e.g. Claude Code, GitHub Copilot, Codex, Cursor) in development workflows
- ▹Advanced prompt engineering skills and the ability to cultivate the system context that makes AI outputs reliable, secure and production-ready
- ▹Understanding of AI/ML concepts applied to operations (e.g. anomaly detection, predictive analysis)
- ▹A related technical degree
Nice to have
- ▹Experience with AI agent frameworks, MCP (Model Context Protocol) or building LLM-powered operational tools
- ▹Contributions to open-source reliability/observability tooling
- ▹AWS/GCP professional-level certifications
- ▹Experience in SRE organizations supporting multi-cloud or hyperscale environments
- ▹Python and Go proficiency for systems-level tooling
- ▹Experience with chaos engineering and game day exercises
Soft skills
Excellent communication skillsLeadership under pressureMentoring and technical coachingGrowth mindset and curiosityManaging multiple priorities under time pressure
About the company
Salesforce describes itself as the number one AI CRM; its Site Reliability organization monitors cloud service availability globally in a follow-the-sun model.
Education: Releváns műszaki diploma
Similar jobs

Job
· Senior
Senior Site Reliability Engineer - Platform Reliability (Resilience)
Elastic
+2
€98,400–€126,900/yr
gross
🏢 On-site
🗣️ EN

Job
· Senior
Senior Site Reliability Engineer
Okta
AI/MLArgocd
+7
€76,000–€104,500/yr
gross
🏢 On-site
Dublin
🗣️ EN

Job
· Senior
Senior Site Reliability Engineer
Fivetran
Argocd
+14
💰 Salary: not specified
🏢 On-site
Dublin
🗣️ EN

Job
· Senior
Senior Site Reliability Engineer (MAAS)
Pragmatike
+6
💰 Salary: not specified
🔀 Hybrid
Armenia
🗣️ EN

Job
· Senior
Senior DevOps Engineer – Observability Platform
Qualcomm
CppCloudformation
+7
💰 Salary: not specified
🏢 On-site
Chennai
🗣️ EN

Job
· Senior
Senior DevOps Engineer (Cloud-Native, AI-Driven Platform)
septeo
AI/MLBitbucket
+16
💰 Salary: not specified
🌍 Remote
España la Vieja
🗣️ EN