← Back to list
Job

Site Reliability Engineer II (NGPOS Operations Support)

Backend Developer • On-site • Full-time • 📍 Cincinnati
Position Overview We are seeking a hands-on Site Reliability Engineer II to support a next-generation Point of Sale (NGPOS) platform in a highly visible production environment. Unlike traditional SRE roles focused primarily on automation or platform engineering, this position emphasizes production reliability, incident leadership, operational excellence, and engineering support . The ideal candidate will lead major incident response efforts, drive Root Cause Analysis (RCA), improve system observability, and collaborate closely with Software Engineering, Platform Engineering, Infrastructure, and Business Operations teams to enhance overall platform reliability. This role is ideal for someone who enjoys solving complex production issues under pressure while contributing to long-term engineering improvements. Location: Blue Ash, OH (Cincinnati) – Onsite (5 Days/Week) Employment Type: W2 – Contract-to-Hire Duration: Full-Time Work Authorization: Permanent Residents only. Must be able to convert to full-time without sponsorship. Experience Required: 3+ Years Nice to Have Enterprise Point of Sale (POS) systems Retail technology experience Automation scripting Monitoring optimization Runbook creation Store technology deployments Key Responsibilities Lead major incident response during production outages Serve as Incident Commander during P1/P2 incidents Coordinate technical bridge calls Communicate outage status to engineering teams and business leadership Lead Root Cause Analysis (RCA) activities Track corrective actions through completion Improve production reliability and system stability Enhance monitoring and observability Reduce alert fatigue Partner with Software Engineering and Platform Engineering teams Support retail store deployments Develop operational documentation, runbooks, and playbooks Participate in after-hours support rotations and maintenance windows Improve service health using SLIs and SLOs Technical Environment Monitoring & Observability Dynatrace Azure Monitor Log Analytics Metrics Dashboards Cloud Microsoft Azure Google Cloud Platform (GCP) Containers Kubernetes Docker Operating Systems Linux Scripting Languages Bash Python Agile Tools Jira Enterprise Environment Retail systems Point of Sale (POS) Production Support Hybrid Infrastructure Ideal Candidate Profile The ideal candidate will demonstrate: Strong leadership during production incidents Excellent troubleshooting and analytical skills Effective communication under pressure Ownership and accountability Experience coordinating multiple engineering teams Strong operational discipline Continuous improvement mindset Passion for reliability engineering Excellent documentation skills

Similar jobs