← Zurück zur Liste
Stelle

Database Administrator - Bengaluru Only (m/f/d)

Database Engineer / DBA • Vor Ort • Vollzeit • Indien Bengaluru, Indien
At CUCULUS , we build intelligent digital solutions that help utilities become more efficient, sustainable, and future-ready. By joining our technology team, you’ll work on business-critical systems that power large-scale utility operations, contribute to high-impact projects, and collaborate with global teams across cloud and hybrid environments. As a Senior Database Administrator, you will be responsible for highly available, mission-critical database environments and ensure their availability, performance, security, resilience and recoverability. You will: Administer and support Oracle 19c databases, including RAC and Data Guard. Manage PostgreSQL databases across on-premise and cloud environments. Administer and support ClickHouse database environments. Support business-critical 24/7/365 production operations within defined SLAs. Perform database performance tuning, SQL optimization and capacity planning. Implement proactive database monitoring and health checks. Plan and execute database migrations, upgrades and patching. Maintain backup, restore and recovery strategies. Support High Availability and Disaster Recovery environments. Deploy and support database workloads on Kubernetes. Troubleshoot database, Linux , storage, connectivity and performance issues. Develop scripts and automation for recurring DBA activities. Perform deep root-cause analysis for database-related incidents. Support critical production incidents and customer escalations. Maintain DBA procedures, runbooks and technical documentation. Work closely with System Engineering, Support and Development teams. Your work will contribute directly to: Production database availability Database reliability and resilience Performance and scalability Faster incident resolution SLA fulfilment Backup and recovery readiness HA/DR readiness Proactive monitoring Automation and operational efficiency Continuous improvement of database operations Your mission is to ensure that our database platforms remain stable, performant, secure, scalable and recoverable. You will take ownership of complex database issues from investigation through recovery and root-cause analysis while proactively identifying risks before they affect production systems or customers We are looking for a Senior Database Administrator with strong expertise in Oracle, PostgreSQL and ClickHouse, backed by solid Linux, Kubernetes and cloud knowledge. You will manage and support highly available, mission-critical production database environments across on-premises, cloud and hybrid platforms. The role requires strong hands-on technical capability, independent troubleshooting, production ownership and the ability to support critical database incidents within strict SLAs. Strong hands-on experience with Oracle 19c, including: Oracle RAC Data Guard Database administration Performance tuning SQL optimization Execution analysis Backup and recovery Capacity management Patching Upgrades Migration Security High Availability Disaster Recovery PostgreSQL Administration Strong production PostgreSQL administration, including: Installation and configuration Database administration Backup and restore Replication PITR Recovery Performance analysis Query troubleshooting Migration Upgrades Capacity management Security High Availability ClickHouse Administration Hands-on ClickHouse administration and operational support, including: Installation and configuration Cluster operations Query troubleshooting Performance monitoring and tuning Storage and capacity management Backup and recovery Replication High Availability Performance & Capacity Management Monitor, analyse and optimize database performance, including: SQL/query optimization Execution analysis Database performance tuning CPU and memory utilization Storage utilization Capacity planning Connection/session analysis Bottleneck identification Storage growth Database scalability Proactively identify capacity and performance risks before they become production incidents. Monitoring & Observability Implement, maintain and continuously improve database monitoring and observability. Experience with MONIT, Grafana, Prometheus or equivalent tools is expected. Monitor: Database availability Database health Performance Capacity Storage Connections and sessions Replication Backup status HA/DR status Logs Alerts Backup, Recovery & Disaster Recovery Design, maintain and support reliable database backup and recovery capabilities, including: Backup strategies Restore procedures PITR Replication Recovery procedures Recovery validation Disaster Recovery Failover/failback Recovery testing Production recovery Linux Administration Strong hands-on Linux administration and troubleshooting skills, including: Processes and services CPU and memory Disk and storage File systems Permissions Networking Logs Security Performance troubleshooting The DBA should be able to determine whether a database issue originates from the database itself or from the underlying operating system, storage, infrastructure or network. Kubernetes Experience supporting databases in production Kubernetes environments, including: Pods Services Networking Storage Persistent volumes Configuration Monitoring Logs Database workload troubleshooting Ability to deploy and support database workloads on Kubernetes. Cloud & Infrastructure Experience with: Oracle Cloud Infrastructure (OCI) AWS and/or Azure Cloud database services RDS/PaaS database services Cloud networking Storage Security Access control HA and multi-AZ concepts Backup and recovery Hybrid environments Scripting & Automation Experience with Shell/Bash, Python or equivalent scripting technologies. Use automation to improve: Database health checks Monitoring Backup validation Capacity checks Diagnostics Log collection Reporting Repetitive DBA activities Operational efficiency Database Security & Reliability Support and maintain: Database access controls Users and roles Permissions Security patches Database hardening Linux security Operational security controls Database reliability standards Incident Management & Root-Cause Analysis Independently handle complex and critical database incidents. Responsibilities include: Incident investigation Database troubleshooting Service recovery Root-cause analysis Corrective actions Preventive actions Technical escalation Post-incident review Documentation of findings and solutions The objective is not only to restore the database service but also to identify the underlying root cause and prevent recurrence. SLA & Production Support Operate within SLA-driven production environments. You will: Prioritize incidents according to severity and customer impact. Support critical production incidents. Identify potential SLA risks early. Escalate when specialist or additional technical support is required. Maintain clear technical documentation and ticket updates. Drive database issues through sustainable resolution. Experience with structured ITSM/ticketing tools, preferably Jira / Jira Service Management, is an advantage. Documentation & Knowledge Management Create, maintain and continuously improve: DBA procedures SOPs Runbooks Troubleshooting guides Backup/recovery procedures Monitoring procedures Known-error documentation Knowledge-base articles Cross-Skilling – ZONOS Knowledge The primary responsibility of this role remains Database Administration. As part of the Support cross-skilling approach, the DBA will progressively acquire operational knowledge of the ZONOS platform to better understand how databases interact with the wider application environment. This includes basic operational understanding of: ZONOS architecture and components Application/database dependencies Application health Application logs Linux/Kubernetes dependencies Database/application connectivity Monitoring Standard documented ZONOS troubleshooting procedures This knowledge enables the DBA to contribute more effectively to L3 incident investigation and work closely with the E2E Support team. The DBA remains the database specialist. Complex ZONOS application troubleshooting, architecture, product defects and code-level investigation remain with the respective E2E/System Engineering/Development specialists. Previous ZONOS knowledge is not required at recruitment and will be developed through structured knowledge transfer and practical experience. 24/7 Production Operations Willingness and capability to participate in 24/7/365 production Support operations according to the defined shift/on-call model. This may include scheduled: Day/night coverage Weekend coverage Public-holiday coverage Critical incident support Structured technical handover is required to ensure service continuity. Required Skills & Qualifications Minimum 5 years of relevant DBA experience, including at least 3 years supporting business-critical production database environments. Strong Oracle 19c administration. Strong Oracle RAC and Data Guard expertise. Strong Oracle performance tuning and SQL optimization. Strong PostgreSQL administration. PostgreSQL backup, replication, PITR, recovery and migration. Hands-on ClickHouse administration. Strong Linux administration and troubleshooting. Kubernetes production knowledge. Experience supporting database workloads on Kubernetes. Shell/Bash/Python or equivalent scripting skills. OCI, AWS, Azure or equivalent cloud experience. Strong monitoring and observability capabilities. Strong backup, restore and recovery expertise. Experience with HA and DR environments. Strong production troubleshooting and RCA skills. Experience operating within SLA-driven production environments. Ability to handle critical production incidents independently. Strong documentation and knowledge-sharing skills. Experience with large-scale enterprise production systems. Hybrid and multi-cloud experience. Jira / Jira Service Management. Grafana, Prometheus, MONIT or equivalent monitoring tools. Terraform. Ansible. Docker. Kubernetes Operators. OKE. Jenkins. GitLab. GitHub Actions. You bring: Strong analytical and problem-solving skills. Deep database troubleshooting capability. Strong ownership and accountability. Ability to work independently on critical production systems. Ability to work effectively under pressure. A proactive approach to performance, security and reliability. Clear technical communication. Strong collaboration across technical teams. Willingness to continuously develop your technical knowledge. Willingness to share knowledge and support cross-skilling. Willingness to support 24/7 production operations. Strong problem-solving skills and ability to handle high-pressure production issues Experience working in high-availability and disaster recovery environments Clear communication and collaboration skills for cross-team coordination A proactive mindset with attention to performance, security, and reliability Willingness to support 24×7 operations and participate in on-call rotations Good to Have Infrastructure as Code (Terraform, Ansible) Docker, Kubernetes Operators, OKE CI/CD tools (Jenkins, GitLab, GitHub Actions)

Ähnliche Stellen