← Back to list
Job · Senior

Cloud Infrastructure Engineer (Open LMS)

Platform Engineer • Senior • Remote • Full-time Hungary Hungary

Senior Cloud Infrastructure Engineer role for LTG's Open LMS, operating a multi-tenant AWS SaaS platform built with Terraform, Puppet, and Python, without Kubernetes.

Responsibilities

  • Design, build, and maintain AWS infrastructure using Terraform (EC2, RDS, S3, SQS, Lambda, ALB, ElastiCache, Route 53, VPC)
  • Write and maintain Puppet modules to configure fleets of EC2 instances across auto-scaling groups
  • Maintain and extend Python-based automation and tooling for platform operations
  • Operate and improve distributed service discovery and configuration management (etcd)
  • Manage and tune a multi-tier caching strategy (Varnish, Redis/Valkey, PHP OPcache)
  • Run and scale the observability stack (Prometheus, Grafana, Loki, Fluentd, PagerDuty), participate in on-call rotations
  • Evaluate and implement distributed storage solutions as the platform evolves
  • Improve deployment workflows and release processes
  • Collaborate with internal teams on API contracts and operational tooling
  • Participate in incident response, root cause analysis, and reliability improvements

Requirements

  • Strong production AWS experience (EC2, RDS, S3, SQS, Lambda, ALB, ElastiCache, Route 53, IAM, VPC)
  • Proficiency authoring and maintaining Terraform modules for production infrastructure
  • Proficiency authoring and maintaining Puppet modules (or equivalent) for fleet management
  • Solid Python skills, writing and maintaining production daemons
  • Deep Linux systems knowledge (Ubuntu): Apache/Nginx, PHP-FPM, Varnish, systemd, filesystem mounts, networking
  • Understanding of distributed systems concepts: consensus, leader election, distributed locking, eventual consistency
  • Proficiency building observability pipelines (Prometheus, Grafana, Loki or equivalent) in production
  • Comfortable in a GitLab-based CI/CD workflow
  • Clear communicator able to document architectural decisions

Nice to have

  • Hands-on experience with distributed storage systems (Ceph, GlusterFS, JuiceFS, CubeFS, AWS EFS)
  • Familiarity with etcd or similar distributed key-value stores (Consul, ZooKeeper)
  • Experience with Varnish and VCL, especially dynamic backend routing or multi-tenancy

Soft skills

Clear communication with technical and non-technical stakeholdersFirst-principles reasoning about distributed systems

About the company

The team builds and scales a multi-tenant SaaS hosting platform on AWS that dynamically provisions, manages, and scales hundreds of Moodle LMS instances for education clients, powered by custom orchestration tooling and infrastructure as code.

Similar jobs