SUMMARY:
-
POSITION INFO:
Location: Remote Employment Type: Full-Time Industry: Cloud Infrastructure | DevOps | Site Reliability Engineering | Data Technology WatersEdge Solutions is partnering with a growing technology business to appoint a DevOps \/ SRE Cloud Engineer (Site Reliability Engineer) to take end-to-end ownership of its Azure-based Kubernetes platform. This is a senior, hands-on engineering role focused on building and operating resilient cloud infrastructure, strengthening CI\/CD, and creating the observability needed to run production workloads reliably at scale. You’ll work across Azure, Kubernetes, Terraform, Docker, GitHub Actions and SRE practices , with the autonomy to own the environment and continuously improve its reliability, security and performance. About the Role As DevOps \/ SRE Cloud Engineer, you’ll take ownership of the production Kubernetes and Azure environment, managing infrastructure as code and ensuring workloads are secure, observable and dependable. You’ll operate Kubernetes clusters, maintain modular Terraform infrastructure, build and improve deployment pipelines, manage cloud identity and secrets, and establish effective monitoring, alerting and SLOs. This role requires someone who is comfortable going beyond maintaining infrastructure. You’ll be expected to understand how the platform behaves under real production load, troubleshoot complex issues alongside engineering teams, respond effectively to incidents and make practical improvements that prevent problems from recurring. Key Responsibilities Operate and maintain production Kubernetes clusters and Azure cloud infrastructure. Manage infrastructure through modular Terraform code across multiple environments. Own Kubernetes deployments, configuration, upgrades and operational troubleshooting. Design and maintain Helm charts and environment-specific deployment configurations. Manage Kubernetes node pools, resource requests and limits, and capacity requirements. Maintain secure and reliable CI\/CD pipelines using GitHub Actions . Implement linting, testing and build quality gates for production deployments. Manage container image build and publishing workflows. Maintain branch protection and required-status-check processes. Own platform observability across dashboards, metrics, alerting and structured logging. Define and maintain meaningful SLOs and alerting strategies. Participate in and improve incident response processes. Perform capacity planning based on measured production workloads. Manage cloud secrets, identity and network security across environments. Maintain secure integration with Azure Key Vault and appropriate identity mechanisms. Manage Azure networking requirements, including VNets, private endpoints and firewall rules. Maintain Docker images, container health checks and local Compose-based development environments. Build and maintain operational tooling and scripts using Bash. Partner with software and data engineering teams to troubleshoot production issues and improve platform reliability. Ensure operational scripts and infrastructure code are maintained with the same discipline as application code. What You’ll Bring 5+ years of DevOps, SRE or platform engineering experience. 2–3+ years of experience operating Kubernetes workloads in production. 2+ years of hands-on Azure experience. Strong production experience with Azure Kubernetes Service (AKS) . Hands-on experience with ADLS Gen2 , Azure Key Vault and Entra ID. Understanding of Azure workload identities and managed identities. Strong Azure networking experience, including VNets, private endpoints and firewall rules. Experience with Azure Service Bus or an equivalent message broker. Strong production Kubernetes knowledge, including Helm chart authoring and environment overlays. Experience working with Kubernetes operators. Understanding of node-pool design, capacity sizing and resource requests\/limits. Strong Kubernetes pod troubleshooting and upgrade experience. Advanced practical experience with Terraform and modular Infrastructure as Code. Experience managing Terraform state and disciplined plan\/apply processes across environments. Strong Docker experience, including image builds, multi-architecture builds, health checks and memory-limit tuning. Experience maintaining Docker Compose-based local development environments. Strong CI\/CD experience using GitHub Actions . Experience implementing automated lint, test and build gates. Hands-on experience with Prometheus and Grafana . Practical understanding of SRE principles, including SLOs, alerting, incident response and capacity planning. Strong Linux administration and Bash scripting skills. Experience implementing secure secrets-management practices. Ability to independently own a production cloud environment. Strong communication skills and the ability to collaborate effectively within hybrid and remote engineering teams. Nice to Have Experience operating Apache Spark on Kubernetes . Knowledge of Spark Operator, Spark Connect and executor\/driver tuning. JupyterHub on Kubernetes experience. Exposure to open data-platform technologies including Hive Metastore, Trino, Apache Ranger and Delta Lake . Experience with S3-compatible object storage such as MinIO. Understanding of governed data and query architectures. OpenTelemetry metrics and tracing experience. OpenLineage or Marquez exposure. Python experience for operational tooling and infrastructure testing. Experience using pytest for environment or infrastructure test suites. Operational experience with SQL Server and PostgreSQL. Database backup and container deployment experience. Exposure to EF migration-driven database schemas. Experience within multi-tenant or regulated-data environments. Understanding of tenant isolation patterns. Exposure to POPIA, GDPR or ISO 27001-aligned controls. Experience with secret scanning, dependency auditing and container image provenance. Azure and\/or Kubernetes certifications such as CKA . Qualifications Bachelor’s degree in Computer Science, Engineering or a related discipline , or equivalent practical experience. Proven experience operating production-grade cloud infrastructure at scale. Azure, Kubernetes or related cloud certifications are advantageous but not essential. What’s On Offer Flexible hybrid and remote working arrangements. End-to-end ownership of a production Azure and Kubernetes environment. Opportunity to work with modern cloud, containerisation and infrastructure technologies. Exposure to complex data-platform infrastructure and distributed workloads. Scope to shape platform reliability, observability, automation and security practices. Wellness initiatives and home-office support. Continuous learning and professional development opportunities. Supportive and inclusive team environment. Regular team-building activities and social events. Performance recognition and employee share ownership opportunities. A culture that values transparency, accountability and work-life balance. Company Culture You’ll be joining a technically ambitious environment where engineers are trusted to take ownership and make meaningful improvements to the systems they manage. The team values reliability, automation and strong engineering discipline, while maintaining a collaborative and supportive working style. You’ll have the autonomy to identify weaknesses, improve infrastructure and tooling, and work closely with engineering teams to create a platform that can scale reliably as the business grows. Please Note: If you have not been contacted within 10 working days, consider your application unsuccessful.