CareersSRE/L3
Full-timeOn-site

SRE/L3

Location
Abu Dhabi, UAE
Experience
5+ years

Team: Platform Engineering / Digital Engineering Location: Abu Dhabi, UAE Core Stack: AWS (EKS/Kubernetes, Lambda) · AWS-native Monitoring & Observability · Incident Management · Python

About the Role

We are looking for an experienced Site Reliability Engineer (SRE) to join our Platform Engineering team in Abu Dhabi.

As an SRE, you will be responsible for troubleshooting and resolving complex production issues, maintaining system uptime, improving observability, and automating operational processes across our AWS environment.

You will work extensively with a Kubernetes- and serverless-based AWS architecture, including services running on Amazon EKS, event-driven workloads using AWS Lambda, cloud storage and processing, and the data paths connecting applications and services.

You will use Python to automate operations, develop platform tooling, build remediation workflows, and continuously eliminate reliability gaps.

This is a build-and-operate role. You will be deeply involved in production incidents when they occur while also building the automation, monitoring, and infrastructure improvements that prevent them from recurring.

What You Will Do

Troubleshoot and Resolve Complex Production Issues

  • Own end-to-end diagnosis of production problems spanning EKS workloads, Lambda functions, ingress and service mesh, queues and streams such as SQS, SNS, Kinesis and EventBridge, databases such as RDS/Aurora and DynamoDB, and S3-based workloads.
  • Troubleshoot complex Kubernetes issues including CrashLoopBackOff, OOMKills, failing readiness/liveness probes, pod evictions, node pressure, scheduling and affinity issues, PVC/CSI failures, CNI/DNS issues, RBAC denials and cluster upgrade issues.
  • Diagnose issues involving latency, partial failures, Lambda cold starts, concurrency limits, throttling, quota exhaustion, VPC networking, IAM/IRSA permissions, stuck queues and failed messages.
  • Identify true root causes and implement permanent fixes across code, configuration, capacity or architecture.
  • Drive corrective and preventive actions through to closure.

Maintain System Uptime and Reliability

  • Define and monitor SLIs, SLOs and error budgets for critical services.
  • Engineer highly available solutions using multi-AZ and, where required, multi-Region architectures.
  • Implement pod disruption budgets, topology spread, anti-affinity, health checks, circuit breakers, retries, graceful degradation and automated recovery mechanisms.
  • Manage capacity, autoscaling and performance using HPA/VPA, KEDA, Karpenter or Cluster Autoscaler.
  • Manage Lambda concurrency and provisioned concurrency.
  • Plan and execute EKS control-plane and node-group upgrades, add-on upgrades, AMI patching and Kubernetes API deprecation remediation.
  • Implement infrastructure-as-code, GitOps and CI/CD practices with safe deployment strategies such as canary and blue/green deployments.
  • Participate in disaster recovery, failover, load and game-day exercises against defined RTO/RPO targets.

Enhance Observability

  • Build and manage AWS-native observability using CloudWatch metrics, Logs, Logs Insights, alarms and dashboards.
  • Work with Container Insights, X-Ray, CloudTrail, AWS Config, CloudWatch Synthetics and Lambda Insights.
  • Manage in-cluster telemetry using CloudWatch Agent, AWS Distro for OpenTelemetry, Fluent Bit and Prometheus.
  • Implement structured logging, distributed tracing and custom metrics.
  • Develop SLO-driven and symptom-based alerting to reduce unnecessary alerts.
  • Build dashboards and golden-signal views for engineering and operational teams.

Incident Response and Operational Excellence

  • Participate in an on-call rotation and act as Incident Commander for high-severity incidents.
  • Work with incident management and escalation platforms such as PagerDuty, Opsgenie or equivalent.
  • Use tools such as AWS Systems Manager Incident Manager, Jira and ServiceNow.
  • Lead blameless postmortems and drive corrective actions.
  • Improve MTTR and MTTD through automation and operational improvements.
  • Develop automated runbooks, diagnostics and remediation workflows using Python and AWS Systems Manager.

Automation and Platform Engineering

  • Develop production-quality Python for operational automation, Lambda workflows, remediation, synthetic checks, log/metric analysis and reporting.
  • Manage infrastructure-as-code using Terraform, CDK or CloudFormation.
  • Manage Kubernetes delivery using Helm, Kustomize, Argo CD or Flux.
  • Build and maintain shared platform capabilities including base charts, cluster add-ons, policy guardrails and production-readiness templates.
  • Identify and eliminate repetitive manual operational tasks.

Security and Compliance

  • Implement secure-by-default infrastructure using least-privilege IAM and Kubernetes RBAC.
  • Work with IRSA/EKS Pod Identity, Secrets Manager, KMS, network policies and pod security standards.
  • Support image scanning/signing, encryption, audit logging and vulnerability management.
  • Ensure appropriate security, privacy and data-residency practices across cloud infrastructure and workloads.

Engineering Culture

  • Coach application teams on production readiness, operability and on-call practices.
  • Develop documentation, dashboards and tooling that make reliability a shared responsibility across engineering teams.

Minimum Qualifications

  • Bachelor's degree in Computer Science, Engineering or a related technical field, or equivalent practical experience.
  • 5+ years of experience in SRE, Production Engineering, DevOps or Software Engineering with significant production ownership.
  • Strong Python programming and automation skills.
  • Deep hands-on AWS experience in production, particularly Amazon EKS/Kubernetes.
  • Experience operating production Kubernetes clusters, including workload/resource management, autoscaling, ingress/load balancing, networking, storage, RBAC and cluster upgrades.
  • Production experience with AWS Lambda, including event sources, concurrency, cold starts, timeouts, error handling and DLQs.
  • Strong knowledge of AWS services including IAM, VPC, S3, ECR, API Gateway, SQS, SNS, EventBridge and at least one managed database.
  • Experience with Helm and/or Kustomize and GitOps or pipeline-based deployment models.
  • Strong experience with AWS monitoring and observability, particularly CloudWatch, X-Ray and CloudTrail.
  • Experience with incident management, on-call processes, incident command, postmortems and corrective actions.
  • Infrastructure-as-code and CI/CD experience with Terraform, CDK, CloudFormation, GitHub Actions, AWS CodePipeline, Argo CD, Flux, Jenkins or equivalent.
  • Strong Linux, containers, networking and distributed-systems fundamentals.
  • Excellent communication and troubleshooting skills.

Location & Availability

📍 Location: Abu Dhabi, UAE

Preference will be given to:

  • Candidates currently based in Abu Dhabi/UAE
  • Candidates who are ready to relocate to Abu Dhabi immediately after selection
  • Immediate joiners to candidates with up to 2 weeks' notice period

Candidates with longer notice periods may be considered based on the profile and business requirement.

Apply Now

SRE/L3

Drop your resume here or browse

PDF or Word (max 10MB)

By submitting, you agree to our privacy policy and consent to being contacted about this position.

Quick Info

Job ID#67
PostedAug 14, 2026
Work ModelOn-site